EDBT 2026 Demo / reviewers in the wild / expert
Bo Liu 0047
dblp:58/2670-47
· DBLP profile ↗
23ranked-venue papers
6as first author
17since 2021 · last 2026
0000-0002-3164-6299ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 15 · 3 first-author · 13 since 2021Artificial intelligence and machine learning · 14 · 4 first-author · 12 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-authorSecurity and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TGDD: Trajectory Guided Dataset Distillation with Balanced DistributionabstractDataset distillation compresses large datasets into compact synthetic ones to reduce storage and computational costs. Among various approaches, distribution matching (DM)-based methods have attracted attention for their high efficiency. However, they often overlook the evolution of feature representations during training, which limits the expressiveness of synthetic data and weakens downstream performance. To address this issue, we propose Trajectory Guided Dataset Distillation (TGDD), which reformulates distribution matching as a dynamic alignment process along the model’s training trajectory. At each training stage, TGDD captures evolving semantics by aligning the feature distribution between the synthetic and original dataset. Meanwhile, it introduces a distribution constraint regularization to reduce class overlap. This design helps synthetic data preserve both semantic diversity and representativeness, improving performance in downstream tasks. Without additional optimization overhead, TGDD achieves a favorable balance between performance and efficiency. Experiments on ten datasets demonstrate that TGDD achieves state-of-the-art performance, notably a 5.0% accuracy gain on high-resolution benchmarks. Fengli Ran, Xiao Pu 0002, Bo Liu 0047, Xiuli Bi, Bin Xiao 0002 |
AAAI | 3 |
| 2025 | CustomTTT: Motion and Appearance Customized Video Generation via Test-Time TrainingabstractBenefiting from large-scale pre-training of text-video pairs, current text-to-video (T2V) diffusion models can generate high-quality videos from the text description. Besides, given some reference images or videos, the parameter-efficient fine-tuning method, i.e. LoRA, can generate high-quality customized concepts, e.g., the specific subject or the motions from a reference video. However, combining the trained multiple concepts from different references into a single network shows obvious artifacts. To this end, we propose CustomTTT, where we can joint custom the appearance and the motion of the given video easily. In detail, we first analyze the prompt influence in the current video diffusion model and find the LoRAs are only needed for the specific layers for appearance and motion customization. Besides, since each LoRA is trained individually, we propose a novel test-time training technique to update parameters after combination utilizing the trained customized models. We conduct detailed experiments to verify the effectiveness of the proposed methods. Our method outperforms several state-of-the-art works in both qualitative and quantitative evaluations. Xiuli Bi, Bo Liu 0047, Xiaodong Cun, Yong Zhang 0034, Weisheng Li 0001, Bin Xiao 0002 |
AAAI | 3 |
| 2025 | Towards Universal AI-Generated Image Detection by Variational Information Bottleneck NetworkabstractThe rapid advancement of generative models has significantly improved the quality of generated images. Mean-while, it challenges information authenticity and credibility. Current generated image detection methods based on large-scale pre-trained multimodal models have achieved impressive results. Although these models provide abundant features, the authentication task-related features are often submerged. Consequently, those authentication task-irrelated features cause models to learn superficial biases, thereby harming their generalization performance across different model genera (e.g., GANs and Diffusion Models). To this end, we proposed VIB-Net, which uses Variational Information Bottlenecks to enforce authentication task-related feature learning. We tested and analyzed the proposed method and existing methods on samples generated by 17 different generative models. Compared to SOTA methods, VIB-Net achieved a 5.55% improvement in mAP and a 9.33% increase in accuracy. Notably, in generalization tests on unseen generative models from different series, VIB-Net improved mAP by 12.48% and accuracy by 23.59% over SOTA methods. The code is available at https://github.com/oceanzhf/VIBAIGCDetect. Qinghui He, Xiuli Bi, Weisheng Li 0001, Bo Liu 0047, Bin Xiao 0002 |
CVPR | 5 |
| 2025 | Covert and Potent: A Weather-Camouflaged Backdoor Attacks on Self-Supervised LearningabstractSelf-supervised learning is widely applied across various domains due to its advantage of learning data representations without the need for labels. However, recent research shows that backdoor attacks on self-supervised learning are achievable by coupling benign features with trigger features without manipulating labels. Existing methods, however, suffer from poor trigger disguise. When designing triggers, more emphasis is placed on attack strength rather than on disguising the triggers, which makes these triggers easily detectable through manual inspection or preprocessing methods. Therefore, we propose a camouflaged self-supervised backdoor attack method from the perspective of visual disguise. Specifically, we design triggers by embedding variable adverse weather information to achieve visual camouflage, which can bypass certain defence methods to some extent. Additionally, since our proposed camouflaged triggers have a global nature, they achieve more efficient backdoor attack capabilities. Experiments demonstrate that our method achieves attack success rates of 83.4% on the CIFAR-100 dataset and 44.8% on the ImageNet-100 dataset, surpassing existing state-of-the-art methods by 14.6% and 24.4%, respectively. At the same time, our method exhibits better stealthiness. Yang Wei 0002, Yonghao Yang, Bo Liu 0047, Bin Xiao 0002 |
ICASSP | 3 |
| 2024 | Focus Stacking with High Fidelity and Superior Visual EffectsabstractFocus stacking is a technique in computational photography, and it synthesizes a single all-in-focus image from different focal plane images. It is difficult for previous works to produce a high-quality all-in-focus image that meets two goals: high-fidelity to its source images and good visual effects without defects or abnormalities. This paper proposes a novel method based on optical imaging process analysis and modeling. Based on a foreground segmentation - diffusion elimination architecture, the foreground segmentation makes most of the areas in full-focus images heritage information from the source images to achieve high fidelity; diffusion elimination models the physical imaging process and is specially used to solve the transition region (TR) problem that is a long-term neglected issue and degrades visual effects of synthesized images. Based on extensive experiments on simulated dataset, existing realistic dataset and our proposed BetaFusion dataset, the results show that our proposed method can generate high-quality all-in-focus images by achieving two goals simultaneously, especially can successfully solve the TR problem and eliminate the visual effect degradation of synthesized images caused by the TR problem. Bo Liu 0047, Xiuli Bi, Weisheng Li 0001, Bin Xiao 0002 |
AAAI | 1 |
| 2024 | Depth-Aware Test-Time Training for Zero-Shot Video Object SegmentationabstractZero-shot Video Object Segmentation (ZSVOS) aims at segmenting the primary moving object without any human annotations. Mainstream solutions mainly focus on learning a single model on large-scale video datasets, which struggle to generalize to unseen videos. In this work, we introduce a test-time training (TTT) strategy to address the problem. Our key insight is to enforce the model to predict consistent depth during the TTT process. In detail, we first train a single network to perform both segmentation and depth prediction tasks. This can be effectively learned with our specifically designed depth modulation layer. Then, for the TTT process, the model is updated by predicting consistent depth maps for the same frame under different data augmentations. In addition, we explore different TTT weight updating strategies. Our empirical results suggest that the momentum-based weight initialization and looping-based training scheme lead to more stable improvements. Experiments show that the proposed method achieves clear improvements on ZSVOS. Our proposed video TTT strategy provides significant superiority over state-of-the-art TTT methods. Our code is available at: https://nifangbaage.github.io/DATTT/. Weihuang Liu, Xi Shen 0001, Haolun Li 0001, Xiuli Bi, Bo Liu 0047, Chi-Man Pun, Xiaodong Cun |
CVPR | 5 |
| 2024 | Using My Artistic Style? You Must Obtain My Authorization
Xiuli Bi, Weisheng Li 0001, Bo Liu 0047, Bin Xiao 0002 |
ECCV (86) | 4 |
| 2024 | PriFU: Capturing Task-Relevant Information Without Adversarial LearningabstractAs machine learning advances, machine learning as a service (MLaaS) in the cloud brings convenience to human lives but also privacy risks, as powerful neural networks used for generation, classification or other tasks can also become privacy snoopers. This motivates privacy preservation in the inference phase. Many approaches for preserving privacy in the inference phase introduce multi-objective functions, training models to remove specific private information from users' uploaded data. Although effective, these adversarial learning-based approaches suffer not only from convergence difficulties, but also from limited generalization beyond the specific privacy for which they are trained. To address these issues, we propose a method for privacy preservation in the inference phase by removing task-irrelevant information, which requires no knowledge of the privacy attacks nor introduction of adversarial learning. Specifically, we introduce a metric to distinguish task-irrelevant information from task-relevant information, and achieve more efficient metric estimation to remove task-irrelevant features. The experiments demonstrate the potential of our method in several tasks. Our code will be available at: https://github.com/iwhoyoung/PriFU. Xiuli Bi, Bo Liu 0047, Weisheng Li 0001, Pamela C. Cosman, Bin Xiao 0002 |
ACM Multimedia | 3 |
| 2024 | D-Net: A dual-encoder network for image splicing forgery detection and localization
Bo Liu 0047, Xiuli Bi, Bin Xiao 0002, Weisheng Li 0001, Guoyin Wang 0001, Xinbo Gao 0001 |
Pattern Recognit. | 2 |
| 2024 | Learning Discriminative Representations From Cross-Scale Features for Camouflaged Object DetectionabstractThe key that hinders the performance improvement of current camouflaged object detection (COD) models is the lack of discriminability of features at fine granularity. We solve this problem from two complementary perspectives. Firstly, complex scenes result in the discriminative feature representations of camouflaged objects being present at different scales and semantic abstraction levels. Therefore, a mechanism is needed to increase the diversity of features to integrate more information potentially beneficial for COD. Second, appearance similarity between objects and environments will inevitably lead to similarity in features. Enhancing feature diversity alone is not enough to solve the above problems. Therefore, it is necessary to give the model semantic perception capabilities to expand the subtle discrepancies between objects and environments in feature embedding. Inspired by the first point, we propose a cross-scale interaction module (CSIM) that utilizes cross-attention between different scales to enhance the diversity of feature representations. Regarding the second point, the semantic guided feature learning (SGFL) is proposed to promote the model to expand feature discrepancies through explicit supervision. Experiments on four popular COD datasets show that our method outperforms recent SOTA methods. In addition, polyp segmentation experiments show that it is also effective for other COD-like tasks. Yongchao Wang 0004, Xiuli Bi, Bo Liu 0047, Yang Wei 0002, Weisheng Li 0001, Bin Xiao 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Self-Supervised Image Local Forgery Detection by JPEG Compression TraceabstractFor image local forgery detection, the existing methods require a large amount of labeled data for training, and most of them cannot detect multiple types of forgery simultaneously. In this paper, we firstly analyzed the JPEG compression traces which are mainly caused by different JPEG compression chains, and designed a trace extractor to learn such traces. Then, we utilized the trace extractor as the backbone and trained self-supervised to strengthen the discrimination ability of learned traces. With its benefits, regions with different JPEG compression chains can easily be distinguished within a forged image. Furthermore, our method does not rely on a large amount of training data, and even does not require any forged images for training. Experiments show that the proposed method can detect image local forgery on different datasets without re-training, and keep stable performance over various types of image local forgery. Xiuli Bi, Wuqing Yan, Bo Liu 0047, Bin Xiao 0002, Weisheng Li 0001, Xinbo Gao 0001 |
AAAI | 3 |
| 2023 | DLBD: A Self-Supervised Direct-Learned Binary DescriptorabstractFor learning-based binary descriptors, the binarization process has not been well addressed. The reason is that the binarization blocks gradient back-propagation. Existing learning-based binary descriptors learn real-valued output, and then it is converted to binary descriptors by their proposed binarization processes. Since their binarizaiion processes are not a component of the network, the learning-based binary descriptor cannot fully utilize the advances of deep learning. To solve this issue, we propose a model-agnostic plugin binary transformation layer (BTL), making the network directly generate binary descriptors. Then, we present the first self-supervised, direct-learned binary descriptor, dubbed DLBD. Furthermore, we propose ultra-wide temperature-scaled crossentropy loss to adjust the distribution of learned descriptors in a larger range. Experiments demonstrate that the proposed BTL can substitute the previous binarization process. Our proposed DLBD outperforms SOTA on different tasks such as image retrieval and classification11Our code is available at: https://github.com/CQUPT-CV/DLBD. Bin Xiao 0002, Bo Liu 0047, Xiuli Bi, Weisheng Li 0001, Xinbo Gao 0001 |
CVPR | 3 |
| 2023 | A Versatile Detection Method for Various Contrast Enhancement ManipulationsabstractContrast enhancement manipulation is a common method to improve the visual effect of an image. Meanwhile, it can also be considered a type of global image forgery because it changes the image’s visual appearance without alerting its semantics. Moreover, for local image forgery, a tampered image may be composited by images with different contrast enhancement manipulations or post-processed by a contrast enhancement manipulation to conceal the trails of tampering. Therefore, contrast enhancement manipulation detection is critical to global image forgery detection. The existing methods can only detect a particular type of contrast enhancement manipulation, such as gamma correction or histogram equalization. To break this limitation, we propose the zero-gap spans (ZGS) as the fingerprint to explore the traces of contrast enhancement manipulations. Based on ZGS, various contrast enhancement manipulations can be distinguished by a simple classification method at image-level and patch-level; different gamma corrections can be identified, and their gamma value can be estimated. Experimental results indicate that the proposed ZGS-based classification method can achieve and maintain good classification performance under different cases (gamma correction, simple histogram equalization, modified histogram equalization techniques). Meanwhile, ZGS can estimate the gamma value with the mean squared error (MSE) below 0.1156. For the local forgery images, the proposed ZGS also can be utilized to locate the regions with different contrast enhancement manipulations. Xiuli Bi, Yixuan Shang, Bo Liu 0047, Bin Xiao 0002, Weisheng Li 0001, Xinbo Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | FF-Net: An End-to-end Feature-Fusion Network for Double JPEG Detection and Localization
Bo Liu 0047, Ranglei Wu, Xiuli Bi, Bin Xiao 0002 |
ACML | 1 |
| 2022 | Detecting Generated Images by Real Images
Bo Liu 0047, Fan Yang 0159, Xiuli Bi, Bin Xiao 0002, Weisheng Li 0001, Xinbo Gao 0001 |
ECCV (14) | 1 |
| 2022 | Privacy-Preserving Color Image Feature Extraction by Quaternion Discrete Orthogonal MomentsabstractTo implement image storage and computation in cloud servers without violating users’ privacy, privacy-preserving feature extraction has been a new research interest. The existing works are mainly designed for grayscale images. For color images, they tend to convert them to grayscale images or obtain the results of the combination of single-channel processes. While the capabilities of features extracted from the encrypted color images will be affected if color information and inter-relationship between color channels are ignored. To fully preserve features of color images, we introduce quaternion theory to encode each color image and propose an improved vector homomorphic encryption scheme (IVHE) to encrypt quaternion-based color images. IVHE helps protect image content and keep vector characteristics of color images. Based on IVHE, the framework for feature extraction of privacy-preserving Quaternion Discrete Orthogonal Moments (PPQDOMs) is presented. Theoretical analyses prove that Quaternion Discrete Orthogonal Moments (QDOMs) can be extracted from the encrypted color images by PPQDOMs. Furthermore, we apply three common Discrete Orthogonal Moments to the proposed framework to evaluate its performance. Experimental results demonstrate that the proposed framework can protect color image content and perform well compared to QDOMs in image reconstruction and image recognition. Xiuli Bi, Chao Shuai, Bo Liu 0047, Bin Xiao 0002, Weisheng Li 0001, Xinbo Gao 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2021 | Fooling deep neural detection networks with adaptive object-oriented adversarial perturbation
Yatie Xiao, Chi-Man Pun, Bo Liu 0047 |
Pattern Recognit. | 3 |
| 2020 | Locating splicing forgery by adaptive-SVD noise estimation and vicinity noise descriptor
Bo Liu 0047, Chi-Man Pun |
Neurocomputing | 1 |
| 2020 | Crafting adversarial example with adaptive root mean square gradient on deep neural networks
Yatie Xiao, Chi-Man Pun, Bo Liu 0047 |
Neurocomputing | 3 |
| 2020 | Exposing splicing forgery in realistic scenes using deep fusion network
Bo Liu 0047, Chi-Man Pun |
Inf. Sci. | 1 |
| 2020 | Adversarial example generation with adaptive gradient search for single and ensemble deep neural network
Yatie Xiao, Chi-Man Pun, Bo Liu 0047 |
Inf. Sci. | 3 |
| 2018 | Locating splicing forgery by fully convolutional networks and conditional random field
Bo Liu 0047, Chi-Man Pun |
Signal Process. Image Commun. | 1 |
| 2016 | Multi-scale noise estimation for image splicing forgery detection
Chi-Man Pun, Bo Liu 0047, Xiaochen Yuan |
J. Vis. Commun. Image Represent. | 2 |