Qichao Ying

dblp:239/5277 · DBLP profile ↗
← Back
20ranked-venue papers
7as first author
20since 2021 · last 2025
0000-0002-6527-2424ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 17 · 6 first-author · 17 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 MoFRR: Mixture of Diffusion Models for Face Retouching Restoration
Qichao Ying, Zhenxing Qian, Sheng Li 0006, Runqi Zhang, Xinpeng Zhang 0001
ICCV2
2025 Radiometric Calibration Using Artificial Intelligence: Constituting Uniform Observing Systems for Infrared Satellites
abstract
Radiometric calibration (RC) is a critical process in aerospace infrared remote sensing that establishes the relationship between the radiation energy of observed objects and the digital number (DN) output from sensors, which is fundamental for ensuring high-precision applications of infrared remote sensing data. At present, source-based RC (SBRC) is the predominant method, relying on a variety of radiometric sources (RSs) including in-orbit blackbodies, or natural targets such as lakes and oceans. This approach, while effective, imposes constraints on remote sensing systems such as space & weight allocation for RS and additional observation time for RC. Moreover, the reliance on physical calibration sources can introduce uncertainties due to factors such as imperfect emissivity of in-orbit blackbodies, lack of data consistency due to varied RS types, and variations in environmental conditions. In this article, we propose a novel RC method named artificial intelligence RC (AIRC), which directly generates RC coefficients for the in-orbit remote sensing satellites using the physical and environmental parameters of the sensor. We first theoretically prove that RC coefficients can be derived as functions of the sensor states. Next, we propose our neural networks for infrared RC (RCNN), to learn this relationship based on historical high-accuracy calibration data, enabling a shift from reference traceability (RT) to states traceability (ST). Then, to verify the feasibility of the proposed scheme, we train and test a multilayered perceptron (MLP) as a simple implementation of RCNN based on our long-term well-curated RC data from our FengYun-4A Advanced Geosynchronous Radiation Imager (FY-4A AGRI), and the experiments show that the proposed method achieves high-accuracy RC comparable with the official RC method applied on FY-4A AGRI that uses an in-orbit blackbody. Our study showcases how to conduct RC using the “reason (the states of sensor)–results (calibration coefficient)” logic, as supplement to the existing “result (observation to RS)–reason (calibration coefficient)” logic, which promotes constituting a uniform observing system for cross-platform infrared satellites.
Boyang Chen 0001, Aiqun Wu, Wen Hui, Peng Rao, Changpei Han, Qichao Ying, Yapeng Wu, Damian Moss, Zhenxing Qian
IEEE Trans. Geosci. Remote. Sens.8
2024 Search, Examine and Early-Termination: Fake News Detection with Annotation-Free Evidences
abstract
Pioneer researches recognize evidences as crucial elements in fake news detection apart from patterns. Existing evidence-aware methods either require laborious pre-processing procedures to assure relevant and high-quality evidence data, or incorporate the entire spectrum of available evidences in all news cases, regardless of the quality and quantity of the retrieved data. In this paper, we propose an approach named SEE that retrieves useful information from web-searched annotation-free evidences with an early-termination mechanism. The proposed SEE is constructed by three main phases: Searching online materials using the news as a query and directly using their titles as evidences without any annotating or filtering procedure, sequentially Examining the news alongside with each piece of evidence via attention mechanisms to produce new hidden states with retrieved information, and allowing Early-termination within the examining loop by assessing whether there is adequate confidence for producing a correct prediction. We have conducted extensive experiments on datasets with unprocessed evidences, i.e., Weibo21, GossipCop, and pre-processed evidences, namely Snopes and PolitiFact. The experimental results demonstrate that the proposed method outperforms state-of-the-art approaches.
Yuzhou Yang, Yangming Zhou, Qichao Ying, Zhenxing Qian, Xinpeng Zhang 0001
ECAI3
2024 Multi-view Feature Extraction via Tunable Prompts is Enough for Image Manipulation Localization
abstract
Deceptive images can quickly spread via social networking services, posing significant risks. The rapid progress in Image Manipulation Localization (IML) seeks to address this issue. However, the scarcity of public training datasets in the IML task directly hampers the performance of models. To address the challenge, we propose a Prompt-IML framework, which leverages the rich prior knowledge of pre-trained models by employing tunable prompts. Specifically, sets of tunable prompts enable the frozen pre-trained model to extract multi-view features, including spatial and high-frequency features. This approach minimizes redundant architecture for feature extraction across different views, resulting in reduced training costs. In addition, we develop a plug-and-play Feature Alignment and Fusion module that seamlessly integrates into the pre-trained models without additional structural modifications. The proposed module reduces noise and uncertainty in features through interactive processing. The experimental results showcase that our proposed method attains superior performance across 6 test datasets, demonstrating exceptional robustness.
Xuntao Liu, Yuzhou Yang, Qichao Ying, Zhenxing Qian, Xinpeng Zhang 0001, Sheng Li 0006
ACM Multimedia4
2024 Establishing Robust Generative Image Steganography via Popular Stable Diffusion
abstract
Generative steganography, a novel paradigm in information hiding, has garnered considerable attention for its potential to withstand steganalysis. However, existing generative steganography approaches suffer from the limited visual quality of generated images and are challenging to apply to lossy transmissions in real-world scenarios with unknown channel attacks. To address these issues, this paper proposes a novel robust generative image steganography scheme, facilitating zero-shot text-driven stego image generation without the need for additional training or fine-tuning. Specifically, we employ the popular Stable Diffusion model as the backbone generative network to establish a covert transmission channel. Our proposed framework overcomes the challenges of numerical instability and perturbation sensitivity inherent in diffusion models. Adhering to Kerckhoff’s principle, we propose a novel mapping module based on dual keys to enhance robustness and security under lossy transmission conditions. Experimental results showcase the superior performance of our method in terms of extraction accuracy, robustness, security, and image quality.
Xiaoxiao Hu, Sheng Li 0006, Qichao Ying, Wanli Peng, Xinpeng Zhang 0001, Zhenxing Qian
IEEE Trans. Inf. Forensics Secur.3
2023 Bootstrapping Multi-View Representations for Fake News Detection
abstract
Previous researches on multimedia fake news detection include a series of complex feature extraction and fusion networks to gather useful information from the news. However, how cross-modal consistency relates to the fidelity of news and how features from different modalities affect the decision-making are still open questions. This paper presents a novel scheme of Bootstrapping Multi-view Representations (BMR) for fake news detection. Given a multi-modal news, we extract representations respectively from the views of the text, the image pattern and the image semantics. Improved Multi-gate Mixture-of-Expert networks (iMMoE) are proposed for feature refinement and fusion. Representations from each view are separately used to coarsely predict the fidelity of the whole news, and the multimodal representations are able to predict the cross-modal consistency. With the prediction scores, we reweigh each view of the representations and bootstrap them for fake news detection. Extensive experiments conducted on typical fake news detection datasets prove that BMR outperforms state-of-the-art schemes.
Qichao Ying, Xiaoxiao Hu, Yangming Zhou, Zhenxing Qian, Dan Zeng 0001, Shiming Ge
AAAI1
2023 DRAW: Defending Camera-shooted RAW against Image Manipulation
abstract
RAW files are the initial measurement of scene radiance widely used in most cameras, and the ubiquitously-used RGB images are converted from RAW data through Image Signal Processing (ISP) pipelines. Nowadays, digital images are risky of being nefariously manipulated. Inspired by the fact that innate immunity is the first line of body defense, we propose DRAW, a novel scheme of defending images against manipulation by protecting their sources, i.e., camera-shooted RAWs. Specifically, we design a lightweight Multi-frequency Partial Fusion Network (MPF-Net) friendly to devices with limited computing resources by frequency learning and partial feature fusion. It introduces invisible watermarks as protective signal into the RAW data. The protection capability can not only be transferred into the rendered RGB images regardless of the applied ISP pipeline, but also is resilient to post-processing operations such as blurring or compression. Once the image is manipulated, we can accurately identify the forged areas with a localization network. Extensive experiments on several famous RAW datasets, e.g., RAISE, FiveK and SIDD, indicate the effectiveness of our method. We hope that this technique can be used in future cameras as an option for image protection, which could effectively restrict image manipulation at the source.
Xiaoxiao Hu, Qichao Ying, Zhenxing Qian, Sheng Li 0006, Xinpeng Zhang 0001
ICCV2
2023 Image Protection for Robust Cropping Localization and Recovery
abstract
Existing image cropping detection schemes ignore that recovering the cropped-out contents can unveil the purpose of the behaved cropping attack. This paper presents CLR-Net, a novel image protection scheme addressing the combined challenge of image Cropping Localization and Recovery. We first protect the original image by introducing imperceptible perturbations. Then, typical image post-processing attacks are simulated to erode the protected image. On the recipient’s side, we predict the cropping mask and recover the original image. Besides, we propose a novel Fine-Grained generative JPEG simulator (FG-JPEG) as well as a feature alignment network to improve the real-world robustness. Comprehensive experiments prove that the quality of the recovered image and the accuracy of crop localization are both satisfactory.
Qichao Ying, Hang Zhou 0007, Xiaoxiao Hu, Zhenxing Qian, Sheng Li 0006, Xinpeng Zhang 0001
ICME1
2023 Multimodal Fake News Detection via CLIP-Guided Learning
abstract
Fake news detection (FND) has attracted much research interests in social forensics. Many existing approaches introduce tailored attention mechanisms to fuse unimodal features. However, they ignore the impact of cross-modal similarity between modalities. Meanwhile, the potential of pretrained multimodal feature learning models in FND has not been well exploited. This paper proposes an FND-CLIP framework, i.e., a multimodal Fake News Detection network based on Contrastive Language-Image Pretraining (CLIP). FND-CLIP extracts the deep representations together from news using two unimodal encoders and two pair-wise CLIP encoders. The CLIP-generated multimodal features are weighted by CLIP similarity of the two modalities. We also introduce a modality-wise attention module to aggregate the features. Extensive experiments are conducted and the results indicate that the proposed framework has a better capability in mining crucial features for fake news detection. The proposed FND-CLIP can achieve better performances than previous works on three typical fake news datasets.
Yangming Zhou, Yuzhou Yang, Qichao Ying, Zhenxing Qian, Xinpeng Zhang 0001
ICME3
2023 Multi-modal Fake News Detection on Social Media via Multi-grained Information Fusion
abstract
The easy sharing of multimedia content on social media has caused a rapid dissemination of fake news, which threatens society’s stability and security. Therefore, fake news detection has garnered extensive research interest in the field of social forensics. Current methods primarily concentrate on the integration of textual and visual features but fail to effectively exploit multi-modal information at both fine-grained and coarse-grained levels. Furthermore, they suffer from an ambiguity problem due to a lack of correlation between modalities or a contradiction between the decisions made by each modality. To overcome these challenges, we present a Multi-grained Multi-modal Fusion Network (MMFN) for fake news detection. Inspired by the multi-grained process of human assessment of news authenticity, we respectively employ two Transformer-based pre-trained models to encode token-level features from text and images. The multi-modal module fuses fine-grained features, taking into account coarse-grained features encoded by the CLIP encoder. To address the ambiguity problem, we design uni-modal branches with similarity-based weighting to adaptively adjust the use of multi-modal features. Experimental results demonstrate that the proposed framework outperforms state-of-the-art methods on three prevalent datasets.
Yangming Zhou, Yuzhou Yang, Qichao Ying, Zhenxing Qian, Xinpeng Zhang 0001
ICMR3
2023 RetouchingFFHQ: A Large-scale Dataset for Fine-grained Face Retouching Detection
abstract
The widespread use of face retouching filters on short-video platforms has raised concerns about the authenticity of digital appearances and the impact of deceptive advertising. To address these issues, there is a pressing need to develop advanced face retouching techniques. However, the lack of large-scale and fine-grained face retouching datasets has been a major obstacle to progress in this field. In this paper, we introduce RetouchingFFHQ, a large-scale and fine-grained face retouching dataset that contains over half a million conditionally-retouched images. RetouchingFFHQ stands out from previous datasets due to its large scale, high quality, fine-grainedness, and customization. By including four typical types of face retouching operations and different retouching levels, we extend the binary face retouching detection into a fine-grained, multi-retouching type, and multi-retouching level estimation problem. Additionally, we propose a Multi-granularity Attention Module (MAM) as a plugin for CNN backbones for enhanced cross-scale representation learning. Extensive experiments using different baselines as well as our proposed method on RetouchingFFHQ show decent performance on face retouching detection.
Qichao Ying, Sheng Li 0006, Haisheng Xu, Zhenxing Qian, Xinpeng Zhang 0001
ACM Multimedia1
2023 Learning to Immunize Images for Tamper Localization and Self-Recovery
abstract
Digital images are vulnerable to nefarious tampering attacks such as content addition or removal that severely alter the original meaning. It is somehow like a person without protection that is open to various kinds of viruses. Image immunization (Imuge) is a technology of protecting the images by introducing trivial perturbation, so that the protected images are immune to the viruses in that the tampered contents can be auto-recovered. This paper presents Imuge+, an enhanced scheme for image immunization. By observing the invertible relationship between image immunization and the corresponding self-recovery, we employ an invertible neural network to jointly learn image immunization and recovery respectively in the forward and backward pass. We also introduce an efficient attack layer that involves both malicious tamper and benign image post-processing, where a novel distillation-based JPEG simulator is proposed for improved JPEG robustness. Our method achieves promising results in real-world tests where experiments show accurate tamper localization as well as high-fidelity content recovery. Additionally, we show superior performance on tamper localization compared to state-of-the-art schemes based on passive forensics.
Qichao Ying, Hang Zhou 0007, Zhenxing Qian, Sheng Li 0006, Xinpeng Zhang 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2022 Hiding Images Into Images with Real-World Robustness
abstract
The existing image embedding networks are basically vulnerable to malicious attacks such as JPEG compression and noise adding, not applicable for real-world copyright protection tasks. To solve this problem, we introduce a generative deep network based method for hiding images into images while assuring high-quality extraction from the destructive synthesized images. An embedding network is sequentially concatenated with an attack layer, a decoupling network and an image extraction network. The addition of decoupling network learns to extract the embedded secret image from the attacked image. We also pinpoint the weaknesses of the adversarial training for robustness in previous works and build our improved real-world attack simulator. Experimental results demonstrate the superiority of the proposed method against typical digital attacks by a large margin, as well as the performance boost of the recovered images with the aid of progressive recovery strategy. Besides, we are the first to robustly hide three secret images.
Qichao Ying, Hang Zhou 0007, Xianhan Zeng, Haisheng Xu, Zhenxing Qian, Xinpeng Zhang 0001
ICIP1
2022 RWN: Robust Watermarking Network for Image Cropping Localization
abstract
Image cropping can be maliciously used to manipulate the layout of an image and alter the underlying meaning. Previous image cropping detection schemes only predict whether an image has been cropped, ignoring which part of the image is cropped. This paper presents a novel robust watermarking network for image cropping localization. We train an anti-cropping processor (ACP) that embeds a watermark into a target image. The visually indistinguishable protected image is then posted on the social network instead of the original image. At the recipient’s side, ACP extracts the watermark from the attacked image, and we conduct feature matching on the original and extracted watermark to locate the position of the cropping. We further extend our scheme to detect tampering attacks on the attacked image, and a simple yet efficient method (JPEG-Mixup) is proposed that noticeably improves the generalization of JPEG robustness. We demonstrate that our scheme is the first to provide high-accuracy and robust image cropping localization.
Qichao Ying, Xiaoxiao Hu, Zhenxing Qian, Sheng Li 0006, Xinpeng Zhang 0001
ICIP1
2022 Invertible Image Dataset Protection
abstract
The security of data storage is a big issue for companies. They must take effective steps to prevent valuable image datasets from being stolen for illegal commercial purposes. While data encryption is a common solution, it drastically down-grades the visual quality and therefore forbids common yet trivial use such as eye-checking without a decryption. We present a novel solution for dataset protection in this scenario by robustly and reversibly transform the images into adver-sarial images. An invertible Image Dataset Protection NET-work (IDP-Net) is developed to introduce slight and acceptable changes to the images within the dataset. The protected images can be published and circulated on the social networks instead of their original version. Malicious attackers can only observe the images but cannot train pirated models based on them. Meanwhile, IDP-Net ensures the performance of au-thorized models, namely, trusted users can revert the protection and retrieve the protected images to their original version. Therefore, the dataset can be stored within the pro-tected version alone to ensure safety. Extensive experiments demonstrate that IDP-Net can better protect the security of image dataset against defensive methods compared to previ-ous methods. Besides, the introduced distortion is acceptable and the original images can be reconstructed nearly error-free.
Kejiang Chen, Xianhan Zeng, Qichao Ying, Sheng Li 0006, Zhenxing Qian, Xinpeng Zhang 0001
ICME3
2022 Image Generation Network for Covert Transmission in Online Social Network
abstract
Online social networks have stimulated communications over the Internet more than ever, making it possible for secret message transmission over such noisy channels. In this paper, we propose a Coverless Image Steganography Network, called CIS-Net, that synthesizes a high-quality image directly conditioned on the secret message to transfer. CIS-Net is composed of four modules, namely, the Generation, Adversarial, Extraction, and Noise Module. The receiver can extract the hidden message without any loss even the images have been distorted by JPEG compression attacks. To disguise the behaviour of steganography, we collected images in the context of profile photos and stickers and train our network accordingly. As such, the generated images are more inclined to escape from malicious detection and attack. The distinctions from previous image steganography methods are majorly the robustness and losslessness against diverse attacks. Experiments over diverse public datasets have manifested the superior ability of anti-steganalysis.
Zhengxin You, Qichao Ying, Sheng Li 0006, Zhenxing Qian, Xinpeng Zhang 0001
ACM Multimedia2
2022 Robust Watermarking for Video Forgery Detection with Improved Imperceptibility and Robustness
abstract
Videos are prone to tampering attacks that alter the meaning and deceive the audience. Previous video forgery detection schemes find tiny clues to locate the tampered areas. However, attackers can successfully evade supervision by destroying such clues using video compression or blurring. This paper proposes a video watermarking network for tampering localization. We jointly train a 3D-UNet-based watermark embedding network and a decoder that predicts the tampering mask under simulated attacks. The perturbation made by watermark embedding is close to imperceptible. Considering that there is no off-the-shelf differentiable video codec simulator, we propose to mimic video compression by ensembling simulation results of other typical attacks, e.g., JPEG compression and blurring, as an approximation. Experimental results demonstrate that our method generates watermarked videos with good imperceptibility and robustly and accurately locates tampered areas within the attacked version.
Yangming Zhou, Qichao Ying, Zhenxing Qian, Xinpeng Zhang 0001
MMSP2
2022 High-Capacity Framework for Reversible Data Hiding in Encrypted Image Using Pixel Prediction and Entropy Encoding
abstract
While the existing reserving room before encryption (RRBE) based reversible data hiding in encrypted image (RDHEI) schemes can achieve decent embedding capacity, the capacity of the existing vacating room by encryption (VRBE) based schemes is relatively low. To address this issue, this paper proposes a generalized framework for high-capacity RDHEI for both the RRBE and VRBE cases. First, an efficient embedding room generation algorithm (ERGA) is designed to produce large embedding room using pixel prediction and entropy encoding. Then, we propose two RDHEI schemes, one for RRBE, another for VRBE. In the RRBE scenario, the image owner generates the embedding room with ERGA and encrypts the preprocessed image using stream cipher with two encryption keys. Then, the data hider locates the embedding room and embeds the additional encrypted data. In the VRBE scenario, the cover image is encrypted by an improved block modulation and permutation encryption algorithm, where the spatial redundancy in the plain-text image is greatly preserved. Then, the data hider applies ERGA on the encrypted image to generate the embedding room and conducts data embedding. For both schemes, receivers with different authentication keys can conduct either error-free data extraction or error-free image recovery. The experimental results show that the two proposed schemes outperform many state-of-the-art RDHEI schemes. Besides, they can ensure high security level, where the original image can be hardly discovered from the encrypted version before or after data hiding by unauthorized users.
Yingqiang Qiu, Qichao Ying, Yuyan Yang, Huanqiang Zeng, Sheng Li 0006, Zhenxing Qian
IEEE Trans. Circuits Syst. Video Technol.2
2021 From Image to Imuge: Immunized Image Generation
abstract
We introduce Imuge, an image tamper resilient generative scheme for image self-recovery. The traditional manner of concealing image content within the image are inflexible and fragile to diverse digital attack, i.e. image cropping and JPEG compression. To address this issue, we jointly train a U-Net backboned encoder, a tamper localization network and a decoder for image recovery. Given an original image, the encoder produces a visually indistinguishable immunized image. At the recipient's side, the verifying network localizes the malicious modifications, and the original content can be approximately recovered by the decoder, despite the presence of the attacks. Several strategies are proposed to boost the training efficiency. We demonstrate that our method can recover the details of the tampered regions with a high quality despite the presence of various kinds of attacks. Comprehensive ablation studies are conducted to validate our network designs.
Qichao Ying, Zhenxing Qian, Hang Zhou 0007, Haisheng Xu, Xinpeng Zhang 0001
ACM Multimedia1
2021 Steganography in animated emoji using self-reference
abstract
Abstract Animated emoji is a kind of GIF image, which is widely used in online social networks (OSN) for its efficiency in transmitting vivid and personalized information. Aiming at realizing covert communication in animated emoji, this paper proposes an improved steganography framework in animated emoji. We propose a self-reference algorithm to improve the steganography security. Meanwhile, the relations between adjacent frames of the cover GIF image are considered to further improve the distortion function. After that we embed the secret message into the GIF image using the popular framework of Syndrome Trellis Coding (STC). Experimental results show that the proposed method can provide better security performances than state-of-the-art works.
Zhiying Zhu 0001, Qichao Ying, Zhenxing Qian, Xinpeng Zhang 0001
Multim. Syst.2