EDBT 2026 Demo / reviewers in the wild / expert
Yuanzhi Yao
dblp:121/1222
· DBLP profile ↗
34ranked-venue papers
7as first author
26since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 15 · 4 first-author · 9 since 2021Databases, data management, data science and information retrieval · 6 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Computer networks · 5 · 5 since 2021Security and privacy · 5 · 1 first-author · 4 since 2021Theory of computation · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Universal Facial Landmark Detection by Landmark-Clustering Relation-Reasoning Transformer
Jun Wan 0005, Yuanzhi Yao, Jiaxing Huang 0001, Xiaoying Ding, Lefei Zhang, Yongsheng Gao 0001, Dacheng Tao |
Int. J. Comput. Vis. | 2 |
| 2026 | Interpretable facial landmark detection by multi-expert collaborative uncertainty-aware deep networks
Jun Wan 0005, Hui Xi, Yuanzhi Yao, Zhihui Lai 0001, Jie Zhou 0009 |
Neural Networks | 3 |
| 2026 | Precise Temporal Forgery Localization via Quantified Audio-Visual AsynchronyabstractTemporal Forgery Localization (TFL) aims to identify the precise temporal boundaries of manipulated segments within videos. This represents a critical advancement beyond binary video-level forgery detection, because the latter is often insufficient for combating sophisticated partial forgeries that insert synthetic content into otherwise authentic media. The generation pipeline of such forgeries introduces two measurable artifacts: (1) audiovisual asynchrony resulting from imperfect lip-speech synchronization, and (2) abrupt transitions occurring at splice points. Current TFL approaches rely on architectures adapted from semantic tasks that implicitly learn forgery cues, limiting their precision in boundary detection. To address this issue, we propose a novel framework that explicitly quantifies audiovisual asynchrony as a direct signal for localization. Our approach utilizes a Coupled Pyramidal Encoder to extract multi-scale synchronized representations across modalities. These features feed into a Multi-Scale Asynchrony Probe that measures the temporal warping cost required for audiovisual alignment, translating desynchronization into a quantifiable forgery indicator. This measured asynchrony then guides our Context-Aware Boundary Pinpointing module to selectively amplify manipulation-related discontinuities while suppressing benign scene changes. Experiments on LAV-DF and Deepfake1M benchmarks demonstrate that our artifact-centric design achieves state-of-the-art performance, improving high-precision localization ([email protected]) by up to 27.5 points over previous methods. These results validate that explicitly quantifying asynchrony provides a powerful guiding signal for precise temporal forgery localization. Yuanzhi Yao, Yunfeng Diao |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2025 | Semi-VFL: Communication-efficient Few-label Vertical Federated Learning with Stacked Generalization and Model-level ConsistencyabstractVertical federated learning (VFL) is a collaborative learning scheme where clients share some overlapping samples but have different feature spaces. Existing VFL schemes are restricted in model performance and deployment feasibility due to the scarcity of overlapping labeled samples and high communication costs. To tackle these issues, we propose a practical VFL scheme Semi-VFL using stacked generalization, which can effectively improve model performance with limited aligned labeled samples and only two communication times. A built-in local semi-supervised learning strategy FewMatch with model-level consistency for few-label VFL setting is designed in our scheme. Extensive experiments indicate the superiority of Semi-VFL in both image and tabular datasets. Specifically, Semi-VFL can achieve accuracy improvement by more than 10.9% and communication cost reduction by more than 380× over the state-of-the-art few-shot VFL scheme on CIFAR-10 with 128 aligned samples. Xuan Jin, Yuanzhi Yao, Caihong Kai, Rui Wang 0043, Nenghai Yu |
ICASSP | 2 |
| 2025 | Text-Infused Audio-Visual Video Parsing with Semantic-Aware Multimodal Contrastive LearningabstractThe Audio-Visual Video Parsing task aims to recognize events occurring in video segments for each modality. Presently, the excellent performance in handling video parsing is shown by generating pseudo labels at the segment level. However, these approaches still suffer from adequate semantic learning of fine-grained segment features, which can cause errors in the prediction of those ambiguous events and affect the predictions of another modality. To tackle this issue, we propose a novel Text-Infused Parsing Network (TIPNet). Specifically, the event text modality, which can offer more precise semantics than the audio and visual modalities, is introduced to enhance the event-related audio/visual segment features by encoding the cross-modal interactions. Furthermore, to obtain more discriminative segment features for each modality, we propose a novel multimodal contrastive loss with a semantic-aware weighting mechanism. Experimental results on the benchmark dataset demonstrate the superior performance of our approach compared to the state-of-the-art methods. Dan Guo 0001, Yuanzhi Yao |
ICASSP | 4 |
| 2025 | Mining Topics towards ChatGPT Using a Disentangled Contextualized-neural Topic ModelabstractMining topics relevant to the advanced AI dialogue system, such as ChatGPT, from short-length posts on social media poses several challenges for existing topic-mining approaches. Firstly, Bag-Of-Words approaches, including probabilistic topic models and their embedding-based variants, may struggle to extract interpretable topics due to insufficient word co-occurrence. Secondly, contextualized based approaches, built on the autoencoding framework, often yield entangled topic spaces, resulting in the mixing of irrelevant words into topics. To address these limitations, we propose a novel Dis entangled Contextualized-neural Topic Model (DisCTM) based on textual representation learning. DisCTM leverages a pre-trained transformer language model to incorporate word sequence information and deal with the sparsity in short text. Additionally, it employs a topic disentangling mechanism to decorrelate dimensions of the latent topic space, effectively separating semantically irrelevant words into different topics. Extensive experiments have been conducted on three publicly available text corpora, and the results demonstrate the effectiveness of DisCTM in extracting high-quality topics, as measured by topic coherence and diversity metrics. Rui Wang 0043, Shuyu Chang, Yuanzhi Yao, Haiping Huang |
WSDM | 5 |
| 2025 | Mining User Preferences from Online Reviews with the Genre-aware Personalized Neural Topic ModelabstractCustomer-generated reviews on e-commerce websites often contain valuable insights into users' interests in product genres and provide a rich source for mining user preferences. However, most existing neural topic models tend to generate meaningless topics that share low correlations with product genres. Furthermore, they often fail to mine user preferences and discover personalized topic profiles due to the absence of explicit user modeling. To address these limitations, we propose a novel Genre-aware Personalized neural Topic Model (GPTM), which incorporates product genre information into the topic modeling process to ensure the relevance between mined topics and product genres. Moreover, it could produce a personalized topic profile for each user by performing user preference modeling. Extensive experimental results on three publicly available Amazon review corpora validate the effectiveness of the proposed GPTM in genre-aware topic modeling. Furthermore, GPTM surpasses state-of-the-art baselines in user preference mining and generates high-quality personalized topic profiles. Rui Wang 0043, Xincheng Lv, Shuyu Chang, Yansheng Wu, Yuanzhi Yao, Haiping Huang, Guozi Sun |
WWW | 6 |
| 2025 | An online tool with Google Earth Engine and cellular automata for seamlessly simulating global urban expansion at high resolutionsabstractProjecting global urban expansion is crucial for environmental assessment under climate change scenarios. However, existing global future urban land products are typically provided at coarse resolutions (1 km) due to data and computing limitations. This hinders the accurate assessment of the impacts of global urban development at finer scales. Thus, we develop the first Cellular Automata (CA) online tool for simulating future global urban expansion in Google Earth Engine (GEE-CA), which can simulate future urban land change at a 30 m resolution under different SSP scenarios. GEE-CA enables seamless simulations of future urban land at high resolution through a partitioned parallel strategy. Seven large urban agglomerations are simulated under shared socioeconomic pathways as representative regions to present our fine-scale results. Comparatively, our datasets preserve significantly more spatial details than existing global urban land products. The improvement in resolution from 1 km to 30 m reduces errors by a range from 7.11% to 21.27% in the estimation of future urban area. So far, the proposed GEE-CA tool allows users to generate urban land projection maps for any defined region with a resolution as high as 30 m. Guohua Hu, Yuanzhi Yao |
Int. J. Geogr. Inf. Sci. | 4 |
| 2025 | Searchable face recognition authentication based on homomorphic encryption
Baiqi Wu, Shuli Zheng, Peiming Dai, Jiazheng Chen, Yuanzhi Yao |
J. Inf. Secur. Appl. | 5 |
| 2025 | IDCNet: Image Decomposition and Cross-View Distillation for Generalizable Deepfake DetectionabstractExisting deepfake detectors predominantly process entire facial images as input, which limits their sensitivity to local forgery cues due to representation bias and information loss through CNN feature aggregation. To address these limitations, we propose IDCNet, a novel deepfake detection framework based on image decomposition and cross-view distillation. Our key insight is that decomposing images into complementary views enables specialized processing of global and local forgery cues, while cross-view distillation facilitates their mutual enhancement. Specifically, the framework employs a lightweight U-Net generator with a dual-objective mechanism to decompose input images into global content and local detail views, optimized through reconstruction and classification losses. A cross-view distillation strategy is then applied to enhance complementary feature learning between views. Furthermore, to integrate local artifact information into existing detection models without architectural modifications, we propose a feature alignment method. Extensive experiments across 14 forgery methods demonstrate the effectiveness of our approach, achieving up to 4.4% AUC improvement on the CDFV2 dataset compared to state-of-the-art methods. The source code is available at: https://github.com/ wangzhiyuan120/idcnet. Yuanzhi Yao, Wenpeng Xing, Meng Li 0006 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2024 | PPIDSG: A Privacy-Preserving Image Distribution Sharing Scheme with GAN in Federated LearningabstractFederated learning (FL) has attracted growing attention since it allows for privacy-preserving collaborative training on decentralized clients without explicitly uploading sensitive data to the central server. However, recent works have revealed that it still has the risk of exposing private data to adversaries. In this paper, we conduct reconstruction attacks and enhance inference attacks on various datasets to better understand that sharing trained classification model parameters to a central server is the main problem of privacy leakage in FL. To tackle this problem, a privacy-preserving image distribution sharing scheme with GAN (PPIDSG) is proposed, which consists of a block scrambling-based encryption algorithm, an image distribution sharing method, and local classification training. Specifically, our method can capture the distribution of a target image domain which is transformed by the block encryption algorithm, and upload generator parameters to avoid classifier sharing with negligible influence on model performance. Furthermore, we apply a feature extractor to motivate model utility and train it separately from the classifier. The extensive experimental results and security analyses demonstrate the superiority of our proposed scheme compared to other state-of-the-art defense methods. The code is available at https://github.com/ytingma/PPIDSG. Yuting Ma 0001, Yuanzhi Yao, Xiaohua Xu 0002 |
AAAI | 2 |
| 2024 | A parallel set-based model on the shortest travel time in long-distance transportation systemsabstractThe shortest travel time in long-distance transportation systems (LDTS) is a key indicator for measuring regional connectivity and mapping accessibility at national and global levels. For traditional methods, it is a great challenge to calculate the shortest travel time with millions of origin-destination pairs and evaluate the overall performance. To fill this technique gap, this study proposed a novel model to get all the trips with different numbers of transfers through a series of set-based methods and then calculate all the shortest travel time between stops. The set-based model was tested for calculating the shortest travel time in LDTS of China. Using the optimization algorithm, we can significantly reduce the number of transfer trips that need to be calculated. For instance, the number of transfer trips decreased by 96.17%, 98.15%, and 79.02%, for conventional railway, high-speed railway, and air transportation, respectively, when the number of transfers was two. The set-based model can be extended to calculate the door-to-door travel time between places, for instance, mapping the fine-scale accessibility at the national level. Furthermore, we proved that this set-based model, proposed in this study, could also be parallelized and applied to any other LDTS in the General Transit Feed Specification format. Han Zhang 0033, Yuanzhi Yao, Xia Li 0001 |
Int. J. Geogr. Inf. Sci. | 2 |
| 2024 | Efficient and Emission-Reducing Blockchain-Enabled Multi-UAV-Assisted MEC System in IoT NetworksabstractIn highly interconnected large-scale event and other Internet of Things (IoT) device-intensive scenarios, traditional terrestrial base stations have difficulty meeting the requirements of IoT devices for network speed and security, and have exacerbated carbon pollution. To this end, a blockchain-enabled unmanned aerial vehicles (UAVs)-assisted mobile edge computing (MEC) system is introduced to enhance communication efficiency and ensure the privacy of IoT devices. In this system, the Byzantine consensus algorithm is applied in the blockchain. Considering the pollution of reducing carbon dioxide emissions, a strategy for jointly optimizing the flight trajectories of UAVs, task offloading scheduling, and MEC computing resource allocation is formulated to minimize the system’s carbon emissions and time delay while meeting MEC and blockchain computing tasks. However, due to the coupling of variables, this problem is very complex. Therefore, the original problem is decoupled into multiple subproblems, and the block coordinate descent method (BCD) and successive convex approximation method (SCA) are used for solving. Specifically, the UAV flight trajectories, task offloading scheduling, and MEC computing resource allocation are alternately optimized until convergence. Simulation results verify the effectiveness and good performance of the proposed algorithm in this article. Lisu Yu, Biao Li 0003, Yuanzhi Yao, Zhen Wang 0022, Zhicheng Dong 0003, Donghong Cai |
IEEE Internet Things J. | 3 |
| 2024 | Vocoder Detection of Spoofing Speech Based on GAN Fingerprints and Domain GeneralizationabstractAs an important part of the text-to-speech (TTS) system, vocoders convert acoustic features into speech waveforms. The difference in vocoders is key to producing different types of forged speech in the TTS system. With the rapid development of general adversarial networks (GANs), an increasing number of GAN vocoders have been proposed. Detectors often encounter vocoders of unknown types, which leads to a decline in the generalization performance of models. However, existing studies lack research on detection generalization based on GAN vocoders. To solve this problem, this study proposes vocoder detection of spoofed speech based on GAN fingerprints and domain generalization. The framework can widen the distance between real speech and forged speech in feature space, improving the detection model’s performance. Specifically, we utilize a fingerprint extractor based on an autoencoder to extract GAN fingerprints from vocoders. We then weight them to the forged speech for subsequent classification to learn the forged speech features with high differentiation. Subsequently, domain generalization is used to further improve the generalization ability of the model for unseen forgery types. We achieve domain generalization using domain-adversarial learning and asymmetric triplet loss to learn a better generalized feature space in which real speech is compact and forged speech synthesized by different vocoders is dispersed. Finally, to optimize the training process, curriculum learning is used to dynamically adjust the contributions of the samples with different difficulties in the training process. Experimental results show that the proposed method achieves the most advanced detection results among four GAN vocoders. The code is available at https://github.com/multimedia-infomation-security/GAN-Vocoder-detection . Zuxing Zhao, Yuanzhi Yao, Xin Liao 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2023 | Image Translation-Based Deniable Encryption against Model Extraction AttackabstractIn cloud storage applications, data owners’ original images are usually encrypted before being outsourced to the cloud for preserving data owners’ privacy. However, in deep learning model-based image encryption methods, an adversary can conduct the model extraction attack to reveal the model parameters and thus restore the privacy information by obtaining numerous encrypted images. In this paper, we propose an image translation-based deniable encryption (ITDE) scheme to achieve encryption deniability and defend against model extraction attacks. Differing from traditional encryption methods in which encrypted images are visually meaningless, ITDE applies image translation to generate encrypted images in the form of human faces. Moreover, ITDE provides deniability for data owners to keep the encryption parameters private. To defend against model extraction attacks, the defense mechanism is introduced in our proposed ITDE to preserve deep learning models. Experimental results demonstrate the superiority of our proposed methods in terms of encryption deniability and privacy preservation. Yiling Chen 0009, Yuanzhi Yao, Nenghai Yu |
ICIP | 2 |
| 2023 | High-fidelity video reversible data hiding using joint spatial and temporal predictionabstractAn efficient predictor is essential in reversible data hiding methods. This paper proposes a video reversible data hiding method, where the correlations between pixels in both spatial and temporal domains are fully considered to improve prediction accuracy. In more detail, for each pixel, both intra prediction and inter prediction are implemented and the better prediction mode is selected according to intra or inter local complexity. A double-layered video partition scheme is applied so that the motion vectors generated during inter prediction are not necessarily embedded as additional information while reversibility is guaranteed. In addition, a cover pixel selection principle based on the probability distribution of the prediction errors is proposed, which prevents the pixels with large prediction errors from being modified. Experimental results demonstrate that the joint spatial and temporal prediction scheme offers satisfactory accuracy, and the proposed video reversible data hiding method can obtain high-quality stego videos. Lincong Li, Yuanzhi Yao, Nenghai Yu |
Signal Process. | 2 |
| 2022 | Privacy-preserving Collaborative Learning with Scalable Image Transformation and AutoencoderabstractCollaborative learning in which local clients jointly train a deep learning model by sharing parameters to the central- ized server has gained great popularity. However, recent works have shown that local private data can be leaked to the server by gradient sharing. In this paper, a privacy-preserving collaborative learning scheme is proposed to defend against gradient-based reconstruction attacks. The sensitive training images are firstly permutated by transformation with scalable block sizes. Then, features of permutated images are extracted by a classification- compliant autoencoder for meaningful representation of high- dimensional images and facilitating classification. The model accuracy constraint is incorporated in the training process to maintain decent classification accuracy. Experimental results demonstrate that the proposed scheme can achieve high privacy preservation with minimal impact on model accuracy. Yuting Ma 0001, Yuanzhi Yao, Nenghai Yu |
GLOBECOM | 2 |
| 2022 | ATDD: Fine-Grained Assured Time-Sensitive Data Deletion Scheme in Cloud StorageabstractWith the rapid development of general cloud services, more and more individuals or collectives use cloud platforms to store data. Assured data deletion deserves investigation in cloud storage. In time-sensitive data storage scenarios, it is necessary for cloud platforms to automatically destroy data after the data owner-specified expiration time. Therefore, assured time-sensitive data deletion should be sought. In this paper, a fine-grained assured time-sensitive data deletion (ATDD) scheme in cloud storage is proposed by embedding the time trapdoor in Ciphertext-Policy Attribute-Based Encryption (CP-ABE). Time-sensitive data is self-destructed after the data owner-specified expiration time so that the authorized users cannot get access to the related data. In addition, a credential is returned to the data owner for data deletion verification. This proposed scheme provides solutions for fine-grained access control and verifiable data self-destruction. Detailed security and performance analysis demonstrate the security and the practicability of the proposed scheme. Zhengyu Yue, Yuanzhi Yao, Weihai Li, Nenghai Yu |
ICC | 2 |
| 2022 | Privacy-preserving Cost-sensitive Federated Learning from Imbalanced DataabstractFederated learning allows multiple clients to collab-oratively train a global deep learning model without revealing their local data to a centralized server. However, the existence of clients whose datasets have imbalanced class distribution has a significant impact on model accuracy. Imbalance makes it challenging for a model to distinguish between the majority and minority classes without accessing clients' local data. In this paper, we aim to tackle this problem by privacy-preserving cost-sensitive federated learning. We design the joint cost-sensitive and differentially private model parameter optimization mechanism which maintains the model accuracy while satisfying differential privacy constraints. Moreover, this mechanism does not alter the original data distribution. Experimental results demonstrate the superior performance of our proposed scheme in terms of model accuracy and privacy preservation. Yuanzhi Yao, Yuting Ma 0001, Nenghai Yu |
ICTAI | 2 |
| 2022 | Fuzzy Keyword Search over Encrypted Cloud Data with Dynamic Fine-grained Access ControlabstractDue to the increasing popularity of cloud computing and privacy preservation concerns, sensitive data should be encrypted before outsourcing to the cloud and data utilization becomes a challenging issue. Searchable encryption (SE) is a promising technique to address this problem. Most existing searchable encryption schemes only support accurate keyword but fuzzy keyword search schemes are appreciated in practice. Moreover, in some application scenarios like the video on demand systems, data owners only hope to give the access of their data to those who have payed but authenticated user identity may often change. Therefore, dynamic user attribute updating should be considered. To solve about issues, we propose a fuzzy secure keyword search scheme over encrypted cloud data with dynamic fine-grained access control. We design a novel fuzzy keyword index to retrieve corresponding documents. To reduce the computation cost, the cloud server selects most relevant top-k documents and return them to the data users. The ciphertext-policy attribute based encryption (CP-ABE) technique is introduced to implement fine-grained access control. Meanwhile, the basic CP-ABE is improved to meet the practical need of user attribute updating. Extensive security and performance analysis demonstrates that our proposed scheme is highly efficient and can satisfy the security requirements for fuzzy keyword search over encrypted cloud data. Boshen Shan, Yuanzhi Yao, Weihai Li, Xiaodong Zuo, Nenghai Yu |
TrustCom | 2 |
| 2022 | An efficient multiple scanning order algorithm for accumulative least-cost surface calculationabstractThe least-cost surface (LCS) calculation is a compute-intensive problem conventionally solved by the queue-based Dijkstra’s algorithm. Alternative raster-based scanning algorithms have also been proposed which use a moving window to scan the whole study area iteratively. Here we propose improvements to the raster-based algorithms. The main improvement is to implement multiple scanning orders (MSO) to replace the conventional single scanning order (SSO, typically from upper-left corner to lower-right corner, row by row). We compared the performance of different algorithms over different cost surfaces and with different numbers of source points. The comparison shows that a raster-based algorithm adopting MSO has a substantially better performance than a conventional raster-based algorithm using SSO. An MSO raster-based algorithm is generally comparable to the queue-based Dijkstra’s algorithm, and surpasses the latter over a relatively simple cost surface (e.g. in which the cost is resampled) and/or when the number of source points is relatively large. Our empirical experiments suggest that MSO reduces the time complexity from to Θ(N2) to Θ(NlogN). Additionally, we found that the MSO raster-based algorithm can be easily parallelized using shared-memory parallel programming. Yuanzhi Yao, Xun Shi |
Int. J. Geogr. Inf. Sci. | 1 |
| 2022 | RIS-assisted secure UAV communications with resource allocation and cooperative jammingabstractAbstract Unmanned aerial vehicles (UAVs) are widely used in wireless communication networks due to their rapid deployment and high mobility. However, in practical scenarios, the existence of obstacles and eavesdroppers will seriously interfere with the communication quality of the UAV network and produce a security risk. Thus, this paper combines reconfigurable intelligent surface (RIS) technology with UAVs to build a secure UAV communication network. Normally, a rotary‐wing UAV (labeled as UAV‐S) acting as a base station sends information signals to a legitimate user on the ground with RIS equipment. However, there is a passive eavesdropper on the ground who can steal the information. Therefore, a friendly UAV jammer (labeled as UAV‐J) with a fixed location is introduced to send jamming signals to confuse the eavesdropper. The goal of this paper is to maximize the average secrecy rate of the communication network by jointly optimizing the flight trajectory, transmit power of the UAV‐S and UAV‐J, and phase shifter of the RIS. Since the constructed problem is highly nonconvex, an alternating optimization algorithm based on successive convex approximation techniques is proposed to solve the problem. Simulation results show that the proposed algorithm can achieve a higher secrecy rate in comparison with other schemes. Jichang Guo, Lisu Yu, Zhiqiong Chen, Yuanzhi Yao, Zhen Wang 0022, Zhenghai Wang, Qingmin Zhao |
IET Commun. | 4 |
| 2021 | Convolutional Neural Network-driven Optimal Prediction for Image Reversible Data HidingabstractReversible data hiding aims to embed data into cover digital media in a reversible way. The key issue of image reversible data hiding is to construct the sharply distributed prediction error histogram using advanced pixel prediction. Inspired by the progress of image super-resolution exploiting convolutional neural networks (CNN), CNN predictors can improve the prediction accuracy compared with conventional predictors generally. However, CNN predictors fail to achieve the best prediction accuracy in some cases due to the dependence on training data. To remedy this drawback, the CNN-driven optimal prediction for image reversible data hiding is proposed in this paper. Instead of only utilizing one specific predictor for prediction error histogram construction, the optimal prediction mechanism is designed by incorporating CNN predictors and conventional predictors. Extensive experiments demonstrate the merits of the proposed method in terms of prediction accuracy and marked image quality. Yuanzhi Yao, Nenghai Yu |
MMSP | 2 |
| 2021 | Nearly Reversible Image-to-Image Translation Using Joint Inter-Frame Coding and EmbeddingabstractImage-to-image translation tasks which have been widely investigated with generative adversarial networks (GAN) aim to map an image from the source domain to the target domain. The translated image can be inversely mapped to the reconstructed source image. However, existing GAN-based schemes lack the ability to accomplish reversible translation. To remedy this drawback, a nearly reversible image-to-image translation scheme where the reconstructed source image is approximately distortion-free compared with the corresponding source image is proposed in this paper. The proposed scheme jointly considers inter-frame coding and embedding. Firstly, we organize the GAN-generated reconstructed source image and the source image into a pseudo video. Furthermore, the bitstream obtained by inter-frame coding is reversibly embedded in the translated image for nearly lossless source image reconstruction. Extensive experimental results and analysis demonstrate that the proposed scheme can achieve a high level of performance in image quality and security. Xinzhu Cao, Yuanzhi Yao, Nenghai Yu |
VCIP | 2 |
| 2021 | Motion vector modification distortion analysis-based payload allocation for video steganography
Yuanzhi Yao, Nenghai Yu |
J. Vis. Commun. Image Represent. | 1 |
| 2021 | Cryptanalysis of the RSA variant based on cubic Pell equation
Mengce Zheng, Noboru Kunihiro, Yuanzhi Yao |
Theor. Comput. Sci. | 3 |
| 2020 | Defining Embedding Distortion for Sample Adaptive Offset-Based HEVC Video SteganographyabstractAs a newly added in-loop filtering technique in High Efficiency Video Coding (HEVC), sample adaptive offset (SAO) can be utilized to embed messages for video steganography. This paper presents a novel SAO-based HEVC video steganographic scheme. The main principle is to design a suitable distortion function which expresses the embedding impacts on offsets based on minimizing embedding distortion. Two factors including the sample rate-distortion cost fluctuation and the sample statistical characteristic are considered in embedding distortion definition. Adaptive message embedding is implemented using syndrome-trellis codes (STC). Experimental results demonstrate the merits of the proposed scheme in terms of undetectability and video coding performance. Yabing Cui, Yuanzhi Yao, Nenghai Yu |
MMSP | 2 |
| 2019 | Content-adaptive reversible visible watermarking in encrypted images
Yuanzhi Yao, Weiming Zhang 0001, Hang Zhou 0007, Nenghai Yu |
Signal Process. | 1 |
| 2019 | Distortion Design for Secure Adaptive 3-D Mesh SteganographyabstractWe propose a novel technique for steganography on 3-D meshes so as to resist steganalysis. The majority of existing methods modulate vertex coordinates to embed messages in a nonadaptive way. We take account of complexity of local regions as joint distortion of a triple unit (vertice) and coding method such as syndrome trellis codes to adaptively embed messages, which owns stronger security with respect to existing steganalysis. Key to the distortion is a novel formulation of adaptive steganography, which relies on some effective steganalytic features such as variation of vertex normal. We provide quantitative and qualitative comparisons of our method with several baselines against steganalytic features LFS64, LFS76, and ensemble classifiers, and show that it outperforms the current state of the art. Meanwhile, we proposed an attacking method on steganography proposed by Chao et al. (2009) with a high detection rate. Hang Zhou 0007, Kejiang Chen, Weiming Zhang 0001, Yuanzhi Yao, Nenghai Yu |
IEEE Trans. Multim. | 4 |
| 2016 | Reversible Data Hiding for Texture Videos and Depth Maps Coding with Quality Scalability
Yuanzhi Yao, Weiming Zhang 0001, Nenghai Yu |
IWDW | 1 |
| 2016 | Inter-frame distortion drift analysis for reversible data hiding in encrypted H.264/AVC video bitstreams
Yuanzhi Yao, Weiming Zhang 0001, Nenghai Yu |
Signal Process. | 1 |
| 2015 | Alternating scanning orders and combining algorithms to improve the efficiency of flow accumulation calculationabstractConventionally, a raster operation that needs to scan the entire image employs only one scanning order (i.e., single scanning order (SSO)), and the scan usually runs from upper left to lower right and row by row. We explore the idea of alternately applying multiple scanning orders (MSO) to raster operations that are based on the local direction, using the flow accumulation (FA) calculation as an example. We constructed several FA methods based on MSO, and compared them with those widely used methods. Our comparison includes experiments over digital elevation models (DEMs) of different landforms and DEMs of different resolutions. For each DEM, we calculated both single-direction FA (SD-FA) and multi-direction FA (MD-FA). In the theoretical aspect, we deducted the time complexity of an MSO sequential algorithm (MSOsq) for FA based on empirical equations in hydrology. Findings from the experiments include the following: (1) an MSO-based method is generally superior to its counterpart SSO-based method. (2) The advantage of MSO is more significant in the SD-FA calculation than in the MD-FA calculation. (3) For SD-FA, the best method among the compared methods is the one that combines the MSOsq and the depth-first algorithm. This method surpasses the commonly recommended dependency graph algorithm, in both speed and memory use. (4) The differences between the compared methods are not sensitive to specific landforms. (5) For SD-FA, the advantage of MSO-based methods is more obvious in a higher DEM resolution, but this does not apply to MD-FA. Yuanzhi Yao, Xun Shi |
Int. J. Geogr. Inf. Sci. | 1 |
| 2015 | Defining embedding distortion for motion vector-based video steganography
Yuanzhi Yao, Weiming Zhang 0001, Nenghai Yu, Xianfeng Zhao |
Multim. Tools Appl. | 1 |
| 2013 | MOS-Based Channel Allocation Schemes for Mixed Services over Cognitive Radio NetworksabstractIn cognitive radio (CR) networks, secondary users (SUs) may have various applications such as multimedia delivery and file download, resulting in different bandwidth requirements. In this paper, we propose a channel allocation scheme for mixed services, especially video streaming, based on mean opinion score (MOS) maximization. MOS is an effective metric of Quality of Experience (QoE) that directly measures the satisfaction of the end users. The cognitive radio network base station (CRNBS) collects all the SUs' application information and allocates available channel resource to the SUs with the overall user perceived MOS maximized and fairness among SUs ensured. The simulation results confirm that the proposed MOS-based channel allocation scheme outperforms the conventional good put-based scheme in terms of overall user satisfaction. Bin Liu 0016, Yuanzhi Yao, Nenghai Yu, Chang Wen Chen |
ICIG | 3 |