Wenjing Gao

dblp:15/8491 · DBLP profile ↗
← Back
17ranked-venue papers
8as first author
12since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 2 since 2021Computer networks · 3 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-authorSystems, architecture and hardware · 2 · 1 first-author · 1 since 2021Security and privacy · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 MEC-Dedup: Secure data deduplication for mobile users in edge-assisted cloud storage systems
Wenjing Gao, Jia Yu 0003
J. Inf. Secur. Appl.2
2025 EIU-IC: Enhancing Interaction Understanding in Text-to-Image Generation Models with Interaction Control
Yonghua Zhu, Wenjing Gao
PRCV (9)3
2025 Towards privacy-preserving compressed sensing reconstruction in cloud
Kaidi Xu, Wenjing Gao
Comput. Secur.3
2025 Enabling Privacy-Preserving Top-k Hamming Distance Query on the Cloud
abstract
The top-k Hamming distance query is to find the k optimal objects with the smallest Hamming distance to the query data. It has a wide range of applications in many domains such as social networks, image retrieval and biological recognition. The existing privacy-preserving protocols do not support the top-k Hamming distance query in practice. To address this issue, we consider letting the user securely query the top-k Hamming distance on the cloud in a secure outsourcing manner. We propose two protocols to realize the privacy-preserving top-k Hamming distance query on the cloud. In the first protocol, two cloud servers are introduced to cooperatively complete the privacy-preserving top-k Hamming distance query. To preserve data privacy, the Paillier encryption and randomization techniques are leveraged to blind the user data, and the ciphertext data is stored on the first cloud server. The second cloud server calculates the Hamming distance on the ciphertexts. After that, the encrypted query results are returned to the query user for recovering the top-k query results. In the second protocol, we adopt the data aggregation strategy to further enhance the efficiency. By packaging data, the computation overhead of each participant is reduced and the communication overhead of the protocol is decreased, remarkably. Security analysis demonstrates that the data privacy is guaranteed in the proposed protocols. Experimental results evaluate the performance of the proposed protocols and confirm the superiority of the second protocol.
Wenjing Gao, Jia Yu 0003
IEEE Trans. Netw. Serv. Manag.1
2025 Disentangled text-driven stylization of 3D faces via directional CLIP losses
abstract
3D face stylization remains challenging due to limited training samples, diverse style domains, and the complex mapping between ambiguous style features and 3D face structures. To address these issues, we propose ClipStyleFace, a text-driven approach for 3D face stylization that leverages CLIP (Contrastive Language-Image Pre-training) knowledge to create style variations in both geometric and texture structures. ClipStyleFace comprises three components. For geometry deformation, a deformable surface is designed to model stylized geometric residuals on the initial mesh. For texture transformation, we construct a compact parameter space enabling style transfer using a pre-trained albedo generator. Both modules are optimized consistently by distilling semantic alignment and domain correction knowledge from the CLIP model. Extensive experiments demonstrate the effectiveness of our approach in generating stylized 3D faces that match target style prompts while preserving identity characteristics and facial details. Our model also holds promise for applications such as animation and image-driven 3D stylized face generation. Our code is released on https://github.com/cutegao715/ClipStyleFace .
Wenjing Gao, JiaoJiao Wang, Dingguo Yu
Vis. Comput.1
2024 Enabling privacy-preserving non-interactive computation for Hamming distance
Wenjing Gao, Wei Liang 0005, Rong Hao
Inf. Sci.1
2024 Enabling Privacy-Preserving Parallel Computation of Linear Regression in Edge Computing Networks
abstract
Linear regression is a classical statistical model with a wide range of applications. The function of linear regression is to predict the value of a dependent variable (the output) given an independent variable (the input). The training of a linear regression model is to find a linear relationship between the input and the output based on data samples. IoT applications usually require real-time data processing. Nonetheless, the existing schemes about privacy-preserving outsourcing of linear regression cannot fully meet the rapid response requirement for computation. To address this issue, we consider employing multiple edge servers to accomplish privacy-preserving parallel computation of linear regression. We propose two novel solutions based on edge servers in edge computing networks and construct two efficient schemes for linear regression. In the first scheme, we present a new blinding technique for data privacy protection. Two edge servers are employed to execute the encrypted linear regression task in parallel. To further enhance the efficiency, we design an adaptive parallel algorithm, which is adopted in the second scheme. Multiple edge servers are employed in the second scheme to achieve higher efficiency. We analyze the correctness, privacy, and verifiability of the proposed schemes. Finally, we assess the computational overhead of the proposed schemes and conduct experiments to validate the performance advantages of the proposed schemes.
Wenjing Gao, Jia Yu 0003, Huaqun Wang
IEEE Trans. Cloud Comput.1
2023 Privacy-Preserving Face Recognition With Multi-Edge Assistance for Intelligent Security Systems
abstract
Face recognition is one of the key technologies in intelligent security systems. Data privacy and identification efficiency have always been concerns about face recognition. Existing privacy-preserving protocols only focus on the training phase of face recognition. Since intelligent security systems mainly complete the calculation of large-scale face data in the identification phase, existing privacy-preserving protocols cannot be well applied to intelligent security systems. In this article, we propose the first privacy-preserving face recognition protocol for the calculations in the identification phase for intelligent security systems. We introduce the Householder matrix to blind user data including model data and face data, which enables the proposed protocol to support privacy-preserving face recognition on semi-trusted edge servers. Utilizing edge computing, fast response for large-scale face recognition can be achieved. The user can offload heavy calculations of matrix multiplication and Euclidean distances to edge servers simultaneously. The proposed protocol supports parallel computing based on multiple edge servers and thus enhances the efficiency of face recognition in intelligent security systems. Moreover, the recognition accuracy in the proposed protocol is the same as that in the original PCA-based face recognition algorithm. The security analysis demonstrates that the protocol protects the privacy of user data. The numerical analysis and simulation experiments are carried out to show the efficiency and feasibility of the proposed protocol.
Wenjing Gao, Jia Yu 0003, Rong Hao, Fanyu Kong 0002
IEEE Internet Things J.1
2023 Privacy-Preserving Parallel Computation of Matrix Determinant With Edge Computing
abstract
With the widespread deployment of secure outsourcing computation, the resource-constrained client can delegate intensive computation tasks to powerful servers. Matrix determinant computation is a fundamental mathematical operation that has been widely used in IoT applications. This operation is computationally expensive. Nevertheless, the existing secure outsourcing protocols for matrix determinant are all designed based on one cloud server, which cannot well meet the low-latency and real-time computing requirements for the client. To address this issue, we explore accelerating the computation of matrix determinant by parallel outsourcing based on two nearby edge servers and propose the first practical protocol. We use the matrix blocking technique to split the computation task into multiple subtasks, which are parallel outsourced to edge servers for accelerating the computation. Moreover, we propose a privacy-preserving matrix transformation technique for data privacy protection. This technique only involves the operations of matrix-vector multiplication and matrix-matrix addition. It achieves the lightweight computation for the client and supports computational indistinguishability for the blinded input and a uniform distribution. The correctness, privacy and verifiability of the proposed protocol are analyzed. Finally, the performance advantage of the proposed protocol is demonstrated through simulation experiments.
Wenjing Gao, Jia Yu 0003
IEEE Trans. Serv. Comput.1
2022 Enabling Privacy-Preserving Parallel Outsourcing Matrix Inversion in IoT
abstract
With the rapid development of Internet of Things (IoT), edge computing has been widely applied as a novel computing paradigm. Securely outsourcing intensive tasks to edge servers is becoming increasingly pervasive. It is a nice approach for resource-limited IoT devices to accomplish heavy computing tasks. Matrix inversion is a basic but time-consuming operation, which has a wide range of applications in IoT. The current privacy-preserving outsourcing schemes for matrix inversion cannot support parallel computing based on multiple edge servers. As a result, they cannot well satisfy the requirement of fast response for computation in IoT. In order to deal with this problem, we propose two privacy-preserving parallel outsourcing schemes for matrix inversion in IoT. In the first scheme, we design a novel method to generate a random matrix, which is used to blind the inputted original matrix. In this scheme, two edge servers compute the inversion of the encrypted matrix in parallel to improve the computational efficiency. To further improve the efficiency, we design a novel subtasks partitioning and assignment strategy and propose the second scheme by balancing the computing load of edge servers. We analyze the correctness, security, and verifiability of the proposed schemes. And we provide theoretical analysis and experimental results to demonstrate the performance advantages of the proposed schemes.
Wenjing Gao, Jia Yu 0003, Ming Yang 0023, Huaqun Wang
IEEE Internet Things J.1
2022 Learning Spatial-Parallax Prior Based on Array Thermal Camera for Infrared Image Enhancement
abstract
In this article, an array thermal camera equipment is developed to capture multiple infrared images with spatial and parallax information. Based on the captured images, an end-to-end method called spatial–parallax prior network (SPPN) is proposed. Specifically, we design a spatial–parallax prior block with two symmetric branches to extract spatial and parallax features in an interactive guidance manner. Then, to effectively integrate spatial and parallax features, we introduce a channel attention mechanism to enable the network to focus on and fuse the most useful information adaptively. In this way, spatial and parallax information can be fully utilized without any explicit alignment operation. Finally, considering the scarcity and poor quality of infrared training data, we leverage transfer learning to better train the network. Extensive experimental results demonstrate that the proposed SPPN consistently outperforms the current state-of-the-art methods, providing a highly effective and scalable solution for the improvement of infrared image quality.
Jiayi Ma 0001, Wenjing Gao, Yong Ma 0001, Jun Huang 0008, Fan Fan 0001
IEEE Trans. Ind. Informatics2
2021 MPIN: a macro-pixel integration network for light field super-resolution
abstract
Most existing light field (LF) super-resolution (SR) methods either fail to fully use angular information or have an unbalanced performance distribution because they use parts of views. To address these issues, we propose a novel integration network based on macro-pixel representation for the LF SR task, named MPIN. Restoring the entire LF image simultaneously, we couple the spatial and angular information by rearranging the four-dimensional LF image into a two-dimensional macro-pixel image. Then, two special convolutions are deployed to extract spatial and angular information, separately. To fully exploit spatial-angular correlations, the integration resblock is designed to merge the two kinds of information for mutual guidance, allowing our method to be angular-coherent. Under the macro-pixel representation, an angular shuffle layer is tailored to improve the spatial resolution of the macro-pixel image, which can effectively avoid aliasing. Extensive experiments on both synthetic and real-world LF datasets demonstrate that our method can achieve better performance than the state-of-the-art methods qualitatively and quantitatively. Moreover, the proposed method has an advantage in preserving the inherent epipolar structures of LF images with a balanced distribution of performance.
Xinya Wang, Jiayi Ma 0001, Wenjing Gao, Junjun Jiang
Frontiers Inf. Technol. Electron. Eng.3
2020 Various syncretic co-attention network for multimodal sentiment analysis
abstract
Summary The multimedia contents shared on social network reveal public sentimental attitudes toward specific events. Therefore, it is necessary to conduct sentiment analysis automatically on abundant multimedia data posted by the public for real‐world applications. However, approaches to single‐modal sentiment analysis neglect the internal connections between textual and visual contents, and current multimodal methods fail to exploit the multilevel semantic relations of heterogeneous features. In this article, the various syncretic co‐attention network is proposed to excavate the intricate multilevel corresponding relations between multimodal data, and combine the unique information of each modality for integrated complementary sentiment classification. Specifically, a multilevel co‐attention module is constructed to explore localized correspondences between each image region and each text word, and holistic correspondences between global visual information and context‐based textual semantics. Then, all the single‐modal features can be fused from different levels, respectively. Except for fused multimodal features, our proposed VSCN also considers unique information of each modality simultaneously and integrates them into an end‐to‐end framework for sentiment analysis. The superior results of experiments on three constructed real‐world datasets and a benchmark dataset of Visual Sentiment Ontology (VSO) prove the effectiveness of our proposed VSCN. Especially qualitative analyses are given for deep explaining of our method.
Yonghua Zhu, Wenjing Gao, Shaoxiu Wang
Concurr. Comput. Pract. Exp.3
2019 A hierarchical recurrent approach to predict scene graphs from a visual-attention-oriented perspective
abstract
Abstract A scene graph provides a powerful intermediate knowledge structure for various visual tasks, including semantic image retrieval, image captioning, and visual question answering. In this paper, the task of predicting a scene graph for an image is formulated as two connected problems, ie, recognizing the relationship triplets, structured as <subject‐predicate‐object>, and constructing the scene graph from the recognized relationship triplets. For relationship triplet recognition, we develop a novel hierarchical recurrent neural network with visual attention mechanism. This model is composed of two attention‐based recurrent neural networks in a hierarchical organization. The first network generates a topic vector for each relationship triplet, whereas the second network predicts each word in that relationship triplet given the topic vector. This approach successfully captures the compositional structure and contextual dependency of an image and the relationship triplets describing its scene. For scene graph construction, an entity localization approach to determine the graph structure is presented with the assistance of available attention information. Then, the procedures for automatically converting the generated relationship triplets into a scene graph are clarified through an algorithm. Extensive experimental results on two widely used data sets verify the feasibility of the proposed approach.
Wenjing Gao, Yonghua Zhu, Honghao Gao
Comput. Intell.1
2017 Bee pose estimation from single images with convolutional neural network
abstract
In this paper, we present a deep convolutional neural network (ConvNet) based framework for estimating the bee pose from a single image. Unlike some existing human pose estimation methods that localize a fixed number of body joints, our method handles the cases with a varying number of targets. Compared to the existing bee pose estimation methods, our framework is more robust and accurate. It is effective even for some challenging images (e.g., when the bee is fed sugar water with a stick). The proposed framework learns a mapping from the global structure and local appearance of a bee to its pose. We evaluated our method on two challenging datasets. Experiments showed that it has achieved significant improvements over the existing insect pose estimation algorithms.
Le Duan, Minmin Shen, Wenjing Gao, Oliver Deussen
ICIP3
2017 Shape recognition by bag of contour fragments with a learned pooling function
abstract
Bag of Contour Fragments (BoCF), derived from the well-known Bag-of-Features (BoF), is an effective framework for shape representation. The feature pooling in this framework is a critical step, while either max pooling or average pooling is not a learnable process. In this paper, we aim at learning a pooling function which is adaptive to the input contour fragment features instead. Towards this end, we formulate our pooling function as a weighted sum of max pooling and average pooling, where the weight is expressed by an activation function of the input contour fragment features. To automatically learn this weight, the output of the pooling function is fed into a SVM classifier and they are trained jointly to minimize a shape classification loss. Experimental results on several standard shape datasets demonstrate the effectiveness of the proposed learned pooling function, which can achieve considerable improvements compared with BoCF.
Wei Shen 0002, Wenjing Gao, Yuan Jiang 0002, Dan Zeng 0001, Zhijiang Zhang
ICIP2
2016 Shape recognition by bag of skeleton-associated contour parts
Wei Shen 0002, Yuan Jiang 0002, Wenjing Gao, Dan Zeng 0001, Xinggang Wang
Pattern Recognit. Lett.3