Fangyuan Gao

dblp:159/4528 · DBLP profile ↗
← Back
11ranked-venue papers
6as first author
9since 2021 · last 2026
0000-0002-1225-2067ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 first-author · 4 since 2021Computer networks · 2Security and privacy · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
4 papers
Image and video processing · 96% Virtual and augmented reality · 4%
Artificial intelligence
3 papers
Representation and self-supervised learning · 48% Probabilistic and Bayesian machine learning · 26% Deep learning architectures and training · 26%
Network and information security
1 paper
Digital forensics and information hiding · 77% Privacy and data protection · 23%

Topics — the 13 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Image and video processing
image restoration
1.822026
AIRPNet: Adaptive Image Restoration With Privacy Protection in Steganographic Domain · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Deep$\mathrm {M^{2}}$M2CDL: Deep Multi-Scale Multi-Modal Convolutional Dictionary Learning Network · IEEE Trans. Pattern Anal. Mach. Intell. 2024
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › sparse coding
convolutional dictionary learning
1.322024
Deep$\mathrm {M^{2}}$M2CDL: Deep Multi-Scale Multi-Modal Convolutional Dictionary Learning Network · IEEE Trans. Pattern Anal. Mach. Intell. 2024
Multi-Modal Convolutional Dictionary Learning · IEEE Trans. Image Process. 2022
Digital forensics and information hiding
steganography
1.012026
AIRPNet: Adaptive Image Restoration With Privacy Protection in Steganographic Domain · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › sparse coding
dictionary learning
0.812024
Deep$\mathrm {M^{2}}$M2CDL: Deep Multi-Scale Multi-Modal Convolutional Dictionary Learning Network · IEEE Trans. Pattern Anal. Mach. Intell. 2024
Image and video processing
image fusion
0.812024
Deep$\mathrm {M^{2}}$M2CDL: Deep Multi-Scale Multi-Modal Convolutional Dictionary Learning Network · IEEE Trans. Pattern Anal. Mach. Intell. 2024
Image and video processing › image fusion
multi-modal image fusion
0.812024
Deep$\mathrm {M^{2}}$M2CDL: Deep Multi-Scale Multi-Modal Convolutional Dictionary Learning Network · IEEE Trans. Pattern Anal. Mach. Intell. 2024
Image and video processing › image restoration
multi-modal image restoration
0.812024
Deep$\mathrm {M^{2}}$M2CDL: Deep Multi-Scale Multi-Modal Convolutional Dictionary Learning Network · IEEE Trans. Pattern Anal. Mach. Intell. 2024
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
bayesian inference
0.612022
Hierarchical Bayesian LSTM for Head Trajectory Prediction on Omnidirectional Images · IEEE Trans. Pattern Anal. Mach. Intell. 2022
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference
hierarchical bayesian inference
0.612022
Hierarchical Bayesian LSTM for Head Trajectory Prediction on Omnidirectional Images · IEEE Trans. Pattern Anal. Mach. Intell. 2022
Machine learning › Deep learning architectures and training › recurrent neural network
LSTM
0.612022
Hierarchical Bayesian LSTM for Head Trajectory Prediction on Omnidirectional Images · IEEE Trans. Pattern Anal. Mach. Intell. 2022
Machine learning › Deep learning architectures and training
recurrent neural network
0.612022
Hierarchical Bayesian LSTM for Head Trajectory Prediction on Omnidirectional Images · IEEE Trans. Pattern Anal. Mach. Intell. 2022
Virtual and augmented reality › immersive media
omnidirectional image
0.212022
Hierarchical Bayesian LSTM for Head Trajectory Prediction on Omnidirectional Images · IEEE Trans. Pattern Anal. Mach. Intell. 2022
Image and video processing
saliency detection
0.212022
Hierarchical Bayesian LSTM for Head Trajectory Prediction on Omnidirectional Images · IEEE Trans. Pattern Anal. Mach. Intell. 2022

Methods — techniques the papers use, named apart from their topics

wavelet lifting · 2.0invertible hiding · 2.0adaptive secure restoration · 2.0sparse coding · 1.5multi-scale convolutional dictionary learning · 1.5deep unfolding · 1.5hierarchical bayesian inference · 1.1future intention estimation · 1.1discrete fourier transform · 1.1alternating direction method of multipliers · 1.1
YearPublicationVenuePosition
2026 Say the image: Auditory masking effect-driven invertible network for progressive image-in-audio steganography
Jinghang Song, Fangyuan Gao, Xin Deng 0002, Shengxi Li, Mai Xu
J. Inf. Secur. Appl.2
2026 AIRPNet: Adaptive Image Restoration With Privacy Protection in Steganographic Domain
abstract
Cloud-based third-party multimedia services have become increasingly popular in last decade, however, they pose serious threats to users' privacy. To address this issue, in this paper, we propose a novel Adaptive Image Restoration network with Privacy protection, namely AIRPNet, which first attempts to perform image restoration in steganographic domain. Compared with existing methods, our method has significant advantages in invisibility, security and flexibility. Specifically, we first propose a wavelet lifting-based Adaptive Invertible Hiding (AIH) module to conceal the low-quality (LQ) secret image into a stego image. Then, instead of performing single type of restoration on the secret image, an adaptive secure restoration (ASR) module is developed to deal with multiple image degradations on the stego image. Finally, a high-quality (HQ) secret image can be extracted from the restored stego image. Here, since the secret image remains hidden throughout the whole image restoration process, the privacy of users can be greatly protected. The framework can be flexibly extended to multiple image restoration, which can restore multiple secret images from the same stego image. Experimental results on various datasets demonstrate that our AIRPNet outperforms existing methods in terms of restoration accuracy, invisibility and security on different image restoration tasks.
Fangyuan Gao, Xin Deng 0002, Junjie Huang 0001, Mai Xu
IEEE Trans. Pattern Anal. Mach. Intell.1
2026 DeepELIC: Deep encrypted lossy image compression network via compressive sensing unfolding
Fangyuan Gao, Yufan Deng, Xin Deng 0002, Zhenyu Guan 0002, Mai Xu
Pattern Recognit.1
2024 Deep$\mathrm {M^{2}}$M2CDL: Deep Multi-Scale Multi-Modal Convolutional Dictionary Learning Network
abstract
For multi-modal image processing, network interpretability is essential due to the complicated dependency across modalities. Recently, a promising research direction for interpretable network is to incorporate dictionary learning into deep learning through unfolding strategy. However, the existing multi-modal dictionary learning models are both single-layer and single-scale, which restricts the representation ability. In this paper, we first introduce a multi-scale multi-modal convolutional dictionary learning (M2CDL) model, which is performed in a multi-layer strategy, to associate different image modalities in a coarse-to-fine manner. Then, we propose a unified framework namely DeepM2CDL derived from the M2CDL model for both multi-modal image restoration (MIR) and multi-modal image fusion (MIF) tasks. The network architecture of DeepM2CDL fully matches the optimization steps of the M2CDL model, which makes each network module with good interpretability. Different from handcrafted priors, both the dictionary and sparse feature priors are learned through the network. The performance of the proposed DeepM2CDL is evaluated on a wide variety of MIR and MIF tasks, which shows the superiority of it over many state-of-the-art methods both quantitatively and qualitatively. In addition, we also visualize the multi-modal sparse features and dictionary filters learned from the network, which demonstrates the good interpretability of the DeepM2CDL network.
Xin Deng 0002, Fangyuan Gao, Xiancheng Sun, Mai Xu
IEEE Trans. Pattern Anal. Mach. Intell.3
2024 Extremely Low Bit-Rate Image Compression via Invertible Image Generation
abstract
Image compression at extremely low bit-rates has always been a challenging task in bandwidth limited scenarios, such as aerospace and deep-sea explorations. Recent years have seen great success of deep learning in image compression, however, few of them are specially designed for extremely low bit-rate conditions. To solve this issue, in this paper, we propose a novel invertible image generation based framework for extremely low bit-rate image compression. The proposed framework is composed of three modules, including an invertible image generation (IIG) module, a generated image compression (GIC) module and a compressed image adjustment (CIA) module. The role of IIG module is to generate a compression-friendly image from the original image. In the IIG module, image generation and restoration are modelled as two mutually reversible processes to avoid the information loss. After the IIG module, the GIC module is employed to compress the generated images to save the coding bit-rates. After that, the CIA module is used to shrink the quality gap between the compressed generated image and the un-compressed image. Finally, the image from the CIA module is sent back to the IIG module to restore the original image. The experimental results on three different datasets show that the proposed framework achieves state-of-the-art performance in image compression with extremely low bit-rates. We also extend the proposed framework to feature compression towards object detection, which saves 90% bit-rates than the VVC standard with the same detection accuracy.
Fangyuan Gao, Xin Deng 0002, Junpeng Jing, Mai Xu
IEEE Trans. Circuits Syst. Video Technol.1
2023 ULcompress: A Unified low bit-rate image Compression Framework via Invertible Image Representation
abstract
In this paper, we propose a unified low bit-rate image compression framework, namely ULCompress, via invertible image representation. The proposed framework is composed of two important modules, including an invertible image rescaling (IIR) module and a compressed quality enhancement (CQE) module. The role of IIR module is to learn a compression-friendly low-resolution (LR) image from the high-resolution (HR) image. Instead of the HR image, we compress the LR image to save the bit-rates. The compression codecs can be any existing codecs. After compression, we propose a CQE module to enhance the quality of the compressed LR image, which is then sent back to the IIR module to restore the original HR image. The network architecture of IIR module is specially designed to ensure the invertibility of LR and HR images, i.e., the downsampling and upsampling processes are invertible. The CQE module works as a buffer between IIR module and the codec, which plays an important role in improving the compatibility of our framework. Experimental results show that our ULCompress is compatible with both standard and learning-based codecs, and is able to significantly improve their performance at low bit-rates.
Fangyuan Gao, Xin Deng 0002, Mai Xu
ICIP1
2022 SFIC: Sparsity-Driven Facial Image Compression Network
abstract
Facial image compression is crucial in many areas like social media and video surveillance. Considering the sparsity of facial features, sparse representation (SR) has been applied to compress facial images, in which each image patch is sparsely represented by a small number of dictionary atoms to save bit-rates. Along this line, we propose the first end-to-end sparsity-driven facial image compression network namely SFIC. In the proposed network, the traditional convolutional sparse coding (CSC) is turned into a learnable CSC block, which is combined with discrete wavelet transform (DWT) to form the sparsity encoding module (SEM). This is the first time that CSC has been explored in facial image compression. In the decoding side, a corresponding sparsity decoding module (SDM) is used to decode the image, and we further propose a quality enhancement module (QEM) to enhance the quality of decoded image. The experimental results verify that the proposed SFIC network achieves 74%, 55%, and 33% bit-rate savings over JPEG, JPEG-2000, and HEVC.
Fangyuan Gao, Xin Deng 0002, Mai Xu
ICIP1
2022 Hierarchical Bayesian LSTM for Head Trajectory Prediction on Omnidirectional Images
abstract
When viewing omnidirectional images (ODIs), viewers can access different viewports via head movement (HM), which sequentially forms head trajectories in spatial-temporal domain. Thus, head trajectories play a key role in modeling human attention on ODIs. In this paper, we establish a large-scale dataset collecting 21,600 head trajectories on 1,080 ODIs. By mining our dataset, we find two important factors influencing head trajectories, i.e., temporal dependency and subject-specific variance. Accordingly, we propose a novel approach integrating hierarchical Bayesian inference into long short-term memory (LSTM) network for head trajectory prediction on ODIs, which is called HiBayes-LSTM. In HiBayes-LSTM, we develop a mechanism of Future Intention Estimation (FIE), which captures the temporal correlations from previous, current and estimated future information, for predicting viewport transition. Additionally, a training scheme called Hierarchical Bayesian inference (HBI) is developed for modeling inter-subject uncertainty in HiBayes-LSTM. For HBI, we introduce a joint Gaussian distribution in a hierarchy, to approximate the posterior distribution over network weights. By sampling subject-specific weights from the approximated posterior distribution, our HiBayes-LSTM approach can yield diverse viewport transition among different subjects and obtain multiple head trajectories. Extensive experiments validate that our HiBayes-LSTM approach significantly outperforms 9 state-of-the-art approaches for trajectory prediction on ODIs, and then it is successfully applied to predict saliency on ODIs.
Li Yang 0014, Mai Xu, Xin Deng 0002, Fangyuan Gao, Zhenyu Guan 0002
IEEE Trans. Pattern Anal. Mach. Intell.5
2022 Multi-Modal Convolutional Dictionary Learning
abstract
Convolutional dictionary learning has become increasingly popular in signal and image processing for its ability to overcome the limitations of traditional patch-based dictionary learning. Although most studies on convolutional dictionary learning mainly focus on the unimodal case, real-world image processing tasks usually involve images from multiple modalities, e.g., visible and near-infrared (NIR) images. Thus, it is necessary to explore convolutional dictionary learning across different modalities. In this paper, we propose a novel multi-modal convolutional dictionary learning algorithm, which efficiently correlates different image modalities and fully considers neighborhood information at the image level. In this model, each modality is represented by two convolutional dictionaries, in which one dictionary is for common feature representation and the other is for unique feature representation. The model is constrained by the requirement that the convolutional sparse representations (CSRs) for the common features should be the same across different modalities, considering that these images are captured from the same scene. We propose a new training method based on the alternating direction method of multipliers (ADMM) to alternatively learn the common and unique dictionaries in the discrete Fourier transform (DFT) domain. We show that our model converges in less than 20 iterations between the convolutional dictionary updating and the CSRs calculation. The effectiveness of the proposed dictionary learning algorithm is demonstrated on various multimodal image processing tasks, achieves better performance than both dictionary learning methods and deep learning based methods with limited training data.
Fangyuan Gao, Xin Deng 0002, Mai Xu, Pier Luigi Dragotti
IEEE Trans. Image Process.1
2017 Events detection and community partition based on probabilistic snapshot for evolutionary social network
Zhongnan Zhang, Ming Qiu, Fangyuan Gao
Peer-to-Peer Netw. Appl.4
2014 Probabilistic Snapshot Based Evolutionary Social Network Events Detection
abstract
Most of the existing researches simply convert associations of nodes within the snapshot of the evolutionary social network to the weight of edges. However, because of the obvious Matthew effect existing in the interactions of nodes in the real social network, the association strength matrices extracted directly by snapshots are extremely uneven. This paper introduces a new evolutionary social network model. Firstly, we generate probabilistic snapshots of the evolutionary social network data. Afterwards, we use the probabilistic factor model to detect the variation points brought by network events. According to experimental results, our proposed probabilistic snapshot model of evolutionary social network is effective for network events detection.
Zhongnan Zhang, Fangyuan Gao
MSN3