Yan Gui

dblp:81/8221 · DBLP profile ↗
← Back
17ranked-venue papers
9as first author
12since 2021 · last 2025
0000-0001-8323-4571ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 6 first-author · 7 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2025 Learnable Bilateral Downsampling and Triple-Augmented Feature Fusion in YOLOv8 for Small Object Detection of UAV Aerial Images
Ruojun Guo, Yan Gui
ICONIP (5)2
2025 Lightweight structure-guided network with hydra interaction attention and global-local gating mechanism for high-resolution image inpainting
Yan Gui, Yaning Liu, Li-Dan Kuang
Expert Syst. Appl.1
2025 Constrained coupled CPD of complex-valued multi-slice multi-subject fMRI data
Li-Dan Kuang, Lei Long, Ting Tang, Yan Gui, Jin Zhang 0018
Signal Process.5
2025 SES-yolov5: small object graphics detection and visualization applications
Yan Gui
Vis. Comput.3
2025 Boosting memory network for video object segmentation in complex scenes
Yuhang Yi, Yan Gui
Vis. Comput.2
2024 Spatio-temporal SiamFC: per-clip visual tracking with siamese non-local 3D convolutional networks and multi-template updating
Yan Gui, Yiru Ou, Jianming Zhang 0003
Pattern Anal. Appl.1
2024 CGAN: lightweight and feature aggregation network for high-performance interactive image segmentation
Yan Gui, Zhengyan Zhang, Jin Zhang 0018
Vis. Comput.1
2023 Enhancing Image Rescaling Using High Frequency Guidance and Attentions in Downscaling and Upscaling Network
Yan Gui, Li-Dan Kuang, Jin Zhang 0018
CGI (1)1
2023 CPG-LS: Causal Perception Guided Linguistic Steganography
abstract
The current lexical substitution-based linguistic steganography primarily determines the substitutions by their linguistic suitability, overlooking the disparity in their capability against steganalysis. This oversight leads to the potential security risk by performing poor substitutions. To address this issue, this letter proposes aCausalPerceptionGuidedLinguisticSteganography(CPG-LS) via elaborate and secure lexical substitutions. CPG-LS constructs a causal perception network by making full use of a trained CNN discriminator to assess the security of each original word in the cover text and its substitutable candidates for controlling the embedding of secret message. Particularly, the causal score of each word is comprehensively measured from two perspectives by the causal perception network, one is word saliency, and the other is anti-steganalysis capability. Finally, the selection of original words is explicitly guided by their causal scores for securely embedding secret message. As the causal score of a word decreases, the perturbation caused by modifying this word becomes smaller, leading to stronger anti-steganalysis capability and lower semantic distortion in the obtained stegotexts. The experimental results demonstrate that CPG-LS achieves better text quality and anti-steganalysis capability than existing similar methods.
Lingyun Xiang, Jiali Xia, Yangfan Liu, Yan Gui
IEEE Signal Process. Lett.4
2022 Optimizing pcsCPD with Alternating Rank-R and Rank-1 Least Squares: Application to Complex-Valued Multi-subject fMRI Data
Li-Dan Kuang, Wenjun Li 0001, Yan Gui
ICONIP (5)3
2022 Group Residual Dense Block for Key-Point Detector with One-Level Feature
Jianming Zhang 0003, Jiajun Tao, Li-Dan Kuang, Yan Gui
PRICAI (2)4
2022 Learning interactive multi-object segmentation through appearance embedding and spatial attention
abstract
Abstract Deep learning approaches to interactive image segmentation are typically formulated as a binary labeling problem. A model trained to make predictions within a fixed set of labels (i.e., foreground and background labels) cannot be used to directly predict the binary masks of multiple objects of interest, which greatly limits its flexibility and adaptivity. The use of different classes of clicks as input is opted for and the first end‐to‐end learning model for multi‐object segmentation, based on a new designed neural network, is developed. The network consists of a visual feature extractor, a recurrent attention module and a dynamic segmentation head, extracts user click‐adapted appearance embedding features and spatial attention features, and then learns to transform this information into a segmentation of multiple objects. It is also proposed to train the network using a joint loss function, taking the embedding learning into account for segmentation. Comprehensive experiments are conducted on three benchmark datasets to demonstrate the effectiveness of the proposed method. It performs favorably against state‐of‐the‐art approaches on the multiple object segmentation task, for example, with 0.15 s per image, 0.06 s per object and mean IoU & F1 score of 84.90% on Pascal VOC 2012 validation set. It is further shown that the method can be used in numerous vision applications such as image recoloring and colorization.
Yan Gui, Bingqiang Zhou, Jianming Zhang 0003, Lingyun Xiang, Jin Zhang 0018
IET Image Process.1
2020 Reliable and Dynamic Appearance Modeling and Label Consistency Enforcing for Fast and Coherent Video Object Segmentation With the Bilateral Grid
abstract
We propose a novel optimization framework for video object segmentation, given the initial annotations of objects in the keyframes of an input video sequence. In this work, video data is represented by a Markov Random Field model, and segmentation is achieved by finding the minimum graph cut label assignment. More specifically, we first create a bilateral representation of the input video sequence which reduces the size of the graph that the min-cut must operate on. We then introduce dynamic appearance models to learn the segmentation likelihoods, and the reliability of likelihoods is measured to identify false likelihoods that may cause segmentation errors. Thus, the model accurately describes changes in the object's appearance that have evolved over time. Furthermore, we augment spatial and temporal connections using a soft higher-order potential, ensuring long-range label consistency in the segmentation. We provide extensive analysis and evaluation with respect to the influence of each component of the framework through the ablation study. Experiments on three benchmark datasets (DAVIS 2016, YouTube-Objects and SegTrack v2) show that our method achieves competitive performance compared to state-of-the-art while having the order of magnitude faster runtime.
Yan Gui, Dao-Jian Zeng, Yiyu Cai
IEEE Trans. Circuits Syst. Video Technol.1
2020 Joint learning of visual and spatial features for edit propagation from a single image
Yan Gui
Vis. Comput.1
2012 Preserving global features of fluid animation from a single image using video examples
abstract
We synthesize animations from a single image by transferring fluid motion of a video example globally. Given a target image of a fluid scene, an alpha matte is required to extract the fluid region. Our method needs to adjust a user-specified video example for producing the fluid motion suitable for the extracted fluid region. Employing the fluid video database, the flow field of the target image is obtained by warping the optical flow of a video frame that has a visually similar scene to the target image according to their scene correspondences, which assigns fluid orientation and speed automatically. Results show that our method is successful in preserving large fluid features in the synthesized animations. In comparison to existing approaches, it is both possible and useful to utilize our method to create flow animations with higher quality.
Yan Gui, Lizhuang Ma
J. Zhejiang Univ. Sci. C1
2012 A gradient-domain-based edge-preserving sharpen filter
Rynson W. H. Lau, Yan Gui, Mingang Chen, Lizhuang Ma
Vis. Comput.3
2010 Periodic pattern of texture analysis and synthesis based on texels distribution
Yan Gui, Lizhuang Ma
Vis. Comput.1