Yuan Yuan 0007

dblp:64/5845-7 · DBLP profile ↗
← Back
13ranked-venue papers
7as first author
6since 2021 · last 2026
0000-0003-3352-0662ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 6 first-author · 5 since 2021Theory of computation · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 PSTAN: A JND-Aware Pairwise Spatio-Temporal Alignment Network for Compressed Videos Quality Enhancement
abstract
Compressed video quality enhancement (CVQE) is crucial for mitigating compression artifacts and improving perceptual visual quality, especially under diverse quantization parameters (QPs) and motion patterns. However, many existing approaches insufficiently exploit long-range temporal dependencies, and their reliance on QP-specific training often leads to limited robustness when compression conditions change. In this work, we propose a just noticeable difference (JND)-aware and perception-driven learning framework for CVQE, termed the Pairwise Spatio-Temporal Alignment Network (PSTAN). PSTAN incorporates perceptual priors primarily through a JND-guided training paradigm rather than relying solely on architectural modifications, where learning is driven by perceptuallypoorvideo segments identified in the VideoSet dataset. This strategy alleviates the reliance on QP-specific supervision and promotes more stable enhancement behavior across varying compression conditions. To effectively capture temporal dependencies, PSTAN employs a pairwise spatio-temporal interaction mechanism that models each reference-target frame pair independently, enabling adaptive utilization of both nearby and distant frames. In addition, a transformer-based alignment module combining temporal mutual attention with cascaded deformable convolution is introduced to handle complex and large motions. Extensive experiments on VideoSet, MFQE 2.0 and our constructed HEVC-comperssed dataset show that PSTAN achieves consistent improvements over state-of-the-art CVQE methods in both objective and perceptual quality metrics. The code of this work is available at https://github.com/leryong/PSTAN.git.
Yuan Yuan 0007, Eryong Li, Jiawei Zhang 0002, Jinchang Ren, Xu Lu 0002
IEEE Trans. Circuits Syst. Video Technol.1
2025 Nu-SAM: A Frequency Decomposition and Channel-Spatial Dual Attention Enhanced SAM for Dense Nuclei Segmentation
Xu Lu 0002, Yexin Huang, Yuan Yuan 0007, Shan Xiong, Wenhua Liang
PRCV (14)4
2025 Deep Augmented Metric Learning Network for Prostate Cancer Classification in Ultrasound Images
abstract
Prostate cancer screening often relies on cost-intensive MRIs and invasive needle biopsies. Transrectal ultrasound imaging, as a more affordable and non-invasive alternative, faces the challenge of high inter-class similarity and intra-class variability between benign and malignant prostate cancers. This complexity requires more stringent differentiation of subtle features for accurate auxiliary diagnosis. In response, we introduce the novel Deep Augmented Metric Learning (DAML) network, specifically tailored for ultrasound-based prostate cancer classification. The DAML network represents a significant innovation in the metric learning space, introducing the Semantic Differences Mining Strategy (SDMS) to effectively discern and represent subtle differences in prostate ultrasound images, thereby enhancing tumor classification accuracy. Additionally, the DAML network strategically addresses class variability and limited sample sizes by combining the Linear Interpolation Augmentation Strategy (LIAS) and Permutation-Aided Reconstruction Loss (PARL). This approach enriches feature representation and introduces variability with straightforward structures, mirroring the efficacy of advanced sample generation techniques. We carried out comprehensive empirical assessments of the DAML model by testing its key components against a range of models, ensuring its effectiveness. Our results demonstrate the enhanced performance of the DAML model, achieving classification accuracies of 0.857 and 0.888 for benign and malignant cancers, respectively, underscoring its effectiveness in prostate cancer classification via medical imaging.
Xu Lu 0002, Yanqi Guo, Shulian Zhang, Yuan Yuan 0007, Chun-Chun Wang
IEEE J. Biomed. Health Informatics4
2022 Self-Guided Video Super-Resolution Based on a Fast Deformable ConvGRU Model
abstract
Video super-resolution (VSR) aims at recovering a natural and realistic high-resolution (HR) video frame from the corre-sponding low-resolution (LR) counterpart and its consecutive neighboring frames. The challenge is how to make full use of spatio-temporal coherence among the input LR frames. In this work, we propose a self-guided deformable convolutional gated recurrent unit (GRU) framework for VSR. Specifically, convolutional GRU can efficiently extract temporal features of the input LR frames. Deformable convolution (DConv) is utilized to spatially align the hidden states of GRU cells with the input feature maps. Moreover, we argue that the reference LR frame itself is efficient to guide the aggregated features learning frame-specific feature maps, which are then used to generate rich and more realistic textures towards the corresponding HR frame. Extensive experimental results on benchmark datasets demonstrate that the proposed framework achieves better performance than state-of-the-art methods and has higher model efficiency.
Jingming Chen, Yuan Yuan 0007, Jiawei Zhang 0002, Jianping Luo
ICME2
2022 Video Super-Resolution with Spatial-Temporal Transformer Encoder
abstract
The challenge of Video super-resolution (VSR) is how to make full use of the spatial-temporal coherence among neigh-bouring LR frames to generate high-resolution (HR) prediction. In this study, we propose to use transformer on VSR to capture long-range temporal dependencies. Specifically, we first spatially divide LR images into patches and split each patch into sub-patches. Transformer encoders are applied to both the patches and sub-patches, such that the self-attention modules can extract both global and local correlations. To accelerate the training process and filter out irrelevant features, we only select top-k similar features for the attention scheme. We then feed the extracted long-range correlations into a temporal, spatial and channel attention fusion mod-ule’ which enhances the useful information along all three di-mensions' respectively. Extensive experiments on benchmark datasets show that the proposed model outperforms state-of-the-art VSR methods in terms of PSNR/SSIM values and vi-sual qualities.
Ruiqi Tan, Yuan Yuan 0007, Jianping Luo
ICME2
2022 Landmarking for Navigational Streaming of Stored High-Dimensional Media
abstract
Modern media data such as 360° videos and light field (LF) images are typically captured in much higher dimensions than the observers’ visual displays. To efficiently browse high-dimensional media, a navigational streaming model is considered: a client navigates the media space by dictating a navigation path to a server, who in response transmits the corresponding pre-encoded media data units (MDU) to the client one-by-one in sequence. Assuming that the MDU quality is pre-chosen and fixed, the problem resides in selecting and storing redundant representations of MDUs at the server in order to best trade off storage and transmission costs, while enabling adequate user’s random access. We address this problem with a landmark-based MDU optimization framework. The media space is divided into neighborhoods, each containing one landmark (a chosen MDU). MDUs in a neighborhood use the associated landmark as a predictor for inter-coding. Thus, for any MDU transition within the same neighborhood, only one inter-coded MDU transmission is required when the landmark resides in the decoder buffer. It results in lower transmission cost and enables navigational random access. To optimize an MDU structure, we employ tree-structured vector quantizer (TSVQ) to first optimize landmark locations, then iteratively add P-MDUs as refinements using a fast branch-and-bound technique. Taking interactive LF images and viewport adaptive 360° images as illustrative applications, and I-, P- and previously proposed merge frames to intra- and inter-code MDUs, we show experimentally that landmarked MDU structures can noticeably reduce the expected transmission cost compared with MDU structures without landmarks.
Yuan Yuan 0007, Gene Cheung, Pascal Frossard, H. Vicky Zhao, Jiwu Huang
IEEE Trans. Circuits Syst. Video Technol.1
2020 Video Super-Resolution using Multi-scale Pyramid 3D Convolutional Networks
abstract
Video super-resolution (SR) aims at generating high-resolution (HR) frames from consecutive low-resolution (LR) frames. The challenge is how to make use of temporal coherence among neighbouring LR frames. Most previous works use motion estimation and compensation based models. However, their performance relies heavily on motion estimation accuracy. In this paper, we propose a multi-scale pyramid 3D convolutional (MP3D) network for video SR, where 3D convolution can explore temporal correlation directly without explicit motion compensation. Specifically, we first apply 3D convolution into a pyramid subnet to extractmulti-scale spatial and temporal features simultaneously from the LR frames, such that it can handle various sizes of motions. We then feed the fused feature maps into an SR reconstruction subnet, where a 3D sub-pixel convolution layer is used for up-sampling. Finally, we append a detail refinement subnet based on the encoder-decoder structure to further enhance texture details of the reconstructed HR frames. Extensive experiments on benchmark datasets and real-world cases show that the proposed MP3D model outperforms state-of-the-art video SR methods in terms of PSNR/SSIM values, visual quality and temporal consistency, respectively.
Jianping Luo, Yuan Yuan 0007
ACM Multimedia3
2020 Multiple Cycle-in-Cycle Generative Adversarial Networks for Unsupervised Image Super-Resolution
abstract
With the help of convolutional neural networks (CNN), the single image super-resolution problem has been widely studied. Most of these CNN based methods focus on learning a model to map a low-resolution (LR) image to a highresolution (HR) image, where the LR image is downsampled from the HR image with a known model. However, in a more general case when the process of the down-sampling is unknown and the LR input is degraded by noises and blurring, it is difficult to acquire the LR and HR image pairs for traditional supervised learning. Inspired by the recent unsupervised imagestyle translation applications using unpaired data, we propose a multiple Cycle-in-Cycle network structure to deal with the more general case using multiple generative adversarial networks (GAN) as the basis components. The first network cycle aims at mapping the noisy and blurry LR input to a noise-free LR space, then a new cycle with a well-trained ×2 network model is orderly introduced to super-resolve the intermediate output of the former cycle. The number of total cycles depends on the different up-sampling factors (×2, ×4, ×8). Finally, all modules are trained in an end-to-end manner to get the desired HR output. Quantitative indexes and qualitative results show that our proposed method achieves comparable performance with the state-of-the-art supervised models.
Yongbing Zhang 0002, Chao Dong 0005, Xinfeng Zhang 0001, Yuan Yuan 0007
IEEE Trans. Image Process.5
2018 Object Shape Approximation and Contour Adaptive Depth Image Coding for Virtual View Synthesis
abstract
A depth image provides partial geometric information of a 3D scene, namely the shapes of physical objects as observed from a particular viewpoint. This information is important when synthesizing images of different virtual camera viewpoints via depth-image-based rendering (DIBR). It has been shown that depth images can be efficiently coded using contour-adaptive codecs that preserve edge sharpness, resulting in visually pleasing DIBR-synthesized images. However, contours are typically losslessly coded as side information, which is expensive if the object shapes are complex. In this paper, we pursue a new paradigm in depth image coding for color-plus-depth representation of a 3D scene: in a pre-processing step, we pro-actively simplify object shapes in a depth and color image pair to reduce depth coding cost, at a penalty of a slight increase in synthesized view distortion. Specifically, we first mathematically derive a distortion upper-bound proxy for 3DSwIM—a quality metric tailored for DIBR-synthesized images. This proxy reduces inter-dependency among pixel rows in a block to ease optimization. We then approximate object contours via a dynamic programming algorithm to optimally tradeoff coding the cost of contours using arithmetic edge coding with our proposed view synthesis distortion proxy. We modify the depth and color images according to the approximated object contours in an inter-view consistent manner. These are then coded, respectively, using a contour-adaptive image codec based on graph Fourier transform for edge preservation and High Efficiency Video Coding (HEVC) intra. Experimental results show that by maintaining sharp but simplified object contours during contour-adaptive coding, for the same visual quality of DIBR-synthesized virtual views, our proposal can reduce depth image coding rate by up to 22% in 3DSwIM and 42% in peak signal-to-noise ratio compared with alternative coding strategies, such as HEVC intra.
Yuan Yuan 0007, Gene Cheung, Patrick Le Callet, Pascal Frossard, H. Vicky Zhao
IEEE Trans. Circuits Syst. Video Technol.1
2017 Optimizing landmark insertions for interactive light field streaming
abstract
Light field imaging enables a user to navigate and observe a static 3D scene from different viewpoints. Downloading the entire data prior to navigation would incur a large startup delay. Instead, previous works propose an interactive light field streaming (ILFS) framework, where a user periodically requests a viewpoint, and in response the server transmits a presynthesized and encoded viewpoint image. Using I-frame, P-frame and previously proposed merge frame that facilitates view-switches, the challenge is how to design and pre-encode a storage-constrained frame structure to enable efficient view navigation. In this paper, we initialize “landmarks” into a structure to improve ILFS performance. A landmark is a designated view with P-frames to/from each neighborhood view, so that any viewpoint image can transition to any other viewpoint image by first visiting a landmark, and then from the landmark to the destination view. This results in a transmission cost of only two P-frames. Using a Lloyd's algorithm variant, we first incrementally insert into a frame structure landmarks one at a time at locally optimal locations. We then employ a greedy algorithm to add / subtract P-frames based on a rate-storage criterion. Experimental results show that our proposed structures have noticeably lower expected transmission cost for the same storage than structures generated by a previous greedy algorithm.
Yuan Yuan 0007, Gene Cheung, Pascal Frossard
ICIP1
2017 Robust optimization approximation for joint chance constrained optimization problem
Yuan Yuan 0007, Zukui Li, Biao Huang 0001
J. Glob. Optim.1
2015 Contour approximation & depth image coding for virtual view synthesis
abstract
A depth image provides geometric information of a 3D scene, namely the shapes of physical objects captured from a particular viewpoint. This information is important for synthesizing images corresponding to different virtual camera viewpoints via depth-image-based rendering (DIBR). Since it has been shown that blurring of object contours in the depth images leads to bleeding artefacts in virtual images. The most effective way to compress depth images relies on edge-adaptive image codecs that preserve contours, which are losslessly coded as side information (SI). However, lossless coding of the exact object contours can be expensive. In this paper, we argue that the contours themselves can be suitably approximated to save bits, while the depth images piecewise smooth (PWS) characteristic stays preserved. Specifically, we first propose a metric that estimates contour coding rate based on edge statistics. Given an initial rate estimate, we then pro-actively approximate object contours in a way that guarantees rate reduction when coded using arithmetic edge coding (AEC) as SI. Given the sharp but approximated contours, we finally encode the image using an edge-adaptive image codec with graph Fourier transform (GFT) for edge preservation. We show in our experiments that by maintaining sharp but slightly inaccurate object contours, the resulting quality of virtual views synthesized via DIBR exceeds those synthesized using depth images compressed with edge-adaptive codecs that losslessly encode object contours as SI, in particular when the total coding rate budget is low. This confirms that optimized coding of depth images results in an effective tradeoff in the representation of contour and respective depth information.
Yuan Yuan 0007, Gene Cheung, Pascal Frossard, Patrick Le Callet, H. Vicky Zhao
MMSP1
2013 Optimizing peer grouping for live free viewpoint video streaming
abstract
In free viewpoint video, a user can pull texture and depth videos captured from two nearby reference viewpoints to synthesize his chosen intermediate virtual view for observation via depth-image-based rendering (DIBR). For users who are observing the same video at the same time but not necessarily from the same virtual viewpoint, they have incentive to pull the same reference views so that the streaming cost can be shared. On the other hand, in general distortion of a synthesized virtual view increases with its distance to the reference views, and so a user also has incentive to select reference views that tightly “sandwich” his chosen virtual view, minimizing distortion. In a previous work, reference view sharing strategies-ones that optimally trade off shared streaming costs with synthesized view distortions-were investigated for the case when users are first divided into groups, and each user group independently pulls two reference views and shares the resulting streaming cost. In this paper, we generalize the previous notion of user group, so that a user can simultaneously belong to two groups, and each group shares the streaming cost of a single view. We also aim to find a Nash Equilibrium (NE) solution of reference view selection, which is stable and from which no one has incentive to unilaterally deviate. Specifically, we first derive a lemma based on known properties of synthesized view distortion functions. We then design a search algorithm to find a NE solution, leveraging on the derived lemma to reduce search complexity. Experimental results show that the stable NE solution increases the overall cost only slightly when compared to the unstable optimal reference selection that gives the lowest overall cost. Further, a larger network will give a lower average cost for each user, and thus, users tend to join large networks for cooperation.
Yuan Yuan 0007, Bo Hu 0036, Gene Cheung, H. Vicky Zhao
ICIP1