Shengrong Zhao

dblp:156/2527 · DBLP profile ↗
← Back
19ranked-venue papers
1as first author
15since 2021 · last 2026
0000-0003-0965-0918ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Systems, architecture and hardware · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 SGR-GS: 3D Gaussian Splatting Reconstruction Enhancement via Structure Consistency and Geometry Refinement
Shengrong Zhao, Hu Liang
ICIC (10)2
2026 Embracing Semantic Friction: A Conflict-Aware Evidential Fusion for Ambiguous Multimodal Sentiment Analysis
Zidan Wang, Shengrong Zhao, Hu Liang
ICIC (8)2
2026 FEMSDAN: Fourier-Enhanced Multi-scale Distillation Attention Network for Lightweight Image Super-Resolution
Hu Liang, Shengrong Zhao
ICIC (10)3
2025 WGMVSNet: An Efficient Dual-branch Self-supervised Multi-view Stereo Network for 3D Reconstruction
Hu Liang, Jiacheng Qu, Shengrong Zhao
ICIC (6)5
2025 Enhancing small satellite image resolution via shrinking rearranged mechanism and multiscale reparameterized attention
Zhibo Zhao, Hu Liang, Shengrong Zhao
Eng. Appl. Artif. Intell.4
2025 LMSFF: Lightweight multi-scale feature fusion network for image recognition under resource-constrained environments
Hu Liang, Shengrong Zhao
Expert Syst. Appl.3
2024 Learning Fine-Grained Information Alignment for Calibrated Cross-Modal Retrieval
abstract
Masked Language Modeling (MLM) and Image-Text Matching (ITM) are always used in fusion encoder to learn the joint representation of images and text. In existing methods, the masking strategy of MLM leads to the neglect of image details during the modeling process. Meanwhile, the sampling strategy of ITM struggles to consistently select high-difficulty hard negative instances, reducing the effectiveness of constraints. This leads to challenges in aligning fine-grained information in cross-modal retrieval. In response to this challenge, a fine-grained information alignment-based visual language model (FAM) is proposed in this paper. On one hand, the attribute-based masking strategy is employed in MLM, helping the model focus on the details of objects in images during modeling. On the other hand, the robust hard negative sample generation strategy provides challenging negative samples for ITM by altering the relationships between objects. This enables the model to align relationships between objects in different modalities and thus calibrates cross-modal retrieval. Extensive experiments demonstrate the effectiveness of the model in cross-modal retrieval tasks.
Jianhua Dong, Shengrong Zhao, Hu Liang
ICASSP2
2024 Lightweight super-resolution via multi-group window self-attention and residual blueprint separable convolution
Hu Liang, Shengrong Zhao
Multim. Syst.4
2023 LDVNet: Lightweight and Detail-Aware Vision Network for Image Recognition Tasks in Resource-Constrained Environments
abstract
In many underwater application scenarios, recognition tasks need to be executed promptly on computationally limited platforms. However, models designed for this field often exhibit spatial locality, and existing works lack the ability to capture crucial details in images. Therefore, a lightweight and detail-aware vision network (LDVNet) for resource-constrained environments is proposed to overcome the limitations of these approaches. Firstly, in order to enhance the accuracy of target image recognition, we introduce transformer modules to acquire global information, thus addressing the issue of spatial locality inherent in traditional convolutional neural networks (CNNs). Secondly, to maintain the network’s lightweight nature, we integrate the transformer module with convolutional operations, thereby mitigating the substantial parameter and floating point operations (FLOPs) overhead. Thirdly, for the efficient extraction of crucial fine-grained details from feature maps, we have devised a channel and spatial attention module (C&SA). This module aids in recognizing intricate and fine-grained visual tasks and enhances image understanding. It is seamlessly integrated into LDVNet with nearly negligible parameter overhead. The experimental results demonstrate that LDVNet outperforms other lightweight networks and hybrid networks in different recognition tasks, while being suitable for resource-constrained environments.
Hu Liang, Ran Qiu, Shengrong Zhao
ICPADS4
2023 GAF-GAN: Gated Attention Feature Fusion Image Inpainting Network Based on Generative Adversarial Network
abstract
Image inpainting, which aims to reconstruct reasonably clear and realistic images from known pixel information, is one of the core problems in computer vision. However, due to the complexity and variability of the underwater environment, the inability to extract valid pixel points and insufficient correlation between feature information in existing image inpainting techniques lead to blurring in the generated images. Therefore, a novel gated attention feature fusion image inpainting network based on generative adversarial networks (GAF-GAN) is proposed. The accuracy of feature similarity matching depends heavily on the validity of the information contained in the features. On the one hand, gating values are dynamically generated by gated convolution to reduce the interference of invalid information. On the other hand, semantic information at distant locations in an image is accurately acquired by the attention mechanism. For these reasons, we designed an improved gated attention mechanism. Gated attention mechanism make the network focus on effective information such as high-frequency texture and color fidelity of restored images. In addition, the dense feature fusion module is added to expand the overall receptive field of the network to fully learn the image features. Experimental results show that the proposed method can effectively repair defective images with complex texture structures and improve the reality and integrity of image details and structures.
Ran Qiu, Hu Liang, Shengrong Zhao
ICPADS4
2023 GA-Net: Gated Attention Mechanism Based Global Refinement Network for Image Inpainting
abstract
Underwater images are often affected by problems such as light attenuation, color distortion, noise and scattering, resulting in image defects. A novel image inpainting method is proposed to intelligently predict and fill damaged areas for complete and continuous visualization of the image. First, in order to effectively solve the problem of color distortion caused by light refraction in underwater environments, the improved gated attention mechanism is used. This mechanism improves the local details by learning and weighting the important features of the image. Second, gated convolution automatically determines the degree of restoration for each pixel based on local features of the original image. It eliminates distractions such as low contrast and scattering, retaining more original detailed information. By doing so, image inpainting techniques improve the quality and visualization of underwater images.
Ran Qiu, Shengrong Zhao, Hu Liang
ICPADS2
2023 Image super resolution via multi-regularization combining hybrid Tikhonov-TV prior and deep denoiser prior
abstract
In a real scenario, the image is often corrupted by complex degradation, and a lot of useful information is lost, which makes super-resolution (SR) reconstruction seriously ill-posed. To effectively solve such a problem, it is crucial to correctly exploit image prior knowledge. Although existing deep learning-based methods can obtain excellent results, they cannot deal with the complex degradation effectively, which would lead to the loss of texture details and the destruction of edge details. In this paper, an efficient multi-regularization method for SR is proposed, which can simultaneously exploit both internal and external image priors within a unified framework. The hybrid Tikhonov-TV prior and deep denoiser prior are introduced to constrain the reconstruction process. That is, the proposed model combines the superiority of the piecewise-smooth prior and deep prior. Moreover, an adaptive weight parameter is employed to make the hybrid component more detail-preserving. Experimental demonstrate that the proposed method achieves better performance in image detail protection than advanced methods.
Shengrong Zhao, Hu Liang, Changchun Wen
ICTAI2
2023 LCCN: A Lightweight Capture Context Network for Image Super-Resolution
abstract
In recent years, with the development of deep learning, many lightweight convolution neural networks (CNN) have achieved remarkable results in the field of single image super-resolution (SISR). However, it is difficult to capture long-range dependencies due to the limited perceptual field of lightweight CNN. To solve this problem, we propose the lightweight capture context network (LCCN), which includes the hierarchical residual attention distillation fusion (HRADF) module and lightweight capture context dependency (LCCD) module. In HRADF, we propose a hierarchical feature fusion structure consisting of multiple residual attention distillation blocks to improve the reconstruction effect by fusing multiple layers of feature maps. Moreover, the tpaconv lrelu block (TLB) and mixed spatial channel attention (MSCA) are applied to make the network lightweight and take full advantage of feature information. In LCCD, we introduce asymmetrical double multi-head attention to achieve low resource consumption and improve the long-range context dependency capture capability of the network. Extensive experiments show that LCCN achieves a good balance between performance and model complexity, and obtains satisfactory results on several benchmark datasets.
Changchun Wen, Hu Liang, Shengrong Zhao
IJCNN3
2023 Cascade Cost Volume Multi-View Stereo Network with Transformer and Pseudo 3D
abstract
Learning-based Multi-view Stereo (MVS) and stereo matching methods typically construct 3D cost volumes based on the camera frustum of the reference view. Regularization and regression of the cost volume are performed to obtain a depth map. However, the resolution of the output depth map is limited by the computational cost, and when performing feature extraction, the characteristics of convolution local perception make it impossible to capture global context information. In this paper, we propose CTPMVSNet by using the Global Feature Aware Transformer (GFT) to aggregate global context information within and across images. In order to make better use of GFT, we use Deformable Convolution Module (DCM) to ensure a smooth transition of the extracted feature range. In addition, in the cost volume regularization stage, to improve efficiency and generation accuracy, we design a lightweight regularization network with integrated pseudo-three-dimensional convolution, and our experiments on multiple dataset have achieved promising results.
Jiacheng Qu, Shengrong Zhao, Hu Liang, Qingmeng Zhang, Tingshuai Li
SMC2
2022 WPNet: Wide Pyramid Network for Recognition of HER2 Expression Levels in Breast Cancer Evaluation
abstract
Among the research methods for HER2 automatic evaluation in recent years, most of the methods using deep learning framework have both segmentation and classification functions. Although these methods provide pathologists with reference lesions, they increase the dependence on dataset and the computational cost. Therefore, we propose a Wide Pyramid Network (WPNet) based on deep learning to solve this problem. Our designed WPNet is different from other neural network models, which is mainly extended in the width of the network, and uses the wide pyramid structure to extract the features of different scales on the image for training. Since HER2 score is determined according to the degree of cell membrane staining and the proportion of cells with different degrees of staining, the WPNet model capable of multi-scale feature extraction can facilitate the determination of HER2 score. Compared with other models, for HER2 score classification based on a small sample set, the proposed model not only accelerates the convergence speed during training, reduces the calculation cost but also improves the classification effect.
Yuanze Zheng, Shengrong Zhao, Hu Liang
IJCNN2
2017 The NAMlet transform: A novel image sparse representation method based on non-symmetry and anti-packing model
Hu Liang, Shengrong Zhao, Chuanbo Chen, Mudar Sarem
Signal Process.2
2016 Multiframe super-resolution based on half-quadratic prior with artifacts suppress
Renchao Jin, Shengrong Zhao, Enmin Song
J. Vis. Commun. Image Represent.2
2016 A Generalized Detail-Preserving Super-Resolution method
Shengrong Zhao, Hu Liang, Mudar Sarem
Signal Process.1
2015 A novel multi-image super-resolution reconstruction method using anisotropic fractional order adaptive norm
Chuanbo Chen, Hu Liang, Shengrong Zhao, Zehua Lyu, Mudar Sarem
Vis. Comput.3