Shigang Wang 0003

dblp:29/3238-3 · DBLP profile ↗
← Back
35ranked-venue papers
0as first author
25since 2021 · last 2026
0000-0002-3598-9352ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 24 · 16 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 5 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Systems, architecture and hardware · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A Brain-Inspired Saliency Prediction Framework for Human-AI Cognitive Consistency in AIGC Content via Multi-Region Liquid Neurons
abstract
In recent years, human-AI cognitive consistency has emerged as a crucial perspective for evaluating the perceptual quality and interpretability of AIGC (Artificial Intelligence Generated Content). This paper proposes a biologically inspired saliency prediction framework that models six core regions of the human visual system—namely V1, V2, V4, MT, LIP, and FEF—using liquid neurons to capture the dynamic saliency features aligned with human gaze behavior. To enable effective alignment between AIGC models and human cognitive mechanisms, we introduce a cross-domain dual-teacher distillation strategy and construct a large-scale multimodal dataset comprising natural images, eye-tracking data, AIGC-generated images, and their corresponding cross-attention maps. Furthermore, we propose HAMCI (Human-AI Mutual Cognitive Index), a novel metric designed to quantitatively assess the spatial and semantic alignment between predicted saliency maps and model attention distributions. The proposed method demonstrates promising performance across various saliency prediction and cognitive alignment tasks, with results comparable to or surpassing recent state-of-the-art methods in several benchmarks. The code and dataset will be released upon acceptance to facilitate future research on cognitively aligned AIGC evaluation.
Yan Zhao 0012, Shigang Wang 0003
AAAI3
2026 Enhancing skin lesion segmentation via martingale feature fusion and adaptive deep semantic modeling
Yan Zhao 0012, Shigang Wang 0003
Multim. Syst.3
2026 Multimodal behavioral analysis for autism spectrum disorder assessment
Yunxiu Zhao, Shigang Wang 0003, Feiyong Jia, Honghua Li, Yan Zhao 0012
Pattern Recognit.2
2026 LKDTNet: Large Kernel Deconstruction Three-Dimensional Network for micro-expression recognition
Zixuan Jie, Qiankun Feng, Shigang Wang 0003
Signal Process. Image Commun.4
2025 KaRF: Weakly-Supervised Kolmogorov-Arnold Networks-based Radiance Fields for Local Color Editing
abstract
Recent advancements have suggested that neural radiance fields (NeRFs) show great potential in color editing within the 3D domain. However, most existing NeRF-based editing methods continue to face significant challenges in local region editing, which usually lead to imprecise local object boundaries, difficulties in maintaining multi-view consistency, and over-reliance on annotated data. To address these limitations, in this paper, we propose a novel weakly-supervised method called KaRF for local color editing, which facilitates high-fidelity and realistic appearance edits in arbitrary regions of 3D scenes. At the core of the proposed KaRF approach is a unified two-stage Kolmogorov-Arnold Networks (KANs)-based radiance fields framework, comprising a segmentation stage followed by a local recoloring stage. This architecture seamlessly integrates geometric priors from NeRF to achieve weakly-supervised learning, leading to superior performance. More specifically, we propose a residual adaptive gating KAN structure, which integrates KAN with residual connections, adaptive parameters, and gating mechanisms to effectively enhance segmentation accuracy and refine specific editing effects. Additionally, we propose a palette-adaptive reconstruction loss, which can enhance the accuracy of additive mixing results. Extensive experiments demonstrate that the proposed KaRF algorithm significantly outperforms many state-of-the-art methods both qualitatively and quantitatively. Our code and more results are available at: https://github.com/PaiDii/KARF.git.
Wudi Chen, Zhiyuan Zha, Shigang Wang 0003, Bihan Wen, Xin Yuan 0002, Jiantao Zhou 0001, Zipei Fan, Ce Zhu
NeurIPS3
2025 Triply Laplacian Scale Mixture Modeling for Seismic Data Noise Suppression
Sirui Pan, Zhiyuan Zha, Shigang Wang 0003, Yue Li 0003, Zipei Fan, Bihan Wen, Ce Zhu
IEEE Trans. Geosci. Remote. Sens.3
2025 Texture-Consistent 3D Scene Style Transfer via Transformer-Guided Neural Radiance Fields
abstract
Recent advancements have suggested that neural radiance fields (NeRFs) show great potential in 3D style transfer. However, most existing NeRF-based style transfer methods still face considerable challenges in generating stylized images that simultaneously preserve clear scene textures and maintain strong cross-view consistency. To address these limitations, in this paper, we propose a novel transformer-guided approach for 3D scene style transfer. Specifically, we first design a transformer-based style transfer network to capture long-range dependencies and generate 2D stylized images with initial consistency, which serve as supervision for the 3D stylized generation. To enable fine-grained control over style, we propose a latent style vector as a conditional feature and design a style network that projects this style information into the 3D space. We further develop a merge network that integrates style features with scene geometry to render 3D stylized images that are both visually coherent and stylistically consistent. In addition, we propose a texture consistency loss to preserve scene structure and enhance texture fidelity across views. Extensive quantitative and qualitative experimental results demonstrate that our proposed approach outperforms many state-of-the-art methods in terms of visual perception, image quality and multi-view consistency. Our code and more results are available at: https://github.com/PaiDii/TGTC-Style.git.
Wudi Chen, Zhiyuan Zha, Shigang Wang 0003, Bihan Wen, Xin Yuan 0002, Jiantao Zhou 0001, Ce Zhu
IEEE Trans. Image Process.3
2024 Fractional Order Spectrum in SAR Image Registration
abstract
SAR image registration is an important processing procedure for change detection and target recognition. However, the registration performance is seriously influenced by Symmetric α Stable (SαS) noises in SAR images. In order to cancel the impact of SαS noises in SAR image applications, a new concept of Fractional Order Spectrum of Cumulant (FOS-C) and SAR image registration based on (FOS-C) are proposed for the first time in this paper. In the proposed method, the images are registered from coarse to precise in three steps. First, the coarse registration based on Fourier Transform is used to estimate the scaling and rotation differences between images. Second, the registration based on FOS-C is designed, and used to provide the rough position of the similar regions. Third, the normalization cross-correlation (NCC) algorithm based on FOS-C is used to achieve the fine registration. Experimental results show that our method outperforms SAR-SIFT and KAZE-SAR.
Yan Zhao 0012, Xinbo Li, Shigang Wang 0003
ICME4
2024 TransDiff: medical image segmentation method based on Swin Transformer with diffusion probabilistic model
Yan Zhao 0012, Shigang Wang 0003
Appl. Intell.3
2024 A study on attention-based fine-grained image recognition: Towards musical instrument performing hand shape assessment
abstract
Automatic identification and professional evaluation makes musical instrument learning more intelligent. Since a proper hand shape is the basis of fingerings in playing instruments, this paper explores an integration of intelligent recognition technique into hand shape assessment of instrument players in an attempt of taking Chinese zither (Zheng) as an example. The fine-grained image recognition is novelly applied to automatically assessing basic hand shapes, as a tentative exploration of interdisciplinary research. First, this paper formulates an assessment scales by combining fine-grained image features with hand shape evaluation indicators in musical instrument learning. Then, an image dataset for hand shapes of Chinese zither performance (CZ-Dataset V2) is established based on free multi-view acquisition. Finally, we propose a fine-grained hand shape image recognition method using attention mechanism . Experimental results show that the basic instrumental hand shapes can be effectively recognized and reasonable suggestions for hand shape assessment can be provided.
Wenting Zhao 0003, Shigang Wang 0003, Yan Zhao 0012, Yecheng Liang, Jiehua Lin
Eng. Appl. Artif. Intell.2
2024 FewarNet: An Efficient Few-Shot View Synthesis Network Based on Trend Regularization
abstract
Novel view synthesis from existing inputs remains a research focus in computer vision. Predicting views becomes more challenging when only a limited number of views are available. This challenge is commonly referred to as the few-shot view synthesis problem. Recently, various strategies have emerged for few-shot view synthesis, such as transfer learning, depth supervision, and regularization constraints. However, transfer learning relies on massive scene data, depth supervision is affected by input depth quality, and regularization causes increased computational costs or impaired generalization. To address these issues, we propose a new few-shot view synthesis framework called FewarNet that introduces trend regularization to leverage depth structural features and a warping loss to supervise depth estimation, possessing the advantages of existing few-shot strategies, enabling high-quality novel view prediction with generalization and efficiency. Specifically, FewarNet consists of three stages: fusion, warping, and rectification. In the fusion stage, a fusion network is introduced to estimate depths using scene priors from coarse depths. In the warping stage, the predicted depths are used to guide the warping of the input views, and a distance-weighted warping loss is proposed to correctly guide depth estimation. To further improve prediction accuracy, we propose trend regularization which imposes penalties on depth variation trends to provide depth structural constraints. In the rectification stage, a rectification network is introduced to refine occluded regions in each warped view to generate novel views. Additionally, a rapid view synthesis strategy that leverages depth interpolation is designed to improve efficiency. We validate the method’s effectiveness and generalization on various datasets. Given the same sparse inputs, our method demonstrates superior performance in quality and efficiency over state-of-the-art few-shot view synthesis methods.
Chenxi Song, Shigang Wang 0003, Yan Zhao 0012
IEEE Trans. Circuits Syst. Video Technol.2
2024 Analysis of DAS Seismic Noise Generation and Elimination Process Based on Mean-SDE Diffusion Model
abstract
Suppressing various noises while achieving precise signal reconstruction in Distributed Acoustic Sensing Vertical Seismic Profiling (DAS VSP) remains a challenge. Existing denoising methods are insufficient due to factors such as the unknown noise-disturbing mechanism, low SNR, and limited training data. Therefore, this study proposes the Mean-Stochastic Differential Equation (SDE) diffusion model as an advanced solution. Built upon the standard diffusion model, which incorporates forward and backward diffusion processes, our model introduced three modifications to enhance performance. 1. Improving the forward diffusion process: Transforming the final state into a combination of the noisy DAS VSP and Gaussian noise. This adjustment allows precise representations of multi-type noise generation and facilitates backward sampling. 2. Enhancing noise prediction between successive steps in the backward process: A Nonlinear Activation Free Network (NAFnet) with a time Multi-Layer Perceptron (MLP) was employed to provide accurate noise predictions at different states. 3. Addressing training instability inherent in standard diffusion: The objective function is modified to seek the optimal trajectory of the best quality of signal reconstruction rather than directly evaluating the noise prediction. The forward diffusion is a dynamic evolution of adding noise to the pure signal, while the backward processing aims to remove the noise step by step. Comprehensive experiments demonstrate the superiority of our method in diverse noise suppression, signal resolution enhancement, and amplitude preservation. Moreover, grounded in physics-based equations, our method exhibits less dependency on training data compared to conventional deep learning methods.
Qiankun Feng, Shigang Wang 0003, Yue Li 0003
IEEE Trans. Geosci. Remote. Sens.2
2024 Fractional Order Spectrum of Cumulant in SAR Image Registration
abstract
Fractional order cumulant (FOC) is a new tool that can suppress symmetric$\alpha $stable (S$\alpha $S) noises in signal processing. Although FOC has obvious advantages for its noise suppression ability, some weaknesses still limit its application in synthetic aperture radar (SAR) image processing. The main challenge is that FOC can only suppress S$\alpha $S noises when all angular frequencies are zeros. In order to provide a statistic with noise robustness for SAR image processing, we propose a new concept of fractional order spectrum of cumulant (FOSC). FOSC can cancel the impact of S$\alpha $S noises without the limitation of angular frequencies in theory, which means FOSC can represent the local features more accurately. Furthermore, a FOSC-based SAR image registration is proposed to verify the advantage of FOSC. First, the 2-D formats of FOSC with different angular frequencies are calculated to provide more image features. Second, a multi-frequency pyramid array is designed to utilize the additional information in FOSC images, which can be used to detect more accurate keypoints. Third, the local descriptors based on FOSC are designed, which are constructed by accumulating orientation histograms of gradients based on FOSC and re-arranging elements in feature vectors regularly based on the orientation assignment. Finally, rotation consistencies are designed and used to eliminate the mismatching points after Euclidean distance-based matching of feature vectors. The proposed method is compared with four state-of-the-art methods. Experimental results show that the proposed method has achieved an impressive SAR image registration with 16.47%–48.37% gains when evaluating using the similarity of stitched areas.
Yan Zhao 0012, Xinbo Li, Shigang Wang 0003
IEEE Trans. Geosci. Remote. Sens.4
2024 5-D Epanechnikov Mixture-of-Experts in Light Field Image Compression
abstract
In this study, we propose a modeling-based compression approach for dense/lenslet light field images captured by Plenoptic 2.0 with square microlenses. This method employs the 5-D Epanechnikov Kernel (5-D EK) and its associated theories. Owing to the limitations of modeling larger image block using the Epanechnikov Mixture Regression (EMR), a 5-D Epanechnikov Mixture-of-Experts using Gaussian Initialization (5-D EMoE-GI) is proposed. This approach outperforms 5-D Gaussian Mixture Regression (5-D GMR). The modeling aspect of our coding framework utilizes the entire EI and the 5D Adaptive Model Selection (5-D AMLS) algorithm. The experimental results demonstrate that the decoded rendered images produced by our method are perceptually superior, outperforming High Efficiency Video Coding (HEVC) and JPEG 2000 at a bit depth below 0.06bpp.
Boning Liu 0001, Yan Zhao 0012, Xiaomeng Jiang, Xingguang Ji, Shigang Wang 0003, Yebin Liu
IEEE Trans. Image Process.5
2023 A Novel Intelligent Assessment Based on Audio-Visual Data for Chinese Zither Fingerings
Wenting Zhao 0003, Shigang Wang 0003, Yan Zhao 0012, Tianshu Li
ICIG (4)2
2023 YOLO-DA: An Efficient YOLO-Based Detector for Remote Sensing Object Detection
abstract
In the past few decades, many efficient object detectors have been proposed for natural scene image object detection. However, due to the complex scenes and high interclass similarity of optical remote sensing (RS) images, applying these detectors to optical RS images directly is not very effective. Most of the recent detectors pursue higher accuracy while ignoring the balance between detection accuracy and speed, which hinders the practical application of these detectors, especially in embedded devices. To meet these challenges, a fast and accurate detector based on YOLO (You Only Look Once) with decoupled attention head (YOLO-DA) is proposed, which effectively improves detection performance while only introducing minimal complexity. Specifically, an attention module at the end of the detector is designed for guiding a neural network to extract more efficient features from the complex background while also minimizing the amount of additional computation. Moreover, a lightweight decoupled detection head with enhanced classification and localization capability is developed to detect objects with high interclass similarity. In the experiments, the proposed method effectively solves the problem of high interclass similarity and improves the mAP by 6.8% on the fine-grained optical RS dataset SIMD, compared with YOLOv5-L. In addition, the proposed method improves the mAP by 1.0%, 1.7% and 0.6% on the other three publicly open optical RS datasets, respectively. Experimental results on detection accuracy and inference time demonstrate that our method achieves the best trade-off between detection performance and speed.
Jiehua Lin, Yan Zhao 0012, Shigang Wang 0003
IEEE Geosci. Remote. Sens. Lett.3
2023 Less Data-Dependent Seismic Noise Suppression Method Based on Transfer Learning With Attention Mechanism
abstract
Deep learning (DL) exhibits excellent performance in seismic noise suppression, and DL successes are attributed to its ability to learn rich representations from a large amount of data. However, obtaining numerous high-quality labeled data is challenging owing to confidentiality, regional sensitivity, and manual labeling, which limits the capability of DL. To reduce data dependency and improve network generalization, this study proposes a novel denoising architecture based on small-sample transfer learning (TL). The proposed architecture uses a fully pretrained model on the source data as a feature extractor, and then copies and transfers the rich features from the extractor to the denoiser for fine-tuning on the target data. Moreover, to reduce the discrepancy between two different data and better reuse the transferred features, a noise attention block (NAB) is proposed to regularize the representations. The results of multiregion experiment indicate that the proposed network leads to a significant improvement in denoising performance, essentially outperforming existing denoising methods; additionally, it exhibits strong generalization for different types and regions of seismic noise. Moreover, the proposed method can effectively address the data dependency issue, thus, providing great potential for real-time processing or small device applications.
Qiankun Feng, Shigang Wang 0003, Yue Li 0003
IEEE Trans. Geosci. Remote. Sens.2
2023 3D Holoscopic Image Compression Based on Gaussian Mixture Model
abstract
We introduce a Gaussian Mixture Model (GMM) framework for 3D holoscopic image compression in this paper. The elemental-images of the 3D holoscopic image are predicted using GMM and the parameters of GMM are estimated using the common Expectation-Maximization (EM) algorithm. GMM Model Optimization (GMO) is used in this framework to select the optimal number of distributions and avoid local optimum of EM at the same time. A three-dimensional distribution-rotation based decomposition is proposed to change covariance parameters to meaningful features and improve the coding efficiency. The features and the remaining parameters of the GMM are encoded using fixed-length bits. A feature-based dictionary is proposed in this framework to match the similar Gaussian distributions utilizing the similar GMM features. And the offsets of the matched distributions are recorded as motion vectors to replace the similar areas in the elemental-images of the 3D holoscopic image. The residual between the original image and the prediction is encoded using Screen Content Coding Extension of High Efficiency Video Coding (HEVC-SCC). Experimental results show that our method performs better than HEVC-SCC, two coding methods based on pseudo-sequences and a state-of-the-art content-based compression method with Gaussian process regression.
Yan Zhao 0012, Shigang Wang 0003
IEEE Trans. Multim.3
2022 An improved augmented-reality method of inserting virtual objects into the scene with transparent objects
abstract
In augmented reality, the insertion of virtual objects into the real scene needs to meet the requirements of visual consistency. The virtual objects rendered by the augmented reality system should be consistent with the illumination of the real scene. However, for complex scenes, it is not enough to just complete the illumination estimation. When there are transparent objects in the real scene, the difference in refractive index and roughness of transparent objects will influence the effect of the virtual and real fusion. To tackle this problem, this paper proposes a new approach to jointly estimate the illumination and transparent material for inserting virtual objects into the real scene. We solve for the material parameters of objects and illumination simultaneously by nesting microfacet model and hemispherical area illumination model into inverse path tracing. Although there is no geometry model of light sources in the recovered geometry model, the proposed hemispherical area illumination model can be used to recover scene appearance. Multiple experiments on both virtual and real-world datasets verify that the proposed approach subjectively and objectively performs better than the state-of-the-art method.
Yan Zhao 0012, Shigang Wang 0003
VR3
2022 4D Epanechnikov Mixture Regression in LF Image Compression
abstract
With the emergence of light field imaging in recent years, the compression of its elementary image array (EIA) has become a significant problem. Our coding framework includes modeling and reconstruction. For the modeling, the covariance-matrix form of the 4D Epanechnikov kernel (4D EK) and its correlated statistics were deduced to obtain the 4-D Epanechnikov mixture models (4-D EMMs). A 4D Epanechnikov mixture regression (4D EMR) was proposed based on this 4D EK, and a 4D adaptive model selection (4D AMLS) algorithm was designed to realize the optimal modeling for a pseudo video sequence (PVS) of the extracted key-EIA. A linear function based reconstruction (LFBR) was proposed based on the correlation between adjacent elementary images (EIs). The decoded images realized a clear outline reconstruction and superior coding efficiency compared to high-efficiency video coding (HEVC) and JPEG 2000 below approximately 0.05 bpp. This work realized an unprecedented theoretical application by (1) proposing the 4D Epanechnikov kernel theory, (2) exploiting the 4D Epanechnikov mixture regression and its application in the modeling of the pseudo video sequence of light field images, (3) using 4D adaptive model selection for the optimal number of models, and (4) employing a linear function-based reconstruction according to the content similarity.
Boning Liu 0001, Yan Zhao 0012, Xiaomeng Jiang, Shigang Wang 0003
IEEE Trans. Circuits Syst. Video Technol.4
2021 Aerial Image Object Detection Based on Superpixel-Related Patch
Jiehua Lin, Yan Zhao 0012, Shigang Wang 0003, Meimei Chen, Hongbo Lin, Zhihong Qian
ICIG (1)3
2021 3-D Epanechnikov Mixture Regression in integral imaging compression
Boning Liu 0001, Yan Zhao 0012, Xiaomeng Jiang, Shigang Wang 0003
J. Vis. Commun. Image Represent.4
2021 Three-dimensional Epanechnikov mixture regression in image coding
abstract
Kernel methods have been studied extensively in recent years. We propose a three-dimensional (3-D) Epanechnikov Mixture Regression (EMR) based on our Epanechnikov Kernel (EK) and realize a complete framework for image coding. In our research, we deduce the covariance-matrix form of 3-D Epanechnikov kernels and their correlated statistics to obtain the Epanechnikov mixture models. To apply our theories to image coding, we propose the 3-D EMR which can better model an image in smaller blocks compared with the conventional Gaussian Mixture Regression (GMR). The regressions are all based on our improved Expectation-Maximization (EM) algorithm with mean square error optimization. Finally, we design an Adaptive Mode Selection (AMS) algorithm to realize the best model pattern combination for coding. Our recovered image has clear outlines and superior coding efficiency compared to JPEG below 0.25bpp. Our work realizes an unprecedented theory application by: (1) enriching the theory of Epanechnikov kernel, (2) improving the EM algorithm using MSE optimization, (3) exploiting the EMR and its application in image coding, and (4) AMS optimal modeling combined with Gaussian and Epanechnikov kernel.
Boning Liu 0001, Yan Zhao 0012, Xiaomeng Jiang, Shigang Wang 0003
Signal Process.4
2021 Image compression based on Gaussian mixture model constrained using Markov random field
abstract
We introduce a Gaussian Mixture Model (GMM) constrained by Markov Random Field (MRF) framework for image compression in this paper. The image is predicted using GMM with MRF and the parameters of the GMM are estimated using an adjusted Expectation-Maximization (EM) algorithm. Mixture Model Optimization (MMO) is used in this framework to select the optimal number of distributions and avoid local optimum of EM at the same time. Parameters are encoded using fixed-length bits. A codebook is used to improve the coding efficiency of the covariance parameters. The residual between the original image and the prediction is encoded using High Efficiency Video Coding (HEVC) intra coding. Experimental results show that our method performs better than our previous work, HEVC, JPEG 2000 and Better Portable Graphics (BPG) which is an improved version of HEVC.
Yan Zhao 0012, Shigang Wang 0003
Signal Process.3
2021 An Improved Augmented-Reality Framework for Differential Rendering Beyond the Lambertian-World Assumption
abstract
In augmented reality, it is important to achieve visual consistency between inserted virtual objects and the real scene. As specular and transparent objects can produce caustics, which affect the appearance of inserted virtual objects, we herein propose a framework for differential rendering beyond the Lambertian-world assumption. Our key idea is to jointly optimize illumination and parameters of specular and transparent objects. To estimate the parameters of transparent objects efficiently, the psychophysical scaling method is introduced while considering visual characteristics of the human eye to obtain the step size for estimating the refractive index. We verify our technique on multiple real scenes, and the experimental results show that the fusion effects are visually consistent.
Yan Zhao 0012, Shigang Wang 0003
IEEE Trans. Vis. Comput. Graph.3
2019 An Image Coding Approach Based on Mixture-of-experts Regression Using Epanechnikov Kernel
abstract
In this paper, we propose an optimal modeling framework for image compression using EMM (Epanechnikov Mixture Model). Epanechnikov Kernel and its correlated statistics are basement of our Epanechnikov Mixture Regression (EMR). In our scheme, the stochastic processes of the pixel values are modelled as an EMM with K experts in three-dimensional space and then we use EMR to search for the optimal solution, whose parameters are determined through EM (Expectation-Maximization) algorithm. In the process of regression, the conditional density is the regression kernel function. Experimental results show that the proposed scheme is effective especially for the image with complex texture without consuming extra bits compared to Gaussian Mixture Regression (GMR).
Boning Liu 0001, Yan Zhao 0012, Xiaomeng Jiang, Shigang Wang 0003
ICASSP4
2019 Image Compression Using GMM Model Optimization
abstract
A Gaussian Mixture Model (GMM)-based framework for image compression is proposed in this paper. The image is predicted using GMM whose parameters are estimated using common Expectation-Maximization (EM) algorithm and encoded with fixed length bits. We introduce a new GMM Model Optimization (GMO) measure to select the optimal number of models and avoid local optimum of EM at the same time. The encoding cost of the residual and parameters are considered in GMO which is demonstrated to be near concave and effective. A parameter dictionary is designed to utilize the correlation of the parameters to improve the coding efficiency. The residual between the original image and the GMM image is encoded using High Efficiency Video Coding (HEVC) intra coding. Experimental results show that our method performs better than HEVC.
Yan Zhao 0012, Shigang Wang 0003
ICASSP3
2019 Video-based, Occlusion-robust Multi-view Stereo Using Inner-boundary Depths of Textureless Areas
abstract
Occlusions and poor textures are two main problems in multi-view stereo reconstruction. This paper presents a video-based solution to address both challenges in depth estimation. We focus on reconstructing accurate inner boundaries of visible textureless areas, particularly for occluded background, by leveraging the reliable depths of object edges. This is done by efficiently respecting two local cues with complementary advantages, i.e. smoothness and density of recovered surfaces. The inner-boundary depths are finally utilized to infer dense geometry without wrong connections between objects. This method only relies on low-level techniques, e.g. intra-view interpolation and inter-view propagation of depths. Experiments indicate its superiority in terms of both depth discontinuities near object silhouettes and surface smoothness in homogeneous regions compared to the state of the art.
Shigang Wang 0003, Yan Zhao 0012
ICASSP2
2019 A New Method to Expand the Showing Range of a Virtual Reconstructed Image in Integral Imaging
Lizhong Zhang, Shigang Wang 0003, Wei Wu 0018, Tianshu Li
ICIG (2)2
2019 Illumination estimation for augmented reality based on a global illumination model
Yan Zhao 0012, Shigang Wang 0003
Multim. Tools Appl.3
2017 Sparse Acquisition Integral Imaging System
Shigang Wang 0003, Wei Wu 0018, Tianshu Li, Lizhong Zhang
ICIG (3)2
2017 A Feature-Based Coding Algorithm for Face Image
Henan Li, Shigang Wang 0003, Yan Zhao 0012, Chuxi Yang, Aobo Wang
ICIG (2)2
2017 Feature-Based Facial Image Coding Method Using Wavelet Transform
Chuxi Yang, Yan Zhao 0012, Shigang Wang 0003
ICIG (2)3
2017 Sar image change detection method based on visual attention
abstract
Change detection is a hot issue and is of great significance in remote sensing. The logarithm operation is a valid way to reduce the influence of multiplicative noise in the Synthetic Aperture Radar (SAR) image. However, changed areas with high gray level values will be weakened due to the nature of the logarithmic function. In this paper, a SAR image change detection framework based on visual attention is proposed. In the proposed method, the SAR image change detection is finished with an extreme method with darkness and brightness on the vision. The main process can be divided into two parts according to dark and bright image patches. The dark changed areas are validly detected via a weighted logarithmic function, which has strong noise immunity. The weak bright changes are taken as noise. The saliency extraction is applied on the initial SAR image patches to enhance the bright changed areas whereas others present murky background. Then bright changed areas can be validly detected using kernel fuzzy c-means (KFCM), in which the cross-time similarities function between image patches is used. Finally, two change maps can be added to obtain final result. The real SAR image pairs of Suzhou area are used to verify proposed change detection method. The experimental results demonstrate the effectiveness of the proposed method.
Yan Zhang 0052, Chao Wang 0004, Shigang Wang 0003, Hong Zhang 0001, Meng Liu 0005
IGARSS3
2011 Spatial error concealment for stereoscopic video coding based on pixel matching
Yan Zhao 0012, Shigang Wang 0003, Hexin Chen
J. Supercomput.3