Anhong Wang

dblp:26/2493 · DBLP profile ↗
← Back
81ranked-venue papers
5as first author
31since 2021 · last 2026
0000-0001-5413-0490ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 52 · 2 first-author · 16 since 2021Artificial intelligence and machine learning · 16 · 1 first-author · 9 since 2021Computer networks · 4 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 2 since 2021Systems, architecture and hardware · 3Databases, data management, data science and information retrieval · 3Security and privacy · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Learning a joint mutual-guidance enhancement network for degraded low-light color image and low-resolution depth map
Bintao Chen, Lijun Zhao 0002, Jinjing Zhang, Anhong Wang, Huihui Bai 0001
Eng. Appl. Artif. Intell.4
2026 DDSF-Net: dual-domain deepfake detection via semantic suppression and spatial-frequency mapping
Zhifeng Xing, Li Liu 0029, Yingchun Wu, Ching-Chun Chang, Anhong Wang, Chin-Chen Chang 0001
Expert Syst. Appl.5
2026 Visual Intelligence-Guided RFID Multitag Spatial Position Measurement and Reading Performance Optimization
abstract
Radio Frequency Identification (RFID) is a core technology in the perception layer of the Passive Internet of Things (Passive IoT). The complex external environments and an increase in the number of tags would cause channel contention and data conflicts during the reading process, which significantly affects the positioning accuracy and reading performance, resulting in the loss of data information stored in the tags. Although there are many methods to improve the reading performance of RFID system, most of them evaluate the reading performance through anti-collision protocol to optimize the position of reader and antenna. However, when moving antennas and readers read multi-tag, the solutions are neither efficient nor reliable. To improve the performance of RFID systems, this paper proposes an improved multi-tag position measuremnet for UNet networks, which can optimize the reading performance of RFID system. Firstly, for the electronic interference in complex environments, an experimental platform assisted visual intelligence is designed to construct the RFID multi-tag position measurement and performance analysis system. Secondly, the image deblurring for the UNet network improved by Implicit Neural Representations (INR) and Residual Fast Fourier Transform (Res FFT) is proposed to improve the image quality of the degraded multi-tag image. Finally, the 3D coordinates of the tags are found by YOLOv9 in the image coordinate system, which are subsequently converted to actual 3D distributions. It can improve the spatial distribution of RFID tags to Combine the prior knowledge of RFID 3D space in visual intelligence with the indicators of reading performance, thereby better guiding more effective physical tags placement methods. Experimental results demonstrate that the PSNR of the proposed method is 30.17 dB, and at least 2% better than the state-of-the-art algorithms, which indicates that the method proposed can accurately obtain the 3D distribution of multi-tag. Our system can capture the multi-tag 3D distribution corresponding to the maximum reading distance, thereby guiding the 3D structure distribution of multi-tag to enhance RFID reading performance.
Lin Li 0052, Lijun Zhao 0002, Qinxiong Lu, Anhong Wang, Junji Li, Xiao Zhuang, Di Zhou 0006, Tianju Yang
IEEE Internet Things J.6
2026 An interpretable depth map super-resolution method via unrolling dual-boundary consistency constrained optimization
Lijun Zhao 0002, Jinjing Zhang, Huihui Bai 0001, Anhong Wang
Pattern Recognit.5
2026 A survey on image compressive sensing: From classical theory to the latest explicable deep learning
Lijun Zhao 0002, Xinlu Wang, Jinjing Zhang, Huihui Bai 0001, Anhong Wang
Pattern Recognit.6
2026 Reflectance Prediction-Based Knowledge Distillation for Robust 3D Object Detection in Compressed Point Clouds
abstract
Regarding intelligent transportation systems, low-bitrate transmission via lossy point cloud compression is vital for facilitating real-time collaborative perception among connected agents, such as vehicles and infrastructures, under restricted bandwidth. In existing compression transmission systems, the sender lossily compresses point coordinates and reflectance to generate a transmission code stream, which faces transmission burdens from reflectance encoding and limited detection robustness due to information loss. To address these issues, this paper proposes a 3D object detection framework with reflectance prediction-based knowledge distillation (RPKD). We compress point coordinates while discarding reflectance during low-bitrate transmission, and feed the decoded non-reflectance compressed point clouds into a student detector. The discarded reflectance is then reconstructed by a geometry-based reflectance prediction (RP) module within the student detector for precise detection. A teacher detector with the same structure as the student detector is designed for performing reflectance knowledge distillation (RKD) and detection knowledge distillation (DKD) from raw to compressed point clouds. Our cross-source distillation training strategy (CDTS) equips the student detector with robustness to low-quality compressed data while preserving the accuracy benefits of raw data through transferred distillation knowledge. Experimental results on the KITTI and DAIR-V2X-V datasets demonstrate that our method can boost detection accuracy for compressed point clouds across multiple code rates. We will release the code publicly at https://github.com/HaoJing-SX/RPKD.
Anhong Wang, Yifan Zhang 0036, Donghan Bu, Junhui Hou
IEEE Trans. Image Process.2
2026 Harnessing Group-Oriented Consistency Constraints for Semi-Supervised Semantic Segmentation in CdZnTe Semiconductors
Man Liu 0003, Huihui Bai 0001, Anhong Wang, Yunchao Wei, Yao Zhao 0001
IEEE Trans. Image Process.5
2025 Dual-Edge Consistency Constrained Unfolding Network for Depth Map Super-Resolution
Lijun Zhao 0002, Jinjing Zhang, Huihui Bai 0001, Anhong Wang
ICIG (1)5
2025 Learning A Decomposition-Driven Two Stages Unfolding Artifact Removal Network for Compressed Images
Lijun Zhao 0002, Jinjing Zhang, Anhong Wang
ICIG (1)6
2025 Learning A Deep Second-Order Unfolding Model for Arbitrary-Scale Depth Map Super-Resolution
Lijun Zhao 0002, Jinjing Zhang, Huihui Bai 0001, Anhong Wang
ICXR5
2025 Energy efficiency optimization of aerial intelligent reflecting surface-assisted communications based on multi-agent deep reinforcement learning
Suyue Li, Guangqian Li, Yunguang Xi, Anhong Wang
Comput. Networks4
2025 Generative adversarial network with circuitous feature collection for image steganographic cost learning
Li Liu 0029, Yingchun Wu, Ching-Chun Chang, Anhong Wang, Chin-Chen Chang 0001
Neurocomputing6
2025 Boosting 3D Object Detection With Semantic-Aware Multi-Branch Framework
abstract
In autonomous driving, LiDAR sensors are vital for acquiring 3D point clouds, providing reliable geometric information. However, traditional sampling methods of preprocessing often ignore semantic features, leading to detail loss and ground point interference in 3D object detection. To address this, we propose a multi-branch two-stage 3D object detection framework using a Semantic-aware Multi-branch Sampling (SMS) module and multi-view consistency constraints. The SMS module includes random sampling, Density Equalization Sampling (DES) for enhancing distant objects, and Ground Abandonment Sampling (GAS) to focus on non-ground points. The sampled multi-view points are processed through a Consistent KeyPoint Selection (CKPS) module to generate consistent keypoint masks for efficient proposal sampling. The first-stage detector uses multi-branch parallel learning with multi-view consistency loss for feature aggregation, while the second-stage detector fuses multi-view data through a Multi-View Fusion Pooling (MVFP) module to precisely predict 3D objects. The experimental results on the KITTI dataset and Waymo Open Dataset show that our method achieves excellent detection performance improvement for a variety of backbones, especially for low-performance backbones with simple network structures. The code will be publicly available at https://github.com/HaoJing-SX/SMS.
Anhong Wang, Lijun Zhao 0002, Yakun Yang, Donghan Bu, Yifan Zhang 0036, Junhui Hou
IEEE Trans. Circuits Syst. Video Technol.2
2025 Joint Deep-Unfolding Optimization Learning for Depth Map Arbitrary-Scale Super-Resolution
Lijun Zhao 0002, Jinjing Zhang, Anhong Wang, Huihui Bai 0001
IEEE Trans. Multim.4
2024 High-capacity multi-MSB predictive reversible data hiding in encrypted domain for triangular mesh models
Guoyou Zhang, Xiaoxue Cheng, Anhong Wang, Xuenan Zhang, Li Liu 0029
J. Vis. Commun. Image Represent.4
2024 Reversible data hiding in encrypted images with block-based bit-plane reallocation
Li Liu 0029, Yingchun Wu, Chin-Chen Chang 0001, Anhong Wang
Multim. Tools Appl.5
2024 Progressive Complementary Knowledge Aggregation for CdZnTe Defect Segmentation
abstract
Automatic quality inspection of industrial products is an indispensable part of modern manufacturing. Cadmium zinc telluride (CdZnTe) crystal is an important industrial raw material, but the special photosensitive properties of CdZnTe make it show different defect boundaries under different lighting angles, which poses challenges for quality inspection. In this article, we propose progressive complementary knowledge aggregation (PCKA) for CdZnTe defect segmentation, which is model-agnostic. First, the 12 images of CdZnTe crystal with different lighting angles are fed into the preliminary aggregation net to aggregate unique pixel-level clues. Second, we use a latent aggregation net to acquire the feature-level complementary clues under the guidance of the pixel-level clues within latent space. Such a learning paradigm is an effective solution for the special photosensitive properties of CdZnTe crystal. Extensive experiments on self-collected dataset demonstrate the effectiveness and efficiency of our PCKA compared with other solutions.
Feng Li 0037, Man Liu 0003, Huihui Bai 0001, Yunchao Wei, Anhong Wang, Shijie Ma, Yao Zhao 0001
IEEE Trans. Ind. Informatics6
2023 Edge-Guided Interpretable Neural Network for Image Compressive Sensing Reconstruction
Xinlu Wang, Lijun Zhao 0002, Jinjing Zhang, Anhong Wang
ICIG (3)5
2023 A Joint Model-Driven Unfolding Network for Degraded Low-Quality Color-Depth Images Enhancement
abstract
Inspired by multi-task learning, degraded low-quality color-depth images enhancement tasks are transformed as a joint color-depth optimization model by using maximum a posteriori estimation. This model is optimized alternatively in an iterative way to get the solutions of CGD-SR task and Low-Brightness Color Image Enhancement (LBC-IE) task. The whole iterative optimization procedure is expanded as a joint model-driven unfolding network. Many experimental results have confirmed that high-resolution reconstruction of the depth map and the enhancement of low-brightness image can be realized simultaneously in one network. Furthermore, the proposed method with network interpretability can exceed that of many inexplicable CGD-SR methods and LBC-IE methods.
Lijun Zhao 0002, Jinjing Zhang, Anhong Wang
ICIP5
2023 Explainable Unfolding Network For Joint Edge-Preserving Depth Map Super-Resolution
abstract
Although color-guided depth map super-resolution methods based on deep learning have achieved great progress, these methods are not explainable and their super-resolution results are suffered from texture-copying and boundary blurring problems. In order to alleviate these problems, we propose an explainable depth map super-resolution network by unfolding edge-constrained optimization model, dubbed Ex-DSRNet. First, we propose an edge reconstruction network to recover accurate edge information, then feed edge information and color information into different depth map reconstruction networks to reconstruct depth information. Finally, a high-fidelity network is proposed to fuse the above two kinds of depth information to obtain high-quality depth map. A large number of experimental results have demonstrated that the proposed Ex-DSRNet can compete against many state-of-the-art depth map super-resolution methods in term of root mean square error.
Lijun Zhao 0002, Jinjing Zhang, Anhong Wang
ICME5
2023 WDU-Net: Wavelet-Guided Deep Unfolding Network for Image Compressed Sensing Reconstruction
Xinlu Wang, Lijun Zhao 0002, Jinjing Zhang, Anhong Wang
PRCV (6)5
2023 Deep Arbitrary-Scale Unfolding Network for Color-Guided Depth Map Super-Resolution
Lijun Zhao 0002, Jinjing Zhang, Bintao Chen, Anhong Wang
PRCV (10)5
2023 Learning deep texture-structure decomposition for low-light image restoration and enhancement
Lijun Zhao 0002, Jinjing Zhang, Anhong Wang, Huihui Bai 0001
Neurocomputing4
2023 Boundary-constrained interpretable image reconstruction network for deep compressive sensing
Lijun Zhao 0002, Xinlu Wang, Jinjing Zhang, Anhong Wang, Huihui Bai 0001
Knowl. Based Syst.4
2023 End-to-end translation of human neural activity to speech with a dual-dual generative adversarial network
Yina Guo, Anhong Wang, Wenwu Wang 0001
Knowl. Based Syst.4
2023 Joint depth map super-resolution method via deep hybrid-cross guidance filter
Lijun Zhao 0002, Jinjing Zhang, Anhong Wang, Huihui Bai 0001
Pattern Recognit.5
2022 Outage Analysis of NOMA-Enabled Backscatter Communications With Intelligent Reflecting Surfaces
abstract
Intelligent reflecting surface (IRS) has emerged as a potential technology to achieve smart wireless communications and high energy efficiency. On the other hand, nonorthogonal multiple access (NOMA)-enabled backscatter communications have shown a great potential in large-scale Internet of Things (IoT) networks. In this article, we consider a downlink IRS-assisted backscatter communication with NOMA. We further consider a two-user scenario with channel disparity from the base station. We first derive the probability density function of the sum of the modulus of reflected channels, where each channel follows the Rayleigh distribution with dissimilar variances. The respective and generalized closed-form outage probability expressions are derived for the considered scenario. Simulation results validate the accuracy of the analytical outage probability expressions. We demonstrate that the far user can achieve a superior performance with the increase of reflecting elements or the reflection coefficients.
Suyue Li, Lina Bariah, Sami Muhaidat, Anhong Wang, Jie Liang 0001
IEEE Internet Things J.4
2022 AS-Net: An attention-aware downsampling network for point clouds oriented to classification tasks
Yakun Yang, Anhong Wang, Donghan Bu, Zewen Feng, Jie Liang 0001
J. Vis. Commun. Image Represent.2
2022 LMDC: Learning a multiple description codec for deep learning-based image compression
Lijun Zhao 0002, Jinjing Zhang, Huihui Bai 0001, Anhong Wang, Yao Zhao 0001
Multim. Tools Appl.4
2021 MRANet: Multi-atrous residual attention Network for stereo image super-resolution
Luyao Ning, Anhong Wang, Lijun Zhao 0002, Weimin Xue, Donghan Bu
J. Vis. Commun. Image Represent.2
2021 Artifact and Detail Attention Generative Adversarial Networks for Low-Dose CT Denoising
abstract
Generative adversarial networks are being extensively studied for low-dose computed tomography denoising. However, due to the similar distribution of noise, artifacts, and high-frequency components of useful tissue images, it is difficult for existing generative adversarial network-based denoising networks to effectively separate the artifacts and noise in the low-dose computed tomography images. In addition, aggressive denoising may damage the edge and structural information of the computed tomography image and make the denoised image too smooth. To solve these problems, we propose a novel denoising network called artifact and detail attention generative adversarial network. First, a multi-channel generator is proposed. Based on the main feature extraction channel, an artifacts and noise attention channel and an edge feature attention channel are added to improve the denoising network's ability to pay attention to the noise and artifacts features and edge features of the image. Additionally, a new structure called multi-scale Res2Net discriminator is proposed, and the receptive field in the module is expanded by extracting the multi-scale features in the same scale of the image to improve the discriminative ability of discriminator. The loss functions are specially designed for each sub-channel of the denoising network corresponding to its function. Through the cooperation of multiple loss functions, the convergence speed, stability, and denoising effect of the network are accelerated, improved, and guaranteed, respectively. Experimental results show that the proposed denoising network can preserve the important information of the low-dose computed tomography image and achieve better denoising effect when compared to the state-of-the-art algorithms.
Xiong Zhang 0006, Zefang Han, Hong Shangguan, Xinglong Han, Xueying Cui, Anhong Wang
IEEE Trans. Medical Imaging6
2020 Lightweight efficient network for defect classification of polarizers
abstract
Summary In industrial production, detecting polarizer defects online and in real time is necessary. Existing methods of detecting polarizer defects based on deep learning can ensure the accuracy of classification; however, there are several issues associated with these methods. These include the models having low detection speed, consuming large amounts of memory, and being difficult to be transplanted into online detection systems. To solve the aforementioned problems, a lightweight efficient network (LWEN) structure based on deep learning was designed, which improves the standard convolution layer and the fully connected (FC) layer to minimize the training model size and increase the speed of classification without reducing the accuracy of classification. First, a new building block, the shunt module, was designed to build the LWEN. Subsequently, a global average pooling layer was used to reduce the spatial resolution to 1 before the FC layer. These key technologies were designed to reduce the number of network parameters and minimize the model size of the network. Experimental results show that the proposed LWEN outperforms the state‐of‐the‐art approaches in terms of classification accuracy, speed, and model size.
Ruizhen Liu, Zhiyi Sun, Anhong Wang, Qianlai Sun
Concurr. Comput. Pract. Exp.3
2020 Light field all-in-focus image fusion based on spatially-guided angular information
Yingchun Wu, Yumei Wang, Jie Liang 0001, Ivan V. Bajic, Anhong Wang
J. Vis. Commun. Image Represent.5
2020 Super resolution of single depth image based on multi-dictionary learning with edge feature regularization
Anhong Wang, Hong Shangguan, Yingchun Wu, Donghong Li, Youcheng Wu, Jie Liang 0001
Multim. Tools Appl.2
2020 Joint Raindrop and Haze Removal From a Single Image
abstract
In a recent study, it was shown that, with adversarial training of an attentive generative network, it is possible to convert a raindrop degraded image into a relatively clean one. However, in real world, raindrop appearance is not only formed by individual raindrops, but also by the distant raindrops accumulation and the atmospheric veiling, namely haze. Current methods are limited in extracting accurate features from a raindrop degraded image with background scene, the blurred raindrop regions, and the haze. In this paper, we propose a new model for an image corrupted by the raindrops and the haze, and introduce an integrated multi-task algorithm to address the joint raindrop and haze removal (JRHR) problem by combining an improved estimate of the atmospheric light, a modified transmission map, a generative adversarial network (GAN) and an optimized visual attention network. The proposed algorithm can extract more accurate features for both sky and non-sky regions. Experimental evaluation has been conducted to show that the proposed algorithm significantly outperforms state-of-the-art algorithms on both synthetic and real-world images in terms of both qualitative and quantitative measures.
Yina Guo, Xiaowen Ren, Anhong Wang, Wenwu Wang 0001
IEEE Trans. Image Process.4
2020 Censor-Based Cooperative Multi-Antenna Spectrum Sensing with Imperfect Reporting Channels
abstract
The present contribution proposes a spectrally efficient censor-based cooperative spectrum sensing (C-CSS) approach in a sustainable cognitive radio network that consists of multiple antenna nodes and experiences imperfect sensing and reporting channels. In this context, exact analytic expressions are first derived for the corresponding probability of detection, probability of false alarm, and secondary throughput, assuming that each secondary user (SU) sends its detection outcome to a fusion center only when it has detected a primary signal. Capitalizing on the findings of the analysis, the effects of critical measures, such as the detection threshold, the number of SUs, and the number of employed antennas, on the overall system performance are also quantified. In addition, the optimal detection threshold for each antenna based on the Neyman-Pearson criterion is derived and useful insights are developed on how to maximize the system throughput with a reduced number of SUs. It is shown that the C-CSS approach provides two distinct benefits compared with the conventional sensing approach, i.e., without censoring: i) the sensing tail problem, which exists in imperfect sensing environments, can be mitigated; and ii) less SUs are ultimately required to obtain higher secondary throughput, rendering the system more sustainable.
Omar Alhussein, Paschalis C. Sofotasios, Sami Muhaidat, Paul D. Yoo, Jie Liang 0001, Anhong Wang
IEEE Trans. Sustain. Comput.7
2019 Deep Multiple Description Coding by Learning Scalar Quantization
abstract
In this paper, we propose a deep multiple description coding framework, whose quantizers are adaptively learned via the minimization of multiple description compressive loss. Firstly, our framework is built upon auto-encoder networks, which have multiple description multi-scale dilated encoder network and multiple description decoder networks. Secondly, two entropy estimation networks are learned to estimate the informative amounts of the quantized tensors, which can further supervise the learning of multiple description encoder network to represent the input image delicately. Thirdly, a pair of scalar quantizers accompanied by two importance-indicator maps is automatically learned in an end-to-end self-supervised way. Finally, multiple description structural dissimilarity distance loss is imposed on multiple description decoded images in pixel domain for diversified multiple description generations rather than on feature tensors in feature domain, in addition to multiple description reconstruction loss. Through testing on two commonly used datasets, it is verified that our method is beyond several state-of-the-art multiple description coding approaches in terms of coding efficiency.
Lijun Zhao 0002, Huihui Bai 0001, Anhong Wang, Yao Zhao 0001
DCC3
2019 Error Analysis of NOMA-Based User Cooperation with SWIPT
abstract
The present contribution analyzes the performance of non-orthogonal multiple access (NOMA)-based user cooperation with simultaneous wireless information and power transfer (SWIPT). In particular, we consider a two-user NOMA-based cooperative SWIPT scenario, in which the near user acts as a SWIPT-enabled relay that assists the farthest user. In this context, we derive analytic expressions for the pairwise error probability (PEP) of both users assuming the both amplify-and-forward (AF) and decode-and-forward (DF) relay protocols. The derived expressions are expressed in closed-form and have a tractable algebraic representation which renders them convenient to handle both analytically and numerically. In addition to this, we derive a simple asymptotic closed-form expression for the PEP in the high signal-to-noise ratio (SNR) regime which provide useful insights on the impact of the involved parameters on the overall system performance. Capitalizing on this, we subsequently quantify the maximum achievable diversity order of both users. It is shown that numerical and simulation results corroborate the derived analytic expressions. Furthermore, the offered results provide interesting insights into the error rate performance of each user, which are expected to be useful in future designs and deployments of NOMA based SWIPT systems.
Suyue Li, Lina Bariah, Sami Muhaidat, Paschalis C. Sofotasios, Jie Liang 0001, Anhong Wang
DCOSS6
2019 Censor-Based Multi-Antenna Cooperative Spectrum Sensing over Erroneous Feedback Channels
abstract
We propose a spectrally efficient censor-based cooperative spectrum sensing (C-CSS) approach for a sustainable cognitive radio network that consists of multiple antenna nodes and experiences imperfect sensing and reporting channels. First, analytic expressions are derived for the corresponding probabilities of detection and false alarm, assuming that each secondary user sends its detection outcome to a fusion center only when it believes to have detected a primary user's signal. Second, we derive lower bounds for the probability of false alarm, where we show that a sensing tail problem, which exist in the conventional (non-censor-based) scheme, can be effectively mitigated with the aid of the proposed C-CSS scheme. Simulation results are presented to corroborate the derived analytic results, and to provide theoretical and technical insights that are useful for the design of cognitive radio networks.
Omar Alhussein, Paschalis C. Sofotasios, Sami Muhaidat, Paul D. Yoo, Jie Liang 0001, Anhong Wang
WCNC7
2019 Learning a virtual codec based on deep convolutional neural network to compress image
Lijun Zhao 0002, Huihui Bai 0001, Anhong Wang, Yao Zhao 0001
J. Vis. Commun. Image Represent.3
2019 Iterative range-domain weighted filter for structural preserving image smoothing and de-noising
Lijun Zhao 0002, Huihui Bai 0001, Anhong Wang, Yao Zhao 0001
Multim. Tools Appl.3
2019 Simultaneous color-depth super-resolution with conditional generative adversarial networks
Lijun Zhao 0002, Huihui Bai 0001, Jie Liang 0001, Bing Zeng 0001, Anhong Wang, Yao Zhao 0001
Pattern Recognit.5
2019 Local activity-driven structural-preserving filtering for noise removal and image smoothing
Lijun Zhao 0002, Huihui Bai 0001, Jie Liang 0001, Anhong Wang, Bing Zeng 0001, Yao Zhao 0001
Signal Process.4
2019 DCSN-Cast: Deep compressed sensing network for wireless video multicast
Hehe Wu, Anhong Wang, Jie Liang 0001, Suyue Li
Signal Process. Image Commun.2
2019 A Depth-Bin-Based Graphical Model for Fast View Synthesis Distortion Estimation
abstract
During 3-D video communication, transmission errors, such as packet loss, could happen to the texture and depth sequences. View synthesis distortion will be generated when these sequences are used to synthesize virtual views according to the depth-image-based rendering method. A depth-value-based graphical model (DVGM) has been employed to achieve the accurate packet-loss-caused view synthesis distortion estimation (VSDE). However, the DVGM models the complicated view synthesis processes at depth-value level, which costs too much computation and is difficult to be applied in practice. In this paper, a depth-bin-based graphical model (DBGM) is developed, in which the complicated view synthesis processes are modeled at depth-bin level so that it can be used for the fast VSDE with 1-D parallel camera configuration. To this end, several depth values are fused into one depth bin, and a depth-bin-oriented rule is developed to handle the warping competition process. Then, the properties of the depth bin are analyzed and utilized to form the DBGM. Finally, a conversion algorithm is developed to convert the per-pixel input depth value probability distribution into the depth-bin format. Experimental results verify that our proposed method is 8-$32\times $ faster and requires 17%-60% less memory than the DVGM, with exactly the same accuracy.
Jie Liang 0001, Yao Zhao 0001, Chunyu Lin, Anhong Wang
IEEE Trans. Circuits Syst. Video Technol.6
2019 Distortion Estimation-Based Adaptive Power Allocation for Hybrid Digital-Analog Video Transmission
abstract
Hybrid digital–analog (HDA) video transmission schemes have shown advantages in avoiding thecliff effect. However, most current HDA schemes assume perfect transmission of the digital signal, which is hardly the case in practice. In this paper, we propose an adaptive recursive distortion estimate for the HDA system (ARDE-HDA), which recursively estimates the decoder-side distortion from the encoder and adaptively allocates the transmission power between the digital and analog signals in HDA. First, we derive the closed-form expression of a recursive distortion estimation (RDE) method, which does not require the digital part of the HDA output to be decoded perfectly, and both transmission error and superposition process of digital and analog parts are taken into consideration. Then, based on the deduced RDE model, an adaptive power allocation is proposed for the digital and analog parts to minimize the decoder-side distortion. Finally, simulation results are presented, which show the accuracy of the proposed RDE model and the ARDE-HDA method. Our method can achieve an average of 4.10 dB gain over existing HDA methods and 15.14-dB gain over the Softcast method in terms of the peak signal-to-noise ratio.
Anhong Wang, Jie Liang 0001, Suyue Li, Xiong Zhang 0006
IEEE Trans. Circuits Syst. Video Technol.2
2019 Multiple Description Convolutional Neural Networks for Image Compression
abstract
Multiple description coding (MDC) is able to stably transmit signal in un-reliable and non-prioritized networks, which has been broadly studied for several decades. However, traditional MDC does not well leverage image's context features to generate multiple descriptions. In this paper, we propose a novel standard-compliant convolutional neural network-based MDC framework, which efficiently leverages image's context information to compress the image. First, multiple description generator network (MDGN) is designed to produce appearance-similar yet feature-different multiple descriptions automatically according to image's content, which are compressed by a standard codec. Second, we present multiple description reconstruction network (MDRN) including side reconstruction networks (SRNs) and central reconstruction network (CRN). When any one of two lossy descriptions is received at decoder, SRN network is used to improve the quality of this decoded lossy description by simultaneously removing compression artifact and up-sampling. Meanwhile, we utilize CRN network with two decoded descriptions as inputs for better reconstruction, if both of lossy descriptions are available. Third, multiple description virtual codec network is proposed to bridge the gap between MDGN network and MDRN network in order to train an end-to-end MDC framework. Here, two learning algorithms are provided to train our whole framework. In addition to structural dis-similarity loss function, the produced descriptions are used as opposing labels with multiple description distance loss function to regularize the training of MDGN network. These losses guarantee that the generated descriptions are structurally similar yet finely diverse. Experimental results show a great deal of objective and subjective quality measurements to validate the effectiveness of our framework.
Lijun Zhao 0002, Huihui Bai 0001, Anhong Wang, Yao Zhao 0001
IEEE Trans. Circuits Syst. Video Technol.3
2018 Recursive Distortion Estimation for Hybrid Digital-Analog Video Transmission
abstract
Recently, hybrid digital-analog (HDA) video transmission scheme has shown advantages in avoiding the cliff effect. However, most HDA schemes assume perfect transmission of digital signals, which is hardly the case in practice. This paper proposes a scheme named RDE-HDA that recursively estimates the decoder-side distortion from the encoder-side for HDA system, and both transmission error and superposition process of digital and analog parts are taken into consideration. Therefore, our method does not require the digital part of the HDA output to be transmitted losslessly, making the HDA more practical. We derive the closed-form expression of the recursive distortion estimation for HDA system. The accuracy of our method is verified by simulation results.
Anhong Wang, Jie Liang 0001
ICASSP2
2018 Deep hashing network for material defect image classification
abstract
Common non‐destructive material testing technology has some well‐known problems such as slow detection, low detection accuracy, and low level of information obtained. To solve these problems, this study applied recent advances in convolution neural networks to propose an effective deep learning network using casting datasets. The approach achieves non‐destructive material testing with automatic, intelligent detection technology. For most existing deep learning networks, an image is eventually transformed into a multidimensional visual feature vector for comparison and classification. However, such vectors may not optimally improve detection precision and speed, and can lead to significant storage problems. A deep hashing network is proposed in which images are mapped into compact binary codes. There are three key components: (i) a sub‐network with multiple convolution‐pooling layers to capture image representations; (ii) a hashing layer to generate compact binary hash codes; (iii) an encoder module to divide the image feature vector from the output of the sub‐network above into multiple branches, each encoded into one hash bit. Extensive experiments using a casting dataset show promising performance compared with the state‐of‐the‐art approach.
Zhiyi Sun, Anhong Wang, Ruizhen Liu, Qianlai Sun
IET Comput. Vis.3
2018 Reversible data hiding for VQ indices using hierarchical state codebook mapping
Binbin Xia, Anhong Wang, Chin-Chen Chang 0001, Li Liu 0029
Multim. Tools Appl.2
2018 Multi-source phase retrieval from multi-channel phaseless STFT measurements
Yina Guo, Anhong Wang, Wenwu Wang 0001
Signal Process.2
2017 Convolutional neural network-based depth image artifact removal
abstract
In 3D video coding and depth-based image rendering, the distortion of the compressed depth image often leads to wrong 3D warpping. In this paper, by generalizing the recent work of convolutional neural network (CNN)-based depth image up-sampling, we propose a CNN-based depth image artifact removal scheme, where both the compressed depth and color images are used to enhance the depth accuracy. The proposed CNN has two sub-networks: joint depth-color sub-network and joint depth sub-network. During the depth and color feature extraction, the gradient of the depth image is used as the input to color image, while the gradient of color image is used as the input of depth feature extraction. Such an exchange of gradient information improves the learned features. Experimental results in terms of both objective and subjective quality of the depth and color images verify the efficiency of the proposed method.
Lijun Zhao 0002, Jie Liang 0001, Huihui Bai 0001, Anhong Wang, Yao Zhao 0001
ICIP4
2017 Single depth image super-resolution with multiple residual dictionary learning and refinement
abstract
Learning-based image super-resolution methods often use large datasets to learn texture features. When these methods are applied to depth images, emphasis should be given on learning the geometrical structures at object boundaries, since depth images do not have much texture information. In this paper, we develop a scheme to learn multiple residual dictionaries from only one external image. After depth image super-resolution, some artifacts may appear. An adaptive depth map refinement method is then proposed to remove these artifacts along the depth edges, based on the shape-adaptive weighted median filtering method. Experimental results demonstrate the advantage of the proposed method over many other methods.
Lijun Zhao 0002, Huihui Bai 0001, Jie Liang 0001, Anhong Wang, Yao Zhao 0001
ICME4
2017 3D saliency detection based on background detection
Hongyun Lin, Chunyu Lin, Yao Zhao 0001, Anhong Wang
J. Vis. Commun. Image Represent.4
2017 Optimized phase-space reconstruction for accurate musical-instrument signal classification
Yina Guo, Qijia Liu, Anhong Wang, Chao-Li Sun, Wenyan Tian, Ganesh R. Naik, Ajith Abraham
Multim. Tools Appl.3
2017 Data hiding based on extended turtle shell matrix construction method
Li Liu 0029, Chin-Chen Chang 0001, Anhong Wang
Multim. Tools Appl.3
2017 Adaptive residual-based distributed compressed sensing for soft video multicasting over wireless networks
Anhong Wang, Suyue Li, Jie Liang 0001
Multim. Tools Appl.2
2017 Two-stage filtering of compressed depth images with Markov Random Field
Lijun Zhao 0002, Huihui Bai 0001, Anhong Wang, Yao Zhao 0001, Bing Zeng 0001
Signal Process. Image Commun.3
2016 Packetization strategies for MVD-based 3D video transmission
abstract
In multi-view video plus depth (MVD) format, virtual views are synthesized by the compressed texture videos and their associated depth through depth-image-based rendering. In this paper, we consider the setup where both the encoded texture and depth bitstreams experience packet losses during transmission. Different packetization strategies are investigated and a novel strategy is developed to improve error resilience of MVD-based video transmission, where texture data and its corresponding depth are put into the same packet. The size of texture plus associated depth data included in each packet needs to be less than the Maximum Transfer Unit (MTU). Experimental results demonstrate that our proposed packetization scheme yields a significant improvement in terms of both texture views and synthesized virtual views quality when fit in H.264/AVC.
Xue Zhang 0008, Yao Zhao 0001, Tammam Tillo, Chunyu Lin, Jimin Xiao, Anhong Wang
VCIP6
2016 Joint iterative guidance filtering for compressed depth images
abstract
In the general 3D scene, the correlation of depth image and corresponding color image exists, so many filtering methods have been proposed to improve the quality of depth images according to this correlation. Unlike the conventional methods, in this paper both depth and color information can be jointly employed to improve the quality of compressed depth image by the way of iterative guidance. Firstly, due to noises and blurring in the compressed image, a depth pre-filtering method is essential to remove artifact noises. Considering that the received geometry structure in the distorted depth image is more reliable than its color image, the color information is merged with depth image to get depth-merged color image. Then the depth image and its corresponding depth-merged color image can be used to refine the quality of the distorted depth image using joint iterative guidance filtering method. Therefore, the efficient depth structural information included in the distorted depth images are preserved relying on depth itself, while the corresponding color structural information are employed to improve the quality of depth image. We demonstrate the efficiency of the proposed filtering method by comparing objective and visual quality of the synthesized image with many existing depth filtering methods.
Lijun Zhao 0002, Huihui Bai 0001, Anhong Wang, Yao Zhao 0001
VCIP3
2016 Reversible data hiding scheme based on histogram shifting of n-bit planes
Li Liu 0029, Chin-Chen Chang 0001, Anhong Wang
Multim. Tools Appl.3
2016 Region-Aware 3-D Warping for DIBR
abstract
In 3-D video (3DV) applications, depth-image-based rendering (DIBR) has been widely employed to synthesize virtual views. However, this approach is performed in a frame-based way, meaning each whole frame is dealt with and the characteristics of different regions in the frame are ignored. As a result, redundant pixels in some regions are abused during the subsequent warping and blending stage. This paper proposes a region-aware 3-D warping approach for DIBR in which warped frames are reasonably divided beforehand so that only the indispensable regions are used. With the proposed scheme, it is possible to avoid noneffective and repeated pixels during the warping stage. In addition, the blending process is also saved. The experimental results show that compared to the state-of-the-art VSRS3.5 and VSRS-1D-fast algorithms, our approach can achieve significant computation savings without sacrificing synthesis quality.
Anhong Wang, Yao Zhao 0001, Chunyu Lin, Bing Zeng 0001
IEEE Trans. Multim.2
2015 A fast region-level 3D-warping method for depth-image-based rendering
abstract
In 3D video, depth-image-based rendering (DIBR) is widely employed in view synthesis to generate virtual views. However, the processing of this algorithm is based on the frame-level, and the characteristics in different regions cannot be fully taken into account before rendering. This drawback will lead to the unnecessary and redundant information in some regions being abused, which increases extra computation. This paper proposes a region-level 3D-warping method for DIBR, where regions are divided according to their characteristics. Then, only the necessary information in some important regions is utilized warping so that the redundant information could be avoided in the computation. Experimental results show that our approach is almost 4 times faster than VSRS-1D-fast, while declines 0.12 dB PSNR in the performance of synthesis views averagely. Hence, our method can achieve a good trade-off between the computation and view synthesis and will be especially useful for applications where the computation is the concern.
Anhong Wang, Yao Zhao 0001, Chunyu Lin
MMSP2
2015 A grouped-scalable secret image sharing scheme
Anhong Wang, Chin-Chen Chang 0001, Li Liu 0029
Multim. Tools Appl.2
2014 Two-Stage Multiview Image Compression Using Interview SIFT Matching
abstract
In this paper, a novel scheme of two-stage multiview image compression is proposed to create two-level reconstructed quality. Differently from the conventional multiview image compression algorithms, SIFT (Scale-Invariant Feature Transform) features matching from interview images are exploited to remove the correlations between multiple views. In the first stage coding, SIFT and RANSAC (RANdom SAmple Consensus) algorithms are combined to calculate the correlation matrix of interview, which then can be developed to obtain the coarse reconstruction of the current view. In the second stage coding, the reconstructed quality can be improved further by using the residual information. The experimental results have shown that at higher compression ratio, the proposed scheme can obtain better rate-distortion performance than intra coding in MVC (Multiview Video Coding). Furthermore, with the change of the compression ratio, the proposed scheme can achieve more stable reconstructed quality.
Huihui Bai 0001, Mengmeng Zhang 0008, Meiqin Liu 0002, Anhong Wang, Yao Zhao 0001
DCC4
2014 Multiple description video coding using correlation optimized temporal sampling
Huihui Bai 0001, Mengmeng Zhang 0008, Anhong Wang, Yao Zhao 0001
Sci. China Inf. Sci.3
2014 A novel real-time and progressive secret image sharing with flexible shadows based on compressive sensing
Li Liu 0029, Anhong Wang, Chin-Chen Chang 0001
Signal Process. Image Commun.2
2014 Wireless multicasting of video signals based on distributed compressed sensing
Anhong Wang, Bing Zeng 0001
Signal Process. Image Commun.1
2014 Multiple Description Video Coding Based on Human Visual System Characteristics
abstract
In this paper, a novel multiple description video coding scheme is proposed based on the characteristics of the human visual system (HVS). Due to the underlying spatial-temporal masking properties, human eyes cannot sense any changes below the just noticeable difference (JND) threshold. Therefore, at an encoder, only the visual information that cannot be predicted well within the JND tolerance needs to be encoded as redundant information, which leads to more effective redundancy allocation according to the HVS characteristics. Compared with the relevant existing schemes, the experimental results exhibit better performance of the proposed scheme at same bit rates, in terms of perceptual evaluation and subjective viewing.
Huihui Bai 0001, Weisi Lin, Mengmeng Zhang 0008, Anhong Wang, Yao Zhao 0001
IEEE Trans. Circuits Syst. Video Technol.4
2013 Directional block compressed sensing for image coding
abstract
Compared with traditional Nyquist sampling, compressed sensing (CS) enables a highly precise reconstruction of the signal from fewer measurements, suggesting great potential for efficient and simplistic data acquisition. In this paper, we propose a directional block-based compressed sensing (DBCS) scheme for image coding, where the directionalities inherently exhibited within image blocks are exploited as the “a priori” information. The image block is first directionally scanned following the dominating direction of its edges/textures. Then the vectorized image block is sampled by a block-based compressed sensing (BCS) method. At the decoder, each image block is recovered and then rearranged by the corresponding inverse-scan to obtain the recovered image. Experimental results show that the proposed DBCS scheme outperforms BCS due to the exploitation of the directional information within image blocks.
Anhong Wang, Kongfen Zhu, Chunyu Lin, Yao Zhao 0001
ISCAS2
2012 Multiple Description Video Coding Using Macro Block Level Correlation of Inter-/Intra-Descriptions
abstract
Multiple description coding (MDC) is a promising technology for robust transmission over error-prone channels, which has attracted a lot research interests. The basic idea of MDC is to how to utilize redundant information of the descriptions for robust transmission. In view of practical applications, many MDC approaches have been proposed compatible with a certain standard codec, especially H.264/AVC. In this paper, we attempt to develop a novel MD video codec with generalized compatibility, which aims to the effective redundancy allocation from inter-/intra-descriptions. In [1], the redundancy allocation may be not enough effective due to frame level. As a result, in this paper, the redundant information will be taken into account at MB level.
Huihui Bai 0001, Mengmeng Zhang 0008, Meiqin Liu 0002, Anhong Wang, Yao Zhao 0001
DCC4
2012 Entropy analysis on multiple description video coding based on pre- and post-processing
abstract
Multiple description (MD) video coding is a promising method to solve real-time video transmission over unreliable network. In the conventional MD video coding, the original video sequence can be split directly into two subsequences by odd and even means. Then the two sub-sequences can be compressed as two descriptions by the standard video encoder. The conventional MD scheme is simple to realize but it may lead to worse reconstructed quality when one description is lost. To solve this problem, the MD scheme based on pre- and post-processing is proposed in this paper. Before odd and even splitting, the original video sequence can be pre-processed by effective redundancy allocation, which is helpful for the estimation of the lost description. Furthermore, the entropy of the descriptions is used to analyse the rate-distortion performance of the two MD schemes. Lastly, the experimental results have shown the proposed MD scheme has better reconstructed quality when information lost has happened, while the conventional MD scheme has better compression efficiency when information can be transmitted accurately. It can be found that the experimental results can be consistent with the entropy analysis.
Huihui Bai 0001, Anhong Wang, Ajith Abraham
HIS2
2012 Stereo video coding using distributed compressive sensing with joint dictionary
abstract
For many practical applications, stereo-paired video is an important special case of multiview video coding (MVC). This paper presents a novel framework of stereo video coding based on distributed compressive sensing, which also can be easily extended to MVC. According to distributed video coding (DVC) at the encoder the video sequences from each view can be compressed independently without any communications between the cameras while at the decoder the inter-view correlation can be exploited for quality enhancement. Furthermore, due to compressive sensing (CS) principles, low complexity at the encoder side can result in low power consumption in the cameras, which may be promising in wireless camera sensor network. Here, the joint dictionary is applied in compressive sensing, which can make good use of inter-view and temporal correlation for better reconstruction quality of convex optimization. The experimental results validate the effectiveness of the proposed scheme with better performance than other compared schemes.
Huihui Bai 0001, Mengmeng Zhang 0008, Anhong Wang, Yao Zhao 0001
ICIP3
2010 GOP-Flexible Distributed Multiview Video Coding with Adaptive Side Information
Lili Meng, Yao Zhao 0001, Jeng-Shyang Pan 0001, Huihui Bai 0001, Anhong Wang
ICCCI (3)5
2009 Distributed Video Coding Based on Multiple Description
abstract
In our paper, we propose a novel distributed video coding (DVC) scheme using the theory of multiple description (MD), in which key frame is encoded by MD codec and transmitted over the corresponding channel. This scheme combines the advantage of DVC as well as robustness of MD, and exploits three different methods to generate multiple descriptions for the key frames that are essential to side information. Experiments demonstrate that it can get better performance than some general DVC methods. Besides, it demonstrates higher robustness in packet-loss channel than general DVC due to the MD algorithm.
Hongxia Ma, Yao Zhao 0001, Chunyu Lin, Anhong Wang
IAS4
2009 A Two-Description Distributed Video Coding
abstract
In this paper, a two-description distributed video coding (2D-DVC) is proposed to address the robust video transmission of low-power captures. The odd/even frame-splitting partitions a video into two subsequences to produce two descriptions. Each description consists of two parts, where part 1 is a zero-motion based H.264 coded bitstream of a subsequence and part 2 is a Wyner-Ziv coded bitstream of the other subsequence. As the redundant part, the Wyner-Ziv coded bitstream guarantees the lost subsequence is recovered when one description is lost. On the other hand, the redundancy degrades the rate-distortion performance as no loss occurs. Therefore, a residual 2D-DVC is employed to reduce the redundancy, where the difference of two subsequences is Wyner-Ziv encoded to generate the part 2 in each description. The experimental results show that the proposed schemes achieve better performance than the referenced one especially when the video motion is low. Moreover, our schemes maintain low-complexity encoding.
Anhong Wang, Yao Zhao 0001
IAS1
2009 Robust multiple description distributed video coding using optimized zero-padding
Anhong Wang, Yao Zhao 0001, Huihui Bai 0001
Sci. China Ser. F Inf. Sci.1
2006 Wavelet-Domain Distributed Video Coding with Motion-Compensated Refinement
abstract
In this paper, we propose a distributed video coding (DVC) paradigm based on lattice vector quantization in wavelet domain. In this framework, we use a fine and a coarse lattice vector quantizer to wavelet coefficients and the difference of two lattice quantizers is coded by turbo encoder. At decoder, side information is gradually updated by motion-compensated refinement. The refinement is built on partially decoded current frame and we give its matching strategy. Due to the refinement, burden of turbo encoding is cut down, bit rate is saved and the reconstruction is improved to some degree.
Anhong Wang, Yao Zhao 0001, Lei Wei 0006
ICIP1
2006 LVQ Based Distributed Video Coding with LDPC in Pixel Domain
Anhong Wang, Yao Zhao 0001
PRICAI1
2004 Non-linear predictor based on ANN in speech coding
abstract
In this paper, non-linear predictors based on different artificial neural network (ANN) are compared. General radical basis function (GERBF) neural network, a modified RBF network, is introduced. From the comparison with BP, RNN and RBF, it is obvious that GERBF has priority in the prediction of speech signal. The experiment results show: the speech coding systems based on ANN have better synthesized speech than ITU's G721 and the speech coding system based on GERBF has higher mean segmental SNR but less computation than that of other systems.
Linsheng Li, Zhiyi Sun, Anhong Wang
ICARCV3
2004 Direct data domain approach to space-time adaptive signal processing
abstract
A novel methodology utilizing the direct data domain approach to space-time adaptive processing (STAP) in airborne radar environments is presented in this paper. The deterministic least squares adaptive signal processing technique operates on a snapshot-by-snapshot basis to determine the adaptive weights for nulling interferences and estimating signal of interest (SOI), and eliminates the requirement for the secondary data from neighboring range cell. This is in contrast to conventional adaptive techniques where processing is done by taking the time averages as opposed to spatial averages. Simulation results illustrate the efficiency of interference suppression in nonhomogeneous environment.
Xiaoqin Wen, Anhong Wang, Linsheng Li, Chongzhao Han
ICARCV2