VLDB 2026 Research / reviewers in the wild / expert
Zulin Wang
dblp:37/10829
· DBLP profile ↗
87ranked-venue papers
0as first author
20since 2021 · last 2026
0000-0002-1328-7739ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 44 · 6 since 2021Computer networks · 23 · 7 since 2021Artificial intelligence and machine learning · 12 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 2 since 2021Theory of computation · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Ambiguity Function Analysis of AFDM Signals for Integrated Sensing and Communications
Haoran Yin 0001, Yanqun Tang, Yuanhan Ni, Zulin Wang, Gaojie Chen 0001, Jun Xiong 0002, Kai Yang 0004, Marios Kountouris, Yong Liang Guan 0001, Yong Zeng 0001 |
IEEE J. Sel. Areas Commun. | 4 |
| 2026 | Ambiguity Function Analysis of AFDM Under Pulse-Shaped Random ISAC SignalingabstractThis paper investigates the ambiguity function (AF) of the emerging affine frequency division multiplexing (AFDM) waveform for random integrated sensing and communication (ISAC) signaling under a pulse shaping regime. Specifically, we first derive the closed-form expression of the average squared discrete period AF (DPAF) for AFDM waveform without pulse shaping, revealing that the AF depends on the parameterc1and the kurtosis of random communication data, while being independent of the parameterc2. As a step further, we conduct a comprehensive analysis on the DPAFs of various waveforms, including AFDM, orthogonal frequency division multiplexing (OFDM) and orthogonal chirp-division multiplexing (OCDM). Our results indicate that all three waveforms exhibit the same number of regular depressions in the sidelobes of their DPAFs, which incurs performance loss for detecting and estimating weak targets. However, the AFDM waveform can flexibly control the positions of depressions by adjusting the parameterc1, which motivates a novel design approach of the AFDM parameters to mitigate the adverse impact of depressions of the strong target on the weak target. Furthermore, the closed-form expressions of the average squared DPAFs for pulse-shaped AFDM, OFDM and OCDM waveforms are derived, which demonstrates that the pulse shaping filter generates the shaped mainlobe along the delay axis and the rapid roll-off sidelobes along the Doppler axis. Numerical results verify the effectiveness of our theoretical analysis and proposed design methodology for the AFDM waveform. Yuanhan Ni, Fan Liu 0005, Haoran Yin 0001, Yanqun Tang, Yuanfang Ma, Zulin Wang |
IEEE Trans. Wirel. Commun. | 6 |
| 2026 | A Secure Affine Frequency Division Multiplexing System for Next-Generation Wireless CommunicationsabstractAffine frequency division multiplexing (AFDM) has garnered significant attention due to its superior performance in high-mobility scenarios, as well as multiple waveform parameters that provide greater degrees of freedom for system design. This paper proposes a novel secure affine frequency division multiplexing (SE-AFDM) system, which enhances physical-layer security by dynamically varying an AFDM pre-chirp parameter across subcarriers and over AFDM symbols. In the SE-AFDM system, the pre-chirp parameter is dynamically generated from a codebook controlled by a long-period pseudo-noise (LPPN) sequence. Instead of applying spreading in the data domain, our parameter-domain spreading approach provides additional security while maintaining reliability and high spectral efficiency. We also propose a synchronization framework to solve the problem of reliably and rapidly synchronizing the dynamic parameter in fast time-varying channels. The theoretical derivations prove that unsynchronized eavesdroppers cannot eliminate the nonlinear impact of the time-varying parameter and further provide useful guidance for codebook design. Simulation results demonstrate the security advantages of the proposed SE-AFDM system in high-mobility scenarios, and the hardware prototype validates the effectiveness of the proposed SE-AFDM system and the proposed synchronization framework. Zulin Wang, Yuanhan Ni, Qu Luo, Yuanfang Ma, Xiaosi Tian, Pei Xiao 0001 |
IEEE Trans. Wirel. Commun. | 2 |
| 2026 | Accurate Characterization and Low-Complexity MMSE Equalization of ISCI in Doubly-Dispersive Channelsabstractthis paper, we investigate the Inter Symbol and Carrier Interference (ISCI) of doubly-dispersive channels in highly dynamic scenarios from the continuous and discrete perspectives, respectively, and thoroughly analyze its impact on the performance of communications systems. Due to its robustness against the Doppler effect in time-varying channels, Orthogonal Time Frequency Space (OTFS) modulation has gained significant research interest, with many equalization algorithms proposed. However, existing studies either fail to fully account for ISCI or suffer from prohibitive complexity. In this study, we accurately quantify the ISCI of doubly-dispersive channels and provide an in-depth analysis from continuous and discrete channel models, respectively. Based on this, we propose a low-complexity Minimum Mean Square Error (MMSE) equalization to equalize ISCI in doubly-dispersive channels. The proposed algorithm demonstrates a reduction in complexity of the MMSE equalization fromO(M3N3)toO(MNL2bw), whereLbwdenotes the bandwidth of the time domain channel matrix. Concurrently, it has been demonstrated to significantly reduce the bit error rate (BER), thereby enhancing communication performance. Ziqin Yan, Fan Jiang 0003, Zulin Wang, Zijun Gong, Cheng Li 0005, Xiaofeng Tao 0001 |
IEEE Trans. Wirel. Commun. | 3 |
| 2026 | An Anti-Interference AFDM System: Interference Impacts Analyses and Parameter OptimizationabstractThis paper proposes an anti-interference affine frequency division multiplexing (AFDM) system to ensure reliability and resource efficiency under malicious high-power interference originating from adversarial devices in high-mobility scenarios. Closed-form expressions of interferences in the discrete affine Fourier transform (DAFT) domain are derived by utilizing the stationary phase principle and the Affine Fourier transform convolution theorem, which indicates that interference impacts can be classified into stationary and non-stationary categories. On this basis, we reveal the analytical relationship between packet throughput and the parameters of spread spectrum and error correction coding in our proposed anti-interference system, which enables the design of a parameter optimization algorithm that maximizes packet throughput. For reception, by jointly utilizing the autocorrelation function of spreading sequence and the cyclic-shift property of AFDM input-output relation, we design a linear-complexity correlation-based DAFT domain detector (CDD) capable of achieving full diversity gain, which performs correlation-based equalization to avoid matrix inversion. Numerical results validate the accuracy of the derived closed-form expressions and verify that the proposed anti-interference AFDM system could achieve high packet throughput under interference in high-mobility scenarios. Peng Yuan 0010, Zulin Wang, Yuanhan Ni |
IEEE Trans. Wirel. Commun. | 2 |
| 2025 | A Secure Affine Frequency Division Multiplexing for Wireless Communication SystemsabstractThis paper introduces a secure affine frequency division multiplexing (SE-AFDM) for wireless communication systems to enhance communication security. Besides configuring the parameter$c_{1}$to obtain communication reliability under doubly selective channels, we also utilize the time-varying parameter$c_{2}$to improve the security of the communications system. The derived input-output relation shows that the legitimate receiver can eliminate the nonlinear impact introduced by the time-varying$c_{2}$without losing the bit error rate (BER) performance. Moreover, it is theoretically proved that the eavesdropper cannot separate the time-varying$c_{2}$and random information symbols, such that the BER performance of the eavesdropper is severely deteriorated. Meanwhile, the analysis of the effective signal-to-interference-plus-noise ratio (SINR) of the eavesdropper illustrates that the SINR decreases as the value range of$c_{2}$expands. Numerical results verify that the proposed SE-AFDM waveform has significant security while maintaining good BER performance in high-mobility scenarios. Zulin Wang, Yuanfang Ma, Xiaosi Tian, Yuanhan Ni |
ICC | 2 |
| 2025 | An Integrated Sensing and Communications System Based on Affine Frequency Division MultiplexingabstractThis paper proposes an integrated sensing and communications (ISAC) system based on affine frequency division multiplexing (AFDM) waveform. To this end, a metric set is designed according to not only the maximum tolerable delay/Doppler, but also the weighted spectral efficiency as well as the outage/error probability of sensing and communications. This enables the analytical investigation of the performance trade-offs of AFDM-ISAC system using the derived analytical relation among metrics and AFDM waveform parameters. Moreover, by revealing that delay and the integral/fractional parts of normalized Doppler can be decoupled in the affine Fourier transform-Doppler domain, an efficient estimation method is proposed for our AFDM-ISAC system, whose unambiguous Doppler can break through the limitation of subcarrier spacing. Theoretical analyses and numerical results verify that our proposed AFDM-ISAC system may significantly enlarge unambiguous delay/Doppler while possessing good spectral efficiency and peak-to-sidelobe level ratio in high-mobility scenarios. Yuanhan Ni, Peng Yuan 0010, Qin Huang 0002, Fan Liu 0005, Zulin Wang |
IEEE Trans. Wirel. Commun. | 5 |
| 2024 | A Zero-Shot NAS Method for SAR Ship Detection Under Polynomial Search ComplexityabstractOne-shot neural architecture search (NAS) has achieved impressive results in the field of synthetic aperture radar (SAR) ship detection. However, it is a challenge to balance resource consumption and search speed. To address this issue, we propose a zero-shot NAS method for searching the backbone of SAR ship detection model, named as ZeroSARNas, which is implemented via a multi-characterization proxy and an integer linear programming (ILP) search algorithm. Specifically, we first design the multi-characterization proxy for network capacity prediction, which takes advantage of information entropy and local intrinsic dimensionality (LID) of feature maps, named as ELID proxy, to obtain a more comprehensive understanding of each candidate module in the search space. We then formulate the NAS problem as a ‘0–1’ ILP problem which maximizes the ELID value under the different constraints such as parameters to quickly identify the optimal network. The experimental results show that the detection accuracy of the networks found by ZeroSARNas on the SSDD, HRSID, and LS-SSDD-v1.0 datasets can reach 98.59%, 91.30%, and 75.11% in mean average precision (mAP) with only 1.23 M, 1.75 M, and 1.29 M parameters, respectively. The proposed method reduces the search time from several GPU days or hours to 10.0 seconds, achieving competitive search efficiency. Hang Wei 0008, Zulin Wang, Gengxin Hua, Yuanhan Ni |
IEEE Signal Process. Lett. | 2 |
| 2023 | Learning Noise-Induced Reward Functions for Surpassing Demonstrations in Imitation LearningabstractImitation learning (IL) has recently shown impressive performance in training a reinforcement learning agent with human demonstrations, eliminating the difficulty of designing elaborate reward functions in complex environments. However, most IL methods work under the assumption of the optimality of the demonstrations and thus cannot learn policies to surpass the demonstrators. Some methods have been investigated to obtain better-than-demonstration (BD) performance with inner human feedback or preference labels. In this paper, we propose a method to learn rewards from suboptimal demonstrations via a weighted preference learning technique (LERP). Specifically, we first formulate the suboptimality of demonstrations as the inaccurate estimation of rewards. The inaccuracy is modeled with a reward noise random variable following the Gumbel distribution. Moreover, we derive an upper bound of the expected return with different noise coefficients and propose a theorem to surpass the demonstrations. Unlike existing literature, our analysis does not depend on the linear reward constraint. Consequently, we develop a BD model with a weighted preference learning technique. Experimental results on continuous control and high-dimensional discrete control tasks show the superiority of our LERP method over other state-of-the-art BD methods. Liangyu Huo, Zulin Wang, Mai Xu |
AAAI | 2 |
| 2023 | Blind VQA on 360° Video via Progressively Learning From Pixels, Frames, and VideoabstractBlind visual quality assessment (BVQA) on 360° video plays a key role in optimizing immersive multimedia systems. When assessing the quality of 360° video, human tends to perceive its quality degradation from the viewport-based spatial distortion of each spherical frame to motion artifact across adjacent frames, ending with the video-level quality score, i.e., a progressive quality assessment paradigm. However, the existing BVQA approaches for 360° video neglect this paradigm. In this paper, we take into account the progressive paradigm of human perception towards spherical video quality, and thus propose a novel BVQA approach (namely ProVQA) for 360° video via progressively learning from pixels, frames and video. Corresponding to the progressive learning of pixels, frames and video, three sub-nets are designed in our ProVQA approach, i.e., the spherical perception aware quality prediction (SPAQ), motion perception aware quality prediction (MPAQ) and multi-frame temporal non-local (MFTN) sub-nets. The SPAQ sub-net first models the spatial quality degradation based on spherical perception mechanism of human. Then, by exploiting motion cues across adjacent frames, the MPAQ sub-net properly incorporates motion contextual information for quality assessment on 360° video. Finally, the MFTN sub-net aggregates multi-frame quality degradation to yield the final quality score, via exploring long-term quality correlation from multiple frames. The experiments validate that our approach significantly advances the state-of-the-art BVQA performance on 360° video over two datasets, the code of which has been public in https://github.com/yanglixiaoshen/ProVQA. Li Yang 0014, Mai Xu, Shengxi Li, Zulin Wang |
IEEE Trans. Image Process. | 5 |
| 2023 | Omnidirectional Image Super-Resolution via Latitude Adaptive NetworkabstractOmnidirectional images (ODI), also known as 360 images, have recently attracted extensive attention from both academia and industry. However, due to storage and transmission limitations, ODIs are usually at extremely low resolution. Thus, it is necessary to restore a high-resolution ODI from a low-resolution ODI, i.e., omnidirectional image super-resolution (ODI-SR). Different from traditional two-dimensional (2D) image SR, the challenge of ODI-SR is the nonuniformly distributed pixel density and geometric distortion across latitudes, which makes traditional SR methods difficult to be applied in ODI-SR. Towards ODI-SR, we propose in this paper a novel latitude-aware upscaling network, namely LAU-Net+, which fully considers the above characteristics of ODIs. In our network, different latitude bands can learn to adopt distinct upscaling factors, which significantly saves the computational resources and improves the SR efficiency. Specifically, a Laplacian multilevel pyramid network is introduced in which the upscaling factor is gradually increased with the number of levels. Each level is composed of a feature enhancement module (FEM), a drop-band decision module (DDM) and a high-latitude enhancement module (HEM). The FEM module serves to enhance the high-level features extracted from the input ODI, while the role of DDM is to dynamically drop the unnecessary high latitude bands and send the remained bands to the next level. The HEM is adopted to further enhance high-level features of dropped latitude bands with a lightweight architecture. In DDM, we develop a reinforcement learning scheme with a latitude adaptive reward to determine which band should be dropped. To the best of our knowledge, our method is the first work which considers the latitude characteristics for ODI-SR task. Extensive experimental results demonstrate that our LAU-Net+ achieves state-of-the-art results on ODI-SR both quantitatively and qualitatively on various ODI datasets. Xin Deng 0002, Hao Wang 0049, Mai Xu, Zulin Wang |
IEEE Trans. Multim. | 5 |
| 2023 | A Task-Agnostic Regularizer for Diverse Subpolicy Discovery in Hierarchical Reinforcement LearningabstractThe automatic subpolicy discovery approach in hierarchical reinforcement learning (HRL) has recently achieved promising performance on sparse reward tasks. This accelerates transfer learning and unsupervised intelligent creatures while eliminating the domain-specific knowledge constraint. Most previously developed approaches are demonstrated to suffer from collapsing into the situation where one subpolicy dominates the whole task, since they cannot ensure the diversity of different subpolicies. In contrast, this article proposes a task-agnostic regularizer (TAR) for learning diverse subpolicies in HRL. Specifically, we first formulate the discovery of diverse subpolicies as a trajectory inference problem and then propose a corresponding information-theoretic objective to encourage diversity. Subsequently, considering computability, we instantiate the objective as two simplifications for discrete and continuous action spaces. We extensively evaluate the proposed diversity-driven regularizer on three HRL task domains: 1) meta reinforcement learning; 2) hierarchical policy learning in the option framework; and 3) unsupervised subpolicy discovery. The extensive results obtained show that our TAR approach can improve upon the state-of-the-art performance on all three HRL domains without modifying any existing hyperparameters, indicating the wide applicability and robustness of our approach. Liangyu Huo, Zulin Wang, Mai Xu, Yuhang Song 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2022 | Viewport-Based CNN: A Multi-Task Approach for Assessing 360° Video QualityabstractFor 360° video, the existing visual quality assessment (VQA) approaches are designed based on either the whole frames or the cropped patches, ignoring the fact that subjects can only access viewports. When watching 360° video, subjects select viewports through head movement (HM) and then fixate on attractive regions within the viewports through eye movement (EM). Therefore, this paper proposes a two-staged multi-task approach for viewport-based VQA on 360° video. Specifically, we first establish a large-scale VQA dataset of 360° video, called VQA-ODV, which collects the subjective quality scores and the HM and EM data on 600 video sequences. By mining our dataset, we find that the subjective quality of 360° video is related to camera motion, viewport positions and saliency within viewports. Accordingly, we propose a viewport-based convolutional neural network (V-CNN) approach for VQA on 360° video, which has a novel multi-task architecture composed of a viewport proposal network (VP-net) and viewport quality network (VQ-net). The VP-net handles the auxiliary tasks of camera motion detection and viewport proposal, while the VQ-net accomplishes the auxiliary task of viewport saliency prediction and the main task of VQA. The experiments validate that our V-CNN approach significantly advances state-of-the-art VQA performance on 360° video and it is also effective in the three auxiliary tasks. Mai Xu, Lai Jiang 0004, Chen Li 0049, Zulin Wang, Xiaoming Tao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | Joint Learning of Multi-Level Tasks for Diabetic Retinopathy Grading on Low-Resolution Fundus ImagesabstractDiabetic retinopathy (DR) is a leading cause of permanent blindness among the working-age people. Automatic DR grading can help ophthalmologists make timely treatment for patients. However, the existing grading methods are usually trained with high resolution (HR) fundus images, such that the grading performance decreases a lot given low resolution (LR) images, which are common in clinic. In this paper, we mainly focus on DR grading with LR fundus images. According to our analysis on the DR task, we find that: 1) image super-resolution (ISR) can boost the performance of both DR grading and lesion segmentation; 2) the lesion segmentation regions of fundus images are highly consistent with pathological regions for DR grading. Based on our findings, we propose a convolutional neural network (CNN)-based method for joint learning of multi-level tasks for DR grading, called DeepMT-DR, which can simultaneously handle the low-level task of ISR, the mid-level task of lesion segmentation and the high-level task of disease severity classification on LR fundus images. Moreover, a novel task-aware loss is developed to encourage ISR to focus on the pathological regions for its subsequent tasks: lesion segmentation and DR grading. Extensive experimental results show that our DeepMT-DR method significantly outperforms other state-of-the-art methods for DR grading over three datasets. In addition, our method achieves comparable performance in two auxiliary tasks of ISR and lesion segmentation. Xiaofei Wang 0004, Mai Xu, Jicong Zhang, Lai Jiang 0004, Liu Li 0001, Mengxian He, Ningli Wang, Hanruo Liu, Zulin Wang |
IEEE J. Biomed. Health Informatics | 9 |
| 2021 | DeepVS2.0: A Saliency-Structured Deep Learning Method for Predicting Dynamic Visual Attention
Lai Jiang 0004, Mai Xu, Zulin Wang, Leonid Sigal |
Int. J. Comput. Vis. | 3 |
| 2021 | MFQE 2.0: A New Approach for Multi-Frame Quality Enhancement on Compressed VideoabstractThe past few years have witnessed great success in applying deep learning to enhance the quality of compressed image/video. The existing approaches mainly focus on enhancing the quality of a single frame, not considering the similarity between consecutive frames. Since heavy fluctuation exists across compressed video frames as investigated in this paper, frame similarity can be utilized for quality enhancement of low-quality frames given their neighboring high-quality frames. This task is Multi-Frame Quality Enhancement (MFQE). Accordingly, this paper proposes an MFQE approach for compressed video, as the first attempt in this direction. In our approach, we first develop a Bidirectional Long Short-Term Memory (BiLSTM) based detector to locate Peak Quality Frames (PQFs) in compressed video. Then, a novel Multi-Frame Convolutional Neural Network (MF-CNN) is designed to enhance the quality of compressed video, in which the non-PQF and its nearest two PQFs are the input. In MF-CNN, motion between the non-PQF and PQFs is compensated by a motion compensation subnet. Subsequently, a quality enhancement subnet fuses the non-PQF and compensated PQFs, and then reduces the compression artifacts of the non-PQF. Also, PQF quality is enhanced in the same way. Finally, experiments validate the effectiveness and generalization ability of our MFQE approach in advancing the state-of-the-art quality enhancement of compressed video. Zhenyu Guan 0002, Qunliang Xing, Mai Xu, Zulin Wang |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2021 | Saliency Prediction on Omnidirectional Image With Generative Adversarial Imitation LearningabstractWhen watching omnidirectional images (ODIs), subjects can access different viewports by moving their heads. Therefore, it is necessary to predict subjects' head fixations on ODIs. Inspired by generative adversarial imitation learning (GAIL), this paper proposes a novel approach to predict saliency of head fixations on ODIs, named SalGAIL. First, we establish a dataset for attention on ODIs (AOI). In contrast to traditional datasets, our AOI dataset is large-scale, which contains the head fixations of 30 subjects viewing 600 ODIs. Next, we mine our AOI dataset and discover three findings: (1) the consistency of head fixations are consistent among subjects, and it grows alongside the increased subject number; (2) the head fixations exist with a front center bias (FCB); and (3) the magnitude of head movement is similar across the subjects. According to these findings, our SalGAIL approach applies deep reinforcement learning (DRL) to predict the head fixations of one subject, in which GAIL learns the reward of DRL, rather than the traditional human-designed reward. Then, multi-stream DRL is developed to yield the head fixations of different subjects, and the saliency map of an ODI is generated via convoluting predicted head fixations. Finally, experiments validate the effectiveness of our approach in predicting saliency maps of ODIs, significantly better than 11 state-of-the-art approaches. Our AOI dataset and code of SalGAIL are available online at https://github.com/yanglixiaoshen/SalGAIL. Mai Xu, Li Yang 0014, Xiaoming Tao 0001, Yiping Duan, Zulin Wang |
IEEE Trans. Image Process. | 5 |
| 2021 | Graftage Coding for Distributed Storage SystemsabstractTo achieve various tradeoffs between storage and repair bandwidth, this article proposes to construct exact repair codes by grafting two codesC1andC2. By replacing certain nonzero entries in the generator matrix ofC1by zero, the repair bandwidth of the resulting grafting part decreases. However, it may no longer keep the maximum-distance-separable (MDS) property. As a result, the grafted codeC2takes these nonzero entries into account such that the entire graftage code can keep the MDS property. The relationship between the bandwidth reduction ofC1and the file size ofC2is derived to optimize graftage codes. Our analysis indicates that these graftage codes may provide better tradeoffs than space-sharing. Jiayi Rui, Qin Huang 0002, Zulin Wang |
IEEE Trans. Inf. Theory | 3 |
| 2021 | Joint Learning of 3D Lesion Segmentation and Classification for Explainable COVID-19 DiagnosisabstractGiven the outbreak of COVID-19 pandemic and the shortage of medical resource, extensive deep learning models have been proposed for automatic COVID-19 diagnosis, based on 3D computed tomography (CT) scans. However, the existing models independently process the 3D lesion segmentation and disease classification, ignoring the inherent correlation between these two tasks. In this paper, we propose a joint deep learning model of 3D lesion segmentation and classification for diagnosing COVID-19, called DeepSC-COVID, as the first attempt in this direction. Specifically, we establish a large-scale CT database containing 1,805 3D CT scans with fine-grained lesion annotations, and reveal 4 findings about lesion difference between COVID-19 and community acquired pneumonia (CAP). Inspired by our findings, DeepSC-COVID is designed with 3 subnets: a cross-task feature subnet for feature extraction, a 3D lesion subnet for lesion segmentation, and a classification subnet for disease diagnosis. Besides, the task-aware loss is proposed for learning the task interaction across the 3D lesion and classification subnets. Different from all existing models for COVID-19 diagnosis, our model is interpretable with fine-grained 3D lesion distribution. Finally, extensive experimental results show that the joint learning framework in our model significantly improves the performance of 3D lesion segmentation and disease classification in both efficiency and efficacy. Xiaofei Wang 0004, Lai Jiang 0004, Liu Li 0001, Mai Xu, Xin Deng 0002, Lisong Dai, Tianyi Li 0004, Zulin Wang, Pier Luigi Dragotti |
IEEE Trans. Medical Imaging | 10 |
| 2021 | Viewport-Dependent Saliency Prediction in 360° VideoabstractSaliency prediction in traditional images and videos has drawn extensive research interests in recent years. Few works have been proposed for saliency prediction over 360° videos. They focus on directly predicting fixations over the whole panorama. When viewing 360° videos, a person can only observe the content in her viewport, which means that only a fraction of the 360° scene can be seen at any given time. In this paper, we study human attention over viewport of 360° videos and propose a novel visual saliency model, dubbed viewport saliency, to predict fixations over 360° videos. Two contributions are introduced. First, we find that where people look is affected by the content and location of the viewport in 360° video. We study this over 200+ 360° videos viewed by 30+ subjects over two recent benchmark databases. Second, we propose a Multi-Task Deep Neural Network (MT-DNN) method for Viewport Saliency (VS) prediction in 360° video, which considers the input content and location of the viewport. Extensive experiments and analyses show that our method outperforms other state-of-the-art methods in this task. In particular, over the two recent 360° video databases, our MT-DNN raises the average CC score by 0.149 and 0.205, compared to SalGAN and DeepVS methods, respectively. Minglang Qiao, Mai Xu, Zulin Wang, Ali Borji |
IEEE Trans. Multim. | 3 |
| 2020 | Learning Diverse Sub-Policies via a Task-Agnostic Regularization on Action DistributionsabstractAutomatic sub-policy discovery has recently received much attention in hierarchical reinforcement learning (HRL). The conventional approaches to learning sub-policies suffer from collapsing into just one sub-policy dominating the whole task, lacking techniques to ensure the diversity of different subpolicies. In this paper, we formulate the discovery of diverse sub-policies as a trajectory inference. Then, we propose an information-theoretic objective based on action distributions to encourage diversity. Moreover, two simplifications are derived on discrete and continuous action space for reducing the computation. Finally, the experimental results show that the proposed approach can further improve the state-of-theart approaches without modifying existing hyperparameters on two different HRL domains, suggesting the wide applicability and robustness of our approach. Liangyu Huo, Zulin Wang, Mai Xu, Yuhang Song 0001 |
ICASSP | 2 |
| 2020 | Construction of Multiple-Burst-Correction Codes in Transform Domain and Its Relation to LDPC CodesabstractThis paper analyzes and explicitly constructs quasi-cyclic (QC) codes for correcting multiple bursts via matrix transformations. Our analysis demonstrates that the multiple-burst-correction capability of QC codes is determined by sub-matrices in the diagonal of their transformed parity-check matrices. By well designing these sub-matrices, the proposed QC codes are able to achieve optimal or asymptotically optimal multiple-burst-correction capability. Moreover, it proves that these codes can be QC low-density parity-check (QC-LDPC) codes, if the diagonal sub-matrices of their transformed parity-check matrices are Hadamard powers of base matrices. Analysis and simulation results show that our QC-LDPC codes perform well over not only random symbol error/erasure channels, but also burst channels. Liyuan Song, Qin Huang 0002, Zulin Wang |
IEEE Trans. Commun. | 3 |
| 2020 | A Meta-Learning Framework for Learning Multi-User Preferences in QoE Optimization of DASHabstractDynamic adaptive video streaming over hypertext transfer protocol (DASH) plays a key role in video transmission over the Internet. The conventional DASH adaptation approaches mainly focus on optimizing the overall quality of experience (QoE) for all client sides, neglecting the QoE diversity of different users. In this paper, we propose a meta-learning framework for multi-user preferences (MLMP) as a new DASH adaptation approach, which is able to optimize the diverse QoE of different users. Specifically, we first design a subjective experiment to analyze the difference of QoE preferences across users, in which QoE refers to the metrics of visual quality, fluctuation, and rebuffering events. Based on our findings, we formulate the QoE optimization of multi-user preferences as a multi-task deep reinforcement learning (DRL) problem. In our formulation, the QoE preference of each user is modeled in the overall QoE calculation via assigning the weights to the three QoE metrics. Then, the MLMP framework is developed to solve the proposed multi-task DRL problem, such that the preferences regarding visual quality, fluctuation, and rebuffering events can be optimized for different users in DASH adaptation. Finally, the simulation results show that the proposed approach outperforms state-of-the-art DASH adaptation approaches in satisfying the different users' QoE preferences regarding visual quality, fluctuation, and rebuffering events. Liangyu Huo, Zulin Wang, Mai Xu, Yong Li 0008, Zhiguo Ding 0001, Hao Wang 0049 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2020 | Understanding and Predicting the Memorability of Outdoor Natural ScenesabstractMemorability measures how easily an image is to be memorized after glancing, which may contribute to designing magazine covers, tourism publicity materials, and so forth. Recent works have shed light on the visual features that make generic images, object images or face photographs memorable. However, these methods are not able to effectively predict the memorability of outdoor natural scene images. To overcome this shortcoming of previous works, in this paper, we provide an attempt to answer: “what exactly makes outdoor natural scenes memorable”. To this end, we first establish a large-scale outdoor natural scene image memorability (LNSIM) database, containing 2,632 outdoor natural scene images with their ground truth memorability scores and the multi-label scene category annotations. Then, similar to previous works, we mine our database to investigate how low-, middle- and high-level handcrafted features affect the memorability of outdoor natural scenes. In particular, we find that the high-level feature of scene category is rather correlated with outdoor natural scene memorability, and the deep features learnt by deep neural network (DNN) are also effective in predicting the memorability scores. Moreover, combining the deep features with the category feature can further boost the performance of memorability prediction. Therefore, we propose an end-to-end DNN based outdoor natural scene memorability (DeepNSM) predictor, which takes advantage of the learned category-related features. Then, the experimental results validate the effectiveness of our DeepNSM model, exceeding the state-of-the-art methods. Finally, we try to understand the reason of the good performance for our DeepNSM model, and also study the cases that our DeepNSM model succeeds or fails to accurately predict the memorability of outdoor natural scenes. Jiaxin Lu 0003, Mai Xu, Zulin Wang |
IEEE Trans. Image Process. | 4 |
| 2020 | A Large-Scale Database and a CNN Model for Attention-Based Glaucoma DetectionabstractGlaucoma is one of the leading causes of irreversible vision loss. Many approaches have recently been proposed for automatic glaucoma detection based on fundus images. However, none of the existing approaches can efficiently remove high redundancy in fundus images for glaucoma detection, which may reduce the reliability and accuracy of glaucoma detection. To avoid this disadvantage, this paper proposes an attention-based convolutional neural network (CNN) for glaucoma detection, called AG-CNN. Specifically, we first establish a large-scale attention-based glaucoma (LAG) database, which includes 11 760 fundus images labeled as either positive glaucoma (4878) or negative glaucoma (6882). Among the 11 760 fundus images, the attention maps of 5824 images are further obtained from ophthalmologists through a simulated eye-tracking experiment. Then, a new structure of AG-CNN is designed, including an attention prediction subnet, a pathological area localization subnet, and a glaucoma classification subnet. The attention maps are predicted in the attention prediction subnet to highlight the salient regions for glaucoma detection, under a weakly supervised training manner. In contrast to other attention-based CNN methods, the features are also visualized as the localized pathological area, which are further added in our AG-CNN structure to enhance the glaucoma detection performance. Finally, the experiment results from testing over our LAG database and another public glaucoma database show that the proposed AG-CNN approach significantly advances the state-of-the-art in glaucoma detection. Liu Li 0001, Mai Xu, Hanruo Liu, Yang Li 0010, Xiaofei Wang 0004, Lai Jiang 0004, Zulin Wang, Ningli Wang |
IEEE Trans. Medical Imaging | 7 |
| 2019 | Image Saliency Prediction in Transformed Domain: A Deep Complex Neural Network MethodabstractThe transformed domain fearures of images show effectiveness in distinguishing salient and non-salient regions. In this paper, we propose a novel deep complex neural network, named SalDCNN, to predict image saliency by learning features in both pixel and transformed domains. Before proposing Sal-DCNN, we analyze the saliency cues encoded in discrete Fourier transform (DFT) domain. Consequently, we have the following findings: 1) the phase spectrum encodes most saliency cues; 2) a certain pattern of the amplitude spectrum is important for saliency prediction; 3) the transformed domain spectrum is robust to noise and down-sampling for saliency prediction. According to these findings, we develop the structure of SalDCNN, including two main stages: the complex dense encoder and three-stream multi-domain decoder. Given the new SalDCNN structure, the saliency maps can be predicted under the supervision of ground-truth fixation maps in both pixel and transformed domains. Finally, the experimental results show that our Sal-DCNN method outperforms other 8 state-of-theart methods for image saliency prediction on 3 databases. Lai Jiang 0004, Mai Xu, Zulin Wang |
AAAI | 4 |
| 2019 | Optimizing QoE of Multiple Users over DASH: A Meta-learning ApproachabstractDynamic adaptive video streaming over HTTP (DASH) plays a key role in video transmission over the Internet. The conventional DASH adaptation approaches concentrate on optimizing the overall quality of experience (QoE) for all client sides, neglecting the QoE diversity of different users. In this paper, we formulate the QoE optimization of multi-user preferences as a multi-task deep reinforcement learning problem, in which QoE refers to the metrics of visual quality, fluctuation and rebuffing events. Then, we propose a meta-learning framework for multi-user preferences (MLMP) as a new DASH adaptation approach. Finally, the simulation results show that the proposed approach outperforms state-of-the-art DASH adaptation approaches in satisfying the different users' QoE preferences regarding the three metrics. Liangyu Huo, Zulin Wang, Mai Xu, Zhiguo Ding 0001, Xiaoming Tao 0001 |
ICASSP | 2 |
| 2019 | A Deep Neural Network Based Maneuvering-target Tracking AlgorithmabstractIn the field of maneuvering-target tracking (MTT), the targets with changeable and uncertain maneuvering movements cannot be tracked precisely because there always exist time delays of maneuvering model estimation with traditional MT-T algorithms. To solve this problem, we propose a deep MTT (DeepMTT) algorithm based on a deep neural network, which can quickly track maneuvering targets once it has been well trained by abundant off-line trajectory data from existent ma-neuvering targets. To this end, we first build a Large-scale trajectory database to offer abundant off-line trajectory data for network training. Second, the DeepMTT algorithm is developed based on a deep neural network, which consists of three bidirectional long short-term memory layers, a filtering layer, a maxout layer and a linear output layer. The simulation results verify that our DeepMTT algorithm outperforms other state-of-the-art MTT algorithms. Jingxian Liu, Zulin Wang, Mai Xu, Jie Ren 0004 |
ICASSP | 2 |
| 2019 | Unsupervised User Clustering in Non-orthogonal Multiple AccessabstractNon-orthogonal multiple access (NOMA) is one of the most promising technologies in fifth-generation mobile communication system for its advantages in serving multiuser simultaneously and enhancing spectrum efficiency. In this paper, we investigate the optimization problem of sum-rate maximization for NOMA-based system, and mainly focus on user clustering. Inspired by the correlation features of users, we introduce machine learning in user clustering. We first develop an expectation maximization (EM) based algorithm for fixed user scenario. Then, the dynamic user scenario is considered and an online EM (OLEM) based clustering algorithm is proposed. Simulation results show that the proposed EM-based and OLEM-based algorithms outperform the state-of-the-art algorithms in fixed and dynamic user scenario, respectively. Jie Ren 0004, Zulin Wang, Mai Xu, Fang Fang 0005, Zhiguo Ding 0001 |
ICASSP | 2 |
| 2019 | Removing Rain in Videos: A Large-Scale Database and a Two-Stream ConvLSTM ApproachabstractRain removal has recently attracted increasing research attention, as it is able to enhance the visibility of rain videos. However, the existing learning based rain removal approaches for videos suffer from insufficient training data, especially when applying deep learning to remove rain. In this paper, we establish a large-scale video database for rain removal (LasVR), which consists of 316 rain videos. Then, we observe from our database that there exist the temporal correlation of clean content and similar patterns of rain across video frames. According to these two observations, we propose a two-stream convolutional long-and short-term memory (ConvLSTM) approach for rain removal in videos. The first stream is composed of the subnet for rain detection, while the second stream is the subnet of rain removal that leverages the features from the rain detection subnet. Finally, the experimental results on both synthetic and real rain videos show the proposed approach performs better than other state-of-the-art approaches. Mai Xu, Zulin Wang |
ICME | 3 |
| 2019 | Querying Policies Based on Sparse Matrices for Noisy 20 QuestionsabstractThis paper shows that the error probability of a querying policy for noisy 20 questions is upper bounded by the minimum Hamming distance of its querying matrix. Following this distance principle, sparse querying matrices with the row-column constraint are constructed for the scenarios, where only limited areas can be detected at each querying round. It demonstrates that the row-column constraint promises a large minimum distance to ensure low error probability of queries. Moreover, the proposed sparse matrices with random block coding provide unequal error-protection capability to further improve the querying accuracy under the scenarios of detecting half of the areas. Simulation results verify that the quantized mean squared errors of our proposed policies outperform those of the existing policies under the above scenarios. Qin Huang 0002, Simeng Zheng, Yuanhan Ni, Zulin Wang |
ISIT | 4 |
| 2019 | Pathology-Aware Deep Network Visualization and Its Application in Glaucoma Image Synthesis
Xiaofei Wang 0004, Mai Xu, Liu Li 0001, Zulin Wang, Zhenyu Guan 0002 |
MICCAI (1) | 4 |
| 2019 | Predicting Head Movement in Panoramic Video: A Deep Reinforcement Learning ApproachabstractPanoramic video provides immersive and interactive experience by enabling humans to control the field of view (FoV) through head movement (HM). Thus, HM plays a key role in modeling human attention on panoramic video. This paper establishes a database collecting subjects' HM in panoramic video sequences. From this database, we find that the HM data are highly consistent across subjects. Furthermore, we find that deep reinforcement learning (DRL) can be applied to predict HM positions, via maximizing the reward of imitating human HM scanpaths through the agent's actions. Based on our findings, we propose a DRL-based HM prediction (DHP) approach with offline and online versions, called offline-DHP and online-DHP. In offline-DHP, multiple DRL workflows are run to determine potential HM positions at each panoramic frame. Then, a heat map of the potential HM positions, named the HM map, is generated as the output of offline-DHP. In online-DHP, the next HM position of one subject is estimated given the currently observed HM position, which is achieved by developing a DRL algorithm upon the learned offline-DHP model. Finally, the experiments validate that our approach is effective in both offline and online prediction of HM positions for panoramic video, and that the learned offline-DHP model can improve the performance of online-DHP. Mai Xu, Yuhang Song 0001, Jianyi Wang, Minglang Qiao, Liangyu Huo, Zulin Wang |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2019 | An EM-Based User Clustering Method in Non-Orthogonal Multiple AccessabstractPower domain non-orthogonal multiple access (NOMA), with the ability to serve multiple users within one resource block, is one of the most promising technologies for the fifth generation. In this paper, we study the downlink millimeter wave (mmWave) NOMA-based system, where the base station sends messages to multiple clusters and serves multiple users simultaneously, and sum-rate maximization problem is investigated. Since users are multiplexed on one resource block, user clustering is important for NOMA and has a great influence on sum-rate optimization problem. Inspired by correlation features of users' spatial distributions in mmWave NOMA-based system, we introduce unsupervised learning method into user clustering. We first develop an Expectation Maximization (EM)-based algorithm in fixed user scenario. Then, the dynamic user scenario is introduced, which includes user reduction, increment and movement situations. After that, an online EM-based clustering algorithm is proposed to fast update user distribution parameters with lower computational complexity compared to the conventional complete re-clustering methods. Simulation results show that the proposed EM-based algorithm can improve the performance of NOMA-based system in fixed user scenario. In addition, the proposed online EM-based algorithm can achieve similar performance as the complete EM-based algorithm with less computational complexity in dynamic user scenario. Jie Ren 0004, Zulin Wang, Mai Xu, Fang Fang 0005, Zhiguo Ding 0001 |
IEEE Trans. Commun. | 2 |
| 2019 | Assessing Visual Quality of Omnidirectional VideosabstractIn contrast with traditional videos, omnidirectional videos enable spherical viewing direction with support for head-mounted displays, providing an interactive and immersive experience. Unfortunately, to the best of our knowledge, there are only a few visual quality assessment (VQA) methods, either subjective or objective, for omnidirectional video coding. This paper proposes both subjective and objective methods for assessing the quality loss in encoding an omnidirectional video. Specifically, we first present a new database, which includes the viewing direction data from several subjects watching omnidirectional video sequences. Then, from our database, we find a high consistency in viewing directions across different subjects. The viewing directions are normally distributed in the center of the front regions, but they sometimes fall into other regions, related to the video content. Given this finding, we present a subjective VQA method for measuring the difference mean opinion score (DMOS) of the whole and regional omnidirectional video, in terms of overall DMOS and vectorized DMOS, respectively. Moreover, we propose two objective VQA methods for the encoded omnidirectional video, in light of the human perception characteristics of the omnidirectional video. One method weighs the distortion of pixels with regard to their distances to the center of front regions, which considers human preference in a panorama. The other method predicts viewing directions according to the video content, and then the predicted viewing directions are leveraged to allocate weights to the distortion of each pixel in our objective VQA method. Finally, our experimental results verify that both the subjective and objective methods proposed in this paper advance the state-of-the-art VQA for omnidirectional videos. Mai Xu, Chen Li 0049, Zhenzhong Chen 0001, Zulin Wang, Zhenyu Guan 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2019 | Enhancing Quality for HEVC Compressed VideosabstractThe latest High Efficiency Video Coding (HEVC) standard has been increasingly applied to generate video streams over the Internet. However, HEVC compressed videos may incur severe quality degradation, particularly at low bit rates. Thus, it is necessary to enhance the visual quality of HEVC videos at the decoder side. To this end, this paper proposes a quality enhancement convolutional neural network (QE-CNN) method that does not require any modification of the encoder to achieve quality enhancement for HEVC. In particular, our QE-CNN method learns QE-CNN-I and QE-CNN-P models to reduce the distortion of HEVC I and P/B frames, respectively. The proposed method differs from the existing CNN-based quality enhancement approaches, which only handle intra-coding distortion and are thus not suitable for P/B frames. Our experimental results validate that our QE-CNN method is effective in enhancing quality for both I and P/B frames of HEVC videos. To apply our QE-CNN method in time-constrained scenarios, we further propose a time-constrained quality enhancement optimization (TQEO) scheme. Our TQEO scheme controls the computational time of QE-CNN to meet a target, meanwhile maximizing the quality enhancement. Next, the experimental results demonstrate the effectiveness of our TQEO scheme from the aspects of time control accuracy and quality enhancement under different time constraints. Finally, we design a prototype to implement our TQEO scheme in a real-time scenario. Mai Xu, Zulin Wang, Zhenyu Guan 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2019 | A Deep Learning Approach for Multi-Frame In-Loop Filter of HEVCabstractAn extensive study on the in-loop filter has been proposed for a high efficiency video coding (HEVC) standard to reduce compression artifacts, thus improving coding efficiency. However, in the existing approaches, the in-loop filter is always applied to each single frame, without exploiting the content correlation among multiple frames. In this paper, we propose a multi-frame in-loop filter (MIF) for HEVC, which enhances the visual quality of each encoded frame by leveraging its adjacent frames. Specifically, we first construct a large-scale database containing encoded frames and their corresponding raw frames of a variety of content, which can be used to learn the in-loop filter in HEVC. Furthermore, we find that there usually exist a number of reference frames of higher quality and of similar content for an encoded frame. Accordingly, a reference frame selector (RFS) is designed to identify these frames. Then, a deep neural network for MIF (known as MIF-Net) is developed to enhance the quality of each encoded frame by utilizing the spatial information of this frame and the temporal information of its neighboring higher-quality frames. The MIF-Net is built on the recently developed DenseNet, benefiting from its improved generalization capacity and computational efficiency. In addition, a novel block-adaptive convolutional layer is designed and applied in the MIF-Net, for handling the artifacts influenced by coding tree unit (CTU) structure in HEVC. Extensive experiments show that our MIF approach achieves on average 11.621% saving of the Bjøntegaard delta bit-rate (BD-BR) on the standard test set, significantly outperforming the standard in-loop filter in HEVC and other state-of-the-art approaches. Tianyi Li 0004, Mai Xu, Ce Zhu, Zulin Wang, Zhenyu Guan 0002 |
IEEE Trans. Image Process. | 5 |
| 2019 | Fast H.264 to HEVC Transcoding: A Deep Learning MethodabstractWith the development of video coding technology, high-efficiency video coding (HEVC) has become a promising alternative, compared with the previous coding standards, for example, H.264. In general, H.264 to HEVC transcoding can be accomplished by fully H.264 decoding and fully HEVC encoding, which suffers from considerable time consumption on the brute-force search of the HEVC coding tree unit (CTU) partition for rate-distortion optimization (RDO). In this paper, we propose a deep learning method to predict the HEVC CTU partition, instead of the brute-force RDO search, for H.264 to HEVC transcoding. First, we build a large-scale H.264 to HEVC transcoding database. Second, we investigate the correlation between the HEVC CTU partition and H.264 features, and analyze both temporal and spatial-temporal similarities of the CTU partition across video frames. Third, we propose a deep learning architecture of a hierarchical long short-term memory (H-LSTM) network to predict the CTU partition of HEVC. Then, the brute-force RDO search of the CTU partition is replaced by the H-LSTM prediction such that the computational time can be significantly reduced for fast H.264 to HEVC transcoding. Finally, the experimental results verify that the proposed H-LSTM method can achieve a better tradeoff between coding efficiency and complexity, compared to the state-of-the-art H.264 to HEVC transcoding methods. Jingyao Xu 0002, Mai Xu, Yanan Wei, Zulin Wang, Zhenyu Guan 0002 |
IEEE Trans. Multim. | 4 |
| 2018 | Multi-Frame Quality Enhancement for Compressed VideoabstractThe past few years have witnessed great success in applying deep learning to enhance the quality of compressed image/video. The existing approaches mainly focus on enhancing the quality of a single frame, ignoring the similarity between consecutive frames. In this paper, we investigate that heavy quality fluctuation exists across compressed video frames, and thus low quality frames can be enhanced using the neighboring high quality frames, seen as Multi-Frame Quality Enhancement (MFQE). Accordingly, this paper proposes an MFQE approach for compressed video, as a first attempt in this direction. In our approach, we firstly develop a Support Vector Machine (SVM) based detector to locate Peak Quality Frames (PQFs) in compressed video. Then, a novel Multi-Frame Convolutional Neural Network (MF-CNN) is designed to enhance the quality of compressed video, in which the non-PQF and its nearest two PQFs are as the input. The MF-CNN compensates motion between the non-PQF and PQFs through the Motion Compensation subnet (MC-subnet). Subsequently, the Quality Enhancement subnet (QE-subnet) reduces compression artifacts of the non-PQF with the help of its nearest PQFs. Finally, the experiments validate the effectiveness and generality of our MFQE approach in advancing the state-of-the-art quality enhancement of compressed video. The code of our MFQE approach is available at https://github.com/ryangBUAA/MFQE.git. Mai Xu, Zulin Wang, Tianyi Li 0004 |
CVPR | 3 |
| 2018 | DeepVS: A Deep Learning Based Video Saliency Prediction Approach
Lai Jiang 0004, Mai Xu, Minglang Qiao, Zulin Wang |
ECCV (14) | 5 |
| 2018 | Graftage Coding for Distributed Storage SystemsabstractRecently, several remarkable works [1]-[5] constructed regenerating codes to offer intermediate tradeoffs between storage and bandwidth. Unlike regenerating codes, this paper proposes to graft codes together to provide various such intermediate tradeoffs. It shows that the linear relations in the generator matrices of grafting codes can be transferred to those of grafted codes without any loss of reconstruction capability. A construction based on minimum storage regenerating codes shows that the resulted graftage codes may provide better tradeoffs than space-sharing and approach cut-set bounds, with the cost of fixed access of helper nodes. Qin Huang 0002, Jiayi Rui, Liyuan Song, Zulin Wang |
GLOBECOM | 4 |
| 2018 | Construction of Multiple-Burst-Correction Codes in Transform DomainabstractThis paper proposes to construct a class of multiple-burst-correction quasi-cyclic (QC) codes via matrix transformations. Due to the diagonal structure of the transformed parity-check matrix of a QC code, multiple bursts can be corrected in the transform domain. By well designing the diagonal submatrices, the constructed QC codes are able to achieve optimal or asymptotically optimal multiple-burst-correction capability. In particular, a subclass of our constructed QC codes are QC low-density parity-check (QC-LDPC) codes which also perform very well over random channels. Simulation results show that our QC-LDPC codes outperform the existing QC-LDPC codes over burst channels. Qin Huang 0002, Liyuan Song, Zulin Wang |
ISIT | 3 |
| 2018 | Encoding of Non-binary Quasi-cyclic Codes by Lin-Chung-Han TransformabstractRecently, Lin, Chung, and Han presented an efficient additive fast Fourier transform based on a novel polynomial basis. This paper explains clearly and proves the convolution theorem of Lin-Chung-Han (LCH) transform. It demonstrates that the corresponding convolutions of LCH transform can be equivalent to cyclic convolutions by preprocessed modulo and polynomial bases conversion. As a result, this paper proposes a fast algorithm for the multiplication of a vector and a circulant matrix. It shows that the algorithm performs very efficient for the encoding of nonbinary quasi-cyclic codes. For an (ne, ke) quasi-cyclic code with circulant size of e, the encoding algorithm needs approximately 1/4n(e+1)log22(e+1)+k(n-k)(e+1) multiplications and additions, which is much less than the number (n-k)ke2of traditional encoding algorithm. Runzhou Li, Qin Huang 0002, Zulin Wang |
ITW | 3 |
| 2018 | Bridge the Gap Between VQA and Human Behavior on Omnidirectional Video: A Large-Scale Dataset and a Deep Learning ModelabstractOmnidirectional video enables spherical stimuli with the $360 \times 180^ \circ$ viewing range. Meanwhile, only the viewport region of omnidirectional video can be seen by the observer through head movement (HM), and an even smaller region within the viewport can be clearly perceived through eye movement (EM). Thus, the subjective quality of omnidirectional video may be correlated with HM and EM of human behavior. To fill in the gap between subjective quality and human behavior, this paper proposes a large-scale visual quality assessment (VQA) dataset of omnidirectional video, called VQA-OV, which collects 60 reference sequences and 540 impaired sequences. Our VQA-OV dataset provides not only the subjective quality scores of sequences but also the HM and EM data of subjects. By mining our dataset, we find that the subjective quality of omnidirectional video is indeed related to HM and EM. Hence, we develop a deep learning model, which embeds HM and EM, for objective VQA on omnidirectional video. Experimental results show that our model significantly improves the state-of-the-art performance of VQA on omnidirectional video. Chen Li 0049, Mai Xu, Xinzhe Du, Zulin Wang |
ACM Multimedia | 4 |
| 2018 | Rate control schemes for panoramic video coding
Yufan Liu 0001, Li Yang 0014, Mai Xu, Zulin Wang |
J. Vis. Commun. Image Represent. | 4 |
| 2018 | A Repair-Efficient Coding for Distributed Storage Systems Under Piggybacking FrameworkabstractPiggybacking is an efficient framework to reduce the repair bandwidth of distributed storage systems, especially, when the systems meet the settings-maximum distance separable (MDS), high code rate, and a small number of substripes. Through an analysis on the repair ratio of a piggybacking construction, this paper reveals that the proportion ppof piggybacking protected stripes is the key to significantly decrease the repair bandwidth. Based on this analysis, this paper proposes a repair efficient coding under piggybacking framework (REPB) by considering various piggybacking protected stripes and MDS only protected stripes. The repair ratio of REPB tends to 0 and is close to theoretical cut-set bound, as the proportion ppis no longer fixed at 1/2. Furthermore, the proposed REPB codes enjoy low computational complexity in repair operation. Shuai Yuan 0017, Qin Huang 0002, Zulin Wang |
IEEE Trans. Commun. | 3 |
| 2018 | Reducing Complexity of HEVC: A Deep Learning ApproachabstractHigh Efficiency Video Coding (HEVC) significantly reduces bit-rates over the preceding H.264 standard but at the expense of extremely high encoding complexity. In HEVC, the quad-tree partition of coding unit (CU) consumes a large proportion of the HEVC encoding complexity, due to the brute-force search for rate-distortion optimization (RDO). Therefore, this paper proposes a deep learning approach to predict the CU partition for reducing the HEVC complexity at both intra-and inter-modes, which is based on convolutional neural network (CNN) and long-and short-term memory (LSTM) network. First, we establish a large-scale database including substantial CU partition data for HEVC intra-and inter-modes. This enables deep learning on the CU partition. Second, we represent the CU partition of an entire coding tree unit (CTU) in the form of a hierarchical CU partition map (HCPM). Then, we propose an early-terminated hierarchical CNN (ETH-CNN) for learning to predict the HCPM. Consequently, the encoding complexity of intra-mode HEVC can be drastically reduced by replacing the brute-force search with ETH-CNN to decide the CU partition. Third, an early-terminated hierarchical LSTM (ETH-LSTM) is proposed to learn the temporal correlation of the CU partition. Then, we combine ETH-LSTM and ETH-CNN to predict the CU partition for reducing the HEVC complexity at inter-mode. Finally, experimental results show that our approach outperforms other state-of-the-art approaches in reducing the HEVC complexity at both intra-and inter-modes. Mai Xu, Tianyi Li 0004, Zulin Wang, Xin Deng 0002, Zhenyu Guan 0002 |
IEEE Trans. Image Process. | 3 |
| 2018 | Closed-Form Optimization on Saliency-Guided Image Compression for HEVC-MSPabstractHigh efficiency video coding (HEVC) is the latest video coding standard, and it has the best performance among all the existing standards. HEVC main still picture profile (HEVC-MSP) also achieves top performance in image compr-ession. In this paper, we propose a closed-form bit allocation approach to optimize the saliency-guided PSNR (viewed as perceptual distortion) such that the coding efficiency of HEVC-based image compression can be significantly improved from a subjective perspective. Specifically, a bit allocation formulation is established to minimize perceptual distortion with a constraint on bit-rates. Then, this formulation is solved using the proposed recursive Taylor expansion method with a closed-form solution. On the basis of our solution, a bit allocation and re-allocation process is developed in our approach to minimize perceptual distortion, meanwhile accurately controlling bit-rates. In addition, we provide both theoretical and numerical analyses of the computational complexity, verifying the little extra time cost of our approach. The experimental results demonstrate the superior performance of our approach over the state-of-the-art HEVC-MSP, and the BD-rate savings are approximately 40% and 24% for face and generic images, respectively. Shengxi Li, Mai Xu, Yun Ren, Zulin Wang |
IEEE Trans. Multim. | 4 |
| 2018 | Saliency Detection in Face Videos: A Data-Driven ApproachabstractRecently, videoconferencing has been popular in multimedia systems, such as FaceTime and Skype. In videoconferencing, almost every frame contains a human face. Therefore, it is important to predict human visual attention on face videos by saliency detection, as saliency may be used as a guide to the region of interest for the content-based applications of face videos. In this paper, we propose a data-driven approach for saliency detection in face videos. From the data-driven perspective, we first establish an eye-tracking database that contains fixations of 76 face videos viewed by 40 subjects. Upon the analysis of our database, we find that visual attention is significantly attracted by faces in videos. More important, the attention distribution within face regions varies with regard to mouth movement. Since previous works have investigated that it is efficient to model face saliency in still images using a Gaussian mixture model (GMM), the variation of visual attention in videos can be modeled by dynamic GMM (DGMM). Accordingly, we propose adopting the particle filter (PF) in modeling DGMM for saliency detection of face videos, which is called PF-DGMM. Finally, the experimental results show that our PF-DGMM approach significantly outperforms other state-of-the-art approaches in saliency detection of face videos. Mai Xu, Yun Ren, Zulin Wang, Jingxian Liu, Xiaoming Tao 0001 |
IEEE Trans. Multim. | 3 |
| 2018 | Efficient Quantization with Linear Index Coding for Deep-Space ImagesabstractDue to inevitable propagation delay involved in deep‐space communication systems, very high cost is associated with the retransmission of erroneous segments. Quantization with linear index coding (QLIC) scheme is known to provide compression along with robust transmission of deep‐space images, and thus the likelihood of retransmissions is significantly reduced. This paper aims to improve its spectral efficiency as well as robustness. First, multiple quantization refinement levels per transmitted source block of QLIC are proposed to increase spectral efficiency. Then, iterative multipass decoding is introduced to jointly decode the subsource symbol‐planes. It achieves better PSNR of the reconstructed image as compared to the baseline one‐pass decoding approach of QLIC. Rehan Mahmood, Zulin Wang, Qin Huang 0007 |
Wirel. Commun. Mob. Comput. | 2 |
| 2018 | A CCM-Based OFDM System with Low PAPR for Sparse SourceabstractOrthogonal frequency division multiplexing (OFDM) usually suffers high peak‐to‐average power ratio (PAPR). As shown in this paper, PAPR becomes even severe for sparse source due to many identical nonzero frequency OFDM symbols. Thus, this paper introduces compressive coded modulation (CCM) in order to restrain PAPR by reducing identical nonzero frequency symbols for sparse source. As a result, the proposed CCM‐based OFDM system, together with iterative clipping and filtering, can efficiently restrain the high PAPR for sparse source. Simulation results show that it outperforms about 4 dB over the traditional OFDM system when source sparsity is 0.1. Qinbiao Yang, Zulin Wang, Qin Huang 0007 |
Wirel. Commun. Mob. Comput. | 2 |
| 2017 | A novel rate control scheme for panoramic video codingabstractThe popularity of multi-view panoramic videos has been considerably increased for producing Virtual Reality (VR) content, due to its immersive visual experience. We argue in this paper that PSNR is less effective in assessing visual quality of compressed panoramic videos than Sphere-based PSNR (S-PNSR), in which sphere-to-plain mapping of panoramic videos is considered. Thus, the conventional rate control (R-C) schemes of 2-Dimensional (2D) video coding, which optimize on PSNR, are not suitable for panoramic video coding. To optimize S-PSNR, we propose in this paper a novel RC scheme for panoramic video coding. Specifically, we develop an S-PSNR optimization formulation with constraint on bit-rate. Then, a solution is provided to the developed formulation, such that bits can be allocated to each coding block for achieving optimal S-PSNR in panoramic video coding. Finally, the experiment results validate the effectiveness of the proposed RC scheme in improving S-PSNR of panoramic video coding. Yufan Liu 0001, Mai Xu, Chen Li 0049, Shengxi Li, Zulin Wang |
ICME | 5 |
| 2017 | Decoder-side HEVC quality enhancement with scalable convolutional neural networkabstractThe latest High Efficiency Video Coding (HEVC) has been increasingly used to generate video streams over Internet. However, the decoded HEVC video streams may incur severe quality degradation, especially at low bit-rates. Thus, it is necessary to enhance visual quality of HEVC videos at the decoder side. To this end, we propose in this paper a Decoder-side Scalable Convolutional Neural Network (DS-CNN) approach to achieve quality enhancement for HEVC, which does not require any modification of the encoder. In particular, our DS-CNN approach learns a model of Convo-lutional Neural Network (CNN) to reduce distortion of both I and B/P frames in HEVC. It is different from the existing CNN-based quality enhancement approaches, which only handle intra coding distortion, thus not suitable for B/P frames. Furthermore, a scalable structure is included in our DS-CNN, suchthat the computational complexity of our DS-CNN approach is adjustable to the changing computational resources. Finally, the experimental results show the effectiveness of our DS-CNN approach in enhancing quality for both I and B/P frames of HEVC. Mai Xu, Zulin Wang |
ICME | 3 |
| 2017 | An LSTM method for predicting CU splitting in H.264 to HEVC transcodingabstractFor H.264 to high efficiency video coding (HEVC) transcoding, this paper proposes a hierarchical Long Short-Term Memory (LSTM) method to predict coding unit (CU) splitting. Specifically, we first analyze the correlation between CU splitting patterns and H.264 features. Upon our analysis, we further propose a hierarchical LSTM architecture for predicting CU splitting of HEVC, with regard to the explored H.264 features. The features of H.264, including residual, macroblock (MB) partition and bit allocation, are employed as the input to our LSTM method. Experimental results demonstrate that the proposed method outperforms the state-of-the-art H.264 to HEVC transcoding methods, in terms of both complexity reduction and PSNR performance. Yanan Wei, Zulin Wang, Mai Xu, Shuhao Qiao |
VCIP | 2 |
| 2017 | Multi-Pass Decoding for the Robust Transmission of Deep-Space ImagesabstractLinear index coding is generally more robust against channel variations as compared to the fixed-to-variable length coding. This paper proposes a novel multi-pass decoding approach to decode linear index coded images. In contrast to the typical one-pass decoding, the proposed scheme harnesses the information recovered in the first decoding pass with the source statistics and utilize it in the subsequent decoding passes. It is demonstrated that the linear index coded image transmitted from deep-space shows an improvement of up to 3.2 dB in terms of reconstructed peak signal to noise ratio by using multi-pass decoding. Rehan Mahmood, Zulin Wang, Qin Huang 0002 |
VTC Spring | 2 |
| 2017 | Road Structure Refined CNN for Road Extraction in Aerial ImageabstractIn this letter, we propose a road structure refined convolutional neural network (RSRCNN) approach for road extraction in aerial images. In order to obtain structured output of road extraction, both deconvolutional and fusion layers are designed in the architecture of RSRCNN. For training RSRCNN, a new loss function is proposed to incorporate the geometric information of road structure in cross-entropy loss, thus called road-structure-based loss function. Experimental results demonstrate that the trained RSRCNN model is able to advance the state-of-the-art road extraction for aerial images, in terms of precision, recall, F-score, and accuracy. Yanan Wei, Zulin Wang, Mai Xu |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2017 | Set Message-Passing Decoding Algorithms for Regular Non-Binary LDPC CodesabstractIn the check node (CN) update of non-binary message-passing algorithms, each element of reliability vectors takes the same computational complexity. However, our analysis indicates that various elements in the same vector have various correct probabilities, thus have different contributions to error performance. In order to match computational complexity with correct probability, all elements in a vector are partitioned into different sets. For the extended min-sum (EMS) decoding, various strategies are applied for sets according to their correct probability. For the trellis-based EMS decoding, it is interesting that set partition only involves fixed paths, thus it does not need to search over the whole trellis of a CN. Complexity analysis and simulation results show that the proposed algorithms efficiently decode non-binary low-density parity-check codes, including ultra-sparse ones. Qin Huang 0002, Liyuan Song, Zulin Wang |
IEEE Trans. Commun. | 3 |
| 2017 | Symbol Flipping Decoding Algorithms Based on Prediction for Non-Binary LDPC CodesabstractThis paper constructs an objective function for symbol flipping decoding algorithms, considering not only soft reliability, but also hard reliability. The maximization of this objective function indicates that the flipping metric should involve both the information before and after symbol flipping, while the existing algorithms consider the information before symbol flipping. Theoretical analysis shows that such prediction mechanism, together with hard reliability, can significantly improve the error performance of symbol flipping algorithms. Simulation results show that the proposed algorithms provide effective tradeoff between error performance and complexity for decoding non-binary LDPC codes. Qin Huang 0002, Zulin Wang |
IEEE Trans. Commun. | 3 |
| 2017 | Corrections on "Symbol Flipping Decoding Algorithms Based on Prediction for Non-Binary LDPC Codes"abstractDue to a production error, an equation in the above paper[1]appeared incorrectly. Below is the correct version. Qin Huang 0002, Zulin Wang |
IEEE Trans. Commun. | 3 |
| 2017 | Optimal Bit Allocation for CTU Level Rate Control in HEVCabstractFor High Efficiency Video Coding (HEVC), the R–$\lambda $scheme is the latest rate control (RC) scheme, which investigates the relationships among allocated bits, the slope of rate-distortion (R-D) curve$\lambda $, and quantization parameter. However, we argue that bit allocation in the existing R–$\lambda $scheme is not optimal. In this paper, we therefore propose an optimal bit allocation (OBA) scheme for coding tree unit level RC in HEVC. Specifically, to achieve the OBA, we first develop an optimization formulation with a novel R-D estimation, instead of the existing R–$\lambda $estimation. Unfortunately, it is intractable to obtain a closed-form solution to the optimization formulation. We thus propose a recursive Taylor expansion (RTE) method to iteratively solve the formulation. As a result, an approximate closed-form solution can be obtained, thus achieving OBA and bit reallocation. Both theoretical and numerical analyses show the fast convergence speed and little computational time of the proposed RTE method. Therefore, our OBA scheme can be achieved at little encoding complexity cost. Finally, the experimental results validate the effectiveness of our scheme in three aspects: R-D performance, RC accuracy, and robustness over dynamic scene changes. Shengxi Li, Mai Xu, Zulin Wang, Xiaoyan Sun 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2017 | Learning to Detect Video Saliency With HEVC FeaturesabstractSaliency detection has been widely studied to predict human fixations, with various applications in computer vision and image processing. For saliency detection, we argue in this paper that the state-of-the-art High Efficiency Video Coding (HEVC) standard can be used to generate the useful features in compressed domain. Therefore, this paper proposes to learn the video saliency model, with regard to HEVC features. First, we establish an eye tracking database for video saliency detection, which can be downloaded from https://github.com/remega/video_database. Through the statistical analysis on our eye tracking database, we find out that human fixations tend to fall into the regions with large-valued HEVC features on splitting depth, bit allocation, and motion vector (MV). In addition, three observations are obtained with the further analysis on our eye tracking database. Accordingly, several features in HEVC domain are proposed on the basis of splitting depth, bit allocation, and MV. Next, a kind of support vector machine is learned to integrate those HEVC features together, for video saliency detection. Since almost all video data are stored in the compressed form, our method is able to avoid both the computational cost on decoding and the storage cost on raw data. More importantly, experimental results show that the proposed method is superior to other state-of-the-art saliency detection methods, either in compressed or uncompressed domain. Mai Xu, Lai Jiang 0004, Xiaoyan Sun 0001, Zhaoting Ye, Zulin Wang |
IEEE Trans. Image Process. | 5 |
| 2016 | Optimizing Subjective Quality in HEVC-MSP: An Approximate Closed-form Image Compression ApproachabstractHEVC, as the latest video coding standard, achieves top performance on image compression. On the basis of this, we propose a novel approach to optimize subjective quality for HEVC-based image compression. Specifically, a bit allocation formulation is established to optimize subjective quality with constraint on bit-rates. Then, we propose a recursive Taylor expansion method to quickly solve such a formulation with an approximate closed-form solution. The experimental results show the superior performance of our approach, with ~40% BD-rate saving over the state-of-the-art HEVC-MSP for face image compression. Shengxi Li, Mai Xu, Yun Ren, Chengzhang Ma, Zulin Wang |
DCC | 5 |
| 2016 | Subjective-quality-optimized complexity control for HEVC decodingabstractThe latest High Efficiency Video Coding (HEVC) standard significantly improves coding efficiency over H.264/AVC, at the cost of heavy encoding and decoding complexity. For reducing HEVC decoding complexity to a target, we propose in this paper a Subjective-Quality-Optimized Complexity Control (SQOCC) approach, which optimizes subjective quality loss caused by the decoding complexity reduction. First, a saliency detection method in HEVC domain is developed as the preliminary of subjective quality metric. Based on detected saliency, we establish a formulation to minimize subjective quality loss at the constraint of specific decoding complexity reduction, via disabling the deblocking filters of some Largest Coding Units (LCUs). Next, we utilize least square fitting to model functions in our formulation. We then provide a solution to our formulation, achieving subjective-quality-optimized complexity control for HEVC decoding. Finally, the experimental results show the effectiveness of our SQOC-C approach in terms of both control accuracy and subjective quality. Mai Xu, Lai Jiang 0004, Zulin Wang |
ICME | 4 |
| 2016 | Set min-sum decoding algorithm for non-binary LDPC codesabstractThis paper reduces the complexity of decoding non-binary low-density parity-check (LDPC) codes by set partition. In the check node update, the input vectors are partitioned into several sets such that different elements in the virtual matrix enjoy various computational strategies. As a result, the proposed algorithm achieves high computational efficiency by setting strategies according to the correct probability of these elements. Simulation results indicate that it significantly decreases the complexity of check node update with negligible performance loss. Liyuan Song, Qin Huang 0002, Zulin Wang |
ISIT | 3 |
| 2016 | Predicting the memorability of natural-scene imagesabstractRecent work has shown that image memorability, in general, can be reliably predicted using some state-of-the-art features. However, all existing methods are not effective in predicting memorability of natural-scene images, far from human. In this paper, we propose a novel method to improve the effectiveness of memorability prediction for natural-scene images. Specifically, we argue that some of HSV colors have either positive or negative impact on memorability of natural-scene images in our Natural-Scene Image Memorability (NSIM) dataset. Then, we develop an HSV-based feature for memorability prediction. Finally, the HSV-based feature is combined with other efficient state-of-the-art features in our approach to predict memorability on natural-scene images. Experimental results validate the effectiveness of our method. Jiaxin Lu 0003, Mai Xu, Zulin Wang |
VCIP | 3 |
| 2016 | A novel double-layer sparse representation approach for unsupervised dictionary learning
Mai Xu, Zulin Wang |
Comput. Vis. Image Underst. | 2 |
| 2016 | Bottom-up saliency detection with sparse representation of learnt texture atoms
Mai Xu, Lai Jiang 0004, Zhaoting Ye, Zulin Wang |
Pattern Recognit. | 4 |
| 2016 | Trimming Soft-Input Soft-Output Viterbi AlgorithmsabstractIn the soft-input soft-output Viterbi algorithm (SOVA), the log-likelihood ratio (LLR) of each bit is determined by the minimum metric difference between the ML path and its competitive paths. This paper proposes to trim large metric differences in order to reduce the complexity of SOVA. By trimming the metric differences, only a small number of backtracking operations are carried out, while many LLRs may be omitted as the result of the lack of metric differences. By revealing the relationship among neighboring LLRs, the omitted LLRs are estimated from its neighoring LLRs as well as intrinsic information. The extrinsic information transfer chart analysis demonstrates that the proposed algorithm has similar convergence behavior as the Log-MAP algorithm, if the trimming factor M is moderate. Other analyses verify that our approach provides good LLR quality with only at most 1/M backtracking operations of SOVA. Simulation results show that it outperforms SOVA and performs as well as its variants and the Log-MAP algorithm. Qin Huang 0002, Qiang Xiao 0001, Li Quan 0002, Zulin Wang, Shafei Wang |
IEEE Trans. Commun. | 4 |
| 2016 | Two Enhanced Reliability-Based Decoding Algorithms for Nonbinary LDPC CodesabstractThe weighted bit-reliability-based (wBRB) algorithm for nonbinary LDPC codes suffers certain loss of symbol-reliability. Thus, this paper enhances its soft-decision version by passing multiple symbol-reliability instead of bit-reliability. Furthermore, it demonstrates that plurality robustly indicates symbol-reliability of extrinsic information-sums. Thus, this paper enhances the hard-decision version by introducing symbol-reliability from plurality. Analysis results show that these two enhanced decoding algorithms significantly outperform the wBRB algorithm with reasonable overhead. Liyuan Song, Qin Huang 0002, Zulin Wang, Mu Zhang 0002, Shafei Wang |
IEEE Trans. Commun. | 3 |
| 2016 | Time-Invariant Quasi-Cyclic Spatially Coupled LDPC Codes Based on PackingsabstractThis paper presents two packings derived from balanced incomplete block designs to construct quasi-cyclic spatially coupled LDPC convolutional codes (SC-LDPC-CCs). The construction gives time-invariant codes, since the sub-blocks corresponding to each time instant of the parity-check matrix are identical. Moreover, it provides flexible design rates and constraint lengths. Simulation results show that the proposed packing-based SC-LDPC-CCs outperform the existing time-invariant codes and perform closely to the time-varying protograph-based codes. Mu Zhang 0002, Zulin Wang, Qin Huang 0002, Shafei Wang |
IEEE Trans. Commun. | 2 |
| 2016 | Subjective-Driven Complexity Control Approach for HEVCabstractThe latest High Efficiency Video Coding (HEVC) standard significantly increases the encoding complexity for improving its coding efficiency, compared with the preceding H.264/Advanced Video Coding (AVC) standard. In this paper, we present a novel subjective-driven complexity control (SCC) approach to reduce and control the encoding complexity of HEVC. Through reasonably adjusting the maximum depth of each largest coding unit (LCU), the encoding complexity can be reduced to a target level with minimal visual distortion. Specifically, the maximum depths of different LCUs can be varied through solving the proposed optimization formulation of complexity control, based on two explored relationships: 1) the relationship between the maximum depth and encoding complexity and 2) the relationship between the maximum depth and visual distortion. Besides, the subjective visual quality is favored with a novel subjective-driven constraint imposed in the formulation, on the basis of a visual attention model. Finally, the experimental results show that our approach can achieve a wide range of encoding complexity control (as low as 20%) for HEVC, with the smallest complexity bias being 0.2%. Meanwhile, our SCC approach outperforms other two state-of-the-art complexity control approaches, in terms of both control accuracy and visual quality. Xin Deng 0002, Mai Xu, Lai Jiang 0004, Xiaoyan Sun 0001, Zulin Wang |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2015 | Low error-floor majority-logic decoding based algorithm for non-binary LDPC codesabstractThe traditional majority-logic decoding (MLgD) based algorithms suffer error-floors for decoding non-binary LDPC codes with small column weights. This paper presents a bit-reliability based MLgD (BRB-MLgD) algorithm with low error-floors for non-binary LDPC codes. The proposed algorithm is carried out based on the binary representations of non-binary symbols. The reliability update along each edge of the Tanner graph of a non-binary LDPC code is in terms of bits rather than symbols. Thus, its computational complexity and memory consumption are less than those of the existing MLgD based algorithms. Simulation results indicate that the proposed algorithm can significantly reduce error-floors with small performance degradation in the waterfall region. Liyuan Song, Mu Zhang 0002, Qin Huang 0002, Zulin Wang |
ICC | 4 |
| 2015 | Learning to Predict Saliency on Face ImagesabstractThis paper proposes a novel method, which learns to detect saliency of face images. To be more specific, we obtain a database of eye tracking over extensive face images, via conducting an eye tracking experiment. With analysis on eye tracking database, we verify that the fixations tend to cluster around facial features, when viewing images with large faces. For modeling attention on faces and facial features, the proposed method learns the Gaussian mixture model (GMM) distribution from the fixations of eye tracking data as the top-down features for saliency detection of face images. Then, in our method, the top-down features (i.e., face and facial features) upon the the learnt GMM are linearly combined with the conventional bottom-up features (i.e., color, intensity, and orientation), for saliency detection. In the linear combination, we argue that the weights corresponding to top-down feature channels depend on the face size in images, and the relationship between the weights and face size is thus investigated via learning from the training eye tracking data. Finally, experimental results show that our learning-based method is able to advance state-of-the-art saliency prediction for face images. The corresponding database and code are available online: www.ee.buaa.edu.cn/xumfiles/saliency_detection.html. Mai Xu, Yun Ren, Zulin Wang |
ICCV | 3 |
| 2015 | A novel method on optimal bit allocation at LCU level for rate control in HEVCabstractIn this paper, we propose a new method, namely recursive Taylor expansion (RTE) method, for optimally allocating bits to each LCU in the R-λ rate control scheme for HEVC. Specifically, we first set up an optimization formulation on optimal bit allocation. Unfortunately, it is intractable to achieve a closed-form solution for this formulation. We therefore propose a RTE solution to iteratively solve the formulation with a fast convergence speed. Then, an approximate closed-form solution can be obtained. This way, the optimal bit allocation can be achieved at little encoding complexity cost. Finally, the experimental results validate the effectiveness of our method in three aspects: compressed distortion, bit-rate control error, and bit fluctuation. Shengxi Li, Mai Xu, Zulin Wang |
ICME | 3 |
| 2015 | Learning Gaussian mixture model for saliency detection on face imagesabstractThe previous work has demonstrated that integrating top-down features in bottom-up saliency methods can improve the saliency prediction accuracy. Therefore, for face images, this paper proposes a saliency detection method based on Gaussian mixture model (GMM), which learns the distribution of saliency over face regions as the top-down feature. Specifically, we verify that fixations tend to cluster around facial features, when viewing images with large faces. Thus, the GMM is learnt from fixations of eye tracking data, for establishing the distribution of saliency in faces. Then, in our method, the top-down feature upon the the learnt GMM is combined with the conventional bottom-up features (i.e., color, intensity, and orientation), for saliency detection. Finally, experimental results validate that our method is capable of improving the accuracy of saliency prediction for face images. Yun Ren, Mai Xu, Ruihan Pan, Zulin Wang |
ICME | 4 |
| 2015 | Weight-based R-λ rate control for perceptual HEVC coding on conversational videos
Shengxi Li, Mai Xu, Xin Deng 0002, Zulin Wang |
Signal Process. Image Commun. | 4 |
| 2015 | A Robust Algorithm for Joint Sparse Recovery in Presence of Impulsive NoiseabstractThis letter presents a robust solution for joint sparse recovery (JSR) under impulsive noise. The unknown measurement noise is endowed with the Student-t distribution, then a novel Bayesian probabilistic model is proposed to describe the JSR problem. To effectively recover the joint row sparse signal, variational Bayes (VB) method is introduced for Bayesian theory based JSR algorithms such that it overcomes the intractable integrations inherent. Simulation results verify that the proposed algorithm significantly outperforms the existing algorithms under impulsive noise. Jiadong Shang, Zulin Wang, Qin Huang 0002 |
IEEE Signal Process. Lett. | 2 |
| 2014 | A novel weight-based URQ scheme for perceptual video coding of conversational video in HEVCabstractIn this paper, we propose a novel weight-based unified rate-quantization (URQ) scheme for rate control in state-of-the-art HEVC standard, to improve its perceived visual quality for conversational videos. In conventional rate control of HEVC, a pixel-wise URQ scheme is proposed by introducing the concept of bit per pixel (bpp). This scheme is able to assign different amounts of bits to the blocks with various sizes, thus well suitable for flexible picture partition of HEVC. However, bpp does not reflect the visual importance of each pixel. Therefore, we propose a novel weight-based URQ scheme to take into account the visual importance for rate control in HEVC. In combination with the weight map acquired from a novel hierarchical perceptual model of face, such a scheme is capable of allocating more bits to the face and much more bits to the facial features, by using bit per weight (bpw) instead of bpp. As a result, the visual quality of face, especially facial features, can be improved such that perceptual video coding is achieved for HEVC. Finally, the experimental results validate such improvement. Shengxi Li, Mai Xu, Xin Deng 0002, Zulin Wang |
ICME | 4 |
| 2014 | Complexity control of HEVC based on region-of-interest attention modelabstractIn this paper, we present a novel complexity control method of HEVC to adjust its encoding complexity. First, a region-of-interest (ROI) attention model is established, which defines different weights for various regions according to their importance. Then, the complexity control algorithm is proposed with a distortion-complexity optimization model, to determine the maximum depth of the largest coding units (LCUs) according to their weights. We can reduce the encoding complexity to a given target level at the cost of little distortion loss. Finally, the experimental results show that the encoding complexity can drop to a pre-defined target complexity as low as 20% with bias less than 7%. Meanwhile, our method is verified to preserve the quality of ROI better than another state-of-the-art approach. Xin Deng 0002, Mai Xu, Shengxi Li, Zulin Wang |
VCIP | 4 |
| 2014 | A novel objective quality assessment method for perceptual video coding in conversational scenariosabstractRecently, numerous perceptual video coding approaches have been proposed to use face as ROI regions, for improving perceived visual quality of compressed conversational videos. However, there exists no objective metric, specialized for efficiently evaluating the perceived visual quality of compressed conversational videos. This paper thus proposes an efficient objective quality assessment method, namely Gaussian mixture model based PSNR (GMM-PSNR), for conversational videos. First, eye tracking experiments, together with a face extraction technique, were carried out to identify importance of the regions of background, face, and facial features, through eye fixation points. Next, assuming that the distribution of some eye fixation points obeys Gaussian mixture model, an importance weight map is generated by introducing a new term, eye fixation points/pixel(efp/p). Finally, GMM-PSNR is computed by assigning different penalties to the distortion of each pixel in a video frame, according to the generated weight map. The experimental results show the effectiveness of our GMM-PSNR by investigating its correlation with subjective quality on several test video sequences. Mai Xu, Jingze Zhang, Zulin Wang |
VCIP | 4 |
| 2014 | Unsupervised dictionary learning with double-layer sparse representationabstractThis paper presents a novel double-layer sparse representation (DLSR) approach for unsupervised dictionary learning. In supervised/unsupervised discriminative dictionary learning, classical approaches usually develop a discriminative term for learning multiple sub-dictionaries, each of which corresponds to one-class training image patches. However, in unsupervised scenario, some of the training patches for learning sub-dictionaries of each class are related to more than one class. Thus, we propose a DLSR formulation, in this paper, to impose the first-layer sparsity on the coefficients and the second-layer sparsity on the classes for each training patch, embedding both the reconstructive (via the first-layer) and discriminative (via the second-layer) abilities in the dictionary. To address the proposed DLSR formulation, a simple yet effective algorithm, called DLSR-OMP, is developed in light of the conventional OMP. Finally, the experimental results show the effectiveness of our approach in the reconstruction task of image denoising and the clustering task of texture segmentation. Mai Xu, Zulin Wang |
WACV | 2 |
| 2014 | Density optimisation of generator matrices of quasi-cyclic low-density parity-check codes and their rank analysisabstractThe efficient encoding of quasi‐cyclic (QC) low‐density parity‐check (LDPC) codes is based on generator matrices in systematic‐circulant (SC) form. The cost of the encoders of QC‐LDPC codes mainly depends on the number of non‐zero entries in the SC generator matrices. This study introduces a novel construction of SC generator matrices based on matrix transformations via Galois Fourier transform. By revealing the structure of SC generator matrices in the transform domain, an algorithm is proposed to reduce the density of the generator matrices of QC‐LDPC codes. Furthermore, a tight upper bound on ranks of QC matrices is derived. Based on the bound, rank distributions of parity‐check matrices and generator matrices in the transform domain illustrate the efficiency of the proposed algorithm. Simulation results show that the density of their SC generator matrices can be significantly decreased with moderate computational complexity. Mu Zhang 0002, Qin Huang 0002, Zulin Wang, Shuai Yuan 0017, Zhe Liu 0018 |
IET Commun. | 3 |
| 2014 | Tower of Knowledge for scene interpretation: A survey
Mai Xu, Zulin Wang, Maria Petrou |
Pattern Recognit. Lett. | 2 |
| 2014 | Low-Complexity Encoding of Quasi-Cyclic Codes Based on Galois Fourier TransformabstractThis paper presents two novel low-complexity encoding algorithms for quasi-cyclic (QC) codes based on Galois Fourier transform. The key idea behind them is making use of the block diagonal structure of the transformed generator matrix. The first one, named encoding by Galois Fourier transform, is equivalent to the fast implementations of the traditional encoding by Galois Fourier transform. The second one, named encoding in the transform domain (ETD), requires much less computational complexity for encoding binary QC codes. It skips the first step of the first algorithm and applies post-processing to save a large number of Galois field multiplications. Its application to QC-LDPC codes is also studied in this paper. Particularly, the hardware cost of the ETD for RS-based LDPC codes can be greatly reduced by short linear-feedback shift registers. Qin Huang 0002, Shanbao He, Zixiang Xiong, Zulin Wang |
IEEE Trans. Commun. | 5 |
| 2014 | Bit-Reliability Based Low-Complexity Decoding Algorithms for Non-Binary LDPC CodesabstractThis paper presents bit-reliability based majority-logic decoding (MLgD) algorithms for non-binary LDPC codes. The proposed algorithms pass only one Galois field element and its reliability along each edge of the Tanner graph of a non-binary LDPC code. Since their reliability updates are in terms of bits rather than symbols, they are more efficient than traditional MLgD based decoding algorithms. By weighting the soft reliability of the extrinsic information-sums based on their hard reliability, the proposed algorithms can achieve good error performance for non-binary LDPC codes with various column weights. Moreover, their computational complexity and memory consumption are remarkably reduced compared with existing MLgD based decoding algorithms. As a result, they provide effective tradeoffs between error performance and complexity for decoding of non-binary LDPC codes. Qin Huang 0002, Mu Zhang 0002, Zulin Wang |
IEEE Trans. Commun. | 3 |
| 2013 | Low-complexity encoding of binary quasi-cyclic codes based on Galois Fourier transformabstractThis paper presents a novel low-complexity encoding algorithm for binary quasi-cyclic (QC) codes based on matrix transformation. First, a message vector is encoded into a transformed codeword in the transform domain. Then, the transmitted codeword is obtained from the transformed codeword by the inverse Galois Fourier transform. Moreover, a simple and fast mapping is devised to post-process the transformed codeword such that the transmitted codeword is binary as well. The complexity of our proposed encoding algorithm is less than ek(n-k)log2e+ne(log22e+log2e)+ n/2 elog32e bit operations for binary codes. This complexity is much lower than its traditional complexity 2e2(n - k)k. In the examples of encoding the binary (4095, 2016) and (15500, 10850) QC codes, the complexities are 12.09% and 9.49% of those of traditional encoding, respectively. Qin Huang 0002, Zulin Wang, Zixiang Xiong |
ISIT | 3 |
| 2012 | Low-density arrays of circulant matrices: Rank and row-redundancy, and QC-LDPC codesabstractThis paper is concerned with general analysis on the rank and row-redundancy of an array of circulants whose null space defines a QC-LDPC code. Based on the Fourier transform and the properties of conjugacy classes and Hadamard products of matrices, tight bounds on rank and row-redundancy are derived, which make it possible to consider row-redundancy in constructions of QC-LDPC codes to achieve better performance. Moreover, a new construction of QC-LDPC codes from random partitions of finite fields, which has flexible code dimensions and is abundant in row-redundancy, is presented and analyzed. Qin Huang 0002, Keke Liu, Zulin Wang |
ISIT | 3 |