VLDB 2026 Research / reviewers in the wild / expert
Long Xu 0001
dblp:37/5269-1
· DBLP profile ↗
71ranked-venue papers
17as first author
23since 2021 · last 2026
0000-0002-9286-2876ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 59 · 15 first-author · 20 since 2021Artificial intelligence and machine learning · 5 · 2 since 2021Systems, architecture and hardware · 3 · 1 first-authorHuman-computer interaction and ubiquitous computing · 3 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CartoonCodec: Generative talking face video coding with cartoon-style customizationabstractThe rapid growth of video-based social applications has intensified the demand for efficient transmission and personalized cartoon-style customization. Existing solutions typically apply talking face video and text codecs, followed by cartoon-style control algorithms, which often result in unsatisfactory compression performance and high inference latency. In this paper, we propose CartoonCodec, an efficient generative framework that unifies face video coding and cartoon-style control into an end-to-end process. Our framework encodes facial video sequences into compact motion feature representations, which are transmitted together with compressed text prompts. To integrate control into the coding process while ensuring efficient compression, a text-guided adaptive layer selection mechanism relying on the compact representations is introduced to dynamically select and optimize the most influential layers in the generators. Furthermore, to facilitate stylization over the decoupled spaces, we propose a self-supervised domain stylization training strategy that constructs both multimodal and unimodal data pairs, enabling the use of diverse loss functions. Extensive experiments demonstrate that CartoonCodec outperforms baselines in compression efficiency for video reconstruction and cartoon-style control tasks while maintaining competitive inference efficiency. CartoonCodec provides key insights for advancing face video communication with cartoon-style customization. The project page can be found at https://github.com/xiaonae/CartoonCodec/tree/main . Xihua Sheng, Meng Wang 0017, Long Xu 0001, Shiqi Wang 0001, Sam Kwong |
Neurocomputing | 4 |
| 2026 | When Video Compression Meets Multimodal Large Language Models: A Unified Paradigm for Cross-Modality Video CompressionabstractTraditional video compression methods perform well at high bitrates but struggle to preserve fine-grained semantic information at low bitrates. Recently, with the blossoming of Multimodal Large Language Models (MLLMs), Cross-modal compression techniques offer prospective solutions for improving video compression under low-bitrate conditions. In this paper, we propose a unified Cross-Modality Video Compression (CMVC) framework that integrates multimodal representations and video generative models. The encoder disentangles video into spatial and temporal components, which are mapped to compact cross modal representations using MLLMs. During decoding, different encoding-decoding modes are employed to acquire various video reconstruction qualities, including Text-Text-to-Video (TT2V) for semantic preservation and Image-Text-to-Video (IT2V) for perceptual consistency. Additionally, we elaborate on an efficient frame interpolation model using Low-Rank Adaptation (LoRA) to improve the perceptual quality. Experimental results demon strate that TT2V achieves effective semantic reconstruction, while IT2V ensures competitive perceptual consistency. These findings suggest the potential of leveraging multimodal priors to improve video compression, offering promising future research directions. Jinlong Li 0003, Kecheng Chen, Meng Wang 0017, Long Xu 0001, Haoliang Li, Nicu Sebe, Sam Kwong, Shiqi Wang 0001 |
IEEE Signal Process. Lett. | 5 |
| 2026 | Learned Reference Picture Resampling Control: A Data-Centric ApproachabstractLearned reference picture resampling control (LRPRC) adaptively adjusts the coding scale for each frame using an offline-trained neural network. It demonstrates promising promising rate-distortion (R-D) performance improvements over traditional methods, particularly in high-resolution, low-bit-rate video coding scenarios. However, existing LRPRC methods rely exclusively on locally optimal decision labels derived from greedy strategies for network training, leading to suboptimal control performance. To address this limitation, we introduce a novel data-centric solution that substantially improves training label quality, thereby enhancing overall LRPRC performance. Specifically, our key contribution is a parallelized beam search-based coding scale labeling algorithm, which captures decision dependencies across coding steps and produces higher-quality training labels with enhanced R-D performance. By fully exploiting the intra-trellis and inter-trellis parallelism of beam search and hierarchical coding, our proposed labeling algorithm achieves logarithmic-squared time complexity, making it highly suitable for large-scale cluster computing. We validate this simple yet effective data-centric LRPRC approach in the Versatile Video Encoder (VVenC) using 4K video sequences. Experimental results demonstrate that merely upgrading the beam search labels (without any neural architecture re-designs) consistently outperforms the state-of-the-art LRPRC method, achieving BD-rate reductions of 5.09%, 3.98%, and 3.59% under thefast,medium, andslowpresets, respectively. Riyu Lu, Yingwen Zhang, Hengyu Man, Meng Wang 0017, Long Xu 0001, Shiqi Wang 0001, Xiaopeng Fan 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | DT-JRD: Deep Transformer-Based Just Recognizable Difference Prediction Model for Video Coding for MachinesabstractJust Recognizable Difference (JRD) represents the minimum visual difference that is detectable by machine vision, which can be exploited to promote machine vision-oriented visual signal processing. In this paper, we propose a Deep Transformer-based JRD (DT-JRD) prediction model for Video Coding for Machines (VCM), where the accurately predicted JRD can be used to reduce the coding bit rate while maintaining the accuracy of machine tasks. Firstly, we model the JRD prediction as a multi-class classification and propose a DT-JRD prediction model that integrates an improved embedding, a content and distortion feature extraction, a multi-class classification, and a novel learning strategy. Secondly, inspired by the perception property that machine vision exhibits a similar response to distortions near JRD, we propose an asymptotic JRD loss by using Gaussian Distribution-based Soft Labels (GDSL), which significantly extends the number of training labels and relaxes classification boundaries. Finally, we propose a DT-JRD-based VCM to reduce the coding bits while maintaining the accuracy of object detection. Extensive experimental results demonstrate that the mean absolute error of the predicted JRD by the DT-JRD is 5.574, outperforming the state-of-the-art JRD prediction model by 13.1%. Coding experiments show that compared with the VVC, the DT-JRD-based VCM achieves an average of 29.58% bit rate reduction while maintaining the object detection accuracy. Yun Zhang 0002, Long Xu 0001, Sam Kwong |
IEEE Trans. Multim. | 4 |
| 2025 | RGBT-Booster: Detail-Boosted Fusion Network for RGB-Thermal Crowd Counting With Local Contrastive LearningabstractWith the swift development of the Internet of Video Things (IOVT), crowd counting has demerged as an indispensable technology in the domains of intelligent transportation and video surveillance. However, due to the insufficient extraction of detail head information and the limited ability to reduce the multimodality differences, the existing methods still have large errors in accurate RGB-thermal (RGB-T) crowd counting. To this end, we propose a novel RGB-T crowd counting network, i.e., RGBT-Booster, to effectively deal with the aforementioned challenges. In RGBT-Booster, by introducing additional detail auxiliary branches for RGB and thermal infrared images and the proposed enhanced detail fusion module (EDFM), we can obtain richer low-level head detail features. In addition, we also propose a local contrastive learning (LCL) to further reduce the multimodality differences for accurate crowd counting. Experimental results on two public RGB-T crowd counting datasets (i.e., RGBT crowd counting (RGBT-CC) and DroneRGBT) and one RGB-Depth (RGB-D) crowd counting dataset (i.e., ShanghaiTechRGBD) show that the proposed RGBT-Booster achieves effective and superior counting performance, compared with previous methods. The source code and datasets used in the experiments will be released athttps://github.com/QSBAOYANGMU/RGBT-Booster. Baoyang Mu, Feng Shao 0001, Zhengxuan Xie, Long Xu 0001, Qiuping Jiang |
IEEE Internet Things J. | 4 |
| 2025 | GADFNet: Geometric Priors Assisted Dual-Projection Fusion Network for Monocular Panoramic Depth EstimationabstractPanoramic depth estimation is crucial for acquiring comprehensive 3D environmental perception information, serving as a foundational basis for numerous panoramic vision tasks. The key challenge in panoramic depth estimation is how to address various distortions in 360° omnidirectional images. Most panoramic images are displayed as 2D equirectangular projections, which exhibit significant distortion, particularly with the severe fisheye effect near the equatorial regions. Traditional depth estimation methods for perspective images are unsuitable for such projections. On the other hand, cubemap projection consists of six distortion-free perspective images, allowing the use of existing depth estimation methods. However, the boundaries between faces of a cubemap projection introduce discontinuities, causing a loss of global information when using cube maps alone. In this work, we propose an innovative geometric priors assisted dual-projection fusion network (GADFNet) that leverages geometric priors of panoramic images and the strengths of both projection types to enhance the accuracy of panoramic depth estimation. Specifically, to better focus the network on key areas, we introduce a distortion perception module (DPM) and incorporate geometric information into the loss function. To more effectively extract global information from the equirectangular projection branch, we propose a scene understanding module (SUM), which captures features from different dimensions. Additionally, to achieve effective fusion of the two projections, we design a dual projection adaptive fusion module (DPAFM) to dynamically adjust the weights of the two branches during fusion. Extensive experiments conducted on four public datasets (including both virtual and real-world scenarios) demonstrate that our proposed GADFNet outperforms existing methods, achieving superior performance. Chengchao Huang, Feng Shao 0001, Hangwei Chen, Baoyang Mu, Long Xu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Bidirectional Patch-Based Correlations With Local Rigidity for Global Nonrigid RegistrationabstractThe registration of time-varying 3D shapes with high degrees of freedom remains a challenging task. Most existing techniques attempt to address this issue by solving an optimization problem defined on deformation graph with as-rigid-as-possible smoothness prior, which usually struggle to capture large scale displacements. Motivated by the insight that a set of points tends to collectively undergo significant rigid motion accompanied by slight nonrigid deformation, we propose a two-step approach to address nonrigid registration in a coarse-to-fine manner. In the first step, coarse correlations between source and target points are constructed by estimating a set of rigid transformations for local patches which are regional clusters of points. To leverage more contextual information, a bidirectional registration module is introduced that estimates both the forward and backward patch-wise rigid transformation fields (PRTFs). Subsequently, in the second step, the source point set is warped by blending both forward and backward PRTFs and fed into a deformation optimization module. Here, unidirectional point-based correspondences are sought to refine the global nonrigid transformation fields (GNTFs) while adhering to local rigidity constraints. To illustrate the efficacy of our method, we conduct tests on challenging scenarios involving human datasets, including large displacements resulting from fast inter-frame motions or pose changes. Both qualitative and quantitative results demonstrate that our approach outperforms several state-of-the-art methods in terms of robustness and registration accuracy. Xuexin Yu, Xinggang Hu, Long Xu 0001, Xiangyang Ji |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Hallucinated-PQA: No reference point cloud quality assessment via injecting pseudo-reference features
Baoyang Mu, Feng Shao 0001, Hangwei Chen, Qiuping Jiang, Long Xu 0001, Yo-Sung Ho |
Expert Syst. Appl. | 5 |
| 2024 | Dynamic Hypergraph Convolutional Network for No-Reference Point Cloud Quality AssessmentabstractWith the rapid advancement of three-dimensional (3D) sensing technology, point cloud has emerged as one of the most important approaches for representing 3D data. However, quality degradation inevitably occurs during the acquisition, transmission, and process of point clouds. Therefore, point cloud quality assessment (PCQA) with automatic visual quality perception is particularly critical. In the literature, the graph convolutional networks (GCNs) have achieved certain performance in point cloud-related tasks. However, they cannot fully characterize the nonlinear high-order relationship of such complex data. In this paper, we propose a novel no-reference (NR) PCQA method with hypergraph learning. Specifically, a dynamic hypergraph convolutional network (DHCN) composing of a projected image encoder, a point group encoder, a dynamic hypergraph generator, and a perceptual quality predictor, is devised. First, a projected image encoder and a point group encoder are used to extract feature representations from projected images and point groups, respectively. Then, using the feature representations obtained by the two encoders, dynamic hypergraphs are generated during each iteration, aiming to constantly update the interactive information between the vertices of hypergraphs. Finally, we design the perceptual quality predictor to conduct quality reasoning on the generated hypergraphs. By leveraging the interactive information among hypergraph vertices, feature representations are well aggregated, resulting in a notable improvement in the accuracy of quality pediction. Experimental results on several point cloud quality assessment databases demonstrate that our proposed DHCN can achieve state-of-the-art performance. The code will be available at:https://github.com/chenwuwq/DHCN. Qiuping Jiang, Wei Zhou 0021, Long Xu 0001, Weisi Lin |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Blind Quality Evaluator of Light Field Images by Group-Based Representations and Multiple Plane-Oriented Perceptual CharacteristicsabstractDue to the emergency of multi-view cameras and commercial Light Field (LF) cameras, the demand of high-performance LF quality evaluator is of great significance for guiding LF acquisition, processing and application and further promoting the visual perceived quality of LF visualizations. However, LF Images (LFIs), as high-dimensional data, suffer from various quality degradations not only in the spatial domain but also in the angular domain. Therefore, it is of great challenge to predict LF quality accurately. An effective LF evaluator should be able to represent these heterogeneous artifacts. In this paper, we provide a novel No-Reference LF Quality Assessment Evaluator (NR LF-QAE) to tackle this problem. Firstly, to measure angular consistency among viewports, we utilize group-based representations to character information similarity of aligned view stacks. Secondly, to better describe the texture information of LFIs, unifying spatial-angular texture statistic measurement is performed via Local Binary Patterns from Three Orthogonal Planes (LBP-TOP). Thirdly, we design 3D Log-Gabor filters to extract LF global structure information in Sub-Aperture Images (SAIs) as spatial feature characterizations and 2D Log-Gabor filters are adopted to characterize ray direction/depth information in Epipolar Plane Images (EPIs) as angular feature characterizations. By comprehensive LF information analyses in angular consistency and spatial-angular feature extraction with texture and structure descriptors, experimental results demonstrate the superiority of the proposed NR LF-QAE over the state-of-the-art comparative models in predicting the quality of LFIs on three available benchmark databases. The code will be released athttps://github.com/zerosola/NR-LF-QAE. Xiongli Chai, Feng Shao 0001, Qiuping Jiang, Xuejin Wang, Long Xu 0001, Yo-Sung Ho |
IEEE Trans. Multim. | 5 |
| 2024 | Progressive Bidirectional Feature Extraction and Enhancement Network for Quality Evaluation of Night-Time ImagesabstractBlind image quality assessment (BIQA) has received increasing attention in the past decades. However, it still remains inadequately researched on BIQA for night-time images suffering from the diverse authentic degradations. Since the intrinsic content degradations of night-time images are highly related to the illumination, how to use the connection between content and illumination to enhance the feature representation ability is the key issue in designing BIQA methods for night-time images. In this article, we first construct an ultra-high-definition night-time image dataset (UHD-NID) with high image resolution and abundant parameter settings. UHD-NID contains 1600 images with a high resolution of 5616 × 3744, and each group of images contains ten exposure levels. Then, we conduct subjective assessment and analyze the subjective data to obtain a mean opinion score to each image in UHD-NID. To enhance the feature representation ability in content and illumination, we propose a Progressive Bidirectional Feature Extraction and Enhancement Network (PBFEE-Net). In addition, we use a decomposition network to decompose the input image into the reflectance and illumination, which can facilitate the ability of feature extraction to some extent. The experimental results show that our proposed method achieves superior performance in evaluating the quality of night-time images. Jiangli Shi, Feng Shao 0001, Chongzhen Tian, Hangwei Chen, Long Xu 0001, Yo-Sung Ho |
IEEE Trans. Multim. | 5 |
| 2024 | Stripe Sensitive Convolution for Omnidirectional Image Dehazingabstractvirtual reality (VR) experience. The recent single image dehazing methods, to date, have been only focused on plane images. In this work, we propose a novel neural network pipeline for single omnidirectional image dehazing. To create the pipeline, we build the first hazy omnidirectional image dataset, which contains both synthetic and real-world samples. Then, we propose a new stripe sensitive convolution (SSConv) to handle the distortion problems due to the equirectangular projections. The SSConv calibrates distortion in two steps: 1) extracting features using different rectangular filters and, 2) learning to select the optimal features by a weighting of the feature stripes (a series of rows in the feature maps). Subsequently, using SSConv, we design an end-to-end network that jointly learns haze removal and depth estimation from a single omnidirectional image. The estimated depth map is leveraged as the intermediate representation and provides global context and geometric information to the dehazing module. Extensive experiments on challenging synthetic and real-world omnidirectional image datasets demonstrate the effectiveness of SSConv, and our network attains superior dehazing performance. The experiments on practical applications also demonstrate that our method can significantly improve the 3-D object detection and 3-D layout performances for hazy omnidirectional images. Dong Zhao 0010, Jia Li 0003, Long Xu 0001 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2023 | Post-Training Quantization for Vision Transformer in Transformed DomainabstractAs a successor to convolutional neural networks (CNNs), transformer-based models have achieved great performance in computer vision tasks. Compressing vision transformers to low-bit brings a number of practical benefits, including higher inference speed, improved memory footprint, and reduced energy consumption. Existing model compression methods, especially quantization techniques, ignore the joint statistics of weights, resulting in sub-optimal task performance at a given quantization bit rate. In this paper, we propose to apply a transform before quantization to decorrelate vision transformer’s weights. And the entire compression flow is optimized in a rate-distortion framework to minimize the network output errors instead of simply optimizing for quantization errors or layer-wise output errors. Extensive experimental results on a variety of vision transformers (e.g. Swin, ViT and DeiT) demonstrate that our proposed method outperforms the state-of-the-art. It can quantize vision transformers (e.g. Swin, ViT and DeiT) on both weights and activations to 6-bit without a significant accuracy drop. Zhuo Chen 0006, Fei Gao 0019, Zhe Wang 0019, Long Xu 0001, Weisi Lin |
ICME | 5 |
| 2023 | Viewport-Sphere-Branch Network for Blind Quality Assessment of Stitched 360° Omnidirectional ImagesabstractCompared with conventional images/videos, omnidirectional data records rich information with higher resolution and wider Field-of-View. Moreover, the stitching distortions introduced in the panoramic content generation process make the quality assessment task more challenging. Targeting at designing an accurate and fast stitched 360° omnidirectional image quality evaluator, we propose a Viewport-Sphere-Branch Network (VSBNet) via dual-branch quality estimation. Specifically, for the viewport quality estimation, we extract distorted viewports around the stitching seams and conduct distortion rectification through a progressively complementary network to obtain pseudo-reference viewports. The qualitative and quantitative experiments validate that pseudo-reference viewports are reliable. Then, the differences between distorted and pseudo-reference viewports are quantified through transformer architecture to obtain quality scores of viewports. The introduction of pseudo-reference viewports can effectively improve the performance of the viewport quality prediction branch. To establish general scenario awareness and accurately evaluate the immersive experience, we extract feature representation through deformable convolutions to eliminate 2D-to-Sphere intrinsic sampling distortions and use multilayer perceptron to predict score of the whole sphere. The final prediction score is obtained by aggregating the quality scores from viewport and sphere branches. We evaluate the proposed VSBNet on two benchmark databases and results demonstrate that the combination of two branches can obtain more accurate results. Overall, our method is superior to existing full reference and no reference models designed for conventional images and 360° omnidirectional images. Chongzhen Tian, Feng Shao 0001, Xiongli Chai, Qiuping Jiang, Long Xu 0001, Yo-Sung Ho |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | Perceptual Quality Assessment of Enhanced Colonoscopy Images: A Benchmark Dataset and an Objective MethodabstractIn colonoscopy, the captured images are usually with low-quality appearance, such as non-uniform illumination, low contrast, etc., due to the specialized imaging environment, which may provide poor visual feedback and bring challenges to subsequent disease analysis. Many low-light image enhancement (LIE) algorithms have recently proposed to improve the perceptual quality. However, how to fairly evaluate the quality of enhanced colonoscopy images (ECIs) generated by different LIE algorithms remains a rarely-mentioned and challenging problem. In this study, we carry out a pioneering investigation on perceptual quality assessment of ECIs. Firstly, considering the lack of specific datasets, we collect 300 low-light images with diverse contents during the real-world colonoscopy and conduct rigorous subjective studies to compare the performance of 8 popular LIE methods, resulting in a benchmark dataset (named ECIQAD) for ECIs. Secondly, in view of the distinctive distortion characteristics of ECIs, we propose an effective no-reference Enhanced Colonoscopy Image Quality (ECIQ) method to automatically evaluate the perceptual quality of ECIs via analysis of brightness, contrast, colorfulness, naturalness, and noise. Extensive experiments on ECIQAD demonstrate the superiority of our proposed ECIQ method over 14 mainstream no-reference image quality assessment methods. Guanghui Yue 0001, Tianwei Zhou, Jingwen Hou, Weide Liu, Long Xu 0001, Tianfu Wang 0001, Jun Cheng 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2023 | Prediction With Visual Evidence: Sketch Classification Explanation via Stroke-Level AttributionsabstractSketch classification models have been extensively investigated by designing a task-driven deep neural network. Despite their successful performances, few works have attempted to explain the prediction of sketch classifiers. To explain the prediction of classifiers, an intuitive way is to visualize the activation maps via computing the gradients. However, visualization based explanations are constrained by several factors when directly applying them to interpret the sketch classifiers: (i) low-semantic visualization regions for human understanding. and (ii) neglecting of the inter-class correlations among distinct categories. To address these issues, we introduce a novel explanation method to interpret the decision of sketch classifiers with stroke-level evidences. Specifically, to achieve stroke-level semantic regions, we first develop a sketch parser that parses the sketch into strokes while preserving their geometric structures. Then, we design a counterfactual map generator to discover the stroke-level principal components for a specific category. Finally, based on the counterfactual feature maps, our model could explain the question of "why the sketch is classified as X" by providing positive and negative semantic explanation evidences. Experiments conducted on two public sketch benchmarks, Sketchy-COCO and TU-Berlin, demonstrate the effectiveness of our proposed model. Furthermore, our model could provide more discriminative and human understandable explanations compared with these existing works. Sixuan Liu, Jingzhi Li 0002, Hua Zhang 0008, Long Xu 0001, Xiaochun Cao |
IEEE Trans. Image Process. | 4 |
| 2023 | MagConv: Mask-Guided Convolution for Image InpaintingabstractStandard convolution applied to image inpainting would lead to color discrepancy and blurriness for treating valid and invalid/hole regions without difference, which was partially amended by partial convolution (PConv). In PConv, a binary/hard mask was maintained as an indicator of valid and invalid pixels, where valid pixels and invalid pixels were treated differently. However, it can not describe validity degree of an impaired pixel. In addition, mask and image paths were separated, without sharing convolution kernel and exchanging information mutually, reducing data utilization efficiency. In this paper, a mask-guided convolution (MagConv) is proposed for image inpainting. In MagConv, mask and image paths share a convolution kernel to interact with each other and form a joint optimization scheme. In addition, a learnable piecewise activation function is raised to replace the reciprocal function of PConv, providing more flexible and adaptable compensation to convolution contaminated by invalid pixels. It also results in a soft mask of floating-point coefficients from 0 to 1 capable of indicating the validity degree of each pixel. Last but not least, MagConv splits the convolution kernel into positive and negative weights so that they can evaluate the validity of each pixel faithfully. Qualitative and quantitative experiments on the CelebA, Paris StreetView and Places2 datasets demonstrate that our method achieves favorable visual quality against state-of-the-art approaches. Xuexin Yu, Long Xu 0001, Jia Li 0003, Xiangyang Ji |
IEEE Trans. Image Process. | 2 |
| 2022 | Channel-Wise Bit Allocation for Deep Visual Feature QuantizationabstractIntermediate deep visual feature compression and transmission is an emerging research topic, which enables a good balance among computing load, bandwidth usage and generalization ability for AI-based visual analysis in edge-cloud collaboration. Quantization and the corresponding rate-distortion optimization are the key techniques in deep feature compression. In this paper, by exploring the feature statistics and a greedy iterative algorithm, we propose a channel-wise bit allocation method for deep feature quantization optimizing for network output error. Given the limited rate and computational power, the proposed method can quantize features with small information loss. Moreover, the method also provides the option to handle the trade-offs between computational cost and quantization performance. Experimental results on ResNet and VGGNet features demonstrate the effectiveness of the proposed bit allocation method. Wei Wang 0283, Zhuo Chen 0006, Zhe Wang 0019, Jie Lin 0001, Long Xu 0001, Weisi Lin |
ICIP | 5 |
| 2022 | VSOIQE: A Novel Viewport-Based Stitched 360° Omnidirectional Image Quality EvaluatorabstractWith the rapid development of virtual reality (VR), 360° omnidirectional images and videos have drawn wide attention. However, the quality assessment of 360° omnidirectional images is a challenging task, especially when the panoramic image contains multiple stitching distortions. We propose a viewport-based stitched 360° omnidirectional image quality evaluator (VSOIQE), by first extracting the features of salient and stitching viewports, and then inferring the overall perceptual quality via multiple linear regression (MLR). Comprehensive image attributes including edge, color, shape and information entropy are considered in the framework. Experimental results on two benchmark databases demonstrate the superiority of the proposed metric over both the state-of-the-art quality models designed for 2D images and the quality models developed for 360° omnidirectional images. Chongzhen Tian, Xiongli Chai, Feng Shao 0001, Qiuping Jiang, Xiangchao Meng, Long Xu 0001, Yo-Sung Ho |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2021 | DehazeFlow: Multi-scale Conditional Flow Network for Single Image DehazingabstractSingle image dehazing is a crucial and preliminary task for many computer vision applications, making progress with deep learning. The dehazing task is an ill-posed problem since the haze in the image leads to the loss of information. Thus, there are multiple feasible solutions for image restoration of a hazy image. Most existing methods learn a deterministic one-to-one mapping between a hazy image and its ground-truth, which ignores the ill-posedness of the dehazing task. To solve this problem, we propose DehazeFlow, a novel single image dehazing framework based on conditional normalizing flow. Our method learns the conditional distribution of haze-free images given a hazy image, enabling the model to sample multiple dehazed results. Furthermore, we propose an attention-based coupling layer to enhance the expression ability of a single flow step, which converts natural images into latent space and fuses features of paired data. These designs enable our model to achieve state-of-the-art performance while considering the ill-posedness of the task. We carry out sufficient experiments on both synthetic datasets and real-world hazy images to illustrate the effectiveness of our method. The extensive experiments indicate that DehazeFlow surpasses the state-of-the-art methods in terms of PSNR, SSIM, LPIPS, and subjective visual effects. Jia Li 0003, Dong Zhao 0008, Long Xu 0001 |
ACM Multimedia | 4 |
| 2021 | Correlation filter via random-projection based CNNs features combination for visual tracking
Mingke Zhang, Long Xu 0001, Xuande Zhang |
J. Vis. Commun. Image Represent. | 2 |
| 2021 | Perceptual Redundancy Estimation of Screen Images via Multi-Domain SensitivitiesabstractVisual redundancy detection is essential for image and video communication. Human visual system (HVS) is difficult to perceive the pixel magnitude change below a certain visibility threshold which is also known as just-noticeable-difference (JND). In this letter, we present an efficient JND estimation approach for screen content images by considering high-frequency sensitivity and orientation sensitivity correction. Specifically, to better quantify the visual redundancy, we investigate the visibility threshold based on the high-frequency distortion sensitivity. To obtain the orientation sensitivity correction, we divide the screen image pixels into three levels based on the oblique effect that considers the sensitive integrity of edges. Compared with several state-of-the-art JNDs, experimental results show that our method tolerates more perceptual redundancy, and delivers better visual quality under the same injected-noise energy. The implementation of the proposed method is publicly available at https://sites.google.com/site/wangmiaohui/. Miaohui Wang, Wuyuan Xie, Long Xu 0001 |
IEEE Signal Process. Lett. | 4 |
| 2021 | Pyramid Global Context Network for Image DehazingabstractHaze caused by atmospheric scattering and absorption would severely affect scene visibility of an image. Thus, image dehazing for haze removal has been widely studied in the literature. Within a hazy image, haze is not confined in a small local patch/position, while widely diffusing in a whole image. Under this circumstance, global context is a crucial factor in the success of dehazing, which was seldom investigated in existing dehazing algorithms. In the literature, the global context (GC) block has been designed to learn point-wise long-range dependencies of an image for global context modeling; however, patch-wise long-range dependencies were ignored. To image dehazing, patch-wise long-range dependencies should be highlighted to cooperate with patch-wise operations of image dehazing. In this paper, we first extend the point-wise GC into a Pyramid Global Context (PGC), which is a multi-scale GC, after undergoing the pyramid pooling. Thus, patch-wise long-range dependencies can be explored by the PGC. Then, the proposed PGC is plugged into a U-Net, getting an attentive U-Net. Further, the attentive U-Net is optimized by importing ResNet's shortcut connection and dilated convolution. Thus, the finalized dehazing model can explore both long-range and patch-wise context dependencies for global context modeling, which is crucial for image dehazing. The extensive experiments on synthetic databases and real-world hazy images demonstrate the superiority of our model over other representative state-of-the-art models from both quantitative and qualitative comparisons. Dong Zhao 0016, Long Xu 0001, Lin Ma 0002, Jia Li 0003, Yihua Yan |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2020 | Rate Constrained Multiple-QP Optimization for HEVCabstractIn High Efficiency Video Coding (HEVC), multiple-QP (quantization parameter) optimization can adapt to a local video content. However, the multiple-QP implementation in the HEVC reference software (HM 16.6) achieves the best QP value for each coding block with a large amount of computational complexity. To address this challenge, we propose a fast rate-constrained multiple-QP optimization approach for the HM platform. We first introduce a template-based transform coefficient selection method which can save the overall complexity of entropy coding. In addition, we model the multiple-QP determination as a new rate-constrained optimization problem, and finally, we get a feasible solution with a lower computation overhead. Experimental results show that our method dramatically reduces the average complexity under the all-intra, low-delay and random-access configuration. Miaohui Wang, Jian Xiong 0005, Long Xu 0001, Wuyuan Xie, King Ngi Ngan, Harry Qin |
IEEE Trans. Multim. | 3 |
| 2019 | Cross-Reference Stitching Quality Assessment for 360° Omnidirectional ImagesabstractAlong with the development of virtual reality (VR), omnidirectional images play an important role in producing multimedia content with an immersive experience. However, despite various existing approaches for omnidirectional image stitching, how to quantitatively assess the quality of stitched images is still insufficiently explored. To address this problem, we first establish a novel omnidirectional image dataset containing stitched images as well as dual-fisheye images captured from standard quarters of 0$^\circ$, 90$^\circ$, 180$^\circ$, and 270$^\circ$. In this manner, when evaluating the quality of an image stitched from a pair of fisheye images (\eg, 0$^\circ$ and 180$^\circ$), the other pair of fisheye images (\eg, 90$^\circ$ and 270$^\circ$) can be used as the cross-reference to provide ground-truth observations of the stitching regions. Based on this dataset, we propose a set of Omnidirectional Stitching Image Quality Assessment (OS-IQA) metrics. In these metrics, the stitching regions are assessed by exploring the local relationships between the stitched image and its cross-reference with histogram statistics, perceptual hash and sparse reconstruction, while the whole stitched images are assessed by the global indicators of color difference and fitness of blind zones.Qualitative and quantitative experiments show our method outperforms the classic IQA metrics and is highly consistent with human subjective evaluations. To the best of our knowledge, it is the first attempt that assesses the stitching quality of omnidirectional images by using cross-references. Jia Li 0003, Kaiwen Yu, Yifan Zhao 0002, Yu Zhang 0035, Long Xu 0001 |
ACM Multimedia | 5 |
| 2019 | Dual-scale weighted structural local sparse appearance model for object trackingabstractIt is a great challenge to develop an effective appearance model for robust visual tracking due to various interfering factors, such as pose change, occlusion, background clutter etc. More and more visual tracking methods tend to exploit the local appearance model to deal with the above challenges. In this study, the authors present a simple yet effective weighted structural local sparse appearance model, which can better describe the target appearance information through patch‐based generative weight. To further improve the robustness of tracking, they implement this appearance model on two‐scale patches. The two derived appearance models are then combined to form a collaborative model to play their advantages. Extensive experiments on the tracking benchmark dataset show that the proposed method performs favourably against several state‐of‐the‐art methods. Xianyou Zeng, Long Xu 0001, Yi-Gang Cen, Ruizhen Zhao, Wanli Feng |
IET Comput. Vis. | 2 |
| 2019 | IDeRs: Iterative dehazing method for single remote sensing image
Long Xu 0001, Dong Zhao 0016, Yihua Yan, Sam Kwong, Jie Chen 0006, Ling-Yu Duan |
Inf. Sci. | 1 |
| 2019 | Multi-scale Optimal Fusion model for single image dehazing
Dong Zhao 0016, Long Xu 0001, Yihua Yan, Jie Chen 0006, Ling-Yu Duan |
Signal Process. Image Commun. | 2 |
| 2018 | Image Quality Assessment Based Label Smoothing in Deep Neural Network LearningabstractFor many computer vision problems, deep neural networks are trained and validated based on the assumption that the input images are pristine (i.e., artifact-free). However, digital images are subject to a wide range of distortions in real application scenarios, while the practical issues regarding image quality in high level visual information understanding have been largely ignored. In this paper, in view of the fact that most widely deployed deep learning models are susceptible to various image distortions, distorted images are involved for data augmentation in the deep neural network training process to learn a reliable model for practical applications. In particular, an image quality assessment based label smoothing method, which aims at regularizing the label distribution of training images, is further proposed to tune the objective functions in learning the neural network. Experimental results show that the proposed method is effective in dealing with both low and high quality images in the typical image classification task. Zhuo Chen 0006, Weisi Lin, Shiqi Wang 0001, Long Xu 0001, Leida Li |
ICASSP | 4 |
| 2018 | Image processing for synthesis imaging of mingantu spectral radioheliograph (MUSER)
Long Xu 0001, Yihua Yan, Lin Ma 0002, Yun Zhang 0002 |
Multim. Tools Appl. | 1 |
| 2017 | Data processing and imaging for MingantU SpEctral Radioheliograph (MUSER)abstractMingantU SpEctral Radioheliograph (MUSER) is a radio synthesis imaging array dedicated to observe the sun, operating on multiple frequencies from decimeter to centimeter range. In the experimental observations MUSER produces about 2.55 Tera Bytes data daily, which will be processed offline to get the solar radio images. In this paper the current data processing and imaging procedure for MUSER will be discussed in details. Some algorithms are presented to solve the problems in the imaging with MUSER. Linjie Chen, Yihua Yan, Long Xu 0001 |
VCIP | 4 |
| 2017 | Auto-flag the baseline for Mingantu Ultrawide Spectral Radioheliograph with LSTMabstractMingantu Ultrawide Spectral Radioheliograph (MUSER) is an aperture synthesis telescope consisting of a group of small antennas to image the Sun. Each two antennas form a baseline contributing a Fourier sampling point for each time of imaging. Both amplitude and phase of a baseline form a time sequence in a time interval. Normally, amplitude/phase should vary linearly in a short time interval. However, many reasons would result in errors of baselines, damaging image quality. This work makes the first attempt to auto-flag baselines so as to delete bad baselines and improve image quality. Inspired by the big success of long short-term memory (LSTM) for time series analysis, we believe that LSTM could accomplish auto-flagging (a binary classification task) of baselines even better by treating them as time sequences. Thus, LSTM is employed to learn the representation of the phase of each baseline for classification, where the interaction and connection within a phase sequence are explored, benefiting classification. The experimental results demonstrate that LSTM can well capture the characteristics of the phase of a baseline, and thus achieves better classification. Jun Cheng 0003, Long Xu 0001, Xuexin Yu, Linjie Chen, Wei Wang 0141, Yihua Yan |
VCIP | 2 |
| 2017 | Learning solar flare forecasting model from magnetogramsabstractSolar flare is one type of violent eruptions from the Sun. Its effects almost immediately arrive to the near-Earth environment, so it is crucial to forecast solar flares in space weather. So far, the physical mechanisms of solar flares are not yet clear, hence we learn a solar flare forecasting model from the historical observational magnetograms by using the deep learning method. Instead of designing the feature extractor by the solar physicist in the traditional solar flare forecasting model, the proposed forecasting model can automatically learn features from input raw data, and followed by a classifier for foretasting from the learned features. The experimental results demonstrate that the proposed model can achieve better performance of solar flare forecasting comparing to traditional solar flare forecasting models. Huaning Wang, Long Xu 0001, Wenqing Sun |
VCIP | 3 |
| 2017 | Bidirectional LSTM for ionospheric vertical Total Electron Content (TEC) forecastingabstractThe ionosphere is a region over earth's atmosphere which is ionized by solar radiation. It plays an important part in atmospheric electricity and forms the inner edge of the magnetosphere, furthermore, it has practical importance for its effects on radio propagation to earth. Accordingly, its cyclic changes and disturbances could influence communication, navigation, radar seriously. Total Electron Content (TEC) is an important parameter reflecting ionospheric significant characteristic. We can analyse ionospheric disturbance and cyclic change by investigating TEC. Now, the time sequence of TEC and relevant parameters is available for our research. In addition, Bidirectional long short-term memory (Bi-LSTM) is believed to have potential in time sequence processing. Hence, we investigate TEC forecast by referring to Bi-LSTM in this work. Experimental results demonstrate that Bi-LSTM can capture the cyclic change feature of TEC, and perform better than LSTM and multi-LSTM for TEC forecast. Wenqing Sun, Long Xu 0001, Tianjiao Yuan, Yihua Yan |
VCIP | 2 |
| 2017 | Learning visual saliency from human fixations for stereoscopic images
Yuming Fang 0001, Jianjun Lei 0001, Jia Li 0003, Long Xu 0001, Weisi Lin, Patrick Le Callet |
Neurocomputing | 4 |
| 2017 | Multimodal deep learning for solar radio burst classification
Lin Ma 0002, Zhuo Chen 0006, Long Xu 0001, Yihua Yan |
Pattern Recognit. | 3 |
| 2017 | Multi-Task Rank Learning for Image Quality AssessmentabstractIn practice, images are distorted by more than one distortion. For image quality assessment (IQA), existing machine learning (ML)-based methods generally establish a unified model for all the distortion types, or each model is trained independently for each distortion type, which is therefore distortion aware. In distortion-aware methods, the common features among different distortions are not exploited. In addition, there are fewer training samples for each model training task, which may result in overfitting. To address these problems, we propose a multi-task learning framework to train multiple IQA models together, where each model is for each distortion type; however, all the training samples are associated with each model training task. Thus, the common features among different distortion types and the said underlying relatedness among all the learning tasks are exploited, which would benefit the generalization ability of trained models and prevent overfitting possibly. In addition, pairwise image quality ranking instead of image quality rating is optimized in our learning task, which is fundamentally departed from traditional ML-based IQA methods toward better performance. The experimental results confirm that the proposed multi-task rank-learning-based IQA metric is prominent against all state-of-the-art nonreference IQA approaches. Long Xu 0001, Jia Li 0003, Weisi Lin, Yongbing Zhang 0002, Lin Ma 0002, Yuming Fang 0001, Yihua Yan |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2016 | A benchmark for robustness analysis of visual tracking algorithmsabstractIn this study, we investigate the robustness of existing visual tracking algorithms with quality-degraded video. A video database including the reference video sequences and their distorted versions is created as the benchmark for robustness analysis of visual tracking algorithms. Ten existing visual tracking algorithms are used to conduct the experiments for robustness analysis based on the benchmark. Our initial investigation demonstrates that all the existing visual tracking algorithms cannot obtain the robust visual tracking results for quality-degraded video sequences. The experimental results in this study show that there is still much room for the design of robust visual tracking algorithms. Yuming Fang 0001, Yuan Yuan 0029, Long Xu 0001, Weisi Lin |
ICASSP | 3 |
| 2016 | Perceptual image quality enhancement for solar radio imageabstractIn solar radio observation, the visualization of data is very important since it can more intuitively and clearly deliver interest information of solar radio activities to astronomers. As to visualization, we highly expect good visual quality of images/videos in favor of the discovery of solar radio events recorded by observation data. The existing imaging system cannot guarantee good visual quality of solar radio data visualization. In this paper, an image quality enhancement algorithm is developed to improve solar radio extreme ultraviolet (EUV) images from Solar Dynamics Observatory (SDO). Firstly, the guided filter is employed to smooth image, which outputs an image with good skeleton and edges. Since the fine structures of solar radio activities are embedded in high frequency components of a solar radio image, we propose a novel structure preserving filtering to amplify the different signal of original input image subtracting smoothed one. Afterwards, fusing the amplified details and smoothed one together, the final enhanced image is generated. The experimental results prove that the image quality is significantly improved by using the proposed image quality enhancement algorithm. Long Xu 0001, Lin Ma 0002, Zhuo Chen 0006, Xianyou Zeng, Yihua Yan |
QoMEX | 1 |
| 2016 | Content adaptive directional transform for high efficiency video codingabstractHEVC is an emerging new standard for digital video compression, which is regarded as a successor to H.264/AVC standard. It still belongs to block-based hybrid video coding framework. The block patterns range from 4×4 to 64×64 blocks, and DCT is extended from 4×4 to 32×32. 2D-DCT for image is performed along the vertical and horizontal directions, so it is good at the energy compaction of residual block with vertical or horizontal edges. However, the edges are usually neither horizontal nor vertical for most cases, such as neither vertical nor horizontal intra prediction, so directional transform was explored in the past several years. In this paper, a directional transform adaptive to image content is proposed. Firstly, the prediction residual blocks are collected from coding a number of video sequences with plenty of image content. Secondly, for each kind of image content, the residual blocks are clustered to form the given number of clusters. Thirdly, each cluster contributes a transform basis after Singular Value Decomposition (SVD). The experimental results in terms of PSNR gains demonstrate the efficiency of the proposed algorithm with the comparison with the standard HM software. Long Xu 0001, Lin Ma 0002, Yun Zhang 0002, Yihua Yan |
VCIP | 1 |
| 2016 | Interest points based collaborative trackingabstractIn this paper, we propose a robust collaborative tracking algorithm based on interest points detection and template matching in sparse representation framework. In the proposed tracker, the target dictionary and the candidate dictionary are constructed with the patches around interest points of the previous frame and the current frame, respectively. The correspondence between target points and candidate points is computed by solving an Li minimization problem. Only the mutually matched target and candidate point pairs are selected, and the displacements of them are measured to generate the candidate targets in current frame. To find the best one among all candidate targets, each of them is sparsely represented by target templates given as benchmarks in initial frame. The candidate target with the smallest projection error produces the final tracking result. The experimental results show that the proposed tracker is superior to the state-of-the-art methods remarkably with respect to the tracking accuracy. Xianyou Zeng, Long Xu 0001, Lin Ma 0002, Ruizhen Zhao |
VCIP | 2 |
| 2016 | Early DIRECT mode decision based on all-zero block and rate distortion cost for multiview video codingabstractThe exhaustive variable‐block‐size mode decision can efficiently remove the redundancies among the multiview videos, while it also leads to significant increase of computational complexity in the multiview video coding (MVC) encoder, and the high encoding complexity becomes a bottleneck for the MVC encoder to achieve real‐time multimedia applications. To address this bottleneck, many fast mode decision methods have been proposed. However, most of them are only suitable for optimising the encoding complexity of the odd views of the MVC encoder. In this study, based on the property of the all‐zero block and rate distortion (RD) cost of the DIRECT mode as well as the correlations between the current macroblock (MB) and its spatial–temporal nearby MBs, an early DIRECT mode decision method is proposed for reducing the encoding complexity of the MVC. Experimental results show that the proposed method achieves 48.25 and 55.64% on average encoding time saving for the even and odd views, respectively, whereas the RD performance degradation is quite acceptable. In summary, the proposed method efficiently reduces the encoding complexity for the MVC encoder. Zhaoqing Pan, Yun Zhang 0002, Jianjun Lei 0001, Long Xu 0001, Xingming Sun |
IET Image Process. | 4 |
| 2016 | Imaging and representation learning of solar radio spectrums for classification
Zhuo Chen 0006, Lin Ma 0002, Long Xu 0001, Chengming Tan, Yihua Yan |
Multim. Tools Appl. | 3 |
| 2016 | A Polynomial Approximation Motion Estimation Model for Motion-Compensated Frame InterpolationabstractMotion-compensated frame interpolation (MCFI) usually finds the most matched blocks by minimizing pixel intensity discrepancies between neighboring frames along the motion trajectory. However, the quality of interpolated frames is susceptible to inaccurate motion vectors possibly for regions with complex texture patterns, irregularly shaped objects, repeated patterns, motion blurring or aliasing, and so on. It is believed that pixel intensity across adjacent frames varies gradually and smoothly, and therefore can be modeled mathematically by a continuous and differentiable function. Thus, the pixel intensity within one frame can be expressed as either a forward polynomial approximation (FWPA) or a backward polynomial approximation (BWPA) by the Taylor expansion in this paper. The discrepancy between the FWPA and BWPA is employed to find the best motion vector. In addition, a motion-aligned partial derivative is proposed to calculate the Taylor expansion along the motion trajectory. The proposed method is applicable to any existing MCFI schemes and achieve superior performance by consuming relatively more buffer memories and computational resources. Extensive experimentation with comparison with previous techniques validates our method in terms of both objective and subjective criteria. Yongbing Zhang 0002, Long Xu 0001, Xiangyang Ji, Qionghai Dai |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2016 | No-Reference Retargeted Image Quality Assessment Based on Pairwise Rank LearningabstractIn this paper, we propose a novel no-reference image quality assessment method for the retargeted image based on the pairwise rank learning approach. Each retargeted image needs to be first represented as a feature vector, which not only captures the image characteristics but also is sensitive to distortions during the retargeting process. As such, we investigate and examine different image representations for their abilities depicting the perceptual quality of retargeted image. Based on the image representations, we resort to the pairwise rank learning approach to discriminate the perceptual quality between the retargeted image pairs. Experimental results demonstrate that the proposed method can effectively depict the perceptual quality of the retargeted image, which can even perform comparably with the full-reference quality assessment methods. Lin Ma 0002, Long Xu 0001, Yichi Zhang 0014, Yihua Yan, King Ngi Ngan |
IEEE Trans. Multim. | 2 |
| 2016 | Free-Energy Principle Inspired Video Quality Metric and Its Use in Video CodingabstractIn this paper, we extend the free-energy principle to video quality assessment (VQA) by incorporating with the recent psychophysical study on human visual speed perception (HVSP). A novel video quality metric, namely the free-energy principle inspired video quality metric (FePVQ), is therefore developed and applied to perceptual video coding optimization. The free-energy principle suggests that the human visual system (HVS) can actively predict “orderly” information and avoid “disorderly” information for image perception. Basically, “orderly” is associated with the skeletons and edges of objects, and “disorderly” mostly concerns textures in images. Based on this principle, an image is separated into orderly and disorderly regions, and processed differently in image quality assessment. For videos, visual attention, or fixation, is associated with the objects with significant motion according to HVSP, resulting in a motion strength factor in the FePVQ so that the free-energy principle is extended into spatio-temporal domain for VQA. In addition, we investigate the application of the FePVQ in perceptual rate distortion optimization (RDO). For this purpose, the FePVQ is realized with low computational cost by using the relative total variation model and the block-wise motion vectors of video coding to simulate the free-energy principle and the HVSP, respectively. The experimental results indicate that the proposed FePVQ is highly consistent with the HVS perception. The linear correlation coefficient and Spearman's rank-order correlation coefficient are up to 0.8324 and 0.8281 on the LIVE video database. Better perceptual quality of encoded video sequences is achieved by FePVQ-motivated RDO in video coding. Long Xu 0001, Weisi Lin, Lin Ma 0002, Yongbing Zhang 0002, Yuming Fang 0001, King Ngi Ngan, Songnan Li, Yihua Yan |
IEEE Trans. Multim. | 1 |
| 2015 | Multi-task rank learning for image quality assessmentabstractIn practice, multiple types of distortions are associated with an image quality degradation process. The existing machine learning (ML) based image quality assessment (IQA) approaches generally established a unified model for all distortion types, or each model is trained independently for each distortion type by using single-task learning, which lead to the poor generalization ability of the models as applied to practical image processing. There are often the underlying cross relatedness amongst these single-task learnings in IQA, which is ignored by the previous approaches. To solve this problem, we propose a multi-task learning framework to train IQA models simultaneously across individual tasks each of which concerns one distortion type. These relatedness can be therefore exploited to improve the generalization ability of IQA models from single-task learning. In addition, pairwise image quality rank instead of image quality rating is optimized in learning task. By mapping image quality rank to image quality rating, a novel no-reference (NR) IQA approach can be derived. The experimental results confirm that the proposed Multi-task Rank Learning based IQA (MRLIQ) approach is prominent among all state-of-the-art NR-IQA approaches. Long Xu 0001, Jia Li 0003, Weisi Lin, Yongbing Zhang 0002, Lin Ma 0002, Yuming Fang 0001, Yun Zhang 0002, Yihua Yan |
ICASSP | 1 |
| 2015 | Multimodal Learning for Classification of Solar Radio SpectrumabstractThis paper proposes the first attempt to utilize multi-modal learning method for the representation learning of the solar radio spectrums. The solar radio signals sensed from differ-ent frequency channels, which present different characteristics, are regarded as different modalities. We employ a multimodal neural network to learn the representations of the solar radio spectrum, which can distinguish the differences and learn the interactions between different modalities. The original solar ra-dio spectrums are firstly pre-processed, including normalization, denoising, channel competition and etc., before being fed into the multimodal learning network. Experimental results have demon-strated that the proposed multimodal learning network can learn the representation of the solar radio spectrum more effectively, and improve the classification accuracy. Zhuo Chen 0006, Lin Ma 0002, Long Xu 0001, Ying Weng, Yihua Yan |
SMC | 3 |
| 2015 | Rank Learning Based No-Reference Quality Assessment of Retargeted ImagesabstractIn this paper, we first propose a novel no-reference (NR) image quality assessment (IQA) method for retargeted image based on the rank learning approach. Firstly, image features for each retargeted image are extracted, which should not only represent the image characteristics but also be sensitive to the retargeted distortions. Specifically, the image feature should be able to capture the shape distortions, which are the commonly encountered distortions of the retargeted image. Based on the extracted image features, the rank learning method is employed to train a model to discriminate the perceptual quality of the retargeted image. Experimental results demonstrate that the proposed method can effectively depict the perceptual quality of the retargeted image, which can even perform comparably with the full-reference (FR) quality assessment methods. Lin Ma 0002, Long Xu 0001, Yichi Zhang 0014, King Ngi Ngan, Yihua Yan |
SMC | 2 |
| 2015 | Machine Learning-Based Coding Unit Depth Decisions for Flexible Complexity Allocation in High Efficiency Video CodingabstractIn this paper, we propose a machine learning-based fast coding unit (CU) depth decision method for High Efficiency Video Coding (HEVC), which optimizes the complexity allocation at CU level with given rate-distortion (RD) cost constraints. First, we analyze quad-tree CU depth decision process in HEVC and model it as a three-level of hierarchical binary decision problem. Second, a flexible CU depth decision structure is presented, which allows the performances of each CU depth decision be smoothly transferred between the coding complexity and RD performance. Then, a three-output joint classifier consists of multiple binary classifiers with different parameters is designed to control the risk of false prediction. Finally, a sophisticated RD-complexity model is derived to determine the optimal parameters for the joint classifier, which is capable of minimizing the complexity in each CU depth at given RD degradation constraints. Comparative experiments over various sequences show that the proposed CU depth decision algorithm can reduce the computational complexity from 28.82% to 70.93%, and 51.45% on average when compared with the original HEVC test model. The Bjøntegaard delta peak signal-to-noise ratio and Bjøntegaard delta bit rate are -0.061 dB and 1.98% on average, which is negligible. The overall performance of the proposed algorithm outperforms those of the state-of-the-art schemes. Yun Zhang 0002, Sam Kwong, Xu Wang 0006, Hui Yuan 0001, Zhaoqing Pan, Long Xu 0001 |
IEEE Trans. Image Process. | 6 |
| 2014 | Rank learning on training set selection and image quality assessmentabstractMachine learning (ML) techniques are widely used in recent no-reference visual quality assessment (NR-VQA) metrics by training on subjective image quality databases. In these metrics, the optimization function is constructed based on L2norm of the distance between subjective image quality and predicted image quality. There are two problems in these L2norm based methods: (1) human's opinion on subjective image quality rating is not reliable at fine-scale level. A small difference between subjective image qualities represented by mean opinion scores (MOSs) of two images may not truly reflect the real quality difference between these two images, but acts as noise. The optimization process should avoid such noise. (2) Generally, human's opinion on pairwise comparison (PC) for image quality is more reliable and believable than MOS. The importance of PC is ignored during the optimization process of existing ML-based studies, which are designed based on the numerical rating system. In this paper, we introduce image quality ranking concept to establish a new optimization objective instead of L2norm optimization, and then a novel NR-VQA is constructed based on ranking learning. The proposed metric firstly suggests a reasonable training set for ML, which is ignored by existing ML-based NR-VQA. The ranking theory is adopted to build optimization function, which reflects the properties of PC over the numerical ranting system used by traditional NR-VQA. By ignoring the small difference between MOSs from two images during the optimization process, the proposed ranking-based NR-VQA can also well address the first problem from the existing related metrics. Experimental results show that the proposed ranking-based NR-VQA can obtain better performance over the state-of-the-art NR-VQA approaches. Long Xu 0001, Weisi Lin, Jia Li 0003, Xu Wang 0006, Yihua Yan, Yuming Fang 0001 |
ICME | 1 |
| 2014 | Reduced-reference image quality assessment with local binary structural patternabstractReduced-reference (RR) image quality assessment (IQA) aims to use less reference data and achieve higher quality prediction accuracy. Recent researches confirm that the human visual system (HVS) is adapted to extract structural information and is sensitive to structure degradation. Therefore, in this paper, we try to represent image contents with several structural patterns, and measure image quality according to the structural degradation on these patterns. The classic local binary patterns (LBPs) are firstly employed to extract image structures and create LBP based structural histogram. And then, the structural degradation is computed as the histogram distance between the reference and distorted images. Experimental results on three large databases demonstrate that the proposed RR IQA method greatly improved the quality prediction accuracy. Jinjian Wu, Weisi Lin, Guangming Shi, Long Xu 0001 |
ISCAS | 4 |
| 2014 | Study on subjective quality assessment of Digital Compound ImagesabstractQuality assessment of digital compound images is a less investigated research topic. In this paper, we present a study for subjective quality assessment of Digital Compound Images (DCIs), and investigate whether existing Image Quality Assessment (IQA) methods are effective to evaluate the quality of distorted DCIs. A new Compound Image Quality Assessment Database (CIQAD) is constructed, including 24 reference DCIs and their 576 distorted versions. The Paired Comparison (PC) method is employed for the subjective viewing, and the Hodgerank decomposition is adopted to generate incomplete but balanced comparison pairs, so as to reduce the execution time while guaranteeing the reliability of the results. In our experiment, correlation of 14 existing IQA methods with the obtained Mean Opinion Score (MOS) values on the CIQAD is calculated, which indicates that the 14 IQA methods are not consistent with human visual perception when judging DCIs in different conditions. Therefore, objective quality assessment metrics should be specifically designed for DCIs. Our subjective study has delivered convincing information to guide the construction of objective metrics. Furthermore, we has also published the database online to favor future research on quality assessment of DCIs. Huan Yang 0001, Weisi Lin, Chenwei Deng, Long Xu 0001 |
ISCAS | 4 |
| 2014 | Generalized Nash Bargaining Solution to Rate Control Optimization for Spatial Scalable Video CodingabstractRate control (RC) optimization is indispensable for scalable video coding (SVC) with respect to bitstream storage and video streaming usage. From the perspective of centralized resource allocation optimization, the inner-layer bit allocation problem is similar to the bargaining problem. Therefore, bargaining game theory can be employed to improve the RC performance for spatial SVC. In this paper, we propose a bargaining game based one-pass RC scheme for spatial H.264/SVC. In each spatial layer (SL), the encoding constraints, such as bit rates, buffer size are jointly modeled as resources in the inner-layer bit allocation bargaining game. The modified rate-distortion (R-D) model incorporated with the inter-layer coding information is investigated. Then the generalized Nash bargaining solution (NBS) is employed to achieve an optimal bit allocation solution. The bandwidth is allocated to the frames from the generalized NBS adaptively based on their own bargaining powers. Experimental results demonstrate that the proposed rate control algorithm achieves appealing image quality improvement and buffer smoothness. The average mismatch of our proposed algorithm is within the range of 0:19%2:63%. Xu Wang 0006, Sam Kwong, Long Xu 0001, Yun Zhang 0002 |
IEEE Trans. Image Process. | 3 |
| 2014 | Efficient H.264/AVC Video Coding with Adaptive TransformsabstractTransform has been widely used to remove spatial redundancy of prediction residuals in the modern video coding standards. However, since the residual blocks exhibit diverse characteristics in a video sequence, conventional transform methods with fixed transform kernels may result in low efficiency. To tackle this problem, we propose a novel content adaptive transform framework for the H.264/AVC-based video coding. The proposed method utilizes pixel rearrangement to dynamically adjust the transform kernels to adapt to the video content. In addition, unlike the traditional adaptive transforms, the proposed method obtains the transform kernels from the reconstructed block, and hence it consumes only one logic indicator for each transform unit. Moreover, a spiral-scanning method is developed to reorder the transform coefficients for better entropy coding. Experimental results on the Key Technical Area (KTA) platform show that the proposed method can achieve an average bitrate reduction of about 7.95% and 7.0% under all-intra and low-delay configurations, respectively. Miaohui Wang, King Ngi Ngan, Long Xu 0001 |
IEEE Trans. Multim. | 3 |
| 2013 | Early termination for TZSearch in HEVC Motion EstimationabstractThe TZSearch algorithm was adopted in the high efficiency video coding reference software HM as a fast Motion Estimation (ME) algorithm for its excellent performance in reducing ME time and maintaining a comparable Rate Distortion (RD) performance. However, the multiple initial search point decision and the hybrid block matching search contribute a relatively high computational complexity to TZSearch. In this paper, based on the statistical analysis of the probability of median predictor to be selected as the final best point in the large Coding Units (CUs) (64×64, 32×32) and small CUs (16×16, 8×8) as well as the center-biased characteristic of the final best search point in ME process, we propose two early terminations for TZSearch. Experimental results show that the proposed early terminations can achieve 38.96% encoding time saving, while the RD performance degradation is quite acceptable. Zhaoqing Pan, Yun Zhang 0002, Sam Kwong, Xu Wang 0006, Long Xu 0001 |
ICASSP | 5 |
| 2013 | Reduced reference video quality assessment based on spatial HVS mutual masking and temporal motion estimationabstractIn this paper, an effective reduced reference (RR) video quality assessment (VQA) is proposed by depicting both the spatial and temporal statistical characteristics of the video signals. For each video frame, spatial information change (SIC) is employed to depict the energy variation. A novel mutual masking strategy based on the extracted SIC is proposed to accurately simulate the human visual system (HVS) texture masking property. For adjacent video frames, the temporal relationship is depicted by block-based motion estimation (BME). The generalized Gaussian density (GGD) function is employed to depict the histogram natural statistic of the residual frame after BME. The city-block distance (CBD) is used to measure the distance between histograms of the original and distorted video sequence. By pooling the measurements from both spatial and temporal perspectives, an efficient RR VQA is constructed. With the evaluations on the public video quality database, the proposed RR VQA demonstrated to be more effective than the representative RR VQAs and even the full-reference (FR) VQAs, such as peak signal-to-noise ratio (PSNR) and structure similarity index (SSIM) in matching the subjective ratings. Furthermore, the proposed RR VQA demonstrated to be much more effective and efficient, requiring only a very small number of bits for the RR feature representation. Lin Ma 0002, King Ngi Ngan, Long Xu 0001 |
ICME | 3 |
| 2013 | High quality image construction from multiple low quality copiesabstractIn this paper, the authors proposed to construct a high quality image based on multiple low quality input images. The relationship between one pixel and its neighbourhood should be consistent between different degraded images. Therefore, the reconstruction method is proposed by enforcing the pixel consistency property, which is ensured by estimating the parameters of piecewise image model for each pixel. Subsequently, the reconstructed coefficients are regularized within a reasonable range. Experimental results on multiple images (with different distortions of different levels) have demonstrated that the proposed method can effectively alleviate the noises meanwhile preserve the detailed information. Better quality images in terms of both objective and subjective measurements can be generated. Lin Ma 0002, Long Xu 0001, Qian Zhang 0001, King Ngi Ngan |
MMSP | 2 |
| 2013 | Visual quality metric for perceptual video codingabstractThe visual quality assessment (VQA) becomes prevailing in the studies of image and video coding. It assesses the quality of image or video more accurately than mean square error (MSE) with respect to the human visual system (HVS). Toward perceptual video coding, MSE is weighted spatially and temporally to simulate the HVS response to visual signal in this paper. Firstly, the image content is depicted by edge strength to compose spatial weighting factors. Secondly, the motion strength calculated from motion vector of each block gives temporal weighting factors. Thirdly, the motion trajectory based saliency map for video signal is integrated as another weighting factor of MSE. The proposed VQM not only efficiently model HVS but also relate to quantization parameter (QP) capable of guiding perceptual video coding. A perceptual rate distortion optimization (RDO) is established on the proposed VQM. The experimental results indicate that the proposed VQM is consistent well with HVS. In addition, the better rate-distortion efficiency and accurate bit rate control can be achieved by the proposed visual quality control algorithm. Long Xu 0001, Lin Ma 0002, King Ngi Ngan, Weisi Lin, Ying Weng |
VCIP | 1 |
| 2013 | Rate control for consistent visual quality of H.264/AVC encoding
Long Xu 0001, Sam Kwong, Hanli Wang, Debin Zhao, Wen Gao 0001 |
Signal Process. Image Commun. | 1 |
| 2013 | Consistent Visual Quality Control in Video CodingabstractVisual quality consistency is one of the most important issues in video quality assessment. When people view a sequential video, they may have an unpleasant perceptual experience if the video has an inconsistent visual quality even though the average visual quality of the video is not compromised. Thus, consistent visual quality control is mostly expected in general video encoding with limited channel bandwidth and buffer resources. However, there still has not been enough study on such an issue. In this paper, a new objective visual quality metric (VQM) is proposed first, which can easily be incorporated into video coding for guiding video coding. Second, a VQM-based window model is proposed to handle the tradeoff between visual quality consistency and buffer constraint in video coding. Third, a window-level rate control algorithm is developed to accomplish visual quality control based on the above two proposals. Finally, experimental results prove that consistent visual quality, high rate-distortion efficiency, accurate bit control, and compliant buffer constraint can be achieved by the proposed rate control algorithm. Long Xu 0001, Songnan Li, King Ngi Ngan, Lin Ma 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2013 | Regional Bit Allocation and Rate Distortion Optimization for Multiview Depth Video Coding With View Synthesis Distortion ModelabstractIn this paper, we propose a view synthesis distortion model (VSDM) that establishes the relationship between depth distortion and view synthesis distortion for the regions with different characteristics: color texture area corresponding depth (CTAD) region and color smooth area corresponding depth (CSAD), respectively. With this VSDM, we propose regional bit allocation (RBA) and rate distortion optimization (RDO) algorithms for multiview depth video coding (MDVC) by allocating more bits on CTAD for rendering quality and fewer bits on CSAD for compression efficiency. Experimental results show that the proposed VSDM based RBA and RDO can improve the coding efficiency significantly for the test sequences. In addition, for the proposed overall MDVC algorithm that integrates VSDM based RBA and RDO, it achieves 9.99% and 14.51% bit rate reduction on average for the high and low bit rate, respectively. It can improve virtual view image quality 0.22 and 0.24 dB on average at the high and low bit rate, respectively, when compared with the original joint multiview video coding model. The RD performance comparisons using five different metrics also validate the effectiveness of the proposed overall algorithm. In addition, the proposed algorithms can be applied to both INTRA and INTER frames. Yun Zhang 0002, Sam Kwong, Long Xu 0001, Sudeng Hu, Gangyi Jiang, C.-C. Jay Kuo |
IEEE Trans. Image Process. | 3 |
| 2012 | Spatial-temporal decorrelation for image/video codingabstractModern image/video compression techniques greatly help to store and transmit digital images and video data. Discrete wavelet transform is used in JPEG2000 because of its scalability and tolerable degradation. In H.264/AVC, predictive coding is employed to remove spatial redundancy before discrete cosine transform. In this paper, we propose a new encoder structure combining both advantages from JPEG2000 [1] and H.264/AVC [2], which firstly utilizes the proposed spatial transform to decompose a single image into several sub-images and then employs motion compensation to convert the conventional spatial decorrelation into temporal decorrelation. Experimental results show that our proposed method outperforms the state-of-the-art image coding algorithms and achieves better rate-distortion performance. Miaohui Wang, King Ngi Ngan, Long Xu 0001 |
PCS | 3 |
| 2012 | Video content dependent directional transform for intra frame codingabstractThe mode-dependent directional transform (MDDT) employed Karhunen-Loève Transform (KLT) for compressing directional residue signal of intra prediction along its direction. The transform bases were derived from the singular value decomposition (SVD) of residue signals coming from all kinds of video sequences, which were expected to be efficient for most of video sequences. However, the advantage of KLT comes from the concept of a “signal content dependent transform”. MDDT and its variants failed to exploit such a concept, so they did not fully exploit the efficiency of KLT. In this paper, a video content feature is firstly defined as the histogram of the residue produced by intra prediction. Secondly, one KLT basis is computed for each feature of each mode from off-line experiments. Thus, multiple KLT bases identified by their features are provided to each mode instead of only one basis in MDDT. One of them is selected during encoding process by matching the feature of signal being processed to the predefined features. The experiments show that the average improvement of 0.17dB PSNR and 2.23% bits saving can be achieved by the proposed video content dependent directional transform (CDDT) comparing to the state-of-the-art MDDT. Long Xu 0001, King Ngi Ngan, Miaohui Wang |
PCS | 1 |
| 2012 | A Universal Rate Control Scheme for Video TranscodingabstractVideo transcoding is proposed for the bitrate adaption, spatial and/or temporal resolutions adaption, and video format conversion. In video streaming application, it converts videos at server to the compatible versions demanded by networks or clients' devices, so that the videos can be delivered over networks and displayed in the clients' devices successfully. This paper provides a universal rate control scheme for various video transcoding purposes. First, a new rate-distortion (R-D) model is established theoretically for better representing the real R-D feature of transcoding. Second, a window-level rate control algorithm is proposed for providing smooth visual quality with compliant buffer constraint by utilizing the two-pass R-D model and a new proposed sliding window buffer control strategy. Finally, a universal rate control scheme for transcoding is developed based on the established R-D model and the proposed window-level rate control algorithm. The extensive experimental results demonstrate that as compared to other state-of-the-art rate control algorithms for transcoding, the proposed scheme can achieve more bit control accuracy with the average mismatch below 0.2%, and much more consistent visual quality with 0.1 dB-0.3 dB peak-to-signal noise ratio improvement in average, while with low computational complexity. Long Xu 0001, Sam Kwong, Hanli Wang, Yun Zhang 0002, Debin Zhao, Wen Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2011 | Hypothesis comparison guided cross validation for unsupervised signer adaptationabstractSigner adaptation is important to sign language recognition systems in that a one-size-fits-all model set can not perform well on all kinds of signers. Supervised signer adaptation must utilize the labeled adaptation data that are collected explicitly. To skip the data collecting process in signer adaptation, we propose an unsupervised adaptation method called hypothesis comparison guided cross validation (HC CV) algorithm. The algorithm not only addresses the problem of overlap between the data set to be labeled and the data set for adaptation, but also employs an additional hypothesis comparison step to decrease the noise rate of the adaptation data set. Experimental results show that the HC CV adaptation algorithm is superior to the CV adaptation algorithm and the conventional self-teaching algorithm. Though the algorithm is proposed for signer adaptation, it can also be applied to speaker adaptation and writer adaptation straightforwardly. Yu Zhou 0015, Xiaokang Yang 0001, Weiyao Lin, Yi Xu 0001, Long Xu 0001 |
ICME | 5 |
| 2011 | Priority pyramid based bit allocation for multiview video codingabstractIn Multivew Video Coding (MVC), Hierarchial B Pictures (HBP) structure is adopted to remove the redundancy of multiview videos. It makes the rate control of MVC more difficult with respect to accurate bit control and good compression efficiency. The existing rate control algorithms of MVC are based on those of monoview video coding standards, which didn't exploit the correlations of multiview video frames fully. In this paper, a Group of pictures from multiple views (called GoGOP) are rearranged in the form of a pyramid. The top layer of the pyramid containing anchor frames receives the highest priority in bits consuming. And the next layer is inferior to it but superior to others in bits consuming. Secondly, a matrix of weighting factors is further introduced to perform the bit allocation for MVC. Thirdly, the parameter updating is handled in a vector operation, which is somewhat robust and has a low level of computational complexity. The experimental results demonstrate that our proposed bit allocation cooperating with the conventional rate-distortion (R-D) model in H.264/AVC is efficient in rate control of MVC. The coding performance of the proposed algorithm is comparable to that of hierarchical quantization scheme (HQS) of MVC. Meanwhile, a small bit control error is obtained by our algorithm. Long Xu 0001, Sam Kwong, Tiesong Zhao, Yu Zhou 0015 |
VCIP | 1 |
| 2011 | Window-Level Rate Control for Smooth Picture Quality and Smooth Buffer OccupancyabstractIn rate control, smooth picture quality and smooth buffer occupancy are both important but contrary to each other at a given bit rate. How to get a good tradeoff between them was not devoted much attention previously. To deal with this problem, a theoretical window model is proposed in this paper, in which several adjacent frames grouped as a window are considered together. The smoothness of both picture quality and buffer occupancy can be gracefully achieved by regulating the size of the window. To illustrate the usage of window model, a window-level rate control algorithm cooperated with the traditional ρ-domain rate-distortion model is further introduced. In experiments, we first show how the proposed window model achieves the tradeoff between picture quality smoothness and buffer smoothness, and then demonstrate the significant PSNR improvement, accuracy of bit control and consistency of visual quality of the proposed window-level rate control algorithm. Long Xu 0001, Debin Zhao, Xiangyang Ji, Lei Deng 0007, Sam Kwong, Wen Gao 0001 |
IEEE Trans. Image Process. | 1 |
| 2009 | Estimating the value of θ in the intra frame for ρ-domain rate control algorithmsabstractThe rho-domain rate control algorithm is very simple and effective. There is only one parameter, thetas, used in the model. thetas reflects the relationship between the rate and the quantization step. Accurately estimating the thetas value of the I (intra) frame is important for an efficient rho-domain rate control algorithm because the I frame will influence the remaining frames in the GOP (group of pictures) greatly. In this paper, we propose to estimate the thetas value based on the energy of the frame. The energy difference of the I frames is used to predict the new thetas. A linear relationship between thetas and the energy is deduced. Experimental results show that using the energy can more accurately predict the thetas value of I frame than only using the thetas value of the previous P (predict) frame. Zhihang Wang, Shengfu Dong, Long Xu 0001, Wen Gao 0001, Qingming Huang |
PCS | 3 |
| 2009 | Window-level rate control for smooth picture quality and smooth buffer occupancyabstractTraditionally, rate control consists of bit allocation and QP decision on R-QP model. In bit allocation, target bits is further confined if buffer overflows. Meanwhile, rate control should also give as smooth as possible picture quality. However, there is no explicit relationship between picture quality and encoding parameters, so the coding result on picture quality is usually unpredictable and uncontrollable. On the condition of smooth picture quality, the smooth buffer occupancy is preferable certainly. In our work, we first proposed a "window model" formulating the size of window and variations of picture quality and buffer occupancy. Thus, given the constraint on picture quality and buffer occupancy, the compliant coding result about them can be expected employing window model. Second, a window-level rate distortion (RD) model inspired by the traditional rho-domain model is introduced. Lastly, the evaluation of our proposal is presented with elaborate experiments. Long Xu 0001, Zhihang Wang, Lei Deng 0007, Xiangyang Ji, Debin Zhao, Wen Gao 0001 |
PCS | 1 |
| 2007 | Rate Control for Hierarchical B-picture Coding with Scaling-factorsabstractThe coding performance can be further improved when the hierarchical B-picture coding is introduced into H.264/AVC. However, the existing rate control schemes can not work efficiently in such new coding framework. This paper proposes a novel rate control algorithm when hierarchical B-picture coding is used in H.264/AVC. Firstly, a set of scaling-factors applied in designing cascaded quantizer for the B frames at different temporal levels is introduced. Based on the designed scaling-factors, an efficient bit-allocation strategy for hierarchical B-picture coding is presented. The experiments show that the proposed rate control algorithm can further improve PSNR up to 0.7dB compared to the existing hierarchical B-picture coding in H.264/AVC, while the mismatch of target bit rate and real bit rate does not exceed 2%. Long Xu 0001, Wen Gao 0001, Xiangyang Ji, Debin Zhao |
ISCAS | 1 |