VLDB 2026 Research / reviewers in the wild / expert
Xu Wang 0006
dblp:w/XuWang6
· DBLP profile ↗
97ranked-venue papers
16as first author
54since 2021 · last 2026
0000-0002-2948-6468ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 73 · 13 first-author · 41 since 2021Artificial intelligence and machine learning · 20 · 2 first-author · 13 since 2021Databases, data management, data science and information retrieval · 7 · 3 since 2021Computer networks · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-authorTheory of computation · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DBGroup: Dual-Branch Point Grouping for Weakly Supervised 3D Semantic Instance SegmentationabstractWeakly supervised 3D instance segmentation is essential for 3D scene understanding, especially as the growing scale of data and high annotation costs associated with fully supervised approaches. Existing methods primarily rely on two forms of weak supervision: one-thing-one-click annotations and bounding box annotations, both of which aim to reduce labeling efforts. However, these approaches still encounter limitations, including labor-intensive annotation processes, high complexity, and reliance on expert annotators. To address these challenges, we propose DBGroup, a two-stage weakly supervised 3D instance segmentation framework that leverages scene-level annotations as a more efficient and scalable alternative. In the first stage, we introduce a Dual-Branch Point Grouping module to generate pseudo labels guided by semantic and mask cues extracted from multi-view images. To further improve label quality, we develop two refinement strategies: Granularity-Aware Instance Merging and Semantic Selection and Propagation. The second stage involves multi-round self-training on an end-to-end instance segmentation network using the refined pseudo-labels. Additionally, we introduce an Instance Mask Filter strategy to address inconsistencies within the pseudo labels. Extensive experiments demonstrate that DBGroup achieves competitive performance compared to sparse-point-level supervised 3D instance segmentation methods, while surpassing state-of-the-art scene-level supervised 3D semantic segmentation approaches. Xuexun Liu, Xiaoxu Xu, Qiudan Zhang, Lin Ma 0002, Xu Wang 0006 |
AAAI | 5 |
| 2026 | Optimal transport-guided multivariable model for point set matching problems
Litao Ma, Xu Wang 0006, Jiqiang Chen |
Signal Process. | 2 |
| 2026 | Asymptotic Insights Into Outage Probability of Multi-Cascaded RISs Over Doubly-Correlated MIMO Fading ChannelsabstractCascaded reconfigurable intelligent surfaces (RISs) greatly improve network coverage and reliability. This paper examines the outage probability (OP) of multi-cascaded RISs (MCRISs)-aided systems over doubly-correlated Rayleigh multiple-input multiple-output (MIMO) channels. The moment-generating function is invoked to convert the OP into a numerical inversion of Laplace integral. To capture profound insights, we conduct an in-depth asymptotic investigation into the outage behavior of MCRIS-aided MIMO communications by leveraging random matrix theory in the high-SNR regime. The asymptotic results reveal that the firsthRISs with the smallest number of reflective elements primarily dictate the bottleneck of reliability performance, wherehis influenced by the variation in the number of reflective elements across the RISs. Specifically, smaller variations result in a largerh, while greater disparities reduce it. Additionally, we identify an “unsaturation effect” that occurs when the performance margin, or residual spatial degree of freedom (DoF), after propagation through the firsthRISs cannot be evenly distributed between the transmitter and the RISs. This effect slows down the decline of the OP with increasing SNR. More cascaded RISs impair the spatial DoF of wireless communications, resulting in the loss of half of the independent fading channels compared to the system without the support of RIS as the cascaded number of RISs increases. Majorization theory is subsequently applied to unveil the negative effect of doubly-spatial correlation on system reliability. Finally, Monte Carlo simulations are carried out for validations. Jintao Wang 0002, Zheng Shi 0001, Xu Wang 0006, Yaru Fu, Guanghua Yang, Shaodan Ma |
IEEE Trans. Commun. | 4 |
| 2026 | Deep Learning-Based Joint Geometry and Attribute Up-Sampling for Large-Scale Colored Point CloudsabstractColored point cloud comprising geometry and attribute components is one of the mainstream representations enabling realistic and immersive 3D applications. To generate large-scale and denser colored point clouds, we propose a deep learning-based Joint Geometry and Attribute Up-sampling (JGAU) method, which learns to model both geometry and attribute patterns and leverages the spatial attribute correlation. Firstly, we establish and release a large-scale dataset for colored point cloud up-sampling, named SYSU-PCUD, which has 121 large-scale colored point clouds with diverse geometry and attribute complexities in six categories and four sampling rates. Secondly, to improve the quality of up-sampled point clouds, we propose a deep learning-based JGAU framework to up-sample the geometry and attribute jointly. It consists of a geometry up-sampling network and an attribute up-sampling network, where the latter leverages the up-sampled auxiliary geometry to model neighborhood correlations of the attributes. Thirdly, we propose two coarse attribute up-sampling methods, Geometric Distance Weighted Attribute Interpolation (GDWAI) and Deep Learning-based Attribute Interpolation (DLAI), to generate coarsely up-sampled attributes for each point. Then, we propose an attribute enhancement module to refine the up-sampled attributes and generate high quality point clouds by further exploiting intrinsic attribute and geometry patterns. Extensive experiments show that Peak Signal-to-Noise Ratio (PSNR) achieved by the proposed JGAU are 33.90 dB, 32.10 dB, 31.10 dB, and 30.39 dB when up-sampling rates are $4\times $ , $8\times $ , $12\times $ , and $16\times $ , respectively. Compared to the state-of-the-art schemes, the JGAU achieves an average of 2.32 dB, 2.47 dB, 2.28 dB and 2.11 dB PSNR gains at four up-sampling rates, respectively, which are significant. The code is released with https://github.com/SYSU-Video/JGAU. Yun Zhang 0002, Feifan Chen, Na Li 0015, Xu Wang 0006, Fen Miao, Sam Kwong |
IEEE Trans. Image Process. | 5 |
| 2026 | Rate-Reconfigurable Deep Point Cloud Compression With Perceptual Bit Allocation OptimizationabstractConventional end-to-end learning-based point cloud compression requires training multiple models to adapt to different target bit rates. Moreover, the rate difference between geometry and attribute components of point clouds is not well-considered. In this paper, we propose an end-to-end Rate-Reconfigurable Deep Point Cloud Compression (RR-DPCC) with on/off-line Perceptual Bit Allocation Optimization (PBAO-ON/OFF), which achieves arbitrary bit rate control with one trained deep model and high efficiency joint geometry and attribute coding. First, we propose the framework of the RR-DPCC using PBAO-ON/OFF, which includes Point Cloud Quality Assessment (PCQA) for perceptual quality measurement, PBAO-ON/OFF modules for bit allocation and RR-DPCC for high efficiency point cloud coding. Second, we propose a one-stream network of the RR-DPCC to encode the attribute and geometry of point clouds jointly. Moreover, in RR-DPCC, a bitrate reconfigurable module is proposed to encode multiple fine-grained bitrate points with one trained model and a rate allocation module is proposed to allocate bits between geometry and attribute. Third, we propose on/off-line PBAO algorithms to maximize the perceptual quality of the reconstructed point cloud, where the bits are properly allocated based on the importance of geometry and attribute. Meanwhile, rate-distortion models (R- $\alpha $ / $\beta $ and D- $\alpha $ / $\beta $ ) are derived for high accuracy rate control and bit allocation. Experimental results show that the proposed RR-DPCC achieves fine-grained bitrate control and allocation through a single trained model. When combined the proposed RR-DPCC with PBAO-ON, it reduces -6.56% and -18.68% bit rate on average as comparing with the state-of-the-art V-PCC and Deep Joint Geometry and Attribute Compression (Deep-JGAC), respectively. When combined with the PBAO-OFF, it achieves -4.90% and -15.34% bit rate reductions on average, and reduces 98.38%/22.05% and 53.75%/10.04% encoding/decoding time on average with respect to V-PCC and Deep-JGAC. Yun Zhang 0002, Lewen Fan, Zixi Guo, Xu Wang 0006, Xiaoxia Huang 0004, Sam Kwong |
IEEE Trans. Image Process. | 4 |
| 2026 | A Visual-Linguistic Approach for Robust RGB-Thermal Tracking With Dynamic Template AdaptationabstractRGB-T tracking seeks to improve tracking robust ness in complex environments by exploiting the complementary advantages of RGB and thermal infrared (TIR) modalities. Nevertheless, current RGB-T tracking methods encounter two critical limitations stemming from their reliance on fixed template images for search region matching. First, the resemblance between the initial fixed template and the target in later frames diminishes as time goes on, causing tracking reliability to decline. Second, background clutter within template bounding boxes introduces disruptive noise that compromises tracking accuracy. To address these challenges, we propose a new referring RGB-T tracking approach that integrates visual-linguistic cues to enhance multimodal data, enabling background noise elimination in the template and dynamic template image updating through adaptive mechanisms. Additionally, to facilitate research in language guided multimodal tracking, we construct a large-scale dataset named Refer-RGBT, which comprises 1508 multimodal video pairs with synchronized RGB/TIR sequences of diverse objects. Each sequence is annotated by descriptive textual captions that detail their appearance and actions. Extensive evaluations on the Refer-RGBT dataset demonstrate that our proposed approach achieves cutting-edge performance compared to the state-of-the art (SOTA) methods. Xu Wang 0006, Huanxin Zheng, Haohong Liao, Qiudan Zhang, Lin Ma 0002, Jianmin Jiang |
IEEE Trans. Multim. | 1 |
| 2025 | CWC-DNERF: Compact Dynamic Neural Radiance Field VIA Discrete Wavelet Transform And Learnable CodebooksabstractNeural radiance fields have significantly advanced dynamic scene reconstruction and novel view synthesis. However, relying on multiple implicit multi-layer perceptrons for reconstructing dynamic scenes is computationally expensive. Recent methods have alleviated this challenge by introducing explicit data structures, such as voxel grids and feature planes, but these significantly increase storage demands and complicate network transmission. We propose Cwc-DNeRF, a compact dynamic NeRF representation that leverages discrete wavelet transform (DWT) and learnable codebooks to achieve superior storage efficiency while maintaining competitive rendering quality compared to K-Planes. In Stage I, DWT and trainable masks are employed to optimize parameter efficiency, resulting in sparse space planes. In Stage II, learnable codebooks are introduced for the space-time planes to merge redundant spatio-temporal features further reducing the storage demand. Additionally, a data compression pipeline is applied to compress both sparse space plane parameters and codebooks. Experimental results on D-NeRF and DyNeRF datasets show that our method achieves state-of-the-art rendering quality within a 10MB storage budget while retaining the benefits of explicit feature planes. Yaojian Xu, Qiudan Zhang, Longhao Zou, Qiong Liu 0001, Xu Wang 0006 |
ICIP | 6 |
| 2025 | Tech-ASan: Two-stage check for Address SanitizerabstractAddress Sanitizer (ASan) is a sharp weapon for detecting memory safety violations, including temporal and spatial errors hidden in C/C++ programs during execution.However, ASan incurs significant runtime overhead, which limits its efficiency in testing large software.The overhead mainly comes from sanitizer checks due to the frequent and expensive shadow memory access.Over the past decade, many methods have been developed to speed up ASan by eliminating and accelerating sanitizer checks, however, they either fail to adequately eliminate redundant checks or compromise detection capabilities.To address this issue, this paper presents Tech-ASan, a two-stage check based technique to accelerate ASan with safety assurance.First, we propose a novel two-stage check algorithm for ASan, which leverages magic value comparison to reduce most of the costly shadow memory accesses.Second, we design an efficient optimizer to eliminate redundant checks, which integrates a novel algorithm for removing checks in loops.Third, we implement Tech-ASan as a memory safety tool based on the LLVM compiler infrastructure.Our evaluation using the SPEC CPU2006 benchmark shows that Tech-ASan outperforms the state-of-theart methods with 33.70% and 17.89% less runtime overhead than ASan and ASan--, respectively.Moreover, Tech-ASan detects 56 fewer false negative cases than ASan and ASan--when testing on the Juliet Test Suite under the same redzone setting. Yixuan Cao 0002, Yuhong Feng, Chongyi Huang, Fangcao Jian, Xu Wang 0006 |
Internetware | 7 |
| 2025 | Q-Doc: Benchmarking Document Image Quality Assessment Capabilities in Multi-modal Large Language Models
Jiaxi Huang, Dongxu Wu, Hanwei Zhu, Lingyu Zhu 0006, Jun Xing, Xu Wang 0006, Baoliang Chen |
PRCV (8) | 6 |
| 2025 | Interpretable Optimization-Inspired Unfolding Network for Low-Light Image EnhancementabstractRetinex model-based methods have shown to be effective in layer-wise manipulation with well-designed priors for low-light image enhancement (LLIE). However, the hand-crafted priors and conventional optimization algorithm adopted to solve the layer decomposition problem result in the lack of adaptivity and efficiency. To this end, this paper proposes a Retinex-based deep unfolding network (URetinex-Net++), which unfolds an optimization problem into a learnable network to decompose a low-light image into reflectance and illumination layers. By formulating the decomposition problem as an implicit priors regularized model, three learning-based modules are carefully designed, responsible for data-dependent initialization, high-efficient unfolding optimization, and fairly-flexible component adjustment, respectively. Particularly, the proposed unfolding optimization module, introducing two networks to adaptively fit implicit priors in the data-driven manner, can realize noise suppression and details preservation for decomposed components. URetinex-Net++ is a further augmented version of URetinex-Net, which introduces a cross-stage fusion block to alleviate the color defect in URetinex-Net. Therefore, boosted performance on LLIE can be obtained in both visual quality and quantitative metrics, where only a few parameters are introduced and little time is cost. Extensive experiments on real-world low-light images qualitatively and quantitatively demonstrate the effectiveness and superiority of the proposed URetinex-Net++ over state-of-the-art methods. Wenhui Wu 0001, Jian Weng 0009, Xu Wang 0006, Wenhan Yang, Jianmin Jiang |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | Learning-Based Compression for Noisy Images in the WildabstractDigital images in real world applications typically undergo a wide variety of quality degradations before compression or re-compression. Existing learning based codecs are typically data-driven, relying on the predefined compression pipeline with pristine or high quality images as the input. However, the images in the wild may exhibit the substantially different characteristics compared to the high quality images, casting major challenges to the learning based image coding. In this paper, we propose a robust noisy image compression framework with the blind assumption on the specific noise type and level. The specifically designed encoder decomposes the representation of visual content into two types of features, including the Features that represent the Intrinsic Content (FIC) and the Features that account for Additive Degradation (FAD). As such, beyond the philosophy of faithfully reconstructing the given image with high fidelity, only FIC needs to be compactly represented and conveyed. The principled disentanglement strategy facilitates the removal of the redundancy from multiple perspectives (e.g., spatial, channel and content), ensuring the handling of a wide variety of noisy images in the wild. Extensive experimental results show that our model can achieve superior performance in terms of the ultimate quality and exhibit the strong generalizability across images degraded by a variety of means. The proposed scheme also points out a new research avenue on learning based compression for images in the wild, which is technically challenging but desirable in practice. Meng Wang 0017, Baoliang Chen, Rongqun Lin, Xu Wang 0006, Shiqi Wang 0001, Sam Kwong |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Geometry-Aware Self-Supervised Indoor 360$^{\circ }$ Depth Estimation via Asymmetric Dual-Domain Collaborative LearningabstractBeing able to estimate monocular depth for spherical panoramas is of fundamental importance in 3D scene perception. However, spherical distortion severely limits the effectiveness of vanilla convolutions. To push the envelope of accuracy, recent approaches attempt to utilize Tangent projection (TP) to estimate the depth of$360 ^{\circ }$images. Yet, these methods still suffer from discrepancies and inconsistencies among patch-wise tangent images, as well as the lack of accurate ground truth depth maps under a supervised fashion. In this paper, we propose a geometry-aware self-supervised$360 ^{\circ }$image depth estimation methodology that explores the complementary advantages of TP and Equirectangular projection (ERP) by an asymmetric dual-domain collaborative learning strategy. Especially, we first develop a lightweight asymmetric dual-domain depth estimation network, which enables to aggregate depth-related features from a single TP domain, and then produce depth distributions of the TP and ERP domains via collaborative learning. This effectively mitigates stitching artifacts and preserves fine details in depth inference without overspending model parameters. In addition, a frequent-spatial feature concentration module is devised to simultaneously capture non-local Fourier features and local spatial features, such that facilitating the efficient exploration of monocular depth cues. Moreover, we introduce a geometric structural alignment module to further improve geometric structural consistency among tangent images. Extensive experiments illustrate that our designed approach outperforms existing self-supervised$360 ^{\circ }$depth estimation methods on three publicly available benchmark datasets. Xu Wang 0006, Ziyan He, Qiudan Zhang, You Yang 0002, Tiesong Zhao, Jianmin Jiang |
IEEE Trans. Multim. | 1 |
| 2025 | Weakly-Supervised 3D Visual Grounding Based on Visual Language AlignmentabstractLearning to ground natural language queries to target objects or regions in 3D point clouds is quite essential for 3D scene understanding. Nevertheless, existing 3D visual grounding approaches require a substantial number of bounding box annotations for text queries, which is time-consuming and labor-intensive to obtain. In this paper, we propose3D-VLA, a weakly supervised approach for3Dvisual grounding based onVisualLanguageAlignment. Our 3D-VLA exploits the superior ability of current large-scale vision-language models (VLMs) on aligning the semantics between texts and 2D images, as well as the naturally existing correspondences between 2D images and 3D point clouds, and thus implicitly constructs correspondences between texts and 3D point clouds with no need for fine-grained box annotations in the training procedure. During the inference stage, the learned text-3D correspondence will help us ground the text queries to the 3D target objects even without 2D images. To the best of our knowledge, this is the first work to investigate 3D visual grounding in a weakly supervised manner by involving large scale vision-language models, and extensive experiments on ReferIt3D and ScanRefer datasets demonstrate that our 3D-VLA achieves comparable and even superior results over the fully supervised methods. Xiaoxu Xu, Yitian Yuan, Qiudan Zhang, Wenhui Wu 0001, Zequn Jie, Lin Ma 0002, Xu Wang 0006 |
IEEE Trans. Multim. | 7 |
| 2025 | Hierarchical Uncertainty-Aware Salient Object Detection for $360 ^{\circ }$ Images via Bi-Projection Collaborative Learningabstract$360^{\circ }$salient object detection has recently received much attention for 3D scene perception owing to its omnidirectional field of view (FoV). The capability of recognizing salient objects of$360^{\circ }$images remains technically challenging due to severe spherical distortion. In this paper, we develop a hierarchical uncertainty-aware$360^{\circ }$image salient object detection methodology that explicitly explores the geometric and spatial complementary coherence of Tangent projection (TP) and Equirectangular projection (ERP) by a collaborative learning strategy. Concretely, to mitigate spherical distortion, we first intend to learn saliency-related features from less-distorted tangent images, in which a deformation-aware attention block is introduced to mitigate the geometric distortion caused by projecting a$360^{\circ }$image onto a 2D plane. However, the discrepancies among tangent images pose a new challenge to$360^{\circ }$image salient object detection. To tackle this issue and achieve accurate localization for salient objects of all sizes, we design a spatial-frequency saliency feature aggregation module to leverage fast Fourier convolution to capture global contextual information from ERP images, such that obtaining more representative saliency features. Moreover, a hierarchical uncertainty-aware bi-projection consistency learning module with strong local-global information embedding capabilities is constructed, which learns the geometric and spatial correlations between tangent images and ERP images via a collaborative learning strategy. Ultimately, salient object maps are produced for$360^{\circ }$images on the basis of the merged saliency features driven by the uncertainty. Extensive experiments show that our developed method improves${\mathrm{F}}_\beta ^{\sigma }$by an average of 31.67% compared to twenty existing advanced methods on the publicly available 360-SOD dataset. Qiudan Zhang, Kaiyu Ji, Xu Wang 0006, Zhaoqing Pan, Jianmin Jiang |
IEEE Trans. Multim. | 4 |
| 2025 | HNR-ISC: Hybrid Neural Representation for Image Set CompressionabstractImage set compression (ISC) refers to compressing the sets of semantically similar images. Traditional ISC methods typically aim to eliminate redundancy among images at either signal or frequency domain, but often struggle to handle complex geometric deformations across different images effectively. Here, we propose a new Hybrid Neural Representation for ISC (HNR-ISC), including an implicit neural representation for Semantically Common content Compression (SCC) and an explicit neural representation for Semantically Unique content Compression (SUC). Specifically, SCC enables the conversion of semantically common contents into a small-and-sweet neural representation, along with embeddings that can be conveyed as a bitstream. SUC is composed of invertible modules for removing intra-image redundancies. The feature level combination from SCC and SUC naturally forms the final image set. Experimental results demonstrate the robustness and generalization capability of HNR-ISC in terms of signal and perceptual quality for reconstruction and accuracy for the downstream analysis task. Shiqi Wang 0001, Meng Wang 0017, Peilin Chen 0001, Wenhui Wu 0001, Xu Wang 0006, Sam Kwong |
IEEE Trans. Multim. | 6 |
| 2025 | RGB-D Data Compression via Bi-Directional Cross-Modal Prior Transfer and Enhanced Entropy ModelingabstractRGB-D data, being homogeneous cross-modal data, demonstrates significant correlations among data elements. However, current research focuses only on a uni-directional pattern of cross-modal contextual information, neglecting the exploration of bi-directional relationships in the compression field. Thus, we propose a joint RGB-D compression scheme, which is combined with Bi-Directional Cross-Modal Prior Transfer (Bi-CPT) modules and a Bi-Directional Cross-Modal Enhanced Entropy (Bi-CEE) model. The Bi-CPT module is designed for compact representations of cross-modal features, effectively eliminating spatial and modality redundancies at different granularity levels. In contrast to the traditional entropy models, our proposed Bi-CEE model not only achieves spatial-channel contextual adaptation through partitioning RGB and depth features but also incorporates information from other modalities as prior to enhance the accuracy of probability estimation for latent variables. Furthermore, this model enables parallel multi-stage processing to accelerate coding. Experimental results demonstrate the superiority of our proposed framework over the current compression scheme, outperforming both rate-distortion performance and downstream tasks, including surface reconstruction and semantic segmentation. The source code will be available at https://github.com/xyy7/Learning-based-RGB-D-Image-Compression . Yuyu Xu, Qiudan Zhang, Wenhui Wu 0001, Yun Zhang 0002, Xu Wang 0006 |
ACM Trans. Multim. Comput. Commun. Appl. | 7 |
| 2024 | 3D Weakly Supervised Semantic Segmentation with 2D Vision-Language Guidance
Xiaoxu Xu, Yitian Yuan, Jinlong Li 0003, Qiudan Zhang, Zequn Jie, Lin Ma 0002, Hao Tang 0005, Nicu Sebe, Xu Wang 0006 |
ECCV (73) | 9 |
| 2024 | Image Coding for Analytics via Adversarially Augmented AdaptationabstractImage Coding for Machine (ICM) aims to compress an image so that the reconstructed one can meet the requirements of both human vision and machine vision. Existing methods apply the constraint from the downstream models to improve machine analytics performance while compromising the visual quality. This paper proposes a novel adversarially augmented adaptation route that achieves a better trade-off between the utility of the human and machine perspectives by making slight changes to the image manifold. In detail, a targeted adversarial attack is employed to generate subtle image perturbations that are nearly imperceptible to humans but significantly improve machine analytic performance. These perturbed images would be subsequently employed as ground truth to guide training/fine-tuning of an end-to-end image compression network. Note that, our method is a plug-and-play framework that does not rely on any change in existing architecture or loss functions. Extensive experimental results demonstrate the superiority of the proposed scheme over conventional ICM frameworks and the effectiveness of our design. Xuelin Shen, Kangsheng Yin, Xu Wang 0006, Yu-Lin He, Shiqi Wang 0001, Wenhan Yang |
ICASSP | 3 |
| 2024 | Sliced Maximal Information Coefficient: A Training-Free Approach for Image Quality Assessment EnhancementabstractFull-reference image quality assessment (FR-IQA) models generally operate by measuring the visual differences between a degraded image and its reference. However, existing FR-IQA models including both the classical ones (e.g., PSNR and SSIM) and deep-learning based measures (e.g., LPIPS and DISTS) still exhibit limitations in capturing the full perception characteristics of the human visual system (HVS). In this paper, instead of designing a new FR-IQA measure, we aim to explore a generalized human visual attention estimation strategy to mimic the process of human quality rating and enhance existing IQA models. In particular, we model human attention generation by measuring the statistical dependency between the degraded image and the reference image. The dependency is captured in a training-free manner by our proposed sliced maximal information coefficient and exhibits surprising generalization in different IQA measures. Experimental results verify the performance of existing IQA models can be consistently improved when our attention module is incorporated. The source code is available at https://github.com/KANGX99/SMIC. Kang Xiao, Xu Wang 0006, Yu-Lin He, Baoliang Chen, Xuelin Shen |
ICME | 2 |
| 2024 | LESS: Label-Efficient and Single-Stage Referring 3D SegmentationabstractReferring 3D Segmentation is a visual-language task that segments all points of the specified object from a 3D point cloud described by a sentence of query. Previous works perform a two-stage paradigm, first conducting language-agnostic instance segmentation then matching with given text query. However, the semantic concepts from text query and visual cues are separately interacted during the training, and both instance and semantic labels for each object are required, which is time consuming and human-labor intensive. To mitigate these issues, we propose a novel Referring 3D Segmentation pipeline, Label-Efficient and Single-Stage, dubbed LESS, which is only under the supervision of efficient binary mask. Specifically, we design a Point-Word Cross-Modal Alignment module for aligning the fine-grained features of points and textual embedding. Query Mask Predictor module and Query-Sentence Alignment module are introduced for coarse-grained alignment between masks and query. Furthermore, we propose an area regularization loss, which coarsely reduces irrelevant background predictions on a large scale. Besides, a point-to-point contrastive loss is proposed concentrating on distinguishing points with subtly similar features. Through extensive experiments, we achieve state-of-the-art performance on ScanRefer dataset by surpassing the previous methods about 3.7% mIoU using only binary labels. Code is available at https://github.com/mellody11/LESS. Xuexun Liu, Xiaoxu Xu, Jinlong Li 0003, Qiudan Zhang, Xu Wang 0006, Nicu Sebe, Lin Ma 0002 |
NeurIPS | 5 |
| 2024 | Multi-exposure embeddings for graph learning: Towards high dynamic range image saliency predictionabstractAbstract Identifying saliency in high dynamic range (HDR) images is a fundamentally important issue in HDR imaging, and plays critical roles towards comprehensive scene understanding. Most of existing studies leverage hand‐crafted features for HDR image saliency prediction, lacking the capabilities of fully exploiting the characteristics of HDR image (i.e. wider luminance range and richer colour gamut). Here, systematical studies are carried out on HDR image saliency prediction by proposing a new framework to single out the contributions from multi‐exposure images. Specifically, inspired by the mechanism of HDR imaging, the method first utilizes graph neural networks to model the relations among multi‐exposure images and the tone‐mapped image obtained from an HDR image, enabling more discriminative saliency‐related feature representations. Subsequently, the saliency features driven by global semantic knowledge are aggregated from the tone‐mapped image through enhancing global context‐aware semantic information. Finally, a fusion module is designed to integrate saliency‐oriented feature representations originated from multi‐exposure images and the tone‐mapped image, producing the saliency maps of HDR images. Moreover, a new challenging HDR eye fixation database (HDR‐EYEFix) is created, expecting to further contribute the research on HDR image saliency prediction. Experiment results show that the method obtains superior performance compared to the state‐of‐the‐art methods. Jun Xing, Qiudan Zhang, Xuelin Shen, Xu Wang 0006 |
IET Image Process. | 4 |
| 2024 | Towards 360$^{\circ }$ image compression for machines via modulating pixel significance
Silin Zheng, Xuelin Shen, Qiudan Zhang, Zhuo Chen 0006, Wenhan Yang, Xu Wang 0006 |
Multim. Tools Appl. | 6 |
| 2024 | Image Intrinsic Components Guided Conditional Diffusion Model for Low-Light Image EnhancementabstractThrough formulating the image restoration as a generation problem, the conditional diffusion model has been applied to low-light image enhancement (LIE) to restore the details in dark regions. However, in the previous diffusion model based LIE methods, the conditions used for guiding generation are degraded images, such as low-light image, signal-to-noise ratio map and color map, which suffer from severe degradation and are simply fed into diffusion model by rigidly concatenating with the noise. To avoid using degraded conditions resulting in sub-optimal performance in recovering details and enhancing brightness, we use the image intrinsic components originating from the Retinex model as guidance, whose multi-scale features are flexibly integrated into the diffusion model, and propose a novel conditional diffusion model for LIE. Specifically, the input low-light image is decomposed into reflectance and illumination by a Retinex decomposition module, where two components contain abundant physical property and lighting conditions of the scene. Then, we extract the latent features from two conditions through a component-dependent feature extraction module, which is designed according to the physical property of components. Finally, instead of previous rigid concatenation manner, a well-designed feature fusion mechanism is equipped to adaptively embed generative conditions into diffusion model. Extensive experimental results demonstrate that our method outperforms the state-of-the-art methods, and is capable of effectively restoring the local details while brightening the dark regions. Our codes are available athttps://github.com/Knossosc/ICCDiff. Sicong Kang, Shuaibo Gao, Wenhui Wu 0001, Xu Wang 0006, Shuoyao Wang, Guoping Qiu |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Revisiting All-Zero Block Detection for Versatile Video CodingabstractThe Versatile Video Coding (VVC) standard adopts a series of new coding tools in transform and quantization, including multiple transform selection, low-frequency non-separable transform, and trellis quantization. These new technologies, which bring significant coding gain, create daunting challenges to optimizing the VVC codec. In this work, we propose a new all-zero block (AZB) detection scheme tailored for VVC, with the collaboration of genuine all-zero block (GAZB) and pseudo all-zero block (PAZB) detection. First, to accommodate the multiple transform sizes in VVC, we develop a GAZB detection method that is apt for square and non-square residual blocks. Meanwhile, a theoretical upper bound is derived to locate the last significant coefficient and detect the potential frequency domain GAZB. Subsequently, a method tailored for trellis-coded quantization in VVC is devised for detecting PAZB. Finally, the GAZB and PAZB detection methods are collaboratively employed for AZB detection in VVC. The proposed method is implemented on the VVC codec Versatile Video Encoder (VVenC), and extensive experimental results show that the proposed method achieves promising time savings for test sequences of different resolutions with negligible rate-distortion performance loss. Zhenhao Sun, Meng Wang 0017, Peilin Chen 0001, Xu Wang 0006, Shiqi Wang 0001, Sam Kwong |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | A New Neurodynamics-Based Model for Fuzzy Convex Optimization Problems With Fuzzy Coefficients and General ConstraintsabstractFuzzy convex optimization problems with fuzzy coefficients (FCOPFCs) arise in many applications. Although many neurodynamics-based models have been proposed for solving FCOPFCs, most of them are designed for FCOPFCs with equality or inequality constraints only. However, in many applications, the FCOPFCs often come with both equality and inequality constraints (general constraints, for short), so most of the neurodynamics-based models no longer work in these situations. Therefore, this article aims to construct a new model for FCOPFCs with general constraints to extend the applications of neurodynamics-based models. First of all, based on fuzzy set theory, the original FCOPFCs with general constraints is transformed into a series of interval programming tasks and further transformed into crisp optimization problems with weights. Then, a novel continuous-time neurodynamics-based model with a single-layer structure is established to solve the crisp optimization problem with weights. Further, we discuss the global existence and prove the stability of state solutions. The theoretical results show that the state solutions reach the feasible region within finite time and converge to the optimal solution with the smallest 2-norm. Simulation results completed for three kinds of FCOPFCs show the validity of the approach, and the results in real-world applications demonstrate the excellent performance of the proposed model. Jiqiang Chen, Litao Ma, Witold Pedrycz, Xu Wang 0006 |
IEEE Trans. Fuzzy Syst. | 4 |
| 2024 | Distortion-Aware Self-Supervised Indoor 360$^{\circ }$ Depth Estimation via Hybrid Projection Fusion and Structural RegularitiesabstractOwing to the rapid development of emerging 360$^{\circ }$panoramic imaging techniques, indoor 360$^{\circ }$depth estimation has aroused extensive attention in the community. Due to the lack of available ground truth depth data, it is extremely urgent to model indoor 360$^{\circ }$depth estimation in self-supervised mode. However, self-supervised 360$^{\circ }$depth estimation suffers from two major limitations. One is the distortion and network training problems caused by Equirectangular projection (ERP), and the other is that texture-less regions are quite difficult to back-propagate in self-supervised mode. Hence, to address the above issues, we introduce spherical view synthesis for learning self-supervised 360$^{\circ }$depth estimation. Specifically, to alleviate the ERP-related problems, we first propose a dual-branch distortion-aware network to produce the coarse depth map, including a distortion-aware module and a hybrid projection fusion module. Subsequently, the coarse depth map is utilized for spherical view synthesis, in which a spherically weighted loss function for view reconstruction and depth smoothing is investigated to optimize the projection distribution problem of 360$^{\circ }$images. In addition, two structural regularities of indoor 360$^{\circ }$scenes are devised as two additional supervisory signals to efficiently optimize our self-supervised 360$^{\circ }$depth estimation model, containing the principal-direction normal constraint and the co-planar depth constraint. The principal-direction normal constraint is designed to align the normal of the 360$^{\circ }$image with the direction of the vanishing points. Meanwhile, we employ the co-planar depth constraint to fit the estimated depth of each pixel through its 3D plane. Finally, a depth map is obtained for the 360$^{\circ }$image. Experimental results illustrate that our proposed method achieves superior performance than the current advanced depth estimation methods on four publicly available datasets. Xu Wang 0006, Weifeng Kong, Qiudan Zhang, You Yang 0002, Tiesong Zhao, Jianmin Jiang |
IEEE Trans. Multim. | 1 |
| 2024 | Weakly-Supervised 3D Scene Graph Generation via Visual-Linguistic Assisted Pseudo-LabelingabstractLearning to build 3D scene graphs is essential for real-world perception in a structured and rich fashion. However, previous 3D scene graph generation methods utilize a fully supervised learning manner and require a large amount of entity-level annotation data of objects and relations, which is extremely resource-consuming and tedious to obtain. To tackle this problem, we propose 3D-VLAP, a weakly-supervised 3D scene graph generation method via Visual-Linguistic Assisted Pseudo-labeling. Specifically, our 3D-VLAP exploits the superior ability of current large-scale visual-linguistic models to align the semantics between texts and 2D images, as well as the naturally existing correspondences between 2D images and 3D point clouds, and thus implicitly constructs correspondences between texts and 3D point clouds. First, we establish the positional correspondence from 3D point clouds to 2D images via camera intrinsic and extrinsic parameters, thereby achieving alignment of 3D point clouds and 2D images. Subsequently, a large-scale cross-modal visual-linguistic model is employed to indirectly align 3D instances with the textual category labels of objects by matching 2D images with object category labels. The pseudo labels for objects and relations are then produced for 3D-VLAP model training by calculating the similarity between visual embeddings and textual category embeddings of objects and relations encoded by the visual-linguistic model, respectively. Ultimately, we design an edge self-attention based graph neural network to generate scene graphs of 3D point clouds. Experiments demonstrate that our 3D-VLAP achieves comparable results with current fully supervised methods, meanwhile alleviating the data annotation pressure. Xu Wang 0006, Qiudan Zhang, Wenhui Wu 0001, Mark Junjie Li, Lin Ma 0002, Jianmin Jiang |
IEEE Trans. Multim. | 1 |
| 2024 | 3DTA: No-Reference 3D Point Cloud Quality Assessment With Twin AttentionabstractPoint clouds are rapidly gaining popularity in many practical applications, and point cloud quality assessment (PCQA) is an important research topic that helps us measure and improve the visual experience in applications using point clouds. Research on full-reference (FR) PCQAs has recently made impressive progress, and research on no-reference (NR) PCQAs has also gradually increased. However, the performance of the prior NR PCQA methods still suffers from weak generalization ability and lower accuracy than the FR metrics in general. In this work, we propose a two-stage sampling method that can reasonably represent a whole point cloud, making it possible to efficiently calculate the point cloud quality. For quality prediction, we designed a twin-attention-based transformer PCQA model (3DTA), which uses the data of the two-stage sampling method as input and directly outputs the predicted quality score. Our model is accurate and widely applicable, and it has a simple and flexible structure. Experimental results show that in most cases, the proposed 3DTA model substantially outperforms the benchmark NR methods. The accuracy of the proposed method is competitive even against that of the FR method, which makes 3DTA a strong candidate for the PCQA task, regardless of the reference availability. The code of the proposed model is publicly available athttps://github.com/philox12358/3DTA-PCQA. Linxia Zhu, Xu Wang 0006, Honglei Su, Huan Yang 0001, Hui Yuan 0001, Jari Korhonen |
IEEE Trans. Multim. | 3 |
| 2023 | Hybrid Prior-Based Diminished Reality for Indoor Panoramic Images
Jiashu Liu, Qiudan Zhang, Xuelin Shen, Wenhui Wu 0001, Xu Wang 0006 |
CGI (3) | 5 |
| 2023 | Lossy LiDAR Point Cloud Compression via Cylindrical 3D Convolution NetworksabstractCompared with object-level and human-level point clouds, LiDAR point clouds have larger data scales and are more sparse, posing a challenge for the existing learning-based lossy compression scheme. In this paper, we resolve this issue by transforming the point cloud into a cylindrical coordinate system. In this way, we can better retain points close to the sensor with a high density while extending the receptive field of convolution in areas of low point density. Following cylindrical quantization, an autoencoder is utilized to progressively downsample voxels. The coordinates and latent features are compressed by G-PCC and hyperprior-based entropy encoding respectively. The results demonstrate that our approach performs better than PCGCv2. The visualization results also show that our algorithm can better retain the shape of objects. Ablation studies further prove the efficiency of the cylindrical coordinates. The code is publicly available at https://github.com/AirManH/cylindrical_pcc. Yelang Gao, Xu Wang 0006 |
ICIP | 3 |
| 2023 | DeepSVC: Deep Scalable Video Coding for Both Machine and Human VisionabstractNowadays, end-to-end video coding for both machine and human vision has become an emerging research topic. In complicated systems such as large-scale internet of video things (IoVT), feature streams and video streams can be separately encoded and delivered for machine judgement and human viewing. In this paper, we propose a deep scalable video codec (DeepSVC) to support three-layer scalability from machine to human vision. First, we design a semantic layer that encodes semantic features extracted from the captured video for machine analysis. This layer employs a conditional semantic compression (CSC) method to remove redundancies between semantic features. Second, we design a structure layer that can be combined with semantic layer to predict the captured video at a low quality. This layer effectively estimates video frames based on semantic layer with an interlayer frame prediction (IFP) network. Third, we design a texture layer that can be combined with the above two layers to reconstruct high-quality video signals. This layer also takes advantage of the IFP network to improve its coding efficiency. In large-scale IoVT systems, DeepSVC can deliver semantic layer for regular use and transmit the other layers on demand. Experimental results indicate that the proposed DeepSVC outperforms popular codecs for machine and human vision. Compared with scalable extension of H.265/HEVC (SHVC), the proposed DeepSVC reduces average bit-per-pixel (bpp) by 25.51%/27.63%/59.87% at the same mAP/PSNR/MS-SSIM. Sourcecode is available at: https://github.com/LHB116/DeepSVC. Zhichen Zhang, Jielian Lin, Xu Wang 0006, Tiesong Zhao |
ACM Multimedia | 5 |
| 2023 | ELFIC: A Learning-based Flexible Image Codec with Rate-Distortion-Complexity OptimizationabstractLearning-based image coding has attracted increasing attentions for its higher compression efficiency than reigning image codecs. However, most existing learning-based codecs do not support variable rates with a single encoder; their decoders are also of fixed, high computational complexity. In this paper, we propose an End-to-end, Learning-based and Flexible Image Codec (ELFIC) that supports variable rate and flexible decoding complexity. First, we propose a general image codec with Nonlinear Feature Fusion Transform (NFFT) as nonlinear transforms to improve its Rate-Distortion (RD) performance. Second, we propose an Instance-aware Decoding Complexity Allocation (IDCA) approach, which exploits image contents for a tradeoff between reconstruction quality and computational complexity in the decoding process. Third, we propose an RD-Complexity (RDC) optimization algorithm, which maximizes the image quality under given rate and complexity constraints for the whole framework. Experimental results show that ELFIC achie-ves variable rate, flexible decoding complexity with the state-of-the-art RD performance. It also supports a more efficient decoding process by focusing on image contents. Source codes are available at https://github.com/Zhichen-Zhang/ELFIC-Image-Compression. Zhichen Zhang, Jielian Lin, Xu Wang 0006, Tiesong Zhao |
ACM Multimedia | 5 |
| 2023 | Salient Object Detection on 360° Omnidirectional Image with Bi-Branch Hybrid Projection NetworkabstractWith the advent of panoramic cameras, modeling saliency in 360° omnidirectional images becomes very urgent and challenging. However, severe distortions limit the prediction accuracy of 360° saliency model. In this paper, we devise a bi-branch hybrid projection network (HPNet), which exploits characteristics of equirectangular projection (ERP) and cubic map projection (CMP) formats to predict salient objects in 360° omnidirectional images. Specifically, an ERP image and a CMP image are first fed into a bi-branch network to aggregate the comprehensive features of the omnidirectional image. Subsequently, to explore the coherence among ERP and CMP images, we design a hybrid projection feature fusion module to efficiently combine CMP and ERP features extracted from different layers. Ultimately, a progressive prediction module is developed to refine the features and locate salient objects incrementally, and then produce the final saliency map for the 360° omnidirectional image. Experimental results illustrate that our model is superior to the existing advanced methods in two publicly available datasets. Qiudan Zhang, Xuelin Shen, Xu Wang 0006 |
MMSP | 4 |
| 2023 | Weakly supervised semantic segmentation via self-supervised destruction learning
Jinlong Li 0003, Zequn Jie, Xu Wang 0006, Yu Zhou 0027, Lin Ma 0002, Jianmin Jiang |
Neurocomputing | 3 |
| 2023 | Atmospheric Scattering Model Induced Statistical Characteristics Estimation for Underwater Image RestorationabstractUnderwater images often suffer from color deviation and low contrast due to selective absorption and light scattering, whose degradation is generally described by an Atmospheric Scattering Model (ASM). However, it is challenging to design hand-craft priors to estimate the transmission map and global light within ASM. To avoid the estimation on these two variables, in this paper, we establish a statistical characteristics relationship between underwater and recovered images based on ASM. With this relationship, a novel lightweight model is proposed for efficient Underwater Image Restoration (UIR). Within our proposed model, the UIR problem is disentangled into global restoration and local compensation, for which two modules are developed. Extensive experimental results demonstrate that our proposed method can effectively improve color deviation and low contrast while preserving details, and outperform state-of-the-art methods. Shuaibo Gao, Wenhui Wu 0001, Hua Li 0012, Linwei Zhu, Xu Wang 0006 |
IEEE Signal Process. Lett. | 5 |
| 2023 | λ-Domain VVC Rate Control Based on Nash EquilibriumabstractWith a significant Rate-Distortion (RD) improvement than H.265/HEVC, Versatile Video Coding (VVC) has set a new milestone in lossy video compression. It also incorporates the emerging$\lambda $-domain rate control technique, aiming at a higher visual quality under a fixed bit constraint. However, the challenge remains how to efficiently allocate bits to all frames and Coding Tree Units (CTUs). In this paper, we propose an effective solution by formulating the above task as a Nash equilibrium problem, where all CTUs are treated as players that bargains with each other. By introducing$\lambda $-domain RD models, a constrained optimization is derived with no closed-form solution. We then propose a two-step strategy to address this issue: a Newton method to iteratively calculate an intermediate variable, and a final solution of Nash equilibrium to obtain an approximately optimal$\lambda $. Finally, we utilize the derived$\lambda $to perform an effective CTU-level bit allocation, which is the very first attempt to introduce Nash equilibrium in$\lambda $-domain rate control. Experimental results with Common Test Conditions (CTC) demonstrate the effectiveness and superiority of our method, which outperforms the state-of-the-art CTU-level rate allocation algorithms for VVC. Jielian Lin, Aiping Huang, Tiesong Zhao, Xu Wang 0006, Sam Kwong |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | Rethinking Semantic Image Compression: Scalable Representation With Cross-Modality TransferabstractThis article proposes the scalable cross-modality compression (SCMC) paradigm, in which the image compression problem is further cast into a representation task by hierarchically sketching the image with different modalities. Herein, we adopt the conceptual organization philosophy to model the overwhelmingly complicated visual patterns, based upon the semantic, structure, and signal level representation accounting for different tasks. The SCMC paradigm that incorporates the representation at different granularities supports diverse application scenarios, such as high-level semantic communication and low-level image reconstruction. The decoder, which enables the recovery of the visual information, benefits from the scalable coding based upon the semantic, structure, and signal layers. Qualitative and quantitative results demonstrate that the SCMC can convey accurate semantic and perceptual information of images, especially at low bitrates, and promising rate-distortion performance has been achieved compared to state-of-the-art methods. The code will be available onlinehttps://github.com/ppingzhang/SCMC. Shiqi Wang 0001, Meng Wang 0017, Jiguo Li 0002, Xu Wang 0006, Sam Kwong |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | Robust and Precise Facial Landmark Detection by Self-Calibrated Pose Attention NetworkabstractCurrent fully supervised facial landmark detection methods have progressed rapidly and achieved remarkable performance. However, they still suffer when coping with faces under large poses and heavy occlusions for inaccurate facial shape constraints and insufficient labeled training samples. In this article, we propose a semisupervised framework, that is, a self-calibrated pose attention network (SCPAN) to achieve more robust and precise facial landmark detection in challenging scenarios. To be specific, a boundary-aware landmark intensity (BALI) field is proposed to model more effective facial shape constraints by fusing boundary and landmark intensity field information. Moreover, a self-calibrated pose attention (SCPA) model is designed to provide a self-learned objective function that enforces intermediate supervision without label information by introducing a self-calibrated mechanism and a pose attention mask. We show that by integrating the BALI fields and SCPA model into a novel SCPAN, more facial prior knowledge can be learned and the detection accuracy and robustness of our method for faces with large poses and heavy occlusions have been improved. The experimental results obtained for challenging benchmark datasets demonstrate that our approach outperforms state-of-the-art methods in the literature. Jun Wan 0005, Hui Xi, Jie Zhou 0009, Zhihui Lai 0001, Witold Pedrycz, Xu Wang 0006 |
IEEE Trans. Cybern. | 6 |
| 2023 | Weakly Supervised Semantic Segmentation Via Progressive Patch LearningabstractMost of the existing semantic segmentation approaches with image-level class labels as supervision, highly rely on the initial class activation map (CAM) generated from the standard classification network. In this paper, a novel “Progressive Patch Learning” approach is proposed to improve the local details extraction of the classification, producing the CAM better covering the whole object rather than only the most discriminative regions as in CAMs obtained in conventional classification models. “Patch Learning” destructs the feature maps into patches and independently processes each local patch in parallel before the final aggregation. Such a mechanism enforces the network to find weak information from the scattered discriminative local parts, achieving enhanced local details sensitivity. “Progressive Patch Learning” further extends the feature destruction and patch learning to multi-level granularities in a progressive manner. Cooperating with a multi-stage optimization strategy, such a “Progressive Patch Learning” mechanism implicitly provides the model with the feature extraction ability across different locality-granularities. As an alternative to the implicit multi-granularity progressive fusion approach, we additionally propose an explicit method to simultaneously fuse features from different granularities in a single model, further enhancing the CAM quality on the full object coverage. Our proposed method achieves outstanding performance on the PASCAL VOC 2012 dataset (e.g., with 69.6$\%$mIoU on thetestset), which surpasses most existing weakly supervised semantic segmentation methods. Jinlong Li 0003, Zequn Jie, Xu Wang 0006, Yu Zhou 0027, Xiaolin Wei, Lin Ma 0002 |
IEEE Trans. Multim. | 3 |
| 2022 | URetinex-Net: Retinex-based Deep Unfolding Network for Low-light Image EnhancementabstractRetinex model-based methods have shown to be effective in layer-wise manipulation with well-designed priors for low-light image enhancement. However, the commonly used handcrafted priors and optimization-driven solutions lead to the absence of adaptivity and efficiency. To address these issues, in this paper, we propose a Retinex-based deep unfolding network (URetinex-Net), which unfolds an optimization problem into a learnable network to decompose a low-light image into reflectance and illumination layers. By formulating the decomposition problem as an implicit priors regularized model, three learning-based modules are carefully designed, responsible for data-dependent initialization, high-efficient unfolding optimization, and user-specified illumination enhancement, respectively. Particularly, the proposed unfolding optimization module, introducing two networks to adaptively fit implicit priors in data-driven manner, can realize noise suppression and details preservation for the final decomposition results. Extensive experiments on real-world low-light images qualitatively and quantitatively demonstrate the effectiveness and superiority of the proposed method over state-of-the-art methods. The code is available at https://github.com/AndersonYong/URetinex-Net. Wenhui Wu 0001, Jian Weng 0009, Xu Wang 0006, Wenhan Yang, Jianmin Jiang |
CVPR | 4 |
| 2022 | Semi-automatic Data Annotation System for Multi-Target Multi-Camera Vehicle TrackingabstractMulti-target multi-camera tracking (MTMCT) plays an important role in intelligent video analysis, surveillance video retrieval, and other application scenarios. Nowadays, the deep-learning-based MTMCT has been the mainstream and has achieved fascinating improvements regarding tracking accuracy and efficiency. However, according to our investigation, the lacking of datasets focusing on real-world application scenarios limits the further improvements for current learning-based MTMCT models. Specifically, the learning-based MTMCT models training by common datasets usually cannot achieve satisfactory results in real-world application scenarios. Motivated by this, this paper presents a semi-automatic data annotation system to facilitate the real-world MTMCT dataset establishment. The proposed system first employs a deep-learning-based single-camera trajectory generation method to automatically extract trajectories from surveillance videos. Subsequently, the system provides a recommendation list in the following manual cross-camera trajectory matching process. The recommendation list is generated based on side information, including camera location, timestamp relation, and background scene. In the experimental stage, extensive results further demonstrate the efficiency of the proposed system. Haohong Liao, Silin Zheng, Xuelin Shen, Mark Junjie Li, Xu Wang 0006 |
DSAA | 5 |
| 2022 | A Learning-based Framework for Multi-View Instance Segmentation in PanoramaabstractThe application of panoramas in computer vision has received a lot of attention due to its ability to represent information about the surrounding environment. Instance segmentation on panoramas can make machines better understand 3D scenes. However, there are few efforts have been made on detecting the instance for panoramas. The main challenge is that in a panorama, objects are subject to three variations: geometric distortion, edge discontinuity and minification of objects. To address the above issues, we propose an instance segmentation method for panoramas based upon the multi-view fusion. First, a set of sub-views are sampled from a panorama by a rectilinear projection, and then the Cascade Mask R-CNN is employed to perform instance segmentation on each sub-view. Subsequently, we reverse mapping the obtained instance segmentation results of sub-views back to the panorama. Finally, we combine the complementary advantages of large and small fields of view, and merge the instance segmentation results of the entire panorama with the reorganized instance segmentation results to obtain high-quality instance segmentation results for panoramas. The quantitative experiments illustrate that our proposed method obtains higher Mean Average Precision (0.28) than existing benchmark methods. Weihao Ye, Ziyang Mai, Qiudan Zhang, Xu Wang 0006 |
DSAA | 4 |
| 2022 | End-To-End Depth Map Compression Framework Via Rgb-To-Depth Structure Priors LearningabstractIn this paper, we propose a novel framework to exploit and utilize the shared information inner RGB-D data for efficient depth map compression. Two main codecs, designed based on the existing end-to-end image compression network, are adopted for RGB image compression and enhanced depth image compression with RGB-to-Depth structure prior, respectively. In particular, we propose a Structure Prior Fusion (SPF) module to extract the structure information from both RGB and depth codecs at multi-scale feature levels and fuse the cross-modal feature to generate more efficient structure priors for depth compression. Extensive experiments show that the proposed framework can achieve competitive rate-distortion performance as well as RGB-D task-specific performance at depth map compression compared with the direct compression scheme. Zhuo Chen 0006, Yun Zhang 0002, Xu Wang 0006, Sam Kwong |
ICIP | 5 |
| 2022 | Learning-Based Multi-Stage Intra Partition for Versatile Video CodingabstractThe latest standard, Versatile Video Coding (VVC), doubles the coding efficiency over the previous generation standard. However, better performance is at the cost of a sharp increase in coding complexity. In order to reduce the complexity of VVC intra coding, this paper proposes a multi-stage block partition decision framework based on deep learning. First, we propose a three-stage redundant modes removal framework that decreases the number of modes checked in the brute-force process. Then, we build a lightweight CNN to complete the classification task of each stage. To reduce the burden of CNN and adapt to different Coding Unit (CU) sizes, we pre-process the luminance component of CU and use the results as input of the network. Finally, the multi-threshold adjusting scheme is proposed for trading off complexity reduction with the bit-rate increase. The experimental results shows our method can reduce the encoding time ranging from 16.93% to 69.40% with the bit-rate increase ranging from 0.31% to 3.59%. Such results demonstrate that our method has superior performance with a wide range of adjustments compared with other state-of-the-art methods. Hongji Zeng, Tiesong Zhao, Weize Feng, Jielian Lin, Xu Wang 0006 |
MMSP | 6 |
| 2022 | Expansion and Shrinkage of Localization for Weakly-Supervised Semantic SegmentationabstractGenerating precise class-aware pseudo ground-truths, a.k.a, class activation maps (CAMs), is essential for Weakly-Supervised Semantic Segmentation. The original CAM method usually produces incomplete and inaccurate localization maps. To tackle with this issue, this paper proposes an Expansion and Shrinkage scheme based on the offset learning in the deformable convolution, to sequentially improve the recall and precision of the located object in the two respective stages. In the Expansion stage, an offset learning branch in a deformable convolution layer, referred to as expansion sampler'', seeks to sample increasingly less discriminative object regions, driven by an inverse supervision signal that maximizes image-level classification loss. The located more complete object region in the Expansion stage is then gradually narrowed down to the final object region during the Shrinkage stage. In the Shrinkage stage, the offset learning branch of another deformable convolution layer referred to as theshrinkage sampler'', is introduced to exclude the false positive background regions attended in the Expansion stage to improve the precision of the localization maps. We conduct various experiments on PASCAL VOC 2012 and MS COCO 2014 to well demonstrate the superiority of our method over other state-of-the-art methods for Weakly-Supervised Semantic Segmentation. The code is available at https://github.com/TyroneLi/ESOL_WSSS. Jinlong Li 0003, Zequn Jie, Xu Wang 0006, Xiaolin Wei, Lin Ma 0002 |
NeurIPS | 3 |
| 2022 | Self-supervised Indoor 360-Degree Depth Estimation via Structural Regularization
Weifeng Kong, Qiudan Zhang, You Yang 0002, Tiesong Zhao, Wenhui Wu 0001, Xu Wang 0006 |
PRICAI (3) | 6 |
| 2022 | Deep stereoscopic image saliency inspired stereoscopic image thumbnail generation
Yu Zhou 0027, Xiaotong Xiao, Qiudan Zhang, Xu Wang 0006, Jianmin Jiang |
Multim. Tools Appl. | 4 |
| 2022 | Adaptive Viewpoint Feature Enhancement-Based Binocular Stereoscopic Image Saliency DetectionabstractModeling 3D visual saliency has received great attention due to the development of emerging 3D display technologies. Traditional methods relying on low-level features may not be efficient in interpreting 3D visual content from high-level semantic perspective. Despite numerous efforts dedicated to this area, existing 3D visual saliency detection methods do not necessarily excel in exploring the stereoscopic image saliency driven by the intra-view and inter-view dependencies among left and right views. In this paper, we propose a visual saliency detection method for stereoscopic images grounded on adaptive viewpoint feature enhancement via binocular vision. More specifically, the correlation among left and right views is investigated through a delicately designed binocular stereoscopic saliency feature aggregation module, enabling the generation of more representative saliency features towards binocular vision. Subsequently, to further aggregate the saliency features in multiple scales, we design a progressive attention-based saliency feature pyramid extraction module to effectively integrate the features from top-level to down-level based on the network hierarchy mechanism. The saliency maps are ultimately produced for stereoscopic images by evaluating the obtained saliency features. In addition, we create a stereoscopic image saliency dataset (SIS-3D) that includes 1086 stereoscopic image pairs with various content and their corresponding human eye fixation annotations, aiming to further facilitate the research on visual saliency detection for stereoscopic images. Extensive experiments demonstrate that our proposed method improves CC by an average of 4.02% compared to representative counterparts on the newly built saliency dataset and another publicly available dataset. Qiudan Zhang, Xiaotong Xiao, Xu Wang 0006, Shiqi Wang 0001, Sam Kwong, Jianmin Jiang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2021 | Game Theory-driven Rate Control for 360-Degree Video CodingabstractThe 360-degree video (omnidirectional video) has become popular recently due to its capability of providing immersive experience, which is generally achieved via spherical moving pictures with freedom of viewpoint changing. Nevertheless, the support of full-view visual contents has inevitably reshaped its perceptual quality metric and dramatically increased its bitrate output after video coding. Therefore in 360-degree video coding, the Rate Control (RC) problem, which aims to maximize the resulted perceptual quality under bitrate constraint, has become a challenging task yet to be addressed. In this paper, we observe a latitude-based bitrate discrepancy in equirectangular-projected 360-degree video coding and further utilize this feature in bitrate allocation under panoramic vision. We introduce game theory to find optimal inter/intra-frame bit allocations that maximize the overall RC performance in terms of utility function. Finally, an overall framework is proposed that is capable of providing both an improved bitrate accuracy and an enhanced perceptual quality. Experimental results demonstrate the efficiency of proposed method, with promising RC performances for 4K and 8K 360-degree videos. Tiesong Zhao, Jielian Lin, Xu Wang 0006, Yuzhen Niu |
ACM Multimedia | 4 |
| 2021 | A problem-specific non-dominated sorting genetic algorithm for supervised feature selection
Yu Zhou 0027, Junhao Kang, Xiao Zhang 0006, Xu Wang 0006 |
Inf. Sci. | 5 |
| 2021 | Active k-labelsets ensemble for multi-label classification
Ran Wang 0001, Sam Kwong, Xu Wang 0006, Yuheng Jia |
Pattern Recognit. | 3 |
| 2021 | Progressive Point Cloud Upsampling via Differentiable RenderingabstractIn this paper, we propose one novel progressive point cloud upsampling framework to tackle the non-uniform distribution issue during the point cloud upsampling process. Specifically, we design an Up-UNet feature expansion module which is capable of learning the local and global point features via a down-feature operator and an up-feature operator, respectively, to alleviate the non-uniform distribution issue and remove the outliers. Moreover, we design a hybrid loss function considering both the multi-scale reconstruction loss and the rendering loss. The multi-scale reconstruction loss enables each upsampling module to generate a denser point cloud, while the rendering loss via point-based differentiable rendering ensures that the proposed model preserves the point cloud structures. Extensive experimental results demonstrate that our proposed model achieves state-of-the-art performance in terms of both qualitative and quantitative evaluations. Xu Wang 0006, Lin Ma 0002, Shiqi Wang 0001, Sam Kwong, Jianmin Jiang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | A Multi-Task Collaborative Network for Light Field Salient Object DetectionabstractBeing able to predict the salient object is of fundamental importance in image processing and computer vision. With numerous approaches proposed for automatic image and video salient object detection, much less work has been dedicated to detecting and segmenting salient objects from light fields. In this article, based on the intrinsic characteristics of light fields, we carefully explore the complementary coherence among multiple cues including spatial, edge and depth information, and elaborately design a multi-task collaborative network for light field salient object detection. More specifically, the correlation mechanisms among edge detection, depth inference and salient object detection are carefully investigated to facilitate the representative saliency features. We first model the coherence among low-level features and heuristic semantic priors, as well as the edge information. Subsequently, the depth-oriented saliency features are derived from the geometry of light fields, in which the 3D convolution operation is leveraged with powerful representation capability to model the disparity correlations among multiple viewpoint images. Finally, a feature-enhanced salient object generator is developed to integrate these complementary saliency features, leading to the final salient object predictions for light fields. Quantitative and qualitative experiments demonstrate the superiority of our proposed model against the state-of-the-art methods over the public light field salient object detection datasets. Qiudan Zhang, Shiqi Wang 0001, Xu Wang 0006, Zhenhao Sun, Sam Kwong, Jianmin Jiang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2021 | Geometry Auxiliary Salient Object Detection for Light Fields via Graph Neural NetworksabstractLight field imaging, originated from the availability of light field capture technology, offers a wide range of applications in the field of computational vision. The capability of predicting salient objects of light fields remains technologically challenging due to its complicated geometry structure. In this paper, we propose a light field salient object detection approach that formulates the geometric coherence among multiple views of light fields as graphs, where the angular/central views represent the nodes and their relations compose the edges. The spatial and disparity correlations between multiple views are effectively explored through multi-scale graph neural networks, enabling the more comprehensive understanding of light field content and more representative and discriminative saliency features generation. Moreover, a multi-scale saliency feature consistency learning module is embedded to enhance the saliency features. Finally, an accurate salient object map is produced for the light field based upon the extracted features. In addition, we establish a new light field salient object detection dataset (CITYU-Lytro) that contains 817 light fields with diverse contents and their corresponding annotations, aiming to further promote the research on light field salient object detection. Quantitative and qualitative experiments demonstrate that the proposed method performs favorably compared with the state-of-the-art methods on the benchmark datasets. Qiudan Zhang, Shiqi Wang 0001, Xu Wang 0006, Zhenhao Sun, Sam Kwong, Jianmin Jiang |
IEEE Trans. Image Process. | 3 |
| 2020 | Lossy Geometry Compression Of 3d Point Cloud Data Via An Adaptive Octree-Guided NetworkabstractIn this paper, we propose a deep learning based framework for point cloud geometry lossy compression via hybrid representation of point cloud. First, the input raw 3D point cloud data is adaptively decomposed into non-overlapping local patches through adaptive Octree decomposition and clustering. Second, a framework of point cloud auto-encoder network with quantization layer is proposed for learning compact latent feature representation from each patch. Specifically, the proposed point cloud auto-encoder networks with different input size are trained for achieving optimal rate-distortion (RD) performance. Final, bitstream specifications of proposed compression systems with additional signaled meta-data and header information are designed to support parallel decoding and successive reconstruction. Experimental results shows that our proposed method can achieve 40.20% bitrate saving in average than the existing standard Geometry based Point Cloud Compression (G-PCC) codec. Xuanzheng Wen, Xu Wang 0006, Junhui Hou, Lin Ma 0002, Yu Zhou 0027, Jianmin Jiang |
ICME | 2 |
| 2020 | Content-Aware Cubemap Projection for Panoramic Image via Deep Q-Learning
Xu Wang 0006, Yu Zhou 0027, Longhao Zou, Jianmin Jiang |
MMM (2) | 2 |
| 2020 | Light Field Salient Object Detection via Hybrid Priors
Junlin Zhang, Xu Wang 0006 |
MMM (2) | 2 |
| 2020 | Multi-Exposure Decomposition-Fusion Model for High Dynamic Range Image Saliency DetectionabstractHigh dynamic range (HDR) imaging techniques have witnessed a great improvement in the past few decades. However, saliency detection task on HDR content is still far from well explored. In this paper, we introduce a multi-exposure decomposition-fusion model for HDR image saliency detection inspired by the brightness adaption mechanism. The proposed model is composed of three modules. Firstly, a decomposition module converts the input raw HDR image into a stack of LDR images by uniformly sampling the exposure time range. Secondly, a saliency region proposal network is employed to generate the candidate saliency maps for each LDR image in the exposure stack. Finally, an uncertainty weighting based fusion algorithm is applied to generate the overall saliency map for the input HDR image by merging the obtained LDR saliency maps. Extensive experiments show that our proposed model achieves superior performance compared with the state-of-the-art methods on the existing HDR eye fixation databases. The source code of the proposed model are made publicly available at https://github.com/sunnycia/DFHSal. Xu Wang 0006, Zhenhao Sun, Qiudan Zhang, Yuming Fang 0001, Lin Ma 0002, Shiqi Wang 0001, Sam Kwong |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2020 | Learning to Explore Saliency for Stereoscopic Videos Via Component-Based InteractionabstractIn this paper, we devise a saliency prediction model for stereoscopic videos that learns to explore saliency inspired by the component-based interactions including spatial, temporal, as well as depth cues. The model first takes advantage of specific structure of 3D residual network (3D-ResNet) to model the saliency driven by spatio-temporal coherence from consecutive frames. Subsequently, the saliency inferred by implicit-depth is automatically derived based on the displacement correlation between left and right views by leveraging a deep convolutional network (ConvNet). Finally, a component-wise refinement network is devised to produce final saliency maps over time by aggregating saliency distributions obtained from multiple components. In order to further facilitate research towards stereoscopic video saliency, we create a new dataset including 175 stereoscopic video sequences with diverse content, as well as their dense eye fixation annotations. Extensive experiments support that our proposed model can achieve superior performance compared to the state-of-the-art methods on all publicly available eye fixation datasets. Qiudan Zhang, Xu Wang 0006, Shiqi Wang 0001, Zhenhao Sun, Sam Kwong, Jianmin Jiang |
IEEE Trans. Image Process. | 2 |
| 2020 | Generative Adversarial Network-Based Intra Prediction for Video CodingabstractIn this paper, a novel intra prediction method is proposed to improve the video coding performance, in which the generative adversarial network (GAN) is adopted to intelligently remove the spatial redundancy with the inference process. The proposed GAN-based method improves the prediction by exploiting more information and generating more flexible prediction patterns. In particular, the intra prediction is modeled as an inpainting task, which is accomplished with the GAN model to fill in the missing part by conditioning on the available reconstructed pixels. As such, the learned GAN model is incorporated into both video encoder and decoder, and the rate-distortion optimization is performed for the competition between GAN-based intra prediction and traditional angular-based intra prediction to achieve better coding performance. The proposed scheme is implemented into the high-efficiency video coding test model (HM 16.17) and the versatile video coding test model (VTM 1.1). The experimental results show that the proposed algorithm can achieve 6.6%, 7.5%, and 7.5% under HM 16.17 and 6.75%, 7.63%, and 7.65% under VTM 1.1 bit rate savings on average for luma and chroma components in the intra coding scenario. Linwei Zhu, Sam Kwong, Yun Zhang 0002, Shiqi Wang 0001, Xu Wang 0006 |
IEEE Trans. Multim. | 5 |
| 2019 | Learning to Explore Intrinsic Saliency for Stereoscopic VideoabstractThe human visual system excels at biasing the stereoscopic visual signals by the attention mechanisms. Traditional methods relying on the low-level features and depth relevant information for stereoscopic video saliency prediction have fundamental limitations. For example, it is cumbersome to model the interactions between multiple visual cues including spatial, temporal, and depth information as a result of the sophistication. In this paper, we argue that the high-level features are crucial and resort to the deep learning framework to learn the saliency map of stereoscopic videos. Driven by spatio-temporal coherence from consecutive frames, the model first imitates the mechanism of saliency by taking advantage of the 3D convolutional neural network. Subsequently, the saliency originated from the intrinsic depth is derived based on the correlations between left and right views in a data-driven manner. Finally, a Convolutional Long Short-Term Memory (Conv-LSTM) based fusion network is developed to model the instantaneous interactions between spatio-temporal and depth attributes, such that the ultimate stereoscopic saliency maps over time are produced. Moreover, we establish a new large-scale stereoscopic video saliency dataset (SVS) including 175 stereoscopic video sequences and their fixation density annotations, aiming to comprehensively study the intrinsic attributes for stereoscopic video saliency detection. Extensive experiments show that our proposed model can achieve superior performance compared to the state-of-the-art methods on the newly built dataset for stereoscopic videos. Qiudan Zhang, Xu Wang 0006, Shiqi Wang 0001, Shikai Li, Sam Kwong, Jianmin Jiang |
CVPR | 2 |
| 2019 | Fast Coding Unit Decision for Intra Screen Content Coding Based on Ensemble LearningabstractThe Screen Content Coding (SCC) is an extension of High Efficiency Video Coding (HEVC), and it achieves significant improvement on compression ratio. However, the obtained coding efficiency is at the cost of high computational complexity. In this paper, to reduce the computation complexity, we propose to use an ensemble classifier for predicting the coding unit (CU) in intra-coding. Firstly, the L1-loss based linear support vector machine (SVM) is employed as basic classifier for its simplicity. Then, a bagging scheme is applied to train the linear classifiers and boost the prediction accuracy by ensemble learning. Compared with the reference software SCM-5.0, the proposed scheme can achieve 30% complexity reduction on average with only 1.64% bit rates increase. Yali Xue, Xu Wang 0006, Linwei Zhu, Zhaoqing Pan, Sam Kwong |
ICASSP | 2 |
| 2019 | Optimizing the Parameters for Post-Processing Consumer Photos via Machine LearningabstractPhoto sharing in social media is a part of everyday life for many, as inexpensive cameras integrated in smartphones are widely available. Unfortunately, low cost consumer devices are often prone to capture artifacts, and this is why there is a growing demand for automatic post-processing to enhance the image quality. Due to the wide range of distortions in non-professional photography, automatic selection of the post-processing methods and parameters is a challenging problem. In this paper, we present a subjective study based on rank-ordering method, comparing the subjective preferences between photos processed with different parameters for image sharpening and denoising. The subjective results are used as a basis to derive the ground truth values for the post-processing parameters for different photos. Then, we apply a pre-trained convolutional neural network (CNN) to extract a set of features from photos, used as input to a regression model to predict the optimal post-processing parameters. Test results show that the learning-based approach can predict post-processing parameters with a satisfactory accuracy. Linlin Bie, Xu Wang 0006, Jari Korhonen |
ICTAI | 2 |
| 2019 | Exploiting Local and Global Structure for Point Cloud Semantic Segmentation with Contextual Point RepresentationsabstractIn this paper, we propose one novel model for point cloud semantic segmentation,which exploits both the local and global structures within the point cloud based onthe contextual point representations. Specifically, we enrich each point represen-tation by performing one novel gated fusion on the point itself and its contextualpoints. Afterwards, based on the enriched representation, we propose one novelgraph pointnet module, relying on the graph attention block to dynamically com-pose and update each point representation within the local point cloud structure.Finally, we resort to the spatial-wise and channel-wise attention strategies to exploitthe point cloud global structure and thereby yield the resulting semantic label foreach point. Extensive results on the public point cloud databases, namely theS3DIS and ScanNet datasets, demonstrate the effectiveness of our proposed model,outperforming the state-of-the-art approaches. Our code for this paper is available at https://github.com/fly519/ELGS. Xu Wang 0006, Jingming He, Lin Ma 0002 |
NeurIPS | 1 |
| 2019 | Bidirectional image-sentence retrieval by local and global deep matching
Lin Ma 0002, Zequn Jie, Xu Wang 0006 |
Neurocomputing | 4 |
| 2018 | Complexity Control for HEVC Inter Coding Based on Two-Level Complexity Allocation and Mode SortingabstractHigh coding complexity is an obstruction when promoting the HEVC standard. In this paper, we propose a complexity control scheme for HEVC, which improves our previous work by employing a finer-grained complexity allocation scheme and mode sorting. In CTU-level complexity allocation, the time budget is allocated to each CTU proportionally to its estimated complexity. Then, the time budget for a CTU is further allocated to each CTU quadtree depth according to its probability of being the dominant depth. At last, a subset of modes are selected to reach the target complexity based on mode sorting. Compared with our previous work of which the least supported complexity ratio is 40%, the proposed scheme further reduces the lower bound of the ratio range to 20%. Jia Zhang 0002, Sam Kwong, Tiesong Zhao, Xu Wang 0006, Shiqi Wang 0001 |
ICIP | 4 |
| 2018 | Subjective Assessment of Post-Processing Methods for Low Light Consumer PhotosabstractConsumer photos taken in low light conditions often suffer from substantial undesired capture artifacts, such as shakiness and sensor noise. In this paper, we use rank ordering method to assess the subjective preferences among different postprocessing methods used to alleviate capture artifacts. The results show that most users prefer sharpened photos, even in the presence of substantial sensor noise. However, there are also systematic differences in individual preferences between users. Therefore, user preferences need to be considered in addition to the image characteristics, when selecting the post-processing algorithms and parameters for photo quality enhancement. Linlin Bie, Xu Wang 0006, Jari Korhonen |
QoMEX | 2 |
| 2018 | Deep intensity guidance based compression artifacts reduction for depth map
Xu Wang 0006, Yun Zhang 0002, Lin Ma 0002, Sam Kwong, Jianmin Jiang |
J. Vis. Commun. Image Represent. | 1 |
| 2018 | Quaternion representation based visual saliency for stereoscopic image quality assessment
Xu Wang 0006, Lin Ma 0002, Sam Kwong, Yu Zhou 0027 |
Signal Process. | 1 |
| 2018 | Two-Stage Fast Inter CU Decision for HEVC Based on Bayesian Method and Conditional Random FieldsabstractIn the latest video coding standard high efficiency video coding (HEVC), a quadtree-based coding unit (CU) partitioning scheme is adopted to better adapt to the characteristics of the video contents. However, the flexible scheme significantly increases the coding complexity, because large amount of possible CU partitioning modes should be traversed. In this paper, we propose a two-stage fast inter CU decision method to reduce the coding complexity of the HEVC encoders. In Stage I, all the CUs are classified into three categories based on the Bayesian method after the prediction unit (PU) mode merge 2N × 2N is checked. Early CU pruning and early CU skipping are then applied to two of the categories, respectively. For the remaining category which is difficult to differentiate by the rate-distortion cost of the PU mode merge 2N × 2N, an early CU pruning scheme based on conditional random fields is performed in Stage II, which takes both the local characteristics of the current CU and the coding information of its neighboring CUs into consideration. Experimental results show that our method can reduce 54.93% and 45.84% of the coding complexity on average with only 1.19% and 1.03% Bjontegaard delta bitrate increment under the random access main and the low delay P configurations, respectively. Jia Zhang 0002, Sam Kwong, Xu Wang 0006 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2018 | Effective Data Driven Coding Unit Size Decision Approaches for HEVC INTRA CodingabstractHigh Efficiency Video Coding (HEVC) INTRA coding improves compression efficiency by adopting advanced coding technologies, such as multi-level quad-tree block partitioning and up to 35-mode INTRA prediction. However, it significantly increases the coding complexity, memory access, and power consumption, which goes against its widely applications, especially for ultra-high definition and/or mobile video applications. To tackle this problem, we propose effective data driven coding unit (CU) size decision approaches for HEVC INTRA coding, which consists of two stages of support vector machine-based fast INTRA CU size decision schemes at four CU decision layers. At the first stage classification, a three output classifier with offline learning is developed to early terminate the CU size decision or early skip checking the current CU depth. As for the samples that neither early skipped nor early terminated, the second stage of binary classification, which learns online from previous coded frames, is proposed to further refine the CU size decision. Representative features for the CU size decision are explored at different decision layers and stages of classifications. Finally, the optimal parameters derived from the training data are achieved to reasonably allocate complexity among different CU layers at given total rate-distortion degradation constraint. Extensive experiments show that the proposed overall algorithm can achieve 27.95%–80.53% and 52.48% on average complexity reduction for the CU size decision as compared with the original HM16.7 model. Meanwhile, the average Bjonteggard delta peak-signal-to-noise ratio degradation is only −0.08 dB, which is negligible. The overall performance of the proposed algorithm outperforms the state-of-the-art benchmark schemes. Yun Zhang 0002, Zhaoqing Pan, Na Li 0015, Xu Wang 0006, Gangyi Jiang, Sam Kwong |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2017 | DCT-based image up-sampling using anchored neighborhood regressionabstractDown-sampling in the discrete cosine transform (DCT) domain is preferable for images coded by DCT transform, such as JPEG/MJPEG/H.264, etc. Recent researches show that the truncated high-frequency DCT coefficients during the DCT down-sampling process can be estimated by learning the correlations between low-frequency and high-frequency DCT coefficients. In this paper, we propose to utilize the powerful super-resolution framework using sparse dictionaries with anchored neighborhood regression to significantly improve the accuracy of the estimated high-frequency DCT coefficients. Experimental results show that the proposed framework outperforms the state-of-the-art DCT-based up-sampling methods in terms of PSNR (0.3-1.63dB) and SSIM values for standard image datasets Set5 and Set14, while the computational time of the proposed method is 23× times faster than the state-of-the-art learning-based method using k-NN MMSE due to pre-computation of the ridge regression during the training process. Kwok-Wai Hung, Jianmin Jiang, Qinglong Chang, Xu Wang 0006 |
ICIP | 4 |
| 2017 | Study of subjective and objective quality assessment for screen content imagesabstractIn this paper, we present the results of a recent large-scale subjective study of image quality on a collection of screen contents distorted by a variety of application-relevant processes. With the development of multi-device interactive multimedia applications, metrics to predict the visual quality of screen content images (SCIs) as perceived by subjects are becoming fundamentally important. For developing the objective image quality assessment (IQA) method, there is a need for large-scale public database with diversity of distorted types and scene contents, and available subjective scores of distorted SCIs. The resulting Immersive Media Laboratory screen content image quality database (IML-SCIQD) contains 1250 distorted SCIs from 25 reference SCIs with 10 distortion types. Each image was rated by 35 human observers, and the different mean opinion scores (DMOS) were obtained after data processing. The performance comparison of 17 state-of-the-arts, publicly available IQA algorithms are evaluated on the new database. The database will be available online in our project website. Xu Wang 0006, Yingying Zhu 0001, Yun Zhang 0002, Jianmin Jiang, Sam Kwong |
ICIP | 1 |
| 2017 | Multi-class ranking based most probable prediction unit selection for HEVC encodingabstractIn this paper, an incremental learning based multi-class Prediction Units (PUs) ranking approach is presented for High Efficiency Video Coding (HEVC) Rate-Distortion-Complexity (RDC) optimization. In particular, the process of PUs selection is formulated as a binary classification plus multi-class ranking task, and incremental learning is applied for classifier training to better exploit the information in the emerging training data. Furthermore, the proposed most probable PUs selection scheme is incorporated into a joint RDC optimization framework, where the complexity can be flexibly allocated targeting at minimizing computational cost under a constrained RD performance degradation. Experimental results demonstrate that the proposed approach can reduce 53.7% and 50.4% computational complexity on average under low delay P and random access configurations with ignorable RD performance degradation, which outperforms the state-of-the-art approaches in terms of RDC performance. Linwei Zhu, Sam Kwong, Yun Zhang 0002, Xu Wang 0006, Shiqi Wang 0001 |
VCIP | 4 |
| 2017 | VideoSet: A large-scale compressed video quality dataset based on JND measurementabstract• A large-scale JND-based coded video quality dataset is presented. • The VideoSet contains 220 5-s sequences in four resolutions coded by H.264/AVC. • The subjective test procedure, JND data cleaning and properties are described. • The significance and implications of the VideoSet are discussed. • This work points out a clear path to data-driven perceptual coding. A new methodology to measure coded image/video quality using the just-noticeable-difference (JND) idea was proposed in Lin et al. (2015). Several small JND-based image/video quality datasets were released by the Media Communications Lab at the University of Southern California in Jin et al. (2016) and Wang et al. (2016) [3]. In this work, we present an effort to build a large-scale JND-based coded video quality dataset. The dataset consists of 220 5-s sequences in four resolutions (i.e., 1920 × 1080 , 1280 × 720 , 960 × 540 and 640 × 360 ). For each of the 880 video clips, we encode it using the H.264/AVC codec with QP = 1 , … , 51 and measure the first three JND points with 30 + subjects. The dataset is called the “VideoSet”, which is an acronym for “Video Subject Evaluation Test (SET)”. This work describes the subjective test procedure, detection and removal of outlying measured data, and the properties of collected JND data. Finally, the significance and implications of the VideoSet to future video coding research and standardization efforts are pointed out. All source/coded video clips as well as measured JND data included in the VideoSet are available to the public in the IEEE DataPort (Wang et al., 2016 [4]). Haiqiang Wang, Ioannis Katsavounidis, Jiantong Zhou, Jeong-Hoon Park, Shawmin Lei, Xin Zhou 0001, Man-On Pun, Xin Jin 0002, Ronggang Wang, Xu Wang 0006, Yun Zhang 0002, Jiwu Huang, Sam Kwong, C.-C. Jay Kuo |
J. Vis. Commun. Image Represent. | 10 |
| 2017 | Motion-Homogeneous-Based Fast Transcoding Method From H.264/AVC to HEVCabstractWith the popularity of high-efficiency video coding (HEVC) standard, a video server usually transcodes a video stream to HEVC for its higher compression ratio. In this paper, a fast H.264/advanced video coding (AVC) to HEVC transcoding method is proposed. In the HEVC encoding procedure, a coding unit (CU), which is a motion-homogeneous block, is first checked based on the analysis of the decoded information from H.264/AVC bit stream. Then, for motion-homogeneous blocks, CU depth and the corresponding prediction unit (PU) mode's early termination strategies are proposed based on the CU size and corresponding prior statistical knowledge. For non-motion-homogeneous blocks, a corresponding PU mode's early termination strategy is also proposed. Experimental results demonstrate the effectiveness of the proposed method. Hui Yuan 0001, Chenglin Guo, Xu Wang 0006, Sam Kwong |
IEEE Trans. Multim. | 4 |
| 2016 | Complex singular value decomposition based stereoscopic image quality assessmentabstractDesigning a reliable and generic perceptual quality metric is a challenging issue in three-dimensional (3D) visual signal processing. Due to the limited knowledge on 3D perceptual, it is difficult to fuse the visual information of left and right views in an effective way. In this paper, we propose a complex singular value decomposition (CSVD) based stereoscopic image quality assessment (SIQA) metric. First, the corresponding blocks of the left/right view are grouped into complex representation (CR) block through the scale-invariant feature transform (SIFT) view matching process. Then we compute the CSVD coefficients of each CR block. Final, a CSVD based quality pooling stage is employed to predict the final visual quality of the distorted 3D image. Experimental results demonstrate that the proposed metric has good consistency with 3D perception of human. Xu Wang 0006, Lin Ma 0002, Yu Zhou 0027, Sam Kwong |
VCIP | 1 |
| 2016 | Breast cancer discriminant feature analysis for diagnosis via jointly sparse learning
Heng Kong, Zhihui Lai 0001, Xu Wang 0006, Feng Liu 0013 |
Neurocomputing | 3 |
| 2016 | Reorganized DCT-based image representation for reduced reference stereoscopic image quality assessment
Lin Ma 0002, Xu Wang 0006, Qiong Liu 0001, King Ngi Ngan |
Neurocomputing | 2 |
| 2016 | Bilevel optimization of block compressive sensing with perceptually nonlocal similarity
Yu Zhou 0027, Sam Kwong, Hainan Guo, Wei Gao 0003, Xu Wang 0006 |
Inf. Sci. | 5 |
| 2016 | A phase congruency based patch evaluator for complexity reduction in multi-dictionary based single-image super-resolution
Yu Zhou 0027, Sam Kwong, Wei Gao 0003, Xu Wang 0006 |
Inf. Sci. | 4 |
| 2016 | User models of subjective image quality assessment on virtual viewpoint in free-viewpoint video system
You Yang 0002, Xu Wang 0006, Qiong Liu 0001, Mingliang Xu 0001 |
Multim. Tools Appl. | 2 |
| 2016 | DCT Coefficient Distribution Modeling and Quality Dependency Analysis Based Frame-Level Bit Allocation for HEVCabstractA frame-level bit allocation optimization method is proposed to improve the rate-distortion performance for High Efficiency Video Coding. First, to avoid the demerits of the mixture Laplacian distribution model on complexity, a new synthesized Laplacian distribution (SynLD) model is proposed to describe the discrete cosine transform transformed coefficients based on Kullback-Leibler-divergence analysis. Second, quality dependencies among frames are investigated, and a linear relationship between quality dependency factor (QDF) and skipmode percentage is proposed for QDF prediction. Based on the proposed SynLD model and QDF prediction method, a p-domain-based frame-level bit allocation method is proposed. Experimental results show that when compared with the state-of-the-art pixel-based unified rate-quantization (URQ) model and R-λ-model-based algorithms, 1.75- and 0.16-dB BD-peak signalto-noise ratio (PSNR) gains can be achieved by the proposed bit allocation method, respectively. For quality consistency, the average PSNR standard deviation shows 0.16 and 0.02 dB lower than URQ and R-λ-model-based algorithms, respectively. The proposed method also has a much more stable buffer control status and works well for scene change cases. Wei Gao 0003, Sam Kwong, Hui Yuan 0001, Xu Wang 0006 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2015 | Smooth View Quality Oriented Bit Allocation Optimization for 3D Video CodingabstractView level bit allocation is an fundamental optimization problem in multiview video plus depth (MVD) based 3D video coding (3DVC). In this paper, we propose a smooth view quality oriented view level bit allocation framework for MVD based 3DVC. The Cauchy-density based rate-distortion model of the texture video and depth map are employed to represent the rate distortion properties. The relationship between the distortion of synthesized view and quantization step size of texture videos and depth maps is approximately fitted as linear model. Final, the bit allocation problem is solved by convex optimization algorithms. Experimental results demonstrated that our proposed algorithm can achieve good performance with acceptable computational complexity comparing to the full search scheme. Xu Wang 0006, Sam Kwong, Wei Gao 0003, Yu Zhou 0027, Hui Yuan 0001, Yun Zhang 0002 |
SMC | 1 |
| 2015 | Natural image statistics based 3D reduced reference image quality assessment in contourlet domain
Xu Wang 0006, Qiong Liu 0001, Ran Wang 0001, Zhuo Chen 0006 |
Neurocomputing | 1 |
| 2015 | A bundled-optimization model of multiview dense depth map synthesis for dynamic scene reconstruction
You Yang 0002, Xu Wang 0006, Qiong Liu 0001, Li Yu 0003 |
Inf. Sci. | 2 |
| 2015 | View synthesis distortion elimination filter for depth video coding in 3D video broadcasting
Linwei Zhu, Yun Zhang 0002, Xu Wang 0006, Sam Kwong |
Multim. Tools Appl. | 3 |
| 2015 | View synthesis distortion model based frame level rate control optimization for multiview depth video coding
Xu Wang 0006, Sam Kwong, Hui Yuan 0001, Yun Zhang 0002, Zhaoqing Pan |
Signal Process. | 1 |
| 2015 | Machine Learning-Based Coding Unit Depth Decisions for Flexible Complexity Allocation in High Efficiency Video CodingabstractIn this paper, we propose a machine learning-based fast coding unit (CU) depth decision method for High Efficiency Video Coding (HEVC), which optimizes the complexity allocation at CU level with given rate-distortion (RD) cost constraints. First, we analyze quad-tree CU depth decision process in HEVC and model it as a three-level of hierarchical binary decision problem. Second, a flexible CU depth decision structure is presented, which allows the performances of each CU depth decision be smoothly transferred between the coding complexity and RD performance. Then, a three-output joint classifier consists of multiple binary classifiers with different parameters is designed to control the risk of false prediction. Finally, a sophisticated RD-complexity model is derived to determine the optimal parameters for the joint classifier, which is capable of minimizing the complexity in each CU depth at given RD degradation constraints. Comparative experiments over various sequences show that the proposed CU depth decision algorithm can reduce the computational complexity from 28.82% to 70.93%, and 51.45% on average when compared with the original HEVC test model. The Bjøntegaard delta peak signal-to-noise ratio and Bjøntegaard delta bit rate are -0.061 dB and 1.98% on average, which is negligible. The overall performance of the proposed algorithm outperforms those of the state-of-the-art schemes. Yun Zhang 0002, Sam Kwong, Xu Wang 0006, Hui Yuan 0001, Zhaoqing Pan, Long Xu 0001 |
IEEE Trans. Image Process. | 3 |
| 2015 | Rate Distortion Optimized Inter-View Frame Level Bit Allocation Method for MV-HEVCabstractIn multi-view video coding, since inter-view prediction has been adopted as an important coding tool which could improve coding efficiency greatly, inter-view dependency is inevitable, i.e., the distortion of the reference view (RV) picture could be propagated to the non-reference view (NRV) pictures . Therefore, in order to achieve higher coding efficiency , the inter-view dependency must be taken into account for inter-view bit allocation. In this paper, the inter-view dependency is analyzed in detail, and a rate-distortion (RD) model for NRVs is derived by taking the distortion of RV into account. Based on the derived RD model, the inter-view bit allocation is represented as a mathematical problem with an analytic form, and is solved by a convex optimization (Lagrangian Multiplier) method. Experimental results demonstrate that the RD performance and the inter-view quality consistency of the proposed method is better than existing methods, while the complexity of the proposed method is comparable with the existing methods. Hui Yuan 0001, Sam Kwong, Xu Wang 0006, Wei Gao 0003, Yun Zhang 0002 |
IEEE Trans. Multim. | 3 |
| 2014 | Rank learning on training set selection and image quality assessmentabstractMachine learning (ML) techniques are widely used in recent no-reference visual quality assessment (NR-VQA) metrics by training on subjective image quality databases. In these metrics, the optimization function is constructed based on L2norm of the distance between subjective image quality and predicted image quality. There are two problems in these L2norm based methods: (1) human's opinion on subjective image quality rating is not reliable at fine-scale level. A small difference between subjective image qualities represented by mean opinion scores (MOSs) of two images may not truly reflect the real quality difference between these two images, but acts as noise. The optimization process should avoid such noise. (2) Generally, human's opinion on pairwise comparison (PC) for image quality is more reliable and believable than MOS. The importance of PC is ignored during the optimization process of existing ML-based studies, which are designed based on the numerical rating system. In this paper, we introduce image quality ranking concept to establish a new optimization objective instead of L2norm optimization, and then a novel NR-VQA is constructed based on ranking learning. The proposed metric firstly suggests a reasonable training set for ML, which is ignored by existing ML-based NR-VQA. The ranking theory is adopted to build optimization function, which reflects the properties of PC over the numerical ranting system used by traditional NR-VQA. By ignoring the small difference between MOSs from two images during the optimization process, the proposed ranking-based NR-VQA can also well address the first problem from the existing related metrics. Experimental results show that the proposed ranking-based NR-VQA can obtain better performance over the state-of-the-art NR-VQA approaches. Long Xu 0001, Weisi Lin, Jia Li 0003, Xu Wang 0006, Yihua Yan, Yuming Fang 0001 |
ICME | 4 |
| 2014 | A multi-dimensional image quality prediction model for user-generated images in social networks
You Yang 0002, Xu Wang 0006, Jialie Shen 0001, Li Yu 0003 |
Inf. Sci. | 2 |
| 2014 | Generalized Nash Bargaining Solution to Rate Control Optimization for Spatial Scalable Video CodingabstractRate control (RC) optimization is indispensable for scalable video coding (SVC) with respect to bitstream storage and video streaming usage. From the perspective of centralized resource allocation optimization, the inner-layer bit allocation problem is similar to the bargaining problem. Therefore, bargaining game theory can be employed to improve the RC performance for spatial SVC. In this paper, we propose a bargaining game based one-pass RC scheme for spatial H.264/SVC. In each spatial layer (SL), the encoding constraints, such as bit rates, buffer size are jointly modeled as resources in the inner-layer bit allocation bargaining game. The modified rate-distortion (R-D) model incorporated with the inter-layer coding information is investigated. Then the generalized Nash bargaining solution (NBS) is employed to achieve an optimal bit allocation solution. The bandwidth is allocated to the frames from the generalized NBS adaptively based on their own bargaining powers. Experimental results demonstrate that the proposed rate control algorithm achieves appealing image quality improvement and buffer smoothness. The average mismatch of our proposed algorithm is within the range of 0:19%2:63%. Xu Wang 0006, Sam Kwong, Long Xu 0001, Yun Zhang 0002 |
IEEE Trans. Image Process. | 1 |
| 2013 | Early termination for TZSearch in HEVC Motion EstimationabstractThe TZSearch algorithm was adopted in the high efficiency video coding reference software HM as a fast Motion Estimation (ME) algorithm for its excellent performance in reducing ME time and maintaining a comparable Rate Distortion (RD) performance. However, the multiple initial search point decision and the hybrid block matching search contribute a relatively high computational complexity to TZSearch. In this paper, based on the statistical analysis of the probability of median predictor to be selected as the final best point in the large Coding Units (CUs) (64×64, 32×32) and small CUs (16×16, 8×8) as well as the center-biased characteristic of the final best search point in ME process, we propose two early terminations for TZSearch. Experimental results show that the proposed early terminations can achieve 38.96% encoding time saving, while the RD performance degradation is quite acceptable. Zhaoqing Pan, Yun Zhang 0002, Sam Kwong, Xu Wang 0006, Long Xu 0001 |
ICASSP | 4 |
| 2013 | Quality Assessment on User Generated Image for Mobile Search Application
Qiong Liu 0001, You Yang 0002, Xu Wang 0006, Liujuan Cao |
MMM (2) | 3 |
| 2011 | Considering binocular spatial sensitivity in stereoscopic image quality assessmentabstractDeveloping reliable and generic perceptual quality metrics is an challenging issue in three-dimensional (3D) visual signals processing, although many two dimensional (2D) image quality metrics have been proposed and work well on 2D images. In this paper, the binocular spatial sensitivity influenced by the binocular fusion and rivalry properties is considered in the quality measurement. Firstly, the binocular spatial sensitivity map is modeled to reflect the properties. Then, a framework of integration of binocular spatial sensitivity map into quality assessment is presented. Experimental results show that the proposed metric correlate well with human perception of quality on a dataset of 3D images and human subjective scores. Xu Wang 0006, Sam Kwong, Yun Zhang 0002 |
VCIP | 1 |
| 2009 | Reduced Reference Image Quality Assessment Based on Contourlet Domain and Natural Image StatisticsabstractReduced-reference (RR) image quality assessment metrics (IQA) evaluate the quality of images by extracting a parameter set from the original reference image and using this set in place of the actual reference image. In this paper, we propose a novel RR-IQA metric based on Contourlet transform. By combining Contourlet transform with a version of the hidden Markov model - Gaussian scale mixtures (GSM), the marginal distributions of neighbor coefficients in the Contourlet domain are modeled. With Contourlet transform as a pre-processing, the marginal histogram of coefficients in each subband can be well fitted by Guassian distribution after divisive normalization transforming. The standard derivation of the fitted Guassian transform and fitted error will be extracted as feature parameters. Experiments show that the proposed metric has good consistency with human subjective perception. Xu Wang 0006, Gangyi Jiang, Mei Yu 0001 |
ICIG | 1 |