Jian Xiong 0005

dblp:48/8081-5 · DBLP profile ↗
← Back
38ranked-venue papers
11as first author
26since 2021 · last 2026
0000-0002-4720-4102ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 30 · 11 first-author · 20 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021Computer networks · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 Point Cloud Quality Assessment via Multi-View Structure-Aware Feature Fusion
abstract
Point cloud quality assessment (PCQA) is essential for reliable 3D visual applications. While point-based methods face challenges in characterizing distortions due to point cloud disorder, projection-based approaches offer better efficiency but suffer from geometric distortion insensitivity and texture representation blind spots. This study proposes SAF-Net, a multi-view structure-aware feature fusion network for PCQA. We first identify two key limitations in projection-based methods: insufficient geometric distortion perception and representation blind spots (RBS) in texture images. To address these issues, SAF-Net innovatively integrates object mask maps and local binary pattern (LBP) maps. The mask maps enhance geometric distortion perception by extracting edge sharpness and curvature variations, while LBP maps capture essential structural information to overcome RBS and align with human visual system (HVS) sensitivity. SAF-Net employs a hybrid CNN-ViT architecture to balance local feature extraction and global context modeling, along with a progressive fusion strategy to optimize cross-modal feature interaction. Extensive experiments demonstrate the superior performance of SAF-Net on multiple benchmarks, establishing new state-of-the-art results in PCQA.
Jian Xiong 0005, Lingxia Jiang, Xianzhong Long, Miaohui Wang, Hao Gao 0005
AAAI1
2026 Region-aware focused masked contrast for self-supervised representation learning
Xianzhong Long, Yun Li 0009, Jian Xiong 0005
Knowl. Based Syst.4
2026 Cellular Aggregation Graph Convolutional Network for Point Cloud Quality Assessment
abstract
Point cloud quality assessment (PCQA) is a challenging task due to the inherently disordered nature of points. Existing point-based methods, such as sparse convolution and PointNet, are limited by local spatial modeling and structural feature extraction. Although 3D graph convolutional networks (GCNs) offer advantages in capturing local structural features through explicit geometric modeling and deformable kernels, their scalability is hindered by the high memory consumption associated with storing neighborhood matrices, particularly for large-scale point clouds. In this paper, to better extract hierarchical structural information and maintain efficiency in computational memory, we propose a novel point-based no-reference PCQA method, namely cellular aggregation network (CANet). The method effectively and efficiently extracts the quality-aware features of large patches in a divide-and-conquer manner. Specifically, a cellular sampling (CS) module is introduced to divide large patches into smaller cells, effectively avoiding the problem of memory explosion. A cellular aggregation (CA) module is proposed to extract intra-cell features and fuse inter-cell features. Moreover, a global aggregation (GA) module is presented to extract global sketch information. Finally, a long-term fusion (LTF) module is introduced to capture long-term dependencies between the features of the CA and GA modules. Experimental results on benchmark datasets demonstrate that the proposed model achieves state-of-the-art performance.
Jian Xiong 0005, Lingxia Jiang, Qiang Hu 0003, Jiucheng Xie, Hao Gao 0005
IEEE Trans. Circuits Syst. Video Technol.1
2026 Foundation Model Empowered Real-Time Video Conference With Semantic Communications
abstract
With the development of real-time video conferences, interactive multimedia services have proliferated, leading to a surge in traffic. Interactivity becomes one of the main features on future multimedia services, which brings a new challenge to Computer Vision (CV) for communications. In addition, many directions for CV in video, like recognition, understanding, saliency segmentation, coding, and so on, do not satisfy the demands of the multiple tasks of interactivity without integration. Meanwhile, with the rapid development of the foundation models, we apply task-oriented semantic communications to handle them. Therefore, we propose a novel framework, called Real-Time Video Conference with Foundation Model (RTVCFM), to satisfy the requirement of interactivity in the multimedia service. Firstly, at the transmitter, we perform the causal understanding and spatiotemporal decoupling on interactive videos, with the Video Time-Aware Large Language Model (VTimeLLM), Iterated Integrated Attributions (IIA) and Segment Anything Model 2 (SAM2), to accomplish the video semantic segmentation. Secondly, in the transmission, we propose a two-stage semantic transmission optimization driven by Channel State Information (CSI), which is also suitable for the weights of asymmetric semantic information in real-time video, so that we achieve a low bit rate and high semantic fidelity in the video transmission. Thirdly, at the receiver, RTVCFM provides multidimensional fusion with the whole semantic segmentation by using the Diffusion Model for Foreground Background Fusion (DMFBF), and then we reconstruct the video streams. Finally, the simulation result demonstrates that RTVCFM can achieve a compression ratio as high as 95.6%, while it guarantees high semantic similarity of 98.73% in Multi-Scale Structural Similarity Index Measure (MS-SSIM) and 98.35% in Structural Similarity (SSIM), which shows that the reconstructed video is relatively similar to the original video.
Mingkai Chen 0001, Mujian Zeng, Xiaoming He 0004, Jian Xiong 0005, Lei Wang 0009, Anwer Adel Al-Dulaimi, Shahid Mumtaz
IEEE Trans. Image Process.5
2026 Effective Gaussian Management for High-Fidelity Scene Reconstruction
abstract
This paper proposes an effective Gaussian management framework for high-fidelity scene reconstruction of both appearance and geometry. Unlike recent Gaussian Splatting (GS) pipelines that treat all primitives uniformly during optimization, our framework explicitly manages the attribute activation, representation and pruning of Gaussian. Specifically, our framework first introduces GauSep, a novel densification strategy that selectively activates Gaussian color or normal attributes to alleviate destructive gradient conflicts arising from dual supervision. We further propose GauRep, an adaptive Gaussian representation that dynamically adjusts spherical harmonics (SHs) orders and performs task-decoupled pruning to reduce redundancy at both the individual and global levels. To provide reliable geometric supervision for above mangement process, we additionally introduce CoRe, an regularized surface reconstruction module that distills robust normal fields from an SDF branch to the Gaussian representation through a confidence mechanism. Notably, the proposed Gaussian management is compatible with various reconstruction architectures and can be seamlessly integrated to improve performance while reducing size of the model. Extensive experiments demonstrate that our approach achieves superior or comparable performance in appearance and geometry reconstruction compared with state-of-the-art methods, while using significantly fewer parameters.
Jiateng Liu, Hao Gao 0005, Jiucheng Xie, Chi-Man Pun, Jian Xiong 0005, Haolun Li 0001, Junxin Chen 0001, Feng Xu 0005
IEEE Trans. Vis. Comput. Graph.5
2025 Adaptive Skeleton Prompt Tuning for Cross-Dataset 3D Human Pose Estimation
abstract
Inconsistency of distributions in human actions and camera viewpoints can lead to significant deviations when the pre-trained 3D pose estimators are tested on cross-datasets. In practical applications, the estimators usually follow the standard full fine-tuning paradigm on the target dataset, which requires updating and saving a complete set of training parameters for different tasks, resulting in a large waste of resources and distorting pre-trained features. Taking inspiration from the widely used prompt learning in NLP, we explore the parameter-efficient fine-tuning solution of 3D pose estimators for the first time and propose the Adaptive Skeleton Prompt Tuning (ASP-Tuning) method, which freezes the backbone of the pre-trained model and generates a series of pose generic promptings as well as adaptive promptings specific to the input skeleton features to learn distribution transformation. Extensive experiments on multiple estimator backbones and datasets show that our method is superior to other fine-tuning methods and achieves state-of-the-art performance.
Haolun Li 0001, Fuchen Zheng, Ye Liu 0005, Jian Xiong 0005, Haidong Hu, Hao Gao 0005
ICASSP4
2025 RFEM: Remote Feature Enhancement Module for Target Detection
abstract
The research and development of dense crowd detection technology have always been one of the hot and challenging topics in the field of computer vision. DETR-like models have shown good performance in both training efficiency and inference capabilities. Nevertheless, as the optimization proceeds, these models can demonstrate sparse long-range feature correlations. This paper presents a specialized long-range feature enhancement module intended for optimizing DETR-like detection models. By utilizing an optimized PVM clustering algorithm, the robustness of the model is enhanced, and linear attention is incorporated into the aggregated tokens to reinforce long-range feature relationships. Besides, our method maintains the connections between occluded segmentation features during both training and inference phases. It also enhances the detection accuracy of small targets without increasing computational overhead. We conducted experiments on the COCO 2017 and the CrowdHuman datasets, and extensive experimental results demonstrate the effectiveness of our proposed method.
Chuangye Wang, Jian Xiong 0005, Haolun Li 0001, Hao Gao 0005
ICASSP3
2025 CANet: Cellular Aggregation Network for Point Cloud Quality Assessment
abstract
The concept of visual masking reveals that human visual perception is influenced by content and distortion information. Existing projection-based methods lose depth information and intrinsic topological structures. Due to the limitations of computational memory, the existing point-based methods tend to deal with small patches with little content information. In this paper, we propose a novel point-based no-reference quality assessment method, namely cellular aggregation network (CANet). The method effectively extracts the quality-aware features of large patches in a divide-and-conquer manner. Specifically, the cellular sampling module is used to divide large patches into smaller cells, which effectively avoids the memory explosion problem. The cellular aggregation module is proposed to obtain more content information from small cells. A global aggregation module is proposed to extract global sketch information. Furthermore, a long-term fusion module is introduced to capture long-term dependencies, which can better receive content-aware semantic features. Experimental results on benchmark databases demonstrate that CANet achieves competitive performances.
Lingxia Jiang, Jian Xiong 0005, Jiucheng Xie, Hao Gao 0005
ISCAS2
2025 Multi-Task Learning Model for V-PCC Geometry Compression Artifact Removal
abstract
In video-based point cloud compression (V-PCC), point clouds are projected as videos using a patch projection method and then compressed using video coding techniques. However, the lossy video compression and the down-sampling of occupancy maps (OMs) can lead to geometry compression artifacts, i.e., depth errors and OM errors, respectively. These errors can significantly affect the reconstruction quality of the point clouds. Existing methods can only eliminate one type of error and therefore have limited quality improvement. In this paper, to improve the quality maximally, a multi-task learning-based geometry compression artifact removal method is proposed to reduce both types of errors simultaneously. Considering the differences between the two tasks, the proposed method deals with the challenges of shared feature extraction and heterogeneous objective optimization. First, we propose a context-aware multi-task learning (CAML) model. The proposed CAML model can extract shared features that are context-aware and satisfy both tasks. Second, an improved optimization scheme is presented to train the proposed model. The improved optimization can fix the gradient imbalance of model updating. Cross-validation experiments show that the proposed method saves an average of over 45% Bjϕntegaard Delta bitrate in terms of the D2 metric.
Jian Xiong 0005, Jiucheng Xie, Hui Yuan 0001, Hao Gao 0005
IEEE Trans. Circuits Syst. Video Technol.1
2024 Geometry Compression Artifact Removal for V-PCC over a Wide Bitrate Range
abstract
In video-based point cloud compression (V-PCC), point clouds are generated as videos via patch projection to be compressed using video coding techniques. However, a large number of filled empty pixels in the videos creates a fake context, which reduces the noise prediction accuracy in compression artifact removal. Moreover, mean square error (MSE)-based trained models perform better on low-bitrates than on high-bitrates due to the unbalanced parameter updates. This paper proposes an learning-based geometry compression artifact removal for V-PCC over a wide range of bitrates. Firstly, an occupancy map-based contextual feature extraction is proposed to eliminate the interference of empty pixels on the neighboring non-empty pixels. Secondly, an incremental Peak Signal to Noise Ratio (PSNR)-based training scheme is presented to balance the error differences. Experimental results show the effectiveness of the proposed method.
Jian Xiong 0005, Jiucheng Xie, Hao Gao 0005
ICASSP1
2024 DeformMLP: Dynamic Large-Scale Receptive Field MLP Networks for Human Motion Prediction
abstract
Predicting human motion requires addressing dependencies and errors for pose forecasting from sequences. The transformer’s self-attention aids this, but its complexity poses computational challenges. We present an efficient DeformMLP network without self-attention, using fully connected layers. DeformMLP includes DeformFCs, DeformFCt, and DeformFCst layers for spatial temporal modeling and calibration. DeformFCs capture semantics, DeformFCt learns relationships by summarizing time tokens, and DeformFCst assigns significance to dimensions to reduce computation. Our method balances efficiency and accuracy through decomposition and weight allocation. Evaluation on Human3.6M, 3DPW, CMU-MoCap datasets shows state-of-the-art prediction performance by benchmarks. The code is publicly available at https://github.com/HHT-98/DeformMLP.
Chi-Man Pun, Haolun Li 0001, Jian Xiong 0005, Hao Gao 0005
ICASSP5
2024 Towards Distortion-Debiased Blind Image Quality Assessment
abstract
Existing blind image quality assessment (BIQA) models are susceptible to biases related to distortion intensity and domain. Intensity bias refers to the relatively accurate perception of severe distortions but larger estimation errors for mild distortions, while domain bias stems from the discrepancies between synthetic and authentic distortion properties. This work introduces a unified learning framework towards addressing these distortion biases. We integrate distortion perception and restoration methods to mitigate intensity bias, where images with minor distortions, which are easily restorable, serve as references for mildly distorted images, while severe distortions benefit directly from distortion perception. The restoration modules employ a combined image-level and feature-level denoising approach, and then an intensity-aware cross-attention mechanism is designed for adaptive handling of intensity bias. To tackle domain bias, we introduce a distortion domain recognition task based on the intrinsic differences between distortion domains and use intra-domain similarity for weighting the quality scores from these domains. Experimental results show that the proposed method achieves state-of-the-art performance on multiple synthetic and authentic distortion datasets. Code and models will be available at https://github.com/xxVENTAZEDxx/Distortion-Debiased-BIQA
Lize Zhou, Jian Xiong 0005, Xianzhong Long, Hao Gao 0005
ACM Multimedia3
2023 An Optimized-Skeleton-Based Parkinsonian Gait Auxiliary Diagnosis Method with Both Monitoring Indicators and Assisted Ratings
abstract
Abnormal gait is one of the indispensable diagnostic sources of Parkinson’s disease (PD) diagnosis, typically presenting as small shuffling steps and gait bradykinesia. However, its diagnostic accuracy is lower due to the subjective judgments of doctors. To assist in improving the accuracy and reducing the subjectivity of the doctors, we propose an optimized-skeleton-based Parkinsonian gait auxiliary diagnosis method with both monitoring indicators and assisted ratings. By inputting a patient gait video captured from the side, our PD symptom-applicable pose trajectory model will extract a more precise and stable 2D skeleton sequence of patients. Next, the sequence will be used to calculate our proposed five monitoring indicators: gait frequency, ankle speed, whole speed, ankle angle speed, ankle acceleration, and previous work indicators: arm swing angle, leg angle, two feet x-axis distance to record the patient’s gait details at every moment. The extracted gait frequency will then be input into a random forest model to obtain the gait rating. Lastly, doctors can make more accurate judgments by referring to our objective monitoring indicators and assisted ratings. Experimental results show that our monitored indicators improve the doctors’ diagnosis accuracy by 16%, the skeleton speed and acceleration error of our optimized-skeleton extraction method achieve 4.18 cm/s and 5.71 cm/s2, and our random forest model has reached a classification accuracy of 95.8%.
Gaoqi Li, Chi-Man Pun, Haolun Li 0001, Jian Xiong 0005, Feng Xu 0005, Hao Gao 0005
BIBM4
2023 Learning Hybrid Representations of Semantics and Distortion for Blind Image Quality Assessment
abstract
Recently, some studies have shown that semantic and distortion representations both benefit the evaluation of image quality. However, the images of existing synthetic distortion databases are annotated with subjective quality scores and distortion types, lacking labels with semantic objects. Therefore, it is virtually infeasible to learn the representations of image semantics and distortion by co-guiding with semantic and distortion labels. To address this issue, we propose a dual-perception network (DPNet) via an end-to-end multi-task learning method, where knowledge distillation is lever-aged as a semantic label-free strategy. Specifically, semantic representation derived from pre-trained ResNet152 is applied to supervise the output of DPNet, while the output is utilized to construct a distortion recognition task. In this way, image semantics and distortion can be hybridly represented in an identical feature map. Finally, image quality is regressed based on the hybrid representations. Experimental results conducted on five benchmark databases validate that the proposed method can achieve state-of-the-art performance.
Jian Xiong 0005, Jin-Li Suo, Hao Gao 0005
ICASSP2
2023 ψ-Net: Point Structural Information Network for No-Reference Point Cloud Quality Assessment
abstract
The human vision system is highly adapted to extract structural information from the viewed scenes. The irregularity of point clouds makes the extraction of structural information containing both color and geometry an important challenge for point cloud quality assessment (PCQA). This paper proposes a point structural information (PSI) network (ψ-Net) for no-reference PCQA. Firstly, a PSI module is proposed to map the position vectors of neighboring points to weights for the calculation of color and geometric structure information. Secondly, a dual-stream network is presented to introduce distortion-related features for PCQA. Experimental results show the effectiveness of the proposed method.
Jian Xiong 0005, Jin-Li Suo, Hao Gao 0005
ICASSP1
2023 Non-Local Geometry and Color Gradient Aggregation Graph Model for No-Reference Point Cloud Quality Assessment
abstract
No-Reference point cloud quality assessment (NR-PCQA) is a challenging task in computer vision due to the irregularity of point cloud structures and the unavailability of reference information. Existing point-based and projection-based NR-PCQA models are limited by the representation of point cloud distortion and the modeling of spatial topological structure. To address these limitations, we first propose two visual quality-related gradients: local-maximum geometry gradient and distance-weighted color gradient, which can effectively represent local variations in terms of spatial structure and color intensities between adjacent points. We further propose a non-local geometry and color gradient aggregation graph model for evaluating the perceptual quality of point clouds. Specifically, local graph convolutions are designed to model the topological relationship across neighboring points by aggregating the geometry and color gradients. Furthermore, a position-adaptive self-attention mechanism is introduced to expand the receptive field for modeling the global dependencies of point clouds. Experimental results on two benchmark databases demonstrate that the proposed model outperforms existing state-of-the-art methods.
Hao Gao 0005, Jian Xiong 0005
ACM Multimedia4
2023 Low-Light Images In-the-Wild: A Novel Visibility Perception-Guided Blind Quality Indicator
abstract
Owing to the increasing deployment of CMOS camera modules, it is inevitable to take photographs under weak illumination. Therefore, low-light imaging quality is one of the most important factors affecting user experience as well as the product values of consumer electronics, automobile, surveillance, factory automation, and other industrial applications. Inspired by human vision, this article jointly considersvisibility perception,luminosity cognition, andcolor sensationand presents a new visibility perception-guided blind quality indicator for low-light images in-the-wild. To excavate effective descriptors for authentic distortions under weak illumination, we utilize maximum ignorable visible difference to characterize the reduced visibility, and employ the luminance statistical properties and color sensation characteristics to represent brightness and colorfulness distortions. Extensive experimental results on the benchmark dataset verify that the proposed blind quality indicator outperforms nine representative methods including general-purpose and distortion-specific methods.
Miaohui Wang, Jian Xiong 0005, Wuyuan Xie
IEEE Trans. Ind. Informatics3
2023 Visual Interaction Perceptual Network for Blind Image Quality Assessment
abstract
In observing images, the perception of the human visual system (HVS) is affected by both image contents and distortions. Obviously, the visual quality of the same image varies under different distortion types and intensities. Furthermore, the visual masking effects reveal that image content and distortion have a visual interaction, where the HVS presents different visibility of the identical distortion for different image contents. Based upon this, we propose a visual interaction perceptual network that can perceive both content and distortion of an image. The proposed model consists of three sub-modules: content perception module (CPM), distortion perception module (DPM), and visual interaction module (VIM). However, the subjective quality score cannot guide the model to explicitly learn the feature representations of image content and distortion. Thus, we perform a two-stage training procedure. In the first stage, we obtain CPM and DPM, where semantic features are extracted to recognize the image content in CPM, and distortion features are extracted to capture the image distortion type and intensity in DPM. In the second stage, the VIM is applied to model the interaction between semantic and distortion features, and the final predicted quality score is given by a fully connected layer. Experimental results demonstrate that the proposed method can achieve state-of-the-art performance on multiple benchmark databases, e.g., CSIQ, TID2013, KADID-10K, and KonIQ-10K.
Jian Xiong 0005, Weisi Lin
IEEE Trans. Multim.2
2023 Efficient Geometry Surface Coding in V-PCC
abstract
In recent video-based point cloud compression (V-PCC), 3D point clouds are projected onto 2D images and compressed by High-Efficiency Video Coding (HEVC). However, HEVC was originally designed for natural visual signals, which is a suboptimal framework for point clouds. Therefore, there are still problems in geometry information compression in V-PCC: (1) The distortion based on the sum of squared error (SSE) in the existing rate-distortion optimization (RDO) is inconsistent with the geometric quality measurement; (2) The existing prediction cannot explore the fixed relationship between the corresponding far layer and near layer depth, which means that the far layer depth can be always not less than the corresponding near layer depth. In this paper, we present an efficient geometry surface coding (EGSC) method for V-PCC to address the problems. Firstly, an error projection (EP) model is designed to establish the relationship between the SSE-based distortion and the geometry quality metric. Secondly, an EP-based RDO is employed to improve the geometry information compression by estimating the point normals with gradients. Finally, an occupancy-map driven scheme is proposed to improve the prediction accuracy of merge modes. Experimental results show that the proposed method achieves an average of over 10% bit-rate saving compared with the V-PCC reference software.
Jian Xiong 0005, Hao Gao 0005, Miaohui Wang, Hongliang Li 0001, King Ngi Ngan, Weisi Lin
IEEE Trans. Multim.1
2022 Encrypted Image Visual Security Index via Non-Local Recognizable Degree Evaluation
abstract
With the development of perceptual image encryption techniques, visual security evaluations of perceptually encrypted images have gained much attention. Current visual security indices isolate neither global nor local visual security evaluations. They ignore the fact that humans can reorganize partial local information to infer global information and local information can possibly be leaked in any position of the encrypted image. To address this problem, we propose a block-based non-local recognizable degree measure with a global structure similarity measure as a visual security index. In our index, the non-local searching strategy is utilized to capture leaked local information in any position of an encrypted image. The recognizable degree is evaluated by appearance recognizability and spatial structural recognizability involving human visual properties. In addition, weighted Minkowski pooling is adopted to evaluate the overall recognizable degree of all blocks depending on highly recognizable blocks. Furthermore, this overall recognizable degree is adjusted by global structure similarity, which is used to alleviate the global region structure distortion problem induced by the block partition. Experimental results demonstrate the sharpness and good robustness of our proposed index on different databases and various encryption types.
Jian Xiong 0005
ICASSP2
2022 Occupancy Map Guided Fast Video-Based Dynamic Point Cloud Coding
abstract
In video-based dynamic point cloud compression (V-PCC), 3D point clouds are projected into patches, and then the patches are padded into 2D images suitable for the video compression framework. However, the patch projection-based method produces a large number of empty pixels; the far and near components are projected to generate different 2D images (video frames), respectively. As a result, the generated video is with high resolutions and double frame rates, so the V-PCC has huge computational complexity. This paper proposes an occupancy map guided fast V-PCC method. Firstly, the relationship between the prediction coding and block complexity is studied based on a local linear image gradient model. Secondly, according to the V-PCC strategies of patch projection and block generation, we investigate the differences of rate-distortion characteristics between different types of blocks, and the temporal correlations between the far and near layers. Finally, by taking advantage of the fact that occupancy maps can explicitly indicate the block types, we propose an occupancy map guided fast coding method, in which coding is performed on the different types of blocks. Experiments have tested typical dynamic point clouds, and shown that the proposed method achieves an average 43.66% time-saving at the cost of only 0.27% and 0.16% Bjontegaard Delta (BD) rate increment under the geometry Point-to-Point (D1) error and attribute Luma Peak-Signal-Noise-Ratio (PSNR), respectively.
Jian Xiong 0005, Hao Gao 0005, Miaohui Wang, Hongliang Li 0001, Weisi Lin
IEEE Trans. Circuits Syst. Video Technol.1
2022 Perceptually Quasi-Lossless Compression of Screen Content Data Via Visibility Modeling and Deep Forecasting
abstract
Screen content data, such as computer-generated photographs, desktop sharing, remote education, video game streaming and screenshot, is one of the most popular visual information carriers in Internet of Video Things. Although lossless compression can guarantee high quality of service for these screen content based industrial applications, it also causes considerable storage space and transmission bandwidth issues. To alleviate these challenges, in this article, we present a visually quasi-lossless coding approach to control the compression distortion belowvisibility thresholdin the human visual system. Specifically, to better quantify the visual redundancy for screen content data, a newvisibility thresholdmethod is designed by incorporating blur sensitivity and oblique correction effects. Then, an end-to-end mapping between thevisibility thresholdand quality control factor is learned and represented as a deep convolutional neural network. The experimental results demonstrate that the proposed method saves the average encoding bits up to 23.15% compared with the latest scheme under the same perceptual quality.
Miaohui Wang, Zhuowei Xu, Jian Xiong 0005, Wuyuan Xie
IEEE Trans. Ind. Informatics4
2022 Objective Object Segmentation Visual Quality Evaluation: Quality Measure and Pooling Method
abstract
Objective object segmentation visual quality evaluation is an emergent member of the visual quality assessment family. It aims to develop an objective measure instead of a subjective survey to evaluate the object segmentation quality in agreement with human visual perception. It is an important benchmark for assessing and comparing the performances of object segmentation methods in terms of visual quality. Despite its essential role, sufficient study compared with other visual quality evaluation studies is still lacking. In this article, we propose a novel full-reference objective measure that includes a two-level single object segmentation visual quality measure and a pooling method for multiple object segmentation overall visual quality. The single object segmentation visual quality measure combines a pixel-level sub-measure and a region-level sub-measure for evaluating the similarity of area, shape, and object completeness between the segmentation result and the ground truth in terms of human visual perception. For the proposed multiple object segmentation overall visual quality pooling method, the rank of each object’s segmentation quality as a novel factor is integrated into the weighted harmonic mean to evaluate the overall quality. To evaluate the performance of our proposed measure, we tested it on an object segmentation subjective visual quality assessment database. The experimental results demonstrate that our proposed two-level measure and pooling method with good robustness perform better in matching subjective assessments compared with other state-of-the-art objective measures.
King Ngi Ngan, Jian Xiong 0005
ACM Trans. Multim. Comput. Commun. Appl.4
2021 Machine Learning-Based Rate Distortion Modeling for VVC/H.266 Intra-Frame
abstract
Rate-distortion (R-D) optimization has been widely adopted to improve the coding efficiency on the video encoder side. However, there are few studies related to modeling the R-D characteristics of the latest Versatile Video Coding (VVC) reference software. In this paper, we investigate the R-D modeling of the intra-frame on the VVC encoder by utilizing four traditional machine learning algorithms. We extract four highly descriptive features to capture the relationship between the video content and the R-D model. Moreover, it is applied to the initial intra-frame rate control of VVC. Experimental results show that our method outperforms VTM-7.0, which improves the accuracy by up to 8.65% with affordable computational complexity increase1.
Miaohui Wang, Lirong Huang, Jian Xiong 0005
ICME4
2021 An improved artificial bee colony algorithm based on elite search strategy with segmentation application on robot vision system
abstract
Summary Aiming at accelerating the convergence speed and enhancing relative poor local search ability of the traditional artificial bee colony algorithm (ABC), this article introduces an ABC with a new elite search strategy. First, we propose a strategy of recording individuals with high performance. Then bees have more chances to learn from a real elite. In the onlooked bee phase, its updating equation is changed for having more opportunities to search in a valuable area. Furthermore, for saving the value of function evaluations, a new learning equation for the best onlooked bee is proposed. The image segmentation of a robot binocular stereo vision system is a key problem in mechanical robot vision system, but the computation time limits its application. The experimental results show that the proposed algorithm achieves better performance on 10 benchmark functions and the image segmentation problem of mechanical robot in comparison with several other state of the art algorithms.
Chuyi Gao, Maolong Xi, Jian Xiong 0005, Chi-Man Pun, Hao Gao 0005
Concurr. Comput. Pract. Exp.6
2021 Robust automated graph regularized discriminative non-negative matrix factorization
Xianzhong Long, Jian Xiong 0005
Multim. Tools Appl.2
2020 Graph Learning Regularized Non-negative Matrix Factorization for Image Clustering
Xianzhong Long, Jian Xiong 0005, Yun Li 0009
ICONIP (5)2
2020 Objective object segmentation visual quality evaluation based on pixel-level and region-level characteristics
abstract
Objective object segmentation visual quality evaluation is an emergent member of the visual quality assessment family. It aims at developing an objective measure instead of a subjective survey to evaluate the object segmentation quality in agreement with human visual perception. It is an important benchmark to assess and compare performances of object segmentation methods in terms of the visual quality. In spite of its essential role, it still lacks of sufficient studying compared with other visual quality evaluation researches. In this paper, we propose a novel full-reference objective measure including a pixel-level sub-measure and a region-level sub-measure. For the pixel-level sub-measure, it assigns proper weights to not only false positive pixels and false negative pixels but also true positive pixels according to their certainty degrees. For the region-level sub-measure, it considers location distribution of the false negative errors and correlations among neighboring pixels. Thus, by combining these two sub-measures, our measure can evaluate similarity of area, shape and object completeness between one segmentation result and its ground truth in terms of human visual perception. In order to evaluate the performance of our proposed measure, we tested it on an object segmentation subjective visual quality assessment database. The experimental results demonstrate that our proposed measure with good robustness performs better in matching subjective assessments compared with other state-of-the-art objective measures.
Jian Xiong 0005
MMAsia2
2020 New multi-view human motion capture framework
abstract
Estimating human pose and shape without markers is a challenging problem. This study proposes a multiple‐view markerless human motion capture framework. Firstly, a multi‐view camera system is built for capturing real‐time images of moving humans on multiple views. Secondly, by employing the OpenPose method, the authors calculate robust 3D key points from 2D key points of the human body, which are estimated from the multi‐view images. And dense 3D point cloud is reconstructed from images. Thirdly, they propose a novel SMPL‐based method to represent human motion by fitting the SMPL model to 3D key points and 3D point clouds. In order to achieve a more accurate human pose, a penalty term is utilised to solve the problem of error accumulation in the process of human motion capture. In addition, they present a dense mesh template‐based SMPL that can be deformed to point cloud to recover a real human body shape. Finally, they map multi‐view colour images onto the human mesh model to acquire rendered mesh. The experimental results show that the proposed method improves the accuracy of human pose and realises the 3D human body model more realistic.
Feiyi Xu, Chi-Man Pun, Wenqi Xiao, Jianhui Nie, Jian Xiong 0005, Hao Gao 0005, Feng Xu 0005
IET Image Process.6
2020 Rate Constrained Multiple-QP Optimization for HEVC
abstract
In High Efficiency Video Coding (HEVC), multiple-QP (quantization parameter) optimization can adapt to a local video content. However, the multiple-QP implementation in the HEVC reference software (HM 16.6) achieves the best QP value for each coding block with a large amount of computational complexity. To address this challenge, we propose a fast rate-constrained multiple-QP optimization approach for the HM platform. We first introduce a template-based transform coefficient selection method which can save the overall complexity of entropy coding. In addition, we model the multiple-QP determination as a new rate-constrained optimization problem, and finally, we get a feasible solution with a lower computation overhead. Experimental results show that our method dramatically reduces the average complexity under the all-intra, low-delay and random-access configuration.
Miaohui Wang, Jian Xiong 0005, Long Xu 0001, Wuyuan Xie, King Ngi Ngan, Harry Qin
IEEE Trans. Multim.2
2019 Lednet: A Lightweight Encoder-Decoder Network for Real-Time Semantic Segmentation
abstract
The extensive computational burden limits the usage of CNNs in mobile devices for dense estimation tasks. In this paper, we present a lightweight network to address this problem, namely LEDNet, which employs an asymmetric encoder-decoder architecture for the task of real-time semantic segmentation. More specifically, the encoder adopts a ResNet as backbone network, where two new operations, channel split and shuffle, are utilized in each residual block to greatly reduce computation cost while maintaining higher segmentation accuracy. On the other hand, an attention pyramid network (APN) is employed in the decoder to further lighten the entire network complexity. Our model has less than 1M parameters, and is able to run at over 71 FPS in a single GTX 1080Ti GPU. The comprehensive experiments demonstrate that our approach achieves state-of-the-art results in terms of speed and accuracy trade-off on CityScapes dataset.
Yu Wang 0109, Quan Zhou 0004, Jian Xiong 0005, Guangwei Gao, Xiaofu Wu, Longin Jan Latecki
ICIP4
2019 ESNet: An Efficient Symmetric Network for Real-Time Semantic Segmentation
Yu Wang 0109, Quan Zhou 0004, Jian Xiong 0005, Xiaofu Wu, Xin Jin 0015
PRCV (2)3
2018 SHAFA: sparse hybrid adaptive filtering algorithm to estimate channels in various SNR environments
abstract
The ‐norm penalised (LP) normalised least mean square algorithm converges faster than the LP normalised least mean fourth algorithm does, but the latter can achieve better steady‐state performance, particularly in regions with low signal‐to‐noise ratios (SNRs). To simultaneously take advantage of both merits, a sparse hybrid adaptive filtering algorithm is proposed in various SNR environments. Specifically, the authors construct a cost function that uses the statistical error term and sparse penalty term. The first term is designed by a hybrid error function of the second‐ and fourth‐order statistical errors, respectively, and the second term is obtained using a sparse constraint function. The hybrid error term can be easily balanced by a proportional parameter . Moreover, they devise a non‐uniform step size in the proposed algorithm to further balance the convergence speed and estimation error. Simulation results are provided to validate the proposed algorithm in various SNR environments.
Jie Wang 0024, Jie Yang 0027, Jian Xiong 0005, Hikmet Sari, Guan Gui 0001
IET Commun.3
2015 Fast HEVC Inter CU Decision Based on Latent SAD Estimation
abstract
The emerging high efficiency video coding (HEVC) standard has improved compression performance significantly in comparison with H.264/AVC. However, more intensive computational complexity has been introduced by adopting a number of new coding tools. In this paper, a fast inter CU decision is proposed based on the latent sum of absolute differences (SAD) estimation. Firstly, a two-layer motion estimation (ME) method is designed to take advantage of the latent SAD cost. The new ME method can obtain the SAD costs for both the upper CU and its sub-CUs. Secondly, a concept of motion compensation rate- distortion (R-D) cost is defined, and an exponential model is proposed to express the relationship between the motion compensation R-D cost and the SAD cost. Then, a fast CU decision approach is designed based on the exponential model. The fast CU decision is implemented by comparing a derived threshold with the SAD cost difference between the upper and sub SAD costs. Experimental results show that the proposed algorithm achieves an average of 52% and 58.4% reductions of the coding time at the cost of 1.61% and 2% bit-rate increases under the low delay and random access conditions, respectively.
Jian Xiong 0005, Hongliang Li 0001, Fanman Meng, Qingbo Wu 0001, King Ngi Ngan
IEEE Trans. Multim.1
2014 Fast and efficient inter CU decision for high efficiency video coding
abstract
In this paper, a graph cut based fast Coding Unit (CU) decision algorithm is proposed for HEVC inter frames. Firstly, a feature called pyramid variance of the absolute difference (PVAD) is designed for the CU selection. Secondly, the CU decision is modeled as a Markov Random Field (MRF) inference problem, which can be optimized by the graph cut algorithm. Thirdly, a maximum a posteriori (MAP) approach based on the R-D cost is conducted to evaluate whether the unsplit CUs should be further split or not. Experimental results show the effectiveness of the proposed method.
Jian Xiong 0005, Hongliang Li 0001, Fanman Meng, Bing Zeng 0001, Shuyuan Zhu, Qingbo Wu 0001
ICIP1
2014 MRF-Based Fast HEVC Inter CU Decision With the Variance of Absolute Differences
abstract
The newly developed High Efficiency Video Coding (HEVC) Standard has improved video coding performance significantly in comparison to its predecessors. However, more intensive computation complexity is introduced by implementing a number of new coding tools. In this paper, a fast coding unit (CU) decision based on Markov random field (MRF) is proposed for HEVC inter frames. First, it is observed that the variance of the absolute difference (VAD) is proportional with the rate-distortion (R-D) cost. The VAD based feature is designed for the CU selection. Second, the decision of CU splittings is modeled as an MRF inference problem, which can be optimized by the Graphcut algorithm. Third, a maximum a posteriori (MAP) approach based on the R-D cost is conducted to evaluate whether the unsplit CUs should be further split or not. Experimental results show that the proposed algorithm can achieve about 53% reduction of the coding time with negligible coding performance degradation, which outperforms the state-of-the-art algorithms significantly.
Jian Xiong 0005, Hongliang Li 0001, Fanman Meng, Shuyuan Zhu, Qingbo Wu 0001, Bing Zeng 0001
IEEE Trans. Multim.1
2014 A Fast HEVC Inter CU Selection Method Based on Pyramid Motion Divergence
abstract
The newly developed HEVC video coding standard can achieve higher compression performance than the previous video coding standards, such as MPEG-4, H.263 and H.264/AVC. However, HEVC's high computational complexity raises concerns about the computational burden on real-time application. In this paper, a fast pyramid motion divergence (PMD) based CU selection algorithm is presented for HEVC inter prediction. The PMD features are calculated with estimated optical flow of the downsampled frames. Theoretical analysis shows that PMD can be used to help selecting CU size. A k nearest neighboring like method is used to determine the CU splittings. Experimental results show that the fast inter prediction method speeds up the inter coding significantly with negligible loss of the peak signal-to-noise ratio.
Jian Xiong 0005, Hongliang Li 0001, Qingbo Wu 0001, Fanman Meng
IEEE Trans. Multim.1
2013 Mode dependent loop filter for intra prediction coding in H.264/AVC
Qingbo Wu 0001, Linfeng Xu 0001, Liaoyuan Zeng, Jian Xiong 0005
J. Vis. Commun. Image Represent.4