Huan Yang 0001

dblp:86/4843-1 · DBLP profile ↗
← Back
37ranked-venue papers
5as first author
21since 2021 · last 2026
0000-0001-5810-0248ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 23 · 3 first-author · 13 since 2021Computer networks · 8 · 5 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Automatic detection of dolphin click signals based on acoustic spectrogram decoupling and fusion
Zongwei Liu, Xiaoke Liu, Yaqian Shi, Baoxiang Huang, Huan Yang 0001
Multim. Syst.7
2026 SpineQFormer: ROI-driven feature fusion and multi-angular regression transformer for spinal image quality assessment
Jianpeng Chen, Yukun Du, Changlin Lv, Yongming Xi, Huan Yang 0001
Signal Process. Image Commun.8
2025 Advancements in medical image quality assessment for spinal CT and MRI images
Yushuo Ling, Yukun Du, Yongming Xi, Huan Yang 0001
J. Vis. Commun. Image Represent.7
2025 Toward a blind quality assessment for underwater images
Guojia Hou, Kunqian Li, Weidong Zhang 0007, Huan Yang 0001, Zhenkuan Pan 0001
Signal Process. Image Commun.5
2024 Probabilistic Model-Based Reinforcement Learning Unmanned Surface Vehicles Using Local Update Sparse Spectrum Approximation
abstract
In this article, we focus on the computational efficiency of probabilistic model-based reinforcement learning (MBRL) in unmanned surface vehicles (USV) under unforeseeable and unobservable external disturbances. A novel MBRL approach, local update spectrum probabilistic model predictive control (LUSPMPC), is proposed to fully release the superiority of the probabilistic model approximated in the frequency domain in computational efficiency while mitigating its risk of overfitting during the learning procedure. It employs a local update strategy to relieve the violation of Bochner's theory, and a frequency clipping trick to encourage the approximated model to focus on the features in the low-frequency domain. Evaluated by the position-keeping task in a real USV data-driven simulation, LUSPMPC shows its significant advantages in computational efficiency while achieving better learning capability, generalization capability, and control performances in a wide range of sparse scales compared with the baseline MBRL approaches that approximate their models in sample space and frequency domain, and therefore becomes an appealing solution for MBRL USV system defending against rapidly changing ocean disturbances.
Yunduan Cui, Huan Yang 0001, Cuiping Shao, Lei Peng 0002, Huiyun Li
IEEE Trans. Ind. Informatics3
2024 3DTA: No-Reference 3D Point Cloud Quality Assessment With Twin Attention
abstract
Point clouds are rapidly gaining popularity in many practical applications, and point cloud quality assessment (PCQA) is an important research topic that helps us measure and improve the visual experience in applications using point clouds. Research on full-reference (FR) PCQAs has recently made impressive progress, and research on no-reference (NR) PCQAs has also gradually increased. However, the performance of the prior NR PCQA methods still suffers from weak generalization ability and lower accuracy than the FR metrics in general. In this work, we propose a two-stage sampling method that can reasonably represent a whole point cloud, making it possible to efficiently calculate the point cloud quality. For quality prediction, we designed a twin-attention-based transformer PCQA model (3DTA), which uses the data of the two-stage sampling method as input and directly outputs the predicted quality score. Our model is accurate and widely applicable, and it has a simple and flexible structure. Experimental results show that in most cases, the proposed 3DTA model substantially outperforms the benchmark NR methods. The accuracy of the proposed method is competitive even against that of the FR method, which makes 3DTA a strong candidate for the PCQA task, regardless of the reference availability. The code of the proposed model is publicly available athttps://github.com/philox12358/3DTA-PCQA.
Linxia Zhu, Xu Wang 0006, Honglei Su, Huan Yang 0001, Hui Yuan 0001, Jari Korhonen
IEEE Trans. Multim.5
2024 Incentivizing Massive Unknown Workers for Budget-Limited Crowdsensing: From Off-Line and On-Line Perspectives
abstract
How to incentivize strategic workers using limited budget is a very fundamental problem for crowdsensing systems; nevertheless, since the sensing abilities of the workers may not always be known as prior knowledge due to the diversities of their sensor devices and behaviors, it is difficult to properly select and pay the unknown workers. Although the uncertainties of the workers can be addressed by the standardCombinatorial Multi-Armed Bandit(CMAB) framework in existing proposals through a trade-off between exploration and exploitation, we may not have sufficient budget to enable the trade-off among the individual workers, especially when the number of the workers is huge while the budget is limited. Moreover, the standard CMAB usually assumes the workers always stay in the system, whereas the workers may join in or depart from the system over time, such that what we have learnt for an individual worker cannot be applied after the worker leaves. To address the above challenging issues, in this paper, we first propose an off-lineContext-Aware CMAB-based Incentive(CACI) mechanism. We innovate in leveraging the exploration-exploitation trade-off in an elaborately partitioned context space instead of the individual workers, to effectively incentivize the massive unknown workers with a very limited budget. We also extend the above basic idea to the on-line setting where unknown workers may join in or depart from the systems dynamically, and propose an on-line version of the CACI mechanism. Specifically, by the exploitation-exploration trade-off in the context space, we learn to estimate the sensing ability of any unknown worker (even it never appeared in the system before) according to its context information. We perform rigorous theoretical analysis to reveal the upper bounds on the regrets of our CACI mechanisms and to prove their truthfulness and individual rationality, respectively. Extensive experiments on both synthetic and real datasets are also conducted to verify the efficacy of our mechanisms.
Feng Li 0002, Yuqi Chai, Huan Yang 0001, Pengfei Hu 0001, Lingjie Duan
IEEE/ACM Trans. Netw.3
2024 VibHead: An Authentication Scheme for Smart Headsets through Vibration
abstract
Recent years have witnessed the fast penetration of Virtual Reality (VR) and Augmented Reality (AR) systems into our daily life, the security and privacy issues of the VR/AR applications have been attracting considerable attention. Most VR/AR systems adopt head-mounted devices (i.e., smart headsets) to interact with users and the devices usually store the users’ private data. Hence, authentication schemes are desired for the head-mounted devices. Traditional knowledge-based authentication schemes for general personal devices have been proved vulnerable to shoulder-surfing attacks, especially considering the headsets may block the sight of the users. Although the robustness of the knowledge-based authentication can be improved by designing complicated secret codes in virtual space, this approach induces a compromise of usability. Another choice is to leverage the users’ biometrics; however, it either relies on highly advanced equipments which may not always be available in commercial headsets or introduce heavy cognitive load to users. In this paper, we propose a vibration-based authentication scheme, VibHead, for smart headsets. Since the propagation of vibration signals through human heads presents unique patterns for different individuals, VibHead employs a CNN-based model to classify registered legitimate users based the features extracted from the vibration signals. We also design a two-step authentication scheme where the above user classifiers are utilized to distinguish the legitimate user from illegitimate ones. We implement VibHead on a Microsoft HoloLens equipped with a linear motor and an IMU sensor which are commonly used in off-the-shelf personal smart devices. According to the results of our extensive experiments, with short vibration signals (≤ 1s ), VibHead has an outstanding authentication accuracy; both FAR and FRR are around 5%.
Feng Li 0002, Huan Yang 0001, Dongxiao Yu, Yuanfeng Zhou, Yiran Shen 0001
ACM Trans. Sens. Networks3
2024 PDSR: A Privacy-Preserving Diversified Service Recommendation Method on Distributed Data
abstract
The last decade has witnessed a tremendous growth of service computing, while efficient service recommendation methods are desired to recommend high-quality services to users. It is well known that collaborative filtering is one of the most popular methods for service recommendation based on QoS, and many existing proposals focus on improving recommendation accuracy, i.e., recommending high-quality redundant services. Nevertheless, users may have different requirements on QoS, and hence diversified recommendation has been attracting increasing attention in recent years to fulfill users’ diverse demands and to explore potential services. Unfortunately, the recommendation performances relies on a large volume of data (e.g., QoS data), whereas the data may be distributed across multiple platforms. Therefore, to enable data sharing across the different platforms for diversified service recommendation, we propose aPrivacy-preserving Diversified Service Recommendation(PDSR) method. Specifically, we innovate in leveraging the Locality-Sensitive Hashing (LSH) mechanism such that privacy-preserved data sharing across different platforms is enabled to construct a service similarity graph. Based on the similarity graph, we propose a novel accuracy-diversity metric and design a 2-approximation algorithm to select$K$services to recommend by maximizing the accuracy-diversity measure. Extensive experiments on real datasets are conducted to verify the efficacy of our PDSR method.
Huan Yang 0001, Yiran Shen 0001, Chao Liu 0008, Lianyong Qi, Xiuzhen Cheng, Feng Li 0002
IEEE Trans. Serv. Comput.2
2023 A no-reference underwater image quality evaluator via quality-aware features
Huan Yang 0001, Guojia Hou
J. Vis. Commun. Image Represent.4
2023 Bitstream-Based Perceptual Quality Assessment of Compressed 3D Point Clouds
abstract
With the increasing demand of compressing and streaming 3D point clouds under constrained bandwidth, it has become ever more important to accurately and efficiently determine the quality of compressed point clouds, so as to assess and optimize the quality-of-experience (QoE) of end users. Here we make one of the first attempts developing a bitstream-based no-reference (NR) model for perceptual quality assessment of point clouds without resorting to full decoding of the compressed data stream. Specifically, we first establish a relationship between texture complexity and the bitrate and texture quantization parameters based on an empirical rate-distortion model. We then construct a texture distortion assessment model upon texture complexity and quantization parameters. By combining this texture distortion model with a geometric distortion model derived from Trisoup geometry encoding parameters, we obtain an overall bitstream-based NR point cloud quality model named streamPCQ. Experimental results show that the proposed streamPCQ model demonstrates highly competitive performance when compared with existing classic full-reference (FR) and reduced-reference (RR) point cloud quality assessment methods with a fraction of computational cost.
Honglei Su, Qi Liu 0029, Hui Yuan 0001, Huan Yang 0001, Zhenkuan Pan 0001, Zhou Wang 0001
IEEE Trans. Image Process.5
2023 UID2021: An Underwater Image Dataset for Evaluation of No-Reference Quality Assessment Metrics
abstract
Achieving subjective and objective quality assessment of underwater images is of high significance in underwater visual perception and image/video processing. However, the development of underwater image quality assessment (UIQA) is limited for the lack of publicly available underwater image datasets with human subjective scores and reliable objective UIQA metrics. To address this issue, we establish a large-scale underwater image dataset, dubbed UID2021, for evaluating no-reference (NR) UIQA metrics. The constructed dataset contains 60 multiply degraded underwater images collected from various sources, covering six common underwater scenes (i.e., bluish scene, blue-green scene, greenish scene, hazy scene, low-light scene, and turbid scene), and their corresponding 900 quality improved versions are generated by employing 15 state-of-the-art underwater image enhancement and restoration algorithms. Mean opinion scores with 52 observers for each image of UID2021 are also obtained by using the pairwise comparison sorting method. Both in-air and underwater-specific NR IQA algorithms are tested on our constructed dataset to fairly compare their performance and analyze their strengths and weaknesses. Our proposed UID2021 dataset enables ones to evaluate NR UIQA algorithms comprehensively and paves the way for further research on UIQA. The dataset is available at https://github.com/Hou-Guojia/UID2021 .
Guojia Hou, Huan Yang 0001, Kunqian Li, Zhenkuan Pan 0001
ACM Trans. Multim. Comput. Commun. Appl.3
2023 Collaborative Learning in General Graphs With Limited Memorization: Complexity, Learnability, and Reliability
abstract
We consider a$K$-armed bandit problem in general graphs where agents are arbitrarily connected and each of them has limited memorizing capabilities and communication bandwidth. The goal is to let each of the agents eventually learn the best arm. Although recent studies show the power of collaboration among the agents in improving the efficacy of learning, it is assumed in these studies that the communication graph should be complete or well-structured, whereas such an assumption is not always valid in practice. Furthermore, limited memorization and communication bandwidth also restrict the collaborations of the agents, since the agents memorize and communicate very few experiences. Additionally, an agent may be corrupted to share falsified experiences to its peers, while the resource limit in terms of memorization and communication may considerably restrict the reliability of the learning process. To address the above issues, we propose a three-staged collaborative learning algorithm. In each step, the agents share their latest experiences with each other through light-weight random walks in a general communication graph, and then make decisions on which arms to pull according to the recommendations received from their peers. The agents finally update their adoptions (i.e., preferences to the arms) based on the reward obtained by pulling the arms. Our theoretical analysis shows that, when there are a sufficient number of agents participating in the collaborative learning process, all the agents eventually learn the best arm with high probability, even with limited memorizing capabilities and light-weight communications. We also reveal in our theoretical analysis the upper bound on the number of corrupted agents our algorithm can tolerate. The efficacy of our proposed three-staged collaborative learning algorithm is finally verified by extensive experiments on both synthetic and real datasets.
Feng Li 0002, Xuyang Yuan, Huan Yang 0001, Dongxiao Yu, Weifeng Lyu, Xiuzhen Cheng
IEEE/ACM Trans. Netw.4
2022 Semantic-aware multi-task learning for image aesthetic quality assessment
abstract
In recent years, image aesthetic quality assessment has attracted considerable attention due to the massive growth of digital images in social platforms and the Internet. However, automatically assessing aesthetic quality of an image is a challenging task, because image aesthetic is affected by various factors, and the criteria for judging the aesthetic of images with diverse semantic information are different. To this end, a Semantic-Aware Multi-task convolution neural network (SAM-CNN) for evaluating image aesthetic quality is proposed in this paper. The network can fuse intermediate features of different layers at different scales in CNN to obtain a more comprehensive and accurate aesthetic expression, under the joint supervision of image aesthetic quality assessment task and semantic classification task in a multi-task learning manner. Besides, by applying the attention mechanism, semantic information with a large receptive field extracted from deep layers is utilised to guide the network to focus on the key parts of features to be fused, to improve the effectiveness of feature fusion. Experimental results on the AVA dataset and Photo.net dataset demonstrate the effectiveness and superiority of the proposed SAM-CNN.
Weiliang Yan, Huan Yang 0001, Baoxiang Huang, Zhenkuan Pan 0001
Connect. Sci.3
2022 IMLBoost for intelligent diagnosis with imbalanced medical records
abstract
Class imbalance of medical records is a critical challenge for disease classification in intelligent diagnosis. Existing machine learning algorithms usually assign equal weights to all classes, which may reduce classification accuracy of imbalanced records. In this paper, a new Imbalance Lessened Boosting (IMLBoost) algorithm is proposed to better classify imbalanced medical records, highlighting the contribution of samples in minor classes as well as hard and boundary samples. A tailored Cost-Fitting Loss (CFL) function is proposed to assign befitting costs to these critical samples. The first and second derivations of the CFL are then derived and embedded into the classical XGBoost framework. In addition, some feature analysis skills are utilized to further improve performance of the IMLBoost, which also can speed up the model training. Experimental results on five UCI imbalanced medical datasets have demonstrated the effectiveness of the proposed algorithm. Compared with other existing classification methods, IMLBoost has improved the classification performance in terms of F1-score, G-mean and AUC.
Xiaofan Chi, Yukun Du, Huan Yang 0001, Yongming Xi
Intell. Data Anal.4
2022 A no-Reference Stereoscopic Image Quality Assessment Network Based on Binocular Interaction and Fusion Mechanisms
abstract
In contemporary society full of stereoscopic images, how to assess visual quality of 3D images has attracted an increasing attention in field of Stereoscopic Image Quality Assessment (SIQA). Compared with 2D-IQA, SIQA is more challenging because some complicated features of Human Visual System (HVS), such as binocular interaction and binocular fusion, must be considered. In this paper, considering both binocular interaction and fusion mechanisms of the HVS, a hierarchical no-reference stereoscopic image quality assessment network (StereoIF-Net) is proposed to simulate the whole quality perception of 3D visual signals in human cortex, including two key modules: BIM and BFM. In particular, Binocular Interaction Modules (BIMs) are constructed to simulate binocular interaction in V2-V5 visual cortex regions, in which a novel cross convolution is designed to explore the interaction details in each region. In the BIMs, different output channel numbers are designed to imitate various receptive fields in V2-V5. Furthermore, a Binocular Fusion Module (BFM) with automatic learned weights is proposed to model binocular fusion of the HVS in higher cortex layers. The verification experiments are conducted on the LIVE 3D, IVC and Waterloo-IVC SIQA databases and three indices including PLCC, SROCC and RMSE are employed to evaluate the assessment consistency between StereoIF-Net and the HVS. The proposed StereoIF-Net achieves almost the best results compared with advanced SIQA methods. Specifically, the metric values on LIVE 3D, IVC and WIVC-I are the best, and are the second-best on the WIVC-II.
Jianwei Si, Baoxiang Huang, Huan Yang 0001, Weisi Lin, Zhenkuan Pan 0001
IEEE Trans. Image Process.3
2022 Attention Gate Based Dual-Pathway Network for Vertebra Segmentation of X-Ray Spine Images
abstract
Automatic spine and vertebra segmentation from X-ray spine images is a critical and challenging problem in many computer-aid spinal image analysis and disease diagnosis applications. In this paper, a two-stage automatic segmentation framework for spine X-ray images is proposed, which can firstly locate the spine regions (including backbone, sacrum and ilium) in the coarse stage and then identify eighteen vertebrae (i.e., cervical vertebra 7, thoracic vertebra 1-12 and lumbar vertebra 1-5) with isolate and clear boundary in the fine stage. A novel Attention Gate based dual-pathway Network (AGNet) composed of context and edge pathways is designed to extract semantic and boundary information for segmentation of both spine and vertebra regions. Multi-scale supervision mechanism is applied to explore comprehensive features and an Edge aware Fusion Mechanism (EFM) is proposed to fuse features extracted from the two pathways. Some other image processing skills, such as centralized backbone clipping, patch cropping and convex hull detection are introduced to further refine the vertebra segmentation results. Experimental validations on spine X-ray images dataset and vertebrae dataset suggest that the proposed AGNet achieves superior performance compared with state-of-the-art segmentation methods, and the coarse-to-fine framework can be implemented in real spinal diagnosis systems.
Tongshuai Xu, Huan Yang 0001, Yongming Xi, Yukun Du, Jinxu Li
IEEE J. Biomed. Health Informatics3
2021 A full-reference stereoscopic image quality assessment index based on stable aggregation of monocular and binocular visual features
abstract
Abstract In stereoscopic image quality assessment, human visual system has been universally taken into account to detect perceptual characteristics. A novel full‐reference stereoscopic image assessment metric by considering both monocular and binocular visual features of human visual system is proposed. In particular, a new region segmentation algorithm is firstly proposed to divide 3D images into occluded and non‐occluded regions. The just noticeable difference model is employed on the occluded regions to formulate the monocular vision, while the binocular just noticeable difference model is applied to the non‐occluded regions to reveal the binocular vision of the human visual system. In the proposed region segmentation, disparity information and Euclidean distance between stereo pairs are both adopted to solve the unstable segmentation problem of traditional methods. A new pooling strategy based on global edge features is then presented to aggregate the just noticeable difference and binocular just noticeable difference evaluation maps. In addition, some local image features as supplementary of just noticeable difference to describe visual characteristics of the human visual system are also extracted. Finally, an overall quality score is calculated based on the above‐mentioned features to measure the visual quality of distorted stereo pairs. Experimental results show that the proposed metric achieves high consistency with the human visual system, and outperforms state‐of‐the‐art algorithms on stereoscopic image quality assessment.
Jianwei Si, Huan Yang 0001, Baoxiang Huang, Zhenkuan Pan 0001, Honglei Su
IET Image Process.2
2021 PQA-Net: Deep No Reference Point Cloud Quality Assessment via Multi-View Projection
abstract
Recently, 3D point cloud is becoming popular due to its capability to represent the real world for advanced content modality in modern communication systems. In view of its wide applications, especially for immersive communication towards human perception, quality metrics for point clouds are essential. Existing point cloud quality evaluations rely on a full or certain portion of the original point cloud, which severely limits their applications. To overcome this problem, we propose a novel deep learning-based no reference point cloud quality assessment method, namely PQA-Net. Specifically, the PQA-Net consists of a multi-view-based joint feature extraction and fusion (MVFEF) module, a distortion type identification (DTI) module, and a quality vector prediction (QVP) module. The DTI and QVP modules share the feature generated from the MVFEF module. By using the distortion type labels, the DTI and the MVFEF modules are first pre-trained to initialize the network parameters, based on which the whole network is then jointly trained to finally evaluate the point cloud quality. Experimental results on the Waterloo Point Cloud dataset show that PQA-Net achieves better or equivalent performance comparing with the state-of-the-art quality assessment methods. The code of the proposed model will be made publicly available to facilitate reproducible researchhttps://github.com/qdushl/PQA-Net.
Qi Liu 0029, Hui Yuan 0001, Honglei Su, Hao Liu 0044, Yu Wang 0106, Huan Yang 0001, Junhui Hou
IEEE Trans. Circuits Syst. Video Technol.6
2021 Reduced Reference Perceptual Quality Model With Application to Rate Control for Video-Based Point Cloud Compression
abstract
In rate-distortion optimization, the encoder settings are determined by maximizing a reconstruction quality measure subject to a constraint on the bitrate. One of the main challenges of this approach is to define a quality measure that can be computed with low computational cost and which correlates well with the perceptual quality. While several quality measures that fulfil these two criteria have been developed for images and videos, no such one exists for point clouds. We address this limitation for the video-based point cloud compression (V-PCC) standard by proposing a linear perceptual quality model whose variables are the V-PCC geometry and color quantization step sizes and whose coefficients can easily be computed from two features extracted from the original point cloud. Subjective quality tests with 400 compressed point clouds show that the proposed model correlates well with the mean opinion score, outperforming state-of-the-art full reference objective measures in terms of Spearman rank-order and Pearson linear correlation coefficient. Moreover, we show that for the same target bitrate, rate-distortion optimization based on the proposed model offers higher perceptual quality than rate-distortion optimization based on exhaustive search with a point-to-point objective quality metric. Our datasets are publicly available at https://github.com/qdushl/Waterloo-Point-Cloud-Database-2.0.
Qi Liu 0029, Hui Yuan 0001, Raouf Hamzaoui, Honglei Su, Junhui Hou, Huan Yang 0001
IEEE Trans. Image Process.6
2021 Full-reference Screen Content Image Quality Assessment by Fusing Multilevel Structure Similarity
abstract
Screen content images (SCIs) usually comprise various content types with sharp edges, in which artifacts or distortions can be effectively sensed by a vanilla structure similarity measurement in a full-reference manner. Nonetheless, almost all of the current state-of-the-art (SOTA) structure similarity metrics are “locally” formulated in a single-level manner, while the true human visual system (HVS) follows the multilevel manner; such mismatch could eventually prevent these metrics from achieving reliable quality assessment. To ameliorate this issue, this article advocates a novel solution to measure structure similarity “globally” from the perspective of sparse representation. To perform multilevel quality assessment in accordance with the real HVS, the abovementioned global metric will be integrated with the conventional local ones by resorting to the newly devised selective deep fusion network. To validate its efficacy and effectiveness, we have compared our method with 12 SOTA methods over two widely used large-scale public SCI datasets, and the quantitative results indicate that our method yields significantly higher consistency with subjective quality scores than the current leading works. Both the source code and data are also publicly available to gain widespread acceptance and facilitate new advancement and validation.
Chenglizhao Chen, Hongmeng Zhao, Huan Yang 0001, Chong Peng 0001, Hong Qin 0001
ACM Trans. Multim. Comput. Commun. Appl.3
2020 Distributed Scheduling Algorithm for Optimizing Age of Information in Wireless Networks
abstract
Age of Information (AoI) is an emerging concept to model information freshness from the perspective of destinations of information deliveries. Serving as a metric to characterize the data delivery timeliness, the peak age indicates the maximum value of the AoI prior to a data packet reception. In this paper, we present a distributed scheduling algorithm for peak age optimization in a wireless network where a number of sensor nodes attempt to deliver their sensed data to a data collector over a wireless channel. In particular, each sensor node accesses the channel for data delivery independently according to an adaptively tuned transmission probability. The beauty of our algorithm lies in that, even with neither centralized infrastructure nor coordinations among the sensor nodes, our algorithm asymptotically approximates the optimal solution by only a constant factor. We perform solid theoretical analysis and extensive simulations to verify the efficacy of our algorithm. To the best of our knowledge, it is the first fully distributed scheduling algorithm for AoI optimization in wireless networks.
Dongxiao Yu, Xinpeng Duan, Feng Li 0002, Huan Yang 0001, Jiguo Yu
IPCCC5
2020 Learning-Aided Mobile Charging for Rechargeable Sensor Networks
Xinpeng Duan, Feng Li 0002, Dongxiao Yu, Huan Yang 0001, Hao Sheng 0001
WASA (1)4
2020 Efficient image structural similarity quality assessment method using image regularised feature
abstract
Image regularised features play a critical role in image processing domain, by integrating regularised feature and structural similarity, a new full‐reference image assessment method (IRF_SSIM) is proposed in this study. As well known, the gradient operator always be used to capture the edge information of the image, while the total variational regularised features can be adopted to calculate the detailed change information of image contrast and texture, as well as noise removal and edge retention. Therefore, the IRF_SSIM method extends the gradient features into the image regularised features to measure the structural changes in the image. In addition, image quality is also affected by variations of luminance and contrast. For a more comprehensive image quality assessment, the IRF_SSIM method considers the changes in structure, luminance and contrast simultaneously. In other words, the total image quality is estimated by structural similarity calculated by integrating the effects of image structure, luminance and contrast changes. Comparing with the representative methods, the experimental results illustrate that the IRF_SSIM method is highly consistent with the subjective assessment results.
Baoxiang Huang, Huan Yang 0001, Guojia Hou, Jinming Duan 0001
IET Image Process.3
2020 A novel dark channel prior guided variational framework for underwater image restoration
Guojia Hou, Jingming Li, Guodong Wang 0001, Huan Yang 0001, Baoxiang Huang, Zhenkuan Pan 0001
J. Vis. Commun. Image Represent.4
2020 Variational level set method for image segmentation with simplex constraint of landmarks
Baoxiang Huang, Zhenkuan Pan 0001, Huan Yang 0001, Li Bai 0001
Signal Process. Image Commun.3
2019 Joint Optimization of Routing and Storage Node Deployment in Heterogeneous Wireless Sensor Networks Towards Reliable Data Storage
Feng Li 0002, Huan Yang 0001, Yifei Zou, Dongxiao Yu, Jiguo Yu
WASA2
2019 An efficient nonlocal variational method with application to underwater image restoration
Guojia Hou, Zhenkuan Pan 0001, Guodong Wang 0001, Huan Yang 0001, Jinming Duan 0001
Neurocomputing4
2017 Content-based bitrate model for perceived compression distortion evaluation of mobile video services
abstract
A novel bitrate model with low complexity is proposed for perceived compression distortion assessment of mobile video with low resolution, which is extremely useful in intermediate network nodes for quality monitoring. Without fully decoding, parameters are extracted by bitstream analysing, such as bitrate, frame type, quantisation parameter, DCT coefficient, motion vector. Bitrate is regarded as an essential parameter meanwhile the bitrate–MOS curve is determined by video content. Respectively, spatial factor is estimated using quantisation parameter and DCT coefficient and temporal factor is estimated using motion vector. Apart from bitrate, the spatial and temporal factors, which reflect the characteristic of video content, are considered in the proposed model to obtain a more accurate evaluation. Experimental results show that the overall performance of proposed model significantly outperforms that of the other five bitrate models in terms of widely used performance criteria, including the Pearson correlation coefficient (PCC), the Spearman rank‐order correlation coefficient (SROCC), the root‐mean‐squared error (RMSE) and the outlier ratio (OR).
Honglei Su, Liu Qi, Huan Yang 0001, Zhenkuan Pan 0001
IET Image Process.5
2016 Saliency-Guided Quality Assessment of Screen Content Images
abstract
With the widespread adoption of multidevice communication, such as telecommuting, screen content images (SCIs) have become more closely and frequently related to our daily lives. For SCIs, the tasks of accurate visual quality assessment, high-efficiency compression, and suitable contrast enhancement have thus currently attracted increased attention. In particular, the quality evaluation of SCIs is important due to its good ability for instruction and optimization in various processing systems. Hence, in this paper, we develop a new objective metric for research on perceptual quality assessment of distorted SCIs. Compared to the classical MSE, our method, which mainly relies on simple convolution operators, first highlights the degradations in structures caused by different types of distortions and then detects salient areas where the distortions usually attract more attention. A comparison of our algorithm with the most popular and state-of-the-art quality measures is performed on two new SCI databases (SIQAD and SCD). Extensive results are provided to verify the superiority and efficiency of the proposed IQA technique.
Ke Gu 0001, Shiqi Wang 0001, Huan Yang 0001, Weisi Lin, Guangtao Zhai, Xiaokang Yang 0001, Wenjun Zhang 0001
IEEE Trans. Multim.3
2015 Modelling Human Factors in Perceptual Multimedia Quality: On The Role of Personality and Culture
abstract
Perception of multimedia quality is shaped by a rich interplay between system, context and human factors. While system and context factors are widely researched, few studies consider human factors as sources of systematic variance. This paper presents an analysis on the influence of personality and cultural traits on the perception of multimedia quality. A set of 144 video sequences (from 12 short movie excerpts) were rated by 114 participants from a cross-cultural population, producing 1232 ratings. On this data, three models are compared: a baseline model that only considers system factors; an extended model that includes personality and culture as human factors; and an optimistic model in which each participant is modelled as a random effect. An analysis shows that personality and cultural traits represent 9.3\% of the variance attributable to human factors while human factors overall predict an equal or higher proportion of variance compared to system factors. In addition, the quality-enjoyment correlation varied across the excerpts. This suggests that human factors play an important role in perceptual multimedia quality, but further research to explore moderation effects and a broader range of human factors is warranted.
Michael 'Adrir' Scott, Sharath Chandra Guntuku, Huan Yang 0001, Weisi Lin, George Ghinea
ACM Multimedia3
2015 Subjective quality evaluation of compressed digital compound images
Huan Yang 0001, Yuming Fang 0001, Yuan Yuan 0029, Weisi Lin
J. Vis. Commun. Image Represent.1
2015 Scale and Orientation Invariant Text Segmentation for Born-Digital Compound Images
abstract
Many recent applications require text segmentation for born-digital compound images. To this end, we propose a coarse-to-fine framework for segmenting texts of arbitrary scales and orientations in born-digital compound images. In the coarse stage, the local image activity measure is designed based upon the variation distribution of characters, to highlight the difference between textual and pictorial regions. This stage outputs a coarse textual layer including textual regions as well as a few pictorial regions with high activity. In the fine stage, a textual connected component (TCC) based refinement is proposed to eliminate the survived pictorial regions. In particular, a scale and orientation invariant grouping algorithm is proposed to adaptively generate TCCs with uniform statistical features. The minimum average distance and morphological operations are employed to assist the formation of candidate TCCs. Then, three string-level features (i.e., shapeness, color similarity, and mean activity level) are designed to distinguish the true TCCs from the false positive ones that are formed by connecting the high activity pictorial components. Extensive experiments show that the proposed framework can segment textual regions precisely from born-digital compound images, while preserving the integrity of texts with varied scales and orientations, and avoiding over-connection of textual regions.
Huan Yang 0001, Shiqian Wu, Chenwei Deng, Weisi Lin
IEEE Trans. Cybern.1
2015 Perceptual Quality Assessment of Screen Content Images
abstract
Research on screen content images (SCIs) becomes important as they are increasingly used in multi-device communication applications. In this paper, we present a study on perceptual quality assessment of distorted SCIs subjectively and objectively. We construct a large-scale screen image quality assessment database (SIQAD) consisting of 20 source and 980 distorted SCIs. In order to get the subjective quality scores and investigate, which part (text or picture) contributes more to the overall visual quality, the single stimulus methodology with 11 point numerical scale is employed to obtain three kinds of subjective scores corresponding to the entire, textual, and pictorial regions, respectively. According to the analysis of subjective data, we propose a weighting strategy to account for the correlation among these three kinds of subjective scores. Furthermore, we design an objective metric to measure the visual quality of distorted SCIs by considering the visual difference of textual and pictorial regions. The experimental results demonstrate that the proposed SCI perceptual quality assessment scheme, consisting of the objective metric and the weighting strategy, can achieve better performance than 11 state-of-the-art IQA methods. To the best of our knowledge, the SIQAD is the first large-scale database published for quality evaluation of SCIs, and this research is the first attempt to explore the perceptual quality assessment of distorted SCIs.
Huan Yang 0001, Yuming Fang 0001, Weisi Lin
IEEE Trans. Image Process.1
2015 Visual Object Tracking by Structure Complexity Coefficients
abstract
Appearance change of moving targets is a challenging problem in visual tracking. In this paper, we present a novel visual object tracking algorithm based on the observation dependent hidden Markov model (OD-HMM) framework. The observation dependency is computed by structure complexity coefficients (SCC) which is defined to predict the target appearance change. Unlike conventional methods addressing the appearance change problem by investigating different online appearance models, we handle this problem by addressing the fundamental reason of motion -related appearance change during visual tracking. Based on the analysis of motion-related appearance change, we investigate the relationship between the structure of the object surface and the appearance stability. The appearance of complex structural regions is easier to change compared with that of smooth structural regions with object moving. Based on this, we define SCC to predict the appearance stability of moving objects. Different from the standard HMM-based tracking algorithms where observations between different frames are assumed to be independent, we consider the observation dependency between consecutive frames with the information provided by SCC. Moreover , we present a novel outlier removing method in appearance model updating which helps to avoid error accumulation. Experimental results on challenging video sequences demonstrate that the proposed visual tracking algorithm with OD-HMM and SCC achieves better performance than existing related tracking algorithms.
Yuan Yuan 0029, Huan Yang 0001, Yuming Fang 0001, Weisi Lin
IEEE Trans. Multim.2
2014 Study on subjective quality assessment of Digital Compound Images
abstract
Quality assessment of digital compound images is a less investigated research topic. In this paper, we present a study for subjective quality assessment of Digital Compound Images (DCIs), and investigate whether existing Image Quality Assessment (IQA) methods are effective to evaluate the quality of distorted DCIs. A new Compound Image Quality Assessment Database (CIQAD) is constructed, including 24 reference DCIs and their 576 distorted versions. The Paired Comparison (PC) method is employed for the subjective viewing, and the Hodgerank decomposition is adopted to generate incomplete but balanced comparison pairs, so as to reduce the execution time while guaranteeing the reliability of the results. In our experiment, correlation of 14 existing IQA methods with the obtained Mean Opinion Score (MOS) values on the CIQAD is calculated, which indicates that the 14 IQA methods are not consistent with human visual perception when judging DCIs in different conditions. Therefore, objective quality assessment metrics should be specifically designed for DCIs. Our subjective study has delivered convincing information to guide the construction of objective metrics. Furthermore, we has also published the database online to favor future research on quality assessment of DCIs.
Huan Yang 0001, Weisi Lin, Chenwei Deng, Long Xu 0001
ISCAS1
2012 Learning based screen image compression
abstract
There are usually two components in computer screen images: textual and pictorial parts. The pictorial part can be compressed efficiently by classical coding approaches (e.g. JPEG, JPEG2000), while the compression of the textual part is still far away from being satisfactory for the reason that the textual content is usually of high-frequency. In this paper, a learning approach is used to construct a tailored dictionary for text representation. Based on the learned dictionary, a novel screen image compression algorithm is proposed through adopting different basis functions for the textual and pictorial components respectively. The screen images are firstly segmented into textual and pictorial parts. Then we employ traditional discrete cosine transformation (DCT) to facilitate the compression of pictorial part, while the learned dictionary is used to represent the textual part in screen images. Experimental results demonstrate the effectiveness of the proposed compression algorithm.
Huan Yang 0001, Weisi Lin, Chenwei Deng
MMSP1