Zhongyuan Lai

dblp:90/4522 · DBLP profile ↗
← Back
29ranked-venue papers
8as first author
15since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 18 · 7 first-author · 8 since 2021Databases, data management, data science and information retrieval · 10 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 9 · 1 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Systems, architecture and hardware · 1Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Counterfactual Augmented Causal Reasoning for Aspect-Based Sentiment Analysis
Liang Hu 0004, Mingzhu Zhou, Tangwei Ye, Xuejie Yang, Zhongyuan Lai, Qi Zhang 0020, Usman Naseem
WWW7
2026 AMID: Model-Agnostic Dataset Distillation by Adversarial Mutual Information Minimization
abstract
The escalating energy consumption and carbon footprint of training large-scale Web AI models pose urgent challenges for sustainable development. Dataset Distillation (DD) offers a promising avenue for green AI by compressing large datasets into small synthetic ones for efficient training. However, most existing DD methods overfit to the inductive biases of specific source architectures (e.g., CNNs or ViTs), resulting in poor cross-model generalization. This limitation necessitates redundant re-distillation processes for different architectures, severely undermining the energy-saving potential of DD. To address this, we introduce Adversarial Mutual Information Distillation (AMID), a rigorous framework designed to create highly reusable and robust synthetic datasets. From an information-theoretic perspective, we cast model-agnosticism as minimizing the mutual information (MI) between the synthetic data and the specific identity of the distillation model. We convert this intractable objective into a tractable two-player adversarial game, which unifies knowledge preservation with adversarial unlearning of architectural bias. Extensive experiments on CIFAR-10 and Tiny ImageNet demonstrate that AMID achieves state-of-the-art cross-architecture generalization across diverse CNNs and ViTs. Crucially, our analysis confirms that AMID significantly reduces the computational overhead and CO2 emissions of downstream training while maintaining robust performance, paving the way for energy-efficient, transferable, and sustainable Web AI ecosystems.
Aoqi Wu, Weiquan Huang, Liang Hu 0004, Yifan Yang 0004, Qi Zhang 0020, Jiaxing Miao, Yuhan Tang, Zhongyuan Lai
WWW10
2026 Accurate Trajectory Recovery in Underserved Areas via Location Inference from Web Crowdsourced Data
Tangwei Ye, Liang Hu 0004, Zhongyuan Lai, Qi Zhang 0020, Jiaxing Miao, Kun Yi 0001
WWW3
2025 STFEnc: A Novel Deep Encoder for Brain-Computer Interface Based on Interpretable Brain Features
abstract
Brain-computer interface (BCI) is a technology that enables direct connection and interaction of brain activity with external devices or systems. The encoding and decoding of neural signals play a crucial role in BCIs. The quality of such encodings are the key to robust and accurate information exchange and control between the brain and external devices. Currently, the limited capabilities of conventional brain signal processing is restricting a wider application of BCIs. In this paper, we propose a deep network encoder, Spatio-temporal frequency domain feature fusion encoder(STFEnc), to robustly and comprehensively encode the original electroencephalogram (EEG) data. The design of STFEnc is based on an interpretable understanding of the brain’s basic structural and connectivity features. The various submodules in STFEnc were designed to ensure the different features were taken into account. We evaluated encoder performance on three motion imagery datasets and one picture stimulus dataset. The results show that our encoder performs better than traditional deep encoders and advanced deep neural network models that have excelled in extracting EEG features for classification in recent years. The code is available at https://github.com/gwkuqgfkqe/Spatio-temporalfrequency-domain-feature-fusion-encoder.
Zonghan Du, Zhongyuan Lai
ICASSP2
2025 RF-DTR: A Multi-Stage DCT Token Regression Network for Progressive Rib Fracture Mask Refinement
abstract
Rib fracture patterns are key indicators of trauma severity. Detecting and locating these fractures is a critical yet time-consuming task, especially in 3D imaging, due to their minute size and irregular geometries. Existing voxel-based spatial methods fail to capture frequency-domain variations inherent in imaging and do not replicate the progressive refinement process used by clinicians during manual annotation, leading to suboptimal results. We propose a novel regression network, RF-DTR, incorporating a gated regressor mechanism and operating entirely in the frequency domain to address these challenges. Specifically, we present an innovative spatial-frequency transform applied to volumes and corresponding masks. Furthermore, we introduce a Mahalanobis regularization technique to enhance the model and learn high-frequency DCT components relevant to clinical tasks. Finally, a hierarchical penalty is proposed to improve the confidence of the prediction. Extensive experiments confirm our method's superiority in handling complex, sparsely annotated medical imaging datasets.
Shouyu Chen, Liang Hu 0008, Juntao Wang 0005, Usman Naseem, Zhongyuan Lai, Qi Zhang 0020
IJCAI5
2025 Recalling The Forgotten Class Memberships: Unlearned Models Can Be Noisy Labelers to Leak Privacy
abstract
Machine Unlearning (MU) technology facilitates the removal of the influence of specific data instances from trained models on request. Despite rapid advancements in MU technology, its vulnerabilities are still underexplored, posing potential risks of privacy breaches through leaks of ostensibly unlearned information. Current limited research on MU attacks requires access to original models containing privacy data, which violates the critical privacy-preserving objective of MU. To address this gap, we initiate the innovative study on recalling the forgotten class memberships from unlearned models (ULMs) without requiring access to the original one. Specifically, we implement a Membership Recall Attack (MRA) framework with a teacher-student knowledge distillation architecture, where ULMs serve as noisy labelers to transfer knowledge to student models. Then, it is translated into a Learning with Noisy Labels (LNL) problem for inferring correct labels of the forgetting instances. Extensive experiments on state-of-the-art MU methods with multiple real datasets demonstrate that the proposed MRA strategy exhibits high efficacy in recovering class memberships of unlearned instances. As a result, our study and evaluation have established a benchmark for future research on MU vulnerabilities.
Zhihao Sui, Liang Hu 0004, Jian Cao 0001, Dora D. Liu, Usman Naseem, Zhongyuan Lai, Qi Zhang 0020
IJCAI6
2025 COLUR: Confidence-Oriented Learning, Unlearning and Relearning with Noisy-Label Data for Model Restoration and Refinement
abstract
Large deep learning models have achieved significant success in various tasks. However, the performance of a model can significantly degrade if it is needed to train on datasets with noisy labels with misleading or ambiguous information. To date, there are limited investigations on how to restore performance when model degradation has been incurred by noisy label data. Inspired by the "forgetting mechanism" in neuroscience, which enables accelerating the relearning of correct knowledge by unlearning the wrong knowledge, we propose a robust model restoration and refinement (MRR) framework COLUR, namely Confidence-Oriented Learning, Unlearning and Relearning. Specifically, we implement COLUR with an efficient co-training architecture to unlearn the influence of label noise, and then refine model confidence on each label for relearning. Extensive experiments are conducted on four real datasets and all evaluation results show that COLUR consistently outperforms other SOTA methods after MRR.
Zhihao Sui, Liang Hu 0008, Jian Cao 0001, Usman Naseem, Zhongyuan Lai, Qi Zhang 0020
IJCAI5
2024 TFSICNet: A Neural Feature-based Encoder for Visual and Motor Imagery EEG Data
abstract
Brain-computer interface (BCI) is a technology that enables direct connection and interaction of brain activity with external devices or systems. The encoding and decoding of neural signals play a crucial role in BCIs. The quality of such encodings are the key to robust and accurate information exchange and control between the brain and external devices. Currently, the limited capabilities of conventional brain signal processing is restricting a wider application of BCIs. In this paper, we propose a deep network encoder, Temporal-Frequency-Spatio-Importance Correlation Network (TFSICNet), to robustly and comprehensively encode the original electroencephalogram (EEG) data. The design of TFSICNet is based on an interpretable understanding of the brain's basic structural and connectivity features. The various submodules in TFSICNet were designed to ensure the different features were taken into account. We evaluated encoder performance on three motion imagery datasets and one picture stimulus dataset. The results show that our encoder performs better than traditional deep encoders and advanced deep neural network models that have excelled in extracting EEG features for classification in recent years.
Zonghan Du, Zhongyuan Lai
BIBM2
2024 FSC: Few-Point Shape Completion
abstract
While previous studies have demonstrated successful 3D object shape completion with a sufficient number of points, they often fail in scenarios when a few points, e.g. tens of points, are observed. Surprisingly, via entropy analysis, we find that even a few points, e.g. 64 points, could retain substantial information to help recover the 3D shape of the object. To address the challenge of shape completion with very sparse point clouds, we then propose Few-point Shape Completion (FSC) model, which contains a novel dual-branch feature extractor for handling extremely sparse inputs, coupled with an extensive branch for maximal point utilization with a saliency branch for dynamic importance assignment. This model is further bolstered by a two-stage revision network that refines both the extracted features and the decoder output, enhancing the detail and authenticity of the completed point cloud. Our experiments demonstrate the feasibility of recovering 3D shapes from a few points. The proposed Few-point Shape Completion (FSC) model outperforms previous methods on both few-point inputs and many-point inputs, and shows good gener-alizability to different object categories. Code is available at https: https://github.com/xianzuwu/FSC.
Xianzu Wu, Xianfeng Wu, Tianyu Luan, Yajing Bai, Zhongyuan Lai, Junsong Yuan 0001
CVPR5
2024 VR-DiagNet: Medical Volumetric and Radiomic Diagnosis Networks with Interpretable Clinician-like Optimizing Visual Inspection
abstract
Interpretable and robust medical diagnoses are essential traits for practicing clinicians. Most computer-augmented diagnostic systems suffer from three major problems: non-interpretability, limited modality analysis, and narrow focus. Existing frameworks can either deal with multimodality to some extent but suffer from non-interpretability or partially interpretable but provide a limited modality and multifaceted capabilities. Our work aims to integrate all these aspects in one complete framework to fully utilize the full spectrum of information offered by multiple modalities and facets. We propose our solution via our novel architecture VR-DiagNet, consisting of a planner and a classifier, optimized iteratively and cohesively. VR-DiagNet simulates the perceptual process of clinicians via the use of volumetric imaging information integrated with radiomic features modality; at the same time, it recreates human thought processes via a customized Monte Carlo Tree Search (MCTS) which constructs a volume-tailored experience tree to identify slices of interest (SoIs) in our multi-slice perception space. We conducted extensive experiments across two diagnostic tasks comprising six public medical volumetric benchmark datasets. Our findings showcase superior performance, as evidenced by heightened accuracy and area under the curve (AUC) metrics, reduced computational overhead, and expedited convergence while conclusively illustrating the immense value of integrating volumetric and radiomic modalities for our current problem setup.
Shouyu Chen, Liang Hu 0004, Tangwei Ye, Zhongyuan Lai, Qi Zhang 0020, Usman Naseem, Nengjun Zhu
ACM Multimedia4
2023 Self-Supervised Learning for Multilevel Skeleton-Based Forgery Detection via Temporal-Causal Consistency of Actions
abstract
Skeleton-based human action recognition and analysis have become increasingly attainable in many areas, such as security surveillance and anomaly detection. Given the prevalence of skeleton-based applications, tampering attacks on human skeletal features have emerged very recently. In particular, checking the temporal inconsistency and/or incoherence (TII) in the skeletal sequence of human action is a principle of forgery detection. To this end, we propose an approach to self-supervised learning of the temporal causality behind human action, which can effectively check TII in skeletal sequences. Especially, we design a multilevel skeleton-based forgery detection framework to recognize the forgery on frame level, clip level, and action level in terms of learning the corresponding temporal-causal skeleton representations for each level. Specifically, a hierarchical graph convolution network architecture is designed to learn low-level skeleton representations based on physical skeleton connections and high-level action representations based on temporal-causal dependencies for specific actions. Extensive experiments consistently show state-of-the-art results on multilevel forgery detection tasks and superior performance of our framework compared to current competing methods.
Liang Hu 0004, Dora D. Liu, Qi Zhang 0020, Usman Naseem, Zhongyuan Lai
AAAI5
2023 A Dynamics and Task Decoupled Reinforcement Learning Architecture for High-Efficiency Dynamic Target Intercept
abstract
Due to the flexibility and ease of control, unmanned aerial vehicles (UAVs) have been increasingly used in various scenarios and applications in recent years. Training UAVs with reinforcement learning (RL) for a specific task is often expensive in terms of time and computation. However, it is known that the main effort of the learning process is made to fit the low-level physical dynamics systems instead of the high-level task itself. In this paper, we study to apply UAVs in the dynamic target intercept (DTI) task, where the dynamics systems equipped by different UAV models are correspondingly distinct. To this end, we propose a dynamics and task decoupled RL architecture to address the inefficient learning procedure, where the RL module focuses on modeling the DTI task without involving physical dynamics, and the design of states, actions, and rewards are completely task-oriented while the dynamics control module can adaptively convert actions from the RL module to dynamics signals to control different UAVs without retraining the RL module. We show the efficiency and efficacy of our results in comparison and ablation experiments against state-of-the-art methods.
Dora D. Liu, Liang Hu 0004, Qi Zhang 0020, Tangwei Ye, Usman Naseem, Zhongyuan Lai
AAAI6
2023 Show Me The Best Outfit for A Certain Scene: A Scene-aware Fashion Recommender System
abstract
Fashion recommendation (FR) has received increasing attention in the research of new types of recommender systems. Existing fashion recommender systems (FRSs) typically focus on clothing item suggestions for users in three scenarios: 1) how to best recommend fashion items preferred by users; 2) how to best compose a complete outfit, and 3) how to best complete a clothing ensemble. However, current FRSs often overlook an important aspect when making FR, that is, the compatibility of the clothing item or outfit recommendations is highly dependent on the scene context. To this end, we propose the scene-aware fashion recommender system (SAFRS), which uncovers a hitherto unexplored avenue where scene information is taken into account when constructing the FR model. More specifically, our SAFRS addresses this problem by encoding scene and outfit information in separation attention encoders and then fusing the resulting feature embeddings via a novel scene-aware compatibility score function. Extensive qualitative and quantitative experiments are conducted to show that our SAFRS model outperforms all baselines for every evaluated metric.
Tangwei Ye, Liang Hu 0004, Qi Zhang 0020, Zhongyuan Lai, Usman Naseem, Dora D. Liu
WWW4
2022 2SFGL: A Simple And Robust Protocol For Graph-Based Fraud Detection
abstract
Financial crime detection using graph learning improves financial safety and efficiency. However, criminals may commit financial crimes across different institutions to avoid detection, which increases the difficulty of detection for financial institutions which use local data for graph learning. As most financial institutions are subject to strict regulations in regards to data privacy protection, the training data is often isolated and conventional learning technology cannot handle the problem. Federated learning (FL) allows multiple institutions to train a model without revealing their datasets to each other, hence ensuring data privacy protection. In this paper, we proposes a novel two-stage approach to federated graph learning (2SFGL): The first stage of 2SFGL involves the virtual fusion of multiparty graphs, and the second involves model training and inference on the virtual graph. We evaluate our framework on a conventional fraud detection task based on the FraudAmazonDataset and FraudYelpDataset. Experimental results show that integrating and applying a GCN (Graph Convolutional Network) with our 2SFGL framework to the same task results in a 17.6%-30.2% increase in performance on several typical metrics compared to the case only using FedAvg, while integrating GraphSAGE with 2SFGL results in a 6%-16.2% increase in performance compared to the case only using FedAvg. We conclude that our proposed framework is a robust and simple protocol which can be simply integrated to pre-existing graph-based fraud detection methods.
Zhirui Pan, Guangzhong Wang, Zhaoning Li, Lifeng Chen, Yang Bian, Zhongyuan Lai
CloudCom6
2022 A Side Information Enhanced Matrix Factorization Approach via Hierarchical Generalized Linear Model
abstract
Matrix factorization (MF) is a popular method for collaborative filtering. Recently, more and more MF methods have been proposed to incorporate side information. However, most of them are vulnerable to changes in data or sub-models. Moreover, data often follows a Pareto distribution and such an imbalance of data leads to a biased global mean, affecting the prediction accuracy. To overcome these defects, we designed a Hierarchical Generalized Linear Model-based MF method (HGLMMF) which can leverage both the original and processed side information. More specifically, HGLMMF utilizes one portion of the side information to construct covariates for fixed effects and the other portion to model the cluster-specific effects to adjust the global-bias problem. In fact, a number of state-of-the-art MF models can be viewed as special cases of HGLMMF. The obtained prediction results from experiments prove that HGLMMF is highly competitive with state-of-the-art methods.
Dora D. Liu, Zhongyuan Lai, Usman Naseem
DSAA2
2016 Image Enhancement Based on Bi-Histogram Equalization with Non-Parametric Modified Technology
abstract
This paper presents a new image enhancement method using histogram equalization called Bi-Histogram Equalization with Non-parametric Modified Technology (BHENMT). Our proposed method consists of three steps: (i) The input original histogram is divided into two parts using the Otsu method. (ii) Then the histogram modification technique is used to control over enhancement and maximize entropy. (iii) Two sub images are enhanced by the traditional histogram equalization method using the corresponding modified histogram respectively and finally are merged into one output enhanced image. The experimental results show that BHENMT is better than other contrast enhancement methods according to subjective evaluation and various image objective evaluation measures, i.e. Entropy, AMBE and PSNR.
Zhijun Yao, Quan Zhou 0004, Zhongyuan Lai, Zhiming Ren
ICPADS3
2016 Fingertips detection and hand gesture recognition based on discrete curve evolution with a kinect sensor
abstract
In this paper, we propose a novel method that can detect fingertips as well as recognize hand gestures. Firstly, we collect the hand curves with a Kinect sensor. Secondly, we detect fingertips based on the discrete curve evolution. Thirdly, we recognize hand gestures using evolved curves partitioned at the detected fingertips. Experimental results show that our method performs well in both fingertips detection and hand gesture recognition.
Zhongyuan Lai, Zhijun Yao, Wu Xia
VCIP1
2016 Shape decomposition and classification by searching optimal part pruning sequence
Zhongyuan Lai
Pattern Recognit.2
2015 B-spline-based shape coding with accurate distortion measurement using analytical model
Zhongyuan Lai, Zhen Zuo, Zhijun Yao, Wenyu Liu 0001
Neurocomputing1
2014 Adaptive Edge Encoding Schemes for the Rate-Distortion Optimal Polygon-Based Shape Coding
abstract
In this paper, we present two adaptive edge encoding schemes for the operational rate-distortion optimal polygon-based shape coding. The encoding edge is represented by an octant number, a major component, and a minor component, where the ranges of the two components are determined at two levels. For the object-level, these ranges are either determined by users or adaptive to the contour characteristics and the predefined admissible distortions using the discrete contour evolution method. For the edge-level, the range of the minor component is further adaptive to the magnitude of the major component. The appropriate code tables are selected for the two components according to their ranges. Experiments on MPEG-4 test sequences showed that our schemes outperform existing schemes in terms of bit-rate at the same distortion level.
Junhuan Zhu, Zhongyuan Lai, Wenyu Liu 0001, Jiebo Luo 0001
DCC2
2014 Operational rate-distortion shape coding with dual error regularization
abstract
Existing operational rate-distortion shape coding aims at finding a polygon which can be encoded with the lowest bit rate under a given upper bound on the edge error. However, this upper bound may cause noticeable errors. Therefore, we add an ℓ2-norm error regularization term to the objective function, and seek the globally optimal solution using a shortest path algorithm for a weighted directed acyclic graph. Experiments confirm the accuracy and robustness of our method.
Zhongyuan Lai, Fan Zhang 0093, Weisi Lin
ICIP1
2014 A Novel Mathematical Morphology Based Antenna Deployment Scheme for Indoor Wireless Coverage
abstract
Coverage is an essential and important issue in wireless networks. In this paper, we propose a novel mathematical morphology based antenna deployment scheme for indoor wireless coverage. We partition the whole indoor area into two types of sub-areas, namely, corridor regions and sub-main regions. For these sub-areas, we extract their skeletons, find initial points on these skeletons and place antennas from these initial points along the skeleton branches in specific distances. At last we take the obstacles into consideration and refine the deployment. We compare the proposed scheme with some other schemes, including the site survey based scheme, coverage prediction based scheme and genetic algorithm based scheme. Simulation results show that the average number of antennas used in our proposed scheme is less than other schemes in each scenario while the computation is not complicated.
Zhongyuan Lai
VTC Fall2
2013 Perceptually friendly shape decomposition by resolving segmentation points with minimum cost
Wenyu Liu 0001, Zhongyuan Lai
J. Vis. Commun. Image Represent.3
2011 A Hybrid Admissible Distortion Checking Algorithm for the B-Spline-Based Operational Rate-Distortion Optimal Shape Coding
abstract
Admissible distortion checking algorithm plays a very important role in both rate distortion performance and computational efficiency of B-spline-based operational rate-distortion optimal shape coding framework under the minimum-maximum criterion. Existing distortion measurement using chord-length parameterization (DMCLP) is fast but results in extra bit-rate problem. In contrast, the up to date accurate distortion measurement using analytical model (ADMAM) can achieve the smallest bit-rate but is very time consuming. It motivates us to develop a hybrid admissible checking algorithm that can take full use of each advantage. Recalling the definitions of both DMCLP and ADMAM for each associated contour point, the authors show that DMCLP is the distance from the parameterized B-spline point while ADMAM is the shortest distance from the approximating B-splines.
Zhongyuan Lai, Zhen Zuo, Wenyu Liu 0001
DCC1
2011 Accurate Distortion Measurement Using Analytical Model for the B-Spline-Based Shape Coding
abstract
Summary form only given. Existing distortion measurements for the B-spline-based shape coding include ap proximation, quantization, or parameterization process, so they are approximate techniques. They may inaccurately predict the actual distortion value, which motivates us to construct a model that can accurately measure the actual distortion. It was reported that the actual distortion for reconstruction quality assessment is the minimal Euclidean distance between each associated contour point and the reconstruction contour.
Zhongyuan Lai, Zhen Zuo, Wenyu Liu 0001
DCC1
2011 Accurate distortion measurement for B-spline-based shape coding
abstract
In this paper, we present a new contour point distortion measurement, called accurate distortion measurement for B-spline-based shape coding (ADMBSC). Different from existing distortion measurements containing approximation, quantization or parameterization, our distortion is defined as the shortest distance from the original B-spline to the associated contour point. This is in line with the subjective-based objective quality metric. Geometric relationships are introduced to simplify computation, followed by a hybrid admissible distortion checking algorithm to reduce execution time. Theoretical analysis and experimental results demonstrate that when the operational rate-distortion optimal shape coding framework under the minimum-maximum criterion is applied, the ADMBSC can lead to the smallest bit-rate among all the distortion measurements that can guarantee the admissible distortion. Moreover, if the original contour has NCpoints, it takes only O(NC) time for segment distortion measuring paradigms, whose computational complexity is the same as the lowest one among the existing distortion measurements.
Zhongyuan Lai, Zhen Zuo, Zhijun Yao, Wenyu Liu 0001
ICIP1
2011 A symmetric KL divergence based spatiogram similarity measure
abstract
Spatiogram is a generalization of histogram to capture higher-order spatial moments information. To apply spatiogram to object tracking, suitable similarity measure is critical. Although there is a series of work on introducing improved distance measure over the original method, their performance in object tracking is very limited due to insufficient discriminative power. In this paper, we present a symmetric KL divergence based spatiogram similarity measure and show both theoretically and experimentally that, the proposed measure gives superior discriminative power than existing methods, and achieved promising performance in tracking object from single or sequence of images.
Zhijun Yao, Zhongyuan Lai, Wenyu Liu 0001
ICIP2
2010 Arbitrary Directional Edge Encoding Schemes for the Operational Rate-Distortion Optimal Shape Coding Framework
abstract
We present two edge encoding schemes, namely 8-sector scheme and 16-sector scheme, for the operational rate-distortion (ORD) optimal shape coding framework. Different from the traditional 8-direction scheme that can only encode edges with angles being an integer multiple of π/4, our proposals can encode edges with arbitrary angles. We partition the digital coordinate plane into 8 and 16 sectors, and design the corresponding differential schemes to encode the short and the long component of each vertex. Experiment results demonstrate that our two proposals can reduce a large number of encoding vertices and therefore reduce 10%~20% bits for the basic ORD optimal algorithms and 10%~30% bits for all the ORD optimal algorithms under the same distortion thresholds, respectively. Moreover, the reconstruction contours are more compact compared with those using the traditional 8-direction edge encoding scheme.
Zhongyuan Lai, Junhuan Zhu, Zhou Ren, Wenyu Liu 0001, Baolan Yan
DCC1
2009 Perceptual Relevance Measure for Generic Shape Coding
abstract
Approximation metric is a significant factor for subjective quality improvement of polygonal vertex-based shape codec. Usually this metric is defined by absolute distance measure (ADM). ADM only considers the shortest absolute distance from candidate vertices on the original object contour segment to corresponding approximating line segment and ignores other visual characteristics of the contour segment. Thus it cannot describe the original contour appropriately in visual aspects. As a result, it may severely degrade the subjective reconstruction quality, especially for contours with sharp salience. We propose perceptual relevance measure (PRM) to address this problem. We first describe the original relative position relationship between candidate vertex C and approximating line segment AB by three parameters having definite visual meanings, namely turn angle of and lengths of two adjacent line segments a and b. And then we provide the following three visual properties to find the exactly expression of PRM.
Zhongyuan Lai, Wenyu Liu 0001
DCC1