VLDB 2026 Research / reviewers in the wild / expert
Miaohui Wang
dblp:83/8841
· DBLP profile ↗
88ranked-venue papers
32as first author
65since 2021 · last 2026
0000-0003-1125-9299ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 66 · 23 first-author · 45 since 2021Artificial intelligence and machine learning · 26 · 7 first-author · 23 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 6 first-author · 12 since 2021Systems, architecture and hardware · 1 · 1 first-authorComputer networks · 1 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MCI-Net: A Robust Multi-Domain Context Integration Network for Point Cloud RegistrationabstractRobust and discriminative feature learning is critical for high-quality point cloud registration. However, existing deep learning–based methods typically rely on Euclidean neighborhood-based strategies for feature extraction, which struggle to effectively capture the implicit semantics and structural consistency in point clouds. To address these issues, we propose a multi-domain context integration network (MCI-Net) that improves feature representation and registration performance by aggregating contextual cues from diverse domains. Specifically, we propose a graph neighborhood aggregation module, which constructs a global graph to capture the overall structural relationships within point clouds. We then propose a progressive context interaction module to enhance feature discriminability by performing intra-domain feature decoupling and inter-domain context interaction. Finally, we design a dynamic inlier selection method that optimizes inlier weights using residual information from multiple iterations of pose estimation, thereby improving the accuracy and robustness of registration. Extensive experiments on indoor RGB-D and outdoor LiDAR datasets show that the proposed MCI-Net significantly outperforms existing state-of-the-art methods, achieving the highest registration recall of 96.4% on 3DMatch. Shuyuan Lin, Wenwu Peng, Qiang Qi, Miaohui Wang, Jian Weng 0001 |
AAAI | 5 |
| 2026 | FreeMem: Enhancing Consistency in Long Video Generation via Tuning-Free MemoryabstractText-to-Video (T2V) generation has advanced greatly, yet maintaining consistency remains challenging, especially for tuning-free long video generation. We attribute the consistency problem to cumulative deviations for long video generation at three levels: the random noise lacking correlation results initial deviation between frames; discrepancy in semantic feature tokens between denoising network blocks gradually accumulates as the frame count grows, leading to greater deviations; attention mechanisms struggle to capture global relationships across distant frames in long videos. To address these, we propose FreeMem, a tuning-free framework leveraging hierarchical memory update and injection: the noise memory stabilizes consistency by manipulating low and high frequency components in the initial noise space; the token memory combats inconsistency through adaptive fusion of historical and current semantic feature tokens between denoising network blocks; and the attention memory establishes persistent cache to model long-range relationships within self attention layers. Evaluated on VBench, FreeMem improves subject and background consistency matrics across various methods, offering a practical solution for low-cost, high-consistency long video generation. Jibin Peng, Di Lin 0002, Zhecheng Xu, Wuyuan Xie, Miaohui Wang, Lingyu Liang, Qing Guo 0005 |
AAAI | 7 |
| 2026 | Perceive More with Less: LiDAR Point Cloud Compression at Just Recognizable Distortion for 3D Scene UnderstandingabstractExisting LiDAR point cloud (LPC) data coding methods primarily focus on balancing compression efficiency and reconstruction quality according to the human vision system (HVS). However, these methods rarely consider the requirements of downstream scene understanding tasks from the perspective of the machine vision system (MVS). To address this challenge, we explore the maximum degree of LPC compression that has negligible impact on perception accuracy, called LPC-based just recognizable compression distortion (lpcJRCD). Specifically, we introduce a novel point-wise quantization approach for constructing a MVS-based LiDAR dataset and present a new lpcJRCD-guided intelligent compression framework tailored for MVS applications. To enhance MVS-based LPC compression efficiency, we develop a dual-feature interaction (DFI) module that fuses point and voxel features. Additionally, we propose a mask-based loss function to ensure accurate point-wise quality level prediction. Experimental results demonstrate the effectiveness of our proposed model in reducing the average bit rate by up to 94.98% while preserving perception accuracy in autonomous vehicles. Miaohui Wang, Runnan Huang, Taojun Liu, Shuyuan Lin, Ye Liu 0005, Yun Song |
AAAI | 1 |
| 2026 | The Last Byte: Learning Just Enough for Machine-Oriented Image CompressionabstractJust recognizable distortion (JRD) has been introduced for image compression for machines, aiming to quantify the maximum coding distortion that can be tolerated by a specific perception model, thereby defining the upper bound of machine vision redundancy (MVR). However, existing JRD-based redundancy estimation methods face three key challenges: limited dataset annotation accuracy, low prediction efficiency, and insufficient perception accuracy, all of which hinder their practical deployment. To address these limitations, we propose a new MVR-Net, a frame-wise efficient JRD prediction method that generates the optimal encoding quantization map in a single inference pass. Furthermore, we refine the annotation standard for JRD datasets based on experimental insights, enhancing the precision of recognizable redundancy measurement. Compared to stateof-the-art methods, MVR-Net achieves a superior balance between bitrate reduction and perception accuracy in JRD-guided compression, while offering up to a 40,000× speed improvement, demonstrating its practicality and efficiency for real-world applications. Wuyuan Xie, Zhenming Li, Ye Liu 0005, Yun Song, Miaohui Wang |
AAAI | 6 |
| 2026 | Firing Bits Where It Matters: Spiking-Guided Just Recognizable Distortion Modeling for Machine-Centric Video CodingabstractJust recognizable distortion (JRD) has emerged as a promising paradigm for machine-centric video coding. However, existing JRD-guided coding methods are limited by coarse annotation granularity and high computational cost, which hinder their deployment. In this paper, we first investigate the impact of different JRD annotation strategies on downstream task performance. By incorporating both instance-level and contextual information, we construct a new JRD dataset with fine-grained annotations compatible with object detection and instance segmentation tasks. To enhance quantization parameter (QP) map prediction while maintaining computational efficiency, we propose a novel spiking neural network (SNN)-based framework that decomposes video frames into spatial structures, channel interactions, and temporal patterns. Furthermore, we introduce a spiking attention mechanism to aggregate task-relevant features and employ adaptive scaling vectors to suppress machine-perceived redundancy, enabling targeted bitrate allocation aligned with task-critical content. Extensive experiments on multiple datasets and backbones demonstrate that our approach consistently outperforms state-of-the-art codec-based and JRD-guided methods in maintaining task performance at ultra-low bitrates, while significantly reducing computational overhead. Wuyuan Xie, Zhenming Li, Yuwu Lu, Di Lin 0002, Yun Song, Miaohui Wang |
AAAI | 6 |
| 2026 | Point Cloud Quality Assessment via Multi-View Structure-Aware Feature FusionabstractPoint cloud quality assessment (PCQA) is essential for reliable 3D visual applications. While point-based methods face challenges in characterizing distortions due to point cloud disorder, projection-based approaches offer better efficiency but suffer from geometric distortion insensitivity and texture representation blind spots. This study proposes SAF-Net, a multi-view structure-aware feature fusion network for PCQA. We first identify two key limitations in projection-based methods: insufficient geometric distortion perception and representation blind spots (RBS) in texture images. To address these issues, SAF-Net innovatively integrates object mask maps and local binary pattern (LBP) maps. The mask maps enhance geometric distortion perception by extracting edge sharpness and curvature variations, while LBP maps capture essential structural information to overcome RBS and align with human visual system (HVS) sensitivity. SAF-Net employs a hybrid CNN-ViT architecture to balance local feature extraction and global context modeling, along with a progressive fusion strategy to optimize cross-modal feature interaction. Extensive experiments demonstrate the superior performance of SAF-Net on multiple benchmarks, establishing new state-of-the-art results in PCQA. Jian Xiong 0005, Lingxia Jiang, Xianzhong Long, Miaohui Wang, Hao Gao 0005 |
AAAI | 4 |
| 2026 | Channel-level feature selection and fusion network for visible-infrared person re-identification
Zelin Deng, Yun Song, Ke Nai, Miaohui Wang |
Multim. Syst. | 5 |
| 2026 | Efficient 3D Surface Super-Resolution via Normal-Based Multimodal RestorationabstractHigh-fidelity 3D surface is essential for vision tasks across various domains such as medical imaging, cultural heritage preservation, quality inspection, virtual reality, and autonomous navigation. However, the intricate nature of 3D data representations poses significant challenges in restoring diverse 3D surfaces while capturing fine-grained geometric details at a low cost. This paper introduces an efficient multimodal normal-based 3D surface super-resolution (mn3DSSR) framework, designed to address the challenges of microgeometry enhancement and computational overhead. Specifically, we have constructed one of the largest normal-based multimodal dataset, ensuring superior data quality and diversity through meticulous subjective selection. Furthermore, we explore a new two-branch multimodal alignment approach along with a multimodal split fusion module to mitigate computational complexity while improving restoration performances. To address the limitations associated with normal-based multimodal learning, we develop novel normal-induced loss functions that facilitate geometric consistency and improve feature alignment. Extensive experiments conducted on seven benchmark datasets across four different 3D data representations demonstrate that mn3DSSR consistently outperforms state-of-the-art super-resolution methods in terms of restoration accuracy with high computational efficiency. Miaohui Wang, Yunheng Liu, Wuyuan Xie, Boxin Shi, Jianmin Jiang |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2026 | DPTracker: Dynamic prompter for RGB-D tracking
Junzhe Zhao, Jintao Su, Ye Liu 0005, Jun Liu 0036, Miaohui Wang |
Pattern Recognit. Lett. | 5 |
| 2026 | APNet: Accurate Prompting Network With Modality Guidance and Structural Awareness for RGB-D Semantic SegmentationabstractParameter-efficient fine-tuning (PEFT) is promising for RGB-D semantic segmentation, as lightweight prompters enable frozen pre-trained RGB backbones to leverage massive RGB pretraining knowledge without full fine-tuning on limited paired RGB-D data. However, existing PEFT methods have two critical limitations: static modal fusion ignores the dynamic reliability of RGB and depth across scenes, leading to suboptimal performance in complex environments; conventional prompts lack structural awareness, causing the loss of edge and texture details essential for dense prediction. To solve these problems, we propose the Accurate Prompting Network (APNet) for precise prompt injection in frozen backbones with two core modules. A Modality Effectiveness Guider (MEG) conducts input-level modal reliability assessment and dynamically generates scene-adaptive modality weights by capturing scene characteristics (e.g., illumination, texture richness). A Structural Awareness Prompter (SAP) injects directional structural priors into prompts via multi-directional gating convolution, endowing prompts with explicit edge and texture information to match semantic segmentation demands. MEG and SAP collaboratively form a precise prompting mechanism that realizes dynamic modal contribution allocation and structural detail preservation, facilitating efficient and accurate cross-modal knowledge transfer to the frozen backbone. Extensive experiments on NYUDv2 and SUN RGB-D show that APNet achieves state-of-the-art mIoU of 59.6% and 52.6% with only 6.2M trainable parameters, realizing a superior trade-off between segmentation accuracy and parameter efficiency. Junzhe Zhao, Jintao Su, Jun Liu 0036, Miaohui Wang, Ye Liu 0005 |
IEEE Signal Process. Lett. | 4 |
| 2026 | Boosting the No-Reference Image Quality Assessment via Low-Quality Pseudo ReferencesabstractNo-reference image quality assessment (NR-IQA) aims to predict perceptual image quality without access to pristine references, which remains challenging due to diverse and complex distortions. Recent pseudo-reference-based methods attempt to mitigate this challenge but often rely on highfidelity pseudo-reference reconstruction. In contrast, this work shows that improving NR-IQA performance does not depend on reconstruction quality, but on effective representation learning, feature alignment, and deviation modeling between distorted images and pseudo references. To this end, we propose a novel NR-IQA framework that leverages low-quality pseudo references generated by a masked autoencoder with a lightweight decoder. Rather than pursuing detailed reconstruction, the pseudo reference is used to facilitate representation-level deviation modeling in a shared latent space via a cross-attention-based mechanism. Extensive experiments on multiple benchmark datasets demonstrate that the proposed method consistently outperforms state-of-the-art NR-IQA approaches while maintaining modest computational complexity. Our source code will be available at: https://github.com/jianjin008/L-IQA. Lili Meng, Yingnan Wang, Miaohui Wang, Guosheng Lin, Cheng Liang 0001, Jiande Sun 0001, Huaxiang Zhang 0001, Weisi Lin |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | mmFAS: Multimodal Face Anti-Spoofing Using Multi-Level Alignment and Switch-Attention FusionabstractThe increasing number of presentation attacks on reliable face matching has raised concerns and garnered attention towards face anti-spoofing (FAS). However, existing methods for FAS modeling commonly fuse multiple visual modalities (e.g., RGB, Depth, and Infrared) in a straightforward manner, disregarding latent feature gaps that can hinder representation learning. To address this challenge, we propose a novel multimodal FAS framework (mmFAS) that focuses on explicit alignment and fusion of latent features across different modalities. Specifically, we develop a multimodal alignment module to alleviate the latent feature gap by using instance-level contrastive learning and class-level matching simultaneously. Further, we explore a new switch-attention based fusion module to automatically aggregate complementary information and control model complexity. To evaluate the anti-spoofing performance more effectively, we adopt a challenging yet meaningful cross-database protocol involving four benchmark multimodal FAS datasets to simulate realworld scenarios. Extensive experimental results demonstrate the effectiveness of mmFAS in improving the accuracy of FAS systems, outperforming 10 representative methods. Geng Chen 0006, Wuyuan Xie, Di Lin 0002, Ye Liu 0005, Miaohui Wang |
AAAI | 5 |
| 2025 | DDJND: Dual Domain Just Noticeable Difference in Multi-Source Content Images with Structural DiscrepancyabstractMost existing just noticeable difference (JND) methods primarily integrate specific masking effects in a single domain. However, these single-domain JND methods struggle with the structural discrepancies in multi-source content images, limiting their effectiveness in visual redundancy estimation. To address this issue, we propose a dual domain encoder that combines spatial and frequency features to comprehensively capture visual patterns. Our design includes spatial pattern balance and frequency detail correction modules to balance global and local patterns and correct low- and high-frequency distributions. Additionally, we develop a dual domain decoder to effectively extract multi-scale pattern redundancies and integrate them with detail redundancies in the frequency domain. Experiments demonstrate the effectiveness and robustness of our proposed method in handling structural discrepancies in multi-source content images. Miaohui Wang, Zhenming Li, Wuyuan Xie |
AAAI | 1 |
| 2025 | MARSNet: Scalable Deep Coding of LiDAR Point Clouds via Multimodal and Residual Learning
Yanji Huang, Runnan Huang, Jianlong Zhou, Yingqi Zhuo, Yanshan Li, Miaohui Wang |
ICIG (1) | 6 |
| 2025 | Wavelet Convolution and Multi-Scale Attention Network for Image Tampering LocalizationabstractConventional tampering localization only extracts features in the image domain, which makes it hard to capture the subtle tampering traces. In this paper, we propose a wavelet convolution and multi-scale attention network (WCMA-Net) for image tampering detection and localization, in which a wavelet convolution module (WCM) branch and a multi-scale attention module (MSAM) branch are integrated following the backbone. In the WCM branch, wavelet decomposition is utilized to enhance high-frequency details and enhance the detection of subtle tampering traces. In the MSAM branch, a multi-scale attention operation is employed to extract global and local features, which are then combined according to their similarity to capture long-range dependencies among pixels. Finally, an adaptive weight strategy is employed to fuse the features from both branches for binary pixel-level tampering mask prediction. Experimental results on various public datasets demonstrate that the proposed method achieves superior precise pixel-level image tampering localization over state-of-the-art methods. Codes and models are available at https://github.com/csust-sonie/WCMA-Net. Yun Song, Yaoyao Xu, Dengyong Zhang, Miaohui Wang |
ICME | 6 |
| 2025 | kgMBQA: Quality Knowledge Graph-driven Multimodal Blind Image AssessmentabstractBlind image assessment aims to simulate human prediction of image quality distortion levels and provide quality scores. However, existing unimodal quality indicators have limited representational ability when facing complex contents and distortion types, and the predicted scores also fail to provide explanatory reasons, which further affects the credibility of their prediction results. To address these challenges, we propose a multimodal quality indicator with explanatory text descriptions, called kgMBQA. Specifically, we construct an image quality knowledge graph and conduct in-depth mining to generate explanatory texts. The text modality is further aligned and fused with the image modality, thereby improving the model performance while also outputting its corresponding quality explanatory description. The experimental results demonstrate that our kgMBQA achieves the best performance compared to recent representative methods on the KonIQ-10k, LIVE Challenge, BIQ2021, TID2013, and AIGC-3K datasets. Wuyuan Xie, Tingcheng Bian, Miaohui Wang |
IJCAI | 3 |
| 2025 | MMVQA: Dual-Path Multimodal Fusion for AI-Generated Video Quality Assessment
Wuyuan Xie, Miaohui Wang |
PRCV (11) | 3 |
| 2025 | Blind Multimodal Quality Assessment of Low-Light Images
Miaohui Wang, Zhuowei Xu, Mai Xu, Weisi Lin |
Int. J. Comput. Vis. | 1 |
| 2025 | CPAL: Cross-Prompting Adapter With LoRAs for RGB+X Semantic SegmentationabstractAs sensor technology evolves, RGB+X systems combine traditional RGB cameras with another type of auxiliary sensor, which enhances perception capabilities and provides richer information for important tasks such as semantic segmentation. However, acquiring massive RGB+X data is difficult due to the need for specific acquisition equipment. Therefore, traditional RGB+X segmentation methods often perform pretraining on relatively abundant RGB data. However, these methods lack corresponding mechanisms to fully exploit the pretrained model, and the scope of the pretraining RGB dataset remains limited. Recent works have employed prompt learning to tap into the potential of pretrained foundation models, but these methods adopt a unidirectional prompting approach i.e., using X or RGB+X modality to prompt pretrained foundation models in RGB modality, neglecting the potential in non-RGB modalities. In this paper, we are dedicated to developing the potential of pretrained foundation models in both RGB and non-RGB modalities simultaneously, which is non-trivial due to the semantic gap between modalities. Specifically, we present the CPAL (Cross-prompting Adapter with LoRAs), a framework that features a novel bi-directional adapter to simultaneously fully exploit the complementarity and bridging the semantic gap between modalities. Additionally, CPAL introduces low-rank adaption (LoRA) to fine-tune the foundation model of each modal. With the support of these elements, we have successfully unleashed the potential of RGB foundation models in both RGB and non-RGB modalities simultaneously. Our method achieves state-of-the-art (SOTA) performance on five multi-modal benchmarks, including RGB+Depth, RGB+Thermal, RGB+Event, and a multi-modal video object segmentation benchmark, as well as four multi-modal salient object detection benchmarks. The code and results are available at:https://github.com/abelny56/CPAL. Ye Liu 0005, Miaohui Wang, Jun Liu 0036 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Prompt Engineering-Assisted Malware Dynamic Analysis Using GPT-4abstractMalware detection remains a critical challenge due to the increasing use of code obfuscation, packing, and wrapping techniques, which hinder traditional static analysis methods. Dynamic analysis, particularly through the examination of Application Programming Interface (API) call sequences, has emerged as an effective approach for identifying malicious behaviors. However, existing deep learning models often struggle to generate high-quality representations of API calls and are unable to handle previously unseen APIs, thereby limiting detection performance and model generalization. To address these challenges, we propose a novel malware dynamic analysis framework that leveragesGPT-4prompt engineering to generate descriptive text for each API call within a sequence. These descriptions are then encoded using a pre-trained BERT model to produce rich, knowledge-enhanced representations of API sequences. Our method not only incorporates external knowledge for improved semantic understanding but also enables the representation of unknown API calls, thus enhancing generalization. We further design a CNN-based classifier to extract features from the enriched representations for malware detection and classification. Extensive experiments on five benchmark datasets demonstrate that our approach outperforms state-of-the-art methods, achieving superior detection accuracy and generalization across different datasets. Especially, the detection accuracy on the Catak dataset increased by 12.38%, which highlights the significant improvement of our method in challenging scenarios. The code is available. Pei Yan, Shunquan Tan, Miaohui Wang, Jiwu Huang |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2025 | Coupled Noise Suppression and Feature Enhancement Network for Skeleton-Based Action RecognitionabstractIn recent years, remarkable progress has been made in skeleton-based action recognition. However, there is a significant amount of noise in skeleton data, which is simply overlooked by most existing methods. Some methods have designed specialized mechanisms to handle noise, but these mechanisms are either based on prior knowledge or require additional supervision information. To overcome these problems, we propose in this article a fully implicit solution, which embeds a soft-thresholding-based denoising module into existing networks, which can automatically learn to remove noise without any prior knowledge or additional supervision information. In addition, by relaxing the nonnegative constraint, the module gains the ability to adaptively enhance key features. Based on this, we further propose a two-staged method for coupled noise suppression and feature enhancement. The proposed method achieves state-of-the-art performance on public datasets. Moreover, on noise polluted datasets, the proposed method demonstrates significant performance advantages over existing methods. Ye Liu 0005, Tianyong Wu, Tianhao Shi, Miaohui Wang, Hao Gao 0005, Jun Liu 0036 |
IEEE Trans. Ind. Informatics | 4 |
| 2025 | suLPCC: A Novel LiDAR Point Cloud Compression Framework for Scene Understanding TasksabstractLight detection and ranging (LiDAR) point cloud compression (LPCC) plays an important role in managing the storage, transmission, and perception of the rapidly expanding volume of LiDAR point cloud (LPC) data. However, there has been a noticeable lack of comprehensive investigation into LPCC methods specifically designed for environmental perception and understanding. To address this gap, we propose a new LPCC framework aimed at meeting the unique requirements of various scene understanding tasks, enhancing the adaptability of LPCCs in real-world scenarios. Specifically, we divide the input LPCs into an object and a scene component through a distinction module, design a new point completion-based method to encode object LPCs, and develop novel structure-aware intracoding and motion-optimized intercoding schemes to compress scene LPCs. Experimental results on three benchmark datasets demonstrate the effectiveness of our proposed method on the localization, mapping, and detection tasks. We believe that the findings presented in this article will contribute to a deeper understanding of LPCCs as well as promote further development of LiDAR sensor-based systems. Miaohui Wang, Runnan Huang, Ye Liu 0005, Yanshan Li, Wuyuan Xie |
IEEE Trans. Ind. Informatics | 1 |
| 2025 | Visual Quality Assessment of Composite Images: A Compression-Oriented Database and MeasurementabstractComposite images (CIs) have experienced unprecedented growth, especially with the prosperity of a large number of generative AI technologies. They are usually created by combining multiple visual elements from different sources to form a single cohesive composition, which have an increasing impact on a variety of vision applications. However, transmission of CIs can degrade their visual quality, especially undergoing lossy compression to reduce bandwidth and storage. To facilitate the development of objective measurements for CIs and investigate the influence of compression distortions on their perception, we establish a compression-oriented image quality assessment (CIQA) database for CIs (called ciCIQA) with 30 typical encoding distortions. Compressed with six representative codecs, we have carried out a large-scale subjective experiment that delivered 3,000 encoded CIs with labeled quality scores, making ciCIQA one of the earliest CI databases with the most compression types. ciCIQA enables us to explore the encoding effects on visual quality from the first five just noticeable difference (JND) points, offering insights for perceptual CI compression and related tasks. Moreover, we have proposed a new multi-masked no-reference CIQA method(called mmCIQA), including a multi-masked quality representation module, a self-supervised quality alignment module, and a multi-masked attentive fusion module. Experimental results demonstrate the outstanding performance of our mmCIQA in assessing the quality of CIs, outperforming 17 competitive approaches. The proposed method and database as well as the collected objective metrics are made publicly available on https://charwill.github.io/mmciqa.html. Miaohui Wang, Zhuowei Xu, Yuming Fang 0001, Weisi Lin |
IEEE Trans. Image Process. | 1 |
| 2025 | CPSR-CLIP: Conditional Prompt-Induced Style Reconstruction for Zero-Shot Domain Adaptation
Jiayu Qian, Yuwu Lu, Wuyuan Xie, Zhihui Lai 0001, Miaohui Wang, Xuelong Li 0001 |
IEEE Trans. Multim. | 5 |
| 2025 | Compression Approaches for LiDAR Point Clouds and Beyond: A SurveyabstractWith the widespread use of LiDAR sensors in autonomous driving, LiDAR point cloud compression (LPCC) plays an important role in effectively managing the storage, transmission, and perception of the growing volume of LiDAR data. Despite this need, there has been a noticeable absence of comprehensive investigations specifically dedicated to LPCC methods. To address this issue, this article presents a systematic survey of existing LPCCs, aiming to summarize recent progress and inspire future research in this field. We begin by providing a general introduction of LPCC fundamentals, covering the latest LiDAR point cloud (LPC) datasets, distinctive attributes, evaluation metrics, and data formats. We then conduct a careful review and comparison of LPCCs, examining image-based, octree-based, deep-learned, and other approaches, offering valuable insights into the strengths and weaknesses of cutting-edge models. Finally, we propose future research directions based on the limitations of recent LPCCs. We believe that the findings presented in this article will contribute to a deeper understanding of LPCCs and promote further development of LiDAR sensor-based systems. Miaohui Wang, Runnan Huang, Wuyuan Xie, Zhan Ma 0001, Siwei Ma 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2024 | msLPCC: A Multimodal-Driven Scalable Framework for Deep LiDAR Point Cloud CompressionabstractLiDAR sensors are widely used in autonomous driving, and the growing storage and transmission demands have made LiDAR point cloud compression (LPCC) a hot research topic. To address the challenges posed by the large-scale and uneven-distribution (spatial and categorical) of LiDAR point data, this paper presents a new multimodal-driven scalable LPCC framework. For the large-scale challenge, we decouple the original LiDAR data into multi-layer point subsets, compress and transmit each layer separately, so as to ensure the reconstruction quality requirement under different scenarios. For the uneven-distribution challenge, we extract, align, and fuse heterologous feature representations, including point modality with position information, depth modality with spatial distance information, and segmentation modality with category information. Extensive experimental results on the benchmark SemanticKITTI database validate that our method outperforms 14 recent representative LPCC methods. Miaohui Wang, Runnan Huang, Hengjin Dong, Di Lin 0002, Yun Song, Wuyuan Xie |
AAAI | 1 |
| 2024 | Visual Redundancy Removal for Composite Images: A Benchmark Dataset and a Multi-Visual-Effects Driven Incremental MethodabstractComposite images (CIs) typically combine various elements from different scenes, views, and styles, which are a very important information carrier in the era of mixed media such as virtual reality, mixed reality, metaverse, etc. However, the complexity of CI content presents a significant challenge for subsequent visual perception modeling and compression. In addition, the lack of benchmark CI databases also hinders the use of recent advanced data-driven methods. To address these challenges, we first establish one of the earliest visual redundancy prediction (VRP) databases for CIs. Moreover, we propose a multi-visual effect (MVE)-driven incremental learning method that combines the strengths of hand-crafted and data-driven approaches to achieve more accurate VRP modeling. Specifically, we design special incremental rules to learn the visual knowledge flow of MVE. To effectively capture the associated features of MVE, we further develop a three-stage incremental learning approach for VRP based on an encoder-decoder network. Extensive experimental results validate the superiority of the proposed method in terms of subjective, objective, and compression experiments. Miaohui Wang, Lirong Huang, Yanshan Li |
AAAI | 1 |
| 2024 | MetaJND: A Meta-Learning Approach for Just Noticeable Difference Estimation
Miaohui Wang, Yukuan Zhu, Wuyuan Xie |
IJCAI | 1 |
| 2024 | SPGNet: A Serial-Parallel Gated Convolutional Network for Image Classification on Small DatasetsabstractVision Transformers (ViTs) pose challenges in training and deploying deep models due to their lack of inductive biases. Previous literature integrated the key ingredients (e.g., long-range relations or input-adaptive weights) of ViTs into convolutional neural networks (CNNs) to address these bias issues on large-scale datasets like ImageNet-1K. However, the performance of these key ingredients on small-scale datasets has received little attention. In this paper, we have decomposed large-kernel convolution in a serial-parallel manner to extract multi-scale image features. By integrating them into a gated convolutional architecture, we have constructed a network backbone for the image classification in the small-scale dataset scenario, called SPGNet. Experiments on public small classification benchmark datasets show that SPGNet achieves a Top-1 accuracy of 86.62% on the CIFAR-100 and 76.57% on the Tiny ImageNet. Moreover, we have conducted experiments on the semantic segmentation task, and our method also achieves promising results under the similar architectures and training configurations. Yun Song, Jinxuan Wang, Miaohui Wang |
IJCNN | 3 |
| 2024 | Voxel Proposal Network via Multi-Frame Knowledge Distillation for Semantic Scene CompletionabstractSemantic scene completion is a difficult task that involves completing the geometry and semantics of a scene from point clouds in a large-scale environment. Many current methods use 3D/2D convolutions or attention mechanisms, but these have limitations in directly constructing geometry and accurately propagating features from related voxels, the completion likely fails while propagating features in a single pass without considering multiple potential pathways. And they are generally only suitable for static scenes and struggle to handle dynamic aspects. This paper introduces Voxel Proposal Network (VPNet) that completes scenes from 3D and Bird's-Eye-View (BEV) perspectives. It includes Confident Voxel Proposal based on voxel-wise coordinates to propose confident voxels with high reliability for completion. This method reconstructs the scene geometry and implicitly models the uncertainty of voxel-wise semantic labels by presenting multiple possibilities for voxels. VPNet employs Multi-Frame Knowledge Distillation based on the point clouds of multiple adjacent frames to accurately predict the voxel-wise labels by condensing various possibilities of voxel relationships. VPNet has shown superior performance and achieved state-of-the-art results on the SemanticKITTI and SemanticPOSS datasets. Lubo Wang, Di Lin 0002, Kairui Yang, Qing Guo 0005, Wuyuan Xie, Miaohui Wang, Lingyu Liang, Ping Li 0016 |
NeurIPS | 7 |
| 2024 | Fast CU Partition for VVC Intra-Frame Coding via Texture ComplexityabstractIn versatile video coding (VVC), the quadtree with nested multi-type tree (QTMT) partition module significantly improves encoding performance compared to other coding tools. However, it also introduces notable computational complexity in intra-frame coding, occupying over 90% of the encoding time. This paper presents a fast coding unit (CU) partition method based on texture complexity to achieve a balance between compression efficiency and computational complexity for VVC intra-frame coding. In particular, the texture complexity of CUs is quantitatively measured by the ratio of horizontal to vertical gradient and that of sub-block variances. Firstly, directions with higher texture complexity are identified as unlikely coding modes and eliminated from the candidate set. Next, the subblock variances of binary and ternary partitions are compared to determine a fine-grained CU partition pattern, avoiding unlikely partition modes. Experimental results show that our method is simple but efficient, and achieves higher computation efficiency compared to recent machine learning-based and handcraftedbased methods. The implementation of the proposed method is publicly available athttps://github.com/csust-sonie/fastCU. Yun Song, Shisheng Cheng, Miaohui Wang, Xiangrong Peng |
IEEE Signal Process. Lett. | 3 |
| 2024 | LLM-Guided Cross-Modal Point Cloud Quality Assessment: A Graph Learning ApproachabstractThis paper addresses the critical need for accurate and reliable point cloud quality assessment (PCQA) in various applications, such as autonomous driving, robotics, virtual reality, and 3D reconstruction. To meet this need, we propose a large language model (LLM)-guided PCQA approach based on graph learning. Specifically, we first utilize the LLM to generate quality description texts for each 3D object, and employ two CLIP-like feature encoders to represent the image and text modalities. Next, we design a latent feature enhancer module to improve contrastive learning, enabling more effective alignment performance. Finally, we develop a graph network fusion module that utilizes a ranking-based loss to adjust the relationship of different nodes, which explicitly considers both modality fusion and quality ranking. Experimental results on three benchmark datasets demonstrate the effectiveness and superiority of our approach over 12 representative PCQA methods, which demonstrate the potential of multi-modal learning, the importance of latent feature enhancement, and the significance of graph-based fusion in advancing the field of PCQA. Wuyuan Xie, Yunheng Liu, Kaiming Wang, Miaohui Wang |
IEEE Signal Process. Lett. | 4 |
| 2024 | Enhanced Dynamic Analysis for Malware Detection With Gradient AttackabstractMalware detection is an effective way to prevent the intrusion of malware into computer systems, and the API-based dynamic analysis method can effectively detect obfuscated and packaged malware. However, existing methods still suffer from limited detection accuracy and weak generalization. To address this issue, this paper presents a gradient attack-based malware dynamic analysis method. Through exerting adversarial noise into the embedding layer, the malware detection model can learn more robust representations of API sequences during training, achieving broader coverage of sample representations. The strategy of normalizing attack noise and recovering attacked representation is designed, which controls the strength of the gradient attack within a reasonable range and prevents a negative impact on the model's detection performance. The proposed method can be applied to existing API-based malware detection models to enhance their detection performance, indicating the strong generality of the proposed method. Experimental results on two benchmark datasets (i.e.,AliyunandCatak) demonstrate the effectiveness of the proposed gradient attack method, which further improves the detection performance of the mainstream API-based models, with an average accuracy increase of 2.80% and 3.66% on these two datasets, respectively. Pei Yan, Shunquan Tan, Miaohui Wang, Jiwu Huang |
IEEE Signal Process. Lett. | 3 |
| 2024 | ReferPose: Distance Optimization-Based Reference Learning for Human Pose Estimation and MonitoringabstractExisting deep learning models for human pose estimation (HPE) have shown satisfactory performance in monitoring human actions. However, they usually face a dilemma between complexity and accuracy. To address this challenge, we propose an effective reference learning method for HPE (namely ReferPose), which is based on a new distance optimization strategy. Specifically, we utilize a reference model for pose learning and representation. The pose representation learned from the entire database is merged into the reference model, providing continuous reference learning guidance for an in-training model. In addition, we design a new cosine annealing-based reference guidance for temporal denoising and further develop a distance optimization strategy to provide joint guidance from pose knowledge, model representation, and temporal experience. Experimental results on two benchmark databases and a human fall monitoring system demonstrate that our ReferPose not only achieves promising accuracy improvement compared with several representative HPE models, but also offers low cost and high efficiency. Miaohui Wang, Zhuowei Xu, Wuyuan Xie |
IEEE Trans. Ind. Informatics | 1 |
| 2024 | BinaryFormer: A Hierarchical-Adaptive Binary Vision Transformer (ViT) for Efficient ComputingabstractVision Transformer (ViT) has recently demonstrated impressive nonlinear modeling capabilities and achieved state-of-the-art performance in various industrial applications, such as object recognition, anomaly detection, and robot control. However, their practical deployment can be hindered by high storage requirements and computational intensity. To alleviate these challenges, we propose a binary transformer called BinaryFormer, which quantizes the learned weights of the ViT module from 32-b precision to 1 b. Furthermore, we propose a hierarchical-adaptive architecture that replaces expensive matrix operations with more affordable addition and bit operations by switching between two attention modes. As a result, BinaryFormer is able to effectively compress the model size as well as reduce the computation cost of ViT. Experimental results on the ImageNet-1K benchmark datasets show that BinaryFormer reduces the size of a typical ViT model by an average of 27.7× and converts over 99% of multiplication operations into bit operations while maintaining reasonable accuracy. Miaohui Wang, Zhuowei Xu, Wuyuan Xie |
IEEE Trans. Ind. Informatics | 1 |
| 2024 | Weight-Based Distributed Formation Control for Networked Marine Surface Vehicles With Hybrid Communication Channel Deception AttacksabstractThe control interaction in networked marine surface vehicles (NMSVs) mainly involves three communication channels, i.e., vehicle-to-vehicle, sensor-to-controller, and controller-to-actuator channels. Existing works have given defense solutions to attacks on a single communication channel while ignoring attacks on hybrid communication channels. Motivated by this observation, this article studies the defense strategy of hybrid communication channel deception attacks (HCCDAs) in distributed NMSV formation control. A technical obstacle is to isolate the influence of HCCDAs without compromising formation integrality. To deal with this, a weight-based adjacency matrix is developed to dynamically evaluate the formation credibility. It is capable of suppressing the spread of the influence from attacked members throughout the entire formation without severing them. Meanwhile, the consensus over the weights regarding the same member and the natural boundedness of the weights render the weight-based adjacency matrix more readily applicable. Besides, a hierarchical observer is developed to estimate both the matched and mismatched uncertainties stemming from HCCDAs such that attacked members can keep working normally. Accordingly, an anti-HCCDA distributed formation control scheme is synthesized. The effectiveness of the proposed control scheme is shown by the robustness and comparison simulations. Cheng Zhu 0001, Miaohui Wang |
IEEE Trans. Ind. Informatics | 2 |
| 2023 | Just Noticeable Visual Redundancy Forecasting: A Deep Multimodal-Driven ApproachabstractJust noticeable difference (JND) refers to the maximum visual change that human eyes cannot perceive, and it has a wide range of applications in multimedia systems. However, most existing JND approaches only focus on a single modality, and rarely consider the complementary effects of multimodal information. In this article, we investigate the JND modeling from an end-to-end homologous multimodal perspective, namely hmJND-Net. Specifically, we explore three important visually sensitive modalities, including saliency, depth, and segmentation. To better utilize homologous multimodal information, we establish an effective fusion method via summation enhancement and subtractive offset, and align homologous multimodal features based on a self-attention driven encoder-decoder paradigm. Extensive experimental results on eight different benchmark datasets validate the superiority of our hmJND-Net over eight representative methods. Wuyuan Xie, Shukang Wang, Sukun Tian, Lirong Huang, Ye Liu 0005, Miaohui Wang |
AAAI | 6 |
| 2023 | CVSformer: Cross-View Synthesis Transformer for Semantic Scene CompletionabstractSemantic scene completion (SSC) requires an accurate understanding of the geometric and semantic relationships between the objects in the 3D scene for reasoning the occluded objects. The popular SSC methods voxelize the 3D objects, allowing the deep 3D convolutional network (3D CNN) to learn the object relationships from the complex scenes. However, the current networks lack the controllable kernels to model the object relationship across multiple views, where appropriate views provide the relevant information for suggesting the existence of the occluded objects. In this paper, we propose Cross-View Synthesis Transformer (CVSformer), which consists of Multi-View Feature Synthesis and Cross-View Transformer for learning cross-view object relationships. In the multi-view feature synthesis, we use a set of 3D convolutional kernels rotated differently to compute the multi-view features for each voxel. In the cross-view transformer, we employ the cross-view fusion to comprehensively learn the cross-view relationships, which form useful information for enhancing the features of individual views. We use the enhanced features to predict the geometric occupancies and semantic labels of all voxels. We evaluate CVSformer on public datasets, where CVS-former yields state-of-the-art results. Our code is available at https://github.com/donghaotian123/CVSformer. Haotian Dong, Enhui Ma, Lubo Wang, Miaohui Wang, Wuyuan Xie, Qing Guo 0005, Ping Li 0016, Lingyu Liang, Kairui Yang, Di Lin 0002 |
ICCV | 4 |
| 2023 | Patch-Wise LiDAR Point Cloud Geometry Compression Based on Autoencoder
Runnan Huang, Miaohui Wang |
ICIG (3) | 2 |
| 2023 | Just Noticeable Difference Estimation for Screen Content Images: A Content Uncertainty-guided ApproachabstractJust-noticeable-difference (JND) effectively describes the threshold of the human visual system (HVS) perceiving visual signal changes, which reflects the visual redundancy. Researches indicate that the information distribution characteristics of images play an important role in the information inference of HVS. Inspired by this, we present a content uncertainty-guided JND estimation approach for screen content images (SCIs). Specifically, we divide SCIs into certain content and uncertain content according to content uncertainty modeling. Based on the analysis of content characteristics, we apply oblique masking (OM) and contrast masking (CM) in uncertain content, and consider luminance adaptation (LA) and blur adaptation (BA) in certain content. Besides, we adjust parameters with subjective experiments to better fit HVS characteristics. Compared with several state-of-the-art JND models, our method tolerates more distortion and owns better perceptual performance under the same injected-noise energy. Lirong Huang, Miaohui Wang |
ICME | 3 |
| 2023 | Rethinking Video Error Concealment: A Benchmark DatasetabstractError concealment is an important technique to restore a damaged video bistream. Although data-driven in-painting methods can be directly applied to video error concealment, existing mask patterns are remarkably different from the practical damaged video bitstream, which causes a great impact on the repairing effect of video quality. To rethink the gap between existing inpainting schemes and practical video transmission characteristics, we have established a new video error concealment (VEC) benchmark dataset. Specifically, different video sequences compressed by different encoders are collected, and various loss types are generated to satisfy different packet loss scenarios. Based on VEC, error-concealed results of existing methods are provided and analyzed, which can serve as a benchmark for video error concealment research. Miaohui Wang |
ICME | 2 |
| 2023 | 3D Surface Super-resolution from Enhanced 2D Normal Images: A Multimodal-driven Variational AutoEncoder Approachabstract3D surface super-resolution is an important technical tool in virtual reality, and it is also a research hotspot in computer vision. Due to the unstructured and irregular nature of 3D object data, it is usually difficult to obtain high-quality surface details and geometry textures via a low-cost hardware setup. In this paper, we establish a multimodal-driven variational autoencoder (mmVAE) framework to perform 3D surface enhancement based on 2D normal images. To fully leverage the multimodal learning, we investigate a multimodal Gaussian mixture model (mmGMM) to align and fuse the latent feature representations from different modalities, and further propose a cross-scale encoder-decoder structure to reconstruct high-resolution normal images. Experimental results on several benchmark datasets demonstrate that our method delivers promising surface geometry structures and details in comparison with competitive advances. Wuyuan Xie, Tengcong Huang, Miaohui Wang |
IJCAI | 3 |
| 2023 | A Method of Micro-Geometric Details Preserving in Surface Reconstruction from GradientabstractSurface from gradient (SfG) is one of the fundamental methods to densely reconstruct 3D object surface in computer vision. However, the reconstruction of micro-geometric details has not been satisfactorily solved in existing SfG methods due to their non-integrability. In this paper, we present an effective discrete geometric approach to reconstruct fine-grained sharp surface feature with non-integrability. Specifically, We investigate the fine-grained structure of surfaces in the micro geometry domain. based on an adaptive projection on vertexes constrained by neighboring gradient vectors, and develop a gradient angle-guided energy optimization to generate a fine-grained surface. Experimental results on various challenging synthetic and real-world data show that the proposed method is able to effectively reconstruct challenging micro-geometric details for general SfG methods. Wuyuan Xie, Miaohui Wang |
ACM Multimedia | 2 |
| 2023 | pmBQA: Projection-based Blind Point Cloud Quality Assessment via Multimodal LearningabstractWith the increasing communication and storage of point cloud data, there is an urgent need for an effective objective method to measure the quality before and after processing. To address this difficulty, we propose a projection-based blind quality indicator via multimodal learning for point cloud data, which can perceive both geometric distortion and texture distortion by using four homogeneous modalities (i.e., texture, normal, depth and roughness). To fully exploit the multimodal information, we further develop a deformable convolutionbased alignment module and a graph-based feature fusion module, and investigate a graph node attention-based evaluation method to forecast the quality score. Extensive experimental results on three benchmark databases show that our method achieves more accurate evaluation performance in comparison with 12 competitive methods. Wuyuan Xie, Kaimin Wang, Yakun Ju, Miaohui Wang |
ACM Multimedia | 4 |
| 2023 | Visual Redundancy Removal of Composite Images via Multimodal LearningabstractComposite images are generated by combining two or more different photographs, and their content is typically heterogeneous. However, existing unimodal visual redundancy prediction methods are difficult to accurately model the complex characteristics of this image type. In this paper, we investigate the visual redundancy modeling of composite images from an end-to-end multimodal perspective, including four cross-media modalities (i.e., text, brightness, color, and segmentation). Specifically, we design a two-stage cross-modal alignment module based on self-attention mechanism and contrastive learning, and develop a fusion module based on a cross-modal augmentation paradigm. Further, we establish the first cross-media visual redundancy dataset for composite images, which contains 413 groups of cross-modal data and generates 13629 realistic compression distortions using the latest versatile video coding (VVC) standard. Experimental results on nine benchmark datasets demonstrate the effectiveness of our method, outperforming seven representative methods. Wuyuan Xie, Shukang Wang, Miaohui Wang |
ACM Multimedia | 4 |
| 2023 | Surface Geometry Processing: An Efficient Normal-Based Detail RepresentationabstractWith the rapid development of high-resolution 3D vision applications, the traditional way of manipulating surface detail requires considerable memory and computing time. To address these problems, we introduce an efficient surface detail processing framework in 2D normal domain, which extracts new normal feature representations as the carrier of micro geometry structures that are illustrated both theoretically and empirically in this article. Compared with the existing state of the arts, we verify and demonstrate that the proposed normal-based representation has three important properties, including detail separability, detail transferability and detail idempotence. Finally, three new schemes are further designed for geometric surface detail processing applications, including geometric texture synthesis, geometry detail transfer, and 3D surface super-resolution. Theoretical analysis and experimental results on the latest benchmark dataset verify the effectiveness and versatility of our normal-based representation, which accepts 30 times of the input surface vertices but at the same time only takes 6.5% memory cost and 14.0% running time in comparison with existing competing algorithms. Wuyuan Xie, Miaohui Wang, Di Lin 0002, Boxin Shi, Jianmin Jiang |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | A Task-Driven Scene-Aware LiDAR Point Cloud Coding Framework for Autonomous VehiclesabstractLiDAR sensors are almost indispensable for autonomous robots to perceive the surrounding environment. However, the transmission of large-scale LiDAR point clouds is highly bandwidth-intensive, which can easily lead to transmission problems, especially for unstable communication networks. Meanwhile, existing LiDAR data compression is mainly based on rate-distortion optimization, which ignores the semantic information of ordered point clouds and the task requirements of autonomous robots. To address these challenges, this article presents a task-driven Scene-Aware LiDAR Point Clouds Coding (SA-LPCC) framework for autonomous vehicles. Specifically, a semantic segmentation model is developed based on multidimension information, in which both 2-D texture and 3-D topology information are fully utilized to segment movable objects. Furthermore, a prediction-based deep network is explored to remove the spatial–temporal redundancy. The experimental results on the benchmark semantic KITTI dataset validate that our SA-LPCC achieves state-of-the-art performance in terms of the reconstruction quality and storage space for downstream tasks. We believe that SA-LPCC jointly considers the scene-aware characteristics of movable objects and removes the spatial–temporal redundancy from an end-to-end learning mechanism, which will boost the related applications from algorithm optimization to industrial products. Xuebin Sun, Miaohui Wang, Jingxin Du, Yuxiang Sun 0002, Shing Shin Cheng, Wuyuan Xie |
IEEE Trans. Ind. Informatics | 2 |
| 2023 | Low-Light Images In-the-Wild: A Novel Visibility Perception-Guided Blind Quality IndicatorabstractOwing to the increasing deployment of CMOS camera modules, it is inevitable to take photographs under weak illumination. Therefore, low-light imaging quality is one of the most important factors affecting user experience as well as the product values of consumer electronics, automobile, surveillance, factory automation, and other industrial applications. Inspired by human vision, this article jointly considersvisibility perception,luminosity cognition, andcolor sensationand presents a new visibility perception-guided blind quality indicator for low-light images in-the-wild. To excavate effective descriptors for authentic distortions under weak illumination, we utilize maximum ignorable visible difference to characterize the reduced visibility, and employ the luminance statistical properties and color sensation characteristics to represent brightness and colorfulness distortions. Extensive experimental results on the benchmark dataset verify that the proposed blind quality indicator outperforms nine representative methods including general-purpose and distortion-specific methods. Miaohui Wang, Jian Xiong 0005, Wuyuan Xie |
IEEE Trans. Ind. Informatics | 1 |
| 2023 | Ship Collision Avoidance Navigation Signal Recognition via Vision Sensing and Machine ForecastingabstractShip collision avoidance (SCA) is an important technique in the field of decision-making in marine navigation. Although some promising solutions have been developed recently, there is still the lack of low-cost and reliable sensing equipment. Inspired by the low-cost of camera sensors and the success of machine learning, this paper designs a vision-based method to recognize ships and their micro-features for SCA navigation planning. Firstly, we develop a vision-based bearing, distance and velocity model based on a wide-field optical imaging system. Secondly, optical information is used to construct the micro-characteristic imaging model of ship navigation signals. Thirdly, we have solved the problem between a large field-of-view (FOV) and high-resolution imaging in vision-based marine navigation. Finally, an improved Adaboost algorithm is designed for the intelligent recognition of an open-sea target (ship types and light patterns). The proposed method has been verified by extensive experiments in a practical environment, and the results show that it can effectively and efficiently identify the navigation signal of a target ship. Qilin Bi, Miaohui Wang, Minling Lai, Xiuying Bi |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2023 | Efficient Geometry Surface Coding in V-PCCabstractIn recent video-based point cloud compression (V-PCC), 3D point clouds are projected onto 2D images and compressed by High-Efficiency Video Coding (HEVC). However, HEVC was originally designed for natural visual signals, which is a suboptimal framework for point clouds. Therefore, there are still problems in geometry information compression in V-PCC: (1) The distortion based on the sum of squared error (SSE) in the existing rate-distortion optimization (RDO) is inconsistent with the geometric quality measurement; (2) The existing prediction cannot explore the fixed relationship between the corresponding far layer and near layer depth, which means that the far layer depth can be always not less than the corresponding near layer depth. In this paper, we present an efficient geometry surface coding (EGSC) method for V-PCC to address the problems. Firstly, an error projection (EP) model is designed to establish the relationship between the SSE-based distortion and the geometry quality metric. Secondly, an EP-based RDO is employed to improve the geometry information compression by estimating the point normals with gradients. Finally, an occupancy-map driven scheme is proposed to improve the prediction accuracy of merge modes. Experimental results show that the proposed method achieves an average of over 10% bit-rate saving compared with the V-PCC reference software. Jian Xiong 0005, Hao Gao 0005, Miaohui Wang, Hongliang Li 0001, King Ngi Ngan, Weisi Lin |
IEEE Trans. Multim. | 3 |
| 2022 | MNSRNet: Multimodal Transformer Network for 3D Surface Super-ResolutionabstractWith the rapid development of display technology, it has become an urgent need to obtain realistic 3D surfaces with as high-quality as possible. Due to the unstructured and irregular nature of 3D object data, it is usually difficult to obtain high-quality surface details and geometry textures at a low cost. In this article, we propose an effective multimodal-driven deep neural network to perform 3D surface super-resolution in 2D normal domain, which is simple, accurate, and robust to the above difficulty. To leverage the multimodal information from different perspectives, we jointly consider the texture, depth, and normal modalities to simultaneously restore fine-grained surface details as well as preserve geometry structures. To better utilize the cross-modality information, we explore a two-bridge normal method with a transformer structure for feature alignment, and investigate an affine transform module for fusing multimodal features. Extensive experimental results on public and our newly constructed photometric stereo dataset demonstrate that the proposed method delivers promising surface geometry details compared with nine competitive schemes. Wuyuan Xie, Tengcong Huang, Miaohui Wang |
CVPR | 3 |
| 2022 | S-CCR: Super-Complete Comparative Representation for Low-Light Image Quality Inference In-the-wildabstractWith the rapid development of weak-illumination imaging technology, low-light images have brought new challenges to quality of experience and service. However, developing a robust quality indicator for authentic low-light distortions in-the-wild remains a major challenge in practical quality control systems. In this paper, we develop a new super-complete comparative representation (S-CCR) for the region-level quality inference of low-light images. Specifically, we excavate the color, luminance, and detail quality evidence for the feature embedding guidance of comparative representation based on the human visual characteristics. Moreover, we decompose the inputs into a super-complete feature group so that the image quality of each region can be fully represented, which allows to preserve the distinctiveness, distinguishability, and consistency. Finally, we further establish a comparative domain alignment method, so that the comparative representation of an unseen image can be aligned with respect to the quality features of already-seen ones. Extensive experiments on the benchmark dataset validate the superiority of our S-CCR over 11 competing methods on authentic distortions. Miaohui Wang, Zhuowei Xu, Yuanhao Gong, Wuyuan Xie |
ACM Multimedia | 1 |
| 2022 | Generative Status Estimation and Information Decoupling for Image Rain RemovalabstractImage rain removal requires the accurate separation between the pixels of the rain streaks and object textures. But the confusing appearances of rains and objects lead to the misunderstanding of pixels, thus remaining the rain streaks or missing the object details in the result. In this paper, we propose SEIDNet equipped with the generative Status Estimation and Information Decoupling for rain removal. In the status estimation, we embed the pixel-wise statuses into the status space, where each status indicates a pixel of the rain or object. The status space allows sampling multiple statuses for a pixel, thus capturing the confusing rain or object. In the information decoupling, we respect the pixel-wise statuses, decoupling the appearance information of rain and object from the pixel. Based on the decoupled information, we construct the kernel space, where multiple kernels are sampled for the pixel to remove the rain and recover the object appearance. We evaluate SEIDNet on the public datasets, achieving state-of-the-art performances of image rain removal. The experimental results also demonstrate the generalization of SEIDNet, which can be easily extended to achieve state-of-the-art performances on other image restoration tasks (e.g., snow, haze, and shadow removal). Di Lin 0002, Xin Wang 0118, Miaohui Wang, Wuyuan Xie, Qing Guo 0005, Ping Li 0016 |
NeurIPS | 6 |
| 2022 | MSCI: A Multi-Source Composite Image Database for Compression Distortion Quality AssessmentabstractWith the rapid development of multi-sensor fusion technology in various industrial fields, many composite images closely related to human life have been produced. To meet the rapidly growing needs of various image-based applications, we have established the first multi-source composite image (MSCI) database for image quality assessment (IQA). Our MSCI database contains 80 reference images and 1600 distorted images, generated by four advanced compression standards with five distortion levels. In particular, these five distortion levels are determined based on the first five just noticeable difference (JND) levels. Moreover, we verify the IQA performance of some representative methods on our MSCI database. The experimental results show that the performance of the existing methods on the MSCI database needs to be further improved. Zhuowei Xu, Zhiheng Lin, Miaohui Wang |
VCIP | 4 |
| 2022 | Occupancy Map Guided Fast Video-Based Dynamic Point Cloud CodingabstractIn video-based dynamic point cloud compression (V-PCC), 3D point clouds are projected into patches, and then the patches are padded into 2D images suitable for the video compression framework. However, the patch projection-based method produces a large number of empty pixels; the far and near components are projected to generate different 2D images (video frames), respectively. As a result, the generated video is with high resolutions and double frame rates, so the V-PCC has huge computational complexity. This paper proposes an occupancy map guided fast V-PCC method. Firstly, the relationship between the prediction coding and block complexity is studied based on a local linear image gradient model. Secondly, according to the V-PCC strategies of patch projection and block generation, we investigate the differences of rate-distortion characteristics between different types of blocks, and the temporal correlations between the far and near layers. Finally, by taking advantage of the fact that occupancy maps can explicitly indicate the block types, we propose an occupancy map guided fast coding method, in which coding is performed on the different types of blocks. Experiments have tested typical dynamic point clouds, and shown that the proposed method achieves an average 43.66% time-saving at the cost of only 0.27% and 0.16% Bjontegaard Delta (BD) rate increment under the geometry Point-to-Point (D1) error and attribute Luma Peak-Signal-Noise-Ratio (PSNR), respectively. Jian Xiong 0005, Hao Gao 0005, Miaohui Wang, Hongliang Li 0001, Weisi Lin |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Perceptually Quasi-Lossless Compression of Screen Content Data Via Visibility Modeling and Deep ForecastingabstractScreen content data, such as computer-generated photographs, desktop sharing, remote education, video game streaming and screenshot, is one of the most popular visual information carriers in Internet of Video Things. Although lossless compression can guarantee high quality of service for these screen content based industrial applications, it also causes considerable storage space and transmission bandwidth issues. To alleviate these challenges, in this article, we present a visually quasi-lossless coding approach to control the compression distortion belowvisibility thresholdin the human visual system. Specifically, to better quantify the visual redundancy for screen content data, a newvisibility thresholdmethod is designed by incorporating blur sensitivity and oblique correction effects. Then, an end-to-end mapping between thevisibility thresholdand quality control factor is learned and represented as a deep convolutional neural network. The experimental results demonstrate that the proposed method saves the average encoding bits up to 23.15% compared with the latest scheme under the same perceptual quality. Miaohui Wang, Zhuowei Xu, Jian Xiong 0005, Wuyuan Xie |
IEEE Trans. Ind. Informatics | 1 |
| 2022 | DCPR-GAN: Dental Crown Prosthesis Restoration Using Two-Stage Generative Adversarial NetworksabstractRestoring the correct masticatory function of broken teeth is the basis of dental crown prosthesis rehabilitation. However, it is a challenging task primarily due to the complex and personalized morphology of the occlusal surface. In this article, we address this problem by designing a new two-stage generative adversarial network (GAN) to reconstruct a dental crown surface in the data-driven perspective. Specifically, in the first stage, a conditional GAN (CGAN) is designed to learn the inherent relationship between the defective tooth and the target crown, which can solve the problem of the occlusal relationship restoration. In the second stage, an improved CGAN is further devised by considering an occlusal groove parsing network (GroNet) and an occlusal fingerprint constraint to enforce the generator to enrich the functional characteristics of the occlusal surface. Experimental results demonstrate that the proposed framework significantly outperforms the state-of-the-art deep learning methods in functional occlusal surface reconstruction using a real-world patient database. Moreover, the standard deviation (SD) and root mean square (RMS) between the generated occlusal surface and the target crown calculated by our method are both less than 0.161 mm. Importantly, the designed dental crown have enough anatomical morphology and higher clinical applicability. Sukun Tian, Miaohui Wang, Luca Fiorenza, Yuchun Sun, Yangmin Li 0001 |
IEEE J. Biomed. Health Informatics | 2 |
| 2021 | Residual Geometric Feature Transform Network for 3D Surface Super-Resolution
Maolin Cui, Wuyuan Xie, Miaohui Wang, Tengcong Huang |
3DV | 3 |
| 2021 | Initial-QP Prediction for Versatile Video Coding: A Multi-domain Feature-Driven Learning Approach
Lirong Huang, Miaohui Wang |
ICIG (1) | 3 |
| 2021 | Deep Human Pose Estimation via Self-guided Learning
Zhuowei Xu, Miaohui Wang |
ICIG (2) | 2 |
| 2021 | Blind Quality Assessment of Night-Time Images Via Weak Illumination AnalysisabstractNight-time images can be generated by various camera sensors in many practical applications, and hence how to effectively evaluate the quality of night-time images is an essential research topic. In this article, we propose a novel blind quality assessment method for night-time images. By analyzing the characteristic of night-time images, we first investigate the statistical properties of local luminance information based on the brightness level division, and then measure the masking effect on color and structure information caused by weak illumination. Finally, all extracted quality-aware features and the associated subjective ratings are trained via support vector regression to build the quality assessment model. Extensive experiments on a real-world night-image database validate the superiority of the proposed method over several state-of-the-art blind image quality metrics1. Miaohui Wang |
ICME | 1 |
| 2021 | Machine Learning-Based Rate Distortion Modeling for VVC/H.266 Intra-FrameabstractRate-distortion (R-D) optimization has been widely adopted to improve the coding efficiency on the video encoder side. However, there are few studies related to modeling the R-D characteristics of the latest Versatile Video Coding (VVC) reference software. In this paper, we investigate the R-D modeling of the intra-frame on the VVC encoder by utilizing four traditional machine learning algorithms. We extract four highly descriptive features to capture the relationship between the video content and the R-D model. Moreover, it is applied to the initial intra-frame rate control of VVC. Experimental results show that our method outperforms VTM-7.0, which improves the accuracy by up to 8.65% with affordable computational complexity increase1. Miaohui Wang, Lirong Huang, Jian Xiong 0005 |
ICME | 1 |
| 2021 | Perceptual Redundancy Estimation of Screen Images via Multi-Domain SensitivitiesabstractVisual redundancy detection is essential for image and video communication. Human visual system (HVS) is difficult to perceive the pixel magnitude change below a certain visibility threshold which is also known as just-noticeable-difference (JND). In this letter, we present an efficient JND estimation approach for screen content images by considering high-frequency sensitivity and orientation sensitivity correction. Specifically, to better quantify the visual redundancy, we investigate the visibility threshold based on the high-frequency distortion sensitivity. To obtain the orientation sensitivity correction, we divide the screen image pixels into three levels based on the oblique effect that considers the sensitive integrity of edges. Compared with several state-of-the-art JNDs, experimental results show that our method tolerates more perceptual redundancy, and delivers better visual quality under the same injected-noise energy. The implementation of the proposed method is publicly available at https://sites.google.com/site/wangmiaohui/. Miaohui Wang, Wuyuan Xie, Long Xu 0001 |
IEEE Signal Process. Lett. | 1 |
| 2021 | SAR Speckle Removal Using Hybrid Frequency ModulationsabstractSynthetic aperture radar (SAR) images often interfere with speckle artifacts that have a great impact on subsequent processing and analysis operations. To remove speckle artifacts, this article introduces a hybrid denoising approach by using a convolutional neural network (CNN) and consistent cycle spinning (CCS) in the nonsubsample shearlet transform (NSST) domain. First, we apply NSST to a noisy SAR image to gain low- and high-frequency coefficients. Second, we adopt a learned deep CNN model to eliminate the speckle noise in the low-frequency coefficients, which retains more contour information. Third, we employ CCS to enhance the high-frequency coefficients, which preserves more details of the original SAR image. Finally, we obtain the denoised image by using inverse NSST applied to the denoised coefficients. Compared with state-of-the-art algorithms, the results of the experiment indicate that our method not only achieves better speckle removal performance but also maintains more detailed information retention. Shuaiqi Liu 0001, Lele Gao, Miaohui Wang, Xiaole Ma, Yudong Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2021 | Efficient Computer-Aided Design of Dental Inlay Restoration: A Deep Adversarial FrameworkabstractRestoring the normal masticatory function of broken teeth is a challenging task primarily due to the defect location and size of a patient's teeth. In recent years, although some representative image-to-image transformation methods (e.g. Pix2Pix) can be potentially applicable to restore the missing crown surface, most of them fail to generate dental inlay surface with realistic crown details (e.g. occlusal groove) that are critical to the restoration of defective teeth with varying shapes. In this article, we design a computer-aided Deep Adversarial-driven dental Inlay reStoration (DAIS) framework to automatically reconstruct a realistic surface for a defective tooth. Specifically, DAIS consists of a Wasserstein generative adversarial network (WGAN) with a specially designed loss measurement, and a new local-global discriminator mechanism. The local discriminator focuses on missing regions to ensure the local consistency of a generated occlusal surface, while the global discriminator aims at defective teeth and adjacent teeth to assess if it is coherent as a whole. Experimental results demonstrate that DAIS is highly efficient to deal with a large area of missing teeth in arbitrary shapes and generate realistic occlusal surface completion. Moreover, the designed watertight inlay prostheses have enough anatomical morphology, thus providing higher clinical applicability compared with more state-of-the-art methods. Sukun Tian, Miaohui Wang, Fulai Yuan, Yuchun Sun, Wuyuan Xie, Harry Qin |
IEEE Trans. Medical Imaging | 2 |
| 2020 | An Advanced LiDAR Point Cloud Sequence Coding Scheme for Autonomous DrivingabstractDue to the huge volume of point cloud data, storing or transmitting it is currently difficult and expensive in autonomous driving. Learning from the high efficiency video coding (HEVC) coding framework, we propose an advanced coding scheme for large-scale LiDAR point cloud sequences, in which several techniques have been developed to remove the spatial and temporal redundancy. The proposed strategy consists mainly of intra-coding and inter-coding. For intra-coding, we utilize a cluster-based prediction method to remove the spatial redundancy. For inter-coding, a predictive recurrent network is designed, which is capable of generating future frames according to the previously encoded frames. By calculating the residual error between the predicted and real point cloud data, the temporal redundancy can be removed. Finally, the residual data is quantized and encoded by lossless coding schemes. Experiments are conducted on the KITTI data set with four different scenes to verify the effectiveness and efficiency of the proposed method. Our approach can deal with multiple types of point cloud data from the simple to more complex, and yields better performance in terms of compression ratio compared with octree, Google Draco, MPEG TMC13 and other recently proposed methods. Xuebin Sun, Sukai Wang, Miaohui Wang, Shing Shin Cheng, Ming Liu 0001 |
ACM Multimedia | 3 |
| 2020 | Surface Reconstruction with Unconnected Normal Maps: An Efficient Mesh-based ApproachabstractNormal integration is a key step in dense 3D reconstruction methods such as shape-from-shading and photometric stereo. However, normal integration cannot be guaranteed between spatially unconnected normal maps, which can ultimately cause a shape deformation in surface-from-normals (SfN). For the first time, this paper presents an efficient approach to address the fundamental problem of surface reconstruction from unconnected normal maps (denoted as "SfN+") using discrete geometry. We first design a normal piece pairing metric to measure the virtually pairing quality between two unconnected normal fragments, which is used as a new constraint for the boundary vertexes during mesh deformation. We then adopt a normal connecting significance indicator to adjust the influence of virtually connected vertexes, which further improves the overall shape deformation. Finally, we model the shape reconstruction of unconnected normal maps as a light-weight energy optimization framework by jointly considering the relaxation of connecting constraints and overall reconstruction error. Experiments show that the proposed SfN+ achieves a robust and efficient performance on dense 3D surface reconstruction. Miaohui Wang, Wuyuan Xie, Maolin Cui |
ACM Multimedia | 1 |
| 2020 | Industrial Applications of Ultrahigh Definition Video Coding With an Optimized Supersample Adaptive Offset FrameworkabstractThis article presents an efficient superblock-based sample adaptive offset (superSAO) that jointly exploits block-wise filter partition and probabilistic band interval segmentation to improve quality-of-experience of industrial video applications. Specifically, we investigate the partition flexibility of a superSAO block whose root size is up to 256 × 256, and propose to optimize the block-wise SAO filter by considering computation complexity and compression efficiency. Furthermore, we segment the band interval by equal probability of the sample intensity distributions, which facilitates the computation of better band offsets to attenuate ringing artifact due to quantization errors or encoded motion vectors. Experimental results show that the proposed superSAO method outperforms state-of-the-art approaches by obtaining 4.6% bandwidth reduction on average for the low delay and high-compression video applications. Miaohui Wang, Wuyuan Xie, Jia Zhang 0002, Harry Qin |
IEEE Trans. Ind. Informatics | 1 |
| 2020 | Rate Constrained Multiple-QP Optimization for HEVCabstractIn High Efficiency Video Coding (HEVC), multiple-QP (quantization parameter) optimization can adapt to a local video content. However, the multiple-QP implementation in the HEVC reference software (HM 16.6) achieves the best QP value for each coding block with a large amount of computational complexity. To address this challenge, we propose a fast rate-constrained multiple-QP optimization approach for the HM platform. We first introduce a template-based transform coefficient selection method which can save the overall complexity of entropy coding. In addition, we model the multiple-QP determination as a new rate-constrained optimization problem, and finally, we get a feasible solution with a lower computation overhead. Experimental results show that our method dramatically reduces the average complexity under the all-intra, low-delay and random-access configuration. Miaohui Wang, Jian Xiong 0005, Long Xu 0001, Wuyuan Xie, King Ngi Ngan, Harry Qin |
IEEE Trans. Multim. | 1 |
| 2019 | Surface Reconstruction From Normals: A Robust DGP-Based Discontinuity Preservation ApproachabstractIn 3D surface reconstruction from normals, discontinuity preservation is an important but challenging task. However, existing studies fail to address the discontinuous normal maps by enforcing the surface integrability in the continuous domain. This paper introduces a robust approach to preserve the surface discontinuity in the discrete geometry way. Firstly, we design two representative normal incompatibility features and propose an efficient discontinuity detection scheme to determine the splitting pattern for a discrete mesh. Secondly, we model the discontinuity preservation problem as a light-weight energy optimization framework by jointly considering the discontinuity detection and the overall reconstruction error. Lastly, we further shrink the feasible solution space to reduce the complexity based on the prior knowledge. Experiments show that the proposed method achieves the best performance on an extensive 3D dataset compared with the state-of-the-arts in terms of mean angular error and computational complexity. Wuyuan Xie, Miaohui Wang, Mingqiang Wei, Jianmin Jiang, Harry Qin |
CVPR | 2 |
| 2019 | Image super-resolution via feature-augmented random forest
Kin-Man Lam 0001, Miaohui Wang |
Signal Process. Image Commun. | 3 |
| 2019 | UHD Video Coding: A Light-Weight Learning-Based Fast Super-Block ApproachabstractThe ultra high-definition (UHD) video format, which has recently become popular, aims to provide high spatial resolution, high temporal frame rate, high sample bit-depth, and wide pixel color gamut. Despite the continued development of global network capacities, it inevitably causes the increased bandwidth cost of catering to the requirement of delivering UHD video services. To address such challenges, this paper presents an improved super coding unit (SCU) method for UHD video coding in High Efficiency Video Coding (HEVC). Initially, the medium coding unit (MCU) is proposed to avoid unnecessary brute-force coding unit (CU) partitions of SCU. Furthermore, the SCU is proposed to be encoded by Direct-MCU and SCU-to-MCU modes: the Direct-MCU mode is intended to better adapt to the texture-rich region, which guarantees the compression efficiency by avoiding extra-size CU partition; the SCU-to-MCU mode is designed for the homogeneous region of UHD content, which saves the encoding time by skipping fine-grained CU partition search. Moreover, a learning-based fast SCU decision approach is proposed to speed up the determination process of Direct-MCU and SCU-to-MCU, where three representative handcrafted features are extracted. Experimental results show that our method achieves an affordable complexity and excellent coding efficiency (up to 7.30% Bjøntegaard Delta rate savings) in UHD video coding compared to recent HEVC reference software. Miaohui Wang, Wuyuan Xie, Xiandong Meng, Huanqiang Zeng, King Ngi Ngan |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2018 | Quality Classified Image Analysis with Application to Face Detection and RecognitionabstractMotion blur, out of focus, insufficient spatial resolution, lossy compression and many other factors can all cause an image to have poor quality. However, image quality is a largely ignored issue in traditional pattern recognition literature. In this paper, we use face detection and recognition as case studies to show that image quality is an essential factor which will affect the performances of traditional algorithms. We demonstrated that it is not the image quality itself that is the most important, but rather the quality of the images in the training set should have similar quality as those in the testing set. To handle real-world application scenarios where images with different kinds and severities of degradation can be presented to the system, we have developed a quality classified image analysis framework to deal with images of mixed qualities adaptively. We use deep neural networks first to classify images based on their quality classes and then design a separate face detector and recognizer for images in each quality class. We will present experimental results to show that our quality classified framework can accurately classify images based on the type and severity of image degradations and can significantly boost the performances of state-of-the-art face detector and recognizer in dealing with image datasets containing mixed quality images. Qian Zhang 0018, Miaohui Wang, Guoping Qiu |
ICPR | 3 |
| 2017 | 3D Surface Detail Enhancement from a Single Normal MapabstractIn 3D reconstruction, the obtained surface details are mainly limited to the visual sensor due to sampling and quantization in the digitalization process. How to get a fine-grained 3D surface with low-cost is still a challenging obstacle in terms of experience, equipment and easyto-obtain. This work introduces a novel framework for enhancing surfaces reconstructed from normal map, where the assumptions on hardware (e.g., photometric stereo setup) and reflection model (e.g., Lambertion reflection) are not necessarily needed. We propose to use a new measure, angle profile, to infer the hidden micro-structure from existing surfaces. In addition, the inferred results are further improved in the domain of discrete geometry processing (DGP) which is able to achieve a stable surface structure under a selectable enhancement setting. Extensive simulation results show that the proposed method obtains significantly improvements over uniform sharpening method in terms of both subjective visual assessment and objective quality metric. Wuyuan Xie, Miaohui Wang, Xianbiao Qi, Lei Zhang 0006 |
ICCV | 2 |
| 2016 | Perceptual sensitivity-based rate control method for high efficiency video coding
Huanqiang Zeng, Aisheng Yang, King Ngi Ngan, Miaohui Wang |
Multim. Tools Appl. | 4 |
| 2016 | Low-Delay Rate Control for Consistent Quality Using Distortion-Based Lagrange MultiplierabstractVideo quality fluctuation plays a significant role in human visual perception, and hence, many rate control approaches have been widely developed to maintain consistent quality for video communication. This paper presents a novel rate control framework based on the Lagrange multiplier in high-efficiency video coding. With the assumption of constant quality control, a new relationship between the distortion and the Lagrange multiplier is established. Based on the proposed distortion model and buffer status, we obtain a computationally feasible solution to the problem of minimizing the distortion variation across video frames at the coding tree unit level. Extensive simulation results show that our method outperforms the rate control used in HEVC Test Model (HM) by providing a more accurate rate regulation, lower video quality fluctuation, and stabler buffer fullness. The average peak signal-to-noise ratio (PSNR) and PSNR deviation improvements are about 0.37 dB and 57.14% in the low-delay (P and B) video communication, where the complexity overhead is ∼ 4.44% . Miaohui Wang, King Ngi Ngan, Hongliang Li 0001 |
IEEE Trans. Image Process. | 1 |
| 2015 | Optimal bit allocation in HEVC for real-time video communicationsabstractRecently, the Lagrange multiplier λ based rate control has been developed in the latest HEVC (High Efficiency Video Coding) codec. It is revealed that while such a rate control approach successfully guarantees a designated bit-rate, the optimal bit allocation has not been investigated or established in the λ-domain. To enhance the overall performance, we propose a novel linear relationship between distortion and λ. With this distortion model, we deduce a closed-form solution to minimize the distortion while satisfying a given rate regulation at the coding tree unit level. Extensive simulation results show that our method can achieve an average 0.20 dB improvement in the low-delay communications. Miaohui Wang, King Ngi Ngan |
ICIP | 1 |
| 2015 | Improved block level adaptive quantization for high efficiency video codingabstractAs the concept of block level adaptivity becomes an important feature in recent video CODECs, block level adaptive quantization (BLAQ) is being considered in the High Efficiency Video Coding (HEVC) standard. The BLAQ is based on the assumption that each block should have its own quantization parameter (QP), which can adapt to the local content of video sequences much better, and hence the video encoder with adaptive QP can perform a better perceptual quality. However, in the HEVC reference software, the BLAQ is required to obtain a proper QP for each block by the rate distortion optimization (RDO) scheme and so the computational complexity of the encoder increases significantly. In this paper, an improved BLAQ algorithm is proposed to obtain the adaptive QP for each block. The simulation results show that the proposed method can save more bits as well as require lower computational complexity, compared to the traditional method. Miaohui Wang, King Ngi Ngan, Hongliang Li 0001, Huanqiang Zeng |
ISCAS | 1 |
| 2015 | An Efficient Frame-Content Based Intra Frame Rate Control for High Efficiency Video CodingabstractRate control plays an important role in the rapid development of high-fidelity video services. As the High Efficiency Video Coding (HEVC) standard has been finalized, many rate control algorithms are being developed to promote its commercial use. The HEVC encoder adopts a new R-lambda based rate control model to reduce the bit estimation error. However, the R-lambda model fails to consider the frame-content complexity that ultimately degrades the performance of the bit rate control. In this letter, a gradient based R-lambda (GRL) model is proposed for the intra frame rate control, where the gradient can effectively measure the frame-content complexity and enhance the performance of the traditional R-lambda method. In addition, a new coding tree unit (CTU) level bit allocation method is developed. The simulation results show that the proposed GRL method can reduce the bit estimation error and improve the video quality in HEVC all intra frame coding. Miaohui Wang, King Ngi Ngan, Hongliang Li 0001 |
IEEE Signal Process. Lett. | 1 |
| 2014 | Efficient H.264/AVC Video Coding with Adaptive TransformsabstractTransform has been widely used to remove spatial redundancy of prediction residuals in the modern video coding standards. However, since the residual blocks exhibit diverse characteristics in a video sequence, conventional transform methods with fixed transform kernels may result in low efficiency. To tackle this problem, we propose a novel content adaptive transform framework for the H.264/AVC-based video coding. The proposed method utilizes pixel rearrangement to dynamically adjust the transform kernels to adapt to the video content. In addition, unlike the traditional adaptive transforms, the proposed method obtains the transform kernels from the reconstructed block, and hence it consumes only one logic indicator for each transform unit. Moreover, a spiral-scanning method is developed to reorder the transform coefficients for better entropy coding. Experimental results on the Key Technical Area (KTA) platform show that the proposed method can achieve an average bitrate reduction of about 7.95% and 7.0% under all-intra and low-delay configurations, respectively. Miaohui Wang, King Ngi Ngan, Long Xu 0001 |
IEEE Trans. Multim. | 1 |
| 2013 | A rate distortion optimized transform for motion compensation residualabstractIn this paper, we propose a rate distortion optimization based content adaptive transform method for motion compensation residuals. The proposed method utilizes pixel rearrangement to dynamically adjust the transform kernels to adapt to the residual content. Comparing with the traditional adaptive transforms, the highlight of this work is that it obtains the transform kernels from the decoded block, and hence it consumes only one overhead bit for each transform unit. Moreover, rate distortion optimization scheme is used to choose the best candidate kernels. Experimental results show that the proposed method achieves an average 0.35 dB gain of PSNR in comparison with the key technical areas (KTA) encoder. Miaohui Wang, King Ngi Ngan, Huanqiang Zeng |
PCS | 1 |
| 2013 | Perceptual adaptive Lagrangian multiplier for high efficiency video codingabstractIn high efficiency video coding (HEVC), Lagrangian rate distortion optimization (RDO) technique is used to optimize the rate distortion (RD) performance. However, the corresponding Lagrangian multiplier does not consider the perceptual characteristic of the input video and thus is not effective for perceptual video coding. To address this problem, an efficient perceptual adaptive Lagrangian multiplier for HEVC is proposed. Based on the human visual system (HVS) observation that the region with less perceptual sensitivity can tolerate more distortion, the Lagrangian multiplier is adaptively adjusted for each coding tree unit (CTU) based on its perceptual sensitivity so that the perceptual quality of the reconstructed video can be improved. The above-mentioned perceptual sensitivity for each CTU is measured according to two perceptual features—spatial energy ratio and temporal motion activity. Experimental results have shown that the proposed method is able to significantly improve the perceptual RD performance, compared with the original HEVC. Huanqiang Zeng, King Ngi Ngan, Miaohui Wang |
PCS | 3 |
| 2013 | An efficient framework for image/video inpainting
Miaohui Wang, Bo Yan 0001, King Ngi Ngan |
Signal Process. Image Commun. | 1 |
| 2012 | Spatial-temporal decorrelation for image/video codingabstractModern image/video compression techniques greatly help to store and transmit digital images and video data. Discrete wavelet transform is used in JPEG2000 because of its scalability and tolerable degradation. In H.264/AVC, predictive coding is employed to remove spatial redundancy before discrete cosine transform. In this paper, we propose a new encoder structure combining both advantages from JPEG2000 [1] and H.264/AVC [2], which firstly utilizes the proposed spatial transform to decompose a single image into several sub-images and then employs motion compensation to convert the conventional spatial decorrelation into temporal decorrelation. Experimental results show that our proposed method outperforms the state-of-the-art image coding algorithms and achieves better rate-distortion performance. Miaohui Wang, King Ngi Ngan, Long Xu 0001 |
PCS | 1 |
| 2012 | Video content dependent directional transform for intra frame codingabstractThe mode-dependent directional transform (MDDT) employed Karhunen-Loève Transform (KLT) for compressing directional residue signal of intra prediction along its direction. The transform bases were derived from the singular value decomposition (SVD) of residue signals coming from all kinds of video sequences, which were expected to be efficient for most of video sequences. However, the advantage of KLT comes from the concept of a “signal content dependent transform”. MDDT and its variants failed to exploit such a concept, so they did not fully exploit the efficiency of KLT. In this paper, a video content feature is firstly defined as the histogram of the residue produced by intra prediction. Secondly, one KLT basis is computed for each feature of each mode from off-line experiments. Thus, multiple KLT bases identified by their features are provided to each mode instead of only one basis in MDDT. One of them is selected during encoding process by matching the feature of signal being processed to the predefined features. The experiments show that the average improvement of 0.17dB PSNR and 2.23% bits saving can be achieved by the proposed video content dependent directional transform (CDDT) comparing to the state-of-the-art MDDT. Long Xu 0001, King Ngi Ngan, Miaohui Wang |
PCS | 3 |
| 2010 | Pyramid model based Down-sampling for image inpaintingabstractImage inpainting is a useful and powerful technique for automatically restoring or removing objects in films and damaged pictures. In the last ten years, many excellent inpainting algorithms have been proposed after Bertalmio et al. [1]. However, no paper systematically and theoretically analyzes the factors, which limit the performances of existing algorithms. Based on extensive experiments, we firstly construct a universal framework for image inpainting, which contains three crucial factors—Area, Shape, and Perimeter (ASP). Then we propose a Pyramid model based Down-sampling Inpainting (PDI) model according to the ASP principles. Experimental results show that the performances of existing methods can be tremendously improved after incorporating the PDI model. Miaohui Wang, Bo Yan 0001, Hamid Gharavi |
ICIP | 1 |
| 2009 | Lagrangian Multiplier Based Joint Three-Layer Rate Control for H.264/AVCabstractLagrangian multiplier (LM) based mode decision is one of the most important technologies in standard H.264/AVC encoder. Based on LM theory, this paper presents a joint three-layer (JTL) model for H.264/AVC rate control. At macroblock (MB) level, we dynamically revise LM for each of MBs by its estimated complexities, which is able to select a better coding mode than the current scheme with the constant LM adopted in H.264/AVC. At frame level, a more flexible and effective quantization parameter (QP) adjustment scheme is designed for I-frame to avoid buffer overflow or underflow. In addition, we also present a new target bits allocation scheme in group of picture (GOP) level. Experimental results show that our JTL model can not only significantly improve the video quality with the average PSNR gain up to 0.97 dB, but also provide a more stable buffer occupancy with respect to other existing rate control methods. Miaohui Wang, Bo Yan 0001 |
IEEE Signal Process. Lett. | 1 |
| 2009 | Adaptive Distortion-Based Intra-Rate Estimation for H.264/AVC Rate ControlabstractRate distortion of Intra-frame (I-frame) plays a significant role in controlling video quality for H.264/AVC. This letter presents an adaptive distortion-based Intra-rate estimation (ADIE) algorithm for H.264/AVC rate control. In this algorithm, a new rate control model is established based on the distortion by taking image complexity, buffer status, and scene change into consideration. After adaptively updating this model, our proposed algorithm is capable of providing more accurate estimation for the quantization step (Qstep) of the I-frame than other conventional methods. Experimental results show that the proposed method can significantly improve video quality (up to 1.62 dB average PSNR performance), while achieving an accurate output bit rate. More importantly, it provides a more stable visual quality which is a result of keeping the buffer occupancy steadier than other existing rate control algorithms. Bo Yan 0001, Miaohui Wang |
IEEE Signal Process. Lett. | 2 |