EDBT 2026 Demo / reviewers in the wild / expert
Wuyuan Xie
dblp:66/8204
· DBLP profile ↗
46ranked-venue papers
15as first author
35since 2021 · last 2026
0000-0001-5791-8267ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 31 · 14 first-author · 23 since 2021Artificial intelligence and machine learning · 22 · 11 first-author · 17 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 7 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FreeMem: Enhancing Consistency in Long Video Generation via Tuning-Free MemoryabstractText-to-Video (T2V) generation has advanced greatly, yet maintaining consistency remains challenging, especially for tuning-free long video generation. We attribute the consistency problem to cumulative deviations for long video generation at three levels: the random noise lacking correlation results initial deviation between frames; discrepancy in semantic feature tokens between denoising network blocks gradually accumulates as the frame count grows, leading to greater deviations; attention mechanisms struggle to capture global relationships across distant frames in long videos. To address these, we propose FreeMem, a tuning-free framework leveraging hierarchical memory update and injection: the noise memory stabilizes consistency by manipulating low and high frequency components in the initial noise space; the token memory combats inconsistency through adaptive fusion of historical and current semantic feature tokens between denoising network blocks; and the attention memory establishes persistent cache to model long-range relationships within self attention layers. Evaluated on VBench, FreeMem improves subject and background consistency matrics across various methods, offering a practical solution for low-cost, high-consistency long video generation. Jibin Peng, Di Lin 0002, Zhecheng Xu, Wuyuan Xie, Miaohui Wang, Lingyu Liang, Qing Guo 0005 |
AAAI | 6 |
| 2026 | The Last Byte: Learning Just Enough for Machine-Oriented Image CompressionabstractJust recognizable distortion (JRD) has been introduced for image compression for machines, aiming to quantify the maximum coding distortion that can be tolerated by a specific perception model, thereby defining the upper bound of machine vision redundancy (MVR). However, existing JRD-based redundancy estimation methods face three key challenges: limited dataset annotation accuracy, low prediction efficiency, and insufficient perception accuracy, all of which hinder their practical deployment. To address these limitations, we propose a new MVR-Net, a frame-wise efficient JRD prediction method that generates the optimal encoding quantization map in a single inference pass. Furthermore, we refine the annotation standard for JRD datasets based on experimental insights, enhancing the precision of recognizable redundancy measurement. Compared to stateof-the-art methods, MVR-Net achieves a superior balance between bitrate reduction and perception accuracy in JRD-guided compression, while offering up to a 40,000× speed improvement, demonstrating its practicality and efficiency for real-world applications. Wuyuan Xie, Zhenming Li, Ye Liu 0005, Yun Song, Miaohui Wang |
AAAI | 1 |
| 2026 | Firing Bits Where It Matters: Spiking-Guided Just Recognizable Distortion Modeling for Machine-Centric Video CodingabstractJust recognizable distortion (JRD) has emerged as a promising paradigm for machine-centric video coding. However, existing JRD-guided coding methods are limited by coarse annotation granularity and high computational cost, which hinder their deployment. In this paper, we first investigate the impact of different JRD annotation strategies on downstream task performance. By incorporating both instance-level and contextual information, we construct a new JRD dataset with fine-grained annotations compatible with object detection and instance segmentation tasks. To enhance quantization parameter (QP) map prediction while maintaining computational efficiency, we propose a novel spiking neural network (SNN)-based framework that decomposes video frames into spatial structures, channel interactions, and temporal patterns. Furthermore, we introduce a spiking attention mechanism to aggregate task-relevant features and employ adaptive scaling vectors to suppress machine-perceived redundancy, enabling targeted bitrate allocation aligned with task-critical content. Extensive experiments on multiple datasets and backbones demonstrate that our approach consistently outperforms state-of-the-art codec-based and JRD-guided methods in maintaining task performance at ultra-low bitrates, while significantly reducing computational overhead. Wuyuan Xie, Zhenming Li, Yuwu Lu, Di Lin 0002, Yun Song, Miaohui Wang |
AAAI | 1 |
| 2026 | Efficient 3D Surface Super-Resolution via Normal-Based Multimodal RestorationabstractHigh-fidelity 3D surface is essential for vision tasks across various domains such as medical imaging, cultural heritage preservation, quality inspection, virtual reality, and autonomous navigation. However, the intricate nature of 3D data representations poses significant challenges in restoring diverse 3D surfaces while capturing fine-grained geometric details at a low cost. This paper introduces an efficient multimodal normal-based 3D surface super-resolution (mn3DSSR) framework, designed to address the challenges of microgeometry enhancement and computational overhead. Specifically, we have constructed one of the largest normal-based multimodal dataset, ensuring superior data quality and diversity through meticulous subjective selection. Furthermore, we explore a new two-branch multimodal alignment approach along with a multimodal split fusion module to mitigate computational complexity while improving restoration performances. To address the limitations associated with normal-based multimodal learning, we develop novel normal-induced loss functions that facilitate geometric consistency and improve feature alignment. Extensive experiments conducted on seven benchmark datasets across four different 3D data representations demonstrate that mn3DSSR consistently outperforms state-of-the-art super-resolution methods in terms of restoration accuracy with high computational efficiency. Miaohui Wang, Yunheng Liu, Wuyuan Xie, Boxin Shi, Jianmin Jiang |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | mmFAS: Multimodal Face Anti-Spoofing Using Multi-Level Alignment and Switch-Attention FusionabstractThe increasing number of presentation attacks on reliable face matching has raised concerns and garnered attention towards face anti-spoofing (FAS). However, existing methods for FAS modeling commonly fuse multiple visual modalities (e.g., RGB, Depth, and Infrared) in a straightforward manner, disregarding latent feature gaps that can hinder representation learning. To address this challenge, we propose a novel multimodal FAS framework (mmFAS) that focuses on explicit alignment and fusion of latent features across different modalities. Specifically, we develop a multimodal alignment module to alleviate the latent feature gap by using instance-level contrastive learning and class-level matching simultaneously. Further, we explore a new switch-attention based fusion module to automatically aggregate complementary information and control model complexity. To evaluate the anti-spoofing performance more effectively, we adopt a challenging yet meaningful cross-database protocol involving four benchmark multimodal FAS datasets to simulate realworld scenarios. Extensive experimental results demonstrate the effectiveness of mmFAS in improving the accuracy of FAS systems, outperforming 10 representative methods. Geng Chen 0006, Wuyuan Xie, Di Lin 0002, Ye Liu 0005, Miaohui Wang |
AAAI | 2 |
| 2025 | DDJND: Dual Domain Just Noticeable Difference in Multi-Source Content Images with Structural DiscrepancyabstractMost existing just noticeable difference (JND) methods primarily integrate specific masking effects in a single domain. However, these single-domain JND methods struggle with the structural discrepancies in multi-source content images, limiting their effectiveness in visual redundancy estimation. To address this issue, we propose a dual domain encoder that combines spatial and frequency features to comprehensively capture visual patterns. Our design includes spatial pattern balance and frequency detail correction modules to balance global and local patterns and correct low- and high-frequency distributions. Additionally, we develop a dual domain decoder to effectively extract multi-scale pattern redundancies and integrate them with detail redundancies in the frequency domain. Experiments demonstrate the effectiveness and robustness of our proposed method in handling structural discrepancies in multi-source content images. Miaohui Wang, Zhenming Li, Wuyuan Xie |
AAAI | 3 |
| 2025 | kgMBQA: Quality Knowledge Graph-driven Multimodal Blind Image AssessmentabstractBlind image assessment aims to simulate human prediction of image quality distortion levels and provide quality scores. However, existing unimodal quality indicators have limited representational ability when facing complex contents and distortion types, and the predicted scores also fail to provide explanatory reasons, which further affects the credibility of their prediction results. To address these challenges, we propose a multimodal quality indicator with explanatory text descriptions, called kgMBQA. Specifically, we construct an image quality knowledge graph and conduct in-depth mining to generate explanatory texts. The text modality is further aligned and fused with the image modality, thereby improving the model performance while also outputting its corresponding quality explanatory description. The experimental results demonstrate that our kgMBQA achieves the best performance compared to recent representative methods on the KonIQ-10k, LIVE Challenge, BIQ2021, TID2013, and AIGC-3K datasets. Wuyuan Xie, Tingcheng Bian, Miaohui Wang |
IJCAI | 1 |
| 2025 | MMVQA: Dual-Path Multimodal Fusion for AI-Generated Video Quality Assessment
Wuyuan Xie, Miaohui Wang |
PRCV (11) | 2 |
| 2025 | suLPCC: A Novel LiDAR Point Cloud Compression Framework for Scene Understanding TasksabstractLight detection and ranging (LiDAR) point cloud compression (LPCC) plays an important role in managing the storage, transmission, and perception of the rapidly expanding volume of LiDAR point cloud (LPC) data. However, there has been a noticeable lack of comprehensive investigation into LPCC methods specifically designed for environmental perception and understanding. To address this gap, we propose a new LPCC framework aimed at meeting the unique requirements of various scene understanding tasks, enhancing the adaptability of LPCCs in real-world scenarios. Specifically, we divide the input LPCs into an object and a scene component through a distinction module, design a new point completion-based method to encode object LPCs, and develop novel structure-aware intracoding and motion-optimized intercoding schemes to compress scene LPCs. Experimental results on three benchmark datasets demonstrate the effectiveness of our proposed method on the localization, mapping, and detection tasks. We believe that the findings presented in this article will contribute to a deeper understanding of LPCCs as well as promote further development of LiDAR sensor-based systems. Miaohui Wang, Runnan Huang, Ye Liu 0005, Yanshan Li, Wuyuan Xie |
IEEE Trans. Ind. Informatics | 5 |
| 2025 | CPSR-CLIP: Conditional Prompt-Induced Style Reconstruction for Zero-Shot Domain Adaptation
Jiayu Qian, Yuwu Lu, Wuyuan Xie, Zhihui Lai 0001, Miaohui Wang, Xuelong Li 0001 |
IEEE Trans. Multim. | 3 |
| 2025 | Compression Approaches for LiDAR Point Clouds and Beyond: A SurveyabstractWith the widespread use of LiDAR sensors in autonomous driving, LiDAR point cloud compression (LPCC) plays an important role in effectively managing the storage, transmission, and perception of the growing volume of LiDAR data. Despite this need, there has been a noticeable absence of comprehensive investigations specifically dedicated to LPCC methods. To address this issue, this article presents a systematic survey of existing LPCCs, aiming to summarize recent progress and inspire future research in this field. We begin by providing a general introduction of LPCC fundamentals, covering the latest LiDAR point cloud (LPC) datasets, distinctive attributes, evaluation metrics, and data formats. We then conduct a careful review and comparison of LPCCs, examining image-based, octree-based, deep-learned, and other approaches, offering valuable insights into the strengths and weaknesses of cutting-edge models. Finally, we propose future research directions based on the limitations of recent LPCCs. We believe that the findings presented in this article will contribute to a deeper understanding of LPCCs and promote further development of LiDAR sensor-based systems. Miaohui Wang, Runnan Huang, Wuyuan Xie, Zhan Ma 0001, Siwei Ma 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2024 | msLPCC: A Multimodal-Driven Scalable Framework for Deep LiDAR Point Cloud CompressionabstractLiDAR sensors are widely used in autonomous driving, and the growing storage and transmission demands have made LiDAR point cloud compression (LPCC) a hot research topic. To address the challenges posed by the large-scale and uneven-distribution (spatial and categorical) of LiDAR point data, this paper presents a new multimodal-driven scalable LPCC framework. For the large-scale challenge, we decouple the original LiDAR data into multi-layer point subsets, compress and transmit each layer separately, so as to ensure the reconstruction quality requirement under different scenarios. For the uneven-distribution challenge, we extract, align, and fuse heterologous feature representations, including point modality with position information, depth modality with spatial distance information, and segmentation modality with category information. Extensive experimental results on the benchmark SemanticKITTI database validate that our method outperforms 14 recent representative LPCC methods. Miaohui Wang, Runnan Huang, Hengjin Dong, Di Lin 0002, Yun Song, Wuyuan Xie |
AAAI | 6 |
| 2024 | MetaJND: A Meta-Learning Approach for Just Noticeable Difference Estimation
Miaohui Wang, Yukuan Zhu, Wuyuan Xie |
IJCAI | 4 |
| 2024 | Voxel Proposal Network via Multi-Frame Knowledge Distillation for Semantic Scene CompletionabstractSemantic scene completion is a difficult task that involves completing the geometry and semantics of a scene from point clouds in a large-scale environment. Many current methods use 3D/2D convolutions or attention mechanisms, but these have limitations in directly constructing geometry and accurately propagating features from related voxels, the completion likely fails while propagating features in a single pass without considering multiple potential pathways. And they are generally only suitable for static scenes and struggle to handle dynamic aspects. This paper introduces Voxel Proposal Network (VPNet) that completes scenes from 3D and Bird's-Eye-View (BEV) perspectives. It includes Confident Voxel Proposal based on voxel-wise coordinates to propose confident voxels with high reliability for completion. This method reconstructs the scene geometry and implicitly models the uncertainty of voxel-wise semantic labels by presenting multiple possibilities for voxels. VPNet employs Multi-Frame Knowledge Distillation based on the point clouds of multiple adjacent frames to accurately predict the voxel-wise labels by condensing various possibilities of voxel relationships. VPNet has shown superior performance and achieved state-of-the-art results on the SemanticKITTI and SemanticPOSS datasets. Lubo Wang, Di Lin 0002, Kairui Yang, Qing Guo 0005, Wuyuan Xie, Miaohui Wang, Lingyu Liang, Ping Li 0016 |
NeurIPS | 6 |
| 2024 | Deep Learning Methods for Calibrated Photometric Stereo and BeyondabstractPhotometric stereo recovers the surface normals of an object from multiple images with varying shading cues, i.e., modeling the relationship between surface orientation and intensity at each pixel. Photometric stereo prevails in superior per-pixel resolution and fine reconstruction details. However, it is a complicated problem because of the non-linear relationship caused by non-Lambertian surface reflectance. Recently, various deep learning methods have shown a powerful ability in the context of photometric stereo against non-Lambertian surfaces. This paper provides a comprehensive review of existing deep learning-based calibrated photometric stereo methods utilizing orthographic cameras and directional light sources. We first analyze these methods from different perspectives, including input processing, supervision, and network architecture. We summarize the performance of deep learning photometric stereo models on the most widely-used benchmark data set. This demonstrates the advanced performance of deep learning-based photometric stereo methods. Finally, we give suggestions and propose future research trends based on the limitations of existing models. Yakun Ju, Kin-Man Lam 0001, Wuyuan Xie, Huiyu Zhou 0001, Junyu Dong, Boxin Shi |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | LLM-Guided Cross-Modal Point Cloud Quality Assessment: A Graph Learning ApproachabstractThis paper addresses the critical need for accurate and reliable point cloud quality assessment (PCQA) in various applications, such as autonomous driving, robotics, virtual reality, and 3D reconstruction. To meet this need, we propose a large language model (LLM)-guided PCQA approach based on graph learning. Specifically, we first utilize the LLM to generate quality description texts for each 3D object, and employ two CLIP-like feature encoders to represent the image and text modalities. Next, we design a latent feature enhancer module to improve contrastive learning, enabling more effective alignment performance. Finally, we develop a graph network fusion module that utilizes a ranking-based loss to adjust the relationship of different nodes, which explicitly considers both modality fusion and quality ranking. Experimental results on three benchmark datasets demonstrate the effectiveness and superiority of our approach over 12 representative PCQA methods, which demonstrate the potential of multi-modal learning, the importance of latent feature enhancement, and the significance of graph-based fusion in advancing the field of PCQA. Wuyuan Xie, Yunheng Liu, Kaiming Wang, Miaohui Wang |
IEEE Signal Process. Lett. | 1 |
| 2024 | ReferPose: Distance Optimization-Based Reference Learning for Human Pose Estimation and MonitoringabstractExisting deep learning models for human pose estimation (HPE) have shown satisfactory performance in monitoring human actions. However, they usually face a dilemma between complexity and accuracy. To address this challenge, we propose an effective reference learning method for HPE (namely ReferPose), which is based on a new distance optimization strategy. Specifically, we utilize a reference model for pose learning and representation. The pose representation learned from the entire database is merged into the reference model, providing continuous reference learning guidance for an in-training model. In addition, we design a new cosine annealing-based reference guidance for temporal denoising and further develop a distance optimization strategy to provide joint guidance from pose knowledge, model representation, and temporal experience. Experimental results on two benchmark databases and a human fall monitoring system demonstrate that our ReferPose not only achieves promising accuracy improvement compared with several representative HPE models, but also offers low cost and high efficiency. Miaohui Wang, Zhuowei Xu, Wuyuan Xie |
IEEE Trans. Ind. Informatics | 3 |
| 2024 | BinaryFormer: A Hierarchical-Adaptive Binary Vision Transformer (ViT) for Efficient ComputingabstractVision Transformer (ViT) has recently demonstrated impressive nonlinear modeling capabilities and achieved state-of-the-art performance in various industrial applications, such as object recognition, anomaly detection, and robot control. However, their practical deployment can be hindered by high storage requirements and computational intensity. To alleviate these challenges, we propose a binary transformer called BinaryFormer, which quantizes the learned weights of the ViT module from 32-b precision to 1 b. Furthermore, we propose a hierarchical-adaptive architecture that replaces expensive matrix operations with more affordable addition and bit operations by switching between two attention modes. As a result, BinaryFormer is able to effectively compress the model size as well as reduce the computation cost of ViT. Experimental results on the ImageNet-1K benchmark datasets show that BinaryFormer reduces the size of a typical ViT model by an average of 27.7× and converts over 99% of multiplication operations into bit operations while maintaining reasonable accuracy. Miaohui Wang, Zhuowei Xu, Wuyuan Xie |
IEEE Trans. Ind. Informatics | 4 |
| 2023 | Just Noticeable Visual Redundancy Forecasting: A Deep Multimodal-Driven ApproachabstractJust noticeable difference (JND) refers to the maximum visual change that human eyes cannot perceive, and it has a wide range of applications in multimedia systems. However, most existing JND approaches only focus on a single modality, and rarely consider the complementary effects of multimodal information. In this article, we investigate the JND modeling from an end-to-end homologous multimodal perspective, namely hmJND-Net. Specifically, we explore three important visually sensitive modalities, including saliency, depth, and segmentation. To better utilize homologous multimodal information, we establish an effective fusion method via summation enhancement and subtractive offset, and align homologous multimodal features based on a self-attention driven encoder-decoder paradigm. Extensive experimental results on eight different benchmark datasets validate the superiority of our hmJND-Net over eight representative methods. Wuyuan Xie, Shukang Wang, Sukun Tian, Lirong Huang, Ye Liu 0005, Miaohui Wang |
AAAI | 1 |
| 2023 | CVSformer: Cross-View Synthesis Transformer for Semantic Scene CompletionabstractSemantic scene completion (SSC) requires an accurate understanding of the geometric and semantic relationships between the objects in the 3D scene for reasoning the occluded objects. The popular SSC methods voxelize the 3D objects, allowing the deep 3D convolutional network (3D CNN) to learn the object relationships from the complex scenes. However, the current networks lack the controllable kernels to model the object relationship across multiple views, where appropriate views provide the relevant information for suggesting the existence of the occluded objects. In this paper, we propose Cross-View Synthesis Transformer (CVSformer), which consists of Multi-View Feature Synthesis and Cross-View Transformer for learning cross-view object relationships. In the multi-view feature synthesis, we use a set of 3D convolutional kernels rotated differently to compute the multi-view features for each voxel. In the cross-view transformer, we employ the cross-view fusion to comprehensively learn the cross-view relationships, which form useful information for enhancing the features of individual views. We use the enhanced features to predict the geometric occupancies and semantic labels of all voxels. We evaluate CVSformer on public datasets, where CVS-former yields state-of-the-art results. Our code is available at https://github.com/donghaotian123/CVSformer. Haotian Dong, Enhui Ma, Lubo Wang, Miaohui Wang, Wuyuan Xie, Qing Guo 0005, Ping Li 0016, Lingyu Liang, Kairui Yang, Di Lin 0002 |
ICCV | 5 |
| 2023 | 3D Surface Super-resolution from Enhanced 2D Normal Images: A Multimodal-driven Variational AutoEncoder Approachabstract3D surface super-resolution is an important technical tool in virtual reality, and it is also a research hotspot in computer vision. Due to the unstructured and irregular nature of 3D object data, it is usually difficult to obtain high-quality surface details and geometry textures via a low-cost hardware setup. In this paper, we establish a multimodal-driven variational autoencoder (mmVAE) framework to perform 3D surface enhancement based on 2D normal images. To fully leverage the multimodal learning, we investigate a multimodal Gaussian mixture model (mmGMM) to align and fuse the latent feature representations from different modalities, and further propose a cross-scale encoder-decoder structure to reconstruct high-resolution normal images. Experimental results on several benchmark datasets demonstrate that our method delivers promising surface geometry structures and details in comparison with competitive advances. Wuyuan Xie, Tengcong Huang, Miaohui Wang |
IJCAI | 1 |
| 2023 | A Method of Micro-Geometric Details Preserving in Surface Reconstruction from GradientabstractSurface from gradient (SfG) is one of the fundamental methods to densely reconstruct 3D object surface in computer vision. However, the reconstruction of micro-geometric details has not been satisfactorily solved in existing SfG methods due to their non-integrability. In this paper, we present an effective discrete geometric approach to reconstruct fine-grained sharp surface feature with non-integrability. Specifically, We investigate the fine-grained structure of surfaces in the micro geometry domain. based on an adaptive projection on vertexes constrained by neighboring gradient vectors, and develop a gradient angle-guided energy optimization to generate a fine-grained surface. Experimental results on various challenging synthetic and real-world data show that the proposed method is able to effectively reconstruct challenging micro-geometric details for general SfG methods. Wuyuan Xie, Miaohui Wang |
ACM Multimedia | 1 |
| 2023 | pmBQA: Projection-based Blind Point Cloud Quality Assessment via Multimodal LearningabstractWith the increasing communication and storage of point cloud data, there is an urgent need for an effective objective method to measure the quality before and after processing. To address this difficulty, we propose a projection-based blind quality indicator via multimodal learning for point cloud data, which can perceive both geometric distortion and texture distortion by using four homogeneous modalities (i.e., texture, normal, depth and roughness). To fully exploit the multimodal information, we further develop a deformable convolutionbased alignment module and a graph-based feature fusion module, and investigate a graph node attention-based evaluation method to forecast the quality score. Extensive experimental results on three benchmark databases show that our method achieves more accurate evaluation performance in comparison with 12 competitive methods. Wuyuan Xie, Kaimin Wang, Yakun Ju, Miaohui Wang |
ACM Multimedia | 1 |
| 2023 | Visual Redundancy Removal of Composite Images via Multimodal LearningabstractComposite images are generated by combining two or more different photographs, and their content is typically heterogeneous. However, existing unimodal visual redundancy prediction methods are difficult to accurately model the complex characteristics of this image type. In this paper, we investigate the visual redundancy modeling of composite images from an end-to-end multimodal perspective, including four cross-media modalities (i.e., text, brightness, color, and segmentation). Specifically, we design a two-stage cross-modal alignment module based on self-attention mechanism and contrastive learning, and develop a fusion module based on a cross-modal augmentation paradigm. Further, we establish the first cross-media visual redundancy dataset for composite images, which contains 413 groups of cross-modal data and generates 13629 realistic compression distortions using the latest versatile video coding (VVC) standard. Experimental results on nine benchmark datasets demonstrate the effectiveness of our method, outperforming seven representative methods. Wuyuan Xie, Shukang Wang, Miaohui Wang |
ACM Multimedia | 1 |
| 2023 | Surface Geometry Processing: An Efficient Normal-Based Detail RepresentationabstractWith the rapid development of high-resolution 3D vision applications, the traditional way of manipulating surface detail requires considerable memory and computing time. To address these problems, we introduce an efficient surface detail processing framework in 2D normal domain, which extracts new normal feature representations as the carrier of micro geometry structures that are illustrated both theoretically and empirically in this article. Compared with the existing state of the arts, we verify and demonstrate that the proposed normal-based representation has three important properties, including detail separability, detail transferability and detail idempotence. Finally, three new schemes are further designed for geometric surface detail processing applications, including geometric texture synthesis, geometry detail transfer, and 3D surface super-resolution. Theoretical analysis and experimental results on the latest benchmark dataset verify the effectiveness and versatility of our normal-based representation, which accepts 30 times of the input surface vertices but at the same time only takes 6.5% memory cost and 14.0% running time in comparison with existing competing algorithms. Wuyuan Xie, Miaohui Wang, Di Lin 0002, Boxin Shi, Jianmin Jiang |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | A Task-Driven Scene-Aware LiDAR Point Cloud Coding Framework for Autonomous VehiclesabstractLiDAR sensors are almost indispensable for autonomous robots to perceive the surrounding environment. However, the transmission of large-scale LiDAR point clouds is highly bandwidth-intensive, which can easily lead to transmission problems, especially for unstable communication networks. Meanwhile, existing LiDAR data compression is mainly based on rate-distortion optimization, which ignores the semantic information of ordered point clouds and the task requirements of autonomous robots. To address these challenges, this article presents a task-driven Scene-Aware LiDAR Point Clouds Coding (SA-LPCC) framework for autonomous vehicles. Specifically, a semantic segmentation model is developed based on multidimension information, in which both 2-D texture and 3-D topology information are fully utilized to segment movable objects. Furthermore, a prediction-based deep network is explored to remove the spatial–temporal redundancy. The experimental results on the benchmark semantic KITTI dataset validate that our SA-LPCC achieves state-of-the-art performance in terms of the reconstruction quality and storage space for downstream tasks. We believe that SA-LPCC jointly considers the scene-aware characteristics of movable objects and removes the spatial–temporal redundancy from an end-to-end learning mechanism, which will boost the related applications from algorithm optimization to industrial products. Xuebin Sun, Miaohui Wang, Jingxin Du, Yuxiang Sun 0002, Shing Shin Cheng, Wuyuan Xie |
IEEE Trans. Ind. Informatics | 6 |
| 2023 | Low-Light Images In-the-Wild: A Novel Visibility Perception-Guided Blind Quality IndicatorabstractOwing to the increasing deployment of CMOS camera modules, it is inevitable to take photographs under weak illumination. Therefore, low-light imaging quality is one of the most important factors affecting user experience as well as the product values of consumer electronics, automobile, surveillance, factory automation, and other industrial applications. Inspired by human vision, this article jointly considersvisibility perception,luminosity cognition, andcolor sensationand presents a new visibility perception-guided blind quality indicator for low-light images in-the-wild. To excavate effective descriptors for authentic distortions under weak illumination, we utilize maximum ignorable visible difference to characterize the reduced visibility, and employ the luminance statistical properties and color sensation characteristics to represent brightness and colorfulness distortions. Extensive experimental results on the benchmark dataset verify that the proposed blind quality indicator outperforms nine representative methods including general-purpose and distortion-specific methods. Miaohui Wang, Jian Xiong 0005, Wuyuan Xie |
IEEE Trans. Ind. Informatics | 4 |
| 2022 | MNSRNet: Multimodal Transformer Network for 3D Surface Super-ResolutionabstractWith the rapid development of display technology, it has become an urgent need to obtain realistic 3D surfaces with as high-quality as possible. Due to the unstructured and irregular nature of 3D object data, it is usually difficult to obtain high-quality surface details and geometry textures at a low cost. In this article, we propose an effective multimodal-driven deep neural network to perform 3D surface super-resolution in 2D normal domain, which is simple, accurate, and robust to the above difficulty. To leverage the multimodal information from different perspectives, we jointly consider the texture, depth, and normal modalities to simultaneously restore fine-grained surface details as well as preserve geometry structures. To better utilize the cross-modality information, we explore a two-bridge normal method with a transformer structure for feature alignment, and investigate an affine transform module for fusing multimodal features. Extensive experimental results on public and our newly constructed photometric stereo dataset demonstrate that the proposed method delivers promising surface geometry details compared with nine competitive schemes. Wuyuan Xie, Tengcong Huang, Miaohui Wang |
CVPR | 1 |
| 2022 | S-CCR: Super-Complete Comparative Representation for Low-Light Image Quality Inference In-the-wildabstractWith the rapid development of weak-illumination imaging technology, low-light images have brought new challenges to quality of experience and service. However, developing a robust quality indicator for authentic low-light distortions in-the-wild remains a major challenge in practical quality control systems. In this paper, we develop a new super-complete comparative representation (S-CCR) for the region-level quality inference of low-light images. Specifically, we excavate the color, luminance, and detail quality evidence for the feature embedding guidance of comparative representation based on the human visual characteristics. Moreover, we decompose the inputs into a super-complete feature group so that the image quality of each region can be fully represented, which allows to preserve the distinctiveness, distinguishability, and consistency. Finally, we further establish a comparative domain alignment method, so that the comparative representation of an unseen image can be aligned with respect to the quality features of already-seen ones. Extensive experiments on the benchmark dataset validate the superiority of our S-CCR over 11 competing methods on authentic distortions. Miaohui Wang, Zhuowei Xu, Yuanhao Gong, Wuyuan Xie |
ACM Multimedia | 4 |
| 2022 | Generative Status Estimation and Information Decoupling for Image Rain RemovalabstractImage rain removal requires the accurate separation between the pixels of the rain streaks and object textures. But the confusing appearances of rains and objects lead to the misunderstanding of pixels, thus remaining the rain streaks or missing the object details in the result. In this paper, we propose SEIDNet equipped with the generative Status Estimation and Information Decoupling for rain removal. In the status estimation, we embed the pixel-wise statuses into the status space, where each status indicates a pixel of the rain or object. The status space allows sampling multiple statuses for a pixel, thus capturing the confusing rain or object. In the information decoupling, we respect the pixel-wise statuses, decoupling the appearance information of rain and object from the pixel. Based on the decoupled information, we construct the kernel space, where multiple kernels are sampled for the pixel to remove the rain and recover the object appearance. We evaluate SEIDNet on the public datasets, achieving state-of-the-art performances of image rain removal. The experimental results also demonstrate the generalization of SEIDNet, which can be easily extended to achieve state-of-the-art performances on other image restoration tasks (e.g., snow, haze, and shadow removal). Di Lin 0002, Xin Wang 0118, Miaohui Wang, Wuyuan Xie, Qing Guo 0005, Ping Li 0016 |
NeurIPS | 7 |
| 2022 | Perceptually Quasi-Lossless Compression of Screen Content Data Via Visibility Modeling and Deep ForecastingabstractScreen content data, such as computer-generated photographs, desktop sharing, remote education, video game streaming and screenshot, is one of the most popular visual information carriers in Internet of Video Things. Although lossless compression can guarantee high quality of service for these screen content based industrial applications, it also causes considerable storage space and transmission bandwidth issues. To alleviate these challenges, in this article, we present a visually quasi-lossless coding approach to control the compression distortion belowvisibility thresholdin the human visual system. Specifically, to better quantify the visual redundancy for screen content data, a newvisibility thresholdmethod is designed by incorporating blur sensitivity and oblique correction effects. Then, an end-to-end mapping between thevisibility thresholdand quality control factor is learned and represented as a deep convolutional neural network. The experimental results demonstrate that the proposed method saves the average encoding bits up to 23.15% compared with the latest scheme under the same perceptual quality. Miaohui Wang, Zhuowei Xu, Jian Xiong 0005, Wuyuan Xie |
IEEE Trans. Ind. Informatics | 5 |
| 2021 | Residual Geometric Feature Transform Network for 3D Surface Super-Resolution
Maolin Cui, Wuyuan Xie, Miaohui Wang, Tengcong Huang |
3DV | 2 |
| 2021 | Perceptual Redundancy Estimation of Screen Images via Multi-Domain SensitivitiesabstractVisual redundancy detection is essential for image and video communication. Human visual system (HVS) is difficult to perceive the pixel magnitude change below a certain visibility threshold which is also known as just-noticeable-difference (JND). In this letter, we present an efficient JND estimation approach for screen content images by considering high-frequency sensitivity and orientation sensitivity correction. Specifically, to better quantify the visual redundancy, we investigate the visibility threshold based on the high-frequency distortion sensitivity. To obtain the orientation sensitivity correction, we divide the screen image pixels into three levels based on the oblique effect that considers the sensitive integrity of edges. Compared with several state-of-the-art JNDs, experimental results show that our method tolerates more perceptual redundancy, and delivers better visual quality under the same injected-noise energy. The implementation of the proposed method is publicly available at https://sites.google.com/site/wangmiaohui/. Miaohui Wang, Wuyuan Xie, Long Xu 0001 |
IEEE Signal Process. Lett. | 3 |
| 2021 | Detail-Preserving Multi-Exposure Fusion With Edge-Preserving Structural Patch DecompositionabstractThe multi-exposure fusion (MEF) methods have received much attention in recent years due to the importance of constructing high dynamic range images. Among most of the existing studies, multi-scale structural-patch-decomposition-based MEF (MSPD-MEF) has achieved state-of-the-art fusion quality and the fastest running time. However, this method still suffers from detail loss in the fused images. To tackle this issue, we first incorporate the edge-preserving factors into this method to preserve the details in the fused images in a single-scale setting. Then, we develop the novel and flexible bell curve function, which can further preserve the details in both bright and dark regions. After that, we also show that our method can seamlessly plug in to this multi-scale framework. Extensive experimental results indicate that the proposed method can produce pleasing fusion results with little artifacts and low computational cost in both static and dynamic scenes. Hui Li 0029, Tsz Nam Chan, Xianbiao Qi, Wuyuan Xie |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2021 | Efficient Computer-Aided Design of Dental Inlay Restoration: A Deep Adversarial FrameworkabstractRestoring the normal masticatory function of broken teeth is a challenging task primarily due to the defect location and size of a patient's teeth. In recent years, although some representative image-to-image transformation methods (e.g. Pix2Pix) can be potentially applicable to restore the missing crown surface, most of them fail to generate dental inlay surface with realistic crown details (e.g. occlusal groove) that are critical to the restoration of defective teeth with varying shapes. In this article, we design a computer-aided Deep Adversarial-driven dental Inlay reStoration (DAIS) framework to automatically reconstruct a realistic surface for a defective tooth. Specifically, DAIS consists of a Wasserstein generative adversarial network (WGAN) with a specially designed loss measurement, and a new local-global discriminator mechanism. The local discriminator focuses on missing regions to ensure the local consistency of a generated occlusal surface, while the global discriminator aims at defective teeth and adjacent teeth to assess if it is coherent as a whole. Experimental results demonstrate that DAIS is highly efficient to deal with a large area of missing teeth in arbitrary shapes and generate realistic occlusal surface completion. Moreover, the designed watertight inlay prostheses have enough anatomical morphology, thus providing higher clinical applicability compared with more state-of-the-art methods. Sukun Tian, Miaohui Wang, Fulai Yuan, Yuchun Sun, Wuyuan Xie, Harry Qin |
IEEE Trans. Medical Imaging | 6 |
| 2020 | Surface Reconstruction with Unconnected Normal Maps: An Efficient Mesh-based ApproachabstractNormal integration is a key step in dense 3D reconstruction methods such as shape-from-shading and photometric stereo. However, normal integration cannot be guaranteed between spatially unconnected normal maps, which can ultimately cause a shape deformation in surface-from-normals (SfN). For the first time, this paper presents an efficient approach to address the fundamental problem of surface reconstruction from unconnected normal maps (denoted as "SfN+") using discrete geometry. We first design a normal piece pairing metric to measure the virtually pairing quality between two unconnected normal fragments, which is used as a new constraint for the boundary vertexes during mesh deformation. We then adopt a normal connecting significance indicator to adjust the influence of virtually connected vertexes, which further improves the overall shape deformation. Finally, we model the shape reconstruction of unconnected normal maps as a light-weight energy optimization framework by jointly considering the relaxation of connecting constraints and overall reconstruction error. Experiments show that the proposed SfN+ achieves a robust and efficient performance on dense 3D surface reconstruction. Miaohui Wang, Wuyuan Xie, Maolin Cui |
ACM Multimedia | 2 |
| 2020 | Fast infrared and visible image fusion with structural decomposition
Hui Li 0029, Xianbiao Qi, Wuyuan Xie |
Knowl. Based Syst. | 3 |
| 2020 | Industrial Applications of Ultrahigh Definition Video Coding With an Optimized Supersample Adaptive Offset FrameworkabstractThis article presents an efficient superblock-based sample adaptive offset (superSAO) that jointly exploits block-wise filter partition and probabilistic band interval segmentation to improve quality-of-experience of industrial video applications. Specifically, we investigate the partition flexibility of a superSAO block whose root size is up to 256 × 256, and propose to optimize the block-wise SAO filter by considering computation complexity and compression efficiency. Furthermore, we segment the band interval by equal probability of the sample intensity distributions, which facilitates the computation of better band offsets to attenuate ringing artifact due to quantization errors or encoded motion vectors. Experimental results show that the proposed superSAO method outperforms state-of-the-art approaches by obtaining 4.6% bandwidth reduction on average for the low delay and high-compression video applications. Miaohui Wang, Wuyuan Xie, Jia Zhang 0002, Harry Qin |
IEEE Trans. Ind. Informatics | 2 |
| 2020 | Rate Constrained Multiple-QP Optimization for HEVCabstractIn High Efficiency Video Coding (HEVC), multiple-QP (quantization parameter) optimization can adapt to a local video content. However, the multiple-QP implementation in the HEVC reference software (HM 16.6) achieves the best QP value for each coding block with a large amount of computational complexity. To address this challenge, we propose a fast rate-constrained multiple-QP optimization approach for the HM platform. We first introduce a template-based transform coefficient selection method which can save the overall complexity of entropy coding. In addition, we model the multiple-QP determination as a new rate-constrained optimization problem, and finally, we get a feasible solution with a lower computation overhead. Experimental results show that our method dramatically reduces the average complexity under the all-intra, low-delay and random-access configuration. Miaohui Wang, Jian Xiong 0005, Long Xu 0001, Wuyuan Xie, King Ngi Ngan, Harry Qin |
IEEE Trans. Multim. | 4 |
| 2019 | Surface Reconstruction From Normals: A Robust DGP-Based Discontinuity Preservation ApproachabstractIn 3D surface reconstruction from normals, discontinuity preservation is an important but challenging task. However, existing studies fail to address the discontinuous normal maps by enforcing the surface integrability in the continuous domain. This paper introduces a robust approach to preserve the surface discontinuity in the discrete geometry way. Firstly, we design two representative normal incompatibility features and propose an efficient discontinuity detection scheme to determine the splitting pattern for a discrete mesh. Secondly, we model the discontinuity preservation problem as a light-weight energy optimization framework by jointly considering the discontinuity detection and the overall reconstruction error. Lastly, we further shrink the feasible solution space to reduce the complexity based on the prior knowledge. Experiments show that the proposed method achieves the best performance on an extensive 3D dataset compared with the state-of-the-arts in terms of mean angular error and computational complexity. Wuyuan Xie, Miaohui Wang, Mingqiang Wei, Jianmin Jiang, Harry Qin |
CVPR | 1 |
| 2019 | Adaptive Multi-Path Aggregation for Human DensePose Estimation in the WildabstractDense human pose "in the wild'' task aims to map all 2D pixels of the detected human body to a 3D surface by establishing surface correspondences, i.e., surface patch index and part-specific UV coordinates. It remains challenging especially under the condition of "in the wild'', where RGB images capture complex, real-world scenes with background, occlusions, scale variations, and postural diversity. In this paper, we propose an end-to-end deep Adaptive Multi-path Aggregation network (AMA-net) for Dense Human Pose Estimation. In the proposed framework, we address two main problems: 1) how to design a simple yet effective pipeline for supporting distinct sub-tasks (e.g., instance segmentation, body part segmentation, and UV estimation); and 2) how to equip this pipeline with the ability of handling "in the wild''. To solve these problems, we first extend FPN by adding a branch for mapping 2D pixels to a 3D surface in parallel with the existing branch for bounding box detection. Then, in AMA-net, we extract variable-sized object-level feature maps (e.g., 7×7, 14×14, and 28×28), named multi-path, from multi-layer feature maps, which capture rich information of objects and are then adaptively utilized in different tasks. AMA-net is simple to train and adds only a small overhead to FPN. We discover that aside from the deep feature map, Adaptive Multi-path Aggregation is of particular importance for improving the accuracy of dense human pose estimation "in the wild''. The experimental results on the challenging Dense-COCO dataset demonstrate that our approach sets a new record for Dense Human Pose Estimation task, and it significantly outperforms the state-of-the-art methods. Our code: \urlhttps://github.com/nobody-g/AMA-net. Yuyu Guo 0001, Lianli Gao, Jingkuan Song, Peng Wang 0023, Wuyuan Xie, Heng Tao Shen |
ACM Multimedia | 5 |
| 2019 | UHD Video Coding: A Light-Weight Learning-Based Fast Super-Block ApproachabstractThe ultra high-definition (UHD) video format, which has recently become popular, aims to provide high spatial resolution, high temporal frame rate, high sample bit-depth, and wide pixel color gamut. Despite the continued development of global network capacities, it inevitably causes the increased bandwidth cost of catering to the requirement of delivering UHD video services. To address such challenges, this paper presents an improved super coding unit (SCU) method for UHD video coding in High Efficiency Video Coding (HEVC). Initially, the medium coding unit (MCU) is proposed to avoid unnecessary brute-force coding unit (CU) partitions of SCU. Furthermore, the SCU is proposed to be encoded by Direct-MCU and SCU-to-MCU modes: the Direct-MCU mode is intended to better adapt to the texture-rich region, which guarantees the compression efficiency by avoiding extra-size CU partition; the SCU-to-MCU mode is designed for the homogeneous region of UHD content, which saves the encoding time by skipping fine-grained CU partition search. Moreover, a learning-based fast SCU decision approach is proposed to speed up the determination process of Direct-MCU and SCU-to-MCU, where three representative handcrafted features are extracted. Experimental results show that our method achieves an affordable complexity and excellent coding efficiency (up to 7.30% Bjøntegaard Delta rate savings) in UHD video coding compared to recent HEVC reference software. Miaohui Wang, Wuyuan Xie, Xiandong Meng, Huanqiang Zeng, King Ngi Ngan |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2017 | 3D Surface Detail Enhancement from a Single Normal MapabstractIn 3D reconstruction, the obtained surface details are mainly limited to the visual sensor due to sampling and quantization in the digitalization process. How to get a fine-grained 3D surface with low-cost is still a challenging obstacle in terms of experience, equipment and easyto-obtain. This work introduces a novel framework for enhancing surfaces reconstructed from normal map, where the assumptions on hardware (e.g., photometric stereo setup) and reflection model (e.g., Lambertion reflection) are not necessarily needed. We propose to use a new measure, angle profile, to infer the hidden micro-structure from existing surfaces. In addition, the inferred results are further improved in the domain of discrete geometry processing (DGP) which is able to achieve a stable surface structure under a selectable enhancement setting. Extensive simulation results show that the proposed method obtains significantly improvements over uniform sharpening method in terms of both subjective visual assessment and objective quality metric. Wuyuan Xie, Miaohui Wang, Xianbiao Qi, Lei Zhang 0006 |
ICCV | 1 |
| 2015 | Photometric stereo with near point lighting: A solution by mesh deformationabstractWe tackle the problem of photometric stereo under near point lighting in this paper. Different from the conventional formulation of photometric stereo that assumes parallel lighting, photometric stereo under the near point lighting condition is a nonlinear problem as the local surface normals are coupled with its distance to the camera as well as the light sources. To solve this non-linear problem of PS with near point lighting, a local/global mesh deformation approach is developed in our work to determine the position and the orientation of a facet simultaneously, where each facet is corresponding to a pixel in the image captured by the camera. Unlike nonlinear optimization schemes, the mesh deformation in our approach is decoupled into an iteration of interlaced steps of local projection and global blending. Experimental results verify that our method can generate accurate estimation of surface shape under near point lighting in a few iterations. Besides, this approach is robust to errors on the positions of light sources and is easy to be implemented. Wuyuan Xie, Chengkai Dai, Charlie C. L. Wang |
CVPR | 1 |
| 2014 | Surface-from-Gradients: An Approach Based on Discrete Geometry ProcessingabstractIn this paper, we propose an efficient method to reconstruct surface-from-gradients (SfG). Our method is formulated under the framework of discrete geometry processing. Unlike the existing SfG approaches, we transfer the continuous reconstruction problem into a discrete space and efficiently solve the problem via a sequence of least-square optimization steps. Our discrete formulation brings three advantages: 1) the reconstruction preserves sharp-features, 2) sparse/incomplete set of gradients can be well handled, and 3) domains of computation can have irregular boundaries. Our formulation is direct and easy to implement, and the comparisons with state-of-the-arts show the effectiveness of our method. Wuyuan Xie, Charlie C. L. Wang, Ronald Chung |
CVPR | 1 |
| 2009 | Unsupervised Human Motion Analysis Using Automatic Label TreesabstractGiving different human motions, there may be similar motion features in some body components, e.g., the motion features of arms when performing jogging and running, which are thus less discriminative for motion classification. In this paper, we consider counting less on these body components that have less discriminative information amongst different human motions. To this end, we present a new topic model, probabilistic latent semantic analysis based on multiple bags (M-PLSA), in which not all body components are considered equally important, i.e., motion features of less discriminative components are made less use of so that their contributions for classification are reduced. We use sparse spatio-temporal features extracted from videos to create visual words which are later assigned to different body components that they are detected from, so that co-occurrence matrices of different components can be calculated based on their corresponding vocabularies. Such label task can be automatically fulfilled by using the query visual words, i.e., words whose component labels are unknown, to traverse an automatic label tree (ALT) that grows from the training words with component labels. We show the performance of our approach on KTH dataset. Kui Jia, Wuyuan Xie |
SMC | 2 |