Weiliang Meng

dblp:58/7599 · DBLP profile ↗
← Back
90ranked-venue papers
2as first author
74since 2021 · last 2026
0000-0002-3221-4981ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 64 · 2 first-author · 50 since 2021Artificial intelligence and machine learning · 22 · 22 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 5 since 2021Systems, architecture and hardware · 3 · 3 since 2021Computer networks · 3 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 2
YearPublicationVenuePosition
2026 Explicit Temporal-Semantic Modeling for Dense Video Captioning via Context-Aware Cross-Modal Interaction
abstract
Dense video captioning jointly localizes and captions salient events in untrimmed videos. Recent methods primarily focus on leveraging additional prior knowledge and advanced multi-task architectures to achieve competitive performance. However, these pipelines rely on implicit modeling that uses frame-level or fragmented video features, failing to capture the temporal coherence across event sequences and comprehensive semantics within visual contexts. To address this, we propose an explicit temporal-semantic modeling framework called Context-Aware Cross-Modal Interaction (CACMI), which leverages both latent temporal characteristics within videos and linguistic semantics from text corpus. Specifically, our model consists of two core components: Cross-modal Frame Aggregation aggregates relevant frames to extract temporally coherent, event-aligned textual features through cross-modal retrieval; and Context-aware Feature Enhancement utilizes query-guided attention to integrate visual dynamics with pseudo-event semantics. Extensive experiments on the ActivityNet Captions and YouCook2 datasets demonstrate that CACMI achieves the state-of-the-art performance on dense video captioning task.
Mingda Jia, Weiliang Meng, Zenghuang Fu, Ju Xin, Rongtao Xu, Jiguang Zhang, Xiaopeng Zhang 0001
AAAI2
2026 Explicit to implicit presentation for 3D unbounded open scenes reconstruction: the survey
Jiguang Zhang, Weiliang Meng, Zhaohui Zhang 0002, Xiaopeng Zhang 0001
Expert Syst. Appl.3
2026 WM-DETR: Dual-branch wavelet-Mamba and sparse attention for robust underwater object detection
Shunpeng Chen, Zenghuang Fu, Longzhao Huang, Shengpeng Xu, Yixian Kong, Changwei Wang 0001, Weiliang Meng, Xiaopeng Zhang 0001
Expert Syst. Appl.9
2026 One-shot motion talking head generation with audio-driven model
Weiliang Meng, Yaonan Wang 0001
Expert Syst. Appl.3
2026 MAFNet: Mamba-based asymmetric fusion network for event-based motion deblurring
Qihang Jiang, Weiliang Meng, Yumeng Ren, Hailong Zou, Shushan Qiao
Inf. Sci.3
2026 FOSP: Feature Orientated Sand Painting Generation
abstract
ABSTRACT Sand painting is a visually distinctive art form characterized by granular textures and diverse expressive techniques, where sand grains are manipulated through various hand movements such as waving, seeping, sweeping, and stroking. However, traditional stylized methods often fail to capture the fine textures and diverse techniques unique to sand painting. In this paper, we propose a Feature‐Oriented Sand Painting (FOSP) inspired by real sand painting techniques, aiming to produce sand paintings with authentic sandy textures and diverse techniques. Our FOSP comprises three main modules: overall sand waving generation, extraction and drawing of sand reduction regions, and simulation of sand grain accumulation with multiple techniques. The first module generates fine sand‐grain textures, whereas the latter two focus on contour rendering to enrich detail expression. Experiments validate that our FOSP can generate high‐quality sand paintings rapidly, outperforming existing methods in sand painting generation tasks.
Meng Yang 0011, Mengting Zhu, Jianglang Kang, Weiliang Meng, Ping Li 0016
Comput. Animat. Virtual Worlds4
2026 MEDP: Multimodal-Enhanced Dynamic Prototype learning for few-shot dynamic scene graph generation
Ziheng Huang, Weiliang Meng, Changbo Wang, Gaoqi He
Knowl. Based Syst.3
2026 Hierarchical segmentation-guided diffusion framework for high-fidelity sonar image generation
Weiliang Meng, Chenghanxue Tang, Longyu Jiang
Multim. Syst.2
2026 FaceEditor: Text-driven and mask-constrained face attribute editing
Lin Zhang 0041, Weiliang Meng, Paul L. Rosin, Yukun Lai, Yaonan Wang 0001
Pattern Recognit.3
2026 P3C-DNet: Pseudo-Groundtruth Contrastive Learning With Color Calibration Dehazing Network
abstract
Existing dehazing methods primarily rely on synthetic hazy images for supervised learning. While effective on synthetic datasets, these methods often struggle to generalize to real-world hazy images, leading to issues such as color distortion and incomplete haze removal. Moreover, their limited adaptability to real-world datasets and inability to handle complex haze scenarios remain significant challenges. To address these limitations, we propose a novel unsupervised framework P3C-DNet (Pseudo-groundtruth Contrastive learning with Color Calibration Dehazing Network). Our P3C-DNet introduces a Pseudo-groundtruth image generation strategy through the Pseudo-groundtruth Contrastive Supervision (PCS) module, which overcomes the lack of real haze-free training data by generating high-quality Pseudo-groundtruth images. To further refine the dehazing process, we incorporate a codebook-based image coding and matching mechanism that aligns Pseudo-groundtruth images with hazy inputs, enhancing the accuracy and detail of the dehazed outputs. To address the prevalent issue of color distortion, especially in complex environments, our P3C-DNet integrates a Dynamic Color Restoration Block (DCRB) to ensure visual quality and color consistency in the dehazed results. Experimental evaluations demonstrate that our P3C-DNet achieves superior performance in haze removal, color fidelity, and detail preservation, significantly outperforming existing methods and setting a new benchmark for real-world dehazing tasks.
Ze Ouyang, Weiliang Meng, Paul L. Rosin, Yukun Lai, Yaonan Wang 0001
IEEE Trans. Image Process.3
2026 Adaptive in Adapter: Boosting Open-Vocabulary Semantic Segmentation With Adaptive Dropout Adapter
abstract
Open-vocabulary semantic segmentation is a challenging multimedia task that requires segmentation and recognition of unseen word classes during the testing phase. Recent works bridge the gap between closed and open-vocabulary recognition by introducing large-scale visual language models such as CLIP with cross-modal alignment capabilities. To preserve multimodal alignment capabilities, it is common to freeze the parameters of the CLIP and then add additional learnable components such as adapters to expand to downstream tasks. However, for the open-vocabulary semantic segmentation task, the plain adapter suffers from overfitting the closed-vocabulary classes and impairs performance on the open-vocabulary unseen classes. In addition, since CLIP is trained to perform image-level alignment can cause the network to over-focus on partially discriminative regions, resulting in incomplete segmentation masks. To alleviate the above problems, we introduce adaptive dropout adapters to release theAdaptiveInAdapter (i.e.AIA) from the following two aspects:i)A Generalization Feature Selection Adapter (GFSA) is proposed to improve the generalization of network over unseen classes.ii)A Discriminative Region Mask Adapter (DRMA) is proposed for retrofitting CLIP backbone, has provided region free biased features for segmentation mask generation. Meanwhile, our proposed AIA achieves the current state-of-the-art performance on several open-vocabulary semantic segmentation benchmarks. Code is available athttps://github.com/clearxu/AIA.
Changwei Wang 0001, Rongtao Xu, Shibiao Xu, Jiguang Zhang, Xiaoqiang Teng, Weiliang Meng, Xiaopeng Zhang 0001
IEEE Trans. Multim.8
2026 Robust detection in complex construction sites: HiPA-DETR with weather-aware and cross-domain generalization
Zenghuang Fu, Muyang Zhang, Changwei Wang 0001, Weiliang Meng, Jiguang Zhang, Xiaopeng Zhang 0001
Vis. Comput.7
2026 A survey of revolutionizing football coaching with virtual reality
Lijuan Mao, Weiliang Meng, Meng Yang 0011
Vis. Comput.3
2025 PanoDiT: Panoramic Videos Generation with Diffusion Transformer
abstract
As immersive experiences become increasingly popular, panoramic video has garnered significant attention in both research and applications. The high cost associated with capturing panoramic video underscores the need for efficient prompt-based generation methods. Although recent text-to-video (T2V) diffusion techniques have shown potential in standard video generation, they face challenges when applied to panoramic videos due to substantial differences in content and motion patterns. In this paper, we propose PanoDiT, a framework that utilizes the Diffusion Transformer (DiT) architecture to generate panoramic videos from text descriptions. Unlike traditional methods that rely on UNet-based denoising, our method leverages a transformer architecture for denoising, incorporating both temporal and global attention mechanisms. This ensures coherent frame generation and smooth motion transitions, offering distinct advantages in long-horizon generation tasks. To further enhance motion and consistency in the generated videos, we introduce DTM-LoRA and two panoramic-specific losses. Compared to previous methods, our PanoDiT achieves state-of-the-art performance across various evaluation metrics and user study, with code is available in the supplementary material.
Muyang Zhang, Yuzhi Chen, Rongtao Xu, Changwei Wang 0001, Weiliang Meng, Jianwei Guo 0003, Xiaopeng Zhang 0001
AAAI6
2025 Multi-granularity Feature Extraction Based on Long-Short Chains for Motion Retargeting
Weiliang Meng, Changbo Wang, Gaoqi He
CGI (3)3
2025 Shape-Preserving and Surface-Fitting Network for 3D Lane Detection
abstract
Current transformer-based 3D lane detection methods typically use instance activation maps (IAM) and point-to-point loss to achieve small geometric deviations of lanes. However, these methods suffer from such lane visibility issues as wrong lane extensions and lane omissions because IAM places the lane vanishing points by mistake and loses the blurred lanes. And their performance is limited by lane continuity issues while using the point-to-point loss. In this paper, we propose a shape-preserving and surface-fitting (SPSF) network to improve the lane visibility and enhance the lane continuity. The proposed SPSF network consists of three key steps: 3D lane preliminary prediction, lane shape-preserving, and 3D lane surface-fitting. First, we design a novel transformer decoder with a mask-guided denoising block to predict preliminary 3D lanes after generating 2D lane masks. Next, after mapping the preliminary 3D lanes to 2D projected lanes, lane shapes are preserved using a two-stage mask-guided strategy to avoid the visibility issues. The two-stage mask-guided strategy includes mask-directed horizontal position adjustment and visibility correction. Finally, after fitting the surface of 3D Lanes, we improve the continuity of lanes through a surface-fitting loss. Various experiments show that our work achieves SOTA performance on two standard benchmarks.
Jianhua Li 0009, Gaoqi He, Weiliang Meng
ICME5
2025 DiffLane: Diffusion Model-Based Lane Mask Generation for Accurate Video Lane Detection
abstract
Mask-based video lane detection methods currently have achieved promising performance. However, they generate irregular lane masks in complex scenes, resulting in inaccurate lane positioning. Diffusion models have achieved notable success in the field of image segmentation because of their ability to restore pixel-level details. In this paper, we propose a novel framework DiffLane, termed Diffusion Model-Based Lane Mask Generation for Accurate Video Lane Detection. The main idea of our work is to exploit the detail-restoring capability of diffusion models to generate high-quality lane masks. DiffLane includes the MultiFrame Fusion Enhancer (MFFE), the MultiScale De-noising Network (MSDN) and the Dynamic Lane Perception Unit (DLPU). In MFFE, the current frame is enhanced with visual information from the past two frames through global matching-based optical flow estimation. This enhanced frame serves as a condition for each denoising step. MSDN predicts noise through a multi-scale fusion strategy, enabling the diffusion model to remove noise and generate regular lane masks precisely. DLPU regresses the coefficient vectors from the generated lane masks with DSConv applied in two directions, completing the accurate video lane detection task. Extensive experiments on the VIL-100 and OpenLane-V datasets demonstrate that our method outperforms other state-of-the-art approaches.
Weiliang Meng, Gaoqi He, Jianhua Li 0009
ICME3
2025 Mask-Guided Transformer with Hybrid Supervision for 3D Instance Segmentation
abstract
3D instance segmentation from point clouds is a classic but strenuous research problem. Recently, Transformer-based methods have dominated 3D instance segmentation, but most previous methods use learnable queries with low instance mask recall. Meanwhile, object queries usually use masked cross-attention involving an iterative optimization process with inaccurate initial instance masks, which may cause the query to fall into sub-optimal situations. In this paper, we propose MHFormer, a novel Mask-guided Transformer with Hybrid supervision for 3D instance segmentation. We initially develop an instance semantic-aware query module to address the issue of low recall. Subsequently, we introduce ground truth masks to proactively refine the prediction masks from the Transformer layer, guiding the query update process. Furthermore, given the scarcity of matched positive samples (e.g., an average of only 13 instances per scene in ScanNet), we introduce a pioneering one-to-many supervision in 3D instance segmentation to enhance matching efficiency. Experiments and ablation studies conducted on ScanNet, ScanNet200, and S3DIS benchmarks validate the efficacy of our approach.
Jianwei Guo 0003, Haobo Qin, Yinchang Zhou, Weiliang Meng, Xiaopeng Zhang 0001
ICME5
2025 Reidentify: Context-Aware Identity Generation for Contextual Multi-Agent Reinforcement Learning
abstract
Generalizing multi-agent reinforcement learning (MARL) to accommodate variations in problem configurations remains a critical challenge in real-world applications, where even subtle differences in task setups can cause pre-trained policies to fail. To address this, we propose Context-Aware Identity Generation (CAID), a novel framework to enhance MARL performance under the Contextual MARL (CMARL) setting. CAID dynamically generates unique agent identities through the agent identity decoder built on a causal Transformer architecture. These identities provide contextualized representations that align corresponding agents across similar problem variants, facilitating policy reuse and improving sample efficiency. Furthermore, the action regulator in CAID incorporates these agent identities into the action-value space, enabling seamless adaptation to varying contexts. Extensive experiments on CMARL benchmarks demonstrate that CAID significantly outperforms existing approaches by enhancing both sample efficiency and generalization across diverse context variants.
Zhiwei Xu 0005, Xin Xin 0003, Weiliang Meng, Yiwei Shi, Hangyu Mao, Bin Zhang 0052, Dapeng Li 0001, Jiangjin Yin
ICML4
2025 DiffusionIMU: Diffusion-Based Inertial Navigation with Iterative Motion Refinement
abstract
Inertial navigation enables self-contained localization using only Inertial Measurement Units (IMUs), making it widely applicable in various domains such as navigation, augmented reality, and robotics. However, existing methods suffer from drift accumulation due to the sensor noise and difficulty capturing long-range temporal dependencies, limiting their robustness and accuracy. To address these challenges, we propose DiffusionIMU, a novel diffusion-based framework for inertial navigation. DiffusionIMU enhances direct velocity regression from IMU data through an iterative generative denoising process, progressively refining motion state estimation. It integrates the noise-adaptive feature modulation for sensor variability handling, the feature alignment mechanism for representation consistency, and the diffusion-based temporal modeling to decrease accumulated drift. Experiments show that DiffusionIMU consistently outperforms existing methods, demonstrating superior generalization to unseen users while alleviating the impact of the sensor noise.
Xiaoqiang Teng, Shibiao Xu, Zhihao Hao, Deke Guo, Hai-Sheng Li 0002, Weiliang Meng, Xiaopeng Zhang 0001
IJCAI8
2025 AccidentX: A Large-Scale Multimodal BEV Dataset for Traffic Accident Analysis and Prevention
abstract
With the rapid development and widespread application of autonomous driving technology, the accurate analysis and prevention of traffic accidents have become critical challenges. However, current traffic accident datasets are often constrained by limited scale and diversity, impeding progress in this field. To address these limitations, we introduce AccidentX, a large-scale multimodal dataset specifically curated for comprehensive traffic accident analysis and prevention. Our AccidentX comprises over 10,000 bird’s-eye view (BEV) videos generated using the CARLA simulator, with detailed annotations covering a wide range of traffic scenarios. In comparison to existing datasets such as nuScenes, our AccidentX offers seven times more video frames and leverages Vision-Language Models (VLMs) and GPT-4o for enhanced scene understanding and decision-making. We also establish a benchmark for state-of-the-art Multimodal Large Language Models (MLLMs) on AccidentX, fostering further research and innovation within the community. AccidentX will be made available as a fully open source resource for the advancement of the autonomous driving safety algorithm community.
Muyang Zhang, Mingda Jia, Weiliang Meng, Jiguang Zhang, Xiaopeng Zhang 0001
IROS5
2025 3D Scene Graph Generation with Cross-Modal Alignment and Adversarial Learning
Yujun Hu, Changbo Wang, Weiliang Meng, Gaoqi He
ICMR4
2025 D3L: Curvature-Constrained Denoising Diffusion Model for 3D Lane Detection
abstract
Monocular 3D lane detection is a challenging task for autonomous driving systems. Recent advances primarily focus on one-step methods for lane detection based on front-view features, which show promising results on straight lanes. However, curved lanes are difficult to handle with one-step prediction, which performs prediction in a single leap without gradual refinement. To address this issue, we propose a novel Denoising Diffusion Model for 3D Lane Detection framework (D3L). The main idea is to leverage the progressive generation capability of the diffusion model to generate accurate 3D curved lanes, and ensuring lane continuity through curvature constraints. The framework includes three creative components: coarse-to-fine denoiser (CFD), curvature-constrained loss (CCL) and multi-sampling aggregation strategy (MSAS). In CFD, both lane-level and point-level transformer blocks are integrated to accurately denoise 3D lanes, which effectively captures both global and local features. CCL is designed to reduce deviations in lane curvature, resulting in smoother lane continuity. This loss enhances both the accuracy and geometric consistency of lane detection, especially in complex curved scenes. MSAS is proposed to select the optimal lane point-by-point from multiple candidates, thus robustness of the lane prediction is significantly improved. Extensive experiments on two popular 3D lane detection benchmarks demonstrate that our D3 L outperforms the state-of-the-art methods.
Weiliang Meng, Gaoqi He, Jianhua Li 0009
ACM Multimedia3
2025 Collaboration Wins More: Dual-Modal Collaborative Attention Reinforcement for Mitigating Large Vision Language Models Hallucination
abstract
Large Vision-Language Models (LVLMs) have demonstrated remarkable capabilities in visual-language understanding for downstream multimodal tasks. However, these models often generate descriptions containing objects or details not present in the input image, a phenomenon commonly referred to as ''hallucination''. Existing methods focus solely on single-side hallucination mitigation: Intra-modal-only reinforcement (e.g. visual attention enhancement) ignores prompt-based guidance; Inter-modal-only correlation correction may introduce low-information visual tokens to mislead reasoning. To tackle this challenge, we propose Dual-Modal Collaborative Attention Reinforcement (DuCAR). Specifically, DuCAR is equipped with intra-visual CLS-driven sampling and cross-modal dynamic sampling, extracting important visual tokens guided by intra- and inter-modal joint information. During the multimodal fusion stage, DuCAR adaptively enhances the attention weights of these visual tokens. Our sampling and enhancement strategies in DuCAR simultaneously reinforces informative visual tokens, and suppresses attention dispersion towards question-irrelevant visual information. We conduct extensive experiments on the POPE and CHAIR hallucination benchmarks, demonstrating that our method outperforms existing state-of-the-art mitigation baselines and effectively reduces hallucinations in text generated by LVLMs. The code is available in the https://github.com/xjy2020/DuCAR.
Jiye Xie, Liangliang You, Zhiqiang Kou, Kexue Fu 0001, Youyang Qu, Wenjie Yang 0005, Jianwei Guo 0003, Weiliang Meng, Longxiang Gao, Haoran Yang 0003, Changwei Wang 0001, Yu Zhang 0133
ACM Multimedia11
2025 Token Masking Transformer for Weakly Supervised Object Localization
abstract
Weakly supervised object localization (WSOL) is both a promising and challenging task that aims to achieve object localization exclusively through image category labels for supervision. Visual transformers have recently been applied to WSOL, demonstrating significant success through the exploitation of long-range feature dependencies in self-attention mechanisms. However, the transformer-based approach suffers from the same partial activation problem as the CNN-based approach due to the use of the classification task to train self-attention map, i.e., only a few discriminative regions are assigned high attention response and thus the localization map does not cover the whole object. To alleviate this problem, we propose a plug-and-play Token Masking Transformer (TMT) method to help transformer-based WSOL methods to obtain a more complete localization map by dynamic discriminative token masking. Specifically, a batch-wise discriminative token selection strategy is first introduced to flexibly determine the tokens to be masked in each image. Then, we design a token masking transformer block to perform token masking and inspire the network to mine more object-related tokens. Besides, we also design an intermediate token activation loss to further improve the performance of TMT by imposing constraints on intermediate tokens. Extensive experiments demonstrate that our TMT can substantially improve the performance of existing transformer-based methods without increasing the computational cost, and achieves state-of-the-art performance on two mainstream benchmarks.
Changwei Wang 0001, Rongtao Xu, Shibiao Xu, Weiliang Meng, Man Zhang 0005, Xiaopeng Zhang 0001
IEEE Trans. Multim.5
2025 ROMOT: Referring-expression-comprehension open-set multi-object tracking
Wei Li 0237, Bowen Li 0014, Weiliang Meng, Jiguang Zhang, Xiaopeng Zhang 0001
Vis. Comput.4
2025 Enhancing sonar image segmentation with random fusion in a diffusion model framework
Weiliang Meng, Xixi Zhao, Longyu Jiang
Vis. Comput.2
2025 Enhanced dual-model framework for precision player tracking and ball detection in soccer videos
Meng Yang 0011, Jianglang Kang, Xiang Suo, Weiliang Meng, Lijuan Mao, Jun Qi 0001
Vis. Comput.6
2025 A survey on soccer player detection and tracking with videos
Meng Yang 0011, Linlu Jiang, Xiang Suo, Lijuan Mao, Weiliang Meng
Vis. Comput.7
2025 Msc-Net: multi-stage colorization network for real-world images with specular highlights
Meng Yang 0011, Weiliang Meng, Ping Li 0016
Vis. Comput.3
2025 Diff-pcg: diffusion point cloud generation conditioned on continuous normalizing flow
Weiliang Meng, Zhongqi Wu, Jianwei Guo 0003, Xiaopeng Zhang 0001
Vis. Comput.2
2025 PDFT: parameter-diminish fine-tuning for transformer-based models
Muyang Zhang, Weiliang Meng, Mingda Jia, Jiaming Gu, Yihua Shao, Changwei Wang 0001, Rongtao Xu, Xiaopeng Zhang 0001
Vis. Comput.2
2024 DefFusion: Deformable Multimodal Representation Fusion for 3D Semantic Segmentation
abstract
The complementarity between camera and LiDAR data makes fusion methods a promising approach to improve 3D semantic segmentation performance. Recent transformer-based methods have also demonstrated superiority in segmentation. However, multimodal solutions incorporating transformers are underexplored and face two key inherent difficulties: over-attention and noise from different modal data. To overcome these challenges, we propose a Deformable Multimodal Representation Fusion (DefFusion) framework consisting mainly of a Deformable Representation Fusion Transformer and Dynamic Representation Augmentation Modules. The Deformable Representation Fusion Transformer introduces the deformable mechanism in multimodal fusion, avoiding over-attention and improving efficiency by adaptively modeling a 2D key/value set for a given 3D query, thus enabling multimodal fusion with higher flexibility. To enhance the 2D representation and 3D representation, the Dynamic Representation Enhancement Module is proposed to dynamically remove noise in the input representation via Dynamic Grouped Representation Generation and Dynamic Mask Generation. Extensive experiments validate that our model achieves the best 3D semantic segmentation performance on SemanticKITTI and NuScenes benchmarks.
Rongtao Xu, Changwei Wang 0001, Duzhen Zhang, Man Zhang 0005, Shibiao Xu, Weiliang Meng, Xiaopeng Zhang 0001
ICRA6
2024 FEKNN: A Wi-Fi Indoor Localization Method Based on Feature Enhancement and KNN
Bowen Li 0014, Weiliang Meng, Jiguang Zhang, Xiaopeng Zhang 0001
WASA (1)4
2024 Arbitrary style transfer via multi-feature correlation
Jin Xiang, Weiliang Meng
Comput. Graph.5
2024 De-NeRF: Ultra-high-definition NeRF with deformable net alignment
abstract
Abstract Neural Radiance Field (NeRF) can render complex 3D scenes with viewpoint‐dependent effects. However, less work has been devoted to exploring its limitations in high‐resolution environments, especially when upscaled to ultra‐high resolution (e.g., 4k). Specifically, existing NeRF‐based methods face severe limitations in reconstructing high‐resolution real scenes, for example, a large number of parameters, misalignment of the input data, and over‐smoothing of details. In this paper, we present a novel and effective framework, called De‐NeRF, based on NeRF and deformable convolutional network, to achieve high‐fidelity view synthesis in ultra‐high resolution scenes: (1) marrying the deformable convolution unit which can solve the problem of misaligned input of the high‐resolution data. (2) Presenting a density sparse voxel‐based approach which can greatly reduce the training time while rendering results with higher accuracy. Compared to existing high‐resolution NeRF methods, our approach improves the rendering quality of high‐frequency details and achieves better visual effects in 4K high‐resolution scenes.
Jianing Hou, Runjie Zhang, Zhongqi Wu, Weiliang Meng, Xiaopeng Zhang 0001, Jianwei Guo 0003
Comput. Animat. Virtual Worlds4
2024 SocialVis: Dynamic social visualization in dense scenes via real-time multi-object tracking and proximity graph construction
abstract
Abstract To monitor and assess social dynamics and risks at large gatherings, we propose “SocialVis,” a comprehensive monitoring system based on multi‐object tracking and graph analysis techniques. Our SocialVis includes a camera detection system that operates in two modes: a real‐time mode, which enables participants to track and identify close contacts instantly, and an offline mode that allows for more comprehensive post‐event analysis. The dual functionality not only aids in preventing mass gatherings or overcrowding by enabling the issuance of alerts and recommendations to organizers, but also allows for the generation of proximity‐based graphs that map participant interactions, thereby enhancing the understanding of social dynamics and identifying potential high‐risk areas. It also provides tools for analyzing pedestrian flow statistics and visualizing paths, offering valuable insights into crowd density and interaction patterns. To enhance system performance, we designed the SocialDetect algorithm in conjunction with the BYTE tracking algorithm. This combination is specifically engineered to improve detection accuracy and minimize ID switches among tracked objects, leveraging the strengths of both algorithms. Experiments on both public and real‐world datasets validate that our SocialVis outperforms existing methods, showing improvement in detection accuracy and reduction in ID switches in dense pedestrian scenarios.
Bowen Li 0014, Wei Li 0237, Weiliang Meng, Jiguang Zhang, Xiaopeng Zhang 0001
Comput. Animat. Virtual Worlds4
2024 Soccer match broadcast video analysis method based on detection and tracking
abstract
Abstract We propose a comprehensive soccer match video analysis pipeline tailored for broadcast footage, which encompasses three pivotal stages: soccer field localization, player tracking, and soccer ball detection. Firstly, we introduce sports camera calibration to seamlessly map soccer field images from match videos onto a standardized two‐dimensional soccer field template. This addresses the challenge of consistent analysis across video frames amid continuous camera angle changes. Secondly, given challenges such as occlusions, high‐speed movements, and dynamic camera perspectives, obtaining accurate position data for players and the soccer ball is non‐trivial. To mitigate this, we curate a large‐scale, high‐precision soccer ball detection dataset and devise a robust detection model, which achieved the of 80.9%. Additionally, we develop a high‐speed, efficient, and lightweight tracking model to ensure precise player tracking. Through the integration of these modules, our pipeline focuses on real‐time analysis of the current camera lens content during matches, facilitating rapid and accurate computation and analysis while offering intuitive visualizations.
Meng Yang 0011, Jianglang Kang, Xiang Suo, Weiliang Meng, Lijuan Mao, Bin Sheng 0001, Jun Qi 0001
Comput. Animat. Virtual Worlds6
2024 HIDE: Hierarchical iterative decoding enhancement for multi-view 3D human parameter regression
abstract
Abstract Parametric human modeling are limited to either single‐view frameworks or simple multi‐view frameworks, failing to fully leverage the advantages of easily trainable single‐view networks and the occlusion‐resistant capabilities of multi‐view images. The prevalent presence of object occlusion and self‐occlusion in real‐world scenarios leads to issues of robustness and accuracy in predicting human body parameters. Additionally, many methods overlook the spatial connectivity of human joints in the global estimation of model pose parameters, resulting in cumulative errors in continuous joint parameters.To address these challenges, we propose a flexible and efficient iterative decoding strategy. By extending from single‐view images to multi‐view video inputs, we achieve local‐to‐global optimization. We utilize attention mechanisms to capture the rotational dependencies between any node in the human body and all its ancestor nodes, thereby enhancing pose decoding capability. We employ a parameter‐level iterative fusion of multi‐view image data to achieve flexible integration of global pose information, rapidly obtaining appropriate projection features from different viewpoints, ultimately resulting in precise parameter estimation. Through experiments, we validate the effectiveness of the HIDE method on the Human3.6M and 3DPW datasets, demonstrating significantly improved visualization results compared to previous methods.
Weitao Lin, Jiguang Zhang, Weiliang Meng, Xianglong Liu 0007, Xiaopeng Zhang 0001
Comput. Animat. Virtual Worlds3
2024 Two-particle debris flow simulation based on SPH
abstract
Abstract Debris flow is a highly destructive natural disaster, necessitating accurate simulation and prediction. Existing simulation methods tend to be overly simplified, neglecting the three‐dimensional complexity and multiphase fluid interactions, and they also lack comprehensive consideration of soil conditions. We propose a novel two‐particle debris flow simulation method based on smoothed particle hydrodynamics (SPH) for enhanced accuracy. Our method employs a sophisticated two‐particle model coupling debris flow dynamics with SPH to simulate fluid‐solid interaction effectively, which considers various soil factors, dividing terrain into variable and fixed areas, incorporating soil impact factors for realistic simulation. By dynamically updating positions and reconstructing surfaces, and employing GPU and hash lookup acceleration methods, we achieve accurate simulation with significantly efficiency. Experimental results validate the effectiveness of our method across different conditions, making it valuable for debris flow risk assessment in natural disaster management.
Jiaxiu Zhang, Meng Yang 0011, Qun'ou Jiang, Weiliang Meng
Comput. Animat. Virtual Worlds6
2024 AG-SDM: Aquascape generation based on stable diffusion model with low-rank adaptation
abstract
Abstract As an amalgamation of landscape design and ichthyology, aquascape endeavors to create visually captivating aquatic environments imbued with artistic allure. Traditional methodologies in aquascape, governed by rigid principles such as composition and color coordination, may inadvertently curtail the aesthetic potential of the landscapes. In this paper, we propose Aquascape Generation based on Stable Diffusion Models (AG‐SDM), prioritizing aesthetic principles and color coordination to offer guiding principles for real artists in Aquascape creation. We meticulously curated and annotated three aquascape datasets with varying aspect ratios to accommodate diverse landscape design requirements regarding dimensions and proportions. Leveraging the Fréchet Inception Distance (FID) metric, we trained AGFID for quality assessment. Extensive experiments validate that our AG‐SDM excels in generating hyper‐realistic underwater landscape images, closely resembling real flora, and achieves state‐of‐the‐art performance in aquascape image generation.
Muyang Zhang, Yuewei Xian, Wei Li 0237, Jiaming Gu, Weiliang Meng, Jiguang Zhang, Xiaopeng Zhang 0001
Comput. Animat. Virtual Worlds6
2024 Cross-lingual font style transfer with full-domain convolutional attention
Tian-le Ji, Paul L. Rosin, Yukun Lai, Weiliang Meng, Yaonan Wang 0001
Pattern Recognit.5
2024 DomainFeat: Learning Local Features With Domain Adaptation
abstract
Accurate and efficient keypoint detection and description is a fundamental step in various computer vision tasks. In this paper, we extract robust descriptors and detect accurate keypoints by learning local Features with Domain adaptation (DomainFeat). Specifically, our Domainfeat includes image-level domain invariance supervision, pixel-level domain consistency supervision, Pixel-Adaptive keypoint Detection(PA-Det), and cross-domain dataset with domain stable point supervision. First, we introduce the image-level domain invariance supervision to make the high-level feature distributions from different domains close by fusing domain-invariant representations in the decoder. Furthermore, to compensate for the inconsistency between descriptors corresponding to the keypoints at the pixel level, we propose the pixel-level domain consistency supervision. Then we present the Pixel-Adaptive keypoint Detection to efficiently detect accurate keypoints, which can improve accuracy by enhancing the local consistency of heatmaps. Finally, we propose an efficient approach to construct data and supervision labels in diverse domains, which can tackle complex application scenarios. With these novel modules and supervision methods, our DomainFeat can make feature detectors more accurate and descriptors more robust. Extensive experiments confirm that Domainfeat achieves state-of-the-art performance on benchmarks such as Aachen-Day-Night localization, HPatches image matching, and the challenging DNIM dataset.
Rongtao Xu, Changwei Wang 0001, Shibiao Xu, Weiliang Meng, Bin Fan 0001, Xiaopeng Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.4
2024 Exploring Intrinsic Discrimination and Consistency for Weakly Supervised Object Localization
abstract
Weakly supervised object localization (WSOL) is a challenging and promising task that aims to localize objects solely based on the supervision of image category labels. In the absence of annotated bounding boxes, WSOL methods must employ the intrinsic properties of the image classification task pipeline to generate object localizations. In this work, we propose a WSOL method for exploring the Intrinsic Discrimination and Consistency in the image classification task pipeline, and call it as IDC. First, we develop a Triplet Metrics Based Foreground Modeling (TMFM) framework to directly predict object foreground regions using intrinsic discrimination. Unlike Class Activation Map (CAM) based methods that also rely on intrinsic discrimination, our TMFM framework alleviates the problem of only focusing on the most discriminative parts by optimizing foreground and background regions synergistically. Second, we design a Dual Geometric Transformation Consistency Constraints (DGTC2) training strategy to introduce additional supervision and regularization constraints for WSOL by leveraging intrinsic geometric transformation consistency. The proposed pixel-wise and object-wise consistency constraint losses cost-effectively provide spontaneous supervision for WSOL. Extensive experiments show that our IDC method achieves significant and consistent performance gains compared to existing state-of-the-art WSOL approaches. Code is available at: https://github.com/vignywang/IDC.
Changwei Wang 0001, Rongtao Xu, Shibiao Xu, Weiliang Meng, Ruisheng Wang 0001, Xiaopeng Zhang 0001
IEEE Trans. Image Process.4
2024 SkinFormer: Learning Statistical Texture Representation With Transformer for Skin Lesion Segmentation
abstract
Accurate skin lesion segmentation from dermoscopic images is of great importance for skin cancer diagnosis. However, automatic segmentation of melanoma remains a challenging task because it is difficult to incorporate useful texture representations into the learning process. Texture representations are not only related to the local structural information learned by CNN, but also include the global statistical texture information of the input image. In this paper, we propose a transFormer network (SkinFormer) that efficiently extracts and fuses statistical texture representation for Skin lesion segmentation. Specifically, to quantify the statistical texture of input features, a Kurtosis-guided Statistical Counting Operator is designed. We propose Statistical Texture Fusion Transformer and Statistical Texture Enhance Transformer with the help of Kurtosis-guided Statistical Counting Operator by utilizing the transformer's global attention mechanism. The former fuses structural texture information and statistical texture information, and the latter enhances the statistical texture of multi-scale features. Extensive experiments on three publicly available skin lesion datasets validate that our SkinFormer outperforms other SOAT methods, and our method achieves 93.2% Dice score on ISIC 2018. It can be easy to extend SkinFormer to segment 3D images in the future.
Rongtao Xu, Changwei Wang 0001, Jiguang Zhang, Shibiao Xu, Weiliang Meng, Xiaopeng Zhang 0001
IEEE J. Biomed. Health Informatics5
2024 DTTCNet: Time-to-Collision Estimation With Autonomous Emergency Braking Using Multi-Scale Transformer Network
abstract
The rapid advancement of autonomous driving technologies has brought the significance of Autonomous Emergency Braking (AEB) systems, which are paramount in mitigating collision risk and elevating road safety by preemptively applying brakes when a potential collision is detected. Within the core mechanisms of AEB systems, the Time-to-Collision (TTC) estimation plays a pivotal role, in quantitatively determining the criticality and timing for initiating braking interventions. However, existing TTC estimation approaches exhibit sensitivity to diverse driving scenarios, compromising the performance of AEB systems, especially in instantaneous situations. To address these issues, this paper presents DTTCNet, a novel supervised deep learning model for TTC estimation that leverages multi-scale transformer architectures and multi-task losses, thereby enhancing precision and boosting system performance. The DTTCNet first extracts spatiotemporal features from raw sensor data and utilizes a supervised training strategy. The multi-scale transformer architecture effectively captures variations across different scales, while the multi-task loss function optimizes the network training performance. Our experimental results on a challenging dataset demonstrate that DTTCNet achieves approximately 20% performance improvements over existing methods in terms of accuracy. This signifies a promising approach to augmenting the safety of autonomous driving systems with the integration of aftermarket mobile devices (e.g., Mobileye and Bosch products).
Xiaoqiang Teng, Shibiao Xu, Deke Guo, Yulan Guo, Weiliang Meng, Xiaopeng Zhang 0001
IEEE Trans. Mob. Comput.5
2024 Wave-Like Class Activation Map With Representation Fusion for Weakly-Supervised Semantic Segmentation
abstract
The Class Activation Map (CAM) is widely used to generate pseudo-labels for Weakly Supervised Semantic Segmentation (WSSS), while it does not adequately consider the modeling of foreground-independent information, resulting in prone to false positive pixels. In this paper, we propose a Wave-like Class Activation Map (WaveCAM) from the perspective of representation fusion and dynamic aggregation representation to alleviate the above problem. Specifically, our WaveCAM includes the foreground-aware representation modeling that enhances perception of foreground information, and the foreground-independent representation modeling that enhances perception of foreground-independent information, and a representation-adaptive fusion module that fuses the two representations. Both representations are expressed as wave functions with amplitude and phase to dynamically aggregate representations and extract semantic information after initialization, and they are fused through the adaptive fusion module to obtain an output containing rich semantic information. Extensive experiments on PASCAL VOC 2012 dataset and MS COCO 2014 dataset validate that our WaveCAM can easily embed multi-stage WSSS and end-to-end WSSS, achieving the state-of-the-art performance.
Rongtao Xu, Changwei Wang 0001, Shibiao Xu, Weiliang Meng, Xiaopeng Zhang 0001
IEEE Trans. Multim.4
2024 Accurate Lung Nodule Segmentation With Detailed Representation Transfer and Soft Mask Supervision
abstract
Accurate lung lesion segmentation from computed tomography (CT) images is crucial to the analysis and diagnosis of lung diseases, such as COVID-19 and lung cancer. However, the smallness and variety of lung nodules and the lack of high-quality labeling make the accurate lung nodule segmentation difficult. To address these issues, we first introduce a novel segmentation mask named " soft mask," which has richer and more accurate edge details description and better visualization, and develop a universal automatic soft mask annotation pipeline to deal with different datasets correspondingly. Then, a novel network with detailed representation transfer and soft mask supervision (DSNet) is proposed to process the input low-resolution images of lung nodules into high-quality segmentation results. Our DSNet contains a special detailed representation transfer module (DRTM) for reconstructing the detailed representation to alleviate the small size of lung nodules images and an adversarial training framework with soft mask for further improving the accuracy of segmentation. Extensive experiments validate that our DSNet outperforms other state-of-the-art methods for accurate lung nodule segmentation, and has strong generalization ability in other accurate medical segmentation tasks with competitive results. Besides, we provide a new challenging lung nodules segmentation dataset for further studies (https://drive.google.com/file/d/15NNkvDTb_0Ku0IoPsNMHezJRTH1Oi1wm/view?usp=sharing).
Changwei Wang 0001, Rongtao Xu, Shibiao Xu, Weiliang Meng, Jun Xiao 0005, Xiaopeng Zhang 0001
IEEE Trans. Neural Networks Learn. Syst.4
2024 Soccer player tracking and data correction based on attention with full-field videos
Meng Yang 0011, Linlu Jiang, Xiang Suo, Weiliang Meng, Lijuan Mao
Vis. Comput.7
2023 Self Correspondence Distillation for End-to-End Weakly-Supervised Semantic Segmentation
abstract
Efficiently training accurate deep models for weakly supervised semantic segmentation (WSSS) with image-level labels is challenging and important. Recently, end-to-end WSSS methods have become the focus of research due to their high training efficiency. However, current methods suffer from insufficient extraction of comprehensive semantic information, resulting in low-quality pseudo-labels and sub-optimal solutions for end-to-end WSSS. To this end, we propose a simple and novel Self Correspondence Distillation (SCD) method to refine pseudo-labels without introducing external supervision. Our SCD enables the network to utilize feature correspondence derived from itself as a distillation target, which can enhance the network's feature learning process by complementing semantic information. In addition, to further improve the segmentation accuracy, we design a Variation-aware Refine Module to enhance the local consistency of pseudo-labels by computing pixel-level variation. Finally, we present an efficient end-to-end Transformer-based framework (TSCD) via SCD and Variation-aware Refine Module for the accurate WSSS task. Extensive experiments on the PASCAL VOC 2012 and MS COCO 2014 datasets demonstrate that our method significantly outperforms other state-of-the-art methods. Our code is available at https://github.com/Rongtao-Xu/RepresentationLearning/tree/main/SCD-AAAI2023.
Rongtao Xu, Changwei Wang 0001, Jiaxi Sun, Shibiao Xu, Weiliang Meng, Xiaopeng Zhang 0001
AAAI5
2023 Audio-Driven Lips and Expression on 3D Human Face
Weiliang Meng, Shibiao Xu, Xiaopeng Zhang 0001
CGI3
2023 Treating Pseudo-labels Generation as Image Matting for Weakly Supervised Semantic Segmentation
abstract
Generating accurate pseudo-labels under the supervision of image categories is a crucial step in Weakly Supervised Semantic Segmentation (WSSS). In this work, we propose a Mat-Label pipeline that provides a fresh way to treat WSSS pseudo-labels generation as an image matting task. By taking a trimap as input which specifies the foreground, background and unknown regions, the image matting task outputs an object mask with fine edges. The intuition behind our Mat-Label is that generating trimap is much easier than generating pseudo-labels directly under weakly supervised setting. Although current CAM-based methods are off-the-shelf solutions for generating a trimap, they suffer from cross-category and foreground-background pixel prediction confusion. To solve this problem, we develop a Double Decoupled Class Activation Map (D2CAM) for Mat-Label to generate a high-quality trimap. By drawing on the idea of metric learning, we explicitly model class activation map with category decoupling and foreground-background decoupling. We also design two simple yet effective refinement constraints for D2CAM to stabilize optimization and eliminate non-exclusive activation. Extensive experiments validate that our Mat-Label achieves substantial and consistent performance gains compared to current state-of-the-art WSSS approaches.
Changwei Wang 0001, Rongtao Xu, Shibiao Xu, Weiliang Meng, Xiaopeng Zhang 0001
ICCV4
2023 FeaCo: Reaching Robust Feature-Level Consensus in Noisy Pose Conditions
abstract
Collaborative perception offers a promising solution to overcome challenges such as occlusion and long-range data processing. However, limited sensor accuracy leads to noisy poses that misalign observations among vehicles. To address this problem, we propose the FeaCo, which achieves robust Feature-level Consensus among collaborating agents in noisy pose conditions without additional training. We design an efficient Pose-error Rectification Module (PRM) to align derived feature maps from different vehicles, reducing the adverse effect of noisy pose and bandwidth requirements. We also provide an effective multi-scale Cross-level Attention Module (CAM) to enhance information aggregation and interaction between various scales. Our FeaCo outperforms all other localization rectification methods, as validated on both the collaborative perception simulation dataset OPV2V and real-world dataset V2V4Real, reducing heading error and enhancing localization accuracy across various error levels. Our code is available at: https://github.com/jmgu0212/FeaCo.git.
Jiaming Gu, Muyang Zhang, Weiliang Meng, Shibiao Xu, Jiguang Zhang, Xiaopeng Zhang 0001
ACM Multimedia4
2023 Sand painting conversion based on detail preservation
Mengting Zhu, Meng Yang 0011, Weiliang Meng, Ping Li 0016
Comput. Graph.3
2023 Automatic polyp segmentation via image-level and surrounding-level context fusion deep neural network
Changwei Wang 0001, Rongtao Xu, Shibiao Xu, Weiliang Meng, Xiaopeng Zhang 0001
Eng. Appl. Artif. Intell.4
2023 Dual-stream Representation Fusion Learning for accurate medical image segmentation
Rongtao Xu, Changwei Wang 0001, Shibiao Xu, Weiliang Meng, Xiaopeng Zhang 0001
Eng. Appl. Artif. Intell.4
2023 RC-Net: Row and Column Network with Text Feature for Parsing Floor Plan Images
Weiliang Meng, Zhengda Lu, Jianwei Guo 0003, Jun Xiao 0005, Xiaopeng Zhang 0001
J. Comput. Sci. Technol.2
2023 Attention Weighted Local Descriptors
abstract
Local features detection and description are widely used in many vision applications with high industrial and commercial demands. With large-scale applications, these tasks raise high expectations for both the accuracy and speed of local features. Most existing studies on local features learning focus on the local descriptions of individual keypoints, which neglect their relationships established from global spatial awareness. In this paper, we present AWDesc with a consistent attention mechanism (CoAM) that opens up the possibility for local descriptors to embrace image-level spatial awareness in both the training and matching stages. For local features detection, we adopt local features detection with feature pyramid to obtain more stable and accurate keypoints localization. For local features description, we provide two versions of AWDesc to cope with different accuracy and speed requirements. On the one hand, we introduce Context Augmentation to address the inherent locality of convolutional neural networks by injecting non-local context information, so that local descriptors can "look wider to describe better". Specifically, well-designed Adaptive Global Context Augmented Module (AGCA) and Diverse Surrounding Context Augmented Module (DSCA) are proposed to construct robust local descriptors with context information from global to surrounding. On the other hand, we design an extremely lightweight backbone network coupled with the proposed special knowledge distillation strategy to achieve the best trade-off in accuracy and speed. What is more, we perform thorough experiments on image matching, homography estimation, visual localization, and 3D reconstruction tasks, and the results demonstrate that our method surpasses the current state-of-the-art local descriptors. Code is available at: https://github.com/vignywang/AWDesc.
Changwei Wang 0001, Rongtao Xu, Ke Lu 0002, Shibiao Xu, Weiliang Meng, Bin Fan 0001, Xiaopeng Zhang 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2023 Toward Accurate and Efficient Road Extraction by Leveraging the Characteristics of Road Shapes
abstract
Automatically extracting roads from very high resolution (VHR) remote sensing images is of great importance in a wide range of remote sensing applications. However, complex shapes of roads (i.e., long, geometrically deformed, and thin) always affected the extraction accuracy, which is one of the challenges of road extraction. Based on the insight into road shape characteristics, we propose a novel road shape aware network (RSANet) to achieve efficient and accurate road extraction. First, we introduce the Efficient Strip Transformer Module (ESTM) to efficiently capture the global context to model the long-distance dependence required by the long roads. Second, we design a Geometric Deformation Estimation Module (GDEM) to adaptively extract the context from the shape deformation caused by shooting roads from different perspectives. Third, we provide a simple but effective Road Edge Focal Loss (REF loss) to make the network focus on optimizing the pixels around the road to alleviate the unbalanced distribution of foreground and background pixels caused by the roads being too thin. Finally, we conduct extensive evaluations on public datasets to verify the effectiveness of RSANet and each of the proposed components. Experiments validate that our RSANet outperforms state-of-the-art methods for road extraction in remote sensing images.
Changwei Wang 0001, Rongtao Xu, Shibiao Xu, Weiliang Meng, Ruisheng Wang 0001, Jiguang Zhang, Xiaopeng Zhang 0001
IEEE Trans. Geosci. Remote. Sens.4
2023 RSSFormer: Foreground Saliency Enhancement for Remote Sensing Land-Cover Segmentation
abstract
High spatial resolution (HSR) remote sensing images contain complex foreground-background relationships, which makes the remote sensing land cover segmentation a special semantic segmentation task. The main challenges come from the large-scale variation, complex background samples and imbalanced foreground-background distribution. These issues make recent context modeling methods sub-optimal due to the lack of foreground saliency modeling. To handle these problems, we propose a Remote Sensing Segmentation framework (RSSFormer), including Adaptive TransFormer Fusion Module, Detail-aware Attention Layer and Foreground Saliency Guided Loss. Specifically, from the perspective of relation-based foreground saliency modeling, our Adaptive Transformer Fusion Module can adaptively suppress background noise and enhance object saliency when fusing multi-scale features. Then our Detail-aware Attention Layer extracts the detail and foreground-related information via the interplay of spatial attention and channel attention, which further enhances the foreground saliency. From the perspective of optimization-based foreground saliency modeling, our Foreground Saliency Guided Loss can guide the network to focus on hard samples with low foreground saliency responses to achieve balanced optimization. Experimental results on LoveDA datasets, Vaihingen datasets, Potsdam datasets and iSAID datasets validate that our method outperforms existing general semantic segmentation methods and remote sensing segmentation methods, and achieves a good compromise between computational overhead and accuracy. Our code is available at https://github.com/Rongtao-Xu/RepresentationLearning/tree/main/RSSFormer-TIP2023.
Rongtao Xu, Changwei Wang 0001, Jiguang Zhang, Shibiao Xu, Weiliang Meng, Xiaopeng Zhang 0001
IEEE Trans. Image Process.5
2023 CNDesc: Cross Normalization for Local Descriptors Learning
abstract
For a long time, the local descriptors learning benefited from the use of L2 normalization, which projects the descriptor space onto the hypersphere. However, there is no free lunch in the world. Although hypersphere description space stabilizes the optimization and improves the repeatability of the descriptors, it causes the descriptors to have a denser distribution, which reduces the discrimination between descriptors and leads to some incorrect matches. To alleviate this problem, we propose the learnablecross normalizationtechnology as an alternative to L2 normalization, which can achieve a consistent improvement in several of the current popular local descriptors. In addition, we propose an ER-Backbone that can efficiently reuse features in descriptors extraction and an IDC Loss that can provide an image-level description space distribution consistency constraint to further stimulate the performance of the local descriptors. Based on the above innovations, we provide a novel local descriptors extraction method named CNDesc. We perform experiments on image matching, homography estimation, 3D reconstruction, and visual localization tasks, and the results demonstrate that our CNDesc surpasses the current state-of-the-art local descriptors. Our code is available athttps://github.com/vignywang/CNDesc.
Changwei Wang 0001, Rongtao Xu, Shibiao Xu, Weiliang Meng, Xiaopeng Zhang 0001
IEEE Trans. Multim.4
2023 HTCViT: an effective network for image classification and segmentation based on natural disaster datasets
Wei Li 0237, Muyang Zhang, Weiliang Meng, Shibiao Xu, Xiaopeng Zhang 0001
Vis. Comput.4
2022 MTLDesc: Looking Wider to Describe Better
abstract
Limited by the locality of convolutional neural networks, most existing local features description methods only learn local descriptors with local information and lack awareness of global and surrounding spatial context. In this work, we focus on making local descriptors ``look wider to describe better'' by learning local Descriptors with More Than Local information (MTLDesc). Specifically, we resort to context augmentation and spatial attention mechanism to make the descriptors obtain non-local awareness. First, Adaptive Global Context Augmented Module and Diverse Local Context Augmented Module are proposed to construct robust local descriptors with context information from global to local. Second, we propose the Consistent Attention Weighted Triplet Loss to leverage spatial attention awareness in both optimization and matching of local descriptors. Third, Local Features Detection with Feature Pyramid is proposed to obtain more stable and accurate keypoints localization. With the above innovations, the performance of the proposed MTLDesc significantly surpasses the current state-of-the-art local descriptors on HPatches, Aachen Day-Night localization and InLoc indoor localization benchmarks. Our code is available at https://github.com/vignywang/MTLDesc.
Changwei Wang 0001, Rongtao Xu, Shibiao Xu, Weiliang Meng, Bin Fan 0001, Xiaopeng Zhang 0001
AAAI5
2022 DOMAINDESC: Learning Local Descriptors With Domain Adaptation
abstract
Robust and efficient local descriptor is crucial in a wide range of applications. In this paper, we propose a novel descriptor DomainDesc which is invariant as much as possible by learning local Descriptor with Domain adaptation. We design the feature-level domain adaptation loss to improve robustness of our DomainDesc by punishing inconsistent high-level feature distributions of different images, while we present the pixel-level cross-domain consistency loss to compensate for the inconsistency between the descriptors corresponding to the keypoints at the pixel level. Besides, we adopt a new architecture to make the descriptor contain as much information as possible, and combine triplet loss and cross-domain consistency loss for descriptor supervision to ensure the distinguished ability of our descriptor. Finally, we give a cross-domain dataset generation strategy to quickly construct our training dataset for diverse domains to adapt to complex application scenarios. Experiments validate that our DomainDesc achieves state-of-the-art performances on HPatches image matching benchmark and Aachen-Day-Night localization benchmark.
Rongtao Xu, Changwei Wang 0001, Bin Fan 0001, Shibiao Xu, Weiliang Meng, Xiaopeng Zhang 0001
ICASSP6
2022 Softgan: Towards Accurate Lung Nodule Segmentation via Soft Mask Supervision
abstract
Accurate lung nodule segmentation from Computed Tomog-raphy (CT) images is crucial to the analysis and diagnosis of lung diseases such as COVID-19 and lung cancer. How-ever, due to the variety of lung nodules and the lack of high-quality labeling, accurate lung nodule segmentation is still a challenging problem. In this paper, we propose a novel paradigm including an automatic accurate annotation pipeline and a segmentation network for this task. First, we introduce a new segmentation mask representation named Soft Mask which has richer and more accurate edge details description and better visualization, and we design a universal automatic Soft Mask annotation pipeline to deal with different datasets. Besides, we provide a new challenging lung nodules segmen-tation dataset with traditional binarized masks and our soft masks for further studies. Second, we propose an effective network called SoftGAN that includes an improved back-bone and an adversarial training framework with Soft Mask, in order to improve the performance of accurate lung nodules segmentation. Extensive experiments validate that our Soft-GAN outperforms the state-of-the-art methods for accurate lung nodule segmentation. [Datasetrelease]
Changwei Wang 0001, Rongtao Xu, Shibiao Xu, Weiliang Meng, Jun Xiao 0005, Qimin Peng, Xiaopeng Zhang 0001
ICME4
2022 GeoROS: Georeferenced Real-time Orthophoto Stitching with Unmanned Aerial Vehicle
abstract
Simultaneous orthophoto stitching during the flight of Unmanned Aerial Vehicles (UAV) can greatly promote the practicability and instantaneity of diverse applications such as emergency disaster rescue, digital agriculture, and cadastral survey, which is of remarkable interest in aerial photogrammetry. However, the inaccurately estimated camera poses and the intuitive fusion strategy of existing methods lead to misalignment and distortion artifacts in orthophoto mosaics. To address these issues, we propose a Georeferenced Real-time Orthophoto Stitching method (GeoROS), which can achieve efficient and accurate camera pose estimation through exploiting geolocation information in monocular visual simultaneous localization and mapping (SLAM) and fuse transformed images via orthogonality-preserving criterion. Specifically, in the SLAM process, georeferenced tracking is employed to acquire high-quality initial camera poses with a geolocation based motion model and facilitate non-linear pose optimization. Meanwhile, we design a georeferenced mapping scheme by introducing robust geolocation constraints in joint optimization of camera poses and the position of landmarks. Finally, aerial images warped with localized cameras are fused by considering both the orthogonality of camera orientation relative to the ground plane and the pixel centrality to fulfill global orthorectification. Besides, we construct two datasets with global navigation satellite system (GNSS) information of different scenarios and validate the superiority of our GeoROS method compared with state-of-the-art methods in accuracy and efficiency.
Guangze Gao, Mengke Yuan, Jiaming Gu, Weiliang Meng, Shibiao Xu, Xiaopeng Zhang 0001
IROS5
2022 DA-Net: Dual Branch Transformer and Adaptive Strip Upsampling for Retinal Vessels Segmentation
Changwei Wang 0001, Rongtao Xu, Shibiao Xu, Weiliang Meng, Xiaopeng Zhang 0001
MICCAI (2)4
2022 Chameleon: A Self-adaptive cache strategy under the ever-changing access frequency in edge network
Pengmiao Li, Yuchao Zhang 0004, Wendong Wang 0003, Weiliang Meng, Ke Xu 0002
Comput. Commun.4
2022 Instance segmentation of biological images using graph convolutional network
Rongtao Xu, Changwei Wang 0001, Shibiao Xu, Weiliang Meng, Xiaopeng Zhang 0001
Eng. Appl. Artif. Intell.5
2022 Driving EEG based multilayer dynamic brain network analysis for steering process
Wenwen Chang, Weiliang Meng, Guanghui Yan
Expert Syst. Appl.2
2022 Triple-strip attention mechanism-based natural disaster images classification and segmentation
Mengke Yuan, Jiaming Gu, Weiliang Meng, Shibiao Xu, Xiaopeng Zhang 0001
Vis. Comput.4
2021 Towards Effective Adversarial Attack Against 3D Point Cloud Classification
abstract
In the domain of 3D point cloud classification, deep learning based classifiers have made significant progress, while they have been also proven to be vulnerable on the adversarial at-tack at the same time. Some recent works employ the attack methods that devised for image classification such as projected gradient descent (PGD) to attack the 3D classifiers, but their performances seem quite limited when faced with statistical operations including point cloud denoising and point cloud upsampling. In this paper, we propose ‘SmoothAttack’, a new attack that can craft adversarial point clouds robust to statistical operations. SmoothAttack can be easily applied in both global constraint and pointwise constraint. Besides, we analyze the directions of perturbations onto the point cloud during the iteration process, where SmoothAttack can some-how stabilize the direction and make full use of the adversarial budgets. Experiments validate that our ‘SmoothAttack’ can raise the attack success rates against statistical defenses up to 98% for untargeted attack and 91% for targeted attack on ModelNet40 database when fooling the classifiers Point-Net and DGCNN.
Chengcheng Ma, Weiliang Meng, Baoyuan Wu, Shibiao Xu, Xiaopeng Zhang 0001
ICME2
2021 DC-Net: Dual Context Network for 2D Medical Image Segmentation
Rongtao Xu, Changwei Wang 0001, Shibiao Xu, Weiliang Meng, Xiaopeng Zhang 0001
MICCAI (1)4
2021 Data-driven floor plan understanding in rural residential buildings via deep recognition
Zhengda Lu, Jianwei Guo 0003, Weiliang Meng, Jun Xiao 0005, Xiaopeng Zhang 0001
Inf. Sci.4
2020 Efficient Joint Gradient Based Attack Against SOR Defense for 3D Point Cloud Classification
abstract
Deep learning based classifiers on 3D point cloud data have been shown vulnerable to adversarial examples, while a defense strategy named Statistical Outlier Removal (SOR) is widely adopted to defend adversarial examples successfully, by discarding outlier points in the point cloud.
Chengcheng Ma, Weiliang Meng, Baoyuan Wu, Shibiao Xu, Xiaopeng Zhang 0001
ACM Multimedia2
2020 Unsupervised Multi-View Constrained Convolutional Network for Accurate Depth Estimation
abstract
Accurate depth estimation from images is a fundamental problem in computer vision. In this paper, we propose an unsupervised learning based method to predict high-quality depth map from multiple images. A novel multi-view constrained DenseDepthNet is designed for this task. Our DenseDepthNet can effectively leverage both the low-level and high-level features of input images and generate appealing results, especially with sharp details. We employ the public datasets KITTI and Cityscapes for training in an end-to-end unsupervised fashion. A novel depth consistency loss based on multi-view geometry constraint is also applied to the corresponding points across pairwise images, which helps to improve the quality of predicted depth maps significantly. We conduct comprehensive evaluations on our DenseDepthNet and our depth consistency loss function. Experiments validate that our method outperforms the state-of-the-art unsupervised methods and produce comparable results with supervised methods.
Shibiao Xu, Baoyuan Wu, Weiliang Meng, Xiaopeng Zhang 0001
IEEE Trans. Image Process.5
2019 Parameter optimization criteria guided 3D point cloud classification
Hongjun Li 0002, Weiliang Meng, Shiming Xiang, Xiaopeng Zhang 0001
Multim. Tools Appl.2
2018 Large-scale 3D Point Cloud Classification Based On Feature Description Matrix By CNN
abstract
Large-scale 3D Point cloud classification is a basic topic for various applications. Traditional geometries features are usually independent of each other and difficult to adapt to a fixed classification model. With the rise of the neural network, deep learning is considered in 3D point cloud application. 3D points are difficult to feed the neural network directly based on deep learning, as they cannot be arranged in a fixed order as image pixels. In this paper, we combine traditional feature-based methods with the Convolutional neural network(CNN) to finish the classification task. The core idea is to construct a feasible structure called Feature Description Matrix(FDM) which encapsulates the local feature of the point to feed CNN for training and testing. By extracting geometry features and designed Feature Description Vectors(FDV) for FDM, a simple mechanism for point cloud classification is given, and experiments validate the effectiveness of our method, with higher classification accuracy compared to state-of-art works.
Lei Wang 0089, Weiliang Meng, Runping Xi, Yanning Zhang 0001, Ling Lu, Xiaopeng Zhang 0001
CASA2
2018 Accurate blind deblurring using salientpatch-based prior for large-size images
Chengcheng Ma, Jiguang Zhang, Shibiao Xu, Weiliang Meng, Runping Xi, G. Hemanth Kumar, Xiaopeng Zhang 0001
Multim. Tools Appl.4
2017 Tree Branch Level of Detail Models for Forest Navigation
abstract
Abstract We present a level of detail (LOD) method designed for tree branches. It can be combined with methods for processing tree foliage to facilitate navigation through large virtual forests. Starting from a skeletal representation of a tree, we fit polygon meshes of various densities to the skeleton while the mesh density is adjusted according to the required visual fidelity. For distant models, these branch meshes are gradually replaced with semi‐transparent lines until the tree recedes to a few lines. Construction of these complete LOD models is guided by error metrics to ensure smooth transitions between adjacent LOD models. We then present an instancing technique for discrete LOD branch models, consisting of polygon meshes plus semi‐transparent lines. Line models with different transparencies are instanced on the GPU by merging multiple tree samples into a single model. Our technique reduces the number of draw calls in GPU and increases rendering performance. Our experiments demonstrate that large‐scale forest scenes can be rendered with excellent detail and shadows in real time.
Xiaopeng Zhang 0001, Guanbo Bao, Weiliang Meng, Marc Jaeger 0002, Hongjun Li 0002, Oliver Deussen, Baoquan Chen
Comput. Graph. Forum3
2017 Shape exploration of 3D heterogeneous models based on cages
Weiliang Meng, Jianwei Guo 0003, Xavier Bonaventura, Mateu Sbert, Xiaopeng Zhang 0001
Multim. Tools Appl.1
2016 Analyzing surface sampling patterns using the localized pair correlation function
abstract
Point distributions with different characteristics have a crucial influence on graphics applications. Various analysis tools have been developed in recent years, mainly for blue noise sampling in Euclidean domains. In this paper, we present a new method to analyze the properties of general sampling patterns that are distributed on mesh surfaces. The core idea is to generalize to surfaces the pair correlation function (PCF) which has successfully been employed in sampling pattern analysis and synthesis in 2D and 3D. Experimental results demonstrate that the proposed approach can reveal correlations of point sets generated by a wide range of sampling algorithms. An acceleration technique is also suggested to improve the performance of the PCF.
Weize Quan, Jianwei Guo 0003, Dong-Ming Yan 0001, Weiliang Meng, Xiaopeng Zhang 0001
Comput. Vis. Media4
2015 3D shape retrieval using viewpoint information-theoretic measures
abstract
Abstract In this paper, we present an information‐theoretic framework to compute the shape similarity between 3D polygonal models. Given a 3D model, an information channel between a sphere of viewpoints around the model and its polygonal mesh is defined to compute the specific information associated with each viewpoint. The obtained information sphere can be seen as a shape descriptor of the model. Then, given two models, their similarity is obtained by performing a registration process between the corresponding information spheres. The distance between the information histograms is also defined as a coarse measure of similarity, as well as the scalar value given by the mutual information of the channel. The performance of all these measures is tested using the Princeton Shape Benchmark database. Copyright © 2013 John Wiley & Sons, Ltd.
Xavier Bonaventura, Jianwei Guo 0003, Weiliang Meng, Miquel Feixas, Xiaopeng Zhang 0001, Mateu Sbert
Comput. Animat. Virtual Worlds3
2014 Perception-motivated multiresolution rendering on sole-cube maps
Bin Sheng 0001, Weiliang Meng, Hanqiu Sun, Wen Wu 0001, Enhua Wu
Multim. Tools Appl.2
2013 Statistical learning based facial animation
abstract
To synthesize real-time and realistic facial animation, we present an effective algorithm which combines image- and geometry-based methods for facial animation simulation. Considering the numerous motion units in the expression coding system, we present a novel simplified motion unit based on the basic facial expression, and construct the corresponding basic action for a head model. As image features are difficult to obtain using the performance driven method, we develop an automatic image feature recognition method based on statistical learning, and an expression image semi-automatic labeling method with rotation invariant face detection, which can improve the accuracy and efficiency of expression feature identification and training. After facial animation redirection, each basic action weight needs to be computed and mapped automatically. We apply the blend shape method to construct and train the corresponding expression database according to each basic action, and adopt the least squares method to compute the corresponding control parameters for facial animation. Moreover, there is a pre-integration of diffuse light distribution and specular light distribution based on the physical method, to improve the plausibility and efficiency of facial rendering. Our work provides a simplification of the facial motion unit, an optimization of the statistical training process and recognition process for facial animation, solves the expression parameters, and simulates the subsurface scattering effect in real time. Experimental results indicate that our method is effective and efficient, and suitable for computer animation and interactive applications.
Shibiao Xu, Guanghui Ma, Weiliang Meng, Xiaopeng Zhang 0001
J. Zhejiang Univ. Sci. C3
2013 Sketch-based design for green geometry and image deformation
Bin Sheng 0001, Weiliang Meng, Hanqiu Sun, Enhua Wu
Multim. Tools Appl.2
2011 Hardware instancing for real-time realistic forest rendering
abstract
Real-time rendering of vegetation is important in many applications, such as video games, internet graphics applications, landscape design and visualization. However, the visualization of large-scale forests has always been a challenge not only due to the high geometric complexity but also due to the small batch problem. Moreover, generating real-time shadows for forests will heavily increase the burden. The batch problem is caused by a large number of graphics API draw calls launched in every frame. Normally, rendering a tree model requires at least one graphics API draw call. As a forest usually consists of thousands of trees and the graphics API invocation is a relatively high cost for CPU, the large-scale forest rendering is often CPU bound.
Guanbo Bao, Weiliang Meng, Hongjun Li 0002, Xiaopeng Zhang 0001
SIGGRAPH Asia Sketches2
2011 MCGIM-Based Model Streaming for Realtime Progressive Rendering
Bin Sheng 0001, Weiliang Meng, Hanqiu Sun, Enhua Wu
J. Comput. Sci. Technol.2
2010 Robust discovery of partial rigid symmetries on 3D models
abstract
The ubiquity of symmetry in nature and man-made artifacts has made symmetry discovery an important tool for numerous applications. The most fundamental and visually prominent kind of symmetry is partial rigid symmetry, which is explained as the invariance between parts of a 3D model under a set of translation, rotation, reflection, and uniform scaling generators. Thus automatic discovery of partial rigid symmetry on general 3D models, with no assumption on the size, shape or location of the symmetric parts, keeps to be a hot topic during recent years. Among such kind of works, the transformation voting technique [Mitra et al. 2006] is most widely used, due to its high efficiency and easiness for understanding and implementation.
Kangying Cai, Weiliang Meng, Wencheng Wang 0001, Zhibo Chen 0001
SIGGRAPH ASIA (Sketches)3
2010 Differential geometry images: remeshing and morphing with local shape preservation
Weiliang Meng, Bin Sheng 0001, Weiwei Lv, Hanqiu Sun, Enhua Wu
Vis. Comput.1