EDBT 2026 Demo / reviewers in the wild / expert
Weiming Wang 0002
dblp:86/3464-2
· DBLP profile ↗
37ranked-venue papers
1as first author
32since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 26 · 1 first-author · 22 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 8 since 2021Artificial intelligence and machine learning · 7 · 5 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Fusing shape descriptors and geometric details for robust category-level object pose estimation
Yun Liu 0002, Weiming Wang 0002, Fu Lee Wang, Haoran Xie 0001, Honghua Chen, Xue Xue, Mingqiang Wei, Harry Qin |
Multim. Tools Appl. | 2 |
| 2026 | FacDNet: A low-rank factorized diffusion network with dual-U compensated attention for low-light enhancement
Peiguang Jing, Ningyuan Zhao, Lijun Lai, Weiming Wang 0002, Yuting Su 0001 |
Signal Process. | 5 |
| 2026 | CrossTracker: Robust Multi-Modal 3D Multi-Object Tracking via Cross CorrectionabstractInaccurate detections remain a critical bottleneck in 3D multi-object tracking (MOT). Recent detection fusion-based methods incorporate camera detections as supplementary to reduce false detections and compensate for missing ones in LiDAR. However, their unidirectional camera-LiDAR correction lacks a feedback mechanism, precluding iterative mutual refinement between modalities for more robust LiDAR-based tracking. Inspired by the coarse-to-fine strategy in two-stage object detection, we introduceCrossTracker, a novel two-stage framework for online multi-modal 3D MOT. CrossTracker first constructs coarse camera and LiDAR trajectories independently, then performs trajectory fusion using both current and historical frames, without requiring future data. This ensures more robust mutual refinement between modalities. Specifically, CrossTracker comprises three core modules: i) the multi-modal modeling (M3) module, which fuses data from images, point clouds, and even planar geometry derived from images to establish a robust tracking constraint; ii) the coarse trajectory generation (C-TG) module, which independently generates coarse trajectories for both modalities using the M3constraint; and iii) the trajectory fusion (TF) module, which applies mutual refinement between coarse LiDAR and camera trajectories through cross correction to ensure robust LiDAR trajectories. Extensive experiments show that CrossTracker outperforms 19 state-of-the-art methods, highlighting its effectiveness in leveraging the synergistic strengths of camera and LiDAR sensors for robust multi-modal 3D MOT. The code is available at https://github.com/lipeng-gu/CrossTracker. Lipeng Gu, Xuefeng Yan 0001, Weiming Wang 0002, Honghua Chen, Dingkun Zhu, Liangliang Nan, Mingqiang Wei |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | LCM-Net: LLM-Driven Cross-Modality MoE Feature Fusion Network for Cancer Survival AnalysisabstractCancer survival analysis aims to predict survival outcomes to evaluate the efficacy and prognosis of treatment. Although current approaches have designed diverse cross-modal learning methods to integrate genetic data and pathology images, they are frequently hindered by data redundancy. Pattern representation in high-dimensional genetic data remains a significant hurdle. Pathology data analysis is computationally intensive because of the giga-pixel resolution. Moreover, the heterogeneity of data types poses a barrier to extending multimodal fusion methods. To address the aforementioned issues, we propose a novel LLM-driven Cross-Modality MoE-feature Fusion Network (LCM-Net) with three innovative modules for boosting cancer survival prediction. Specifically, the Genomic Language Alignment (GLA) module integrates genomic features with learnable prompts. Utilizing large language models, it encodes genomic information into concise and semantically relevant representations. Then, we devise the Pathological Feature Refinement (PFR) module to serve as a plug-and-play component that filters out irrelevant regions in pathology images. Finally, we propose a Multimodal Expert Integration (MEI) module to effectively leverage the capabilities of different experts, integrating the processed features from both the genomic and pathological domains. Extensive experiments on five public datasets demonstrate that our approach outperforms state-of-the-art methods, and the ablation study confirms the effectiveness of the proposed modules. Our code is publicly available at https://github.com/script-Yang/LCM-Net. Sicheng Yang 0001, Haipeng Zhou, Weiming Wang 0002, Shifu Chen, Guang Yang 0006, Huazhu Fu, Lei Zhu 0003 |
IEEE Trans. Medical Imaging | 4 |
| 2025 | SnowMaster: Comprehensive Real-world Image Desnowing via MLLM with Multi-Model Feedback OptimizationabstractSnowfall presents significant challenges for visual data processing, necessitating specialized desnowing algorithms. However, existing models often fail to generalize effectively due to their heavy reliance on synthetic datasets. Furthermore, current real-world snowfall datasets are limited in scale and lack dedicated evaluation metrics designed specifically for snowfall degradation, thus hindering the effective integration of real snowy images into model training to reduce domain gaps. To address these challenges, we first introduce RealSnow10K, a large-scale, high-quality dataset consisting of over 10,000 annotated real-world snowy images. In addition, we curate a preference dataset comprising 36,000 expert-ranked image pairs, enabling the adaptation of multimodal large language models (MLLMs) to better perceive snowy image quality through our innovative Multi-Model Preference Optimization (MMPO). Finally, we propose the SnowMaster, which employs MMPO-enhanced MLLM to perform accurate snowy image evaluation and pseudo-label filtering for semi-supervised training. Experiments demonstrate that SnowMaster delivers superior desnowing performance under real-world conditions. Jianyu Lai, Sixiang Chen, Yunlong Lin, Tian Ye 0001, Yun Liu 0002, Song Fei, Zhaohu Xing, Weiming Wang 0002, Lei Zhu 0003 |
CVPR | 9 |
| 2025 | P2AW-YOLO11n: Improving Drone-Based Small Object Detection via AFPN and Enhanced Loss FunctionsabstractIn recent years, the decreasing production costs and rapid technological advancements of Unmanned Aerial Vehicles (UAVs) have led to their widespread adoption in critical domains such as agricultural plant protection, power line inspection, emergency rescue, and homeland security. Despite their growing utility, UAVs face significant challenges when performing tasks that require real-time object detection. The images captured by their onboard vision systems often feature densely clustered, small-sized targets, which strain the capabilities of existing detection algorithms. For instance, the YOLO11n model struggles with substantial feature loss due to its downsampling operation, limiting its effectiveness in detecting small objects in complex scenarios. To address these limitations, this paper introduces P2AW-YOLO11n, an enhanced model derived from YOLO11n, specifically designed to improve small object detection. The proposed model incorporates three key innovations. First, Adaptive Feature Pyramid Network (AFPN) is integrated into the neck of the model to enable more flexible feature fusion through adaptive mechanisms such as dynamic weight learning and cross-scale interactions, effectively mitigating feature dilution. Second, the P2 detection head utilizes large-scale$160 \times 160$feature maps to capture finer target details, significantly enhancing the model's ability to detect small objects. Third, a refined version of the Wise-IoU (WIoU) loss function is introduced to address challenges posed by low-quality samples, thereby improving the model's generalization capabilities. Extensive comparative and ablation experiments conducted on the VisDrone dataset demonstrate the efficacy of the proposed algorithm. P2AWYOLO11n achieves a 7.5% improvement in Precision, a 5.5% increase in Recall, and a 6.9% boost in$\text{mAP} @ 50$compared to the baseline YOLO11n model. These results underscore the robustness and effectiveness of P2AW-YOLO11n in overcoming the challenges associated with small object detection, paving the way for enhanced UAV applications in complex environments. Xilin Li, Weiming Wang 0002, Fu Lee Wang, Xue Xue, Mingqiang Wei |
CW | 2 |
| 2025 | Source-Free Active Domain Adaptation for Efficient Medical Video Polyp Segmentation
Hongqiu Wang, Weiming Wang 0002, Harry Qin, Qiong Wang 0001, Lei Zhu 0003 |
MICCAI (10) | 3 |
| 2025 | Adversarial neighbor perception network with feature distillation for anomaly detection
Yuting Su 0001, Enqi Su, Weiming Wang 0002, Peiguang Jing, Dubuke Ma, Fu Lee Wang |
Expert Syst. Appl. | 3 |
| 2025 | Lost in UNet: Improving Infrared Small Target Detection by Underappreciated Local FeaturesabstractInfrared small target detection (ISTD) is a challenging task due to the low contrast and small size of the targets, which are often affected by complex backgrounds. UNet and its variants, known for their encoder–decoder structures, are widely used in such tasks since they can capture both local and global features. However, a significant drawback of UNet-based networks is the irreversible loss of crucial local features during downsampling, leading to missed detections and false positives, especially for small targets. Compared to other architectures like feature pyramid networks, UNet provides a more symmetric and efficient structure, allowing it to handle dense pixel-wise predictions effectively. However, standard UNet models still struggle to fully retain small target details, motivating the need for further improvements. To address this issue, we propose HintU, a novel network to recover the local features lost by various UNet-based methods for effective ISTD. HintU has two key contributions. First, it introduces the “Hint” mechanism for the first time, i.e., leveraging the prior knowledge of target locations to highlight critical local features. Second, it improves the mainstream UNet-based architecture to preserve target pixels even after downsampling. HintU can shift the focus of various networks (e.g., vanilla UNet, UNet++, UIUNet, MiM+, and HCFNet) from the irrelevant background pixels to a more restricted area from the beginning. Experimental results on three datasets NUDT-SIRST, SIRSTv2, and IRSTD1K demonstrate that HintU enhances the performance of existing methods with only an additional 1.88-ms cost (on RTX Titan). Additionally, the explicit constraints of HintU enhance the generalization ability of UNet-based methods. Code is available athttps://github.com/Wuzhou-Quan/HintU. Wuzhou Quan, Wei Zhao 0039, Weiming Wang 0002, Haoran Xie 0001, Fu Lee Wang, Mingqiang Wei |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | AMDANet: Augmented Multiscale Difference Aggregation Network for Image Change DetectionabstractThe field of remote sensing image change detection (CD) has made significant improvements with the rapid development of deep learning techniques. However, current methods often inadequately utilize difference features of bitemporal images, resulting in biased focus and insensitivity to change information. Furthermore, the classic challenges of pseudo-CD and edge recognition in complex scenes have also weakened CD performance. In this article, we propose an augmented multiscale difference aggregation network (AMDANet) for image CD, which incorporates a difference feature extractor (DFE) within a Siamese feature extractor to perceive changes by capturing differences between bitemporal features. To address the issue of biased focus, we propose a hierarchical feature aggregator (HFA) that captures intrascale interactions in parallel for multigranularity change perception while personalizing coarse-grained and fine-grained features to highlight the attention to change regions. To deepen the perception of complex dependency relationships, we further design an O-shape feature augmentor (OFA) that leverages an information feedback loop to achieve precise alignment of multigranularity features. The integration of information across different granularities improves the recognition of pseudo-changes and edges. Experimental results on three publicly available datasets demonstrate the superiority of AMDANet over current state-of-the-art (SOTA) methods. Our source code will be publicly available athttps://github.com/mp-st/AMDANet. Yuting Su 0001, Peng Ma, Weiming Wang 0002, Shaochu Wang, Yun Li 0006, Peiguang Jing |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | Federated Pseudo Modality Generation for Incomplete Multi-Modal MRI ReconstructionabstractWhile multi-modal learning has been widely used for MRI reconstruction, it relies on paired multi-modal data, which is difficult to acquire in real clinical scenarios. Especially in the federated setting, there is a common issue that several medical institutions suffer from missing modalities or even only have single-modal data. Therefore, it is infeasible to deploy a standard federated learning framework in such conditions. In this paper, we propose a novel communication-efficient federated learning framework (namely Fed-PMG) to address the missing modality challenge in federated multi-modal MRI reconstruction. Specifically, we utilize a pseudo modality generation mechanism to recover the missing modality for each single-modal client by sharing the distribution information of the amplitude spectrum in frequency space. However, the step of sharing the original amplitude spectrum leads to heavy communication costs. To reduce the communication cost, we introduce a clustering scheme to project the set of amplitude spectrum into a finite number of cluster centroids and share them among the clients. With such an elaborate design, our approach can effectively complete the missing modality within an acceptable communication cost. Extensive experimental results demonstrate that our proposed method can outperform state-of-the-art methods and reach a performance similar to the ideal scenario (i.e., all clients have the full set of modalities). Yunlu Yan, Chun-Mei Feng 0001, Yuexiang Li, Ping Li 0016, Rick Siow Mong Goh, Bai Ying Lei, Weiming Wang 0002, David Dagan Feng, Lei Zhu 0003 |
IEEE J. Biomed. Health Informatics | 7 |
| 2025 | Multimodal Dual-Graph Collaborative Network With Serial Attentive Aggregation Mechanism for Micro-Video Multi-Label ClassificationabstractThe increasing commercial value of micro-videos has spurred a rising demand for grasping their contents. The abundant multimodal cues in micro-videos exhibit substantial potential in enhancing content comprehension. However, effectively harnessing the collaborative characteristics across different modalities remains a significant challenge, especially in multi-label scenarios due to inconsistent behaviors regarding label correlations. To better tackle this issue, in this paper, we first introduce a multimodal dual-graph collaborative network with serial attentive aggregation mechanism (MDGCN) for micro-video multi-label classification. In MDGCN, we exploit an asymmetric encoder-decoder framework, which incorporates multiple parallel encoders with complementary representations and a decoder to ensure the completeness of encoded results. Meanwhile, an adversarial constraint is used to ensure individual differences prominently featured within each modality. Furthermore, considering the inconsistency of label correlations across various modalities, we then construct a serial attentive graph convolutional network that employs an interactive dual-graph attention paradigm to sequentially integrate multimodal representations and dynamically explore label correlations. The experiments conducted on two datasets demonstrate that our proposed method outperforms state-of-the-art approaches. Wei Lu 0026, Peiguang Jing, Weiming Wang 0002, Yuting Su 0001 |
IEEE Trans. Multim. | 4 |
| 2025 | PADNet: Progressive-Difference-Aware Feature Reconstruction Mechanism for Anomaly Detection
Peiguang Jing, Weiming Wang 0002, Fu Lee Wang, Yuting Su 0001 |
IEEE Trans. Multim. | 3 |
| 2025 | HADiff: hierarchy aggregated diffusion model for pathology image segmentation
Zhaohu Xing, Feng Gao 0023, Yuandong Tao, Zhenyan Han, Weiming Wang 0002, Lei Zhu 0003 |
Vis. Comput. | 7 |
| 2024 | Shape Descriptor Guided Learning for Category-Level Object Pose Estimation
Yun Liu 0002, Weiming Wang 0002, Fu Lee Wang, Haoran Xie 0001, Honghua Chen, Mingqiang Wei, Harry Qin |
CGI (3) | 2 |
| 2024 | RainMamba: Enhanced Locality Learning with State Space Models for Video DerainingabstractThe outdoor vision systems are frequently contaminated by rain streaks and raindrops, which significantly degenerate the performance of visual tasks and multimedia applications. The nature of videos exhibits redundant temporal cues for rain removal with higher stability. Traditional video deraining methods heavily rely on optical flow estimation and kernel-based manners, which have a limited receptive field. Yet, transformer architectures, while enabling long-term dependencies, bring about a significant increase in computational complexity. Recently, the linear-complexity operator of the state space models (SSMs) has contrarily facilitated efficient long-term temporal modeling, which is crucial for rain streaks and raindrops removal in videos. Unexpectedly, its uni-dimensional sequential process on videos destroys the local correlations across the spatio-temporal dimension by distancing adjacent pixels. To address this, we present an improved SSMs-based video deraining network (RainMamba) with a novel Hilbert scanning mechanism to better capture sequence-level local information. We also introduce a difference-guided dynamic contrastive locality learning strategy to enhance the patch-level self-similarity learning ability of the proposed network. Extensive experiments on four synthesized video deraining datasets and real-world rainy videos demonstrate the superiority of our network in the removal of rain streaks and raindrops. Our code and results are available at https://github.com/TonyHongtaoWu/RainMamba. Weiming Wang 0002, Jinni Zhou, Lei Zhu 0003 |
ACM Multimedia | 4 |
| 2024 | PointeNet: A lightweight framework for effective and efficient point cloud analysis
Lipeng Gu, Xuefeng Yan 0001, Liangliang Nan, Dingkun Zhu, Honghua Chen, Weiming Wang 0002, Mingqiang Wei |
Comput. Aided Geom. Des. | 6 |
| 2024 | Real-World Image Deraining Using Model-Free Unsupervised LearningabstractWe propose a novel model‐free unsupervised learning paradigm to tackle the unfavorable prevailing problem of real‐world image deraining, dubbed MUL‐Derain. Beyond existing unsupervised deraining efforts, MUL‐Derain leverages a model‐free Multiscale Attentive Filtering (MSAF) to handle multiscale rain streaks. Therefore, formulation of any rain imaging is not necessary, and it requires neither iterative optimization nor progressive refinement operations. Meanwhile, MUL‐Derain can efficiently compute spatial coherence and global interactions by modeling long‐range dependencies, allowing MSAF to learn useful knowledge from a larger or even global rain region. Furthermore, we formulate a novel multiloss function to constrain MUL‐Derain to preserve both color and structure information from the rainy images. Extensive experiments on both synthetic and real‐world datasets demonstrate that our MUL‐Derain obtains state‐of‐the‐art performance over un/semisupervised methods and exhibits competitive advantages over the fully‐supervised ones. Rongwei Yu, Jingyi Xiang, Ni Shu, Peihao Zhang, Yizhan Li, Yiyang Shen, Weiming Wang 0002, Lina Wang 0001 |
Int. J. Intell. Syst. | 7 |
| 2024 | Point Transformer-Based Salient Object Detection Network for 3-D Measurement Point CloudsabstractWhile salient object detection (SOD) on 2D images has been extensively studied, there is very little SOD work on 3D measurement surfaces. We propose an effective point transformer-based SOD network for 3D measurement point clouds, termed PSOD-Net. PSOD-Net is an encoder-decoder network that takes full advantage of transformers to model the contextual information in both multi-scale point- and scene-wise manners. In the encoder, we develop a Point Context Transformer (PCT) module to capture region contextual features at the point level; PCT contains two different transformers to excavate the relationship among points. In the decoder, we develop a Scene Context Transformer (SCT) module to learn context representations at the scene level; SCT contains both Upsampling-and-Transformer blocks and Multi-context Aggregation units to integrate the global semantic and multi-level features from the encoder into the global scene context. Experiments show clear improvements of PSOD-Net over its competitors and validate that PSOD-Net is more robust to challenging cases such as small objects, multiple objects, and objects with complex structures. Code is available at: https://github.com/ZeyongWei/PSOD-Net. Zeyong Wei, Baian Chen, Weiming Wang 0002, Honghua Chen, Mingqiang Wei, Jonathan Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | RegiFormer: Unsupervised Point Cloud Registration via Geometric Local-to-Global Transformer and Self-AugmentationabstractRepresentation learning for two partially overlapping point clouds remains an open challenge in unsupervised point cloud registration (U-PCR). In this article, we introduce RegiFormer, a geometric local-to-global transformer (GLGT)-based unsupervised framework equipped with a self-augmentation (SA) strategy, for point cloud registration. The GLGT not only aggregates features from local neighborhoods but also extracts global intrarelationships within the entire point cloud using a transformation-invariant geometry embedding. In addition, it enhances the interrelationships between paired point clouds. To overcome the limited ability of U-PCR methods to learn alignment knowledge, we design an SA strategy that can be flexibly integrated into advanced models, significantly boosting their registration performance. Extensive experiments, conducted on five popular synthetic and real-scanned benchmarks, demonstrate the superior performance of RegiFormer compared to state-of-the-art methods, both qualitatively and quantitatively. Mengjiao Ma, Zhilei Chen, Honghua Chen, Weiming Wang 0002, Mingqiang Wei |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | Structure-preserving image smoothing via contrastive learning
Dingkun Zhu, Weiming Wang 0002, Xue Xue, Haoran Xie 0001, Gary Cheng 0001, Fu Lee Wang |
Vis. Comput. | 2 |
| 2023 | Video Adverse-Weather-Component Suppression Network via Weather Messenger and Adversarial BackpropagationabstractAlthough convolutional neural networks (CNNs) have been proposed to remove adverse weather conditions in single images using a single set of pre-trained weights, they fail to restore weather videos due to the absence of temporal information. Furthermore, existing methods for removing adverse weather conditions (e.g., rain, fog, and snow) from videos can only handle one type of adverse weather. In this work, we propose the first framework for restoring videos from all adverse weather conditions by developing a video adverse-weather-component suppression network (ViWS-Net). To achieve this, we first devise a weather-agnostic video transformer encoder with multiple transformer stages. Moreover, we design a long short-term temporal modeling mechanism for weather messenger to early fuse input adjacent video frames and learn weather-specific information. We further introduce a weather discriminator with gradient reversion, to maintain the weather-invariant common information and suppress the weather-specific information in pixel features, by adversarially predicting weather types. Finally, we develop a messenger-driven video transformer decoder to retrieve the residual weather-specific feature, which is spatiotemporally aggregated with hierarchical pixel features and refined to predict the clean target frame of input videos. Experimental results, on benchmark datasets and real-world weather videos, demonstrate that our ViWS-Net outperforms current state-of-the-art methods in terms of restoring videos degraded by any weather condition. Angelica I. Avilés-Rivero, Huazhu Fu, Weiming Wang 0002, Lei Zhu 0003 |
ICCV | 5 |
| 2023 | SVDFormer: Complementing Point Cloud via Self-view Augmentation and Self-structure Dual-generatorabstractIn this paper, we propose a novel network, SVDFormer, to tackle two specific challenges in point cloud completion: understanding faithful global shapes from incomplete point clouds and generating high-accuracy local structures. Current methods either perceive shape patterns using only 3D coordinates or import extra images with well-calibrated intrinsic parameters to guide the geometry estimation of the missing parts. However, these approaches do not always fully leverage the cross-modal self-structures available for accurate and high-quality point cloud completion. To this end, we first design a Self-view Fusion Network that leverages multiple-view depth image information to observe incomplete self-shape and generate a compact global shape. To reveal highly detailed structures, we then introduce a refinement module, called Self-structure Dual-generator, in which we incorporate learned shape priors and geometric self-similarities for producing new points. By perceiving the incompleteness of each point, the dual-path design disentangles refinement strategies conditioned on the structural type of each point. SVDFormer absorbs the wisdom of self-structures, avoiding any additional paired information such as color images with precisely calibrated camera intrinsic parameters. Comprehensive experiments indicate that our method achieves state-of-the-art performance on widely-used benchmarks. Code is available at https://github.com/czvvd/SVDFormer. Zhe Zhu, Honghua Chen, Weiming Wang 0002, Harry Qin, Mingqiang Wei |
ICCV | 4 |
| 2023 | Deep Texture-Aware Features for Camouflaged Object DetectionabstractCamouflaged object detection is a challenging task that aims to identify objects having similar texture to the surroundings. This paper presents to amplify the subtle texture difference between camouflaged objects and the background for camouflaged object detection by formulating multiple texture-aware refinement modules to learn the texture-aware features in a deep convolutional neural network. The texture-aware refinement module computes the biased co-variance matrices of feature responses to extract the texture information, adopts an affinity loss to learn a set of parameter maps that help to separate the texture between camouflaged objects and the background, and leverages a boundary-consistency loss to explore the structures of object details. We evaluate our network on the benchmark datasets for camouflaged object detection both qualitatively and quantitatively. Experimental results show that our approach outperforms various state-of-the-art methods by a large margin. Xiaowei Hu 0001, Lei Zhu 0003, Xuemiao Xu, Yangyang Xu 0003, Weiming Wang 0002, Zijun Deng, Pheng-Ann Heng |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2023 | Contrastive Learning Models for Sentence RepresentationsabstractSentence representation learning is a crucial task in natural language processing, as the quality of learned representations directly influences downstream tasks, such as sentence classification and sentiment analysis. Transformer-based pretrained language models such as bidirectional encoder representations from transformers (BERT) have been extensively applied to various natural language processing tasks, and have exhibited moderately good performance. However, the anisotropy of the learned embedding space prevents BERT sentence embeddings from achieving good results in the semantic textual similarity tasks. It has been shown that contrastive learning can alleviate the anisotropy problem and significantly improve sentence representation performance. Therefore, there has been a surge in the development of models that utilize contrastive learning to fine-tune BERT-like pretrained language models to learn sentence representations. But no systematic review of contrastive learning models for sentence representations has been conducted. To fill this gap, this article summarizes and categorizes the contrastive learning based sentence representation models, common evaluation tasks for assessing the quality of learned representations, and future research directions. Furthermore, we select several representative models for exhaustive experiments to illustrate the quantitative improvement of various strategies on sentence representations. Haoran Xie 0001, Zongxi Li, Fu Lee Wang, Weiming Wang 0002, Qing Li 0001 |
ACM Trans. Intell. Syst. Technol. | 5 |
| 2023 | CF-YOLO: Cross Fusion YOLO for Object Detection in Adverse Weather With a High-Quality Real Snow DatasetabstractSnow is one of the toughest adverse weather conditions for object detection (OD). Currently, not only there is a lack of snowy OD datasets to train cutting-edge detectors, but also these detectors have difficulties of learning latent information beneficial for detection in snow. To alleviate the two above problems, we first establish a real-world snowy OD dataset, named RSOD. Besides, we develop an unsupervised training strategy with a distinctive activation function, called$Peak Act$, to quantitatively evaluate the effect of snow on each object. Peak Act helps grade the images in RSOD into four-difficulty levels. To our knowledge, RSOD is the first quantitatively evaluated and graded real-world snowy OD dataset. Then, we propose a novel Cross Fusion (CF) block to construct a lightweight OD network based on YOLOv5s (called CF-YOLO). CF is a plug-and-play feature aggregation module, which integrates the advantages of Feature Pyramid Network and Path Aggregation Network in a simpler yet more flexible form. Both RSOD and CF lead our CF-YOLO to possess an optimization ability for OD in real-world snow. That is, CF-YOLO can handle unfavorable detection problems of vagueness, distortion and covering of snow. Experiments show that our CF-YOLO achieves better detection results on RSOD, compared to SOTAs. The code and dataset are available athttps://github.com/qqding77/CF-YOLO-and-RSOD. Qiqi Ding, Peng Li 0064, Xuefeng Yan 0001, Ding Shi, Luming Liang, Weiming Wang 0002, Haoran Xie 0001, Jonathan Li 0001, Mingqiang Wei |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2023 | S $^3$ Net: Self-Supervised Self-Ensembling Network for Semi-Supervised RGB-D Salient Object DetectionabstractRGB-D salient object detection aims to detect visually distinctive objects or regions from a pair of the RGB image and the depth image. State-of-the-art RGB-D saliency detectors are mainly based on convolutional neural networks but almost suffer from an intrinsic limitation relying on the labeled data, thus degrading detection accuracy in complex cases. In this work, we present a self-supervised self-ensembling network (S$^3$Net) for semi-supervised RGB-D salient object detection by leveraging the unlabeled data and exploring a self-supervised learning mechanism. To be specific, we first build a self-guided convolutional neural network (SG-CNN) as a baseline model by developing a series of three-layer cross-model feature fusion (TCF) modules to leverage complementary information among depth and RGB modalities and formulating an auxiliary task that predicts a self-supervised image rotation angle. After that, to further explore the knowledge from unlabeled data, we assign SG-CNN to a student network and a teacher network, and encourage the saliency predictions and self-supervised rotation predictions from these two networks to be consistent on the unlabeled data. Experimental results on seven widely-used benchmark datasets demonstrate that our network quantitatively and qualitatively outperforms the state-of-the-art methods. Lei Zhu 0003, Xiaoqiang Wang 0007, Ping Li 0016, Xin Yang 0011, Qing Zhang 0006, Weiming Wang 0002, Carola-Bibiane Schönlieb, C. L. Philip Chen |
IEEE Trans. Multim. | 6 |
| 2022 | MBA-RainGAN: A Multi-Branch Attention Generative Adversarial Network for Mixture of Rain RemovalabstractRain severely degrades the visibility of scene objects, especially when images are captured through the glass under rainy weather. We observe three intriguing phenomena: 1) rain is a mixture of raindrops, rain streaks and rainy haze; 2) the depth from the camera determines the degree of object visibility, where objects nearby and far away are visually blocked by rain streaks and rainy haze, respectively; and 3) raindrops on the glass randomly affect the object visibility of the whole image space. However, existing solutions and benchmark datasets lack full consideration of the mixture of rain (MOR). In this paper, we originally consider that the overall object visibility is determined by MOR, and enrich the RainCityscapes by considering real-world raindrops to construct the MOR dataset, named RainCityscapes++. To solve the practical rain removal problem arisen from MOR, we formulate a new rain imaging model and propose a multi-branch attention generative adversarial network (MBA-RainGAN). Extensive experiments show clear improvements of our approach over SOTAs on RainCityscapes++. Yiyang Shen, Yidan Feng, Weiming Wang 0002, Dong Liang 0008, Harry Qin, Haoran Xie 0001, Mingqiang Wei |
ICASSP | 3 |
| 2022 | SPCNet: Stepwise Point Cloud Completion NetworkabstractAbstract How will you repair a physical object with large missings? You may first recover its global yet coarse shape and stepwise increase its local details. We are motivated to imitate the above physical repair procedure to address the point cloud completion task. We propose a novel stepwise point cloud completion network (SPCNet) for various 3D models with large missings. SPCNet has a hierarchical bottom‐to‐up network architecture. It fulfills shape completion in an iterative manner, which 1) first infers the global feature of the coarse result; 2) then infers the local feature with the aid of global feature; and 3) finally infers the detailed result with the help of local feature and coarse result. Beyond the wisdom of simulating the physical repair, we newly design a cycle loss to enhance the generalization and robustness of SPCNet. Extensive experiments clearly show the superiority of our SPCNet over the state‐of‐the‐art methods on 3D point clouds with large missings. Code is available at https://github.com/1127368546/SPCNet . Honghua Chen, Xuequan Lu, Zhe Zhu, Jun Wang 0039, Weiming Wang 0002, Fu Lee Wang, Mingqiang Wei |
Comput. Graph. Forum | 6 |
| 2022 | SO(3)-Pose: SO(3)-Equivariance Learning for 6D Object Pose EstimationabstractAbstract 6D pose estimation of rigid objects from RGB‐D images is crucial for object grasping and manipulation in robotics. Although RGB channels and the depth (D) channel are often complementary, providing respectively the appearance and geometry information, it is still non‐trivial on how to fully benefit from the two cross‐modal data. From the simple yet new observation, when an object rotates, its semantic label is invariant to the pose while its keypoint offset direction is variant to the pose. To this end, we present SO(3)‐Pose, a new representation learning network to explore SO(3)‐equivariant and SO(3)‐invariant features from the depth channel for pose estimation. The SO(3)‐invariant features facilitate to learn more distinctive representations for segmenting objects with similar appearance from RGB channels. The SO(3)‐equivariant features communicate with RGB features to deduce the (missed) geometry for detecting keypoints of an object with the reflective surface from the depth channel. Unlike most of existing pose estimation methods, our SO(3)‐Pose not only implements the information communication between the RGB and depth channels, but also naturally absorbs the SO(3)‐equivariance geometry knowledge from depth images, leading to better appearance and geometry representation learning. Comprehensive experiments show that our method achieves the state‐of‐the‐art performance on three benchmarks. Code is available at https://github.com/phaoran9999/SO3-Pose . Haoran Pan, Jun Zhou 0007, Xuequan Lu, Weiming Wang 0002, Xuefeng Yan 0001, Mingqiang Wei |
Comput. Graph. Forum | 5 |
| 2022 | FDDL-Net: frequency domain decomposition learning for speckle reduction in ultrasound images
Tongda Yang, Weiming Wang 0002, Gary Cheng 0001, Mingqiang Wei, Haoran Xie 0001, Fu Lee Wang |
Multim. Tools Appl. | 2 |
| 2022 | HDRD-Net: High-resolution detail-recovering image deraining network
Dingkun Zhu, Weiming Wang 0002, Gary Cheng 0001, Mingqiang Wei, Fu Lee Wang, Haoran Xie 0001 |
Multim. Tools Appl. | 3 |
| 2018 | Feature-preserving ultrasound speckle reduction via L0 minimization
Lei Zhu 0003, Weiming Wang 0002, Xiaomeng Li 0001, Qiong Wang 0001, Harry Qin, Kin Hong Wong, Kup-Sze Choi, Chi-Wing Fu, Pheng-Ann Heng |
Neurocomputing | 2 |
| 2017 | Fast feature-preserving speckle reduction for ultrasound images via phase congruency
Lei Zhu 0003, Weiming Wang 0002, Harry Qin, Kin Hong Wong, Kup-Sze Choi, Pheng-Ann Heng |
Signal Process. | 2 |
| 2016 | Ultrasound Speckle Reduction via L_0 Minimization
Lei Zhu 0003, Weiming Wang 0002, Xiaomeng Li 0001, Qiong Wang 0001, Harry Qin, Kin Hong Wong, Pheng-Ann Heng |
ACCV (3) | 2 |
| 2013 | Coarse-to-Fine Normal Filtering for Feature-Preserving Mesh Denoising Based on Isotropic SubneighborhoodsabstractAbstract State‐of‐theart normal filters usually denoise each face normal using its entire anisotropic neighborhood. However, enforcing these filters indiscriminately on the anisotropic neighborhood will lead to feature blurring, especially in challenging regions with shallow features. We develop a novel mesh denoising framework which can effectively preserve features with various sizes. Our idea is inspired by the observation that the underlying surface of a noisy mesh is piecewise smooth. In this regard, it is more desirable that we denoise each face normal within its piecewise smooth region (we call such a region as an isotropic subneighborhood) instead of using the anisotropic neighborhood. To achieve this, we first classify mesh faces into several types using a face normal tensor voting and then perform a normal filter to obtain a denoised coarse normal field. Based on the results of normal classification and the denoised coarse normal field, we segment the anisotropic neighborhood of every feature face into a number of isotropic subneighborhoods via local spectral clustering. Thus face normal filtering can be performed again on the isotropic subneighborhoods and produce a more accurate normal field. Extensive tests on various models demonstrate that our method can achieve better performance than state‐of‐theart normal filters, especially in challenging regions with features. Lei Zhu 0003, Mingqiang Wei, Jinze Yu 0001, Weiming Wang 0002, Harry Qin, Pheng-Ann Heng |
Comput. Graph. Forum | 4 |
| 2009 | A Physically-Based Modeling and Simulation Framework for Facial AnimationabstractRealistic facial animation is important in many graphics applications, like animated feature films and computer games, to enrich human computer interaction. In this paper, we propose a physically-based facial animation approach employing knowledge from the anatomy and biomechanics of human facial muscles. First, the 3D face mesh is generated automatically by commercial software and the facial skin is represented by a nonlinear mass-spring system which simulates realistic elastic dynamics of human dermis. Then, a structure skull is attached to fit the face mesh. A set of anatomically consistent facial muscles are incorporated to model the forces deforming the face mesh. Finally, we extend Waters' muscle model to improve the combination of multiple muscle actions and to generate realistic expression. Experiments show that our method is superior to the traditional geometric model and can achieve comparative results with the commercial software FaceGen. Weiming Wang 0002, Xiaoqi Yan, Yongming Xie, Harry Qin, Wai-Man Pang, Pheng-Ann Heng |
ICIG | 1 |