EDBT 2026 Demo / reviewers in the wild / expert
Weimin Wang 0007
dblp:20/3553-7
· DBLP profile ↗
37ranked-venue papers
5as first author
32since 2021 · last 2026
0000-0001-6557-7175ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 26 · 3 first-author · 25 since 2021Artificial intelligence and machine learning · 9 · 1 first-author · 8 since 2021Systems, architecture and hardware · 5 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Learning 3D Occupancy from Beam Overlap in 2D Rotating mmWave RadarabstractRobust 3D perception under adverse weather is critical for autonomous systems. While mmWave Radars are inherently weather-resistant, conventional 2D rotating Radar sensors lack direct elevation resolution, limiting their 3D perception ability. Although 4D imaging radars can provide elevation information, they typically suffer from limited coverage and range. In this work, we exploit a key observation about mechanically rotating 2D mmWave Radars: in each sweep, an overlap exists between adjacent azimuth beam coverage due to the width of the main lobe, which makes the reflected intensity difference imply object materials and geometric shapes, including elevation. With this observation, we propose a method that learns 3D occupancy by disentangling bird’s-eye view (BEV) layout and elevation estimation from one frame Radar scan. Specifically, we partition one sweep into two interleaved subsets, corresponding to overlapping beam directions, and utilize them to infer coarse geometric structure through spatial differences and intensity patterns. Extensive quantitative and qualitative evaluations on two real-world datasets demonstrate that our proposed method outperforms existing baselines. The codes will be publicly available. Ruifeng Nie, Long Ma 0002, Chengpei Xu, Yu Liu 0012, Weimin Wang 0007 |
AAAI | 6 |
| 2026 | HiEn: Hierarchical ensemble learning for semi-supervised medical image segmentation
Long Ma 0002, Xinwei Xue, Chengpei Xu, Weimin Wang 0007, Yi Wang 0037 |
Neurocomputing | 6 |
| 2026 | SWG-Fusion: Soft weather-guided multimodal fusion with VLM-assistance for BEV object detection under harsh weather
Weimin Wang 0007, Ruifeng Nie, Yingchi Liu, Long Ma 0002, Chengpei Xu, Qi Jia 0001, Yu Liu 0012, Na Lei |
Pattern Recognit. | 1 |
| 2026 | Model-aware ellipse detection via parametric correlation learningabstractEllipse detection presents a significant challenge in computer vision and pattern recognition, often hindered by traditional parameter regression methods that fail to account for the unique geometric characteristics and complex parameter interactions of ellipses. These limitations frequently result in imprecise detections, notably with small or partially occluded ellipses. To overcome these challenges, we propose EDNet, a novel ellipse detection network that exploits the geometric properties of ellipses, thus moving beyond the reliance on internal textures. EDNet improves ellipse detection by refining the loss function to better capture the relationship between the error and each parameters during training. It features a LoG-like Edge Detection Module (LEDM) and an Edge Guided Module (EGM) for precise boundary extraction and multi-scale feature enhancement. Additionally, an auxiliary component estimates ellipse vertices, boosting accuracy for occluded ellipses. Experimental results on two wildly-used benchmark datasets demonstrate that EDNet achieves significant improvements, with an average detection accuracy increase of 6% and 10% over leading state-of-the-art models. • We concentrate on the geometric characteristics of ellipse detection via Edge Detection Module and Edge Guided Module. • We design an auxiliary head for the estimation of four ellipse vertices, invoking additional feature attention on these pivotal points. • We establish the relations between the error and geometric characteristics of the ellipse by a model-aware loss function. Qi Jia 0001, Zezheng Liu, Yu Liu 0012, Yi Wang 0037, Xinwei Xue, Weimin Wang 0007 |
Signal Process. | 6 |
| 2026 | Enhancing Underwater Images via Resonant FusionabstractRecent advances in learning-based underwater image enhancement have achieved remarkable progress. However, the inherent diversity and complexity of underwater scenes still limit the ability of existing approaches to simultaneously restore fine structural details and global image layouts. To address this challenge, we propose a Resonant Fusion (ReFu) framework that explicitly leverages complementary information in both spatial and frequency domains. Specifically, we design a frequency decomposer and a spatial decomposer to capture high- and low-frequency cues from different perspectives. A resonant fuser is then introduced to adaptively integrate high-frequency resonances for detail refinement and low-frequency resonances for structural consistency. This fine-grained cross-domain fusion significantly improves structural preservation and detail enhancement, thereby generating visually more natural and perceptually friendly underwater images. Extensive quantitative and qualitative evaluations across diverse underwater benchmarks show that ReFu consistently surpasses state-of-the-art methods by a clear margin. Comprehensive ablation studies further validate the effectiveness of each module and prove the necessity of the proposed ReFu mechanism. Our code is available at https://github.com/CircleQa/ReFu-main. Xinwei Xue, Zimeng Xu, Jincheng Yuan, Jingchun Zhou, Chengpei Xu, Xiaoke Shang, Long Ma 0002, Weimin Wang 0007 |
IEEE Trans. Image Process. | 9 |
| 2025 | CoA: Towards Real Image Dehazing via Compression-and-AdaptationabstractLearning-based image dehazing algorithms have shown remarkable success in synthetic domains. However, real image dehazing is still in suspense due to computational resource constraints and the diversity of real-world scenes. Therefore, there is an urgent need for an algorithm that excels in both efficiency and adaptability to address real image dehazing effectively. This work proposes a Compression-and-Adaptation (CoA) computational flow to tackle these challenges from a divide-and-conquer perspective. First, model compression is performed in the synthetic domain to develop a compact dehazing parameter space, satisfying efficiency demands. Then, a bilevel adaptation in the real domain is introduced to be fearless in unknown real environments by aggregating the synthetic dehazing capabilities during the learning process. Leveraging a succinct design free from additional constraints, our CoA exhibits domain-irrelevant stability and model-agnostic flexibility, effectively bridging the model chasm between synthetic and real domains to further improve its practical utility. Extensive evaluations and analyses underscore the approach's superiority and effectiveness. The code is publicly available at https://github.com/fyxnl/COA. Long Ma 0002, Yan Zhang 0002, Jinyuan Liu 0001, Weimin Wang 0007, Guang-Yong Chen, Chengpei Xu, Zhuo Su 0001 |
CVPR | 5 |
| 2025 | NoPain: No-box Point Cloud Attack via Optimal Transport Singular BoundaryabstractAdversarial attacks exploit the vulnerability of deep models against adversarial samples. Existing point cloud attackers are tailored to specific models, iteratively optimizing perturbations based on gradients in either a white-box or black-box setting. Despite their promising attack performance, they often struggle to produce transferable adversarial samples due to overfitting to the specific parameters of surrogate models. To overcome this issue, we shift our focus to the data distribution itself and introduce a novel approach named NoPain, which employs optimal transport (OT) to identify the inherent singular boundaries of the data manifold for cross-network point cloud attacks. Specifically, we first calculate the OT mapping from noise to the target feature space, then identify singular boundaries by locating non-differentiable positions. Finally, we sample along singular boundaries to generate adversarial point clouds. Once the singular boundaries are determined, NoPain can efficiently produce adversarial samples without the need of iterative updates or guidance from the surrogate classifiers. Extensive experiments demonstrate that the proposed end-to-end method outperforms baseline approaches in terms of both transferability and efficiency, while also maintaining notable advantages even against defense strategies. Code and model are available at https://github.com/cognaclee/nopain. Zezeng Li, Na Lei, Liming Chen 0002, Weimin Wang 0007 |
CVPR | 5 |
| 2025 | Bright to Dark: Stage-wise Bilevel Knowledge Transfer for Seeing Text in the DarkabstractLocalizing text under low-light conditions has gained attention, with typical approaches relying on two stage cascading modules that combine low-light enhancement and text localization. However, these often require additional enhancement modules and cause inefficiency in joint optimization. In this work, we address the challenge by adopting a novel approach: tailoring the detector for low light conditions through knowledge distillation from normal light conditions, without relying on any enhancement module. First, we design a Graph Topological Aggregation (GTA) model that utilizes the message passing mechanism of graph neural networks to structurally represent text topology and facilitate structured feature expression in knowledge transfer. We then introduce two specially designed knowledge transfer constraints aimed at enhancing the learning of text's multi-scale features and topological knowledge. Finally,we propose a Stage-wise Bilevel Knowledge Transfer learning strategy that designates the low-light learning process as the upper-level task, while treating normal light learning as the lower-level task, effectively addressing the coupling issues and sequential dependencies prevalent during the distillation process. Extensive experiments underscore the approach's superiority. Chengpei Xu, Long Ma 0002, Weimin Wang 0007, Feng Xia 0001, Binghao Li, Wenjie Zhang 0001 |
ACM Multimedia | 4 |
| 2025 | 3D-NLM: Voxel-based non-local means for 3D point cloud noise detection and smoothing
Weimin Wang 0007, Yu Liu 0012, Qiong Chang |
Comput. Graph. | 1 |
| 2025 | Accelerating Nearest Neighbor Search in 3D Point Cloud Registration on GPUsabstractThe Iterative Closest Points (ICP) algorithm is the most widely used method for estimating rigid transformation in 3D point cloud registration. However, the ICP relies on repeatedly performing computationally intensive nearest neighbor searches (NNS) within 3D space. This dependency becomes a significant bottleneck when processing large datasets, thereby hindering the practical deployment of point cloud technologies in real-world applications. To address this issue, we propose two approximate nearest neighbor search (ANNS) acceleration strategies for efficient improvement of the processing speed of the NNS. Our strategies first voxelize target cloud points and then fill voxels in the 3D coordinate space around the source point cloud in two different ways, which can convert the global nearest neighbor search to a local search. Both the proposed methods are suited to be parallelized on GPUs with a low computational load. Extensive experiments show that our methods significantly accelerate NNS processing while maintaining high accuracy, outperforming most of the currently known approaches. Qiong Chang, Weimin Wang 0007, Jun Miyazaki |
ACM Trans. Archit. Code Optim. | 2 |
| 2025 | Point2Quad: Generating Quad Meshes From Point Clouds via Face PredictionabstractQuad meshes are essential in geometric modeling and computational mechanics. Although learning-based methods for triangle mesh demonstrate considerable advancements, quad mesh generation remains less explored due to the challenge of ensuring coplanarity, convexity, and quad-only meshes. In this paper, we presentPoint2Quad, the first learning-based method for quad-only mesh generation from point clouds. The key idea is learning to identify quad mesh with fused pointwise and facewise features. Specifically, Point2Quad begins with a k-NN-based candidate generation considering the coplanarity and squareness. Then, two encoders are followed to extract geometric and topological features that address the challenge of quad-related constraints, especially by combining in-depth quadrilaterals-specific characteristics. Subsequently, the extracted features are fused to train the classifier with a designed compound loss. The final results are derived after the refinement by a quad-specific post-processing. Extensive experiments on both clear and noise data demonstrate the effectiveness and superiority of Point2Quad, compared to baseline methods under comprehensive metrics. The code and dataset are available athttps://github.com/cognaclee/Point2Quad. Zezeng Li, Zhihui Qi, Weimin Wang 0007, Junyi Duan, Na Lei |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Faster than Fast: Accelerating Oriented FAST Feature Detection on Low-end Embedded GPUsabstractThe visual-based SLAM (Simultaneous Localization and Mapping) is a technology widely used in applications such as robotic navigation and virtual reality, which primarily focuses on detecting feature points from visual images to construct an unknown environmental map and simultaneously determines its own location. It usually imposes stringent requirements on hardware power consumption, processing speed, and accuracy. Currently, the ORB (Oriented FAST and Rotated BRIEF)-based SLAM systems have exhibited superior performance in terms of processing speed and robustness. However, they still fall short of meeting the demands for real-time processing on mobile platforms. This limitation is primarily due to the time-consuming Oriented FAST calculations accounting for approximately half of the entire SLAM system. This article presents two methods to accelerate the Oriented FAST feature detection on low-end embedded GPUs. These methods optimize the most time-consuming steps in Oriented FAST feature detection: FAST feature point detection and Harris corner detection, which is achieved by implementing a binary-level encoding strategy to determine candidate points quickly and a separable Harris detection strategy with efficient low-level GPU hardware-specific instructions. Extensive experiments on a Jetson TX2 embedded GPU demonstrate an average speedup of over 7.3 times compared to widely used OpenCV with GPU support. This significant improvement highlights its effectiveness and potential for real-time applications in mobile and resource-constrained environments. Qiong Chang, Xiang Li 0110, Weimin Wang 0007, Jun Miyazaki |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2025 | Rectangling for Stitched Image via Pixel-Wise Deformation LearningabstractImage rectangling involves filling in the blanks created during image stitching through deformation techniques. However, existing methods still struggle with incomplete filling and distortion of content, ultimately affecting the overall visual impression and potentially hindering subsequent tasks such as recognition. In this work, we design a pixel-wise deformation framework that utilizes explicit edge guidance to maintain consistency of texture and structure, yielding rectangular images with natural structure. Specifically, we decouple motion into region-level and pixel-level components through uniform mesh warping and pixel-wise deformation to precisely rearrange the spatial distribution of all pixels. Uniform deformation preserves local structure within divided patches, while pixel-wise motion coordinates the consistency between patches. Their combination provides robust and accurate pixel-wise offsets for structure-preserved rectangling. To further bolster the consistency of structure and texture, we leverage edge information to establish structural constraints and design an edge-guided enhancement module to aid in restoring fine texture details. Additionally, stitched images encompass both meaningful content and blank spaces, we innovatively incorporate a mask predictor, which acts as a guiding beacon, directing the network's attention solely towards content-rich regions to facilitate precise pixel-wise motion estimation. Experimental results demonstrate that our approach achieves state-of-the-art performance in rectifying irregular boundaries while contributing to downstream visual perception tasks. Xiaomei Feng, Qi Jia 0001, Yu Liu 0012, Weimin Wang 0007, Yuqing Liu 0001, Xinwei Xue |
IEEE Trans. Multim. | 4 |
| 2025 | Hyper-Spherical Optimal Transport for Semantic Alignment in Text-to-3D End-to-End GenerationabstractRecent CLIP-guided 3D generation methods have achieved promising results but struggle with generating faithful 3D shapes that conform with input text due to the gap between text and image embeddings. To this end, this paper proposes HOTS3D which makes the first attempt to effectively bridge this gap by aligning text features to the image features with spherical optimal transport (SOT). However, in high-dimensional situations, solving the SOT remains a challenge. To obtain the SOT map for high-dimensional features obtained from CLIP encoding of two modalities, we mathematically formulate and derive the solution based on Villani's theorem, which can directly align two hyper-sphere distributions without manifold exponential maps. Furthermore, we implement it by leveraging input convex neural networks (ICNNs) for the optimal Kantorovich potential. With the optimally mapped features, a diffusion-based generator is utilized to decode them into 3D shapes. Extensive quantitative and qualitative comparisons with state-of-the-art methods demonstrate the superiority of HOTS3D for text-to-3D generation, especially in the consistency with text semantics. Zezeng Li, Weimin Wang 0007, WenHai Li, Na Lei, Xianfeng Gu |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2024 | Novel Class Discovery for Ultra-Fine-Grained Visual CategorizationabstractUltra-fine-grained visual categorization (Ultra-FGVC) aims at distinguishing highly similar sub-categories within fine-grained objects, such as different soybean cultivars. Compared to traditional fine-grained visual categorization, Ultra-FGVC encounters more hurdles due to the small inter-class and large intra-class variation. Given these challenges, relying on human annotation for Ultra-FGVC is impractical. To this end, our work introduces a novel task termed Ultra-Fine-Grained Novel Class Discovery (UFG-NCD), which leverages partially annotated data to identify new categories of unlabeled images for Ultra-FGVC. To tackle this problem, we devise a Region-Aligned Proxy Learning (RAPL) framework, which comprises a Channel-wise Region Alignment (CRA) module and a Semi-Supervised Proxy Learning (SemiPL) strategy. The CRA module is designed to extract and utilize discriminative features from local regions, facilitating knowledge transfer from labeled to unlabeled classes. Furthermore, SemiPL strengthens representation learning and knowledge transfer with proxy-guided supervised learning and proxy-guided contrastive learning. Such techniques leverage class distribution information in the embedding space, improving the mining of subtle differences between labeled and unlabeled ultra-fine-grained classes. Extensive experiments demonstrate that RAPL significantly outperforms baselines across various datasets, indicating its effectiveness in handling the challenges of UFG-NCD. Code is available at https://github.com/SSDUT-Caiyq/UFG-NCD. Yu Liu 0012, Yaqi Cai, Qi Jia 0001, Binglin Qiu, Weimin Wang 0007, Nan Pu |
CVPR | 5 |
| 2024 | Attribution-Based Scanline Perturbation Attack on 3d Detectors of Lidar Point CloudsabstractLiDAR point cloud data is widely utilized in autonomous driving systems and has significantly improved the 3D detection performance with well-designed deep neural network models. However, due to the complexity of real-world environments and model vulnerability, false detections or malicious attacks may cause severe accidents in unseen situations. In this paper, we propose a novel attack approach, Attribution-based Scanline Perturbation (ASP), an efficient and physically possible adversarial attack method for 3D detectors. ASP first utilizes attribution methods to identify critical points for the detection model and perturbs them along the laser beams by simulating the situation in which particles exist between the LiDAR sensor and objects, which can actually occur in snow or sandstorm weather. Extensive experiments on practical 3D detectors validate the effectiveness of our approach in misleading the model and causing both false and missed detections. Ziyang Yu 0004, Qiong Chang, Yu Liu 0012, Weimin Wang 0007 |
ICASSP | 5 |
| 2024 | Sketch-Based 3D Shape Retrieval With Multi-View Fusion TransformerabstractSketch-based 3D shape retrieval aims to retrieve similar 3D shapes given a 2D sketch query. Although this task has been studied for years, the inherent cross-modal gap and data imbalance between 2D sketches and 3D shapes remain challenging. To address the problems, we propose a simple and effective framework based on Multi-view Fusion Transformer. To be specific, we project 3D shapes into twelve distinct views, and their CNN features are combined with position embeddings, passing together into a transformer encoder to learn view weights. Then we process them through average pooling and MLPs to obtain the final 3D shape representation. Furthermore, to narrow the data imbalance between 2D sketches and 3D shapes, affine transformation and elastic deformation are fully utilized for sketch augmentation, so as to extract more comprehensive sketch features for feature matching with the multi-view 3D shape representation. Extensive experiments on SHREC13, SHREC14 and PART-SHREC14 datasets demonstrate our method achieves superior performance than previous competitive methods. Cunjuan Zhu, Dongdong Cui, Qi Jia 0001, Weimin Wang 0007, Yu Liu 0012, Michael S. Lew |
ICASSP | 4 |
| 2024 | Point Cloud Compression via Constrained Optimal TransportabstractThis paper presents a novel point cloud compression method COT-PCC by formulating the task as a constrained optimal transport (COT) problem. COT-PCC takes the bitrate of compressed features as an extra constraint of optimal transport (OT) which learns the distribution transformation between original and reconstructed points. Specifically, the formulated COT is implemented with a generative adversarial network (GAN) and a bitrate loss for training. The discriminator measures the Wasserstein distance between input and reconstructed points, and a generator calculates the optimal mapping between distributions of input and reconstructed point cloud. Moreover, we introduce a learnable sampling module for downsampling in the compression procedure. Extensive results on both sparse and dense point cloud datasets demonstrate that COT-PCC outperforms state-of-the-art methods in terms of both CD and PSNR metrics. Source codes are available at https://github.com/cognaclee/PCC-COT. Zezeng Li, Weimin Wang 0007, Na Lei |
ICME | 2 |
| 2024 | MergeNet: Explicit Mesh Reconstruction from Sparse Point Clouds via Edge PredictionabstractThis paper introduces a novel method for reconstructing meshes from sparse point clouds by predicting edge connection. Existing implicit methods usually produce superior smooth and watertight meshes due to the isosurface extraction algorithms (e.g., Marching Cubes). However, these methods become memory and computationally intensive with increasing resolution. Explicit methods are more efficient by directly forming the face from points. Nevertheless, the challenge of selecting appropriate faces from enormous candidates often leads to undesirable faces and holes. Moreover, the reconstruction performance of both approaches tends to degrade when the point cloud gets sparse. To this end, we propose MEsh Reconstruction via edGE (MergeNet), which converts mesh reconstruction into local connectivity prediction problems. Specifically, MergeNet learns to extract the features of candidate edges and regress their distances to the underlying surface. Consequently, the predicted distance is utilized to filter out edges that lay on surfaces. Finally, the meshes are reconstructed by refining the triangulations formed by these edges. Extensive experiments on synthetic and real-scanned datasets demonstrate the superiority of MergeNet to SoTA explicit methods. Weimin Wang 0007, Yingxu Deng, Zezeng Li, Yu Liu 0012, Na Lei |
ICME | 1 |
| 2024 | Unseen No More: Unlocking the Potential of CLIP for Generative Zero-shot HOI DetectionabstractZero-shot human-object interaction (HOI) detector is capable of generalizing to HOI categories even not encountered during training. Inspired by the impressive zero-shot capabilities offered by CLIP, latest methods strive to leverage CLIP embeddings for improving zero-shot HOI detection. However, these embedding-based methods train the classifier on seen classes only, inevitably resulting in seen-unseen confusion for the model during inference. Besides, we find that using prompt-tuning and adapters further increases the gap between seen and unseen accuracy. To tackle this challenge, we present the first generation-based model using CLIP for zero-shot HOI detection, coined HOIGen. It allows to unlock the potential of CLIP for feature generation instead of feature extraction only. To achieve it, we develop a CLIP-injected feature generator in accordance with the generation of human, object and union features. Then, we extract realistic features of seen samples and mix them with synthetic features together, allowing the model to train seen and unseen classes jointly. To enrich the HOI scores, we construct a generative prototype bank in a pairwise HOI recognition branch, and a multi-knowledge prototype bank in an image-wise HOI recognition branch, respectively. Extensive experiments on HICO-DET benchmark demonstrate our HOIGen achieves superior performance for both seen and unseen classes under various zero-shot settings, compared with other top-performing methods. Code is available at: https://github.com/soberguo/HOIGen Yu Liu 0012, Weimin Wang 0007, Qi Jia 0001 |
ACM Multimedia | 4 |
| 2024 | PMGNet: Disentanglement and entanglement benefit mutually for compositional zero-shot learning
Yu Liu 0012, Yanyi Zhang, Qi Jia 0001, Weimin Wang 0007, Nan Pu, Nicu Sebe |
Comput. Vis. Image Underst. | 5 |
| 2023 | VAN-ICP: GPU-Accelerated Approximate Nearest Neighbor Search for ICP Registration via Voxel DilationabstractThe Iterative Closest Points (ICP) algorithm and its variants have been widely applied for 3D point cloud registration which estimates the rigid transformation. As the most computationally intensive step in ICP, nearest neighbor search (NNS) takes up most of the execution time, hindering the practical applications of ICP registration. To overcome the bottleneck, we propose a novel GPU-friendly approximate nearest neighbor search (ANNS) acceleration scheme, named Voxel dilAtioN (VAN), which can efficiently convert the global search to local $\left. {\mathcal{O}(n)} \right)$. Extensive experiments demonstrate that our VAN can drastically boost the NNS processing while keeping high registration accuracy. Specifically, our GPU-based VAN-ICP achieves 2.7x, 7.6x, and 13.4x speedup on three datasets compared with the CPU-based ICP implementation of Point Cloud Library (PCL). Source codes are available at https://github.com/mfxox/VAN-ICP. Weimin Wang 0007, Qiong Chang |
ICASSP | 1 |
| 2023 | MSECNet: Accurate and Robust Normal Estimation for 3D Point Clouds by Multi-Scale Edge ConditioningabstractEstimating surface normals from 3D point clouds is critical for various applications, including surface reconstruction and rendering. While existing methods for normal estimation perform well in regions where normals change slowly, they tend to fail where normals vary rapidly. To address this issue, we propose a novel approach called MSECNet, which improves estimation in normal varying regions by treating normal variation modeling as an edge detection problem. MSECNet consists of a backbone network and a multi-scale edge conditioning (MSEC) stream. The MSEC stream achieves robust edge detection through multi-scale feature fusion and adaptive edge detection. The detected edges are then combined with the output of the backbone network using the edge conditioning module to produce edge-aware representations. Extensive experiments show that MSECNet outperforms existing methods on both synthetic (PCPNet) and real-world (SceneNN) datasets while running significantly faster. We also conduct various analyses to investigate the contribution of each component in the MSEC stream. Finally, we demonstrate the effectiveness of our approach in surface reconstruction. Haoyi Xiu, Xin Liu 0020, Weimin Wang 0007, Kyoung-Sook Kim 0001, Masashi Matsuoka |
ACM Multimedia | 3 |
| 2023 | Optimizing Local Feature Representations of 3D Point Clouds with Anisotropic Edge Modeling
Haoyi Xiu, Xin Liu 0020, Weimin Wang 0007, Kyoung-Sook Kim 0001, Takayuki Shinohara, Qiong Chang, Masashi Matsuoka |
MMM (1) | 3 |
| 2023 | Broaden Your Positives: A General Rectification Approach for Novel Class Discovery
Yaqi Cai, Nan Pu, Qi Jia 0001, Weimin Wang 0007, Yu Liu 0012 |
PRCV (4) | 4 |
| 2023 | Enhancing Lidar and Radar Fusion for Vehicle Detection in Adverse Weather via Cross-Modality Semantic Consistency
Qiong Chang, Weimin Wang 0007 |
PRCV (3) | 5 |
| 2023 | Diffusion unit: Interpretable edge enhancement and suppression learning for 3D point cloud segmentationabstract3D point clouds are discrete samples of continuous surfaces which can be used for various applications. However, the lack of true connectivity information, i.e., edge information, makes point cloud recognition challenging. Recent edge-aware methods incorporate edge modeling into network designs to better describe local structures. Although these methods show that incorporating edge information is beneficial, how edge information helps remains unclear, making it difficult for users to analyze its usefulness. To shed light on this issue, in this study, we propose a new algorithm called Diffusion Unit (DU) that handles edge information in a principled and interpretable manner while providing decent improvement. First, we theoretically show that DU learns to perform task-beneficial edge enhancement and suppression. Second, we experimentally observe and verify the edge enhancement and suppression behavior. Third, we empirically demonstrate that this behavior contributes to performance improvement. Extensive experiments and analyses performed on challenging benchmarks verify the effectiveness of DU. Specifically, our method achieves state-of-the-art performance in object part segmentation using ShapeNet part and scene segmentation using S3DIS. Our source code is available at https://github.com/martianxiu/DiffusionUnit. Haoyi Xiu, Xin Liu 0020, Weimin Wang 0007, Kyoung-Sook Kim 0001, Takayuki Shinohara, Qiong Chang, Masashi Matsuoka |
Neurocomputing | 3 |
| 2022 | Weakly Supervised Point Cloud Upsampling VIA Optimal TransportabstractExisting learning-based methods usually train a point cloud upsampling model with synthesized, paired sparse-dense point clouds. However, the distribution gap between synthesized and real data limits the performance and generalization. To solve this problem, we innovatively regard the upsamplig task as an optimal transport (OT) problem from sparse to dense point cloud. Further we propose PU-CycGAN, a cycle network that consists of a Densifier, Sparsifier and two discriminators. It can be directly trained for upsampling with unpaired real sparse point clouds, so that the distribution gap can be filled via the learning. Especially, quadratic Wasserstein distance is introduced for the stable training. Extensive experiments on both synthetic and real-scanned datasets validate the effectiveness and advantages in terms of distribution uniformity, underlying surface representation and applicability to real data. The source code is available at https://github.com/cognaclee/PU-CycGAN. Zezeng Li, Weimin Wang 0007, Na Lei |
ICASSP | 2 |
| 2022 | Surgical Skill Assessment via Video Semantic Aggregation
Zhenqiang Li 0002, Lin Gu 0003, Weimin Wang 0007, Ryosuke Nakamura, Yoichi Sato 0001 |
MICCAI (8) | 3 |
| 2022 | Efficient stereo matching on embedded GPUs with zero-means cross correlation
Qiong Chang, Aolong Zha, Weimin Wang 0007, Xin Liu 0020, Masaki Onishi, Meng Joo Er, Tsutomu Maruyama |
J. Syst. Archit. | 3 |
| 2022 | Spatio-Temporal Perturbations for Video AttributionabstractThe attribution method provides a direction for interpreting opaque neural networks in a visual way by identifying and visualizing the input regions/pixels that dominate the output of a network. Regarding the attribution method for visually explaining video understanding networks, it is challenging because of the unique spatiotemporal dependencies existing in video inputs and the special 3D convolutional or recurrent structures of video understanding networks. However, most existing attribution methods focus on explaining networks taking a single image as input and a few works specifically devised for video attribution come short of dealing with diversified structures of video understanding networks. In this paper, we investigate a generic perturbation-based attribution method that is compatible with diversified video understanding networks. Besides, we propose a novel regularization term to enhance the method by constraining the smoothness of its attribution results in both spatial and temporal dimensions. In order to assess the effectiveness of different video attribution methods without relying on manual judgement, we introduce reliable objective metrics which are checked by a newly proposed reliability measurement. We verified the effectiveness of our method by both subjective and objective evaluation and comparison with multiple significant attribution methods. Zhenqiang Li 0002, Weimin Wang 0007, Zuoyue Li, Yifei Huang 0002, Yoichi Sato 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | Towards Visually Explaining Video Understanding Networks with Perturbationabstract"Making black box models explainable " is a vital problem that accompanies the development of deep learning networks. For networks taking visual information as input, one basic but challenging explanation method is to identify and visualize the input pixels/regions that dominate the network's prediction. However, most existing works focus on explaining networks taking a single image as input and do not consider the temporal relationship that exists in videos. Providing an easy-to-use visual explanation method that is applicable to diversified structures of video understanding networks still remains an open challenge. In this paper, we investigate a generic perturbation-based method for visually explaining video understanding networks. Besides, we propose a novel loss function to enhance the method by constraining the smoothness of its results in both spatial and temporal dimensions. The method enables the comparison of explanation results between different network structures to become possible and can also avoid generating the pathological adversarial explanations for video inputs. Experimental comparison results verified the effectiveness of our method. Zhenqiang Li 0002, Weimin Wang 0007, Zuoyue Li, Yifei Huang 0002, Yoichi Sato 0001 |
WACV | 2 |
| 2020 | Z2-ZNCC: ZigZag Scanning based Zero-means Normalized Cross Correlation for Fast and Accurate Stereo Matching on Embedded GPUabstractMobile stereo matching systems are becoming more important in many applications such as auto-driving and autonomous robots. However, to maintain its low power consumption, mobile platforms have only limited hardware resources. Accurate stereo matching methods require a high computational complexity, and it is difficult to maintain both acceptable accuracy and processing speed on the mobile platforms. To solve this trade-off, in this paper, we propose a novel acceleration approach for a well-known matching algorithm Zero-means Normalized Cross Correlation (ZNCC), and show its effectiveness on a Jetson TX2 embedded GPU. By combining our new approach, Z2- ZNCC, with the Semi-Global Matching (SGM) algorithm, our system achieves a low error rate of 7.76% while keeping 28 fps for 1242×375 pixels images with the maximum disparity of 128 on the KITTI 2015 dataset. This performance is higher than previous state-of-the-art system on the same hardware platform. Qiong Chang, Aolong Zha, Weimin Wang 0007, Xin Liu 0020, Masaki Onishi, Tsutomu Maruyama |
ICCD | 3 |
| 2020 | Weakly Supervised Silhouette-based Semantic Scene Change DetectionabstractThis paper presents a novel semantic scene change detection scheme with only weak supervision. A straightforward approach for this task is to train a semantic change detection network directly from a large-scale dataset in an end-to-end manner. However, a specific dataset for this task, which is usually labor-intensive and time-consuming, becomes indispensable. To avoid this problem, we propose to train this kind of network from existing datasets by dividing this task into change detection and semantic extraction. On the other hand, the difference in camera viewpoints, for example, images of the same scene captured from a vehicle-mounted camera at different time points, usually brings a challenge to the change detection task. To address this challenge, we propose a new siamese network structure with the introduction of correlation layer. In addition, we create a publicly available dataset for semantic change detection to evaluate the proposed method. The experimental results verified both the robustness to viewpoint difference in change detection task and the effectiveness for semantic change detection of the proposed networks. Our code and dataset are available at https://github.com/xdspacelab/sscdnet. Ken Sakurada, Mikiya Shibuya, Weimin Wang 0007 |
ICRA | 3 |
| 2020 | Densification of Airborne Lidar Point Cloud with Fused Encoder-Decoder NetworksabstractThis paper presents a density enhancement method for airborne LiDAR point cloud with the corresponding image based on a fused Encoder-Decoder network. Different from terrestrial indoor or outdoor scenes, the variance of objects and depth ranges in the large scale airborne data is challenging. To address the problem of objects at different scales, we propose a RGB and depth fused Encoder-Decoder structure inspired by UNet. In addition, we propose a heuristic method for refining the result if instance segmentation labels are available. Both quantitative and qualitative evaluations are performed on a dataset covering 24km2area of Osaka in Japan validates the feasibility of the proposed method for densification of point cloud in large scale environment. Weimin Wang 0007, Vinayaraj Poliyapram, Ryosuke Nakamura |
IGARSS | 1 |
| 2020 | P2Net: A Post-Processing Network for Refining Semantic Segmentation of LiDAR Point Cloud based on Consistency of Consecutive FramesabstractWe present a lightweight post-processing method to refine the semantic segmentation results of point cloud sequences. Most existing methods usually segment frame by frame and encounter the inherent ambiguity of the problem: based on a measurement in a single frame, labels are sometimes difficult to predict even for humans. To remedy this problem, we propose to explicitly train a network to refine these results predicted by an existing segmentation method. The network, which we call the P2Net, learns the consistency constraints between "coincident" points from consecutive frames after registration. We evaluate the proposed post-processing method both qualitatively and quantitatively on the SemanticKITTI dataset that consists of real outdoor scenes. The effectiveness of the proposed method is validated by comparing the results predicted by two representative networks with and without the refinement by the post-processing network. Specifically, qualitative visualization validates the key idea that labels of the points that are difficult to predict can be corrected with P2Net. Quantitatively, overall mIoU is improved from 10.5% to 11.7% for PointNet [1] and from 10.8% to 15.9% for PointNet++ [2]. Yutaka Momma, Weimin Wang 0007, Edgar Simo-Serra, Satoshi Iizuka, Ryosuke Nakamura, Hiroshi Ishikawa 0002 |
SMC | 2 |
| 2020 | PSNet: A Style Transfer Network for Point Cloud Stylization on Geometry and ColorabstractWe propose a neural style transfer method for colored point clouds which allows stylizing the geometry and/or color property of a point cloud from another. The stylization is achieved by manipulating the content representations and Gram-based style representations extracted from a pretrained PointNet-based classification network for colored point clouds. As Gram-based style representation is invariant to the number or the order of points, the style can also be an image in the case of stylizing the color property of a point cloud by merely treating the image as a set of pixels. Experimental results and analysis demonstrate the capability of the proposed method for stylizing a point cloud either from another point cloud or an image. Weimin Wang 0007, Katashi Nagao, Ryosuke Nakamura |
WACV | 2 |