VLDB 2026 Research / reviewers in the wild / expert
Moyun Liu
dblp:268/3069
· DBLP profile ↗
15ranked-venue papers
2as first author
15since 2021 · last 2026
0000-0002-4530-2606ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LaSSM: Efficient Semantic-Spatial Query Decoding via Local Aggregation and State Space Models for 3D Instance SegmentationabstractQuery-based 3D scene instance segmentation from point clouds has attained notable performance. However, existing methods suffer from the query initialization dilemma due to the sparse nature of point clouds and rely on computationally intensive attention mechanisms in query decoders. We accordingly introduceLaSSM, prioritizing simplicity and efficiency while maintaining competitive performance. Specifically, we propose a hierarchical semantic-spatial query initializer to derive the query set from superpoints by considering both semantic cues and spatial distribution, achieving comprehensive scene coverage and accelerated convergence. We further present a coordinate-guided state space model (SSM) decoder that progressively refines queries. The novel decoder features a local aggregation scheme that restricts the model to focus on geometrically coherent regions and a spatial dual-path SSM block to capture underlying dependencies within the query set by integrating associated coordinates information. Our design enables efficient instance prediction, avoiding the incorporation of noisy information and reducing redundant computation. LaSSM ranksfirst placeon the latest ScanNet++ V2 leaderboard, outperforming the previous best method by 2.5% mAP with only 1/3 FLOPs, demonstrating its superiority in challenging large-scale scene instance segmentation. LaSSM also achieves competitive performance on ScanNet V2, ScanNet200, S3DIS and ScanNet++ V1 benchmarks with less computational cost. Extensive ablation studies and qualitative results validate the effectiveness of our design. The code and weights are available at https://github.com/RayYoh/LaSSM. Yi Wang 0068, Yawen Cui, Moyun Liu, Lap-Pui Chau |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | Prompt-Driven Lightweight Foundation Model for Instance Segmentation-Based Fault Detection in Freight TrainsabstractAccurate visual fault detection in freight trains remains a critical challenge for intelligent transportation system maintenance, due to complex operational environments, structurally repetitive components, and frequent occlusions or contaminations in safety-critical regions. Conventional instance segmentation methods based on convolutional neural networks and Transformers often suffer from poor generalization and limited boundary accuracy under such conditions. To address these challenges, we propose a lightweight self-prompted instance segmentation framework tailored for freight train fault detection. Our method leverages the Segment Anything Model by introducing a self-prompt generation module that automatically produces task-specific prompts, enabling effective knowledge transfer from foundation models to domain-specific inspection tasks. In addition, we adopt a Tiny Vision Transformer backbone to reduce computational cost, making the framework suitable for real-time deployment on edge devices in railway monitoring systems. We construct a domain-specific dataset collected from real-world freight inspection stations and conduct extensive evaluations. Experimental results show that our method achieves$74.6~AP^{\text {box}}$and$74.2~AP^{\text {mask}}$on the dataset, outperforming existing state-of-the-art methods in both accuracy and robustness while maintaining low computational overhead. This work offers a deployable and efficient vision solution for automated freight train inspection, demonstrating the potential of foundation model adaptation in industrial-scale fault diagnosis scenarios. Project page:https://github.com/MVME-HBUT/SAM_FTI-FDet Guodong Sun 0002, Qihang Liang, Xingyu Pan, Moyun Liu, Yang Zhang 0053 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2026 | T-Mamba: A Unified Framework With Long-Range Dependency in Dual-Domain for 2D & 3D Tooth SegmentationabstractTooth segmentation is an essential step in digital dental diagnosis and treatment planning, playing a crucial role across various dental specialties. Despite its importance, this process is fraught with challenges due to the high noise and low contrast inherent in 2D and 3D dental radiographic images. Both convolutional neural networks (CNNs) and transformers have shown promise in tooth image segmentation, yet each method has limitations in handling long-range dependencies and computational complexity. To address this issue, this paper introduces T-Mamba, integrating frequency-based features and shared bi-positional encoding into vision mamba to address limitations in modeling global features efficiently. Besides, we designed a gate selection unit to integrate two features in the spatial domain and one feature in the frequency domain adaptively. T-Mamba is the first to integrate frequency-based features into vision mamba, offering the flexibility to process both 2D and 3D tooth image data without the need for separate modules. Additionally, the TED, a large-scale public 2D dental X-ray dataset, has been presented in this paper. Extensive experiments demonstrated that T-Mamba achieved new SOTA performance on both the 3D dental CBCT dataset and our proposed 2D TED dataset. The code and models are publicly available athttps://github.com/isbrycee/T-Mamba. Yonghui Zhu, Moyun Liu, James K. H. Tsoi, Kuo Feng Hung |
IEEE Trans. Multim. | 4 |
| 2025 | GaussianCross: Cross-modal Self-supervised 3D Representation Learning via Gaussian SplattingabstractThe significance of informative and robust point representations has been widely acknowledged for 3D scene understanding. Despite existing self-supervised pre-training counterparts demonstrating promising performance, the model collapse and structural information deficiency remain prevalent due to insufficient point discrimination difficulty, yielding unreliable expressions and suboptimal performance. In this paper, we present GaussianCross, a novel cross-modal self-supervised 3D representation learning architecture integrating feed-forward 3D Gaussian Splatting (3DGS) techniques to address current challenges. GaussianCross seamlessly converts scale-inconsistent 3D point clouds into a unified cuboid-normalized Gaussian representation without missing details, enabling stable and generalizable pre-training. Subsequently, a tri-attribute adaptive distillation splatting module is incorporated to construct a 3D feature field, facilitating synergetic feature capturing of appearance, geometry, and semantic cues to maintain cross-modal consistency. To validate GaussianCross, we perform extensive evaluations on various benchmarks, including ScanNet, ScanNet200, and S3DIS. In particular, GaussianCross shows a prominent parameter and data efficiency, achieving superior performance through linear probing (<0.1% parameters) and limited data training (1% of scenes) compared to state-of-the-art methods. Furthermore, GaussianCross demonstrates strong generalization capabilities, improving the full fine-tuning accuracy by 9.3% mIoU and 6.1% AP50 on ScanNet200 semantic and instance segmentation tasks, respectively, supporting the effectiveness of our approach. The code, weights, and visualizations are publicly available at https://rayyoh.github.io/GaussianCross/. Yi Wang 0068, Moyun Liu, Lap-Pui Chau |
ACM Multimedia | 4 |
| 2025 | Generalized Probabilistic Graphical Modeling for Multi-View Bipartite Graph ClusteringabstractMulti-view bipartite graph clustering (MVBGC) is an active pipeline in unsupervised learning to tackle the limited scalability issue of traditional graph clustering. Despite improved performance, numerous variants still fall under conventional modeling that plugs additional modules, which however induces increasingly intricate models and fails to reveal the inherent variable relationship. We make the first attempt to introduce probabilistic graphical models for modeling the multi-view bipartite graph clustering task, reformulating it as a maximum likelihood estimation (MLE) problem. Such a setting uncovers the underlying probabilistic correlations among the commonality, view-specific variables, and noisy components. By pruning redundancy and disturbance collectively referred to as noise, we prove that minimizing the total noise is an approximation of the lower bound of MLE for multi-view data observations. We further generalize the MLE setting with clustering-suited constraints, deriving a Generalized Probabilistic Graphical Modeling framework (GProM), achieving an interpretable, concise, and flexible MVBGC framework. Extensive experiments verify the effectiveness of our framework. Furthermore, statistical significance analysis reveals the effectiveness of different distribution assumptions, providing valuable insights for model design. Liang Li 0041, Yuangang Pan, Yinghua Yao, Junpu Zhang, Moyun Liu, Xueling Zhu, Xinwang Liu 0002, Kenli Li 0001, Ivor W. Tsang, Keqin Li 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | SGIFormer: Semantic-Guided and Geometric-Enhanced Interleaving Transformer for 3D Instance SegmentationabstractIn recent years, transformer-based models have exhibited considerable potential in point cloud instance segmentation. Despite the promising performance achieved by existing methods, they encounter challenges such as instance query initialization problems and excessive reliance on stacked layers, rendering them incompatible with large-scale 3D scenes. This paper introduces a novel method, named SGIFormer, for 3D instance segmentation, which is composed of the Semantic-guided Mix Query (SMQ) initialization and the Geometric-enhanced Interleaving Transformer (GIT) decoder. Specifically, the principle of our SMQ initialization scheme is to leverage the predicted voxel-wise semantic information to implicitly generate the scene-aware query, yielding adequate scene prior and compensating for the learnable query set. Subsequently, we feed the formed overall query into our GIT decoder to alternately refine instance query and global scene features for further capturing fine-grained information and reducing complex design intricacies simultaneously. To emphasize geometric property, we consider bias estimation as an auxiliary task and progressively integrate shifted point coordinates embedding to reinforce instance localization. SGIFormer attains state-of-the-art performance on ScanNet V2, ScanNet200, S3DIS datasets, and the challenging high-fidelity ScanNet ++ benchmark, striking a balance between accuracy and efficiency. The code, weights, and demo videos are publicly available athttps://rayyoh.github.io/SGIFormer/. Yi Wang 0068, Moyun Liu, Lap-Pui Chau |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | GEM: Boost Simple Network for Glass Surface Segmentation via Vision Foundation ModelsabstractDetecting glass regions is a challenging task due to the inherent ambiguity in their transparency and reflective characteristics. Current solutions in this field remain rooted in conventional deep learning paradigms, requiring the construction of annotated datasets and the design of network architectures. However, the evident drawback with these mainstream solutions lies in the time-consuming and labor-intensive process of curating datasets, alongside the increasing complexity of model structures. In this paper, we propose to address these issues by fully harnessing the capabilities of two existing vision foundation models (VFMs): Stable Diffusion and Segment Anything Model (SAM). Firstly, we construct a Synthetic but photorealistic large-scale Glass Surface Detection dataset, dubbed S-GSD, without any labour cost via Stable Diffusion. This dataset consists of four different scales, consisting of 168 k images totally with precise masks. Besides, based on the powerful segmentation ability of SAM, we devise a simpleGlass surface sEgMentor named GEM, which follows the simple query-based encoder-decoder architecture. Comprehensive experiments are conducted on the large-scale glass segmentation dataset GSD-S. Our GEM establishes a new state-of-the-art performance with the help of these two VFMs, surpassing the best-reported method GlassSemNet with an IoU improvement of 2.1%. Additionally, extensive experiments demonstrate that our synthetic dataset S-GSD exhibits remarkable performance in zero-shot and transfer learning settings. Moyun Liu, Kuo Feng Hung |
IEEE Trans. Multim. | 2 |
| 2024 | Multiple prior representation learning for self-supervised monocular depth estimation via hybrid transformer
Guodong Sun 0002, Mingxuan Liu 0002, Moyun Liu, Yang Zhang 0053 |
Eng. Appl. Artif. Intell. | 4 |
| 2024 | A concise but high-performing network for image guided depth completion in autonomous driving
Moyun Liu, Youping Chen, Jingming Xie, Yang Zhang 0053, Joey Tianyi Zhou |
Knowl. Based Syst. | 1 |
| 2024 | Efficient Visual Fault Detection for Freight Train via Neural Architecture Search With Data Volume RobustnessabstractDeep learning-based fault detection methods have achieved significant success. In visual fault detection of freight trains, there exists a large characteristic difference between interclass components (scale variance) but intraclass on the contrary, which entails scale-awareness for detectors. Moreover, the design of task-specific networks heavily relies on human expertise. As a consequence, neural architecture search (NAS) that automates the model design process gains considerable attention because of its promising performance. However, NAS is computationally intensive due to the large search space and huge data volume. In this work, we propose an efficient NAS-based framework for visual fault detection of freight trains to search for the task-specific detection head with capacities of multiscale representation. First, we design a scale-aware search space for discovering an effective receptive field in the head. Second, we explore the robustness of data volume to reduce search costs based on the specifically designed search space, and a novel sharing strategy is proposed to reduce memory and further improve search efficiency. Extensive experimental results demonstrate the effectiveness of our method with data volume robustness, which achieves 46.8 and 47.9 mAP on the bottom view and side view datasets, respectively. Our framework outperforms the state-of-the-art approaches and linearly decreases the search costs with reduced data volumes. Yang Zhang 0053, Mingying Li, Huilin Pan, Moyun Liu |
IEEE Trans. Ind. Informatics | 4 |
| 2024 | MENet: Multi-Modal Mapping Enhancement Network for 3D Object Detection in Autonomous DrivingabstractTo achieve more accurate perception performance, LiDAR and camera are gradually chosen to improve 3D object detection simultaneously. However, it is still a non-trivial task to build an effective fusion mechanism, and this is hindering the development of multi-modal based method. Especially, the mapping relationship construction between two modalities is far from fully explored. Canonical cross-modal mapping suffers from failure when the calibration matrix is incorrect, and it also greatly wastes the amount and density of RGB image information. This paper aims to extend the traditional one-to-one alignment relationship between LiDAR and camera. For all projected point clouds, we enhance their cross-modal mapping relationship through aggregating color-texture related feature and shape-contour related feature. Further, a mapping pyramid is proposed to leverage the semantic representation of the image feature at different stages. Based on the above mapping enhancement strategies, our method increases the engagement rate of image. Finally, we design a fusion module based on an attention mechanism to improve the point cloud feature with the auxiliary image feature. Extensive experiments on the KITTI dataset and SUN-RGBD dataset show that our model achieves satisfactory 3D object detection, especially for categories with sparse point clouds compared with other multi-modal fusion networks. Moyun Liu, Youping Chen, Jingming Xie, Yang Zhang 0053, Zhenshan Bing, Genghang Zhuang, Kai Huang 0001, Joey Tianyi Zhou |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2023 | BF3D: Bi-directional fusion 3D detector with semantic sampling and geometric mapping
Jingming Xie, Moyun Liu, Youping Chen |
Image Vis. Comput. | 3 |
| 2023 | Adaptive fusion affinity graph with noise-free online low-rank representation for natural image segmentation
Yang Zhang 0053, Moyun Liu, Guodong Sun 0002, Jingwu He |
Pattern Recognit. | 2 |
| 2022 | Affinity Fusion Graph-Based Framework for Natural Image SegmentationabstractThis paper proposes an affinity fusion graph framework to effectively connect different graphs with highly discriminating power and nonlinearity for natural image segmentation. The proposed framework combines adjacency-graphs and kernel spectral clustering based graphs (KSC-graphs) according to a new definition named affinity nodes of multi-scale superpixels. These affinity nodes are selected based on a better affiliation of superpixels, namely subspace-preserving representation which is generated by sparse subspace clustering based on subspace pursuit. Then a KSC-graph is built via a novel kernel spectral clustering to explore the nonlinear relationships among these affinity nodes. Moreover, an adjacency-graph at each scale is constructed, which is further used to update the proposed KSC-graph at affinity nodes. The fusion graph is built across different scales, and it is partitioned to obtain final segmentation result. Experimental results on the Berkeley segmentation dataset and Microsoft Research Cambridge dataset show the superiority of our framework in comparison with the state-of-the-art methods. The code is available athttps://github.com/Yangzhangcst/AF-graph. Yang Zhang 0053, Moyun Liu, Jingwu He, Yanwen Guo 0001 |
IEEE Trans. Multim. | 2 |
| 2021 | A Unified Light Framework for Real-Time Fault Detection of Freight Train ImagesabstractReal-time fault detection for freight trains plays a vital role in guaranteeing the security and optimal operation of railway transportation under stringent resource requirements. Despite the promising results for deep-learning-based approaches, the performance of these fault detectors on freight train images is far from satisfactory in both accuracy and efficiency. This article proposes a unified light framework to improve detection accuracy while supporting a real-time operation with a low-resource requirement. We first design a novel lightweight backbone (real-time fault detection network-RFDNet) to improve the accuracy and reduce computational cost. Then, we propose a multiregion proposal network using multiscale feature maps generated from the RFDNet to improve the detection performance. Finally, we present multilevel position-sensitive score maps and region of interest pooling to further improve accuracy with few redundant computations. Extensive experimental results on public benchmark datasets suggest that our RFDNet can significantly improve the performance of the baseline network with higher accuracy and efficiency. Experiments on six fault datasets show that our method is capable of real-time detection at over 38 frames/s and achieves competitive accuracy and lower computation than the state-of-the-art detectors. Yang Zhang 0053, Moyun Liu, Yang Yang 0092, Yanwen Guo 0001 |
IEEE Trans. Ind. Informatics | 2 |