Jing Bai 0004

dblp:35/328-4 · DBLP profile ↗
← Back
34ranked-venue papers
4as first author
23since 2021 · last 2026
0000-0003-4247-6210ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 18 · 4 first-author · 9 since 2021Artificial intelligence and machine learning · 11 · 9 since 2021Databases, data management, data science and information retrieval · 4 · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 since 2021
YearPublicationVenuePosition
2026 UVConv: A UV grid-based convolutional neural network for classification and segmentation of 3D CAD models
Zenghui Su, Jing Bai 0004, Fei-wei Qin
Comput. Aided Geom. Des.2
2026 A cross-scale interaction framework combining Mamba and Convolutional Neural Networks for Arbitrary-Scale Super-Resolution of infrared images
Fei-wei Qin, Changmiao Wang, Kai Zhang 0008, Yong Peng 0001, Jing Bai 0004
Eng. Appl. Artif. Intell.6
2026 PVSTrans: Patch-view-shape progressive interaction transformer for 3D shape recognition
Jing Bai 0004, Zenghui Su
Inf. Process. Manag.2
2026 Bridging cross-modalities: Deep learning approaches for sketch-based 3D shape retrieval
Maozhu Xiang, Jing Bai 0004, Fei-wei Qin
Inf. Process. Manag.2
2026 Progressive Fusion of Multi-Scale Mamba Context and Local Detail Priors for Infrared Small Target Detection
abstract
Infrared Small Target Detection (IRSTD) requires strong target-level detection capability, which depends on effective modeling of long-range global dependencies. This demand has driven the transition from CNN-based approaches to Transformer-based architectures. Although Transformers improve global context modeling, their high computational cost limits practical deployment. Recent advances in Mamba enable efficient long-range dependency modeling with reduced complexity, offering a promising alternative that alleviates the efficiency limitations of Transformers while preserving target-level detection performance. However, Mamba is not inherently tailored for IRSTD, as it lacks explicit mechanisms for capturing fine-grained local details and modeling background variations across multiple spatial scales. To address these limitations, we propose MCFNet, an encoder-decoder framework that integrates Mamba to enhance target-level detection performance with moderate computational cost. MCFNet introduces a Detail-Capturable Convolution Block to strengthen local detail perception and a Multi-scale Contextual Mamba Block to improve background modeling across different scales. While the resulting dual-branch design enhances both global semantics and local details, it also introduces challenges in feature fusion. To this end, a Feature Fusion Decoding Module is further proposed to enable effective collaboration between global and local representations. Extensive experiments on multiple public IRSTD benchmark datasets demonstrate that MCFNet consistently outperforms existing methods in both pixel-level and target-level metrics, achieving higher detection accuracy with reduced false alarms. The code of our model is available at: https://github.com/Fihven/MCFNet.
Xiangjun Zhu, Fei-wei Qin, Changmiao Wang, Jin Fan 0003, Fei Lin 0006, Jing Bai 0004, Chenglong Zhang 0001, David Zhang 0001
IEEE Trans. Image Process.6
2026 CMI-Net: Cross-View Message Token Interaction Network for 3D Shape Recognition
Jing Bai 0004, Zenghui Su
IEEE Trans. Multim.2
2025 Glance and Gaze: Progressive end-to-end deep learning for fine-grained 3D shape classification
abstract
Fine-grained 3D shape classification (FGSC) poses a unique challenge due to subtle inter-class differences and intricate structural information, often leading to suboptimal performance. This paper introduces GGView-Net, a novel approach for FGSC utilizing a progressive Glance and Gaze mechanism inspired by human object recognition processes. GGView-Net comprises three key stages - Glance, Gaze, and Joint Classification. The Glance phase mimics the holistic perception of a 3D shape by extracting global features, while the Gaze stage focuses on generating intra-view attention features guided by the global context. Finally, the Joint Classification step tailors losses to individual views and FGSC requirements. Through comprehensive exper-iments on FG3D, GGView-Net demonstrates superior FGSC performance using only class labels, with visualization results confirming its interpretability. Moreover, GGView-Net achieves state-of-the-art results in meta-category 3D shape classification on the ModelNet40 dataset, showcasing remarkable versatility.
Jing Bai 0004, Jinzhe Jiang
CSCWD2
2025 Multi-View Text Enhancement for Parameter-Free Zero-Shot 3D Model Classification
abstract
Large-scale pre-trained models have demonstrated significant advantages when dealing with visual and language tasks in open-world scenes. However, recent studies have identified certain limitations when utilizing comparative language-visual pretrained models for zero-shot 3D model classification, mainly in the form of neglecting the multi-view independent information of the 3D model. In addition, existing textual cues usually rely only on category semantics and fail to fully exploit the rich contextual structure of the 3D model itself, thus limiting the effective matching of visual and semantic information. To this end, this paper proposes a multi-view text-enhanced zero-shot 3D model classification based on a parameter-free network. The method utilizes a pre-trained image encoder to extract multi-view visual features and combines the semantic comprehension capability of a large language model to mine 3D model structure information from multiple perspectives. Fine-grained associations between views and textual cues are made through view-by-view interaction, and decision-level fusion is used to integrate the independent decision results of each view to achieve zero-shot classification. Without any 3D training, the method in this paper achieves 67.4% accuracy on the ZS3D dataset and 61.3%, 43.3% and 28.1% accuracy on the three sub-datasets of the Ali dataset, which verifies the generalizability and effectiveness of the proposed method.
Shuting Xi, Jing Bai 0004
CSCWD2
2025 FG3DFormer: Fine-Grained 3D Shape Classification Based on Vision Transformer
abstract
Fine-grained 3D shape classification (FGSC) remains challenging due to the difficulty of adaptively capturing global structure differences and subtle inter-class distinctions. This paper directly extends Vision Transformer (ViT) to FGSC, proposing a pure Transformer network FG3DFormer that fully leverages ViT’s global correlation and local attention abilities. FG3Dformer comprises the Hierarchical Feature Extraction (HFE) and the Hierarchical Feature Refinement (HFR), interconnected through the Adaptive View Region Selection (AVRS). Firstly, the HFE comprehensively evaluates the significance of intra-view patches and views driven by inter-view and intraview attention. Then, the AVRS adaptively selects crucial patch Tokens from different views to serve as sources of subtle local features. Finally, the HFR refines the 3D shape descriptor, capturing more discriminative global and subtle local features by leveraging both the view and selected crucial patch Tokens. Extensive experiments on FG3D and ModelNet40 demonstrate the superiority of FG3Dformer in FGSC and meta-category 3D shape classification tasks.
Jing Bai 0004, Jinzhe Jiang
ICASSP2
2025 DKD2L: Dual Knowledge Distillation Dynamic Learning for sketch-based 3D shape retrieval
abstract
Sketch-based 3D shape retrieval has become a prominent area of research in computer vision, confronting challenges related to the inherent diversity and abstraction of sketches, as well as inter-domain discrepancies. This paper introduces a novel approach called Dual Knowledge Distillation Dynamic Learning (DKD2L) aimed at enhancing the extraction of spatio-temporal features for sketch-based 3D shape retrieval. We develop a temporal feature extraction network to effectively capture the dynamic temporal characteristics of sketches and improve retrieval efficiency through temporal knowledge distillation. Additionally, to tackle intra-class variation and inter-class imbalance, we apply semantic knowledge distillation, enabling the 3D shape network to guide the sketch network in capturing common semantic information. This approach facilitates precise cross-modal alignment and enhances retrieval accuracy. Extensive experiments on two benchmark datasets demonstrate that DKD2L surpasses existing state-of-the-art methods.
Yawen Su, Jing Bai 0004, Gan Lin
ICASSP2
2025 D2Former: Dynamic Sketch Decomposition and Selection Transformer for Sketch-Based 3D Shape Retrieval
abstract
Sketch-based 3D shape retrieval (SBSR) is an active research area in the computer vision community, but it is still very challenging. One main reason is that existing methods usually treat sketches as regular images, neglecting their unique abstraction and sparsity. This paper introduces a novel SBSR Transformer framework, D2Former, which fully exploits the uniqueness of hand-drawn sketches composed of strokes. For sketches, we first decompose the dynamical sketch into a sequence of partially overlapping sub-sketches and design a Dynamic Sub-sketch Select Module (DSSM) to select key sub-sketch features. The Transformer architecture is then employed to integrate global sketch features while highlighting their representative discriminative features. For 3D shapes, we construct an Inter-View Fusion Module (IVFM), which can incorporate features from multiple views and accurately capture the associations between them. Experiments on two benchmark datasets demonstrate that our method exhibits superior retrieval performance compared to the current state-of-the-art SBSR methods.
Gan Lin, Jing Bai 0004, Yawen Su
IJCNN2
2025 MDT-Net: A Mask Decoder Tuning Strategy for CLIP-Based Zero-Shot 3D Classification
Jing Bai 0004
MMM (2)2
2025 SKD-SBSR: Structural Knowledge Distillation for Sketch-Based 3D Shape Retrieval
Yawen Su, Jing Bai 0004, Gan Lin
Knowl. Based Syst.3
2025 $\hbox {C}^2$DFL: cross-view cross-layer discriminative feature learning for fine-grained 3D shape classification
Jinzhe Jiang, Jing Bai 0004
Neural Comput. Appl.2
2025 DLS-HCAN: Duplex Label Smoothing Based Hierarchical Context-Aware Network for Fine-Grained 3D Shape Classification
abstract
Fine-grained 3D shape classification (FGSC) has garnered significant attention recently and has made notable advancements. However, due to high inter-class similarity and intra-class diversity, it is still a challenge for existing methods to capture subtle differences between different subcategories for FGSC. On the one hand, one-hot labels in loss function are too hard to describe the above data characteristics, and on the other hand, local details are submerged in the global features extraction process and final network constraints, impacting classification results. In this paper, we propose a duplex label smoothing-based hierarchical context-aware network for fine-grained 3D shape classification, named DLS-HCAN. Specifically, DLS-HCAN firstly employs a hierarchical context-aware network (HCAN), in which the intra-view context attention mechanism (intra-ATT) and the inter-view context multilayer perceptron (inter-MLP) are designed to focus on and discern the beneficial local details. Subsequently, we propose a novel duplex label smoothing (DLS) regularization in which shape-level and view-level smooth labels are separately applied in two improved loss functions, adapting to the fine-grained data characteristics and considering the varying uniqueness of different views. Notably, our approach does not require additional annotation information. Experimental results and comparison with state-of-the-art methods demonstrate the superiority of our proposed DLS-HCAN for FGSC. In addition, our approach also achieves comparable performance for the coarse-grained dataset on ModelNet40.
Shaojin Bai, Jing Bai 0004
IEEE Trans. Multim.3
2024 D2GL: Dual-level dual-scale graph learning for sketch-based 3D shape retrieval
Jing Bai 0004
Pattern Recognit.2
2024 Corrigendum to "FGPNet: A weakly supervised fine-grained 3D point clouds classification network" [Pattern Recognition 139 (2023) 109509]
Huihui Shao, Jing Bai 0004, Rusong Wu, Jinzhe Jiang, Hongbo Liang
Pattern Recognit.2
2024 DCNet: exploring fine-grained vision classification for 3D point clouds
Rusong Wu, Jing Bai 0004, Jinzhe Jiang
Vis. Comput.2
2024 V2MLP: an accurate and simple multi-view MLP network for fine-grained 3D shape recognition
Jing Bai 0004, Shaojin Bai
Vis. Comput.2
2023 PAGML: Precise Alignment Guided Metric Learning for sketch-based 3D shape retrieval
Shaojin Bai, Jing Bai 0004, Jiwen Tuo, Min Liu 0018
Image Vis. Comput.2
2023 HDA2L: Hierarchical Domain-Augmented Adaptive Learning for sketch-based 3D shape retrieval
Shaojin Bai, Jing Bai 0004
Knowl. Based Syst.2
2023 FGPNet: A weakly supervised fine-grained 3D point clouds classification network
Huihui Shao, Jing Bai 0004, Rusong Wu, Jinzhe Jiang, Hongbo Liang
Pattern Recognit.2
2022 3D CAD model retrieval based on sketch and unsupervised variational autoencoder
Fei-wei Qin, Shuming Gao, Jing Bai 0004
Adv. Eng. Informatics4
2020 Feature-attention module for context-aware image-to-image translation
Jing Bai 0004, Min Liu 0018
Vis. Comput.1
2018 Ellipse Detection on Images Using Conic Power of Two Points
Min Liu 0018, Bodi Yuan, Jing Bai 0004
BMVC3
2017 WireFab: Mix-Dimensional Modeling and Fabrication for 3D Mesh Models
abstract
Many rapid fabrication technologies are directed towards layer wise printing or laser based prototyping. We propose WireFab, a rapid modeling and prototyping system that uses bent metal wires as the structure framework. WireFab approximates both the skeletal articulation and the skin appearance of the corresponding virtual skin meshes, and it allows users to personalize the designs by (1) specifying joint positions and part segmentations, (2) defining joint types and motion ranges to build a wire-based skeletal model, and (3) abstracting the segmented meshes into mixed-dimensional appearance patterns or attachments.
Min Liu 0018, Jing Bai 0004, Yuanzhi Cao, Jeffrey M. Alperovich, Karthik Ramani
CHI3
2016 Deep Learning 3D Shape Surfaces Using Geometry Images
Ayan Sinha, Jing Bai 0004, Karthik Ramani
ECCV (6)2
2016 An ontology-based semantic retrieval approach for heterogeneous 3D CAD models
Fei-wei Qin, Shuming Gao, Ming Li 0017, Jing Bai 0004
Adv. Eng. Informatics5
2015 Constructive generation of the medial axis for solid models
Housheng Zhu, Yusheng Liu 0006, Jing Bai 0004, Xiaoping Ye
Comput. Aided Des.3
2012 A flexible assembly retrieval approach for model reuse
Shuming Gao, Jing Bai 0004
Comput. Aided Des.4
2010 Design reuse oriented partial retrieval of CAD models
Jing Bai 0004, Shuming Gao, Weihua Tang, Yusheng Liu 0006
Comput. Aided Des.1
2009 Semantic-based partial retrieval of CAD models for design reuse
abstract
In this paper, we present a semantic-based partial retrieval approach of CAD models for design reuse. Based on the observation of reusable regions for design of 3D CAD models, for each model in the model library, we propose an algorithm that automatically extracts its reusable regions for partial retrieval. To further effectively support the partial retrieval of these reusable regions through simple queries, we represent each reusable region by all its local matching regions and describe each local matching region in a hierarchical way. Based on the hierarchical descriptor, a partial retrieval method for reusable regions is put forward. The approach proposed is implemented and tested by hundreds of mechanical parts. Preliminary results show that the method can effectively support partial retrieval for design reuse.
Jing Bai 0004, Shuming Gao, Weihua Tang, Yusheng Liu 0006
Symposium on Solid and Physical Modeling1
2007 Multi-mode Solid Model Retrieval Based on Partial Matching
abstract
In this paper, a multi-mode retrieval approach of solid models is presented. First, a local feature determination method based on wavelet transform is given, which makes the local feature agree with the human perception of resemblance; then a uniform and efficient DBMS (dilation based multiresolutional skeleton) generation method for global model and local feature using adaptive voxelization and local dilation is proposed; finally a flexible similarity assessment is proposed which supports both global matching and partial matching effectively and further supports different retrieval requires. Preliminary experiment results show that a multi-mode retrieval approach can describe user's retrieval intent more precisely and return retrieval results more properly.
Jing Bai 0004, Yusheng Liu 0006, Shuming Gao
CAD/Graphics1
2006 Multiresolutional similarity assessment and retrieval of solid models based on DBMS
Shuming Gao, Yusheng Liu 0006, Jing Bai 0004, B. K. Hu
Comput. Aided Des.4