VLDB 2026 Research / reviewers in the wild / expert
Shaojin Bai
dblp:342/1533
· DBLP profile ↗
7ranked-venue papers
4as first author
7since 2021 · last 2025
0000-0002-7444-6115ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
3D vision · 64% Representation and self-supervised learning · 28% Generative modeling · 8% | |
| Computer graphics and multimedia
1 paper |
Multimedia analysis and retrieval · 100% |
Topics — the 7 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › 3D vision › 3d object recognition
3d object classification |
0.9 | 1 | 2025 | DLS-HCAN: Duplex Label Smoothing Based Hierarchical Context-Aware Network for Fine-Grained 3D Shape Classification · IEEE Trans. Multim. 2025 |
Machine learning › Representation and self-supervised learning › representation learning
disentangled representation learning |
0.9 | 1 | 2025 | DCDL: Dual Causal Disentangled Learning for Zero-Shot Sketch-Based Image Retrieval · IEEE Trans. Multim. 2025 |
Computer vision › 3D vision › 3d object recognition › 3d object classification
fine-grained 3d shape classification |
0.9 | 1 | 2025 | DLS-HCAN: Duplex Label Smoothing Based Hierarchical Context-Aware Network for Fine-Grained 3D Shape Classification · IEEE Trans. Multim. 2025 |
Multimedia analysis and retrieval › image retrieval
sketch-based image retrieval |
0.9 | 1 | 2025 | DCDL: Dual Causal Disentangled Learning for Zero-Shot Sketch-Based Image Retrieval · IEEE Trans. Multim. 2025 |
Multimedia analysis and retrieval › image retrieval › sketch-based image retrieval
zero-shot sketch-based image retrieval |
0.9 | 1 | 2025 | DCDL: Dual Causal Disentangled Learning for Zero-Shot Sketch-Based Image Retrieval · IEEE Trans. Multim. 2025 |
Computer vision › 3D vision
3d shape analysis |
0.3 | 1 | 2025 | DLS-HCAN: Duplex Label Smoothing Based Hierarchical Context-Aware Network for Fine-Grained 3D Shape Classification · IEEE Trans. Multim. 2025 |
Machine learning › Generative modeling
variational autoencoder |
0.3 | 1 | 2025 | DCDL: Dual Causal Disentangled Learning for Zero-Shot Sketch-Based Image Retrieval · IEEE Trans. Multim. 2025 |
Methods — techniques the papers use, named apart from their topics
variational autoencoder · 1.7dual alignment · 1.7causal disentanglement · 1.7label smoothing · 0.9hierarchical context-aware network · 0.9attention mechanism · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SCDL: Sketch Causal Disentangled Learning for Sketch-Based 3D Shape RetrievalabstractSketch-based 3D shape retrieval (SBSR) has been a challenging task for decades, crucially depending on aligning shared semantic attributes between sketches and 3D shapes. Previous efforts mainly aimed at creating a common embedding space to bridge domain gaps. However, sketches’ subjective and abstract nature, known as confounders, potentially reduces learning performance of matching with 3D shapes. To address this issue, in this paper, we propose a sketch causal disentangled learning for SBSR, named SCDL, which introduce causal intervention to explicitly disentangle sketches into the inherent shared semantic part, and other unrelated confounders to classification (styles, abstraction levels, etc.) for the first time. Specifically, we construct a structural causal model (SCM) in the sketch branch under the dual variational autoencoder (VAE) architectures to alleviate confounders negative impact through learning the semantic attributes in the latent variable space. Next, we adopt a learning strategy on the separated semantic latent variables to construct a shared semantic embedding space further to make cross-modal features of the same class more similar, alleviating the cross-modality discrepancies effectively and establishing new state-of-the-art on three benchmarks. Comprehensive experiment results, ablation studies, and visualization validate the effectiveness of our approach. Shaojin Bai, Yalu Li, Rihao Chang, Qi Liang 0004, Weizhi Nie |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | DLS-HCAN: Duplex Label Smoothing Based Hierarchical Context-Aware Network for Fine-Grained 3D Shape ClassificationabstractFine-grained 3D shape classification (FGSC) has garnered significant attention recently and has made notable advancements. However, due to high inter-class similarity and intra-class diversity, it is still a challenge for existing methods to capture subtle differences between different subcategories for FGSC. On the one hand, one-hot labels in loss function are too hard to describe the above data characteristics, and on the other hand, local details are submerged in the global features extraction process and final network constraints, impacting classification results. In this paper, we propose a duplex label smoothing-based hierarchical context-aware network for fine-grained 3D shape classification, named DLS-HCAN. Specifically, DLS-HCAN firstly employs a hierarchical context-aware network (HCAN), in which the intra-view context attention mechanism (intra-ATT) and the inter-view context multilayer perceptron (inter-MLP) are designed to focus on and discern the beneficial local details. Subsequently, we propose a novel duplex label smoothing (DLS) regularization in which shape-level and view-level smooth labels are separately applied in two improved loss functions, adapting to the fine-grained data characteristics and considering the varying uniqueness of different views. Notably, our approach does not require additional annotation information. Experimental results and comparison with state-of-the-art methods demonstrate the superiority of our proposed DLS-HCAN for FGSC. In addition, our approach also achieves comparable performance for the coarse-grained dataset on ModelNet40. Shaojin Bai, Jing Bai 0004 |
IEEE Trans. Multim. | 1 |
| 2025 | DCDL: Dual Causal Disentangled Learning for Zero-Shot Sketch-Based Image RetrievalabstractZero-shot sketch-based image retrieval (ZS-SBIR) is a challenging task that hinges on overcoming the cross-domain differences between sketches and images. Previous methods primarily address cross-domain differences by creating a common embedding space, improving final retrieval results. However, most previous approaches have overlooked a critical aspect: sketch-based image retrieval task actually requires only the cross-domain invariant information relevant to the retrieval. Irrelevant information (such as posture, expression, background, and specificity) may detract from retrieval accuracy. In addition, most previous methods perform well on traditional SBIR datasets but lack corresponding research on generalization and extensibility in the face of more diverse and complex data. To address these issues, we propose a Dual Causal Disentangled Learning (DCDL) for ZS-SBIR. This approach can mitigate the negative impact of irrelevant features by separating retrieval-relevant features in the latent variable space. Specifically, we constructed a causal disentanglement model using two Variational Autoencoders (VAE), each applied to the sketch and image domains, to obtain disentangled variables with exchangeable attributes. Our framework effectively integrates causal intervention with disentangled representation learning, enabling a clearer separation of cross-domain retrieval-relevant and intra-class irrelevant features, which can be recombined into new reconstructed samples. Concurrently, we designed a Dual Alignment Module (DAM), leveraging the accurate and comprehensive semantic features provided by a text encoder pre-trained on large-scale datasets to supplement semantic associations and align disentangled retrieval-relevant features. The Dual Alignment Module enhances the model's ability to generalize across diverse datasets by effectively aligning retrieval-relevant information from different domains. Extensive experiments demonstrate that our method achieves state-of-the-art (SOTA) performance on the Sketchy and TU–Berlin datasets. Additionally, more experiments on larger scale dataset QuickDraw, fine-grained datasets, Shoe-V2 and Chair-V2, as well as an inter-dataset further validate the generalization and extensibility of DCDL. Qiang Li 0048, Wei Zhang 0390, Shaojin Bai, Weizhi Nie, Anan Liu |
IEEE Trans. Multim. | 4 |
| 2024 | Multi-modal fusion network guided by prior knowledge for 3D CAD model recognition
Qiang Li 0048, Zibo Xu, Shaojin Bai, Weizhi Nie, Anan Liu |
Neurocomputing | 3 |
| 2024 | V2MLP: an accurate and simple multi-view MLP network for fine-grained 3D shape recognition
Jing Bai 0004, Shaojin Bai |
Vis. Comput. | 3 |
| 2023 | PAGML: Precise Alignment Guided Metric Learning for sketch-based 3D shape retrieval
Shaojin Bai, Jing Bai 0004, Jiwen Tuo, Min Liu 0018 |
Image Vis. Comput. | 1 |
| 2023 | HDA2L: Hierarchical Domain-Augmented Adaptive Learning for sketch-based 3D shape retrieval
Shaojin Bai, Jing Bai 0004 |
Knowl. Based Syst. | 1 |