Boyan Wan

dblp:266/8359 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
3since 2021 · last 2026
0000-0003-1055-6822ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
3D vision · 76% Generative modeling · 16% Information extraction and text analysis · 8%

Topics — the 10 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision › object pose estimation › 6d object pose estimation
category-level object pose estimation
2.532026
Learning Positive-Incentive Point Sampling in Neural Implicit Fields for Object Pose Estimation · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Equivariant Diffusion Model With A5-Group Neurons for Joint Pose Estimation and Shape Reconstruction · IEEE Trans. Pattern Anal. Mach. Intell. 2025
SOCS: Semantically-aware Object Coordinate Space for Category-Level 6D Object Pose Estimation under Large Shape Variations · ICCV 2023
Computer vision › 3D vision
object pose estimation
2.532026
Learning Positive-Incentive Point Sampling in Neural Implicit Fields for Object Pose Estimation · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Equivariant Diffusion Model With A5-Group Neurons for Joint Pose Estimation and Shape Reconstruction · IEEE Trans. Pattern Anal. Mach. Intell. 2025
SOCS: Semantically-aware Object Coordinate Space for Category-Level 6D Object Pose Estimation under Large Shape Variations · ICCV 2023
Computer vision › 3D vision
3d shape reconstruction
1.222026
Equivariant Diffusion Model With A5-Group Neurons for Joint Pose Estimation and Shape Reconstruction · IEEE Trans. Pattern Anal. Mach. Intell. 2025
Learning Positive-Incentive Point Sampling in Neural Implicit Fields for Object Pose Estimation · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Computer vision › 3D vision
implicit neural representation
1.012026
Learning Positive-Incentive Point Sampling in Neural Implicit Fields for Object Pose Estimation · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Machine learning › Generative modeling
diffusion model
0.912025
Equivariant Diffusion Model With A5-Group Neurons for Joint Pose Estimation and Shape Reconstruction · IEEE Trans. Pattern Anal. Mach. Intell. 2025
Machine learning › Generative modeling › diffusion model › geometric diffusion model
equivariant diffusion model
0.912025
Equivariant Diffusion Model With A5-Group Neurons for Joint Pose Estimation and Shape Reconstruction · IEEE Trans. Pattern Anal. Mach. Intell. 2025
Computer vision › 3D vision › 3d shape reconstruction
object shape reconstruction
0.912025
Equivariant Diffusion Model With A5-Group Neurons for Joint Pose Estimation and Shape Reconstruction · IEEE Trans. Pattern Anal. Mach. Intell. 2025
Natural language and speech › Information extraction and text analysis › named entity recognition
chinese named entity recognition
0.412020
Word-Character Graph Convolution Network for Chinese Named Entity Recognition · IEEE ACM Trans. Audio Speech Lang. Process. 2020
Natural language and speech › Information extraction and text analysis
named entity recognition
0.412020
Word-Character Graph Convolution Network for Chinese Named Entity Recognition · IEEE ACM Trans. Audio Speech Lang. Process. 2020
Computer vision › 3D vision
point cloud
0.312025
Equivariant Diffusion Model With A5-Group Neurons for Joint Pose Estimation and Shape Reconstruction · IEEE Trans. Pattern Anal. Mach. Intell. 2025

Methods — techniques the papers use, named apart from their topics

SO(3)-equivariant convolution · 1.9teacher-student pseudo-labeling · 1.0positive-incentive point sampling · 1.0geometry-based plausibility measure · 0.9equivariant feature extraction · 0.9a5-group neurons · 0.9multi-scale attention network · 0.7coordinate regression · 0.7attention · 0.4BiLSTM-CRF · 0.4
YearPublicationVenuePosition
2026 Learning Positive-Incentive Point Sampling in Neural Implicit Fields for Object Pose Estimation
abstract
Learning neural implicit fields of 3D shapes is a rapidly emerging field that enables shape representation at arbitrary resolutions. Due to the flexibility, neural implicit fields have succeeded in many research areas, including shape reconstruction, novel view image synthesis, and more recently, object pose estimation. Neural implicit fields enable learning dense correspondences between the camera space and the object's canonical space - including unobserved regions in camera space - significantly boosting object pose estimation performance in challenging scenarios like highly occluded objects and novel shapes. Despite progress, predicting canonical coordinates for unobserved camera-space regions remains challenging due to the lack of direct observational signals. This necessitates heavy reliance on the model's generalization ability, resulting in high uncertainty. Consequently, densely sampling points across the entire camera space may yield inaccurate estimations that hinder the learning process and compromise performance. To alleviate this problem, we propose a method combining an SO(3)-equivariant convolutional implicit network and a positive-incentive point sampling (PIPS) strategy. The SO(3)-equivariant convolutional implicit network estimates point-level attributes with SO(3)-equivariance at arbitrary query locations, demonstrating superior performance compared to most existing baselines. The PIPS strategy dynamically determines sampling locations based on the input, thereby boosting the network's accuracy and training efficiency. The PIPS strategy is implemented with a PIPS estimation network which generates sparse sample points with distinctive features capable of determining all object pose DoFs with high certainty. To collect the training data of the PIPS estimation network, we propose to automatically generate the pseudo ground-truth with a teacher model. Our method outperforms the state-of-the-art on three pose estimation datasets. It achieves 0.63 in the $5^{\circ }2$5∘2 cm metric on NOCS-REAL275, 0.62 in the $5^{\circ }5$5∘5 cm metric on ShapeNet-C, and 77.3 in the AR metric on LineMOD-O. Notably, it demonstrates significant improvements in challenging scenarios, such as objects captured with unseen pose, high occlusion, novel geometry, and severe noise.
Boyan Wan, Xin Xu 0001, Kai Xu 0004
IEEE Trans. Pattern Anal. Mach. Intell.2
2025 Equivariant Diffusion Model With A5-Group Neurons for Joint Pose Estimation and Shape Reconstruction
abstract
Object pose estimation and shape reconstruction are inherently coupled tasks although they have so far been studied separately in most existing approaches. A few recent works addressed the problem of joint pose estimation and shape reconstruction, but they found difficulties in handling partial observations and shape ambiguities. An open challenge in this area is to design a mechanism that has the two tasks benefit each other and boost the performance and robustness of both. In this work, we advocate the use of diffusion models for joint estimation of category-level object poses and reconstruction of object geometry. Diffusion models formulate shape reconstruction as a generation process conditioned on input observations. It has two main advantages. First, the iterative inference of diffusion models provides a mechanism for iterative optimization for both pose estimation and shape reconstruction. Second, diffusion models allow multiple outputs starting from different input noises, which would address the problem of ambiguity caused by partial observations. To achieve this, we propose equivariant diffusion model for joint pose estimation and shape reconstruction. The approach consists of an equivariant feature extractor to aggregate features of the input point cloud and a ShapePose diffusion model to generate object pose and shape simultaneously. To avoid training the model on all possible shape poses in the SO(3) space, we propose to augment the diffusion model with A5-group neurons where the neurons are converted into 5D vectors and can be rotated with the alternating group A5. Based on the A5-group neurons, we implement SO(3)-equivariant 3D point convolution and SO(3)-equivariant concatenation, making the entire network SO(3)-equivariant. Moreover, to select the most plausible combination of pose and shape from the generated ones, we propose a geometry-based measure of plausibility for an estimated pose along with a reconstructed shape. Extensive experiments demonstrate the effectiveness of the proposed method. Specifically, our method achieves the state-of-the-art on two public datasets and a new dataset with stacked objects, in terms of shape reconstruction and pose estimation. In particular, we show the proposed method could provide multiple plausible outputs under partial observations and shape ambiguities.
Boyan Wan, Kai Xu 0004
IEEE Trans. Pattern Anal. Mach. Intell.1
2023 SOCS: Semantically-aware Object Coordinate Space for Category-Level 6D Object Pose Estimation under Large Shape Variations
abstract
Most learning-based approaches to category-level 6D pose estimation are design around normalized object coordinate space (NOCS). While being successful, NOCS-based methods become inaccurate and less robust when handling objects of a category containing significant intra-category shape variations. This is because the object coordinates induced by global and rigid alignment of objects are semantically incoherent, making the coordinate regression hard to learn and generalize. We propose Semantically-aware Object Coordinate Space (SOCS) built by warping-and-aligning the objects guided by a sparse set of keypoints with semantically meaningful correspondence. SOCS is semantically coherent: Any point on the surface of a object can be mapped to a semantically meaningful location in SOCS, allowing for accurate pose and size estimation under large shape variations. To learn effective coordinate regression to SOCS, we propose a novel multi-scale coordinatebased attention network. Evaluations demonstrate that our method is easy to train, well-generalizing for large intracategory shape variations and robust to inter-object occlusions. Code is provided at: https://github.com/wanboyan/SOCS.
Boyan Wan, Kai Xu 0004
ICCV1
2020 Word-Character Graph Convolution Network for Chinese Named Entity Recognition
abstract
Recent researches try to integrate word information into the character-based Chinese NER by modifying the structure of the standard BiLSTM-CRF model. They follow the paradigm of explicitly modeling forward and backward sequences, adopting an LSTM variant that takes both characters and words as input for each direction. Though enriching the representations, these models cannot fully exploit the interaction between future and past contexts. In this paper, we propose a novel word-character graph convolution network (WC-GCN) which uses a cross GCN block to simultaneously process the word-character directed acyclic graphs (DAGs) of two directions. To improve the capture of long-distance dependency, a global attention GCN block is introduced to learn node representations conditioned on a global context. In both blocks, unlike previous works where each word is attached to its associated character or taken as a shortcut between LSTM cells, words and characters are treated equally as nodes in the graph and have their instance-specific representations. Experiments on four widely used datasets show that our proposed model can work standalone or with the standard BiLSTM. Both forms can outperform previous LSTM-based models without training on extra corpora while only an external lexicon and its corresponding pretrained character and word embeddings are needed.
Zhuo Tang, Boyan Wan, Li Yang 0012
IEEE ACM Trans. Audio Speech Lang. Process.2