Jun Xiao 0005

dblp:71/2308-5 · DBLP profile ↗
← Back
70ranked-venue papers
2as first author
59since 2021 · last 2026
0000-0002-1799-3948ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 38 · 32 since 2021Artificial intelligence and machine learning · 24 · 22 since 2021Applied, interdisciplinary, general and emerging computing · 15 · 1 first-author · 11 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Security and privacy · 1 · 1 first-author
YearPublicationVenuePosition
2026 PSPO: Trainable Potential-Based Reward Shaping with Internal Model Signals for Post-Training Policy Optimization of Large Language Models
Miaobo Hu, Bokun Wang, Shuhao Hu, Xin Wang 0086, Daren Zha, Jun Xiao 0005
ICIC (5)8
2026 AF-BEV: Object-Aware Adaptive Frustum-based BEV Aggregation for 3D Object Detection
abstract
Multi-view 3D detection in bird’s-eye-view (BEV) has attracted significant attention from both industry and academia due to its unified, intuitive representation and low cost. By aggregating features along cast rays based on their spatial order, existing methods aim to address the false positive (FP) artifacts commonly observed in detections. However, this strategy may also suppress features in occluded regions along the path, which leads to both inaccurate localization and missed detections. In this paper, we propose AF-BEV, an Adaptive Frustum-based BEV detector, which leverages visible object information to compensate for occluded regions, thereby achieving more accurate localization. Our framework first extends each ray into a frustum and employs an offset prediction module to guide the adaptive frustum angle based on the perceived spatial range of objects. A learnable occlusion-aware aggregation module is then introduced to effectively aggregate frustum features for 3D detection. Finally, we incorporate a unique object-level supervision to guide the generation of appropriate frustum ranges. Extensive experiments on nuScenes demonstrate the effectiveness of our proposed AF-BEV, achieving consistent improvements over the baseline. Additionally, qualitative visualizations provide interpretable evidence of the contributions.
Bingyu Zhu, Dongbo Yu, Yunbiao Wang, Jun Xiao 0005, Haiyong Jiang
ICMR4
2026 MIFFNet: A multidimensional image feature fusion network for lettuce fresh weight estimation
abstract
• We introduce an innovative multidimensional image feature fusion (MIFF) network specifically designed for lettuce fresh weight estimation. The core component of the proposed MIFF Module effectively integrates image features with RGB image data. • The prediction results across the three experimental datasets demonstrated optimal performance, achieving R² values of 0.929, 0.940, and 0.943, respectively.These results significantly outperform baseline methods and all other comparative approaches. • Extensive comparative experiments also demonstrate that the proposed network offers advantages in terms of low computational cost, and lightweight architecture. The estimation of lettuce fresh weight is critical for assessing growth status and optimizing cultivation. Traditional methods are often inefficient, error-prone, and costly. Computer vision offers opportunities for image-based non-destructive fresh weight estimation. This paper introduces MIFFNet, an end-to-end network integrating RGB images, visible light vegetation indices, geometric features, and color features for lettuce fresh weight estimation. This model employs the Inception structure as the multi-scale feature extraction (MSFE) block, alternating with the proposed multidimensional image feature fusion (MIFF) module to form the network backbone. This design enhances the model’s ability to capture multiscale features while thoroughly integrating multidimensional image features. Comparative experiments were conducted with 10 competitors, including classical convolutional neural networks, and existing lettuce fresh weight estimation models, across three lettuce datasets. Experimental results demonstrated that MIFFNet outperforms others across all three datasets. On the self-built dataset, it achieved an R 2 of 0.929, with RMSE and MAE values of 28.544g and 14.446g, respectively. On two public datasets, the R 2 values reached 0.94 and 0.943, with lower RMSE and MAE than competitors. Furthermore, MIFFNet exhibited significant advantages in terms of model complexity and parameter efficiency. These results highlight MIFFNet’s superior capability of accurate and efficient lettuce fresh weight estimation.
Jun Xiao 0005, Jinmeng Zhang, Rupeng Luan, Qian Zhang 0110
Pattern Recognit.2
2026 Hi-RWKV: Hierarchical RWKV Modeling for Hyperspectral Image Classification
abstract
Hyperspectral image (HSI) classification demands models that can jointly capture long-range spatial relations and high-dimensional spectral structures while remaining scalable to large scenes and robust under limited supervision. Existing CNN-, Transformer-, and state-space-based approaches either suffer from restricted receptive fields, quadratic attention complexity, or directional biases that hinder dense pixel-wise prediction. To address these limitations, we propose Hi-RWKV, a hierarchical recurrent weighted key-value framework tailored for hyperspectral analysis. Hi-RWKV introduces three key innovations: 1) a spatial structure-guided bidirectional propagation mechanism that integrates global spatial context while preserving boundary fidelity via edge-aware gating; 2) a spectral identity-driven channel mixing module that incorporates learnable band embeddings and whitening transforms to enhance cross-band discriminability; and 3) a multi-stage hierarchical encoder that progressively refines spectral-spatial representations with strictly linear complexity. Together, these designs enable efficient, direction-free spectral-spatial reasoning essential for large-scale HSI interpretation. Extensive experiments on four benchmarks demonstrate that Hi-RWKV consistently achieves state-of-the-art accuracy under diverse training regimes. Ablation studies confirm that each proposed module offers complementary gains in boundary preservation, spectral discrimination, and data efficiency. By unifying scalable recurrence with hyperspectral-specific structural modeling, Hi-RWKV establishes a strong and efficient paradigm for high-resolution remote sensing. The logs and source data of this article are available at https://github.com/HSI-Lab/Hi-RWKV.
Yunbiao Wang, Dongbo Yu, Hengyu Niu, Daifeng Xiao, Lupeng Liu, Jun Xiao 0005
IEEE Trans. Image Process.7
2026 HierRelTriple: Guiding Indoor Layout Generation With Hierarchical Relationship Triplet Losses
abstract
We present a hierarchical triplet-based indoor relationship learning method, coined HierRelTriple, with a focus on spatial relationship learning. Existing approaches often depend on attention-based network designs and manually defined training objectives using handcrafted spatial rules or simplified pairwise relationships. However, these methods fail to capture complex, multi-object relationships found in real scenarios, leading to overcrowded or physically implausible arrangements. We introduce HierRelTriple, a hierarchical framework for modeling relational triplets, which first partitions functional regions and then automatically extracts three levels of spatial relationships: object-to-region (O2R), object-to-object (O2O), and corner-to-corner (C2C). By representing these relationships as geometric triplets and employing approaches based on Delaunay Triangulation to establish spatial priors, we derive IoU-based losses between denoised and ground-truth triplets and integrate them seamlessly into the diffusion denoising process. The joint formulation of inter-object distances, angular orientations, and spatial relationships enhances the physical realism of the generated scenes. Extensive experiments on unconditional layout synthesis, floorplan-conditioned layout generation, and scene rearrangement demonstrate that HierRelTriple improves spatial-relation metrics by over 15% and substantially reduces collisions and boundary violations compared to state-of-the-art methods.
Kaifan Sun, Bingchen Yang, Peter Wonka, Jun Xiao 0005, Haiyong Jiang
IEEE Trans. Vis. Comput. Graph.4
2025 An Industrial Multi-machining Feature Dataset and Contrastive Learning-Based Network for Feature Recognition
Haochen He, Zhengda Lu, Haiyong Jiang, Yiqun Wang 0001, Jun Xiao 0005
CGI (2)6
2025 Empowering Vector Graphics with Consistently Arbitrary Viewing and View-dependent Visibility
abstract
This work presents a novel text-to-vector graphics generation approach, Dream3DVG, allowing for arbitrary viewpoint viewing, progressive detail optimization, and view-dependent occlusion awareness. Our approach is a dual-branch optimization framework, consisting of an auxiliary 3D Gaussian Splatting optimization branch and a 3D vector graphics optimization branch. The introduced 3DGS branch can bridge the domain gaps between text prompts and vector graphics with more consistent guidance. Moreover, 3DGS allows for progressive detail control by scheduling classifier-free guidance, facilitating guiding vector graphics with coarse shapes at the initial stages and finer details at later stages. We also improve the view-dependent occlusions by devising a visibility-awareness rendering module. Extensive results on 3D sketches and 3D iconographies, demonstrate the superiority of the method on different abstraction levels of details, cross-view consistency, and occlusion-aware stroke culling. Code is available at https://github.com/chenxinl/Dream3DVG.git.
Jun Xiao 0005, Zhengda Lu, Yiqun Wang 0001, Haiyong Jiang
CVPR2
2025 Activating Sparse Part Concepts for 3D Class Incremental Learning
abstract
This work tackles the challenge of 3D Class-Incremental Learning (CIL), where a model must learn to classify new 3D objects while retaining knowledge of previously learned classes. Existing methods often struggle with catastrophic forgetting, misclassifying old objects due to overreliance on shortcut local features. Our approach addresses this issue by learning a set of part concepts for part-aware features. Particularly, we only activate a small subset of part concepts for the feature representation of each part-aware feature. This facilitates better generalization across categories and mitigates catastrophic forgetting. We further improve the task-wise classification through a part relation-aware Transformer design. At last, we devise learnable affinities to fuse task-wise classification heads and avoid confusion among different tasks. We evaluate our method on three 3D CIL benchmarks, achieving state-of-the-art performance. Code is available at https://github.com/zhenyatian/ILPC.
Zhenya Tian, Jun Xiao 0005, Lupeng Liu, Haiyong Jiang
CVPR2
2025 D^3CTTA: Domain-Dependent Decorrelation for Continual Test-Time Adaption of 3D LiDAR Segmentation
abstract
Adapting pre-trained LiDAR segmentation models to dynamic domain shifts during testing is of paramount importance for the safety of autonomous driving. Most existing methods neglect the influence of domain changes and point density in continual test-time adaption (CTTA), relying on backpropagation and large batch sizes for stability. We approach this problem with three insights: 1) Point clouds at different distances usually have different densities resulting in distribution disparities; 2) The feature distribution of different domains varies, and domain-aware parameters can alleviate domain gaps; 3) Features are highly correlated and make segmentation of different labels confusing. To this end, this work presents D3CTTA, an online backpropagation-free framework for 3D continual test-time adaption for LiDAR segmentation. D3CTTA consists of a distance-aware prototype learning module to integrate LiDAR-based geometry prior and a domain-dependent decorrelation module to reduce feature correlations among different domains and different categories. Extensive experiments on three benchmarks showcase that our method achieves a state-of-the-art performance compared to both backpropagation-based methods and backpropagation-free methods. Code is available at https://github.com/ZhaoJichun1/D3CTTA.
Jichun Zhao, Haiyong Jiang, Haoxuan Song, Jun Xiao 0005, Dong Gong
CVPR4
2025 FlexFFN: Hierarchical Dynamic Selection of Feedforward Networks for Large Language Models
abstract
Optimizing the efficiency and adaptability of large language models (LLMs) for diverse downstream tasks remains a critical challenge. We propose FlexFFN (Flexible Feedforward Network), a novel framework that introduces hierarchical dynamic selection to enhance computational efficiency, flexibility, and performance in LLMs. At the macro level, FlexFFN leverages a Mixture of Experts (MoE) architecture to dynamically activate distinct FFN modules based on input characteristics. At the micro level, within each FFN module, a dynamic switching mechanism selects between KAN and traditional MLP, combining the rapid convergence capabilities of MLPs with the compositional learning and interpretability strengths of KAN. Additionally, FlexFFN integrates QLoRA (Quantized Low-Rank Adaptation) to significantly reduce memory requirements and computational costs during fine-tuning. By introducing these innovations, FlexFFN achieves a fine-grained balance between computational cost and model expressiveness, making it well-suited for large-scale training and deployment. Experimental results demonstrate that FlexFFN outperforms traditional architectures by reducing computational overhead while improving task-specific adaptability and model efficiency.
Miaobo Hu, Bokun Wang, Haoyuan Teng, Daren Zha, Xin Wang 0086, Jun Xiao 0005, Lei Wang 0135
IJCNN7
2025 DynaSketch: Abstracting Coherent Sketches for Dynamic Image Sequences
abstract
In this work, we address the problem of dynamic sketch abstraction for an image sequence without any supervised learning. Most existing works focus on sketch abstraction and generation for a static image and are not competent for dynamic sketch abstraction. This is because dynamic sketches demand coherent and seamless transitions between frames, posing unique challenges for the abstraction process. In response to these challenges, we introduce DynSketch, a novel solution capable of generating sketch sequences from video sequences, maintaining semantic alignment, and ensuring coherent frame transitions. While previous works achieve single-frame semantic alignment between sketch and image through the semantic mapping of CLIP model, we extend it to dynamic sketch and video alignment by formulating a global optimization problem. In particular, based on temporal correspondences along frames for both image and sketch, we search for more accurate and coherent tracking of semantics by strokes than single-frame abstraction. We also leverage a rigidity prior for each sketch stroke among nearby frames to ensure the overall smoothness of dynamic sketches. Finally, we adaptively adjust stroke numbers to account for the dynamic emergence and disappearance of key features throughout the sequence. By optimizing per-frame sketches sequentially, our method can produce coherent dynamic sketches as illustrated in Fig. 1. Extensive experiments and visual results show that our method outperforms previous methods and is promising for creative sketch generation and editing.
Jiakai Wu, Jun Xiao 0005
IJCNN2
2025 FC-MonoDETR: A Monocular 3D Object Detection Network Based on Foreground Constraint
abstract
Estimating 3D information about objects from a single image is a challenging problem in computer vision due to the lack of multi-view information for depth estimation. The Transformer-based methods propagate the target's 3D center depth within its 2D bounding box to construct object-level depth labels. By treating the 2D box area as a unified entity, these methods can perform sufficient feature sampling and parsing within the aforementioned area. This rough but robust detection strategy effectively avoids the dependence of 3D detection on accurate depth estimation of the target center and a few key points nearby. However, due to the lack of effective foreground constraints, these methods struggle to ensure that sampling points are located inside the target, while exterior points lack valid depth values for supervision, which affects both detection accuracy and stability. To address this issue, we propose a Foreground-Constrained Monocular 3D Object Detector (FC-MonoDETR). First, we leverage 2D annotations to generate target segmentation masks using the Segment Anything Model (SAM), directly establishing depth supervision under foreground constraints. Second, we design an attention-based feature fusion module that utilizes contour information to refine visual features and emphasizes the role of foreground regions in depth estimation, guiding the network to focus more effectively on the foreground during holistic 3D information parsing. Finally, we model the relative depth relationships between targets and optimize the estimation of the target's center depth through a specially designed target center depth loss function. Considering the stability issues of Transformer-based methods, we recommend using a more comprehensive evaluation strategy. The sufficient migration experiments have verified the effectiveness of our constructed foreground-constrained depth supervision and feature fusion module in optimizing Transformer-based methods.
Daifeng Xiao, Dongbo Yu, Yunbiao Wang, Jun Xiao 0005, Ying Wang 0030, Lupeng Liu
ICMR4
2025 GraphSplat: Sparse-View Generalizable 3D Gaussian Splatting is Worth Graph of Nodes
abstract
Generalizable 3D Gaussian Splatting (G-3DGS) has recently emerged as a promising solution for efficient 3D scene representation and novel view synthesis. However, sparse-view scenarios pose a critical challenge for accurate depth estimation. In such cases, viewpoint overlaps are minimal, and many regions are visible from only a single view. As a result, reliable multi-view matching is unavailable in these areas, leading to significant reconstruction quality degradation. To tackle this bottleneck, we propose GraphSplat, a feed-forward framework for novel view synthesis that dynamically incorporates both cross-view and monocular cues through a graph-based feature aggregation strategy. Central to our approach is a Multi-view Aggregate Graph Attention (MAGA) mechanism, which adaptively reweights intra-view and inter-view node connections to compensate for unreliable multi-view correspondences with robust single-view depth priors. In addition, we design a Hierarchical Depth Fusion Estimator (HDFE) module to integrate monocular and multi-view depth cues, effectively reducing ghosting artifacts and improving geometric consistency. Extensive evaluations on RealEstate10K and ACID benchmarks show that GraphSplat achieves competitive performance against prior SOTA methods, with improvements in appearance fidelity and cross-dataset generalization particularly under challenging sparse-view conditions.
Zeyang Bai, Yunbiao Wang, Dongbo Yu, Jun Xiao 0005, Lupeng Liu
ACM Multimedia4
2025 SegGraph: Leveraging Graphs of SAM Segments for Few-Shot 3D Part Segmentation
abstract
This work presents a novel framework for few-shot 3D part segmentation. Recent advances have demonstrated the significant potential of 2D foundation models for low-shot 3D part segmentation. However, it is still an open problem that how to effectively aggregate 2D knowledge from foundation models to 3D. Existing methods either ignore geometric structures for 3D feature learning or neglects the high-quality grouping clues from SAM, leading to under-segmentation and inconsistent part labels. We devise a novel SAM segment graph-based propagation method, named SegGraph, to explicitly learn geometric features encoded within SAM's segmentation masks. Our method encodes geometric features by modeling mutual overlap and adjacency between segments while preserving intra-segment semantic consistency. We construct a segment graph, conceptually similar to an atlas, where nodes represent segments and edges capture their spatial relationships (overlap/adjacency). Each node adaptively modulates 2D foundation model features, which are then propagated via a graph neural network to learn global geometric structures. To enforce intra-segment semantic consistency, we map segment features to 3D points with a novel view-direction-weighted fusion attenuating contributions from low-quality segments. Extensive experiments on PartNet-E demonstrate that our method outperforms all competing baselines by at least 6.9% mIoU. Further analysis reveals that SegGraph achieves particularly strong performance on small components and part boundaries, demonstrating its superior geometric understanding.
Yueyang Hu, Haiyong Jiang, Haoxuan Song, Jun Xiao 0005, Hao Pan 0001
NeurIPS4
2025 LSF-Animation: Label-Free Speech-Driven Facial Animation via Implicit Feature Representation
abstract
Speech-driven 3D facial animation has attracted increasing interest since its potential to generate expressive and temporally synchronized digital humans. While recent works have begun to explore emotion-aware animation, they still depend on explicit one-hot encodings to represent identity and emotion with given emotion and identity labels, which limits their ability to generalize to unseen speakers. Moreover, the emotional cues inherently present in speech are often neglected, limiting the naturalness and adaptability of generated animations. In this work, we propose LSF-Animation, a novel framework that eliminates the reliance on explicit emotion and identity feature representations. Specifically, LSF-Animation implicitly extracts emotion information from speech and captures the identity features from a neutral facial mesh, enabling improved generalization to unseen speakers and emotional states without requiring manual labels. Furthermore, we introduce a Hierarchical Interaction Fusion Block (HIFB), which employs a fusion token to integrate dual transformer features and more effectively integrate emotional, motion-related and identity-related cues. Extensive experiments conducted on the 3DMEAD dataset demonstrate that our method surpasses recent state-of-the-art approaches in terms of emotional expressiveness, identity generalization, and animation realism. The source code will be released at: https://github.com/Dogter521/LSF-Animation.
Chuanqing Zhuang, Chenxi Jin, Zhengda Lu, Yiqun Wang 0001, Wu Liu 0005, Jun Xiao 0005
SIGGRAPH Asia7
2025 Parallel Control With Event-Based Adaptive Critic Implementation for Robust Optimal Tracking of Uncertain Nonlinear Systems
abstract
This paper investigates event-based robust optimal parallel tracking control for a class of uncertain nonlinear systems via adaptive dynamic programming (ADP). Analysis reveals that optimal control of the nominal system with sufficient feedback gain leads to robust optimal control. Then, optimal control is implemented online employing a critic neural network (NN), and event-triggered mechanisms (ETMs) are explored to update the weights intermittently. Among them, a unique dynamic event-triggered mechanism (DETM), which releases data in the light of an auxiliary term designed for stability verification, merits significant focus, and the comparison emphasizes its potential for better handling practical control challenges. Finally, experimental findings highlight the feasibility of the proposed robust control method while validating the characteristics and superiority of the ETMs. Note to Practitioners—As a significant issue in practical applications, event-based robust optimal control motivates this study. Distinguished from the previous efforts, this study develops the practical control difficulty into virtual space via parallel control and proportionally increases the nominal system’s optimal control rule for robust optimal control. The relaxation of both the system’s prior knowledge and the assumption related to the input dynamics allows the proposed method to be more compatible with practical systems. Moreover, ETMs are adopted to better balance control performance and resource occupation, with segmented DETM providing a fresh insight to effectively address practical control issues, i.e., releasing more data to prevent the system from being damaged by the disturbance while invoking less data to improve resource efficiency when the system is stable. Finally, the experimental comparisons validate the theoretical results.
Shanshan Jiao 0001, Qinglai Wei, Jun Xiao 0005
IEEE Trans Autom. Sci. Eng.3
2025 Robust Vegetation Filtering for Rock-Mass Scene Point Clouds via Bidirectional Mamba and Adaptive Triplane Feature Representation
abstract
The irregular and chaotically interwoven distributions of vegetation and rock mass in natural environments pose significant challenges for vegetation filtering of rock point clouds, as existing methods struggle with inefficient contextual modeling and insufficient feature discriminability. To address these limitations, we propose a robust vegetation filtering method for rock mass scene point clouds featuring two key innovations: (1) Bidirectional Point Cloud Mamba module that adopts an alternating allocation strategy to divide the sampled superpoints into two complementary groups, and constructs optimized sequences from edge to center for each group via the Traveling Salesman Problem (TSP) algorithm to enable spatially continuous context capture while eliminating dependency on regular geometric priors; (2) Structure-based adaptive Tri-Plane feature representation includes Principal Component Analysis (PCA)-based support plane selection, projective 2D feature mapping, and learnable multi-scale aggregation. By exploiting the distribution differences of rock mass and vegetation structure information in 2D projections while emulating human visual perception mechanisms that rely on optimal viewing perspectives and different scales, this module enables more discriminative feature extraction in complex natural rock-mass scenarios. Extensive comparative and ablation experiments on real-world datasets demonstrate that our method achieves superior accuracy, robustness, and generalization, and excels in preserving intricate details compared to existing methods.
Shuaichen Guo, Lupeng Liu, Daifeng Xiao, Wenniu Zhang, Ying Wang 0030, Jun Xiao 0005, Dongbo Yu
IEEE Trans. Geosci. Remote. Sens.6
2025 Effective Spatial-Spectral Feature Representation for Hyperspectral Image Classification
abstract
Hyperspectral image (HSI), with their rich spectral information and spatial details, have demonstrated significant potential for classification tasks in fields such as remote sensing, agriculture, and environmental monitoring. However, existing methods still exhibit limitations in feature representation, primarily manifested in insufficient contextual modeling and the inability to effectively address spectral redundancy and significant variations in spatial scales. To address these challenges, this paper proposes an enhanced MambaHSI-based framework for HSI classification, focusing on improving the representation capability of spatial-spectral features. The proposed method consists of three key innovations: (1) A hierarchical DualGroupMamba module that progressively models intra-group and inter-group spectral dependencies to enhance fine-grained spectral discrimination and global contextual awareness; (2) a lightweight Hyperspectral Channel Attention (HCA) that dynamically adjusts the importance of feature channels based on the spatial–spectral information of different bands, effectively suppressing redundant information and highlighting discriminative features. (3) a Hybrid Feature Enhancer (HFE) module that effectively represents and fuses multi-scale spatial features by extracting local texture details and perceiving the overall spatial distribution of scenes, thereby enhancing the model’s adaptability to complex spatial structures. Through a systematic evaluation on four benchmark hyperspectral datasets, the proposed method achieved an average overall classification accuracy of 95.64%, outperforming the current best method by 1.89%. The experimental results validate the superior performance of the proposed approach in enhancing the representation of spatial–spectral features. The latest logs are now available at https://github.com/Tomyaya/EFR.
Dongbo Yu, Yunbiao Wang, Ying Wang 0030, Jun Xiao 0005, Lupeng Liu
IEEE Trans. Geosci. Remote. Sens.5
2025 MambaHSI+: Multidirectional State Propagation for Efficient Hyperspectral Image Classification
abstract
Hyperspectral image classification faces significant challenges due to high-dimensional spectral redundancy and complex spatial-spectral dependencies. While existing MambaHSI models leverage state-space modeling to enhance representation learning, their unidirectional formulation fails to fully capture bidirectional spatial interactions and cross-band contextual dependencies. Moreover, the uniform projection mechanism struggles to effectively distinguish spectral variations across different wavelengths. To address these limitations, we propose MambaHSI+, a novel framework that integrates bidirectional state-space modeling with spectral trajectory learning. The proposed architecture introduces three key innovations: (1) A bidirectional context modeling module, enhanced by reverse-order modeling, enables multi-directional spatial information propagation through recursive state transitions, facilitating comprehensive local-global feature aggregation while maintaining linear computational efficiency; (2) a spectral trajectory learning paradigm that formulates spectral evolution as continuous state-space process with bidirectional propagation, effectively encoding cross-band relationships; and (3) a Mamba-enhanced channel attention mechanism that adaptively emphasizes discriminative spectral features via selective state-space transformations. Extensive experiments on four benchmark datasets demonstrate state-of-the-art performance, achieving an average accuracy improvement of 2.51% over MambaHSI. By integrating state-space systems with advanced Mamba mechanisms, MambaHSI+ establishes a new paradigm for spectral-spatial representation learning, significantly advancing classification accuracy while ensuring computational efficiency. The latest logs are now available at https://github.com/RockAilab/MambaHSI_Plus.
Yunbiao Wang, Lupeng Liu, Jun Xiao 0005, Dongbo Yu, Wenniu Zhang
IEEE Trans. Geosci. Remote. Sens.3
2025 BGPSeg: Boundary-Guided Primitive Instance Segmentation of Point Clouds
abstract
Point cloud primitive instance segmentation is critical for understanding the geometric shapes of man-made objects. Existing learning-based methods mainly focus on learning high-dimensional feature representations of points and further perform clustering or region growing to obtain corresponding primitive instances. However, these features generally cannot accurately represent the discriminability between instances, especially near the boundaries or in regions with small differences in geometric properties. This limitation often leads to over- or under-segmentation of geometric primitives. On the other hand, the boundaries of different primitives are the direct features that distinguish them and thus utilizing boundary information to guide feature learning and clustering is crucial for this task. In this paper, we propose a novel framework BGPSeg for point cloud primitive instance segmentation that utilizes boundary-guided feature extraction and clustering. Specifically, we first introduce a boundary-guided feature extractor with the additional input of a boundary probability map, which utilizes boundary-guided sampling and a boundary transformer to enhance feature discrimination among points crossing geometric boundaries. Furthermore, we propose a boundary-guided primitive clustering module, which combines boundary clues and geometric feature discrimination for clustering to further improve the segmentation performance. Finally, we demonstrate the effectiveness of our BGPSeg with a series of comparison and ablation experiments while achieving the state-of-the-art primitive instance segmentation. Our code is available at https://github.com/fz-20/BGPSeg.
Chuanqing Zhuang, Zhengda Lu, Yiqun Wang 0001, Lupeng Liu, Jun Xiao 0005
IEEE Trans. Image Process.6
2025 Multi-Scale Spatio-Temporal Attention Network for Epileptic Seizure Prediction
abstract
Epilepticseizure prediction from electroencephalogram (EEG) data has attracted much attention in the clinical diagnosis and treatment of epilepsy. Most of the existing methods in literature extract either spatial or temporal features at a single scale from EEG data, however, their learned features are generally less discriminative since the EEG data is complex and severely noisy in general, leading to low-accuracy predictions. To address this problem, we propose a Multi-scale Spatio-temporal Attention Network to learn discriminative features for seizure prediction, called MSAN, which contains a backbone module, a spatial pyramid module, and a multi-scale sequential aggregation module. The backbone module is to extract initial spatial features from the input EEG spectrograms, and the pyramid module is introduced to learn multi-scale features from the initial features. Then by taking these multi-scale features as input temporal features, the sequential aggregation module employs multiple Long Short-Term Memory(LSTM) blocks to aggregate these features. In addition, a dual-loss function is introduced to alleviate the class imbalance problem. The proposed method achieves an average sensitivity of 96.27% with a mean false prediction rate of 0.00/h on the CHB-MIT dataset and an average sensitivity of 93.57% with a mean false prediction rate of 0.044/h on the Kaggle dataset. The comparative results demonstrate that the proposed method outperforms 10 state-of-the-art epileptic seizure prediction models.
Qiulei Dong, Han Zhang 0059, Jun Xiao 0005, Jiayin Sun
IEEE J. Biomed. Health Informatics3
2025 PS-CAD: Local Geometry Guidance via Prompting and Selection for CAD Reconstruction
abstract
Reverse engineering CAD models from raw geometry is a classic but challenging research problem. In particular, reconstructing the CAD modeling sequence from point clouds provides great interpretability and convenience for editing. Analyzing previous work, we observed that a CAD modeling sequence represented by tokens and processed by a generative model does not have an immediate geometric interpretation. To improve upon this problem, we introduce geometric guidance into the reconstruction network. Our proposed model, PS-CAD, reconstructs the CAD modeling sequence one step at a time as illustrated in Figure 1 . At each step, we provide three forms of geometric guidance. First, we provide the geometry of surfaces where the current reconstruction differs from the complete model as a point cloud. This helps the framework to focus on regions that still need work. Second, we use geometric analysis to extract a set of planar prompts, that correspond to candidate surfaces where a CAD extrusion step could be started. Third, we present a step-wise sampling to generate multiple complete candidate CAD modeling steps instead of single-tokens without direct geometric interpretation. Our framework has three major components. Geometric guidance computation extracts the first two types of geometric guidance. Single-step reconstruction computes a single candidate CAD modeling step for each provided prompt. Single-step selection selects among the candidate CAD modeling steps. The process continues until the reconstruction is completed. Our quantitative results show a significant improvement across all metrics. For example, on the dataset DeepCAD, PS-CAD improves upon the best published SOTA method by reducing the geometry errors (CD and HD) by 10%, and the structural error (ECD metric) by about 13%.
Bingchen Yang, Haiyong Jiang, Hao Pan 0001, Guosheng Lin, Jun Xiao 0005, Peter Wonka
ACM Trans. Graph.5
2025 Exploring Structural Lines for Interior Floorplan Segmentation
Bingchen Yang, Haiyong Jiang, Zhengda Lu, Jun Xiao 0005
Vis. Comput.4
2024 Fully Data-Driven Pseudo Label Estimation for Pointly-Supervised Panoptic Segmentation
abstract
The core of pointly-supervised panoptic segmentation is estimating accurate dense pseudo labels from sparse point labels to train the panoptic head. Previous works generate pseudo labels mainly based on hand-crafted rules, such as connecting multiple points into polygon masks, or assigning the label information of labeled pixels to unlabeled pixels based on the artificially defined traversing distance. The accuracy of pseudo labels is limited by the quality of the hand-crafted rules (polygon masks are rough at object contour regions, and the traversing distance error will result in wrong pseudo labels). To overcome the limitation of hand-crafted rules, we estimate pseudo labels with a fully data-driven pseudo label branch, which is optimized by point labels end-to-end and predicts more accurate pseudo labels than previous methods. We also train an auxiliary semantic branch with point labels, it assists the training of the pseudo label branch by transferring semantic segmentation knowledge through shared parameters. Experiments on Pascal VOC and MS COCO demonstrate that our approach is effective and shows state-of-the-art performance compared with related works. Codes are available at https://github.com/BraveGroup/FDD.
Jing Li 0112, Junsong Fan, Yuran Yang, Shuqi Mei, Jun Xiao 0005, Zhaoxiang Zhang 0001
AAAI5
2024 World to Code: Multi-modal Data Generation via Self-Instructed Compositional Captioning and Filtering
abstract
Recent advances in Vision-Language Models (VLMs) and the scarcity of high-quality multimodal alignment data have inspired numerous researches on synthetic VLM data generation.The conventional norm in VLM data construction uses a mixture of specialists in caption and OCR, or stronger VLM APIs and expensive human annotation.In this paper, we present World to Code (W2C), a meticulously curated multi-modal data construction pipeline that organizes the final generation output into a Python code format.The pipeline leverages the VLM itself to extract cross-modal information via different prompts and filter the generated outputs again via a consistency filtering strategy.Experiments have demonstrated the high quality of W2C by improving various existing visual question answering and visual grounding benchmarks across different VLMs.Further analysis also demonstrates that the new code parsing ability of VLMs presents better crossmodal equivalence than the commonly used detail caption ability.Our code is available at https://github.com/foundation-multimodal- models/World2Code.
Jiacong Wang, Bohong Wu, Haiyong Jiang, Haoyuan Guo, Jun Xiao 0005
EMNLP7
2024 FC-4DFS: Frequency-controlled Flexible 4D Facial Expression Synthesizing
abstract
4D facial expression synthesizing is a critical problem in the fields of computer vision and graphics. Current methods lack flexibility and smoothness when simulating the inter-frame motion of expression sequences. In this paper, we propose a frequency-controlled 4D facial expression synthesizing method, FC-4DFS. Specifically, we introduce a frequency-controlled LSTM network to generate 4D facial expression sequences frame by frame from a given neutral landmark with a given length. Meanwhile, we propose a temporal coherence loss to enhance the perception of temporal sequence motion and improve the accuracy of relative displacements. Furthermore, we designed a Multi-level Identity-Aware Displacement Network based on a cross-attention mechanism to reconstruct the 4D facial expression sequences from landmark sequences. Finally, our FC-4DFS achieves flexible and SOTA generation results of 4D facial expression sequences with different lengths on CoMA and Florence4D datasets. The code will be available on GitHub.
Chuanqing Zhuang, Zhengda Lu, Yiqun Wang 0001, Jun Xiao 0005
ACM Multimedia5
2024 DepthGAN: GAN-based depth generation from semantic layouts
abstract
Existing GAN-based generative methods are typically used for semantic image synthesis. We pose the question of whether GAN-based architectures can generate plausible depth maps and find that existing methods have difficulty in generating depth maps which reasonably represent 3D scene structure due to the lack of global geometric correlations. Thus, we propose DepthGAN, a novel method of generating a depth map using a semantic layout as input to aid construction, and manipulation of well-structured 3D scene point clouds. Specifically, we first build a feature generation model with a cascade of semantically-aware transformer blocks to obtain depth features with global structural information. For our semantically aware transformer block, we propose a mixed attention module and a semantically aware layer normalization module to better exploit semantic consistency for depth features generation. Moreover, we present a novel semantically weighted depth synthesis module, which generates adaptive depth intervals for the current scene. We generate the final depth map by using a weighted combination of semantically aware depth weights for different depth ranges. In this manner, we obtain a more accurate depth map. Extensive experiments on indoor and outdoor datasets demonstrate that DepthGAN achieves superior results both quantitatively and visually for the depth generation task.
Jun Xiao 0005, Yiqun Wang 0001, Zhengda Lu
Comput. Vis. Media2
2024 PartCom: Part Composition Learning for 3D Open-Set Recognition
Tingyu Weng, Jun Xiao 0005, Hao Pan 0001, Haiyong Jiang
Int. J. Comput. Vis.2
2024 Observer-Based Optimal Backstepping Security Control for Nonlinear Systems Using Reinforcement Learning Strategy
abstract
This article considers an observer-based optimal backstepping security control for nonlinear systems using reinforcement learning (RL) strategy. The main challenge faced is the design of optimal contoller under the deception attacks. Therefore, this article introduces an improved security RL algorithm based on neural network technology under the design framework of critic-actor to resist attacks and optimize the entire system. Second, compared with some existing results, how to relax the general assumption about deception attack is also a difficult research topic. In this article, an unusual observer that uses the attacked system output is designed to estimate the real unavailable states caused by deception attacks, so that the impact of deception attacks is eliminated and the output feedback control is also achieved. By selecting the virtual controllers and the real controller as corresponding optimized controllers within the framework of the RL algorithm, the control strategy can ensure that all signals in the closed-loop system are semi-globally ultimately bounded. Finally, two simulation experiments will be run to demonstrate the effectiveness of the strategy.
Qinglai Wei, Xiangmin Tan, Jun Xiao 0005, Qi Dong 0005
IEEE Trans. Cybern.4
2024 Isoperimetric Constraint Inference for Discrete-Time Nonlinear Systems Based on Inverse Optimal Control
abstract
In this article, the problem of inferring unknown isoperimetric constraints is considered given optimal state and control trajectories that solve the optimal control problem with isoperimetric constraints. By exploiting Pontryagin's principle, the recovery equations for unknown isoperimetric constraints are established. Under verifiable dimensionality condition and matrix rank condition, the proposed method is guaranteed to infer the unknown isoperimetric constraints exactly. Furthermore, the proposed method is extended to multiple trajectory setting. Finally, the effectiveness of the proposed method is illustrated by two simulation examples with various settings.
Qinglai Wei, Tao Li 0058, Jie Zhang 0116, Hongyang Li 0002, Xin Wang 0137, Jun Xiao 0005
IEEE Trans. Cybern.6
2024 Efficient Single Correspondence Voting for Point Cloud Registration
abstract
3D point cloud registration is a crucial task in a variety of fields, including remote sensing mapping, computer vision, virtual reality, and autonomous driving. However, this task is still challenging due to the challenges of noise, non-uniformity, partial overlap, and repeated local features in large scene point clouds. In this paper, we propose an efficient single correspondence voting method for large scene point cloud registration. Specifically, we first propose an efficient hypothetical transformation prediction method called SCVC, which determines the 5 degrees of freedom of the transformation through one correspondence, and then uses Hough voting to determine the last degree of freedom. This algorithm can significantly improve the accuracy of registration in both indoor and outdoor scenes. On the other hand, we propose a more robust transformation verification function called VDIR, which can obtain the optimal registration result of two raw point clouds. Finally, we conduct a series of experiments that demonstrate that our method achieves state-of-the-art performance on four real-world datasets: 3DMatch, 3DLoMatch, KITTI, and WHU-TLS. Our code is available at https://github.com/xingxuejun1989/SCVC.
Xuejun Xing, Zhengda Lu, Yiqun Wang 0001, Jun Xiao 0005
IEEE Trans. Image Process.4
2024 Accurate Lung Nodule Segmentation With Detailed Representation Transfer and Soft Mask Supervision
abstract
Accurate lung lesion segmentation from computed tomography (CT) images is crucial to the analysis and diagnosis of lung diseases, such as COVID-19 and lung cancer. However, the smallness and variety of lung nodules and the lack of high-quality labeling make the accurate lung nodule segmentation difficult. To address these issues, we first introduce a novel segmentation mask named " soft mask," which has richer and more accurate edge details description and better visualization, and develop a universal automatic soft mask annotation pipeline to deal with different datasets correspondingly. Then, a novel network with detailed representation transfer and soft mask supervision (DSNet) is proposed to process the input low-resolution images of lung nodules into high-quality segmentation results. Our DSNet contains a special detailed representation transfer module (DRTM) for reconstructing the detailed representation to alleviate the small size of lung nodules images and an adversarial training framework with soft mask for further improving the accuracy of segmentation. Extensive experiments validate that our DSNet outperforms other state-of-the-art methods for accurate lung nodule segmentation, and has strong generalization ability in other accurate medical segmentation tasks with competitive results. Besides, we provide a new challenging lung nodules segmentation dataset for further studies (https://drive.google.com/file/d/15NNkvDTb_0Ku0IoPsNMHezJRTH1Oi1wm/view?usp=sharing).
Changwei Wang 0001, Rongtao Xu, Shibiao Xu, Weiliang Meng, Jun Xiao 0005, Xiaopeng Zhang 0001
IEEE Trans. Neural Networks Learn. Syst.5
2024 A Self-Attention-Based Deep Reinforcement Learning Approach for AGV Dispatching Systems
abstract
The automated guided vehicle (AGV) dispatching problem is to develop a rule to assign transportation tasks to certain vehicles. This article proposes a new deep reinforcement learning approach with a self-attention mechanism to dynamically dispatch the tasks to AGV. The AGV dispatching system is modeled as a less complicated Markov decision process (MDP) using vehicle-initiated rules to dispatch a workcenter to an idle AGV. In order to deal with the highly dynamical environment, the self-attention mechanism is introduced to calculate the importance of different information. The invalid action masking technique is performed to alleviate false actions. A multimodal structure is employed to mix the features of various sources. Comparative experiments are performed to show the effectiveness of the proposed method. The properties of the learned policies are also investigated under different environment settings. It is discovered that the policies explore and learn the properties of different systems, and also smooth the traffic congestion. Under certain environment settings, the policy converges to a heuristic rule that assigns the idle AGV to the workcenter with the shortest queue length, which shows the adaptiveness of the proposed method.
Qinglai Wei, Yutian Yan, Jie Zhang 0116, Jun Xiao 0005
IEEE Trans. Neural Networks Learn. Syst.4
2023 Unsupervised 3D Articulated Object Correspondences with Part Approximation and Shape Refinement
Junqi Diao, Haiyong Jiang, Feilong Yan, Jinhui Luan, Jun Xiao 0005
CAD/Graphics6
2023 VectorFloorSeg: Two-Stream Graph Attention Network for Vectorized Roughcast Floorplan Segmentation
abstract
Vector graphics (VG) are ubiquitous in industrial designs. In this paper, we address semantic segmentation of a typical VG, i.e., roughcast floorplans with bare wall structures, whose output can be directly used for further applications like interior furnishing and room space modeling. Previous semantic segmentation works mostly process well-decorated floorplans in raster images and usually yield aliased boundaries and outlier fragments in segmented rooms, due to pixel-level segmentation that ignores the regular elements (e.g. line segments) in vector floor-plans. To overcome these issues, we propose to fully utilize the regular elements in vector floorplans for more integral segmentation. Our pipeline predicts room segmentation from vector floorplans by dually classifying line segments as room boundaries, and regions partitioned by line segments as room segments. To fully exploit the structural relationships between lines and regions, we use two-stream graph neural networks to process the line segments and partitioned regions respectively, and devise a novel modulated graph attention layer to fuse the heterogeneous information from one stream to the other. Extensive experiments show that by directly operating on vector floorplans, we outper-form image-based methods in both mIoU and mAcc. In addition, we propose a new metric that captures room integrity and boundary regularity, which confirms that our method produces much more regular segmentations. Source code is available at https://github.com/DrZiji/VecFloorSeg.
Bingchen Yang, Haiyong Jiang, Hao Pan 0001, Jun Xiao 0005
CVPR4
2023 Informative Data Mining for One-shot Cross-Domain Semantic Segmentation
abstract
Contemporary domain adaptation offers a practical solution for achieving cross-domain transfer of semantic segmentation between labelled source data and unlabeled target data. These solutions have gained significant popularity; however, they require the model to be retrained when the test environment changes. This can result in unbearable costs in certain applications due to the time-consuming training process and concerns regarding data privacy. One-shot domain adaptation methods attempt to overcome these challenges by transferring the pre-trained source model to the target domain using only one target data. Despite this, the referring style transfer module still faces issues with computation cost and over-fitting problems. To address this problem, we propose a novel framework called Informative Data Mining (IDM) that enables efficient one-shot domain adaptation for semantic segmentation. Specifically, IDM provides an uncertainty-based selection criterion to identify the most informative samples, which facilitates quick adaptation and reduces redundant training. We then perform a model adaptation method using these selected samples, which includes patch-wise mixing and prototype-based information maximization to update the model. This approach effectively enhances adaptation and mitigates the overfitting problem. In general, we provide empirical evidence of the effectiveness and efficiency of IDM. Our approach outperforms existing methods and achieves a new state-of-the-art one-shot performance of 56.7%/55.4% on the GTA5/SYNTHIA to Cityscapes adaptation tasks, respectively. The code will be released at https://github.com/yxiwang/IDM.
Yuxi Wang 0001, Jian Liang 0001, Jun Xiao 0005, Shuqi Mei, Yuran Yang, Zhaoxiang Zhang 0001
ICCV3
2023 SSF: Accelerating Training of Spiking Neural Networks with Stabilized Spiking Flow
abstract
Surrogate gradient (SG) is one of the most effective approaches for training spiking neural networks (SNNs). While assisting SNNs to achieve classification performance comparable to artificial neural networks, SG suffers from the problem of time-consuming training, preventing it from efficient learning. In this paper, we formally analyze the backward process of classic SG and find that the membrane accumulation through time leads to exponential growth of training time. With this discovery, we propose Stabilized Spiking Flow (SSF), a simple yet effective approach to accelerate training of SG-based SNNs. For each spiking neuron, SSF averages its input and output activations over time to yield stabilized input and output, respectively. Then, instead of back propagating all errors that are related to current neuron and inherently entangled in time domain, the auxiliary gradient is directly propagated from the stabilized output to input through a devised relationship mapping. Additionally, SSF method is suitable to different neuron models. Extensive experiments on both static and neuromorphic datasets demonstrate that SNNs trained with SSF approach can achieve performance comparable to the original counterparts, while reducing the training time significantly. In particular, SSF speeds up the training process of state-of-the-art SNN models up to 10× when time steps equal to 80.
Zengjie Song, Yuxi Wang 0001, Jun Xiao 0005, Yuran Yang, Shuqi Mei, Zhaoxiang Zhang 0001
ICCV4
2023 Decompose Novel into Known: Part Concept Learning For 3D Novel Class Discovery
abstract
In this work, we address 3D novel class discovery (NCD) that discovers novel classes from an unlabeled dataset by leveraging the knowledge of disjoint known classes. The key challenge of 3D NCD is that learned features by known class recognition are heavily biased and hinder generalization to novel classes. Since geometric parts are more generalizable across different classes, we propose to decompose novel into known parts, coined DNIK, to mitigate the above problems. DNIK learns a part concept bank encoding rich part geometric patterns from known classes so that novel 3D shapes can be represented as part concept compositions to facilitate cross-category generalization. Moreover, we formulate three constraints on part concepts to ensure diverse part concepts without collapsing. A part relation encoding module (PRE) is also developed to leverage part-wise spatial relations for better recognition. We construct three 3D NCD tasks for evaluation and extensive experiments show that our method achieves significantly superior results than SOTA baselines (+11.7%, +14.1%, and +16.3% improvements on average for three tasks, respectively). Code and data will be released.
Tingyu Weng, Jun Xiao 0005, Haiyong Jiang
NeurIPS2
2023 Combating Spurious Correlations in Loose-fitting Garment Animation Through Joint-Specific Feature Learning
abstract
Abstract We address the 3D animation of loose‐fitting garments from a sequence of body motions. State‐of‐the‐art approaches treat all body joints as a whole to encode motion features, which usually gives rise to learned spurious correlations between garment vertices and irrelevant joints as shown in Fig. 1. To cope with the issue, we encode temporal motion features in a joint‐wise manner and learn an association matrix to map human joints only to most related garment regions by encouraging its sparsity. In this way, spurious correlations are mitigated and better performance is achieved. Furthermore, we devise the joint‐specific pose space deformation (PSD) to decompose the high‐dimensional displacements as the combination of dynamic details caused by individual joint poses. Extensive experiments show that our method outperforms previous works in most indicators. Moreover, garment animations are not interfered with by artifacts caused by spurious correlations, which further validates the effectiveness of our approach. The code is available at https://github.com/qiji77/JointNet .
Junqi Diao, Jun Xiao 0005, Yihong He, Haiyong Jiang
Comput. Graph. Forum2
2023 Joint specular highlight detection and removal in single images via Unet-Transformer
abstract
Specular highlight detection and removal is a fundamental problem in computer vision and image processing. In this paper, we present an efficient end-to-end deep learning model for automatically detecting and removing specular highlights in a single image. In particular, an encoder—decoder network is utilized to detect specular highlights, and then a novel Unet-Transformer network performs highlight removal; we append transformer modules instead of feature maps in the Unet architecture. We also introduce a highlight detection module as a mask to guide the removal task. Thus, these two networks can be jointly trained in an effective manner. Thanks to the hierarchical and global properties of the transformer mechanism, our framework is able to establish relationships between continuous self-attention layers, making it possible to directly model the mapping between the diffuse area and the specular highlight area, and reduce indeterminacy within areas containing strong specular highlight reflection. Experiments on public benchmark and real-world images demonstrate that our approach outperforms state-of-the-art methods for both highlight detection and removal tasks.
Zhongqi Wu, Jianwei Guo 0003, Chuanqing Zhuang, Jun Xiao 0005, Dong-Ming Yan 0001, Xiaopeng Zhang 0001
Comput. Vis. Media4
2023 RC-Net: Row and Column Network with Text Feature for Parsing Floor Plan Images
Weiliang Meng, Zhengda Lu, Jianwei Guo 0003, Jun Xiao 0005, Xiaopeng Zhang 0001
J. Comput. Sci. Technol.5
2023 SPDET: Edge-Aware Self-Supervised Panoramic Depth Estimation Transformer With Spherical Geometry
abstract
Panoramic depth estimation has become a hot topic in 3D reconstruction techniques with its omnidirectional spatial field of view. However, panoramic RGB-D datasets are difficult to obtain due to the lack of panoramic RGB-D cameras, thus limiting the practicality of supervised panoramic depth estimation. Self-supervised learning based on RGB stereo image pairs has the potential to overcome this limitation due to its low dependence on datasets. In this work, we propose the SPDET, an edge-aware self-supervised panoramic depth estimation network that combines the transformer with a spherical geometry feature. Specifically, we first introduce the panoramic geometry feature to construct our panoramic transformer and reconstruct high-quality depth maps. Furthermore, we introduce the pre-filtered depth-image-based rendering method to synthesize the novel view image for self-supervision. Meanwhile, we design an edge-aware loss function to improve the self-supervised depth estimation for panorama images. Finally, we demonstrate the effectiveness of our SPDET with a series of comparison and ablation experiments while achieving the state-of-the-art self-supervised monocular panoramic depth estimation. Our code and models are available at https://github.com/zcq15/SPDET.
Chuanqing Zhuang, Zhengda Lu, Yiqun Wang 0001, Jun Xiao 0005, Ying Wang 0030
IEEE Trans. Pattern Anal. Mach. Intell.4
2023 NR-MVSNet: Learning Multi-View Stereo Based on Normal Consistency and Depth Refinement
abstract
Multi-view Stereo (MVS) aims to reconstruct a 3D point cloud model from multiple views. In recent years, learning-based MVS methods have received a lot of attention and achieved excellent performance compared with traditional methods. However, these methods still have apparent shortcomings, such as the accumulative error in the coarse-to-fine strategy and the inaccurate depth hypotheses based on the uniform sampling strategy. In this paper, we propose the NR-MVSNet, a coarse-to-fine structure with the depth hypotheses based on the normal consistency (DHNC) module, and the depth refinement with reliable attention (DRRA) module. Specifically, we design the DHNC module to generate more effective depth hypotheses, which collects the depth hypotheses from neighboring pixels with the same normals. As a result, the predicted depth can be smoother and more accurate, especially in texture-less and repetitive-texture regions. On the other hand, we update the initial depth map in the coarse stage by the DRRA module, which can combine attentional reference features and cost volume features to improve the depth estimation accuracy in the coarse stage and address the accumulative error problem. Finally, we conduct a series of experiments on the DTU, BlendedMVS, Tanks & Temples, and ETH3D datasets. The experimental results demonstrate the efficiency and robustness of our NR-MVSNet compared with the state-of-the-art methods. Our implementation is available at https://github.com/wdkyh/NR-MVSNet.
Jingliang Li, Zhengda Lu, Yiqun Wang 0001, Jun Xiao 0005, Ying Wang 0030
IEEE Trans. Image Process.4
2023 Context-Aware 3D Point Cloud Semantic Segmentation With Plane Guidance
abstract
Point cloud segmentation is fundamental in under- standing 3D environments. However, most existing methods usually perform poorly on identifying boundaries of touching objects and large surfaces of objects. Planes in a scene usually act as supporting surfaces to separate touching objects and provide geometry priors to group points on a large surface as shown in Fig. 1. Besides, planes can roughly represent the structure of a scene, and are more efficient to encode holistic scene contexts than large scale point clouds. In light of the above advantages, we advise a plane-assisted module, coined3D-PAM, to enhance semantic segmentation of touching objects and large surface objects.3D-PAMconsists of a plane separation network (PS-Net) and a plane relation network (PR-Net).PS-Netfocuses on learning features that can robustly separate touching objects, e.g., a chair on a floor, as well as capture plane-based geometry priors to group points on a large plane, e.g., points of a desk.PR-Netencodes mutual plane relations as a proxy of a scene structure to capture holistic contexts.3D-PAMis designed as a plug-and-play module so that it can be easily plugged into any off-the-shelf semantic segmentation network. Extensive experiments demonstrate that the method achieves large segmentation improvements on several backbones, and accomplishes superior results on most categories when using a RandLA-Net backbone ($11/13$categories on S3DIS dataset and$15/20$categories on ScanNetv2 dataset). The project is available at GitHubhttps://github.com/windmillknight/Context-Aware-3D-Point-Cloud-Semantic-Segmentation-With-Plane-Guidance
Tingyu Weng, Jun Xiao 0005, Feilong Yan, Haiyong Jiang
IEEE Trans. Multim.2
2023 Continuous-Time Stochastic Policy Iteration of Adaptive Dynamic Programming
abstract
In this article, we study the optimal control problem of continuous-time (CT) time-invariant nonlinear systems with stochastic nonlinear disturbances. A new stochastic adaptive dynamic programming (ADP) method is developed to solve the Hamilton–Jacobi–Bellman equation (HJBE). Under the conditional expectation, the value function and the control law are successively approximated simultaneously. The asymptotic stability of the closed-loop stochastic system in probability is analyzed by the stochastic Lyapunov direct method, and the convergence of the developed ADP method is given. Finally, four simulations illustrate the effectiveness of the developed method.
Qinglai Wei, Tianmin Zhou, Jingwei Lu, Yu Liu 0078, Shuai Su, Jun Xiao 0005
IEEE Trans. Syst. Man Cybern. Syst.6
2022 ACDNet: Adaptively Combined Dilated Convolution for Monocular Panorama Depth Estimation
abstract
Depth estimation is a crucial step for 3D reconstruction with panorama images in recent years. Panorama images maintain the complete spatial information but introduce distortion with equirectangular projection. In this paper, we propose an ACDNet based on the adaptively combined dilated convolution to predict the dense depth map for a monocular panoramic image. Specifically, we combine the convolution kernels with different dilations to extend the receptive field in the equirectangular projection. Meanwhile, we introduce an adaptive channel-wise fusion module to summarize the feature maps and get diverse attention areas in the receptive field along the channels. Due to the utilization of channel-wise attention in constructing the adaptive channel-wise fusion module, the network can capture and leverage the cross-channel contextual information efficiently. Finally, we conduct depth estimation experiments on three datasets (both virtual and real-world) and the experimental results demonstrate that our proposed ACDNet substantially outperforms the current state-of-the-art (SOTA) methods. Our codes and model parameters are accessed in https://github.com/zcq15/ACDNet.
Chuanqing Zhuang, Zhengda Lu, Yiqun Wang 0001, Jun Xiao 0005, Ying Wang 0030
AAAI4
2022 Softgan: Towards Accurate Lung Nodule Segmentation via Soft Mask Supervision
abstract
Accurate lung nodule segmentation from Computed Tomog-raphy (CT) images is crucial to the analysis and diagnosis of lung diseases such as COVID-19 and lung cancer. How-ever, due to the variety of lung nodules and the lack of high-quality labeling, accurate lung nodule segmentation is still a challenging problem. In this paper, we propose a novel paradigm including an automatic accurate annotation pipeline and a segmentation network for this task. First, we introduce a new segmentation mask representation named Soft Mask which has richer and more accurate edge details description and better visualization, and we design a universal automatic Soft Mask annotation pipeline to deal with different datasets. Besides, we provide a new challenging lung nodules segmen-tation dataset with traditional binarized masks and our soft masks for further studies. Second, we propose an effective network called SoftGAN that includes an improved back-bone and an adversarial training framework with Soft Mask, in order to improve the performance of accurate lung nodules segmentation. Extensive experiments validate that our Soft-GAN outperforms the state-of-the-art methods for accurate lung nodule segmentation. [Datasetrelease]
Changwei Wang 0001, Rongtao Xu, Shibiao Xu, Weiliang Meng, Jun Xiao 0005, Qimin Peng, Xiaopeng Zhang 0001
ICME5
2022 DS-MVSNet: Unsupervised Multi-view Stereo via Depth Synthesis
abstract
In recent years, supervised or unsupervised learning-based MVS methods achieved excellent performance compared with traditional methods. However, these methods only use the probability volume computed by cost volume regularization to predict reference depths and this manner cannot mine enough information from the probability volume. Furthermore, the unsupervised methods usually try to use two-step or additional inputs for training which make the procedure more complicated. In this paper, we propose the DS-MVSNet, an end-to-end unsupervised MVS structure with the source depths synthesis. To mine the information in probability volume, we creatively synthesize the source depths by splattering the probability volume and depth hypotheses to source views. Meanwhile, we propose the adaptive Gaussian sampling and improved adaptive bins sampling approach that improve the depths hypotheses accuracy. On the other hand, we utilize the source depths to render the reference images and propose depth consistency loss and depth smoothness loss. These can provide additional guidance according to photometric and geometric consistency in different views without additional inputs. Finally, we conduct a series of experiments on the DTU dataset and Tanks $&$ Temples dataset that demonstrate the efficiency and robustness of our DS-MVSNet compared with the state-of-the-art methods.
Jingliang Li, Zhengda Lu, Yiqun Wang 0001, Ying Wang 0030, Jun Xiao 0005
ACM Multimedia5
2022 Dynamic video mix-up for cross-domain action recognition
Chunfeng Song, Shaolong Yue, Zhenyu Wang 0012, Jun Xiao 0005, Yanyang Liu
Neurocomputing5
2022 A Novel Rock-Mass Point Cloud Registration Method Based on Feature Line Extraction and Feature Point Matching
abstract
Registration will directly affect the quality of overall rock-mass point cloud, which is the basis of 3-D reconstruction for rock mass. Advanced methods establish correspondence by extracting various features that remain unchanged. Although these methods have made great progress, they analyze the local characteristics of each sample point, which leads to be inefficient. In this article, we select registration interesting points from feature lines that were extracted based on supervoxel and innovatively introduce the “clustering, primary matching, and coarse registration” strategy, which effectively reduces the complexity of calculating the corresponding relationship during point cloud registration. Finally, the iterative closest point (ICP) algorithm is used to optimize the result of coarse registration. By selecting registration interesting points from the extracted feature lines, the proposed method inherits the robustness of feature lines to noise, initial position, and so on. The experimental results prove that the coarse registration and refined registration results of the proposed method both have high accuracy and efficiency.
Lupeng Liu, Jun Xiao 0005, Yunbiao Wang, Zhengda Lu, Ying Wang 0030
IEEE Trans. Geosci. Remote. Sens.2
2022 Registration Method for Point Clouds of Complex Rock Mass Based on Dual Structure Information
abstract
Obtaining complete point cloud data is the basis of rock surface segmentation and related rock-mass numerical simulation. The existing rock-mass point cloud registration methods usually extract point features as the key information for registration. However, the calculation results of point features may be affected by many factors such as data integrity, point density, and noise, which limits the application of existing algorithms in some complex rock scenes (complex collection conditions and complex surface structures). In this article, we propose to divide the rock-mass registration task into two stages: global matching and local matching. During global matching, we extract structure-level features (interrelationships between planes) and shape features (point distribution information in a specific region) instead of point features as the basis for establishing preliminary correspondence between point clouds, achieving robust and efficient region-to-region matching. In the local matching stage, the method based on feature point extraction and matching is proposed to establish accurate point-to-point correspondences in the local region, thus effectively solving the influence of the error of structure-level feature matching on the registration accuracy. In this article, the registration accuracy and efficiency of our method are fully validated by using both repository datasets and real-scene datasets. The robustness of the method is also demonstrated under various conditions. The experimental results show that the root-mean-square error of this method is less than 0.05 m when dealing with mountain data with a length greater than 200 m, which is obviously better than the existing best method.
Dongbo Yu, Jun Xiao 0005, Ying Wang 0030
IEEE Trans. Geosci. Remote. Sens.2
2022 Efficient Pairwise 3-D Registration of Urban Scenes via Hybrid Structural Descriptors
abstract
Automatic registration of point clouds captured by terrestrial laser scanning (TLS) plays an important role in many fields including remote sensing (e.g., transportation management, 3-D reconstruction in large-scale urban areas and environment monitoring), computer vision, and virtual reality and robotics. However, noise, outliers, nonuniform point density, and small overlaps are inevitable when collecting multiple views of data, which poses great challenges to 3-D registration of point clouds. Since conventional registration methods aim to find point correspondences and estimate transformation parameters directly in the original point space, the traditional way to address these difficulties is to introduce many restrictions during the scanning process (e.g., more scanning and careful selection of scanning positions), thus making the data acquisition more difficult. In this article, we present a novel 3-D registration framework that performs in a “middle-level structural space” and is capable of robustly and efficiently reconstructing urban, semiurban, and indoor scenes, despite disturbances introduced in the scanning process. The new structural space is constructed by extracting multiple types of middle-level geometric primitives (planes, spheres, cylinders, and cones) from the 3-D point cloud. We design a robust method to find effective primitive combinations corresponding to the 6-D poses of the raw point clouds and then construct hybrid-structure-based descriptors. By matching descriptors and computing rotation and translation parameters, successful registration is achieved. Note that the whole process of our method is performed in the structural space, which has the advantages of capturing geometric structures (the relationship between primitives) and semantic features (primitive types and parameters) in larger fields. Experiments show that our method achieves state-of-the-art performance in several point cloud registration benchmark datasets at different scales and even obtains good registration results for data without overlapping areas.
Jianwei Guo 0003, Zhanglin Cheng, Jun Xiao 0005, Xiaopeng Zhang 0001
IEEE Trans. Geosci. Remote. Sens.4
2022 Single-Image Specular Highlight Removal via Real-World Dataset Construction
abstract
Specular reflections pose great challenges on various multimedia and computer vision tasks,e.g., image segmentation, detection and matching. In this paper, we build a large-scale Paired Specular-Diffuse (PSD) image dataset, where the images are carefully captured by using real-world objects and the ground-truth specular-free diffuse images are provided. To the best of our knowledge, this is the first real-world benchmark dataset for specular highlight removal task, which is useful for evaluating and encouraging new deep learning-based approaches. Given this dataset, we present a novel Generative Adversarial Network (GAN) for specular highlight removal from a single image by introducing the detection of specular reflection information as a guidance. Our network also makes full use of the attention mechanism and is able to directly model the mapping relation between the diffuse area and the specular highlight area without any explicit estimation of the illumination. Experimental results demonstrate that the proposed network is more effective to remove specular reflection components with the guidance of specular highlight detection than recent state-of-the-art methods.
Zhongqi Wu, Chuanqing Zhuang, Jianwei Guo 0003, Jun Xiao 0005, Xiaopeng Zhang 0001, Dong-Ming Yan 0001
IEEE Trans. Multim.5
2022 Model-Free Adaptive Optimal Control for Unknown Nonlinear Multiplayer Nonzero-Sum Game
abstract
In this article, an online adaptive optimal control algorithm based on adaptive dynamic programming is developed to solve the multiplayer nonzero-sum game (MP-NZSG) for discrete-time unknown nonlinear systems. First, a model-free coupled globalized dual-heuristic dynamic programming (GDHP) structure is designed to solve the MP-NZSG problem, in which there is no model network or identifier. Second, in order to relax the requirement of systems dynamics, an online adaptive learning algorithm is developed to solve the Hamilton-Jacobi equation using the system states of two adjacent time steps. Third, a series of critic networks and action networks are used to approximate value functions and optimal policies for all players. All the neural network (NN) weights are updated online based on real-time system states. Fourth, the uniformly ultimate boundedness analysis of the NN approximation errors is proved based on the Lyapunov approach. Finally, simulation results are given to demonstrate the effectiveness of the developed scheme.
Qinglai Wei, Liao Zhu, Ruizhuo Song, Pinjia Zhang, Derong Liu 0001, Jun Xiao 0005
IEEE Trans. Neural Networks Learn. Syst.6
2022 Blending Surface Segmentation and Editing for 3D Models
abstract
Recognizing and fitting shape primitives from underlying 3D models are key components of many computer graphics and computer vision applications. Although a vast number of structural recovery methods are available, they usually fail to identify blending surfaces, which corresponds to small transitional regions among relatively large primary patches. To address this issue, we present a novel approach for automatic segmentation and surface fitting with accurate geometric parameters from 3D models, especially mechanical parts. Overall, we formulate the structural segmentation as a Markov random field (MRF) labeling problem. In contrast to existing techniques, we first propose a new clustering algorithm to build superfacets by incorporating 3D local geometric information. This algorithm extracts the general quadric and rolling-ball blending regions, and improves the robustness of further segmentation. Next, we apply a specially designed MRF framework to efficiently partition the original model into different meaningful patches of known surface types by defining the multilabel energy function on the superfacets. Furthermore, we present an iterative optimization algorithm based on skeleton extraction to fit rolling-ball blending patches by recovering the parameters of the rolling center trajectories and ball radius. Experiments on different complex models demonstrate the effectiveness and robustness of the proposed method, and the superiority of our method is also verified through comparisons with state-of-the-art approaches. We further apply our algorithm in applications such as mesh editing by changing the radius of the rolling balls.
Jianwei Guo 0003, Jun Xiao 0005, Xiaopeng Zhang 0001, Dong-Ming Yan 0001
IEEE Trans. Vis. Comput. Graph.3
2021 Neighborhood-based Neural Implicit Reconstruction from Point Clouds
abstract
Neural implicit reconstruction is emerging as a promising approach to constructing 3D geometry from point clouds due to its ability to model geometry with complicated topology and unrestricted resolution. Current methods in this category usually deliver smooth and good quality results, but suffer from defective details and generalization issues. The major reason is that these methods use either a global code or interpolated feature on 3D grids of limited resolution to estimate implicit surface, therefore may cause distortion in feature discretization. This paper presents a neighborhood-aware neural implicit reconstruction framework that consists of an encoder network, a feature aggregation module, and a decoder network to learn implicit surface. The method can easily incorporate an off-the-shelf 3D point-based or volume-based neural network as an encoder. At the heart of our framework is the aggregation module that fuses the learnt contextual features on neighbor inputs so that the method can directly exploit local features of neighboring inputs for geometry detail recovery as well as cross-domain generalization. Experimental results demonstrate that our method significantly outperforms the state-of-the-art methods (about 4.0 points IoU improvements in ShapeNet dataset and 9.0 points IoU improvements in DFAUST dataset). Furthermore, our method preserves finer shape details and can be successfully transferred to a novel category without fine-tuning.
Haiyong Jiang, Jianfei Cai 0001, Jianmin Zheng, Jun Xiao 0005
3DV4
2021 Extracting Cycle-aware Feature Curve Networks from 3D Models
Zhengda Lu, Jianwei Guo 0003, Jun Xiao 0005, Ying Wang 0030, Xiaopeng Zhang 0001, Dong-Ming Yan 0001
Comput. Aided Des.3
2021 Data-driven floor plan understanding in rural residential buildings via deep recognition
Zhengda Lu, Jianwei Guo 0003, Weiliang Meng, Jun Xiao 0005, Xiaopeng Zhang 0001
Inf. Sci.5
2021 Accurate Rock-Mass Extraction From Terrestrial Laser Point Clouds via Multiscale and Multiview Convolutional Feature Representation
abstract
Existing 3-D object extraction methods on terrestrial laser point clouds are further developed through filtering and labeling. However, such predefined features are heuristically designed to process generic object point clouds. Thus, existing abilities are insufficient to handle specific rock-mass point clouds. Given the complexity and diversity of terrestrial environments, the effective removal of vegetation points from rock-mass point clouds is particularly challenging. To address such problems, this study presents a novel approach for 3-D rock-mass point clouds labeling by using convolutional feature learning based on distribution priors with multiple scales and views. First, to extract discriminative features of each point for classification, we propose novel multiview supporting planes to analyze the spatial distribution and structure of its neighboring points for each category. Second, we define the multiscale spatial distribution matrix on a grid representation (e.g., the number of points projected into each cell). Last, the statistical information of points is nonlinearly combined and hierarchically compressed to generate a compact and effective convolutional feature representation for classification. The effectiveness of the proposed method is evaluated via experiments on rock-mass point clouds from different scenes. Compared with existing extraction approaches, experimental results indicate the superiority of the proposed method in terms of the precision and recall.
Yunbiao Wang, Shibiao Xu, Jun Xiao 0005, Ying Wang 0030, Lupeng Liu
IEEE Trans. Geosci. Remote. Sens.3
2020 CIAN: Cross-Image Affinity Net for Weakly Supervised Semantic Segmentation
abstract
Weakly supervised semantic segmentation with only image-level labels saves large human effort to annotate pixel-level labels. Cutting-edge approaches rely on various innovative constraints and heuristic rules to generate the masks for every single image. Although great progress has been achieved by these methods, they treat each image independently and do not take account of the relationships across different images. In this paper, however, we argue that the cross-image relationship is vital for weakly supervised segmentation. Because it connects related regions across images, where supplementary representations can be propagated to obtain more consistent and integral regions. To leverage this information, we propose an end-to-end cross-image affinity module, which exploits pixel-level cross-image relationships with only image-level labels. By means of this, our approach achieves 64.3% and 65.3% mIoU on Pascal VOC 2012 validation and test set respectively, which is a new state-of-the-art result by only using image-level labels for weakly supervised semantic segmentation, demonstrating the superiority of our approach.
Junsong Fan, Zhaoxiang Zhang 0001, Tieniu Tan, Chunfeng Song, Jun Xiao 0005
AAAI5
2020 End-to-End 3D Point Cloud Instance Segmentation Without Detection
abstract
3D instance segmentation plays a predominant role in environment perception of robotics and augmented reality. Many deep learning based methods have been presented recently for this task. These methods rely on either a detection branch to propose objects or a grouping step to assemble same-instance points. However, detection based methods do not ensure a consistent instance label for each point, while the grouping step requires parameter-tuning and is computationally expensive. In this paper, we introduce a novel framework to enable end-to-end instance segmentation without detection and a separate step of grouping. The core idea is to convert instance segmentation to a candidate assignment problem. At first, a set of instance candidates is sampled. Then we propose an assignment module for candidate assignment and a suppression module to eliminate redundant candidates. A mapping between instance labels and instance candidates is further sought to construct an instance grouping loss for the network training. Experimental results demonstrate that our method is more effective and efficient than previous approaches.
Haiyong Jiang, Feilong Yan, Jianfei Cai 0001, Jianmin Zheng, Jun Xiao 0005
CVPR5
2020 Efficient and automatic plane detection approach for 3-D rock mass point clouds
Jun Xiao 0005, Ying Wang 0030
Multim. Tools Appl.2
2020 An automatic 3D registration method for rock mass point clouds based on plane detection and polygon matching
Jun Xiao 0005, Ying Wang 0030
Vis. Comput.2
2019 Help LabelMe: A Fast Auxiliary Method for Labeling Image and Using It in ChangE's CCD Data
Yunfan Lu, Yifan Hu 0012, Jun Xiao 0005
ICIG (1)3
2019 A fast registration algorithm of rock point cloud based on spherical projection and feature extraction
Yaru Xian, Jun Xiao 0005, Ying Wang 0030
Frontiers Comput. Sci.2
2019 Fast and Error-Bounded Space-Variant Bilateral Filtering
Mengke Yuan, Longquan Dai, Dong-Ming Yan 0001, Liqiang Zhang 0001, Jun Xiao 0005, Xiaopeng Zhang 0001
J. Comput. Sci. Technol.5
2019 Efficient Rock-Mass Point Cloud Registration Using n-Point Complete Graphs
abstract
The surfaces of rock masses are arbitrary and complex. Moreover, the point clouds of rock-mass surfaces acquired via terrestrial laser scanning typically span large distances and have high resolutions. These characteristics cause difficulties in registration between scans. To address these difficulties, an efficient method using$n$-point complete graphs is proposed. To handle massive point clouds, a step-by-step strategy is adopted to reduce the number of points involved in the computation. First, the Gaussian curvature of each point of the initial data is estimated, and points with low Gaussian curvatures are filtered out such that only the interesting points are preserved. Second, these interesting points are clustered, and the centroid of each cluster is calculated. Finally, a descriptor is built from the$n$-point complete graph formed by each centroid and its$n-1$nearest neighbors. By matching the descriptors generated from two point clouds, corresponding point pairs can be obtained, thus achieving alignment. In addition, this strategy inherently incorporates denoising, outlier handling, and filtration, thereby endowing the method with strong adaptability to various conditions without incurring any additional cost. Experiments on data sets with varying degrees of outliers, noise and overlap were conducted to demonstrate the robustness of the proposed method. The results show that, with point span$r~\approx ~1$cm, the output root mean square error is around 0.5 cm, which is comparable with that of the Iterative Closest Point algorithm. A runtime analysis shows that the total processing time of the proposed method grows nearly linearly with increasing data size.
Jun Xiao 0005, Ying Wang 0030
IEEE Trans. Geosci. Remote. Sens.2
2018 Filtering method of rock points based on BP neural network and principal component analysis
Jun Xiao 0005, Sidong Liu, Ying Wang 0030
Frontiers Comput. Sci.1
2018 A survey on algorithms of hole filling in 3D surface reconstruction
Xiaoyuan Guo, Jun Xiao 0005, Ying Wang 0030
Vis. Comput.2
2008 Multiple Watermarking with Side Information
Jun Xiao 0005, Ying Wang 0030
IWDW1