Hongshan Yu

dblp:04/4793 · DBLP profile ↗
← Back
34ranked-venue papers
3as first author
24since 2021 · last 2027
0000-0003-1973-6766ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 21 · 3 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 11 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Systems, architecture and hardware · 2 · 2 since 2021Computer networks · 1Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2027 SFA-DiffNet: Spatial-frequency aware diffusion network for robust medical image segmentation
Tongtong Xie, Hongshan Yu, Yan Zheng 0003, Yong He 0012, Zhengeng Yang, Zechuan Li, Naveed Akhtar
Expert Syst. Appl.2
2026 P-MLP: Language-conditioned task planning with multimodal lexical priors over labels
Wei Sun 0028, Yan Zheng 0003, Jian Liu 0014, Genwei Zhang, Hongshan Yu, Ajmal Mian
Knowl. Based Syst.8
2026 Efficient Point Cloud Processing With High-Dimensional Positional Encoding and Non-Local MLPs
abstract
Multi-Layer Perceptron (MLP) models are the foundation of contemporary point cloud processing. However, their complex network architectures obscure the source of their strength and limit the application of these models. In this article, we develop a two-stage abstraction and refinement (ABS-REF) view for modular feature extraction in point cloud processing. This view elucidates that whereas the early models focused on ABS stages, the more recent techniques devise sophisticated REF stages to attain performance advantages. Then, we propose a High-dimensional Positional Encoding (HPE) module to explicitly utilize intrinsic positional information, extending the "positional encoding" concept from Transformer literature. HPE can be readily deployed in MLP-based architectures and is compatible with transformer-based methods. Within our ABS-REF view, we rethink local aggregation in MLP-based methods and propose replacing time-consuming local MLP operations, which are used to capture local relationships among neighbors. Instead, we use non-local MLPs for efficient non-local information updates, combined with the proposed HPE for effective local information representation. We leverage our modules to develop HPENets, a suite of MLP networks that follow the ABS-REF paradigm, incorporating a scalable HPE-based REF stage. Extensive experiments on seven public datasets across four different tasks show that HPENets deliver a strong balance between efficiency and effectiveness. Notably, HPENet surpasses PointNeXt, a strong MLP-based counterpart, by 1.1% mAcc, 4.0% mIoU, 1.8% mIoU and 0.2% Cls. mIoU, with only 50.0%, 21.5%, 23.1%, 44.4% of FLOPs on ScanObjectNN, S3DIS, ScanNet, and ShapeNetPart, respectively.
Yanmei Zou, Hongshan Yu, Yaonan Wang 0001, Zhengeng Yang, Xieyuanli Chen, Kailun Yang 0001, Naveed Akhtar
IEEE Trans. Pattern Anal. Mach. Intell.2
2026 Soft-Masked Transformer for Point Cloud Processing With Skip Attention-Based Upsampling
abstract
Point cloud processing methods leverage local and global point features to cater to downstream tasks, yet they often overlook the task-level context inherent in point clouds during the encoding stage. We argue that integrating task-level information into the encoding stage significantly enhances performance. To that end, we propose SMTransformer which incorporates task-level information into a vector-based transformer by utilizing a soft mask generated from task-level queries and keys to learn the attention weights. Additionally, to facilitate effective communication between features from the encoding and decoding layers in high-level tasks such as segmentation, we introduce a skip-attention-based up-sampling block. This block dynamically fuses features from various resolution points across the encoding and decoding layers. To mitigate the increase in network parameters and training time resulting from the complexity of the aforementioned blocks, we propose a novel shared point position encoding strategy. This strategy allows various transformer blocks to share the same position information over the same resolution points, thereby reducing network parameters and training time without compromising accuracy. Experimental comparisons with existing methods on multiple datasets demonstrate the efficacy of SMTransformer and skip-attention-based up-sampling for semantic segmentation task. In particular, we achieve state-of-the-art semantic segmentation results of 73.9% mIoU on S3DIS Area 5 and 62.4% mIoU on SWAN dataset. Note to Practitioners—Point cloud processing underpins automation tasks such as robotic perception, navigation, and inspection, where accurate 3D understanding is essential. Existing methods often prioritize vision benchmarks while overlooking automation needs like efficiency on limited hardware and robustness in real-world environments. The proposed SMTransformer embeds task-level guidance into feature learning and employs skip-attention up-sampling to improve segmentation accuracy with practical efficiency. It is well-suited for robotic manipulation, autonomous driving, and inspection applications. Current limitations include reliance on GPUs and sensitivity to extreme density variations. Future work will target edge-device deployment and multi-task extensions.
Yong He 0012, Hongshan Yu, Chaoxu Mu, Mingtao Feng, Tongjia Chen, Zechuan Li, Anwaar Ulhaq, Ajmal Mian
IEEE Trans Autom. Sci. Eng.2
2026 Multi-Modal Shape Encoding for 3-D Object Detection
Zechuan Li, Hongshan Yu, Niu Zhang, Jinhao Qiao, Wei Sun 0028, Naveed Akhtar
IEEE Trans Autom. Sci. Eng.2
2026 LPL3D: LVLM-Driven Pseudo-Labeling for 3D Object Detection
abstract
Effective 3D object detection requires large-scale annotated datasets, which are expensive and time-consuming to produce - especially in indoor environments containing dense object arrangements. To address this, we propose an Large Vision-Language Model (LVLM)-driven automatic high-quality pseudo-label generation technique for 3D object detection in single- and multi-view scenarios. We propose an Auto3DLabeler that introduces the first-ever text-to-3D Bounding Box transformation. Its pipeline employs a text-based detector, a segmenter and an LVLM to generate annotation estimates, which are further refined by our IoU-guided iterative Box Aggregator and layout-aware prompt Class Refiner modules. We also introduce a semantic-enhanced multi-modal fusion module that integrates image-level semantic information into point cloud representations for precise detections. Collectively, our contributions provide a remarkable boost to the 3D object detection state-of-the-art. Extensive experiments on SUN RGB-D and ScanNet datasets show our unsupervised detector variant outperforming existing semi-supervised detectors, and our semi-supervised variant achieving up to 28.2% absolute gain in challenging scenarios - all this while maintaining considerable compute advantage over existing label-efficient methods. Our code and models will be made public for the community. Our code and model will be made public after acceptance.
Zechuan Li, Hongshan Yu, Yihao Ding, Shuai Yuan 0013, Naveed Akhtar
IEEE Trans. Circuits Syst. Video Technol.2
2025 Med-SER: Enhancing Reasoning Interpretability in Medical Visual Question Answering via Structured Chain-of-Thought
abstract
Existing Medical Visual Question Answering (Med-VQA) methods typically rely on either direct answer generation or Chain-of-Thought (CoT) reasoning, both of which suffer from limited interpretability or logical inconsistency. To overcome these challenges, we propose Med-SER, a novel framework featuring Structured Chain-of-Thought (SCoT), which decomposes reasoning into four clinically grounded stages: Summary, Caption, Reasoning, and Conclusion. To facilitate training, we construct VQA-RAD-SCoT, the first Med-VQA dataset annotated with structured reasoning chains. Med-SER further introduces a Dual-Channel Visual Projection (DCVP) module to extract both holistic and stage-specific visual features, and a Dual Dynamic Supervision (DDS) mechanism combining adaptive stage-aware weighting and logical consistency loss. Experiments on VQA-RAD-SCoT demonstrate that Med-SER demonstrates the potential of Med-SER to establish a new interpretable and trustworthy paradigm for Med-VQA.
Jinhao Qiao, Yi Xiao 0004, Hongshan Yu, Yan Zheng 0003
BIBM6
2025 GO-N3RDet: Geometry Optimized NeRF-enhanced 3D Object Detector
abstract
We propose GO-N3RDet, a scene-geometry optimized multi-view 3D object detector enhanced by neural radiance fields. The key to accurate 3D object detection is in effective voxel representation. However, due to occlusion and lack of 3D information, constructing 3D features from multi-view 2D images is challenging. Addressing that, we introduce a unique 3D positional information embedded voxel optimization mechanism to fuse multi-view features. To prioritize neural field reconstruction in object regions, we also devise a double importance sampling scheme for the NeRF branch of our detector. We additionally propose an opacity optimization module for precise voxel opacity prediction by enforcing multi-view consistency constraints. Moreover, to further improve voxel density consistency across multiple perspectives, we incorporate ray distance as a weighting factor to minimize cumulative ray errors. Our unique modules synergetically form an end-to-end neural model that establishes new state-of-the-art in NeRF-based multi-view 3D detection, verified with extensive experiments on ScanNet and ARKITScenes. Code will be available at https://github.com/ZechuanLi/GO-N3RDet.
Zechuan Li, Hongshan Yu, Yihao Ding, Jinhao Qiao, Basim Azam, Naveed Akhtar
CVPR2
2025 InsCMPR: Efficient Cross-Modal Place Recognition via Instance-Aware Hybrid Mamba-Transformer
abstract
Place recognition is an important technique for autonomous mobile robotic applications. While single-modal sensor-based approaches have shown satisfactory performance, cross-modal place recognition remains underexplored due to the challenge of bridging the cross-modal heterogeneity gap. In this work, we introduce an instance-aware cross-modal place recognition approach, named InsCMPR. We design a novel instance-aware modality alignment module, which aligns multi-modal data at both pixel-level and instance-level by leveraging a pre-trained vision foundation model SAM. Then a novel dual-branch hybrid Mamba-Transformer network is proposed to efficiently enhance the distinctiveness of the produced descriptors by integrating global features with local instance features. Experimental results on the KITTI, NCLT, and HAOMO datasets show that our proposed methods achieve state-of-the-art performance while operating in real time. We will open source the implementation of our method at: https://github.com/nubot-nudt/InsCMPR.
Shuaifeng Jiao, Zhuoqun Su, Lun Luo, Hongshan Yu, Zongtan Zhou, Huimin Lu 0002, Xieyuanli Chen
ICRA4
2025 MRMT-PR: A Multi-Scale Reverse-View Mamba-Transformer for LiDAR Place Recognition
abstract
Place recognition is a fundamental technology of high relevance for autonomous robot navigation. Existing methods encounter significant challenges arising from scene variations (e.g., illumination changes, dynamic objects), view-point shifts, and difficulties in data fusion and alignment. These factors often lead to a substantial drop in recognition recall, which is typically addressed in the literature by training deep neural networks to learn invariant feature representations. In this paper, we propose MRMT-PR, a novel multi-scale reverse-view Mamba-Transformer architecture for LiDAR-based place recognition that uses a single-frame point cloud as its input. Our MRMT-PR framework consists of a multi-scale reverse-view preprocessing module for LiDAR point clouds, a Mamba-Transformer feature encoder, and a global feature fusion module. This architecture effectively mitigates the impact of perspective and illumination variations, enhances the global representational capacity of LiDAR features, and significantly improves recognition robustness under challenging conditions such as viewpoint changes and long-term localization. Experiments conducted on NCLT dataset with challenging scenarios demonstrate that MRMT-PR outperforms existing LiDAR-based place recognition baselines in terms of overall performance.
Jingwen Wang 0009, Hongshan Yu, Yaonan Wang 0001, Javier Civera 0001, Xieyuanli Chen
IROS3
2025 NPC-SPU: Nonlinear Phase Coding-Based Stereo Phase Unwrapping for Efficient 3D Measurement
abstract
3D imaging based on phase-shifting structured light is widely used in industrial measurement due to its non-contact nature. However, it typically requires a large number of additional images (multi-frequency heterodyne (M-FH) method) or introduces intensity features that compromise accuracy (space domain modulation phase-shifting (SDM-PS) method) for phase unwrapping, and it remains sensitive to motion. To overcome these issues, this article proposes a nonlinear phase coding-based stereo phase unwrapping (NPC-SPU) method that requires no additional patterns while maintaining measurement accuracy. In the encoding stage, a novel nonlinear distortion feature is introduced, while the signal-to-noise ratio of the phase codeword is preserved. In the decoding stage, a local phase unwrapping method that does not require additional auxiliary information is first proposed, closely associating the distortion information in the local wrapped phase. Then, a pre-calibrated stereo constraint system is used to filter potential matching phases, significantly reducing phase ambiguity and computational costs. Finally, to avoid the time-consuming and complex intensity kernel matching used in traditional methods, we propose a local phase correlation matching (LPCM) technique that enables lightweight and robust phase unwrapping. Experimental results demonstrate that this algorithm significantly enhances 3D reconstruction performance in scenarios with large depth, large disparity, complex colored structures, and dynamic scenes. Specifically, in dynamic environments (20mm/s), the proposed method achieves a lower measurement error rate (0.7829% vs. 6.4962%) with only 3 patterns, compared to the traditional three-frequency heterodyne (T-FH) method (using 9 patterns). Additionally, its measurement accuracy outperforms the advanced SDM-PS method, which also uses 3 patterns (0.1102 mm vs. 0.3232 mm).
Ruiming Yu, Hongshan Yu, Wei Sun 0028, Yaonan Wang 0001, Naveed Akhtar, Kemao Qian
IEEE Trans. Image Process.2
2025 Exploring Hierarchical Spatial Layout Cues for 3D Point Cloud Based Scene Graph Prediction
abstract
3D scene graph prediction is important for intelligent agents to gather information and perceive semantics of their environments. However, constructing an effective graph is nontrivial given the complexity of natural scenes. Existing solutions for graph representation of 3D scenes still distinguish each detailed discrepancy among all the relationships as flat thinking, ignoring the mechanism used by humans to perform this task. Inspired by the role of the prefrontal cortex in hierarchical reasoning, we analyze this problem from a novel perspective: exploring hierarchical spatial layout cues in 3D space and navigating that hierarchy to make the 3D scene graph more accurate in a vertical division to horizontal propagation strategy. To this end, we first encode the contextual object features for fine-gained object category classification. Next, we build a bottom-up hierarchical graph to predict remarkably diverse support relationships in a single concept regardless of numerous irrelevant relationships. Finally, equipped with the spatially-true and semantically-meaningful support relationships, we focus on the local region layout to propagate the semantic features to predict the additional non-support relationships under the guidance of the given referred hierarchical graph nodes. Experiments on the challenging 3DSSG benchmark show that our algorithm outperforms existing state-of-the-art, and can also alleviate the impact of the long-tailed distribution of training data. Our code is available athttps://github.com/HHrEtvP/HSLC-3DSG/.
Mingtao Feng, Haoran Hou, Liang Zhang 0010, Yulan Guo, Hongshan Yu, Yaonan Wang 0001, Ajmal Mian
IEEE Trans. Multim.5
2025 Full Point Encoding for Local Feature Aggregation in 3-D Point Clouds
abstract
Point cloud processing methods exploit local point features and global context through aggregation which does not explicitly model the internal correlations between local and global features. To address this problem, we propose full point encoding which is applicable to convolution and transformer architectures. Specifically, we propose full point convolution (FuPConv) and full point transformer (FPTransformer) architectures. The key idea is to adaptively learn the weights from local and global geometric connections, where the connections are established through local and global correlation functions, respectively. FuPConv and FPTransformer simultaneously model the local and global geometric relationships as well as their internal correlations, demonstrating strong generalization ability and high performance. FuPConv is incorporated in classical hierarchical network architectures to achieve local and global shape-aware learning. In FPTransformer, we introduce full point position encoding in self-attention, that hierarchically encodes each point position in the global and local receptive field. We also propose a shape-aware downsampling block that takes into account the local shape and the global context. Experimental comparison to existing methods on benchmark datasets shows the efficacy of FuPConv and FPTransformer for semantic segmentation, object detection, classification, and normal estimation tasks. In particular, we achieve state-of-the-art semantic segmentation results of 76.8% mIoU on S3DIS sixfold and 73.1% on S3DIS Area 5. Our code is available at https://github.com/hnuhyuwa/FullPointTransformer.
Yong He 0012, Hongshan Yu, Zhengeng Yang, Wei Sun 0028, Ajmal Mian
IEEE Trans. Neural Networks Learn. Syst.2
2024 Improved MLP Point Cloud Processing with High-Dimensional Positional Encoding
abstract
Multi-Layer Perceptron (MLP) models are the bedrock of contemporary point cloud processing. However, their complex network architectures obscure the source of their strength. We first develop an “abstraction and refinement” (ABS-REF) view for the neural modeling of point clouds. This view elucidates that whereas the early models focused on the ABS stage, the more recent techniques devise sophisticated REF stages to attain performance advantage in point cloud processing. We then borrow the concept of “positional encoding” from transformer literature, and propose a High-dimensional Positional Encoding (HPE) module, which can be readily deployed to MLP based architectures. We leverage our module to develop a suite of HPENet, which are MLP networks that follow ABS-REF paradigm, albeit with a sophisticated HPE based REF stage. The developed technique is extensively evaluated for 3D object classification, object part segmentation, semantic segmentation and object detection. We establish new state-of-the-art results of 87.6 mAcc on ScanObjectNN for object classification, and 85.5 class mIoU on ShapeNetPart for object part segmentation, and 72.7 and 78.7 mIoU on Area-5 and 6-fold experiments with S3DIS for semantic segmentation. The source code for this work is available at https://github.com/zouyanmei/HPENet.
Yanmei Zou, Hongshan Yu, Zhengeng Yang, Zechuan Li, Naveed Akhtar
AAAI2
2024 OST: Refining Text Knowledge with Optimal Spatio-Temporal Descriptor for General Video Recognition
abstract
Due to the resource-intensive nature of training vision- language models on expansive video data, a majority of studies have centered on adapting pre-trained image- language models to the video domain. Dominant pipelines propose to tackle the visual discrepancies with additional temporal learners while overlooking the substantial discrepancy for web-scaled descriptive narratives and concise action category names, leading to less distinct semantic space and potential performance limitations. In this work, we prioritize the refinement of text knowledge to facilitate generalizable video recognition. To address the limitations of the less distinct semantic space of category names, we prompt a large language model (LLM) to augment action class names into Spatio-Temporal Descriptors thus bridging the textual discrepancy and serving as a knowledge base for general recognition. Moreover, to assign the best descriptors with different video instances, we propose Optimal Descriptor Solver, forming the video recognition problem as solving the optimal matching flow across frame-level representations and descriptors. Comprehensive evaluations in zero-shot, few-shot, and fully supervised video recognition highlight the effectiveness of our approach. Our best model achieves a state-of-the-art zero-shot accuracy of 75.1% on Kinetics-600.
Tom Tongjia Chen, Hongshan Yu, Zhengeng Yang, Zechuan Li, Wei Sun 0028, Chen Chen 0001
CVPR2
2024 VAG: Voxel Attenuation Grid For Sparse-View CBCT Reconstruction
abstract
Sparse view CBCT reconstruction has become one of the important research fields to reduce the radiation impact of CT scanning. However, the reconstruction of high-quality 3D CT volumes from sparse and noisy CBCT data still faces challenges such as slow convergence, long computation time, and increased noise. In light of these issues, we propose a voxel attenuation grid representation to explicitly model the attenuation field of the 3D CT volume. Since this representation does not involve the implementation of neural networks, our method for reconstruction is extremely fast. Furthermore, trim regularization and total variation regularization terms are introduced on top of the mean square error loss to optimize the voxel attenuation grid and significantly reduce the noise in the reconstructed 3D CT volume. Experiments on the NSCLC dataset demonstrate the superiority of our method and its potential in clinical applications. Our code will be available at: https://github.com/qiaodongxing/VAG.
Jinhao Qiao, Yi Xiao 0004, Hongshan Yu, Yan Zheng 0003
ICIP5
2024 EFRNet-VL: An end-to-end feature refinement network for monocular visual localization in dynamic environments
Jingwen Wang 0009, Hongshan Yu, Xuefei Lin, Zechuan Li, Wei Sun 0028, Naveed Akhtar
Expert Syst. Appl.2
2024 LPL-VIO: monocular visual-inertial odometry with deep learning-based point and line features
Changxiang Liu, Qinhan Yang, Hongshan Yu, Qiang Fu 0013, Naveed Akhtar
Neural Comput. Appl.3
2024 Domain-Invariant Prototypes for Semantic Segmentation
abstract
Deep learning has greatly advanced the performance of semantic segmentation, however, its success relies on the availability of large amounts of annotated data for training. Hence, many efforts have been devoted to domain adaptive semantic segmentation that focuses on transferring semantic knowledge from a labeled source domain to an unlabeled target domain. Existing self-training methods typically require multiple rounds of training, while another popular framework based on adversarial training is known to be sensitive to hyper-parameters. We propose an easy-to-train framework that learns domain-invariant prototypes for domain adaptive semantic segmentation. In particular, we show that domain adaptation shares a common character with few-shot learning in that both aim to recognize some types of unseen data with knowledge learned from large amounts of seen data. Thus, we propose a unified framework for domain adaptation and few-shot learning. The core idea is to use the class prototypes extracted from few-shot annotated target images to classify pixels of both source images and target images. Our method involves only one-stage training and does not need to be trained on large-scale un-annotated target images. Moreover, our method can be extended to variants of both domain adaptation and few-shot learning. Competitive performances achieved on GTA5-to-Cityscapes and SYNTHIA-to-Cityscapes adaptation tasks show the effectiveness of the proposed novel while simple domain adaptation framework. The source code used in this paper is available at https://github.com/zgyang-hnu/DIP-hunnu.
Zhengeng Yang, Hongshan Yu, Wei Sun 0028, Li Cheng 0001, Ajmal Mian
IEEE Trans. Circuits Syst. Video Technol.2
2024 Fully Convolutional Network-Based Self-Supervised Learning for Semantic Segmentation
abstract
Although deep learning has achieved great success in many computer vision tasks, its performance relies on the availability of large datasets with densely annotated samples. Such datasets are difficult and expensive to obtain. In this article, we focus on the problem of learning representation from unlabeled data for semantic segmentation. Inspired by two patch-based methods, we develop a novel self-supervised learning framework by formulating the jigsaw puzzle problem as a patch-wise classification problem and solving it with a fully convolutional network. By learning to solve a jigsaw puzzle comprising 25 patches and transferring the learned features to semantic segmentation task, we achieve a 5.8% point improvement on the Cityscapes dataset over the baseline model initialized from random values. It is noted that we use only about 1/6 training images of Cityscapes in our experiment, which is designed to imitate the real cases where fully annotated images are usually limited to a small number. We also show that our self-supervised learning method can be applied to different datasets and models. In particular, we achieved competitive performance with the state-of-the-art methods on the PASCAL VOC2012 dataset using significantly fewer time costs on pretraining.
Zhengeng Yang, Hongshan Yu, Yong He 0012, Wei Sun 0028, Zhi-Hong Mao, Ajmal Mian
IEEE Trans. Neural Networks Learn. Syst.2
2023 AShapeFormer : Semantics-Guided Object-Level Active Shape Encoding for 3D Object Detection via Transformers
abstract
3D object detection techniques commonly follow a pipeline that aggregates predicted object central point features to compute candidate points. However, these candidate points contain only positional information, largely ignoring the object-level shape information. This eventually leads to sub-optimal 3D object detection. In this work, we propose AShapeFormer, a semantics-guided object-level shape encoding module for 3D object detection. This is a plug-n-play module that leverages multi-head attention to encode object shape information. We also propose shape tokens and object-scene positional encoding to ensure that the shape information is fully exploited. Moreover, we introduce a semantic guidance sub-module to sample more foreground points and suppress the influence of background points for a better object shape perception. We demonstrate a straightforward enhancement of multiple existing methods with our AShapeFormer. Through extensive experiments on the popular SUN RGB-D and ScanNetV2 dataset, we show that our enhanced models are able to outperform the baselines by a considerable absolute margin of up to 8.1%. Code will be available at https://github.com/ZechuanLi/AShapeFormer
Zechuan Li, Hongshan Yu, Zhengeng Yang, Tom Tongjia Chen, Naveed Akhtar
CVPR2
2022 Learning from Pixel-Level Noisy Label : A New Perspective for Light Field Saliency Detection
abstract
Saliency detection with light field images is becoming attractive given the abundant cues available, however, this comes at the expense of large-scale pixel level annotated data which is expensive to generate. In this paper, we propose to learn light field saliency from pixel-level noisy labels obtained from unsupervised hand crafted featured-based saliency methods. Given this goal, a natural question is: can we efficiently incorporate the relationships among light field cues while identifying clean labels in a unified framework? We address this question by formulating the learning as a joint optimization of intra light field features fusion stream and inter scenes correlation stream to generate the predictions. Specially, we first introduce a pixel forgetting guided fusion module to mutually enhance the light field features and exploit pixel consistency across iterations to identify noisy pixels. Next, we introduce a cross scene noise penalty loss for better reflecting latent structures of training data and enabling the learning to be invariant to noise. Extensive experiments on multiple benchmark datasets demonstrate the superiority of our framework showing that it learns saliency prediction comparable to state-of-the-art fully supervised light field saliency methods. Our code is available at h t tps://github.com/ OLobbCode/NoiseLF.
Mingtao Feng, Kendong Liu, Liang Zhang 0010, Hongshan Yu, Yaonan Wang 0001, Ajmal Mian
CVPR4
2022 Fast ORB-SLAM Without Keypoint Descriptors
abstract
Indirect methods for visual SLAM are gaining popularity due to their robustness to environmental variations. ORB-SLAM2 (Mur-Artal and Tardós, 2017) is a benchmark method in this domain, however, it consumes significant time for computing descriptors that never get reused unless a frame is selected as a keyframe. To overcome these problems, we present FastORB-SLAM which is light-weight and efficient as it tracks keypoints between adjacent frames without computing descriptors. To achieve this, a two stage descriptor-independent keypoint matching method is proposed based on sparse optical flow. In the first stage, we predict initial keypoint correspondences via a simple but effective motion model and then robustly establish the correspondences via pyramid-based sparse optical flow tracking. In the second stage, we leverage the constraints of the motion smoothness and epipolar geometry to refine the correspondences. In particular, our method computes descriptors only for keyframes. We test FastORB-SLAM on TUM and ICL-NUIM RGB-D datasets and compare its accuracy and efficiency to nine existing RGB-D SLAM methods. Qualitative and quantitative results show that our method achieves state-of-the-art accuracy and is about twice as fast as the ORB-SLAM2.
Qiang Fu 0013, Hongshan Yu, Xiaolong Wang 0005, Zhengeng Yang, Yong He 0012, Hong Zhang 0013, Ajmal Mian
IEEE Trans. Image Process.2
2021 NDNet: Narrow While Deep Network for Real-Time Semantic Segmentation
abstract
The rapid development of autonomous driving in recent years presents many challenges for scene understanding. As an essential step towards scene understanding, semantic segmentation has received increased attention in the past few years. Although deep learning based approaches have achieved great success in improving the segmentation accuracy, most of them suffer from an inefficiency problem and can hardly be applied to real-time applications. In this paper, we analyze the computational cost of Convolutional Neural Network (CNN) and find that the inefficiency of CNNs is mainly caused by their wide structure rather than deep structure. In addition, the success of pruning based model compression methods proves that there are many redundant channels in CNNs. Thus, we design a narrow while deep backbone network to improve the efficiency of semantic segmentation. By casting our network to fully convolutional network (FCN32) segmentation architecture, the basic structure of most segmentation methods, we achieve 61.5% mIoU on Cityscapes validation dataset with only 4.2G floating-point operations (FLOPs) on 1024×2048 inputs, which already outperforms one of the earliest real-time deep learning based segmentation methods: ENet (58.3% mIoU, 3.8G FLOPs on 640×360 inputs). By further refining the output resolution of our network to the 1/8 of the input resolution with a simple encoder-decoder structure, we achieve 65.3% mIoU on Cityscapes test set with 14.0G FLOPs and 39.9 frames per second (FPS) on Titan X card. We have made our model publicly available at https://github.com/zgyang-hnu/NDNet.
Zhengeng Yang, Hongshan Yu, Qiang Fu 0013, Wei Sun 0028, Wenyan Jia, Mingui Sun, Zhi-Hong Mao
IEEE Trans. Intell. Transp. Syst.2
2020 A multirobot target searching method based on bat algorithm in unknown environments
Hongwei Tang, Wei Sun 0028, Hongshan Yu, Anping Lin
Expert Syst. Appl.3
2020 Small Object Augmentation of Urban Scenes for Real-Time Semantic Segmentation
abstract
Semantic segmentation is a key step in scene understanding for autonomous driving. Although deep learning has significantly improved the segmentation accuracy, current highquality models such as PSPNet and DeepLabV3 are inefficient given their complex architectures and reliance on multi-scale inputs. Thus, it is difficult to apply them to real-time or practical applications. On the other hand, existing real-time methods cannot yet produce satisfactory results on small objects such as traffic lights, which are imperative to safe autonomous driving. In this paper, we improve the performance of real-time semantic segmentation from two perspectives, methodology and data. Specifically, we propose a real-time segmentation model coined Narrow Deep Network (NDNet) and build a synthetic dataset by inserting additional small objects into the training images. The proposed method achieves 65.7% mean intersection over union (mIoU) on the Cityscapes test set with only 8.4G floatingpoint operations (FLOPs) on 1024×2048 inputs. Furthermore, by re-training the existing PSPNet and DeepLabV3 models on our synthetic dataset, we obtained an average 2% mIoU improvement on small objects.
Zhengeng Yang, Hongshan Yu, Mingtao Feng, Wei Sun 0028, Xuefei Lin, Mingui Sun, Zhi-Hong Mao, Ajmal Mian
IEEE Trans. Image Process.2
2019 A novel hybrid algorithm based on PSO and FOA for target searching in unknown environments
Hongwei Tang, Wei Sun 0028, Hongshan Yu, Anping Lin, Yuxue Song
Appl. Intell.3
2019 Locate the Mobile Device by Enhancing the WiFi-Based Indoor Localization Model
abstract
Due to the advent and pervasive deployment of wireless local area networks, WiFi-based indoor localization systems have received increasing attention in the last few years. However, their localization accuracy has always been a challenging issue. In addition, because of diverse interference such as multipath effects, the block of signals, an unstable or weak signal in itself, etc., not all the access points (APs) are informative for the localization. Faced with these problems, we propose a WiFi-based localization model by modifying the large localization errors and enhancing the Gaussian process regression (MEGPR). 1) To select the AP subsets that contribute more to the localization and further reduce the computational load, the AP discrimination criterion (APDC) is introduced to quantify the discernibility of the APs detected in the workspace and filter out the APs with low discrimination. 2) Second, to enhance the localization model, the localization residual is fed and learnt by the model. 3) Furthermore, the large localization errors are mitigated by the location modification method (LMM). Experiments were conducted in a real environment with an area of more than 1200 m2and the results show that compared with other existing localization models, the average localization error of the proposed MEGPR model is minimum, which further verifies the effectiveness of the proposed MEGPR localization model.
Wei Sun 0028, Hongshan Yu, Hongwei Tang, Anping Lin, Roger Zimmermann
IEEE Internet Things J.3
2018 Methods and datasets on semantic segmentation: A review
Hongshan Yu, Zhengeng Yang, Yaonan Wang 0001, Wei Sun 0028, Mingui Sun, Yandong Tang
Neurocomputing1
2017 All-dimension neighborhood based particle swarm optimization with randomly selected neighbors
Wei Sun 0028, Anping Lin, Hongshan Yu, Qiaokang Liang, Guohua Wu 0001
Inf. Sci.3
2012 Robot Navigation Based on Fuzzy Behavior Controller
Hongshan Yu, Yaonan Wang 0001
ISNN (2)1
2011 Tracking Multiple Persons Based on Attributed Relational Graph
abstract
The appearance model is very effective in tracking multiple persons. The main difficulty in tracking persons is to represent appearance reliably and effectively, especially in the presence of occlusions. In this paper, an effective Attributed Relational Graph (ARG) based tracking algorithm is presented to track multiple persons even under occlusions. The appearance of each person is expressed by an ARG model which not only combines color feature with spatial information but also illustrates the relations among body parts. The similarity of ARG models is computed to build a matching matrix in consecutive frames. Four tracking situations are determined according to the matching matrix. In addition, to track persons under occlusions, probabilistic relaxation labeling in the ARG models of body parts is deduced to label occluded persons optimally. Experimental validation of the proposed tracking method is verified and presented on indoor and outdoor sequences.
Qin Wan 0001, Yaonan Wang 0001, Hongshan Yu, Xiaofang Yuan
Int. J. Pattern Recognit. Artif. Intell.3
2007 Neural Network-Based Robust Tracking Control for Nonholonomic Mobile Robot
Jinzhu Peng, Yaonan Wang 0001, Hongshan Yu
ISNN (1)3
2007 An Occupancy Grids Building Method with Sonar Sensors Based on Improved Neural Network Model
Hongshan Yu, Yaonan Wang 0001, Jinzhu Peng
ISNN (1)1