EDBT 2026 Demo / reviewers in the wild / expert
Wei Sun 0028
dblp:09/5042-28
· DBLP profile ↗
41ranked-venue papers
4as first author
31since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 21 · 3 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 1 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 5 since 2021Computer networks · 5 · 4 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Deep Learning-Based Object Pose Estimation: A Comprehensive Survey
Jian Liu 0014, Wei Sun 0028, Chongpei Liu, Hossein Rahmani 0001, Nicu Sebe, Ajmal Mian |
Int. J. Comput. Vis. | 2 |
| 2026 | P-MLP: Language-conditioned task planning with multimodal lexical priors over labels
Wei Sun 0028, Yan Zheng 0003, Jian Liu 0014, Genwei Zhang, Hongshan Yu, Ajmal Mian |
Knowl. Based Syst. | 2 |
| 2026 | Multi-Modal Shape Encoding for 3-D Object Detection
Zechuan Li, Hongshan Yu, Niu Zhang, Jinhao Qiao, Wei Sun 0028, Naveed Akhtar |
IEEE Trans Autom. Sci. Eng. | 5 |
| 2026 | CT-UIO: Continuous-Time UWB-Inertial-Odometer Localization Using Non-Uniform B-Spline With Fewer AnchorsabstractUltra-wideband (UWB) based positioning with fewer anchors has attracted significant research interest in recent years, especially under energy-constrained conditions. However, most existing methods rely on discrete-time representations and smoothness priors to infer a robot's motion states, which often struggle with ensuring multi-sensor data synchronization. In this article, we present a continuous-time UWB-Inertial-Odometer localization system (CT-UIO), utilizing a non-uniform B-spline framework with fewer anchors. Unlike traditional uniform B-spline-based continuous-time methods, we introduce an adaptive knot-span adjustment strategy for non-uniform continuous-time trajectory representation. This is accomplished by adjusting control points dynamically based on movement speed. To enable efficient fusion of inertial measurement unit (IMU) and odometer data, we propose an improved extended kalman filter (EKF) with innovation-based adaptive estimation to provide short-term accurate motion prior. Furthermore, to address the challenge of achieving a fully observable UWB localization system under few-anchor conditions, the virtual anchor (VA) generation method based on multiple hypotheses is proposed. At the backend, we propose an adaptive sliding window strategy for global trajectory estimation. Comprehensive experiments are conducted on three self-collect datasets with different UWB anchor numbers and motion modes. The result shows that the proposed CT-UIO achieves$0.403m$,$0.150m$, and$0.189m$localization accuracy in corridor, exhibition hall, and office environments, yielding$17.2\%$,$26.1\%$, and$15.2\%$improvements compared with competing state-of-the-art UIO systems, respectively. The codebase and datasets of this work will be open-sourced athttps://github.com/JasonSun623/CT-UIO. Wei Sun 0028, Genwei Zhang, Kailun Yang 0001, Xiangqi Meng, Na Deng, Chongbin Tan |
IEEE Trans. Mob. Comput. | 2 |
| 2026 | Scalable Unseen Objects 6-DoF Absolute Pose Estimation With Robotic Integration
Jian Liu 0014, Wei Sun 0028, Hossein Rahmani 0001, Ajmal Mian, Lin Wang 0025 |
IEEE Trans. Robotics | 2 |
| 2025 | MonoDiff9D: Monocular Category-Level 9D Object Pose Estimation via Diffusion ModelabstractObject pose estimation is a core means for robots to understand and interact with their environment. For this task, monocular category-level methods are attractive as they require only a single RGB camera. However, current methods rely on shape priors or CAD models of the intra-class known objects. We propose a diffusion-based monocular category-level 9D object pose generation method, MonoDiff9D. Our motivation is to leverage the probabilistic nature of diffusion models to alleviate the need for shape priors, CAD models, or depth sensors for intra-class unknown object pose estimation. We first estimate coarse depth via DINOv2 from the monocular image in a zero-shot manner and convert it into a point cloud. We then fuse the global features of the point cloud with the input image and use the fused features along with the encoded time step to condition MonoDiff9D. Finally, we design a transformer-based denoiser to recover the object pose from Gaussian noise. Extensive experiments on two popular benchmark datasets show that MonoDiff9D achieves state-of-the-art monocular category-level 9D object pose estimation accuracy without the need for shape priors or CAD models at any stage. Our code will be made public at https://github.com/CNJianLiu/MonoDiff9D. Jian Liu 0014, Wei Sun 0028, Zichen Geng, Hossein Rahmani 0001, Ajmal Mian |
ICRA | 2 |
| 2025 | Visual-tactile fusion learning for material recognition based on channel switching and dual cross-attention
Wei Sun 0028, Qiaokang Liang, Yudong Yang |
Eng. Appl. Artif. Intell. | 2 |
| 2025 | UWB-IMU-Odometer Fusion for Simultaneous Calibration and LocalizationabstractThe location accuracy of fixed anchors plays a pivotal role in ultrawideband (UWB) positioning. However, existing calibration methods for calculating the anchor positions require anchor-to-anchor communication or known initial values for anchor locations. Moreover, the calibration accuracy is adversely affected by non line-of-sight (NLOS) conditions. We propose a two-stage calibration scheme to conduct simultaneous calibration and localization (SCAL) based on UWB-inertial measurement unit (IMU)–odometer sensor fusion without the need for anchor-to-anchor communication or manual intervention. Our method performs IMU-odometer-aided multidimensional scaling (IO-MDS) to provide the initial calibration value without anchor-to-anchor ranging measurements. This is followed by a novel factor-graph-based framework to achieve coarse-to-fine calibration based on UWB, inertial, and odometer measurements. In existing works, a single UWB range measurement is regarded as a weak constraint, as it often leads to incorrect estimates. Our method uses the derived radial velocity (DRV) and IO-MDS factors as additional strong constraints for reliable estimates. To minimize the NLOS influence on the calibration process, we introduce an improved least square-support vector machine (ILS-SVM) based on adaptive weight parameter and multikernel function. Experimental results on two field collected datasets show enhancements in NLOS identification and SCAL. In the lab dataset, identification accuracy increased by 5.5%, with improvements of 0.124 m and 0.309 m in root mean square error (RMSE) for robot and anchor locations, respectively. In the parking lot dataset, identification accuracy improved by 5.0%, with RMSE improvements of 0.223 m and 0.317 m for robot and anchor locations, respectively. Wei Sun 0028, Jian Liu 0014, Ajmal Mian |
IEEE Internet Things J. | 2 |
| 2025 | Diff9D: Diffusion-Based Domain-Generalized Category-Level 9-DoF Object Pose EstimationabstractNine-degrees-of-freedom (9-DoF) object pose and size estimation is crucial for enabling augmented reality and robotic manipulation. Category-level methods have received extensive research attention due to their potential for generalization to intra-class unknown objects. However, these methods require manual collection and labeling of large-scale real-world training data. To address this problem, we introduce a diffusion-based paradigm for domain-generalized category-level 9-DoF object pose estimation. Our motivation is to leverage the latent generalization ability of the diffusion model to address the domain generalization challenge in object pose estimation. This entails training the model exclusively on rendered synthetic data to achieve generalization to real-world scenes. We propose an effective diffusion model to redefine 9-DoF object pose estimation from a generative perspective. Our model does not require any 3D shape priors during training or inference. By employing the Denoising Diffusion Implicit Model, we demonstrate that the reverse diffusion process can be executed in as few as 3 steps, achieving near real-time performance. Finally, we design a robotic grasping system comprising both hardware and software components. Through comprehensive experiments on two benchmark datasets and the real-world robotic system, we show that our method achieves state-of-the-art domain generalization performance. Jian Liu 0014, Wei Sun 0028, Pengchao Deng, Chongpei Liu, Nicu Sebe, Hossein Rahmani 0001, Ajmal Mian |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | NPC-SPU: Nonlinear Phase Coding-Based Stereo Phase Unwrapping for Efficient 3D Measurementabstract3D imaging based on phase-shifting structured light is widely used in industrial measurement due to its non-contact nature. However, it typically requires a large number of additional images (multi-frequency heterodyne (M-FH) method) or introduces intensity features that compromise accuracy (space domain modulation phase-shifting (SDM-PS) method) for phase unwrapping, and it remains sensitive to motion. To overcome these issues, this article proposes a nonlinear phase coding-based stereo phase unwrapping (NPC-SPU) method that requires no additional patterns while maintaining measurement accuracy. In the encoding stage, a novel nonlinear distortion feature is introduced, while the signal-to-noise ratio of the phase codeword is preserved. In the decoding stage, a local phase unwrapping method that does not require additional auxiliary information is first proposed, closely associating the distortion information in the local wrapped phase. Then, a pre-calibrated stereo constraint system is used to filter potential matching phases, significantly reducing phase ambiguity and computational costs. Finally, to avoid the time-consuming and complex intensity kernel matching used in traditional methods, we propose a local phase correlation matching (LPCM) technique that enables lightweight and robust phase unwrapping. Experimental results demonstrate that this algorithm significantly enhances 3D reconstruction performance in scenarios with large depth, large disparity, complex colored structures, and dynamic scenes. Specifically, in dynamic environments (20mm/s), the proposed method achieves a lower measurement error rate (0.7829% vs. 6.4962%) with only 3 patterns, compared to the traditional three-frequency heterodyne (T-FH) method (using 9 patterns). Additionally, its measurement accuracy outperforms the advanced SDM-PS method, which also uses 3 patterns (0.1102 mm vs. 0.3232 mm). Ruiming Yu, Hongshan Yu, Wei Sun 0028, Yaonan Wang 0001, Naveed Akhtar, Kemao Qian |
IEEE Trans. Image Process. | 3 |
| 2025 | Full Point Encoding for Local Feature Aggregation in 3-D Point CloudsabstractPoint cloud processing methods exploit local point features and global context through aggregation which does not explicitly model the internal correlations between local and global features. To address this problem, we propose full point encoding which is applicable to convolution and transformer architectures. Specifically, we propose full point convolution (FuPConv) and full point transformer (FPTransformer) architectures. The key idea is to adaptively learn the weights from local and global geometric connections, where the connections are established through local and global correlation functions, respectively. FuPConv and FPTransformer simultaneously model the local and global geometric relationships as well as their internal correlations, demonstrating strong generalization ability and high performance. FuPConv is incorporated in classical hierarchical network architectures to achieve local and global shape-aware learning. In FPTransformer, we introduce full point position encoding in self-attention, that hierarchically encodes each point position in the global and local receptive field. We also propose a shape-aware downsampling block that takes into account the local shape and the global context. Experimental comparison to existing methods on benchmark datasets shows the efficacy of FuPConv and FPTransformer for semantic segmentation, object detection, classification, and normal estimation tasks. In particular, we achieve state-of-the-art semantic segmentation results of 76.8% mIoU on S3DIS sixfold and 73.1% on S3DIS Area 5. Our code is available at https://github.com/hnuhyuwa/FullPointTransformer. Yong He 0012, Hongshan Yu, Zhengeng Yang, Wei Sun 0028, Ajmal Mian |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2025 | MH6D: Multi-Hypothesis Consistency Learning for Category-Level 6-D Object Pose EstimationabstractSix-degree-of-freedom (6DoF) object pose estimation is a crucial task for virtual reality and accurate robotic manipulation. Category-level 6DoF pose estimation has recently become popular as it improves generalization to a complete category of objects. However, current methods focus on data-driven differential learning, which makes them highly dependent on the quality of the real-world labeled data and limits their ability to generalize to unseen objects. To address this problem, we propose multi-hypothesis (MH) consistency learning (MH6D) for category-level 6-D object pose estimation without using real-world training data. MH6D uses a parallel consistency learning structure, alleviating the uncertainty problem of single-shot feature extraction and promoting self-adaptation of domain to reduce the synthetic-to-real domain gap. Specifically, three randomly sampled pose transformations are first performed in parallel on the input point cloud. An attention-guided category-level 6-D pose estimation network with channel attention (CA) and global feature cross-attention (GFCA) modules is then proposed to estimate the three hypothesized 6-D object poses by extracting and fusing the global and local features effectively. Finally, we propose a novel loss function that considers both the process and the final result information allowing MH6D to perform robust consistency learning. We conduct experiments under two different training data settings (i.e., only synthetic data and synthetic and real-world data) to verify the generalization ability of MH6D. Extensive experiments on benchmark datasets demonstrate that MH6D achieves state-of-the-art (SOTA) performance, outperforming most data-driven methods even without using any real-world data. The code is available at https://github.com/CNJianLiu/MH6D. Jian Liu 0014, Wei Sun 0028, Chongpei Liu, Ajmal Mian |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | OST: Refining Text Knowledge with Optimal Spatio-Temporal Descriptor for General Video RecognitionabstractDue to the resource-intensive nature of training vision- language models on expansive video data, a majority of studies have centered on adapting pre-trained image- language models to the video domain. Dominant pipelines propose to tackle the visual discrepancies with additional temporal learners while overlooking the substantial discrepancy for web-scaled descriptive narratives and concise action category names, leading to less distinct semantic space and potential performance limitations. In this work, we prioritize the refinement of text knowledge to facilitate generalizable video recognition. To address the limitations of the less distinct semantic space of category names, we prompt a large language model (LLM) to augment action class names into Spatio-Temporal Descriptors thus bridging the textual discrepancy and serving as a knowledge base for general recognition. Moreover, to assign the best descriptors with different video instances, we propose Optimal Descriptor Solver, forming the video recognition problem as solving the optimal matching flow across frame-level representations and descriptors. Comprehensive evaluations in zero-shot, few-shot, and fully supervised video recognition highlight the effectiveness of our approach. Our best model achieves a state-of-the-art zero-shot accuracy of 75.1% on Kinetics-600. Tom Tongjia Chen, Hongshan Yu, Zhengeng Yang, Zechuan Li, Wei Sun 0028, Chen Chen 0001 |
CVPR | 5 |
| 2024 | EFRNet-VL: An end-to-end feature refinement network for monocular visual localization in dynamic environments
Jingwen Wang 0009, Hongshan Yu, Xuefei Lin, Zechuan Li, Wei Sun 0028, Naveed Akhtar |
Expert Syst. Appl. | 5 |
| 2024 | Domain-Invariant Prototypes for Semantic SegmentationabstractDeep learning has greatly advanced the performance of semantic segmentation, however, its success relies on the availability of large amounts of annotated data for training. Hence, many efforts have been devoted to domain adaptive semantic segmentation that focuses on transferring semantic knowledge from a labeled source domain to an unlabeled target domain. Existing self-training methods typically require multiple rounds of training, while another popular framework based on adversarial training is known to be sensitive to hyper-parameters. We propose an easy-to-train framework that learns domain-invariant prototypes for domain adaptive semantic segmentation. In particular, we show that domain adaptation shares a common character with few-shot learning in that both aim to recognize some types of unseen data with knowledge learned from large amounts of seen data. Thus, we propose a unified framework for domain adaptation and few-shot learning. The core idea is to use the class prototypes extracted from few-shot annotated target images to classify pixels of both source images and target images. Our method involves only one-stage training and does not need to be trained on large-scale un-annotated target images. Moreover, our method can be extended to variants of both domain adaptation and few-shot learning. Competitive performances achieved on GTA5-to-Cityscapes and SYNTHIA-to-Cityscapes adaptation tasks show the effectiveness of the proposed novel while simple domain adaptation framework. The source code used in this paper is available at https://github.com/zgyang-hnu/DIP-hunnu. Zhengeng Yang, Hongshan Yu, Wei Sun 0028, Li Cheng 0001, Ajmal Mian |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Domain-Generalized Robotic Picking via Contrastive Learning-Based 6-D Pose EstimationabstractVision-guided robotic picking in 3-D space is a key technology for industrial automation and intelligent manufacturing. However, existing methods rely on labeled real-world data for learning, significantly limiting their ability to generalize to novel objects and robustness to challenging scenes containing occlusions and clutter. To address these problems, we propose a domain-generalized robotic picking method (DGPF6D) that builds on contrastive learning-based 6-D pose estimation. DGPF6D generalizes to real-world scenes by training only on synthetic data and without using shape priors. Specifically, we first perform continuous data augmentations on the synthetic RGB and point cloud images such that they can better simulate real-world scenes with occlusions and clutter. We then feed the augmented images in parallel to a two-stage (i.e., 3-D shape reconstruction and 6-D pose estimation) contrastive learning framework, thereby enhancing the domain-generalization ability and robustness of DGPF6D. Moreover, we propose a point cloud cross attention-guided intracategory unknown object 3-D shape reconstruction network, which can effectively fuse the observed and the unit random point clouds and explicitly highlight their differences, thus avoiding the dependence of DGPF6D on shape priors. Finally, we build a robotic picking system employing DGPF6D to realize domain-generalized robotic picking in 3-D space. Extensive experiments on two benchmarks and real-world scenes show that DGPF6D achieves state-of-the-art performance, and can be effectively applied for domain-generalized robotic picking. Jian Liu 0014, Wei Sun 0028, Chongpei Liu, Ajmal Mian |
IEEE Trans. Ind. Informatics | 2 |
| 2024 | A Sequential-Multi-Decision Scheme for WiFi Localization Using Vision-Based RefinementabstractCurrently, most mobile devices have WIFI and camera modules to locate their position. However, there are two main challenges in large, highly similar indoor environments (localization accuracy and localization time). Aiming to balance these problems, we propose a sequential-multi-decision integrated system that combines WIFI and vision to acquire users’ locations. This system has two phases: sequential fusion localization and adaptive multi-decision fusion localization. The former employs WIFI-based localization first, then image-based localization and fusion localization are used within the constraints of WIFI-based localization. In the WIFI-based localization phase, the gaussian process regression (GPR) model is used to construct a WIFI indoor map. Subsequently, we propose to apply the hybrid whale optimization algorithm (HWOA) to WIFI-based localization to improve its accuracy and stability. The latter uses an adaptive multi-decision fusion mechanism that integrates WIFI-based localization, image-based localization, and fusion localization to obtain the users’ location finally. The experiments show the effectiveness of HWOA applied to WIFI-based localization. We also experimentally evaluate the proposed fusion algorithm with other state-of-the-art fusion algorithms (e.g., accuracy and time) in a real environment (an area larger than 10,000$m^{2}$). The experimental results show that the proposed fusion system is competitive. Chenjun Tang, Wei Sun 0028, Chongpei Liu |
IEEE Trans. Mob. Comput. | 2 |
| 2024 | Wi-Fi-Based Indoor Localization With Interval Random Analysis and Improved Particle Swarm OptimizationabstractThe rise of the Internet of Things has spurred the growth of wireless applications, particularly Wi-Fi-based indoor localization, which is gaining prominence owing to its cost-effectiveness. Nevertheless, the accuracy of Wi-Fi-based indoor localization is hindered by signal instability. To address this limitation, we introduce an interval random analysis approach for uncertain Wi-Fi-based indoor localization. Specifically, this approach employs an interval random parameter lognormal shadowing model for radio map enhancement and adaptive Bayesian comprehensive learning (IRPLS-ABCL) particle swarm optimization (PSO) for location estimation accuracy enhancement. The process comprises two stages: offline training and online localization. During the offline phase, we establish the interval random parameter lognormal shadowing model, considering the parameters as interval random variables, rather than precise values, in a sparse reference point scenario. In the online phase, we use a double-panel fingerprint homogeneity model to assess fingerprint similarity and apply the adaptive Bayesian comprehensive learning PSO algorithm to enhance localization precision. The experimental results show that the proposed algorithm can achieve the best performance in terms of localization accuracy based on the predicted average received signal strength (RSS), reaching 1.89 m. Wei Sun 0028, Anping Lin, Jian Liu 0014, Shuzhi Sam Ge |
IEEE Trans. Mob. Comput. | 2 |
| 2024 | Fully Convolutional Network-Based Self-Supervised Learning for Semantic SegmentationabstractAlthough deep learning has achieved great success in many computer vision tasks, its performance relies on the availability of large datasets with densely annotated samples. Such datasets are difficult and expensive to obtain. In this article, we focus on the problem of learning representation from unlabeled data for semantic segmentation. Inspired by two patch-based methods, we develop a novel self-supervised learning framework by formulating the jigsaw puzzle problem as a patch-wise classification problem and solving it with a fully convolutional network. By learning to solve a jigsaw puzzle comprising 25 patches and transferring the learned features to semantic segmentation task, we achieve a 5.8% point improvement on the Cityscapes dataset over the baseline model initialized from random values. It is noted that we use only about 1/6 training images of Cityscapes in our experiment, which is designed to imitate the real cases where fully annotated images are usually limited to a small number. We also show that our self-supervised learning method can be applied to different datasets and models. In particular, we achieved competitive performance with the state-of-the-art methods on the PASCAL VOC2012 dataset using significantly fewer time costs on pretraining. Zhengeng Yang, Hongshan Yu, Yong He 0012, Wei Sun 0028, Zhi-Hong Mao, Ajmal Mian |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2023 | Fine segmentation and difference-aware shape adjustment for category-level 6DoF object pose estimation
Chongpei Liu, Wei Sun 0028, Jian Liu 0014, Shimeng Fan, Qiang Fu 0013 |
Appl. Intell. | 2 |
| 2023 | Accurate edge-preserving stereo matching by enhancing anisotropy
Shimeng Fan, Wei Sun 0028, Qiang Fu 0013, Wei Wu 0022 |
Signal Process. Image Commun. | 2 |
| 2023 | Robotic Continuous Grasping System by Shape Transformer-Guided Multiobject Category-Level 6-D Pose EstimationabstractRobotic grasping is one of the key functions for realizing industrial automation and human–machine interaction. However, current robotic grasping methods for unknown objects mainly focus on generating the 6-D grasp poses, which cannot obtain rich object pose information and are not robust in challenging scenes. Based on this, in this article, we propose a robotic continuous grasping system that achieves end-to-end robotic grasping of intraclass unknown objects in 3-D space by accurate category-level 6-D object pose estimation. Specifically, to achieve object pose estimation, first, we propose a global shape extraction network (GSENet) based on ResNet1D to extract the global shape of an object category from the 3-D models of intraclass known objects. Then, with the global shape as the prior feature, we propose a transformer-guided network to reconstruct the shape of intraclass unknown object. The proposed network can effectively introduce internal and mutual communication between the prior feature, current feature, and their difference feature. The internal communication is performed by self-attention. The mutual communication is performed by cross attention to strengthen their correlation. To achieve robotic grasping for multiple objects, we propose a low-computation and effective grasping strategy based on the predefined vector orientation, and develop a graphical user interface for monitoring and control. Experiments on two benchmark datasets demonstrate that our system achieves state-of-the-art 6-D pose estimation accuracy. Moreover, the real-world experiments show that our system also achieves superior robotic grasping performance, with a grasping success rate of 81.6$\%$for multiple objects. Jian Liu 0014, Wei Sun 0028, Chongpei Liu, Qiang Fu 0013 |
IEEE Trans. Ind. Informatics | 2 |
| 2022 | Self-supervised part segmentation via motion imitation
Qiaokang Liang, Kunlin Zou, Wei Sun 0028, Yaonan Wang 0001 |
Image Vis. Comput. | 5 |
| 2022 | A hybrid whale optimization algorithm with artificial bee colony
Chenjun Tang, Wei Sun 0028, Hongwei Tang, Wei Wu 0022 |
Soft Comput. | 2 |
| 2022 | HFF6D: Hierarchical Feature Fusion Network for Robust 6D Object Pose TrackingabstractTracking the 6-degree-of-freedom (6D) object pose in video sequences is gaining attention because it has a wide application in multimedia and robotic manipulation. However, current methods often perform poorly in challenging scenes, such as incorrect initial pose, sudden re-orientation, and severe occlusion. In contrast, we present a robust 6D object pose tracking method with a novel hierarchical feature fusion network, refer it as HFF6D, which aims to predict the object’s relative pose between adjacent frames. Instead of extracting features from adjacent frames separately, HFF6D establishes sufficient spatial-temporal information interaction between adjacent frames. In addition, we propose a novel subtraction feature fusion (SFF) module with attention mechanism to leverage feature subtraction during feature fusion. It explicitly highlights the feature differences between adjacent frames, thus improving the robustness of relative pose estimation in challenging scenes. Besides, we leverage data augmentation technology to make HFF6D be used more effectively in the real world by training only with synthetic data, thereby reducing manual effort in data annotation. We evaluate HFF6D on the well-known YCB-Video and YCBInEOAT datasets. Quantitative and qualitative results demonstrate that HFF6D outperforms state-of-the-art (SOTA) methods in both accuracy and efficiency. Moreover, it is also proved to achieve high-robustness tracking under the above-mentioned challenging scenes. Jian Liu 0014, Wei Sun 0028, Chongpei Liu, Shimeng Fan, Wei Wu 0022 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Multimodal Vigilance Estimation Using Deep LearningabstractThe phenomenon of increasing accidents caused by reduced vigilance does exist. In the future, the high accuracy of vigilance estimation will play a significant role in public transportation safety. We propose a multimodal regression network that consists of multichannel deep autoencoders with subnetwork neurons (MCDAE$_{sn}$). After we define two thresholds of “0.35” and “0.70” from the percentage of eye closure, the output values are in the continuous range of 0–0.35, 0.36–0.70, and 0.71–1 representing the awake state, the tired state, and the drowsy state, respectively. To verify the efficiency of our strategy, we first applied the proposed approach to a single modality. Then, for the multimodality, since the complementary information between forehead electrooculography and electroencephalography features, we found the performance of the proposed approach using features fusion significantly improved, demonstrating the effectiveness and efficiency of our method. Wei Wu 0022, Wei Sun 0028, Q. M. Jonathan Wu, Yimin Yang 0001, Hui Zhang 0023, Wei-Long Zheng, Bao-Liang Lu |
IEEE Trans. Cybern. | 2 |
| 2021 | Automatic Leaf Diseases Detection System Based on Multi-stage Recognition
Songyun Deng, Lekai Cheng, Wei Sun 0028, Yaonan Wang 0001, Qiaokang Liang |
ICIG (1) | 4 |
| 2021 | A GWO-based multi-robot cooperation method for target searching in unknown environments
Hongwei Tang, Wei Sun 0028, Anping Lin |
Expert Syst. Appl. | 2 |
| 2021 | EdgeGAN: One-way mapping generative adversarial network based on the edge information for unpaired training set
Qiaokang Liang, Youcheng Lei, Wei Sun 0028, Yaonan Wang 0001, Dan Zhang 0006 |
J. Vis. Commun. Image Represent. | 5 |
| 2021 | Sequence-tracker: Multiple object tracking with sequence features in severe occlusion scene
Qiaokang Liang, Wei Sun 0028, Yaonan Wang 0001, Dan Zhang 0006 |
J. Vis. Commun. Image Represent. | 4 |
| 2021 | NDNet: Narrow While Deep Network for Real-Time Semantic SegmentationabstractThe rapid development of autonomous driving in recent years presents many challenges for scene understanding. As an essential step towards scene understanding, semantic segmentation has received increased attention in the past few years. Although deep learning based approaches have achieved great success in improving the segmentation accuracy, most of them suffer from an inefficiency problem and can hardly be applied to real-time applications. In this paper, we analyze the computational cost of Convolutional Neural Network (CNN) and find that the inefficiency of CNNs is mainly caused by their wide structure rather than deep structure. In addition, the success of pruning based model compression methods proves that there are many redundant channels in CNNs. Thus, we design a narrow while deep backbone network to improve the efficiency of semantic segmentation. By casting our network to fully convolutional network (FCN32) segmentation architecture, the basic structure of most segmentation methods, we achieve 61.5% mIoU on Cityscapes validation dataset with only 4.2G floating-point operations (FLOPs) on 1024×2048 inputs, which already outperforms one of the earliest real-time deep learning based segmentation methods: ENet (58.3% mIoU, 3.8G FLOPs on 640×360 inputs). By further refining the output resolution of our network to the 1/8 of the input resolution with a simple encoder-decoder structure, we achieve 65.3% mIoU on Cityscapes test set with 14.0G FLOPs and 39.9 frames per second (FPS) on Titan X card. We have made our model publicly available at https://github.com/zgyang-hnu/NDNet. Zhengeng Yang, Hongshan Yu, Qiang Fu 0013, Wei Sun 0028, Wenyan Jia, Mingui Sun, Zhi-Hong Mao |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2020 | A multirobot target searching method based on bat algorithm in unknown environments
Hongwei Tang, Wei Sun 0028, Hongshan Yu, Anping Lin |
Expert Syst. Appl. | 2 |
| 2020 | Small Object Augmentation of Urban Scenes for Real-Time Semantic SegmentationabstractSemantic segmentation is a key step in scene understanding for autonomous driving. Although deep learning has significantly improved the segmentation accuracy, current highquality models such as PSPNet and DeepLabV3 are inefficient given their complex architectures and reliance on multi-scale inputs. Thus, it is difficult to apply them to real-time or practical applications. On the other hand, existing real-time methods cannot yet produce satisfactory results on small objects such as traffic lights, which are imperative to safe autonomous driving. In this paper, we improve the performance of real-time semantic segmentation from two perspectives, methodology and data. Specifically, we propose a real-time segmentation model coined Narrow Deep Network (NDNet) and build a synthetic dataset by inserting additional small objects into the training images. The proposed method achieves 65.7% mean intersection over union (mIoU) on the Cityscapes test set with only 8.4G floatingpoint operations (FLOPs) on 1024×2048 inputs. Furthermore, by re-training the existing PSPNet and DeepLabV3 models on our synthetic dataset, we obtained an average 2% mIoU improvement on small objects. Zhengeng Yang, Hongshan Yu, Mingtao Feng, Wei Sun 0028, Xuefei Lin, Mingui Sun, Zhi-Hong Mao, Ajmal Mian |
IEEE Trans. Image Process. | 4 |
| 2019 | A novel hybrid algorithm based on PSO and FOA for target searching in unknown environments
Hongwei Tang, Wei Sun 0028, Hongshan Yu, Anping Lin, Yuxue Song |
Appl. Intell. | 2 |
| 2019 | Locate the Mobile Device by Enhancing the WiFi-Based Indoor Localization ModelabstractDue to the advent and pervasive deployment of wireless local area networks, WiFi-based indoor localization systems have received increasing attention in the last few years. However, their localization accuracy has always been a challenging issue. In addition, because of diverse interference such as multipath effects, the block of signals, an unstable or weak signal in itself, etc., not all the access points (APs) are informative for the localization. Faced with these problems, we propose a WiFi-based localization model by modifying the large localization errors and enhancing the Gaussian process regression (MEGPR). 1) To select the AP subsets that contribute more to the localization and further reduce the computational load, the AP discrimination criterion (APDC) is introduced to quantify the discernibility of the APs detected in the workspace and filter out the APs with low discrimination. 2) Second, to enhance the localization model, the localization residual is fed and learnt by the model. 3) Furthermore, the large localization errors are mitigated by the location modification method (LMM). Experiments were conducted in a real environment with an area of more than 1200 m2and the results show that compared with other existing localization models, the average localization error of the proposed MEGPR model is minimum, which further verifies the effectiveness of the proposed MEGPR localization model. Wei Sun 0028, Hongshan Yu, Hongwei Tang, Anping Lin, Roger Zimmermann |
IEEE Internet Things J. | 2 |
| 2019 | Weakly Supervised Biomedical Image Segmentation by Reiterative LearningabstractRecent advances in deep learning have produced encouraging results for biomedical image segmentation; however, outcomes rely heavily on comprehensive annotation. In this paper, we propose a neural network architecture and a new algorithm, known as overlapped region forecast, for the automatic segmentation of gastric cancer images. To the best of our knowledge, this report for the first time describes that deep learning has been applied to the segmentation of gastric cancer images. Moreover, a reiterative learning framework that achieves superior performance without pretraining or further manual annotation is presented to train a simple network on weakly annotated biomedical images. We customize the loss function to make the model converge faster while avoiding becoming trapped in local minima. Patch boundary errors were eliminated by our overlapped region forecast algorithm. By studying the characteristics of the model trained using two different patch extraction methods, we train iteratively and integrate predictions and weak annotations to improve the quality of the training data. Using these methods, a mean Intersection over Union coefficient of 0.883 and a mean accuracy of 91.09% were achieved on the partially labeled dataset, thereby securing a win in the 2017 China Big Data and Artificial Intelligence Innovation and Entrepreneurship Competition. Qiaokang Liang, Yang Nan 0002, Gianmarc Coppola, Kunglin Zou, Wei Sun 0028, Dan Zhang 0006, Yaonan Wang 0001, Guanzhen Yu |
IEEE J. Biomed. Health Informatics | 5 |
| 2018 | Methods and datasets on semantic segmentation: A review
Hongshan Yu, Zhengeng Yang, Yaonan Wang 0001, Wei Sun 0028, Mingui Sun, Yandong Tang |
Neurocomputing | 5 |
| 2018 | Modified clustering-based differential evolution with a flexible combination of exploration and exploitation
Wei Sun 0028, Yuxue Song, Anping Lin, Hongwei Tang |
Soft Comput. | 1 |
| 2017 | All-dimension neighborhood based particle swarm optimization with randomly selected neighbors
Wei Sun 0028, Anping Lin, Hongshan Yu, Qiaokang Liang, Guohua Wu 0001 |
Inf. Sci. | 1 |
| 2006 | Adaptive Control Based on Recurrent Fuzzy Wavelet Neural Network and Its Application on Robotic Tracking Control
Wei Sun 0028, Yaonan Wang 0001, Xiaohua Zhai |
ISNN (2) | 1 |
| 2004 | An adaptive fuzzy control for robotic manipulatorsabstractIn this paper, an adaptive fuzzy control strategy is developed for robotic manipulators to guarantee both global stability and performance. A controller output error method (COEM) is introduced and applied to the design of adaptive fuzzy control system. The proposed control strategy employs a gradient descent algorithm to minimize a cost function which is based on the error of the controller output and is minimized by tuning some or all of the parameters of fuzzy controller. The simulation results show that the proposed control is effective and yields superior tracking performance. Wei Sun 0028, Yaonan Wang 0001 |
ICARCV | 1 |