Zongming Yang

dblp:235/2679 · DBLP profile ↗
← Back
9ranked-venue papers
1as first author
9since 2021 · last 2024
0000-0003-0876-7487ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021
YearPublicationVenuePosition
2024 SPD: Semi-Supervised Learning and Progressive Distillation for 3-D Detection
abstract
Current learning-based 3-D object detection accuracy is heavily impacted by the annotation quality. It is still a challenge to expect an overall high detection accuracy for all classes under different scenarios given the dataset sparsity. To mitigate this challenge, this article proposes a novel method called semi-supervised learning and progressive distillation (SPD), which uses semi-supervised learning (SSL) and knowledge distillation to improve label efficiency. The SPD uses two big backbones to hand the unlabeled/labeled input data augmented by the periodic IO augmentation (PA). Then the backbones are compressed using progressive distillation (PD). Precisely, PA periodically shifts the data augmentation operations between the input and output of the big backbone, aiming to improve the network's generalization of the unseen and unlabeled data. Using the big backbone can benefit from large-scale augmented data better than the small one. And two backbones are trained by the data scale and ratio-sensitive loss (data-loss). It solves the over-flat caused by the large-scale unlabeled data from PA and helps the big backbone prevent overfitting on the limited-scale labeled data. Hence, using the PA and data loss during SSL training dramatically improves the label efficiency. Next, the trained big backbone set as the teacher CNN is progressively distilled to obtain a small student model, referenced as PD. PD mitigates the problem that student CNN performance degrades when the gap between the student and the teacher is oversized. Extensive experiments are conducted on the indoor datasets SUN RGB-D and ScanNetV2 and outdoor dataset KITTI. Using only 50% labeled data and a 27% smaller model size, SPD performs 0.32 higher than the fully supervised VoteNet [1] which is adopted as our backbone. Besides, using only 2% labeled data, compared to the other fully supervised backbone PV-RCNN [2], SPD accomplishes a similar accuracy (84.1 and 84.83) and 30% less inference time.
Bangquan Xie, Zongming Yang, Ruifa Luo, Ailin Wei, Xiaoxiong Weng, Bing Li 0008
IEEE Trans. Neural Networks Learn. Syst.2
2023 FourStr: When Multi-sensor Fusion Meets Semi-supervised Learning
abstract
This research proposes a novel semi-supervised learning framework FourStr (Four-Stream formed by two two-stream models) that focuses on the improvement of fusion and labeling efficiency for 3D multi-sensor detector. FourStr adopts a multi-sensor single-stage detector named adaptive fusion network (AFNet) as the backbone and trains it through the semi-supervision learning (SSL) strategy Stereo Fusion. Note that multi-sensor AFNet and SSL Stereo Fusion can benefit each other. On the one hand, the Four-stream composed of two AFNets naturally provides rich inputs and large models for SSL Stereo Fusion. While other SSL works have to use massive augmentation to obtain rich inputs, and deepen and widen the network for large models. On the other hand, by the novel three fusion stages and Loss Pruning, Stereo Fusion improves the fusion and labeling efficiency for AFNet. Finally, extensive experiments demonstrate that FourStr performs excellently on outdoor dataset (KITTI and Waymo Open Dataset) and indoor dataset (SUN RGB-D), especially for the small contour objects. And compared to the fully-supervised methods, FourStr achieves similar accuracy with only 2% labeled data on KITTI (or with 50% labeled data on SUN RGB-D).
Bangquan Xie, Zongming Yang, Ailin Wei, Xiaoxiong Weng, Bing Li 0008
ICRA3
2023 ANAS: Asymptotic NAS for large-scale proxyless search and multi-task transfer learning
Bangquan Xie, Zongming Yang, Ruifa Luo, Ailin Wei, Xiaoxiong Weng, Bing Li 0008
Pattern Recognit.2
2023 AMMF: Attention-Based Multi-Phase Multi-Task Fusion for Small Contour Object 3D Detection
abstract
Recently significant progress has been made in 3D detection. However, it is still challenging to detect small contour objects under complex scenes. This paper proposes a novel Attention-based Multi-phase Multi-task Fusion (AMMF) that uses point-level, RoI-level, and multi-task fusions to complement the disadvantages of LiDAR and camera, to solve this challenge. First, at the feature extraction phase, AMMF uses the Low and High-level Fusion with Matching Attention (LHF-MA) and efficient FPN (eFPN) to perform point-level fusion for cross sensors and single sensor, respectively. Instead of merging each level and using expensive 3D CNN like other methods, LHF-MA fuses low-level spatial location and high-level contextual feature of 2D CNN customized feature extractors and ignores the fusion of middle levels, reducing the computational cost. Then, at the proposal generation phase, Progressive Proposal Fusion (PPF) with learned attention map is used to perform coarse-to-fine RoI-level fusion, instead of only combining coarse-grained features at high-level of network. PPF using progressively increasing IoU thresholds could avoid overfitting and improve the performance. Note that the matching attentions and learned attention maps are utilized to weigh the priority of different sensors. Moreover, to solve the sparseness of point-wise fusion between LiDAR BEV and RGB image, AMMF uses multi-task fusion that generates pseudo-LiDAR from camera by depth estimation task, to guide this point-wise fusion. Finally, AMMF performs excellently for detecting small contour objects like pedestrians, cyclists, and distant cars. On the KITTI, AMMF finishes 3.62% improvements in the moderate instance for pedestrians. It achieves a 2.21% improvement in the$>$50 instance of LEVEL$_{-}$2 level for vehicle on the Waymo Open Dataset. And AMMF is further verified on our customized dataset consisting of challenging scenarios like strong illumination and heavy shadow cases.
Bangquan Xie, Zongming Yang, Ailin Wei, Xiaoxiong Weng, Bing Li 0008
IEEE Trans. Intell. Transp. Syst.2
2022 FocusTR: Focusing on Valuable Feature by Multiple Transformers for Fusing Feature Pyramid on Object Detection
abstract
The feature pyramid, which is a vital component of the convolutional neural networks, plays a significant role in several perception tasks, including object detection for autonomous driving. However, how to better fuse multi-level and multi-sensor feature pyramids is still a significant challenge, especially for object detection. This paper presents a FocusTR (Focusing on the valuable features by multiple Transformers), which is a simple yet effective architecture, to fuse feature pyramid for the single-stream 2D detector and two-stream 3D detector. Specifically, FocusTR encompasses several novel self-attention mechanisms, including the spatial-wise boxAlign attention (SB) for low-level spatial locations, context-wise affinity attention (CA) for high-level context information, and level-wise attention for the multi-level feature. To alleviate self-attention's computational complexity and slow training convergence, Fo-cusTR introduces a low and high-level fusion (LHF) to reduce the computational parameters, and the Pre- Ln [1]to accelerate the training convergence.
Bangquan Xie, Zongming Yang, Ailin Wei, Xiaoxiong Weng, Bing Li 0008
IROS3
2022 Embodied-AI Wheelchair Framework with Hands-free Interface and Manipulation
abstract
Assistive robots can be found in hospitals and rehabilitation clinics, where they help patients maintain a positive disposition. Our proposed robotic mobility solution combines state of the art hardware and software to provide a safer, more independent, and more productive lifestyle for people with some of the most severe disabilities. New hardware includes, a retractable roof, manipulator arm, a hard backpack, a number of sensors that collect environmental data and processors that generate 3D maps for a hands-free human-machine interface.The proposed new system receives input from the user via head tracking or voice command, and displays information through augmented reality into the user’s field of view. The software algorithm will use a novel cycle of self-learning artificial intelligence that achieves autonomous navigation while avoiding collisions with stationary and dynamic objects. The prototype will be assembled and tested over the next three years and a publicly available version could be ready two years thereafter.
Jesse F. Leaman, Zongming Yang, Yasmine N. El-Glaly, Hung Manh La, Bing Li 0008
SMC2
2022 SeeWay: Vision-Language Assistive Navigation for the Visually Impaired
abstract
Assistive navigation for blind or visually impaired (BVI) individuals is of significance to extend their mobility and safety in traveling, enhancing their employment opportunities and fostering personal fulfillment. Conventional research is mainly based on robotic navigation approaches through localization, mapping, and path planning frameworks. They require heavy manual annotation of semantic information in maps and its alignment with sensor mapping. Inspired by the fact that we human beings naturally rely on language instruction inquiry and visual scene understanding to navigate in an unfamiliar environment, this paper proposes a novel vision-language model-based approach for BVI navigation. It does not need heavy-labeled indoor maps and provides a Safe and Efficient E-Wayfinding (SeeWay) assistive solution for BVI individuals. The system consists of a scene-graph map construction module, a navigation path generation module for global path inference by vision-language navigation (VLN), and a navigation with obstacle avoidance module for real-time local navigation. The SeeWay system was deployed on portable iPhone devices with cloud computing assistance for the VLN model inference. The field tests show the effectiveness of the VLN global path finding and local path re-planning. Experiments and quantitative results reveal that heuristic-style instruction outperforms direction/detailed-style instructions for VLN success rate (SR), and the SR decreases as the navigation length increases.
Zongming Yang, Liren Kong, Ailin Wei, Jesse Leaman, Johnell O. Brooks, Bing Li 0008
SMC1
2022 Modeling and Prediction of User Stability and Comfortability on Autonomous Wheelchairs With 3-D Mapping
abstract
Traditional manual wheelchairs have a fixed seat with no movement or angle adjustment, which can seriously affect the user's comfort and greatly limit user experience. However, the electric wheelchair relies on strong intelligence and automatic features; it can not only realize the multidegree freedom adjustment of the human body and the seat but also has a rich and powerful man–machine control interface, which greatly facilitates and improves the user experience. This study upgraded a Permobil C400-powered wheelchair with multisensor data fusion technology to enrich its terrain recognition, tipping stability, and comfortability prediction. The tipping stability modeling of the wheelchair dummy system is carried out using multibody dynamics and vibration mechanics to obtain the tipping stability limit and the comfort evaluation of the wheelchair vibration acceleration on the human body during travel. Based on the elevation mapping method, the wheelchair can estimate the terrain from the local point of view at any point in time. At the same time, the RGB-D depth camera is connected to the robot operating system (ROS) system, and the open-source algorithm package RTAB-MAP is used to complete the MAP construction and collect the 3-D point-cloud terrain data. Then, the real 3-D terrain files are generated through the point-cloud stitching technology for stability simulation of the wheelchair–human system. The tipping stability and comfort indexes of the wheelchair–human system when passing over different physical terrains can be obtained. The experimental results show that the IMU data located on the human chest agree well with the simulation analysis data and are suitable for a variety of complex real-terrain conditions, verifying the accuracy of the wheelchair–human system dynamics model and the feasibility of the simulation analysis process. Thus, this modeling and simulation method can predict wheelchair stability and user comfortability well and ensure a high-performance experience.
Zongming Yang, Peng Yin 0001, Johnell O. Brooks, Bing Li 0008
IEEE Trans. Hum. Mach. Syst.2
2022 Multi-Scale Fusion With Matching Attention Model: A Novel Decoding Network Cooperated With NAS for Real-Time Semantic Segmentation
abstract
This paper proposes a real-time multi-scale semantic segmentation network (MsNet). MsNet is a combination of our novel multi-scale fusion with matching attention model (MFMA) as the decoding network and the network searched by asymptotic neural architecture search (ANAS) or MobileNetV3 as the encoding network. The MFMA not only extracts low level spatial features from multi-scale inputs but also decodes the contextual features extracted by ANAS. Specifically, considering the advantages and disadvantages of the addition fusion and concatenation fusion, we design multi-scale fusion (MF) that balances speed and accuracy. Then we creatively design two matching attention mechanisms (MA), including matching attention with low calculation (MALC) mechanism and matching attention with strong global context modeling (MASG) mechanism, to match varying resolutions and information of features at different levels of a network. Besides, the ANAS performs the deep neural network search by employing an asymptotic method and provide an efficient encoding network for MsNet, releasing researchers from those tedious mechanical trials. Through extensive experiments, we prove that MFMA, which can be applied to numerous recognition tasks, possesses excellent decoding ability. And we demonstrate the effectiveness and necessity of implementing the “matching” attention mechanism. Finally, the proposed two versions, MsNet_ANAS and MsNet_M achieve a new state-of-the-art trade-off between accuracy and speed on the CamVid and Cityscapes datasets. More remarkably, on the Nvidia Tesla V100 GPU, our MsNet_ANAS achieves 74.1% mIoU with the speed of 184.2 FPS on the CamVid while 72.9% mIoU with the speed of 119.9 FPS on the Cityscapes.
Bangquan Xie, Zongming Yang, Ruifa Luo, Ailin Wei, Xiaoxiong Weng, Bing Li 0008
IEEE Trans. Intell. Transp. Syst.2