Boyang Gao

dblp:42/8677 · DBLP profile ↗
← Back
17ranked-venue papers
1as first author
9since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 3 since 2021Systems, architecture and hardware · 4 · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Wavelet-enhanced transformers for real-time event-based object detection in low-light environments
Yangjie Cui, Zhan Tu, Daochun Li, Boyang Gao, Jinwu Xiang, Xin Dong 0020
Neurocomputing4
2025 Link Prediction Research Based on Visual Analysis: Expansion of the LERTR Index and System Validation
abstract
Link prediction is widely used in areas like social networks, bioinformatics, recommendation systems, IoT, and healthcare to uncover latent information in complex networks. While similarity-based methods are simple and interpretable, they often overlook high-order structures and additional attributes, limiting their performance. To address this, we developed LPExplorer, a system that integrates the LERTR index (combining RA and LCP principles) for interpretable and accurate predictions across three-hop paths. LPExplorer visually represents network structures, resource flows, and parameter impacts, allowing users to sort and filter predictions effectively. Experimental results demonstrate its strong performance in analyzing and exploring complex networks.
Yilei He, Jiansu Pu, Hanlin Lan, Boyang Gao, Yanlin Zhu
PacificVis5
2024 Generalizing 6-DoF Grasp Detection via Domain Prior Knowledge
abstract
We focus on the generalization ability of the 6-DoF grasp detection method in this paper. While learning-based grasp detection methods can predict grasp poses for unseen ob-jects using the grasp distribution learned from the training set, they often exhibit a significant performance drop when encountering objects with diverse shapes and struc-tures. To enhance the grasp detection methods' general-ization ability, we incorporate domain prior knowledge of robotic grasping, enabling better adaptation to objects with significant shape and structure differences. More specifi-cally, we employ the physical constraint regularization during the training phase to guide the model towards predicting grasps that comply with the physical rule on grasping. For the unstable grasp poses predicted on novel objects, we design a contact-score joint optimization using the pro-jection contact map to refine these poses in cluttered sce-narios. Extensive experiments conducted on the GraspNet-1 billion benchmark demonstrate a substantial performance gain on the novel object set and the real-world grasping experiments also demonstrate the effectiveness of our gen-eralizing 6-DoF grasp detection method. Code is available at https://github.com/mahaoxiang822/Generalizing-Grasp.
Modi Shi, Boyang Gao, Di Huang 0001
CVPR3
2024 Sim-to-Real Grasp Detection with Global-to-Local RGB-D Adaptation
abstract
This paper focuses on the sim-to-real issue of RGB-D grasp detection and formulates it as a domain adaptation problem. In this case, we present a global-to-local method to address hybrid domain gaps in RGB and depth data and insufficient multi-modal feature alignment. First, a self-supervised rotation pre-training strategy is adopted to deliver robust initialization for RGB and depth networks. We then propose a global-to-local alignment pipeline with individual global domain classifiers for scene features of RGB and depth images as well as a local one specifically working for grasp features in the two modalities. In particular, we propose a grasp prototype adaptation module, which aims to facilitate fine-grained local feature alignment by dynamically updating and matching the grasp prototypes from the simulation and real-world scenarios throughout the training process. Due to such designs, the proposed method substantially reduces the domain shift and thus leads to consistent performance improvements. Extensive experiments are conducted on the GraspNet-Planar benchmark and physical environment, and superior results are achieved which demonstrate the effectiveness of our method. Code is available at https://github.com/mahaoxiang822/GL-MSDA.
Ran Qin, Modi Shi, Boyang Gao, Di Huang 0001
ICRA4
2024 Active Perception for Grasp Detection via Neural Graspness Field
abstract
This paper tackles the challenge of active perception for robotic grasp detection in cluttered environments. Incomplete 3D geometry information can negatively affect the performance of learning-based grasp detection methods, and scanning the scene from multiple views introduces significant time costs. To achieve reliable grasping performance with efficient camera movement, we propose an active grasp detection framework based on the Neural Graspness Field (NGF), which models the scene incrementally and facilitates next-best-view planning. Constructed in real-time as the camera moves, the NGF effectively models the grasp distribution in 3D space by rendering graspness predictions from each view. For next-best-view planning, we aim to reduce the uncertainty of the NGF through a graspness inconsistency-guided policy, selecting views based on discrepancies between NGF outputs and a pre-trained graspness network. Additionally, we present a neural graspness sampling method that decodes graspness values from the NGF to improve grasp pose detection results. Extensive experiments on the GraspNet-1Billion benchmark demonstrate significant performance improvements compared to previous works. Real-world experiments show that our method achieves a superior trade-off between grasping performance and time costs.
Modi Shi, Boyang Gao, Di Huang 0001
NeurIPS3
2023 RGB-D Grasp Detection via Depth Guided Learning with Cross-modal Attention
abstract
Planar grasp detection is one of the most fundamental tasks to robotic manipulation, and the recent progress of consumer-grade RGB-D sensors enables delivering more comprehensive features from both the texture and shape modalities. However, depth maps are generally of a relatively lower quality with much stronger noise compared to RGB images, making it challenging to acquire grasp depth and fuse multi-modal clues. To address the two issues, this paper proposes a novel learning based approach to RGB-D grasp detection, namely Depth Guided Cross-modal Attention Network (DGCAN). To better leverage the geometry information recorded in the depth channel, a complete 6-dimensional rectangle representation is adopted with the grasp depth dedicatedly considered in addition to those defined in the common 5-dimensional one. The prediction of the extra grasp depth substantially strengthens feature learning, thereby leading to more accurate results. Moreover, to reduce the negative impact caused by the discrepancy of data quality in two modalities, a Local Cross-modal Attention (LCA) module is designed, where the depth features are refined according to cross-modal relations and concatenated to the RGB ones for more sufficient fusion. Extensive simulation and physical evaluations are conducted and the experimental results highlight the superiority of the proposed approach.
Ran Qin, Boyang Gao, Di Huang 0001
ICRA3
2022 matExplorer: Visual Exploration on Predicting Ionic Conductivity for Solid-state Electrolytes
abstract
Lithium ion batteries (LIBs) are widely used as important energy sources for mobile phones, electric vehicles, and drones. Experts have attempted to replace liquid electrolytes with solid electrolytes that have wider electrochemical window and higher stability due to the potential safety risks, such as electrolyte leakage, flammable solvents, poor thermal stability, and many side reactions caused by liquid electrolytes. However, finding suitable alternative materials using traditional approaches is very difficult due to the incredibly high cost in searching. Machine learning (ML)-based methods are currently introduced and used for material prediction. However, learning tools designed for domain experts to conduct intuitive performance comparison and analysis of ML models are rare. In this case, we propose an interactive visualization system for experts to select suitable ML models and understand and explore the predication results comprehensively. Our system uses a multifaceted visualization scheme designed to support analysis from various perspectives, such as feature distribution, data similarity, model performance, and result presentation. Case studies with actual lab experiments have been conducted by the experts, and the final results confirmed the effectiveness and helpfulness of our system.
Jiansu Pu, Boyang Gao, Zhengguo Zhu, Yanlin Zhu, Yunbo Rao
IEEE Trans. Vis. Comput. Graph.3
2021 Visual Analysis on Machine Learning Assisted Prediction of Ionic Conductivity for Solid-State Electrolytes
abstract
Lithium ion batteries (LIBs) are widely used as the important energy sources in our daily life such as mobile phones, electric vehicles, and drones etc. Due to the potential safety risks caused by liquid electrolytes, the experts have tried to replace liquid electrolytes with solid ones. However, it is very difficult to find suitable alternatives materials in traditional ways for its incredible high cost in searching. Machine learning (ML) based methods are currently introduced and used for material prediction. But there is rarely an assisting learning tools designed for domain experts for institutive performance comparison and analysis of ML model. In this case, we propose an interactive visualization system for experts to select suitable ML models, understand and explore the predication results comprehensively. Our system employs a multi-faceted visualization scheme designed to support analysis from the perspective of feature composition, data similarity, model performance, and results presentation. A case study with real experiments in lab has been taken by the expert and the results of confirmed the effectiveness and helpfulness of our system.
Jiansu Pu, Yanlin Zhu, Boyang Gao, Zhengguo Zhu, Yunbo Rao
PacificVis4
2021 Double-Dot Network for Antipodal Grasp Detection
abstract
This paper proposes a new deep learning approach to antipodal grasp detection, named Double-Dot Network (DD-Net). It follows the recent anchor-free object detection framework, which does not depend on empirically pre-set anchors and thus allows more generalized and flexible prediction on unseen objects. Specifically, unlike the widely used 5-dimensional rectangle, the gripper configuration is defined as a pair of fingertips. An effective CNN architecture is introduced to localize such fingertips, and with the help of auxiliary centers for refinement, it accurately and robustly infers grasp candidates. Additionally, we design a specialized loss function to measure the quality of grasps, and in contrast to the IoU scores of bounding boxes adopted in object detection, it is more consistent to the grasp detection task. Both the simulation and robotic experiments are executed and state of the art accuracies are achieved, showing that DD-Net is superior to the counterparts in handling unseen objects.
Yangtao Zheng, Boyang Gao, Di Huang 0001
IROS3
2018 An AR system for artistic creativity education
abstract
Creativity and innovation training is the core of the art education. Modern technology provides more effective tools to help students obtain artistic creativity. In this paper, we propose to employ augmented reality technology to assist artistic creativity education. We first analyze the inefficiency of traditional artistic creation training. We then introduce our AR-based smartphone app with technical detail and explain how it can improve accelerate artistic creativity training. We finally show 3 examples created by our AR app to demonstrate the effectiveness of our proposed method.
Jiajia Tan, Boyang Gao, Xiaobo Lu
VRST2
2018 Visual and Semantic Knowledge Transfer for Large Scale Semi-Supervised Object Detection
abstract
Deep CNN-based object detection systems have achieved remarkable success on several large-scale object detection benchmarks. However, training such detectors requires a large number of labeled bounding boxes, which are more difficult to obtain than image-level annotations. Previous work addresses this issue by transforming image-level classifiers into object detectors. This is done by modeling the differences between the two on categories with both image-level and bounding box annotations, and transferring this information to convert classifiers to detectors for categories without bounding box annotations. We improve this previous work by incorporating knowledge about object similarities from visual and semantic domains during the transfer process. The intuition behind our proposed method is that visually and semantically similar categories should exhibit more common transferable properties than dissimilar categories, e.g. a better detector would result by transforming the differences between a dog classifier and a dog detector onto the cat class, than would by transforming from the violin class. Experimental results on the challenging ILSVRC2013 detection dataset demonstrate that each of our proposed object similarity based knowledge transfer methods outperforms the baseline methods. We found strong evidence that visual similarity and semantic relatedness are complementary for the task, and when combined notably improve detection, achieving state-of-the-art detection performance in a semi-supervised setting.
Yuxing Tang, Josiah Wang, Boyang Gao, Emmanuel Dellandréa, Robert J. Gaizauskas, Liming Chen 0002
IEEE Trans. Pattern Anal. Mach. Intell.4
2016 Large Scale Semi-Supervised Object Detection Using Visual and Semantic Knowledge Transfer
abstract
Deep CNN-based object detection systems have achieved remarkable success on several large-scale object detection benchmarks. However, training such detectors requires a large number of labeled bounding boxes, which are more difficult to obtain than image-level annotations. Previous work addresses this issue by transforming image-level classifiers into object detectors. This is done by modeling the differences between the two on categories with both imagelevel and bounding box annotations, and transferring this information to convert classifiers to detectors for categories without bounding box annotations. We improve this previous work by incorporating knowledge about object similarities from visual and semantic domains during the transfer process. The intuition behind our proposed method is that visually and semantically similar categories should exhibit more common transferable properties than dissimilar categories, e.g. a better detector would result by transforming the differences between a dog classifier and a dog detector onto the cat class, than would by transforming from the violin class. Experimental results on the challenging ILSVRC2013 detection dataset demonstrate that each of our proposed object similarity based knowledge transfer methods outperforms the baseline methods. We found strong evidence that visual similarity and semantic relatedness are complementary for the task, and when combined notably improve detection, achieving state-of-the-art detection performance in a semi-supervised setting.
Yuxing Tang, Josiah Wang, Boyang Gao, Emmanuel Dellandréa, Robert J. Gaizauskas, Liming Chen 0002
CVPR3
2016 Enhancing bilateral teleoperation using camera-based online virtual fixtures generation
abstract
In this paper we present an interactive system to enhance bilateral teleoperation through online virtual fixtures generation and task switching. This is achieved using a stereo camera system which provides accurate information of the surrounding environment of the robot and of the tasks that have to be performed in it. The use of the proposed approach aims at improving the performances of bilateral teleoperation systems by reducing the human operator workload and increasing both the implementation and the execution efficiency. In fact, using our method virtual guidances do not need to be programmed a priori but they can be instead automatically generated and updated making the system suitable for unstructured environments. We strengthen the proposed method using passivity control in order to safely switch between different tasks while teleoperating under active constraints. A series of experiments emulating real industrial scenarios are used to show that the switch between multiple tasks can be passively and safely achieved and handled by the system.
Mario Selvaggio, Gennaro Notomista, Fei Chen 0007, Boyang Gao, Francesco Trapani, Darwin G. Caldwell
IROS4
2016 Active colloids segmentation and tracking
Boyang Gao, Simon Masnou, Liming Chen 0002, Isaac Theurkauff, Cécile Cottin-Bizonne, Frank Y. Shih
Pattern Recognit.2
2011 Improved Prosody Generation by Maximizing Joint Probability of State and Longer Units
abstract
The current state-of-the-art hidden Markov model (HMM)-based text-to-speech (TTS) can produce highly intelligible, synthesized speech with decent segmental quality. However, its prosody, especially at phrase or sentence level, still tends to be bland. This blandness is partially due to the fact that the state-based HMM is inadequate in capturing global, hierarchical suprasegmental information in speech signals. In this paper, to improve the TTS prosody, longer units are first explicitly modeled with appropriate parametric distributions. The resultant models are then integrated with the state-based baseline models in generating better prosody by maximizing the joint probability. Experimental results in both Mandarin and English show consistent improvements over our baseline system with only state-based prosody model. The improvements are both objectively measurable and subjectively perceivable.
Yao Qian, Zhizheng Wu 0001, Boyang Gao, Frank K. Soong
IEEE Trans. Speech Audio Process.3
2010 Study on the Recognition of Objectionable Audio
abstract
In this paper, a novel method from the feature — porno-sounds recognition — point of view is proposed to detect adult video sequences automatically which may serve as a verification step, a supplementary method or an independent detector. To the specificity of erotic sound, its feature analysis is given. Based on the popular features, histograms and contours are introduced as new sets of features. At the same time due to the complexity of outside data, a general framework called in-class clustering is proposed which selects the most representative subclass for training and classification. All these efforts increase the recall rate and decrease the false positive rate. Experiments on real data from the Internet indicate that the proposed method yields superior performance with 89.17% recall rate and 10.78% false positive rate being achieved.
Ziqiang Shi, Boyang Gao, Tieran Zheng, Jiqing Han 0001
Int. J. Pattern Recognit. Artif. Intell.2
2008 Duration refinement by jointly optimizing state and longer unit likelihood
Boyang Gao, Yao Qian, Zhizheng Wu 0001, Frank K. Soong
INTERSPEECH1