VLDB 2026 Research / reviewers in the wild / expert
Xiangbo Kong
dblp:235/8349
· DBLP profile ↗
6ranked-venue papers
0as first author
6since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Computer networks · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Poster: A Comparative Analysis of Machine Learning Models for SAT Runtime Prediction
Tomohisa Kawakami, Tomoyasu Shimada, Xiangbo Kong, Hiroyuki Tomiyama, Shigeru Yamashita |
RTCSA | 3 |
| 2025 | YOLO-EMR: Efficient Multi-Scale and Rotated Object Detection in UAV Aerial ImageryabstractWith the rapid development of UAVs, object detection in aerial imagery has become an important techniques. However, deploying real-time detection models on UAV platforms remains highly challenging due to limited computational costs, as well as the multi-scale objects and dense distribution caused by high-altitude imaging. Numerous studies have contributed valuable insights to address these challenges, yet opportunities remain for improving the balance between model efficiency and detection accuracy. Moreover, current research mainly focuses on small object detection, without fully considering the detection requirements for multi-scale and rotated objects in high-altitude imagery. To address these issues, this paper proposes a lightweight object detection model specifically designed for multi-scale and rotated object detection in high-altitude imagery. Furthermore, to better evaluate the model’s performance in UAV-based object detection in high-altitude imagery, this work uses VisDrone 2019 dataset to assess the model’s real-world performance. As a results, compared to existing object detection approaches, the proposed model achieves a strong balance between detection accuracy, model efficiency, and inference speed, reducing parameters by approximately 31% and improving inference speed by 25%, while maintaining the same detection accuracy as the best baseline methods. Haimin Yan, Xiangbo Kong, Tomoyasu Shimada, Hiroyuki Tomiyama |
SMC | 2 |
| 2024 | YOLO-UTS: Lightweight YOLOv5 for UAV Traffic Monitoring and SurveillanceabstractIn the application of drone-based target detection, both the lightweighting of the model and its accuracy are crucial. Therefore, this paper aims to achieve model lightweighting and accuracy improvement through structural improvements to the YOLO model. To meet this objective, this work optimizes the structure of the existing model to address the challenge of operational limitations in small drones, which have restricted computational capacities. By integrating lightweight modules as the backbone of the model, it enhances the feasibility of deployment on devices with limited processing power. Furthermore, the introduction of a self-attention mechanism, which is placed in the Neck of the model improves the ability to prioritize critical regions within the image. This enhancement is crucial for accurately detecting overlapping or blurred objects encountered by the moving drone. Additionally, the modification of the Intersection over Union (IoU) metric, which now considers the aspect ratio and shape of the bounding boxes further refines the target detection capabilities, ensuring more precise and reliable object localization. Since drones do not remain at a constant height or position, the angles and distances of the vehicles which they capture are bound to change. This adjustment, which modifies how the IoU quantifies overlap, allows the IoU to more accurately quantify the degree of overlap of vehicle detection bounding boxes at different angles and distances, thereby providing more reliable target detection performance. Experimental results indicate that compared to existing YOLO models, our method achieves 30% reduction in model size and 1% improvement in accuracy. Haimin Yan, Xiangbo Kong, Hiroyuki Tomiyama |
SMC | 3 |
| 2024 | YOLO-ELD: Efficient and Lightweight Detection for UAV Aerial ImageryabstractObject detection in UAV imagery has become a hot topic in recent years. However, deploying real-time detection models on UAV platforms is highly challenging due to limited computational power and memory. Moreover, the large size of UAV-captured images, the small size of objects, and their dense distribution all impact detection efficiency. Many researchers have made a series of improvements to address these issues, but they have not maintained a good balance among model size, inference speed, and accuracy. To address above difficulties, this paper proposes an efficient and lightweight model that maintains moderate detection accuracy while achieving model lightweighting and reduced inference time. Concretely, considering the higher resolution of drone-captured images, we have designed a backbone with more lightweight downsampling modules, enhancing deployment efficiency on devices with limited resources. Additionally, this work incorporates a self-attention mechanism in the feature extraction component of the model, which significantly improves the ability to process critical areas in the image, crucial for detecting small-scale and densely distributed targets. Moreover, this work designs an IoU tailored for drone aerial images, which calculates losses by focusing on the shape and scale of bounding boxes, thereby enhancing the accuracy of bounding box regression. Additionally, it uses a ratio of scale factors to control the generation of auxiliary bounding boxes, which aids in loss calculation and accelerates convergence. Experimental results on the VisDrone 2019 dataset show that compared to existing detection methods used for drones, our model is more lightweight and efficient, while also achieving medium accuracy. Also, compared to our baseline method, YOLO-ELD reduces the number of model parameters by about 40%, increases the inference speed of the model by 10%, and also improves the precision of model by 3%. Haimin Yan, Xiangbo Kong, Hiroyuki Tomiyama |
SMC | 3 |
| 2024 | Table Tennis Stroke Classification from Game Videos Using 3D Human KeypointsabstractIn this paper, we propose a classification method for table tennis strokes using player’s 3D joint coordinates. In existing studies on stroke classification, classification is performed based on videos taken by a camera installed on a table tennis table or data obtained by attaching an inertial sensor to a player. However, in these existing methods, sensors and cameras interfere with the game, and it is difficult to adapt to the game videos. Therefore, in this paper, we classify strokes from videos that can be taken during games under more practical conditions. For the classification, we use the player’s 3D joint coordinates as input to classify the game video. In our method, we use deep learning to learn two kinds of information, the recorded video and the player’s joint coordinates obtained by 3D pose estimation and perform classification. As a result of the experiment, in the classification of the rally video dataset, the accuracy is improved by 4.2~15.5% in the validation data, which is the stroke of the learned player, and by 1.8~15.0% in the test data, which is the stroke of the unlearned player, compared with the existing method which input the videos and 2D joint coordinates. Yuta Fujihara, Xiangbo Kong, Ami Tanaka, Hiroki Nishikawa, Hiroyuki Tomiyama |
VCIP | 2 |
| 2024 | Dynamic Point-Pixel Feature Alignment for Multimodal 3-D Object DetectionabstractDetection of small or distant objects is a major challenge in 3-D object detection in autonomous driving either through RGB images or LiDAR point clouds. Despite the growing popularity of sensor fusion in this task, existing fusion methods have not adequately taken into account the challenges associated with 3-D small object detection, such as semantic misalignment of small objects, caused by occlusion and calibration errors. To address this issue, we propose dynamic point-pixel feature alignment network (DPPFA-Net) for multimodal 3-D small object detection by introducing memory-based point-pixel fusion (MPPF) modules, deformable point-pixel fusion (DPPF) modules, and semantic alignment evaluator (SAE) modules. More concretely, the proposed MPPF module automatically performs intramodal and cross-modal feature interactions. The intramodal interaction reduces sensitivity to noise points, while the explicit cross-modal feature interaction based on the memory bank facilitates easier network learning and enables a more comprehensive and discriminative feature representation. The DPPF module establishes interactions exclusively with key position pixels based on a sampling strategy. This design not only guarantees a low-computational complexity but also enables adaptive fusion functionality, especially beneficial for high-resolution images. The SAE module guarantees semantic alignment of the fused features, thereby enhancing the robustness and reliability of the fusion process. Furthermore, we construct a simulated multimodal noise data set, which enables quantitative analysis of the robustness of multimodal methods under varying degrees of multimodal noise. Extensive experiments on the KITTI benchmark and challenging multimodal noisy cases show that DPPFA-Net achieves a new state-of-the-art, highlighting its effectiveness in detecting small objects. Our proposed method is compared to the first place on the KITTI leaderboard and achieves better performance by 2.07%, 6.52%, 7.18%, and 6.22% of the average precision on the varying degrees of multimodal noise cases. Xiangbo Kong, Hiroki Nishikawa, Qiuyou Lian, Hiroyuki Tomiyama |
IEEE Internet Things J. | 2 |