Pinhao Song

dblp:263/1928 · DBLP profile ↗
← Back
13ranked-venue papers
2as first author
12since 2021 · last 2025
0000-0002-3950-0403ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 2 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 6 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 See&Trek: Training-Free Spatial Prompting for Multimodal Large Language Model
abstract
We introduce See&Trek, the first training-free prompting framework tailored to enhance the spatial understanding of Multimodal Large Language Models (MLLMs) under vision-only constraints. While prior efforts have incorporated modalities like depth or point clouds to improve spatial reasoning, purely visual-spatial understanding remains underexplored. See&Trek addresses this gap by focusing on two core principles: increasing visual diversity and motion reconstruction. For visual diversity, we conduct Maximum Semantic Richness Sampling, which employs an off-the-shell perception model to extract semantically rich keyframes that capture scene structure. For motion reconstruction, we simulate visual trajectories and encode relative spatial positions into keyframes to preserve both spatial relations and temporal coherence. Our method is training&GPU-free, requiring only a single forward pass, and can be seamlessly integrated into existing MLLMs. Extensive experiments on the VSI-Bench and STI-Bench show that See&Trek consistently boosts various MLLMs performance across diverse spatial reasoning tasks with the most +3.5% improvement, offering a promising path toward stronger spatial intelligence.
Pengteng Li, Pinhao Song, Wuyang Li, Huizai Yao, Weiyu Guo, Yijie Xu, Dugang Liu, Hui Xiong 0001
NeurIPS2
2024 Robot Trajectron: Trajectory Prediction-based Shared Control for Robot Manipulation
abstract
We address the problem of (a) predicting the trajectory of an arm reaching motion, based on a few seconds of the motion’s onset, and (b) leveraging this predictor to facilitate shared-control manipulation tasks, by reducing the operator’s cognitive load through assistance in their anticipated direction of motion. Our novel intent estimator, dubbed the Robot Trajectron (RT), produces a probabilistic representation of the robot’s anticipated trajectory based on its recent position, velocity and acceleration history. By taking arm dynamics into account, RT can capture the operator’s intent better than other SOTA models that only use the arm’s position, making it particularly well-suited to assist in tasks where the operator’s intent is susceptible to change. We derive a novel shared-control solution that combines RT’s predictive capacity to a representation of the locations of potential reaching targets. Our experiments demonstrate RT’s effectiveness in both intent estimation and shared-control tasks. We will make the code and data supporting our experiments publicly available at https://gitlab.kuleuven.be/detry-lab/public/robot-trajectron
Pinhao Song, Pengteng Li, Erwin Aertbeliën, Renaud Detry
ICRA1
2024 OTOcc: Optimal Transport for Occupancy Prediction
Pengteng Li, Ying He 0006, F. Richard Yu, Pinhao Song, Xingchen Zhou, Guang Zhou
IJCAI4
2024 A gated cross-domain collaborative network for underwater object detection
Linhui Dai, Hong Liu 0008, Pinhao Song, Mengyuan Liu 0001
Pattern Recognit.3
2024 Neural ordinary differential equation for irregular human motion prediction
Hong Liu 0008, Pinhao Song, Wenhao Li 0002
Pattern Recognit. Lett.3
2023 Bagging R-CNN: Ensemble for Object Detection in Complex Traffic Scenes
abstract
Generic object detection methods have achieved preferable results, but it is still challenging to detect objects from complicated traffic scenes like extreme illumination and adverse weather. The existing methods are not robust enough to be extended to new complex traffic scenes. To address this issue, we leverage the idea of ensemble learning for strong robustness and propose a novel Bagging R-CNN framework. Specially, we design a bagging classification branch that uses adaptive sampling to train base learners and make them different from each other. The final predictions are the ensemble of the base learners, achieving strong robustness to the challenging objects. For localizing more accurately, a progressive regression branch is proposed in which bounding boxes are continuously optimized for high quality. Extensive experiment results on TJU-DHD-traffic and Pascal VOC datasets show that our Bagging R-CNN achieves superior detection accuracy over state-of-the-art methods. The source code can be found at https://github.com/PungTeng/BaggingRCNN.
Pengteng Li, Ying He 0006, Dongfu Yin, F. Richard Yu, Pinhao Song
ICASSP5
2023 IGG: Improved Graph Generation for Domain Adaptive Object Detection
abstract
Domain Adaptive Object Detection (DAOD) transfers an object detector from a labeled source domain to a novel unlabeled target domain. Recent works bridge the domain gap by aligning cross-domain pixel-pairs in the non-euclidean graphical space and minimizing the domain discrepancy for adapting semantic distribution. Though great successes, these methods model graphs roughly with coarse semantic sampling due to ignoring the non-informative noises and failing to concentrate on precise semantics alignment. Besides, the coarse graph generation inevitably contains abnormal nodes. These challenges result in biased domain adaptation. Therefore, we propose an Improved Graph Generation (IGG) framework which conducts high-quality graph generation for DAOD. Specifically, we design an Intensive Node Refinement (INR) module that reconstructs the noisy sampled nodes with a memory bank, and contrastively regularizes the noisy features. For better semantics alignment, we decouple the domain-specific style and category-invariant content encoded in graph covariance and selectively eliminate only the domain-specific style. Then, a Precision Graph Optimization (PGO) adaptor is proposed which utilizes the variational inference to down-weight abnormal nodes. Comprehensive experiments on three adaptation benchmarks demonstrate that IGG achieves state-of-the-art results in unsupervised domain adaptation.
Pengteng Li, Ying He 0006, F. Richard Yu, Pinhao Song, Dongfu Yin, Guang Zhou
ACM Multimedia4
2023 Achieving domain generalization for underwater object detection by domain mixup and contrastive learning
Pinhao Song, Hong Liu 0008, Linhui Dai, Xiaochuan Zhang, Runwei Ding, Shengquan Li 0001
Neurocomputing2
2023 Boosting R-CNN: Reweighting R-CNN samples by RPN's error for underwater object detection
Pinhao Song, Pengteng Li, Linhui Dai
Neurocomputing1
2023 AO2-DETR: Arbitrary-Oriented Object Detection Transformer
abstract
Arbitrary-oriented object detection (AOOD) is a challenging task to detect objects in the wild with arbitrary orientations and cluttered arrangements. Existing approaches are mainly based on anchor-based boxes or dense points, which rely on complicated hand-designed processing steps and inductive bias, such as anchor generation, transformation, and non-maximum suppression reasoning. Recently, the emerging transformer-based approaches view object detection as a direct set prediction problem that effectively removes the need for hand-designed components and inductive biases. In this paper, we propose an Arbitrary-Oriented Object DEtection TRansformer framework, termed AO2-DETR, which comprises three dedicated components. More precisely, an oriented proposal generation mechanism is proposed to explicitly generate oriented proposals, which provides better positional priors for pooling features to modulate the cross-attention in the transformer decoder. An adaptive oriented proposal refinement module is introduced to extract rotation-invariant region features and eliminate the misalignment between region features and objects. And a rotation-aware set matching loss is used to ensure the one-to-one matching process for direct set prediction without duplicate predictions. Our method considerably simplifies the overall pipeline and presents a new AOOD paradigm. Comprehensive experiments on several challenging datasets show that our method achieves superior performance on the AOOD task.
Linhui Dai, Hong Liu 0008, Hao Tang 0005, Pinhao Song
IEEE Trans. Circuits Syst. Video Technol.5
2022 Pose-Guided Feature Disentangling for Occluded Person Re-identification Based on Transformer
abstract
Occluded person re-identification is a challenging task as human body parts could be occluded by some obstacles (e.g. trees, cars, and pedestrians) in certain scenes. Some existing pose-guided methods solve this problem by aligning body parts according to graph matching, but these graph-based methods are not intuitive and complicated. Therefore, we propose a transformer-based Pose-guided Feature Disentangling (PFD) method by utilizing pose information to clearly disentangle semantic components (e.g. human body or joint parts) and selectively match non-occluded parts correspondingly. First, Vision Transformer (ViT) is used to extract the patch features with its strong capability. Second, to preliminarily disentangle the pose information from patch information, the matching and distributing mechanism is leveraged in Pose-guided Feature Aggregation (PFA) module. Third, a set of learnable semantic views are introduced in transformer decoder to implicitly enhance the disentangled body part features. However, those semantic views are not guaranteed to be related to the body without additional supervision. Therefore, Pose-View Matching (PVM) module is proposed to explicitly match visible body parts and automatically separate occlusion features. Fourth, to better prevent the interference of occlusions, we design a Pose-guided Push Loss to emphasize the features of visible body parts. Extensive experiments over five challenging datasets for two tasks (occluded and holistic Re-ID) demonstrate that our proposed PFD is superior promising, which performs favorably against state-of-the-art methods. Code is available at https://github.com/WangTaoAs/PFD_Net
Hong Liu 0008, Pinhao Song, Tianyu Guo 0001, Wei Shi 0009
AAAI3
2022 Excavating RoI Attention for Underwater Object Detection
abstract
Self-attention is one of the most successful designs in deep learning, which calculates the similarity of different tokens and reconstructs the feature based on the attention matrix. Originally designed for NLP, self-attention is also popular in computer vision, and can be categorized into pixel-level attention and patch-level attention. In object detection, RoI features can be seen as patches from base feature maps. This paper aims to apply the attention module to RoI features to improve performance. Instead of employing an original self-attention module, we choose the external attention module, a modified self-attention with reduced parameters. With the proposed Double Head structure and the Positional Encoding module, our method can achieve promising performance in object detection. The comprehensive experiments show that it achieves promising performance, especially in the under-water object detection dataset. The code will be avaiable in: https://github.com/zsyasd/Excavating-RoI-Attention-for-Underwater-Object-Detection
Xutao Liang, Pinhao Song
ICIP2
2020 Towards Domain Generalization In Underwater Object Detection
abstract
A General Underwater Object Detector (GUOD) should perform well on most of underwater circumstances. However, with limited underwater dataset, conventional object detection methods suffer from domain shift severely. This paper aims to build a GUOD using small underwater dataset with limited types of water quality. First, we propose a data augmentation method Water Quality Transfer (WQT) to increase domain diversity of the original small dataset. Second, for mining the semantic information from data generated by WQT, Domain Generalization YOLO (DG-YOLO) is proposed, which consists of three parts: YOLOv3, Domain Invariant Module and Invariant Risk Minimization penalty. Finally, experiments on original and synthetic URPC2019 dataset prove that WQT combined with DG-YOLO achieves promising performance of domain generalization in underwater object detection. The source code can be found at https://github.com/mousecpn/DG-YOLO.
Hong Liu 0008, Pinhao Song, Runwei Ding
ICIP2