Haiming Gang

dblp:237/9503 · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
4since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Systems, architecture and hardware · 2 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Autonomous driving · 34% 3D vision · 33% Vision and language · 15%

Topics — the 10 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision
3d scene understanding
0.912025
MM-Spatial: Exploring 3D Spatial Understanding in Multimodal LLMs · ICCV 2025
Knowledge, reasoning and agents › Knowledge representation and reasoning › spatial reasoning
3d spatial reasoning
0.912025
MM-Spatial: Exploring 3D Spatial Understanding in Multimodal LLMs · ICCV 2025
Computer vision › Vision and language › vision-language model
multimodal large language model
0.912025
MM-Spatial: Exploring 3D Spatial Understanding in Multimodal LLMs · ICCV 2025
Computer vision › 3D vision
spatial understanding
0.912025
MM-Spatial: Exploring 3D Spatial Understanding in Multimodal LLMs · ICCV 2025
Robotics › Autonomous driving
behavior prediction
0.612022
Important Object Identification with Semi-Supervised Learning for Autonomous Driving · ICRA 2022
Robotics › Autonomous driving
trajectory prediction
0.512021
LOKI: Long Term and Key Intentions for Trajectory Prediction · ICCV 2021
Robotics › Autonomous driving
perception
0.412019
The H3D Dataset for Full-Surround 3D Multi-Object Detection and Tracking in Crowded Urban Scenes · ICRA 2019
Machine learning › Learning paradigms
semi-supervised learning
0.212022
Important Object Identification with Semi-Supervised Learning for Autonomous Driving · ICRA 2022
Computer vision › 3D vision
3d object detection
0.112019
The H3D Dataset for Full-Surround 3D Multi-Object Detection and Tracking in Crowded Urban Scenes · ICRA 2019
Computer vision › 3D vision › 3d object detection › point cloud object detection
LiDAR-based 3D object detection
0.112019
The H3D Dataset for Full-Surround 3D Multi-Object Detection and Tracking in Crowded Urban Scenes · ICRA 2019

Methods — techniques the papers use, named apart from their topics

multimodal large language model · 0.9semi-supervised learning · 0.6relational reasoning · 0.6attention mechanism · 0.6recurrent reasoning · 0.5intention modeling · 0.5labeling methodology · 0.4benchmark evaluation · 0.4
YearPublicationVenuePosition
2025 MM-Spatial: Exploring 3D Spatial Understanding in Multimodal LLMs
Erik A. Daxberger, Nina Wenzel, David Griffiths, Haiming Gang, Justin Lazarow, Gefen Kohavi, Marcin Eichner, Yinfei Yang, Afshin Dehghan, Peter Grasch
ICCV4
2022 Important Object Identification with Semi-Supervised Learning for Autonomous Driving
abstract
Accurate identification of important objects in the scene is a prerequisite for safe and high-quality decision making and motion planning of intelligent agents (e.g., autonomous vehicles) that navigate in complex and dynamic environments. Most existing approaches attempt to employ attention mechanisms to learn importance weights associated with each object indirectly via various tasks (e.g., trajectory prediction), which do not enforce direct supervision on the importance estimation. In contrast, we tackle this task in an explicit way and formulate it as a binary classification (“important” or “unimportant”) problem. We propose a novel approach for important object identification in egocentric driving scenarios with relational reasoning on the objects in the scene. Besides, since human annotations are limited and expensive to obtain, we present a semi-supervised learning pipeline to enable the model to learn from unlimited unlabeled data. Moreover, we propose to leverage the auxiliary tasks of ego vehicle behavior prediction to further improve the accuracy of importance estimation. The proposed approach is evaluated on a public egocentric driving dataset (H3D) collected in complex traffic scenarios. A detailed ablative study is conducted to demonstrate the effectiveness of each model component and the training strategy. Our approach also outperforms rule-based baselines by a large margin.
Jiachen Li 0001, Haiming Gang, Hengbo Ma, Masayoshi Tomizuka, Chiho Choi
ICRA2
2021 Semi-supervised 3D Object Detection via Temporal Graph Neural Networks
abstract
3D object detection plays an important role in autonomous driving and other robotics applications. However, these detectors usually require training on large amounts of annotated data that is expensive and time-consuming to collect. Instead, we propose leveraging large amounts of unlabeled point cloud videos by semi-supervised learning of 3D object detectors via temporal graph neural networks. Our insight is that temporal smoothing can create more accurate detection results on unlabeled data, and these smoothed detections can then be used to retrain the detector. We learn to perform this temporal reasoning with a graph neural network, where edges represent the relationship between candidate detections in different time frames. After semi-supervised learning, our method achieves state-of-the-art detection performance on the challenging nuScenes [3] and H3D [19] benchmarks, compared to baselines trained on the same amount of labeled data. Project and code are released at https://www.jianrenw.com/SOD-TGNN/.
Jianren Wang, Haiming Gang, Siddarth Ancha, Yi-Ting Chen 0001, David Held
3DV2
2021 LOKI: Long Term and Key Intentions for Trajectory Prediction
abstract
Recent advances in trajectory prediction have shown that explicit reasoning about agents’ intent is important to accurately forecast their motion. However, the current research activities are not directly applicable to intelligent and safety critical systems. This is mainly because very few public datasets are available, and they only consider pedestrian-specific intents for a short temporal horizon from a restricted egocentric view. To this end, we propose LOKI (LOng term and Key Intentions), a novel large-scale dataset that is designed to tackle joint trajectory and intention prediction for heterogeneous traffic agents (pedestrians and vehicles) in an autonomous driving setting. The LOKI dataset is created to discover several factors that may affect intention, including i) agent’s own will, ii) social interactions, iii) environmental constraints, and iv) contextual information. We also propose a model that jointly performs trajectory and intention prediction, showing that recurrently reasoning about intention can assist with trajectory prediction. We show our method outperforms state-of-the-art trajectory prediction methods by upto 27% and also provide a baseline for frame-wise intention estimation. The dataset is available at https://usa.honda-ri.com/loki
Harshayu Girase, Haiming Gang, Srikanth Malla, Jiachen Li 0001, Akira Kanehara, Karttikeya Mangalam, Chiho Choi
ICCV2
2019 The H3D Dataset for Full-Surround 3D Multi-Object Detection and Tracking in Crowded Urban Scenes
abstract
3D multi-object detection and tracking are crucial for traffic scene understanding. However, the community pays less attention to these areas due to the lack of a standardized benchmark dataset to advance the field. Moreover, existing datasets (e.g., KITTI [1]) do not provide sufficient data and labels to tackle challenging scenes where highly interactive and occluded traffic participants are present. To address the issues, we present the Honda Research Institute 3D Dataset (H3D), a large-scale full-surround 3D multi-object detection and tracking dataset collected using a 3D LiDAR scanner. H3D comprises of 160 crowded and highly interactive traffic scenes with a total of 1 million labeled instances in 27,721 frames. With unique dataset size, rich annotations, and complex scenes, H3D is gathered to stimulate research on full-surround 3D multi-object detection and tracking. To effectively and efficiently annotate a large-scale 3D point cloud dataset, we propose a labeling methodology to speed up the overall annotation cycle. A standardized benchmark is created to evaluate full-surround 3D multi-object detection and tracking algorithms. 3D object detection and tracking algorithms are trained and tested on H3D. Finally, sources of errors are discussed for the development of future algorithms.
Abhishek Patil, Srikanth Malla, Haiming Gang, Yi-Ting Chen 0001
ICRA3