EDBT 2026 Demo / reviewers in the wild / expert
Hao Cheng 0008
dblp:09/5158-8
· DBLP profile ↗
13ranked-venue papers
5as first author
9since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 4 first-author · 8 since 2021Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | 4DSTR: Advancing Generative 4D Gaussians with Spatial-Temporal Rectification for High-Quality and Consistent 4D GenerationabstractRemarkable advances in recent 2D image and 3D shape generation have induced a significant focus on dynamic 4D content generation. However, previous 4D generation methods commonly struggle to maintain spatial-temporal consistency and adapt poorly to rapid temporal variations, due to the lack of effective spatial-temporal modeling. To address these problems, we propose a novel 4D generation network called 4DSTR, which modulates generative 4D Gaussian Splatting with spatial-temporal rectification. Specifically, temporal correlation across generated 4D sequences is designed to rectify deformable scales and rotations and guarantee temporal consistency. Furthermore, an adaptive spatial densification and pruning strategy is proposed to address significant temporal variations by dynamically adding or deleting Gaussian points with the awareness of their pre-frame movements. Extensive experiments demonstrate that our 4DSTR achieves state-of-the-art performance in video-to-4D generation, excelling in reconstruction quality, spatial-temporal consistency, and adaptation to rapid temporal movements. Jiuming Liu, Michael Ying Yang, Francesco Nex, Hao Cheng 0008 |
AAAI | 7 |
| 2025 | DVLO4D: Deep Visual-Lidar Odometry with Sparse Spatial-Temporal FusionabstractVisual-LiDAR odometry is a critical component for autonomous system localization, yet achieving high accuracy and strong robustness remains a challenge. Traditional approaches commonly struggle with sensor misalignment, fail to fully leverage temporal information, and require extensive manual tuning to handle diverse sensor configurations. To address these problems, we introduce DVLO4D, a novel visual-LiDAR odometry framework that leverages sparse spatial-temporal fusion to enhance accuracy and robustness. Our approach proposes three key innovations: (1) Sparse Query Fusion, which utilizes sparse LiDAR queries for effective multi-modal data fusion; (2) a Temporal Interaction and Update module that integrates temporally-predicted positions with current frame data, providing better initialization values for pose estimation and enhancing model's robustness against accumulative errors; and (3) a Temporal Clip Training strategy combined with a Collective Average Loss mechanism that aggregates losses across multiple frames, enabling global optimization and reducing the scale drift over long sequences. Extensive experiments on the KITTI and Argoverse Odometry dataset demonstrate the superiority of our proposed DVLO4D, which achieves state-of-the-art performance in terms of both pose accuracy and robustness. Additionally, our method has high efficiency, with an inference time of 82 ms, possessing the potential for the real-time deployment. Michael Ying Yang, Jiuming Liu, Sander Oude Elberink, George Vosselman, Hao Cheng 0008 |
ICRA | 8 |
| 2025 | TopoLiDM: Topology-Aware LiDAR Diffusion Models for Interpretable and Realistic LiDAR Point Cloud GenerationabstractLiDAR scene generation is critical for mitigating real-world LiDAR data collection costs and enhancing the robustness of downstream perception tasks in autonomous driving. However, existing methods commonly struggle to capture geometric realism and global topological consistency. Recent LiDAR Diffusion Models (LiDMs) predominantly embed LiDAR points into the latent space for improved generation efficiency, which limits their interpretable ability to model detailed geometric structures and preserve global topological consistency. To address these challenges, we propose TopoLiDM, a novel framework that integrates graph neural networks (GNNs) with diffusion models under topological regularization for high-fidelity LiDAR generation. Our approach first trains a topological-preserving VAE to extract latent graph representations by graph construction and multiple graph convolutional layers. Then we freeze the VAE and generate novel latent topological graphs through the latent diffusion models. We also introduce 0-dimensional persistent homology (PH) constraints, ensuring the generated LiDAR scenes adhere to real-world global topological structures. Extensive experiments on the KITTI-360 dataset demonstrate TopoLiDM’s superiority over state-of-the-art methods, achieving improvements of 22.6% lower Fréchet Range Image Distance (FRID) and 9.2% lower Minimum Matching Distance (MMD). Notably, our model also enables fast generation speed with an average inference time of 1.68 samples/s, showcasing its scalability for real-world applications. We will release the related codes at https://github.com/IRMVLab/TopoLiDM. Jiuming Liu, Tianchen Deng, Francesco Nex, Hao Cheng 0008, Hesheng Wang 0001 |
IROS | 6 |
| 2025 | Inspiring External Human-Machine Interface Designs for Autonomous Personal Mobility Vehicle: Causal Discovering the Influence of Passengers' Personality Traits on User ExperienceabstractAs autonomous personal mobility vehicles (APMVs) are increasingly integrated into shared spaces, short-distance interactions between pedestrians and APMVs will become more frequent. To facilitate communication in shared spaces, APMVs equipped with external human-machine interfaces (eHMIs). Although the eHMI is primarily designed to communicate with pedestrians, its communication also affects the APMV passenger due to the short-distance interaction. This paper focused on the effect of passengers’ personality traits on their user experience when the APMV exhibits different eHMIs. An experiment was conducted in the field with 24 participants as APMV passengers who experienced three distinct eHMI types: eHMI-T (text-based), eHMI-NV (neutral voice-based), and eHMI-AV (affective voice-based). Through causal discovery analysis, our findings revealed that when the APMV is equipped with eHMI-T, various personality traits of passengers collectively influenced their user experience. In contrast, the eHMI-NV design demonstrated that personality traits had no direct influence on user experience. The eHMI-AV design primarily showed that agreeableness and extraversion negatively influenced concerns about drawing attention, which subsequently affected other user experience. Based on the results, this paper recommends designing different eHMIs based on the APMV ownerships, such as private or public shared APMVs. Hailong Liu 0001, Zhe Zeng 0002, Yang Li 0169, Hao Cheng 0008, Takahiro Wada |
IROS | 4 |
| 2025 | Is Silent External Human-Machine Interface (eHMI) Enough? A Passenger-Centric Study on Effective eHMI for Autonomous Personal Mobility Vehicles in the FieldabstractAutonomous personal mobility vehicle (APMV) is a miniaturized autonomous vehicle designed for short-distance mobility to everyone. Due to its open design, APMV’s passengers are exposed to communications between the external human-machine interface (eHMI) on APMV and pedestrians. Therefore, effective eHMI designs for APMV need to consider potential impacts of APMV-pedestrian interactions on passengers’ subjective feelings. This study from the perspective of APMV passengers discussed three eHMI designs: (1) graphical user interface (GUI)-based eHMI with text message (eHMI-T), (2) multimodal user interface (MUI)-based eHMI with neutral voice (eHMI-NV), and (3) MUI-based eHMI with affective voice (eHMI-AV). In a riding field experiment (N = 24), eHMI-T made passengers feel awkward during the “silent time” when eHMI-T conveyed information exclusively to pedestrians, not passengers. MUI-based eHMIs with voice cues showed advantages, with eHMI-NV excelling in pragmatic quality and eHMI-AV in hedonic quality. Considering passengers’ personalities and genders in APMV eHMI design is also highlighted. Hailong Liu 0001, Yang Li 0169, Zhe Zeng 0002, Hao Cheng 0008, Takahiro Wada |
Int. J. Hum. Comput. Interact. | 4 |
| 2024 | Controllable Diverse Sampling for Diffusion Based Motion Behavior ForecastingabstractIn autonomous driving tasks, trajectory prediction in complex traffic environments requires adherence to real-world context conditions and behavior multimodalities. Existing methods predominantly rely on prior assumptions or generative models trained on curated data to learn road agents’ stochastic behavior bounded by scene constraints. However, they often face mode averaging issues due to data imbalance and simplistic priors, and could even suffer from mode collapse due to unstable training and single ground truth supervision. These issues lead the existing methods to a loss of predictive diversity and adherence to the scene constraints. To address these challenges, we introduce a novel trajectory generator named Controllable Diffusion Trajectory (CDT), which integrates map information and social interactions into a Transformer-based conditional denoising diffusion model to guide the prediction of future trajectories. To ensure multimodality, we incorporate behavioral tokens to direct the trajectory’s modes, such as going straight, turning right or left. Moreover, we incorporate the predicted endpoints as an alternative behavioral token into the CDT model to facilitate the prediction of accurate trajectories. Extensive experiments on the Argoverse 2 benchmark demonstrate that CDT excels in generating diverse and scene-compliant trajectories in complex urban settings. Yiming Xu 0005, Hao Cheng 0008, Monika Sester |
IV | 2 |
| 2023 | ForceFormer: Exploring Social Force and Transformer for Pedestrian Trajectory PredictionabstractPredicting trajectories of pedestrians based on goal information in highly interactive scenes is a crucial step toward Intelligent Transportation Systems and Autonomous Driving. The challenges of this task come from two key sources: (1) complex social interactions in high pedestrian density scenarios and (2) limited utilization of goal information to effectively associate with past motion information. To address these difficulties, we integrate social forces into a Transformer-based stochastic generative model backbone and propose a new goal-based trajectory predictor called ForceFormer. Differentiating from most prior works that simply use the destination position as an input feature, we leverage the driving force from the destination to efficiently simulate the guidance of a target on a pedestrian. Additionally, repulsive forces are used as another input feature to describe the avoidance action among neighboring pedestrians. Extensive experiments show that our proposed method achieves on-par performance measured by distance errors with the state-of-the-art models but evidently decreases collisions, especially in dense pedestrian scenarios on widely used pedestrian datasets. Weicheng Zhang, Hao Cheng 0008, Fatema T. Johora, Monika Sester |
IV | 2 |
| 2023 | Hybrid POI group recommender system based on group type in LBSN
Zahra Bahari Sojahrood, Mohammad Taleai, Hao Cheng 0008 |
Expert Syst. Appl. | 3 |
| 2021 | Exploring Dynamic Context for Multi-path Trajectory PredictionabstractTo accurately predict future positions of different agents in traffic scenarios is crucial for safely deploying intelligent autonomous systems in the real-world environment. However, it remains a challenge due to the behavior of a target agent being affected by other agents dynamically and there being more than one socially possible paths the agent could take. In this paper, we propose a novel framework, named Dynamic Context Encoder Network (DCENet). In our framework, first, the spatial context between agents is explored by using self-attention architectures. Then, the two-stream encoders are trained to learn temporal context between steps by taking the respective observed trajectories and the extracted dynamic spatial context as input. The spatial-temporal context is encoded into a latent space using a Conditional Variational Auto-Encoder (CVAE) module. Finally, a set of future trajectories for each agent is predicted conditioned on the learned spatial-temporal context by sampling from the latent space, repeatedly. DCENet is evaluated on one of the most popular challenging benchmarks for trajectory forecasting Trajnet and reports a new state-of-the-art performance. It also demonstrates superior performance evaluated on the benchmark inD for mixed traffic at intersections. A series of ablation studies is conducted to validate the effectiveness of each proposed module. Our code is available at https://github.com/wtliao/DCENet. Hao Cheng 0008, Wentong Liao, Xuejiao Tang, Michael Ying Yang, Monika Sester, Bodo Rosenhahn |
ICRA | 1 |
| 2020 | Automatic Interaction Detection Between Vehicles and Vulnerable Road Users During Turning at an IntersectionabstractInteraction detection between vehicles and vulnerable road users (e.g. pedestrians and cyclists) is important for e.g. safety control and autonomous driving. However, there are many challenges for automatically detecting interactions, such as the ambiguity of defining when interaction is required in dynamic traffic activities among different road users and the lack of labeled data for training a machine learning detector. To overcome the challenges, we introduce a way to define whether or not interaction is required in various traffic scenes and create a large real-world dataset from a very challenging intersection. A sequence-to-sequence method that uses the object information and motion information of the traffic scenes extracted by a state-of-the-art object detector and from optical flow, respectively, is proposed for automatic interaction detection. The proposed method generates a probability of interaction at each short interval (<; 0.1 s) that represents the changing of interaction along a sequence. We obtain a baseline model that differentiates no interaction from interaction on the basis of the location and road user type from the detected object information. Compared with the baseline model, the empirical results of the proposed method demonstrate very accurate predictions for vehicle turning sequences with varying length. Hao Cheng 0008, Hailong Liu 0001, Fumito Shinmura, Naoki Akai, Hiroshi Murase, Takatsugu Hirayama |
IV | 1 |
| 2020 | Trajectory Modelling in Shared Spaces: Expert-Based vs. Deep Learning Approach?
Hao Cheng 0008, Fatema T. Johora, Monika Sester, Jörg P. Müller |
MABS | 1 |
| 2019 | Pedestrian Group Detection in Shared SpaceabstractIn shared space, pedestrians are often found walking in groups and behaving differently than individual pedestrians. However, automatically detecting pedestrian groups with high accuracy is not trivial given the dynamic environment and interactions in mixed traffic. Instead of tedious manual work and in order to cope with large scales of data, we propose a time-sequence Density-Based Spatial Clustering of Applications with Noise (DBSCAN) for pedestrian group detection. It is based on coexisting time and Euclidean distance between pedestrians. Our approach outputs reliable results with high IoU values. It can be easily adapted to other groups, e.g., cyclists and animals. In addition to individual behavior, the output data with differentiation of group behavior can be used in further studies in intent detection and motion prediction. Hao Cheng 0008, Yao Li 0033, Monika Sester |
IV | 1 |
| 2017 | The Influence of City Size on Dietary Choices and Food RecommendationabstractContextual features have been leveraged by recommender systems in many different domains. Traditional contextual features -- such as location and time -- have successfully been combined with collaborative filtering or content-based features. However, it is likely that there are other -- domain-specific -- features that may have even more impact. In this paper, we focus on the influence of city size on food preferences. Apart from location and time, our results show that city size can significantly boost the performance of food recommendation. Hao Cheng 0008, Markus Rokicki, Eelco Herder |
UMAP | 1 |