Han Jiang 0003

dblp:80/10002-3 · DBLP profile ↗
← Back
19ranked-venue papers
3as first author
18since 2021 · last 2026
0000-0003-3222-4140ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 1 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 LLM-enabled universal traffic signal control across different intersections and traffic flows
Haiyang Yu 0002, Han Jiang 0003, Minda Li, Zhiyong Cui, Yilong Ren
Knowl. Based Syst.3
2026 OccSora: 4D Occupancy Generation Models as World Simulators for Autonomous Driving
abstract
Understanding the evolution of 3D scenes is crucial for autonomous driving. While conventional methods describe scene development through individual instance motions, world models provide a generative framework for modeling overall scene dynamics. However, most existing approaches rely on autoregressive next-token prediction, which suffers from error accumulation and limited global spatiotemporal reasoning, leading to degraded long-term consistency. To address these issues, we propose a diffusion-based 4D occupancy generation model, OccSora, to simulate 3D world evolution for autonomous driving. A 4D scene tokenizer is introduced to obtain compact spatiotemporal representations and enable high-quality reconstruction of long occupancy sequences. We then train a diffusion transformer on these representations to generate 4D occupancy conditioned on trajectory prompts. Experiments on the nuScenes dataset with Occ3D annotations show that OccSora can generate 16s videos with authentic 3D layout and strong temporal consistency. With trajectory-aware 4D generation, OccSora has the potential to serve as a world simulator for autonomous driving decision-making. Project page: https://wzzheng.net/OccSora.
Wenzhao Zheng, Yilong Ren, Han Jiang 0003, Zhiyong Cui, Haiyang Yu 0002, Jiwen Lu
IEEE Trans. Image Process.4
2026 OAlight: Overflow-Aware Adaptive Traffic Signal Control via State-Wise Reinforcement Learning
abstract
Adaptive traffic signal control dynamically adapts to real-time traffic flow variations, effectively mitigating traffic congestion through efficient decision-making. Current ATSC methods primarily focus on vehicle mobility metrics, such as average waiting time and average speed. However, urban road networks require not only high traffic efficiency but also proactive prevention of critical scenarios like overflow, which, if not intervened, may degrade network-wide traffic conditions and compromise safety. Therefore, in order to a balance between overflow prevention and efficiency, we model traffic signal control as a Constrained Markov Decision Process (CMDP), and propose OAlight, an overflow-aware adaptive traffic signal control framework based on reinforcement learning (RL). Unlike existing RL approaches that merely incorporate overflow-related features into the reward function, OAlight integrates overflow awareness throughout the entire RL lifecycle — environment interaction, state representation, and policy learning. Specifically, considering the explicit overflow factor of excessive queue lengths and the implicit overflow factor of excessive waiting time for individual vehicles, we construct a dual safeguard mechanism for overflow prevention through environmental interactions by developing a comprehensive reward-cost function framework. In this framework, the reward function simultaneously captures both explicit and implicit factors, while two specialized cost functions assess behaviour that violates threshold performance metrics of explicit and implicit factors, respectively. We then design a state encoder and a state-action joint encoder to extract overflow-aware representations from the state space. Finally, we employ a Proximal Policy Optimization with Lagrangian method to guide policy learning, ensuring compliance with overflow constraints while optimizing control actions. The experimental results indicate that our approach achieves superior performance on our proposed overflow cost metric without compromising traffic efficiency metrics on both synthetic and real-world datasets, compared to the current state-of-the-art traffic signal control methods. Furthermore, the environment interaction module exhibits broad compatibility with diverse RL methods, significantly reducing overflow risks.
Haiyang Yu 0002, Han Jiang 0003, Zhiyong Cui, Yilong Ren
IEEE Trans. Intell. Transp. Syst.3
2025 Authentic 4D Driving Simulation with a Video Generation Model
Wenzhao Zheng, Dalong Du, Yilong Ren, Han Jiang 0003, Zhiyong Cui, Haiyang Yu 0002, Jie Zhou 0001, Shanghang Zhang
ICCV6
2025 Capturing The Temporal Dynamics of Learner Interactions In Moocs: A Comprehensive Approach With Longitudinal And Inferential Network Analysis
abstract
While research on social network analysis is abundant and less frequently so temporal network analysis, research that uses inferential temporal network methods is barely existent. This paper aims to fill this gap by conducting a comparative analysis of temporal networks and inferential longitudinal network methods in the context of learner interactions in Massive Open Online Courses (MOOCs). We focus on three prominent methods: Temporal Network Analysis (TNA), Temporal Exponential Random Graph Models (TERGM) and Simulation Investigation for Empirical Network Analysis (SIENA). Using a five-week Nature Education MOOC as a case study, we compared the features, metrics of each method as well as their understanding of using network to analyze learner interactions. TNA focuses on describing and visualizing temporal changes in network structure, while TERGM and SIENA view networks as evolving systems influenced by individual behaviors and structural dependencies. TERGM treats network changes as a joint of random processes, while SIENA emphasizes the agency of learners and analyzes continuous network evolution. The findings provide guidelines for researchers and educators to select appropriate network analysis methods for temporal studies in educational contexts.
Mengtong Xiang, Mohammed Saqr, Han Jiang 0003, Wei Liu 0020
LAK4
2025 AC-DiT: Adaptive Coordination Diffusion Transformer for Mobile Manipulation
abstract
Recently, mobile manipulation has attracted increasing attention for enabling language-conditioned robotic control in household tasks. However, existing methods still face challenges in coordinating mobile base and manipulator, primarily due to two limitations. On the one hand, they fail to explicitly model the influence of the mobile base on manipulator control, which easily leads to error accumulation under high degrees of freedom. On the other hand, they treat the entire mobile manipulation process with the same visual observation modality (e.g., either all 2D or all 3D), overlooking the distinct multimodal perception requirements at different stages during mobile manipulation. To address this, we propose the Adaptive Coordination Diffusion Transformer (AC-DiT), which enhances mobile base and manipulator coordination for end-to-end mobile manipulation. First, since the motion of the mobile base directly influences the manipulator's actions, we introduce a mobility-to-body conditioning mechanism that guides the model to first extract base motion representations, which are then used as context prior for predicting whole-body actions. This enables whole-body control that accounts for the potential impact of the mobile base’s motion. Second, to meet the perception requirements at different stages of mobile manipulation, we design a perception-aware multimodal conditioning strategy that dynamically adjusts the fusion weights between various 2D visual images and 3D point clouds, yielding visual features tailored to the current perceptual needs. This allows the model to, for example, adaptively rely more on 2D inputs when semantic information is crucial for action prediction, while placing greater emphasis on 3D geometric information when precise spatial understanding is required. We empirically validate AC-DiT through extensive experiments on both simulated and real-world mobile manipulation tasks, demonstrating superior performance compared to existing methods.
Sixiang Chen, Jiaming Liu 0003, Siyuan Qian, Han Jiang 0003, Zhuoyang Liu, Chenyang Gu, Xiaoqi Li 0009, Chengkai Hou, Pengwei Wang 0004, Zhongyuan Wang 0006, Renrui Zhang, Shanghang Zhang
NeurIPS4
2025 CLlight: Enhancing representation of multi-agent reinforcement learning with contrastive learning for cooperative traffic signal control
Yilong Ren, Han Jiang 0003, Zhiyong Cui, Haiyang Yu 0002
Expert Syst. Appl.3
2025 Deep Reinforcement Learning With Fuzzy Feature Fusion for Cooperative Control in Traffic Light and Connected Autonomous Vehicles
abstract
A mixed traffic environment of manual driving and automatic driving will become the norm in future intelligent transportation systems. The deep reinforcement learning (DRL) method has shown significant promise in cooperative control for traffic lights and connected autonomous vehicles (CAV) in a mixed-traffic environment. However, the uncertainty and noise in integrating agents' observations can lead to inadequate exploration of environmental data by DRL algorithms. Consequently, these algorithms are prone to overfitting and becoming trapped in local optimal, which limits the performance of control strategies. To more effectively harness the gathered environmental data and thereby facilitate improved decision-making by agents, a DRL-based cooperative control method with fuzzy feature fusion (F3DRL) was proposed in this article. First, the adaptive fuzzy inference module is implemented to adaptively mitigate information uncertainty as the data from CAV is aggregated. Then, a deep information extraction module was introduced and integrated with the output of the adaptive fuzzy inference module to establish a parallel feature fusion module. The adaptive fuzzy inference module mitigates uncertainty in the extracted traffic environmental states, while the deep information extraction module facilitates the extraction of a more comprehensive environmental representation. The fusion of features derived from these two distinct modules aids DRL agents in making better action selections, which significantly enhances the effectiveness and stability of the F3DRL method. In simulations, F3DRL significantly reduced travel and delay times, fuel consumption, and CO$_{2}$emissions, outperforming both traditional and state-of-the-art methods.
Zhengyang Zhang, Han Jiang 0003, Haiyang Yu 0002, Yilong Ren
IEEE Trans. Fuzzy Syst.3
2025 Toward City-Scale Vehicular Crowd Sensing: A Decentralized Framework for Online Participant Recruitment
abstract
As an emerging urban computing paradigm, vehicle crowd sensing (VCS) leverages ubiquitous vehicles as basic sensing units to achieve more efficient data collection. However, with the expansion of the sensing range, the tens of thousands of vehicles and the openness of urban road networks pose a huge challenge for real-time participant recruitment in online VCS systems. To achieve efficient city-scale VCS, this paper proposes Dec-Recruiter, a decentralized framework for online participant recruitment. Specifically, Dec-Recruiter adopts a novel decision-making mode based on virtual grid agents, where vehicles traveling in the same direction within the same grid are considered homogeneous, simplifying the recruitment of specific vehicles to the selection of the number of vehicles in each direction. Meanwhile, through policy sharing among grid agents with the same geographic features, the complexity of city-scale VCS participant recruitment is further reduced. The core of Dec-Recruiter is a multi-agent contextual double-deep Q-network algorithm, which enables grid agents with different geographic features to collaborate on network-wide sensing tasks through their asynchronous decision-making. In this process, the Gaussian function is employed to adjust the reward distribution to address cold-start and data integrity issues in VCS. In addition, to ensure the convergence and training efficiency of the model on large-scale road networks, a pre-training-based transfer learning paradigm is also introduced. We conduct extensive experiments on both synthetic and real-world datasets. The results demonstrate that Dec-Recruiter can effectively recruit appropriate participants in the large-scale VCS and outperforms all baselines.
Han Jiang 0003, Yilong Ren, Yanan Zhao 0002, Zhiyong Cui, Haiyang Yu 0002
IEEE Trans. Intell. Transp. Syst.1
2025 RM2Occ: Re-Projection Multi-Task Multi-Sensor Fusion for Autonomous Driving 3D Object Detection and Occupancy Perception
abstract
Occupancy prediction plays a crucial role in supporting autonomous driving planning and decision-making. Existing methods typically rely on modular stacking and fusion techniques of object detection, semantic segmentation, and depth estimation to achieve 3D occupancy. However, they fail to deeply explore the transformation relationships between 2D and 3D spaces and to efficiently fuse the different characteristics of multi-source sensors. We propose R$M^{2}$Occ, the first 3D occupancy perception network that integrates multi-sensor fusion based on different sensor principles and achieves multi-task learning. To leverage the rich 2D semantic information captured by cameras and elevate it to the 3D domain, we begin by querying and populating predefined empty voxels with multi-view image features. Subsequently, we progressively fuse 3D LiDAR point clouds with these populated voxels through an unbalanced fusion strategy that effectively supplements missing information and suppresses noise. Leveraging IMU data and calibration parameters, we then re-project the enriched voxels back onto the 2D image plane according to camera coordinates, performing a secondary query using the semantic segmentation results to recover semantic details potentially lost due to radar fusion limitations and incomplete voxel querying. Finally, supported by a multi-task detection head, R$M^{2}$Occ simultaneously accomplishes 3D object detection, semantic segmentation, Bird’s Eye View (BEV) detection, and full-scene grid occupancy prediction, enabling comprehensive multi-task output. Extensive experiments and ablation studies on the nuScenes dataset demonstrate that R$M^{2}$Occ significantly outperforms existing state-of-the-art methods, establishing a new paradigm for accurate and efficient multi-sensor fusion and multi-task perception in autonomous driving scenarios.
Yilong Ren, Minda Li, Han Jiang 0003, Zhiyong Cui, Mengmeng Yang 0001, Haiyang Yu 0002, Diange Yang
IEEE Trans. Intell. Transp. Syst.4
2024 DenseKoopman: A Plug-and-Play Framework for Dense Pedestrian Trajectory Prediction
Xianbang Li, Yilong Ren, Han Jiang 0003, Haiyang Yu 0002, Yanlei Cui
IJCAI3
2024 MLF3D: Multi-Level Fusion for Multi-Modal 3D Object Detection
abstract
Recently, 3D object detection techniques based on the fusion of camera and LiDAR sensor modalities have received much attention due to their complementary capabilities. How-ever, prevalent multi-modal models are relatively homogeneous in terms of feature fusion strategies, making their performance being strictly limited to the detection results of one of the modalities. While the latest data-level fusion models based on virtual point clouds do not make further use of image features, resulting in a large amount of noise in depth estimation. To address the above issues, this paper integrates the advantages of data-level and feature-level sensor fusion, and proposes MLF3D, a 3D object detection based on multi-level fusion. MLF3D generates virtual point clouds to realize the data-level fusion, and implements feature-level fusion through two key designs: VIConv3D and ASFA. VIConv3D reduces the noise problem and realizes deep interactive enhancement of features through cross-modal fusion, noise sensing, and cross-space fusion. ASFA refines the bounding box by adaptively fusing cross-layer spatial semantic information. Our MLF3D achieves 92.91%, 87.71% AP and 85.25% AP in easy, medium and hard scenarios on the KITTI’s 3D Car Detection Leaderboard, realizing excellent performance.
Han Jiang 0003, Jianru Xiao, Yanan Zhao 0002, Wanqing Chen, Yilong Ren, Haiyang Yu 0002
IV1
2024 E-MLP: Effortless Online HD Map Construction with Linear Priors
abstract
Online High-definition map (HD-map) construction based on vehicle sensors has garnered widespread attention recently. While state-of-the-art methods achieve remarkable accuracy, most of them overlook the importance of inference speed and the inherent linear priors of map elements. Concretely, slow inference speed impacts the safety of autonomous vehicles, making it challenging for applications. Additionally, the absence of linear priors in map element predictions results in distorted or blurry outcomes. To address these issues, we propose E-MLP, an effortless online HD-map construction method that relies solely on camera sensors and incorporates the linear priors of map elements. Specifically, we first introduce a novel Principal Feature Analysis (PFA) module, designed to efficiently reduce the time cost of view transformation. Then, two thoughtfully crafted loss functions are introduced to incorporate the natural linear priors of map elements as constraints in the map construction process. Extensive experiments conducted on the nuScenes dataset revealed that, compared to the baseline method, our approach achieved a remarkable 34.9% increase in inference speed with virtually no loss in accuracy.
Ruikai Li, Hao Shan, Han Jiang 0003, Jianru Xiao, Yizhuo Chang, Haiyang Yu 0002, Yilong Ren
IV3
2024 AccidentGPT: A V2X Environmental Perception Multi-modal Large Model for Accident Analysis and Prevention
abstract
Traffic accidents are a significant factor leading to injuries and property losses, prompting extensive research in the field of traffic safety. However, previous studies, whether focused on static environment assessment, dynamic driving analysis, pre-accident prediction, or post-accident rule checks, have often been conducted independently. Our introduces V2X Environmental Perception Multi-modal Large Model AccidentGPT for accident analysis and prevention. AccidentGPT establishes a multi-modal information interaction framework based on multisensory perception. It adopts a holistic approach to address traffic safety issues, providing environmental perception for autonomous vehicles to avoid collisions and maintain control. In human-driven vehicles, it offers proactive safety warnings, blind spot alerts, and driving suggestions through human-machine dialogue. Additionally, it aids traffic police and management agencies in considering factors such as pedestrians, vehicles, roads, and the environment for intelligent real-time analysis of traffic safety. The system also conducts a thorough analysis of accident causes and post-accident liabilities, making it the first large-scale model to integrate comprehensive scene understanding into traffic safety research. Project page: https://accidentgpt.github.io
Yilong Ren, Han Jiang 0003, Pinlong Cai, Daocheng Fu, Zhiyong Cui, Haiyang Yu 0002, Xuesong Wang 0006, Hanchu Zhou, Helai Huang, Yinhai Wang
IV3
2024 MuGIL: A Multi-Graph Interaction Learning Network for Multi-Task Traffic Prediction
Haiyang Yu 0002, Han Jiang 0003, Zhenliang Ma, Zhiyong Cui, Yilong Ren
Knowl. Based Syst.3
2024 SHIP: A State-Aware Hybrid Incentive Program for Urban Crowd Sensing With for-Hire Vehicles
abstract
Benefiting from unified sensors and long-term traffic engagement, for-hire vehicles (FHVs) are widely considered the mainstay for vehicular crowd sensing (VCS) tasks. However, incentivizing FHVs to participate in sensing tasks remains a fundamental challenge for FHV-enabled VCS systems: for one thing, the distribution diversity of orders and tasks limits FHVs from participating in VCS; for another, FHVs’ operating states determine whether they are free to execute sensing tasks. To address the above issues, this article proposes SHIP, a State-aware Hybrid Incentive Program for FHV-enabled VCS systems. Our proposal finely categorizes FHVs’ operating states and provides a hybrid incentive scheme that incorporates both opportunistic and participatory approaches. We also introduce coverage diversity to reflect the distribution of FHVs and sensing tasks. By combining coverage diversity with vehicle revenue, we establish a dynamic multi-objective optimization model to select appropriate FHVs for sensing to achieve a multi-win situation. Experiments based on real-world datasets show that our proposal can effectively utilize FHVs with different operating states to improve the quality of sensing tasks while increasing FHVs’ revenue.
Han Jiang 0003, Yilong Ren, Yang Yang 0148, Haiyang Yu 0002
IEEE Trans. Intell. Transp. Syst.1
2022 An RSU Deployment Strategy Based on Traffic Demand in Vehicular Ad Hoc Networks (VANETs)
abstract
The rapid development of connected automatic vehicle (CAV) technology makes vehicularad hocnetworks (VANETs) an urgently needed research field. It includes vehicle-to-vehicle (V2V) and vehicle-to-infrastructure (V2I) message flows. A roadside unit (RSU) is an important infrastructure for V2I communication and provides roadside information services for CAVs. However, an unoptimal RSU deployment may result in RSUs failing to improve the efficiency of VANETs and compromising the capability of service to most vehicles. Motivated by this observation, this study focuses on balancing the two objectives of efficiency and coverage and establishing an RSU deployment strategy based on traffic demand. In detail, this model optimizes both the average data delivery delay in VANETs and the number of vehicles covered by RSUs. The effectiveness of the method is verified by simulation in a 4 km${\times }4$km virtual road network. We also found that: 1) if 25% of the road segments in the road network are covered by RSUs, most vehicles can be served, and the delay of VANETs can be reduced; 2) compared with the road network with low traffic demand, more RSUs need to be deployed in the road network with high traffic demand to achieve the same effect; and 3) early RSU investment is more cost effective. Our method can provide a reference for the areas where RSU investments should be made and the priority of the areas.
Haiyang Yu 0002, Runkun Liu, Zhiheng Li 0001, Yilong Ren, Han Jiang 0003
IEEE Internet Things J.5
2022 TBSM: A traffic burst-sensitive model for short-term prediction under special events
Yilong Ren, Han Jiang 0003, Haiyang Yu 0002
Knowl. Based Syst.2
2016 A Two-Layer Model for Taxi Customer Searching Behaviors Using GPS Trajectory Data
abstract
This paper proposes a two-layer decision framework to model taxi drivers' customer-search behaviors within urban areas. The first layer models taxi drivers' pickup location choice decisions, and a Huff model is used to describe the attractiveness of pickup locations. Then, a path size logit (PSL) model is used in the second layer to analyze route choice behaviors considering information such as path size, path distance, travel time, and intersection delay. Global Positioning System data are collected from more than 36 000 taxis in Beijing, China, at the interval of 30 s during six months. The Xidan district with a large shopping center is selected to validate the proposed model. Path travel time is estimated based on probe taxi vehicles on the network. The validation results show that the proposed Huff model achieved high accuracy to estimate drivers' pickup location choices. The PSL outperforms traditional multinomial logit in modeling drivers' route choice behaviors. The findings of this paper can help understand taxi drivers' customer searching decisions and provide strategies to improve the system services.
Jinjun Tang, Han Jiang 0003, Zhibin Li 0003, Meng Li 0017, Fang Liu 0021, Yinhai Wang
IEEE Trans. Intell. Transp. Syst.2