VLDB 2026 Research / reviewers in the wild / expert
Bokui Chen
dblp:116/6220
· DBLP profile ↗
17ranked-venue papers
2as first author
16since 2021 · last 2026
0000-0002-4947-5619ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 7 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | EvolveNav: Empowering LLM-Based Vision-Language Navigation via Self-Improving Embodied ReasoningabstractRecent studies have revealed the potential of training open-source Large Language Models (LLMs) to unleash LLMs' reasoning ability for enhancing vision-language navigation (VLN) performance, and simultaneously mitigate the domain gap between LLMs' training corpus and the VLN task. However, these approaches predominantly adopt straightforward input-output mapping paradigms, causing the mapping learning difficult and the navigational decisions unexplainable. Chain-of-Thought (CoT) training is a promising way to improve both navigational decision accuracy and interpretability, while the complexity of the navigation task makes the perfect CoT labels unavailable and may lead to overfitting through pure CoT supervised fine-tuning. To address these issues, we propose EvolveNav, a novel sElf-improving embodied reasoning paradigm that realizes adaptable and generalizable navigational reasoning for boosting LLM-based vision-language Navigation. Specifically, EvolveNav involves a two-stage training process: (1) Formalized CoT Supervised Fine-Tuning, where we train the model with curated formalized CoT labels to first activate the model's navigational reasoning capabilities, and simultaneously increase the reasoning speed; (2) Self-Reflective Post-Training, where the model is iteratively trained with its own reasoning outputs as self-enriched CoT labels to enhance the supervision diversity. A self-reflective auxiliary task is also designed to encourage the model to learn correct reasoning patterns by contrasting with wrong ones. Experimental results under both task-specific and cross-task training paradigms demonstrate the consistent superiority of EvolveNav over previous LLM-based VLN approaches on various popular benchmarks, including R2R, REVERIE, CVDN, and SOON. EvolveNav open avenues for exploring effective self-improving reasoning paradigms, enabling building agents capable of self-evolving for promoting LLM-based embodied AI research. Bingqian Lin, Yunshuang Nie, Khun Loun Zai, Ziming Wei 0001, Mingfei Han 0002, Rongtao Xu, Minzhe Niu, Jianhua Han, Hanwang Zhang, Liang Lin 0004, Bokui Chen, Cewu Lu, Xiaodan Liang |
IEEE Trans. Pattern Anal. Mach. Intell. | 11 |
| 2025 | Learning to Stabilize Column GenerationabstractColumn generation is a widely adopted technique for solving linear programming problems with a large number of variables. However, standard column generation often suffers from slow convergence due to the dual solution instability. In this paper, we present a novel learning-based stabilization approach for column generation. Unlike traditional methods that address dual solution stabilization at each iteration in isolation, our method adopts a holistic perspective, leveraging its learning-based nature to explore for optimal stabilization policies that lead to faster overall convergence. We frame dual solution stabilization as a sequential decision-making problem and cast column generation as a Markov decision process. A graph convolutional neural network-based agent is employed to improve dual solution quality at each iteration. Additionally, we introduce a two-stage training scheme that combines supervised learning and reinforcement learning, ensuring stable and efficient training of the agent. Experimental evaluations on cutting stock and vertex coloring problems demonstrate that our approach outperforms several well-known stabilization methods in terms of iteration efficiency and exhibits competitive performance in terms of total runtime. Furthermore, our method shows strong generalization capabilities, performing well on significantly larger problem instances and diverse benchmarks. Lichang Fang, Haofeng Yuan, Shiji Song, Bokui Chen |
IJCNN | 4 |
| 2025 | PhyBlock: A Progressive Benchmark for Physical Understanding and Planning via 3D Block AssemblyabstractWhile vision-language models (VLMs) have demonstrated promising capabilities in reasoning and planning for embodied agents, their ability to comprehend physical phenomena, particularly within structured 3D environments, remains severely limited. To close this gap, we introduce PhyBlock, a progressive benchmark designed to assess VLMs on physical understanding and planning through robotic 3D block assembly tasks. PhyBlock integrates a novel four-level cognitive hierarchy assembly task alongside targeted Visual Question Answering (VQA) samples, collectively aimed at evaluating progressive spatial reasoning and fundamental physical comprehension, including object properties, spatial relationships, and holistic scene understanding. PhyBlock includes 2600 block tasks (400 assembly tasks, 2200 VQA tasks) and evaluates models across three key dimensions: partial completion, failure diagnosis, and planning robustness. We benchmark 23 state-of-the-art VLMs, highlighting their strengths and limitations in physically grounded, multi-step planning. Our empirical findings indicate that the performance of VLMs exhibits pronounced limitations in high-level planning and reasoning capabilities, leading to a notable decline in performance for the growing complexity of the tasks.Error analysis reveals persistent difficulties in spatial orientation and dependency reasoning.We position PhyBlock as a unified testbed to advance embodied reasoning, bridging vision-language understanding and real-world physical problem-solving. Jiajun Wen 0003, Rongtao Xu, Xiwen Liang, Bingqian Lin, Ziming Wei 0001, Haokun Lin, Mingfei Han 0002, Meng Cao 0002, Bokui Chen, Ivan Laptev, Xiaodan Liang |
NeurIPS | 13 |
| 2025 | Embodied Cognition Augmented End2End Autonomous DrivingabstractIn recent years, vision-based end-to-end autonomous driving has emerged as a new paradigm. However, popular end-to-end approaches typically rely on visual feature extraction networks trained under label supervision. This limited supervision framework restricts the generality and applicability of driving models. In this paper, we propose a novel paradigm termed $E^{3}AD$, which advocates for comparative learning between visual feature extraction networks and the general EEG large model, in order to learn latent human driving cognition for enhancing end-to-end planning. In this work, we collected a cognitive dataset for the mentioned contrastive learning process. Subsequently, we investigated the methods and potential mechanisms for enhancing end-to-end planning with human driving cognition, using popular driving models as baselines on publicly available autonomous driving datasets. Both open-loop and closed-loop tests are conducted for a comprehensive evaluation of planning performance. Experimental results demonstrate that the $E^{3}AD$ paradigm significantly enhances the end-to-end planning performance of baseline models. Ablation studies further validate the contribution of driving cognition and the effectiveness of comparative learning process. To the best of our knowledge, this is the first work to integrate human driving cognition for improving end-to-end autonomous driving planning. It represents an initial attempt to incorporate embodied cognitive data into end-to-end autonomous driving, providing valuable insights for future brain-inspired autonomous driving systems. Our code will be made available at https://github.com/AIR-DISCOVER/E-cubed-AD. Ling Niu, Xiaoji Zheng, Ziyuan Yang 0005, Bokui Chen, Jiangtao Gong |
NeurIPS | 6 |
| 2025 | ChordPrompt: Orchestrating Cross-Modal Prompt Synergy for Multi-domain Incremental Learning in CLIP
Bokui Chen |
ECML/PKDD (8) | 2 |
| 2025 | CPIR: Multimodal Industrial Anomaly Detection via Latent Bridged Cross-modal Prediction and Intra-modal Reconstruction
Wen Shangguan, Hongqiang Wu, Yanchang Niu, Haonan Yin, Bokui Chen, Biqing Huang |
Adv. Eng. Informatics | 6 |
| 2025 | Eco-Driving Decision Making Based on V2X Communication and Spatio-Temporal Prediction of PedestriansabstractThe operational dynamics of vehicular transportation significantly influence energy expenditure and contribute to the escalation of global warming. However, a noticeable gap exists in the availability of Eco-driving methodologies tailored to mitigate conflicts between pedestrians and vehicles. In response, this study proposes a Vehicle-to-Pedestrian communication Eco-driving (V2P Eco-driving) strategy that operates without traffic lights and incorporates collaborative pedestrian trajectory prediction. Its performance is evaluated through a comparative study with the Ecological Intelligent Traffic Lights System (Eco-ITLS) strategy, which adjusts traffic light phases based on pedestrian and vehicle flow detection. To enhance the generalization of the prediction model, pedestrian social interactions are modeled using relative displacement and velocity metrics, while Kalman filtering mitigates systemic discrepancies in vehicular and infrastructural components. A modified distance-discrete dynamic programming (D-DDP) algorithm, accounting for remaining travel time, is introduced to optimize eco-friendly vehicle actions. The algorithm is benchmarked against other Eco-driving algorithms in terms of solution quality, memory consumption, and computational efficiency. Experimental results demonstrate that the proposed model achieves a balance between computational efficiency and solution quality. Real-world data validation and parameter calibration confirm its practicality. Simulations further highlight the V2P Eco-driving strategy’s significant potential for reducing energy consumption and emissions compared to conventional traffic light-based Eco-driving strategies. Ling Niu, Qi Wang 0081, Bokui Chen, Yingping Zhao, Yi Zhang 0029 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2025 | Driving Risk Field Model and Its Application in Trajectory Planning: A New PerspectiveabstractDriving risk field (DRF) emerges as an effective way to assess the driving safety of connected and automated vehicles (CAVs). Most existing DRF models are established from the so-called birds-eye-view (BEV), which limits their accuracy for distributed vehicle-level tasks such as trajectory planning since the interactions between ego vehicle (EV) and its surrounding traffic environment have not been fully considered. To fill this research gap, we establish a novel DRF model from ego-vehicle-view (EVV) and apply it in trajectory planning in this paper. Firstly, the collision boundary between EV and its surrounding obstacles is defined by introducing the elliptical model to fully consider the geometry characteristics of vehicles. Secondly, the relative motion influence coefficient is designed to accurately characterize the relative motion between EV and obstacles, instead of using only basic driving state information such as location and velocity. On this basis, the unified DRF is established from EVV for driving safety assessment, which contains vehicle risk field (VRF) and lane marking risk field (LMRF). Based on the established DRF model, we then design a rolling trajectory planning method (RTPM) with a rolling horizon strategy, which not only ensures a long prediction horizon but also effectively reduces the computational complexity. Multiple simulation results under different traffic scenarios jointly verify the accuracy and applicability of the proposed RTPM and DRF model established from this new perspective. Huaxin Pei, Yi Zhang 0029, Danya Yao, Li Xiao 0006, Bokui Chen |
IEEE Trans. Intell. Transp. Syst. | 7 |
| 2024 | INCPrompt: Task-Aware Incremental Prompting for Rehearsal-Free Class-Incremental LearningabstractThis paper introduces INCPrompt, an innovative continual learning solution that effectively addresses catastrophic forgetting. INCPrompt’s key innovation lies in its use of adaptive key-learner and task-aware prompts that capture task-relevant information. This unique combination encapsulates general knowledge across tasks and encodes task-specific knowledge. Our comprehensive evaluation across multiple continual learning benchmarks demonstrates INCPrompt’s superiority over existing algorithms, showing its effectiveness in mitigating catastrophic forgetting while maintaining high performance. These results highlight the significant impact of task-aware incremental prompting on continual learning performance. Xiaoyang Qu, Jing Xiao 0006, Bokui Chen, Jianzong Wang |
ICASSP | 4 |
| 2024 | P2DT: Mitigating Forgetting in Task-Incremental Learning with Progressive Prompt Decision TransformerabstractCatastrophic forgetting poses a substantial challenge for managing intelligent agents controlled by a large model, causing performance degradation when these agents face new tasks. In our work, we propose a novel solution - the Progressive Prompt Decision Transformer (P2DT). This method enhances a transformer-based model by dynamically appending decision tokens during new task training, thus fostering task-specific policies. Our approach mitigates forgetting in continual and offline reinforcement learning scenarios. Moreover, P2DT leverages trajectories collected via traditional reinforcement learning from all tasks and generates new taskspecific tokens during training, thereby retaining knowledge from previous studies. Preliminary results demonstrate that our model effectively alleviates catastrophic forgetting and scales well with increasing task environments. Xiaoyang Qu, Jing Xiao 0006, Bokui Chen, Jianzong Wang |
ICASSP | 4 |
| 2024 | Task-agnostic Decision Transformer for Multi-type Agent Control with Federated Split TrainingabstractWith the rapid advancements in artificial intelligence, the development of knowledgeable and personalized agents has become increasingly prevalent. However, the inherent variability in state variables and action spaces among personalized agents poses significant aggregation challenges for traditional federated learning algorithms. To tackle these challenges, we introduce the Federated Split Decision Transformer (FSDT), an innovative framework designed explicitly for AI agent decision tasks. The FSDT framework excels at navigating the intricacies of personalized agents by harnessing distributed data for training while preserving data privacy. It employs a two-stage training process, with local embedding and prediction models on client agents and a global transformer decoder model on the server. Our comprehensive evaluation using the benchmark D4RL dataset highlights the superior performance of our algorithm in federated split learning for personalized agents, coupled with significant reductions in communication and computational overhead compared to traditional centralized training approaches. The FSDT framework demonstrates strong potential for enabling efficient and privacy-preserving collaborative learning in applications such as autonomous driving decision systems. Our findings underscore the efficacy of the FSDT framework in effectively leveraging distributed offline reinforcement learning data to enable powerful multi-type agent decision systems. Bokui Chen, Xiaoyang Qu, Zhenhou Hong, Jing Xiao 0006, Jianzong Wang |
IJCNN | 2 |
| 2024 | Large Language Models Powered Context-aware Motion Prediction in Autonomous DrivingabstractMotion prediction is among the most fundamental tasks in autonomous driving. Traditional methods of motion forecasting primarily encode vector information of maps and historical trajectory data of traffic participants, lacking a comprehensive understanding of overall traffic semantics, which in turn affects the performance of prediction tasks. In this paper, we utilized Large Language Models (LLMs) to enhance the global traffic context understanding for motion prediction tasks. We first conducted systematic prompt engineering, visualizing complex traffic environments and historical trajectory information of traffic participants into image prompts— Transportation Context Map (TC-Map), accompanied by corresponding text prompts. Through this approach, we obtained rich traffic context information from the LLM. By integrating this information into the motion prediction model, we demonstrate that such context can enhance the accuracy of motion predictions. Furthermore, considering the cost associated with LLMs, we propose a cost-effective deployment strategy: enhancing the accuracy of motion prediction tasks at scale with 0.7% LLM-augmented datasets. Our research offers valuable insights into enhancing the understanding of traffic scenes of LLMs and the motion prediction performance of autonomous driving. The source code is available at https://github.com/AIR-DISCOVER/LLM-Augmented-MTR and https://aistudio.baidu.com/projectdetail/7809548. Xiaoji Zheng, Lixiu Wu, Zhijie Yan, Yuanrong Tang, Hao Zhao 0002, Bokui Chen, Jiangtao Gong |
IROS | 7 |
| 2024 | An extended self-representation model of complex networks for link prediction
Yuxuan Xiu, Xinglu Liu, Kexin Cao, Bokui Chen, Wai Kin Chan |
Inf. Sci. | 4 |
| 2023 | Transportation Internet: A Sustainable Solution for Intelligent Transportation SystemsabstractNew challenges such as automation, connection, electrification, and sharing (ACES) have brought disruptive changes to vehicles, transportation, and mobility services, which urgently requires an ideal solution for sustainable transportation. This paper introduces the Internet as a paradigm and, for the first time, proposes the Transportation Internet (TI), inspired by the similarity between the Internet and transportation. Referring to the construction ideas of the Internet, this paper establishes the framework of TI, proposes the transportation router based on the transportation switching and routing models, and preliminarily forms a large-scale automatic transportation solution. Following the latest technologies of the Internet, this paper further presents the software-defined transportation (SDT) by separating the control plane and transport plane of the transportation router, which can enhance transportation routing and provide Internet-like capabilities such as centralized intelligent control, terminals plug-and-play, and open application ecology. The evaluation of the prototype system shows promising results. The software-defined signals (SDS) can save 36% energy compared to signal machines, and the software-defined vehicles (SDV) automatic driving can save 24% energy compared to manual driving. Overall, TI brings innovations to sustainable transportation, and provides a framework for a new generation of Intelligent Transportation Systems (ITS). Hui Li 0107, Yongquan Chen, Keqiang Li 0002, Chong Wang 0017, Bokui Chen |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2021 | An optimal global algorithm for route guidance in advanced traveler information systems
Bokui Chen, Zhong-Jun Ding, Jun Zhou 0014, Yongquan Chen |
Inf. Sci. | 1 |
| 2021 | Towards Effective Classification of aMCI Based on Resting-State Multiscale Brain Features and Machine Learning ApproachesabstractSmart healthcare has undergone new opportunities and challenges with the arrival of the Industry 4.0 era. The intelligent imaging diagnosis system is a staple part of smart healthcare, helping doctors make clinical decisions. Nevertheless, intelligent diagnosis analysis is still confronted with the issue that it is challenging to extract effective features from the limited and high‐dimensional data, particularly in resting‐state data of amnesic mild cognitive impairment (aMCI). Furthermore, the intelligent imaging diagnosis system for aMCI is conductive to make timely predicting groups that may convert to Alzheimer’s disease (AD). To improve the system’s detection performance and reduce its data redundancy, we first develop an adaptive structure feature generation strategy (ASFGS) based on the Laplacian matrix and sparse autoencoder to obtain the structural features of brain functional network (BFN). Concurrently, we present a multiscale local feature detection strategy (MLFDS) to overcome the low utilization of local features of BFN. And finally, multiscale features, including structural features and multiscale local features, are fused by concatenation method to further improve the detection performance of aMCI system. Support vector machine based on radial basis function (RBF‐SVM) for small data learning is adopted to evaluate the effectiveness of the proposed features. Besides, we employ leave‐one‐out cross‐validation strategy to avoid the overfitting problem of classifier training process. The experiment results elucidate that the accuracy (ACC) and the area under the curve (AUC) in this work provide 86.57% and 86.36%, respectively, which outperforms the traditional methods and offers new insights for accuracy requirements of the aMCI system. Chunting Cai, Jiqiang Yan, Wuyang Zheng, Chenhui Yang, Zhemin Zhang, Bokui Chen, Dan Hong |
Wirel. Commun. Mob. Comput. | 7 |
| 2020 | A future intelligent traffic system with mixed autonomous vehicles and human-driven vehicles
Bokui Chen, Duo Sun, Jun Zhou 0014, Weng-Fai Wong, Zhong-Jun Ding |
Inf. Sci. | 1 |