VLDB 2026 Research / reviewers in the wild / expert
Xingyuan Dai
dblp:203/8062
· DBLP profile ↗
16ranked-venue papers
1as first author
13since 2021 · last 2026
0000-0001-7517-5049ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Evaluating the Perceptual Robustness of Vision-Language Models for Autonomous Driving in Corner Cases
Peizhe Gong, Enming Zhang, Ruixi Qiao, Xingyuan Dai, Xiaoyan Gong, Qinghai Miao |
IV | 4 |
| 2026 | Predictive reinforcement learning based on heterogeneous graph model for trajectory planning
Xingyuan Dai, Hub Ali, Fenghua Zhu |
Expert Syst. Appl. | 2 |
| 2025 | Leveraging Heterogeneous Experts with Advantageous Pattern Memory Learning for Traffic PredictionabstractAccurate traffic prediction is essential for mitigating congestion and enabling convenient trip arrangements. However, a single modeling approach often struggles to excel across diverse traffic patterns due to the inherent complexities and external influences in traffic scenarios. To address these issues, we propose a method named Memory-enhanced Heterogeneous Mixture of Experts (MH-MoE), which leverages memory-enhanced gating to integrate multiple pretrained models. The proposed method first obtains spatio-temporal embeddings from historical traffic sequences, followed by a traffic pattern extractor to capture representative patterns. Furthermore, a memory gating module memorizes each expert's advantageous patterns and learns to allocate traffic patterns to suitable experts. Finally, by combining predictions from these experts, MH-MoE effectively leverages the strengths of heterogeneous modeling to excel across traffic patterns. Experiments on multiple traffic datasets demonstrate that MH-MoE outperforms existing methods by leveraging diverse expert strengths, improving predictive accuracy, and offering scalability and efficiency for complex traffic prediction tasks. Yueyang Yao, Xingyuan Dai |
ICDE | 2 |
| 2025 | MiniDrive: More Efficient Vision-Language Models with Multi-level 2D Features as Text Tokens for Autonomous Driving
Enming Zhang, Xingyuan Dai, Min Huang 0009, Qinghai Miao |
PRCV (11) | 2 |
| 2025 | TransRAG for parallel transportation: toward reliable and trustworthy transportation systems via retrieval-augmented generationabstract平行交通是一种实现智能交通管理与控制的综合性范式,致力于解决人类行为和社会因素的复杂性问题。近年来,基础模型(foundational models, FMs)的崛起为平行交通的实现提供了新的可能。但这种模型固有的知识陈旧、“幻觉”现象以及“黑盒”特性削弱了其决策的可靠性和可信度。为解决这一问题,提出一种基于检索增强生成与思维链提示(chain-of-thought prompting)的平行交通框架TransRAG。该框架由紧密协作的存储层、管理层和执行层组成,旨在为用户提供个性且多样化的交通服务。其中,存储层引入的外部知识增强了管理层中基础模型的性能,以实现复杂的计算实验。执行层中人工交通系统与实际交通系统的虚实交互使得管理层的决策得到持续优化,从而实现动态知识更新和灵活的策略调整,以适应不断变化的交通环境。此外,TransRAG通过区块链、智能合约和缓存技术的集成,能够有效应对单点故障、隐私泄露以及数据访问延迟等问题,从而加速推进向“6S”交通5.0的全面迈进。 Jing Yang 0044, Xingyuan Dai, Levente Kovács, Fei-Yue Wang 0001 |
Frontiers Inf. Technol. Electron. Eng. | 2 |
| 2024 | SC- Tune: Unleashing Self-Consistent Referential Comprehension in Large Vision Language ModelsabstractRecent trends in Large Vision Language Models (LVLMs) research have been increasingly focusing on ad-vancing beyond general image understanding towards more nuanced, object-level referential comprehension. In this paper, we present and delve into the self-consistency ca-pability of LVLMs, a crucial aspect that reflects the mod-els' ability to both generate informative captions for spe-cific objects and subsequently utilize these captions to ac-curately re-identify the objects in a closed-loop process. This capability significantly mirrors the precision and reli-ability of fine- grained visual-language understanding. Our findings reveal that the self-consistency level of existing LVLMs falls short of expectations, posing limitations on their practical applicability and potential. To address this gap, we introduce a novel fine-tuning paradigm named Self-Consistency Tuning (SC-Tune). It features the syn-ergistic learning of a cyclic describer-locator system. This paradigm is not only data-efficient but also exhibits gener-alizability across multiple LVLMs. Through extensive ex-periments, we demonstrate that SC- Tune significantly ele-vates performance across a spectrum of object-level vision-language benchmarks and maintains competitive or im-proved performance on image-level vision-language bench-marks. Both our model and code will be publicly available at https://github.com/ivattyue/SC-Tune. Tongtian Yue, Jie Cheng 0009, Longteng Guo, Xingyuan Dai, Zijia Zhao, Xingjian He, Gang Xiong 0001, Jing Liu 0001 |
CVPR | 4 |
| 2024 | RIME: Robust Preference-based Reinforcement Learning with Noisy PreferencesabstractPreference-based Reinforcement Learning (PbRL) circumvents the need for reward engineering by harnessing human preferences as the reward signal. However, current PbRL methods excessively depend on high-quality feedback from domain experts, which results in a lack of robustness. In this paper, we present RIME, a robust PbRL algorithm for effective reward learning from noisy preferences. Our method utilizes a sample selection-based discriminator to dynamically filter out noise and ensure robust training. To counteract the cumulative error stemming from incorrect selection, we suggest a warm start for the reward model, which additionally bridges the performance gap during the transition from pre-training to online training in PbRL. Our experiments on robotic manipulation and locomotion tasks demonstrate that RIME significantly enhances the robustness of the state-of-the-art PbRL method. Code is available at https://github.com/CJReinforce/RIME_ICML2024. Jie Cheng 0009, Gang Xiong 0001, Xingyuan Dai, Qinghai Miao, Fei-Yue Wang 0001 |
ICML | 3 |
| 2024 | Open-Set Entity Alignment using Large Language Models with Retrieval AugmentationabstractRecent years have witnessed remarkable advance-ments in entity alignment, which endeavors to identify entities that represent the same real-world objects across different knowledge graphs (KGs). Nonetheless, prevailing approaches predominantly operate within closed-domain scenarios, rendering them inadequate for handling unmatchable entities. To address this challenge, we propose a retrieval augmented large language model framework (RALLM) to leverage the reasoning capacities of large language models (LLMs) to achieve open-set entity alignment, which not only enables the identification of equivalent entities for matchable entities but also addresses the identification of unmatchable ones. Specifically, we propose a novel retrieval augmentation method that leverages both textual and structural information of entities to retrieve potential equivalent candidates. Subsequently, we employ an iterative process to prompt the LLM to discern the equivalence between the retrieved candidate entity and the entity requiring alignment. To mitigate issues related to many-to-one alignment prediction and enhance alignment efficacy, we devise a memory mechanism to store highly confident aligned entity pairs and provide reminders to the LLM when a candidate entity has been matched. Our experimental findings underscore the superior performance of RALLM, highlighting the potential of LLMs in facilitating open-set entity alignment tasks. Linyao Yang, Hongyang Chen 0001, Xiao Wang 0002, Yonglin Tian, Xingyuan Dai, Fei-Yue Wang 0001 |
SMC | 6 |
| 2024 | ArtCap: A Dataset for Image Captioning of Fine Art PaintingsabstractThe image captioning of fine art paintings aims at generating content descriptions for the paintings. Due to the complexity of modeling both image and language, this task usually needs sufficient training data. However, different from photographic image captioning, there are few satisfactory datasets for painting captioning. In this article, we introduce a painting captioning dataset (named the ArtCap dataset), which contains 3606 paintings and five descriptions for each painting. We present the carefully designed construction pipeline of our dataset and further evaluate our dataset from two aspects of annotation quality and application effectiveness, respectively. For the annotation quality, we compare the global characteristics, annotation content, and annotation consistency of our dataset with other painting descriptions datasets. For application effectiveness, we employ our dataset and other painting descriptions datasets to train image captioning models and analyze the captioning performances. The results demonstrate the promising annotation quality and application effectiveness of our dataset. Chao Guo 0006, Xingyuan Dai, Fei-Yue Wang 0001 |
IEEE Trans. Comput. Soc. Syst. | 3 |
| 2023 | ArtVerse: A Paradigm for Parallel Human-Machine Collaborative Painting Creation in MetaversesabstractCurrently, the development of the foundation model, metaverse, nonfungible token (NFT), and other emerging technologies has brought profound effects on the whole art field, including art creation, dissemination, transaction, etc. However, there is no research focusing on the framework, methodologies, and applications of the human–machine collaborative creation in the metaverse era. Based on parallel theory, this article proposes a novel human–machine collaborative creation paradigm called ArtVerse, in which machines take on the roles of humans to perform creation exploration and evolution and build decentralized art organizations. Besides, the operational processes involving several key technologies are designed to achieve the proposed ArtVerse. Then, a prototype system of ArtVerse, our long-term efforts toward the human–machine collaborative painting, is presented. Finally, a new ecology of artistic creation in the metaverse era is demonstrated through the applications of the ArtVerse. Chao Guo 0006, Yong Dou, Tianxiang Bai, Xingyuan Dai, Chunfa Wang |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2023 | Decentralized Autonomous Operations and Organizations in TransVerse: Federated Intelligence for Smart MobilityabstractHuman and social factors are essential to transportation systems, yet top-down management fails to consider them sufficiently. Consequently, management strategies are not tailored to human needs and are inadequate in providing transportation intelligence. This article investigates a management architecture based on decentralized/distributed autonomous operations/organizations (DAOs) that considers both the technical and societal aspects in our transportation metaverse, TransVerse. This design maps people’s transportation needs in physical space to their digital counterparts in cyberspace, utilizing blockchain technology to guarantee the secure exchange of information and ultimately bring about the Internet of Minds (IoM). With the federated intelligence that emerged in IoM, we can devise reliable and prompt traffic decisions by incorporating consensus, community voting, and smart contracts into the organizational, coordination, and execution structure. Details on operational procedures and key technologies are also covered. To demonstrate the efficacy of DAOs-based management, a case study of world model-driven cooperative signal control is provided, indicating its promising application in future transportation management. Chen Zhao 0016, Xingyuan Dai, Jinglong Niu, Yilun Lin 0002 |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2022 | Data-efficient image captioning of fine art paintings via virtual-real semantic alignment training
Chao Guo 0006, Xingyuan Dai, Fei-Yue Wang 0001 |
Neurocomputing | 3 |
| 2022 | Image-based traffic signal control via world modelsabstractTraffic signal control is shifting from passive control to proactive control, which enables the controller to direct current traffic flow to reach its expected destinations. To this end, an effective prediction model is needed for signal controllers. What to predict, how to predict, and how to leverage the prediction for control policy optimization are critical problems for proactive traffic signal control. In this paper, we use an image that contains vehicle positions to describe intersection traffic states. Then, inspired by a model-based reinforcement learning method, DreamerV2, we introduce a novel learning-based traffic world model. The traffic world model that describes traffic dynamics in image form is used as an abstract alternative to the traffic environment to generate multi-step planning data for control policy optimization. In the execution phase, the optimized traffic controller directly outputs actions in real time based on abstract representations of traffic states, and the world model can also predict the impact of different control behaviors on future traffic conditions. Experimental results indicate that the traffic world model enables the optimized real-time control policy to outperform common baselines, and the model achieves accurate image-based prediction, showing promising applications in futuristic traffic signal control. Xingyuan Dai, Chen Zhao 0016, Xiao Wang 0002, Yilun Lin 0002, Fei-Yue Wang 0001 |
Frontiers Inf. Technol. Electron. Eng. | 1 |
| 2020 | Investigating the dynamic memory effect of human drivers via ON-LSTM
Shengzhe Dai, Zhiheng Li 0001, Li Li 0013, Dongpu Cao, Xingyuan Dai, Yilun Lin 0002 |
Sci. China Inf. Sci. | 5 |
| 2019 | Pattern Sensitive Prediction of Traffic Flow Based on Generative Adversarial FrameworkabstractTraffic flow prediction is one of the most popular topics in the field of the intelligent transportation system due to its importance. Powered by advanced machine learning techniques, especially the deep learning method, prediction accuracy noticeably increases in recent years. However, most existing methods applied a data-driven paradigm and tend to ignore the outliers, which result in poor performance while handling burst phenomena in the traffic system. To overcome this problem, the prediction model needs to recognize different patterns and handle them in different ways. In this paper, we propose a new prediction model (called pattern sensitive network) that can handle different traffic patterns automatically. By using adversarial training, our model can make more accurate predictions in unusual states without compromising its performance in usual states. Experiments demonstrate that our method can work well in both usual traffic states and unusual traffic states. Yilun Lin 0002, Xingyuan Dai, Li Li 0013, Fei-Yue Wang 0001 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2017 | Master general parking skill via deep learningabstractParking is one basic function of autonomous vehicles. However, parking still remains difficult to be implemented, since it requires to generate a relatively long-term series of actions to reach a certain objective under complicated constraints. One recently proposed method used deep neural networks(DNN) to learn the relationship between the actual parking trajectories and the corresponding steering actions, so as to find the best parking trajectory via direct recalling. However, this method can only handle a special vehicle whose dynamic parameters are well known. In this paper, we use transfer learning technique to further extend this direct trajectory planning method and master general parking skills. We aim to mimic how human drivers make parking by using a specially designed deep neural network. The first few layers of this DNN contain the general parking trajectory planning knowledge for all kinds of vehicles; while the last few layers of this DNN can be quickly tuned to adapt various kinds of vehicles. Numerical tests show that, combining transfer learning and direct trajectory planning solution, our new approach enables automated vehicles to convey the knowledge of trajectory planning from one vehicle to another with a few try-and-tests. Yilun Lin 0002, Li Li 0013, Xingyuan Dai, Nanning Zheng 0001, Fei-Yue Wang 0001 |
Intelligent Vehicles Symposium | 3 |