Yirong Chen

dblp:220/0124 · DBLP profile ↗
← Back
17ranked-venue papers
1as first author
15since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Systems, architecture and hardware · 1Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 ReBrain: Brain MRI Reconstruction from Sparse CT Slice via Retrieval-Augmented Diffusion
abstract
Magnetic Resonance Imaging (MRI) plays a crucial role in brain disease diagnosis, but it is not always feasible for certain patients due to physical or clinical constraints. Recent studies attempt to synthesize MRI from Computed Tomography (CT) scans; however, low-dose protocols often result in highly sparse CT volumes with poor throughplane resolution, making accurate reconstruction of the full brain MRI volume particularly challenging. To address this, we propose ReBrain, a retrieval-augmented diffusion framework for brain MRI reconstruction. Given any 3D CT scan with limited slices, we first employ a Brownian Bridge Diffusion Model (BBDM) to synthesize MRI slices along the 2D dimension. Simultaneously, we retrieve structurally and pathologically similar CT slices from a comprehensive prior database via a fine-tuned retrieval model. These retrieved slices are used as references, incorporated through a ControlNet branch to guide the generation of intermediate MRI slices and ensure structural continuity. We further account for rare retrieval failures when the database lacks suitable references and apply spherical linear interpolation to provide supplementary guidance. Extensive experiments on SynthRAD2023 and BraTS demonstrate that ReBrain achieves state-of-the-art performance in cross-modal reconstruction under sparse conditions.
Weihua Cheng, Yujin Kang, Yirong Chen, Ding Wang 0001, Guosun Zeng
WACV5
2026 Mosaic: Data-free knowledge distillation via mixture-of-experts for heterogeneous distributed environments
Yanting Gao, Siyuan Meng, Aoqi Wu, Yirong Chen, Shiping Wen 0001
Knowl. Based Syst.7
2026 OmniPathoVQA: Benchmarking pathology vision-language models with Encyclopedia-scale knowledge
Kaitao Chen, Linda Wei, Shaohao Rui, Xialing Zhang, Zunguo Du, Mianxin Liu, Mu Zhou, Yirong Chen
Medical Image Anal.11
2025 PsyDT: Using LLMs to Construct the Digital Twin of Psychological Counselor with Personalized Counseling Style for Psychological Counseling
abstract
Currently, large language models (LLMs) have made significant progress in the field of psychological counseling.However, existing mental health LLMs overlook a critical issue where they do not consider the fact that different psychological counselors exhibit different personal styles, including linguistic styles and therapeutic types, etc.As a result, these LLMs fail to satisfy the individual needs of clients who seek different counseling styles.To help bridge this gap, we propose PsyDT, a novel framework using LLMs to construct the Digital Twin of Psychological counselor with personalized counseling style.Compared to the timeconsuming and costly approach of collecting a large number of real-world counseling cases to create a specific counselor's digital twin, our framework offers a faster and more costeffective solution.To construct PsyDT, we utilize dynamic one-shot learning by using GPT-4 to capture counselor's unique counseling style, mainly focusing on linguistic style and therapy technique.Subsequently, using existing singleturn long-text dialogues with client personality, GPT-4 is guided to synthesize multi-turn dialogues of specific counselor.Finally, we finetune the LLMs on the synthesized dataset, Psy-DTCorpus, to achieve the digital twin of psychological counselor with personalized counseling style.Experimental results indicate that our proposed PsyDT framework can synthesize multi-turn dialogues that closely resemble realworld counseling cases and demonstrate better performance compared to other baselines, thereby show that our framework can effectively construct the digital twin of psychological counselor with a specific counseling style. 1
Haojie Xie, Yirong Chen, Xiaofen Xing, Jingkai Lin, Xiangmin Xu 0001
ACL (1)2
2025 Ontology-based Adaptive Knowledge System (OAKS): Adaptive and Consistent Knowledge Acquisition through LLMs for Diverse User Backgrounds
abstract
Although most of the research on large language models (LLMs) focuses on their development and validation against datasets, significant gaps remain in their application to real-world knowledge-intensive tasks. This research addresses key challenges in using LLMs for extracting and synthesizing knowledge from unstructured sources, focusing on applications where the validity and consistency of the results are critical. We propose Ontology-based Adaptive Knowledge System (OAKS) as a holistic approach to manage the complexities of acquiring unstructured knowledge, varying user expertise, and dynamic query formulation. This research provides practical value for enabling the domain user community to leverage their technical documentation and expertise and accelerate ongoing working projects through improved literature review and cross-disciplinary insight discovery. Validated through empirical studies, our findings offer insight into best practices for the deployment of OAKS, bridging the gaps between AI capabilities and real-world needs in knowledge acquisition and research.
Muran Yu, Jie Wang 0006, Yirong Chen, Michael D. Lepech, Ying Liu 0039, Kincho H. Law
COMPSAC3
2025 Aligning Vision to Language: Annotation-Free Multimodal Knowledge Graph Construction for Enhanced LLMs Reasoning
abstract
Multimodal reasoning in Large Language Models (LLMs) struggles with incomplete knowledge and hallucination artifacts, challenges that textual Knowledge Graphs (KGs) only partially mitigate due to their modality isolation. While Multimodal Knowledge Graphs (MMKGs) promise enhanced cross-modal understanding, their practical construction is impeded by semantic narrowness of manual text annotations and inherent noise in visual-semantic entity linkages. In this paper, we propose Vision-align-to-Language integrated Knowledge Graph (VaLiK), a novel approach for constructing MMKGs that enhances LLMs reasoning through cross-modal information supplementation. Specifically, we cascade pre-trained Vision-Language Models (VLMs) to align image features with text, transforming them into descriptions that encapsulate image-specific information. Furthermore, we developed a cross-modal similarity verification mechanism to quantify semantic consistency, effectively filtering out noise introduced during feature alignment. Even without manually annotated image captions, the refined descriptions alone suffice to construct the MMKG. Compared to conventional MMKGs construction paradigms, our approach achieves substantial storage efficiency gains while maintaining direct entity-to-image linkage capability. Experimental results on multimodal reasoning tasks demonstrate that LLMs augmented with VaLiK outperform previous state-of-the-art models. Our code is published at https://github.com/Wings-Of-Disaster/VaLiK.
Siyuan Meng, Yanting Gao, Song Mao, Pinlong Cai, Guohang Yan, Yirong Chen, Zilin Bian, Ding Wang 0001, Botian Shi
ICCV7
2025 VLMLight: Safety-Critical Traffic Signal Control via Vision-Language Meta-Control and Dual-Branch Reasoning Architecture
abstract
Traffic signal control (TSC) is a core challenge in urban mobility, where real-time decisions must balance efficiency and safety. Existing methods—ranging from rule-based heuristics to reinforcement learning (RL)—often struggle to generalize to complex, dynamic, and safety-critical scenarios. We introduce \textbf{VLMLight}, a novel TSC framework that integrates vision-language meta-control with dual-branch reasoning. At the core of VLMLight is the first image-based traffic simulator that enables multi-view visual perception at intersections, allowing policies to reason over rich cues such as vehicle type, motion, and spatial density. A large language model (LLM) serves as a safety-prioritized meta-controller, selecting between a fast RL policy for routine traffic and a structured reasoning branch for critical cases. In the latter, multiple LLM agents collaborate to assess traffic phases, prioritize emergency vehicles, and verify rule compliance. Experiments show that VLMLight reduces waiting times for emergency vehicles by up to 65% over RL-only systems, while preserving real-time performance in standard conditions with less than 1% degradation. VLMLight offers a scalable, interpretable, and safety-aware solution for next-generation traffic signal control.
Maonan Wang, Yirong Chen, Aoyu Pang, Chung Shue Chen, Yuheng Kan, Man-On Pun
NeurIPS2
2025 MPF: A Multi-Noise Perception Framework to Enhance Online Map Matching Algorithms
abstract
Map matching is crucial to facilitating location-based services, and recent advancements in map matching have demonstrated excellent performance with high-quality data. However, the use of low-precision devices often introduces high measurement noise, and the slow update rate of maps may result in errors in digital maps. Consequently, multiple types of noise significantly impact the performance of map matching algorithms. To tackle this issue, this paper presents a novel multi-noise perception framework, named MPF, aiming to enhance the performance and robustness of existing map matching algorithms. The main challenge lies in detecting anomalies during map matching, identifying the root causes, and devising appropriate solutions. Firstly, we propose a matching quality assessment (MQA) method that assesses abnormal variance in matching probability. Secondly, we introduce a multiple noise discrimination (MND) mechanism to effectively differentiate between measurement noise and map errors. Thirdly, we present a missing segment generation (MSG) scheme that dynamically fills in map gaps to prevent significant detours. To validate the effectiveness of MPF, we conduct experiments using real-world taxi trajectories from four cities, covering a total distance of 79,670.6 km. MPF is compare with seven online map matching algorithms and is used to optimize their performance. The experiments show that MPF outperforms the top baselines by 15.6%-26.9% and enhances their performance by 18.7%-38.2%.
Hanwen Hu, Shiyou Qian, Jian Cao 0001, Yirong Chen, Jie Wang 0006
IEEE Trans. Intell. Transp. Syst.6
2025 PREFER: A Pre-trained Model Recommendation Framework for Edge Computing Enabled Traffic Flow Prediction
abstract
The recent years have witnessed a surge in the development of traffic flow prediction methods, often deployed on cloud platforms to offer predictive services for entire transportation networks. However, the processes of training and executing a model for the entire traffic network are both time-consuming and computationally expensive. As a result, the utilization of edge servers for local sub-network prediction services has gained prominence. Nevertheless, training prediction models for numerous sub-networks within the extensive traffic network remains a time-intensive and computing resource-consuming task. To tackle this challenge, this article introduces the Pre-trained model REcommendation Framework for Edge computing enabled tRaffic flow prediction (PREFER). PREFER trains a set of traffic flow prediction models on selected sub-networks, then recommends optimal pre-trained models for edge servers. The recommendation is specifically based on performance prediction, integrating neural collaborative filtering and traffic flow characteristics. Experiments conducted on real datasets reveal that the pre-trained models recommended by PREFER perform close to the actual optimal ones and significantly outperform existing recommendation algorithms.
Qiqi Cai, Jian Cao 0001, Yirong Chen, Shiyou Qian, Liangxiao Yuan, Jie Wang 0006
ACM Trans. Knowl. Discov. Data3
2024 Domain Adaption and Unified Knowledge Base Motivate Better Retrieval Models in Dialog Systems With RAG
abstract
Retrieval augmented generation (RAG) has emerged as a paradigm to address problems like hallucination in dialog systems based on large language model (LLM). Retrieval model is a key component in RAG framework for recalling relevant information. This paper describes our solution for FutureDial-RAG Challenge Track 1. We identify two primary challenges in this track: domain specificity and heterogeneity of knowledge base. To address the two challenges, we first adopt continual pre-training of a pre-trained retrieval model on both labeled and unlabeled data for domain adaption. Subsequently, we modify and expand the knowledge base, ensuring that each piece of knowledge is uniformly structured in a question-answer (QA) format. Finally, we construct negative samples based on the labeled data and the unified knowledge base, and fine-tune the retrieval model using contrastive learning. Our solution achieves a score of 2.023 on the dev set, which significantly outperforms the baseline.
Huadong Lin, Yirong Chen, Wenyu Tao, Mingyu Chen 0016, Xiangmin Xu 0001, Xiaofen Xing
SLT2
2024 Multi-agent reinforcement learning for vehicular task offloading with multi-step trajectory prediction
Yanmin Zhu 0006, Chunyang Wang 0001, Jian Cao 0001, Yirong Chen, Jie Wang 0006
CCF Trans. Pervasive Comput. Interact.5
2024 Traffic Signal Cycle Control With Centralized Critic and Decentralized Actors Under Varying Intervention Frequencies
abstract
Traffic congestion in urban areas is a significant problem, leading to prolonged travel times, reduced efficiency, and increased environmental concerns. Effective traffic signal control (TSC) is a key strategy for reducing congestion. Unlike most TSC systems that rely on high-frequency control, this study introduces an innovative joint phase traffic signal cycle control method that operates effectively with varying control intervals. Our method features an adjust all phases action design, enabling simultaneous phase changes within the signal cycle, which fosters both immediate stability and sustained TSC effectiveness, especially at lower frequencies. The approach also integrates decentralized actors to handle the complexity of the action space, with a centralized critic to ensure coordinated phase adjusting. Extensive testing on both synthetic and real-world data across different intersection types and signal setups shows that our method significantly outperforms other popular techniques, particularly at high control intervals. Case studies of policies derived from traffic data further illustrate the robustness and reliability of our proposed method.
Maonan Wang, Yirong Chen, Yuheng Kan, Chengcheng Xu 0005, Michael D. Lepech, Man-On Pun
IEEE Trans. Intell. Transp. Syst.2
2023 Speaker-Aware Hierarchical Transformer For Personality Recognition In Multiparty Dialogues
abstract
Personality recognition is one of the core technologies in human-machine interaction, which has received increasing attention. Previous works mainly focus on essays or monologues, while personality traits reveal more in the interactions with others. Due to the lack of appropriate datasets, a few approaches aim to recognize personality traits in conversations, and most of them ignore interdependence between speakers and connection between conversations. In this paper, we create a multiparty dialogue-based personality dataset derived from CPED containing 1,195 data samples. We center on one speaker and extract related dialogues to compose each data sample annotated with speaker’s Big-Five traits, which is conducive to fully describe a center speaker using diverse cues of personality in different dialogues. Along the same lines, we propose a Speaker-aware Hierarchical Transformer named SH-Transformer to address above concerns, in which Personalized Embeddings (PE) adopt special tokens to distinguish center speakers in complete conversations and hierarchical Transformer capture diverse cues in utterances and conversations. Experimental results show that our method outperforms the non-interactive baseline by 1.38%, which confirms the necessity of considering both interactive information and diverse cues among dialogues. Our code will be released at github.com/Chloehxxx/SH-Transformer.
Wenjing Han, Yirong Chen, Xiaofen Xing, Guohua Zhou, Xiangmin Xu 0001
ICASSP2
2023 AMM: An Adaptive Online Map Matching Algorithm
abstract
Online map matching is essential for some location-based services, such as car navigation. However, due to GPS measurement errors and/or the lack of sufficient information, the performance of most existing online algorithms will degrade in the increasingly complex traffic environment. In this paper, we propose an adaptive online map matching algorithm called AMM. The basic idea is that AMM should be able to calibrate GPS observation data for various measurement errors and under complex urban conditions. First, we establish a collaborative evaluation model between GPS points and candidate points to effectively filter low-quality GPS measurement points, dynamically set the weights of different features and comprehensively select the best candidate points. Second, we propose a retrospective correction mechanism to correct the previous matching results when more information is available, which will help improve the accuracy of future GPS points. Furthermore, we define parameter self-tuning rules for AMM to enhance its portability by avoiding time-consuming parameter tuning steps. We conduct extensive experiments to evaluate the performance of AMM on real vehicle trajectory datasets. The experiment results show that AMM outperforms its counterparts by up to 32% in terms of accuracy and its performance in different traffic conditions is more stable.
Hanwen Hu, Shiyou Qian, Jingchao Ouyang, Jian Cao 0001, Jie Wang 0006, Yirong Chen
IEEE Trans. Intell. Transp. Syst.7
2022 @ME: A Fine-grained Route Recommendation System to Grab Impatient Passengers
abstract
Data analysis reveals that passengers can only endure a few minutes before taking a taxi. However, most existing route recommendation systems are not adequate to satisfy impatient customers due to two shortcomings: inaccurate demand forecast and the lack of an efficient supply-demand balance mechanism. In this paper, we propose a recommendation system called @ Me, which aims to dispatch vacant taxis to the vicinity of potential customers at the right time. To achieve minute-level demand forecasting, we ensemble a contextualized spatial-temporal network (CSTN) with an LSTM network to optimize prediction accuracy. In addition, we characterize the attractiveness of the road grid to vacant taxis as a force model, on which a taxi scheduling algorithm is proposed to dynamically balance supply and demand. Extensive experiments on real datasets clearly indicate that our method is superior to the selected baselines. Vacant taxis that follow the routes suggested by @ME can catch more impatient customers in a shorter cruising time. The 7 -day experimental results on the Manhattan dataset show that @Me can carry an additional 48,332 passengers and increase drivers' revenue by $570,317.
Hanwen Hu, Yirong Chen, Jingchao Ouyang, Shiyou Qian, Jian Cao 0001, Jie Wang 0006, Michael D. Lepech
IJCNN2
2020 Multi-Agent Context-Aware Dynamic-Scheduling for Large-scale Processing Networks
abstract
Optimally scheduling jobs in processing networks to meet multiple objectives for economic considerations and operational efficiencies has been a hot topic. However, most data-driven methods are rarely applied. One of the critical reasons is that the scheduling policy generated by these methods tends to bias toward the specific environment. In order to better deal with the discrepancy, this work in progress paper presents a context-aware dynamic scheduling method (CADS) that can adaptively select a specific policy based on the on-demand context. The CADS has two components: 1) The evaluation module that evaluates the performances of the policies that are learned from each operational context; 2) The decision-making module that maintains the knowledge of each policy’s performance under each context and constructs the weighted best-fit policy based on the identified context. The promising preliminary result using numerical simulation that demonstrates the effectiveness of CADS is presented. CADS outperforms traditional scheduling methods in various kinds of processing network environments.
Shuhui Qu, Yirong Chen, Jürgen Jasperneite, Michael D. Lepech, Jie Wang 0006
ETFA2
2018 The utility of LASSO-based models for real time forecasts of endemic infectious diseases: A cross country comparison
Yirong Chen, Collins Wenhan Chu, Mark Chen 0002, Alex R. Cook
J. Biomed. Informatics1