Wendi Li

dblp:224/6352 · DBLP profile ↗
← Back
12ranked-venue papers
6as first author
12since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 6 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Diverse Human Driving Vehicle Simulation in Background Traffic for Autonomous Driving Tests
abstract
Realistic background traffic is critical to the simulation platforms for autonomous driving (AD) testing. Given that most vehicles in reality are driven by human beings, introducing human driving (HD) vehicles to the background traffic is necessary to be able to discover more problems of the tested AD vehicle in the simulation stage. However, existing methods rely on ad-hoc rules or data-driven training to mimic partial human driver behaviors, which are not comprehensive and lack transparency. In this work, we design a smart human driving vehicle simulator HDSim which is empowered by cognitively inspired modeling and AI models. HDSim enables diverse, realistic, and scalable HD traffic simulation on AD testing platforms like CARLA in a non-intrusive manner. There are two novel components in HDSim. First, we introduce a driver model to guide the generation of diverse human driving styles by using different combinations of latent cognitive factors in a hierarchy. Second, we design a Perception-Mediated Behavior Influence (PMBI) mechanism to use LLM-assisted perceptual transformations to indirectly fuse driving actions with driving styles. Experiments show that HDSim traffic can help simulation platforms like CARLA to reveal 68% more failures of tested AD vehicles, and the explainability of reported accidents is also improved.
Wendi Li, Hao Wu 0067, Bing Mao 0001, Fengyuan Xu, Sheng Zhong 0002
AAAI1
2026 Uncertainty Quantification in LLM Agents: Foundations, Emerging Challenges, and Opportunities
abstract
Changdae Oh, Seongheon Park, To Eun Kim, Jiatong Li, Wendi Li, Samuel Yeh, Sean Du, Hamed Hassani, Paul Bogdan, Dawn Song, Sharon Li. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Changdae Oh, Seongheon Park, To Eun Kim, Wendi Li, Samuel Yeh 0001, Sean Du, Seyed Hamed Hassani, Paul Bogdan, Dawn Song, Yixuan Li 0001
ACL (1)5
2026 SlimFit-Gens: Toward Low Bandwidth One-on-One Video Calls on COTS Smartphones
abstract
Mobile video calls play an essential role in our daily lives. However, in bandwidth-limited scenarios (e.g., inadequate cellular coverage, congested satellite links, and metered connections), users often experience poor quality of experience (QoE) during video calls. While recent advances in deep learning have demonstrated significant improvements in video compression over traditional methods, existing approaches are ill-suited for bidirectional video streaming on smartphones. The primary challenge lies in simultaneously achieving high video quality, computational and bandwidth efficiency, and practical usability on constrained mobile devices. In this work, we present SlimFit-Gens, the first practical video calling system for smartphones capable of delivering real-time 480p video at as low as 30 kbps. SlimFit-Gens addresses the challenge with joint algorithm and system-level optimizations. The core technique is a fine-grained model personalization design tailored for mobile video calling, enabling high-fidelity video generation at low model complexity. SlimFit-Gens achieves effective personalized adaptation through a novel two-stage personalization mechanism working upon an optimized model architecture. It also incorporates a privacy preserving, resource-efficient system design, featuring TEE-based (e.g., Confidential VM/NVIDIA Confidential Computing) fine-tuning on the server side and heterogeneity-aware inference on the device side. We implement SlimFit-Gens on four commercial off-the-shelf (COTS) smartphones with different system-on-chip (SoC) configurations and conduct extensive evaluations. Compared to prior work, SlimFit-Gens simultaneously improves generation quality with a 0.09-0.12 reduction in LPIPS and system efficiency through a 1.6-1.8× increase in video frame rate.
Jingzhou Zhu, Lizhi Sun, Peiwen Dong, Wendi Li, Yixin Xu 0003, Hao Wu 0067, Fengyuan Xu, Sheng Zhong 0002
IEEE Trans. Mob. Comput.6
2025 Process Reward Model with Q-value Rankings
abstract
Process Reward Modeling (PRM) is critical for complex reasoning and decision-making tasks where the accuracy of intermediate steps significantly influences the overall outcome. Existing PRM approaches, primarily framed as classification problems, employ cross-entropy loss to independently evaluate each step's correctness. This method can lead to suboptimal reward distribution and does not adequately address the interdependencies among steps. To address these limitations, we introduce the Process Q-value Model (PQM), a novel framework that redefines PRM in the context of a Markov Decision Process. PQM optimizes Q-value rankings based on a novel comparative loss function, enhancing the model's ability to capture the intricate dynamics among sequential decisions. This approach provides a more granular and theoretically grounded methodology for process rewards. Our extensive empirical evaluations across various sampling policies, language model backbones, and multi-step reasoning benchmarks show that PQM outperforms classification-based PRMs. The effectiveness of the comparative loss function is highlighted in our comprehensive ablation studies, confirming PQM’s practical efficacy and theoretical advantage.
Wendi Li, Yixuan Li 0001
ICLR1
2025 Free Process Rewards without Process Labels
abstract
Different from its counterpart outcome reward models (ORMs), which evaluate the entire responses, a process reward model (PRM) scores a reasoning trajectory step by step, providing denser and more fine-grained rewards. However, training a PRM requires labels annotated at every intermediate step, presenting significant challenges for both manual and automatic data collection. This paper aims to address this challenge. Both theoretically and empirically, we show that an implicit PRM can be obtained at no additional cost, by simply training an ORM on the cheaper response-level labels. The only assumption is to parameterize the outcome reward as the log-likelihood ratios of the policy and reference models r$\phi$(y) = $\beta$ log $\pi$$\phi$(y) $\pi$ref(y) , which can be optimized regardless of the specific choice of loss objectives. In experiments, we instantiate our implicit PRMs with various objectives and evaluate their performance on MATH. We show that our implicit PRM outperforms a strong MCTS-based baseline á la Math-Shepherd (Wang et al., 2023) using less than 1/38 of the training data. Its performance can be further improved with majority voting. We further find that scaling up instructions and responses benefits our implicit PRM, and the latter brings a larger gain. Particularly, we find that our implicit PRM, when instantiated with the cross-entropy (CE) loss, is more data-efficient and can keep improving generation models even when trained with only one response per instruction, the setup that suffers from extreme data scarcity and imbalance. Further, instructions should be relevant to downstream tasks while the diversity of responses does not bring gains. Surprisingly, training on extra Math-Shepherd step labels brings no further improvements to our implicit PRM trained on only outcome data. We hope that our work will encourage a rethinking of PRM training approaches and contribute to making training PRMs more accessible.
Lifan Yuan, Wendi Li, Huayu Chen, Ganqu Cui, Ning Ding 0002, Bowen Zhou 0002, Zhiyuan Liu 0001, Hao Peng 0001
ICML2
2024 Position Debiasing Fine-Tuning for Causal Perception in Long-Term Dialogue
Shixuan Fan, Wei Wei 0002, Wendi Li, Xianling Mao, Wenfeng Xie, Dangyang Chen
IJCAI3
2024 Fire danger forecasting using machine learning-based models and meteorological observation: a case study in Northeastern China
Zhenyu Chen 0001, Wendi Li, Lanyu Gao, Changsheng Zhang 0001
Multim. Tools Appl.3
2023 TREA: Tree-Structure Reasoning Schema for Conversational Recommendation
abstract
Conversational recommender systems (CRS) aim to timely trace the dynamic interests of users through dialogues and generate relevant responses for item recommendations.Recently, various external knowledge bases (especially knowledge graphs) are incorporated into CRS to enhance the understanding of conversation contexts.However, recent reasoning-based models heavily rely on simplified structures such as linear structures or fixed-hierarchical structures for causality reasoning, hence they cannot fully figure out sophisticated relationships among utterances with external knowledge.To address this, we propose a novel Treestructure Reasoning schEmA named TREA.TREA constructs a multi-hierarchical scalable tree as the reasoning structure to clarify the causal relationships between mentioned entities, and fully utilizes historical conversations to generate more reasonable and suitable responses for recommended results.Extensive experiments on two public CRS datasets have demonstrated the effectiveness of our approach.Our
Wendi Li, Wei Wei 0002, Xiaoye Qu, Xianling Mao, Wenfeng Xie, Dangyang Chen
ACL (1)1
2023 Towards Hierarchical Policy Learning for Conversational Recommendation with Hypergraph-based Reinforcement Learning
abstract
Conversational recommendation systems (CRS) aim to timely and proactively acquire user dynamic preferred attributes through conversations for item recommendation. In each turn of CRS, there naturally have two decision-making processes with different roles that influence each other: 1) director, which is to select the follow-up option (i.e., ask or recommend) that is more effective for reducing the action space and acquiring user preferences; and 2) actor, which is to accordingly choose primitive actions (i.e., asked attribute or recommended item) to estimate the effectiveness of the director’s option. However, existing methods heavily rely on a unified decision-making module or heuristic rules, while neglecting to distinguish the roles of different decision procedures, as well as the mutual influences between them. To address this, we propose a novel Director-Actor Hierarchical Conversational Recommender (DAHCR), where the director selects the most effective option, followed by the actor accordingly choosing primitive actions that satisfy user preferences. Specifically, we develop a dynamic hypergraph to model user preferences and introduce an intrinsic motivation to train from weak supervision over the director. Finally, to alleviate the bad effect of model bias on the mutual influence between the director and actor, we model the director’s option by sampling from a categorical distribution. Extensive experiments demonstrate that DAHCR outperforms state-of-the-art methods.
Sen Zhao 0001, Wei Wei 0002, Yifan Liu 0004, Wendi Li, Xianling Mao, Shuai Zhu, Zujie Wen
IJCAI5
2023 Population-Based Hyperparameter Tuning With Multitask Collaboration
abstract
Population-based optimization methods are widely used for hyperparameter (HP) tuning for a given specific task. In this work, we propose the population-based hyperparameter tuning with multitask collaboration (PHTMC), which is a general multitask collaborative framework with parallel and sequential phases for population-based HP tuning methods. In the parallel HP tuning phase, a shared population for all tasks is kept and the intertask relatedness is considered to both yield a better generalization ability and avoid data bias to a single task. In the sequential HP tuning phase, a surrogate model is built for each new-added task so that the metainformation from the existing tasks can be extracted and used to help the initialization for the new task. Experimental results show significant improvements in generalization abilities yielded by neural networks trained using the PHTMC and better performances achieved by multitask metalearning. Moreover, a visualization of the solution distribution and the autoencoder's reconstruction of both the PHTMC and a single-task population-based HP tuning method is compared to analyze the property with the multitask collaboration.
Wendi Li, Ting Wang 0015, Wing W. Y. Ng
IEEE Trans. Neural Networks Learn. Syst.1
2022 DDG-DA: Data Distribution Generation for Predictable Concept Drift Adaptation
abstract
In many real-world scenarios, we often deal with streaming data that is sequentially collected over time. Due to the non-stationary nature of the environment, the streaming data distribution may change in unpredictable ways, which is known as the concept drift in the literature. To handle concept drift, previous methods first detect when/where the concept drift happens and then adapt models to fit the distribution of the latest data. However, there are still many cases that some underlying factors of environment evolution are predictable, making it possible to model the future concept drift trend of the streaming data, while such cases are not fully explored in previous work. In this paper, we propose a novel method DDG-DA, that can effectively forecast the evolution of data distribution and improve the performance of models. Specifically, we first train a predictor to estimate the future data distribution, then leverage it to generate training samples, and finally train models on the generated data. We conduct experiments on three real-world tasks (forecasting on stock price trend, electricity load and solar irradiance) and obtained significant improvement on multiple widely-used models.
Wendi Li, Weiqing Liu, Yingce Xia, Jiang Bian 0002
AAAI1
2021 HELP: An LSTM-based approach to hyperparameter exploration in neural network learning
Wendi Li, Wing W. Y. Ng, Ting Wang 0015, Marcello Pelillo, Sam Kwong
Neurocomputing1