Jiashuo Wang

dblp:204/7570 · DBLP profile ↗
← Back
11ranked-venue papers
2as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 2 first-author · 8 since 2021Systems, architecture and hardware · 4 · 4 since 2021
YearPublicationVenuePosition
2026 ACOPT: Adaptive continuity-aware address translation for performance optimization of MCM-GPU architectures
Jingweijia Tan, Zhanyuntian Li, Weiren Wang, Jiashuo Wang, Kaige Yan
Future Gener. Comput. Syst.4
2025 Towards Dynamic Theory of Mind: Evaluating LLM Adaptation to Temporal Evolution of Human States
abstract
As Large Language Models (LLMs) increasingly participate in human-AI interactions, evaluating their Theory of Mind (ToM) capabilities - particularly their ability to track dynamic mental states - becomes crucial. While existing benchmarks assess basic ToM abilities, they predominantly focus on static snapshots of mental states, overlooking the temporal evolution that characterizes real-world social interactions. We present DynToM, a novel benchmark specifically designed to evaluate LLMs’ ability to understand and track the temporal progression of mental states across interconnected scenarios. Through a systematic four-step framework, we generate 1,100 social contexts encompassing 5,500 scenarios and 78,100 questions, each validated for realism and quality. Our comprehensive evaluation of ten state-of-the-art LLMs reveals that their average performance underperforms humans by 44.7%, with performance degrading significantly when tracking and reasoning about the shift of mental states. This performance gap highlights fundamental limitations in current LLMs’ ability to model the dynamic nature of human mental states.
Jiashuo Wang, Qiancheng Xu, Changhe Song, Chunpu Xu, Wenjie Li 0002, Pengfei Liu 0003
ACL (1)2
2025 Brain-Inspired Spatial Continuous State Encoding for Efficient Spiking-Based Navigation
abstract
Spiking neural networks (SNNs) show great potential in mapless navigation tasks due to their low power consumption, but the continuous representation of spatial information poses a challenge to SNN training. Neuroscience findings reveal that spatial cognition cells encode spatial information through population spike patterns. Inspired by this, we propose a navigation method based on SNNs, leveraging spatial cognition cells, which include grid cells (GCs), head direction cells (HDCs), and boundary vector cells (BVCs). Our method integrates spike-based information to achieve precise navigation goal encoding and egocentric environment perception, significantly improving SNN navigation capabilities in complex environments. Simulation and real-world experiments demonstrate that our method achieves significant improvements in navigation success rate and energy efficiency, showcasing superior adaptability across environments. Our work provides a novel approach to developing efficient brain-inspired navigation systems.
Qingao Chai, Jiashuo Wang, Runhao Jiang, Huajin Tang
ICRA2
2025 LIMOPro: Reasoning Refinement for Efficient and Effective Test-time Scaling
abstract
Large language models (LLMs) have demonstrated remarkable reasoning capabilities through test-time scaling approaches, particularly when fine-tuned with chain-of-thought (CoT) data distilled from more powerful large reasoning models (LRMs). However, these reasoning chains often contain verbose elements that mirror human problem-solving, categorized as progressive reasoning (the essential solution development path) and functional elements (verification processes, alternative solution approaches, and error corrections). While progressive reasoning is crucial, the functional elements significantly increase computational demands during test-time inference. We introduce PIR (Perplexity-based Importance Refinement), a principled framework that quantitatively evaluates the importance of each reasoning step based on its impact on answer prediction confidence. PIR systematically identifies and selectively prunes only low-importance functional steps while preserving all progressive reasoning components, creating optimized training data that maintains the integrity of the core solution path while reducing verbosity. Models fine-tuned on PIR-optimized data exhibit superior test-time scaling properties, generating more concise reasoning chains while achieving improved accuracy (+0.9\% to +6.6\%) with significantly reduced token usage (-3\% to -41\%) across challenging reasoning benchmarks (AIME, AMC, and GPQA Diamond). Our approach demonstrates strong generalizability across different model sizes, data sources, and token budgets, offering a practical solution for deploying reasoning-capable LLMs in scenarios where efficient test-time scaling, response time, and computational efficiency are valuable constraints. Code and dataset are available at the [LIMOPro GitHub repository.](https://github.com/GAIR-NLP/LIMOPro)
Jiashuo Wang, Ruifeng Yuan, Chunpu Xu, Kaishuai Xu, Wenjie Li 0002, Pengfei Liu 0003
NeurIPS2
2025 Evaluating GPU's Instruction-Level Error Characteristics Under Low Supply Voltages
abstract
Supply voltage underscaling has been an effective approach to improve the energy-efficiency of modern high-performance processors, such as GPUs. However, energy efficiency and reliability are two sides of a trade-off. Undervolting will inevitably undermine reliability, since it reduces chip manufacturers’ voltage guardbands that is designed to ensure correct operations under worst-case scenarios. To achieve optimal energy efficiency while maintaining enough reliability, it is necessary to deeply understand the error characteristics caused by undervolting. Unlike previous works which focus mostly on program level, we perform the first comprehensive instruction-level voltage margin and error characteristics evaluation for GPU architectures. We systematically measure the error probability and patterns of GPU instructions during undervolting. Then, we also analyze the impact of locations (SMs, threads, and bits) and operand data values on the error characteristics. Based on our observations, we reduce the voltage to the minimum safe limit for different instructions which achieves 18.37% energy saving, and we further propose an error detection strategy which reduces the performance and energy overhead by 14.8% with negligible 0.01% degradation for error detection rate.
Jingweijia Tan, Jiashuo Wang, Kaige Yan, Xiaohui Wei 0002, Xin Fu 0001
IEEE Trans. Computers2
2025 DirecGeo-GS: direction-aware and geometry-enhanced Gaussian splatting for Tibetan temple symbols
Mingqiang Zhou, Kexin Qiu, Jiashuo Wang
J. Supercomput.5
2024 Instruct Once, Chat Consistently in Multiple Rounds: An Efficient Tuning Framework for Dialogue
abstract
Tuning language models for dialogue generation has been a prevalent paradigm for building capable dialogue agents.Yet, traditional tuning narrowly views dialogue generation as resembling other language generation tasks, ignoring the role disparities between two speakers and the multi-round interactive process that dialogues ought to be.Such a manner often leads to unsatisfactory chat consistency for the built agent.In this work, we emphasize the interactive, communicative nature of dialogue and argue that it is more feasible to model the speaker roles of agent and user separately, enabling the agent to adhere to its role consistently.With this in mind, we propose an efficient Multi-round Interactive Dialogue Tuning (MIDI-Tuning) framework 1 .It models the agent and user individually with two adapters built upon large language models.The adapters make use of respective utterances round by round in alternating order and they are tuned via a round-level memory caching mechanism.Extensive experiments demonstrate that, our framework performs superior to traditional finetuning and harbors the tremendous potential for improving dialogue consistency.
Jian Wang 0054, Chak Tou Leong, Jiashuo Wang, Dongding Lin, Wenjie Li 0002, Xiaoyong Wei
ACL (1)3
2023 Self-Detoxifying Language Models via Toxification Reversal
abstract
Language model detoxification aims to minimize the risk of generating offensive or harmful content in pretrained language models (PLMs) for safer deployment.Existing methods can be roughly categorized as finetuning-based and decoding-based.However, the former is often resource-intensive, while the latter relies on additional components and potentially compromises the generation fluency.In this paper, we propose a more lightweight approach that enables the PLM itself to achieve "selfdetoxification".Our method is built upon the observation that prepending a negative steering prompt can effectively induce PLMs to generate toxic content.At the same time, we are inspired by the recent research in the interpretability field, which formulates the evolving contextualized representations within the PLM as an information stream facilitated by the attention layers.Drawing on this idea, we devise a method to identify the toxification direction from the normal generation process to the one prompted with the negative prefix, and then steer the generation to the reversed direction by manipulating the information movement within the attention layers.Experimental results show that our approach, without any fine-tuning or extra components, can achieve comparable performance with state-of-the-art methods. 1 A simple approach to controlled text generation.
Chak Tou Leong, Jiashuo Wang, Jian Wang 0054, Wenjie Li 0002
EMNLP3
2023 Aligning Language Models with Human Preferences via a Bayesian Approach
abstract
In the quest to advance human-centric natural language generation (NLG) systems, ensuring alignment between NLG models and human preferences is crucial. For this alignment, current popular methods leverage a reinforcement learning (RL) approach with a reward model trained on feedback from humans. However, inherent disagreements due to the subjective nature of human preferences pose a significant challenge for training the reward model, resulting in a deterioration of the NLG performance. To tackle this issue, previous approaches typically rely on majority voting or averaging to consolidate multiple inconsistent preferences into a merged one. Although straightforward to understand and execute, such methods suffer from an inability to capture the nuanced degrees of disaggregation among humans and may only represent a specialized subset of individuals, thereby lacking the ability to quantitatively disclose the universality of human preferences. To address this challenge, this paper proposes a novel approach, which employs a Bayesian framework to account for the distribution of disagreements among human preferences as training a preference model, and names it as $\textbf{d-PM}$. Besides, considering the RL strategy's inefficient and complex training process over the training efficiency, we further propose utilizing the contrastive learning strategy to train the NLG model with the preference scores derived from the d-PM model. Extensive experiments on two human-centric NLG tasks, i.e., emotional support conversation and integrity ``Rule-of-Thumb'' generation, show that our method consistently exceeds previous SOTA models in both automatic and human evaluations.
Jiashuo Wang, Haozhao Wang, Shichao Sun, Wenjie Li 0002
NeurIPS1
2022 Improving Multi-turn Emotional Support Dialogue Generation with Lookahead Strategy Planning
abstract
Providing Emotional Support (ES) to soothe people in emotional distress is an essential capability in social interactions.Most existing research on building ES conversation systems only considers single-turn interactions with users, which is over-simplified.In comparison, multi-turn ES conversation systems can provide ES more effectively, but face several new technical challenges, including: i) how to conduct support strategy planning that could lead to the best supporting effects; ii) how to dynamically model the user's state.In this paper, we propose a novel system named MultiESC to address these issues.For strategy planning, drawing inspiration from the A* search algorithm, we propose lookahead heuristics to estimate the future user feedback after using particular strategies, which helps to select strategies that can lead to the best long-term effects.For user state modeling, MultiESC focuses on capturing users' subtle emotional expressions and understanding their emotion causes.Extensive experiments show that MultiESC significantly outperforms competitive baselines in both strategy planning and dialogue generation.Our codes are available at https: //github.com/lwgkzl/MultiESC.
Wenge Liu, Wenjie Li 0002, Jiashuo Wang, Ruihui Zhao, Bang Liu 0003, Xiaodan Liang, Yefeng Zheng 0001
EMNLP4
2021 Empathetic Response Generation through Graph-based Multi-hop Reasoning on Emotional Causality
Jiashuo Wang, Wenjie Li 0002, Peiqin Lin, Feiteng Mu
Knowl. Based Syst.1