Zijing Shi

dblp:313/5293 · DBLP profile ↗
← Back
13ranked-venue papers
5as first author
13since 2021 · last 2026
0009-0001-4872-9234ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 5 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2026 Benchmarking Web Agent Safety under E-commerce Deceptive Interfaces
abstract
As autonomous web agents are increasingly deployed to perform real-world tasks, ensuring their safety has become a critical concern.In this work, we study web agent behavior under realistic deceptive interfaces in the ecommerce domain.We introduce WebDecept, a lightweight and configurable plugin framework that enables controlled injection of deceptive interface patterns into existing web environments.Using WebDecept, we instantiate seven deceptive patterns commonly observed on the open web, including targeted advertisements, domain redirection, and shopping manipulation.By injecting these patterns into the frontend during task execution, we perform controlled evaluation of multiple multimodal web agents.Our results show that current web agents are highly susceptible to multiple classes of deceptive interfaces, and that prompt-based constraints are often insufficient to mitigate these failures.We further analyze how the design choices of deceptive patterns influence the success of such manipulations.These findings highlight safety challenges that should be addressed as web agents are scaled toward realworld deployment.
Zijing Shi, Ling Chen 0006
ACL (1)1
2026 Quantifying and mitigating the spiral of silence in recommender systems: A modular probabilistic framework
abstract
Online rating systems are an indispensable building block for many web applications, yet they are pervasively affected by data biases that distort our understanding of user preferences. A key source of this bias is the “Spiral of Silence” (SoS) phenomenon, where users who perceive their opinions to be in the minority tend to withhold them, leading to non-random missing data. While initial research has begun to explore this issue, existing SoS models for recommender systems fail to capture real-world complexities, such as rating recency, user heterogeneity (e.g., “hardcore” users immune to social pressure), and asymmetric responses to conforming versus dissenting opinions. Furthermore, these models often lack formal theoretical guarantees. To fill this gap, we introduce Bi-MCP, a modular probabilistic framework to quantify and mitigate SoS effects. Bi-MCP models a bifurcated user population (hardcore vs. conformist) and their bidirectional (asymmetric) response to the perceived opinion climate. The framework is (i) Expressive, capturing novel factors missed by prior work; (ii) Modular, allowing its components to be extended or replaced to support future research; and (iii) Theoretically grounded, with a formal proof of convergence for its Generalized Expectation-Maximization (GEM) inference algorithm. Experimental results on four large-scale public datasets demonstrate that our model significantly improves recommendation accuracy over state-of-the-art baselines, showcasing the practical benefit of explicitly modeling these nuanced SoS effects.
Mingze Zhong, Hong Xie 0004, Zijing Shi, Ling Chen 0006
Knowl. Based Syst.4
2025 Monte Carlo Planning with Large Language Model for Text-Based Game Agents
abstract
Text-based games provide valuable environments for language-based autonomous agents. However, planning-then-learning paradigms, such as those combining Monte Carlo Tree Search (MCTS) and reinforcement learning (RL), are notably time-consuming due to extensive iterations. Additionally, these algorithms perform uncertainty-driven exploration but lack language understanding and reasoning abilities. In this paper, we introduce the Monte Carlo planning with Dynamic Memory-guided Large language model (MC-DML) algorithm. MC-DML leverages the language understanding and reasoning capabilities of Large Language Models (LLMs) alongside the exploratory advantages of tree search algorithms. Specifically, we enhance LLMs with in-trial and cross-trial memory mechanisms, enabling them to learn from past experiences and dynamically adjust action evaluations during planning. We conduct experiments on a series of text-based games from the Jericho benchmark. Our results demonstrate that the MC-DML algorithm significantly enhances performance across various games at the initial planning phase, outperforming strong contemporary methods that require multiple iterations. This demonstrates the effectiveness of our algorithm, paving the way for more efficient language-grounded planning in complex environments.
Zijing Shi, Ling Chen 0006
ICLR1
2025 Hierarchical Multi-Agent Framework for Dynamic Macroeconomic Modelling Using Large Language Models
Zhixun Chen, Zijing Shi, Yaodong Yang 0001, Yali Du 0001
AAMAS2
2025 Hazards in Daily Life? Enabling Robots to Proactively Detect and Resolve Anomalies
abstract
Zirui Song, Guangxian Ouyang, Meng Fang, Hongbin Na, Zijing Shi, Zhenhao Chen, Fu Yujie, Zeyu Zhang, Shiyu Jiang, Miao Fang, Ling Chen, Xiuying Chen. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Zirui Song, Guangxian Ouyang, Hongbin Na, Zijing Shi, Zhenhao Chen, Yujie Fu, Zeyu Zhang 0006, Miao Fang 0001, Ling Chen 0006, Xiuying Chen
NAACL (Long Papers)5
2025 Bridging Confidence and Competence: Evaluating Self-assessment Alignment in LLM Mathematical Reasoning
Mingze Zhong, Zijing Shi, Ling Chen 0006
PRICAI2
2024 Large Language Models Are Neurosymbolic Reasoners
abstract
A wide range of real-world applications is characterized by their symbolic nature, necessitating a strong capability for symbolic reasoning. This paper investigates the potential application of Large Language Models (LLMs) as symbolic reasoners. We focus on text-based games, significant benchmarks for agents with natural language capabilities, particularly in symbolic tasks like math, map reading, sorting, and applying common sense in text-based worlds. To facilitate these agents, we propose an LLM agent designed to tackle symbolic challenges and achieve in-game objectives. We begin by initializing the LLM agent and informing it of its role. The agent then receives observations and a set of valid actions from the text-based games, along with a specific symbolic module. With these inputs, the LLM agent chooses an action and interacts with the game environments. Our experimental results demonstrate that our method significantly enhances the capability of LLMs as automated agents for symbolic reasoning, and our LLM agent is effective in text-based games involving symbolic tasks, achieving an average performance of 88% across all tasks.
Shilong Deng, Yudi Zhang 0006, Zijing Shi, Ling Chen 0006, Mykola Pechenizkiy, Jun Wang 0012
AAAI4
2024 Human-Guided Moral Decision Making in Text-Based Games
abstract
Training reinforcement learning (RL) agents to achieve desired goals while also acting morally is a challenging problem. Transformer-based language models (LMs) have shown some promise in moral awareness, but their use in different contexts is problematic because of the complexity and implicitness of human morality. In this paper, we build on text-based games, which are challenging environments for current RL agents, and propose the HuMAL (Human-guided Morality Awareness Learning) algorithm, which adaptively learns personal values through human-agent collaboration with minimal manual feedback. We evaluate HuMAL on the Jiminy Cricket benchmark, a set of text-based games with various scenes and dense morality annotations, using both simulated and actual human feedback. The experimental results demonstrate that with a small amount of human feedback, HuMAL can improve task performance and reduce immoral behavior in a variety of games and is adaptable to different personal values.
Zijing Shi, Ling Chen 0006, Yali Du 0001, Jun Wang 0012
AAAI1
2024 Enhancing Retrieval and Managing Retrieval: A Four-Module Synergy for Improved Quality and Efficiency in RAG Systems
abstract
Retrieval-augmented generation (RAG) techniques leverage the in-context learning capabilities of large language models (LLMs) to produce more accurate and relevant responses. Originating from the simple ‘retrieve-then-read’ approach, the RAG framework has evolved into a highly flexible and modular paradigm. A critical component, the Query Rewriter module, enhances knowledge retrieval by generating a search-friendly query. This method aligns input questions more closely with the knowledge base. Our research identifies opportunities to enhance the Query Rewriter module to Query Rewriter+ by generating multiple queries to overcome the Information Plateaus associated with a single query and by rewriting questions to eliminate Ambiguity, thereby clarifying the underlying intent. We also find that current RAG systems exhibit issues with Irrelevant Knowledge; to overcome this, we propose the Knowledge Filter. These two modules are both based on the instruction-tuned Gemma-2B model, which together enhance response quality. The final identified issue is Redundant Retrieval; we introduce the Memory Knowledge Reservoir and the Retriever Trigger to solve this. The former supports the dynamic expansion of the RAG system’s knowledge base in a parameter-free manner, while the latter optimizes the cost for accessing external knowledge, thereby improving resource utilization and response efficiency. These four RAG modules synergistically improve the response quality and efficiency of the RAG system. The effectiveness of these modules has been validated through experiments and ablation studies across six common QA datasets. The source code can be accessed at https://github.com/Ancientshi/ERM4.
Yunxiao Shi, Xing Zi, Zijing Shi, Haimin Zhang 0001, Qiang Wu 0001, Min Xu 0001
ECAI3
2023 CHBias: Bias Evaluation and Mitigation of Chinese Conversational Language Models
abstract
Warning: This paper contains content that may be offensive or upsetting.Pretrained conversational agents have been exposed to safety issues, exhibiting a range of stereotypical human biases such as gender bias.However, there are still limited bias categories in current research, and most of them only focus on English.In this paper, we introduce a new Chinese dataset, CHBias, for bias evaluation and mitigation of Chinese conversational language models.Apart from those previous well-explored bias categories, CHBias includes under-explored bias categories, such as ageism and appearance biases, which received less attention.We evaluate two popular pretrained Chinese conversational models, CDial-GPT and EVA2.0, using CHBias.Furthermore, to mitigate different biases, we apply several debiasing methods to the Chinese pretrained models.Experimental results show that these Chinese pretrained models are potentially risky for generating texts that contain social biases, and debiasing methods using the proposed dataset can make response generation less biased while preserving the models' conversational capabilities.
Jiaxu Zhao 0002, Zijing Shi, Ling Chen 0006, Mykola Pechenizkiy
ACL (1)3
2023 Self-imitation Learning for Action Generation in Text-based Games
abstract
In this work, we study reinforcement learning (RL) in solving text-based games.We address the challenge of combinatorial action space, by proposing a confidence-based self-imitation model to generate action candidates for the RL agent.Firstly, we leverage the self-imitation learning to rank and exploit past valuable trajectories to adapt a pre-trained language model (LM) towards a target game.Then, we devise a confidence-based strategy to measure the LM's confidence with respect to a state, thus adaptively pruning the generated actions to yield a more compact set of action candidates.In multiple challenging games, our model demonstrates promising performance in comparison to the baselines.
Zijing Shi, Yunqiu Xu, Ling Chen 0006
EACL1
2023 Stay Moral and Explore: Learn to Behave Morally in Text-based Games
Zijing Shi, Yunqiu Xu, Ling Chen 0006, Yali Du 0001
ICLR1
2022 Cross-sectional analysis and data-driven forecasting of confirmed COVID-19 cases
Nan Jing, Zijing Shi, Ji Yuan
Appl. Intell.2