Jian Wang 0054

dblp:39/449-54 · DBLP profile ↗
← Back
17ranked-venue papers
7as first author
15since 2021 · last 2026
0000-0002-8992-8336ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 6 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2026 OptScale: Probabilistic Optimality for Inference-time Scaling
abstract
Inference-time scaling has emerged as a powerful technique for enhancing the reasoning performance of Large Language Models (LLMs). However, existing approaches often rely on heuristic strategies for parallel sampling, lacking a principled foundation. To address this gap, we propose a probabilistic framework that formalizes the optimality of inference-time scaling under the assumption that parallel samples are independently and identically distributed (i.i.d.), and where the Best-of-N selection strategy follows a probability distribution that can be estimated. Within this framework, we derive a theoretical lower bound on the required number of samples to achieve a target performance level, providing the first principled guidance for compute-efficient scaling. Leveraging this insight, we develop OptScale, a practical algorithm that dynamically determines the optimal number of sampled responses. OptScale employs a language model-based predictor to estimate probabilistic prior parameters, enabling the decision of the minimal number of samples needed that satisfy predefined performance thresholds and confidence levels. Extensive experiments on representative reasoning benchmarks (including MATH-500, GSM8K, AIME, and AMC) demonstrate that OptScale significantly reduces sampling overhead while remaining better or on par with state-of-the-art reasoning performance. Our work offers both a theoretical foundation and a practical solution for principled inference-time scaling, addressing a critical gap in the efficient deployment of LLMs for complex reasoning.
Youkang Wang, Jian Wang 0054, Rubing Chen, Xiaoyong Wei
AAAI2
2026 Where and What: Reasoning Dynamic and Implicit Preferences in Situated Conversational Recommendation
abstract
Situated conversational recommendation (SCR), which utilizes visual scenes grounded in specific environments and natural language dialogue to deliver contextually appropriate recommendations, has emerged as a promising research direction due to its close alignment with real-world scenarios.Compared to traditional recommendations, SCR requires a deeper understanding of dynamic and implicit user preferences, as the surrounding scene often influences users' underlying interests, while both may evolve across conversations.This complexity significantly impacts the timing and relevance of recommendations.To address this, we propose situated preference reasoning (SiPeR), a novel framework that integrates two core mechanisms: (i) Scene transition estimation, which estimates whether the current scene satisfies user needs, and guides the user toward a more suitable scene when necessary; and (ii) Bayesian inverse inference, which leverages the likelihood of multimodal large language models (MLLMs) to predict user preferences about candidate items within the scene.Extensive experiments on two representative benchmarks demonstrate SiPeR's superiority in both recommendation accuracy and response generation quality.The code and data are available at https://github.com/DongdingLin/SiPeR.
Dongding Lin, Jian Wang 0054, Yongqi Li 0001, Wenjie Li 0002
ACL (1)2
2026 Foresight Optimization for Strategic Reasoning in Large Language Models
abstract
Jessie Wang, Jiawen Duan, Jian Wang, Kaitao Song, Chunpu Xu, Johnny K. W. Ho, YU Fenggang, Johan F. Hoorn, Wenjie Li. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Jessie Wang 0004, Jiawen Duan, Jian Wang 0054, Kaitao Song, Chunpu Xu, Johnny K. W. Ho, Fenggang Yu, Johan F. Hoorn, Wenjie Li 0002
ACL (1)3
2025 Why Safeguarded Ships Run Aground? Aligned Large Language Models' Safety Mechanisms Tend to Be Anchored in The Template Region
abstract
The safety alignment of large language models (LLMs) remains vulnerable, as their initial behavior can be easily jailbroken by even relatively simple attacks.Since infilling a fixed template between the input instruction and initial model output is a common practice for existing LLMs, we hypothesize that this template is a key factor behind their vulnerabilities: LLMs' safety-related decision-making overly relies on the aggregated information from the template region, which largely influences these models' safety behavior.We refer to this issue as template-anchored safety alignment.In this paper, we conduct extensive experiments and verify that template-anchored safety alignment is widespread across various aligned LLMs.Our mechanistic analyses demonstrate how it leads to models' susceptibility when encountering inference-time jailbreak attacks.Furthermore, we show that detaching safety mechanisms from the template region is promising in mitigating vulnerabilities to jailbreak attacks.We encourage future research to develop more robust safety alignment techniques that reduce reliance on the template region.
Chak Tou Leong, Qingyu Yin, Jian Wang 0054, Wenjie Li 0002
ACL (1)3
2025 Towards Harmless Multimodal Assistants with Blind Preference Optimization
abstract
Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities in multimodal understanding, reasoning, and interaction. Given the extensive applications of MLLMs, the associated safety issues have become increasingly critical. Due to the effectiveness of preference optimization in aligning MLLMs with human preferences, there is an urgent need for safety-related preference data for MLLMs. To address this, we construct the MMSafe-PO preference dataset towards harmless multimodal assistants, featuring multimodal instructions, the conversational format, and ranked paired responses from human feedback. We also identify two insightful observations: modality co-defense and modality cheating, which illustrate that MLLMs possess a certain level of inherent defense while still presenting unique safety challenges. Based on these observations, we propose the Blind Preference Optimization (BPO) approach. Comprehensive experiments on three benchmarks show that BPO effectively enhances the safety capabilities of MLLMs. Notably, BPO significantly improves the safety rate of the base MLLM by 45.0%, outperforming the DPO approach. Additionally, applying BPO to the MMSafe-PO dataset greatly reduces the base MLLM's unsafe rate on other safety benchmarks (14.5% on MM-SafetyBench and 82.9% on HarmEval), demonstrating the effectiveness and robustness of both the dataset and the approach.
Yongqi Li 0001, Lu Yang 0008, Jian Wang 0054, Runyang You, Wenjie Li 0002, Liqiang Nie
ACM Multimedia3
2024 Cooper: Coordinating Specialized Agents towards a Complex Dialogue Goal
abstract
In recent years, there has been a growing interest in exploring dialogues with more complex goals, such as negotiation, persuasion, and emotional support, which go beyond traditional service-focused dialogue systems. Apart from the requirement for much more sophisticated strategic reasoning and communication skills, a significant challenge of these tasks lies in the difficulty of objectively measuring the achievement of their goals in a quantifiable way, making it difficult for existing research to directly optimize the dialogue procedure towards them. In our work, we emphasize the multifaceted nature of complex dialogue goals and argue that it is more feasible to accomplish them by comprehensively considering and jointly promoting their different aspects. To this end, we propose a novel dialogue framework, Cooper, which coordinates multiple specialized agents, each dedicated to a specific dialogue goal aspect separately, to approach the complex objective. Through this divide-and-conquer manner, we make complex dialogue goals more approachable and elicit greater intelligence via the collaboration of individual agents. Experiments on persuasion and emotional support dialogues demonstrate the superiority of our method over a set of competitive baselines. Our codes are available at https://github.com/YiCheng98/Cooper.
Wenge Liu, Jian Wang 0054, Chak Tou Leong, Wenjie Li 0002, Xian Wu 0001, Yefeng Zheng 0001
AAAI3
2024 Instruct Once, Chat Consistently in Multiple Rounds: An Efficient Tuning Framework for Dialogue
abstract
Tuning language models for dialogue generation has been a prevalent paradigm for building capable dialogue agents.Yet, traditional tuning narrowly views dialogue generation as resembling other language generation tasks, ignoring the role disparities between two speakers and the multi-round interactive process that dialogues ought to be.Such a manner often leads to unsatisfactory chat consistency for the built agent.In this work, we emphasize the interactive, communicative nature of dialogue and argue that it is more feasible to model the speaker roles of agent and user separately, enabling the agent to adhere to its role consistently.With this in mind, we propose an efficient Multi-round Interactive Dialogue Tuning (MIDI-Tuning) framework 1 .It models the agent and user individually with two adapters built upon large language models.The adapters make use of respective utterances round by round in alternating order and they are tuned via a round-level memory caching mechanism.Extensive experiments demonstrate that, our framework performs superior to traditional finetuning and harbors the tremendous potential for improving dialogue consistency.
Jian Wang 0054, Chak Tou Leong, Jiashuo Wang, Dongding Lin, Wenjie Li 0002, Xiaoyong Wei
ACL (1)1
2024 SCREEN: A Benchmark for Situated Conversational Recommendation
abstract
Engaging in conversational recommendations within a specific scenario represents a promising paradigm in the real world. Scenario-relevant situations often affect conversations and recommendations from two closely related aspects: varying the appealingness of items to users, namely situated item representation, and shifting user interests in the targeted items, namely situated user preference. We highlight that considering those situational factors is crucial, as this aligns with the realistic conversational recommendation process in the physical world. However, it is challenging yet under-explored. In this work, we are pioneering to bridge this gap and introduce a novel setting: Situated Conversational Recommendation Systems (SCRS). We observe an emergent need for high-quality datasets, and building one from scratch requires tremendous human effort. To this end, we construct a new benchmark, named SCREEN, via a role-playing method based on multimodal large language models. We take two multimodal large language models to play the roles of a user and a recommender, simulating their interactions in a co-observed scene. Our SCREEN comprises over 20k dialogues across 1.5k diverse situations, providing a rich foundation for exploring situational influences on conversational recommendations. Based on the SCREEN, we propose three worth-exploring subtasks and evaluate several representative baseline models. Our evaluations suggest that the benchmark is high quality, establishing a solid experimental basis for future research. The code and data are available at https://github.com/DongdingLin/SCREEN.
Dongding Lin, Jian Wang 0054, Chak Tou Leong, Wenjie Li 0002
ACM Multimedia2
2024 A Target-Driven Planning Approach for Goal-Directed Dialog Systems
abstract
Existing dialog systems mainly build social bonds reactively with users for chitchat or assist users with specific tasks. In this work, we push forward to a promising yet under-explored proactive dialog paradigm called goal-directed dialog systems, where the "goal" refers to achieving the recommendation for a predetermined target topic through social conversations. We focus on how to make plans that naturally lead users to achieve the goal through smooth topic transitions. To this end, we propose a target-driven planning network (TPNet) to drive the system to transit between different conversation stages. Built upon the widely used transformer architecture, TPNet frames the complicated planning process as a sequence generation task, which plans a dialog path consisting of dialog actions and topics. We then apply our TPNet with planned content to guide dialog generation using various backbone models. Extensive experiments show that our approach obtains the state-of-the-art performance in automatic and human evaluations. The results demonstrate that TPNet affects the improvement of goal-directed dialog systems significantly.
Jian Wang 0054, Dongding Lin, Wenjie Li 0002
IEEE Trans. Neural Networks Learn. Syst.1
2024 Target-constrained Bidirectional Planning for Generation of Target-oriented Proactive Dialogue
abstract
Target-oriented proactive dialogue systems aim at leading conversations from a dialogue context toward a pre-determined target, such as making recommendations on designated items or introducing new specific topics. To this end, it is critical for such dialogue systems to plan reasonable actions to drive the conversation proactively, and meanwhile, to plan appropriate topics to move the conversation forward to the target topic smoothly. In this work, we mainly focus on effective dialogue planning for target-oriented dialogue generation. Inspired by decision-making theories in cognitive science, we propose a novel target-constrained bidirectional planning (TRIP) approach, which plans an appropriate dialogue path by looking ahead and looking back. By formulating the planning as a generation task, our TRIP bidirectionally generates a dialogue path consisting of a sequence of pairs using two Transformer decoders. They are expected to supervise each other and converge on consistent actions and topics by minimizing the decision gap and contrastive generation of targets. Moreover, we propose a target-constrained decoding algorithm with a bidirectional agreement to better control the planning process. Subsequently, we adopt the planned dialogue paths to guide dialogue generation in a pipeline manner, where we explore two variants: prompt-based generation and plan-controlled generation. Extensive experiments are conducted on two challenging dialogue datasets, which are re-purposed for exploring target-oriented dialogue. Our automatic and human evaluations demonstrate that the proposed methods significantly outperform various baseline models.
Jian Wang 0054, Dongding Lin, Wenjie Li 0002
ACM Trans. Inf. Syst.1
2023 COLA: Improving Conversational Recommender Systems by Collaborative Augmentation
abstract
Conversational recommender systems (CRS) aim to employ natural language conversations to suggest suitable products to users. Understanding user preferences for prospective items and learning efficient item representations are crucial for CRS. Despite various attempts, earlier studies mostly learned item representations based on individual conversations, ignoring item popularity embodied among all others. Besides, they still need support in efficiently capturing user preferences since the information reflected in a single conversation is limited. Inspired by collaborative filtering, we propose a collaborative augmentation (COLA) method to simultaneously improve both item representation learning and user preference modeling to address these issues. We construct an interactive user-item graph from all conversations, which augments item representations with user-aware information, i.e., item popularity. To improve user preference modeling, we retrieve similar conversations from the training corpus, where the involved items and attributes that reflect the user's potential interests are used to augment the user representation through gate control. Extensive experiments on two benchmark datasets demonstrate the effectiveness of our method. Our code and data are available at https://github.com/DongdingLin/COLA.
Dongding Lin, Jian Wang 0054, Wenjie Li 0002
AAAI2
2023 Self-Detoxifying Language Models via Toxification Reversal
abstract
Language model detoxification aims to minimize the risk of generating offensive or harmful content in pretrained language models (PLMs) for safer deployment.Existing methods can be roughly categorized as finetuning-based and decoding-based.However, the former is often resource-intensive, while the latter relies on additional components and potentially compromises the generation fluency.In this paper, we propose a more lightweight approach that enables the PLM itself to achieve "selfdetoxification".Our method is built upon the observation that prepending a negative steering prompt can effectively induce PLMs to generate toxic content.At the same time, we are inspired by the recent research in the interpretability field, which formulates the evolving contextualized representations within the PLM as an information stream facilitated by the attention layers.Drawing on this idea, we devise a method to identify the toxification direction from the normal generation process to the one prompted with the negative prefix, and then steer the generation to the reversed direction by manipulating the information movement within the attention layers.Experimental results show that our approach, without any fine-tuning or extra components, can achieve comparable performance with state-of-the-art methods. 1 A simple approach to controlled text generation.
Chak Tou Leong, Jiashuo Wang, Jian Wang 0054, Wenjie Li 0002
EMNLP4
2023 Target-oriented Proactive Dialogue Systems with Personalization: Problem Formulation and Dataset Curation
abstract
Target-oriented dialogue systems, designed to proactively steer conversations toward predefined targets or accomplish specific system-side goals, are an exciting area in conversational AI.In this work, by formulating a pair as the conversation target, we explore a novel problem of personalized targetoriented dialogue by considering personalization during the target accomplishment process.However, there remains an emergent need for high-quality datasets, and building one from scratch requires tremendous human effort.To address this, we propose an automatic dataset curation framework using a role-playing approach.Based on this framework, we construct a large-scale personalized target-oriented dialogue dataset, TOPDIAL 1 , which comprises about 18K multi-turn dialogues.The experimental results show that this dataset is of high quality and could contribute to exploring personalized target-oriented dialogue.
Jian Wang 0054, Dongding Lin, Chak Tou Leong, Wenjie Li 0002
EMNLP1
2023 Majority-to-minority resampling for boosting-based classification under imbalanced data
Gaoshan Wang, Jian Wang 0054, Kejing He 0001
Appl. Intell.2
2021 Template-guided Clarifying Question Generation for Web Search Clarification
abstract
Clarification has attracted much attention because of its many potential applications especially in Web search. Since search queries are very short, the underlying user intents are often ambiguous. This makes it challenging for search engines to return the appropriate results that pertain to the users' actual information needs. To address this issue, asking clarifying questions has been recognized as a critical technique. Although previous studies have analyzed the importance of asking to clarify, generating clarifying questions for Web search remains under-explored. In this paper, we tackle this problem in a template-guided manner. Our objective is jointly learning to select question templates and fill question slots, using Transformer-based networks. We conduct experiments on MIMICS, a collection of datasets containing real Web search queries sampled from Bing's search logs. Our method is demonstrated to achieve significant improvements over various competitive baselines.
Jian Wang 0054, Wenjie Li 0002
CIKM1
2020 Improving Knowledge-Aware Dialogue Generation via Knowledge Base Question Answering
abstract
Neural network models usually suffer from the challenge of incorporating commonsense knowledge into the open-domain dialogue systems. In this paper, we propose a novel knowledge-aware dialogue generation model (called TransDG), which transfers question representation and knowledge matching abilities from knowledge base question answering (KBQA) task to facilitate the utterance understanding and factual knowledge selection for dialogue generation. In addition, we propose a response guiding attention and a multi-step decoding strategy to steer our model to focus on relevant features for response generation. Experiments on two benchmark datasets demonstrate that our model has robust superiority over compared methods in generating informative and fluent dialogues. Our code is available at https://github.com/siat-nlp/TransDG.
Jian Wang 0054, Junhao Liu 0001, Wei Bi, Xiaojiang Liu, Kejing He 0001, Ruifeng Xu 0001, Min Yang 0007
AAAI1
2020 Dual Dynamic Memory Network for End-to-End Multi-turn Task-oriented Dialog Systems
abstract
Existing end-to-end task-oriented dialog systems struggle to dynamically model long dialog context for interactions and effectively incorporate knowledge base (KB) information into dialog generation. To conquer these limitations, we propose a Dual Dynamic Memory Network (DDMN) for multi-turn dialog generation, which maintains two core components: dialog memory manager and KB memory manager. The dialog memory manager dynamically expands the dialog memory turn by turn and keeps track of dialog history with an updating mechanism, which encourages the model to filter irrelevant dialog history and memorize important newly coming information. The KB memory manager shares the structural KB triples throughout the whole conversation, and dynamically extracts KB information with a memory pointer at each turn. Experimental results on three benchmark datasets demonstrate that DDMN significantly outperforms the strong baselines in terms of both automatic evaluation and human evaluation. Our code is available at https://github.com/siat-nlp/DDMN.
Jian Wang 0054, Junhao Liu 0001, Wei Bi, Xiaojiang Liu, Kejing He 0001, Ruifeng Xu 0001, Min Yang 0007
COLING1