Minqian Liu

dblp:193/2086 · DBLP profile ↗
← Back
15ranked-venue papers
5as first author
14since 2021 · last 2026
0009-0001-6014-3949ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 5 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Domain Generalizable AI Guardrails with Augmented Policy Training
abstract
Minqian Liu, Ioana Baldini, David Rabinowitz, David S Rosenberg, Sebastian Gehrmann, Mark Dredze. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Minqian Liu, Ioana Baldini, David Rabinowitz, David S. Rosenberg, Sebastian Gehrmann, Mark Dredze
ACL (1)1
2026 From Vulnerable to Resilient: Examining Parent and Teen Perceptions on How to Respond to Unwanted Cybergrooming Advances
abstract
Cybergrooming is a form of online abuse that threatens teens’ mental health and physical safety. Yet, most prior work has focused on detecting perpetrators’ behaviors, leaving a limited understanding of how teens might respond to such unwanted advances. To address this gap, we conducted an online survey with 74 participants—51 parents and 23 teens—who responded to simulated cybergrooming scenarios in two ways: responses that they think would make teens more vulnerable or resilient to unwanted sexual advances. Through a mixed-methods analysis, we identified four types of vulnerable responses (encouraging escalation, accepting an advance, displaying vulnerability, and negating risk concern) and four types of protective strategies (setting boundaries, directly declining, signaling risk awareness, and leveraging avoidance techniques). As the cybergrooming risk escalated, both vulnerable responses and protective strategies showed a corresponding progression. This study contributes a teen-centered understanding of cybergrooming, a labeled dataset, and a stage-based taxonomy of perceived protective strategies, while offering implications for educational programs and sociotechnical interventions.
Xinyi Zhang 0007, Mamtaj Akter, Heajun An, Minqian Liu, Qi Zhang 0104, Lifu Huang, Jin-Hee Cho, Pamela J. Wisniewski, Sang Won Lee 0002
CHI4
2026 StagePilot: Stage-Level Planning for Long-Horizon Dialogue Simulation in Cybergrooming
abstract
Cybergrooming is an evolving threat to youth, requiring proactive educational interventions. We address this by modeling dialogue progression as a structured planning problem over stage-wise interactions. We propose StagePilot, a dialogue framework that separates stage-level planning from response generation, in which the model selects the next stage under constrained transitions and generates responses conditioned on it, enabling coherent and realistic progression. Reinforcement learning is used to learn stage-level policies from offline data, optimizing for both emotional alignment and goal-consistent progression. Our empirical experiments show that StagePilot generates more structured, coherent dialogue trajectories and reduces conversational stagnation compared to baselines; notably, the IQL+AWAC variant reaches the final stage more often while maintaining over 70% positive or neutral responses, yielding a 43% relative improvement.
Heajun An, Qi Zhang 0104, Minqian Liu, Xinyi Zhang 0007, Sang Won Lee 0002, Lifu Huang, Pamela J. Wisniewski, Jin-Hee Cho
SIGDIAL3
2025 Modality-Specialized Synergizers for Interleaved Vision-Language Generalists
abstract
Recent advancements in Vision-Language Models (VLMs) have led to the emergence of Vision-Language Generalists (VLGs) capable of understanding and generating both text and images. However, seamlessly generating an arbitrary sequence of text and images remains a challenging task for the current VLGs. One primary limitation lies in applying a unified architecture and the same set of parameters to simultaneously model discrete text tokens and continuous image features. Recent works attempt to tackle this fundamental problem by introducing modality-aware expert models. However, they employ identical architectures to process both text and images, disregarding the intrinsic inductive biases in these two modalities. In this work, we introduce Modality-Specialized Synergizers (MoSS), a novel design that efficiently optimizes existing unified architectures of VLGs with modality-specialized adaptation layers, i.e., a Convolutional LoRA for modeling the local priors of image patches and a Linear LoRA for processing sequential text. This design enables more effective modeling of modality-specific features while maintaining the strong cross-modal integration gained from pretraining. In addition, to improve the instruction-following capability on interleaved text-and-image generation, we introduce LeafInstruct, the first open-sourced interleaved instruction tuning dataset comprising 184,982 high-quality instances on more than 10 diverse domains. Extensive experiments show that VLGs integrated with MoSS achieve state-of-the-art performance, significantly surpassing baseline VLGs in complex interleaved generation tasks. Furthermore, our method exhibits strong generalizability on different VLGs.
Zhiyang Xu, Minqian Liu, Ying Shen 0006, Joy Rimchala, Jiaxin Zhang 0005, Qifan Wang 0001, Lifu Huang
ICLR2
2025 ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding
abstract
Structured image understanding, such as interpreting tables and charts, requires strategically refocusing across various structures and texts within an image, forming a reasoning sequence to arrive at the final answer. However, current multimodal large language models (LLMs) lack this multihop selective attention capability. In this work, we introduce ReFocus, a simple yet effective framework that equips multimodal LLMs with the ability to generate ``visual thoughts'' by performing visual editing on the input image through code, shifting and refining their visual focuses. Specifically, ReFocus enables multimodal LLMs to generate Python codes to call tools and modify the input image, sequentially drawing boxes, highlighting sections, and masking out areas, thereby enhancing the visual reasoning process. We experiment upon a wide range of structured image understanding tasks involving tables and charts. ReFocus largely improves performance on all tasks over GPT-4o without visual editing, yielding an average gain of 11.0% on table tasks and 6.8% on chart tasks. We present an in-depth analysis of the effects of different visual edits, and reasons why ReFocus can improve the performance without introducing additional information. Further, we collect a 14k training set using ReFocus, and prove that such visual chain-of-thought with intermediate information offers a better supervision than standard VQA data, reaching a 8.0% average gain over the same model trained with QA pairs and 2.6% over CoT.
Minqian Liu, Zhengyuan Yang, John Corring, Yijuan Lu, Dan Roth 0001, Dinei A. F. Florêncio, Cha Zhang
ICML2
2024 MULTISCRIPT: Multimodal Script Learning for Supporting Open Domain Everyday Tasks
abstract
Automatically generating scripts (i.e. sequences of key steps described in text) from video demonstrations and reasoning about the subsequent steps are crucial to the modern AI virtual assistants to guide humans to complete everyday tasks, especially unfamiliar ones. However, current methods for generative script learning rely heavily on well-structured preceding steps described in text and/or images or are limited to a certain domain, resulting in a disparity with real-world user scenarios. To address these limitations, we present a new benchmark challenge – MULTISCRIPT, with two new tasks on task-oriented multimodal script learning: (1) multimodal script generation, and (2) subsequent step prediction. For both tasks, the input consists of a target task name and a video illustrating what has been done to complete the target task, and the expected output is (1) a sequence of structured step descriptions in text based on the demonstration video, and (2) a single text description for the subsequent step, respectively. Built from WikiHow, MULTISCRIPT covers multimodal scripts in videos and text descriptions for over 6,655 human everyday tasks across 19 diverse domains. To establish baseline performance on MULTISCRIPT, we propose two knowledge-guided multimodal generative frameworks that incorporate the task-related knowledge prompted from large language models such as Vicuna. Experimental results show that our proposed approaches significantly improve over the competitive baselines.
Jingyuan Qi, Minqian Liu, Ying Shen 0006, Zhiyang Xu, Lifu Huang
AAAI2
2024 Ameli: Enhancing Multimodal Entity Linking with Fine-Grained Attributes
abstract
Barry Yao, Sijia Wang, Yu Chen, Qifan Wang, Minqian Liu, Zhiyang Xu, Licheng Yu, Lifu Huang. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Barry Menglong Yao, Yu Chen 0022, Qifan Wang 0001, Minqian Liu, Zhiyang Xu, Licheng Yu, Lifu Huang
EACL (1)5
2024 Holistic Evaluation for Interleaved Text-and-Image Generation
abstract
Interleaved text-and-image generation has been an intriguing research direction, where the models are required to generate both images and text pieces in an arbitrary order.Despite the emerging advancements in interleaved generation, the progress in its evaluation still significantly lags behind.Existing evaluation benchmarks do not support arbitrarily interleaved images and text for both inputs and outputs, and they only cover a limited number of domains and use cases.Also, current works predominantly use similarity-based metrics which fall short in assessing the quality in open-ended scenarios.To this end, we introduce INTER-LEAVEDBENCH, the first benchmark carefully curated for the evaluation of interleaved textand-image generation.INTERLEAVEDBENCH features a rich array of tasks to cover diverse real-world use cases.In addition, we present INTERLEAVEDEVAL, a strong reference-free metric powered by GPT-4o to deliver accurate and explainable evaluation.We carefully define five essential evaluation aspects for IN-TERLEAVEDEVAL, including text quality, perceptual quality, image coherence, text-image coherence, and helpfulness, to ensure a comprehensive and fine-grained assessment.Through extensive experiments and rigorous human evaluation, we show that our benchmark and metric can effectively evaluate the existing models with a strong correlation with human judgments surpassing previous reference-based metrics.We also provide substantial findings and insights to foster future research in interleaved generation and its evaluation. 1
Minqian Liu, Zhiyang Xu, Zihao Lin 0003, Trevor Ashby, Joy Rimchala, Jiaxin Zhang 0005, Lifu Huang
EMNLP1
2024 Towards Effective Long Conversation Generation with Dynamic Topic Tracking and Recommendation
abstract
During conversations, the human flow of thoughts may result in topic shifts and evolution.In open-domain dialogue systems, it is crucial to track the topics discussed and recommend relevant topics to be included in responses to have effective conversations.Furthermore, topic evolution is needed to prevent stagnation as conversation length increases.Existing open-domain dialogue systems do not pay sufficient attention to topic evolution and shifting, resulting in performance degradation due to ineffective responses as conversation length increases.To address the shortcomings of existing approaches, we propose EVOLV-CONV.EVOLVCONV conducts real-time conversation topic and user preference tracking and utilizes the tracking information to evolve and shift topics depending on conversation status.We conduct extensive experiments to validate the topic evolving and shifting capabilities of EVOLVCONV as conversation length increases.Un-referenced evaluation metric UniEval compare EVOLVCONV with the baselines.Experimental results show that EVOLV-CONV maintains a smooth conversation flow without abruptly shifting topics; the probability of topic shifting ranges between 5%-8% throughout the conversation.EVOLVCONV recommends 4.77% more novel topics than the baselines, and the topic evolution follows balanced topic groupings.Furthermore, we conduct user surveys to test the practical viability of EVOLVCONV.User survey results reveal that responses generated by EVOLVCONV are preferred 47.8% of the time compared to the baselines and comes second to real human responses.
Trevor Ashby, Adithya Kulkarni, Jingyuan Qi, Minqian Liu, Eunah Cho, Vaibhav Kumar, Lifu Huang
INLG4
2024 X-Eval: Generalizable Multi-aspect Text Evaluation via Augmented Instruction Tuning with Auxiliary Evaluation Aspects
abstract
Minqian Liu, Ying Shen, Zhiyang Xu, Yixin Cao, Eunah Cho, Vaibhav Kumar, Reza Ghanadan, Lifu Huang. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Minqian Liu, Ying Shen 0006, Zhiyang Xu, Yixin Cao 0002, Eunah Cho, Vaibhav Kumar, Reza Ghanadan, Lifu Huang
NAACL-HLT1
2023 The Art of SOCRATIC QUESTIONING: Recursive Thinking with Large Language Models
abstract
Chain-of-Thought (CoT) prompting enables large language models to solve complex reasoning problems by generating intermediate steps.However, confined by its inherent singlepass and sequential generation process, CoT heavily relies on the initial decisions, causing errors in early steps to accumulate and impact the final answers.In contrast, humans adopt recursive thinking when tackling complex reasoning problems, i.e., iteratively breaking the original problem into approachable subproblems and aggregating their answers to resolve the original one.Inspired by the human cognitive process, we propose SOCRATIC QUESTIONING, a divide-and-conquer style algorithm that mimics the recursive thinking process.Specifically, SOCRATIC QUESTIONING leverages large language models to raise and answer sub-questions until collecting enough information to tackle the original question.Unlike CoT, SOCRATIC QUESTIONING explicitly navigates the thinking space, stimulates effective recursive thinking, and is more robust towards errors in the thinking process.Extensive experiments on several complex reasoning tasks, including MMLU, MATH, LogiQA, and visual question-answering demonstrate significant performance improvements over the stateof-the-art prompting methods, such as CoT, and Tree-of-Thought.The qualitative analysis clearly shows that the intermediate reasoning steps elicited by SOCRATIC QUESTIONING are similar to humans' recursively thinking process of complex reasoning problems 12 .
Jingyuan Qi, Zhiyang Xu, Ying Shen 0006, Minqian Liu, Qifan Wang 0001, Lifu Huang
EMNLP4
2022 Incremental Prompting: Episodic Memory Prompt for Lifelong Event Detection
abstract
Lifelong event detection aims to incrementally update a model with new event types and data while retaining the capability on previously learned old types. One critical challenge is that the model would catastrophically forget old types when continually trained on new data. In this paper, we introduce Episodic Memory Prompts (EMP) to explicitly retain the learned task-specific knowledge. Our method adopts continuous prompt for each task and they are optimized to instruct the model prediction and learn event-specific representation. The EMPs learned in previous tasks are carried along with the model in subsequent tasks, and can serve as a memory module that keeps the old knowledge and transferring to new tasks. Experiment results demonstrate the effectiveness of our method. Furthermore, we also conduct a comprehensive analysis of the new and old event types in lifelong learning.
Minqian Liu, Shiyu Chang, Lifu Huang
COLING1
2022 Co-attention network with label embedding for text classification
Minqian Liu, Lizhao Liu, Junyi Cao
Neurocomputing1
2021 Progressive Dialogue State Tracking for Multi-Domain Dialogue Systems
Minqian Liu, Xiaojun Quan
ICASSP2
2020 Dynamic Extension Nets for Few-shot Semantic Segmentation
abstract
Semantic segmentation requires a large amount of densely annotated data for training and may generalize poorly to novel categories. In real-world applications, we have an urgent need for few-shot semantic segmentation which aims to empower a model to handle unseen object categories with limited data. This task is non-trivial due to several challenges. First, it is difficult to extract the class-relevant information to handle the novel class as only a few samples are available. Second, since the image content can be very complex, the novel class information may be suppressed by the base categories due to limited data. Third, one may easily learn promising base classifiers based on a large amount of training data, but it is non-trivial to exploit the knowledge to train the novel classifiers. More critically, once a novel classifier is built, the output probability space will change. How to maintain the base classifiers and dynamically include the novel classifiers remains an open question. To address the above issues, we propose a Dynamic Extension Network (DENet) in which we dynamically construct and maintain a classifier for the novel class by leveraging the knowledge from the base classes and the information from novel data. More importantly, to overcome the information suppression issue, we design a Guided Attention Module (GAM), which can be plugged into any framework to help learn class-relevant features. Last, rather than directly train the model with limited data, we propose a dynamic extension training algorithm to predict the weights of novel classifiers, which is able to exploit the knowledge of base classifiers by dynamically extending classes during training. The extensive experiments show that our proposed method achieves state-of-the-art performance on the PASCAL-5i and COCO-20i datasets. The source code is available at https://github.com/lizhaoliu-Lec/DENet.
Lizhao Liu, Junyi Cao, Minqian Liu, Qi Chen 0014, Mingkui Tan
ACM Multimedia3