Qicheng Li

dblp:71/28 · DBLP profile ↗
← Back
20ranked-venue papers
1as first author
15since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 6 since 2021Software engineering, systems software and programming languages · 2Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 AgentCDM: Enhancing Multi-Agent Collaborative Decision-Making via ACH-Inspired Structured Reasoning
abstract
Multi-agent systems (MAS) powered by large language models (LLMs) hold significant promise for solving complex decision-making tasks. However, the core process of collaborative decision-making (CDM) within these systems remains underexplored. Existing approaches often rely on either "dictatorial" strategies that are vulnerable to the cognitive biases of a single agent, or "voting-based" methods that fail to fully harness collective intelligence. To address these limitations, we propose AgentCDM, a structured framework for enhancing collaborative decision-making in LLM-based multi-agent systems. Drawing inspiration from the Analysis of Competing Hypotheses (ACH) in cognitive science, AgentCDM introduces a structured reasoning paradigm that systematically mitigates cognitive biases and shifts decision-making from passive answer selection to active hypothesis evaluation and construction. To internalize this reasoning process, we develop a two-stage training paradigm: the first stage uses explicit ACH-inspired scaffolding to guide the model through structured reasoning, while the second stage progressively removes this scaffolding to encourage autonomous generalization. Experiments on multiple benchmark datasets demonstrate that AgentCDM achieves state-of-the-art performance and exhibits strong generalization, validating its effectiveness in improving the quality and robustness of collaborative decisions in MAS.
Shiwan Zhao, Hualong Yu, Qicheng Li
AAAI5
2026 RealTalk-CN: A Realistic Chinese Speech Task-Oriented Dialogue Benchmark with Cross-Modal Analysis
abstract
Recent advances in speech large language models (e.g., GPT-4o) have enabled end-to-end spoken interactions, yet their robustness in realworld applications remains unclear, where systems must assist users in completing specific tasks under complex conditions such as multiturn, ambiguous, and often spontaneous speech, as well as natural alternation between speech and text.Task-oriented dialogue (TOD) offers a realistic scenario to evaluate whether models can effectively help users accomplish such task-oriented goals, but existing benchmarks are mainly text-based, and the few speech datasets are limited to English and often neglect spontaneous disfluencies and speaker diversity.To address this gap, we introduce RealTalk-CN, the first Chinese multi-turn, multi-domain speech-text TOD dataset, containing 5.4k dialogues (60K turns, ~150 hours) of real human-to-human recordings with detailed annotations for dialogue states, disfluency types, and speaker characteristics.Based on this dataset, we propose a cross-modal interaction task supporting dynamic speech-text switching and a comprehensive evaluation protocol assessing robustness to disfluencies, sensitivity to speaker variation, and cross-domain generalization.Experiments on state-of-the-art models demonstrate the challenges posed by RealTalk-CN and establish its value as a benchmark for developing reliable and fair Speech LLMs in real-world deployments.The dataset and evaluation framework are available 1 to encourage further research.
Enzhi Wang, Jiaming Zhou 0001, Yuhang Jia, Aobo Kong, Qicheng Li
ACL (1)5
2026 Text-to-3D City: Plan-then-Execute Urban Generation With LLM Planners and Procedural Synthesis
abstract
ABSTRACT City‐scale 3D urban generation requires planning‐level semantic grounding from user intent and scalable geometric synthesis with structural validity and editability. Procedural content generation (PCG) offers controllability and scalability, but is hard to author due to high‐dimensional parameters and nonintuitive workflows. Meanwhile, directly generating city geometry or scripts from text with LLMs can suffer from weak large‐scale consistency and limited geometric validity, hindering downstream editing and engine deployment. We present Text‐to‐3D City, a plan‐then‐execute framework that couples an LLM‐based City Planner with a PCG‐based Implementer. Given a natural language description, the Planner grounds textual intent into a structured city plan by composing PCG parameters via a schema and in‐context exemplars. The Implementer deterministically executes road generation, block extraction, lot subdivision, and asset placement with validity checks and reproducible seeding to synthesize an engine‐ready 3D city. Experiments on multi‐view renderings evaluate text‐scene alignment, diversity, realism, and runtime, demonstrating rapid generation and scalability to large urban scenes.
Xiaohang Dong, Hualong Yu, Xu Zhang 0081, Jianye Wang, Qicheng Li
Comput. Animat. Virtual Worlds5
2026 Superpixel attention-based method for real-world image super-Resolution reconstruction
Qicheng Li, Shuobing Li, Hongxia Deng
Pattern Recognit.2
2026 Corrigendum to "Superpixel Attention-Based Method for Real-World Image Super-Resolution Reconstruction" [Pattern Recognition 177 (2026) 113298]
Qicheng Li, Shuobing Li, Hongxia Deng
Pattern Recognit.2
2025 SDPO: Segment-Level Direct Preference Optimization for Social Agents
abstract
Social agents powered by large language models (LLMs) can simulate human social behaviors but fall short in handling complex social dialogues. Direct Preference Optimization (DPO) has proven effective in aligning LLM behavior with human preferences across various agent tasks. However, standard DPO focuses solely on individual turns, which limits its effectiveness in multi-turn social interactions. Several DPO-based multi-turn alignment methods with session-level data have shown potential in addressing this problem. While these methods consider multiple turns across entire sessions, they are often overly coarse-grained, introducing training noise, and lack robust theoretical support. To resolve these limitations, we propose Segment-Level Direct Preference Optimization (SDPO), which dynamically select key segments within interactions to optimize multi-turn agent behavior. SDPO minimizes training noise and is grounded in a rigorous theoretical framework. Evaluations on the SOTOPIA benchmark demonstrate that SDPO-tuned agents consistently outperform both existing DPO-based methods and proprietary LLMs like GPT-4o, underscoring SDPO’s potential to advance the social intelligence of LLM-based agents. We release our code and data at https://anonymous.4open.science/r/SDPO-CE8F.
Aobo Kong, Shiwan Zhao, Yongbin Li 0001, Yuchuan Wu, Qicheng Li, Fei Huang 0002
ACL (1)8
2025 GUM-DiT: A Foundation Model for Generating Urban Morphological Layouts
Xiaohang Dong, Qicheng Li, Hualong Yu
CGI (1)2
2025 Non-Pharmacological Interventions: A Virtual Training Framework for Fine Motor Learning
abstract
In contemporary rehabilitation treatments, it is common to integrate advanced technologies such as virtual reality and human-computer interaction. However, there is insufficient research on their efficacy in enhancing the fine motor skills of children with developmental coordination disorder (DCD). This paper proposes a virtual motion training system that employs gesture manipulation to improve the hand movement skills of young children. The system utilizes virtual reality devices to provide a more immersive and experiential training method compared to traditional physical rehabilitation. Finally, we recruited 60 volunteers for the study, and the results demonstrated that this training approach had promising effects on the hand motor coordination abilities of children with developmental coordination issues.
Xiaohang Dong, Qicheng Li
ICASSP2
2025 kNN-CL: Enhancing Continual Learning with Nearest Neighbor Retrieval
abstract
Continual learning aims to learn new tasks sequentially without forgetting previously acquired knowledge. However, catastrophic forgetting remains a significant challenge. In this paper, we introduce kNN-CL, a simple yet effective approach that harnesses k-nearest neighbors (kNN) to mitigate forgetting in continual learning. Specifically, kNN-CL identifies the k most similar instances (key-value pairs) from the previous tasks to refine model predictions, enabling the model to adapt to the relevant task for a given test instance. Notably, kNN-CL can be seamlessly integrated into existing continual learning frameworks in a plug-and-play manner, without any additional training. Modern deep neural networks have achieved remarkable progress on a wide range of tasks. However, they encounter difficulties when processing sequential data streams. As these networks recalibrate their parameters to assimilate new data, they inadvertently compromise on previously acquired knowledge, leading to an issue known as catastrophic forgetting. Experimental results show that kNN-CL substantially improves accuracy in both settings, demonstrating the effectiveness of kNN-CL in mitigating catastrophic forgetting.
Enzhi Wang, Qicheng Li
ICASSP2
2025 Enhancing Continual Learning for Medical Imaging: Efficient Knowledge Transfer and Multi-Disease Prediction
abstract
Deep learning models for medical disease detection require extensive labeled data, which is often scarce and expensive. Transfer learning can help by leveraging knowledge from large source domains, but directly fine-tuning these models can lead to catastrophic forgetting, making it impossible to reuse it for new diseases that continue to emerge. Current continual learning methods can mitigate catastrophic forgetting but fail to support efficient knowledge transfer and use mutually exclusive classifiers, which are inadequate for multi-disease prediction in medical imaging. To solve these issues, we propose a Continual Learning Multi-Disease Prediction Framework with class-specific adapters to independently model each disease, supporting multi-disease prediction and feature integration. Additionally, we introduce a two-stage knowledge integration method based on self-attention, which concatenates features from previous tasks and computes attention scores to utilize historical knowledge effectively. Experimental results demonstrate that our framework supports multi-disease prediction and knowledge transfer effectively. They also confirm the effectiveness of our self-attention-based method, providing robust support for medical disease detection.
Enzhi Wang, Qicheng Li
ICASSP2
2025 Self-Prompt Tuning: Enable Autonomous Role-Playing in LLMs
Aobo Kong, Shiwan Zhao, Qicheng Li, Jiaming Zhou 0001, Haoqin Sun
NLPCC (1)4
2024 Better Zero-Shot Reasoning with Role-Play Prompting
abstract
Aobo Kong, Shiwan Zhao, Hao Chen, Qicheng Li, Yong Qin, Ruiqi Sun, Xin Zhou, Enzhi Wang, Xiaohang Dong. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Aobo Kong, Shiwan Zhao, Qicheng Li, Enzhi Wang, Xiaohang Dong
NAACL-HLT4
2023 PromptRank: Unsupervised Keyphrase Extraction Using Prompt
abstract
The keyphrase extraction task refers to the automatic selection of phrases from a given document to summarize its core content.Stateof-the-art (SOTA) performance has recently been achieved by embedding-based algorithms, which rank candidates according to how similar their embeddings are to document embeddings.However, such solutions either struggle with the document and candidate length discrepancies or fail to fully utilize the pretrained language model (PLM) without further fine-tuning.To this end, in this paper, we propose a simple yet effective unsupervised approach, PromptRank, based on the PLM with an encoder-decoder architecture.Specifically, PromptRank feeds the document into the encoder and calculates the probability of generating the candidate with a designed prompt by the decoder.We extensively evaluate the proposed PromptRank on six widely used benchmarks.PromptRank outperforms the SOTA approach MDERank, improving the F 1 score relatively by 34.18%, 24.87%, and 17.57% for 5, 10, and 15 returned results, respectively.This demonstrates the great potential of using prompt for unsupervised keyphrase extraction.We release our code at this url.
Aobo Kong, Shiwan Zhao, Qicheng Li, Xiaoyan Bai
ACL (1)4
2023 Anomaly-Based Insider Threat Detection via Hierarchical Information Fusion
Enzhi Wang, Qicheng Li, Shiwan Zhao, Xue Han 0018
ICANN (3)2
2023 CyFormer: Accurate State-of-Health Prediction of Lithium-Ion Batteries via Cyclic Attention
abstract
Predicting the State-of-Health (SoH) of lithium-ion batteries is a fundamental task of battery management systems on electric vehicles. It aims at estimating future SoH based on historical aging data. Most existing deep learning methods rely on filter-based feature extractors (e.g., CNN or Kalman filters) and recurrent time sequence models. Though efficient, they generally ignore cyclic features and the domain gap between training and testing batteries. To address this problem, we present CyFormer, a transformer-based cyclic time sequence model for SoH prediction. Instead of the conventional CNN-RNN structure, we adopt an encoder-decoder architecture. In the encoder, row-wise and column-wise attention blocks effectively capture intra-cycle and inter-cycle connections and extract cyclic features. In the decoder, the SoH queries cross-attend to these features to form the final predictions. We further utilize a transfer learning strategy to narrow the domain gap between the training and testing set. To be specific, we use fine-tuning to shift the model to a target working condition. Finally, we made our model more efficient by pruning. The experiment shows that our method attains an MAE of 0.75% with only 10% data for fine-tuning on a testing battery, surpassing prior methods by a large margin. Effective and robust, our method provides a potential solution for all cyclic time sequence prediction tasks.
Zhiqiang Nie, Jiankun Zhao, Qicheng Li
IJCNN3
2015 A Service-Based Framework for Mobile Social Messaging in PaaS Systems
abstract
With the rapid growth of cloud computing, developing business applications on Platform-as-a-Service (PaaS) systems is increasingly popular among industry companies. Various services are developed to support different business requirements on PaaS systems. However, to the best of our knowledge, currently there is no service that provides mobile social messaging services to enable users of their apps to share messages in their social networks (e.g., We Chat, Whats App, Kaka Talk). Lack of such mobile social messaging services prevents industry companies from succeeding in drastic market competitions (e.g., Capture high customer satisfactory). In this paper, we propose a service-based framework to enable the mobile social messaging in PaaS systems (e.g., IBM Blue mix). Using this framework, developers can focus on the service encapsulation of existing applications, and define their business process flows via the conversation management in our platform (no coding work is needed). As such, our framework can effectively reduce the development workload for mobile social messaging in PaaS systems.
Lijun Mei, Hao Chen 0021, Shaochun Li, Qicheng Li, Guangtai Liang, Jeaha Yang
ICWS4
2015 Second Order-Based Real-Time Anomaly Detection for Application Maintenance Services
abstract
Application Maintenance Services (AMS) is essential for applications executed on servers to function properly. Its objective is to reduce the application incidents happened and quickly recover services from application failures/issues. The application incidents defined as events when there are some application failures/issues happened are major concerns of AMS, therefore we propose a second order-based anomaly detection method to describe and predict application incidents based on analysis of monitored server traffic metrics. The proposed method first detects anomalies for each metric, second builds the linkage between detected anomalies for all metrics of the server and application incidents, and then predicts potential application incidents. Through the experiments, we find that the presented method provides satisfactory results for identify application incident, which gives more than 90 percentage recall rate while about 65 percentage precision rate.
Qicheng Li, Lijun Mei, Shaochun Li, Liu Rong, Weiye Chen, Fenfei Wang
ICSS1
2011 A Service-Oriented Framework for Hybrid Immersive Web Applications
abstract
Immersive Web (IW for short) applications such as Second Life and SimCity are increasingly popular among individual users and companies. However, it is difficult to build a hybrid Immersive Web application, because of the different architectures used for building such IW applications. In this paper, we propose a service-oriented framework to support the mashup of heterogeneous IW applications for hybrid IW applications. We first model each object in IW (e.g., a virtual album or a virtual room) as an IW Object Service (IWOS) located by a unique URL with standardized service interfaces. As such, one IW application can use an IWOS instance in another IW application. Moreover, non-IW applications (e.g., an online album system) can also access an IWOS instance via its service interface. We further propose a service-oriented framework to manipulate these IWOS instances, to support the transition of these objects mainly between different IW applications, and to enable the transition between IW applications and non-IW applications. Our framework also includes the fundamental services to build hybrid IW applications. Finally, we provide examples to demonstrate the advantages of our framework, and conduct a qualitative analysis to show the effectiveness.
Lijun Mei, Qicheng Li
ICWS3
2008 Network patterns: designing effective user interfaces for connections management at work
abstract
New collaborative technologies are making modern work a highly social process in the era of Web 2.0, yet as a result, it is increasingly challenging for workers to maintain knowledge of and manage their connections. Former netWORKing studies have suggested integrating communication and information with models of social networks. To elucidate these netWORK processes and explore the viability of organizing connections at work in terms of a social network of contacts, we conducted field studies investigating patterns of people's subjective ego-centered netWORKs. Different netWORK categorizing and structuring strategies are identified and user interface design implications are proposed for connections management systems.
Qinying Liao, Qicheng Li
CSCW2
2007 An Efficient Approach to Real-Time Sky Simulation
abstract
Real-time sky rendering plays a great role on visualization in many 3D applications. It is difficult to simulate a realistic sky without considering the effect of clouds and the scattering effect of the atmosphere. In this paper, we use a physical scattering model to simulate the sky, and a Perlin noise texture to create highly realistic clouds. Different effects of the sky are also generated by combining the lighting model with various atmospheric constituents.
Qicheng Li, Zhangjin Huang
CAD/Graphics2