Qingkai Zeng 0001

dblp:66/3005-1 · DBLP profile ↗
← Back
19ranked-venue papers
2as first author
13since 2021 · last 2026
0000-0002-0858-937XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 2 first-author · 10 since 2021Databases, data management, data science and information retrieval · 8 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Instant Personalized Large Language Model Adaptation via Hypernetwork
abstract
Zhaoxuan Tan, Zixuan Zhang, Haoyang Wen, Zheng Li, Rongzhi Zhang, Pei Chen, Fengran Mo, Zheyuan Liu, Qingkai Zeng, Qingyu Yin, Meng Jiang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Zhaoxuan Tan, Haoyang Wen, Zheng Li 0018, Rongzhi Zhang, Fengran Mo, Zheyuan Liu 0010, Qingkai Zeng 0001, Qingyu Yin, Meng Jiang 0001
ACL (1)9
2025 Enhancing Mathematical Reasoning in LLMs by Stepwise Correction
abstract
Best-of-N decoding methods instruct large language models (LLMs) to generate multiple solutions, score each using a scoring function, and select the highest scored as the final answer to mathematical reasoning problems.However, this repeated independent process often leads to the same mistakes, making the selected solution still incorrect.We propose a novel prompting method named Stepwise Correction (STEPCO) that helps LLMs identify and revise incorrect steps in their generated reasoning paths.It iterates verification and revision phases that employ a process-supervised verifier.The verifythen-revise process not only improves answer correctness but also reduces token consumption with fewer paths needed to generate.With STEPCO, a series of LLMs demonstrate exceptional performance.Notably, using GPT-4o as the backend LLM, STEPCO achieves an average accuracy of 94.1 across eight datasets, significantly outperforming the state-of-the-art Best-of-N method by +2.4, while reducing token consumption by 77.8%.Our implementation is made publicly available at https: //wzy6642.github.io/stepco.github.io.
Zhenyu Wu 0004, Qingkai Zeng 0001, Zhihan Zhang 0001, Zhaoxuan Tan, Meng Jiang 0001
ACL (1)2
2025 Protecting Privacy in Multimodal Large Language Models with MLLMU-Bench
abstract
Zheyuan Liu, Guangyao Dou, Mengzhao Jia, Zhaoxuan Tan, Qingkai Zeng, Yongle Yuan, Meng Jiang. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Zheyuan Liu 0010, Guangyao Dou, Mengzhao Jia, Zhaoxuan Tan, Qingkai Zeng 0001, Yongle Yuan, Meng Jiang 0001
NAACL (Long Papers)5
2024 Chain-of-Layer: Iteratively Prompting Large Language Models for Taxonomy Induction from Limited Examples
abstract
Automatic taxonomy induction is crucial for web search, recommendation systems, and question answering. Manual curation of taxonomies is expensive in terms of human effort, making automatic taxonomy construction highly desirable. In this work, we introduce Chain-of-Layer which is an in-context learning framework designed to induct taxonomies from a given set of entities. Chain-of-Layer breaks down the task into selecting relevant candidate entities in each layer and gradually building the taxonomy from top to bottom. To minimize errors, we introduce the Ensemble-based Ranking Filter to reduce the hallucinated content generated at each iteration. Through extensive experiments, we demonstrate that Chain-of-Layer achieves state-of-the-art performance on four real-world benchmarks. Source code available at: https://github.com/qingkaizeng/chain-of-layer.
Qingkai Zeng 0001, Yuyang Bai, Zhaoxuan Tan, Shangbin Feng, Zhenwen Liang, Zhihan Zhang 0001, Meng Jiang 0001
CIKM1
2024 ChatEL: Entity Linking with Chatbots
abstract
Entity Linking (EL) is an essential and challenging task in natural language processing that seeks to link some text representing an entity within a document or sentence with its corresponding entry in a dictionary or knowledge base. Most existing approaches focus on creating elaborate contextual models that look for clues the words surrounding the entity-text to help solve the linking problem. Although these fine-tuned language models tend to work, they can be unwieldy, difficult to train, and do not transfer well to other domains. Fortunately, Large Language Models (LLMs) like GPT provide a highly-advanced solution to the problems inherent in EL models, but simply naive prompts to LLMs do not work well. In the present work, we define ChatEL, which is a three-step framework to prompt LLMs to return accurate results. Overall the ChatEL framework improves the average F1 performance across 10 datasets by more than 2%. Finally, a thorough error analysis shows many instances with the ground truth labels were actually incorrect, and the labels predicted by ChatEL were actually correct. This indicates that the quantitative results presented in this paper may be a conservative estimate of the actual performance. All data and code are available as an open-source package on GitHub at https://github.com/yifding/In_Context_EL.
Yifan Ding 0001, Qingkai Zeng 0001, Tim Weninger
LREC/COLING2
2024 Large Language Models Can Self-Correct with Key Condition Verification
abstract
Intrinsic self-correct was a method that instructed large language models (LLMs) to verify and correct their responses without external feedback.Unfortunately, the study concluded that the LLMs could not self-correct reasoning yet.We find that a simple yet effective prompting method enhances LLM performance in identifying and correcting inaccurate answers without external feedback.That is to mask a key condition in the question, add the current response to construct a verification question, and predict the condition to verify the response.The condition can be an entity in an open-domain question or a numerical value in an arithmetic question, which requires minimal effort (via prompting) to identify.We propose an iterative verify-then-correct framework to progressively identify and correct (probably) false responses, named PROCO.We conduct experiments on three reasoning tasks.On average, PROCO, with GPT-3.5-Turbo-1106 as the backend LLM, yields +6.8 exact match on four open-domain question answering datasets, +14.1 accuracy on three arithmetic reasoning datasets, and +9.6 accuracy on a commonsense reasoning dataset, compared to Self-Correct.Our implementation is made publicly avail
Zhenyu Wu 0004, Qingkai Zeng 0001, Zhihan Zhang 0001, Zhaoxuan Tan, Meng Jiang 0001
EMNLP2
2024 Democratizing Large Language Models via Personalized Parameter-Efficient Fine-tuning
abstract
Personalization in large language models (LLMs) is increasingly important, aiming to align the LLMs' interactions, content, and recommendations with individual user preferences.Recent advances have highlighted effective prompt design by enriching user queries with non-parametric knowledge through behavior history retrieval and textual profiles.However, these methods faced limitations due to a lack of model ownership, resulting in constrained customization and privacy issues, and often failed to capture complex, dynamic user behavior patterns.To address these shortcomings, we introduce One PEFT Per User (OPPU) 1 , employing personalized parameter-efficient finetuning (PEFT) modules to store user-specific behavior patterns and preferences.By plugging in personal PEFT parameters, users can own and use their LLMs individually.OPPU integrates parametric user knowledge in the personal PEFT parameters with non-parametric knowledge from retrieval and profiles, adapting LLMs to user behavior shifts.Experimental results demonstrate that OPPU significantly outperforms existing prompt-based methods across seven diverse tasks in the LaMP benchmark.Further studies reveal OPPU's enhanced capabilities in handling user behavior shifts, modeling users at different activity levels, maintaining robustness across various user history formats, and displaying versatility with different PEFT methods.
Zhaoxuan Tan, Qingkai Zeng 0001, Yijun Tian 0001, Zheyuan Liu 0010, Meng Jiang 0001
EMNLP2
2022 Automatic Controllable Product Copywriting for E-Commerce
abstract
Automatic product description generation for e-commerce has witnessed significant advancement in the past decade. Product copy- writing aims to attract users' interest and improve user experience by highlighting product characteristics with textual descriptions. As the services provided by e-commerce platforms become diverse, it is necessary to adapt the patterns of automatically-generated descriptions dynamically. In this paper, we report our experience in deploying an E-commerce Prefix-based Controllable Copywriting Generation (EPCCG) system into the JD.com e-commerce product recommendation platform. The development of the system contains two main components: 1) copywriting aspect extraction; 2) weakly supervised aspect labelling; 3) text generation with a prefix-based language model; and 4) copywriting quality control. We conduct experiments to validate the effectiveness of the proposed EPCCG. In addition, we introduce the deployed architecture which cooperates the EPCCG into the real-time JD.com e-commerce recommendation platform and the significant payoff since deployment. The codes for implementation are provided at https://github.com/xguo7/Automatic-Controllable-Product-Copywriting-for-E-Commerce.git.
Xiaojie Guo 0002, Qingkai Zeng 0001, Meng Jiang 0001, Bo Long, Lingfei Wu 0001
KDD2
2021 Traceability Transformed: Generating more Accurate Links with Pre-Trained BERT Models
abstract
Software traceability establishes and leverages associations between diverse development artifacts. Researchers have proposed the use of deep learning trace models to link natural language artifacts, such as requirements and issue descriptions, to source code; however, their effectiveness has been restricted by availability of labeled data and efficiency at runtime. In this study, we propose a novel framework called Trace BERT (T-BERT) to generate trace links between source code and natural language artifacts. To address data sparsity, we leverage a three-step training strategy to enable trace models to transfer knowledge from a closely related Software Engineering challenge, which has a rich dataset, to produce trace links with much higher accuracy than has previously been achieved. We then apply the T-BERT framework to recover links between issues and commits in Open Source Projects. We comparatively evaluated accuracy and efficiency of three BERT architectures. Results show that a Single-BERT architecture generated the most accurate links, while a Siamese-BERT architecture produced comparable results with significantly less execution time. Furthermore, by learning and transferring knowledge, all three models in the framework outperform classical IR trace models. On the three evaluated real-word OSS projects, the best T-BERT stably outperformed the VSM model with average improvements of 60.31% measured using Mean Average Precision (MAP). RNN severely underperformed on these projects due to insufficient training data, while T-BERT overcame this problem by using pretrained language models and transfer learning.
Jinfeng Lin, Yalin Liu, Qingkai Zeng 0001, Meng Jiang 0001, Jane Cleland-Huang
ICSE3
2021 Enhancing Taxonomy Completion with Concept Generation via Fusing Relational Representations
abstract
Automatic construction of a taxonomy supports many applications in e-commerce, web search, and question answering. Existing taxonomy expansion or completion methods assume that new concepts have been accurately extracted and their embedding vectors learned from the text corpus. However, one critical and fundamental challenge in fixing the incompleteness of taxonomies is the incompleteness of the extracted concepts, especially for those whose names have multiple words and consequently low frequency in the corpus. To resolve the limitations of extraction-based methods, we propose GenTaxo to enhance taxonomy completion by identifying positions in existing taxonomies that need new concepts and then generating appropriate concept names. Instead of relying on the corpus for concept embeddings, GenTaxo learns the contextual embeddings from their surrounding graph-based and language-based relational information, and leverages the corpus for pre-training a concept name generator. Experimental results demonstrate that GenTaxo improves the completeness of taxonomies over existing methods.
Qingkai Zeng 0001, Jinfeng Lin, Wenhao Yu 0002, Jane Cleland-Huang, Meng Jiang 0001
KDD1
2021 Enhancing Factual Consistency of Abstractive Summarization
abstract
Chenguang Zhu, William Hinthorn, Ruochen Xu, Qingkai Zeng, Michael Zeng, Xuedong Huang, Meng Jiang. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Chenguang Zhu 0001, William Hinthorn, Ruochen Xu, Qingkai Zeng 0001, Michael Zeng 0001, Xuedong Huang 0001, Meng Jiang 0001
NAACL-HLT4
2021 Biomedical Knowledge Graphs Construction From Conditional Statements
abstract
Conditions play an essential role in biomedical statements. However, existing biomedical knowledge graphs (BioKGs) only focus on factual knowledge, organized as a flat relational network of biomedical concepts. These BioKGs ignore the conditions of the facts being valid, which loses essential contexts for knowledge exploration and inference. We consider both facts and their conditions in biomedical statements and proposed a three-layered information-lossless representation of BioKG. The first layer has biomedical concept nodes, attribute nodes. The second layer represents both biomedical fact and condition tuples by nodes of the relation phrases, connecting to the subject and object in the first layer. The third layer has nodes of statements connecting to a set of fact tuples and/or condition tuples in the second layer. We transform the BioKG construction problem into a sequence labeling problem based on a novel designed tag schema. We design a Multi-Input Multi-Output sequence labeling model (MIMO) that learns from multiple input signals and generates proper number of multiple output sequences for tuple extraction. Experiments on a newly constructed dataset show that MIMO outperforms the existing methods. Further case study demonstrates that the BioKGs constructed provide a good understanding of the biomedical statements.
Tianwen Jiang, Qingkai Zeng 0001, Tong Zhao 0003, Bing Qin 0001, Ting Liu 0001, Nitesh V. Chawla, Meng Jiang 0001
IEEE ACM Trans. Comput. Biol. Bioinform.2
2021 Modeling Complementarity in Behavior Data with Multi-Type Itemset Embedding
abstract
People are looking for complementary contexts, such as team members of complementary skills for project team building and/or reading materials of complementary knowledge for effective student learning, to make their behaviors more likely to be successful. Complementarity has been revealed by behavioral sciences as one of the most important factors in decision making. Existing computational models that learn low-dimensional context representations from behavior data have poor scalability and recent network embedding methods only focus on preserving the similarity between the contexts. In this work, we formulate a behavior entry as a set of context items and propose a novel representation learning method, Multi-type Itemset Embedding , to learn the context representations preserving the itemset structures. We propose a measurement of complementarity between context items in the embedding space. Experiments demonstrate both effectiveness and efficiency of the proposed method over the state-of-the-art methods on behavior prediction and context recommendation. We discover that the complementary contexts and similar contexts are significantly different in human behaviors.
Daheng Wang, Qingkai Zeng 0001, Nitesh V. Chawla, Meng Jiang 0001
ACM Trans. Intell. Syst. Technol.2
2020 Crossing Variational Autoencoders for Answer Retrieval
abstract
Answer retrieval is to find the most aligned answer from a large set of candidates given a question.Learning vector representations of questions/answers is the key factor.Questionanswer alignment and question/answer semantics are two important signals for learning the representations.Existing methods learned semantic representations with dual encoders or dual variational auto-encoders.The semantic information was learned from language models or question-to-question (answer-to-answer) generative processes.However, the alignment and semantics were too separate to capture the aligned semantics between question and answer.In this work, we propose to cross variational auto-encoders by generating questions with aligned answers and generating answers with aligned questions.Experiments show that our method outperforms the state-of-theart answer retrieval method on SQuAD.Question Answer Decoder 𝑝(𝑞|𝒛 𝒂 ) 𝑝(𝑎|𝒛 𝒒 ) 𝑝(𝑦|𝑧 !, 𝑧 " ) 𝑝(𝑦|𝑧 !, 𝑧 " ) Question Answer Question Answer Decoder Encoder Encoder Decoder Decoder Encoder 𝑝(𝑧 !|𝑞) 𝑝(𝑧 " |𝑎) Encoder Encoder (a) Dual-Encoders (Yang et al., 2019)Question Answer Decoder 𝑝(𝑞|𝒛 𝒂 ) 𝑝(𝑎|𝒛 𝒒 ) 𝑝(𝑦|𝑧 !, 𝑧 " ) 𝑝(𝑦|𝑧 !, 𝑧 " ) Question Answer Question Answer Decoder Encoder Encoder Decoder Decoder Encoder 𝑝(𝑧 !|𝑞) 𝑝(𝑧 " |𝑎) Encoder Encoder (b) Dual-VAEs (Shen et al., 2018) 𝑧 !~𝑝(𝑧 ! ) 𝑧 " ~𝑝(𝑧 " ) Question Answer 𝑧 !~𝑝(𝑧 ! ) 𝑝(𝑧 !|𝑞) 𝑧 " ~𝑝(𝑧 " ) 𝑝(𝑧 " |𝑎) Question Answer 𝑝(𝑦|𝑧 !, 𝑧 " ) 𝑝(𝑞|𝑧 ! ) 𝑝(𝑎|𝑧 " )
Wenhao Yu 0002, Lingfei Wu 0001, Qingkai Zeng 0001, Shu Tao, Yu Deng 0004, Meng Jiang 0001
ACL3
2020 Towards Semantically Guided Traceability
abstract
In many regulated domains, traceability is established across diverse artifacts such as requirements, design, code, test cases, and hazards - either manually or with the help of supporting tools, and the resulting trace links are used to support activities such as impact analysis, compliance verification, and safety inspections. Automated tracing techniques need to leverage the semantics of underlying artifacts in order to establish more accurate trace links and to provide explanations of links that have been created in either a manual or automated fashion. To support this, we propose an automated technique which leverages source code, project artifacts and an external domain corpus to generate a domain-specific concept model. We then use the generated concept model to improve traceability results and to provide explanations of the results. Our approach overcomes existing problems with deep-learning traceability algorithms, as it does not require a training set of existing trace links. Finally, as an initial proof-of-concept, we apply our semantically-guided approach to the Dronology project, and show that it improves over other tracing techniques that do not use a concept model.
Yalin Liu, Jinfeng Lin, Qingkai Zeng 0001, Meng Jiang 0001, Jane Cleland-Huang
RE3
2020 Experimental Evidence Extraction System in Data Science with Hybrid Table Features and Ensemble Learning
abstract
Data Science has been one of the most popular fields in higher education and research activities. It takes tons of time to read the experimental section of thousands of papers and figure out the performance of the data science techniques. In this work, we build an experimental evidence extraction system to automate the integration of tables (in the paper PDFs) into a database of experimental results. First, it crops the tables and recognizes the templates. Second, it classifies the column names and row names into “method”, “dataset”, or “evaluation metric”, and then unified all the table cells into (method, dataset, metric, score)-quadruples. We propose hybrid features including structural and semantic table features as well as an ensemble learning approach for column/row name classification and table unification. SQL statements can be used to answer questions such as whether a method is the state-of-the-art or whether the reported numbers are conflicting.
Wenhao Yu 0002, Yu Shu, Qingkai Zeng 0001, Meng Jiang 0001
WWW4
2019 Through the eyes of a poet: classical poetry recommendation with visual input on social media
abstract
With the increasing popularity of portable devices with cameras (e.g., smartphones and tablets) and ubiquitous Internet connectivity, travelers can share their instant experience during the travel by posting photos they took to social media platforms. In this paper, we present a new image-driven poetry recommender system that takes a traveler's photo as input and recommends classical poems that can enrich the photo with aesthetically pleasing quotes from the poems. Three critical challenges exist to solve this new problem: i) how to extract the implicit artistic conception embedded in both poems and images? ii) How to identify the salient objects in the image without knowing the creator's intent? iii) How to accommodate the diverse user perceptions of the image and make a diversified poetry recommendation? The proposed iPoemRec system jointly addresses the above challenges by developing heterogeneous information network and neural embedding techniques. Evaluation results from real-world datasets and a user study demonstrate that our system can recommend highly relevant classical poems for a given photo and receive significantly higher user ratings compared to the state-of-the-art baselines.
Daniel Yue Zhang, Bo Ni, Qiyu Zhi, Thomas Plummer, Qi Li 0016, Hao Zheng 0006, Qingkai Zeng 0001, Yang Zhang 0031, Dong Wang 0002
ASONAM7
2019 Tablepedia: Automating PDF Table Reading in an Experimental Evidence Exploration and Analytic System
abstract
Web research, data science, and artificial intelligence have been rapidly changing our life and society. Researchers and practitioners in the fields take a large amount of time to read literature and compare existing approaches. It would significantly improve their efficiency if there was a system that extracted and managed experimental evidences (say, a specific method achieves a score of a specific metric on a specific dataset) from tables of paper PDFs for search, exploration, and analytic. We build such a demonstration system, called Tablepedia, that use rule-based and learning-based methods to automate the “reading” of PDF tables. It has three modules: template recognition, unification, and SQL operations. We implement three functions to facilitate research and practice: (1) finding related methods and datasets, (2) finding top-performing baseline methods, and (3) finding conflicting reported numbers. A pointer to a screencast on Vimeo: https://vimeo.com/310162310
Wenhao Yu 0002, Qingkai Zeng 0001, Meng Jiang 0001
WWW3
2018 Multi-Type Itemset Embedding for Learning Behavior Success
abstract
Contextual behavior modeling uses data from multiple contexts to discover patterns for predictive analysis. However, existing behavior prediction models often face difficulties when scaling for massive datasets. In this work, we formulate a behavior as a set of context items of different types (such as decision makers, operators, goals and resources), consider an observable itemset as a behavior success, and propose a novel scalable method, "multi-type itemset embedding", to learn the context items' representations preserving the success structures. Unlike most of existing embedding methods that learn pair-wise proximity from connection between a behavior and one of its items, our method learns item embeddings collectively from interaction among all multi-type items of a behavior, based on which we develop a novel framework, LearnSuc, for (1) predicting the success rate of any set of items and (2) finding complementary items which maximize the probability of success when incorporated into an itemset. Extensive experiments demonstrate both effectiveness and efficency of the proposed framework.
Daheng Wang, Meng Jiang 0001, Qingkai Zeng 0001, Zachary Eberhart, Nitesh V. Chawla
KDD3