EDBT 2026 Demo / reviewers in the wild / expert
Qingcheng Zeng
dblp:84/1451
· DBLP profile ↗
26ranked-venue papers
6as first author
26since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 20 · 6 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | On the Effect of Hyperparameters in Language Modeling for Computational LinguisticsabstractTraining language models and examining their linguistic behaviors have been a common protocol in computational linguistics for studying linguistic phenomena and modeling human language processing.However, work in this area is often limited to proof-of-concept demonstrations with arbitrary model configurations, without considering hyperparameter sensitivity, an important source of variation in model performance.In this work, we replicate three prior studies (Chang and Bergen, 2022; Hu et al., 2020b;Kuribayashi et al., 2024) with hyperparameters varied within a practical range, and show that modest hyperparameter changes can alter some qualitative conclusions about models' linguistic abilities and even reverse the ranking of model performance.Our results highlight the risk that prior work may have reflected optimization artifacts rather than the genuine inductive biases of model classes, and that hyperparameter sensitivity should receive more attention as a factor that can meaningfully influence model behavior.We suggest future work to report the variation of performance across the configuration space to enhance the reliability and generalizability of conclusions. Ruoxi Ning, Yongpeng Zhu, Qingcheng Zeng, Tatsuki Kuribayashi, Freda Shi |
ACL (1) | 3 |
| 2026 | How to Improve LLMs' Performance on Specific Languages: A Perspective on LLM-Derived Language SimilarityabstractLarge language models (LLMs) exhibit uneven performance across languages. In language-specific applications, practitioners often rely on target-language corpora or cross-lingual transfer to achieve better performance. However, traditional linguistic typology, commonly used as a transfer language selection strategy in previous studies, may not align with LLM’s perception of language similarity. This work proposes LLM-based language similarity as a novel perspective for selecting effective fine-tuning languages. We construct a framework to quantify the similarity within each language pair through both the lenses of language-specific performance patterns and cross-lingual transferability, ultimately deriving three similarity score matrices. Moreover, we observe a counter-intuitive phenomenon: super-additive transfer effect, where fine-tuning on a certain language yields higher performance than fine-tuning directly on the target language. Additionally, due to the absence of an existing dataset meeting our experimental requirements, we construct and release M4CQ-Pro dataset, which features domain-diverse distribution of 135 tasks and content consistency across 31 languages (including over 20 medium- and low-resource languages), with 61518 manually reviewed high-quality questions per language. We evaluate our approach on representative multilingual LLMs and results show that all three LLM-based similarity measures effectively guide fine-tuning language selection, outperforming traditional linguistic similarity, with the integrated measure achieving the best results. Our approach provides not only a novel perspective on language similarity, but also practical baselines for selecting fine-tuning languages. Xinhe Shi, Qingcheng Zeng, Weihao Xuan, Linchao Zhu |
ACL (1) | 2 |
| 2026 | The Confidence Dichotomy: Analyzing and Mitigating Miscalibration in Tool-Use AgentsabstractAutonomous agents based on large language models (LLMs) are rapidly evolving to handle multi-turn tasks, but ensuring their trustworthiness remains a critical challenge.A fundamental pillar of this trustworthiness is calibration, which refers to an agent's ability to express confidence that reliably reflects its actual performance.While calibration is well-established for static models, its dynamics in tool-integrated agentic workflows remain under-explored.In this work, we systematically investigate verbalized calibration in tooluse agents, revealing a fundamental confidence dichotomy driven by tool type.Specifically, our pilot study identifies that evidence tools (e.g., web search) systematically induce severe overconfidence due to inherent noise in retrieved information, while verification tools (e.g., code interpreters) can ground reasoning through deterministic feedback and mitigate miscalibration.To robustly improve calibration across tool types, we propose a reinforcement learning (RL) fine-tuning framework that jointly optimizes task accuracy and calibration, supported by a holistic benchmark of reward designs.We demonstrate that our trained agents not only achieve superior calibration but also exhibit robust generalization from local training environments to noisy web settings and to distinct domains such as mathematical reasoning.Our results highlight the necessity of domain-specific calibration strategies for tooluse agents.More broadly, this work establishes a foundation for building self-aware agents that can reliably communicate uncertainty in highstakes, real-world deployments. Weihao Xuan, Qingcheng Zeng, Heli Qi, Yunze Xiao, Naoto Yokoya |
ACL (1) | 2 |
| 2026 | Integrated AGV Positioning and Scheduling Using Simulation-Based Reinforcement Learning and Combinatorial Optimization in Automated Container TerminalsabstractAutomated Guided Vehicle (AGV) positioning and scheduling critically impact automated container terminal efficiency, yet operational uncertainties significantly challenge effective decision-making. This study addresses the integrated AGV positioning and scheduling problem (AGV-PSP) under uncertainty through a novel framework integrating reinforcement learning (RL), combinatorial optimization, and high-fidelity simulation. We formulate the AGV-PSP as a Markov Decision Process where RL learns the state value function, estimating the long-term operational value of terminal locations. Based on the value function, our positioning policy strategically locates AGVs in high-value areas, while our scheduling policy balances immediate and long-term rewards using integer linear programming solved by an improved Hungarian algorithm. Experiments demonstrate 35.6% higher rewards and 36.6% lower costs compared to benchmarks on average. Sensitivity analyses validate the framework’s robustness across varying uncertainty. Qingcheng Zeng, Xingchun Li, Haobin Li |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2025 | Fact or Facsimile? Evaluating the Factual Robustness of Modern RetrieversabstractDense retrievers and rerankers are central to retrieval-augmented generation (RAG) pipelines, where accurately retrieving factual information is crucial for maintaining system trustworthiness and defending against RAG poisoning.However, little is known about how much factual competence these components inherit or lose from the large language models (LLMs) they are based on.We pair 12 publicly released embedding checkpoints with their original base LLMs and evaluate both sets on a factuality benchmark.Across every model evaluated, the embedding variants achieve markedly lower accuracy than their bases, with absolute drops ranging from 12 to 43 percentage points (median 28 pts) and typical retriever accuracies collapsing into the 25-35 % band versus the 60-70 % attained by the generative models.This degradation intensifies under a more demanding condition: when the candidate pool per question is expanded from four options to one thousand, the strongest retriever's top-1 accuracy falls from 33 % to 26 %, revealing acute sensitivity to distractor volume.Statistical tests further show that, for every embedding model, cosine-similarity scores between queries and correct completions are significantly higher than those for incorrect ones (𝑝 < 0.01), indicating decisions driven largely by surface-level semantic proximity rather than factual reasoning.To probe this weakness, we employed GPT-4.1 to paraphrase each correct completion, creating a rewritten test set that preserved factual truth while masking lexical cues, and observed that over two-thirds of previously correct predictions flipped to wrong, reducing overall accuracy to roughly one-third of its original level.Taken together, these findings reveal a systematic trade-off introduced by contrastive learning for retrievers: gains in semantic retrieval are paid for with losses in parametric factual knowledge, and the resulting models remain highly vulnerable to adversarial or even benign rephrasings.Our study underscores the need for retrieval objectives that balance similarity with factual fidelity to safeguard next-generation RAG systems against both misinformation and targeted attacks. Qingcheng Zeng, Kaize Ding |
CIKM | 2 |
| 2025 | Uncertainty Quantification for Multiple-Choice Questions is Just One-Token DeepabstractMultiple-choice question (MCQ) benchmarks such as MMLU and GPQA are widely used to assess the capabilities of large language models (LLMs). While accuracy remains the standard evaluation metric, recent work has introduced uncertainty quantification (UQ) methods, such as entropy, conformal prediction, and verbalized confidence, as complementary measures of model reliability and calibration. However, we find that these UQ methods, when applied to MCQ tasks, are unexpectedly fragile. Specifically, we show that fine-tuning a model on just 1,000 examples to adjust the probability of the first generated token, under the common prompting setup where the model is instructed to output only a single answer choice, can systematically distort a broad range of UQ methods across models, prompts, and domains, all while leaving answer accuracy unchanged. We validate this phenomenon through extensive experiments on five instruction-tuned LLMs, tested under standard prompting, zero-shot chain-of-thought reasoning, and a biomedical question answering setting. In all cases, models retain similar accuracy but exhibit significantly degraded calibration. These results suggest that current UQ practices for MCQs are ''one-token deep'', driven more by first-token decoding behavior than by any deeper representation of uncertainty, and are easily manipulated through minimal interventions. Our findings call for more robust and interpretable approaches to uncertainty estimation, particularly in structured formats like MCQs, where confidence signals are often reduced to token-level heuristics. Qingcheng Zeng, Mingyu Jin, Qinkai Yu, Zhenting Wang, Wenyue Hua, Guangyan Sun, Yanda Meng, Shiqing Ma, Qifan Wang 0001, Felix Juefei-Xu, Fan Yang 0023, Kaize Ding, Ruixiang Tang, Yongfeng Zhang 0003 |
CIKM | 1 |
| 2025 | Exploring Concept Depth: How Large Language Models Acquire Knowledge and Concept at Different Layers?abstractLarge language models (LLMs) have shown remarkable performances across a wide range of tasks. However, the mechanisms by which these models encode tasks of varying complexities remain poorly understood. In this paper, we explore the hypothesis that LLMs process concepts of varying complexities in different layers, introducing the idea of “Concept Depth” to suggest that more complex concepts are typically acquired in deeper layers. Specifically, we categorize concepts based on their level of abstraction, defining them in the order of increasing complexity within factual, emotional, and inferential tasks. We conduct extensive probing experiments using layer-wise representations across various LLM families (Gemma, LLaMA, Qwen) on various datasets spanning the three domains of tasks. Our findings reveal that models could efficiently conduct probing for simpler tasks in shallow layers, and more complex tasks typically necessitate deeper layers for accurate understanding. Additionally, we examine how external factors, such as adding noise to the input and quantizing the model weights, might affect layer-wise representations. Our findings suggest that these factors can impede the development of a conceptual understanding of LLMs until deeper layers are explored. We hope that our proposed concept and experimental insights will enhance the understanding of the mechanisms underlying LLMs. Our codes are available at https://github.com/Luckfort/CD. Mingyu Jin, Qinkai Yu, Qingcheng Zeng, Zhenting Wang, Wenyue Hua, Haiyan Zhao 0003, Kai Mei, Yanda Meng, Kaize Ding, Fan Yang 0023, Mengnan Du, Yongfeng Zhang 0003 |
COLING | 4 |
| 2025 | Good Intentions Beyond ACL: Who Does NLP for Social Good, and Where?abstractThe social impact of Natural Language Processing (NLP) is increasingly important, with a rising community focus on initiatives related to NLP for Social Good (NLP4SG).Indeed, in recent years, almost 20% of all papers in the ACL Anthology address topics related to social good as defined by the UN Sustainable Development Goals (Adauto et al., 2023).In this study, we take an author-and venue-level perspective to map the landscape of NLP4SG, quantifying the proportion of work addressing social good concerns both within and beyond the ACL community, by both core ACL contributors and non-ACL authors.With this approach we discover two surprising facts about the landscape of NLP4SG.First, ACL authors are dramatically more likely to do work addressing social good concerns when publishing in venues outside of ACL.Second, the vast majority of publications using NLP techniques to address concerns of social good are done by non-ACL authors in venues outside of ACL.We discuss the implications of these findings on agendasetting considerations for the ACL community related to NLP4SG. Grace LeFevre, Qingcheng Zeng, Adam Leif, Jason Jewell, Denis Peskoff, Rob Voigt |
EMNLP | 2 |
| 2025 | MMLU-ProX: A Multilingual Benchmark for Advanced Large Language Model EvaluationabstractWeihao Xuan, Rui Yang, Heli Qi, Qingcheng Zeng, Yunze Xiao, Aosong Feng, Dairui Liu, Yun Xing, Junjue Wang, Fan Gao, Jinghui Lu, Yuang Jiang, Huitao Li, Xin Li, Kunyu Yu, Ruihai Dong, Shangding Gu, Yuekang Li, Xiaofei Xie, Felix Juefei-Xu, Foutse Khomh, Osamu Yoshie, Qingyu Chen, Douglas Teodoro, Nan Liu, Randy Goebel, Lei Ma, Edison Marrese-Taylor, Shijian Lu, Yusuke Iwasawa, Yutaka Matsuo, Irene Li. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Weihao Xuan, Rui Yang 0016, Heli Qi, Qingcheng Zeng, Yunze Xiao, Aosong Feng, Dairui Liu, Yun Xing 0001, Jinghui Lu, Yuang Jiang, Huitao Li, Xin Li 0079, Kunyu Yu, Ruihai Dong, Shangding Gu, Yuekang Li, Xiaofei Xie, Felix Juefei-Xu, Foutse Khomh, Osamu Yoshie, Qingyu Chen 0001, Douglas Teodoro, Nan Liu 0003, Randy Goebel, Lei Ma 0003, Edison Marrese-Taylor, Shijian Lu, Yusuke Iwasawa, Yutaka Matsuo, Irene Li |
EMNLP | 4 |
| 2025 | Seeing is Believing, but How Much? A Comprehensive Analysis of Verbalized Calibration in Vision-Language ModelsabstractUncertainty quantification is essential for assessing the reliability and trustworthiness of modern AI systems.Among existing approaches, verbalized uncertainty, where models express their confidence through natural language, has emerged as a lightweight and interpretable solution in large language models (LLMs).However, its effectiveness in vision-language models (VLMs) remains insufficiently studied.In this work, we conduct a comprehensive evaluation of verbalized confidence in VLMs, spanning three model categories, four task domains, and three evaluation scenarios.Our results show that current VLMs often display notable miscalibration across diverse tasks and settings.Notably, visual reasoning models (i.e., thinking with images) consistently exhibit better calibration, suggesting that modality-specific reasoning is critical for reliable uncertainty estimation.To further address calibration challenges, we introduce VI-SUAL CONFIDENCE-AWARE PROMPTING, a two-stage prompting strategy that improves confidence alignment in multimodal settings.Overall, our study highlights the inherent miscalibration in VLMs across modalities.More broadly, our findings underscore the fundamental importance of modality alignment and model faithfulness in advancing reliable multimodal systems.Mirko Borszukovszki, Ivo Pascal De Jong, and Matias Valdenegro-Toro.2025.Know what you do not know: Verbalized uncertainty estimation robustness on corrupted images in vision-language models.In Weihao Xuan, Qingcheng Zeng, Heli Qi, Naoto Yokoya |
EMNLP | 2 |
| 2025 | Thinking Out Loud: Do Reasoning Models Know When They're Right?abstractLarge reasoning models (LRMs) have recently demonstrated impressive capabilities in complex reasoning tasks by leveraging increased test-time computation and exhibiting behaviors reminiscent of human-like self-reflection.While LRMs show a clear capacity for valuable self-reflection, how this ability interacts with other model behaviors remains underexplored.We investigate this connection by analyzing verbalized confidence, how models articulate their certainty, as a lens into the nature of self-reflection in LRMs.We find that supervised fine-tuning on reasoning traces (i.e., distillation) and reinforcement learning can improve verbalized calibration in reasoningintensive settings in a progressive, laddered fashion.However, our results also indicate that reasoning models may possess a diminished awareness of their own knowledge boundaries, as evidenced by significantly lower "I don't know" response rates on factuality benchmarks.Moreover, we examine the relationship between verbalized confidence and reasoning chains, finding that models tend to express higher confidence when providing shorter or less elaborate reasoning.Our findings highlight how reasoning-oriented training can enhance performance in reasoning-centric tasks while potentially incurring a reasoning tax, a cost reflected in the model's reduced ability to accurately recognize the limits of its own knowledge in small-scale models.More broadly, our work showcases how this erosion of knowledge boundaries can compromise model faithfulness, as models grow more confident without a commensurate understanding of when they should abstain. Qingcheng Zeng, Weihao Xuan, Leyang Cui, Rob Voigt |
EMNLP | 1 |
| 2025 | Distributed Invariant Kalman Filter for Object-Level Multi-Robot Pose SLAMabstractCooperative localization and target tracking are essential for multi-robot systems to implement high-level tasks. To this end, we propose a distributed invariant Kalman filter (KF) based on covariance intersection (CI) for effective multi-robot pose estimation. The paper utilizes the object-level measurement models, which have condensed information further reducing the communication burden. Besides, by modeling states on special Lie groups, and representing uncertainty in corresponding Lie algebras, better linearity and consistency are obtained under the invariant KF framework. We also combine CI and invariant KF to avoid overly confident or conservative estimates in multi-robot systems with intricate and unknown correlations, and some level of robot degradation is acceptable through multi-robot collaboration. The simulation and real data experiment validate the practicability and superiority of the proposed algorithm. The source code is publicly available11https://github.com/LIAS-CUHKSZ/Distributed-object-based-SLAM. Haoying Li, Qingcheng Zeng, Yanglin Zhang, Junfeng Wu 0001 |
ICRA | 2 |
| 2025 | SCORE: Saturated Consensus Relocalization in Semantic Line MapsabstractWe present SCORE, a visual relocalization system that achieves unprecedented map compactness through semantically labeled 3D line maps. SCORE requires only 0.01%-0.1% of the storage needed by structure-based or learning-based baselines, while maintaining practical accuracy and comparable runtime. The key innovation is a novel robust mechanism, Saturated Consensus Maximization (Sat-CM), which generalizes classical Consensus Maximization (CM) by assigning diminishing weights to inlier associations with probabilistic justification. Under extreme outlier ratios (up to 99.5%) arising from one-to-many ambiguity in semantic matching, Sat-CM enables accurate estimation when CM fails. To ensure computational efficiency, we propose an accelerating framework for globally solving Sat-CM formulations and specialize it for the Perspective-n-Lines problem at the core of SCORE. Haodong Jiang, Yanglin Zhang, Qingcheng Zeng, Yiqian Li, Ziyang Hong 0001, Junfeng Wu 0001 |
IROS | 4 |
| 2025 | CARES: Comprehensive Evaluation of Safety and Adversarial Robustness in Medical LLMsabstractLarge language models (LLMs) are increasingly deployed in medical contexts, raising critical concerns about safety, alignment, and susceptibility to adversarial manipulation. While prior benchmarks assess model refusal capabilities for harmful prompts, they often lack clinical specificity, graded harmfulness levels, and coverage of jailbreak-style attacks. We introduce CARES (Clinical Adversarial Robustness and Evaluation of Safety), a benchmark for evaluating LLM safety in healthcare. CARES includes over 18,000 prompts spanning eight medical safety principles, four harm levels, and four prompting styles—direct, indirect, obfuscated, and role-play—to simulate both malicious and benign use cases. We propose a three-way response evaluation protocol (Accept, Caution, Refuse) and a fine-grained Safety Score metric to assess model behavior. Our analysis reveals that many state-of-the-art LLMs remain vulnerable to jailbreaks that subtly rephrase harmful prompts, while also over-refusing safe but atypically phrased queries. Finally, we propose a mitigation strategy using a lightweight classifier to detect jailbreak attempts and steer models toward safer behavior via reminder-based conditioning. CARES provides a rigorous framework for testing and improving medical LLM safety under adversarial and ambiguous conditions. Mengxue Zhang, Eric Hanchen Jiang, Qingcheng Zeng, Chen-Hsiang Yu |
NeurIPS | 5 |
| 2025 | ThinkBench: Dynamic Out-of-Distribution Evaluation for Robust LLM ReasoningabstractEvaluating large language models (LLMs) poses significant challenges, particularly due to issues of data contamination and the leakage of correct answers. To address these challenges, we introduce ThinkBench, a novel evaluation framework designed to robustly evaluate the reasoning capability of LLMs. ThinkBench proposes a dynamic data generation method for constructing out-of-distribution (OOD) datasets and offers an OOD dataset that contains 2,912 samples drawn from reasoning tasks. ThinkBench unifies the evaluation of reasoning models and non-reasoning models. We evaluate 16 LLMs and 4 PRMs under identical experimental conditions and show that most of the LLMs' performance are far from robust and they face a certain level of data leakage. By dynamically generating OOD datasets, ThinkBench effectively provides a reliable evaluation of LLMs and reduces data contamination impact. Our data and codes are available at https://github.com/huangshulin123/ThinkBench. Shulin Huang, Linyi Yang, Yan Song 0003, Shawn Chen, Leyang Cui, Ziyu Wan, Qingcheng Zeng, Ying Wen 0001, Kun Shao, Weinan Zhang 0001, Jun Wang 0012, Yue Zhang 0004 |
NeurIPS | 7 |
| 2025 | Consistent and Optimal Solution to Camera Motion EstimationabstractGiven 2D point correspondences between an image pair, inferring the camera motion is a fundamental issue in the computer vision community. The existing works generally set out from the epipolar constraint and estimate the essential matrix, which is not optimal in the maximum likelihood (ML) sense. In this paper, we dive into the original measurement model with respect to the rotation matrix and normalized translation vector and formulate the ML problem. We then propose an optimal two-step algorithm to solve it: In the first step, we estimate the variance of measurement noises and devise a consistent estimator based on bias elimination; In the second step, we execute a one-step Gauss-Newton iteration on manifold to refine the consistent estimator. We prove that the proposed estimator achieves the same asymptotic statistical properties as the ML estimator: The first is consistency, i.e., the estimator converges to the ground truth as the point number increases; The second is asymptotic efficiency, i.e., the mean squared error of the estimator converges to the theoretical lower bound - Cramer-Rao bound. In addition, we show that our algorithm has linear time complexity. These appealing characteristics endow our estimator with a great advantage in the case of dense point correspondences. Experiments on both synthetic data and real images demonstrate that when the point number reaches the order of hundreds, our estimator outperforms the state-of-the-art ones in terms of estimation accuracy and CPU time. Guangyang Zeng, Qingcheng Zeng, Xinghan Li, Biqiang Mu, Jiming Chen 0001, Ling Shi 0001, Junfeng Wu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | A Tailored Two-Stage Algorithm for Quay Crane and Automated Guided Vehicle Scheduling ProblemsabstractThe integrated scheduling problem of cranes and automated guided vehicles (AGVs) in automated container terminals is a crucial area of concern for ports. In the terminal with AGV-supports in the yard, AGVs can autonomously place or pick up containers without waiting for yard cranes. Therefore, in such a terminal, meticulous scheduling and coordination between quay cranes (QCs) and AGVs are core for efficient and orderly operations. However, managing the operation of QCs and AGVs is complex as numerous factors affect the operational performance, such as QC interference, vehicle congestion, and limited capacity of handover points. To address the problem, we formulate a mixed-integer linear programming model that explicitly considers the above realistic factors. As the model is computationally inefficient even for small-scale instances, we develop a tailored two-stage algorithm, where the first stage is branch-and-bound for QC operations and the second is column generation for AGV operations. To validate the solution quality, we compare the proposed algorithm with some benchmark methods, and the numerical experiments confirm the effectiveness of the proposed approach. Qingcheng Zeng, Baoli Liu, Chenrui Qu |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2024 | Adaptive Axes: A Pipeline for In-domain Social Stereotype AnalysisabstractPrior work has explored the possibility of using the semantic information obtained from embedding representations to quantify social stereotypes, leveraging techniques such as word embeddings combined with a list of traits (Garg et al., 2018;Charlesworth et al., 2022) or semantic axes (An et al., 2018;Lucy et al., 2022).However, these approaches have struggled to fully capture the variability in stereotypes across different conceptual domains for the same social group (e.g., black in science, health, and art), in part because the identity of a word and the associations formed during pretraining can dominate its contextual representation (Field and Tsvetkov, 2019).This study explores the ability to recover stereotypes from the contexts surrounding targeted entities by utilizing state-of-the-art text embedding models and adaptive semantic axes enhanced by large language models (LLMs).Our results indicate that the proposed pipeline not only surpasses token-based methods in capturing in-domain framing but also effectively tracks stereotypes over time and along domain-specific semantic axes for in-domain texts.Our research highlights the potential of employing text embedding models to achieve a deeper understanding of nuanced social stereotypes. Qingcheng Zeng, Mingyu Jin, Rob Voigt |
EMNLP | 1 |
| 2024 | Integrated energy management and operations planning in oil-electric hybrid container terminals considering multi-energy supply
Chenrui Qu, Qingcheng Zeng, Xinyun Qu |
Adv. Eng. Informatics | 3 |
| 2024 | A comprehensive knowledge map for AI improving security management of cyber-physical system enabled smart manufacturing
Yu Cao 0021, Hanning Li, Qingcheng Zeng, Jing Gao 0001 |
Comput. Secur. | 4 |
| 2023 | Masked Spectrogram Prediction for Self-Supervised Audio Pre-TrainingabstractTransformer-based models attain excellent results and generalize well when trained on sufficient amounts of data. However, constrained by the limited data available in the audio domain, most transformer-based models for audio tasks are finetuned from pre-trained models in other domains (e.g. image), which has a notable gap with the audio domain. Other methods explore the self-supervised learning approaches directly in the audio domain but currently do not perform well in the downstream tasks. In this paper, we present a novel self-supervised learning method for transformer-based audio models, called masked spectrogram prediction (MaskSpec), to learn powerful audio representations from unlabeled audio data (AudioSet used in this paper). Our method masks random patches of the input spectrogram and reconstructs the masked regions with an encoder-decoder architecture. Experimental results demonstrate MaskSpec reaches the performance of 0.471 (mAP) on AudioSet, 0.854 (mAP) on Open-MIC2018, 0.982 (accuracy) on ESC-50, 0.976 (accuracy) on SCV2, and 0.823 (accuracy) on DCASE2019 Task1A. The source code and pre-trained models have been released.1 Dading Chong, Helin Wang, Peilin Zhou, Qingcheng Zeng |
ICASSP | 4 |
| 2023 | GreenPLM: Cross-Lingual Transfer of Monolingual Pre-Trained Language Models at Almost No CostabstractLarge pre-trained models have revolutionized natural language processing (NLP) research and applications, but high training costs and limited data resources have prevented their benefits from being shared equally amongst speakers of all the world's languages. To address issues of cross-linguistic access to such models and reduce energy consumption for sustainability during large-scale model training, this study proposes an effective and energy-efficient framework called GreenPLM that uses bilingual lexicons to directly ``translate'' pre-trained language models of one language into another at almost no additional cost. We validate this approach in 18 languages' BERT models and show that this framework is comparable to, if not better than, other heuristics with high training costs. In addition, given lightweight continued pre-training on limited data where available, this framework outperforms the original monolingual language models in six out of seven tested languages with up to 200x less pre-training efforts. Aiming at the Leave No One Behind Principle (LNOB), our approach manages to reduce inequalities between languages and energy consumption greatly. We make our codes and models publicly available at https://github.com/qcznlp/GreenPLMs. Qingcheng Zeng, Lucas Garay, Peilin Zhou, Dading Chong, Yining Hua, Jiageng Wu, Yikang Pan, Han Zhou 0010, Rob Voigt, Jie Yang 0039 |
IJCAI | 1 |
| 2023 | Task Scheduling of Real-Time Traffic Information Processing Based on Digital TwinsabstractThe Intelligent Transportation System under Digital Twins can provide accurate data sources for traffic control. The present work focuses on the real-time information processing and task scheduling problems of the Internet of Vehicles (IoV) system based on Virtual Reality. They are the Quality/Distance Algorithm (QDA), Task Density Algorithm, Distance Balance Algorithm (DBA), and Bionic-DBA (B-DBA). The simulation experiment analysis suggests that the DBA algorithm takes the balance of travel distance into account and effectively improves task quality. The Utility Function in B-DBA and the Biological Heuristic Search Algorithm in Pareto Ant Colony Optimization play a critically important role in enhancing the overall task quality. In addition, a Transmission based on Privacy Protection (TPP) algorithm is designed to protect the attribute-based privacy information in the traffic information transmission system. This algorithm ensures that the real-time traffic information processing system resists various attacks from malicious nodes. It has been verified that when the number of selfish nodes accounts for 30%, the transmission efficiency of the TPP algorithm reaches 0.77. The research content has a practical reference value for providing users with continuous and high-quality IoV network services. Yang Liu 0231, Qingcheng Zeng, Yuhui Sun, Jing Gao 0001, Zhihan Lyu |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2022 | A Survey in Automatic Irony Processing: Linguistic, Cognitive, and Multi-X PerspectivesabstractIrony is a ubiquitous figurative language in daily communication. Previously, many researchers have approached irony from linguistic, cognitive science, and computational aspects. Recently, some progress have been witnessed in automatic irony processing due to the rapid development in deep neural models in natural language processing (NLP). In this paper, we will provide a comprehensive overview of computational irony, insights from linguisic theory and cognitive science, as well as its interactions with downstream NLP tasks and newly proposed multi-X irony processing perspectives. Qingcheng Zeng, Anran Li 0004 |
COLING | 1 |
| 2022 | Low-resource Accent Classification in Geographically-proximate Settings: A Forensic and Sociophonetics PerspectiveabstractAccented speech recognition and accent classification are relatively under-explored research areas in speech technology.Recently, deep learning-based methods and Transformer-based pretrained models have achieved superb performances in both areas.However, most accent classification tasks focused on classifying different kinds of English accents and little attention was paid to geographically-proximate accent classification, especially under a low-resource setting where forensic speech science tasks usually encounter.In this paper, we explored three main accent modelling methods combined with two different classifiers based on 105 speaker recordings retrieved from five urban varieties in Northern England.Although speech representations generated from pretrained models generally have better performances in downstream classification, traditional methods like Mel Frequency Cepstral Coefficients (MFCCs) and formant measurements are equipped with specific strengths.These results suggest that in forensic phonetics scenario where data are relatively scarce, a simple modelling method and classifier could be competitive with state-of-the-art pretrained speech models as feature extractors, which could enhance a sooner estimation for the accent information in practices.Besides, our findings also cross-validated a new methodology in quantifying sociophonetic changes. Qingcheng Zeng, Dading Chong, Peilin Zhou, Jie Yang 0039 |
INTERSPEECH | 1 |
| 2022 | Calibrate and Refine! A Novel and Agile Framework for ASR Error Robust Intent DetectionabstractThe past ten years have witnessed the rapid development of textbased intent detection, whose benchmark performances have already been taken to a remarkable level by deep learning techniques.However, automatic speech recognition (ASR) errors are inevitable in real-world applications due to the environment noise, unique speech patterns and etc, leading to sharp performance drop in state-of-the-art text-based intent detection models.Essentially, this phenomenon is caused by the semantic drift brought by ASR errors and most existing works tend to focus on designing new model structures to reduce its impact, which is at the expense of versatility and flexibility.Different from previous one-piece model, in this paper, we propose a novel and agile framework called CR-ID for ASR error robust intent detection with two plug-and-play modules, namely semantic drift calibration module (SDCM) and phonemic refinement module (PRM), which are both model-agnostic and thus could be easily integrated to any existing intent detection models without modifying their structures.Experimental results on SNIPS dataset show that, our proposed CR-ID framework achieves competitive performance and outperform all the baseline methods on ASR outputs, which verifies that CR-ID can effectively alleviate the semantic drift caused by ASR errors. Peilin Zhou, Dading Chong, Helin Wang, Qingcheng Zeng |
INTERSPEECH | 4 |