VLDB 2026 Research / reviewers in the wild / expert
Qihuang Zhong
dblp:272/6439
· DBLP profile ↗
15ranked-venue papers
8as first author
14since 2021 · last 2026
0009-0001-0118-5217ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 4 first-author · 10 since 2021Databases, data management, data science and information retrieval · 3 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Achieving >97% on GSM8K: deeply understanding the problems makes LLMs better solvers for math word problems
Qihuang Zhong, Liang Ding 0006, Juhua Liu, Bo Du 0001 |
Frontiers Comput. Sci. | 1 |
| 2026 | MKDS: Multi-source knowledge-driven data synthesis framework for effective domain adaptation of large language models
Qihuang Zhong, Jinzhao Gong, Juhua Liu, Bo Du 0001 |
Knowl. Based Syst. | 1 |
| 2026 | SFA: Scan, Focus, and Amplify toward guidance-aware answering for Video TextVQA
Haibin He 0001, Qihuang Zhong, Juhua Liu, Bo Du 0001, Peng Wang 0076, Jing Zhang 0037 |
Pattern Recognit. | 2 |
| 2024 | Revisiting Knowledge Distillation for Autoregressive Language ModelsabstractKnowledge distillation (KD) is a common approach to compress a teacher model to reduce its inference cost and memory footprint, by training a smaller student model.However, in the context of autoregressive language models (LMs), we empirically find that larger teachers might dramatically result in a poorer student.In response to this problem, we conduct a series of analyses and reveal that different tokens have different teaching modes, neglecting which will lead to performance degradation.Motivated by this, we propose a simple yet effective adaptive teaching approach (ATKD) to improve the KD.The core of ATKD is to reduce rote learning and make teaching more diverse and flexible.Extensive experiments on 8 LM tasks show that, with the help of ATKD, various baseline KD methods can achieve consistent and significant performance gains (up to +3.04% average score) across all model types and sizes.More encouragingly, ATKD can improve the student model generalization effectively. Qihuang Zhong, Liang Ding 0006, Li Shen 0008, Juhua Liu, Bo Du 0001, Dacheng Tao |
ACL (1) | 1 |
| 2024 | AdaSAM: Boosting sharpness-aware minimization with adaptive learning rate and momentum for training deep neural networks
Hao Sun 0019, Li Shen 0008, Qihuang Zhong, Liang Ding 0006, Shixiang Chen, Jingwei Sun 0001, Jing Li 0047, Guangzhong Sun, Dacheng Tao |
Neural Networks | 3 |
| 2024 | PanDa: Prompt Transfer Meets Knowledge Distillation for Efficient Model AdaptationabstractPrompt Transfer (PoT) is a recently-proposed approach to improve prompt-tuning, by initializing the target prompt with the existing prompt trained on similar source tasks. However, such a vanilla PoT approach usually achieves sub-optimal performance, as (i) the PoT is sensitive to the similarity of source-target pair and (ii) directly fine-tuning the prompt initialized with source prompt on target task might lead to forgetting of the useful general knowledge learned from source task. To tackle these issues, we propose a new metric to accurately predict the prompt transferability (regarding (i)), and a novel PoT approach (namelyPanDa) that leverages the knowledge distillation technique to alleviate the knowledge forgetting effectively (regarding (ii)). Extensive and systematic experiments on 189 combinations of 21 source and 9 target datasets across 5 scales of PLMs demonstrate that: 1)our proposed metric works well to predict the prompt transferability; 2)ourPanDaconsistently outperforms the vanilla PoT approach by 2.3% average score (up to 24.1%) among all tasks and model sizes; 3)with ourPanDaapproach, prompt-tuning can achieve competitive and even better performance than model-tuning in various PLM scales scenarios. Qihuang Zhong, Liang Ding 0006, Juhua Liu, Bo Du 0001, Dacheng Tao |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2024 | E2S2: Encoding-Enhanced Sequence-to-Sequence Pretraining for Language Understanding and GenerationabstractSequence-to-sequence (seq2seq) learning is a popular fashion for large-scale pretraining language models. However, the previous seq2seq pretraining models generally focus on reconstructive objectives on the decoder side and neglect the effect of encoder-side supervision, which we argue may lead to sub-optimal performance. To verify our hypothesis, we first empirically study the functionalities of the encoder and decoder in seq2seq pretrained language models, and find that the encoder takes an important but under-exploitation role than the decoder regarding the downstream performance and neuron activation. Therefore, we propose an encoding-enhanced seq2seq pretraining strategy, namelyE2S2, which improves the seq2seq models via integrating more efficient self-supervised information into the encoders. Specifically, E2S2 adopts two self-supervised objectives on the encoder side from two aspects: 1) locally denoising the corrupted sentence (denoising objective); and 2) globally learning better sentence representations (contrastive objective). With the help of both objectives, the encoder can effectively distinguish the noise tokens and capture high-level (i.e., syntactic and semantic) knowledge, thus strengthening the ability of seq2seq model to accurately achieve the conditional generation. On a large diversity of downstream natural language understanding and generation tasks, E2S2 dominantly improves the performance of its powerful backbone models, e.g., BART and T5. For example, upon BART backbone, we achieve +1.1% averaged gain on the general language understanding evaluation (GLUE) benchmark and +1.75%$F_{0.5}$score improvement on CoNLL2014 dataset. We also provide in-depth analyses to show the improvement stems from better linguistic representation. We hope that our work will foster future self-supervision research on seq2seq language model pretraining. Qihuang Zhong, Liang Ding 0006, Juhua Liu, Bo Du 0001, Dacheng Tao |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2023 | Revisiting Token Dropping Strategy in Efficient BERT PretrainingabstractToken dropping is a recently-proposed strategy to speed up the pretraining of masked language models, such as BERT, by skipping the computation of a subset of the input tokens at several middle layers.It can effectively reduce the training time without degrading much performance on downstream tasks.However, we empirically find that token dropping is prone to a semantic loss problem and falls short in handling semantic-intense tasks ( §2).Motivated by this, we propose a simple yet effective semantic-consistent learning method (SCTD) to improve the token dropping.SCTD aims to encourage the model to learn how to preserve the semantic information in the representation space.Extensive experiments on 12 tasks show that, with the help of our SCTD, token dropping can achieve consistent and significant performance gains across all task types and model sizes.More encouragingly, SCTD saves up to 57% of pretraining time and brings up to +1.56% average improvement over the vanilla token dropping. Qihuang Zhong, Liang Ding 0006, Juhua Liu, Xuebo Liu 0002, Min Zhang 0005, Bo Du 0001, Dacheng Tao |
ACL (1) | 1 |
| 2023 | Self-Evolution Learning for Mixup: Enhance Data Augmentation on Few-Shot Text Classification TasksabstractText classification tasks often encounter fewshot scenarios with limited labeled data, and addressing data scarcity is crucial.Data augmentation with mixup merges sample pairs to generate new pseudos, which can relieve the data deficiency issue in text classification.However, the quality of pseudo-samples generated by mixup exhibits significant variations.Most of the mixup methods fail to consider the varying degree of learning difficulty in different stages of training.And mixup generates new samples with one-hot labels, which encourages the model to produce a high prediction score for the correct class that is much larger than other classes, resulting in the model's over-confidence.In this paper, we propose a self-evolution learning (SE) based mixup approach for data augmentation in text classification, which can generate more adaptive and model-friendly pseudo samples for the model training.SE caters to the growth of the model learning ability and adapts to the ability when generating training samples.To alleviate the model over-confidence, we introduce an instance-specific label smoothing regularization approach, which linearly interpolates the model's output and one-hot labels of the original samples to generate new soft labels for label mixing up.Through experimental analysis, experiments show that our SE brings consistent and significant improvements upon different mixup methods.In-depth analyses demonstrate that SE enhances the model's generalization ability. Haoqi Zheng, Qihuang Zhong, Liang Ding 0006, Zhiliang Tian, Xin Niu 0002, Dongsheng Li 0001, Dacheng Tao |
EMNLP | 2 |
| 2023 | Zero-shot Sharpness-Aware Quantization for Pre-trained Language ModelsabstractQuantization is a promising approach for reducing memory overhead and accelerating inference, especially in large pre-trained language model (PLM) scenarios.While having no access to original training data due to security and privacy concerns has emerged the demand for zero-shot quantization.Most of the cuttingedge zero-shot quantization methods primarily ❶ apply to computer vision tasks, and ❷ neglect of overfitting problem in the generative adversarial learning process, leading to sub-optimal performance.Motivated by this, we propose a novel zero-shot sharpness-aware quantization (ZSAQ) framework for the zeroshot quantization of various PLMs.The key algorithm in solving ZSAQ is the SAM-SGA optimization, which aims to improve the quantization accuracy and model generalization via optimizing a minimax problem.We theoretically prove the convergence rate for the minimax optimization problem and this result can be applied to other nonconvex-PL minimax optimization frameworks.Extensive experiments on 11 tasks demonstrate that our method brings consistent and significant performance gains on both discriminative and generative PLMs, i.e., up to +6.98 average score.Furthermore, we empirically validate that our method can effectively improve the model generalization. Miaoxi Zhu, Qihuang Zhong, Li Shen 0008, Liang Ding 0006, Juhua Liu, Bo Du 0001, Dacheng Tao |
EMNLP | 2 |
| 2023 | Joint image and feature adaptative attention-aware networks for cross-modality semantic segmentation
Qihuang Zhong, Fanzhou Zeng, Juhua Liu, Bo Du 0001, Jedi S. Shang |
Neural Comput. Appl. | 1 |
| 2023 | Unified Instance and Knowledge Alignment Pretraining for Aspect-Based Sentiment AnalysisabstractThe goal of aspect-based sentiment analysis (ABSA) is to determine the sentiment polarity towards an aspect. Because of the expensive and limited amounts of labelled data, the pretraining strategy has become the de facto standard for ABSA. However, there always exists a severe domain shift between the pretraining and downstream ABSA datasets, which hinders effective knowledge transfer when directly fine-tuning, making the downstream task suboptimal. To mitigate this domain shift, we introduce a unified alignment pretraining framework into the vanilla pretrain-finetune pipeline, that has both instance- and knowledge-level alignments. Specifically, we first devise a novel coarse-to-fine retrieval sampling approach to select target domain-related instances from the large-scale pretraining dataset, thus aligning the instances between pretraining and the target domains (First Stage). Then, we introduce a knowledge guidance-based strategy to further bridge the domain gap at the knowledge level. In practice, we formulate the model pretrained on the sampled instances into a knowledge guidance model and a learner model. On the target dataset, we design an on-the-fly teacher-student joint fine-tuning approach to progressively transfer the knowledge from the knowledge guidance model to the learner model (Second Stage). Therefore, the learner model can maintain more domain-invariant knowledge when learning new knowledge from the target dataset. In theThird Stage,the learner model is finetuned to better adapt its learned knowledge to the target dataset. Extensive experiments and analyses on several ABSA benchmarks demonstrate the effectiveness and universality of our proposed pretraining framework. Juhua Liu, Qihuang Zhong, Liang Ding 0006, Bo Du 0001, Dacheng Tao |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2023 | Knowledge Graph Augmented Network Towards Multiview Representation Learning for Aspect-Based Sentiment AnalysisabstractAspect-based sentiment analysis (ABSA) is a fine-grained task of sentiment analysis. To better comprehend long complicated sentences and obtain accurate aspect-specific information, linguistic and commonsense knowledge are generally required in this task. However, most current methods employ complicated and inefficient approaches to incorporate external knowledge, e.g., directly searching the graph nodes. Additionally, the complementarity between external knowledge and linguistic information has not been thoroughly studied. To this end, we propose a knowledge graph augmented network (KGAN), which aims to effectively incorporate external knowledge with explicitly syntactic and contextual information. In particular, KGAN captures the sentiment feature representations from multiple different perspectives,i.e., context-, syntax- and knowledge-based. First, KGAN learns the contextual and syntactic representations in parallel to fully extract the semantic features. Then, KGAN integrates the knowledge graphs into the embedding space, based on which the aspect-specific knowledge representations are further obtained via an attention mechanism. Last, we propose a hierarchical fusion module to complement these multi-view representations in alocal-to-globalmanner. Extensive experiments on five popular ABSA benchmarks demonstrate the effectiveness and robustness of our KGAN. Notably, with the help of the pretrained model of RoBERTa, KGAN achieves a new record of state-of-the-art performance among all datasets. Qihuang Zhong, Liang Ding 0006, Juhua Liu, Bo Du 0001, Dacheng Tao |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2022 | A Contrastive Cross-Channel Data Augmentation Framework for Aspect-Based Sentiment AnalysisabstractAspect-based sentiment analysis (ABSA) is a fine-grained sentiment analysis task, which focuses on detecting the sentiment polarity towards the aspect in a sentence. However, it is always sensitive to the multi-aspect challenge, where features of multiple aspects in a sentence will affect each other. To mitigate this issue, we design a novel training framework, called Contrastive Cross-Channel Data Augmentation (C3 DA), which leverages an in-domain generator to construct more multi-aspect samples and then boosts the robustness of ABSA models via contrastive learning on these generated data. In practice, given a generative pretrained language model and some limited ABSA labeled data, we first employ some parameter-efficient approaches to perform the in-domain fine-tuning. Then, the obtained in-domain generator is used to generate the synthetic sentences from two channels, i.e., Aspect Augmentation Channel and Polarity Augmentation Channel, which generate the sentence condition on a given aspect and polarity respectively. Specifically, our C3 DA performs the sentence generation in a cross-channel manner to obtain more sentences, and proposes an Entropy-Minimization Filter to filter low-quality generated samples. Extensive experiments show that our C3 DA can outperform those baselines without any augmentations by about 1% on accuracy and Macro- F1. Code and data are released in https://github.com/wangbing1416/C3DA. Bing Wang 0018, Liang Ding 0006, Qihuang Zhong, Ximing Li 0002, Dacheng Tao |
COLING | 3 |
| 2020 | SemiText: Scene text detection with semi-supervised learning
Juhua Liu, Qihuang Zhong, Hai Su, Bo Du 0001 |
Neurocomputing | 2 |