EDBT 2026 Demo / reviewers in the wild / expert
Zhe Zhao 0006
dblp:28/6429-6
· DBLP profile ↗
35ranked-venue papers
6as first author
21since 2021 · last 2026
0009-0008-9496-5917ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 22 · 2 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 1 first-author · 7 since 2021Databases, data management, data science and information retrieval · 7 · 4 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Self-Evolving LLMs via Continual Instruction TuningabstractIn real-world industrial scenarios, large language models (LLMs) require Continuous Learning (CL) to adapt to diverse tasks as opera- tional requirements diversify, demanding self-evolution capabilities to autonomously refine their knowledge and adapt to dynamic envi- ronments. However, existing CL approaches, such as replay-based and parameter isolation techniques, struggle with the catastrophic forgetting problem: new task training degrades performance on prior tasks due to the model's adaptation to new data distributions, which weakens its generalization to old tasks. To address this issue, we propose a novel parameter-efficient adversarial MoE framework, MoE-CL, for industrial-scale self-evolving continual instruction tuning of LLMs. Specifically, MoE-CL employs a dual-expert archi- tecture to enable self-evolution: a dedicated LoRA expert for each task to preserve task-specific knowledge, ensuring parameter inde- pendence and mitigating forgetting, and a shared LoRA expert to facilitate cross-task knowledge transfer. Specifically, a task-aware discriminator within a Generative Adversarial Network (GAN) is integrated into the shared expert to suppress task-irrelevant noise, ensuring only task-aligned knowledge is transferred during se- quential task training. Through adversarial training, the shared ex- pert learns generalized representations that mimic the task-aware discriminator, while dedicated experts retain task-specific details, balancing knowledge retention and cross-task generalization—key to the model's self-evolution by autonomously optimizing knowl- edge integration across tasks. Extensive experiments on a public MTL5 benchmark and an industrial Tencent3 benchmark validate MoE-CL's effectiveness in self-evolving continual learning. In real- world A/B testing on content compliance review in the Tencent Video Platform, MoE-CL reduced manual review costs by 15.3%,demonstrating its applicability for large-scale industrial deploy- ment where self-evolution is critical for adapting to evolving op- erational demands. Implementation code is publicly available at https://github.com/BAI-LAB/MoE-CL. Jiazheng Kang, Cheng Hou, Zhe Zhao 0006, Zhenxiang Yan, Ting Bai 0004 |
WWW | 4 |
| 2026 | M2HF: Multi-Branch Multi-Modal Hybrid Fusion for Text-Video RetrievalabstractVideos contain multi-modal content, and exploring multi-branch cross-modal interactions with natural language queries can be of benefit to the text-video retrieval task (TVR). However, recent methods applying the large-scale pre-trained CLIP model for TVR only focus on visual cues in videos. Furthermore, traditional methods of simply concatenating multimodal features do not exploit fine-grained cross-modal information in videos. In this paper, we propose a multi-branch multi-modal hybrid fusion (M2HF) network to hierarchically explore interaction between text queries and other modality content in videos. Specifically, M2HF first fuses visual features extracted by CLIP with audio and motion features extracted from videos to obtain fused audio-visual features and motion-visual features respectively. The multi-modal completion problem is also considered and solved in this process. Then, visual features, audio-visual features, motion-visual features, and text extracted from the video are used to establish cross-modal relationships with caption text queries using a multibranch approach. The retrieval outputs from all branches are then fused to obtain the final text-video retrieval results. Our framework provides two kinds of training strategies, using an ensemble approach and an end-to-end approach. Moreover, a novel multi-modal loss function is proposed to balance the contributions of each modality for efficient end-to-end training. M2HF allows us to obtain state-of-the-art results on various benchmarks: Rank@1 of 66.0%, 68.6%, 33.9%, 57.4%, and 57.3% on MSR-VTT, MSVD, LSMDC, DiDeMo, and ActivityNet, respectively. Weize Quan, Zhe Zhao 0006, Kimmo Yan, Chen Chen 0001, Dong-Ming Yan 0001 |
Comput. Vis. Media | 6 |
| 2025 | Joint Knowledge Editing for Information Enrichment and Probability PromotionabstractKnowledge stored in large language models requires timely updates to reflect the dynamic nature of real-world information. To update the knowledge, most knowledge editing methods focus on the low layers, since recent probes into the knowledge recall process reveal that the answer information is enriched in low layers. However, these probes only and could only reveal critical recall stages for the original answers, while the goal of editing is to rectify model's prediction for the target answers. This inconsistency indicates that both the probe approaches and the associated editing methods are deficient. To mitigate the inconsistency and identify critical editing regions, we propose a contrast-based probe approach, and locate two crucial stages where the model behavior diverges between the original and target answers: Information Enrichment in low layers and Probability Promotion in high layers. Building upon the insights, we develop the Joint knowledge Editing for information Enrichment and probability Promotion (JEEP) method, which jointly edits both the low and high layers to modify the two critical recall stages. Considering the mutual interference and growing forgetting due to dual modifications, JEEP is designed to ensure that updates to distinct regions share the same objectives and are complementary. We rigorously evaluate JEEP by editing up to thousands of facts on various models, i.e., GPT-J (6B) and LLaMA (7B), and addressing diverse editing objectives, i.e., adding factual and counterfactual knowledge. In all tested scenarios, JEEP achieves best performances, validating the effectiveness of the revealings of our probe approach and the designs of our editing method. Wenhang Shi, Shuqing Bian, Xinyi Zhang 0002, Zhe Zhao 0006, Wei Lu 0015, Xiaoyong Du 0001 |
AAAI | 5 |
| 2025 | Edge-free but Structure-aware: Prototype-Guided Knowledge Distillation from GNNs to MLPsabstractDistilling high-accuracy Graph Neural Networks (GNNs) to low-latency multilayer perceptrons (MLPs) on graph tasks has become a hot research topic. However, conventional MLP learning relies almost exclusively on graph nodes and fails to effectively capture the graph structural information. Previous methods address this issue by processing graph edges into extra inputs for MLPs, but such graph structures may be unavailable for various scenarios. To this end, we propose Prototype-Guided Knowledge Distillation (PGKD), which does not require graph edges (edge-free setting) yet learns structure-aware MLPs. Our insight is to distill graph structural information from GNNs. Specifically, we first employ the class prototypes to analyze the impact of graph structures on GNN teachers, and then design two losses to distill such information from GNNs to MLPs. Experimental results on popular graph benchmarks demonstrate the effectiveness and robustness of the proposed PGKD. Taiqiang Wu, Zhe Zhao 0006, Jiahao Wang 0005, Xingyu Bai, Ngai Wong 0001, Yujiu Yang 0001 |
COLING | 2 |
| 2025 | Rethinking Kullback-Leibler Divergence in Knowledge Distillation for Large Language ModelsabstractKullback-Leiber divergence has been widely used in Knowledge Distillation (KD) to compress Large Language Models (LLMs). Contrary to prior assertions that reverse Kullback-Leibler (RKL) divergence is mode-seeking and thus preferable over the mean-seeking forward Kullback-Leibler (FKL) divergence, this study empirically and theoretically demonstrates that neither mode-seeking nor mean-seeking properties manifest in KD for LLMs. Instead, RKL and FKL are found to share the same optimization objective and both converge after a sufficient number of epochs. However, due to practical constraints, LLMs are seldom trained for such an extensive number of epochs. Meanwhile, we further find that RKL focuses on the tail part of the distributions, while FKL focuses on the head part at the beginning epochs. Consequently, we propose a simple yet effective Adaptive Kullback-Leiber (AKL) divergence method, which adaptively allocates weights to combine FKL and RKL. Metric-based and GPT-4-based evaluations demonstrate that the proposed AKL outperforms the baselines across various tasks and improves the diversity and quality of generated responses. Taiqiang Wu, Chaofan Tao, Jiahao Wang 0005, Runming Yang, Zhe Zhao 0006, Ngai Wong 0001 |
COLING | 5 |
| 2025 | No Loss, No Gain: Gated Refinement and Adaptive Compression for Prompt OptimizationabstractPrompt engineering is crucial for leveraging the full potential of large language models (LLMs). While automatic prompt optimization offers a scalable alternative to costly manual design, generating effective prompts remains challenging. Existing methods often struggle to stably generate improved prompts, leading to low efficiency, and overlook that prompt optimization easily gets trapped in local optima. Addressing this, we propose GRACE, a framework that integrates two synergistic strategies: Gated Refinement and Adaptive Compression, achieving Efficient prompt optimization. The gated refinement strategy introduces a feedback regulation gate and an update rejection gate, which refine update signals to produce stable and effective prompt improvements. When optimization stagnates, the adaptive compression strategy distills the prompt’s core concepts, restructuring the optimization trace and opening new paths. By strategically introducing information loss through refinement and compression, GRACE delivers substantial gains in performance and efficiency. In extensive experiments on 11 tasks across three practical domains, including BIG-Bench Hard (BBH), domain-specific, and general NLP tasks, GRACE achieves significant average relative performance improvements of 4.7\%, 4.4\% and 2.7\% over state-of-the-art methods, respectively. Further analysis shows that GRACE achieves these gains using only 25\% of the prompt generation budget required by prior methods, highlighting its high optimization efficiency and low computational overhead. Our code is available at https://github.com/Eric8932/GRACE. Wenhang Shi, Shuqing Bian, Xinyi Zhang 0002, Zhe Zhao 0006, Wei Lu 0015, Xiaoyong Du 0001 |
NeurIPS | 7 |
| 2025 | Efficient Multi-task Prompt Tuning for RecommendationabstractWith the expansion of business scenarios, real recommender systems are facing challenges in dealing with the constantly emerging new tasks in multi-task learning frameworks. In this article, we attempt to improve the generalization ability of multi-task recommendations when dealing with new tasks. A novel two-stage prompt-tuning MTL framework (MPT-Rec) is proposed to address task irrelevance and training efficiency problems in multi-task recommender systems. Specifically, we disentangle the task-specific and task-sharing information in the multi-task pre-training stage and then use task-aware prompts to transfer knowledge from other tasks to the new task effectively. By freezing parameters in the pre-training tasks, MPT-Rec solves the negative impacts that may be brought by the new task and greatly reduces the training costs. Extensive experiments on three real-world datasets show the effectiveness of our proposed multi-task learning framework. MPT-Rec achieves the best performance compared to the SOTA multi-task learning method on three real-world datasets. Besides, it maintains comparable model performance but vastly improves the training efficiency (i.e., with up to 10% parameters in the full-training way) in the new task learning. Our code is publicly available at https://github.com/BAI-LAB/MPT-Rec . Ting Bai 0004, Yue Yu 0007, Cheng Yang 0002, Cheng Hou, Zhe Zhao 0006, Chuan Shi 0001 |
ACM Trans. Inf. Syst. | 6 |
| 2024 | Mixture-of-Subspaces in Low-Rank AdaptationabstractIn this paper, we introduce a subspace-inspired Low-Rank Adaptation (LoRA) method, which is computationally efficient, easy to implement, and readily applicable to large language, multimodal, and diffusion models. Initially, we equivalently decompose the weights of LoRA into two subspaces, and find that simply mixing them can enhance performance. To study such a phenomenon, we revisit it through a fine-grained subspace lens, showing that such modification is equivalent to employing a fixed mixer to fuse the subspaces. To be more flexible, we jointly learn the mixer with the original LoRA weights, and term the method as Mixture-of-Subspaces LoRA (MoSLoRA). MoSLoRA consistently outperforms LoRA on tasks in different modalities, including commonsense reasoning, visual instruction tuning, and subject-driven text-to-image generation, demonstrating its effectiveness and robustness. Taiqiang Wu, Jiahao Wang 0005, Zhe Zhao 0006, Ngai Wong 0001 |
EMNLP | 3 |
| 2024 | Dynamic Data Sampler for Cross-Language Transfer Learning in Large Language ModelsabstractLarge Language Models (LLMs) have gained significant attention in the field of natural language processing (NLP) due to their wide range of applications. However, training LLMs for languages other than English poses significant challenges, due to the difficulty in acquiring large-scale corpus and the requisite computing resources. In this paper, we propose ChatFlow, a cross-language transfer-based LLM, to address these challenges and train large Chinese language models in a cost-effective manner. We employ a mix of Chinese, English, and parallel corpus to continuously train the LLaMA2 model, aiming to align cross-language representations and facilitate the knowledge transfer specifically to the Chinese language model. In addition, we use a dynamic data sampler to progressively transition the model from unsupervised pre-training to supervised fine-tuning. Experimental results demonstrate that our approach accelerates model convergence and achieves superior performance. We evaluate ChatFlow on popular Chinese and English benchmarks, the results indicate that it outperforms other Chinese models post-trained on LLaMA-2-7B. Yudong Li 0001, Zhe Zhao 0006, LinLin Shen, Cheng Hou, Xianxu Hou |
ICASSP | 4 |
| 2024 | RMCBench: Benchmarking Large Language Models' Resistance to Malicious CodeabstractWarning: Please note that this article contains potential harmful or offensive content. This content is only for the evaluating and analysis of LLMs and does not imply any intention to promote criminal activities. Jiachi Chen, Qingyuan Zhong, Yanlin Wang 0001, Kaiwen Ning, Yongkun Liu, Zenan Xu, Zhe Zhao 0006, Ting Chen 0002, Zibin Zheng |
ASE | 7 |
| 2024 | FLIP-80M: 80 Million Visual-Linguistic Pairs for Facial Language-Image Pre-TrainingabstractWhile significant progress has been made in multi-modal learning driven by large-scale image-text datasets, there is still a noticeable gap in the availability of such datasets within the facial domain. To facilitate and advance the field of facial representation learning, we present FLIP-80M, a large-scale visual-linguistic dataset comprising over 80 million face images paired with text descriptions. FLIP-80M is constructed by leveraging the large openly available image-text-pair dataset LAION-5B and a mixed-method approach to filter face-related pairs from both visual and linguistic perspectives. Our curation process involves face detection, face caption classification, text de-noising, and synthesis-based image augmentation. As a result, FLIP-80M stands as the largest face-text dataset to date. To evaluate the potential of our dataset, we fine-tune the CLIP model using the proposed FLIP-80M, to create FLIP (Facial Language-Image Pretraining) and assess its representation capabilities across various downstream tasks. Our experiments demonstrate that our FLIP model achieves state-of-the-art results in a range of face analysis tasks, including face parsing, face alignment, and face attribute classification. The dataset and models are available at https://github.com/ydli-ai/FLIP. Yudong Li 0001, Xianxu Hou, Dezhi Zheng, LinLin Shen, Zhe Zhao 0006 |
ACM Multimedia | 5 |
| 2023 | Document-Level Event Argument Extraction With a Chain Reasoning ParadigmabstractDocument-level event argument extraction aims to identify event arguments beyond sentence level, where a significant challenge is to model long-range dependencies.Focusing on this challenge, we present a new chain reasoning paradigm for the task, which can generate decomposable first-order logic rules for reasoning.This paradigm naturally captures long-range interdependence due to the chains' compositional nature, which also improves interpretability by explicitly modeling the reasoning process.We introduce T-norm fuzzy logic for optimization, which permits end-toend learning and shows promise for integrating the expressiveness of logical reasoning with the generalization of neural networks.In experiments, we show that our approach outperforms previous methods by a significant margin on two standard benchmarks (over 6 points in F1).Moreover, it is data-efficient in lowresource scenarios and robust enough to defend against adversarial attacks. Jian Liu 0032, Jin An Xu, Haoyan Liu 0001, Zhe Zhao 0006 |
ACL (1) | 5 |
| 2023 | Learning with Partial Annotations for Event DetectionabstractEvent detection (ED) seeks to discover and classify event instances in plain texts.Previous methods for ED typically adopt supervised learning, requiring fully labeled and high-quality training data.However, in a realworld application, we may not obtain clean training data but only partially labeled one, which could substantially impede the learning process.In this work, we conduct a seminal study for learning with partial annotations for ED.We propose a new trigger localization formulation using contrastive learning to distinguish ground-truth triggers from contexts, showing a decent robustness for addressing partial annotation noise.Impressively, in an extreme scenario where more than 90% of events are unlabeled, our approach achieves an F1 score of over 60%.In addition, we reannotate and make available two fully annotated subsets of ACE 2005 to serve as an unbiased benchmark for event detection.We hope our approach and data will inspire future studies on this vital yet understudied problem. Jian Liu 0032, Dianbo Sui, Kang Liu 0001, Haoyan Liu 0001, Zhe Zhao 0006 |
ACL (1) | 5 |
| 2023 | Create and Find Flatness: Building Flat Training Spaces in Advance for Continual LearningabstractCatastrophic forgetting remains a critical challenge in the field of continual learning, where neural networks struggle to retain prior knowledge while assimilating new information. Most existing studies emphasize mitigating this issue only when encountering new tasks, overlooking the significance of the pre-task phase. Therefore, we shift the attention to the current task learning stage, presenting a novel framework, C&F (Create and Find Flatness), which builds a flat training space for each task in advance. Specifically, during the learning of the current task, our framework adaptively creates a flat region around the minimum in the the loss landscape. Subsequently, it finds the parameters’ importance to the current task based on their flatness degrees. When adapting the model to a new task, constraints are applied according to the flatness and a flat space is simultaneously prepared for the impending task. We theoretically demonstrate the consistency between the created and found flatness. In this manner, our framework not only accommodates ample parameter space for learning new tasks but also preserves the preceding knowledge of earlier tasks. Experimental results exhibit C&F’s state-of-the-art performance as a standalone continual learning approach and its efficacy as a framework incorporating other methods. Our work is available at https://github.com/Eric8932/Create-and-Find-Flatness. Wenhang Shi, Zhe Zhao 0006, Wei Lu 0015, Kimmo Yan, Xiaoyong Du 0001 |
ECAI | 3 |
| 2023 | Addressing NER Annotation Noises with Uncertainty-Guided Tree-Structured CRFsabstractReal-world named entity recognition (NER) datasets are notorious for their noisy nature, attributed to annotation errors, inconsistencies, and subjective interpretations.Such noises present a substantial challenge for traditional supervised learning methods.In this paper, we present a new and unified approach to tackle annotation noises for NER.Our method considers NER as a constituency tree parsing problem, utilizing a tree-structured Conditional Random Fields (CRFs) with uncertainty evaluation for integration.Through extensive experiments conducted on four realworld datasets, we demonstrate the effectiveness of our model in addressing both partial and incorrect annotation errors.Remarkably, our model exhibits superb performance even in extreme scenarios with 90% annotation noise. Jian Liu 0032, Weichang Liu, Yufeng Chen 0005, Jin An Xu, Zhe Zhao 0006 |
EMNLP | 5 |
| 2023 | Recouple Event Field via Probabilistic Bias for Event ExtractionabstractEvent Extraction (EE), aiming to identify and classify event triggers and arguments from event mentions, has benefited from pre-trained language models (PLMs). However, existing PLM-based methods ignore the information of trigger/argument fields, which is crucial for understanding event schemas. To this end, we propose a Probabilistic reCoupling model enhanced Event extraction framework (ProCE). Specifically, we first model the syntactic-related event fields as probabilistic biases, to clarify the event fields from ambiguous entanglement. Furthermore, considering multiple occurrences of the same triggers/arguments in EE, we explore probabilistic interaction strategies among multiple fields of the same triggers/arguments, to recouple the corresponding clarified distributions and capture more latent information fields. Experiments on EE datasets demonstrate the effectiveness and generalization of our proposed approach. Xingyu Bai, Taiqiang Wu, Zhe Zhao 0006, Xuefeng Yang, Jiayi Li 0002, Weijie Liu 0002, Qi Ju 0002, Weigang Guo, Yujiu Yang 0001 |
ICASSP | 4 |
| 2023 | An Empirical Study on Adaptive Inference for Pretrained Language ModelabstractAdaptive inference has been proven to improve bidirectional encoder representations from transformers (BERT)'s inference speed with minimal loss of accuracy. However, current work only focuses on the BERT model and lacks exploration of other pretrained language models (PLMs). Therefore, this article conducts an empirical study on the application of adaptive inference mechanism in various PLMs, including generative pretraining (GPT), GCNN, ALBERT, and TinyBERT. This mechanism is verified on both English and Chinese benchmarks, and experimental results demonstrated that it is able to speed up by a wide range from 1 to 10 times if given different speed thresholds. In addition, its application on ALBERT shows that adaptive inference can work with parameter sharing, achieving model compression and acceleration simultaneously, while the application on TinyBERT proves that it can further accelerate the distilled small model. As for the problem that too many labels make adaptive inference invalid, this article also proposes a solution, namely label reduction. Finally, this article open-sources an easy-to-use toolkit called FastPLM to help developers adopt pretrained models with adaptive inference capabilities in their applications. Weijie Liu 0002, Zhe Zhao 0006, Qi Ju 0002, Xuefeng Yang, Wei Lu 0015 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2022 | A Simple and Effective Method to Improve Zero-Shot Cross-Lingual Transfer LearningabstractExisting zero-shot cross-lingual transfer methods rely on parallel corpora or bilingual dictionaries, which are expensive and impractical for low-resource languages. To disengage from these dependencies, researchers have explored training multilingual models on English-only resources and transferring them to low-resource languages. However, its effect is limited by the gap between embedding clusters of different languages. To address this issue, we propose Embedding-Push, Attention-Pull, and Robust targets to transfer English embeddings to virtual multilingual embeddings without semantic loss, thereby improving cross-lingual transferability. Experimental results on mBERT and XLM-R demonstrate that our method significantly outperforms previous works on the zero-shot cross-lingual text classification task and can obtain a better multilingual alignment. Kunbo Ding, Weijie Liu 0002, Yuejian Fang, Weiquan Mao, Zhe Zhao 0006, Haoyan Liu 0001, Rong Tian |
COLING | 5 |
| 2022 | CSL: A Large-scale Chinese Scientific Literature DatasetabstractScientific literature serves as a high-quality corpus, supporting a lot of Natural Language Processing (NLP) research. However, existing datasets are centered around the English language, which restricts the development of Chinese scientific NLP. In this work, we present CSL, a large-scale Chinese Scientific Literature dataset, which contains the titles, abstracts, keywords and academic fields of 396k papers. To our knowledge, CSL is the first scientific document dataset in Chinese. The CSL can serve as a Chinese corpus. Also, this semi-structured data is a natural annotation that can constitute many supervised NLP tasks. Based on CSL, we present a benchmark to evaluate the performance of models across scientific domain tasks, i.e., summarization, keyword generation and text classification. We analyze the behavior of existing text-to-text models on the evaluation tasks and reveal the challenges for Chinese scientific NLP tasks, which provides a valuable reference for future research. Data and code will be publicly available. Yudong Li 0001, Zhe Zhao 0006, LinLin Shen, Weijie Liu 0002, Weiquan Mao |
COLING | 3 |
| 2022 | Talk2Face: A Unified Sequence-based Framework for Diverse Face Generation and Analysis TasksabstractFacial analysis is an important domain in computer vision and has received extensive research attention. For numerous downstream tasks with different input/output formats and modalities, existing methods usually design task-specific architectures and train them using face datasets collected in the particular task domain. In this work, we proposed a single model, Talk2Face, to simultaneously tackle a large number of face generation and analysis tasks, e.g. text guided face synthesis, face captioning and age estimation. Specifically, we cast different tasks into a sequence-to-sequence format with the same architecture, parameters and objectives. While text and facial images are tokenized to sequences, the annotation labels of faces for different tasks are also converted to natural languages for unified representation. We collect a set of 2.3M face-text pairs from available datasets across different tasks, to train the proposed model. Uniform templates are then designed to enable the model to perform different downstream tasks, according to the task context and target. Experiments on different tasks show that our model achieves better face generation and caption performances than SOTA approaches. On age estimation and multi-attribute classification, our model reaches competitive performance with those models specially designed and trained for these particular tasks. In practice, our model is much easier to be deployed to different facial analysis related tasks. Code and dataset will be available at https://github.com/ydli-ai/Talk2Face. Yudong Li 0001, Xianxu Hou, Zhe Zhao 0006, LinLin Shen, Xuefeng Yang, Kimmo Yan |
ACM Multimedia | 3 |
| 2021 | Overview of the NLPCC 2021 Shared Task: AutoIE2
Weigang Guo, Xuefeng Yang, Xingyu Bai, Taiqiang Wu, Weijie Liu 0002, Zhe Zhao 0006, Qi Ju 0002, Yujiu Yang 0001 |
NLPCC (2) | 6 |
| 2020 | K-BERT: Enabling Language Representation with Knowledge GraphabstractPre-trained language representation models, such as BERT, capture a general language representation from large-scale corpora, but lack domain-specific knowledge. When reading a domain text, experts make inferences with relevant knowledge. For machines to achieve this capability, we propose a knowledge-enabled language representation model (K-BERT) with knowledge graphs (KGs), in which triples are injected into the sentences as domain knowledge. However, too much knowledge incorporation may divert the sentence from its correct meaning, which is called knowledge noise (KN) issue. To overcome KN, K-BERT introduces soft-position and visible matrix to limit the impact of knowledge. K-BERT can easily inject domain knowledge into the models by being equipped with a KG without pre-training by itself because it is capable of loading model parameters from the pre-trained BERT. Our investigation reveals promising results in twelve NLP tasks. Especially in domain-specific tasks (including finance, law, and medicine), K-BERT significantly outperforms BERT, which demonstrates that K-BERT is an excellent choice for solving the knowledge-driven problems that require experts. Weijie Liu 0002, Zhe Zhao 0006, Zhiruo Wang 0001, Qi Ju 0002, Haotang Deng, Ping Wang 0003 |
AAAI | 3 |
| 2020 | FastBERT: a Self-distilling BERT with Adaptive Inference TimeabstractPre-trained language models like BERT have proven to be highly performant.However, they are often computationally expensive in many practical scenarios, for such heavy models can hardly be readily implemented with limited resources.To improve their efficiency with an assured model performance, we propose a novel speed-tunable FastBERT with adaptive inference time.The speed at inference can be flexibly adjusted under varying demands, while redundant calculation of samples is avoided.Moreover, this model adopts a unique selfdistillation mechanism at fine-tuning, further enabling a greater computational efficacy with minimal loss in performance.Our model achieves promising results in twelve English and Chinese datasets.It is able to speed up by a wide range from 1 to 12 times than BERT if given different speedup thresholds to make a speed-performance tradeoff. Weijie Liu 0002, Zhiruo Wang 0001, Zhe Zhao 0006, Haotang Deng, Qi Ju 0002 |
ACL | 4 |
| 2020 | CLUE: A Chinese Language Understanding Evaluation BenchmarkabstractLiang Xu, Hai Hu, Xuanwei Zhang, Lu Li, Chenjie Cao, Yudong Li, Yechen Xu, Kai Sun, Dian Yu, Cong Yu, Yin Tian, Qianqian Dong, Weitang Liu, Bo Shi, Yiming Cui, Junyi Li, Jun Zeng, Rongzhao Wang, Weijian Xie, Yanting Li, Yina Patterson, Zuoyu Tian, Yiwen Zhang, He Zhou, Shaoweihua Liu, Zhe Zhao, Qipeng Zhao, Cong Yue, Xinrui Zhang, Zhengliang Yang, Kyle Richardson, Zhenzhong Lan. Proceedings of the 28th International Conference on Computational Linguistics. 2020. Liang Xu 0011, Hai Hu 0001, Xuanwei Zhang, Chenjie Cao, Yudong Li 0001, Yechen Xu, Kai Sun 0006, Dian Yu 0001, Cong Yu 0010, Yin Tian, Qianqian Dong, Weitang Liu, Yiming Cui 0001, Rongzhao Wang, Weijian Xie, Yina Patterson, Zuoyu Tian, Shaoweihua Liu, Zhe Zhao 0006, Qipeng Zhao, Cong Yue, Zhengliang Yang, Kyle Richardson 0001, Zhen-Zhong Lan |
COLING | 26 |
| 2017 | Neural Bag-of-NgramsabstractBag-of-ngrams (BoN) models are commonly used for representing text. One of the main drawbacks of traditional BoN is the ignorance of n-gram's semantics. In this paper, we introduce the concept of Neural Bag-of-ngrams (Neural-BoN), which replaces sparse one-hot n-gram representation in traditional BoN with dense and rich-semantic n-gram representations. We first propose context guided n-gram representation by adding n-grams to word embeddings model. However, the context guided learning strategy of word embeddings is likely to miss some semantics for text-level tasks. Text guided n-gram representation and label guided n-gram representation are proposed to capture more semantics like topic or sentiment tendencies. Neural-BoN with the latter two n-gram representations achieve state-of-the-art results on 4 document-level classification datasets and 6 semantic relatedness categories. They are also on par with some sophisticated DNNs on 3 sentence-level classification datasets. Similar to traditional BoN, Neural-BoN is efficient, robust and easy to implement. We expect it to be a strong baseline and be used in more real-world applications. Bofang Li, Tao Liu 0001, Zhe Zhao 0006, Puwei Wang, Xiaoyong Du 0001 |
AAAI | 3 |
| 2017 | Investigating Different Syntactic Context Types and Context Representations for Learning Word EmbeddingsabstractThe number of word embedding models is growing every year.Most of them are based on the co-occurrence information of words and their contexts.However, it is still an open question what is the best definition of context.We provide a systematical investigation of 4 different syntactic context types and context representations for learning word embeddings.Comprehensive experiments are conducted to evaluate their effectiveness on 6 extrinsic and intrinsic tasks.We hope that this paper, along with the published code, would be helpful for choosing the best context type and representation for a given task. Bofang Li, Tao Liu 0001, Zhe Zhao 0006, Buzhou Tang, Aleksandr Drozd, Anna Rogers, Xiaoyong Du 0001 |
EMNLP | 3 |
| 2017 | Initializing Convolutional Filters with Semantic Features for Text ClassificationabstractConvolutional Neural Networks (CNNs) are widely used in NLP tasks.This paper presents a novel weight initialization method to improve the CNNs for text classification.Instead of randomly initializing the convolutional filters, we encode semantic features into them, which helps the model focus on learning useful features at the beginning of the training.Experiments demonstrate the effectiveness of the initialization technique on seven text classification tasks, including sentiment analysis and topic classification. Zhe Zhao 0006, Tao Liu 0001, Renfen Hu, Xiaoyong Du 0001 |
EMNLP | 2 |
| 2017 | Ngram2vec: Learning Improved Word Representations from Ngram Co-occurrence StatisticsabstractThe existing word representation methods mostly limit their information source to word co-occurrence statistics.In this paper, we introduce ngrams into four representation methods: SGNS, GloVe, PPMI matrix, and its SVD factorization.Comprehensive experiments are conducted on word analogy and similarity tasks.The results show that improved word representations are learned from ngram cooccurrence statistics.We also demonstrate that the trained ngram representations are useful in many aspects such as finding antonyms and collocations.Besides, a novel approach of building co-occurrence matrix is proposed to alleviate the hardware burdens brought by ngrams. Zhe Zhao 0006, Tao Liu 0001, Bofang Li, Xiaoyong Du 0001 |
EMNLP | 1 |
| 2017 | Guiding the Training of Distributed Text Representation with Supervised Weighting Scheme for Sentiment AnalysisabstractWith the rapid growth of social media, sentiment analysis has received growing attention from both academic and industrial fields. One line of researches for sentiment analysis is to feed bag-of-words (BOW) text representation into classifiers. Usually, raw BOW requires weighting schemes to obtain better performance, where important words are given more weights while unimportant ones are given less weights. Another line of researches focuses on neural models, where distributed text representations are learned from raw texts automatically. In this paper, we take advantages of techniques in both lines of researches. We use words’ weights to guide neural models to focus on important words. Various supervised weighting schemes are explored in this work. We discover that better text features are learned for sentiment analysis when suitable weighting schemes are applied upon neural models. Zhe Zhao 0006, Tao Liu 0001, Bofang Li, Xiaoyong Du 0001 |
Data Sci. Eng. | 1 |
| 2017 | Correction to: Guiding the Training of Distributed Text Representation with Supervised Weighting Scheme for Sentiment AnalysisabstractIn the originally published article, the acknowledgment section is missing. Please find it as follows. Zhe Zhao 0006, Tao Liu 0001, Bofang Li, Xiaoyong Du 0001 |
Data Sci. Eng. | 1 |
| 2016 | Classifying Relation via Bidirectional Recurrent Neural Network Based on Local Information
Xiaoyun Hou, Zhe Zhao 0006, Tao Liu 0001, Xiaoyong Du 0001 |
APWeb (1) | 2 |
| 2016 | A Text Retrieval System Based on Distributed Representations
Zhe Zhao 0006, Tao Liu 0001, Jun Chen 0021, Bofang Li, Xiaoyong Du 0001 |
APWeb (2) | 1 |
| 2016 | Distributed Text Representation with Weighting Scheme Guidance for Sentiment Analysis
Zhe Zhao 0006, Tao Liu 0001, Xiaoyun Hou, Bofang Li, Xiaoyong Du 0001 |
APWeb (1) | 1 |
| 2016 | Weighted Neural Bag-of-n-grams Model: New Baselines for Text ClassificationabstractNBSVM is one of the most popular methods for text classification and has been widely used as baselines for various text representation approaches. It uses Naive Bayes (NB) feature to weight sparse bag-of-n-grams representation. N-gram captures word order in short context and NB feature assigns more weights to those important words. However, NBSVM suffers from sparsity problem and is reported to be exceeded by newly proposed distributed (dense) text representations learned by neural networks. In this paper, we transfer the n-grams and NB weighting to neural models. We train n-gram embeddings and use NB weighting to guide the neural models to focus on important words. In fact, our methods can be viewed as distributed (dense) counterparts of sparse bag-of-n-grams in NBSVM. We discover that n-grams and NB weighting are also effective in distributed representations. As a result, our models achieve new strong baselines on 9 text classification datasets, e.g. on IMDB dataset, we reach performance of 93.5% accuracy, which exceeds previous state-of-the-art results obtained by deep neural models. All source codes are publicly available at https://github.com/zhezhaoa/neural_BOW_toolkit. Bofang Li, Zhe Zhao 0006, Tao Liu 0001, Puwei Wang, Xiaoyong Du 0001 |
COLING | 2 |
| 2016 | Cluster-Driven Model for Improved Word and Text EmbeddingabstractMost of the existing word embedding models only consider the relationships between words and their local contexts (e.g. ten words around the target word). However, information beyond local contexts (global contexts), which reflect the rich semantic meanings of words, are usually ignored. In this paper, we present a general framework for utilizing global information to learn word and text representations. Our models can be easily integrated into existing local word embedding models, and thus introduces global information of varying degrees according to different downstream tasks. Moreover, we view our models in the co-occurrence matrix perspective, based on which a novel weighted term-document matrix is factorized to generate text representations. We conduct a range of experiments to evaluate word and text representations learned by our models. Experimental results show that our models outperform or compete with state-of-the-art models. Source code of the paper is available at https://github.com/zhezhaoa/cluster-driven. Zhe Zhao 0006, Tao Liu 0001, Bofang Li, Xiaoyong Du 0001 |
ECAI | 1 |