EDBT 2026 Demo / reviewers in the wild / expert
Yaming Yang 0001
dblp:204/3789-1
· DBLP profile ↗
23ranked-venue papers
4as first author
20since 2021 · last 2025
0009-0005-8744-7693ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 20 · 3 first-author · 17 since 2021Databases, data management, data science and information retrieval · 7 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MTL-LoRA: Low-Rank Adaptation for Multi-Task LearningabstractParameter-efficient fine-tuning (PEFT) has been widely employed for domain adaptation, with LoRA being one of the most prominent methods due to its simplicity and effectiveness. However, in multi-task learning (MTL) scenarios, LoRA tends to obscure the distinction between tasks by projecting sparse high-dimensional features from different tasks into the same dense low-dimensional intrinsic space. This leads to task interference and suboptimal performance for LoRA and its variants. To tackle this challenge, we propose MTL-LoRA, which retains the advantages of low-rank adaptation while significantly enhancing MTL capabilities. MTL-LoRA augments LoRA by incorporating additional task-adaptive parameters that differentiate task-specific information and capture shared knowledge across various tasks within low-dimensional spaces. This approach enables pretrained models to jointly adapt to different target domains with a limited number of trainable parameters. Comprehensive experimental results, including evaluations on public academic benchmarks for natural language understanding, commonsense reasoning, and image-text understanding, as well as real-world industrial text Ads relevance datasets, demonstrate that MTL-LoRA outperforms LoRA and its various variants with comparable or even fewer learnable parameters in MTL setting. Yaming Yang 0001, Dilxat Muhtar, Yelong Shen, Yuefeng Zhan, Yujing Wang 0002, Hao Sun 0015, Feng Sun 0008, Qi Zhang 0066, Weizhu Chen, Yunhai Tong |
AAAI | 1 |
| 2025 | Token-level Proximal Policy Optimization for Query GenerationabstractYichen Ouyang, Lu Wang, Fangkai Yang, Pu Zhao, Chenghua Huang, Jianfeng Liu, Bochen Pang, Yaming Yang, Yuefeng Zhan, Hao Sun, Qingwei Lin, Saravan Rajmohan, Weiwei Deng, Dongmei Zhang, Feng Sun. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Yichen Ouyang, Lu Wang 0029, Fangkai Yang, Pu Zhao 0004, Chenghua Huang, Bochen Pang, Yaming Yang 0001, Yuefeng Zhan, Hao Sun 0015, Qingwei Lin, Saravan Rajmohan, Dongmei Zhang 0001, Feng Sun 0008 |
EMNLP | 8 |
| 2025 | Direct Preference Optimization for LLM-Enhanced Recommendation SystemsabstractLarge Language Models (LLMs) have exhibited remarkable performance across a wide range of domains, motivating research into their potential for recommendation systems. Early efforts have leveraged LLMs’ rich knowledge and strong generalization capabilities via in-context learning, where recommendation tasks are framed as prompts. However, LLM performance in recommendation scenarios remains limited due to the mismatch between their pretraining objectives and recommendation tasks, as well as the lack of recommendation-specific data during pretraining. To address these challenges, we propose DPO4Rec, a novel framework that integrates Direct Preference Optimization (DPO) into LLM-enhanced recommendation systems. First, we prompt the LLM to infer user preferences from historical interactions, which are then used to augment traditional ID-based sequential recommendation models. Next, we train a reward model based on knowledge-augmented recommendation architectures to assess the quality of LLM-generated reasoning. Using this, we select the highest- and lowest-ranked responses from N samples to construct a dataset for LLM fine-tuning. Finally, we apply a structure alignment strategy via DPO to align the LLM’s outputs with desirable recommendation behavior. Extensive experiments show that DPO4Rec significantly improves re-ranking performance over strong baselines, demonstrating enhanced instruction-following capabilities of LLMs in recommendation tasks. Yaobo Liang, Yaming Yang 0001, Shilin Xu 0001, Tianmeng Yang, Yunhai Tong |
ICME | 3 |
| 2024 | You Can't Ignore Either: Unifying Structure and Feature Denoising for Robust Graph LearningabstractRecent research on the robustness of Graph Neural Networks (GNNs) under noises or attacks has attracted great attention due to its importance in real-world applications. Most previous methods explore a single noise source, recovering corrupt node embedding by reliable structures bias or developing structure learning with reliable node features. However, the noises and attacks may come from both structures and features in graphs, making the graph denoising a dilemma and challenging problem. In this paper, we develop a unified graph denoising (UGD) framework to unravel the deadlock between structure and feature denoising. Specifically, a high-order neighborhood proximity evaluation method is proposed to recognize noisy edges, considering features may be perturbed simultaneously. Moreover, we propose to refine noisy features with reconstruction based on a graph auto-encoder. An iterative updating algorithm is further designed to optimize the framework and acquire a clean graph, thus enabling robust graph learning for downstream tasks. Our UGD framework is self-supervised and can be easily implemented as a plug-and-play module. We carry out extensive experiments, which proves the effectiveness and advantages of our method. Code is avalaible at https://github.com/YoungTimmy/UGD. Tianmeng Yang, Jiahao Meng, Min Zhou 0006, Yaming Yang 0001, Yujing Wang 0002, Xiangtai Li, Yunhai Tong |
CIKM | 4 |
| 2024 | Collaborative Multi-Task Representation for Natural Language UnderstandingabstractMulti-task learning has shown large benefits in Natural Language Understanding (NLU). However, current state-of-the-arts (SOTAs) like MT-DNN and MMoE do not model task relationships explicitly and fail to obtain effective task alignment. In this paper, we propose a Collaborative Multi-Task Representation (CMTR) framework to tackle this problem. We capture instance-level task relations through a task interaction layer, which helps guide the fusion of task-oriented representations into the final representation. Moreover, tailored loss functions are proposed to facilitate the learning of task alignment. Specifically, we leverage knowledge distillation as an auxiliary loss to assist the adaptation layers in generating task-oriented representations. We also introduce a regularization loss to learn better gating functions for multi-task fusion. Empirically, CMTR outperforms SOTA multi-task learning frameworks on most natural language understanding tasks in the GLUE benchmark. Furthermore, it achieves better task alignment and demonstrates good interpretability. Yaming Yang 0001, Defu Cao, Ming Zeng 0009, Jing Yu 0007, Yunhai Tong, Yujing Wang 0002 |
IJCNN | 1 |
| 2024 | Characteristic-Aware Time-Series Representation Learning for Unsupervised Anomaly DetectionabstractTime-series anomaly detection is an important research topic in data mining, popular in both academia and industry. Recently, unsupervised anomaly detection draws considerable attention, since it can detect anomalies without parameter tuning on labels and meets the demands of industrial applications. Time-series representation learning plays a vital role in addressing unsupervised anomaly detection. However, it remains challenging to learn a unified representation model with diverse distributions and handle multivariate times-series with various features.To alleviate these challenges, we propose a novel representation strategy, termed CAT-AD, for unsupervised time-series anomaly detection. It learns characteristic-aware priors for representations of time-series by incorporating embeddings of broad characteristics and is capable of handling diverse anomaly detection tasks, regardless of their lengths and dimensions.Our proposed strategy is simple yet effective, which has been verified on two univariate datasets and five multivariate datasets from public sources. Yaming Yang 0001, Pingping Lin, Juanyong Duan, Tianmeng Yang, Congrui Huang, Zhengjie Lin, Yunhai Tong |
IJCNN | 1 |
| 2023 | MMDialog: A Large-scale Multi-turn Dialogue Dataset Towards Multi-modal Open-domain ConversationabstractJiazhan Feng, Qingfeng Sun, Can Xu, Pu Zhao, Yaming Yang, Chongyang Tao, Dongyan Zhao, Qingwei Lin. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Jiazhan Feng, Qingfeng Sun, Can Xu 0002, Pu Zhao 0004, Yaming Yang 0001, Chongyang Tao, Dongyan Zhao 0001, Qingwei Lin |
ACL (1) | 5 |
| 2023 | Multiple Connectivity Views for Session-based RecommendationabstractSession-based recommendation (SBR), which makes the next-item recommendation based on previous anonymous actions, has drawn increasing attention. The last decade has seen multiple deep learning-based modeling choices applied on SBR successfully, e.g., recurrent neural networks (RNNs), convolutional neural networks (CNNs), graph neural networks (GNNs), and each modeling choice has its intrinsic superiority and limitation. We argue that these modeling choices differentiate from each other by (1) the way they capture the interactions between items within a session and (2) the operators they adopt for composing the neural network, e.g., convolutional operator or self-attention operator. Yaming Yang 0001, Jieyu Zhang 0001, Yujing Wang 0005, Zheng Miao, Yunhai Tong |
RecSys | 1 |
| 2023 | Convolution-Enhanced Evolving Attention NetworksabstractAttention-based neural networks, such as Transformers, have become ubiquitous in numerous applications, including computer vision, natural language processing, and time-series analysis. In all kinds of attention networks, the attention maps are crucial as they encode semantic dependencies between input tokens. However, most existing attention networks perform modeling or reasoning based on representations, wherein the attention maps of different layers are learned separately without explicit interactions. In this paper, we propose a novel and generic evolving attention mechanism, which directly models the evolution of inter-token relationships through a chain of residual convolutional modules. The major motivations are twofold. On the one hand, the attention maps in different layers share transferable knowledge, thus adding a residual connection can facilitate the information flow of inter-token relationships across layers. On the other hand, there is naturally an evolutionary trend among attention maps at different abstraction levels, so it is beneficial to exploit a dedicated convolution-based module to capture this process. Equipped with the proposed mechanism, the convolution-enhanced evolving attention networks achieve superior performance in various applications, including time-series representation, natural language understanding, machine translation, and image classification. Especially on time-series representation tasks, Evolving Attention-enhanced Dilated Convolutional (EA-DC-) Transformer outperforms state-of-the-art models significantly, achieving an average of 17% improvement compared to the best SOTA. To the best of our knowledge, this is the first work that explicitly models the layer-wise evolution of attention maps. Our implementation is available at https://github.com/pkuyym/EvolvingAttention. Yujing Wang 0002, Yaming Yang 0001, Jiangang Bai, Mingliang Zhang 0004, Xiangtai Li, Jing Yu 0007, Ce Zhang 0001, Gao Huang 0001, Yunhai Tong |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | Graph Pointer Neural NetworksabstractGraph Neural Networks (GNNs) have shown advantages in various graph-based applications. Most existing GNNs assume strong homophily of graph structure and apply permutation-invariant local aggregation of neighbors to learn a representation for each node. However, they fail to generalize to heterophilic graphs, where most neighboring nodes have different labels or features, and the relevant nodes are distant. Few recent studies attempt to address this problem by combining multiple hops of hidden representations of central nodes (i.e., multi-hop-based approaches) or sorting the neighboring nodes based on attention scores (i.e., ranking-based approaches). As a result, these approaches have some apparent limitations. On the one hand, multi-hop-based approaches do not explicitly distinguish relevant nodes from a large number of multi-hop neighborhoods, leading to a severe over-smoothing problem. On the other hand, ranking-based models do not joint-optimize node ranking with end tasks and result in sub-optimal solutions. In this work, we present Graph Pointer Neural Networks (GPNN) to tackle the challenges mentioned above. We leverage a pointer network to select the most relevant nodes from a large amount of multi-hop neighborhoods, which constructs an ordered sequence according to the relationship with the central node. 1D convolution is then applied to extract high-level features from the node sequence. The pointer-network-based ranker in GPNN is joint-optimized with other parts in an end-to-end manner. Extensive experiments are conducted on six public node classification datasets with heterophilic graphs. The results show that GPNN significantly improves the classification performance of state-of-the-art methods. In addition, analyses also reveal the privilege of the proposed GPNN in filtering out irrelevant neighbors and reducing over-smoothing. Tianmeng Yang, Zhihan Yue, Yaming Yang 0001, Yunhai Tong, Jing Bai 0010 |
AAAI | 4 |
| 2022 | Multimodal Dialogue Response GenerationabstractQingfeng Sun, Yujing Wang, Can Xu, Kai Zheng, Yaming Yang, Huang Hu, Fei Xu, Jessica Zhang, Xiubo Geng, Daxin Jiang. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Qingfeng Sun, Can Xu 0002, Kai Zheng 0021, Yaming Yang 0001, Huang Hu, Xiubo Geng, Daxin Jiang |
ACL (1) | 5 |
| 2022 | Binary Classification with Positive Labeling SourcesabstractTo create a large amount of training labels for machine learning models effectively and efficiently, researchers have turned to Weak Supervision (WS), which uses programmatic labeling sources rather than manual annotation. Existing works of WS for binary classification typically assume the presence of labeling sources that are able to assign both positive and negative labels to data in roughly balanced proportions. However, for many tasks of interest where there is a minority positive class, negative examples could be too diverse for developers to generate indicative labeling sources. Thus, in this work, we study the application of WS on binary classification tasks with positive labeling sources only. We propose WEAPO, a simple yet competitive WS method for producing training labels without negative labeling sources. On 10 benchmark datasets, we show WEAPO achieves the highest averaged performance in terms of both the quality of synthesized labels and the performance of the final classifier supervised with these labels. We incorporated the implementation of WEAPO into WRENCH, an existing benchmarking platform. Jieyu Zhang 0001, Yujing Wang 0005, Yaming Yang 0001, Alexander Ratner |
CIKM | 3 |
| 2022 | Privacy-preserving Online AutoML for Domain-Specific Face DetectionabstractDespite the impressive progress of general face detection, the tuning of hyper-parameters and architectures is still critical for the performance of a domain-specific face detector. Though existing AutoML works can speedup such process, they either require tuning from scratch for a new scenario or do not consider data privacy. To scale up, we derive a new AutoML setting from a platform perspective. In such setting, new datasets sequentially arrive at the platform, where an architecture and hyper-parameter configuration is recommended to train the optimal face detector for each dataset. This, however, brings two major challenges: (1) how to predict the best configuration for any given dataset without touching their raw images due to the privacy concern? and (2) how to continuously improve the AutoML algorithm from previous tasks and offer a better warm-up for future ones? We introduce “HyperFD”, a new privacy-preserving online AutoML framework for face detection. At its core part, a novel meta-feature representation of a dataset as well as its learning paradigm is proposed. Thanks to HyperFD, each local task (client) is able to effectively leverage the learning “experience” of previous tasks without uploading raw images to the platform; meanwhile, the meta-feature extractor is continuously learned to better trade off the bias and variance. Extensive experiments demonstrate the effectiveness and efficiency of our design. Chenqian Yan, Yuge Zhang, Quanlu Zhang, Yaming Yang 0001, Xinyang Jiang, Yuqing Yang 0001, Baoyuan Wang |
CVPR | 4 |
| 2022 | Attentive Knowledge-aware Graph Convolutional Networks with Collaborative Guidance for Personalized RecommendationabstractTo alleviate data sparsity and cold-start problems of traditional recommender systems (RSs), incorporating knowledge graphs (KGs) to supplement auxiliary information has attracted considerable attention recently. However, simply integrating KGs in current KG-based RS models is not necessarily a guarantee to improve the recommendation performance, which may even weaken the holistic model capability. This is because the construction of these KGs is independent of the collection of historical user-item interactions; hence, information in these KGs may not always be helpful for recommendation to all users. In this paper, we propose attentive Knowledge-aware Graph convolutional networks with Collaborative Guidance for personalized Recommendation (CG-KGR). CG-KGR is a novel knowledge-aware recommendation model that enables ample and coherent learning of KGs and user-item interactions, via our proposed Collaborative Guidance Mechanism. Specifically, CG-KGR first encapsulates historical interactions to interactive information summarization. Then CG-KGR utilizes it as guidance to extract information out of KGs, which eventually provides more precise personalized recommendation. We conduct extensive experiments on four real-world datasets over two recommendation tasks, i.e., Top-K recommendation and Click-Through rate (CTR) prediction. The experimental results show that the CG-KGR model significantly outperforms recent state-of-the-art models by 1.4-27.0% in terms of Recall metric on Top-K recommendation. Yankai Chen 0001, Yaming Yang 0001, Jing Bai 0010, Xiangchen Song, Irwin King |
ICDE | 2 |
| 2022 | Creating Training Sets via Weak Indirect Supervision
Jieyu Zhang 0001, Xiangchen Song, Yujing Wang 0005, Yaming Yang 0001, Jing Bai 0010, Alexander Ratner |
ICLR | 5 |
| 2022 | Enhancing Self-Attention with Knowledge-Assisted Attention MapsabstractJiangang Bai, Yujing Wang, Hong Sun, Ruonan Wu, Tianmeng Yang, Pengfei Tang, Defu Cao, Mingliang Zhang1, Yunhai Tong, Yaming Yang, Jing Bai, Ruofei Zhang, Hao Sun, Wei Shen. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Jiangang Bai, Yujing Wang 0002, Ruonan Wu, Tianmeng Yang, Defu Cao, Mingliang Zhang 0004, Yunhai Tong, Yaming Yang 0001, Jing Bai 0010, Ruofei Zhang, Hao Sun 0015 |
NAACL-HLT | 10 |
| 2022 | Learning Multi-granularity Consecutive User Intent Unit for Session-based RecommendationabstractSession-based recommendation aims to predict a user's next action based on previous actions in the current session. The major challenge is to capture authentic and complete user preferences in the entire session. Recent work utilizes graph structure to represent the entire session and adopts Graph Neural Network (GNN) to encode session information. This modeling choice has been proved to be effective and achieved remarkable results. However, most of the existing studies only consider each item within the session independently and do not capture session semantics from a high-level perspective. Such limitation often leads to severe information loss and increases the difficulty of capturing long-range dependencies within a session. Intuitively, compared with individual items, a session snippet, i.e., a group of locally consecutive items, is able to provide supplemental user intents which are hardly captured by existing methods. In this work, we propose to learn multi-granularity consecutive user intent unit to improve the recommendation performance. Specifically, we creatively propose Multi-granularity Intent Heterogeneous Session Graph (MIHSG) which captures the interactions between different granularity intent units and relieves the burden of long-dependency. Moreover, we propose the Intent Fusion Ranking (IFR) module to compose the recommendation results from various granularity user intents. Compared with current methods that only leverage intents from individual items, IFR benefits from different granularity user intents to generate more accurate and comprehensive session representation, thus eventually boosting recommendation performance. We conduct extensive experiments on five session-based recommendation datasets and the results demonstrate the effectiveness of our method. Compared to current state-of-the-art methods, we achieve as large as 10.21% gain on [email protected] and 15.53% gain on [email protected] Jiayan Guo, Yaming Yang 0001, Xiangchen Song, Yuan Zhang 0024, Yujing Wang 0002, Jing Bai 0010, Yan Zhang 0004 |
WSDM | 2 |
| 2021 | Syntax-BERT: Improving Pre-trained Transformers with Syntax TreesabstractJiangang Bai, Yujing Wang, Yiren Chen, Yaming Yang, Jing Bai, Jing Yu, Yunhai Tong. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021. Jiangang Bai, Yujing Wang 0002, Yaming Yang 0001, Jing Bai 0010, Jing Yu 0007, Yunhai Tong |
EACL | 4 |
| 2021 | Evolving Attention with Residual ConvolutionsabstractTransformer is a ubiquitous model for natural language processing and has attracted wide attentions in computer vision. The attention maps are indispensable for a transformer model to encode the dependencies among input tokens. However, they are learned independently in each layer and sometimes fail to capture precise patterns. In this paper, we propose a novel and generic mechanism based on evolving attention to improve the performance of transformers. On one hand, the attention maps in different layers share common knowledge, thus the ones in preceding layers can instruct the attention in succeeding layers through residual connections. On the other hand, low-level and high-level attentions vary in the level of abstraction, so we adopt convolutional layers to model the evolutionary process of attention maps. The proposed evolving attention mechanism achieves significant performance improvement over various state-of-the-art models for multiple tasks, including image classification, natural language understanding and machine translation. Yujing Wang 0002, Yaming Yang 0001, Jiangang Bai, Mingliang Zhang 0004, Jing Bai 0010, Jing Yu 0007, Ce Zhang 0001, Gao Huang 0001, Yunhai Tong |
ICML | 2 |
| 2021 | DeGNN: Improving Graph Neural Networks with Graph DecompositionabstractMining from graph-structured data is an integral component of graph data management. A recent trending technique, graph convolutional network (GCN), has gained momentum in the graph mining field, and plays an essential part in numerous graph-related tasks. Although the emerging GCN optimization techniques bring improvements to specific scenarios, they perform diversely in different applications and introduce many trial-and-error costs for practitioners. Moreover, existing GCN models often suffer from oversmoothing problem. Besides, the entanglement of various graph patterns could lead to non-robustness and harm the final performance of GCNs. In this work, we propose a simple yet efficient graph decomposition approach to improve the performance of general graph neural networks. We first empirically study existing graph decomposition methods and propose an automatic connectivity-ware graph decomposition algorithm, DeGNN. To provide a theoretical explanation, we then characterize GCN from the information-theoretic perspective and show that under certain conditions, the mutual information between the output after l layers and the input of GCN converges to 0 exponentially with respect to l. On the other hand, we show that graph decomposition can potentially weaken the condition of such convergence rate, alleviating the information loss when GCN becomes deeper. Extensive experiments on various academic benchmarks and real-world production datasets demonstrate that graph decomposition generally boosts the performance of GNN models. Moreover, our proposed solution DeGNN achieves state-of-the-art performances on almost all these tasks. Xupeng Miao, Nezihe Merve Gürel, Wentao Zhang 0001, Zhichao Han 0001, Bo Li 0026, Wei Min, Susie Xi Rao, Hansheng Ren, Yinan Shan, Yingxia Shao, Fan Wu 0011, Hui Xue 0004, Yaming Yang 0001, Zitao Zhang, Shuai Zhang 0007, Yujing Wang 0002, Bin Cui 0001, Ce Zhang 0001 |
KDD | 14 |
| 2020 | TextNAS: A Neural Architecture Search Space Tailored for Text RepresentationabstractLearning text representation is crucial for text classification and other language related tasks. There are a diverse set of text representation networks in the literature, and how to find the optimal one is a non-trivial problem. Recently, the emerging Neural Architecture Search (NAS) techniques have demonstrated good potential to solve the problem. Nevertheless, most of the existing works of NAS focus on the search algorithms and pay little attention to the search space. In this paper, we argue that the search space is also an important human prior to the success of NAS in different applications. Thus, we propose a novel search space tailored for text representation. Through automatic search, the discovered network architecture outperforms state-of-the-art models on various public datasets on text classification and natural language inference tasks. Furthermore, some of the design principles found in the automatic network agree well with human intuition. Yujing Wang 0002, Yaming Yang 0001, Jing Bai 0010, Ce Zhang 0001, Guinan Su, Xiaoyu Kou, Yunhai Tong, Mao Yang 0004, Lidong Zhou |
AAAI | 2 |
| 2020 | AutoADR: Automatic Model Design for Ad RelevanceabstractLarge-scale pre-trained models have attracted extensive attention in the research community and shown promising results on various tasks of natural language processing. However, these pre-trained models are memory and computation intensive, hindering their deployment into industrial online systems like Ad Relevance. Meanwhile, how to design an effective yet efficient model architecture is another challenging problem in online Ad Relevance. Recently, AutoML shed new lights on architecture design, but how to integrate it with pre-trained language models remains unsettled. In this paper, we propose AutoADR (Automatic model design for AD Relevance) --- a novel end-to-end framework to address this challenge, and share our experience to ship these cutting-edge techniques into online Ad Relevance system at Microsoft Bing. Specifically, AutoADR leverages a one-shot neural architecture search algorithm to find a tailored network architecture for Ad Relevance. The search process is simultaneously guided by knowledge distillation from a large pre-trained teacher model (e.g. BERT), while taking the online serving constraints (e.g. memory and latency) into consideration. We add the model designed by AutoADR as a sub-model into the production Ad Relevance model. This additional sub-model improves the Precision-Recall AUC (PR AUC) on top of the original Ad Relevance model by 2.65X of the normalized shipping bar. More importantly, adding this automatically designed sub-model leads to a statistically significant 4.6% Bad-Ad ratio reduction in online A/B testing. This model has been shipped into Microsoft Bing Ad Relevance Production model. Yaming Yang 0001, Yujing Wang 0002, Yunhai Tong, Jing Bai 0010, Ruofei Zhang |
CIKM | 2 |
| 2020 | LadaBERT: Lightweight Adaptation of BERT through Hybrid Model CompressionabstractBERT is a cutting-edge language representation model pre-trained by a large corpus, which achieves superior performances on various natural language understanding tasks. However, a major blocking issue of applying BERT to online services is that it is memory-intensive and leads to unsatisfactory latency of user requests, raising the necessity of model compression. Existing solutions leverage the knowledge distillation framework to learn a smaller model that imitates the behaviors of BERT. However, the training procedure of knowledge distillation is expensive itself as it requires sufficient training data to imitate the teacher model. In this paper, we address this issue by proposing a tailored solution named LadaBERT (Lightweight adaptation of BERT through hybrid model compression), which combines the advantages of different model compression methods, including weight pruning, matrix factorization and knowledge distillation. LadaBERT achieves state-of-the-art accuracy on various public datasets while the training overheads can be reduced by an order of magnitude. Yihuan Mao, Yujing Wang 0002, Chufan Wu, Chen Zhang 0001, Yang Wang 0053, Quanlu Zhang, Yaming Yang 0001, Yunhai Tong, Jing Bai 0010 |
COLING | 7 |