VLDB 2026 Research / reviewers in the wild / expert
Dianhai Yu
dblp:118/7157
· DBLP profile ↗
25ranked-venue papers
1as first author
16since 2021 · last 2026
0000-0002-0163-2603ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 10 since 2021Databases, data management, data science and information retrieval · 4 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RRAtention: Dynamic Block Sparse Attention via Per-Head Round-Robin Shifts for Long-Context InferenceabstractThe quadratic complexity of attention mechanisms poses a critical bottleneck for large language models processing long contexts. While dynamic sparse attention methods offer input-adaptive efficiency, they face fundamental trade-offs: requiring preprocessing, lacking global evaluation, violating query independence, or incurring high computational overhead. We present RRAttention, a novel dynamic sparse attention method that simultaneously achieves all desirable properties through a head round-robin (RR) sampling strategy. By rotating query sampling positions across attention heads within each stride, RRAttention maintains query independence while enabling efficient global pattern discovery with stride-level aggregation. Our method reduces complexity from O(L^2) to O(L^2/S^2) and employs adaptive Top-\tau selection for optimal sparsity. Extensive experiments on natural language understanding (HELMET) and multimodal video comprehension (Video-MME) demonstrate that RRAttention recovers over 99% of full attention performance while computing only half of the attention blocks, achieving 2.4\times speedup at 128K context length and outperforming existing dynamic sparse attention methods. The code is available at https://github.com/PaddlePaddle/PaddleFleet (see ‘Research/RRAttention‘). Siran Liu, Guoxia Wang, Sa Wang, Jinle Zeng, Haoyang Xie, Siyu Lou, Jiabin Yang, Dianhai Yu |
ACL (1) | 8 |
| 2025 | FlashMask: Efficient and Rich Mask Extension of FlashAttentionabstractThe computational and memory demands of vanilla attention scale quadratically with the sequence length $N$, posing significant challenges for processing long sequences in Transformer models. FlashAttention alleviates these challenges by eliminating the $\mathcal{O}(N^2)$ memory dependency and reducing attention latency through IO-aware memory optimizations. However, its native support for certain attention mask types is limited, and it does not inherently accommodate more complex masking requirements. Previous approaches resort to using dense masks with $\mathcal{O}(N^2)$ memory complexity, leading to inefficiencies. In this paper, we propose \ours{}, an extension of FlashAttention that introduces a column-wise sparse representation of attention masks. This approach efficiently represents a wide range of mask types and facilitates the development of optimized kernel implementations. By adopting this novel representation, \ours{} achieves linear memory complexity $\mathcal{O}(N)$, making it suitable for modeling long-context sequences. Moreover, this representation enables kernel optimizations that eliminate unnecessary computations by leveraging sparsity in the attention mask, without sacrificing computational accuracy, resulting in higher computational efficiency. We evaluate \ours{}'s performance in fine-tuning and alignment training of LLMs such as SFT, LoRA, DPO, and RM. \ours{} achieves significant throughput improvements, with end-to-end speedups ranging from 1.65x to 3.22x compared to existing FlashAttention dense method. Additionally, our kernel-level comparisons demonstrate that \ours{} surpasses the latest counterpart, FlexAttention, by 12.1\% to 60.7\% in terms of kernel TFLOPs/s, achieving 37.8\% to 62.3\% of the theoretical maximum FLOPs/s on the A100 GPU. The code is open-sourced on PaddlePaddle\footnote{\url{https://github.com/PaddlePaddle/Paddle}} and integrated into PaddleNLP\footnote{\url{https://github.com/PaddlePaddle/PaddleNLP}}, supporting models with over 100 billion parameters for contexts extending up to 128K tokens. Guoxia Wang, Jinle Zeng, Xiyuan Xiao, Siming Wu, Jiabin Yang, Lujing Zheng, Dianhai Yu |
ICLR | 9 |
| 2024 | NACL: A General and Effective KV Cache Eviction Framework for LLM at Inference TimeabstractYilong Chen, Guoxia Wang, Junyuan Shang, Shiyao Cui, Zhenyu Zhang, Tingwen Liu, Shuohuan Wang, Yu Sun, Dianhai Yu, Hua Wu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Guoxia Wang, Junyuan Shang, Shiyao Cui, Zhenyu Zhang 0006, Tingwen Liu, Shuohuan Wang, Dianhai Yu, Hua Wu 0003 |
ACL (1) | 9 |
| 2024 | Spectral Heterogeneous Graph Convolutions via Positive Noncommutative PolynomialsabstractHeterogeneous Graph Neural Networks (HGNNs) have gained significant popularity in various heterogeneous graph learning tasks. However, most existing HGNNs rely on spatial domain-based methods to aggregate information, i.e., manually selected meta-paths or some heuristic modules, lacking theoretical guarantees. Furthermore, these methods cannot learn arbitrary valid heterogeneous graph filters within the spectral domain, which have limited expressiveness. To tackle these issues, we present a positive spectral heterogeneous graph convolution via positive noncommutative polynomials. Then, using this convolution, we propose PSHGCN, a novel Positive Spectral Heterogeneous Graph Convolutional Network. PSHGCN offers a simple yet effective method for learning valid heterogeneous graph filters. Moreover, we demonstrate the rationale of PSHGCN in the graph optimization framework. We conducted an extensive experimental study to show that PSHGCN can learn diverse heterogeneous graph filters and outperform all baselines on open benchmarks. Notably, PSHGCN exhibits remarkable scalability, efficiently handling large real-world graphs comprising millions of nodes and edges. Our codes are available at https://github.com/ivam-he/PSHGCN. Mingguo He, Zhewei Wei, Shikun Feng, Zhengjie Huang, Weibin Li 0004, Yu Sun 0029, Dianhai Yu |
WWW | 7 |
| 2024 | Exploiting Cross-Modal Prediction and Relation Consistency for Semisupervised Image CaptioningabstractThe task of image captioning aims to generate captions directly from images via the automatically learned cross-modal generator. To build a well-performing generator, existing approaches usually need a large number of described images (i.e., supervised image-sentence pairs), requiring a huge effects on manual labeling. However, in real-world applications, a more general scenario is that we only have limited amount of described images and a large number of undescribed images. Therefore, a resulting challenge is how to effectively combine the undescribed images into the learning of cross-modal generator (i.e., semisupervised image captioning). To solve this problem, we propose a novel image captioning method by exploiting the cross-modal prediction and relation consistency (CPRC), which aims to utilize the raw image input to constrain the generated sentence in the semantic space. In detail, considering that the heterogeneous gap between modalities always leads to the supervision difficulty while using the global embedding directly, CPRC turns to transform both the raw image and corresponding generated sentence into the shared semantic space, and measure the generated sentence from two aspects: 1) prediction consistency: CPRC utilizes the prediction of raw image as soft label to distill useful supervision for the generated sentence, rather than employing the traditional pseudo labeling and 2) relation consistency: CPRC develops a novel relation consistency between augmented images and corresponding generated sentences to retain the important relational knowledge. In result, CPRC supervises the generated sentence from both the informativeness and representativeness perspectives, and can reasonably use the undescribed images to learn a more effective generator under the semisupervised scenario. The experiments show that our method outperforms state-of-the-art comparison methods on the MS-COCO "Karpathy" offline test split under complex nonparallel scenarios, for example, CPRC achieves at least 6% improvements on the CIDEr-D score. Yang Yang 0074, Hongchen Wei, Hengshu Zhu, Dianhai Yu, Hui Xiong 0001, Jian Yang 0003 |
IEEE Trans. Cybern. | 4 |
| 2024 | MoESys: A Distributed and Efficient Mixture-of-Experts Training and Inference System for Internet ServicesabstractWhile modern internet services, such as chatbots, search engines, and online advertising, demand the use of large-scale deep neural networks (DNNs), distributed training and inference over heterogeneous computing systems are desired to facilitate these DNN models. Mixture-of-Experts (MoE) is one the most common strategies to lower the cost of training subject to the overall size of models/data through gating and parallelism in a divide-and-conquer fashion. While DeepSpeed [1] has made efforts in carrying out large-scale MoE training over heterogeneous infrastructures, the efficiency of training and inference could be further improved from several system aspects, including load balancing, communication/computation efficiency, and memory footprint limits. In this work, we present a novel MoESys that boosts efficiency in both large-scale training and inference. Specifically, in the training procedure, the proposed MoESys adopts an Elastic MoE training strategy with 2D prefetch and Fusion communication over Hierarchical storage, so as to enjoy efficient parallelisms. For scalable inference in a single node, especially when the model size is larger than GPU memory, MoESys builds the CPU-GPU memory jointly into a ring of sections to load the model, and executes the computation tasks across the memory sections in a round-robin manner for efficient inference. We carried out extensive experiments to evaluate MoESys, where MoESys successfully trains a Unified Feature Optimization [2] (UFO) model with a Sparsely-Gated Mixture-of-Experts model of 12B parameters in 8 days on 48 A100 GPU cards. The comparison against the state-of-the-art shows that MoESys outperformed DeepSpeed with 33 unbalanced MoE Tasks, e.g., UFO, MoESys achieved 64 Dianhai Yu, Hongxiang Hao, Weibao Gong, HuaChao Wu, Jiang Bian 0003, Li-Rong Dai 0001, Haoyi Xiong |
IEEE Trans. Serv. Comput. | 1 |
| 2023 | PGLBox: Multi-GPU Graph Learning Framework for Web-Scale RecommendationabstractWhile having been used widely for large-scale recommendation and online advertising, the Graph Neural Network (GNN) has demonstrated its representation learning capacity to extract embeddings of nodes and edges through passing, transforming, and aggregating information over the graph. In this work, we propose PGLBox1 - a multi-GPU graph learning framework based on PaddlePaddle [24], incorporating with optimized storage, computation, and communication strategies, to train deep GNNs based on web-scale graphs for the recommendation. Specifically, PGLBox adopts a hierarchical storage system with three layers to facilitate I/O, where graphs and embeddings are stored in the HBMs and SSDs, respectively, with MEMs as the cache. To fully utilize multi-GPUs and I/O bandwidth, PGLBox proposes an asynchronous pipeline with three stages - it first samples the subgraphs from the input graph, then pulls & updates embeddings and trains GNNs on the subgraph with parameters updating queued at the end of the pipeline. Thanks to the capacity of PGLBox in handling web-scale graphs, it becomes feasible to unify the view of GNN-based recommendation tasks for multiple advertising verticals and fuse all these graphs into a unified yet huge one. We evaluate PGLBox using a bucket of realistic GNN training tasks for the recommendation, and compare the performance of PGLBox on top of a multi-GPU server (Tesla A100×8) and the legacy training system based on a 40-node MPI cluster at Baidu. The overall comparisons show that PGLBox could save up to 55% monetary cost for training GNN models, and achieve up to 14× training speedup with the same accuracy as the legacy trainer. The open-source implementation of PGLBox is available at https://github.com/PaddlePaddle/PGL/tree/main/apps/PGLBox. Xuewu Jiao, Weibin Li 0004, Xinxuan Wu, Jiang Bian 0003, Siming Dai, Xinsheng Luo, Mingqing Hu, Zhengjie Huang, Danlei Feng, Junchao Yang 0001, Shikun Feng, Haoyi Xiong, Dianhai Yu, Shuanglong Li, Jingzhou He, Yanjun Ma |
KDD | 15 |
| 2023 | Label Information Enhanced Fraud Detection against Low Homophily in GraphsabstractNode classification is a substantial problem in graph-based fraud detection. Many existing works adopt Graph Neural Networks (GNNs) to enhance fraud detectors. While promising, currently most GNN-based fraud detectors fail to generalize to the low homophily setting. Besides, label utilization has been proved to be significant factor for node classification problem. But we find they are less effective in fraud detection tasks due to the low homophily in graphs. In this work, we propose GAGA, a novel Group AGgregation enhanced TrAnsformer, to tackle the above challenges. Specifically, the group aggregation provides a portable method to cope with the low homophily issue. Such an aggregation explicitly integrates the label information to generate distinguishable neighborhood information. Along with group aggregation, an attempt towards end-to-end trainable group encoding is proposed which augments the original feature space with the class labels. Meanwhile, we devise two additional learnable encodings to recognize the structural and relational context. Then, we combine the group aggregation and the learnable encodings into a Transformer encoder to capture the semantic information. Experimental results clearly show that GAGA outperforms other competitive graph-based fraud detectors by up to 24.39% on two trending public datasets and a real-world industrial dataset from Baidu. Even more, the group aggregation is demonstrated to outperform other label utilization methods (e.g., C&S, BoT/UniMP) in the low homophily setting. Jinghui Zhang 0001, Zhengjie Huang, Weibin Li 0004, Shikun Feng, Ziheng Ma, Yu Sun 0029, Dianhai Yu, Fang Dong 0001, Jiahui Jin 0001, Beilun Wang, Junzhou Luo |
WWW | 8 |
| 2023 | Large-scale knowledge distillation with elastic heterogeneous computing resourcesabstractAbstract Although more layers and more parameters generally improve the accuracy of the models, such big models generally have high computational complexity and require big memory, which exceed the capacity of small devices for inference and incurs long training time. In addition, it is difficult to afford long training time and inference time of big models even in high performance servers, as well. As an efficient approach to compress a large deep model (a teacher model) to a compact model (a student model), knowledge distillation emerges as a promising approach to deal with the big models. Existing knowledge distillation methods cannot exploit the elastic available computing resources and correspond to low efficiency. In this paper, we propose an Elastic Deep Learning framework for knowledge Distillation, that is, EDL‐Dist. The advantages of EDL‐Dist are threefold. First, the inference and the training process is separated. Second, elastic available computing resources can be utilized to improve the efficiency. Third, fault‐tolerance of the training and inference processes is supported. We take extensive experimentation to show that the throughput of EDL‐Dist is up to 3.125 times faster than the baseline method (online knowledge distillation) while the accuracy is similar or higher. Ji Liu 0003, Daxiang Dong, An Qin 0001, Xingjian Li 0002, Patrick Valduriez, Dejing Dou, Dianhai Yu |
Concurr. Comput. Pract. Exp. | 8 |
| 2023 | HeterPS: Distributed deep learning with reinforcement learning based scheduling in heterogeneous environments
Ji Liu 0003, Danlei Feng, Minxu Zhang, Xinxuan Wu, Xuefeng Yao, Dianhai Yu, Yanjun Ma, Dejing Dou |
Future Gener. Comput. Syst. | 7 |
| 2023 | Graph neural networks meet with distributed graph partitioners and reconciliations
Zongshen Mu, Siliang Tang, Chang Zong, Dianhai Yu, Yueting Zhuang |
Neurocomputing | 4 |
| 2023 | Attribute-driven streaming edge partitioning with reconciliations for distributed graph neural network training
Zongshen Mu, Siliang Tang, Yueting Zhuang, Dianhai Yu |
Neural Networks | 4 |
| 2022 | Simple and Effective Relation-based Embedding Propagation for Knowledge Representation LearningabstractRelational graph neural networks have garnered particular attention to encode graph context in knowledge graphs (KGs). Although they achieved competitive performance on small KGs, how to efficiently and effectively utilize graph context for large KGs remains an open problem. To this end, we propose the Relation-based Embedding Propagation (REP) method. It is a post-processing technique to adapt pre-trained KG embeddings with graph context. As relations in KGs are directional, we model the incoming head context and the outgoing tail context separately. Accordingly, we design relational context functions with no external parameters. Besides, we use averaging to aggregate context information, making REP more computation-efficient. We theoretically prove that such designs can avoid information distortion during propagation. Extensive experiments also demonstrate that REP has significant scalability while improving or maintaining prediction quality. Particularly, it averagely brings about 10% relative improvement to triplet-based embedding methods on OGBL-WikiKG2 and takes 5%-83% time to achieve comparable results as the state-of-the-art GC-OTE. Siming Dai, Weiyue Su, Zeyang Fang, Zhengjie Huang, Shikun Feng, Yu Sun 0029, Dianhai Yu |
IJCAI | 10 |
| 2022 | mmLayout: Multi-grained MultiModal Transformer for Document UnderstandingabstractRecent efforts of multimodal Transformers have improved Visually Rich Document Understanding (VrDU) tasks via incorporating visual and textual information. However, existing approaches mainly focus on fine-grained elements such as words and document image patches, making it hard for them to learn from coarse-grained elements, including natural lexical units like phrases and salient visual regions like prominent image regions. In this paper, we attach more importance to coarse-grained elements containing high-density information and consistent semantics, which are valuable for document understanding. At first, a document graph is proposed to model complex relationships among multi-grained multimodal elements, in which salient visual regions are detected by a cluster-based method. Then, a multi-grained multimodal Transformer called mmLayout is proposed to incorporate coarse-grained information into existing pre-trained fine-grained multimodal Transformers based on the graph. In mmLayout, coarse-grained information is aggregated from fine-grained, and then, after further processing, is fused back into fine-grained for final prediction. Furthermore, common sense enhancement is introduced to exploit the semantic information of natural lexical units. Experimental results on four tasks, including information extraction and document question answering, show that our method can improve the performance of multimodal Transformers based on fine-grained elements and achieve better performance with fewer parameters. Qualitative analyses show that our method can capture consistent semantics in coarse-grained elements. Wenjin Wang 0003, Zhengjie Huang, Qianglong Chen, Qiming Peng, Yinxu Pan, Weichong Yin, Shikun Feng, Yu Sun 0029, Dianhai Yu, Yin Zhang 0006 |
ACM Multimedia | 10 |
| 2022 | TA-MoE: Topology-Aware Large Scale Mixture-of-Expert TrainingabstractSparsely gated Mixture-of-Expert (MoE) has demonstrated its effectiveness in scaling up deep neural networks to an extreme scale. Despite that numerous efforts have been made to improve the performance of MoE from the model design or system optimization perspective, existing MoE dispatch patterns are still not able to fully exploit the underlying heterogeneous network environments. In this paper, we propose TA-MoE, a topology-aware routing strategy for large-scale MoE trainging, from a model-system co-design perspective, which can dynamically adjust the MoE dispatch pattern according to the network topology. Based on communication modeling, we abstract the dispatch problem into an optimization objective and obtain the approximate dispatch pattern under different topologies. On top of that, we design a topology-aware auxiliary loss, which can adaptively route the data to fit in the underlying topology without sacrificing the model accuracy. Experiments show that TA-MoE can substantially outperform its counterparts on various hardware and model configurations, with roughly 1.01x-1.61x, 1.01x-4.77x, 1.25x-1.54x improvements over the popular DeepSpeed-MoE, FastMoE and FasterMoE systems. Dianhai Yu |
NeurIPS | 4 |
| 2021 | Rethinking Label-Wise Cross-Modal Retrieval from A Semantic Sharing PerspectiveabstractThe main challenge of cross-modal retrieval is to learn the consistent embedding for heterogeneous modalities. To solve this problem, traditional label-wise cross-modal approaches usually constrain the inter-modal and intra-modal embedding consistency relying on the label ground-truths. However, the experiments reveal that different modal networks actually have various generalization capacities, thereby end-to-end joint training with consistency loss usually leads to sub-optimal uni-modal model, which in turn affects the learning of consistent embedding. Therefore, in this paper, we argue that what really needed for supervised cross-modal retrieval is a good shared classification model. In other words, we learn the consistent embedding by ensuring the classification performance of each modality on the shared model, without the consistency loss. Specifically, we consider a technique called Semantic Sharing, which directly trains the two modalities interactively by adopting a shared self-attention based classification model. We evaluate the proposed approach on three representative datasets. The results validate that the proposed semantic sharing can consistently boost the performance under NDCG metric. Yang Yang 0074, Chubing Zhang, Yi-Chu Xu, Dianhai Yu, De-Chuan Zhan, Jian Yang 0003 |
IJCAI | 4 |
| 2019 | RLTM: An Efficient Neural IR Framework for Long DocumentsabstractDeep neural networks have achieved significant improvements in information retrieval (IR). However, most existing models are computational costly and can not efficiently scale to long documents. This paper proposes a novel End-to-End neural ranking framework called Reinforced Long Text Matching (RLTM) which matches a query with long documents efficiently and effectively. The core idea behind the framework can be analogous to the human judgment process which firstly locates the relevance parts quickly from the whole document and then matches these parts with the query carefully to obtain the final label. Firstly, we select relevant sentences from the long documents by a coarse and efficient matching model. Secondly, we generate a relevance score by a more sophisticated matching model based on the sentence selected. The whole model is trained jointly with reinforcement learning in a pairwise manner by maximizing the expected score gaps between positive and negative examples. Experimental results demonstrate that RLTM has greatly improved the efficiency and effectiveness of the states-of-the-art models. Chen Zheng 0006, Shengxian Wan, Dianhai Yu |
IJCAI | 4 |
| 2018 | Multi-Turn Response Selection for Chatbots with Deep Attention Matching NetworkabstractXiangyang Zhou, Lu Li, Daxiang Dong, Yi Liu, Ying Chen, Wayne Xin Zhao, Dianhai Yu, Hua Wu. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2018. Xiangyang Zhou, Daxiang Dong, Ying Chen 0011, Wayne Xin Zhao, Dianhai Yu, Hua Wu 0003 |
ACL (1) | 7 |
| 2018 | A New Method of Region Embedding for Text Classification
Chao Qiao, Guocheng Niu, Daren Li, Daxiang Dong, Wei He 0014, Dianhai Yu, Hua Wu 0003 |
ICLR (Poster) | 7 |
| 2016 | Multi-view Response Selection for Human-Computer Conversation
Xiangyang Zhou, Daxiang Dong, Hua Wu 0003, Dianhai Yu, Hao Tian 0005, Rui Yan 0001 |
EMNLP | 5 |
| 2015 | Multi-Task Learning for Multiple Language TranslationabstractDaxiang Dong, Hua Wu, Wei He, Dianhai Yu, Haifeng Wang. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Daxiang Dong, Hua Wu 0003, Wei He 0014, Dianhai Yu, Haifeng Wang 0001 |
ACL (1) | 4 |
| 2014 | Improve Statistical Machine Translation with Context-Sensitive Bilingual Semantic Embedding ModelabstractWe investigate how to improve bilingual embedding which has been successfully used as a feature in phrase-based statistical machine translation (SMT). Despite bilingual embedding’s success, the contextual information, which is of critical importance to translation quality, was ignored in previous work. To employ the contextual information, we propose a simple and memory-efficient model for learning bilingual embedding, taking both the source phrase and context around the phrase into account. Bilingual translation scores generated from our proposed bilingual embedding model are used as features in our SMT system. Experimental results show that the proposed method achieves significant improvements on large-scale Chinese-English translation task. Daxiang Dong, Xiaoguang Hu, Dianhai Yu, Wei He 0014, Hua Wu 0003, Haifeng Wang 0001, Ting Liu 0001 |
EMNLP | 4 |
| 2013 | Compound Embedding Features for Semi-supervised Learning
Mo Yu, Tiejun Zhao, Daxiang Dong, Dianhai Yu |
HLT-NAACL | 5 |
| 2012 | Enlister: baidu's recommender system for the biggest chinese Q&A websiteabstractIn this paper, we describe the concept & design of a real-time question RS (recommender system), the Enlister project, for the biggest Chinese Q&A (Questions and Answers) website and evaluate its performance on massive data from this real-world practice. We demonstrate how we weigh in among different recommendation algorithms and optimization methods. To enhance recommendation accuracy and handling time-sensitive questions, we propose a large scale real-time RS based on the combination of machine learning algorithms and the stream computing technology. Considering of algorithm flexibility and performance, we use the maximum entropy model as the fundamental model design in the CTR (click-through rate) prediction of recommendation items. In the perspective of the Enlister system architecture, we illustrate how we divide and conquer massive data processing problem with a novel stream computing design which reduces the data process latency down to seconds. Finally we analyze the online test result and prove our design concept by achieving a series of significant improvements. Qiwen Liu, Tianjian Chen, Dianhai Yu |
RecSys | 4 |
| 2008 | An Improved CRF based Chinese Language Processing System for SIGHAN Bakeoff 2007
Xihong Wu, Xiaojun Lin 0002, Chunyao Wu, Dianhai Yu |
IJCNLP | 6 |