Tao Wang 0056

dblp:12/5838-56 · DBLP profile ↗
← Back
11ranked-venue papers
0as first author
9since 2021 · last 2022
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 9 since 2021
YearPublicationVenuePosition
2022 ITA: Image-Text Alignments for Multi-Modal Named Entity Recognition
abstract
Xinyu Wang, Min Gui, Yong Jiang, Zixia Jia, Nguyen Bach, Tao Wang, Zhongqiang Huang, Kewei Tu. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Xinyu Wang 0013, Min Gui, Yong Jiang 0005, Zixia Jia, Nguyen Bach, Tao Wang 0056, Zhongqiang Huang, Kewei Tu
NAACL-HLT6
2021 Multi-View Cross-Lingual Structured Prediction with Minimum Supervision
abstract
Zechuan Hu, Yong Jiang, Nguyen Bach, Tao Wang, Zhongqiang Huang, Fei Huang, Kewei Tu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Zechuan Hu, Yong Jiang 0005, Nguyen Bach, Tao Wang 0056, Zhongqiang Huang, Fei Huang 0002, Kewei Tu
ACL/IJCNLP (1)4
2021 Risk Minimization for Zero-shot Sequence Labeling
abstract
Zechuan Hu, Yong Jiang, Nguyen Bach, Tao Wang, Zhongqiang Huang, Fei Huang, Kewei Tu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Zechuan Hu, Yong Jiang 0005, Nguyen Bach, Tao Wang 0056, Zhongqiang Huang, Fei Huang 0002, Kewei Tu
ACL/IJCNLP (1)4
2021 Improving Named Entity Recognition by External Context Retrieving and Cooperative Learning
abstract
Xinyu Wang, Yong Jiang, Nguyen Bach, Tao Wang, Zhongqiang Huang, Fei Huang, Kewei Tu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Xinyu Wang 0013, Yong Jiang 0005, Nguyen Bach, Tao Wang 0056, Zhongqiang Huang, Fei Huang 0002, Kewei Tu
ACL/IJCNLP (1)4
2021 Automated Concatenation of Embeddings for Structured Prediction
abstract
Xinyu Wang, Yong Jiang, Nguyen Bach, Tao Wang, Zhongqiang Huang, Fei Huang, Kewei Tu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Xinyu Wang 0013, Yong Jiang 0005, Nguyen Bach, Tao Wang 0056, Zhongqiang Huang, Fei Huang 0002, Kewei Tu
ACL/IJCNLP (1)4
2021 Structural Knowledge Distillation: Tractably Distilling Information for Structured Predictor
abstract
Xinyu Wang, Yong Jiang, Zhaohui Yan, Zixia Jia, Nguyen Bach, Tao Wang, Zhongqiang Huang, Fei Huang, Kewei Tu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Xinyu Wang 0013, Yong Jiang 0005, Zhaohui Yan 0001, Zixia Jia, Nguyen Bach, Tao Wang 0056, Zhongqiang Huang, Fei Huang 0002, Kewei Tu
ACL/IJCNLP (1)6
2021 Word Reordering for Zero-shot Cross-lingual Structured Prediction
abstract
Adapting word order from one language to another is a key problem in cross-lingual structured prediction.Current sentence encoders (e.g., RNN, Transformer with position embeddings) are usually word order sensitive.Even with uniform word form representations (MUSE, mBERT), word order discrepancies may hurt the adaptation of models.This paper builds structured prediction models with bag-of-words inputs.It introduces a new reordering module to organize words following the source language order, which learns taskspecific reordering strategies from a generalpurpose order predictor model.Experiments on zero-shot cross-lingual dependency parsing, POS tagging, and morphological tagging show that our model can significantly improve target language performances, especially for languages that are distant from the source language.1
Yong Jiang 0005, Tao Wang 0056, Zhongqiang Huang, Fei Huang 0002, Yuanbin Wu
EMNLP (1)3
2021 A Unified Encoding of Structures in Transition Systems
abstract
Transition systems usually contain various dynamic structures (e.g., stacks, buffers).An ideal transition-based model should encode these structures completely and efficiently.Previous works relying on templates or neural network structures either only encode partial structure information or suffer from computation efficiency.In this paper, we propose a novel attention-based encoder unifying representation of all structures in a transition system.Specifically, we separate two views of items on structures, namely structure-invariant view and structure-dependent view.With the help of parallel-friendly attention network, we are able to encoding transition states with O(1) additional complexity (with respect to basic feature extractors).Experiments on the PTB and UD show that our proposed method significantly improves the test speed and achieves the best transition-based model, and is comparable to state-of-the-art methods. 1
Yong Jiang 0005, Tao Wang 0056, Zhongqiang Huang, Fei Huang 0002, Yuanbin Wu
EMNLP (1)3
2021 MuVER: Improving First-Stage Entity Retrieval with Multi-View Entity Representations
abstract
Entity retrieval, which aims at disambiguating mentions to canonical entities from massive KBs, is essential for many tasks in natural language processing.Recent progress in entity retrieval shows that the dual-encoder structure is a powerful and efficient framework to nominate candidates if entities are only identified by descriptions.However, they ignore the property that meanings of entity mentions diverge in different contexts and are related to various portions of descriptions, which are treated equally in previous works.In this work, we propose Multi-View Entity Representations (MuVER), a novel approach for entity retrieval that constructs multi-view representations for entity descriptions and approximates the optimal view for mentions via a heuristic searching method.Our method achieves the state-ofthe-art performance on ZESHEL and improves the quality of candidates on three standard Entity Linking datasets 1 .
Xinyin Ma, Yong Jiang 0005, Nguyen Bach, Tao Wang 0056, Zhongqiang Huang, Fei Huang 0002, Weiming Lu 0001
EMNLP (1)4
2020 Structure-Level Knowledge Distillation For Multilingual Sequence Labeling
abstract
Multilingual sequence labeling is a task of predicting label sequences using a single unified model for multiple languages.Compared with relying on multiple monolingual models, using a multilingual model has the benefit of a smaller model size, easier in online serving, and generalizability to low-resource languages.However, current multilingual models still underperform individual monolingual models significantly due to model capacity limitations.In this paper, we propose to reduce the gap between monolingual models and the unified multilingual model by distilling the structural knowledge of several monolingual models (teachers) to the unified multilingual model (student).We propose two novel KD methods based on structure-level information:(1) approximately minimizes the distance between the student's and the teachers' structurelevel probability distributions, (2) aggregates the structure-level knowledge to local distributions and minimizes the distance between two local probability distributions.Our experiments on 4 multilingual tasks with 25 datasets show that our approaches outperform several strong baselines and have stronger zero-shot generalizability than both the baseline model and teacher models.
Xinyu Wang 0013, Yong Jiang 0005, Nguyen Bach, Tao Wang 0056, Fei Huang 0002, Kewei Tu
ACL4
2020 AIN: Fast and Accurate Sequence Labeling with Approximate Inference Network
abstract
The linear-chain Conditional Random Field (CRF) model is one of the most widely-used neural sequence labeling approaches.Exact probabilistic inference algorithms such as the forward-backward and Viterbi algorithms are typically applied in training and prediction stages of the CRF model.However, these algorithms require sequential computation that makes parallelization impossible.In this paper, we propose to employ a parallelizable approximate variational inference algorithm for the CRF model.Based on this algorithm, we design an approximate inference network that can be connected with the encoder of the neural CRF model to form an end-to-end network, which is amenable to parallelization for faster training and prediction.The empirical results show that our proposed approaches achieve a 12.7-fold improvement in decoding speed with long sentences and a competitive accuracy compared with the traditional CRF approach.
Xinyu Wang 0013, Yong Jiang 0005, Nguyen Bach, Tao Wang 0056, Zhongqiang Huang, Fei Huang 0002, Kewei Tu
EMNLP (1)4