Yile Wang 0001

dblp:32/1915-1 · DBLP profile ↗
← Back
14ranked-venue papers
6as first author
11since 2021 · last 2025
0000-0001-8705-9598ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 6 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021
YearPublicationVenuePosition
2025 Leveraging Language-based Representations for Better Solving Symbol-related Problems with Large Language Models
abstract
Symbols such as numerical sequences, chemical formulas, and table delimiters exist widely, playing important roles in symbol-related tasks such as abstract reasoning, chemical property prediction, and tabular question-answering. Compared to tasks based on natural language expressions, large language models (LLMs) have limitations in understanding and reasoning on symbol-based representations, making it difficult for them to handle symbol-related problems. In this paper, we propose symbol-to-language (S2L), a method that converts symbol-based representations to language-based representations, providing valuable information for language models during reasoning. We found that, for both closed-source and open-source LLMs, the capability to solve symbol-related problems can be largely enhanced by incorporating such language-based representations. For example, by employing S2L for GPT-4, there can be substantial improvements of +21.9% and +9.5% accuracy for 1D-ARC and Dyck language tasks, respectively. There is also a consistent improvement in other six general symbol-related tasks such as table understanding and Tweet analysis. We release the GPT logs in https://github.com/THUNLP-MT/symbol2language.
Yile Wang 0001, Sijie Cheng, Zixin Sun, Peng Li 0030, Yang Liu 0005
COLING1
2025 Pre-Training a Graph Recurrent Network for Text Understanding
abstract
Transformer-based pre-trained models have gained much advance in recent years, Transformer architecture also becomes one of the most important backbones in natural language processing. Recent works show that the attention mechanism inside Transformer may not be necessary, and Transformer alternatives such as convolutional neural networks, multi-layer perceptron, and state space model have also been investigated. Transformer-based models have two main limitations: First, they have quadratic time complexity due to the full attention mechanism, which leads to high computational costs. Second, they rely on representation of a special token such as [CLS] to encode entire text, which limits its sentence-level expressiveness. In this paper, we consider a graph recurrent network with linear time complexity for language model pre-training, which builds a graph structure for each sequence with local token-level communications, together with a sentence-level representation detached from other normal tokens. On both English and Chinese text understanding tasks, our model can achieve comparable performance to existing pre-trained models while also achieving higher inference efficiency. Furthermore, we discovered that the representations generated by our model are more diverse and uniform compared to that of Transformer, which alleviates the problems in existing pre-trained models such as representation degradation.
Yile Wang 0001, Linyi Yang, Zhiyang Teng, Ming Zhou 0001, Yue Zhang 0004
IEEE Trans. Pattern Anal. Mach. Intell.1
2024 Reasoning in Conversation: Solving Subjective Tasks through Dialogue Simulation for Large Language Models
abstract
Xiaolong Wang, Yile Wang, Yuanchi Zhang, Fuwen Luo, Peng Li, Maosong Sun, Yang Liu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Xiaolong Wang 0014, Yile Wang 0001, Yuanchi Zhang, Fuwen Luo, Peng Li 0030, Maosong Sun 0001, Yang Liu 0005
ACL (1)2
2024 Enhancing Multilingual Capabilities of Large Language Models through Self-Distillation from Resource-Rich Languages
abstract
Yuanchi Zhang, Yile Wang, Zijun Liu, Shuo Wang, Xiaolong Wang, Peng Li, Maosong Sun, Yang Liu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Yuanchi Zhang, Yile Wang 0001, Shuo Wang 0013, Xiaolong Wang 0014, Peng Li 0030, Maosong Sun 0001, Yang Liu 0005
ACL (1)2
2024 DEEM: Dynamic Experienced Expert Modeling for Stance Detection
abstract
Recent work has made a preliminary attempt to use large language models (LLMs) to solve the stance detection task, showing promising results. However, considering that stance detection usually requires detailed background knowledge, the vanilla reasoning method may neglect the domain knowledge to make a professional and accurate analysis. Thus, there is still room for improvement of LLMs reasoning, especially in leveraging the generation capability of LLMs to simulate specific experts (i.e., multi-agents) to detect the stance. In this paper, different from existing multi-agent works that require detailed descriptions and use fixed experts, we propose a Dynamic Experienced Expert Modeling (DEEM) method which can leverage the generated experienced experts and let LLMs reason in a semi-parametric way, making the experts more generalizable and reliable. Experimental results demonstrate that DEEM consistently achieves the best results on three standard benchmarks, outperforms methods with self-consistency reasoning, and reduces the bias of LLMs.
Xiaolong Wang 0014, Yile Wang 0001, Sijie Cheng, Peng Li 0030, Yang Liu 0005
LREC/COLING2
2024 Position: Towards Unified Alignment Between Agents, Humans, and Environment
abstract
The rapid progress of foundation models has led to the prosperity of autonomous agents, which leverage the universal capabilities of foundation models to conduct reasoning, decision-making, and environmental interaction. However, the efficacy of agents remains limited when operating in intricate, realistic environments. In this work, we introduce the principles of Unified Alignment for Agents (UA$^2$), which advocate for the simultaneous alignment of agents with human intentions, environmental dynamics, and self-constraints such as the limitation of monetary budgets. From the perspective of UA$^2$, we review the current agent research and highlight the neglected factors in existing agent benchmarks and method candidates. We also conduct proof-of-concept studies by introducing realistic features to WebShop, including user profiles demonstrating intentions, personalized reranking reflecting complex environmental dynamics, and runtime cost statistics as self-constraints. We then follow the principles of UA$^2$ to propose an initial design of our agent and benchmark its performance with several candidate baselines in the retrofitted WebShop. The extensive experimental results further prove the importance of the principles of UA$^2$. Our research sheds light on the next steps of autonomous agent research with improved general problem-solving abilities.
Zonghan Yang, Kaiming Liu, Fangzhou Xiong, Yile Wang 0001, Zeyuan Yang 0002, Zhenhe Zhang, Fuwen Luo, Zhicheng Guo, Peng Li 0030, Yang Liu 0005
ICML6
2024 Lost in Context? On the Sense-Wise Variance of Contextualized Word Embeddings
abstract
Contextualized word embeddings in language models have given much advance to NLP. Intuitively, sentential information is integrated into the representation of words, which can help model polysemy. However, context sensitivity also leads to the variance of representations, which may break the semantic consistency for synonyms. Previous works that investigate contextualized sensitivity focus on thetokenlevel representations, while we are taking a deeper dive into exploring representations at the fine-grainedsenselevel. In particular, we quantify how much the contextualized embeddings of each word sense vary across contexts in typical pre-trained models, the results show that contextualized embeddings can be highly consistent across contexts, even for two different words with the same sense. In addition, part-of-speech, number of word senses, and sentence length have an influence on the variance of sense representations. Interestingly, we find that word representations are position-biased, where the first words in different contexts tend to be more similar. We analyze such a phenomenon and also propose a prompt-augmentation method to alleviate such bias in distance-based word sense disambiguation settings. Finally, we investigate the influence of sense-level pre-training on the performance of different downstream tasks, results show that such external tasks can improve the sense- and syntactic-related tasks, while not necessarily benefiting general language understanding tasks.
Yile Wang 0001, Yue Zhang 0004
IEEE ACM Trans. Audio Speech Lang. Process.1
2024 Gradual Syntactic Label Replacement for Language Model Pre-Training
abstract
Pre-training serves as a foundation of recent NLP models, where language modeling tasks are performed over large texts. Typical models like BERT and GPT take the corpus as a whole and treat each word equally for language modeling. However, recent works show that the naturally existing frequency bias in the raw corpus may limit the power of the language model. In this article, we propose a multi-stage training strategy that gradually increases the training vocabulary by modifying the training data. Specifically, we leverage the syntactic structure as a bridge for infrequent words and replace them with the corresponding syntactic labels, then we recover their original lexical surface for further training. Such strategy results in an easy-to-hard curriculum learning process, where the model learns the most common words and some basic syntax concepts, before recognizing a large number of uncommon words via their specific usages and the previously learned category knowledge. Experimental results show that such a method can improve the performance of both discriminative and generative pre-trained language models on benchmarks and various downstream tasks.
Yile Wang 0001, Yue Zhang 0004, Peng Li 0030, Yang Liu 0005
IEEE ACM Trans. Audio Speech Lang. Process.1
2023 CVT-SLR: Contrastive Visual-Textual Transformation for Sign Language Recognition with Variational Alignment
abstract
Sign language recognition (SLR) is a weakly supervised task that annotates sign videos as textual glosses. Recent studies show that insufficient training caused by the lack of large-scale available sign datasets becomes the main bottleneck for SLR. Most SLR works thereby adopt pretrained visual modules and develop two mainstream solutions. The multi-stream architectures extend multi-cue visual features, yielding the current SOTA performances but requiring complex designs and might introduce potential noise. Alternatively, the advanced single-cue SLR frameworks using explicit cross-modal alignment between visual and textual modalities are simple and effective, potentially competitive with the multi-cue framework. In this work, we propose a novel contrastive visual-textual transformation for SLR, CVT-SLR, to fully explore the pretrained knowledge of both the visual and language modalities. Based on the single-cue cross-modal alignment framework, we propose a variational autoencoder (VAE) for pretrained contextual knowledge while introducing the complete pretrained language module. The VAE implicitly aligns visual and textual modalities while benefiting from pretrained contextual knowledge as the traditional contextual module. Meanwhile, a contrastive cross-modal alignment algorithm is designed to explicitly enhance the consistency constraints. Extensive experiments on public datasets (PHOENIX-2014 and PHOENIX-2014T) demonstrate that our proposed CVT-SLR consistently outperforms existing single-cue methods and even outperforms SOTA multi-cue methods. The source codes and models are available at https://github.com/binbinjiang/CVT-SLR.
Jiangbin Zheng 0002, Yile Wang 0001, Cheng Tan 0012, Siyuan Li 0002, Jun Xia 0001, Yidong Chen 0001, Stan Z. Li
CVPR2
2022 Using Context-to-Vector with Graph Retrofitting to Improve Word Embeddings
abstract
Jiangbin Zheng, Yile Wang, Ge Wang, Jun Xia, Yufei Huang, Guojiang Zhao, Yue Zhang, Stan Li. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Jiangbin Zheng 0002, Yile Wang 0001, Jun Xia 0001, Yufei Huang 0002, Guojiang Zhao, Yue Zhang 0004, Stan Z. Li
ACL (1)2
2021 Improving Skip-Gram Embeddings Using BERT
abstract
Contextualized embeddings such as BERT and GPT have been shown to give significant improvement in NLP tasks. On the other hand, static embeddings such as skip-gram and GloVe still have desirable characteristics such as low computational cost, easy deployment and freedom from severe contextualized variation in representation. There has been some recent attempt enhancing the skip-gram model by adding syntactic information of context using GCN. We investigate the use of BERT embeddings instead for stronger context representation, which contains not only syntactic and surface features, but also rich knowledge from large-scale pre-training. Results show that BERT-enhanced skip-gram embeddings outperform GCN-enhanced embeddings on a range of tasks. Such embeddings also outperform recent effort distilling BERT embeddings into context-independent vectors.
Yile Wang 0001, Leyang Cui, Yue Zhang 0004
IEEE ACM Trans. Audio Speech Lang. Process.1
2020 Does Chinese BERT Encode Word Structure?
abstract
Contextualized representations give significantly improved results for a wide range of NLP tasks.Much work has been dedicated to analyzing the features captured by representative models such as BERT.Existing work finds that syntactic, semantic and word sense knowledge are encoded in BERT.However, little work has investigated word features for character-based languages such as Chinese.We investigate Chinese BERT using both attention weight distribution statistics and probing tasks, finding that (1) word information is captured by BERT; (2) word-level features are mostly in the middle representation layers; (3) downstream tasks make different use of word features in BERT, with POS tagging and chunking relying the most on word features, and natural language inference relying the least on such features.
Yile Wang 0001, Leyang Cui, Yue Zhang 0004
COLING1
2020 LogiQA: A Challenge Dataset for Machine Reading Comprehension with Logical Reasoning
abstract
Machine reading is a fundamental task for testing the capability of natural language understand- ing, which is closely related to human cognition in many aspects. With the rising of deep learning techniques, algorithmic models rival human performances on simple QA, and thus increasingly challenging machine reading datasets have been proposed. Though various challenges such as evidence integration and commonsense knowledge have been integrated, one of the fundamental capabilities in human reading, namely logical reasoning, is not fully investigated. We build a comprehensive dataset, named LogiQA, which is sourced from expert-written questions for testing human Logical reasoning. It consists of 8,678 QA instances, covering multiple types of deductive reasoning. Results show that state-of-the-art neural models perform by far worse than human ceiling. Our dataset can also serve as a benchmark for reinvestigating logical AI under the deep learning NLP setting. The dataset is freely available at https://github.com/lgw863/LogiQA-dataset.
Jian Liu 0030, Leyang Cui, Hanmeng Liu, Dandan Huang, Yile Wang 0001, Yue Zhang 0004
IJCAI5
2020 Lattice LSTM for Chinese Sentence Representation
abstract
Words provide a useful source of information for Chinese NLP, and word segmentation has been taken as a pre-processing step for most downstream tasks. For many NLP tasks, however, word segmentation can introduce noise and lead to error propagation. The rise of neural representation learning models allows sentence-level semantic information to be collected from characters directly. As a result, it is an empirical question whether a fully character-based model should be used instead of first performing word segmentation. We investigate a neural representation that simultaneously encodes character and word information without the need for segmentation. In particular, candidate words are found in a sentence by matching with a pre-defined lexicon. A lattice structured LSTM is used to encode the resulting word-character lattice, where gate vectors are used to control information flow through words, so that the more useful words can be automatically identified by end-to-end training. We compare the performance of the resulting lattice LSTM and baseline sequence LSTM structures over both character sequences and automatically segmented word sequences. Results on NER show that the character-word lattice model can significantly improve the performance. In addition, as a general sentence representation architecture, character-word lattice LSTM can also be used for learning contextualized representations. To this end, we compare lattice LSTM structure with its sequential LSTM counterpart, namely ELMo. Results show that our lattice version of ELMo gives better language modeling performances. On Chinese POS-tagging, chunking and syntactic parsing tasks, the resulting contextualized Chinese embeddings also give better performance than ELMo trained on the same data.
Yue Zhang 0004, Yile Wang 0001, Jie Yang 0039
IEEE ACM Trans. Audio Speech Lang. Process.2