Zhiyang Teng

dblp:136/8660 · DBLP profile ↗
← Back
26ranked-venue papers
4as first author
13since 2021 · last 2025
0000-0001-6757-6897ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 24 · 4 first-author · 12 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Pre-Training a Graph Recurrent Network for Text Understanding
abstract
Transformer-based pre-trained models have gained much advance in recent years, Transformer architecture also becomes one of the most important backbones in natural language processing. Recent works show that the attention mechanism inside Transformer may not be necessary, and Transformer alternatives such as convolutional neural networks, multi-layer perceptron, and state space model have also been investigated. Transformer-based models have two main limitations: First, they have quadratic time complexity due to the full attention mechanism, which leads to high computational costs. Second, they rely on representation of a special token such as [CLS] to encode entire text, which limits its sentence-level expressiveness. In this paper, we consider a graph recurrent network with linear time complexity for language model pre-training, which builds a graph structure for each sequence with local token-level communications, together with a sentence-level representation detached from other normal tokens. On both English and Chinese text understanding tasks, our model can achieve comparable performance to existing pre-trained models while also achieving higher inference efficiency. Furthermore, we discovered that the representations generated by our model are more diverse and uniform compared to that of Transformer, which alleviates the problems in existing pre-trained models such as representation degradation.
Yile Wang 0001, Linyi Yang, Zhiyang Teng, Ming Zhou 0001, Yue Zhang 0004
IEEE Trans. Pattern Anal. Mach. Intell.3
2024 Refining Latent Homophilic Structures over Heterophilic Graphs for Robust Graph Convolution Networks
abstract
Graph convolution networks (GCNs) are extensively utilized in various graph tasks to mine knowledge from spatial data. Our study marks the pioneering attempt to quantitatively investigate the GCN robustness over omnipresent heterophilic graphs for node classification. We uncover that the predominant vulnerability is caused by the structural out-of-distribution (OOD) issue. This finding motivates us to present a novel method that aims to harden GCNs by automatically learning Latent Homophilic Structures over heterophilic graphs. We term such a methodology as LHS. To elaborate, our initial step involves learning a latent structure by employing a novel self-expressive technique based on multi-node interactions. Subsequently, the structure is refined using a pairwisely constrained dual-view contrastive learning approach. We iteratively perform the above procedure, enabling a GCN model to aggregate information in a homophilic way on heterophilic graphs. Armed with such an adaptable structure, we can properly mitigate the structural OOD threats over heterophilic graphs. Experiments on various benchmarks show the effectiveness of the proposed LHS approach for robust GCNs.
Chenyang Qiu 0001, Guoshun Nan, Tianyu Xiong, Wendi Deng, Di Wang 0011, Zhiyang Teng, Qimei Cui, Xiaofeng Tao 0001
AAAI6
2024 Fast-DetectGPT: Efficient Zero-Shot Detection of Machine-Generated Text via Conditional Probability Curvature
abstract
Large language models (LLMs) have shown the ability to produce fluent and cogent content, presenting both productivity opportunities and societal risks. To build trustworthy AI systems, it is imperative to distinguish between machine-generated and human-authored content. The leading zero-shot detector, DetectGPT, showcases commendable performance but is marred by its intensive computational costs. In this paper, we introduce the concept of **conditional probability curvature** to elucidate discrepancies in word choices between LLMs and humans within a given context. Utilizing this curvature as a foundational metric, we present **Fast-DetectGPT**, an optimized zero-shot detector, which substitutes DetectGPT's perturbation step with a more efficient sampling step. Our evaluations on various datasets, source models, and test conditions indicate that Fast-DetectGPT not only surpasses DetectGPT by a relative around 75\% in both the white-box and black-box settings but also accelerates the detection process by a factor of 340, as detailed in Table 1.
Guangsheng Bao, Yanbin Zhao, Zhiyang Teng, Linyi Yang, Yue Zhang 0004
ICLR3
2024 Exploring Self-supervised Logic-enhanced Training for Large Language Models
abstract
Fangkai Jiao, Zhiyang Teng, Bosheng Ding, Zhengyuan Liu, Nancy Chen, Shafiq Joty. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Fangkai Jiao, Zhiyang Teng, Bosheng Ding, Zhengyuan Liu, Nancy F. Chen, Shafiq R. Joty
NAACL-HLT2
2024 HSDirSniper: A New Attack Exploiting Vulnerabilities in Tor's Hidden Service Directories
Zhiyang Teng, Yue Gao 0003, Qingyun Liu 0001, Jinqiao Shi
WWW2
2023 Target-Side Augmentation for Document-Level Machine Translation
abstract
Document-level machine translation faces the challenge of data sparsity due to its long input length and a small amount of training data, increasing the risk of learning spurious patterns.To address this challenge, we propose a target-side augmentation method, introducing a data augmentation (DA) model to generate many potential translations for each source document.Learning on these wider range translations, an MT model can learn a smoothed distribution, thereby reducing the risk of data sparsity.We demonstrate that the DA model, which estimates the posterior distribution, largely improves the MT performance, outperforming the previous best system by 2.30 s-BLEU on News and achieving new state-of-the-art on News and Europarl benchmarks.Our code is available at https://github.com/baoguangsheng/ target-side-augmentation.
Guangsheng Bao, Zhiyang Teng, Yue Zhang 0004
ACL (1)2
2023 LogiQA 2.0 - An Improved Dataset for Logical Reasoning in Natural Language Understanding
abstract
NLP research on logical reasoning regains momentum with the recent releases of a handful of datasets, notably LogiQA and Reclor. Logical reasoning is exploited in many probing tasks over large Pre-trained Language Models (PLMs) and downstream tasks like question-answering and dialogue systems. In this paper, we release LogiQA 2.0. The dataset is an amendment and re-annotation of LogiQA in 2020, a large-scale logical reasoning reading comprehension dataset adapted from the Chinese Civil Service Examination. We increase the data size, refine the texts with manual translation by professionals, and improve the quality by removing items with distinctive cultural features like Chinese idioms. Furthermore, we conduct a fine-grained annotation on the dataset and turn it into a two-way natural language inference (NLI) task, resulting in 35k premise-hypothesis pairs with gold labels, making it the first large-scale NLI dataset for complex logical reasoning. Compared to Question Answering, Natural Language Inference excels in generalizability and helps downstream tasks better. We establish a baseline for logical reasoning in NLI and incite further research.
Hanmeng Liu, Jian Liu 0030, Leyang Cui, Zhiyang Teng, Nan Duan 0001, Ming Zhou 0001, Yue Zhang 0004
IEEE ACM Trans. Audio Speech Lang. Process.4
2022 Discrete Opinion Tree Induction for Aspect-based Sentiment Analysis
abstract
Dependency trees have been intensively used with graph neural networks for aspect-based sentiment classification.Though being effective, such methods rely on external dependency parsers, which can be unavailable for low-resource languages or perform worse in low-resource domains.In addition, dependency trees are also not optimized for aspect-based sentiment classification.In this paper, we propose an aspect-specific and language-agnostic discrete latent opinion tree model as an alternative structure to explicit dependency trees.To ease the learning of complicated structured latent variables, we build a connection between aspect-to-context attention scores and syntactic distances, inducing trees from the attention scores.Results on six English benchmarks, one Chinese dataset and one Korean dataset show that our model can achieve competitive performance and interpretability.
Chenhua Chen, Zhiyang Teng, Yue Zhang 0004
ACL (1)2
2022 METS-CoV: A Dataset of Medical Entity and Targeted Sentiment on COVID-19 Related Tweets
abstract
The COVID-19 pandemic continues to bring up various topics discussed or debated on social media. In order to explore the impact of pandemics on people's lives, it is crucial to understand the public's concerns and attitudes towards pandemic-related entities (e.g., drugs, vaccines) on social media. However, models trained on existing named entity recognition (NER) or targeted sentiment analysis (TSA) datasets have limited ability to understand COVID-19-related social media texts because these datasets are not designed or annotated from a medical perspective. In this paper, we release METS-CoV, a dataset containing medical entities and targeted sentiments from COVID-19 related tweets. METS-CoV contains 10,000 tweets with 7 types of entities, including 4 medical entity types (Disease, Drug, Symptom, and Vaccine) and 3 general entity types (Person, Location, and Organization). To further investigate tweet users' attitudes toward specific entities, 4 types of entities (Person, Organization, Drug, and Vaccine) are selected and annotated with user sentiments, resulting in a targeted sentiment dataset with 9,101 entities (in 5,278 tweets). To the best of our knowledge, METS-CoV is the first dataset to collect medical entities and corresponding sentiments of COVID-19 related tweets. We benchmark the performance of classical machine learning models and state-of-the-art deep learning models on NER and TSA tasks with extensive experiments. Results show that this dataset has vast room for improvement for both NER and TSA tasks. With rich annotations and comprehensive benchmark results, we believe METS-CoV is a fundamental resource for building better medical social media understanding tools and facilitating computational social science research, especially on epidemiological topics. Our data, annotation guidelines, benchmark models, and source code are publicly available (\url{https://github.com/YLab-Open/METS-CoV}) to ensure reproducibility.
Peilin Zhou, Zeqiang Wang, Dading Chong, Zhijiang Guo, Yining Hua, Zichang Su, Zhiyang Teng, Jiageng Wu, Jie Yang 0039
NeurIPS7
2022 Contrastive latent variable models for neural text generation
abstract
Deep latent variable models such as variational autoencoders and energy-based models are widely used for neural text generation. Most of them focus on matching the prior distribution with the posterior distribution of the latent variable for text reconstruction. In addition to instance-level reconstruction, this paper aims to integrate contrastive learning in the latent space, forcing the latent variables to learn high-level semantics by exploring inter-instance relationships. Experiments on various text generation benchmarks show the effectiveness of our proposed method. We also empirically show that our method can mitigate the posterior collapse issue for latent variable based text generation models.
Zhiyang Teng, Chenhua Chen, Yan Zhang 0004, Yue Zhang 0004
UAI1
2022 SemGloVe: Semantic Co-Occurrences for GloVe From BERT
abstract
GloVe learns word embeddings by leveraging statistical information from word co-occurrence matrices. However, word pairs in the matrices are extracted from a predefined local context window, which might lead to limited word pairs and potentially semantic irrelevant word pairs. In this paper, we proposeSemGloVe, which distillssemantic co-occurrencesfrom BERT into static GloVe word embeddings. Particularly, we propose two models to extract co-occurrence statistics based on either the masked language model or the multi-head attention weights of BERT. Our methods can extract word pairs limited by the local window assumption, and can define the co-occurrence weights by directly considering the semantic distance between word pairs. Experiments on several word similarity datasets and external tasks show that SemGloVe can outperform GloVe.
Leilei Gan, Zhiyang Teng, Yue Zhang 0004, Linchao Zhu, Fei Wu 0001, Yi Yang 0001
IEEE ACM Trans. Audio Speech Lang. Process.2
2021 G-Transformer for Document-Level Machine Translation
abstract
Guangsheng Bao, Yue Zhang, Zhiyang Teng, Boxing Chen, Weihua Luo. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Guangsheng Bao, Yue Zhang 0004, Zhiyang Teng, Boxing Chen, Weihua Luo
ACL/IJCNLP (1)3
2021 Solving Aspect Category Sentiment Analysis as a Text Generation Task
abstract
Aspect category sentiment analysis has attracted increasing research attention.The dominant methods make use of pre-trained language models by learning effective aspect category-specific representations, and adding specific output layers to its pre-trained representation.We consider a more direct way of making use of pre-trained language models, by casting the ACSA tasks into natural language generation tasks, using natural language sentences to represent the output.Our method allows more direct use of pre-trained knowledge in seq2seq language models by directly following the task setting during pre-training.Experiments on several benchmarks show that our method gives the best reported results, having large advantages in few-shot and zero-shot settings.
Jian Liu 0030, Zhiyang Teng, Leyang Cui, Hanmeng Liu, Yue Zhang 0004
EMNLP (1)2
2020 Inducing Target-Specific Latent Structures for Aspect Sentiment Classification
abstract
Aspect-level sentiment analysis aims to recognize the sentiment polarity of an aspect or a target in a comment.Recently, graph convolutional networks based on linguistic dependency trees have been studied for this task.However, the dependency parsing accuracy of commercial product comments or tweets might be unsatisfactory.To tackle this problem, we associate linguistic dependency trees with automatically induced aspectspecific graphs.We propose gating mechanisms to dynamically combine information from word dependency graphs and latent graphs which are learned by self-attention networks.Our model can complement supervised syntactic features with latent semantic dependencies.Experimental results on five benchmarks show the effectiveness of our proposed latent models, giving significantly better results than models without using latent graphs.
Chenhua Chen, Zhiyang Teng, Yue Zhang 0004
EMNLP (1)2
2020 Lightweight, Dynamic Graph Convolutional Networks for AMR-to-Text Generation
abstract
AMR-to-text generation is used to transduce Abstract Meaning Representation structures (AMR) into text.A key challenge in this task is to efficiently learn effective graph representations.Previously, Graph Convolution Networks (GCNs) were used to encode input AMRs, however, vanilla GCNs are not able to capture non-local information and additionally, they follow a local (first-order) information aggregation scheme.To account for these issues, larger and deeper GCN models are required to capture more complex interactions.In this paper, we introduce a dynamic fusion mechanism, proposing Lightweight Dynamic Graph Convolutional Networks (LDGCNs) that capture richer non-local interactions by synthesizing higher order information from the input graphs.We further develop two novel parameter saving strategies based on the group graph convolutions and weight tied convolutions to reduce memory usage and model complexity.With the help of these strategies, we are able to train a model with fewer parameters while maintaining the model capacity.Experiments demonstrate that LDGCNs outperform stateof-the-art models on two benchmark datasets for AMR-to-text generation with significantly fewer parameters.
Yan Zhang 0004, Zhijiang Guo, Zhiyang Teng, Wei Lu 0011, Shay B. Cohen, Zuozhu Liu, Lidong Bing
EMNLP (1)3
2020 Dialogue State Induction Using Neural Latent Variable Models
abstract
Dialogue state modules are a useful component in a task-oriented dialogue system. Traditional methods find dialogue states by manually labeling training corpora, upon which neural models are trained. However, the labeling process can be costly, slow, error-prone, and more importantly, cannot cover the vast range of domains in real-world dialogues for customer service. We propose the task of dialogue state induction, building two neural latent variable models that mine dialogue states automatically from unlabeled customer service dialogue records. Results show that the models can effectively find meaningful dialogue states. In addition, equipped with induced dialogue states, a state-of-the-art dialogue system gives better performance compared with not using a dialogue state module.
Qingkai Min, Libo Qin 0001, Zhiyang Teng, Xiao Liu 0029, Yue Zhang 0004
IJCAI3
2019 Densely Connected Graph Convolutional Networks for Graph-to-Sequence Learning
abstract
We focus on graph-to-sequence learning, which can be framed as transducing graph structures to sequences for text generation. To capture structural information associated with graphs, we investigate the problem of encoding graphs using graph convolutional networks (GCNs). Unlike various existing approaches where shallow architectures were used for capturing local structural information only, we introduce a dense connection strategy, proposing a novel Densely Connected Graph Convolutional Network (DCGCN). Such a deep architecture is able to integrate both local and non-local features to learn a better structural representation of a graph. Our model outperforms the state-of-the-art neural models significantly on AMR-to-text generation and syntax-based neural machine translation.
Zhijiang Guo, Yan Zhang 0004, Zhiyang Teng, Wei Lu 0011
Trans. Assoc. Comput. Linguistics3
2019 DeepClue: Visual Interpretation of Text-Based Deep Stock Prediction
abstract
The recent advance of deep learning has enabled trading algorithms to predict stock price movements more accurately. Unfortunately, there is a significant gap in the real-world deployment of this breakthrough. For example, professional traders in their long-term careers have accumulated numerous trading rules, the myth of which they can understand quite well. On the other hand, deep learning models have been hardly interpretable. This paper presents DeepClue, a system built to bridge text-based deep learning models and end users through visually interpreting the key factors learned in the stock price prediction model. We make three contributions in DeepClue. First, by designing the deep neural network architecture for interpretation and applying an algorithm to extract relevant predictive factors, we provide a useful case on what can be interpreted out of the prediction model for end users. Second, by exploring hierarchies over the extracted factors and displaying these factors in an interactive, hierarchical visualization interface, we shed light on how to effectively communicate the interpreted model to end users. Specially, the interpretation separates the predictables from the unpredictables for stock prediction through the use of intercept model parameters and a risk visualization design. Third, we evaluate the integrated visualization system through two case studies in predicting the stock price with financial news and company-related tweets from social media. Quantitative experiments comparing the proposed neural network architecture with state-of-the-art models and the human baseline are conducted and reported. Feedbacks from an informal user study with domain experts are summarized and discussed in details. The study results demonstrate the effectiveness of DeepClue in helping to complete stock market investment and analysis tasks.
Lei Shi 0002, Zhiyang Teng, Yue Zhang 0004, Alexander Binder
IEEE Trans. Knowl. Data Eng.2
2018 Two Local Models for Neural Constituent Parsing
abstract
Non-local features have been exploited by syntactic parsers for capturing dependencies between sub output structures. Such features have been a key to the success of state-of-the-art statistical parsers. With the rise of deep learning, however, it has been shown that local output decisions can give highly competitive accuracies, thanks to the power of dense neural input representations that embody global syntactic information. We investigate two conceptually simple local neural models for constituent parsing, which make local decisions to constituent spans and CFG rules, respectively. Consistent with previous findings along the line, our best model gives highly competitive results, achieving the labeled bracketing F1 scores of 92.4% on PTB and 87.3% on CTB 5.1.
Zhiyang Teng, Yue Zhang 0004
COLING1
2017 Head-Lexicalized Bidirectional Tree LSTMs
abstract
Sequential LSTMs have been extended to model tree structures, giving competitive results for a number of tasks. Existing methods model constituent trees by bottom-up combinations of constituent nodes, making direct use of input word information only for leaf nodes. This is different from sequential LSTMs, which contain references to input words for each node. In this paper, we propose a method for automatic head-lexicalization for tree-structure LSTMs, propagating head words from leaf nodes to every constituent node. In addition, enabled by head lexicalization, we build a tree LSTM in the top-down direction, which corresponds to bidirectional sequential LSTMs in structure. Experiments show that both extensions give better representations of tree structures. Our final model gives the best results on the Stanford Sentiment Treebank and highly competitive results on the TREC question type classification task.
Zhiyang Teng, Yue Zhang 0004
Trans. Assoc. Comput. Linguistics1
2016 Combining Discrete and Neural Features for Sequence Labeling
Jie Yang 0039, Zhiyang Teng, Meishan Zhang, Yue Zhang 0004
CICLing (1)2
2016 Measuring the Information Content of Financial News
abstract
Measuring the information content of news text is useful for decision makers in their investments since news information can influence the intrinsic values of companies. We propose a model to automatically measure the information content given news text, trained using news and corresponding cumulative abnormal returns of listed companies. Existing methods in finance literature exploit sentiment signal features, which are limited by not considering factors such as events. We address this issue by leveraging deep neural models to extract rich semantic features from news text. In particular, a novel tree-structured LSTM is used to find target-specific representations of news text given syntax structures. Empirical results show that the neural models can outperform sentiment-based models, demonstrating the effectiveness of recent NLP technology advances for computational finance.
Ching-Yun Chang, Yue Zhang 0004, Zhiyang Teng, Zahn Bozanic, Bin Ke
COLING3
2016 Exploiting Mutual Benefits between Syntax and Semantic Roles using Neural Network
abstract
We investigate mutual benefits between syntax and semantic roles using neural network models, by studying a parsing→SRL pipeline, a SRL→parsing pipeline, and a simple joint model by embedding sharing.The integration of syntactic and semantic features gives promising results in a Chinese Semantic Treebank, demonstrating large potentials of neural models for joint parsing and semantic role labeling.
Peng Shi 0010, Zhiyang Teng, Yue Zhang 0004
EMNLP2
2016 Context-Sensitive Lexicon Features for Neural Sentiment Analysis
abstract
Sentiment lexicons have been leveraged as a useful source of features for sentiment analysis models, leading to the state-of-the-art accuracies.On the other hand, most existing methods use sentiment lexicons without considering context, typically taking the count, sum of strength, or maximum sentiment scores over the whole input.We propose a context-sensitive lexicon-based method based on a simple weighted-sum model, using a recurrent neural network to learn the sentiments strength, intensification and negation of lexicon sentiments in composing the sentiment value of sentences.Results show that our model can not only learn such operation details, but also give significant improvements over state-of-the-art recurrent neural network baselines without lexical features, achieving the best results on a Twitter benchmark.
Zhiyang Teng, Duy-Tin Vo, Yue Zhang 0004
EMNLP1
2016 LibN3L: A Lightweight Package for Neural NLP
Meishan Zhang, Jie Yang 0039, Zhiyang Teng, Yue Zhang 0004
LREC3
2016 Expectation-Regulated Neural Model for Event Mention Extraction
Ching-Yun Chang, Zhiyang Teng, Yue Zhang 0004
HLT-NAACL2