Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Xingyi Cheng

dblp:206/6376 · DBLP profile ↗
← Back
17ranked-venue papers
5as first author
10since 2021 · last 2025
0009-0005-4510-8693ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 4 first-author · 8 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Language models and text generation · 41% Deep learning architectures and training · 14% Trustworthy machine learning · 13%
Interdisciplinary, comprehensive, and emerging computing
5 papers
Bioinformatics and computational biology · 100%

Topics — the 30 heaviest of 35, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology
protein structure prediction
1.422024
MSAGPT: Neural Prompting Protein Structure Prediction via MSA Generative Pre-Training · NeurIPS 2024
Injecting Multimodal Information into Rigid Protein Docking via Bi-level Optimization · NeurIPS 2023
Bioinformatics and computational biology
protein function prediction
0.912025
ProtGO: universal protein function prediction utilizing multi-modal gene ontology knowledge · Bioinform. 2025
Machine learning › Efficient and distributed learning › efficient training
compute-optimal training
0.812024
Training Compute-Optimal Protein Language Models · NeurIPS 2024
Natural language and speech › Language models and text generation › neural language model
protein language model
0.812024
Training Compute-Optimal Protein Language Models · NeurIPS 2024
Machine learning › Deep learning architectures and training
scaling laws
0.812024
Training Compute-Optimal Protein Language Models · NeurIPS 2024
Bioinformatics and computational biology
multiple sequence alignment
0.812024
MSAGPT: Neural Prompting Protein Structure Prediction via MSA Generative Pre-Training · NeurIPS 2024
Bioinformatics and computational biology › protein sequence analysis › protein sequence representation
protein language model
0.812024
MSAGPT: Neural Prompting Protein Structure Prediction via MSA Generative Pre-Training · NeurIPS 2024
Natural language and speech › Language models and text generation
large language model evaluation
0.712023
Revisiting Out-of-distribution Robustness in NLP: Benchmarks, Analysis, and LLMs Evaluations · NeurIPS 2023
Natural language and speech › Language models and text generation › text generation › argument generation
rebuttal generation
0.712023
Won't Get Fooled Again: Answering Questions with False Premises · ACL (1) 2023
Machine learning › Trustworthy machine learning
robustness
0.712023
Revisiting Out-of-distribution Robustness in NLP: Benchmarks, Analysis, and LLMs Evaluations · NeurIPS 2023
Machine learning › Trustworthy machine learning › robustness › distribution shift
robustness to distribution shift
0.712023
Revisiting Out-of-distribution Robustness in NLP: Benchmarks, Analysis, and LLMs Evaluations · NeurIPS 2023
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning
self-supervised representation learning
0.712023
xTrimoGene: An Efficient and Scalable Representation Learner for Single-Cell RNA-Seq Data · NeurIPS 2023
Bioinformatics and computational biology
genomics
0.712023
xTrimoGene: An Efficient and Scalable Representation Learner for Single-Cell RNA-Seq Data · NeurIPS 2023
Bioinformatics and computational biology › machine learning for biology
multimodal protein representation
0.712023
Injecting Multimodal Information into Rigid Protein Docking via Bi-level Optimization · NeurIPS 2023
Bioinformatics and computational biology › protein structure prediction
protein-protein docking
0.712023
Injecting Multimodal Information into Rigid Protein Docking via Bi-level Optimization · NeurIPS 2023
Bioinformatics and computational biology › protein structure prediction › protein-protein docking
rigid-body docking
0.712023
Injecting Multimodal Information into Rigid Protein Docking via Bi-level Optimization · NeurIPS 2023
Bioinformatics and computational biology › single-cell analysis › single-cell RNA sequencing
single-cell RNA-seq analysis
0.712023
xTrimoGene: An Efficient and Scalable Representation Learner for Single-Cell RNA-Seq Data · NeurIPS 2023
Machine learning › Efficient and distributed learning › model compression
knowledge distillation
0.512021
Dual-View Distilled BERT for Sentence Embedding · SIGIR 2021
Machine learning › Representation and self-supervised learning › text embedding
sentence embedding
0.512021
Dual-View Distilled BERT for Sentence Embedding · SIGIR 2021
Machine learning › Deep learning architectures and training
siamese network
0.512021
Dual-View Distilled BERT for Sentence Embedding · SIGIR 2021
Natural language and speech › Language models and text generation › text correction › spelling correction
chinese spelling check
0.412020
SpellGCN: Incorporating Phonological and Visual Similarities into Language Models for Chinese Spelling Check · ACL 2020
Machine learning › Graph learning › graph neural network › attention-based graph neural network
graph attention network
0.412020
Question Directed Graph Attention Network for Numerical Reasoning over Text · EMNLP (1) 2020
Natural language and speech › Language models and text generation › large language model training › language model pretraining
next sentence prediction
0.412020
Symmetric Regularization based BERT for Pair-wise Semantic Reasoning · SIGIR 2020
Natural language and speech › Language models and text generation › large language model training
pre-training objectives
0.412020
Symmetric Regularization based BERT for Pair-wise Semantic Reasoning · SIGIR 2020
Knowledge, reasoning and agents › Knowledge representation and reasoning
semantic reasoning
0.412020
Symmetric Regularization based BERT for Pair-wise Semantic Reasoning · SIGIR 2020
Natural language and speech › Language models and text generation › text correction
spelling correction
0.412020
SpellGCN: Incorporating Phonological and Visual Similarities into Language Models for Chinese Spelling Check · ACL 2020
Bioinformatics and computational biology › protein sequence analysis
protein sequence modeling
0.212024
Training Compute-Optimal Protein Language Models · NeurIPS 2024
Machine learning › Trustworthy machine learning › robustness
distribution shift
0.212023
Revisiting Out-of-distribution Robustness in NLP: Benchmarks, Analysis, and LLMs Evaluations · NeurIPS 2023
Machine learning › Deep learning architectures and training › transformer
efficient transformer
0.212023
xTrimoGene: An Efficient and Scalable Representation Learner for Single-Cell RNA-Seq Data · NeurIPS 2023
Machine learning › Deep learning architectures and training
transformer
0.212023
xTrimoGene: An Efficient and Scalable Representation Learner for Single-Cell RNA-Seq Data · NeurIPS 2023

Methods — techniques the papers use, named apart from their topics

masked language modeling · 1.5causal language modeling · 1.5fine-tuning · 1.3BERT · 0.9text alignment · 0.9taxonomy encoding · 0.9protein language model · 0.9gene ontology graph embedding · 0.9transfer learning · 0.8rejective fine-tuning · 0.8reinforcement learning from feedback · 0.8generative pre-training · 0.82d positional encoding · 0.8sparse attention · 0.7replay training · 0.7pre-training · 0.7in-context learning · 0.7asymmetric encoder-decoder transformer · 0.7
YearPublicationVenuePosition
2025 Protein Inverse Folding From Structure Feedback
abstract
The inverse folding problem, aiming to design amino acid sequences that fold into desired three-dimensional structures, is pivotal for various biotechnological applications. Here, we introduce a novel approach leveraging Direct Preference Optimization (DPO) to fine-tune an inverse folding model using feedback from a protein folding model. Given a target protein structure, we begin by sampling candidate sequences from the inverse‐folding model, then predict the three‐dimensional structure of each sequence with the folding model to generate pairwise structural‐preference labels. These labels are used to fine‐tune the inverse‐folding model under the DPO objective. Our results on the CATH 4.2 test set demonstrate that DPO fine-tuning not only improves sequence recovery of baseline models but also leads to a significant improvement in average TM-Score from 0.77 to 0.81, indicating enhanced structure similarity. Furthermore, iterative application of our DPO-based method on challenging protein structures yields substantial gains, with an average TM-Score increase of 79.5\% with regard to the baseline model. This work establishes a promising direction for enhancing protein sequence design ability from structure feedback by effectively utilizing preference optimization.
Junde Xu, Zijun Gao, Xinyi Zhou 0010, Xingyi Cheng, Guangyong Chen, Pheng-Ann Heng, Jiezhong Qiu
NeurIPS5
2025 ProtGO: universal protein function prediction utilizing multi-modal gene ontology knowledge
abstract
MOTIVATION: As one of the recalcitrant challenges in life sciences and biomedicine, protein function prediction suffers from a deluge of AI-designed proteins, particularly having to face multi-modal information in the era of big data. Importing the high-throughput neural-network-based prediction framework to replace the low-throughput biological experiments, a universal multi-modal method is straightforward in addressing the growing gap between known sequences and predicting functions. RESULTS: To bridge the gap, we propose ProtGO, a three-step framework for predicting protein function, which leverages the credible Gene Ontology (GO) knowledge base and integrates four common modalities. Specifically, we first introduce frontier pre-trained protein language models (PLMs) for representation learning of mainstay functional protein sequences. For the remaining multi-modal data, we design a text alignment module for explainable text descriptions, a taxonomy encoding module for species-specific taxonomy, and a GO graph embedding module for biological GO relations. Each module is independent and adaptive for the referenced modalities. By harnessing these four knowledge representations, ProtGO maximizes the potential of GO resources, enhancing the performance of vanilla PLMs and biological language models (LMs) in downstream GO prediction tasks. Extensive experiments demonstrate that ProtGO significantly advances the abilities of state-of-the-art PLMs to predict protein functions: approximately 8% to 27% increase in the maximum F1 measure (Fmax) compared to base models. These comprehensive studies confirm ProtGO's capability to deliver outstanding performance in protein function prediction by utilizing a rich blend of functional and evolutionary knowledge. AVAILABILITY AND IMPLEMENTATION: Our source code and all the data are available at https://github.com/sunyatawang/ProtGO.
Xingyi Cheng, Bo Chen 0026, Zhilei Bei, Wei Wang 0074, Jie Tang 0001
Bioinform.3
2024 MSAGPT: Neural Prompting Protein Structure Prediction via MSA Generative Pre-Training
abstract
Multiple Sequence Alignment (MSA) plays a pivotal role in unveiling the evolutionary trajectories of protein families. The accuracy of protein structure predictions is often compromised for protein sequences that lack sufficient homologous information to construct high-quality MSA. Although various methods have been proposed to generate high-quality MSA under these conditions, they fall short in comprehensively capturing the intricate co-evolutionary patterns within MSA or require guidance from external oracle models. Here we introduce MSAGPT, a novel approach to prompt protein structure predictions via MSA generative pre-training in a low-MSA regime. MSAGPT employs a simple yet effective 2D evolutionary positional encoding scheme to model the complex evolutionary patterns. Endowed by this, the flexible 1D MSA decoding framework facilitates zero- or few-shot learning. Moreover, we demonstrate leveraging the feedback from AlphaFold2 (AF2) can further enhance the model’s capacity via Rejective Fine-tuning (RFT) and Reinforcement Learning from AF2 Feedback (RLAF). Extensive experiments confirm the efficacy of MSAGPT in generating faithful and informative MSA (up to +8.5% TM-Score on few-shot scenarios). The transfer learning also demonstrates its great potential for the wide range of tasks resorting to the quality of MSA.
Bo Chen 0026, Zhilei Bei, Xingyi Cheng, Jie Tang 0001
NeurIPS3
2024 Training Compute-Optimal Protein Language Models
abstract
We explore optimally training protein language models, an area of significant interest in biological research where guidance on best practices is limited. Most models are trained with extensive compute resources until performance gains plateau, focusing primarily on increasing model sizes rather than optimizing the efficient compute frontier that balances performance and compute budgets. Our investigation is grounded in a massive dataset consisting of 939 million protein sequences. We trained over 300 models ranging from 3.5 million to 10.7 billion parameters on 5 to 200 billion unique tokens, to investigate the relations between model sizes, training token numbers, and objectives. First, we observed the effect of diminishing returns for the Causal Language Model (CLM) and that of overfitting for Masked Language Model (MLM) when repeating the commonly used Uniref database. To address this, we included metagenomic protein sequences in the training set to increase the diversity and avoid the plateau or overfitting effects. Second, we obtained the scaling laws of CLM and MLM on Transformer, tailored to the specific characteristics of protein sequence data. Third, we observe a transfer scaling phenomenon from CLM to MLM, further demonstrating the effectiveness of transfer through scaling behaviors based on estimated Effectively Transferred Tokens. Finally, to validate our scaling laws, we compare the large-scale versions of ESM-2 and PROGEN2 on downstream tasks, encompassing evaluations of protein generation as well as structure- and function-related tasks, all within less or equivalent pre-training compute budgets.
Xingyi Cheng, Bo Chen 0026, Jie Tang 0001
NeurIPS1
2023 Won't Get Fooled Again: Answering Questions with False Premises
abstract
Pre-trained language models (PLMs) have shown unprecedented potential in various fields, especially as the backbones for questionanswering (QA) systems.However, they tend to be easily deceived by tricky questions such as "How many eyes does the sun have?".Such frailties of PLMs often allude to the lack of knowledge within them.In this paper, we find that the PLMs already possess the knowledge required to rebut such questions, and the key is how to activate the knowledge.To systematize this observation, we investigate the PLMs' responses to one kind of tricky questions, i.e., the false premises questions (FPQs).We annotate a FalseQA dataset containing 2365 human-written FPQs, with the corresponding explanations for the false premises and the revised true premise questions.Using FalseQA, we discover that PLMs are capable of discriminating FPQs by fine-tuning on moderate numbers (e.g., 256) of examples.PLMs also generate reasonable explanations for the false premise, which serve as rebuttals.Further replaying a few general questions during training allows PLMs to excel on FPQs and general questions simultaneously.Our work suggests that once the rebuttal ability is stimulated, knowledge inside the PLMs can be effectively utilized to handle FPQs, which incentivizes the research on PLM-based QA systems.The FalseQA dataset and code are available at https://github.com/thunlp/FalseQA.
Shengding Hu, Xingyi Cheng, Zhiyuan Liu 0001, Maosong Sun 0001
ACL (1)4
2023 xTrimoGene: An Efficient and Scalable Representation Learner for Single-Cell RNA-Seq Data
abstract
Advances in high-throughput sequencing technology have led to significant progress in measuring gene expressions at the single-cell level. The amount of publicly available single-cell RNA-seq (scRNA-seq) data is already surpassing 50M records for humans with each record measuring 20,000 genes. This highlights the need for unsupervised representation learning to fully ingest these data, yet classical transformer architectures are prohibitive to train on such data in terms of both computation and memory. To address this challenge, we propose a novel asymmetric encoder-decoder transformer for scRNA-seq data, called xTrimoGene$^\alpha$ (or xTrimoGene for short), which leverages the sparse characteristic of the data to scale up the pre-training. This scalable design of xTrimoGene reduces FLOPs by one to two orders of magnitude compared to classical transformers while maintaining high accuracy, enabling us to train the largest transformer models over the largest scRNA-seq dataset today. Our experiments also show that the performance of xTrimoGene improves as we scale up the model sizes, and it also leads to SOTA performance over various downstream tasks, such as cell type annotation, perturb-seq effect prediction, and drug combination prediction. xTrimoGene model is now available for use as a service via the following link: https://api.biomap.com/xTrimoGene/apply.
Minsheng Hao, Xingyi Cheng, Chiming Liu, Jianzhu Ma, Xuegong Zhang, Taifeng Wang
NeurIPS3
2023 Injecting Multimodal Information into Rigid Protein Docking via Bi-level Optimization
abstract
The structure of protein-protein complexes is critical for understanding binding dynamics, biological mechanisms, and intervention strategies. Rigid protein docking, a fundamental problem in this field, aims to predict the 3D structure of complexes from their unbound states without conformational changes. In this scenario, we have access to two types of valuable information: sequence-modal information, such as coevolutionary data obtained from multiple sequence alignments, and structure-modal information, including the 3D conformations of rigid structures. However, existing docking methods typically utilize single-modal information, resulting in suboptimal predictions. In this paper, we propose xTrimoBiDock (or BiDock for short), a novel rigid docking model that effectively integrates sequence- and structure-modal information through bi-level optimization. Specifically, a cross-modal transformer combines multimodal information to predict an inter-protein distance map. To achieve rigid docking, the roto-translation transformation is optimized to align the docked pose with the predicted distance map. In order to tackle this bi-level optimization problem, we unroll the gradient descent of the inner loop and further derive a better initialization for roto-translation transformation based on spectral estimation. Compared to baselines, BiDock achieves a promising result of a maximum 234% relative improvement in challenging antibody-antigen docking problem.
YiWu Sun, Shaochuan Li, Cheng Yang 0002, Xingyi Cheng, Chuan Shi 0001
NeurIPS6
2023 Revisiting Out-of-distribution Robustness in NLP: Benchmarks, Analysis, and LLMs Evaluations
abstract
This paper reexamines the research on out-of-distribution (OOD) robustness in the field of NLP. We find that the distribution shift settings in previous studies commonly lack adequate challenges, hindering the accurate evaluation of OOD robustness. To address these issues, we propose a benchmark construction protocol that ensures clear differentiation and challenging distribution shifts. Then we introduceBOSS, a Benchmark suite for Out-of-distribution robustneSS evaluation covering 5 tasks and 20 datasets. Based on BOSS, we conduct a series of experiments on pretrained language models for analysis and evaluation of OOD robustness. First, for vanilla fine-tuning, we examine the relationship between in-distribution (ID) and OOD performance. We identify three typical types that unveil the inner learningmechanism, which could potentially facilitate the forecasting of OOD robustness, correlating with the advancements on ID datasets. Then, we evaluate 5 classic methods on BOSS and find that, despite exhibiting some effectiveness in specific cases, they do not offer significant improvement compared to vanilla fine-tuning. Further, we evaluate 5 LLMs with various adaptation paradigms and find that when sufficient ID data is available, fine-tuning domain-specific models outperform LLMs on ID examples significantly. However, in the case of OOD instances, prioritizing LLMs with in-context learning yields better results. We identify that both fine-tuned small models and LLMs face challenges in effectively addressing downstream tasks. The code is public at https://github.com/lifan-yuan/OOD_NLP.
Lifan Yuan, Yangyi Chen, Ganqu Cui, Hongcheng Gao, Fangyuan Zou, Xingyi Cheng, Heng Ji 0001, Zhiyuan Liu 0001, Maosong Sun 0001
NeurIPS6
2021 K-AID: Enhancing Pre-trained Language Models with Domain Knowledge for Question Answering
abstract
Knowledge enhanced pre-trained language models (K-PLMs) are shown to be effective for many public tasks in the literature, but few of them have been successfully applied in practice. To address this problem, we propose K-AID, a systematic approach that includes a low-cost knowledge acquisition process for acquiring domain knowledge, an effective knowledge infusion module for improving model performance, and a knowledge distillation component for reducing the model size and deploying K-PLMs on resource-restricted devices (e.g., CPU) for real-world application. Importantly, instead of capturing entity knowledge like the majority of existing K-PLMs, our approach captures relational knowledge, which contributes to better improving sentence-level text classification and text matching tasks that play a key role in question answering (QA). We conducted a set of experiments on five text classification tasks and three text matching tasks from three domains, namely E-commerce, Government, and Film&TV, and performed online A/B tests in E-commerce. Experimental results show that our approach is able to achieve substantial improvement on sentence-level question answering tasks and bring beneficial business value in industrial settings.
Fu Sun, Feng-Lin Li, Qianglong Chen, Xingyi Cheng, Ji Zhang 0011
CIKM5
2021 Dual-View Distilled BERT for Sentence Embedding
abstract
Recently, BERT realized significant progress for sentence matching via word-level cross sentence attention. However, the performance significantly drops when using siamese BERT-networks to derive two sentence embeddings, which fall short in capturing the global semantic since the word-level attention between two sentences is absent. In this paper, we propose a Dual-view distilled BERT~(DvBERT) for sentence matching with sentence embeddings. Our method deals with a sentence pair from two distinct views, i.e., Siamese View and Interaction View. Siamese View is the backbone where we generate sentence embeddings. Interaction View integrates the cross sentence interaction as multiple teachers to boost the representation ability of sentence embeddings. Experiments on six STS tasks show that our method outperforms the state-of-the-art sentence embedding methods.
Xingyi Cheng
SIGIR1
2020 SpellGCN: Incorporating Phonological and Visual Similarities into Language Models for Chinese Spelling Check
abstract
Chinese Spelling Check (CSC) is a task to detect and correct spelling errors in Chinese natural language.Existing methods have made attempts to incorporate the similarity knowledge between Chinese characters.However, they take the similarity knowledge as either an external input resource or just heuristic rules.This paper proposes to incorporate phonological and visual similarity knowledge into language models for CSC via a specialized graph convolutional network (SpellGCN).The model builds a graph over the characters, and SpellGCN is learned to map this graph into a set of inter-dependent character classifiers.These classifiers are applied to the representations extracted by another network, such as BERT, enabling the whole network to be end-to-end trainable.Experiments 1 are conducted on three human-annotated datasets.Our method achieves superior performance against previous models by a large margin.
Xingyi Cheng, Weidi Xu, Kunlong Chen, Shaohua Jiang, Taifeng Wang, Yuan Qi 0001
ACL1
2020 Towards Fast and Accurate Neural Chinese Word Segmentation with Multi-Criteria Learning
abstract
The ambiguous annotation criteria lead to divergence of Chinese Word Segmentation (CWS) datasets in various granularities.Multi-criteria Chinese word segmentation aims to capture various annotation criteria among datasets and leverage their common underlying knowledge.In this paper, we propose a domain adaptive segmenter to exploit diverse criteria of various datasets.Our model is based on Bidirectional Encoder Representations from Transformers (BERT), which is responsible for introducing open-domain knowledge.Private and shared projection layers are proposed to capture domain-specific knowledge and common knowledge, respectively.We also optimize computational efficiency via distillation, quantization, and compiler optimization.Experiments show that our segmenter outperforms the previous state of the art (SOTA) models on 10 CWS datasets with superior efficiency.
Weipeng Huang, Xingyi Cheng, Kunlong Chen, Taifeng Wang
COLING2
2020 Question Directed Graph Attention Network for Numerical Reasoning over Text
abstract
Kunlong Chen, Weidi Xu, Xingyi Cheng, Zou Xiaochuan, Yuyu Zhang, Le Song, Taifeng Wang, Yuan Qi, Wei Chu. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020.
Kunlong Chen, Weidi Xu, Xingyi Cheng, Zou Xiaochuan, Yuyu Zhang, Taifeng Wang, Yuan Qi 0001
EMNLP (1)3
2020 Symmetric Regularization based BERT for Pair-wise Semantic Reasoning
abstract
The ability of semantic reasoning over the sentence pair is essential for many natural language understanding tasks, e.g., natural language inference and machine reading comprehension. A recent significant improvement in these tasks comes from BERT. As reported, the next sentence prediction (NSP) in BERT is of great significance for downstream problems with sentence-pair input. Despite its effectiveness, NSP still lacks the essential signal to distinguish between entailment and shallow correlation. To remedy this, we propose to augment the NSP task to a multi-class categorization task, which includes previous sentence prediction (PSP). This task encourages the model to learn the subtle semantics, thereby improves the ability of semantic understanding. Furthermore, by using a smoothing technique, the scopes of NSP and PSP are expanded into a broader range which includes close but nonsuccessive sentences. This simple method yields remarkable improvement against vanilla BERT. Our method consistently improves the performance on the NLI and MRC benchmarks by a large margin, including the challenging HANS dataset.
Weidi Xu, Xingyi Cheng, Kunlong Chen, Taifeng Wang
SIGIR2
2019 Variational Semi-Supervised Aspect-Term Sentiment Analysis via Transformer
abstract
Aspect-term sentiment analysis (ATSA) is a long-standing challenge in natural language processing.It requires fine-grained semantical reasoning about a target entity appeared in the text.As manual annotation over the aspects is laborious and time-consuming, the amount of labeled data is limited for supervised learning.This paper proposes a semisupervised method for the ATSA problem by using the Variational Autoencoder based on Transformer.The model learns the latent distribution via variational inference.By disentangling the latent representation into the aspect-specific sentiment and the lexical context, our method induces the underlying sentiment prediction for the unlabeled data, which then benefits the ATSA classifier.Our method is classifier-agnostic, i.e., the classifier is an independent module and various supervised models can be integrated.Experimental results are obtained on the SemEval 2014 task 4 and show that our method is effective with different five specific classifiers and outperforms these models by a significant margin.
Xingyi Cheng, Weidi Xu, Taifeng Wang, Weipeng Huang, Kunlong Chen
CoNLL1
2019 BERT-Based Multi-head Selection for Joint Entity-Relation Extraction
Weipeng Huang, Xingyi Cheng, Taifeng Wang
NLPCC (2)2
2018 DeepTransport: Learning Spatial-Temporal Dependency for Traffic Condition Forecasting
abstract
Predicting traffic conditions has been recently explored as a way to relieve traffic congestion. Several pioneering approaches have been proposed based on traffic observations of the target location as well as its adjacent regions, but they obtain somewhat limited accuracy due to lack of mining road topology. To address the effect attenuation problem, we propose to take account of the traffic of surrounding locations. We propose an end-to-end framework called DeepTransport, in which Convolutional Neural Networks (CNN) and Recurrent Neural Networks (RNN) are utilized to obtain spatial-temporal traffic information within a transport network topology. In addition, attention mechanism is introduced to align spatial and temporal information. Moreover, we constructed and released a real-world large traffic condition dataset with 5-minute resolution. Our experiments on this dataset demonstrate our method captures the complex relationship in both temporal and spatial domain. It significantly outperforms traditional statistical methods and a state-of-the-art deep learning method.
Xingyi Cheng, Ruiqing Zhang, Jie Zhou 0025, Wei Xu 0017
IJCNN1