Gehui Shen

dblp:205/9017 · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
4since 2021 · last 2022
0000-0002-2457-3810ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 4 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Language models and text generation · 40% Probabilistic and Bayesian machine learning · 26% Information extraction and text analysis · 13%

Topics — the 12 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation › text generation
constrained text generation
0.612022
Switch-GPT: An Effective Method for Constrained Text Generation under Few-Shot Settings (Student Abstract) · AAAI 2022
Natural language and speech › Language models and text generation › text generation
few-shot text generation
0.612022
Switch-GPT: An Effective Method for Constrained Text Generation under Few-Shot Settings (Student Abstract) · AAAI 2022
Natural language and speech › Language models and text generation › large language model
GPT-2
0.612022
Switch-GPT: An Effective Method for Constrained Text Generation under Few-Shot Settings (Student Abstract) · AAAI 2022
Machine learning › Probabilistic and Bayesian machine learning › deep probabilistic models › bayesian deep learning
bayesian neural networks
0.412020
Variational Learning of Bayesian Neural Networks via Bayesian Dark Knowledge · IJCAI 2020
Machine learning › Efficient and distributed learning › model compression
knowledge distillation
0.412020
Variational Learning of Bayesian Neural Networks via Bayesian Dark Knowledge · IJCAI 2020
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods
markov chain monte carlo
0.412020
Variational Learning of Bayesian Neural Networks via Bayesian Dark Knowledge · IJCAI 2020
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
variational inference
0.412020
Variational Learning of Bayesian Neural Networks via Bayesian Dark Knowledge · IJCAI 2020
Machine learning › Deep learning architectures and training
recurrent neural network
0.412019
Leap-LSTM: Enhancing Long Short-Term Memory for Text Categorization · IJCAI 2019
Natural language and speech › Information extraction and text analysis
text classification
0.412019
Leap-LSTM: Enhancing Long Short-Term Memory for Text Categorization · IJCAI 2019
Natural language and speech › Language models and text generation › natural language understanding
sentence pair modeling
0.312017
Inter-Weighted Alignment Network for Sentence Pair Modeling · EMNLP 2017
Natural language and speech › Information extraction and text analysis › text similarity › semantic similarity
sentence similarity
0.312017
Inter-Weighted Alignment Network for Sentence Pair Modeling · EMNLP 2017
Natural language and speech › Machine translation › statistical machine translation
word alignment
0.312017
Inter-Weighted Alignment Network for Sentence Pair Modeling · EMNLP 2017

Methods — techniques the papers use, named apart from their topics

LSTM · 0.7lexical constraints · 0.6attention module · 0.6variational inference · 0.4markov chain monte carlo · 0.4knowledge distillation · 0.4feature encoding · 0.4word-level similarity matrix · 0.3attention weighting · 0.3
YearPublicationVenuePosition
2022 Switch-GPT: An Effective Method for Constrained Text Generation under Few-Shot Settings (Student Abstract)
abstract
In real-world applications of natural language generation, target sentences are often required to satisfy some lexical constraints. However, the success of most neural-based models relies heavily on data, which is infeasible for data-scarce new domains. In this work, we present FewShotAmazon, the first benchmark for the task of Constrained Text Generation under few-shot settings on multiple domains. Further, we propose the Switch-GPT model, in which we utilize the strong language modeling capacity of GPT-2 to generate fluent and well-formulated sentences, while using a light attention module to decide which constraint to attend to at each step. Experiments show that the proposed Switch-GPT model is effective and remarkably outperforms the baselines. Codes will be available at https://github.com/chang-github-00/Switch-GPT.
Gehui Shen, Zhi-Hong Deng 0001
AAAI3
2021 Generative Feature Replay with Orthogonal Weight Modification for Continual Learning
abstract
The ability of intelligent agents to learn and remember multiple tasks sequentially is crucial to achieving artificial general intelligence. Many continual learning (CL) methods have been proposed to overcome catastrophic forgetting which results from non i.i.d data in the sequential learning of neural networks. In this paper we focus on class incremental learning, a challenging CL scenario. For this scenario, generative replay is a promising strategy which generates and replays pseudo data for previous tasks to alleviate catastrophic forgetting. However, it is hard to train a generative model continually for relatively complex data. Based on recently proposed orthogonal weight modification (OWM) algorithm which can approximately keep previously learned feature invariant when learning new tasks, we propose to 1) replay penultimate layer feature with a generative model; 2) leverage a self-supervised auxiliary task to further enhance the stability of feature. Empirical results on several datasets show our method always achieves substantial improvement over powerful OWM while conventional generative replay always results in a negative effect. Meanwhile our method beats several strong baselines including one based on real data storage. In addition, we conduct experiments to study why our method is effective.
Gehui Shen, Zhi-Hong Deng 0001
IJCNN1
2021 Towards unsupervised text multi-style transfer with parameter-sharing scheme
Gehui Shen, Zhi-Hong Deng 0001, Unil Yun
Neurocomputing3
2021 Sequence generative adversarial nets with a conditional discriminator
Yongfei Yan, Gehui Shen, Zhi-Hong Deng 0001, Unil Yun
Neurocomputing2
2020 Variational Learning of Bayesian Neural Networks via Bayesian Dark Knowledge
abstract
Bayesian neural networks (BNNs) have received more and more attention because they are capable of modeling epistemic uncertainty which is hard for conventional neural networks. Markov chain Monte Carlo (MCMC) methods and variational inference (VI) are two mainstream methods for Bayesian deep learning. The former is effective but its storage cost is prohibitive since it has to save many samples of neural network parameters. The latter method is more time and space efficient, however the approximate variational posterior limits its performance. In this paper, we aim to combine the advantages of above two methods by distilling MCMC samples into an approximate variational posterior. On the basis of an existing distillation technique we first propose variational Bayesian dark knowledge method. Moreover, we propose Bayesian dark prior knowledge, a novel distillation method which considers MCMC posterior as the prior of a variational BNN. Two proposed methods both not only can reduce the space overhead of the teacher model so that are scalable, but also maintain a distilled posterior distribution capable of modeling epistemic uncertainty. Experimental results manifest our methods outperform existing distillation method in terms of predictive accuracy and uncertainty modeling.
Gehui Shen, Zhi-Hong Deng 0001
IJCAI1
2020 Learning to compose over tree structures via POS tags for sentence representation
Gehui Shen, Zhi-Hong Deng 0001
Expert Syst. Appl.1
2020 A Window-Based Self-Attention approach for sentence encoding
Zhi-Hong Deng 0001, Gehui Shen
Neurocomputing3
2019 Leap-LSTM: Enhancing Long Short-Term Memory for Text Categorization
abstract
Recurrent Neural Networks (RNNs) are widely used in the field of natural language processing (NLP), ranging from text categorization to question answering and machine translation. However, RNNs generally read the whole text from beginning to end or vice versa sometimes, which makes it inefficient to process long texts. When reading a long document for a categorization task, such as topic categorization, large quantities of words are irrelevant and can be skipped. To this end, we propose Leap-LSTM, an LSTM-enhanced model which dynamically leaps between words while reading texts. At each step, we utilize several feature encoders to extract messages from preceding texts, following texts and the current word, and then determine whether to skip the current word. We evaluate Leap-LSTM on several text categorization tasks: sentiment analysis, news categorization, ontology classification and topic classification, with five benchmark data sets. The experimental results show that our model reads faster and predicts better than standard LSTM. Compared to previous models which can also skip words, our model achieves better trade-offs between performance and efficiency.
Gehui Shen, Zhi-Hong Deng 0001
IJCAI2
2017 Inter-Weighted Alignment Network for Sentence Pair Modeling
abstract
Sentence pair modeling is a crucial problem in the field of natural language processing.In this paper, we propose a model to measure the similarity of a sentence pair focusing on the interaction information.We utilize the word level similarity matrix to discover fine-grained alignment of two sentences.It should be emphasized that each word in a sentence has a different importance from the perspective of semantic composition, so we exploit two novel and efficient strategies to explicitly calculate a weight for each word.Although the proposed model only use a sequential LSTM for sentence modeling without any external resource such as syntactic parser tree and additional lexicon features, experimental results show that our model achieves state-of-the-art performance on three datasets of two tasks.
Gehui Shen, Yunlun Yang, Zhi-Hong Deng 0001
EMNLP1