VLDB 2026 Research / reviewers in the wild / expert
Aliaksei Severyn
dblp:86/8354
· DBLP profile ↗
29ranked-venue papers
14as first author
4since 2021 · last 2025
0009-0003-2954-4167ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 21 · 8 first-author · 4 since 2021Databases, data management, data science and information retrieval · 12 · 9 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
15 papers |
Language models and text generation · 53% Information extraction and text analysis · 12% Reinforcement learning · 8% | |
| Databases, data mining, and information retrieval
4 papers |
Information retrieval · 86% Machine learning and data management · 14% |
Topics — the 30 heaviest of 34, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning › reinforcement learning from human feedback › learning from human feedback
RLHF |
0.9 | 1 | 2025 | BOND: Aligning LLMs with Best-of-N Distillation · ICLR 2025 |
Information retrieval › ranking
learning to rank |
0.8 | 4 | 2017 | Neural Ranking Models with Weak Supervision · SIGIR 2017 Learning to Rank Short Text Pairs with Convolutional Deep Neural Networks · SIGIR 2015 A syntax-aware re-ranker for microblog retrieval · SIGIR 2014 |
Natural language and speech › Language models and text generation › controllable text generation
text editing |
0.8 | 2 | 2020 | Unsupervised Text Style Transfer with Padded Masked Language Models · EMNLP (1) 2020 Encode, Tag, Realize: High-Precision Text Editing · EMNLP/IJCNLP (1) 2019 |
Natural language and speech › Information extraction and text analysis
sentiment analysis |
0.7 | 3 | 2017 | Leveraging Large Amounts of Weakly Supervised Data for Multi-Language Sentiment Classification · WWW 2017 Twitter Sentiment Analysis with Deep Convolutional Neural Networks · SIGIR 2015 Opinion Mining on YouTube · ACL (1) 2014 |
Machine learning › Efficient and distributed learning
inference acceleration |
0.7 | 1 | 2023 | Fast Text Generation with Text-Editing Models · KDD 2023 |
Natural language and speech › Language models and text generation › decoding
constrained decoding |
0.5 | 1 | 2021 | Controlled Text Generation as Continuous Optimization with Multiple Constraints · NeurIPS 2021 |
Natural language and speech › Language models and text generation
controllable text generation |
0.5 | 1 | 2021 | Controlled Text Generation as Continuous Optimization with Multiple Constraints · NeurIPS 2021 |
Natural language and speech › Language models and text generation
decoding |
0.5 | 1 | 2021 | Controlled Text Generation as Continuous Optimization with Multiple Constraints · NeurIPS 2021 |
Natural language and speech › Language models and text generation › controllable text generation
text style transfer |
0.4 | 1 | 2020 | Unsupervised Text Style Transfer with Padded Masked Language Models · EMNLP (1) 2020 |
Natural language and speech › Language models and text generation › controllable text generation › text style transfer
unsupervised text style transfer |
0.4 | 1 | 2020 | Unsupervised Text Style Transfer with Padded Masked Language Models · EMNLP (1) 2020 |
Machine learning › Deep learning architectures and training
output layer |
0.4 | 1 | 2019 | Breaking the Softmax Bottleneck via Learnable Monotonic Pointwise Non-linearities · ICML 2019 |
Natural language and speech › Language models and text generation › language modeling
softmax bottleneck |
0.4 | 1 | 2019 | Breaking the Softmax Bottleneck via Learnable Monotonic Pointwise Non-linearities · ICML 2019 |
Natural language and speech › Information extraction and text analysis › sentiment analysis › sentiment classification
cross-lingual sentiment classification |
0.3 | 1 | 2017 | Leveraging Large Amounts of Weakly Supervised Data for Multi-Language Sentiment Classification · WWW 2017 |
Natural language and speech › Language models and text generation
text generation |
0.3 | 1 | 2017 | A Hybrid Convolutional Variational Autoencoder for Text Generation · EMNLP 2017 |
Machine learning › Generative modeling
variational autoencoder |
0.3 | 1 | 2017 | A Hybrid Convolutional Variational Autoencoder for Text Generation · EMNLP 2017 |
Information retrieval › retrieval models › neural retrieval
neural ranking model |
0.3 | 1 | 2017 | Neural Ranking Models with Weak Supervision · SIGIR 2017 |
Information retrieval
retrieval models |
0.3 | 1 | 2017 | Neural Ranking Models with Weak Supervision · SIGIR 2017 |
Machine learning and data management
weak supervision |
0.3 | 1 | 2017 | Neural Ranking Models with Weak Supervision · SIGIR 2017 |
Information retrieval › web search › web information retrieval › social media retrieval
microblog retrieval |
0.3 | 2 | 2015 | A syntax-aware re-ranker for microblog retrieval · SIGIR 2014 Learning to Rank Short Text Pairs with Convolutional Deep Neural Networks · SIGIR 2015 |
Natural language and speech › Question answering and dialogue systems › community question answering
answer ranking |
0.2 | 1 | 2015 | Learning to Rank Short Text Pairs with Convolutional Deep Neural Networks · SIGIR 2015 |
Natural language and speech › Information extraction and text analysis › sentiment analysis › sentiment classification
twitter sentiment classification |
0.2 | 1 | 2015 | Twitter Sentiment Analysis with Deep Convolutional Neural Networks · SIGIR 2015 |
Machine learning › Efficient and distributed learning › model compression
knowledge distillation |
0.2 | 1 | 2023 | Fast Text Generation with Text-Editing Models · KDD 2023 |
Machine learning › Representation and self-supervised learning
automated feature generation |
0.2 | 1 | 2013 | Automatic Feature Engineering for Answer Selection and Extraction · EMNLP 2013 |
Natural language and speech › Information extraction and text analysis
feature engineering |
0.2 | 1 | 2013 | Automatic Feature Engineering for Answer Selection and Extraction · EMNLP 2013 |
Natural language and speech › Language models and text generation › text generation › surface realization
linearization |
0.2 | 1 | 2013 | Fast Linearization of Tree Kernels over Large-Scale Data · IJCAI 2013 |
Machine learning › Kernel, tree and ensemble methods › kernel function
tree kernel |
0.2 | 1 | 2013 | Fast Linearization of Tree Kernels over Large-Scale Data · IJCAI 2013 |
Natural language and speech › Machine translation
controllable machine translation |
0.1 | 1 | 2021 | Controlled Text Generation as Continuous Optimization with Multiple Constraints · NeurIPS 2021 |
Machine learning › Generative modeling
style transfer |
0.1 | 1 | 2021 | Controlled Text Generation as Continuous Optimization with Multiple Constraints · NeurIPS 2021 |
Natural language and speech › Question answering and dialogue systems
answer re-ranking |
0.1 | 1 | 2012 | Structural relationships for large-scale learning of answer re-ranking · SIGIR 2012 |
Information retrieval › reranking
answer re-ranking |
0.1 | 1 | 2012 | Structural relationships for large-scale learning of answer re-ranking · SIGIR 2012 |
Methods — techniques the papers use, named apart from their topics
convolutional neural network · 0.9reinforcement learning from human feedback · 0.9jeffreys divergence · 0.9distribution matching · 0.9text-editing · 0.7seq2seq · 0.7knowledge distillation · 0.7lagrangian multipliers · 0.5gradient descent · 0.5continuous relaxation · 0.5word embeddings · 0.3weak supervision · 0.3feed-forward neural network · 0.3structural kernel · 0.2kernel learning · 0.2tree kernels · 0.2linearization · 0.2supervised learning · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | BOND: Aligning LLMs with Best-of-N DistillationabstractReinforcement learning from human feedback (RLHF) is a key driver of quality and safety in state-of-the-art large language models.
Yet, a surprisingly simple and strong inference-time strategy is Best-of-N sampling that selects the best generation among N candidates.
In this paper, we propose Best-of-N Distillation (BOND), a novel RLHF algorithm that seeks to emulate Best-of-N but without its significant computational overhead at inference time. Specifically, BOND is a distribution matching algorithm that forces the distribution of generations from the policy to get closer to the Best-of-N distribution. We use the Jeffreys divergence (a linear combination of forward and backward KL) to balance between mode-covering and mode-seeking behavior, and derive an iterative formulation that utilizes a moving anchor for efficiency. We demonstrate the effectiveness of our approach and several design choices through experiments on abstractive summarization and Gemma models. Pier Giuseppe Sessa, Robert Dadashi, Léonard Hussenot, Johan Ferret, Nino Vieillard, Alexandre Ramé, Bobak Shahriari, Sarah Perrin, Abram L. Friesen, Geoffrey Cideron, Sertan Girgin, Piotr Stanczyk, Andrea Michi, Danila Sinopalnikov, Sabela Ramos, Amélie Héliou, Aliaksei Severyn, Matt Hoffman 0001, Nikola Momchev, Olivier Bachem |
ICLR | 17 |
| 2024 | Small Language Models Improve Giants by Rewriting Their OutputsabstractGiorgos Vernikos, Arthur Brazinskas, Jakub Adamek, Jonathan Mallinson, Aliaksei Severyn, Eric Malmi. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Giorgos Vernikos, Arthur Brazinskas, Jakub Adámek, Jonathan Mallinson, Aliaksei Severyn, Eric Malmi |
EACL (1) | 5 |
| 2023 | Fast Text Generation with Text-Editing ModelsabstractText-editing models have recently become a prominent alternative to seq2seq models for monolingual text-generation tasks such as grammatical error correction, simplification, and style transfer. These tasks share a common trait -- they exhibit a large amount of textual overlap between the source and target texts. Text-editing models take advantage of this observation and learn to generate the output by predicting edit operations applied to the source sequence. In contrast, seq2seq models generate outputs word-by-word from scratch thus making them slow at inference time. Text-editing models provide several benefits over seq2seq models including faster inference speed, higher sample efficiency, and better control and explainability of the outputs. This tutorial provides a comprehensive overview of text-editing models and discusses how they can be used to mitigate hallucination and bias, both pressing challenges in the field of text generation. Finally, we discuss how to optimize latency of large language models via distillation to text-editing models and other means. Eric Malmi, Yue Dong 0002, Jonathan Mallinson, Aleksandr Chuklin, Jakub Adámek, Daniil Mirylenka, Felix Stahlberg, Sebastian Krause, Shankar Kumar, Aliaksei Severyn |
KDD | 10 |
| 2021 | Controlled Text Generation as Continuous Optimization with Multiple ConstraintsabstractAs large-scale language model pretraining pushes the state-of-the-art in text generation, recent work has turned to controlling attributes of the text such models generate. While modifying the pretrained models via fine-tuning remains the popular approach, it incurs a significant computational cost and can be infeasible due to a lack of appropriate data. As an alternative, we propose \textsc{MuCoCO}---a flexible and modular algorithm for controllable inference from pretrained models. We formulate the decoding process as an optimization problem that allows for multiple attributes we aim to control to be easily incorporated as differentiable constraints. By relaxing this discrete optimization to a continuous one, we make use of Lagrangian multipliers and gradient-descent-based techniques to generate the desired text. We evaluate our approach on controllable machine translation and style transfer with multiple sentence-level attributes and observe significant improvements over baselines. Sachin Kumar 0009, Eric Malmi, Aliaksei Severyn, Yulia Tsvetkov |
NeurIPS | 3 |
| 2020 | Unsupervised Text Style Transfer with Padded Masked Language ModelsabstractWe propose MASKER, an unsupervised textediting method for style transfer.To tackle cases when no parallel source-target pairs are available, we train masked language models (MLMs) for both the source and the target domain.Then we find the text spans where the two models disagree the most in terms of likelihood.This allows us to identify the source tokens to delete to transform the source text to match the style of the target domain.The deleted tokens are replaced with the target MLM, and by using a padded MLM variant, we avoid having to predetermine the number of inserted tokens.Our experiments on sentence fusion and sentiment transfer demonstrate that MASKER performs competitively in a fully unsupervised setting.Moreover, in lowresource settings, it improves supervised methods' accuracy by over 10 percentage points when pre-training them on silver training data generated by MASKER. Eric Malmi, Aliaksei Severyn, Sascha Rothe |
EMNLP (1) | 2 |
| 2020 | Leveraging Pre-trained Checkpoints for Sequence Generation TasksabstractUnsupervised pre-training of large neural models has recently revolutionized Natural Language Processing. By warm-starting from the publicly released checkpoints, NLP practitioners have pushed the state-of-the-art on multiple benchmarks while saving significant amounts of compute time. So far the focus has been mainly on the Natural Language Understanding tasks. In this paper, we demonstrate the efficacy of pre-trained checkpoints for Sequence Generation. We developed a Transformer-based sequence-to-sequence model that is compatible with publicly available pre-trained BERT, GPT-2, and RoBERTa checkpoints and conducted an extensive empirical study on the utility of initializing our model, both encoder and decoder, with these checkpoints. Our models result in new state-of-the-art results on Machine Translation, Text Summarization, Sentence Splitting, and Sentence Fusion. Sascha Rothe, Shashi Narayan, Aliaksei Severyn |
Trans. Assoc. Comput. Linguistics | 3 |
| 2019 | Encode, Tag, Realize: High-Precision Text EditingabstractEric Malmi, Sebastian Krause, Sascha Rothe, Daniil Mirylenka, Aliaksei Severyn. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Eric Malmi, Sebastian Krause, Sascha Rothe, Daniil Mirylenka, Aliaksei Severyn |
EMNLP/IJCNLP (1) | 5 |
| 2019 | Breaking the Softmax Bottleneck via Learnable Monotonic Pointwise Non-linearitiesabstractThe Softmax function on top of a final linear layer is the de facto method to output probability distributions in neural networks. In many applications such as language models or text generation, this model has to produce distributions over large output vocabularies. Recently, this has been shown to have limited representational capacity due to its connection with the rank bottleneck in matrix factorization. However, little is known about the limitations of Linear-Softmax for quantities of practical interest such as cross entropy or mode estimation, a direction that we explore here. As an efficient and effective solution to alleviate this issue, we propose to learn parametric monotonic functions on top of the logits. We theoretically investigate the rank increasing capabilities of such monotonic functions. Empirically, our method improves in two different quality metrics over the traditional Linear-Softmax layer in synthetic and real language model experiments, adding little time or memory overhead, while being comparable to the more computationally expensive mixture of Softmaxes. Octavian-Eugen Ganea, Sylvain Gelly, Gary Bécigneul, Aliaksei Severyn |
ICML | 4 |
| 2017 | A Hybrid Convolutional Variational Autoencoder for Text GenerationabstractIn this paper we explore the effect of architectural choices on learning a variational autoencoder (VAE) for text generation. In contrast to the previously introduced VAE model for text where both the encoder and decoder are RNNs, we propose a novel hybrid architecture that blends fully feed-forward convolutional and deconvolutional components with a recurrent language model. Our architecture exhibits several attractive properties such as faster run time and convergence, ability to better handle long sequences and, more importantly, it helps to avoid the issue of the VAE collapsing to a deterministic model. Stanislau Semeniuta, Aliaksei Severyn, Erhardt Barth |
EMNLP | 2 |
| 2017 | Neural Ranking Models with Weak SupervisionabstractDespite the impressive improvements achieved by unsupervised deep neural networks in computer vision and NLP tasks, such improvements have not yet been observed in ranking for information retrieval. The reason may be the complexity of the ranking problem, as it is not obvious how to learn from queries and documents when no supervised signal is available. Hence, in this paper, we propose to train a neural ranking model using weak supervision, where labels are obtained automatically without human annotators or any external resources (e.g., click data). To this aim, we use the output of an unsupervised ranking model, such as BM25, as a weak supervision signal. We further train a set of simple yet effective ranking models based on feed-forward neural networks. We study their effectiveness under various learning scenarios (point-wise and pair-wise models) and using different input representations (i.e., from encoding query-document pairs into dense/sparse vectors to using word embedding representation). We train our networks using tens of millions of training instances and evaluate it on two standard collections: a homogeneous news collection (Robust) and a heterogeneous large-scale web collection (ClueWeb). Our experiments indicate that employing proper objective functions and letting the networks to learn the input representation based on weakly supervised data leads to impressive performance, with over 13% and 35% MAP improvements over the BM25 model on the Robust and the ClueWeb collections. Our findings also suggest that supervised neural ranking models can greatly benefit from pre-training on large amounts of weakly labeled data that can be easily obtained from unsupervised IR models. Mostafa Dehghani 0001, Hamed Zamani, Aliaksei Severyn, Jaap Kamps, W. Bruce Croft |
SIGIR | 3 |
| 2017 | Leveraging Large Amounts of Weakly Supervised Data for Multi-Language Sentiment ClassificationabstractThis paper presents a novel approach for multi-lingual sentiment classification in short texts. This is a challenging task as the amount of training data in languages other than English is very limited. Previously proposed multi-lingual approaches typically require to establish a correspondence to English for which powerful classifiers are already available. In contrast, our method does not require such supervision. We leverage large amounts of weakly-supervised data in various languages to train a multi-layer convolutional network and demonstrate the importance of using pre-training of such networks. We thoroughly evaluate our approach on various multi-lingual datasets, including the recent SemEval-2016 sentiment prediction benchmark (Task 4), where we achieved state-of-the-art performance. We also compare the performance of our model trained individually for each language to a variant trained for all languages at once. We show that the latter model reaches slightly worse - but still acceptable - performance when compared to the single language model, while benefiting from better generalization properties across languages. Jan Deriu, Aurélien Lucchi, Valeria De Luca, Aliaksei Severyn, Mark Cieliebak, Thomas Hofmann 0001, Martin Jaggi |
WWW | 4 |
| 2016 | Globally Normalized Transition-Based Neural NetworksabstractDaniel Andor, Chris Alberti, David Weiss, Aliaksei Severyn, Alessandro Presta, Kuzman Ganchev, Slav Petrov, Michael Collins. Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2016. Daniel Andor, Christopher Alberti, David Weiss 0001, Aliaksei Severyn, Alessandro Presta, Kuzman Ganchev, Slav Petrov, Michael Collins 0001 |
ACL (1) | 4 |
| 2016 | Recurrent Dropout without Memory LossabstractThis paper presents a novel approach to recurrent neural network (RNN) regularization. Differently from the widely adopted dropout method, which is applied to forward connections of feedforward architectures or RNNs, we propose to drop neurons directly in recurrent connections in a way that does not cause loss of long-term memory. Our approach is as easy to implement and apply as the regular feed-forward dropout and we demonstrate its effectiveness for the most effective modern recurrent network – Long Short-Term Memory network. Our experiments on three NLP benchmarks show consistent improvements even when combined with conventional feed-forward dropout. Stanislau Semeniuta, Aliaksei Severyn, Erhardt Barth |
COLING | 2 |
| 2016 | Multi-lingual opinion mining on YouTube
Aliaksei Severyn, Alessandro Moschitti, Olga Uryupina, Barbara Plank, Katja Filippova |
Inf. Process. Manag. | 1 |
| 2015 | On the Automatic Learning of Sentiment LexiconsabstractThis paper describes a simple and princi-pled approach to automatically construct sen-timent lexicons using distant supervision. We induce the sentiment association scores for the lexicon items from a model trained on a weakly supervised corpora. Our empiri-cal findings show that features extracted from such a machine-learned lexicon outperform models using manual or other automatically constructed sentiment lexicons. Finally, our system achieves the state-of-the-art in Twitter Sentiment Analysis tasks from Semeval-2013 and ranks 2nd best in Semeval-2014 according to the average rank. Aliaksei Severyn, Alessandro Moschitti |
HLT-NAACL | 1 |
| 2015 | Learning to Rank Short Text Pairs with Convolutional Deep Neural NetworksabstractLearning a similarity function between pairs of objects is at the core of learning to rank approaches. In information retrieval tasks we typically deal with query-document pairs, in question answering -- question-answer pairs. However, before learning can take place, such pairs needs to be mapped from the original space of symbolic words into some feature space encoding various aspects of their relatedness, e.g. lexical, syntactic and semantic. Feature engineering is often a laborious task and may require external knowledge sources that are not always available or difficult to obtain. Recently, deep learning approaches have gained a lot of attention from the research community and industry for their ability to automatically learn optimal feature representation for a given task, while claiming state-of-the-art performance in many tasks in computer vision, speech recognition and natural language processing. In this paper, we present a convolutional neural network architecture for reranking pairs of short texts, where we learn the optimal representation of text pairs and a similarity function to relate them in a supervised way from the available training data. Our network takes only words in the input, thus requiring minimal preprocessing. In particular, we consider the task of reranking short text pairs where elements of the pair are sentences. We test our deep learning system on two popular retrieval tasks from TREC: Question Answering and Microblog Retrieval. Our model demonstrates strong performance on the first task beating previous state-of-the-art systems by about 3\% absolute points in both MAP and MRR and shows comparable results on tweet reranking, while enjoying the benefits of no manual feature engineering and no additional syntactic parsers. Aliaksei Severyn, Alessandro Moschitti |
SIGIR | 1 |
| 2015 | Twitter Sentiment Analysis with Deep Convolutional Neural NetworksabstractThis paper describes our deep learning system for sentiment analysis of tweets. The main contribution of this work is a new model for initializing the parameter weights of the convolutional neural network, which is crucial to train an accurate model while avoiding the need to inject any additional features. Briefly, we use an unsupervised neural language model to train initial word embeddings that are further tuned by our deep learning model on a distant supervised corpus. At a final stage, the pre-trained parameters of the network are used to initialize the model. We train the latter on the supervised training data recently made available by the official system evaluation campaign on Twitter Sentiment Analysis organized by Semeval-2015. A comparison between the results of our approach and the systems participating in the challenge on the official test sets, suggests that our model could be ranked in the first two positions in both the phrase-level subtask A (among 11 teams) and on the message-level subtask B (among 40 teams). This is an important evidence on the practical value of our solution. Aliaksei Severyn, Alessandro Moschitti |
SIGIR | 1 |
| 2014 | Opinion Mining on YouTubeabstractAliaksei Severyn, Alessandro Moschitti, Olga Uryupina, Barbara Plank, Katja Filippova. Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2014. Aliaksei Severyn, Alessandro Moschitti, Olga Uryupina, Barbara Plank, Katja Filippova |
ACL (1) | 1 |
| 2014 | Encoding Semantic Resources in Syntactic Structures for Passage RerankingabstractIn this paper, we propose to use semantic knowledge from Wikipedia and largescale structured knowledge datasets available as Linked Open Data (LOD) for the answer passage reranking task.We represent question and candidate answer passages with pairs of shallow syntactic/semantic trees, whose constituents are connected using LOD.The trees are processed by SVMs and tree kernels, which can automatically exploit tree fragments.The experiments with our SVM rank algorithm on the TREC Question Answering (QA) corpus show that the added relational information highly improves over the state of the art, e.g., about 15.4% of relative improvement in P@1. Kateryna Tymoshenko, Alessandro Moschitti, Aliaksei Severyn |
EACL | 3 |
| 2014 | SenTube: A Corpus for Sentiment Analysis on YouTube Social Media
Olga Uryupina, Barbara Plank, Aliaksei Severyn, Agata Rotondi, Alessandro Moschitti |
LREC | 3 |
| 2014 | A syntax-aware re-ranker for microblog retrievalabstractWe tackle the problem of improving microblog retrieval algorithms by proposing a robust structural representation of (query, tweet) pairs. We employ these structures in a principled kernel learning framework that automatically extracts and learns highly discriminative features. We test the generalization power of our approach on the TREC Microblog 2011 and 2012 tasks. We find that relational syntactic features generated by structural kernels are effective for learning to rank (L2R) and can easily be combined with those of other existing systems to boost their accuracy. In particular, the results show that our L2R approach improves on almost all the participating systems at TREC, only using their raw scores as a single feature. Our method yields an average increase of 5% in retrieval effectiveness and 7 positions in system ranks. Aliaksei Severyn, Alessandro Moschitti, Manos Tsagkias, Richard Berendsen, Maarten de Rijke |
SIGIR | 1 |
| 2013 | Building structures from classifiers for passage rerankingabstractThis paper shows that learning to rank models can be applied to automatically learn complex patterns, such as relational semantic structures occurring in questions and their answer passages. This is achieved by providing the learning algorithm with a tree representation derived from the syntactic trees of questions and passages connected by relational tags, where the latter are again provided by the means of automatic classifiers, i.e., question and focus classifiers and Named Entity Recognizers. This way effective structural relational patterns are implicitly encoded in the representation and can be automatically utilized by powerful machine learning models such as kernel methods. Aliaksei Severyn, Massimo Nicosia, Alessandro Moschitti |
CIKM | 1 |
| 2013 | Learning Adaptable Patterns for Passage Reranking
Aliaksei Severyn, Massimo Nicosia, Alessandro Moschitti |
CoNLL | 1 |
| 2013 | Automatic Feature Engineering for Answer Selection and ExtractionabstractThis paper proposes a framework for automatically engineering features for two important tasks of question answering: answer sentence selection and answer extraction.We represent question and answer sentence pairs with linguistic structures enriched by semantic information, where the latter is produced by automatic classifiers, e.g., question classifier and Named Entity Recognizer.Tree kernels applied to such structures enable a simple way to generate highly discriminative structural features that combine syntactic and semantic information encoded in the input trees.We conduct experiments on a public benchmark from TREC to compare with previous systems for answer sentence selection and answer extraction.The results show that our models greatly improve on the state of the art, e.g., up to 22% on F1 (relative improvement) for answer extraction, while using no additional resources and no manual feature engineering. Aliaksei Severyn, Alessandro Moschitti |
EMNLP | 1 |
| 2013 | Fast Linearization of Tree Kernels over Large-Scale Data
Aliaksei Severyn, Alessandro Moschitti |
IJCAI | 1 |
| 2012 | Structural relationships for large-scale learning of answer re-rankingabstractSupervised learning applied to answer re-ranking can highly improve on the overall accuracy of question answering (QA) systems. The key aspect is that the relationships and properties of the question/answer pair composed of a question and the supporting passage of an answer candidate, can be efficiently compared with those captured by the learnt model. Aliaksei Severyn, Alessandro Moschitti |
SIGIR | 1 |
| 2012 | Fast support vector machines for convolution tree kernels
Aliaksei Severyn, Alessandro Moschitti |
Data Min. Knowl. Discov. | 1 |
| 2011 | Fast Support Vector Machines for Structural Kernels
Aliaksei Severyn, Alessandro Moschitti |
ECML/PKDD (3) | 1 |
| 2010 | Large-Scale Support Vector Learning with Structural Kernels
Aliaksei Severyn, Alessandro Moschitti |
ECML/PKDD (3) | 1 |