VLDB 2026 Research / reviewers in the wild / expert
Paul-Alexis Dray
dblp:259/2999
· DBLP profile ↗
6ranked-venue papers
0as first author
2since 2021 · last 2021
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Language models and text generation · 57% Generative modeling · 26% Information extraction and text analysis · 8% | |
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 100% |
Topics — the 13 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling › generative adversarial network
GAN training |
0.9 | 2 | 2021 | To Beam Or Not To Beam: That is a Question of Cooperation for Language GANs · NeurIPS 2021 ColdGANs: Taming Language GANs with Cautious Sampling Strategies · NeurIPS 2020 |
Natural language and speech › Language models and text generation
text generation |
0.9 | 2 | 2021 | To Beam Or Not To Beam: That is a Question of Cooperation for Language GANs · NeurIPS 2021 ColdGANs: Taming Language GANs with Cautious Sampling Strategies · NeurIPS 2020 |
Natural language and speech › Language models and text generation
text summarization |
0.9 | 2 | 2020 | Discriminative Adversarial Search for Abstractive Summarization · ICML 2020 MLSUM: The Multilingual Summarization Corpus · EMNLP (1) 2020 |
Natural language and speech › Information extraction and text analysis
fact-based evaluation |
0.5 | 1 | 2021 | QuestEval: Summarization Asks for Fact-based Evaluation · EMNLP (1) 2021 |
Natural language and speech › Language models and text generation › text summarization
summarization evaluation |
0.5 | 1 | 2021 | QuestEval: Summarization Asks for Fact-based Evaluation · EMNLP (1) 2021 |
Natural language and speech › Language models and text generation › text summarization
abstractive summarization |
0.4 | 1 | 2020 | Discriminative Adversarial Search for Abstractive Summarization · ICML 2020 |
Machine learning › Trustworthy machine learning › robustness › adversarial robustness
adversarial training |
0.4 | 1 | 2020 | Discriminative Adversarial Search for Abstractive Summarization · ICML 2020 |
Natural language and speech › Language models and text generation
decoding |
0.4 | 1 | 2020 | Discriminative Adversarial Search for Abstractive Summarization · ICML 2020 |
Machine learning › Generative modeling
generative adversarial network |
0.4 | 1 | 2020 | Discriminative Adversarial Search for Abstractive Summarization · ICML 2020 |
Natural language and speech › Language models and text generation › text summarization
multilingual summarization |
0.4 | 1 | 2020 | MLSUM: The Multilingual Summarization Corpus · EMNLP (1) 2020 |
Machine learning › Generative modeling › diffusion model › controllable generation
reward-guided generation |
0.3 | 2 | 2021 | To Beam Or Not To Beam: That is a Question of Cooperation for Language GANs · NeurIPS 2021 ColdGANs: Taming Language GANs with Cautious Sampling Strategies · NeurIPS 2020 |
Natural language and speech › Question answering and dialogue systems › question generation
question generation for evaluation |
0.1 | 1 | 2021 | QuestEval: Summarization Asks for Fact-based Evaluation · EMNLP (1) 2021 |
Information retrieval
evaluation |
0.1 | 1 | 2020 | MLSUM: The Multilingual Summarization Corpus · EMNLP (1) 2020 |
Methods — techniques the papers use, named apart from their topics
reinforcement learning · 0.9cooperative discriminator-generator training · 0.5beam search · 0.5maximum likelihood estimation · 0.4generative adversarial network · 0.4discriminator-guided decoding · 0.4cold sampling · 0.4adversarial search · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | QuestEval: Summarization Asks for Fact-based EvaluationabstractThomas Scialom, Paul-Alexis Dray, Sylvain Lamprier, Benjamin Piwowarski, Jacopo Staiano, Alex Wang, Patrick Gallinari. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021. Thomas Scialom, Paul-Alexis Dray, Sylvain Lamprier, Benjamin Piwowarski, Jacopo Staiano, Patrick Gallinari |
EMNLP (1) | 2 |
| 2021 | To Beam Or Not To Beam: That is a Question of Cooperation for Language GANsabstractDue to the discrete nature of words, language GANs require to be optimized from rewards provided by discriminator networks, via reinforcement learning methods. This is a much harder setting than for continuous tasks, which enjoy gradient flows from discriminators to generators, usually leading to dramatic learning instabilities. However, we claim that this can be solved by making discriminator and generator networks cooperate to produce output sequences during training. These cooperative outputs, inherently built to obtain higher discrimination scores, not only provide denser rewards for training but also form a more compact artificial set for discriminator training, hence improving its accuracy and stability.In this paper, we show that our SelfGAN framework, built on this cooperative principle, outperforms Teacher Forcing and obtains state-of-the-art results on two challenging tasks, Summarization and Question Generation. Thomas Scialom, Paul-Alexis Dray, Jacopo Staiano, Sylvain Lamprier, Benjamin Piwowarski |
NeurIPS | 2 |
| 2020 | MLSUM: The Multilingual Summarization CorpusabstractWe present MLSUM, the first large-scale Mul-tiLingual SUMmarization dataset.Obtained from online newspapers, it contains 1.5M+ article/summary pairs in five different languages -namely, French, German, Spanish, Russian, Turkish.Together with English news articles from the popular CNN/Daily mail dataset, the collected data form a large scale multilingual dataset which can enable new research directions for the text summarization community.We report cross-lingual comparative analyses based on state-of-the-art systems.These highlight existing biases which motivate the use of a multi-lingual dataset. Thomas Scialom, Paul-Alexis Dray, Sylvain Lamprier, Benjamin Piwowarski, Jacopo Staiano |
EMNLP (1) | 2 |
| 2020 | Discriminative Adversarial Search for Abstractive SummarizationabstractWe introduce a novel approach for sequence decoding, Discriminative Adversarial Search (DAS), which has the desirable properties of alleviating the effects of exposure bias without requiring external metrics. Inspired by Generative Adversarial Networks (GANs), wherein a discriminator is used to improve the generator, our method differs from GANs in that the generator parameters are not updated at training time and the discriminator is used to drive sequence generation at inference time. We investigate the effectiveness of the proposed approach on the task of Abstractive Summarization: the results obtained show that a naive application of DAS improves over the state-of-the-art methods, with further gains obtained via discriminator retraining. Moreover, we show how DAS can be effective for cross-domain adaptation. Finally, all results reported are obtained without additional rule-based filtering strategies, commonly used by the best performing systems available: this indicates that DAS can effectively be deployed without relying on post-hoc modifications of the generated outputs. Thomas Scialom, Paul-Alexis Dray, Sylvain Lamprier, Benjamin Piwowarski, Jacopo Staiano |
ICML | 2 |
| 2020 | What BERT Sees: Cross-Modal Transfer for Visual Question GenerationabstractPre-trained language models have recently contributed to significant advances in NLP tasks.Recently, multi-modal versions of BERT have been developed, using heavy pretraining relying on vast corpora of aligned textual and image data, primarily applied to classification tasks such as VQA.In this paper, we are interested in evaluating the visual capabilities of BERT out-of-the-box, by avoiding pre-training made on supplementary data.We choose to study Visual Question Generation, a task of great interest for grounded dialog, that enables to study the impact of each modality (as input can be visual and/or textual).Moreover, the generation aspect of the task requires an adaptation since BERT is primarily designed as an encoder.We introduce BERT-gen, a BERT-based architecture for text generation, able to leverage on either monoor multi-modal representations.The results reported under different configurations indicate an innate capacity for BERT-gen to adapt to multi-modal data and text generation, even with few data available, avoiding expensive pre-training.The proposed model obtains substantial improvements over the state-of-the-art on two established VQG datasets.txt1 Thomas Scialom, Patrick Bordes, Paul-Alexis Dray, Jacopo Staiano, Patrick Gallinari |
INLG | 3 |
| 2020 | ColdGANs: Taming Language GANs with Cautious Sampling StrategiesabstractTraining regimes based on Maximum Likelihood Estimation (MLE) suffer from known limitations, often leading to poorly generated text sequences that lack of coherence, factualness, and are prone to repetitions. At the root of these limitations is the mismatch between training and inference, i.e. the so-called exposure bias. Another problem lies in considering only the reference text as correct, while in practice several alternative formulations could be as good. Generative Adversarial Networks (GANs) could mitigate those limitations. Nonetheless, the discrete nature of text has hindered their application to language generation: the approaches proposed so far, based on Reinforcement Learning, have been shown to under-perform MLE. In this context, the exploration is known to be critical, while surprisingly being under-studied. In this work, we show how the most popular sampling method results in unstable training for language GANs. We propose alternative exploration strategies that we named Cold-GANs. By forcing the sampling to be close to the distribution mode, the learning dynamic becomes smoother. We report experimental results obtained on three tasks: unconditional text generation, question generation, and abstractive summarization. For the first time, to the best of our knowledge, the proposed language GANs compare favorably to MLE, and obtain improvements over the state-of-the-art on the considered tasks. Thomas Scialom, Paul-Alexis Dray, Sylvain Lamprier, Benjamin Piwowarski, Jacopo Staiano |
NeurIPS | 2 |