EDBT 2026 Demo / reviewers in the wild / expert
Moez Baccouche
dblp:92/8338
· DBLP profile ↗
19ranked-venue papers
4as first author
4since 2021 · last 2022
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 3 first-author · 3 since 2021Security and privacy · 1Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Trustworthy machine learning · 41% Vision and language · 35% Efficient and distributed learning · 14% | |
| Computer graphics and multimedia
1 paper |
Visualization and visual analytics · 100% |
Topics — the 11 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Vision and language
visual question answering |
1.7 | 4 | 2022 | Supervising the Transfer of Reasoning Patterns in VQA · NeurIPS 2021 How Transferable Are Reasoning Patterns in VQA? · CVPR 2021 Roses Are Red, Violets Are Blue... but Should VQA Expect Them To? · CVPR 2021 |
Machine learning › Trustworthy machine learning
dataset bias |
1.0 | 2 | 2021 | How Transferable Are Reasoning Patterns in VQA? · CVPR 2021 Roses Are Red, Violets Are Blue... but Should VQA Expect Them To? · CVPR 2021 |
Machine learning › Trustworthy machine learning
robustness |
1.0 | 2 | 2021 | How Transferable Are Reasoning Patterns in VQA? · CVPR 2021 Roses Are Red, Violets Are Blue... but Should VQA Expect Them To? · CVPR 2021 |
Visualization and visual analytics › information visualization
attention visualization |
0.6 | 1 | 2022 | VisQA: X-raying Vision and Language Reasoning in Transformers · IEEE Trans. Vis. Comput. Graph. 2022 |
Visualization and visual analytics › explainable AI
model interpretation |
0.6 | 1 | 2022 | VisQA: X-raying Vision and Language Reasoning in Transformers · IEEE Trans. Vis. Comput. Graph. 2022 |
Visualization and visual analytics
visual analytics |
0.6 | 1 | 2022 | VisQA: X-raying Vision and Language Reasoning in Transformers · IEEE Trans. Vis. Comput. Graph. 2022 |
Machine learning › Transfer learning and domain adaptation
knowledge transfer |
0.5 | 1 | 2021 | Supervising the Transfer of Reasoning Patterns in VQA · NeurIPS 2021 |
Machine learning › Efficient and distributed learning › automated machine learning › neural architecture search › task-specific architecture search
multimodal fusion architecture search |
0.4 | 1 | 2019 | MFAS: Multimodal Fusion Architecture Search · CVPR 2019 |
Machine learning › Efficient and distributed learning › automated machine learning
neural architecture search |
0.4 | 1 | 2019 | MFAS: Multimodal Fusion Architecture Search · CVPR 2019 |
Machine learning › Trustworthy machine learning
interpretability |
0.1 | 1 | 2021 | Roses Are Red, Violets Are Blue... but Should VQA Expect Them To? · CVPR 2021 |
Computer vision › Vision and language
multimodal reasoning |
0.1 | 1 | 2021 | How Transferable Are Reasoning Patterns in VQA? · CVPR 2021 |
Methods — techniques the papers use, named apart from their topics
transformer models · 1.1attention map · 1.1visual oracle · 0.5self-supervised pretraining · 0.5regularization · 0.5out-of-distribution evaluation · 0.5fine-tuning · 0.5bias reduction · 0.5attention analysis · 0.5PAC learning · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | VisQA: X-raying Vision and Language Reasoning in TransformersabstractVisual Question Answering systems target answering open-ended textual questions given input images. They are a testbed for learning high-level reasoning with a primary use in HCI, for instance assistance for the visually impaired. Recent research has shown that state-of-the-art models tend to produce answers exploiting biases and shortcuts in the training data, and sometimes do not even look at the input image, instead of performing the required reasoning steps. We present VisQA, a visual analytics tool that explores this question of reasoning vs. bias exploitation. It exposes the key element of state-of-the-art neural models - attention maps in transformers. Our working hypothesis is that reasoning steps leading to model predictions are observable from attention distributions, which are particularly useful for visualization. The design process of VisQA was motivated by well-known bias examples from the fields of deep learning and vision-language reasoning and evaluated in two ways. First, as a result of a collaboration of three fields, machine learning, vision and language reasoning, and data analytics, the work lead to a better understanding of bias exploitation of neural models for VQA, which eventually resulted in an impact on its design and training through the proposition of a method for the transfer of reasoning patterns from an oracle model. Second, we also report on the design of VisQA, and a goal-oriented evaluation of VisQA targeting the analysis of a model decision process from multiple experts, providing evidence that it makes the inner workings of models accessible to users. Theo Jaunet, Corentin Kervadec, Romain Vuillemot, Grigory Antipov, Moez Baccouche, Christian Wolf 0001 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2021 | Roses Are Red, Violets Are Blue... but Should VQA Expect Them To?abstractModels for Visual Question Answering (VQA) are notorious for their tendency to rely on dataset biases, as the large and unbalanced diversity of questions and concepts involved and tends to prevent models from learning to "reason", leading them to perform "educated guesses" instead. In this paper, we claim that the standard evaluation metric, which consists in measuring the overall in-domain accuracy, is misleading. Since questions and concepts are unbalanced, this tends to favor models which exploit subtle training set statistics. Alternatively, naively introducing artificial distribution shifts between train and test splits is also not completely satisfying. First, the shifts do not reflect real-world tendencies, resulting in unsuitable models; second, since the shifts are handcrafted, trained models are specifically designed for this particular setting, and do not generalize to other configurations. We propose the GQAOOD benchmark designed to overcome these concerns: we measure and compare accuracy over both rare and frequent question-answer pairs, and argue that the former is better suited to the evaluation of reasoning abilities, which we experimentally validate with models trained to more or less exploit biases. In a large-scale study involving 7 VQA models and 3 bias reduction techniques, we also experimentally demonstrate that these models fail to address questions involving infrequent concepts and provide recommendations for future directions of research. Corentin Kervadec, Grigory Antipov, Moez Baccouche, Christian Wolf 0001 |
CVPR | 3 |
| 2021 | How Transferable Are Reasoning Patterns in VQA?abstractSince its inception, Visual Question Answering (VQA) is notoriously known as a task, where models are prone to exploit biases in datasets to find shortcuts instead of performing high-level reasoning. Classical methods address this by removing biases from training data, or adding branches to models to detect and remove biases. In this paper, we argue that uncertainty in vision is a dominating factor preventing the successful learning of reasoning in vision and language problems. We train a visual oracle and in a large scale study provide experimental evidence that it is much less prone to exploiting spurious dataset biases compared to standard models. We propose to study the attention mechanisms at work in the visual oracle and compare them with a SOTA Transformer-based model. We provide an in-depth analysis and visualizations of reasoning patterns obtained with an online visualization tool which we make publicly available1. We exploit these insights by transferring reasoning patterns from the oracle to a SOTA Transformer-based VQA model taking standard noisy visual inputs via fine-tuning. In experiments we report higher overall accuracy, as well as accuracy on infrequent answers for each question type, which provides evidence for improved generalization and a decrease of the dependency on dataset biases. Corentin Kervadec, Theo Jaunet, Grigory Antipov, Moez Baccouche, Romain Vuillemot, Christian Wolf 0001 |
CVPR | 4 |
| 2021 | Supervising the Transfer of Reasoning Patterns in VQAabstractMethods for Visual Question Anwering (VQA) are notorious for leveraging dataset biases rather than performing reasoning, hindering generalization. It has been recently shown that better reasoning patterns emerge in attention layers of a state-of-the-art VQA model when they are trained on perfect (oracle) visual inputs. This provides evidence that deep neural networks can learn to reason when training conditions are favorable enough. However, transferring this learned knowledge to deployable models is a challenge, as much of it is lost during the transfer.We propose a method for knowledge transfer based on a regularization term in our loss function, supervising the sequence of required reasoning operations.We provide a theoretical analysis based on PAC-learning, showing that such program prediction can lead to decreased sample complexity under mild hypotheses. We also demonstrate the effectiveness of this approach experimentally on the GQA dataset and show its complementarity to BERT-like self-supervised pre-training. Corentin Kervadec, Christian Wolf 0001, Grigory Antipov, Moez Baccouche, Madiha Nadri Wolf |
NeurIPS | 4 |
| 2020 | Weak Supervision Helps Emergence of Word-Object Alignment and Improves Vision-Language TasksabstractThe large adoption of the self-attention (i.e. transformer model) and BERT-like training principles has recently resulted in a number of high performing models on a large panoply of vision-and-language problems (such as Visual Question Answering (VQA), image retrieval, etc.). In this paper we claim that these State-Of-The-Art (SOTA) approaches perform reasonably well in structuring information inside a single modality but, despite their impressive performances , they tend to struggle to identify fine-grained inter-modality relationships. Indeed, such relations are frequently assumed to be implicitly learned during training from application-specific losses, mostly cross-entropy for classification. While most recent works provide inductive bias for inter-modality relationships via cross attention modules, in this work, we demonstrate (1) that the latter assumption does not hold, i.e. modality alignment does not necessarily emerge automatically, and (2) that adding weak supervision for alignment between visual objects and words improves the quality of the learned models on tasks requiring reasoning. In particular , we integrate an object-word alignment loss into SOTA vision-language reasoning models and evaluate it on two tasks VQA and Language-driven Comparison of Images. We show that the proposed fine-grained inter-modality supervision significantly improves performance on both tasks. In particular, this new learning signal allows obtaining SOTA-level performances on GQA dataset (VQA task) with pre-trained models without finetuning on the task, and a new SOTA on NLVR2 dataset (Language-driven Comparison of Images). Finally, we also illustrate the impact of the contribution on the models reasoning by visualizing attention distributions. Corentin Kervadec, Grigory Antipov, Moez Baccouche, Christian Wolf 0001 |
ECAI | 3 |
| 2019 | MFAS: Multimodal Fusion Architecture SearchabstractWe tackle the problem of finding good architectures for multimodal classification problems. We propose a novel and generic search space that spans a large number of possible fusion architectures. In order to find an optimal architecture for a given dataset in the proposed search space, we leverage an efficient sequential model-based exploration approach that is tailored for the problem. We demonstrate the value of posing multimodal fusion as a neural architecture search problem by extensive experimentation on a toy dataset and two other real multimodal datasets. We discover fusion architectures that exhibit state-of-the-art performance for problems with different domain and dataset size, including the \ntu~dataset, the largest multimodal action recognition dataset available. Juan-Manuel Pérez-Rúa, Valentin Vielzeuf, Stéphane Pateux, Moez Baccouche, Frédéric Jurie |
CVPR | 4 |
| 2018 | Efficient Progressive Neural Architecture Search
Juan-Manuel Pérez-Rúa, Moez Baccouche, Stéphane Pateux |
BMVC | 2 |
| 2017 | Boosting cross-age face verification via generative age normalizationabstractDespite the tremendous progress in face verification performance as a result of Deep Learning, the sensitivity to human age variations remains an Achilles' heel of the majority of the contemporary face verification software. A promising solution to this problem consists in synthetic aging/rejuvenation of the input face images to some predefined age categories prior to face verification. We recently proposed [3] Age-cGAN aging/rejuvenation method based on generative adversarial neural networks allowing to synthesize more plausible and realistic faces than alternative non-generative methods. However, in this work, we show that Age-cGAN cannot be directly used for improving face verification due to its slightly imperfect preservation of the original identities in aged/rejuvenated faces. We therefore propose Local Manifold Adaptation (LMA) approach which resolves the stated issue of Age-cGAN resulting in the novel Age-cGAN+LMA aging/rejuvenation method. Based on Age-cGAN+LMA, we design an age normalization algorithm which boosts the accuracy of an off-the-shelf face verification software in the cross-age evaluation scenario. Grigory Antipov, Moez Baccouche, Jean-Luc Dugelay |
IJCB | 2 |
| 2017 | Face aging with conditional generative adversarial networksabstractIt has been recently shown that Generative Adversarial Networks (GANs) can produce synthetic images of exceptional visual fidelity. In this work, we propose the first GAN-based method for automatic face aging. Contrary to previous works employing GANs for altering of facial attributes, we make a particular emphasize on preserving the original person's identity in the aged version of his/her face. To this end, we introduce a novel approach for “Identity-Preserving” optimization of GAN's latent vectors. The objective evaluation of the resulting aged and rejuvenated face images by the state-of-the-art face recognition and age estimation solutions demonstrate the high potential of the proposed method. Grigory Antipov, Moez Baccouche, Jean-Luc Dugelay |
ICIP | 2 |
| 2017 | Effective training of convolutional neural networks for face-based gender and age prediction
Grigory Antipov, Moez Baccouche, Sid-Ahmed Berrani, Jean-Luc Dugelay |
Pattern Recognit. | 2 |
| 2016 | Boosting face recognition via neural Super-Resolution
Guillaume Berger, Clément Peyrard, Moez Baccouche |
ESANN | 3 |
| 2016 | Blind Super-Resolution with Deep Convolutional Neural Networks
Clément Peyrard, Moez Baccouche, Christophe Garcia |
ICANN (2) | 2 |
| 2015 | ICDAR2015 competition on Text Image Super-ResolutionabstractThis paper presents the first international competition on Text Image Super-Resolution (SR) and the ICDAR2015-TextSR dataset. We describe the core of the competition: interest, dataset generation and evaluation procedure, together with participating teams and their respective methods. The obtained results, along with baseline image upscaling schemes and state-of-the-art SR approaches are reported and commented. The main conclusion of this competition is that SR systems may improve OCR performances by up to 16.55 points in accuracy compared with bicubic interpolation for the proposed low resolution images. Clément Peyrard, Moez Baccouche, Franck Mamalet, Christophe Garcia |
ICDAR | 2 |
| 2014 | Deep learning of split temporal context for automatic speech recognitionabstractThis paper follows the recent advances in speech recognition which recommend replacing the standard hybrid GMM/HMM approach by deep neural architectures. These models were shown to drastically improve recognition performances, due to their ability to capture the underlying structure of data. However, they remain particularly complex since the entire temporal context of a given phoneme is learned with a single model, which must therefore have a very large number of trainable weights. This work proposes an alternative solution that splits the temporal context into blocks, each learned with a separate deep model. We demonstrate that this approach significantly reduces the number of parameters compared to the classical deep learning procedure, and obtains better results on the TIMIT dataset, among the best of state-of-the-art (with a 20.20% PER). We also show that our approach is able to assimilate data of different nature, ranging from wide to narrow bandwidth signals. Moez Baccouche, Benoit Besset, Patrice Collen, Olivier Le Blouch |
ICASSP | 1 |
| 2014 | Evaluation of video activity localizations integrating quality and quantity measurements
Christian Wolf 0001, Eric Lombardi, Julien Mille, Oya Çeliktutan, Mingyuan Jiu, Emre Dogan, Gonen Eren, Moez Baccouche, Emmanuel Dellandréa, Charles-Edmond Bichot, Christophe Garcia, Bülent Sankur |
Comput. Vis. Image Underst. | 8 |
| 2012 | Spatio-Temporal Convolutional Sparse Auto-Encoder for Sequence ClassificationabstractInternational audience Moez Baccouche, Franck Mamalet, Christian Wolf 0001, Christophe Garcia, Atilla Baskurt |
BMVC | 1 |
| 2012 | Sparse shift-invariant representation of local 2D patterns and sequence learning for human action recognition
Moez Baccouche, Franck Mamalet, Christian Wolf 0001, Christophe Garcia, Atilla Baskurt |
ICPR | 1 |
| 2010 | Action Classification in Soccer Videos with Long Short-Term Memory Recurrent Neural Networks
Moez Baccouche, Franck Mamalet, Christian Wolf 0001, Christophe Garcia, Atilla Baskurt |
ICANN (2) | 1 |
| 2009 | Monitoring Slow Ground Movements around Tunis City by Different SAR Interferometric MeasuresabstractThis paper presents an application of DInSAR techniques for the assessment of ground subsidences around Tunis City. A longterm analysis using two interferometric techniques were carried out to attempt reliable measurements. In this work, some aspects of interferometric processing softwares are reviewed and possible improvements are proposed in order to get better results. The convergence of the two different interferometric techniques done simultaneously and independently confirms results accuracy. Ferdaous Chaabane, Khaoula Elagouni, Moez Baccouche, Nadine Pourthié, Céline Tison, Pierre Briole |
IGARSS (3) | 3 |