Alexis Lechervy

dblp:95/8614 · DBLP profile ↗
← Back
23ranked-venue papers
4as first author
13since 2021 · last 2026
0000-0002-9441-0187ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 17 · 3 first-author · 9 since 2021Artificial intelligence and machine learning · 10 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Visually Informed Text Representations for Visual Contextual Classification of Arts
Raphaëlle Lemaire, Jérémie Pantin, Alexis Lechervy, Azamat Kaibaldiyev, Fabrice Maurel, Gaël Dias, Youssef Chahir
ICPR (3)3
2025 UNETRSal: Saliency Prediction with Hybrid Transformer-Based Architecture
Azamat Kaibaldiyev, Jérémie Pantin, Alexis Lechervy, Fabrice Maurel, Youssef Chahir, Gaël Dias
ACIVS3
2025 Inclusive Easy-to-Read Text Generation for Individuals with Cognitive Impairments
abstract
Ensuring accessibility for individuals with cognitive impairments is essential for autonomy, self-determination, and full citizenship. However, manual Easy-to-Read (ETR) text adaptations are slow, costly, and difficult to scale, limiting access to crucial information in healthcare, education, and civic life. AI-driven ETR generation offers a scalable solution but faces key challenges, including dataset scarcity, domain adaptation, and balancing lightweight learning of Large Language Models (LLMs). In this paper, we introduce ETR-fr, the first dataset for ETR text generation fully compliant with European ETR guidelines. We implement parameter-efficient fine-tuning on PLMs and LLMs to establish generative baselines. To ensure high-quality and accessible outputs, we introduce an evaluation framework based on automatic metrics supplemented by human assessments. The latter is conducted using a 36-question evaluation form that is aligned with the guidelines. Overall results show that PLMs perform comparably to LLMs and adapt effectively to out-of-domain texts. Code and datasets are available at https://github.com/FrLdy/ETR-fr.
François Ledoyen, Gaël Dias, Alexis Lechervy, Jérémie Pantin, Fabrice Maurel, Youssef Chahir, Elisa Gouzonnat, Mélanie Berthelot, Stanislas Moravac, Armony Altinier, Amy Khairalla
ECAI3
2025 Facilitating Cognitive Accessibility with LLMs: A Multi-Task Approach to Easy-to-Read Text Generation
abstract
Simplifying complex texts is essential to ensure equitable access to information, particularly for individuals with cognitive impairments.The Easy-to-Read (ETR) initiative provides a framework to make content more accessible for these individuals.However, manually creating such texts remains time-consuming and resourceintensive.In this work, we investigate the potential of large language models (LLMs) to automate the generation of ETR content.To address the scarcity of aligned corpora and the specific constraints of ETR, we propose a multitask learning (MTL) approach that trains models jointly on text summarization, text simplification, and ETR generation.We explore two complementary strategies: multi-task retrievalaugmented generation (RAG) for in-context learning (ICL), and MTL-LoRA for parameterefficient fine-tuning (PEFT).Our experiments with Mistral-7B and LLaMA-3-8B, conducted on ETR-fr, a new high-quality dataset, show that MTL-LoRA consistently outperforms all other strategies in in-domain settings, while the MTL-RAG-based approach achieves better generalization in out-of-domain scenarios.
François Ledoyen, Gaël Dias, Jérémie Pantin, Alexis Lechervy, Fabrice Maurel, Youssef Chahir
EMNLP4
2025 CHASE: Channel-Wise and Spatial Attention for Early Exiting in Image Classification
abstract
Dynamic early-exiting neural networks have been proposed for image classification to balance the trade-off between classification performance and inference cost. In this context, we propose a multi-exit neural network architecture that exploits the power of attention mechanisms, which improve performance but incur significant computational overhead. In CHASE, we introduce two attention-like mechanisms to go beyond existing multi-exit architectures. The first mechanism dynamically adjusts the importance of different feature channels and spatial locations, recalibrating channel-wise feature responses. The second mechanism, based on self-attention, aggregates features from different spatial locations at the end of the network. We evaluate the proposed architecture on the CIFAR and ImageNet datasets, comparing it with the original network and other state-of-the-art approaches. Our results show that the proposed architecture achieves competitive performance in terms of accuracy and computational efficiency.
Youva Addad, Alexis Lechervy, Frédéric Jurie
ICASSP2
2025 WYSIWYG: What You See Is Where Your Gaze
abstract
As Picasso said, a painting lives only through the one who looks at it. To materialize this thought, we propose to automatically produce artworks that visually transform paintings by amplifying and distorting the most observed areas by viewers. Our work is based on a study conducted at the Caen Museum of Fine Arts in France. During the study, 151 participants were equipped with eye-tracking glasses, and observed various paintings, first alone and then in pairs. Based on the fixation and gaze path stored data, we first generate saliency maps that reflect the visual attention given to each painting. These maps are then used to fine-tune the UNETRSal model, a neural network designed to predict saliency maps, in order to align its outputs with human visual patterns observed during the experiment. The saliency maps generated are subsequently used to create deformations of the original painting. This overall process gives rise to a new artwork born from the interaction between human gaze and AI-prediction.
Raphaëlle Lemaire, Azamat Kaibaldiyev, Eléonore Mariette, Débora Viglieri, Alexis Lechervy, Fabrice Maurel, Gaël Dias, Jérémie Pantin, Gaëtane Blaizot, Véronique Agin, Nicolas Poirel, Eric Bui, Hervé Platel, Denis Vivien, Youssef Chahir
ACM Multimedia5
2025 A Conflict-Guided Evidential Multimodal Fusion for Semantic Segmentation
abstract
International audience
Lucas Deregnaucourt, Hind Laghmara, Alexis Lechervy, Samia Ainouz 0001
WACV3
2025 Toward simplicity in dynamic inference: a critical study and redesign of early-exit networks
Youva Addad, Alexis Lechervy, Frédéric Jurie
Mach. Vis. Appl.2
2024 Balancing Accuracy and Efficiency in Budget-Aware Early-Exiting Neural Networks
Youva Addad, Alexis Lechervy, Frédéric Jurie
ICPR (6)2
2023 Automatic Sleep Stage Classification on EEG Signals Using Time-Frequency Representation
Paul Dequidt, Mathieu Seraphim, Alexis Lechervy, Ivan Igor Gaez, Luc Brun, Olivier Etard
AIME3
2023 Temporal Sequences of EEG Covariance Matrices for Automated Sleep Stage Scoring with Attention Mechanisms
Mathieu Seraphim, Paul Dequidt, Alexis Lechervy, Florian Yger, Luc Brun, Olivier Etard
CAIP (2)3
2022 Combining Vision and Language Representations for Patch-based Identification of Lexico-Semantic Relations
abstract
Although a wide range of applications have been proposed in the field of multimodal natural language processing, very few works have been tackling multimodal relational lexical semantics. In this paper, we propose the first attempt to identify lexico-semantic relations with visual clues, which embody linguistic phenomena such as synonymy, co-hyponymy or hypernymy. While traditional methods take advantage of the paradigmatic approach or/and the distributional hypothesis, we hypothesize that visual information can supplement the textual information, relying on the apperceptum subcomponent of the semiotic textology linguistic theory. For that purpose, we automatically extend two gold-standard datasets with visual information, and develop different fusion techniques to combine textual and visual modalities following the patch-based strategy. Experimental results over the multimodal datasets show that the visual information can supplement the missing semantics of textual encodings with reliable performance improvements.
Prince Jha, Gaël Dias, Alexis Lechervy, José G. Moreno 0001, Anubhav Jangra, Sebastião Pais, Sriparna Saha 0001
ACM Multimedia3
2021 Improving Neural Text Style Transfer by Introducing Loss Function Sequentiality
abstract
Text style transfer is an important issue for conversational agents as it may adapt utterance production to specific dialogue situations. It consists in introducing a given style within a sentence while preserving its semantics. Within this scope, different strategies have been proposed that either rely on parallel data or take advantage of non-supervised techniques. In this paper, we follow the latter approach and show that the sequential introduction of different loss functions into the learning process can boost the performance of a standard model. We also evidence that combining different style classifiers that either focus on global or local textual information improves sentence generation. Experiments on the Yelp dataset show that our methodology strongly competes with the current state-of-the-art models across style accuracy, grammatical correctness, and content preservation.
Chinmay Rane, Gaël Dias, Alexis Lechervy, Asif Ekbal
SIGIR3
2018 CAKE: a Compact and Accurate K-dimensional representation of Emotion
Corentin Kervadec, Valentin Vielzeuf, Stéphane Pateux, Alexis Lechervy, Frédéric Jurie
BMVC4
2018 TS-NET: Combining Modality Specific and Common Features for Multimodal Patch Matching
abstract
Multimodal patch matching addresses the problem of finding the correspondences between image patches from two different modalities, e.g. RGB vs sketch or RGB vs near-infrared. The comparison of patches of different modalities can be done by discovering the information common to both modalities (Siamese like approaches) or the modality-specific information (Pseudo-Siamese like approaches). We observed that none of these two scenarios is optimal. This motivates us to propose a three-stream architecture, dubbed as TS-Net, combining the benefits of the two. In addition, we show that adding extra constraints in the intermediate layers of such networks further boosts the performance. Experimentations on three multimodal datasets show significant performance gains in comparison with Siamese and Pseudo-Siamese networks†.
Sovann En, Alexis Lechervy, Frédéric Jurie
ICIP2
2018 An Occam's Razor View on Learning Audiovisual Emotion Recognition with Small Training Sets
abstract
This paper presents a light-weight and accurate deep neural model for audiovisual emotion recognition. To design this model, the authors followed a philosophy of simplicity, drastically limiting the number of parameters to learn from the target datasets, always choosing the simplest learning methods: i) transfer learning and low-dimensional space embedding allows to reduce the dimensionality of the representations, ii) visual temporal information handled by a simple score-per-frame selection process averaged across time, iii) simple frame selection mechanism for weighting images within sequences, iv) fusion of the different modalities at prediction level (late fusion). The paper also highlights the inherent challenges of the AFEW dataset and the difficulty of model selection with as few as 383 validation sequences. The proposed real-time emotion classifier achieved a state-of-the-art accuracy of 60.64 % on the test set of AFEW, and ranked 4th at the Emotion in the Wild 2018 challenge.
Valentin Vielzeuf, Corentin Kervadec, Stéphane Pateux, Alexis Lechervy, Frédéric Jurie
ICMI4
2016 MLBoost Revisited: A Faster Metric Learning Algorithm for Identity-Based Face Retrieval
Romain Negrel, Alexis Lechervy, Frédéric Jurie
BMVC2
2016 A joint learning approach for cross domain age estimation
abstract
We propose a novel joint learning method for cross domain age estimation, a domain adaptation problem. The proposed method learns a low dimensional projection along with a re-gressor, in the projection space, in a joint framework. The projection aligns the features from two different domains, i.e. source and target, to the same space, while the regressor predicts the age from the domain aligned features. After this alignment, a regressor trained with only a few examples from the target domain, along with more examples from the source domain, can predict very well the ages of the target domain face images. We provide empirical validation on the largest publicly available dataset for age estimation i.e. MORPH-II. The proposed method improves performance over several strong baselines and the current state-of-the-art methods.
Binod Bhattarai, Gaurav Sharma 0004, Alexis Lechervy, Frédéric Jurie
ICASSP3
2015 Boosted Metric Learning for Efficient Identity-Based Face Retrieval
abstract
International audience
Romain Negrel, Alexis Lechervy, Frédéric Jurie
BMVC2
2014 Boosted kernel for image categorization
Alexis Lechervy, Philippe Henri Gosselin, Frédéric Precioso
Multim. Tools Appl.1
2012 Linear kernel combination using boosting
Alexis Lechervy, Philippe Henri Gosselin, Frédéric Precioso
ESANN1
2012 Boosting kernel combination for multi-class image categorization
abstract
In this paper, we propose a novel algorithm to design multi-class kernel functions based on an iterative combination of weak kernels in a scheme inspired from boosting framework. The method proposed in this article aims at building a new feature where the centroid for each class are optimally located. We evaluate our method for image categorization by considering a state-of-the-art image database and by comparing our results with reference methods. We show that on the Oxford Flower databases our approach achieves better results than previous state-of-the-art methods.
Alexis Lechervy, Philippe Henri Gosselin, Frédéric Precioso
ICIP1
2010 Active Boosting for Interactive Object Retrieval
abstract
This paper presents a new algorithm based on boosting for interactive object retrieval in images. Recent works propose ”online boosting” algorithms where weak classifier sets are iteratively trained from data. These algorithms are proposed for visual tracking in videos, and are not well adapted to ”online boosting” for interactive retrieval. We propose in this paper to iteratively build weak classifiers from images, labeled as positive by the user during a retrieval session. A novel active learning strategy for the selection of images for user annotation is also proposed. This strategy is used to enhance the strong classifier resulting from ”boosting” process, but also to build new weak classifiers. Experiments have been carried out on a generalist database in order to compare the proposed method to a SVM based reference approach.
Alexis Lechervy, Philippe Henri Gosselin, Frédéric Precioso
ICPR1