Catherine Achard

dblp:40/812 · DBLP profile ↗
← Back
41ranked-venue papers
4as first author
18since 2021 · last 2026
0000-0002-5790-0830ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 27 · 3 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 19 · 3 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 8 · 7 since 2021Systems, architecture and hardware · 2 · 1 since 2021
YearPublicationVenuePosition
2026 iMatcher: Improve matching in point cloud registration via local-to-global geometric consistency learning
Karim Slimani, Catherine Achard, Brahim Tamadazte
Pattern Recognit.2
2024 LoGDesc: Local Geometric Features Aggregation for Robust Point Cloud Registration
Karim Slimani, Brahim Tamadazte, Catherine Achard
ACCV (9)3
2024 Learning to Estimate Motion Between Non-adjacent Frames in Cardiac Cine MRI Data: A Fusion Approach
Nicolas Portal, Thomas Dietenbeck, Saud Khan, Mikaël Prigent, Mohamed Zarai, Khaoula Bouazizi-Verdier, Johanne Sylvain, Alban Redheuil, Gilles Montalescot, Nadjia Kachenoura, Catherine Achard
ICPR (11)12
2024 Adaptive virtual agent: Design and evaluation for real-time human-agent interaction
abstract
When we converse, we adapt our behaviors to our interlocutors. The adaptation can serve to indicate our engagement which can also elicit enhancement of the involvement of others. Virtual agents (or socially interactive virtual agents) that play the role of interaction partners can improve the human users’ interaction experience by displaying continuous and adaptive behaviors in real time. Virtual agents have been used in multiple domains to improve user interaction and performance. The promising results of the endowment of adaptation to agents in increasing the agents’ perception and user experience were shown in previous studies. In this paper, we develop an adaptive virtual agent that renders real-time adaptive behaviors based on the behaviors shown by its human interlocutor. The ASAP model rendering reciprocally adaptive agent behavior was employed to realize the system. The system consists of four main parts: perception of social signals, agent adaptive behavior generation, agent visualization (i.e. rendering of the agent’s verbal and nonverbal behavior), and communication of signals. To showcase the usefulness of our adaptive agent, as a proof-of-concept we choose the e-health application of cognitive behavior therapy (CBT), which identifies and rectifies biased and irrational thoughts (or automatic thoughts). Through this study, we show the importance of giving the agent reciprocal adaptation capability notably in enhancing the user experience and the effectiveness of the CBT session. We validate the importance of endowing such adaptation capability by studying the difference between agents that are reciprocal adaptive, solely expressive (with mismatched behavior), and inexpressive (in a still posture) via questionnaires and measures related to the agent perception (naturalness, human-likeliness, synchrony, and engagement) for user experience and the CBT effectiveness (mood, anxiety, stress, and cognitive change). These results highlight the value of making virtual agents adapt in real time. This could lead to agents being capable of providing more personalized and interactive experiences for a wide range of applications. Also, we have collected a new human-agent interaction (HAI) database, HAI-CBT database, which is publicly available to the research community.
Jieyeon Woo, Kazuhiro Shidara, Catherine Achard, Hiroki Tanaka, Satoshi Nakamura 0001, Catherine Pelachaud
Int. J. Hum. Comput. Stud.3
2024 Rocnet: 3D robust registration of points clouds using deep learning
Karim Slimani, Brahim Tamadazte, Catherine Achard
Mach. Vis. Appl.3
2024 RoCNet++: Triangle-based descriptor for accurate and robust point cloud registration
Karim Slimani, Catherine Achard, Brahim Tamadazte
Pattern Recognit.2
2023 Are we in sync during turn switch?
abstract
During an interaction, people exchange speaking turns by coordinating with their partners. Exchanges can be done smoothly, with pauses between turns or through interruptions. Previous studies have analyzed various modalities to investigate turn shifts and their types (smooth turn exchange, overlap, and interruption). Modality analyses were also done to study the interpersonal synchronization which is observed throughout the whole interaction. Likewise, we intend to analyze different modalities to find a relationship between the different turn switch types and interpersonal synchrony. In this study, we provide an analysis of multimodal features, focusing on prosodic features (F0 and loudness), head activity, and facial action units, to characterize different switch types.
Jieyeon Woo, Catherine Achard, Catherine Pelachaud
FG3
2023 Reciprocal Adaptation Measures for Human-Agent Interaction Evaluation
abstract
International audience
Jieyeon Woo, Catherine Pelachaud, Catherine Achard
ICAART (1)3
2023 ASAP: Endowing Adaptation Capability to Agent in Human-Agent Interaction
abstract
Socially Interactive Agents (SIAs) offer users with interactive face-to-face conversations. They can take the role of a speaker and communicate verbally and nonverbally their intentions and emotional states; but they should also act as active listener and be an interactive partner. In human-human interaction, interlocutors adapt their behaviors reciprocally and dynamically. The endowment of such adaptation capability can allow SIAs to show social and engaging behaviors. In this paper, we focus on modelizing the reciprocal adaptation to generate SIA behaviors for both conversational roles of speaker and listener. We propose the Augmented Self-Attention Pruning (ASAP) neural network model. ASAP incorporates recurrent neural network, attention mechanism of transformers, and pruning technique to learn the reciprocal adaptation via multimodal social signals. We evaluate our work objectively, via several metrics, and subjectively, through a user perception study where the SIA behaviors generated by ASAP is compared with those of other state-of-the-art models. Our results demonstrate that ASAP significantly outperforms the state-of-the-art models and thus shows the importance of reciprocal adaptation modeling.
Jieyeon Woo, Catherine Pelachaud, Catherine Achard
IUI3
2023 IAVA: Interactive and Adaptive Virtual Agent
abstract
During an interaction, partners adapt their behaviors to each other. Adaptation can have several functions such as being a sign of engagement and enhancing human users' interaction experience. It is important that virtual agents acting as interaction partners should continuously adapt their behaviors to those of their interlocutors in real time. This paper focuses on creating an interactive virtual agent that is capable of rendering real-time adaptive behaviors in response to its human interlocutor. It ensures the two aspects: generating real-time adaptive behavior and managing natural dialogue. We propose a system of an adaptive virtual agent and choose the e-health application of Cognitive Behavioral Therapy (CBT), which is a mental health treatment that restructures automatic thoughts into balanced thoughts, as a proof-of-concept to showcase the benefit of endowing behavior adaptation to the agent. The virtual agent adapts to the user via the display of nonverbal behaviors, which are generated via a deep learning model, throughout the whole interaction while acting as a therapist helping human users to detect their negative automatic thoughts.
Jieyeon Woo, Michele Grimaldi, Catherine Pelachaud, Catherine Achard
IVA4
2023 Conducting Cognitive Behavioral Therapy with an Adaptive Virtual Agent
abstract
When conversing, people adapt their behaviors to one another to show their engagement. Virtual agents, acting as interaction partners, should also adapt to their interlocutors in real time. In this paper, we introduce a virtual agent delivering Cognitive Behavioral Therapy (CBT) and adapting its behaviors in real time. The system focuses on the real-time generation of adaptive behavior and management of natural CBT dialogue.
Jieyeon Woo, Michele Grimaldi, Catherine Pelachaud, Catherine Achard
IVA4
2023 Now or When?: Interruption timing prediction in dyadic interaction
abstract
Interruptions are an important aspect of human-human communication. They help to adjust the conversation flow. Our aim is to equip virtual agents with the ability to handle interruptions, that is to decide when and how to interrupt their human interlocutor. In this paper, we focus on predicting when interruptions may occur during the conversation using multimodal features only from the speaker and propose a model trained on a corpus of dyadic interactions. To assess the model's accuracy, we conduct a perceptual study where we compare different timings (ground truth, randomly chosen or predicted by our model).
Catherine Achard, Catherine Pelachaud
IVA2
2022 Multimodal classification of interruptions in humans' interaction
abstract
During an interaction interruptions occur frequently. Interruptions may arise to fulfill different goals such as changing the topic of conversation abruptly, asking for clarification, completing the current speaker’s turn. Interruptions may be cooperative or competitive depending on the interrupter’s intention. Our main goal is to endow a Socially Interactive Agent with the capacity to handle user interruptions in dyadic interaction. It requires the agent to detect an interruption and recognize its type (cooperative/competitive), and then to plan its behaviours to respond appropriately. As a first step towards this goal, we developed a multimodal classification model using acoustic features, facial expression, head movement, and gaze direction from both, the interrupter and the interruptee. The classification model learns from the sequential information to automatically identify interruptions type. We also present studies we conducted to measure the shortest delay needed (0.6s) for our classification model to identify interruption types with a high classification accuracy (81%). On average, most interruption overlaps last longer than 0.6s, so a Socially Interactive Agent has time to detect and recognize an interruption type and can respond in a timely manner to its human interlocutor’s interruption.
Catherine Achard, Catherine Pelachaud
ICMI2
2022 Annotating Interruption in Dyadic Human Interaction
abstract
Integrating the existing interruption and turn switch classification methods, we propose a new annotation schema to annotate different types of interruptions through timeliness, switch accomplishment and speech content level. The proposed method is able to distinguish smooth turn exchange, backchannel and interruption (including interruption types) and to annotate dyadic conversation. We annotated the French part of NoXi corpus with the proposed structure and use these annotations to study the probability distribution and duration of each turn switch type.
Catherine Achard, Catherine Pelachaud
LREC2
2021 Providing Automatic Feedback to Trainees after Automatic Evaluation
abstract
Learning how to perform precise and controlled gestures is difficult, especially when feedback about made errors is sparse. Therefore, some works try to facilitate learning by providing virtual "coaches". Most of them propose to automatically score task quality. But simply assessing quality through a score is not enough. Indeed, it is essential to provide explanations on assigned scores just like experts do when supervising trainees. However when quality assessment is done automatically, such explanations are rare and computing an automatic feedback is complex. In this work, we propose to address this problem by providing an automatic feedback based on neural network explanation. Contrary to previous state of the art methods, which are focused on neural networks explicability for classification tasks, we want to explain network decision on a regression problem (quality score prediction). Thus, we propose to use gradient-based approaches and adapt them to a regression task. Moreover, to address the problem of noise present in sensitivity maps, we propose a solution that leads to more robust gradients. To test our approach, since automatic quality assessment datasets do not contain ground truth about errors position and amplitude, a synthetic dataset representing a simple temporal task has been created, with its associated ground truth. Once the method has been validated on this synthetic dataset, we apply it on real data composed of robotic surgical gestures.
Mégane Millan, Catherine Achard
ICRA2
2021 Interruptions in Human-Agent Interaction
abstract
Turn management is one of the necessary social interactions skills. In human-human interactions, turn changes are naturally completed by interruption, "cooperatively" or "competitively". Interruptions are inherent in conversation. They can be considered disruptive at first glance, but can also be cooperative and participate to enriching the interaction. To create natural human-agent interaction, Embodied Conversational Agent (ECA) should be able to communicate autonomously with humans both verbally and nonverbally. A challenge is then to handle interruptions during their interaction. This article presents our ongoing work to endow ECA to manage interruption during the interaction with a human partner. In order to achieve this goal, we start by analyzing human-human interaction data.
Catherine Achard, Catherine Pelachaud
IVA2
2021 SALAD: Self-Assessment Learning for Action Detection
abstract
Literature on self-assessment in machine learning mainly focuses on the production of well-calibrated algorithms through consensus frameworks i.e. calibration is seen as a problem. Yet, we observe that learning to be properly confident could behave like a powerful regularization and thus, could be an opportunity to improve performance.Precisely, we show that used within a framework of action detection, the learning of a self-assessment score is able to improve the whole action localization process. Experimental results show that our approach outperforms the state-of-the-art on two action detection benchmarks. On THUMOS14 dataset, the mAP at [email protected] is improved from 42.8% to 44.6%, and from 50.4% to 51.7% on ActivityNet1.3 dataset. For lower tIoU values, we achieve even more significant improvements on both datasets.
Guillaume Vaudaux-Ruth, Adrien Chan-Hon-Tong, Catherine Achard
WACV3
2021 Single-shot 3D multi-person pose estimation in complex images
Abdallah Benzine, Bertrand Luvison, Quoc Cuong Pham, Catherine Achard
Pattern Recognit.4
2020 PandaNet: Anchor-Based Single-Shot Multi-Person 3D Pose Estimation
abstract
Recently, several deep learning models have been proposed for 3D human pose estimation. Nevertheless, most of these approaches only focus on the single-person case or estimate 3D pose of a few people at high resolution. Furthermore, many applications such as autonomous driving or crowd analysis require pose estimation of a large number of people possibly at low-resolution. In this work, we present PandaNet (Pose estimAtioN and Dectection Anchor-based Network), a new single-shot, anchor-based and multi-person 3D pose estimation approach. The proposed model performs bounding box detection and, for each detected person, 2D and 3D pose regression into a single forward pass. It does not need any post-processing to regroup joints since the network predicts a full 3D pose for each bounding box and allows the pose estimation of a possibly large number of people at low resolution. To manage people overlapping, we introduce a Pose-Aware Anchor Selection strategy. Moreover, as imbalance exists between different people sizes in the image, and joints coordinates have different uncertainties depending on these sizes, we propose a method to automatically optimize weights associated to different people scales and joints for efficient training. PandaNet surpasses previous single-shot methods on several challenging datasets: a multi-person urban virtual but very realistic dataset (JTA Dataset), and two real world 3D multi-person datasets (CMU Panoptic and MuPoTS-3D).
Abdallah Benzine, Florian Chabot, Bertrand Luvison, Quoc Cuong Pham, Catherine Achard
CVPR5
2020 ActionSpotter: Deep Reinforcement Learning Framework for Temporal Action Spotting in Videos
abstract
Action spotting has recently been proposed as an alternative to action detection and key frame extraction. However, the current state-of-the-art method of action spotting requires an expensive ground truth composed of the search sequences employed by human annotators spotting actions - a critical limitation. In this article, we propose to use a reinforcement learning algorithm to perform efficient action spotting using only the temporal segments from the action detection annotations, thus opening an interesting solution for video understanding. Experiments performed on THUMOS14 and ActivityNet datasets show that the proposed method, named ActionSpotter, leads to good results and outperforms state-of-the-art detection outputs redrawn for this application. In particular, the spotting mean Average Precision on THUMOS14 is significantly improved from 59.7% to 65.6% while skipping 23% of video.
Guillaume Vaudaux-Ruth, Adrien Chan-Hon-Tong, Catherine Achard
ICPR3
2019 Deep, Robust and Single Shot 3D Multi-Person Human Pose Estimation from Monocular Images
abstract
In this paper, we propose a new single shot method for multi-person 3D pose estimation, from monocular RGB images. Our model jointly learns to locate the human joints in the image, to estimate their 3D coordinates and to group these predictions into full human skeletons. Our approach leverages and extends the Stacked Hourglass Network and its multi-scale feature learning to manage multi-person situations. Thus, we exploit the Occlusions Robust Pose Maps (ORPM) to fully describe several 3D human poses even in case of strong occlusions or cropping. Then, joint grouping and human pose estimation for an arbitrary number of people are performed using associative embedding. We evaluate our method on the challenging CMU Panoptic dataset, and demonstrate that it achieves better results than the state of the art.
Abdallah Benzine, Bertrand Luvison, Quoc Cuong Pham, Catherine Achard
ICIP4
2019 Atomic force microscope tip localization and tracking through deep learning based vision inside an electron microscope
abstract
Scanning Electron Microscopy (SEM) is an ideal observation tool for small scales robotics. It has the potential to achieve automated nano-robotic tasks such as nano-handling and nano-assembly. Path following control of nano-robot end effectors using SEM vision feedback is a key for an intuitive programming of elementary robotic tasks sequences. It requires the ability to track end effectors under various SEM scan speeds. SEM suffers however from tricky issues that limits robotic tracking capabilities. This paper focuses on one specific issue related to the compromise between the scan speed and the image quality. This restriction seriously limits the performance of conventional vision tracking algorithms when used with electron images. At high scan speed, the image quality is very noisy making very difficult to differentiate the robot end effector from the background, hence limiting the tracking capabilities. The work related in this paper explores for the first time the potential value of Convolutional Neural Networks (ConvNet) in the context of nano-robotic vision tracking inside SEM. The aim is to localize an end-effector, AFM cantilever in the case of the study, from SEM images for any scan speed configuration and despite of low images quality. For that purpose, a data set of AFM tip images is build up from SEM images for the learning algorithm. Network performances are estimated under different SEM scan speeds. Thanks to the learning algorithm, experimental results show robust AFM tip tracking capabilities inside the SEM under various scan speed conditions.
Mokrane Boudaoud, Catherine Achard, Weibin Rong, Stéphane Régnier
IROS3
2018 Modeling the synchrony between interacting people: application to role recognition
Sheng Fang 0003, Catherine Achard, Séverine Dubuisson
Multim. Tools Appl.2
2018 Time-series averaging using constrained dynamic time warping with tolerance
Marion Morel, Catherine Achard, Richard Kulpa, Séverine Dubuisson
Pattern Recognit.2
2017 The DAily Home LIfe Activity Dataset: A High Semantic Activity Dataset for Online Recognition
abstract
In this article, we introduce the DAily Home LIfe Activity (DAHLIA) Dataset, a new dataset adapted to the context of smart-home or video-assistance. Videos were recorded in realistic conditions, with 3 KinectTMv2 sensors located as they would be in a real context. The very long-range activities were performed in an unconstrained way (participants received few instructions), and in a continuous (untrimmed) sequence, resulting in long videos (39 min in average per subject). Contrary to previously published databases, in which labeled actions are very short and with low-semantic level, this new database focuses on high-level semantic activities such as ”Preparing lunch” or ”House Working”. As a baseline, we evaluated several metrics on three different algorithms designed for online action recognition or detection.
Geoffrey Vaquette, Astrid Orcesi, Laurent Lucat, Catherine Achard
FG4
2017 Automatic evaluation of sports motion: A generic computation of spatial and temporal errors
Marion Morel, Catherine Achard, Richard Kulpa, Séverine Dubuisson
Image Vis. Comput.2
2016 Bidirectional sparse representations for multi-shot person re-identification
abstract
With the development of surveillance cameras, person re-identification has gained much interest, however re-identifying people across cameras remains a challenging problem which not only requires a good feature description but also a reliable matching scheme. Our method can be applied with any feature and focuses on the second requirement. We propose a robust bidirectional sparse coding method that improves simple sparse coding performances. Some recent work have already explored sparse representation for the re-identification task but none has considered the problem from both the probe and the gallery perspectives. We propose a bidirectional sparse representations method which searches for the most likely match for the test element in the gallery set and makes sure that the selected gallery match is indeed closely related to the probe. Extensive experiments on two datasets, CUHK03 and iLIDS-VID, show the effectiveness of our approach.
Solene Chan-Lang, Quoc Cuong Pham, Catherine Achard
AVSS3
2016 Personality classification and behaviour interpretation: an approach based on feature categories
abstract
This paper focuses on recognizing and understanding social dimensions (the personality traits and social impressions) during small group interactions. We extract a set of audio and visual features, which are divided into three categories: intra-personal features (i.e. related to only one participant), dyadic features (i.e. related to a pair of participants) and one vs all features (i.e. related to one participant versus the other members of the group). First, we predict the personality traits (PT) and social impressions (SI) by using these three feature categories. Then, we analyse the interplay be- tween groups of features and the personality traits/social impressions of the interacting participants. The prediction is done by using Support Vector Machine and Ridge Regression which allows to determine the most dominant features for each social dimension. Our experiments show that the combination of intra-personal and one vs all features can greatly improve the prediction accuracy of personality traits and social impressions. Prediction accuracy reaches 81.37% for the social impression named ’Rank of Dominance’. Finally, we draw some interesting conclusions about the relationship between personality traits/social impressions and social features.
Sheng Fang 0003, Catherine Achard, Séverine Dubuisson
ICMI2
2015 Exploiting 3D geometric primitives for multicamera pedestrian detection
abstract
In this paper, we present an approach for multicamera pedestrian detection exploiting the concepts of multiview geometry and the shapes of 3D geometric primitives. Multicamera occupancy maps provide peak responses corresponding to the object detection but suffer from several false detections known as ghosts. The novelty of this paper is the introduction of shape patterns which can model the objects, such as pedestrians, by defining a kernel function in the projected occupancy space. This kernel depends upon the geometry of the 3D primitives and also varies in relation to their position with respect to the cameras in the real world configuration. For multiple objects visible across several cameras, we define a formation model which is the convolution of this spatially varying kernel with the set of possible object locations. The locations corresponding to detections can thus be obtained through a deconvolution process. For efficient computations, we further propose an estimated deconvolution process specific to our kernel responses which can also be heavily parallelized. We show the application of this process towards pedestrian detection by studying various 3D cylindrical primitives. Experiments on two public dataset sequences, including comparison with another approach, show the efficiency of the proposed method in terms of pedestrian detection and ghost pruning, including in adverse and challenging conditions.
Muhammad Owais Mehmood, Sebastien Ambellouis, Catherine Achard
AVSS3
2014 Simultaneous segmentation and classification of human actions in video streams using deeply optimized Hough transform
Adrien Chan-Hon-Tong, Catherine Achard, Laurent Lucat
Pattern Recognit.2
2010 People Reacquisition across Multiple Cameras with Disjoint Views
Dung Nghi Truong Cong, Louahdi Khoudour, Catherine Achard
ICISP3
2010 A Hough Transform with projection for velocity estimation
Ghilès Mostafaoui, Catherine Achard, Maurice Milgram
Mach. Vis. Appl.2
2010 People re-identification by spectral classification of silhouettes
Dung Nghi Truong Cong, Louahdi Khoudour, Catherine Achard, Cyril Meurie, Olivier Lézoray
Signal Process.3
2008 A novel approach for recognition of human actions with semi-global features
Catherine Achard, Xingtai Qu, Arash Mokhber, Maurice Milgram
Mach. Vis. Appl.1
2008 Recognition of human behavior by space-time silhouette characterization
Arash Mokhber, Catherine Achard, Maurice Milgram
Pattern Recognit. Lett.2
2007 Action Recognition with Semi-global Characteristics and Hidden Markov Models
Catherine Achard, Xingtai Qu, Arash Mokhber, Maurice Milgram
ACIVS1
2005 Real Time Tracking of Multiple Persons on Colour Image Sequences
Ghilès Mostafaoui, Catherine Achard, Maurice Milgram
ACIVS2
2005 Real time tracking of multiple persons using elementary tracks
abstract
We propose a real time tracking algorithm for a counting person application. Our algorithm needs any a priori knowledge neither on the model of person, nor on their size or their number, which can evolve with time. It manages several problems such as occlusion and under or over-segmentations especially thanks to a shape model obtained with a homography. The first step consisting in motion detection, leads to regions that have to be assigned to trajectories. This tracking step is achieved using a new concept: elementary tracks. They allow on the one hand to manage the tracking and on the other hand, to detect the output of occlusion by introducing coherent sets of regions. Those sets enable to define temporal kinematical, shape and colour models. Significant results will be presented on several sequences with ground truth.
Ghilès Mostafaoui, Catherine Achard, Maurice Milgram
AVSS2
2000 A Sub-Pixel and Multispectral Corner Detector
abstract
Proposes in a corner detector algorithm, which leads to results both on mono-spectral and multispectral images. To validate the method, we compare its mono-spectral version to the Harris detector, which is the most frequently used in literature. This study shows that the proposed method gives generally more efficient results. However, bad localisations appear for very blurred images (as for most corner detectors). Therefore, we have implemented a sub-pixel detector able to find the exact corner position.
Catherine Achard, Erwan Bigorgne, Jean Devars
ICPR1
2000 Object Image Retrieval with Image Compactness Vectors
abstract
We present in this paper a new global measure to characterise an image: the compactness vector. This measure considers both object shape and grey level distribution function and does not require any preliminary segmentation. It is invariant to rotation, translation, scale and luminance and is then a powerful tool for image retrieval from a query image. We present here some object retrieval examples from large database images.
Catherine Achard, Jean Devars, Lionel Lacassagne
ICPR1
2000 An Invariant Local Vector for Content-Based Image Retrieval
abstract
In this paper, we present the use of Full-Zernike moments as a local characterization of the image signal. Their computation allows us to construct a locally invariant vector, of which the projection in an index table provides a vote for some model-image. This approach is based on the quasi-invariant theory applied to perspective transformation. Then it requires a characterization being invariant to translation, rotation and change of scale in the image; in other respect, an appropriate normalization of the signal delivers an invariance to illuminance conditions.
Erwan Bigorgne, Catherine Achard, Jean Devars
ICPR2