EDBT 2026 Demo / reviewers in the wild / expert
Nicolae-Catalin Ristea
dblp:253/8663 · also Catalin Nicolae Ristea
· DBLP profile ↗
24ranked-venue papers
14as first author
21since 2021 · last 2025
0000-0002-7880-9307ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 7 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 7 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 5 first-author · 6 since 2021Systems, architecture and hardware · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Learning Rate CurriculumabstractAbstract Most curriculum learning methods require an approach to sort the data samples by difficulty, which is often cumbersome to perform. In this work, we propose a novel curriculum learning approach termed Learning Rate Curriculum (LeRaC), which leverages the use of a different learning rate for each layer of a neural network to create a data-agnostic curriculum during the initial training epochs. More specifically, LeRaC assigns higher learning rates to neural layers closer to the input, gradually decreasing the learning rates as the layers are placed farther away from the input. The learning rates increase at various paces during the first training iterations, until they all reach the same value. From this point on, the neural model is trained as usual. This creates a model-level curriculum learning strategy that does not require sorting the examples by difficulty and is compatible with any neural network, generating higher performance levels regardless of the architecture. We conduct comprehensive experiments on 12 data sets from the computer vision (CIFAR-10, CIFAR-100, Tiny ImageNet, ImageNet-1K, Food-101, UTKFace, PASCAL VOC), language (BoolQ, QNLI, RTE) and audio (ESC-50, CREMA-D) domains, considering various convolutional (ResNet-18, Wide-ResNet-50, DenseNet-121, YOLOv5), recurrent (LSTM) and transformer (CvT, BERT, SepTr) architectures. We compare our approach with the conventional training regime, as well as with Curriculum by Smoothing (CBS), a state-of-the-art data-agnostic curriculum learning approach. Unlike CBS, our performance improvements over the standard training regime are consistent across all data sets and models. Furthermore, we significantly surpass CBS in terms of training time (there is no additional cost over the standard training regime for LeRaC). Our code is freely available at: https://github.com/CroitoruAlin/LeRaC . Florinel-Alin Croitoru, Nicolae-Catalin Ristea, Radu Tudor Ionescu, Nicu Sebe |
Int. J. Comput. Vis. | 2 |
| 2024 | Self-Distilled Masked Auto-Encoders are Efficient Video Anomaly DetectorsabstractWe propose an efficient abnormal event detection model based on a lightweight masked auto-encoder (AE) applied at the video frame level. The novelty of the proposed model is threefold. First, we introduce an approach to weight tokens based on motion gradients, thus shifting the focus from the static background scene to the foreground objects. Second, we integrate a teacher decoder and a student decoder into our architecture, leveraging the discrepancy between the outputs given by the two decoders to improve anomaly detection. Third, we generate synthetic abnormal events to augment the training videos, and task the masked AE model to jointly reconstruct the original frames (without anomalies) and the corresponding pixel-level anomaly maps. Our design leads to an efficient and effective model, as demonstrated by the extensive experiments carried out on four benchmarks: Avenue, Shanghai Tech, UBnormal and UCSD Ped2. The empirical results show that our model achieves an excellent trade-off between speed and accuracy, obtaining competitive AUC scores, while processing 1655 FPS. Hence, our model is between 8 and 70 times faster than competing methods. We also conduct an ablation study to justify our design. Our code is freely available at: https://github.com/ristea/aed-mae. Nicolae-Catalin Ristea, Florinel-Alin Croitoru, Radu Tudor Ionescu, Marius Popescu, Fahad Shahbaz Khan, Mubarak Shah |
CVPR | 1 |
| 2024 | Multi-Dimensional Speech Quality Assessment in CrowdsourcingabstractSubjective speech quality assessment is the gold standard for evaluating speech enhancement processing and telecommunication systems. The commonly used standard ITU-T Rec. P.800 defines how to measure speech quality in lab environments, and ITU-T Rec. P.808 extended it for crowdsourcing. ITU-T Rec. P.835 extends P.800 to measure the quality of speech in the presence of noise. ITU-T Rec. P.804 targets the conversation test and introduces perceptual speech quality dimensions which are measured during the listening phase of the conversation. The perceptual dimensions are noisiness, coloration, discontinuity, and loudness. We create a crowd-sourcing implementation of a multi-dimensional subjective test following the scales from P.804 and extend it to include reverberation, the speech signal, and overall quality. We show the tool is both accurate and reproducible. The tool has been used in the ICASSP 2023 Speech Signal Improvement challenge and we show the utility of these speech quality dimensions in this challenge. The tool will be publicly available as open-source at https://github.com/microsoft/P.808. Babak Naderi, Ross Cutler, Nicolae-Catalin Ristea |
ICASSP | 3 |
| 2024 | Multi-Head Transposed Attention Transformer for Sea Ice Segmentation in Sar ImageryabstractSea ice plays a pivotal role in the Earth’s climate system and exhibits high sensitivity to shifts in temperature and atmospheric conditions. The precise and timely assessment of sea ice parameters is essential for comprehending and forecasting the climate changes. However, the vast volume of satellite data covering ice-covered regions is impractical to be subjectively assessed. Hence, the utilization of automated algorithms becomes mandatory to fully exploit the continuous data streams from satellites. In this paper, we propose a UNet transformer-based architecture, called UT-MHTA, to sea ice segmentation using SAR satellite imagery. Our UT-MHTA network replaces the conventional multi-head attention (MHA) block with a multi-head transposed attention (MHTA) which can capture long-range pixel interactions, while still remaining suitable for large images. Our method demonstrates superior performance compared to state-of-the-art methods, without drastically raising the computational complexity. In particular, UT-MHTA achieves a mean intersection over union (mIoU) of 68.76% on the AI4Arctic data set, with an inference time of 865ms for a 400 km2product. Nicolae-Catalin Ristea, Andrei Anghel, Alexis Mouche, Frédéric Nouguier, Antoine Grouazel, Mihai Datcu |
IGARSS | 1 |
| 2024 | CL-MAE: Curriculum-Learned Masked AutoencodersabstractMasked image modeling has been demonstrated as a powerful pretext task for generating robust representations that can be effectively generalized across multiple downstream tasks. Typically, this approach involves randomly masking patches (tokens) in input images, with the masking strategy remaining unchanged during training. In this paper, we propose a curriculum learning approach that updates the masking strategy to continually increase the complexity of the self-supervised reconstruction task. We conjecture that, by gradually increasing the task complexity, the model can learn more sophisticated and transferable representations. To facilitate this, we introduce a novel learnable masking module that possesses the capability to generate masks of different complexities, and integrate the proposed module into masked autoencoders (MAE). Our module is jointly trained with the MAE, while adjusting its behavior during training, transitioning from a partner to the MAE (optimizing the same reconstruction loss) to an adversary (optimizing the opposite loss), while passing through a neutral state. The transition between these behaviors is smooth, being regulated by a factor that is multiplied with the reconstruction loss of the masking module. The resulting training procedure generates an easy-to-hard curriculum. We train our Curriculum-Learned Masked Autoencoder (CL-MAE) on ImageNet and show that it exhibits superior representation learning capabilities compared to MAE. The empirical results on five downstream tasks confirm our conjecture, demonstrating that curriculum learning can be successfully used to self-supervise masked autoencoders. We release our code at https://github.com/ristea/cl-mae. Neelu Madan, Nicolae-Catalin Ristea, Kamal Nasrollahi, Thomas B. Moeslund, Radu Tudor Ionescu |
WACV | 2 |
| 2024 | Lightning fast video anomaly detection via multi-scale adversarial distillationabstractWe propose a very fast frame-level model for anomaly detection in video, which learns to detect anomalies by distilling knowledge from multiple highly accurate object-level teacher models. To improve the fidelity of our student, we distill the low-resolution anomaly maps of the teachers by jointly applying standard and adversarial distillation, introducing an adversarial discriminator for each teacher to distinguish between target and generated anomaly maps. We conduct experiments on three benchmarks (Avenue, ShanghaiTech, UCSD Ped2), showing that our method is over 7 times faster than the fastest competing method, and between 28 and 62 times faster than object-centric models, while obtaining comparable results to recent methods. Our evaluation also indicates that our model achieves the best trade-off between speed and accuracy, due to its previously unheard-of speed of 1480 FPS. In addition, we carry out a comprehensive ablation study to justify our architectural design choices. Our code is freely available at: https://github.com/ristea/fast-aed. Florinel-Alin Croitoru, Nicolae-Catalin Ristea, Dana Dascalescu, Radu Tudor Ionescu, Fahad Shahbaz Khan, Mubarak Shah |
Comput. Vis. Image Underst. | 2 |
| 2024 | Self-Supervised Masked Convolutional Transformer Block for Anomaly DetectionabstractAnomaly detection has recently gained increasing attention in the field of computer vision, likely due to its broad set of applications ranging from product fault detection on industrial production lines and impending event detection in video surveillance to finding lesions in medical scans. Regardless of the domain, anomaly detection is typically framed as a one-class classification task, where the learning is conducted on normal examples only. An entire family of successful anomaly detection methods is based on learning to reconstruct masked normal inputs (e.g. patches, future frames, etc.) and exerting the magnitude of the reconstruction error as an indicator for the abnormality level. Unlike other reconstruction-based methods, we present a novel self-supervised masked convolutional transformer block (SSMCTB) that comprises the reconstruction-based functionality at a core architectural level. The proposed self-supervised block is extremely flexible, enabling information masking at any layer of a neural network and being compatible with a wide range of neural architectures. In this work, we extend our previous self-supervised predictive convolutional attentive block (SSPCAB) with a 3D masked convolutional layer, a transformer for channel-wise attention, as well as a novel self-supervised objective based on Huber loss. Furthermore, we show that our block is applicable to a wider variety of tasks, adding anomaly detection in medical images and thermal videos to the previously considered tasks based on RGB images and surveillance videos. We exhibit the generality and flexibility of SSMCTB by integrating it into multiple state-of-the-art neural models for anomaly detection, bringing forth empirical results that confirm considerable performance improvements on five benchmarks: MVTec AD, BRATS, Avenue, ShanghaiTech, and Thermal Rare Event. Neelu Madan, Nicolae-Catalin Ristea, Radu Tudor Ionescu, Kamal Nasrollahi, Fahad Shahbaz Khan, Thomas B. Moeslund, Mubarak Shah |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Sea Ice Segmentation from SAR Data by Convolutional Transformer NetworksabstractSea ice is a crucial component of the Earth’s climate system and is highly sensitive to changes in temperature and atmospheric conditions. Accurate and timely measurement of sea ice parameters is important for understanding and predicting the impacts of climate change. Nevertheless, the amount of satellite data acquired over ice areas is huge, making the subjective measurements ineffective. Therefore, automated algorithms must be used in order to fully exploit the continuous data feeds coming from satellites. In this paper, we present a novel approach for sea ice segmentation based on SAR satellite imagery using hybrid convolutional transformer (ConvTr) networks. We show that our approach outperforms classical convolutional networks, while being considerably more efficient than pure transformer models. ConvTr obtained a mean intersection over union (mIoU) of 63.68% on the AI4Arctic data set, assuming an inference time of 120ms for a 400×400 km2product. Nicolae-Catalin Ristea, Andrei Anghel, Mihai Datcu |
IGARSS | 1 |
| 2023 | DeepVQE: Real Time Deep Voice Quality Enhancement for Joint Acoustic Echo Cancellation, Noise Suppression and Dereverberation
Nicolae-Catalin Ristea, Evgenii Indenbom, Ando Saabas, Tanel Pärnamaa, Jegor Guzvin, Ross Cutler |
INTERSPEECH | 1 |
| 2023 | Cascaded Cross-Modal Transformer for Request and Complaint DetectionabstractWe propose a novel cascaded cross-modal transformer (CCMT) that combines speech and text transcripts to detect customer requests and complaints in phone conversations. Our approach leverages a multimodal paradigm by transcribing the speech using automatic speech recognition (ASR) models and translating the transcripts into different languages. Subsequently, we combine language-specific BERT-based models with Wav2Vec2.0 audio features in a novel cascaded cross-attention transformer model. We apply our system to the Requests Sub-Challenge of the ACM Multimedia 2023 Computational Paralinguistics Challenge, reaching unweighted average recalls (UAR) of 65.41% and 85.87% for the complaint and request classes, respectively. Nicolae-Catalin Ristea, Radu Tudor Ionescu |
ACM Multimedia | 1 |
| 2023 | Multimodal Multi-Head Convolutional Attention with Various Kernel Sizes for Medical Image Super-ResolutionabstractSuper-resolving medical images can help physicians in providing more accurate diagnostics. In many situations, computed tomography (CT) or magnetic resonance imaging (MRI) techniques capture several scans (modes) during a single investigation, which can jointly be used (in a multimodal fashion) to further boost the quality of super-resolution results. To this end, we propose a novel multi-modal multi-head convolutional attention module to super-resolve CT and MRI scans. Our attention module uses the convolution operation to perform joint spatial-channel attention on multiple concatenated input tensors, where the kernel (receptive field) size controls the reduction rate of the spatial attention, and the number of convolutional filters controls the reduction rate of the channel attention, respectively. We introduce multiple attention heads, each head having a distinct receptive field size corresponding to a particular reduction rate for the spatial attention. We integrate our multimodal multi-head convolutional attention (MMHCA) into two deep neural architectures for super-resolution and conduct experiments on three data sets. Our empirical results show the superiority of our attention module over the state-of-the-art attention mechanisms used in super-resolution. Moreover, we conduct an ablation study to assess the impact of the components involved in our attention module, e.g. the number of inputs or the number of heads. Our code is freely available at https://github.com/lilygeorgescu/MHCA. Mariana-Iuliana Georgescu, Radu Tudor Ionescu, Andreea-Iuliana Miron, Olivian Savencu, Nicolae-Catalin Ristea, Nicolae Verga, Fahad Shahbaz Khan |
WACV | 5 |
| 2023 | Nonlinear neurons with human-like apical dendrite activations
Mariana-Iuliana Georgescu, Radu Tudor Ionescu, Nicolae-Catalin Ristea, Nicu Sebe |
Appl. Intell. | 3 |
| 2023 | CyTran: A cycle-consistent transformer with multi-level consistency for non-contrast to contrast CT translation
Nicolae-Catalin Ristea, Andreea-Iuliana Miron, Olivian Savencu, Mariana-Iuliana Georgescu, Nicolae Verga, Fahad Shahbaz Khan, Radu Tudor Ionescu |
Neurocomputing | 1 |
| 2023 | Guided Unsupervised Learning by Subaperture Decomposition for Ocean SAR Image RetrievalabstractSpaceborne synthetic aperture radar (SAR) can provide accurate images of the ocean surface roughness day-or-night in nearly all weather conditions, being an unique asset for many geophysical applications. Considering the huge amount of data daily acquired by satellites, automated techniques for physical features extraction are needed. Even if supervised deep learning methods attain state-of-the-art results, they require a great amount of labelled data, which are difficult and excessively expensive to acquire for ocean SAR imagery. To this end, we use the subaperture decomposition (SD) algorithm to enhance the unsupervised learning retrieval on the ocean surface, empowering ocean researchers to search into large ocean databases. We empirically prove that SD improves the retrieval precision with over 20% for an unsupervised transformer auto-encoder network. Moreover, we show that SD brings an important performance boost when Doppler centroid images are used as input data, leading the way to new unsupervised physics guided retrieval algorithms. Nicolae-Catalin Ristea, Andrei Anghel, Mihai Datcu, Bertrand Chapron |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Self-Supervised Predictive Convolutional Attentive Block for Anomaly DetectionabstractAnomaly detection is commonly pursued as a one-class classification problem, where models can only learn from normal training samples, while being evaluated on both normal and abnormal test samples. Among the successful approaches for anomaly detection, a distinguished category of methods relies on predicting masked information (e.g. patches, future frames, etc.) and leveraging the reconstruction error with respect to the masked information as an abnormality score. Different from related methods, we propose to integrate the reconstruction-based functionality into a novel self-supervised predictive architectural building block. The proposed self-supervised block is generic and can easily be incorporated into various state-of-the-art anomaly detection methods. Our block starts with a convolutional layer with dilated filters, where the center area of the receptive field is masked. The resulting activation maps are passed through a channel attention module. Our block is equipped with a loss that minimizes the reconstruction error with respect to the masked area in the receptive field. We demonstrate the generality of our block by integrating it into several state-of-the-art frameworks for anomaly detection on image and video, providing empirical evidence that shows considerable performance improvements on MVTec AD, Avenue, and ShanghaiTech. We release our code as open source at: https://github.com/ristea/sspcab. Nicolae-Catalin Ristea, Neelu Madan, Radu Tudor Ionescu, Kamal Nasrollahi, Fahad Shahbaz Khan, Thomas B. Moeslund, Mubarak Shah |
CVPR | 1 |
| 2022 | Convolutional Transformersl for Aerial Image Classification: a General to Specific Learning CurveabstractRemote sensing image classification is at the center of many tasks in the remote sensing domain. However, the complexity and content variety of aerial images contribute to making the task still challenging. Transformers have recently achieved state-of-the-art performances for numerous natural language processing and image processing tasks. In this paper, we pro-pose a novel solution towards remote sensing image classification based on a general-to-specific learning curve achieved through a cascaded chain of Vision Transformers (Vit). First, a standard pre-trained Vision Transformer (ViT) is used to provide general information regarding the remote sensing scenes, whereas the specific details are learned by means of a Convolutional Vision Transformer (CvT) which is trained end-to-end. The experiments conducted over two benchmark datasets of high resolution remote sensing images show the effectiveness of the proposed technique. Comparisons to other methods in the literature are also provided. Mihail-Antonio Chirtu, Nicolae-Catalin Ristea, Anamaria Radoi |
IGARSS | 2 |
| 2022 | Guided Deep Learning by Subaperture Decomposition: Ocean Patterns from SAR ImageryabstractSpaceborne synthetic aperture radar (SAR) can provide meters-scale images of the ocean surface roughness day-or-night in nearly all weather conditions. This makes it a unique asset for many geophysical applications. Sentinel-l SAR wave mode (WV) vignettes have made possible to capture many important oceanic and atmospheric phenomena since 2014. However, considering the amount of data provided, expanding applications requires a strategy to automatically process and extract geophysical parameters. In this study, we propose to apply subaperture decomposition (SD) as a preprocessing stage for SAR deep learning models. Our data-centring approach surpassed the baseline by 0.7%, obtaining state-of-the-art on the TenGeoP-SARwv data set. In addition, we empirically showed that SD could bring additional information over the original vignette, by rising the number of clusters for an unsupervised segmentation method. Overall, we encourage the development of data-centring approaches, showing that, data preprocessing could bring significant performance improvements over existing deep learning models. Nicolae-Catalin Ristea, Andrei Anghel, Mihai Datcu, Bertrand Chapron |
IGARSS | 1 |
| 2022 | SepTr: Separable Transformer for Audio Spectrogram ProcessingabstractFollowing the successful application of vision transformers in multiple computer vision tasks, these models have drawn the attention of the signal processing community. This is because signals are often represented as spectrograms (e.g. through Discrete Fourier Transform) which can be directly provided as input to vision transformers. However, naively applying transformers to spectrograms is suboptimal. Since the axes represent distinct dimensions, i.e. frequency and time, we argue that a better approach is to separate the attention dedicated to each axis. To this end, we propose the Separable Transformer (SepTr), an architecture that employs two transformer blocks in a sequential manner, the first attending to tokens within the same time interval, and the second attending to tokens within the same frequency bin. We conduct experiments on three benchmark data sets, showing that our separable architecture outperforms conventional vision transformers and other state-of-the-art methods. Unlike standard transformers, SepTr linearly scales the number of trainable parameters with the input size, thus having a lower memory footprint. Our code is available as open source at https://github.com/ristea/septr. Nicolae-Catalin Ristea, Radu Tudor Ionescu, Fahad Shahbaz Khan |
INTERSPEECH | 1 |
| 2022 | Complex Neural Networks for Estimating Epicentral Distance, Depth, and Magnitude of Seismic WavesabstractTaking advantage of the latest advances in deep learning for seismology, we address earthquake characterization from a data-driven perspective. Many of the usual procedures for extracting information from seismograms require processing a large volume of data using empirical and physics rule-based techniques. In this letter, we propose a novel approach for estimating epicentral distance, depth, and magnitude directly from individual raw three-component seismograms of 1-min length observed by single stations. Our convolutional neural network-based method is able to handle complex-valued representations of the seismic data in the time–frequency domain by using dedicated convolutional and activation functions. In this way, our method benefits both from extracting relevant information through time–frequency domain analysis and from designing a single architecture that deals with complex information. The proposed method achieves a mean absolute error of 4.51 km for epicentral distance, 6.15 km for depth, and 0.26 for magnitude estimation. The experiments were conducted over a publicly available and large database, STanford EArthquake data set (STEAD), and the comparisons with current state-of-the-art approaches show the effectiveness of the proposed approach. Source code and best model are available athttps://github.com/ristea/stead-earthquake-cnn. Nicolae-Catalin Ristea, Anamaria Radoi |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2021 | Programmable Systems for Intelligence in Automobiles (PRYSTINE): Final results after Year 3abstractAutonomous driving is disrupting the automotive industry as we know it today. For this, fail-operational behavior is essential in the sense, plan, and act stages of the automation chain in order to handle safety-critical situations on its own, which currently is not reached with state-of-the-art approaches.The European ECSEL research project PRYSTINE realizes Fail-operational Urban Surround perceptION (FUSION) based on robust Radar and LiDAR sensor fusion and control functions in order to enable safe automated driving in urban and rural environments. This paper showcases some of the key exploitable results (e.g., novel Radar sensors, innovative embedded control and E/E architectures, pioneering sensor fusion approaches, AI-controlled vehicle demonstrators) achieved until its final year 3. Norbert Druml, Anna Ryabokon, Rupert Schorn, Jochen Koszescha, Kaspars Ozols, Aleksandrs Levinskis, Rihards Novickis, Ethiopia Nigussie, Jouni Isoaho, Selim Solmaz, Georg Stettinger, Sergio E. Diaz, Mauricio Marcano, Jorge Villagra, Juan Medina, Martina Schwarz, Antonio Artuñedo, Mauro Comi, Rutger Beekelaar, Onur Özçelik, Elif Aksu Tasdelen, Yesim Gürbüz, Jan Saijets, Jukka Kyynäräinen, Dmitry Morits, Björn Debaillie, Maxim Rykunov, Joan Escamilla, Jarno Vanne, Tomi Korhonen, Kalle Holma, Eva-Maria Matzhold, Carlo Novara, Fabio Tango, Paolo Burgio, Giuseppe Carlo Calafiore, Milad Karimshoushtari, Emilie Boulay, Miguel Dhaens, Kylian Praet, Han Zwijnenberg, Henri Palm, David Aledo Ortega, Ercan Kalali, Tuomas Pensala, Arto Kyytinen, Morten Larsen, Omar Veledar, Georg Macher, Michael Lafer, Lorenzo Giraudi, Jakob Reckenzaun, Daniel Hammer, Naveen Mohan, Josef Schmid, Alfred Höß, Shai Ophir, Anand Dubey, Jonas Fuchs, Maximilian Lübke, Andrei Anghel, Nicolae-Catalin Ristea, Martin Törngren, Alua Musralina, Marlene Harter, Joseena Memadathil Jose, George Dimitrakopoulos 0001 |
DSD | 62 |
| 2021 | Self-Paced Ensemble Learning for Speech and Audio ClassificationabstractCombining multiple machine learning models into an ensemble is known to provide superior performance levels compared to the individual components forming the ensemble. This is because models can complement each other in taking better decisions. Instead of just combining the models, we propose a self-paced ensemble learning scheme in which models learn from each other over several iterations. During the self-paced learning process based on pseudo-labeling, in addition to improving the individual models, our ensemble also gains knowledge about the target domain. To demonstrate the generality of our self-paced ensemble learning (SPEL) scheme, we conduct experiments on three audio tasks. Our empirical results indicate that SPEL significantly outperforms the baseline ensemble models. We also show that applying self-paced learning on individual models is less effective, illustrating the idea that models in the ensemble actually learn from each other. Nicolae-Catalin Ristea, Radu Tudor Ionescu |
Interspeech | 1 |
| 2020 | Programmable Systems for Intelligence in Automobiles (PRYSTINE): Technical Progress after Year 2abstractAutonomous driving has the potential to disruptively change the automotive industry as we know it today. For this, fail-operational behavior is essential in the sense, plan, and act stages of the automation chain in order to handle safety-critical situations by its own, which currently is not reached with state-of-the-art approaches.The European ECSEL research project PRYSTINE realizes Fail-operational Urban Surround perceptION (FUSION) based on robust Radar and LiDAR sensor fusion and control functions in order to enable safe automated driving in urban and rural environments. This paper showcases some of the key results (e.g., novel Radar sensors, innovative embedded control and E/E architectures, pioneering sensor fusion approaches, AI controlled vehicle demonstrators) achieved until year 2. Norbert Druml, Björn Debaillie, Andrei Anghel, Nicolae-Catalin Ristea, Jonas Fuchs, Anand Dubey, Torsten Reissland, Maike Hartstem, Viktor Rack, Anna Ryabokon, Kaspars Ozols, Rihards Novickis, Aleksandrs Levinskis, Omar Veledar, Georg Macher, Johannes Jany-Luig, Selim Solmaz, Jakob Reckenzaun, Naveen Mohan, Shai Ophir, Georg Stettinger, Sergio E. Diaz, Mauricio Marcano, Jorge Villagra, Andrea Castellano, Rutger Beekelaar, Fabio Tango, Jarno Vanne, Kalle Holma, Oguz Icoglu, George Dimitrakopoulos 0001 |
DSD | 4 |
| 2020 | Are you Wearing a Mask? Improving Mask Detection from Speech Using Augmentation by Cycle-Consistent GANsabstractThe task of detecting whether a person wears a face mask from speech is useful in modelling speech in forensic investigations, communication between surgeons or people protecting themselves against infectious diseases such as COVID-19. In this paper, we propose a novel data augmentation approach for mask detection from speech. Our approach is based on (i) training Generative Adversarial Networks (GANs) with cycle-consistency loss to translate unpaired utterances between two classes (with mask and without mask), and on (ii) generating new training utterances using the cycle-consistent GANs, assigning opposite labels to each translated utterance. Original and translated utterances are converted into spectrograms which are provided as input to a set of ResNet neural networks with various depths. The networks are combined into an ensemble through a Support Vector Machines (SVM) classifier. With this system, we participated in the Mask Sub-Challenge (MSC) of the INTERSPEECH 2020 Computational Paralinguistics Challenge, surpassing the baseline proposed by the organizers by 2.8%. Our data augmentation technique provided a performance boost of 0.9% on the private test set. Furthermore, we show that our data augmentation approach yields better results than other baseline and state-of-the-art augmentation methods. Nicolae-Catalin Ristea, Radu Tudor Ionescu |
INTERSPEECH | 1 |
| 2020 | Fully Convolutional Neural Networks for Automotive Radar Interference MitigationabstractThe interest of the automotive industry has progressively focused on subjects related to driver assistance systems as well as autonomous cars. Cars combine a variety of sensors to perceive their surroundings robustly. Among them, radar sensors are indispensable because of their independence of lighting conditions and the possibility to directly measure velocity. However, radar interference is an issue that becomes prevalent with the increasing amount of radar systems in automotive scenarios. In this paper, we address this issue for frequency modulated continuous wave (FMCW) radars with fully convolutional neural networks (FCNs), a state-of-the-art deep learning technique. We propose two FCNs that take spectrograms of the beat signals as input, and provide the corresponding clean range profiles as output. We propose two architectures for interference mitigation which outperform the classical zeroing technique. Moreover, considering the lack of databases for this task, we release as open source a large scale data set that closely replicates real world automotive scenarios for single-interference cases, allowing others to objectively compare their future work in this domain. The data set is available for download at: http://github.com/ristea/arim. Nicolae-Catalin Ristea, Andrei Anghel, Radu Tudor Ionescu |
VTC Fall | 1 |