Fan Yang 0019

dblp:29/3081-19 · DBLP profile ↗
← Back
24ranked-venue papers
2as first author
7since 2021 · last 2025
0000-0002-6844-2521ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 2 first-author · 5 since 2021Systems, architecture and hardware · 7 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Steering Prediction via a Multi-Sensor System for Autonomous Racing
abstract
Autonomous racing has rapidly gained research attention. Traditionally, racing cars rely on 2D LiDAR as their primary visual system. In this work, we explore the integration of an event camera with the existing system to provide enhanced temporal information. Our goal is to fuse the 2D LiDAR data with event data in an end-to-end learning framework for steering prediction, which is crucial for autonomous racing. To the best of our knowledge, this is the first study addressing this challenging research topic. We start by creating a multisensor dataset specifically for steering prediction. Using this dataset, we establish a benchmark by evaluating various SOTA fusion methods. Our observations reveal that existing methods often incur substantial computational costs. To address this, we apply low-rank techniques to propose a novel, efficient, and effective fusion design. We introduce a new fusion learning policy to guide the fusion process, enhancing robustness against misalignment. Our fusion architecture provides better steering prediction than LiDAR alone, significantly reducing the RMSE from 7.72 to 1.28. Compared to the second-best fusion method, our work represents only 11% of the learnable parameters while achieving better accuracy. The source code and dataset are publicly available at: https://github.com/ZZY-Zhou/F1Tenth-Steering.
Zhuyun Zhou, Zongwei Wu, Florian Bolli, Rémi Boutteau, Fan Yang 0019, Radu Timofte, Dominique Ginhac, Tobi Delbruck
ICRA5
2025 Tree-Based Personalized Clustered Federated Learning: A Driver Stress Monitoring Through Physiological Data Case Study
abstract
Recent advancements in wearable biosensor technology have significantly enhanced the precision and ease of collecting physiological data. This has proven particularly useful in monitoring stress among drivers. However, the sensitive nature of this data raises significant privacy concerns. To tackle these challenges, federated learning (FL) has emerged as a novel solution to safeguard data privacy by decentralizing model training to individual devices, eliminating the need for data sharing. This approach is compliant with strict privacy regulations like the GDPR (EU) and HIPAA (US), significantly reducing data breach risks and data transfer costs. Despite its benefits, FL struggles with nonindependent and identically distributed (non-IID) data. This issue hampers the FL model performance and adaptability by complicating convergence with a generalized model. To address this limitation, we introduce in this article an innovative tree-based personalized clustered FL (TPCFL) approach. TPCFL effectively exploits similarities in drivers’ private data characteristics to assign each driver a personalized model relying on a tree-based clustering approach. Grounded in the realm of individuals clustering in FL to address non-IID data challenges, notably recognized as clustered FL (CFL), TPCFL is augmented with a novel tree-based clustering approach and a tailored cluster selection technique, enabling it to adeptly address core CFL challenges, such as hyperparameter optimization for cluster selection and integration of new unlabeled drivers. Experiments demonstrated the superior performance of the proposed clustering method and highlighted TPCFL’s efficiency in achieving an optimal balance between personalized and generalized learning, showcasing its effectiveness on the two public data sets.
Houda Rafi, Yannick Benezeth, Fan Yang 0019, Philippe Reynaud, Emmanuel Arnoux, Cédric Demonceaux
IEEE Internet Things J.3
2024 Event-Free Moving Object Segmentation from Moving Ego Vehicle
abstract
Moving object segmentation (MOS) in dynamic scenes is an important, challenging, but under-explored research topic for autonomous driving, especially for sequences obtained from moving ego vehicles. Most segmentation methods leverage motion cues obtained from optical flow maps. However, since these methods are often based on optical flows that are pre-computed from successive RGB frames, this neglects the temporal consideration of events occurring within the inter-frame, consequently constraining its ability to discern objects exhibiting relative staticity but genuinely in motion. To address these limitations, we propose to exploit event cameras for better video understanding, which provide rich motion cues without relying on optical flow. To foster research in this area, we first introduce a novel large-scale dataset called DSEC-MOS for moving object segmentation from moving ego vehicles, which is the first of its kind. For benchmarking, we select various mainstream methods and rigorously evaluate them on our dataset. Subsequently, we devise EmoFormer, a novel network able to exploit the event data. For this purpose, we fuse the event temporal prior with spatial semantic maps to distinguish genuinely moving objects from the static background, adding another level of dense supervision around our object of interest. Our proposed network relies only on event data for training but does not require event input during inference, making it directly comparable to frame-only methods in terms of efficiency and more widely usable in many application cases. The exhaustive comparison highlights a significant performance improvement of our method over all other methods. The source code and dataset are publicly available at: https://github.com/ZZYZhou/DSEC-MOS.
Zhuyun Zhou, Zongwei Wu, Danda Pani Paudel, Rémi Boutteau, Fan Yang 0019, Luc Van Gool, Radu Timofte, Dominique Ginhac
IROS5
2023 RGB-Event Fusion for Moving Object Detection in Autonomous Driving
abstract
Moving Object Detection (MOD) is a critical vision task for successfully achieving safe autonomous driving. Despite plausible results of deep learning methods, most existing approaches are only frame-based and may fail to reach reasonable performance when dealing with dynamic traffic participants. Recent advances in sensor technologies, especially the Event camera, can naturally complement the conventional camera approach to better model moving objects. However, event-based works often adopt a pre-defined time window for event representation, and simply integrate it to estimate image intensities from events, neglecting much of the rich temporal information from the available asynchronous events. Therefore, from a new perspective, we propose RENet, a novel RGB-Event fusion Network, that jointly exploits the two complementary modalities to achieve more robust MOD under challenging scenarios for autonomous driving. Specifically, we first design a temporal multi-scale aggregation module to fully leverage event frames from both the RGB exposure time and larger intervals. Then we introduce a bi-directional fusion module to attentively calibrate and fuse multi-modal features. To evaluate the performance of our network, we carefully select and annotate a sub-MOD dataset from the commonly used DSEC dataset. Extensive experiments demonstrate that our proposed method performs significantly better than the state-of-the-art RGB-Event fusion alternatives. The source code and dataset are publicly available at: https://github.com/ZZY-Zhou/RENet.
Zhuyun Zhou, Zongwei Wu, Rémi Boutteau, Fan Yang 0019, Cédric Demonceaux, Dominique Ginhac
ICRA4
2023 Accumulated micro-motion representations for lightweight online action detection in real-time
Yu Liu 0060, Fan Yang 0019, Dominique Ginhac
J. Vis. Commun. Image Represent.2
2023 UBFC-Phys: A Multimodal Database For Psychophysiological Studies of Social Stress
abstract
As humans, we experience social stress in countless everyday-life situations. Giving a speech in front of an audience, passing a job interview, and similar experiences all lead us to go through stress states that impact both our psychological and physiological states. Therefore, studying the link between stress and physiological responses had become a critical societal issue, and recently, research in this field has grown in popularity. However, publicly available datasets have limitations. In this article, we propose a new dataset, UBFC-Phys, collected with and without contact from participants living social stress situations. A wristband was used to measure contact blood volume pulse (BVP) and electrodermal activity (EDA) signals. Video recordings allowed to compute remote pulse signals, using remote photoplethysmography (RPPG), and facial expression features. Pulse rate variability (PRV) was extracted from BVP and RPPG signals. Our dataset permits to evaluate the possibility of using video-based physiological measures compared to more conventional contact-based modalities. The goal of this article is to present both the dataset, which we make publicly available, and experimental results of contact and non-contact data comparison, as well as stress recognition. We obtained a stress state recognition accuracy of 85.48 percent, achieved by remote PRV features.
Rita Meziati Sabour, Yannick Benezeth, Pierre De Oliveira, Julien Chappé, Fan Yang 0019
IEEE Trans. Affect. Comput.5
2021 ACDnet: An action detection network for real-time edge computing based on flow-guided feature approximation and memory aggregation
Yu Liu 0060, Fan Yang 0019, Dominique Ginhac
Pattern Recognit. Lett.2
2019 Automatic Assessment of Depression Based on Visual Cues: A Systematic Review
abstract
Automatic depression assessment based on visual cues is a rapidly growing research domain. The present exhaustive review of existing approaches as reported in over sixty publications during the last ten years focuses on image processing and machine learning algorithms. Visual manifestations of depression, various procedures used for data collection, and existing datasets are summarized. The review outlines methods and algorithms for visual feature extraction, dimensionality reduction, decision methods for classification and regression approaches, as well as different fusion strategies. A quantitative meta-analysis of reported results, relying on performance metrics robust to chance, is included, identifying general trends and key unresolved issues to be considered in future studies of automatic depression assessment utilizing visual cues alone or in combination with vocal or verbal cues.
Anastasia Pampouchidou, Panagiotis G. Simos, Kostas Marias, Fabrice Mériaudeau, Fan Yang 0019, Matthew Pediaditis, Manolis Tsiknakis
IEEE Trans. Affect. Comput.5
2018 Detection of H. pylori Induced Gastric Inflammation by Diffuse Reflectance Analysis
abstract
Spectral acquisitions contain rich information and thus, are promising modalities for early detection of gastric diseases. In this study, we analyze the diffuse reflectance of the gastric inflammatory lesions induced by the bacterium H. pylori in the mouse stomach. A pipeline has been designed to characterize and classify spectra acquired on mice. The pipeline is based on a band clustering algorithm followed by the computation of meaningful division and subtraction features and by classification with a linear SVM classifier. Currently, the pipeline is able to recognize inflamed stomach's spectra with an accuracy of 98%. These results are promising and the same pipeline could be adapted for the study of gastric pathologies in humans.
Alexandre Krebs, Vania Camilo, Eliette Touati, Yannick Benezeth, Valerie Michel, Gregory Jouvion, Fan Yang 0019, Dominique Lamarque, Franck Marzani
BIBE7
2018 Comparison of Region of Interest Segmentation Methods for Video-Based Heart Rate Measurements
abstract
Conventional contact photoplethysmography (PPG) sensors are not suitable in situations of skin damage or when unconstrained movement is required. As a consequence, remote photoplethysmography (rPPG) has recently emerged because it provides remote physiological measurements without expensive hardware and improves comfort for long term monitoring. RPPG estimation methods use the spatially averaged RGB values of pixels in a Region Of Interest (ROI) to generate a temporal RGB signal. The selection of ROI is a critical first step to obtain reliable pulse signals and must contain as many skin pixels as possible with a low percentage of non-skin pixels. In this paper, we experimentally compare seven ROI segmentation methods in the perspective of heart rate (HR) measurements with dedicated metrics. The algorithms are compared using our in-house database UBFC-RPPG, comprising of 53 videos specifically geared towards rPPG analysis.
Peixi Li, Yannick Benezeth, Keisuke Nakamura, Randy Gomez, Chao Li 0005, Fan Yang 0019
BIBE6
2018 A robust multispectral palmprint matching algorithm and its evaluation for FPGA applications
Chao Li 0005, Yannick Benezeth, Keisuke Nakamura, Randy Gomez, Fan Yang 0019
J. Syst. Archit.5
2016 High-dynamic range image generation from single low-dynamic range image
abstract
Due to the growing popularity of high‐dynamic range (HDR) image and the high complexity to capture HDR image, researchers focus on converting low‐dynamic range (LDR) content to HDR, which gives rise to a number of dynamic range expansion methods. Most of the existing methods try their best to tackle highlight areas during the expanding, however, in some cases, they cannot achieve approving results. In this study, a novel LDR image expansion technique is presented. The technique first detects the highlight areas in image; then preprocesses them and reconstructs the information of these regions; finally, expands the LDR image to HDR. Unlike the existing schemes, the proposed approach escapes the complicated treatment to highlight areas in the process of expansion, which makes the expansion straightforward; at the same time, it facilitates the expansion scheme and minimises the formation of the artefacts. The experimental results show that the proposed method performs well; the tone mapped versions of the produced HDR images are popular. The results of the image quality metric also illustrate that the novel approach can recover more image details with minimised contrast loss and reversal, compared with the existing schemes considered in the comparison.
Yongqing Huo, Fan Yang 0019
IET Image Process.2
2016 Shape-constrained level set segmentation for hybrid CPU-GPU computers
Souleymane Balla-Arabé, Xinbo Gao 0001, Dominique Ginhac, Fan Yang 0019
Neurocomputing4
2016 Embedded multi-spectral image processing for real-time medical application
Chao Li 0005, Souleymane Balla-Arabé, Fan Yang 0019
J. Syst. Archit.3
2016 Architecture-Driven Level Set Optimization: From Clustering to Subpixel Image Segmentation
abstract
Thanks to their effectiveness, active contour models (ACMs) are of great interest for computer vision scientists. The level set methods (LSMs) refer to the class of geometric active contours. Comparing with the other ACMs, in addition to subpixel accuracy, it has the intrinsic ability to automatically handle topological changes. Nevertheless, the LSMs are computationally expensive. A solution for their time consumption problem can be hardware acceleration using some massively parallel devices such as graphics processing units (GPUs). But the question is: which accuracy can we reach while still maintaining an adequate algorithm to massively parallel architecture? In this paper, we attempt to push back the compromise between, speed and accuracy, efficiency and effectiveness, to a higher level, comparing with state-of-the-art methods. To this end, we designed a novel architecture-aware hybrid central processing unit (CPU)-GPU LSM for image segmentation. The initialization step, using the well-known k -means algorithm, is fast although executed on a CPU, while the evolution equation of the active contour is inherently local and therefore suitable for GPU-based acceleration. The incorporation of local statistics in the level set evolution allowed our model to detect new boundaries which are not extracted by the used clustering algorithm. Comparing with some cutting-edge LSMs, the introduced model is faster, more accurate, less subject to giving local minima, and therefore suitable for automatic systems. Furthermore, it allows two-phase clustering algorithms to benefit from the numerous LSM advantages such as the ability to achieve robust and subpixel accurate segmentation results with smooth and closed contours. Intensive experiments demonstrate, objectively and subjectively, the good performance of the introduced framework both in terms of speed and accuracy.
Souleymane Balla-Arabé, Xinbo Gao 0001, Dominique Ginhac, Vincent Brost, Fan Yang 0019
IEEE Trans. Cybern.5
2014 Image boundaries detection: from thresholding to implicit curve evolution
abstract
The development of high dimensional large-scale imaging devices increases the need of fast, robust and accurate image segmentation methods. Due to its intrinsic advantages such as the ability to extract complex boundaries, while handling topological changes automatically, the level set method (LSM) has been widely used in boundaries detection. Nevertheless, their computational complexity limits their use for real time systems. Furthermore, most of the LSMs share the limit of leading very often to a local minimum, while the effectiveness of many computer vision applications depends on the whole image boundaries. In this paper, using the image thresholding and the implicit curve evolution frameworks, we design a novel boundaries detection model which handles the above related drawbacks of the LSMs. In order to accelerate the method using the graphics processing units, we use the explicit and highly parallelizable lattice Boltzmann method to solve the level set equation. The introduced algorithm is fast and achieves global image segmentation in a spectacular manner. Experimental results on various kinds of images demonstrate the effectiveness and the efficiency of the proposed method.
Souleymane Balla-Arabé, Vincent Brost, Fan Yang 0019
ICMV3
2014 Multi-Kernel Implicit Curve Evolution for Selected Texture Region Segmentation in VHR Satellite Images
abstract
Very high resolution (VHR) satellite images provide a mass of detailed information which can be used for urban planning, mapping, security issues, or environmental monitoring. Nevertheless, the processing of this kind of image is timeconsuming, and extracting the needed information from among the huge quantity of data is a real challenge. For some applications such as natural disaster prevention and monitoring (typhoon, flood, bushfire, etc.), the use of fast and effective processing methods is demanded. Furthermore, such methods should be selective in order to extract only the information required to allow an efficient interpretation. For this purpose, we propose a texture region segmentation method using the level set algorithm and the multi-kernel theory. We design a selective and local multi-kernel stop function for which the regularization term depends on the fuzzy membership degree of a given pixel to be on the boundary or not. Favored by its local nature, the method is accelerated by means of an NVIDIA graphics processing unit programming. The new algorithm is selective, effective, and fast. Experimental results on VHR satellite images demonstrate subjectively and objectively the effectiveness of the proposed method.
Souleymane Balla-Arabé, Xinbo Gao 0001, Bin Wang 0027, Fan Yang 0019, Vincent Brost
IEEE Trans. Geosci. Remote. Sens.4
2014 Physiological inverse tone mapping based on retina response
Yongqing Huo, Fan Yang 0019, Vincent Brost
Vis. Comput.2
2009 Fast and Robust Face Detection on a Parallel Optimized Architecture Implemented on FPGA
abstract
In this paper, we present a parallel architecture for fast and robust face detection implemented on FPGA hardware. We propose the first implementation that meets both real-time requirements in an embedded context and face detection robustness within complex backgrounds. The chosen face detection method is the Convolutional Face Finder (CFF) algorithm, which consists of a pipeline of convolution and subsampling operations, followed by a multilayer perceptron. We present the design methodology of our face detection processor element (PE). This methodology was followed in order to optimize our implementation in terms of memory usage and parallelization efficiency. We then built a parallel architecture composed of a PE ring and an FIFO memory, resulting in a scalable system capable of processing images of different sizes. A ring of 25 PEs running at 80 MHz is able to process 127 QVGA images per second and performing real-time face detection on VGA images (35 images per second).
Nicolas Farrugia, Franck Mamalet, Sébastien Roux, Fan Yang 0019, Michel Paindavoine
IEEE Trans. Circuits Syst. Video Technol.4
2007 A modular VLIW Processor
abstract
Recent FPGA chips, with their large capacity memory and reconfigurability potential, have opened new ways for rapid prototyping of embedded systems. With the advent of high density FPGAs it is now feasible to implement a high-performance VLIW processor core in an FPGA. In this paper, we describe research results of enabling the DSP TMS320 C6201 model for real-time image processing application, by exploiting FPGA technology. The goals are, firstly, to keep the flexibility of DSP in order to shorten the development cycle, and secondly to use powerful available resources on FPGA to increase real-time performance. We present a modular DSP C6201 VHDL model which contains the hardware just necessary for each target application. The application development cycle proposed is validated with some common algorithms of image. Our results demonstrate that an algorithm can easily be, in an optimal manner, specified and then automatically converted to VHDL language and implemented on an FPGA device with system level software.
Vincent Brost, Fan Yang 0019, Michel Paindavoine
ISCAS2
2007 A Parallel Face Detection System Implemented on FPGA
abstract
In this paper, we introduce a methodology for designing a system for face detection and its implementation on FPGA. The chosen face detection method is the well-known convolutional face finder (CFF) algorithm, which consists in a pipeline of convolutions and subsampling operations. Our goal is to define a parallel architecture able to process efficiently this algorithm. We present a dataflow based architecture algorithm adequation (AAA) methodology implemented using the SynDEx software, in order to find the best compromise between the processing power and functionality requirement of each processor element (PE), and the efficiency of algorithm parallelization. We describe a first implementation of a PE on a Virtex 4 FPGA using the DSP48 dedicated blocks. This PE is able to run at a maximum frequency of 352 MHz and occupies only 2% of a Virtex 4 SX35 device.
Nicolas Farrugia, Franck Mamalet, Sébastien Roux, Fan Yang 0019, Michel Paindavoine
ISCAS4
2003 Implementation of an RBF neural network on embedded systems: real-time face tracking and identity verification
abstract
This paper describes a real time vision system that allows us to localize faces in video sequences and verify their identity. These processes are image processing techniques based on the radial basis function (RBF) neural network approach. The robustness of this system has been evaluated quantitatively on eight video sequences. We have adapted our model for an application of face recognition using the Olivetti Research Laboratory (ORL), Cambridge, UK, database so as to compare the performance against other systems. We also describe three hardware implementations of our model on embedded systems based on the field programmable gate array (FPGA), zero instruction set computer (ZISC) chips, and digital signal processor (DSP) TMS320C62, respectively. We analyze the algorithm complexity and present results of hardware implementations in terms of the resources used and processing speed. The success rates of face tracking and identity verification are 92% (FPGA), 85% (ZISC), and 98.2% (DSP), respectively. For the three embedded systems, the processing speeds for images size of 288 /spl times/ 352 are 14 images/s, 25 images/s, and 4.8 images/s, respectively.
Fan Yang 0019, Michel Paindavoine
IEEE Trans. Neural Networks1
1998 Face recognition: pre-processing techniques for linear autoassociators
E. Drege, Fan Yang 0019, Michel Paindavoine, Hervé Abdi
ESANN2
1997 A Pre-Processing Technique Based on the Wavelet Transform for Linear Autoassociators with Applications to Face Recognition
Fan Yang 0019, Michel Paindavoine, Hervé Abdi
ICANN1