Ilya Kiselev

dblp:184/4258 · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
3since 2021 · last 2025
0000-0002-5055-4314ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 4 first-author · 3 since 2021Artificial intelligence and machine learning · 2Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2025 Online Prediction of Core Body Temperature from Sweat Wearable Printed Sensors Using Recurrent Neural Network
abstract
Real-time monitoring of core body temperature (CBT) is important for preventing heat-related physiological problems during work or exercise. This paper shows how sweat biomarkers (sodium and potassium concentrations) as measured by printed sensors on a sweat wearable patch can be used for online continual prediction of CBT. These sensor measurements and additional biomarker data (heart rate and regional sweat rate) were collected from two healthy male athletes during controlled cycling sessions. The biomarker data was used to continuously predict CBT with one of three models: a recurrent neural network (RNN), a multilayer perceptron, and a linear regression model. The results show that with a window size of 30 seconds, sweat sodium and potassium concentrations outperform other biomarker pairs in predicting CBT. Out of all three models, an RNN model with only 70.8 K parameters achieved the lowest prediction error of 0.04 °C using the two sweat biomarkers. These findings support the use of sweat-based non-invasive monitoring systems for reliable online CBT prediction.
Silvia Demuru, Céline Lafaye, Brince Paul Kunnel, Cyril Besson, Ilya Kiselev, Vincent Gremeaux, Mathieu Saubade, Danick Briand, Shih-Chii Liu
ISCAS7
2022 Spiking Cochlea With System-Level Local Automatic Gain Control
abstract
Including local automatic gain control (AGC) circuitry into a silicon cochlea design has been challenging because of transistor mismatch and model complexity. To address this, we present an alternative system-level algorithm that implements channel-specific AGC in a silicon spiking cochlea by measuring the output spike activity of individual channels. The bandpass filter gain of a channel is adapted dynamically to the input amplitude so that the average output spike rate stays within a defined range. Because this AGC mechanism only needs counting and adding operations, it can be implemented at low hardware cost in a future design. We evaluate the impact of the local AGC algorithm on a classification task where the input signal varies over 32dB input range. Two classifier types receiving cochlea spike features were tested on a speech versus noise classification task. The logistic regression classifier achieves an average of 6% improvement and 40.8% relative improvement in accuracy when the AGC is enabled. The deep neural network classifier shows a similar improvement for the AGC case and achieves a higher mean accuracy of 96% compared to the best accuracy of 91% from the logistic regression classifier.
Ilya Kiselev, Chang Gao 0002, Shih-Chii Liu
IEEE Trans. Circuits Syst. I Regul. Pap.1
2021 Event-Driven Local Gain Control on a Spiking Cochlea Sensor
abstract
Including local automatic gain control (AGC) circuitry into a silicon cochlea design can be challenging because of transistor mismatch and model complexity. To address this, we present an alternative system-level algorithm that implements channel-specific AGC by using the output spikes of a spiking silicon cochlea. By measuring the output spike activity of each channel, the bandpass filter gain of a channel is adapted dynamically to the input sound amplitude so that the average output spike rate stays within a defined range. We evaluate the effect of our local AGC algorithm on a classification task where the input signal varies over a large amplitude range. Results on a task to classify speech versus noise show that a classifier trained on spike responses of a cochlea with local AGC maintains an average of 25% higher accuracy over a 32 dB input dynamic range, compared to the case when the AGC is disabled.
Ilya Kiselev, Shih-Chii Liu
ISCAS1
2020 Evaluating Multi-Channel Multi-Device Speech Separation Algorithms in the Wild: A Hardware-Software Solution
abstract
Evaluation methods for multi-channel speech separation algorithms in the real world are becoming increasingly important as the number of applications involving audio assistants and hearing aid devices continues to grow. To make such evaluations easier, this paper presents a multi-microphone hardware platform, WHISPER, built specifically for this purpose and its subsequent use for evaluating speech processing algorithms. The platform can also be constructed as an ad-hoc wireless acoustic sensor network (WASN) with high synchronization precision. Using WHISPER, we describe real-world experiments where an example speech separation algorithm is applied to mixtures of varying number of talkers and signal-to-noise ratios. The results when compared with those from a simulated environment, show the usefulness of WASNs and that simulations tend to underestimate the difficulty of speech separation in real-world scenarios. This work represents an important step towards developing a hardware-software framework for evaluating speech processing algorithms in the wild.
Enea Ceolini, Ilya Kiselev, Shih-Chii Liu
IEEE ACM Trans. Audio Speech Lang. Process.2
2019 Live Demonstration: Real-Time Spoken Digit Recognition using the DeltaRNN Accelerator
abstract
This demonstration shows a real-time continuous speech recognition hardware system using our previously published DeltaRNN accelerator that enables low latency recurrent neural network (RNN) computation. The network is trained on augmented audio samples from the TIDIGITS dataset to achieve a label error rate (LER) of 2.31%. It is implemented on a Xilinx Zynq-7100 FPGA running at 1 MHz. The incremental RNN power consumption is 30 mW. Visitors interact with the system by speaking digits into a microphone connected to the FPGA system and the classification outputs of the network are continuously displayed on a laptop screen in real time.
Chang Gao 0002, Stefan Braun 0005, Ilya Kiselev, Jithendar Anumula, Tobi Delbruck, Shih-Chii Liu
ISCAS3
2019 Real-Time Speech Recognition for IoT Purpose using a Delta Recurrent Neural Network Accelerator
abstract
This paper describes a continuous speech recognition hardware system that uses a delta recurrent neural network accelerator (DeltaRNN) implemented on a Xilinx Zynq-7100 FPGA to enable low latency recurrent neural network (RNN) computation. The implemented network consists of a single-layer RNN with 256 gated recurrent unit (GRU) neurons and is driven by input features generated either from the output of a filter bank running on the ARM core of the FPGA in a PmodMic3 microphone setup or from the asynchronous outputs of a spiking silicon cochlea circuit. The microphone setup achieves 7.1 ms minimum latency and 177 frames-per-second (FPS) maximum throughput while the cochlea setup achieves 2.9 ms minimum latency and 345 FPS maximum throughput. The low latency and 70 mW power consumption of the DeltaRNN makes it suitable as an IoT computing platform.
Chang Gao 0002, Stefan Braun 0005, Ilya Kiselev, Jithendar Anumula, Tobi Delbruck, Shih-Chii Liu
ISCAS3
2018 Speaker Activity Detection and Minimum Variance Beamforming for Source Separation
Enea Ceolini, Jithendar Anumula, Adrian E. G. Huber, Ilya Kiselev, Shih-Chii Liu
INTERSPEECH4
2016 Live demonstration: Event-driven deep neural network hardware system for sensor fusion
abstract
We demonstrate an interactive digit recognition system using a spiking Deep Neural Network (DNN) FPGA-based system connected to two event-driven sensors: a Dynamic Vision Sensor (DVS) and a Dynamic Audio Sensor (DAS). Sensor fusion is demonstrated on a digit classification task using a DNN trained on the MNIST dataset supplemented by assignment of a unique pure ton e for each digit.
Ilya Kiselev, Daniel Neil, Shih-Chii Liu
ISCAS1
2016 Event-driven deep neural network hardware system for sensor fusion
abstract
This paper presents a real-time multi-modal spiking Deep Neural Network (DNN) implemented on an FPGA platform. The hardware DNN system, called n-Minitaur, demonstrates a 4-fold improvement in computational speed over the previous DNN FPGA system. The proposed system directly interfaces two different event-based sensors: a Dynamic Vision Sensor (DVS) and a Dynamic Audio Sensor (DAS). The DNN for this bimodal hardware system is trained on the MNIST digit dataset and a set of unique audio tones for each digit. When tested on the spikes produced by each sensor alone, the classification accuracy is around 70% for DVS spikes generated in response to displayed MNIST images, and 60% for DAS spikes generated in response to noisy tones. The accuracy increases to 98% when spikes from both modalities are provided simultaneously. In addition, the system shows a fast latency response of only 5ms.
Ilya Kiselev, Daniel Neil, Shih-Chii Liu
ISCAS1