Jakub Tkaczuk

dblp:119/7322 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
2since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Efficient and distributed learning · 67% Probabilistic and Bayesian machine learning · 33%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Embedded and real-time systems · 100%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
inference acceleration
0.812024
SOI: Scaling Down Computational Complexity by Estimating Partial States of the Model · NeurIPS 2024
Machine learning › Efficient and distributed learning
model compression
0.812024
SOI: Scaling Down Computational Complexity by Estimating Partial States of the Model · NeurIPS 2024
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference
online inference
0.812024
SOI: Scaling Down Computational Complexity by Estimating Partial States of the Model · NeurIPS 2024
Embedded and real-time systems › embedded processor
microcontroller
0.212024
SOI: Scaling Down Computational Complexity by Estimating Partial States of the Model · NeurIPS 2024

Methods — techniques the papers use, named apart from their topics

scattered online inference · 1.5partial state estimation · 1.5
YearPublicationVenuePosition
2024 SOI: Scaling Down Computational Complexity by Estimating Partial States of the Model
abstract
Consumer electronics used to follow the miniaturization trend described by Moore’s Law. Despite increased processing power in Microcontroller Units (MCUs), MCUs used in the smallest appliances are still not capable of running even moderately big, state-of-the-art artificial neural networks (ANNs) especially in time-sensitive scenarios. In this work, we present a novel method called Scattered Online Inference (SOI) that aims to reduce the computational complexity of ANNs. SOI leverages the continuity and seasonality of time-series data and model predictions, enabling extrapolation for processing speed improvements, particularly in deeper layers. By applying compression, SOI generates more general inner partial states of ANN, allowing skipping full model recalculation at each inference.
Grzegorz Stefanski, Pawel Daniluk, Artur Szumaczuk, Jakub Tkaczuk
NeurIPS4
2022 As We Speak: Real-Time Visually Guided Speaker Separation and Localization
abstract
Real-time speaker separation and localization is crucial to enable applications for video call enhancement, automatic subtitles localization, as well as spatial voice generation/panning. The common approach to perform speaker localization and separation is to detect candidate faces and then perform visual guided voice separation for each. There are two methods used for face detection: with face detector on static video frames [1], [2] or with audio visual sequence processing for active speaker detection [3]. In this work, we propose improvements for the visual guided speaker separation model to make it real-time. The described model follows the approach with a face detector. The model extends real-time models known for speech enhancement [4], [5] by adding face processing to ultimately perform visual guided speaker separation. Our system is lightweight with 0.6M trainable parameters. It performs speaker separation near instantaneously with the delay of a single input audio frame. To our knowledge, it is the first real-time system for visual guided speaker separation. From the application point of view it is important that the model performs both tasks at the time: speech separation and active speaker localization.
Piotr Czarnecki, Jakub Tkaczuk
MMSP2