VLDB 2026 Research / reviewers in the wild / expert
Jörg Bitzer
dblp:02/7155
· DBLP profile ↗
12ranked-venue papers
2as first author
6since 2021 · last 2026
0000-0002-0595-0262ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 5 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Towards privacy-preserving conversation analysis in everyday life: Exploring the privacy-utility trade-offabstractRecordings in everyday life provide valuable insights for health-related applications, such as analyzing conversational behavior as an indicator of social interaction and well-being. However, these recordings require privacy preservation of both the speech content and the speaker’s identity of all persons involved. This article investigates privacy-preserving features feasible for power-constrained recording devices by combining smoothing and subsampling in the frequency and time domain with a low-cost speaker anonymization technique. A speech recognition and a speaker verification system are used to evaluate privacy protection, whereas a voice activity detection and a speaker diarization model are used to assess the utility for analyzing conversations. The evaluation results demonstrate that combining speaker anonymization with the aforementioned smoothing and subsampling protects speech privacy, albeit at the expense of utility performance. Overall, our privacy-preserving methods offer various trade-offs between privacy and utility, reflecting the requirements of different application scenarios. Jule Pohlhausen, Francesco Nespoli, Jörg Bitzer |
Comput. Speech Lang. | 3 |
| 2025 | Influence of Room Acoustics on Objective Voice Assessment Methods in the Context of Speech and Language Therapy
Sven Franz, Tanja Grewe, Bernd T. Meyer, Jörg Bitzer |
INTERSPEECH | 4 |
| 2024 | Microphone Subset Selection for the Weighted Prediction Error Algorithm Using a Group Sparsity PenaltyabstractReverberation can severely degrade the quality of speech signals recorded using microphones in an enclosure. In acoustic sensor networks with spatially distributed microphones, a similar dereverberation performance may be achieved using only a subset of all available microphones. Using the popular convex relaxation method, in this paper we propose to perform microphone subset selection for the weighted prediction error (WPE) multi-channel dereverberation algorithm by introducing a group sparsity penalty on the prediction filter coefficients. The resulting problem is shown to be solved efficiently using the accelerated proximal gradient algorithm. Experimental evaluation using measured impulse responses shows that the performance of the proposed method is close to the optimal performance obtained by exhaustive search, both for frequency-dependent as well as frequency-independent microphone subset selection. Furthermore, the performance using only a few microphones for frequency-independent microphone subset selection is only marginally worse than using all available microphones. Anselm Lohmann, Toon van Waterschoot, Jörg Bitzer, Simon Doclo |
ICASSP | 3 |
| 2023 | Geometry-Aware DOA Estimation Using a Deep Neural Network with Mixed-Data Input FeaturesabstractUnlike model-based direction of arrival (DoA) estimation algorithms, supervised learning-based DoA estimation algorithms based on deep neural networks (DNNs) are usually trained for one specific microphone array geometry, resulting in poor performance when applied to a different array geometry. In this paper we illustrate the fundamental difference between supervised learning-based and model-based algorithms leading to this sensitivity. Aiming at designing a supervised learning-based DoA estimation algorithm that generalizes well to different array geometries, in this paper we propose a geometry-aware DoA estimation algorithm. The algorithm uses a fully connected DNN and takes mixed data as input features, namely the time lags maximizing the generalized cross-correlation with phase transform and the microphone coordinates, which are assumed to be known. Experimental results in a reverberant scenario demonstrate the flexibility of the proposed algorithm towards different array geometries and show that the proposed algorithm outperforms model-based algorithms such as steered response power with phase transform. Ulrik Kowalk, Simon Doclo, Jörg Bitzer |
ICASSP | 3 |
| 2023 | Dereverberation in Acoustic Sensor Networks Using weighted Prediction Error with Microphone-Dependent Prediction DelaysabstractIn the last decades several multi-microphone speech dereverberation algorithms have been proposed, among which the weighted prediction error (WPE) algorithm. In the WPE algorithm, a prediction delay is required to reduce the correlation between the prediction signals and the direct component in the reference microphone signal. In compact arrays with closely-spaced microphones, the prediction delay is often chosen microphone-independent. In acoustic sensor networks with spatially distributed microphones, large time-differences-of-arrival (TDOAs) of the speech source between the reference microphone and other microphones typically occur. Hence, when using a microphone-independent prediction delay1the reference and prediction signals may still be significantly correlated, leading to distortion in the dereverberated output signal. In order to decorrelate the signals, in this paper we propose to apply TDOA compensation with respect to the reference microphone, resulting in microphone-dependent prediction delays for the WPE algorithm. We consider both optimal TDOA compensation using crossband filtering in the short-time Fourier transform domain as well as band-to-band and integer delay approximations. Simulation results for different reverberation times using oracle as well as estimated TDOAs clearly show the benefit of using microphone-dependent prediction delays. Anselm Lohmann, Toon van Waterschoot, Jörg Bitzer, Simon Doclo |
ICASSP | 3 |
| 2023 | Two-Stage Voice Anonymization for Enhanced Privacy
Francesco Nespoli, Daniel Barreda, Jörg Bitzer, Patrick A. Naylor |
INTERSPEECH | 3 |
| 2020 | Adaptive Compressive Onset-Enhancement for Improved Speech Intelligibility in Noise and ReverberationabstractNear-end listening enhancement (NELE) algorithms aim to pre-process speech prior to playback via loudspeakers so as to maintain high speech intelligibility even when listening conditions are not optimal, e.g., due to noise or reverberation. Often NELE algorithms are designed for scenarios considering either only the detrimental effect of noise or only reverberation, but not both disturbances. In many typical applications scenarios, however, both factors are present. In this paper, we evaluate a new combination of a noise-dependent and a reverberation-dependent algorithm implemented in a common framework. Specifically, we use instrumental measures as well as subjective ratings of listening effort for acoustic scenarios with different reverberation times and realistic signal-to-noise ratios. The results show that the noise-dependent algorithm also performs well in reverberation, and that the combination of both algorithms can yield slightly better performance than the individual algorithms alone. This benefit appears to depend strongly on the specific acoustic condition, indicating that further work is required to optimize the adaptive algorithm behavior. Felicitas Bederna, Henning F. Schepker, Christian Rollwage, Simon Doclo, Arne Pusch, Jörg Bitzer, Jan Rennies |
INTERSPEECH | 6 |
| 2015 | Reduction of Gaussian, Supergaussian, and Impulsive Noise by Interpolation of the Binary Mask ResidualabstractIn this paper, we present a new approach for noise reduction. A binary time-frequency (T-F) masking threshold criterion is proposed and analyzed with respect to the average spectra of music and noise disturbances. Modified autoregressive (AR) detection and AR interpolation are then applied to the residual signal of the binary masking process. The proposed method is able to reduce supergaussian and impulsive noise while ensuring preservation of the desired signal, which is crucial for professional high-quality audio restoration, and it is also suitable for Gaussian noise to a certain extent. The approach is compared to a state-of-the-art restoration algorithm by means of the objective measures signal-to-noise ratio (SNR) improvement and perceptual quality, and by subjective listening tests. The objective results as well as the listening tests show that the proposed algorithm is especially suited for supergaussian, grainy-sounding noise types, e.g., optical soundtrack noise of celluloid movie footage, or rain noise. Marco Ruhland, Jörg Bitzer, Matthias Brandt, Stefan Goetze |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2003 | Multi-microphone residual echo estimationabstractPost-filters are a powerful extension to improve echo attenuation when combined with the well-known echo canceller. In order to guarantee high quality of the transmitted speech signal, the primary purpose of a post-filtering system is to estimate the power spectral density (PSD) of the residual echo at the output of the echo canceller as accurately as possible. We introduce a novel technique to estimate the residual echo by using a microphone array. The robustness against double-talk and other additive interferences is reached by means of minimum statistics and further enhanced by exploiting spatial information. Simulation results show that the new methods are able to estimate the residual echo even under adverse conditions. Markus Kallinger, Karl-Dirk Kammeyer, Jörg Bitzer |
ICASSP (5) | 3 |
| 2001 | Multi-microphone noise reduction techniques as front-end devices for speech recognition
Jörg Bitzer, Klaus Uwe Simmer, Karl-Dirk Kammeyer |
Speech Commun. | 1 |
| 2000 | Study on combining multi-channel echo cancellers with beamformersabstractIn this paper, we compare different combinations of a multi-channel non-adaptive noise reduction unit (NRU) and an acoustic echo cancellation unit (AEC) for a standard single-channel voice transmission. The results show that the NRU and the AEC-unit can be interchanged without increasing the computational complexity of the combined system, The length of an AEC's adaptive filter can be greatly reduced, if a multi-channel AEC-unit in front of a multichannel NRU is used. Furthermore, the control information, like the step-size of the NLMS-algorithm, has to be computed only once. Markus Kallinger, Jörg Bitzer, Karl-Dirk Kammeyer |
ICASSP | 2 |
| 1999 | Theoretical noise reduction limits of the generalized sidelobe canceller (GSC) for speech enhancementabstractWe present an analysis of the generalized sidelobe canceller (GSC). It can be shown that the theoretical limits of the noise reduction performance depend only on the auto- and cross-spectral densities of the input signals. Furthermore, we compute the limits of the noise reduction performance for the theoretically determined diffuse noise field, which is an approximation for reverberant rooms. Our results show that the GSC cannot reduce noise further than 1 dB. These results were verified by simulation of reverberant environments. Only in sound-proof rooms with a reverberation time less than 100 ms the GSC performs well. Jörg Bitzer, Klaus Uwe Simmer, Karl-Dirk Kammeyer |
ICASSP | 1 |