EDBT 2026 Demo / reviewers in the wild / expert
Dejan Markovic
dblp:17/1368
· DBLP profile ↗
43ranked-venue papers
10as first author
15since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 17 · 3 first-author · 12 since 2021Artificial intelligence and machine learning · 13 · 5 first-author · 9 since 2021Systems, architecture and hardware · 10 · 1 first-authorComputer networks · 7 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 2 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SoundVista: Novel-View Ambient Sound Synthesis via Visual-Acoustic BindingabstractWe introduce SoundVista, a method to generate the ambient sound of an arbitrary scene at novel viewpoints. Given a pre-acquired recording of the scene from sparsely distributed microphones, SoundVista can synthesize the sound of that scene from an unseen target viewpoint. The method learns the underlying acoustic transfer function that relates the signals acquired at the distributed microphones to the signal at the target viewpoint, using a limited number of known recordings. Unlike existing works, our method does not require constraints or prior knowledge of sound source details. Moreover, our method efficiently adapts to diverse room layouts, reference microphone configurations and unseen environments. To enable this, we introduce a visual-acoustic binding module that learns visual embeddings linked with local acoustic properties from panoramic RGB and depth data. We first leverage these embeddings to optimize the placement of reference microphones in any given scene. During synthesis, we leverage multiple embeddings extracted from reference locations to get adaptive weights for their contribution, conditioned on target viewpoint. We benchmark the task on both publicly available data and real-world settings. We demonstrate significant improvements over existing methods. Mingfei Chen, Israel D. Gebru, Ishwarya Ananthabhotla, Christian Richardt, Dejan Markovic, Jake Sandakly, Steven Krenn, Todd Keebler, Eli Shlizerman, Alexander Richard |
CVPR | 5 |
| 2025 | A2B: Neural Rendering of Ambisonic Recordings to BinauralabstractThis paper introduces a novel neural network model for rendering binaural audio directly from ambisonic recordings. We optimized the model end-to-end to learn a direct mapping between ambisonic and binaural signals. Our approach eliminates traditional processing steps that were required to mitigate artifacts due to spherical harmonic order truncation and spatial aliasing, as well as other complex filtering needed to compensate for near-field sound sources. To showcase the advantage of neural network-based rendering over traditional signal processing approaches, we introduce a new dataset that includes challenging near-field sound sources, including speech and background noises. We demonstrate that our model can produce binaural audio results that closely match the fidelity of ground truth binaural recordings. Our comprehensive validation shows that the proposed method outperforms existing methods on several error metrics as well as in subjective evaluations. Model code, demos and datasets are available on our project webpage. Israel D. Gebru, Todd Keebler, Jake Sandakly, Steven Krenn, Dejan Markovic, Julia Buffalini, Samuel Hassel, Alexander Richard |
ICASSP | 5 |
| 2025 | ComplexDec: A Domain-robust High-fidelity Neural Audio Codec with Complex Spectrum ModelingabstractNeural audio codecs have been widely adopted in audio-generative tasks because their compact and discrete representations are suitable for both large-language-model-style and regression-based generative models. However, most neural codecs struggle to model out-of-domain audio, resulting in error propagations to downstream generative tasks. In this paper, we first argue that information loss from codec compression degrades out-of-domain robustness. Then, we propose full-band 48 kHz ComplexDec with complex spectral input and output to ease the information loss while adopting the same 24 kbps bitrate as the baseline AuidoDec and ScoreDec. Objective and subjective evaluations demonstrate the out-of-domain robustness of ComplexDec trained using only the 30-hour VCTK corpus. Yi-Chiao Wu, Dejan Markovic, Steven Krenn, Israel D. Gebru, Alexander Richard |
ICASSP | 2 |
| 2025 | BinauralFlow: A Causal and Streamable Approach for High-Quality Binaural Speech Synthesis with Flow Matching ModelsabstractBinaural rendering aims to synthesize binaural audio that mimics natural hearing based on a mono audio and the locations of the speaker and listener. Although many methods have been proposed to solve this problem, they struggle with rendering quality and streamable inference. Synthesizing high-quality binaural audio that is indistinguishable from real-world recordings requires precise modeling of binaural cues, room reverb, and ambient sounds. Additionally, real-world applications demand streaming inference. To address these challenges, we propose a flow matching based streaming binaural speech synthesis framework called BinauralFlow. We consider binaural rendering to be a generation problem rather than a regression problem and design a conditional flow matching model to render high-quality audio. Moreover, we design a causal U-Net architecture that estimates the current audio frame solely based on past information to tailor generative models for streaming inference. Finally, we introduce a continuous inference pipeline incorporating streaming STFT/ISTFT operations, a buffer bank, a midpoint solver, and an early skip schedule to improve rendering continuity and speed. Quantitative and qualitative evaluations demonstrate the superiority of our method over SOTA approaches. A perceptual study further reveals that our model is nearly indistinguishable from real-world recordings, with a 42% confusion rate. Susan Liang, Dejan Markovic, Israel D. Gebru, Steven Krenn, Todd Keebler, Jake Sandakly, Frank Yu, Samuel Hassel, Chenliang Xu, Alexander Richard |
ICML | 2 |
| 2024 | Modeling and Driving Human Body Soundfields Through Acoustic Primitives
Chao Huang 0033, Dejan Markovic, Chenliang Xu, Alexander Richard |
ECCV (10) | 2 |
| 2024 | ScoreDec: A Phase-Preserving High-Fidelity Audio Codec with a Generalized Score-Based Diffusion Post-FilterabstractAlthough recent mainstream waveform-domain end-to-end (E2E) neural audio codecs achieve impressive coded audio quality with a very low bitrate, the quality gap between the coded and natural audio is still significant. A generative adversarial network (GAN) training is usually required for these E2E neural codecs because of the difficulty of direct phase modeling. However, such adversarial learning hinders these codecs from preserving the original phase information. To achieve human-level naturalness with a reasonable bitrate, preserve the original phase, and get rid of the tricky and opaque GAN training, we develop a score-based diffusion post-filter (SPF) in the complex spectral domain and combine our previous AudioDec with the SPF to propose ScoreDec, which can be trained using only spectral and score-matching losses. Both the objective and subjective experimental results show that ScoreDec with a 24 kbps bitrate encodes and decodes full-band 48 kHz speech with human-level naturalness and well-preserved phase information. Yi-Chiao Wu, Dejan Markovic, Steven Krenn, Israel D. Gebru, Alexander Richard |
ICASSP | 2 |
| 2023 | Nord: Non-Matching Reference Based Relative Depth Estimation from Binaural SpeechabstractWe propose NORD: a novel framework for estimating the relative depth between two binaural speech recordings. In contrast to existing depth estimation techniques, ours only requires audio signals as input. We trained the framework to solve depth preference (i.e. which input perceptually sounds closer to the listener’s head), and quantification tasks (i.e. quantifying the depth difference between the inputs). In addition, training leverages recent advances in metric and multi-task learning, which allows the framework to be invariant to both signal content (i.e. non-matched reference) and directional cues (i.e. azimuth and elevation). Our framework has additional useful qualities that make it suitable for use as an objective metric to benchmark binaural audio systems, particularly depth perception and sound externalization, which we demonstrate through experiments. We also show that NORD generalizes well under different reverberation and environments. The results from preference and quantification tasks correlate well with measured results. Pranay Manocha, Israel D. Gebru, Anurag Kumar 0003, Dejan Markovic, Alexander Richard |
ICASSP | 4 |
| 2023 | Audiodec: An Open-Source Streaming High-Fidelity Neural Audio CodecabstractA good audio codec for live applications such as telecommunication is characterized by three key properties: (1) compression, i.e. the bitrate that is required to transmit the signal should be as low as possible; (2) latency, i.e. encoding and decoding the signal needs to be fast enough to enable communication without or with only minimal noticeable delay; and (3) reconstruction quality of the signal. In this work, we propose an open-source, streamable, and real-time neural audio codec that achieves strong performance along all three axes: it can reconstruct highly natural sounding 48 kHz speech signals while operating at only 12 kbps and running with less than 6 ms (GPU)/10 ms (CPU) latency. An efficient training paradigm is also demonstrated for developing such neural audio codecs for real-world scenarios. Both objective and subjective evaluations using the VCTK corpus are provided. To sum up, AudioDec is a well-developed plug-and-play benchmark for audio codec applications. Yi-Chiao Wu, Israel D. Gebru, Dejan Markovic, Alexander Richard |
ICASSP | 3 |
| 2023 | Spatialization Quality Metric for Binaural Speech
Pranay Manocha, Israel D. Gebru, Anurag Kumar 0003, Dejan Markovic, Alexander Richard |
INTERSPEECH | 4 |
| 2023 | Sounding Bodies: Modeling 3D Spatial Sound of Humans Using Body Pose and AudioabstractWhile 3D human body modeling has received much attention in computer vision, modeling the acoustic equivalent, i.e. modeling 3D spatial audio produced by body motion and speech, has fallen short in the community. To close this gap, we present a model that can generate accurate 3D spatial audio for full human bodies. The system consumes, as input, audio signals from headset microphones and body pose, and produces, as output, a 3D sound field surrounding the transmitter's body, from which spatial audio can be rendered at any arbitrary position in the 3D space. We collect a first-of-its-kind multimodal dataset of human bodies, recorded with multiple cameras and a spherical array of 345 microphones. In an empirical evaluation, we demonstrate that our model can produce accurate body-induced sound fields when trained with a suitable loss. Dataset and code are available online. Xudong Xu, Dejan Markovic, Jake Sandakly, Todd Keebler, Steven Krenn, Alexander Richard |
NeurIPS | 2 |
| 2022 | Audio-Visual Speech Codecs: Rethinking Audio-Visual Speech Enhancement by Re-SynthesisabstractSince facial actions such as lip movements contain significant information about speech content, it is not surprising that audio-visual speech enhancement methods are more accurate than their audio-only counterparts. Yet, state-of-the-art approaches still struggle to generate clean, realistic speech without noise artifacts and unnatural distortions in challenging acoustic environments. In this paper, we propose a novel audio-visual speech enhancement framework for high-fidelity telecommunications in AR/VR. Our approach leverages audio-visual speech cues to generate the codes of a neural speech codec, enabling efficient synthesis of clean, realistic speech from noisy signals. Given the importance of speaker-specific cues in speech, we focus on developing personalized models that work well for individual speakers. We demonstrate the efficacy of our approach on a new audio-visual speech dataset collected in an unconstrained, large vocabulary setting, as well as existing audio-visual datasets, outperforming speech enhancement baselines on both quantitative metrics and human evaluation studies. Please see the supplemental video for qualitative results11https://github.com/facebookresearch/facestar/releases/download/paper_materials/video.mp4. Karren Yang, Dejan Markovic, Steven Krenn, Vasu Agrawal, Alexander Richard |
CVPR | 2 |
| 2022 | End-to-End Binaural Speech SynthesisabstractIn this work, we present an end-to-end binaural speech synthesis system that combines a low-bitrate audio codec with a powerful binaural decoder that is capable of accurate speech binauralization while faithfully reconstructing environmental factors like ambient noise or reverb.The network is a modified vectorquantized variational autoencoder, trained with several carefully designed objectives, including an adversarial loss.We evaluate the proposed system on an internal binaural dataset with objective metrics and a perceptual study.Results show that the proposed approach matches the ground truth data more closely than previous methods.In particular, we demonstrate the capability of the adversarial loss in capturing environment effects needed to create an authentic auditory scene. Wen-Chin Huang, Dejan Markovic, Alexander Richard, Israel D. Gebru, Anjali Menon |
INTERSPEECH | 2 |
| 2022 | Implicit Neural Spatial Filtering for Multichannel Source Separation in the Waveform Domain
Dejan Markovic, Alexandre Défossez, Alexander Richard |
INTERSPEECH | 1 |
| 2021 | Implicit HRTF Modeling Using Temporal Convolutional NetworksabstractEstimation of accurate head-related transfer functions (HRTFs) is crucial to achieve realistic binaural acoustic experiences. HRTFs depend on source/listener locations and are therefore expensive and cumbersome to measure; traditional approaches require listener-dependent measurements of HRTFs at thousands of distinct spatial directions in an anechoic chamber. In this work, we present a data-driven approach to learn HRTFs implicitly with a neural network that achieves state of the art results compared to traditional approaches but relies on a much simpler data capture that can be performed in arbitrary, non-anechoic rooms. Despite that simpler and less acoustically ideal data capture, our deep learning based approach learns HRTF of high quality. We show in a perceptual study that the produced binaural audio is ranked on par with traditional DSP approaches by humans and illustrate that interaural time differences (ITDs), interaural level differences (ILDs) and spectral clues are accurately estimated. Israel D. Gebru, Dejan Markovic, Alexander Richard, Steven Krenn, Gladstone Alexander Butler, Fernando De la Torre, Yaser Sheikh |
ICASSP | 2 |
| 2021 | Neural Synthesis of Binaural Speech From Mono Audio
Alexander Richard, Dejan Markovic, Israel D. Gebru, Steven Krenn, Gladstone Alexander Butler, Fernando De la Torre, Yaser Sheikh |
ICLR | 2 |
| 2019 | Soundfield Reconstruction in Reverberant Environments Using Higher-order Microphones and Impulse Response MeasurementsabstractThis paper addresses the problem of soundfield reconstruction over a large area using a distributed array of higher-order microphones. Given an area enclosed by the array, one can distinguish between two components of the soundfield: the interior soundfield generated by sources outside of the enclosed area and the exterior soundfield generated by sources inside the enclosed area. These components form an indistinguishable mixture and, despite the existence of theoretical solutions to separate them, practical implementation is challenging due to high number of microphones needed for large regions and high frequencies. In this work, we consider a scenario where the interior soundfield is characterized by reverberation and show how a set of RIR measurements can be used to parametrize the interior component as a function of the exterior component, effectively reducing the unknowns of the problem. Federico Borra, Israel D. Gebru, Dejan Markovic |
ICASSP | 3 |
| 2018 | A Weighted Least Squares Beam Shaping Technique for Sound Field ControlabstractA weighted least squares beam shaping technique for sound field control using a loudspeaker array is proposed. Given a desired spatial response at prescribed control points, the space-time filter is designed by solving a least squares minimization problem. To reduce the computational effort, we propose to place control points only along an arc of circumference centered at the center of the array and passing through a region of interest. Furthermore, we adopt a weighted least squares approach for the design of the space-time filter, so that control points at directions towards which we admit a looser control of the sound field are less relevant in the filter design. The choice of the weights depends on the specific application and we demonstrate the feasibility of the proposed approach for sound zones scenario with one bright and one dark zone. Antonio Canclini, Dejan Markovic, Martin Schneider 0009, Fabio Antonacci, Emanuël A. P. Habets, Andreas Walther 0001, Augusto Sarti |
ICASSP | 2 |
| 2016 | Reconfigure your RTL with EFLX join the SoC revolutionabstractPresents a collection of slides covering the following topics: RTL; EFLX; SoC revolution; FPGA; and 128-tap programmable FIR. Cheng C. Wang, Dejan Markovic |
Hot Chips Symposium | 2 |
| 2016 | A linear operator for the computation of soundfield mapsabstractIn the process of soundfield imaging, as defined in the literature, a microphone array is subdivided into overlapping sub-arrays and soundfield images are obtained by juxtaposition of spatial spectra computed from individual subarray data. In this paper we show that the whole process can be conveniently seen as a linear transformation applied to array data. This linear transformation embeds a nonlinear mapping to cast the directional information in a more convenient domain: the ray space. We show by simulations that the proposed formulation is suitable for fast implementation of the soundfield imaging operation, and, more specifically, for the localization of acoustic sources. Lucio Bianchi, V. Baldini Anastasio, Dejan Markovic, Fabio Antonacci, Augusto Sarti, Stefano Tubaro |
ICASSP | 3 |
| 2016 | Extraction of Acoustic Sources Through the Processing of Sound Field Maps in the Ray SpaceabstractOur goal is to develop a model-based approach to acoustic source extraction from microphone array data, which is suitable for both near-field and far-field sources. A signal representation based on plane-wave (PW) decomposition is suitable for acoustic sources in the far field as the resulting spectrum turns out to be impulsive. When the source approaches the array, however, the curvature of the wavefront causes the spectrum of the PW components to depart from impulsive behavior, thus making source extraction harder to attain. In this paper, we adopt a sound field representation based on the local estimation of the plenacoustic function along the array line. This approach consists of dividing the array into subarrays, and applying the PW analysis on individual subarrays. This has the immediate result of extending the range of validity of the far-field hypothesis, as a source that enters the near-field range of the extended array is still in the far-field range of the subarrays. PW analysis on subarrays allows us to construct the so-called sound field map in a domain of acoustic visibility called ray space. The extraction of the desired source is accomplished through spatial filtering of the sound field map. The design of the spatial filter relies on a linear minimum mean square error criterion defined on the sound field map. The effectiveness of the proposed methodology is proven through an extensive simulation campaign as well as real experiments. Dejan Markovic, Fabio Antonacci, Lucio Bianchi, Stefano Tubaro, Augusto Sarti |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2016 | 3D Beam Tracing Based on Visibility Lookup for Interactive Acoustic ModelingabstractWe present a method for accelerating the computation of specular reflections in complex 3D enclosures, based on acoustic beam tracing. Our method constructs the beam tree on the fly through an iterative lookup process of a precomputed data structure that collects the information on the exact mutual visibility among all reflectors in the environment (region-to-region visibility). This information is encoded in the form of visibility regions that are conveniently represented in the space of acoustic rays using the Plücker coordinates. During the beam tracing phase, the visibility of the environment from the source position (the beam tree) is evaluated by traversing the precomputed visibility data structure and testing the presence of beams inside the visibility regions. The Plücker parameterization simplifies this procedure and reduces its computational burden, as it turns out to be an iterative intersection of linear subspaces. Similarly, during the path determination phase, acoustic paths are found by testing their presence within the nodes of the beam tree data structure. The simulations show that, with an average computation time per beam in the order of a dozen of microseconds, the proposed method can compute a large number of beams at rates suitable for interactive applications with moving sources and receivers. Dejan Markovic, Fabio Antonacci, Augusto Sarti, Stefano Tubaro |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2015 | Multiview Soundfield Imaging in the Projective Ray SpaceabstractA soundfield image is a data structure that efficiently encodes and represents the wave field as captured by a microphone array. Its representation is based on the directional plenacoustic function, which is defined as the radiance of the acoustic paths (rays) that cross the segment that the array lies upon. The soundfield image can be processed “as is” to develop a variety of applications. In its original formulation, the soundfield image is based on a Euclidean parameterization that can accommodate a limited range of rays and is suitable for managing a single array only. In this paper, we generalize this methodology to the case of multiple microphone arrays deployed in space. The use of multiple arrays allows us to capture truly global information on the sound field but requires us to rethink the ray space, and adopt a global representation of the acoustic rays based on projective geometry. After introducing the new parameterization, we present two examples of applications: the estimation of the mutual poses of two or more arrays (self-calibration); and the localization of multiple acoustic sources. The effectiveness of these applications is proven through simulations as well as real data experiments. Dejan Markovic, Fabio Antonacci, Augusto Sarti, Stefano Tubaro |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2014 | A scalable sparse matrix-vector multiplication kernel for energy-efficient sparse-blas on FPGAsabstractSparse Matrix-Vector Multiplication (SpMxV) is a widely used mathematical operation in many high-performance scientific and engineering applications. In recent years, tuned software libraries for multi-core microprocessors (CPUs) and graphics processing units (GPUs) have become the status quo for computing SpMxV. However, the computational throughput of these libraries for sparse matrices tends to be significantly lower than that of dense matrices, mostly due to the fact that the compression formats required to efficiently store sparse matrices mismatches traditional computing architectures. This paper describes an FPGA-based SpMxV kernel that is scalable to efficiently utilize the available memory bandwidth and computing resources. Benchmarking on a Virtex-5 SX95T FPGA demonstrates an average computational efficiency of 91.85%. The kernel achieves a peak computational efficiency of 99.8%, a >50x improvement over two Intel Core i7 processors (i7-2600 and i7-4770) and showing a >300x improvement over two NVIDA GPUs (GTX 660 and GTX Titan), when running the MKL and cuSPARSE sparse-BLAS libraries, respectively. In addition, the SpMxV FPGA kernel is able to achieve higher performance than its CPU and GPU counterparts, while using only 64 single-precision processing elements, with an overall 38-50x improvement in energy efficiency. Richard Dorrance, Fengbo Ren, Dejan Markovic |
FPGA | 3 |
| 2014 | Logarithmic quantization scheme for reduced hardware cost and improved error floor in non-binary LDPC decodersabstractNon-binary low-density parity-check (NB-LDPC) codes exhibit excellent error correction performance at the cost of high computational complexity of the decoding algorithm. A logarithmic quantization scheme is proposed to reduce the VLSI implementation cost of the Min-Max decoding algorithm, by scaling down the complexity of the check node calculations that are the prime bottleneck in NB-LDPC decoding. The proposed scheme is also shown to be robust against certain types of errors, and thus enables excellent error correction capabilities even for aggressively reduced wordlengths, relative to traditional, uniform quantization schemes that exhibit either poor waterfall region performance or high error floors for a similar number of bits. The proposed scheme is directly applicable to existing architectures with few modifications and is shown to reduce the computational complexity by up to 40%, especially for large field orders. Yuta Toriyama, Behzad Amiri, Lara Dolecek, Dejan Markovic |
GLOBECOM | 4 |
| 2014 | Estimation of Acoustic Reflection Coefficients Through Pseudospectrum MatchingabstractEstimating the geometric and reflective properties of the environment is important for a wide range of applications of space-time audio processing, from acoustic scene analysis to room equalization and spatial audio rendering. In this manuscript, we propose a methodology for frequency-subband in-situ estimation of the reflection coefficients of planar surfaces. This is a rather challenging task, as the reflection coefficients depend on the frequency and the angle of incidence and their estimate is highly sensitive to background noise and interfering sources. Our method is based on the assumption that we know the geometry of the reflectors; the position and the radiation pattern of the source; the position and the spatial response of the array. Applying beamforming algorithms on a single set of measured sensor data, we estimate the angular distribution of the acoustic energy (angular pseudospectrum) that impinges on a microphone array. We then apply a two-step iterative estimation technique based on an Expectation-Maximization (EM) algorithm. The first step estimates the scaling factors. The second one infers the reflection coefficients from the scaling factors. Under the assumption of additive white Gaussian noise, we finally determine the reflection coefficients with a Maximum Likelihood (ML) estimation method. The effectiveness and the accuracy of the proposed technique are assessed through experiments based on measured data. Dejan Markovic, Konrad Kowalczyk, Fabio Antonacci, Christian Hofmann 0001, Augusto Sarti, Walter Kellermann |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2013 | A single-precision compressive sensing signal reconstruction engine on FPGAsabstractCompressive sensing (CS) is a promising technology for the low-power and cost-effective data acquisition in wireless healthcare systems. However, its efficient realtime signal reconstruction is still challenging, and there is a clear demand for hardware acceleration. In this paper, we present the first single-precision floating-point CS reconstruction engine implemented a Kintex-7 FPGA using the orthogonal matching pursuit (OMP) algorithm. In order to achieve high performance with maximum hardware utilization, we propose a highly parallel architecture that shares the computing resources among different tasks of OMP by using configurable processing elements (PEs). By fully utilizing the FPGA recourses, our implementation has 128 PEs in parallel and operates at 53.7 MHz. In addition, it can support 2x larger problem size and 10x more sparse coefficients than prior work, which enables higher reconstruction accuracy by adding finer details to the recovered signal. Hardware results from the ECG reconstruction tests show the same level of accuracy as the double-precision C program. Compared to the execution time of a 2.27 GHz CPU, the FPGA reconstruction achieves an average speed-up of 41x. Fengbo Ren, Richard Dorrance, Wenyao Xu, Dejan Markovic |
FPL | 4 |
| 2013 | An area-efficient minimum-time FFT schedule using single-ported memoryabstractFFT design requires an exhaustive recoupling of data across successive stages of computation. The resulting memory access patterns have constantly-changing strides, making it hard to interleave the data for reliable conflict-free access of operand pairs. We modify an existing method of “swizzling” data locations so as to guarantee conflict-free access within any given stage and, with minimal support for buffering, we provide conflict-free access across the boundaries of adjoining stages as well. As a result, implementations that would naively require either a fully-associative, or at the very least a multiported register file, can be implemented using four single-ported banks of memory per butterfly unit, plus one bypass buffer. Because fewer ports means less area, and given that a butterfly must read two inputs and write two results for each cycle of operation, this solution should represent the least-area memory configuration for a resource-constrained FFT. Using this scheme, we show examples including a minimal one-butterfly FFT having 9% less area versus a competing equal-performance design and 20% better throughput versus a competing equal-area design. Stephen Richardson, Ofer Shacham, Dejan Markovic, Mark Horowitz |
VLSI-SoC | 3 |
| 2013 | Soundfield Imaging in the Ray SpaceabstractIn this work we propose a general approach to acoustic scene analysis based on a novel data structure (ray-space image) that encodes the directional plenacoustic function over a line segment (Observation Window, OW). We define and describe a system for acquiring a ray-space image using a microphone array and refer to it as ray-space (or “soundfield”) camera. The method consists of acquiring the pseudo-spectra corresponding to a grid of sampling points over the OW, and remapping them onto the ray space, which parameterizes acoustic paths crossing the OW. The resulting ray-space image displays the information gathered by the sensors in such a way that the elements of the acoustic scene (sources and reflectors) will be easy to discern, recognize and extract. The key advantage of this method is that ray-space images, irrespective of the application, are generated by a common (and highly parallelizable) processing layer, and can be processed using methods coming from the extensive literature of pattern analysis. After defining the ideal ray-space image in terms of the directional plenacoustic function, we show how to acquire it using a microphone array. We also discuss resolution and aliasing issues and show two simple examples of applications of ray-space imaging. Dejan Markovic, Fabio Antonacci, Augusto Sarti, Stefano Tubaro |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2012 | Challenges and directions of ultra low energy wireless sensor nodes for biosignal monitoringabstractThis paper discusses design challenges and strategies for aggressively increasing energy efficiency of biosignal monitoring sensors. For the holistic understanding of energy efficiency, we introduce Energy Efficiency metric for all sensor communication blocks which include not only Rx/Tx RF&Analog, PLL and DSP/Modem but also Antenna and Power Management. Based on the metric, an ultra-low energy sensor node design at 2.36~2.5GHz is addressed from RFIC, DSP/Modem to Antenna. To tackle the stringent power requirements, we theoretically revisit the technology, circuits, architecture and system and explore the cross-layer power minimization algorithm. Seong Joong Kim, Bumman Kim, Sangwook Nam, Dejan Markovic, Sang-Gug Lee 0001, Jaesup Lee |
ISCAS | 4 |
| 2011 | Compressive Sensing of Neural Action Potentials Using a Learned Union of SupportsabstractWireless neural recording systems are subject to stringent power consumption constraints to support long-term recordings and to allow for implantation inside the brain. In this paper, we propose using a combination of on-chip detection of action potentials ("spikes") and compressive sensing (CS) techniques to reduce the power consumption of the neural recording system by reducing the power required for wireless transmission. We empirically verify that spikes are compressible in the wavelet domain and show that spikes from different neurons acquired from the same electrode have subtly different sparsity patterns or supports. We exploit the latter fact to further enhance the sparsity by incorporating a union of these supports learned over time into the spike recovery procedure. We show, using extra cellular recordings from human subjects, that this mechanism improves the SNDR of the recovered spikes over conventional basis pursuit recovery by up to 9.5 dB (6 dB mean) for the same number of CS measurements. Though the compression ratio in our system is contingent on the spike rate at the electrode, for the datasets considered here, the mean ratio achieved for 20-dB SNDR recovery is improved from 26:1 to 43:1 using the learned union of supports. Zainul Charbiwala, Vaibhav Karkare, Sarah Gibson, Dejan Markovic, Mani Srivastava 0001 |
BSN | 4 |
| 2011 | An Energy-Efficient VLSI Architecture for Cognitive Radio Wideband Spectrum SensingabstractSpectrum sensing over a wide bandwidth increases the probability of finding unutilized spectrum for cognitive radios. However, energy-efficient VLSI realization of wideband sensing algorithms is challenging due to complex signal processing and real-time requirement. In addition, strong primary users introduce spectral leakage in adjacent unused bands, resulting in sensing performance degradation. To address these challenges, we propose a cascaded filter-bank channelization scheme and its reconfigurable VLSI architecture that can be optimized for power, area, sensing time, and detection performance. In addition, the spectral leakage due to strong interference is compensated with an interference cancellation method. Compared to the conventional PSD-based energy detection, the proposed channelization scheme supports reliable wideband signal detection with 2.5× less power. Given a 0.5ms sensing time under 30-dB adjacent-band interference-to-noise ratio, a 30× sensing-time improvement is achieved while maintaining a false- alarm probability of 0.1 and a detection probability of 0.9. The energy efficiency is improved from lower power consumption and reduced sensing time. Tsung-Han Yu, Chia-Hsiang Yang, Dejan Markovic, Danijela Cabric |
GLOBECOM | 3 |
| 2011 | A Hardware-Efficient VLSI Architecture for Hybrid Sphere-MCMC DetectionabstractThis paper presents a hybrid soft-output MIMO detector that searches reliable soft-information in both deterministic and probabilistic ways. The fixed-complexity sphere detector (FSD) is first applied to provide near maximum-likelihood (ML) solutions. The solutions are next used to initialize the Markov Chain Monte Carlo (MCMC) detector that uses parallel Gibbs samplers (GSs) for remaining candidate enumeration. A low-complexity VLSI architecture is proposed to demonstrate the feasibility of hardware realization for high-throughput applications. Simulation results indicate that the hybrid detector has a 2.3x complexity reduction and a 2× throughput improvement compared to individual soft-output FSD and MCMC detectors. Fang-Li Yuan, Chia-Hsiang Yang, Dejan Markovic |
GLOBECOM | 3 |
| 2011 | Effects of quantization on neural spike sortingabstractWireless neural recording systems require the data-acquisition and signal-processing hardware to be moved to the transmit side. The strict power-density constraints on implanted devices require new ideas for system power minimization. Minimizing the number of bits of the ADC would have a significant impact on the total system power by reducing the power of the ADC, the DSP, and the transmitter. In this paper we examine the effects of quantization on the performance of spike sorting. We derive the resolution required of uniform quantizers to ensure the most accurate spike detection and clustering, and compare this to simulation results. We then provide evidence that optimal quantizers are well suited for neural data, and show that optimal quantizers provide a savings of at least 2 bits compared to uniform quantizers. Sarah Gibson, Victoria Wang, Dejan Markovic |
ISCAS | 3 |
| 2010 | Cognitive Radio Wideband Spectrum Sensing Using Multitap Windowing and Power Detection with Threshold AdaptationabstractA common technique for cognitive radio wideband spectrum sensing is energy/power detection of primary users (PU) in frequency domain. Specifically, power spectrum estimation methods are combined with power detection statistics to test the PU presence. However, when detecting in a particular band of interest these techniques suffer from energy leakage and adjacent channel interference. In this paper, we derive a common matrix framework for the analytical performance of power detectors when FFT, windowed FFT, or multitap windowed FFT are used. Our matrix model is verified by simulations of modulated PU signals. We further propose a low-complexity compensation method to adapt the thresholds in the presence of large power difference between channels. By using both the multitap windowing and the constant false-alarm-rate method in the presence of strong signals, we demonstrate a 2-times increase in the detection rate performance as compared to existing methods. The proposed algorithm achieves similar PFAand PDas FFT at lower sample complexity, leading to reduced sensing times. Tsung-Han Yu, Santiago Rodriguez-Parera, Dejan Markovic, Danijela Cabric |
ICC | 3 |
| 2010 | Visibility-based beam tracing for soundfield renderingabstractIn this paper we present a visibility-based beam tracing solution for the simulation of the acoustics of environment that makes use of a projective geometry representation. More specifically, projective geometry turns out to be useful for the pre-computation of the visibility among all the reflectors in the environment. The simulation engine has a straightforward application in the rendering of the acoustics of virtual environments using loudspeaker arrays. More specifically, the acoustic wavefield is conceived as a superposition of acoustic beams, whose parameters (i.e. origin, orientation and aperture) are computed using the fast beam tracing methodology presented here. This information is processed by the rendering engine to compute spatial filters to be applied to the loudspeakers within the array. Simulative results show that an accurate simulation of the acoustic wavefield can be obtained using this approach. Dejan Markovic, Antonio Canclini, Fabio Antonacci, Augusto Sarti, Stefano Tubaro |
MMSP | 1 |
| 2010 | Ultralow-Power Design in Near-Threshold RegionabstractOperation in the subthreshold region most often is synonymous to minimum-energy operation. Yet, the penalty in performance is huge. In this paper, we explore how design in the moderate inversion region helps to recover some of that lost performance, while staying quite close to the minimum-energy point. An energy-delay modeling framework that extends over the weak, moderate, and strong inversion regions is developed. The impact of activity and design parameters such as supply voltage and transistor sizing on the energy and performance in this operational region is derived. The quantitative benefits of operating in near-threshold region are established using some simple examples. The paper shows that a 20% increase in energy from the minimum-energy point gives back ten times in performance. Based on these observations, a pass-transistor based logic family that excels in this operational region is introduced. The logic family operates most of its logic in the above-threshold mode (using low-threshold transistors), yet containing leakage to only those in subthreshold. Operation below minimum-energy point of CMOS is demonstrated. In leakage-dominated ultralow-power designs, time-multiplexing will be shown to yield not only area, but also energy reduction due to lower leakage. Finally, the paper demonstrates the use of ultralow-power design techniques in chip synthesis. Dejan Markovic, Cheng C. Wang, Louis P. Alarcón, Tsung-Te Liu, Jan M. Rabaey |
Proc. IEEE | 1 |
| 2008 | A Multi-Core Sphere Decoder VLSI Architecture for MIMO CommunicationsabstractThe sphere decoding algorithm finds applications in multi-input multi-output (MIMO) decoding, because it achieves near maximum likelihood (ML) detection performance with significantly reduced computational complexity. Previous work has focused on implementations based on K-best or depth-first search, limiting the BER performance or the search speed. This paper presents a scalable multi-core sphere decoder architecture that can combine the advantages of the K-best and depth-first search methods. The proposed architecture demonstrated a 3-5 dB improvement in the BER performance for 16times16 systems using 16 processing elements (PEs) compared to the architecture with one PE. An improved search speed of the multi-core architecture also enables a 10times energy efficiency improvement over the single core architecture for the same data rate. Chia-Hsiang Yang, Dejan Markovic |
GLOBECOM | 2 |
| 2008 | A Flexible VLSI Architecture for Extracting Diversity and Spatial Multiplexing Gains in MIMO ChannelsabstractThe sphere decoding algorithm is able to approach maximum likelihood (ML) detection with significantly reduced computational complexity for multi-input multi-output (MIMO) communications. The computational reduction makes it attractive for hardware implementation. This paper presents a unified sphere decoder architecture that deploys diversity-multiplexing tradeoff in MIMO channels by taking advantage of the flexibility in the number of antennas and modulation schemes. Several signal processing and circuit techniques are constructively combined to reduce the hardware complexity: a 20 times area reduction is achieved even without interleaving of sub-carriers compared to the direct-mapped architecture. The proposed flexible architecture supports antenna arrays from 2x2 to 16x16, modulations from BPSK to 64-QAM, over 16 to 128 sub-carriers. The peak estimated data rate exceeds 1.5 Gbps over a 16 MHz bandwidth in just 0.55 mm2in a standard 90 nm CMOS process. Chia-Hsiang Yang, Dejan Markovic |
ICC | 2 |
| 2008 | Integrated circuit design with NEM relaysabstractTo overcome the energy-efficiency limitations imposed by finite sub-threshold slope in CMOS transistors, this paper explores the design of integrated circuits based on nano-electro-mechanical (NEM) relays. A dynamical Verilog-A model of the NEM relay is described and correlated to device measurements. Using this model we explore NEM relay design strategies for digital logic and I/O that can significantly improve the energy efficiency of the whole VLSI system. By exploiting the low effective threshold voltage and zero leakage achievable with these relays, we show that NEM relay-based adders can achieve an order of magnitude or more improvement in energy efficiency over CMOS adders with ns-range delays and with no area penalty. By applying parallelism, this improvement in energy-efficiency can be achieved at higher throughputs as well, at the cost of increased area. Similar improvements in high-speed I/O energy are also predicted by making use of the relays to implement highly energy-efficient digital-to-analog and analog-to-digital converters. Fred Chen, Hei Kam, Dejan Markovic, Tsu-Jae King Liu, Vladimir Stojanovic, Elad Alon |
ICCAD | 3 |
| 2008 | Linear analysis of random process variabilityabstractThis paper describes an alternate method to Monte Carlo for calculating circuit node voltage and branch current variances due to random process variability. Recent results show that the complex models traditionally used to describe random process variations of a transistor can be replaced by a single independent current noise source with a variance dependent on the transistor’s size and operating points. As a result, each transistor affected by random process variability can be modeled as a deterministic device in parallel with a current noise source. By replacing all the transistors in a circuit with this model, the spatial voltage variances of circuit nodes can be calculated through linear small-signal analysis. The idea is presented in this paper and a tool implemented for Berkeley SPICE is described. The results of the variability SPICE tool match the results from Monte Carlo run in SPECTRE, with an accuracy determined by the accuracy of the random process variability model. For example, the standard deviations computed by the tool are within a 5.0% accuracy of those calculated through measured silicon data which has a fitting error of 5.4%. The Monte Carlo method computes node variances in a time proportional to the number of circuit nodes and the number of iterations, whereas the computation time required by the variability tool is only a function of the number of circuit nodes. For large analog designs this results in a significant speed-up in the amount of time required to calculate circuit node variances. Victoria Wang, Dejan Markovic |
ICCAD | 2 |
| 2006 | Power and Area Efficient VLSI Architectures for Communication Signal ProcessingabstractA methodology for VLSI realization of signal processing algorithms for wireless communications is presented that optimizes architecture for reduced power and area. When power is limited, optimal architecture represents a point on the best power-area tradeoff curve that is obtained by balancing the algorithm throughput with the power-performance tradeoff of the underlying building blocks. Architectural optimization is done in the graphical Matlab/Simulink environment, which is also used for algorithm verification. Hardware description language produced by Simulink enables algorithm emulation on the FPGA and also serves as design entry for the chip realization. This is illustrated on complex multi-dimensional algorithms such as wideband MIMO channel decoupling through singular value decomposition (SVD) using 16 sub-carriers. Dejan Markovic, Borivoje Nikolic, Robert W. Brodersen |
ICC | 1 |
| 2002 | Methods for true power minimizationabstractThis paper presents methods for efficient power minimization at circuit and micro-architectural levels. The potential energy savings are strongly related to the energy profile of a circuit. These savings are obtained by using gate sizing, supply voltage, and threshold voltage optimization, to minimize energy consumption subject to a delay constraint. The true power minimization is achieved when the energy reduction potentials of all tuning variables are balanced. We derive the sensitivity of energy to delay for each of the tuning variables connecting its energy saving potential to the physical properties of the circuit. This helps to develop understanding of optimization performance and identify the most efficient techniques for energy reduction. The optimizations are applied to some examples that span typical circuit topologies including inverter chains, SRAM decoders, and adders. At a delay of 20% larger than the minimum, energy savings of 40% to 70% are possible, indicating that achieving peak performance is expensive in terms of energy. Energy savings of about 50% can be achieved without delay penalty with the balancing of sizes, supplies, and thresholds. Robert W. Brodersen, Mark Horowitz, Dejan Markovic, Borivoje Nikolic, Vladimir Stojanovic |
ICCAD | 3 |
| 2001 | Analysis and design of low-energy flip-flopsabstractThis paper develops a methodology for selecting and optimizing flip-flops for low-energy systems with constant throughput. Characterization metrics, relevant to low-energy systems are discussed, providing insight into timing and energy parameters at both the circuit and system levels. Transistor sizes are optimized for minimal delay under constrained energy consumption. This methodology is applied to characterization of various flip-flop styles and their comparison in 0.25µm CMOS technology under scaled supply voltages. A transmission-gate master-slave latchpair has the largest internal race margin, lowest energy consumption, and has energy-delay product comparable to much faster pulse-triggered latches. Dejan Markovic, Borivoje Nikolic, Robert W. Brodersen |
ISLPED | 1 |