Matija Marolt

dblp:96/6903 · DBLP profile ↗
← Back
13ranked-venue papers
3as first author
8since 2021 · last 2026
0000-0002-0619-8789ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 5 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Artificial intelligence and machine learning · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 A Two-Case Study on Extending Conventional Music Practice Using Mobile Applications
abstract
Nasl. z nasl. zaslona.
Matevz Pesek, Jaka Kuzner, Emir Hodzic, Klara Znidersic, Matija Marolt
CSEDU (1)5
2025 Storytelling in Gamified Rhythmic Training
Matevz Pesek, Zala Pregelj, Klara Znidersic, Matija Marolt
CSEDU (1)4
2025 Language Learning with VR: The Effects of Immersive Gamification on Student Motivation and Knowledge
Klara Znidersic, Nik Jan Spruk, Matija Marolt, Matevz Pesek
CSEDU (1)3
2024 Troubadour: Inverse Dictation Games for Ear Training
Klara Znidersic, Matija Podbreznik, Ziga Klun, Peter Savli, Matija Marolt, Matevz Pesek
CSEDU (1)5
2024 Anomalous Sound Detection by Feature-Level Anomaly Simulation
abstract
Recently a growing number of works focus on machine defect detection from anomalous audio patterns. The datasets for the machine audio domain are scarce and recent methods that perform well on benchmarks such as DCASE2020 Task 2, rely on auxiliary information such as annotated data from other training classes in the domain to extract information that can be used in deep-learning classification-based anomaly detection approaches. However, in practical scenarios, annotated data from the same domain may not be readily available so annotation-free methods that can learn appropriate audio representations from unannotated data are needed. We propose AudDSR, a simulation-based anomaly detection method that learns to detect anomalies without additional annotated data and instead focuses on a discrete feature space sampling method for an anomaly simulation process. AudDSR outperforms competing methods that do not rely on annotated data on the DCASE2020 anomalous sound detection benchmark and even matches the performance of some methods that utilize additional annotation information.
Vitjan Zavrtanik, Matija Marolt, Matej Kristan, Danijel Skocaj
ICASSP2
2024 Evaluation of depth perception in crowded volumes
abstract
Depth perception in volumetric visualization plays a crucial role in the understanding and interpretation of volumetric data. Numerous visualization techniques, many of which rely on physically based optical effects, promise to improve depth perception but often do so without considering camera movement or the content of the volume. As a result, the findings from previous studies may not be directly applicable to crowded volumes, where a large number of contained structures disrupts spatial perception. Crowded volumes therefore require special analysis and visualization tools with sparsification capabilities. Interactivity is an integral part of visualizing and exploring crowded volumes, but has received little attention in previous studies. To address this gap, we conducted a study to assess the impact of different rendering techniques on depth perception in crowded volumes, with a particular focus on the effects of camera movement. The results show that depth perception considering camera motion depends much more on the content of the volume than on the chosen visualization technique. Furthermore, we found that conventional non-photorealistic rendering techniques, which have often performed poorly in previous studies, showed comparable performance to modern photorealistic techniques in our study. The source code for the visualization system, survey, and analysis, as well as the data set used in the study and the participants’ responses, have been made publicly available.
Ziga Lesar, Ciril Bohak, Matija Marolt
Comput. Graph.3
2024 Volume conductor: interactive visibility management for crowded volumes
abstract
Abstract We present a novel smart visibility system for visualizing crowded volumetric data containing many object instances. The presented approach allows users to form groups of objects through membership predicates and to individually control the visibility of the instances in each group. Unlike previous smart visibility approaches, our approach controls the visibility on a per-instance basis and decides which instances are displayed or hidden based on the membership predicates and the current view. Thus, cluttered and dense volumes that are notoriously difficult to explore effectively are automatically sparsified so that the essential information is extracted and presented to the user. The proposed system is generic and can be easily integrated into existing volume rendering applications and applied to many different domains. We demonstrate the use of the volume conductor for visualizing fiber-reinforced polymers and intracellular organelle structures.
Ziga Lesar, Ruwayda Alharbi, Ciril Bohak, Ondrej Strnad, Christoph Heinzl, Matija Marolt, Ivan Viola
Vis. Comput.6
2024 Combined volume and surface rendering with global illumination caching
abstract
Abstract We present a combined volume and surface rendering technique with global illumination caching. Our approach uses volumetric path tracing to compute the global illumination volume and local shading models for rendering the isosurface. By joining both visualization approaches, we have enhanced the display and illumination of the surfaces while preserving physically realistic illumination of the participating media. To achieve real-time performance and avoid recomputing the image when the camera view changes, we compute the global illumination volume incrementally and defer the projection to a later step. We evaluated our technique by comparing different local shading models for isosurface rendering with the result of full volumetric path tracing and with the non-caching variant of our technique. Results show that the caching and non-caching variants perform comparably well, while the caching variant has the added benefit of being camera-view-independent. Additionally, we show that our approach emphasizes the surfaces within volumes better than volumetric path tracing.
Uros Smajdek, Ziga Lesar, Matija Marolt, Ciril Bohak
Vis. Comput.3
2019 Prediction of music pairwise preferences from facial expressions
abstract
Users of a recommender system may be requested to express their preferences about items either with evaluations of items (e.g. a rating) or with comparisons of item pairs. In this work we focus on the acquisition of pairwise preferences in the music domain. Asking the user to explicitly compare music, i.e., which, among two listened tracks, is preferred, requires some user effort. We have therefore developed a novel approach for automatically extracting these preferences from the analysis of the facial expressions of the users while listening to the compared tracks. We have trained a predictor that infers user's pairwise preferences by using features extracted from these data. We show that the predictor performs better than a commonly used baseline, which leverages the user's listening duration of the tracks to infer pairwise preferences. Furthermore, we show that there are differences in the accuracy of the proposed method between users with different personalities and we have therefore adapted the trained model accordingly. Our work shows that by introducing a low user effort preference elicitation approach, which, however, requires to access information that may raise potential privacy issues (face expression), one can obtain good prediction accuracy of pairwise music preferences.
Marko Tkalcic, Nima Maleki, Matevz Pesek, Mehdi Elahi, Francesco Ricci 0001, Matija Marolt
IUI6
2017 A Research Tool for User Preferences Elicitation with Facial Expressions
abstract
We present a research tool for user preference elicitation that collects both explicit user feedback and unobtrusively acquired facial expressions. The concrete implementation is a web-based user interface where the user is presented with two music excerpts. After listening to both, the user provides a pairwise score (i.e. which of the two items is preferred) for each pair of music excerpts. The novelty of the demo is the integration of the unobtrusive acquisition of facial expressions through the webcam. During the listening of the music excerpts, the system extracts features related to the facial expressions of the user several times per second. The interaction runs as a web application, which allows for a large-scale remote acquisition of emotional data. Up to now, such acquisitions were usually done in controlled environments with few subjects, hence being of little use for the recommender systems community.
Marko Tkalcic, Nima Maleki, Matevz Pesek, Mehdi Elahi, Francesco Ricci 0001, Matija Marolt
RecSys6
2012 Automatic Transcription of Bell Chiming Recordings
abstract
Bell chiming is a folk music tradition that involves performers playing rhythmic patterns on church bells. The paper presents a method for automatic transcription of bell chiming recordings, where the goal is to detect the bells that were played and their onset times. We first present an algorithm that estimates the number of bells in a recording and their approximate spectra. The algorithm uses a modified version of the intelligent k-means algorithm, as well as some prior knowledge of church bell acoustics to find clusters of partials with synchronous onsets in the time-frequency representation of a recording. Cluster centers are used to initialize non-negative matrix factorization that factorizes the time-frequency representation into a set of basis vectors (bell spectra) and their activations. To transcribe a recording, we propose a probabilistic framework that integrates factorization and onset detection data with prior knowledge of bell chiming performance rules. Both parts of the algorithm are evaluated on a set of bell chiming field recordings.
Matija Marolt
IEEE Trans. Speech Audio Process.1
2008 A Mid-Level Representation for Melody-Based Retrieval in Audio Collections
abstract
Searching audio collections using high-level musical descriptors is a difficult problem, due to the lack of reliable methods for extracting melody, harmony, rhythm, and other such descriptors from unstructured audio signals. In this paper, we present a novel approach to melody-based retrieval in audio collections. Our approach supports audio, as well as symbolic queries and ranks results according to melodic similarity to the query. We introduce a beat-synchronous melodic representation consisting of salient melodic lines, which are extracted from the analyzed audio signal. We propose the use of a 2D shift-invariant transform to extract shift-invariant melodic fragments from the melodic representation and demonstrate how such fragments can be indexed and stored in a song database. An efficient search algorithm based on locality-sensitive hashing is used to perform retrieval according to similarity of melodic fragments. On the cover song detection task, good results are achieved for audio, as well as for symbolic queries, while fast retrieval performance makes the proposed system suitable for retrieval in large databases.
Matija Marolt
IEEE Trans. Multim.1
2004 A connectionist approach to automatic transcription of polyphonic piano music
abstract
In this paper, we present a connectionist approach to automatic transcription of polyphonic piano music. We first compare the performance of several neural network models on the task of recognizing tones from time-frequency representation of a musical signal. We then propose a new partial tracking technique, based on a combination of an auditory model and adaptive oscillator networks. We show how synchronization of adaptive oscillators can be exploited to track partials in a musical signal. We also present an extension of our technique for tracking individual partials to a method for tracking groups of partials by joining adaptive oscillators into networks. We show that oscillator networks improve the accuracy of transcription with neural networks. We also provide a short overview of our entire transcription system and present its performance on transcriptions of several synthesized and real piano recordings. Results show that our approach represents a viable alternative to existing transcription systems.
Matija Marolt
IEEE Trans. Multim.1