Marc A. Kastner 0001

dblp:203/3524-1 · also Marc Aurel Kastner · DBLP profile ↗
← Back
26ranked-venue papers
5as first author
19since 2021 · last 2025
0000-0002-9193-5973ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 24 · 5 first-author · 18 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Systems, architecture and hardware · 1
YearPublicationVenuePosition
2025 ICDAR 25: Intelligent Cross-Data Analysis and Retrieval
abstract
The sixth edition of the Intelligent Cross-Data Analysis and Retrieval (ICDAR) workshop continues to serve as a forum for researchers and practitioners addressing the integration, analysis, and retrieval of heterogeneous data sources. While individual modalities such as wearable sensors, lifelogging cameras, and social media have been well studied, analyzing cross-data that incorporates multiple perspectives remains a crucial yet challenging task for advancing human-centered applications. In 2025, the workshop received 19 submissions, of which 7 were accepted following a careful peer-review process, resulting in an acceptance rate of 37%. The accepted papers covered a wide range of topics, including zero-shot composed image retrieval, vision-language scene understanding, adaptive modality fusion, lightweight fine-tuning with truncated SVD, and real-world federated split learning on mobile devices. By fostering interdisciplinary collaboration across domains such as well-being, disaster mitigation, mobility, food computing, and smart cities, the workshop continues to highlight emerging challenges and solutions for building intelligent, sustainable, and human-centric systems driven by cross-modal and multimodal data analytics.
Takahiro Komamizu, Marc A. Kastner 0001, Minh-Son Dao, Michael Riegler 0001, Duc-Tien Dang-Nguyen, Son N. Tran
ICMR2
2025 MUWS 2025: The 4th International Workshop on Multimodal Human Understanding for the Web and Social Media
abstract
Multimodal human understanding is an evolving interdisciplinary field integrating computer science, psychology, and social sciences to model human perception, behaviour, and biases in multimodal data. While recent advancements in multimodal learning excel in tasks like image-text synthesis, they often overlook nuanced human-centric dynamics---such as cultural, political, and individual influences on how modalities (e.g., text and images) interact, complement, or contradict each other. The 4th International Workshop on Multimodal Human Understanding (MUWS) aims at addressing these challenges, fostering novel solutions that explicitly model human perception, behaviour, and biases in multimodal data, with a particular emphasis on real-world challenges in web and social media analysis. This year edition covers two tracks: (1) human-centred multimodal understanding, such as quantifying social biases, analysing sentiment and hate speech, and modelling cross-modal interactions through interdisciplinary theories (e.g., semiotics, gestalt psychology); and (2) Multimodal understanding of global events, supported by a newly curated dataset covering news articles with diverse stances, which facilitates research on cultural framing, societal impact, and bias mitigation in vision-language models. The event features two keynotes from renowned experts from journalism and computer science, research presentations for six accepted papers, and interactive discussions to explore and discuss cutting-edge methodologies and applications in multimodal human understanding. The workshop proceedings can be found at: https://dl.acm.org/doi/proceedings/10.1145/3728481
Sherzod Hakimov, David Semedo, Eric Müller-Budack, Marc A. Kastner 0001, Takahiro Komamizu
ACM Multimedia4
2025 IntentVC 2025: The ACM Multimedia Grand Challenge on Intention-Oriented Controllable Video Captioning
abstract
The IntentVC Challenge, held in conjunction with ACM Multimedia 2025, introduces a novel benchmark for intention-oriented controllable video captioning. Unlike conventional captioning methods that generate generic, scene-level summaries, IntentVC focuses on intention-specific generation. Participants are required to produce captions explicitly conditioned on user-defined intentions, such as emphasizing a specific object tracked within a video. To support this task, the challenge provides an extended version of the LaSOT dataset annotated with intention-focused captions across 70 object categories. A standardized evaluation protocol and public leaderboard enable fair and reproducible comparison among submitted methods. By advancing research in personalized and adaptive video understanding, IntentVC offers a platform for exploring controllable vision-language modeling with practical relevance for accessibility, retrieval, and human-AI interaction. As a result, a total of 23 teams and 58 active participants have participated, and a total of 1,443 entries have been submitted. More information and resources are available at https://sites.google.com/view/intentvc/.
Takahiro Komamizu, Marc A. Kastner 0001, Yasutomo Kawanishi, Trung Thanh Nguyen 0006, Junan Chen 0004
ACM Multimedia2
2025 Transformer-Based Audio Generation Conditioned by 2D Latent Maps: A Demonstration
Christian Limberg, Marc A. Kastner 0001
MMM (5)3
2025 Quantifying Image-Adjective Associations by Leveraging Large-Scale Pretrained Models
Chihaya Matsuhira, Marc A. Kastner 0001, Takahiro Komamizu, Takatsugu Hirayama, Ichiro Ide
MMM (4)2
2025 Towards Visual Storytelling by Understanding Narrative Context Through Scene-Graphs
Itthisak Phueaksri, Marc A. Kastner 0001, Yasutomo Kawanishi, Takahiro Komamizu, Ichiro Ide
MMM (4)2
2024 MUWS 2024: The 3rd International Workshop on Multimodal Human Understanding for the Web and Social Media
abstract
Multimodal human understanding and analysis are emerging research areas that cut through several disciplines like Computer Vision (CV), Natural Language Processing (NLP), Speech Processing, Human-Computer Interaction (HCI), and Multimedia. Several multimodal learning techniques have recently shown the benefit of combining multiple modalities in image-text, audio-visual and video representation learning and various downstream multimodal tasks. At the core, these methods focus on modelling the modalities and their complex interactions by using large amounts of data, different loss functions and deep neural network architectures. However, for many Web and Social media applications, there is the need to model the human, including the understanding of human behaviour and perception. For this, it becomes important to consider interdisciplinary approaches, including social sciences and psychology. The core is understanding various cross-modal relations, quantifying bias such as social biases, and the applicability of models to real-world problems. Interdisciplinary theories such as semiotics or gestalt psychology can provide additional insights on perceptual understanding through signs and symbols across multiple modalities. In general, these theories provide a compelling view of multimodality and perception that can further expand computational research and multimedia applications on the Web and Social media.
Marc A. Kastner 0001, Gullal Singh Cheema, Sherzod Hakimov, Noa Garcia
ICMR1
2024 Investigating Conceptual Blending of a Diffusion Model for Improving Nonword-to-Image Generation
abstract
Text-to-image diffusion models sometimes depict blended concepts in the generated images. One promising use case of this effect would be the nonword-to-image generation task which attempts to generate images intuitively imaginable from a non-existing word (nonword). To realize nonword-to-image generation, an existing study focused on associating nonwords with similar-sounding words. Since each nonword can have multiple similar-sounding words, generating images containing their blended concepts would increase intuitiveness, facilitating creative activities and promoting computational psycholinguistics. Nevertheless, no existing study has quantitatively evaluated this effect in either diffusion models or the nonword-to-image generation paradigm. Therefore, this paper first analyzes the conceptual blending in a pretrained diffusion model, Stable Diffusion. The analysis reveals that a high percentage of generated images depict blended concepts when inputting an embedding interpolating between the text embeddings of two text prompts referring to different concepts. Next, this paper explores the best text embedding space conversion method of an existing nonword-to-image generation framework to ensure both the occurrence of conceptual blending and image generation quality. We compare the conventional direct prediction approach with the proposed method that combines k-nearest neighbor search and linear regression. Evaluation reveals that the enhanced accuracy of the embedding space conversion by the proposed method improves the image generation quality, while the emergence of conceptual blending could be attributed mainly to the specific dimensions of the high-dimensional text embedding space.
Chihaya Matsuhira, Marc A. Kastner 0001, Takahiro Komamizu, Takatsugu Hirayama, Ichiro Ide
ACM Multimedia2
2024 Computational measurement of perceived pointiness from pronunciation
abstract
Abstract Sound symbolism is a well-researched topic of psycholinguistics, which tries to comprehend the connection between the sound of a word and its meanings. The Bouba-Kiki effect , one form of sound symbolism, claims that people perceive the pronunciation of “Kiki” as pointier than that of “Bouba.” There is no research that focuses on modeling such perception, i.e., how pointy a pronunciation sounds to humans, through computational and data-driven approaches. To address this, this paper first proposes the novel concept of “phonetic pointiness” defined as how pointy a shape humans are most likely to associate with a given pronunciation. We then model this phonetic pointiness from computational and data-driven approaches to calculate a score for an arbitrary pronunciation. There are three proposed models: a referential model, an expressive model, and a combined model, which integrates the previous two. The idea comes from an existing psycholinguistic classification of two types of sound symbolisms: referential symbolism and expressive symbolism , where the former relates to vocabulary knowledge, while the latter is based on pure human intuition. The proposed models are constructed only with image and language data available on the Web, therefore not requiring task-specific human annotations. We evaluate these models through a crowd-sourced user study, finding a promising correlation between human perception and the phonetic pointiness calculated by the proposed models. The results indicate that human perception can be modeled better by combining both types of sound symbolisms. Furthermore, by observing the behaviors of the models, we show several possible use-cases, such as product naming and psycholinguistic research, which can be a useful insight to further studies and applications.
Chihaya Matsuhira, Marc A. Kastner 0001, Takahiro Komamizu, Ichiro Ide, Takatsugu Hirayama, Yasutomo Kawanishi, Keisuke Doman, Daisuke Deguchi
Multim. Tools Appl.2
2024 Correction to: Computational measurement of perceived pointiness from pronunciation
Chihaya Matsuhira, Marc A. Kastner 0001, Takahiro Komamizu, Ichiro Ide, Takatsugu Hirayama, Yasutomo Kawanishi, Keisuke Doman, Daisuke Deguchi
Multim. Tools Appl.2
2023 MUWS'2023: The 2nd International Workshop on Multimodal Human Understanding for the Web and Social Media
abstract
Multimodal human understanding and analysis is an emerging research area that cuts through several disciplines like Computer Vision, Natural Language Processing (NLP), Speech Processing, Human-Computer Interaction, and Multimedia. Several multimodal learning techniques have recently shown the benefit of combining multiple modalities in image-text, audio-visual and video representation learning and various downstream multimodal tasks. At the core, these methods focus on modelling the modalities and their complex interactions by using large amounts of data, different loss functions and deep neural network architectures. However, for many Web and Social media applications, there is the need to model the human, including the understanding of human behaviour and perception. For this, it becomes important to consider interdisciplinary approaches, including social sciences, semiotics and psychology. The core is understanding various cross-modal relations, quantifying bias such as social biases, and the applicability of models to real-world problems. Interdisciplinary theories such as semiotics or gestalt psychology can provide additional insights and analysis on perceptual understanding through signs and symbols via multiple modalities. In general, these theories provide a compelling view of multimodality and perception that can further expand computational research and multimedia applications on the Web and Social media.
Gullal Singh Cheema, Sherzod Hakimov, Marc A. Kastner 0001, Noa Garcia
CIKM3
2023 Towards Captioning an Image Collection from a Combined Scene Graph Representation Approach
Itthisak Phueaksri, Marc A. Kastner 0001, Yasutomo Kawanishi, Takahiro Komamizu, Ichiro Ide
MMM (1)2
2022 Detection of Birds in a 3D Environment Referring to Audio-Visual Information
abstract
We propose a method to detect birds in a 3D environment referring to both audio information observed from a microphone array and visual information observed from a panorama camera. In general, in panorama images, birds appear relatively too small to be detected accurately even with the state-of-the-art deep learning models. Thus, the proposed method takes a two step approach where the birds are first roughly located referring to audio information by Sound Source Localization (SSL), and then image detection is applied within its vicinity. Through evaluation on a dataset annotated with bounding boxes surrounding the birds, we show that the proposed method improves detection performance of birds that appear in relatively small sizes in the image, in both accuracy and processing speed.
Yasutomo Kawanishi, Ichiro Ide, Baidong Chu, Chihaya Matsuhira, Marc A. Kastner 0001, Takahiro Komamizu, Daisuke Deguchi
AVSS5
2022 Personalized Fashion Recommendation Using Pairwise Attention
Donnaphat Trakulwaranont, Marc A. Kastner 0001, Shin'ichi Satoh 0001
MMM (2)2
2022 On Assisting Diagnoses of Pareidolia by Emulating Patient Behavior
Zhaohui Zhu, Marc A. Kastner 0001, Shin'ichi Satoh 0001
MMM (1)2
2021 Reproducibility Companion Paper: Kalman Filter-Based Head Motion Prediction for Cloud-Based Mixed Reality
abstract
In our MM'20 paper,, we presented a Kalman filter-based approach for prediction of head motion in 6DoF. The proposed approach was employed in our cloud-based volumetric video streaming system to reduce the interaction latency experienced by the user. In this companion paper, we present the dataset collected for our experiments and our simulation framework that reproduces the obtained experimental results. Our implementation is freely available on Github to facilitate further research.
Serhan Gul, Sebastian Bosse, Dimitri Podborski, Thomas Schierl, Cornelius Hellge, Marc A. Kastner 0001, Jan Zahálka
ACM Multimedia6
2021 Reproducibility Companion Paper: Describing Subjective Experiment Consistency by p-Value P-P Plot
abstract
In this paper we reproduce experimental results presented in our earlier work titled "Describing Subjective Experiment Consistency by p-Value P-P Plot" that was presented in the course of the 28th ACM International Conference on Multimedia. The paper aims at verifying the soundness of our prior results and helping others understand our software framework. We present artifacts that help reproduce tables, figures and all the data derived from raw subjective responses that were included in our earlier work. Using the artifacts we show that our results are reproducible. We invite everyone to use our software framework for subjective responses analyses going beyond reproducibility efforts.
Jakub Nawala, Lucjan Janowski, Bogdan Cmiel, Krzysztof Rusek, Marc A. Kastner 0001, Jan Zahálka
ACM Multimedia5
2021 Pose-aware Outfit Transfer between Unpaired in-the-wild Fashion Images
abstract
Virtual try-on systems became popular for visualizing outfits, due to the importance of individual fashion in many communities. The objective of such a system is to transfer a piece of clothing to another person while preserving its detail and characteristics. To generate a realistic in-the-wild image, it needs visual optimization of the clothing, background and target person, making this task still very challenging. In this paper, we develop a method that generates realistic try-on images with unpaired images from in-the-wild datasets. Our proposed method starts with generating a mock-up paired image using geometric transfer. Then, the target’s pose information is adjusted using a modified pose-attention module. We combine a reconstruction and a content loss to preserve the detail and style of the transferred clothing, background and the target person. We evaluate the approach on the Fashionpedia dataset and can show a promising performance over a baseline approach.
Donnaphat Trakulwaranont, Marc A. Kastner 0001, Shin'ichi Satoh 0001
MMAsia2
2021 Tell as You Imagine: Sentence Imageability-Aware Image Captioning
Kazuki Umemura, Marc A. Kastner 0001, Ichiro Ide, Yasutomo Kawanishi, Takatsugu Hirayama, Keisuke Doman, Daisuke Deguchi, Hiroshi Murase
MMM (2)2
2020 FashionGraph: Understanding fashion data using scene graph generation
abstract
Fashion analysis is an attractive domain for vision research due to its direct applications in e-commerce contexts. However, fashion datasets are commonly rather demanding, as both objects and attributes tend to be fine-grained and thus result in very long-tailed datasets. Furthermore, relationships between objects and attributes are often dense, but are crucial for the performance of fashion applications. In this paper, we propose to generate scene graphs for existing fashion datasets. By detecting relationships between fashion objects, their parts, and their attributes we gain a better understanding of the scenes. As no current fashion dataset provides scene graphs, we generate relationships between fashion objects from existing annotations. The output is post-processed and filtered to generate a meaningful scene graph for each image. In the experiments, we can show existing applications like image retrieval benefiting from the scene graph understanding. We first evaluate the accuracy of the generated scene graphs. Then, we employ scene graphs to fashion image retrieval in order to showcase their performance in real applications. The results show various benefits for fashion applications by exploiting scene graph knowledge. The resources and model for the proposed method are available on GitHub1.
Shabnam Sadegharmaki, Marc A. Kastner 0001, Shin'ichi Satoh 0001
ICPR2
2020 Imageability Estimation using Visual and Language Features
abstract
Imageability is a concept from Psycholinguistics quantizing the human perception of words. However, existing datasets are created through subjective experiments and are thus very small. Therefore, methods to automatically estimate the imageability can be helpful. For an accurate automatic imageability estimation, we extend the idea of a psychological hypothesis called Dual-Coding Theory, that discusses the connection of our perception towards visual information and language information, and also focus on the relationship between the pronunciation of a word and its imageability. In this research, we propose a method to estimate imageability of words using both visual and language features extracted from corresponding data. For the estimation, we use visual features extracted from low- and high-level image features, and language features extracted from textual features and phonetic features of words. Evaluations show that our proposed method can estimate imageability more accurately than comparative methods, implying the contribution of each feature to the imageability.
Chihaya Matsuhira, Marc A. Kastner 0001, Ichiro Ide, Yasutomo Kawanishi, Takatsugu Hirayama, Keisuke Doman, Daisuke Deguchi, Hiroshi Murase
ICMR2
2020 Browsing Visual Sentiment Datasets Using Psycholinguistic Groundings
Marc A. Kastner 0001, Ichiro Ide, Yasutomo Kawanishi, Takatsugu Hirayama, Daisuke Deguchi, Hiroshi Murase
MMM (2)1
2020 Estimating the imageability of words by mining visual characteristics from crawled image data
Marc A. Kastner 0001, Ichiro Ide, Frank Nack, Yasutomo Kawanishi, Takatsugu Hirayama, Daisuke Deguchi, Hiroshi Murase
Multim. Tools Appl.1
2019 On Quantizing the Mental Image of Concepts for Visual Semantic Analyses
abstract
With the rise of multi-modal applications, the need for better understanding of the relationship between language and vision becomes prominent. While modern applications often consider both text and image, human perception is often only of secondary consideration. In my doctoral studies, I research the quantization of visual differences between concepts regarding human perception. Initially, I looked at local visual differences between concepts and their subordinate concepts, measuring the variety gap between images of, e.g. car and vehicle. In the following study, I applied data-mining on Web-crawled images to estimate psycholinguistics metrics like the imageability of words. In this way, the tendency of low- vs. high-imageability can be estimated on a dictionary-level, defining the gap between words like peace and car. Going forward, I want to create visualization demos to analyze psycholinguistic relationships in image datasets.
Marc A. Kastner 0001
ACM Multimedia1
2019 Estimating the visual variety of concepts by referring to Web popularity
Marc A. Kastner 0001, Ichiro Ide, Yasutomo Kawanishi, Takatsugu Hirayama, Daisuke Deguchi, Hiroshi Murase
Multim. Tools Appl.1
2018 Dependable Non-Volatile Memory
abstract
Recent advances in persistent memory (PM) enable fast, byte-addressable main memory that maintains its state across power cycling events. To survive power outages and prevent inconsistent application state, current approaches introduce persistent logs and require expensive cache flushes. Thus, these solutions cause a performance penalty of up to 10x for write operations on PM. With respect to wear-out effects, and a significantly lower write performance compared to read operations, we identify this as a major flaw that impacts performance and lifetime of PM. In addition, most PM technologies are susceptible to soft-errors that cause corrupted data, which implies a high risk of a permanently inconsistent system state.
Arthur Martens, Rouven Scholz, Phil Lindow, Niklas Lehnfeld, Marc A. Kastner 0001, Rüdiger Kapitza
SYSTOR5