Abhinaba Roy

dblp:199/9373 · DBLP profile ↗
← Back
9ranked-venue papers
6as first author
4since 2021 · last 2026
0000-0002-3290-3322ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 4 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Generative modeling · 42% Reinforcement learning · 30% Language models and text generation · 15%
Computer graphics and multimedia
2 papers
Audio and music processing · 100%

Topics — the 10 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
music generation
1.012026
Aligning Generative Music AI with Human Preferences: Methods and Challenges · AAAI 2026
Natural language and speech › Language models and text generation › alignment
preference alignment
1.012026
Aligning Generative Music AI with Human Preferences: Methods and Challenges · AAAI 2026
Machine learning › Reinforcement learning
preference learning
1.012026
Aligning Generative Music AI with Human Preferences: Methods and Challenges · AAAI 2026
Machine learning › Reinforcement learning
reinforcement learning from human feedback
1.012026
Aligning Generative Music AI with Human Preferences: Methods and Challenges · AAAI 2026
Machine learning › Generative modeling
autoregressive model
0.912025
Text2midi: Generating Symbolic Music from Captions · AAAI 2025
Machine learning › Generative modeling › music generation
conditional music generation
0.912025
Text2midi: Generating Symbolic Music from Captions · AAAI 2025
Audio and music processing › music generation
symbolic music generation
0.912025
Text2midi: Generating Symbolic Music from Captions · AAAI 2025
Knowledge, reasoning and agents › Knowledge representation and reasoning
commonsense reasoning
0.412020
KinGDOM: Knowledge-Guided DOMain Adaptation for Sentiment Analysis · ACL 2020
Natural language and speech › Information extraction and text analysis › sentiment analysis › sentiment classification
cross-domain sentiment classification
0.412020
KinGDOM: Knowledge-Guided DOMain Adaptation for Sentiment Analysis · ACL 2020
Audio and music processing
music generation
0.312026
Aligning Generative Music AI with Human Preferences: Methods and Challenges · AAAI 2026

Methods — techniques the papers use, named apart from their topics

preference optimization · 2.0inference-time optimization · 2.0large language model · 1.7autoregressive transformer decoder · 1.7graph convolutional autoencoder · 0.9domain-adversarial training · 0.4domain adversarial training · 0.4
YearPublicationVenuePosition
2026 Aligning Generative Music AI with Human Preferences: Methods and Challenges
abstract
Recent advances in generative AI for music have achieved remarkable fidelity and stylistic diversity, yet these systems often fail to align with nuanced human preferences due to the specific loss functions they use. This paper advocates for the systematic application of preference alignment techniques to music generation, addressing the fundamental gap between computational optimization and human musical appreciation. Drawing on recent breakthroughs including MusicRL's large-scale preference learning, multi-preference alignment frameworks like diffusion-based preference optimization in DiffRhythm+, and inference-time optimization techniques like Text2midi-InferAlign, we discuss how these techniques can address music's unique challenges: temporal coherence, harmonic consistency, and subjective quality assessment. We identify key research challenges including scalability to long-form compositions, reliability amongst others in preference modelling. Looking forward, we envision preference-aligned music generation enabling transformative applications in interactive composition tools and personalized music services. This work calls for sustained interdisciplinary research combining advances in machine learning, music-theory to create music AI systems that truly serve human creative and experiential needs.
Dorien Herremans, Abhinaba Roy
AAAI2
2025 Text2midi: Generating Symbolic Music from Captions
abstract
This paper introduces text2midi, an end-to-end model to generate MIDI files from textual descriptions. Leveraging the growing popularity of multimodal generative approaches, text2midi capitalizes on the extensive availability of textual data and the success of large language models (LLMs). Our end-to-end system harnesses the power of LLMs to generate symbolic music in the form of MIDI files. Specifically, we utilize a pretrained LLM encoder to process captions, which then condition an autoregressive transformer decoder to produce MIDI sequences that accurately reflect the provided descriptions. This intuitive and user-friendly method significantly streamlines the music creation process by allowing users to generate music pieces using text prompts. We conduct comprehensive empirical evaluations, incorporating both automated and human studies, that show our model generates MIDI files of high quality that are indeed controllable by text captions that may include music theory terms such as chords, keys, and tempo.
Keshav Bhandari, Abhinaba Roy, Kyra Wang, Geeta Puri, Simon Colton, Dorien Herremans
AAAI2
2025 JamendoMaxCaps: A Large Scale Music-caption Dataset with Imputed Metadata
abstract
We introduce JamendoMaxCaps, a large-scale music-caption dataset featuring over 362,000 freely licensed instrumental tracks from the renowned Jamendo platform. The dataset includes captions generated by a state-of-the-art captioning model, enhanced with imputed metadata. We also introduce a retrieval system that leverages both musical features and metadata to identify similar songs, which are then used to fill in missing metadata using a local large language model (LLLM). This approach allows us to provide a more comprehensive and informative dataset for researchers working on music-language understanding tasks. We validate this approach quantitatively with five different measurements. By making the JamendoMaxCaps dataset publicly available, we provide a high-quality resource to advance research in music-language understanding tasks such as music retrieval, multimodal representation learning, and generative music models.
Abhinaba Roy, Renhang Liu, Tongyu Lu, Dorien Herremans
IJCNN1
2022 Soft labeling constraint for generalizing from sentiments in single domain
Abhinaba Roy, Erik Cambria
Knowl. Based Syst.1
2020 KinGDOM: Knowledge-Guided DOMain Adaptation for Sentiment Analysis
abstract
Cross-domain sentiment analysis has received significant attention in recent years, prompted by the need to combat the domain gap between different applications that make use of sentiment analysis.In this paper, we take a novel perspective on this task by exploring the role of external commonsense knowledge.We introduce a new framework, KinGDOM, which utilizes the ConceptNet knowledge graph to enrich the semantics of a document by providing both domain-specific and domain-general background concepts.These concepts are learned by training a graph convolutional autoencoder that leverages inter-domain concepts in a domain-invariant manner.Conditioning a popular domain-adversarial baseline method with these learned concepts helps improve its performance over state-of-the-art approaches, demonstrating the efficacy of our proposed framework.
Deepanway Ghosal, Devamanyu Hazarika, Abhinaba Roy, Navonil Majumder, Rada Mihalcea, Soujanya Poria
ACL3
2018 Visually-Driven Semantic Augmentation for Zero-Shot Learning
Abhinaba Roy, Jacopo Cavazza, Vittorio Murino
BMVC1
2018 Discriminative Latent Visual Space For Zero-Shot Object Classification
abstract
In this paper We deal with the problem of zero-shot visual recognition. The standard zero-shot learning (ZSL) pipeline is based on the idea of learning a functional mapping from a visual embedding space to an auxiliary semantic space for a set of seen categories. In the testing phase, the task is to recognize a set of novel categories which are semantically linked to the already known ones. Although such a pipeline is inherently supervised, there exists very few endeavours in the context of ZSL that enforce discrimination in learning this mapping. In this work, we propose a novel encoder-decoder network to explore the possibility of learning an intermediate latent space for the visual features, which is deemed to be simultaneously reconstructive and discriminative. By reaching a trade-off between the joint (re)construction of the visual and the semantic embedding spaces, while ensuring separability among the known classes, the proposed model better generalizes to the unknown categories. Experimental results obtained on challenging datasets, such as AwA, CUB, and ImageNet-2, establish the efficacy of such a discriminative latent space for the standard ZSL setup.
Abhinaba Roy, Biplab Banerjee, Vittorio Murino
ICPR1
2018 Discriminative body part interaction mining for mid-level action representation and classification
Abhinaba Roy, Biplab Banerjee, Vittorio Murino
J. Vis. Commun. Image Represent.1
2017 A Novel Dictionary Learning based Multiple Instance Learning Approach to Action Recognition from Videos
abstract
In this paper we deal with the problem of action recognition from unconstrained videos under the notion of multiple instance learning (MIL).The traditional MIL paradigm considers the data items as bags of instances with the constraint that the positive bags contain some class-specific instances whereas the negative bags consist of instances only from negative classes.A classifier is then further constructed using the bag level annotations and a distance metric between the bags.However, such an approach is not robust to outliers and is time consuming for a moderately large dataset.In contrast, we propose a dictionary learning based strategy to MIL which first identifies class-specific discriminative codewords, and then projects the bag-level instances into a probabilistic embedding space with respect to the selected codewords.This essentially generates a fixedlength vector representation of the bags which is specifically dominated by the properties of the class-specific instances.We introduce a novel exhaustive search strategy using a support vector machine classifier in order to highlight the class-specific codewords.The standard multiclass classification pipeline is followed henceforth in the new embedded feature space for the sake of action recognition.We validate the proposed framework on the challenging KTH and Weizmann datasets, and the results obtained are promising and comparable to representative techniques from the literature.
Abhinaba Roy, Biplab Banerjee, Vittorio Murino
ICPRAM1