Nikita Srivatsan

dblp:227/3475 · DBLP profile ↗
← Back
5ranked-venue papers
5as first author
3since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 5 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Vision and language · 35% Language models and text generation · 35% Information extraction and text analysis · 30%
Computer graphics and multimedia
3 papers
Visual content generation and editing · 55% Geometric modeling and processing · 31% Audio and music processing · 14%
Databases, data mining, and information retrieval
2 papers
Information retrieval · 70% Web and social media mining · 30%
Human-computer interaction and pervasive computing
1 paper
Accessibility and assistive technology · 100%

Topics — the 8 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language
image captioning
0.812024
Alt-Text with Context: Improving Accessibility for Images on Twitter · ICLR 2024
Accessibility and assistive technology › image accessibility
alt-text generation
0.812024
Alt-Text with Context: Improving Accessibility for Images on Twitter · ICLR 2024
Geometric modeling and processing
shape analysis
0.512021
Scalable Font Reconstruction with Dual Latent Manifolds · EMNLP (1) 2021
Visual content generation and editing
font generation
0.412019
A Deep Factorization of Style and Structure in Fonts · EMNLP/IJCNLP (1) 2019
Natural language and speech › Information extraction and text analysis › topic model
neural topic model
0.312018
Modeling Online Discourse with Coupled Distributed Topics · EMNLP 2018
Natural language and speech › Information extraction and text analysis
topic model
0.312018
Modeling Online Discourse with Coupled Distributed Topics · EMNLP 2018
Web and social media mining › social media analysis
online discourse analysis
0.312018
Modeling Online Discourse with Coupled Distributed Topics · EMNLP 2018
Audio and music processing › music information retrieval
music understanding
0.212024
Retrieval Guided Music Captioning via Multimodal Prefixes · IJCAI 2024

Methods — techniques the papers use, named apart from their topics

retrieval · 2.3multimodal prefixes · 2.3deep generative model · 1.5multimodal model · 1.5BLEU@4 · 1.5dual latent manifolds · 0.5factorization · 0.4mean-field inference · 0.3mean field inference · 0.3
YearPublicationVenuePosition
2024 Alt-Text with Context: Improving Accessibility for Images on Twitter
abstract
In this work we present an approach for generating alternative text (or alt-text) descriptions for images shared on social media, specifically Twitter. More than just a special case of image captioning, alt-text is both more literally descriptive and context-specific. Also critically, images posted to Twitter are often accompanied by user-written text that despite not necessarily describing the image may provide useful context that if properly leveraged can be informative. We address this task with a multimodal model that conditions on both textual information from the associated social media post as well as visual signal from the image, and demonstrate that the utility of these two information sources stacks. We put forward a new dataset of 371k images paired with alt-text and tweets scraped from Twitter and evaluate on it across a variety of automated metrics as well as human evaluation. We show that our approach of conditioning on both tweet text and visual information significantly outperforms prior work, by more than 2x on BLEU@4.
Nikita Srivatsan, Sofía Samaniego, Omar Florez, Taylor Berg-Kirkpatrick
ICLR1
2024 Retrieval Guided Music Captioning via Multimodal Prefixes
Nikita Srivatsan, Ke Chen 0021, Shlomo Dubnov, Taylor Berg-Kirkpatrick
IJCAI1
2021 Scalable Font Reconstruction with Dual Latent Manifolds
abstract
We propose a deep generative model that performs typography analysis and font reconstruction by learning disentangled manifolds of both font style and character shape.Our approach enables us to massively scale up the number of character types we can effectively model compared to previous methods.Specifically, we infer separate latent variables representing character and font via a pair of inference networks which take as input sets of glyphs that either all share a character type, or belong to the same font.This design allows our model to generalize to characters that were not observed during training time, an important task in light of the relative sparsity of most fonts.We also put forward a new loss, adapted from prior work that measures likelihood using an adaptive distribution in a projected space, resulting in more natural images without requiring a discriminator.We evaluate on the task of font reconstruction over various datasets representing character types of many languages, and compare favorably to modern style transfer systems according to both automatic and manually-evaluated metrics.
Nikita Srivatsan, Jonathan T. Barron, Taylor Berg-Kirkpatrick
EMNLP (1)1
2019 A Deep Factorization of Style and Structure in Fonts
abstract
Nikita Srivatsan, Jonathan Barron, Dan Klein, Taylor Berg-Kirkpatrick. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Nikita Srivatsan, Jonathan T. Barron, Daniel Klein 0001, Taylor Berg-Kirkpatrick
EMNLP/IJCNLP (1)1
2018 Modeling Online Discourse with Coupled Distributed Topics
abstract
In this paper, we propose a deep, globally normalized topic model that incorporates structural relationships connecting documents in socially generated corpora, such as online forums.Our model (1) captures discursive interactions along observed reply links in addition to traditional topic information, and (2) incorporates latent distributed representations arranged in a deep architecture, which enables a GPU-based mean-field inference procedure that scales efficiently to large data.We apply our model to a new social media dataset consisting of 13M comments mined from the popular internet forum Reddit, a domain that poses significant challenges to models that do not account for relationships connecting user comments.We evaluate against existing methods across multiple metrics including perplexity and metadata prediction, and qualitatively analyze the learned interaction patterns.
Nikita Srivatsan, Zachary Wojtowicz, Taylor Berg-Kirkpatrick
EMNLP1