Deepali Aneja

dblp:165/9425 · DBLP profile ↗
← Back
16ranked-venue papers
6as first author
9since 2021 · last 2025
0000-0001-9610-5648ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 8 · 4 first-author · 4 since 2021Artificial intelligence and machine learning · 7 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2025 2D Neural Fields with Learned Discontinuities
abstract
Abstract Effective representation of 2D images is fundamental in digital image processing, where traditional methods like raster and vector graphics struggle with sharpness and textural complexity, respectively. Current neural fields offer high fidelity and resolution independence but require predefined meshes with known discontinuities, restricting their utility. We observe that by treating all mesh edges as potential discontinuities, we can represent the discontinuity magnitudes as continuous variables and optimize. We further introduce a novel discontinuous neural field model that jointly approximates the target image and recovers discontinuities. Through systematic evaluations, our neural field outperforms other methods that fit unknown discontinuities with discontinuous representations, exceeding Field of Junction and Boundary Attention by over 11dB in both denoising and super‐resolution tasks and achieving 3.5× smaller Chamfer distances than Mumford–Shah‐based methods. It also surpasses InstantNGP with improvements of more than 5dB (denoising) and 10dB (super‐resolution). Additionally, our approach shows remarkable capability in approximating complex artistic and natural images and cleaning up diffusion‐generated depth maps.
Chenxi Liu 0004, Siqi Wang 0003, Matthew Fisher, Deepali Aneja, Alec Jacobson
Comput. Graph. Forum4
2024 How do video content creation goals impact which concepts people prioritize for generating B-roll imagery?
abstract
B-roll is vital when producing high-quality videos, but finding the right images can be difficult and time-consuming. Moreover, what B-roll is most effective can depend on a video content creator’s intent—is the goal to entertain, to inform, or something else? While new text-to-image generation models provide promising avenues for streamlining B-roll production, it remains unclear how these tools can provide support for content creators with different goals. To close this gap, we aimed to understand how video content creator’s goals guide which visual concepts they prioritize for B-roll generation. Here we introduce a benchmark containing judgments from > 800 people as to which terms in 12 video transcripts should be assigned highest priority for B-roll imagery accompaniment. We verified that participants reliably prioritized different visual concepts depending on whether their goal was help produce informative or entertaining videos. We next explored how well several algorithms, including heuristic approaches and large language models (LLMs), could predict systematic patterns in human judgments. We found that none of these methods fully captured human judgments in either goal condition, with state-of-the-art LLMs (i.e., GPT-4) even underperforming a baseline that sampled only nouns or nouns and adjectives. Overall, our work identifies opportunities to develop improved algorithms to support video production workflows.
Holly Huey, Mackenzie Leake, Deepali Aneja, Matthew Fisher, Judith E. Fan
Creativity & Cognition3
2024 Elastica: Adaptive Live Augmented Presentations with Elastic Mappings Across Modalities
abstract
Augmented presentations offer compelling storytelling by combining speech content, gestural performance, and animated graphics in a congruent manner. The expressiveness of these presentations stems from the harmonious coordination of spoken words and graphic elements, complemented by smooth animations aligned with the presenter’s gestures. However, achieving such desired congruence in a live presentation poses significant challenges due to the unpredictability and imprecision inherent in presenters’ real-time actions. Existing methods either leveraged rigid mapping without predefined states or required the presenters to conform to predefined animations. We introduce adaptive presentations that dynamically adjust predefined graphic animations to real-time speech and gestures. Our approach leverages script following and motion warping to establish elastic mappings that generate runtime graphic parameters coordinating speech, gesture, and predefined animation state. Our evaluation demonstrated that the proposed adaptive presentation can effectively mitigate undesired visual artifacts caused by performance deviations and enhance the expressiveness of resulting presentations.
Yining Cao, Rubaiat Habib Kazi, Li-Yi Wei, Deepali Aneja, Haijun Xia
CHI4
2024 A Closer Look at the Limitations of Instruction Tuning
abstract
Instruction Tuning (IT), the process of training large language models (LLMs) using instruction-response pairs, has emerged as the predominant method for transforming base pre-trained LLMs into open-domain conversational agents. While IT has achieved notable success and widespread adoption, its limitations and shortcomings remain underexplored. In this paper, through rigorous experiments and an in-depth analysis of the changes LLMs undergo through IT, we reveal various limitations of IT. In particular, we show that (1) IT fails to enhance knowledge or skills in LLMs. LoRA fine-tuning is limited to learning response initiation and style tokens, and full-parameter fine-tuning leads to knowledge degradation. (2) Copying response patterns from IT datasets derived from knowledgeable sources leads to a decline in response quality. (3) Full-parameter fine-tuning increases hallucination by inaccurately borrowing tokens from conceptually similar instances in the IT dataset for generating responses. (4) Popular methods to improve IT do not lead to performance improvements over a simple LoRA fine-tuned model. Our findings reveal that responses generated solely from pre-trained knowledge consistently outperform responses by models that learn any form of new knowledge from IT on open-source datasets. We hope the insights and challenges revealed in this paper inspire future work in related directions.
Sreyan Ghosh, Chandra Kiran Reddy Evuru, Sonal Kumar, Ramaneswaran S., Deepali Aneja, Zeyu Jin, Ramani Duraiswami, Dinesh Manocha
ICML5
2024 Deep Sketch Vectorization via Implicit Surface Extraction
abstract
We introduce an algorithm for sketch vectorization with state-of-the-art accuracy and capable of handling complex sketches. We approach sketch vectorization as a surface extraction task from an unsigned distance field, which is implemented using a two-stage neural network and a dual contouring domain post processing algorithm. The first stage consists of extracting unsigned distance fields from an input raster image. The second stage consists of an improved neural dual contouring network more robust to noisy input and more sensitive to line geometry. To address the issue of under-sampling inherent in grid-based surface extraction approaches, we explicitly predict undersampling and keypoint maps. These are used in our post-processing algorithm to resolve sharp features and multi-way junctions. The keypoint and undersampling maps are naturally controllable, which we demonstrate in an interactive topology refinement interface. Our proposed approach produces far more accurate vectorizations on complex input than previous approaches with efficient running time.
Chuan Yan, Deepali Aneja, Matthew Fisher, Edgar Simo-Serra, Yotam I. Gingold
ACM Trans. Graph.3
2023 SoundToons: Exemplar-Based Authoring of Interactive Audio-Driven Animation Sprites
abstract
Animations can come to life when they are synchronized with relevant sounds. Yet, synchronizing animations to audio requires tedious key-framing or programming, which is difficult for novice creators. There are existing tools that support audio-driven live animation, but they focus primarily on speech and have little or no support for non-speech sounds. We present SoundToons, an exemplar-based authoring tool for interactive, audio-driven animation focusing on non-speech sounds. Our tool enables novice creators to author live animations to a wide variety of non-speech sounds, such as clapping and instrumental music. We support two types of audio interactions: (1) discrete interaction, which triggers animations when a discrete sound event is detected, and (2) continuous, which synchronizes an animation to continuous audio parameters. By employing an exemplar-based iterative authoring approach, we empower novice creators to design and quickly refine interactive animations. User evaluations demonstrate that novice users can author and perform live audio-driven animation intuitively. Moreover, compared to other input modalities such as trackpads or foot pedals, users preferred using audio as an intuitive way to drive animation.
Toby Chong, Hijung Shin, Deepali Aneja, Takeo Igarashi
IUI3
2022 APES: Articulated Part Extraction from Sprite Sheets
abstract
Rigged puppets are one of the most prevalent representations to create 2D character animations. Creating these puppets requires partitioning characters into independently moving parts. In this work, we present a method to automatically identify such articulated parts from a small set of character poses shown in a sprite sheet, which is an illustration of the character that artists often draw before puppet creation. Our method is trained to infer articulated parts, e.g. head, torso and limbs, that can be reassembled to best reconstruct the given poses. Our results demonstrate significantly better performance than alternatives qualitatively and quantitatively. Our project page https://zhan-xu.github.io/parts/ includes our code and data.
Matthew Fisher, Yang Zhou 0009, Deepali Aneja, Rushikesh Dudhat, Li Yi 0001, Evangelos Kalogerakis
CVPR4
2022 Audio-driven Neural Gesture Reenactment with Video Motion Graphs
abstract
Human speech is often accompanied by body gestures including arm and hand gestures. We present a method that reenacts a high-quality video with gestures matching a target speech audio. The key idea of our method is to split and re-assemble clips from a reference video through a novel video motion graph encoding valid transitions between clips. To seamlessly connect different clips in the reenactment, we propose a pose-aware video blending network which synthesizes video frames around the stitched frames between two clips. Moreover, we developed an audio-based gesture searching algorithm to find the optimal order of the reenacted frames. Our system generates reen-actments that are consistent with both the audio rhythms and the speech content. We evaluate our synthesized video quality quantitatively, qualitatively, and with user studies, demonstrating that our method produces videos of much higher quality and consistency with the target audio compared to previous work and baselines. Our project page https://github.com/yzhou359/vid-reenact includes code and data.
Yang Zhou 0009, Jimei Yang, Dingzeyu Li, Jun Saito, Deepali Aneja, Evangelos Kalogerakis
CVPR5
2021 Understanding Conversational and Expressive Style in a Multimodal Embodied Conversational Agent
abstract
Embodied conversational agents have changed the ways we can interact with machines. However, these systems often do not meet users’ expectations. A limitation is that the agents are monotonic in behavior and do not adapt to an interlocutor. We present SIVA (a Socially Intelligent Virtual Agent), an expressive, embodied conversational agent that can recognize human behavior during open-ended conversations and automatically align its responses to the conversational and expressive style of the other party. SIVA leverages multimodal inputs to produce rich and perceptually valid responses (lip syncing and facial expressions) during the conversation. We conducted a user study (N=30) in which participants rated SIVA as being more empathetic and believable than the control (agent without style matching). Based on almost 10 hours of interaction, participants who preferred interpersonal involvement evaluated SIVA as significantly more animate than the participants who valued consideration and independence.
Deepali Aneja, Jessie Hoegen, Daniel McDuff, Mary Czerwinski
CHI1
2020 Conversational Error Analysis in Human-Agent Interaction
abstract
Conversational Agents (CAs) present many opportunities for changing how we interact with information and computer systems in a more natural, accessible way. Building on research in machine learning and HCI, it is now possible to design and test multi-turn CA that is capable of extended interactions. However, there are many ways in which these CAs can "fail" and fall short of human expectations. We systematically analyzed how five different types of conversational errors impacted perceptions of an embodied CA. Not all errors negatively impacted perceptions of the agent. Repetitions by the agent and clarifications by the human significantly decreased the perceived intelligence and anthropomorphism of the agent. Turn-taking errors significantly decreased the likability of the agent. However, coherence errors significantly positively increased likability, and these errors were also associated with positive valence via facial expressions, suggesting that the users found them amusing. We believe this work is the first to identify that turn-taking, repetition, clarification, and coherence errors directly affect users' acceptance of an embodied CA, and are worth taking note by designers of such systems during dialog configuration. We release the Agent Conversational Error (ACE) dataset, a set of transcripts and error annotations of human-agent conversations. The dataset can be found at the GITHUB link: https://github.com/deepalianeja/ACE-dataset
Deepali Aneja, Daniel McDuff, Mary Czerwinski
IVA1
2019 A Facial Affect Analysis System for Autism Spectrum Disorder
abstract
In this paper, we introduce an end-to-end machine learning-based system for classifying autism spectrum disorder (ASD) using facial attributes such as expressions, action units, arousal, and valence. Our system classifies ASD using representations of different facial attributes from convolutional neural networks, which are trained on images in the wild. Our experimental results show that different facial attributes used in our system are statistically significant and improve sensitivity, specificity, and F1 score of ASD classification by a large margin. In particular, the addition of different facial attributes improves the performance of ASD classification by about 7% which achieves a F1 score of 76%.
Beibin Li, Sachin Mehta, Deepali Aneja, Claire E. Foster, Pamela Ventola, Frédérick Shic, Linda G. Shapiro
ICIP3
2019 A High-Fidelity Open Embodied Avatar with Lip Syncing and Expression Capabilities
abstract
Embodied avatars as virtual agents have many applications and provide benefits over disembodied agents, allowing nonverbal social and interactional cues to be leveraged, in a similar manner to how humans interact with each other. We present an open embodied avatar built upon the Unreal Engine that can be controlled via a simple python programming interface. The avatar has lip syncing (phoneme control), head gesture and facial expression (using either facial action units or cardinal emotion categories) capabilities. We release code and models to illustrate how the avatar can be controlled like a puppet or used to create a simple conversational agent using public application programming interfaces (APIs). GITHUB link: https://github.com/danmcduff/AvatarSim
Deepali Aneja, Daniel McDuff, Shital Shah
ICMI1
2019 An End-to-End Conversational Style Matching Agent
abstract
We present an end-to-end voice-based conversational agent that is able to engage in naturalistic multi-turn dialogue and align with the interlocutor's conversational style. The system uses a series of deep neural network components for speech recognition, dialogue generation, prosodic analysis and speech synthesis to generate language and prosodic expression with qualities that match those of the user. We conducted a user study (N=30) in which participants talked with the agent for 15 to 20 minutes, resulting in over 8 hours of natural interaction data. Users with high consideration conversational styles reported the agent to be more trustworthy when it matched their conversational style. Whereas, users with high involvement conversational styles were indifferent. Finally, we provide design guidelines for multi-turn dialogue interactions using conversational style adaptation.
Jessie Hoegen, Deepali Aneja, Daniel McDuff, Mary Czerwinski
IVA2
2018 Learning to Generate 3D Stylized Character Expressions from Humans
abstract
We present ExprGen, a system to automatically generate 3D stylized character expressions from humans in a perceptually valid and geometrically consistent manner. Our multi-stage deep learning system utilizes the latent variables of human and character expression recognition convolutional neural networks to control a 3D animated character rig. This end-to-end system takes images of human faces and generates the character rig parameters that best match the human's facial expression. ExprGen generalizes to multiple characters, and allows expression transfer between characters in a semi-supervised manner. Qualitative and quantitative evaluation of our method based on Mechanical Turk tests show the high perceptual accuracy of our expression transfer results.
Deepali Aneja, Bindita Chaudhuri, Alex Colburn, Gary Faigin, Linda G. Shapiro, Barbara Mones
WACV1
2016 Modeling Stylized Character Expressions via Deep Learning
Deepali Aneja, Alex Colburn, Gary Faigin, Linda G. Shapiro, Barbara Mones
ACCV (2)1
2015 Automated Detection of 3D Landmarks for the Elimination of Non-Biological Variation in Geometric Morphometric Analyses
abstract
Landmark-based morphometric analyses are used by anthropologists, developmental and evolutionary biologists to understand shape and size differences (eg. in the cranioskeleton) between groups of specimens. The standard, labor intensive approach is for researchers to manually place landmarks on 3D image datasets. As landmark recognition is subject to inaccuracies of human perception, digitization of landmark coordinates is typically repeated (often by more than one person) and the mean coordinates are used. In an attempt to improve efficiency and reproducibility between researchers, we have developed an algorithm to locate landmarks on CT mouse hemi-mandible data. The method is evaluated on 3D meshes of 28-day old mice, and results compared to landmarks manually identified by experts. Quantitative shape comparison between two inbred mouse strains demonstrate that data obtained using our algorithm also has enhanced statistical power when compared to data obtained by manual landmarking.
Deepali Aneja, Siddharth R. Vora, Esra D. Camci, Linda G. Shapiro, Timothy C. Cox
CBMS1