VLDB 2026 Research / reviewers in the wild / expert
Yulia Kumar
dblp:289/2262
· DBLP profile ↗
11ranked-venue papers
2as first author
11since 2021 · last 2026
0000-0002-7621-2734ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Data Compression for AI Model TrainingabstractTraining AI models requires ingesting massive amounts of training data. We define Parametric Matching (PM), a grammar-based compression technique that parses structured text into abstract syntax trees, replaces format patterns with compact tokens, and moves numerical parameters into homogeneous streams of tokens that can additionally be bit compressed using LZMA. The results are input streams that are equivalent to the original interleaved text and data in formats such as SVG, 3D-models such as OBJ, PDF, but highly compressed and much more amenable to processing. PM has many advantages. Loading data is much faster (a factor of 20 for 3D models, and 100 to 200 times faster to load). Just as important, by preprocessing the semantics, ambiguous cases can be resolved once, and can be turned into embedding vectors more efficiently, and potentially more accurately. Our experiments with SVG, G-Code, OBJ and PDF files show large compression. The results in Table 1 show that the PM algorithm can create data objects that are not only more compressible but far faster to load because the binary format can be instantly used, in this case to load OpenGL and render without parsing ASCII data. When compressing a binary object that is half the size of the original ASCII representation, by definition LZMA takes half the time to compress because there are fewer bytes to process. Dov Kruger, Yulia Kumar, J. Jenny Li 0001 |
DCC | 2 |
| 2025 | Real-Time Object Detection and Skeletonization for Motion Prediction in Video StreamingabstractThe increasing demand for real-time analysis in video streaming has driven significant advancements in object detection and motion prediction. This paper presents SkelAI, an innovative application that combines YOLOv8, OpenCV, OpenAI API, and our own innovative algorithms to achieve real-time object detection and medial axis skeletonization tailored explicitly for live video streaming environments. In addition, SkelAI integrates AI-generated image capabilities through the DALL-E 3 model, enabling the extraction of skeletons from synthetic content that simulates streaming scenarios. The application supports exporting skeleton data in PyTorch-compatible formats, facilitating the training of sequence predicting deep learning models. Comprehensive evaluations demonstrate SkelAI’s enhanced accuracy, efficiency, and versatility compared to existing tools, underscoring its potential applications in digital animation, biomechanical research and robotics, human-computer interaction, and video compression within streaming platforms. Gavin Wong, Yulia Kumar, J. Jenny Li 0001, Dov Kruger |
AAAI | 2 |
| 2025 | Parametric Matching for Improved Data CompressionabstractModern general-purpose compressors can compress a wide variety of files but do not achieve high compression ratios on files that contain short sequences of delimiters with interleaved numeric data and generally with interleaved data where each sequence is not well correlated to the previous bytes. We demonstrate Parametric Matching (PM), which vastly improves the compression of various structured languages, including PDF, SVG, and G-code files. By de-interleaving and coalescing delimiters and storing data as delta-encoded, discretized binary, compressions of a factor of 10 or more are possible. A Python prototype compresses files to a binary representation, which is then compressed using Lempel-Ziv-Markov (LZMA) to efficiently store the binary tokens in a minimal number of bits. Table 1 shows a ratio of 6 for PDF files containing only text, which are first parsed, and recompressed using PM. For SVG, we demonstrate a factor of 8 to 10 for files including a randomized spiral and a US county map. For the G-code, we compressed the Statue of Liberty, demonstrating that even when the layers are different, a high degree of compression can be achieved. Times are all less than 250ms, even in our Python prototype. Dov Kruger, Yulia Kumar, J. Jenny Li 0001 |
DCC | 2 |
| 2024 | Adversarial Testing of LLMs Across Multiple LanguagesabstractThis study builds on prior research in the field of jailbreaking, focusing on the vulnerabilities of state-of-the-art Large Language Models (LLMs) and their associated chatbots. The primary objective is to evaluate these vulnerabilities through a multilingual lens, extending the scope of previous monolingual studies. The main chatbots under examination include ChatGPT (both legacy ChatGPT -3.5 and the latest ChatGPT -40 model), Gemini, Microsoft Copilot, and Perplexity. Researchers conducted multilingual adversarial attacks, facilitating cross-language and cross-model comparisons, and explored different modalities, such as text versus speech inputs. The findings reveal significant weaknesses in these major models, particularly in their susceptibility to adversarial attacks conducted in languages such as Spanish, Russian, and Traditional Chinese. Given the global proliferation and accessibility of these models, it is imperative to rigorously assess the robustness of LLMs against adversarial inputs across multiple languages and modalities. Yulia Kumar, C. Paredes, J. Jenny Li 0001, Patricia Morreale |
ISNCC | 1 |
| 2024 | An AI-Powered Digital Foundation Recommender SystemabstractThis paper addresses the significant challenge of selecting suitable foundation shades, particularly for darker skin tones, which have been inadequately represented and catered to in the beauty industry. It introduces an innovative system designed to offer individualized makeup recommendations, significantly diminishing the time and effort traditionally demanded of consumers when searching for appropriate products. Utilizing machine learning methodologies and computer vision techniques, the system analyzes image data to precisely identify a range of skin tones, enabling it to propose compatible foundation shades automatically. Unlike sophisticated Large Language Models, such as ChatGPT, Gemini, Microsoft Copilot, or Claude, which are not equipped to undertake such visually driven tasks due to ethical guidelines that prohibit them from processing personal images without explicit consent and the absence of face image processing capabilities, this technological development represents a step forward in fostering inclusivity, illustrating the transformative potential of AI in accommodating the unique beauty preferences of all individuals. Dahana Moz-Ruiz, A. Watson, Yulia Kumar, J. Jenny Li 0001, Patricia Morreale |
ISNCC | 3 |
| 2023 | AssureAIDoctor- A Bias-Free AI BotabstractThe researchers introduce the AssureAIDoctor (AAID) App - a pioneering application that aims to revolutionize healthcare by integrating the latest artificial intelligence (AI) features into a mobile-native product. The app leverages the OpenAI API to simulate virtual doctor-patient interactions, offering users potential remedies for their symptoms. The distinguishing feature of AAID is its use of DALL-E, an advanced image generator, and the forthcoming OpenAI's Code interpreter. This allows users to enhance their interactions with the AI by uploading images, thereby personalizing their healthcare experience. The app's user interface is designed to support this advanced functionality. Preliminary tests show promising results, with AAID accurately responding to various symptom inputs. While scalability is a key focus, the app addresses potential challenges such as increased operational costs associated with Microsoft Azure AI cloud and OpenAI API services. Despite these challenges, AAID is committed to making healthcare accessible for underrepresented and uninsured individuals. The app embodies the potential of AI in healthcare, promising to make healthcare more equitable and accessible. Yulia Kumar, Justin Delgado, E. Kupershtein, Brendan Hannon, Zachary Gordon, J. Jenny Li 0001, Patricia Morreale |
ISNCC | 1 |
| 2023 | Implementing Inclusive Software Design in the CS Curriculum
Pankati Patel, Jean Chu, Yulia Kumar, Daehan Kwak, Patricia Morreale, Rosalinda Garcia, Margaret M. Burnett |
SIGCSE (2) | 3 |
| 2023 | Embedding Equitable Design in the CS Computing CurriculaabstractComputer science (CS) students' curricula is heavily focused on technical skills, and CS ethics, usability, equity, and people/society considerations are not well-integrated into the CS curriculum. If these topics are introduced, they are disconnected from the core courses. As a result, students do not learn to incorporate inclusive practices into their software designs. Thus, students create software through the perspective of a computer scientist - when the important perspective is that of the intended users. As a result, students entering the workforce are inclined to design software that is non-inclusive. We propose the integration of inclusive design in the undergraduate curriculum will result in students creating inclusive software. This research, based on the foundations of inclusive design methods, investigates a new approach to teaching CS. Inclusive software design is embedded into computing courses for all four years of the undergraduate CS curriculum. This new approach is "minimally invasive", occupying very little classroom time, instead it is integrated into the course work that is already assigned. With this work, we hope to answer the following questions: (1) Will this new approach improve students' ability to design inclusive software? (2) Will this approach create an inclusive climate among peers? (3) Will it affect student's success or lack thereof? (4) How and to what extent is this embedded inclusive design curriculum feasible to use? Pankati Patel, Patricia Morreale, Yulia Kumar, Daehan Kwak, Jean Chu, Rose Garcia, Margaret M. Burnett |
SIGCSE (2) | 3 |
| 2022 | An Assure AI Bot (AAAI bot)abstractArtificial Intelligence (AI) bots receive much attention and usage in industry manufacturing and even store cashier applications. Our research is to train AI bots to be software engineering assistants, specifically to detect biases and errors inside AI software applications. An example application is an AI machine learning system that sorts and classifies people according to various attributes, such as the algorithms involved in criminal sentencing, hiring, and admission practices. Biases, unfair decisions, and flaws in terms of the diversity, equity, and inclusion (DEI), in such systems could have severe consequences. As a Hispanic-Serving Institution, we are concerned about underrepresented groups and devoted an extended amount of our time to implementing “An Assure AI” (AAAI) Bot to detect biases and errors in AI applications. Our state-of-the-art AI Bot was developed based on our previous accumulated research in AI and Deep Learning (DL). The key differentiator is that we are taking a unique approach: instead of cleaning the input data, filtering it out and minimizing its biases, we trained our deep Neural Networks (NN) to detect and mitigate biases of existing AI models. The backend of our bot uses the Detection Transformer (DETR) framework, developed by Facebook, to monitor and detect the deep learning model’s internal biases. Neil Tellez, J. Serra, U. Ebreso, K. Opara, Yulia Kumar, J. Jenny Li 0001, Patricia Morreale |
ISNCC | 5 |
| 2022 | Evaluation of the Use of Growth Mindset in the CS ClassroomabstractWithin computer science education, a growth mindset is encouraged. However, faculty development on the use of growth mindset in the classroom is rare and resources to support the use of a growth mindset are limited. A framework for a computer science growth mindset classroom, which includes faculty development, lesson plans, and vocabulary for use with students, has been developed. The objective is to determine if faculty development in growth mindset and active use of the growth mindset cues in the CS0 and CS1 classroom result in superior academic outcomes. Comparative study results are presented for two semesters of virtual classroom environments: one semester without Growth Mindset, and one semester with Growth Mindset. Female students demonstrated the most growth, as measured by academic grades, in CS0, and maintained that growth in CS1. Males demonstrated growth as well, with both males and females converging at the same high point of accomplishment at the end of CS1. Race and ethnicity gaps between students were reduced, improving academic equity. Daehan Kwak, Patricia Morreale, Sarah Hug, Yulia Kumar, Jean Chu, Ching-Yu Huang, J. Jenny Li 0001, Paolien Wang |
SIGCSE (1) | 4 |
| 2021 | Framework for a Growth Mindset ClassroomabstractA growth mindset encourages the development of intelligence, in contrast to a fixed mindset, which considers intelligence to be fixed and unable to be changed. Within computer science education, there is an awareness of a growth mindset, but resources to support the use of a growth mindset in the computer science classroom are limited, and faculty development in the use of growth mindset in the classroom is rare. Researchers have developed a framework for a computer science growth mindset classroom, which includes faculty development, lesson plans, and vocabulary for use with students. In addition, the faculty have formed a community of practice which identifies areas where growth mindset techniques can be used and has implemented these strategies in their classrooms. The framework and materials developed are presented here, with early results from this ongoing work. The objective is to determine if faculty development on growth mindset and active use of the framework for a growth mindset classroom results in superior academic outcomes in CS0 and CS1. Preliminary results are for faculty in both face-to-face and virtual classroom environments. Patricia Morreale, J. Jenny Li 0001, Ching-Yu Huang, Daehan Kwak, Jean Chu, Yulia Kumar, Paolien Wang |
SIGCSE | 6 |