VLDB 2026 Research / reviewers in the wild / expert
Zheng Ning
dblp:294/8443
· DBLP profile ↗
11ranked-venue papers
5as first author
11since 2021 · last 2025
0000-0001-7374-7453ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 8 · 4 first-author · 8 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | AROMA: Mixed-Initiative AI Assistance for Non-Visual Cooking by Grounding Multimodal Information Between Reality and VideosabstractVideos offer rich audiovisual information that can support people in performing activities of daily living (ADLs), but they remain largely inaccessible to blind or low-vision (BLV) individuals.In cooking, BLV people often rely on non-visual cues-such as touch, taste, and smell-to navigate their environment, making it difficult to follow UIST '25, September 28-October 01, 2025, Busan, Republic of Korea Ning et al.the predominantly audiovisual instructions found in video recipes.To address this problem, we introduce Aroma, an AI system that provides timely responses to the user based on real-time, contextaware assistance by integrating non-visual cues perceived by the user, a wearable camera feed, and video recipe content.Aroma uses a mixed-initiative approach: it responds to user requests while also proactively monitoring the video stream to offer timely alerts and guidance.This collaborative design leverages the complementary strengths of the user and AI system to align the physical environment with the video recipe, helping the user interpret their current state and make sense of the steps.We evaluated Aroma through a study with eight BLV participants and offered insights for designing interactive AI systems to support BLV individuals in performing ADLs. Zheng Ning, Leyang Li, Daniel Killough, JooYoung Seo, Patrick Carrington, Yapeng Tian, Yuhang Zhao 0001, Franklin Mingzhe Li, Toby Jia-Jun Li |
UIST | 1 |
| 2025 | AgentPbD: Interactive Agentic Workflow Generation from User Demonstration on Web BrowsersabstractProgramming by Demonstration (PbD) enables users to automate tasks through examples, but traditional systems generate low-level scripts that are hard to generalize or reuse. Recent advances in Large Language Models (LLMs) offer the potential to infer higher-level task structures, but rely on ambiguous natural language input. We present AgentPbD, a system that synthesizes task-level agentic workflows from a single user demonstration. By capturing browser actions and contextual metadata, AgentPbD automatically infers user goals and intentions, transforming user demonstrations into an editable, modular LLM agent workflow, and displays it on the browser extension interface. Users can further review and modify the workflow through visual programming. We demonstrate how AgentPbD bridges PbD and LLM planning, enabling interpretable and generalizable automation of complex web tasks. Zheng Ning, Toby Jia-Jun Li |
VL/HCC | 2 |
| 2024 | PodReels: Human-AI Co-Creation of Video Podcast TeasersabstractVideo podcast teasers are short videos that can be shared on social media platforms to capture interest in full episodes of a video podcast. These teasers enable long-form podcasters to reach new audiences and gain more followers. However, creating a compelling teaser from an hour-long episode can be challenging. Selecting interesting clips requires significant mental effort; editing the chosen clips into a cohesive, well-produced teaser is time-consuming. To support the creation of video podcast teasers, we first investigated what makes a good teaser. We combined insights from audience comments and creator interviews to identify key ingredients. We also identified a common workflow used by creators during this process. Based on these findings, we developed a human-AI co-creative tool called PodReels to assist video podcasters in crafting teasers. Our user study demonstrated that PodReels significantly reduces creators’ mental demand and improves their efficiency in producing video podcast teasers. Sitong Wang 0001, Zheng Ning, Anh Truong, Mira Dontcheva, Dingzeyu Li, Lydia B. Chilton |
Conference on Designing Interactive Systems | 2 |
| 2024 | MIMOSA: Human-AI Co-Creation of Computational Spatial Audio Effects on VideosabstractSpatial audio offers more immersive video consumption experiences to viewers; however, creating and editing spatial audio often expensive and requires specialized hardware equipment and skills, posing a high barrier for amateur video creators. We present Mimosa, a human-AI co-creation tool that enables amateur users to computationally generate and manipulate spatial audio effects. For a video with only monaural or stereo audio, Mimosa automatically grounds each sound source to the corresponding sounding object in the visual scene and enables users to further validate and fix errors in the location of the sounding objects. Users can also augment the spatial audio effect by flexibly manipulating the sounding source positions and creatively customizing the audio effect. The design of Mimosa exemplifies a human-AI collaboration approach that, instead of utilizing state-of-art end-to-end “black-box” ML models, uses a multistep pipeline that aligns its interpretable intermediate results with the user’s workflow. A lab user study with 15 participants demonstrates Mimosa’s usability, usefulness, expressiveness, and capability in creating immersive spatial audio effects in collaboration with users. Zheng Ning, Zheng Zhang 0043, Jerrick Ban, Ruohong Gan, Yapeng Tian, Toby Jia-Jun Li |
Creativity & Cognition | 1 |
| 2024 | SPICA: Interactive Video Content Exploration through Augmented Audio Descriptions for Blind or Low-Vision ViewersabstractBlind or Low-Vision (BLV) users often rely on audio descriptions (AD) to access video content. However, conventional static ADs can leave out detailed information in videos, impose a high mental load, neglect the diverse needs and preferences of BLV users, and lack immersion. To tackle these challenges, we introduce Spica, an AI-powered system that enables BLV users to interactively explore video content. Informed by prior empirical studies on BLV video consumption, Spica offers interactive mechanisms for supporting temporal navigation of frame captions and spatial exploration of objects within key frames. Leveraging an audio-visual machine learning pipeline, Spica augments existing ADs by adding interactivity, spatial sound effects, and individual object descriptions without requiring additional human annotation. Through a user study with 14 BLV participants, we evaluated the usability and usefulness of Spica and explored user behaviors, preferences, and mental models when interacting with augmented ADs. Zheng Ning, Brianna L. Wimer, Keyi Chen 0008, Jerrick Ban, Yapeng Tian, Yuhang Zhao 0001, Toby Jia-Jun Li |
CHI | 1 |
| 2024 | Developer Behaviors in Validating and Repairing LLM-Generated Code Using IDE and Eye TrackingabstractThe increasing use of large language model (LLM)-powered code generation tools, such as GitHub Copilot, is transforming software engineering practices. This paper investigates how developers validate and repair code generated by Copilot and examines the impact of code provenance awareness during these processes. We conducted a lab study with 28 participants tasked with validating and repairing Copilot-generated code in three software projects. Participants were randomly divided into two groups: one informed about the provenance of LLM-generated code and the other not. We collected data on IDE interactions, eye-tracking, cognitive workload assessments, and conducted semi-structured interviews. Our results indicate that, without explicit information, developers often fail to identify the LLM origin of the code. Developers exhibit LLM-specific behaviors such as frequent switching between code and comments, different attentional focus, and a tendency to delete and rewrite code. Being aware of the code’s provenance led to improved performance, increased search efforts, more frequent Copilot usage, and higher cognitive workload. These findings enhance our understanding of developer interactions with LLM-generated code and inform the design of tools for effective human-LLM collaboration in software development. Ningzhi Tang, Meng Chen 0020, Zheng Ning, Aakash Bansal, Yu Huang 0015, Collin McMillan, Toby Jia-Jun Li |
VL/HCC | 3 |
| 2024 | Insights into Natural Language Database Query Errors: from Attention Misalignment to User Handling StrategiesabstractQuerying structured databases with natural language (NL2SQL) has remained a difficult problem for years. Recently, the advancement of machine learning (ML), natural language processing (NLP), and large language models (LLM) have led to significant improvements in performance, with the best model achieving ∼85% percent accuracy on the benchmark Spider dataset. However, there is a lack of a systematic understanding of the types, causes, and effectiveness of error-handling mechanisms of errors for erroneous queries nowadays. To bridge the gap, a taxonomy of errors made by four representative NL2SQL models was built in this work, along with an in-depth analysis of the errors. Second, the causes of model errors were explored by analyzing the model-human attention alignment to the natural language query. Last, a within-subjects user study with 26 participants was conducted to investigate the effectiveness of three interactive error-handling mechanisms in NL2SQL. Findings from this article shed light on the design of model structure and error discovery and repair strategies for natural language data query interfaces in the future. Zheng Ning, Zheng Zhang 0043, Tianyi Zhang 0001, Toby Jia-Jun Li |
ACM Trans. Interact. Intell. Syst. | 1 |
| 2023 | Interactive Text-to-SQL Generation via Editable Step-by-Step ExplanationsabstractRelational databases play an important role in business, science, and more.However, many users cannot fully unleash the analytical power of relational databases, because they are not familiar with database languages such as SQL.Many techniques have been proposed to automatically generate SQL from natural language, but they suffer from two issues: (1) they still make many mistakes, particularly for complex queries, and (2) they do not provide a flexible way for non-expert users to validate and refine incorrect queries.To address these issues, we introduce a new interaction mechanism that allows users to directly edit a stepby-step explanation of a query to fix errors.Our experiments on multiple datasets, as well as a user study with 24 participants, demonstrate that our approach can achieve better performance than multiple SOTA approaches. Zheng Zhang 0043, Zheng Ning, Toby Jia-Jun Li, Jonathan K. Kummerfeld, Tianyi Zhang 0001 |
EMNLP | 3 |
| 2023 | An Empirical Study of Model Errors and User Error Discovery and Repair Strategies in Natural Language Database QueriesabstractRecent advances in machine learning (ML) and natural language processing (NLP) have led to significant improvement in natural language interfaces for structured databases (NL2SQL). Despite the great strides, the overall accuracy of NL2SQL models is still far from being perfect (∼ 75% on the Spider benchmark). In practice, this requires users to discern incorrect SQL queries generated by a model and manually fix them when using NL2SQL models. Currently, there is a lack of comprehensive understanding about the common errors in auto-generated SQLs and the effective strategies to recognize and fix such errors. To bridge the gap, we (1) performed an in-depth analysis of errors made by three state-of-the-art NL2SQL models; (2) distilled a taxonomy of NL2SQL model errors; and (3) conducted a within-subjects user study with 26 participants to investigate the effectiveness of three representative interactive mechanisms for error discovery and repair in NL2SQL. Findings from this paper shed light on the design of future error discovery and repair strategies for natural language data query interfaces. Zheng Ning, Zheng Zhang 0043, Tianyi Zhang 0001, Toby Jia-Jun Li |
IUI | 1 |
| 2023 | PEANUT: A Human-AI Collaborative Tool for Annotating Audio-Visual DataabstractAudio-visual learning seeks to enhance the computer’s multi-modal perception leveraging the correlation between the auditory and visual modalities. Despite their many useful downstream tasks, such as video retrieval, AR/VR, and accessibility, the performance and adoption of existing audio-visual models have been impeded by the availability of high-quality datasets. Annotating audio-visual datasets is laborious, expensive, and time-consuming. To address this challenge, we designed and developed an efficient audio-visual annotation tool called Peanut. Peanut’s human-AI collaborative pipeline separates the multi-modal task into two single-modal tasks, and utilizes state-of-the-art object detection and sound-tagging models to reduce the annotators’ effort to process each frame and the number of manually-annotated frames needed. A within-subject user study with 20 participants found that Peanut can significantly accelerate the audio-visual data annotation process while maintaining high annotation accuracy. Zheng Zhang 0043, Zheng Ning, Chenliang Xu, Yapeng Tian, Toby Jia-Jun Li |
UIST | 2 |
| 2023 | Exploring Contrast Consistency of Open-Domain Question Answering Systems on Minimally Edited QuestionsabstractAbstract Contrast consistency, the ability of a model to make consistently correct predictions in the presence of perturbations, is an essential aspect in NLP. While studied in tasks such as sentiment analysis and reading comprehension, it remains unexplored in open-domain question answering (OpenQA) due to the difficulty of collecting perturbed questions that satisfy factuality requirements. In this work, we collect minimally edited questions as challenging contrast sets to evaluate OpenQA models. Our collection approach combines both human annotation and large language model generation. We find that the widely used dense passage retriever (DPR) performs poorly on our contrast sets, despite fitting the training set well and performing competitively on standard test sets. To address this issue, we introduce a simple and effective query-side contrastive loss with the aid of data augmentation to improve DPR training. Our experiments on the contrast sets demonstrate that DPR’s contrast consistency is improved without sacrificing its accuracy on the standard test sets.1 Zhihan Zhang 0001, Wenhao Yu 0002, Zheng Ning, Mingxuan Ju, Meng Jiang 0001 |
Trans. Assoc. Comput. Linguistics | 3 |