EDBT 2026 Demo / reviewers in the wild / expert
Shujing Dong
dblp:215/8315
· DBLP profile ↗
7ranked-venue papers
1as first author
7since 2021 · last 2025
0009-0008-0240-681XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | LentEx: Generalizable Latent Entity Extraction via Synthetic Data and Instruction-Tuned LLMsabstractLatent entity extraction (LEE) tackles the challenge of identifying implicit, contextually inferred entities within free text—an area where traditional entity extraction methods fall short. In this paper, we introduce LentEx, a novel framework for latent entity extraction that leverages synthetic data generation and instruction fine-tuning to optimize smaller, efficient large language models (LLMs). Latent entities, which are often abstract and thematic, are crucial for applications such as retrieval-augmented generation (RAG), customer persona analysis, and knowledge graph enrichment. LentEx addresses the scarcity of labeled datasets by employing a template-based approach to generate diverse, contextually rich synthetic data, ensuring high variability and alignment with real-world distributions. To our knowledge, LentEx is the first to systematically approach LEE through the lens of LLMs. LentEx demonstrates significant performance improvements across multiple tasks, notably surpassing state-of-the-art models on the MTEB Clustering Benchmark. Furthermore, our methodology enables robust generalization to unseen domains, making LentEx highly applicable in real-world NLP tasks, including RAG and clustering, thereby establishing a new paradigm for latent entity understanding and extraction in natural language processing. Umesh Bodhwani, Yuan Ling, Cibi Chakravarthy Senthilkumar, Shujing Dong, Yarong Feng, Ayush Goyal |
IJCNN | 4 |
| 2025 | KDD Workshop on Evaluation and Trustworthiness of Agentic and Generative AIabstractThe rapid deployment of Generative and Agentic AI systems-ranging from large language models to autonomous agents-has created a critical need for rigorous and trustworthy evaluation methodologies. As these models influence real-world decision-making, traditional performance metrics alone fall short in capturing issues of safety, ethical alignment, misinformation, and human-centered usability. This workshop addresses these challenges by fostering interdisciplinary discussions and innovations in evaluation strategies that go beyond conventional benchmarks. Topics include holistic and multi-perspective assessments, scalable evaluation pipelines, reasoning and goal alignment in agentic behavior, misinformation detection, cross-modal generation, and trust calibration. By advancing robust, user-centric, and societally grounded evaluation practices, this workshop contributes to expanding KDD's methodological frontier into the emerging domain of responsible AI systems. Yuan Ling, Shujing Dong, Zheng Chen 0010, Yarong Feng, Sadid A. Hasan, George Karypis, Chandan K. Reddy |
KDD (2) | 2 |
| 2024 | Context-Aware and User Intent-Aware Follow-Up Question Generation (CA-UIA-QG): Mimicking User Behavior in Multi-Turn SettingabstractThis paper introduces a Context-Aware and User Intent-Aware follow-up Question Generation (CA-UIA-QG) method in multi-turn conversational settings. Our CA-UIA-QG model is designed to simultaneously consider the evolving context of a conversation and identify user intent. By integrating these aspects, it generates relevant follow-up questions, which can better mimic user behavior and align well with users’ conversational goals. When assessed using public Shopping datasets on Fashion domain, our approach demonstrates significant enhancements over CA-QG baseline models. Specifically, it achieves an improvement of up to 3% in BLEU, 7% in METEOR, and 8% in ROUGE-Lsum. Additionally, our findings show the efficacy of fine-tuning in enhancing the model’s capacity to better mimic user behavior, CoT prompting with fine-tuned model yields superior performance compared to the ensemble method. Furthermore, we investigate the impact of model size, model type, and intent granularity, highlighting their impact to overall model performance. The importance of our work lies in its effectiveness to improve follow-up question generation from the user’s perspective and application in developing user-centric conversational AI systems. Shujing Dong, Yuan Ling, Shunyan Luo, Yarong Feng, Zongyi Joe Liu, Ayush Goyal, Bruce Ferry |
IEEE Big Data | 1 |
| 2024 | KDD workshop on Evaluation and Trustworthiness of Generative AI ModelsabstractThe KDD workshop on Evaluation and Trustworthiness of Generative AI Models aims to address the critical need for reliable generative AI technologies by exploring comprehensive evaluation strategies. This workshop will delve into various aspects of assessing generative AI models, including Large Language Models (LLMs) and diffusion models, focusing on trustworthiness, safety, bias, fairness, and ethical considerations. With an emphasis on interdisciplinary collaboration, the workshop will feature invited talks, peer-reviewed paper presentations, and panel discussions to advance the state of the art in generative AI evaluation. Yuan Ling, Shujing Dong, Yarong Feng, Zongyi Joe Liu, George Karypis, Chandan K. Reddy |
KDD | 2 |
| 2024 | Detecting Content Segments from Online Sports Streaming Events: Challenges and SolutionsabstractDeveloping a client-side segmentation algorithm for on-line sports streaming holds significant importance. For instance, in order to assess the video quality from an end-user perspective such as artifact detection, it is important to initially segment the content within the streaming playback. The challenge lies in localizing the content due to the intricate scene changes between content and non-content sections in popular sports like football, tennis, baseball, and more. Client-side content detection can be implemented in two ways: intrusively, involving the interception of network traffic and parsing service provider data and logs, or non-intrusively, which entails capturing streamed videos from content providers and subjecting them to analysis using computer vision technologies. In this paper, we introduce a non-intrusive framework that leverages a combination of traditional machine learning algorithms and deep neural networks (DNN) to distinguish content sections from noncontent sections across various online sports streaming services. Our algorithm has demonstrated a remarkable level of accuracy and effectiveness in sports broadcasting events, effectively overcoming the complexities introduced by intricate non-content insertion methods during the games. Yarong Feng, Shunyan Luo, Yuan Ling, Shujing Dong |
WACV | 5 |
| 2023 | International Workshop on Multimodal Learning - 2023 Theme: Multimodal Learning with Foundation ModelsabstractThe recent advancements in machine learning and artificial intelligence (particularly foundation models such as BERT, GPT-3, T5, ResNet, etc.) have demonstrated remarkable capabilities and driven significant revolutionary changes to the way we make inferences from complex data. These models represent a fundamental shift in the way data are approached and offer exciting new research directions and opportunities for multimodal learning and data fusion. Given the potential of foundation models to transform the field of multimodal learning, there is a need to bring together experts and researchers to discuss the latest developments in this area, exchange ideas, and identify key research questions and challenges that need to be addressed. By hosting this workshop, we aim to create a forum for researchers to share their insights and expertise on multimodal data fusion and learning using foundation models, and to explore potential new research directions and applications in the rapidly evolving field. We expect contributions from interdisciplinary researchers to study and model interactions between (but not limited to) modalities of language, graphs, time-series, vision, tabular data, sensors, and more. Our workshop will emphasize interdisciplinary work and aim at seeding cross-team collaborations around new tasks, datasets, and models. Yuan Ling, Fanyou Wu, Shujing Dong, Yarong Feng, George Karypis, Chandan K. Reddy |
KDD | 3 |
| 2022 | Detect Audio-Video Temporal Synchronization Errors in Advertisements (Ads)abstractDetecting audio-video (A/V) synchronization error is important to measure end user experience. Today, researches in this domain are mainly focused on contents such as movies or sports. The state of art algorithms usually first detect a specific type of events and then correlate the A/V data within during these events, e.g., find the human chatting events and then correlate the vocals with the lip shapes. Detecting A/V sync errors during Ads, on the other hand, has not received a lot of attentions. Compared with contents, an Ads section do not contain a particular type of events that can be used to detect A/V sync error. For example, many vocals in Ads are either from background narrators or have a very short period of time, so that the popular lip-sync based algorithms won’t work accurately. In this paper, we present a novel algorithm that uses the scene change time features: we first segment out individual Ad from a playback. Then for each pair of temporal adjacent Ads, we compute the scene change time for the video data and the audio data separately, and then build their time difference histogram. Next, we aggregate the histograms from all Ads pairs within one Ads section. Finally, we combine the aggregated histogram to compute the A/V off-sync time values. We show that compared with the traditional lip-sync based algorithms, the new algorithm not only significantly improves the prediction rate, but also increases the prediction accuracy. Zongyi Joe Liu, Devin Chen, Yarong Feng, Yuan Ling, Shunyan Luo, Shujing Dong, Bruce Ferry |
ICPR | 6 |