EDBT 2026 Demo / reviewers in the wild / expert
Qiyu Li 0001
dblp:254/4412-1
· DBLP profile ↗
5ranked-venue papers
1as first author
5since 2021 · last 2026
0009-0004-5174-9844ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 3 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Language models and text generation · 63% Generative modeling · 28% Question answering and dialogue systems · 10% | |
| Network and information security
2 papers |
Privacy and data protection · 36% Usable security · 36% Web and mobile security · 28% | |
| Human-computer interaction and pervasive computing
3 papers |
Games and playful interaction · 62% Human-AI interaction · 38% | |
| Databases, data mining, and information retrieval
1 paper |
Data integration and cleaning · 100% |
Topics — the 7 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Usable security
developer-centered security |
1.0 | 1 | 2026 | PrivacyAkinator: Articulating Key Privacy Design Decisions by Answering LLM-Generated Multiple-choice Questions · CHI 2026 |
Privacy and data protection
privacy risk assessment |
1.0 | 1 | 2026 | PrivacyAkinator: Articulating Key Privacy Design Decisions by Answering LLM-Generated Multiple-choice Questions · CHI 2026 |
Natural language and speech › Language models and text generation
large language model evaluation |
0.9 | 1 | 2025 | GameArena: Evaluating LLM Reasoning through Live Computer Games · ICLR 2025 |
Natural language and speech › Language models and text generation › evaluation of language models
reasoning evaluation |
0.9 | 1 | 2025 | GameArena: Evaluating LLM Reasoning through Live Computer Games · ICLR 2025 |
Machine learning › Generative modeling
diffusion model |
0.8 | 1 | 2024 | Moderator: Moderating Text-to-Image Diffusion Models through Fine-grained Context-based Policies · CCS 2024 |
Web and mobile security
content moderation policy |
0.8 | 1 | 2024 | Moderator: Moderating Text-to-Image Diffusion Models through Fine-grained Context-based Policies · CCS 2024 |
Human-AI interaction › AI-assisted creativity
LLM-assisted design |
0.3 | 1 | 2026 | PrivacyAkinator: Articulating Key Privacy Design Decisions by Answering LLM-Generated Multiple-choice Questions · CHI 2026 |
Methods — techniques the papers use, named apart from their topics
reverse fine-tuning · 2.3fine-tuning · 2.3user study · 2.0observational study · 2.0large language model · 2.0human evaluation · 1.7dynamic benchmarking · 1.7task vector · 1.5task vectors · 0.8pretraining objective · 0.7hierarchical attention network · 0.7coherent chunk · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PrivacyAkinator: Articulating Key Privacy Design Decisions by Answering LLM-Generated Multiple-choice QuestionsabstractNIST’s Privacy Risk Assessment Methodology (PRAM) provides a structured framework for privacy experts to assess privacy risks. However, its complexity and reliance on expert knowledge make it difficult for novice developers to use effectively. This paper explores methods to lower these barriers. We first performed an observational study with 12 participants using PRAM in real-world scenarios, and found that novice developers struggled most with articulating privacy-related design decisions. We then developed PrivacyAkinator, an interactive tool that helps developers articulate key privacy decisions by answering LLM-generated multiple-choice questions. PrivacyAkinator introduces three innovations: a universal privacy representation that abstracts privacy-related design decisions into data flows and stakeholder interactions; a domain-aware design space mined from 10K privacy-related news articles; and a dynamic question-generation workflow to prioritize relevant questions. Our user study with 24 participants suggests that developers using PrivacyAkinator identified 47% more key decisions in 73% less time compared to PRAM. Qiyu Li 0001, Yuen Sum Wong, Yuen Kei Wong, Longxuan Yu, Haojian Jin |
CHI | 1 |
| 2025 | GameArena: Evaluating LLM Reasoning through Live Computer GamesabstractEvaluating the reasoning abilities of large language models (LLMs) is challenging. Existing benchmarks often depend on static datasets, which are vulnerable to data contamination and may get saturated over time, or on binary live human feedback that conflates reasoning with other abilities. As the most prominent dynamic benchmark, Chatbot Arena evaluates open-ended questions in real-world settings, but lacks the granularity in assessing specific reasoning capabilities. We introduce GameArena, a dynamic benchmark designed to evaluate LLM reasoning capabilities through interactive gameplay with humans. GameArena consists of three games designed to test specific reasoning capabilities (e.g., deductive and inductive reasoning), while keeping participants entertained and engaged. We analyze the gaming data retrospectively to uncover the underlying reasoning processes of LLMs and measure their fine-grained reasoning capabilities. We collect over 2000 game sessions and provide detailed assessments of various reasoning capabilities for five state-of-the-art LLMs. Our user study with 100 participants suggests that GameArena improves user engagement compared to Chatbot Arena. For the first time, GameArena enables the collection of step-by-step LLM reasoning data in the wild. Lanxiang Hu, Qiyu Li 0001, Anze Xie, Ion Stoica, Haojian Jin, Hao Zhang 0025 |
ICLR | 2 |
| 2024 | Moderator: Moderating Text-to-Image Diffusion Models through Fine-grained Context-based PoliciesabstractWe present Moderator, a policy-based model management system that allows administrators to specify fine-grained content moderation policies and modify the weights of a text-to-image (TTI) model to make it significantly more challenging for users to produce images that violate the policies. In contrast to existing general-purpose model editing techniques, which unlearn concepts without considering the associated contexts, Moderator allows admins to specify what content should be moderated, under which context, how it should be moderated, and why moderation is necessary. Given a set of policies, Moderator first prompts the original model to generate images that need to be moderated, then uses these self-generated images to reverse fine-tune the model to compute task vectors for moderation and finally negates the original model with the task vectors to decrease its performance in generating moderated content. We evaluated Moderator with 14 participants to play the role of admins and found they could quickly learn and author policies to pass unit tests in approximately 2.29 policy iterations. Our experiment with 32 stable diffusion users suggested that Moderator can prevent 65% of users from generating moderated content under 15 attempts and require the remaining users an average of 8.3 times more attempts to generate undesired content. Peiran Wang, Qiyu Li 0001, Longxuan Yu, Ang Li 0005, Haojian Jin |
CCS | 2 |
| 2023 | SheetPT: Spreadsheet Pre-training Based on Hierarchical Attention NetworkabstractSpreadsheets are an important and unique type of business document for data storage, analysis and presentation. The distinction between spreadsheets and most other types of digital documents lies in that spreadsheets provide users with high flexibility of data organization on the grid. Existing related techniques mainly focus on the tabular data and are incompetent in understanding the entire sheet. On the one hand, spreadsheets have no explicit separation across tabular data and other information, leaving a gap for the deployment of such techniques. On the other hand, pervasive data dependence and semantic relations across the sheet require comprehensive modeling of all the information rather than only the tables. In this paper, we propose SheetPT, the first pre-training technique on spreadsheets to enable effective representation learning under this scenario. For computational effectiveness and efficiency, we propose the coherent chunk, an intermediate semantic unit of sheet structure; and we accordingly devise a hierarchical attention-based architecture to capture contextual information across different structural granularities. Three pre-training objectives are also designed to ensure sufficient training against millions of spreadsheets. Two representative downstream tasks, formula prediction and sheet structure recognition are utilized to evaluate its capability and the prominent results reveal its superiority over existing state-of-the-art methods. Ran Jia, Qiyu Li 0001, Xiaoyuan Jin, Lun Du, Haoyu Dong 0001, Shi Han, Dongmei Zhang 0001 |
AAAI | 2 |
| 2023 | Knowledge-inspired Subdomain Adaptation for Cross-Domain Knowledge TransferabstractMost state-of-the-art deep domain adaptation techniques align source and target samples in a global fashion. That is, after alignment, each source sample is expected to become similar to any target sample. However, global alignment may not always be optimal or necessary in practice. For example, consider cross-domain fraud detection, where there are two types of transactions: credit and non-credit. Aligning credit and non-credit transactions separately may yield better performance than global alignment, as credit transactions are unlikely to exhibit patterns similar to non-credit transactions. To enable such fine-grained domain adaption, we propose a novel Knowledge-Inspired Subdomain Adaptation (KISA) framework. In particular, (1) We provide the theoretical insight that KISA minimizes the shared expected loss which is the premise for the success of domain adaptation methods. (2) We propose the knowledge-inspired subdomain division problem that plays a crucial role in fine-grained domain adaption. (3) We design a knowledge fusion network to exploit diverse domain knowledge. Extensive experiments demonstrate that KISA achieves remarkable results on fraud detection and traffic demand prediction tasks. Liyue Chen, Linian Wang, Weiqiang Wang 0002, Wenbiao Zhao, Qiyu Li 0001, Leye Wang |
CIKM | 7 |