EDBT 2026 Demo / reviewers in the wild / expert
Meng Ling
dblp:268/1480
· DBLP profile ↗
7ranked-venue papers
2as first author
7since 2021 · last 2026
0000-0001-6597-5448ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer graphics and multimedia
3 papers |
Visualization and visual analytics · 100% |
Topics — the 4 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Visualization and visual analytics
visualization evaluation |
1.0 | 1 | 2026 | What Makes a Visualization Image Complex? · IEEE Trans. Vis. Comput. Graph. 2026 |
Visualization and visual analytics
visualization design |
0.8 | 1 | 2024 | Eleven Years of Gender Data Visualization: A Step Towards More Inclusive Gender Representation · IEEE Trans. Vis. Comput. Graph. 2024 |
Visualization and visual analytics
visualization dataset |
0.5 | 1 | 2021 | VIS30K: A Collection of Figures and Tables From IEEE Visualization Conference Publications · IEEE Trans. Vis. Comput. Graph. 2021 |
Visualization and visual analytics
visual analytics |
0.1 | 1 | 2021 | VIS30K: A Collection of Figures and Tables From IEEE Visualization Conference Publications · IEEE Trans. Vis. Comput. Graph. 2021 |
Methods — techniques the papers use, named apart from their topics
information-theoretic metrics · 1.0crowdsourcing experiment · 1.0content analysis · 0.8coding · 0.8convolutional neural network · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LLM-Based Content Tagging at The Washington PostabstractWe present a production LLM-based taxonomy classification system deployed at The Washington Post that tags news content across five schemas (Subject, Person, Company, Organization, Geography) using a proprietary taxonomy of ∼ 20,400 entries across seven hierarchical levels. For the Subject schema, we employ embedding-based candidate filtering followed by LLM selection. For other schemas, we combine LLM-based named entity extraction with fuzzy n-gram matching, followed by LLM selection. Comparison of post-production F1 scores against commercial vendor baselines demonstrates significant improvements across all five schemas, with the most substantial gain in Subject schema (+29.3%, p < 0.001). The system processes hundreds to thousands of articles and news items daily with a mean latency of 3–4 seconds per request and supports zero-downtime taxonomy updates. Meng Ling, Himanshu Jahagirdar, Janith Weerasinghe, Han Jun Yoon, Suja Thomas, Anuradha Uduwage, Eui-Hong Han |
UMAP | 1 |
| 2026 | A Case Study of Offline Reinforcement Learning for Paywall DecisioningabstractWe describe how The Washington Post deployed an offline reinforcement learning (RL) system to optimize paywall decisioning at production scale. We cast each non-subscriber article access attempt as a sequential decision with three actions: free access, registration wall, or subscription paywall, and learn policies from logged data collected via a small-traffic randomized controlled trial and subsequent production logging. We iterated from a tabular Q-learning baseline to a deep offline RL model trained with Conservative Q-Learning (CQL), using off-policy evaluation primarily to screen and rank candidates before online testing. The system was rolled out with guardrails and a persistent randomized holdout to manage risk in a revenue-critical setting. In year-long online experiments, the learned policies outperformed the legacy rules-based metering policy and improved a stakeholder-weighted value metric; the CQL policy delivered a +3% lift versus the randomized baseline while increasing subscriptions (+6%) and reducing the registration gap relative to earlier RL iterations. This case study highlights the practical steps needed to safely train, evaluate, and deploy offline RL for high-stakes personalization. Janith Weerasinghe, Han Jun Yoon, Meng Ling, Himanshu Jahagirdar, Suja Thomas, Anuradha Uduwage, Sam Han |
UMAP | 3 |
| 2026 | Uncertainty-Aware Reinforcement Learning for Conversion-Optimized Content GatingabstractPublishers increasingly rely on access gates to drive registrations and subscriptions. Determining when to present these gates is a sequential decision problem well suited to reinforcement learning (RL). However, online exploration is costly and risky due to delayed conversion signals. We introduce Uncertainty-Aware Advantage-Weighted Actor–Critic (UA-AWAC), an offline RL method that learns from logged traffic to produce conversion-ready policies. UA-AWAC optimizes a multi-objective reward incorporating subscriptions, registrations, and engagement, while mitigating distribution shift through epistemic uncertainty modeling and pessimistic value targets. The policy is trained using advantage-weighted behavioral cloning with Kullback–Leibler (KL) regularization to remain close to historical gating behavior. Experiments show that UA-AWAC improves subscription rate by up to 10% and registration rate by 62% compared to baseline and state-of-the-art offline RL methods, demonstrating a practical and stable solution for intelligent content gating where exploration risks are high. Han Jun Yoon, Janith Weerasinghe, Himanshu Jahagirdar, Meng Ling, Suja Thomas, Anuradha Uduwage, Sam Han |
UMAP | 4 |
| 2026 | What Makes a Visualization Image Complex?abstractWe investigate the perceived visual complexity (VC) in data visualizations using objective image-based metrics. We collected VC scores through a large-scale crowdsourcing experiment involving 349 participants and 1,800 visualization images. We then examined how these scores align with 12 image-based metrics spanning pixel-based and statistic-information-theoretic (clutter), color, shape, and our two new object-based metrics (meaningful-color-count (MeC) and text-to-ink ratio (TiR)). Our results show that both low-level edges and high-level elements affect perceived VC in visualization images; the number of corners and distinct colors are robust metrics across visualizations. Second, feature congestion, a statistical information-theoretic metric capturing color and texture patterns, is the strongest predictor of perceived complexity in visualizations rich in the same continuous color/texture stimuli; edge density effectively explains VC in node-link diagrams. Additionally, we observe a bell-curve effect for texts: increasing TiR initially reduces complexity, reaching an optimal point, beyond which further text increases VC. Our quantification model is also interpretable-enabling metric-based explanations-grounded in the VisComplexity2K dataset, bridging computational metrics with human perceptual responses. The preregistration is available at osf.io/5xe8a. osf.io/bdet6 has the dataset and analysis code. Mengdi Chu, Zefeng Qiu, Meng Ling, Shuning Jiang, Robert S. Laramee, Michael Sedlmair, Jian Chen 0006 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2024 | Eleven Years of Gender Data Visualization: A Step Towards More Inclusive Gender RepresentationabstractWe present an analysis of the representation of gender as a data dimension in data visualizations and propose a set of considerations around visual variables and annotations for gender-related data. Gender is a common demographic dimension of data collected from study or survey participants, passengers, or customers, as well as across academic studies, especially in certain disciplines like sociology. Our work contributes to multiple ongoing discussions on the ethical implications of data visualizations. By choosing specific data, visual variables, and text labels, visualization designers may, inadvertently or not, perpetuate stereotypes and biases. Here, our goal is to start an evolving discussion on how to represent data on gender in data visualizations and raise awareness of the subtleties of choosing visual variables and words in gender visualizations. In order to ground this discussion, we collected and coded gender visualizations and their captions from five different scientific communities (Biology, Politics, Social Studies, Visualisation, and Human-Computer Interaction), in addition to images from Tableau Public and the Information Is Beautiful awards showcase. Overall we found that representation types are community-specific, color hue is the dominant visual channel for gender data, and nonconforming gender is under-represented. We end our paper with a discussion of considerations for gender visualization derived from our coding and the literature and recommendations for large data collection bodies. A free copy of this paper and all supplemental materials are available at https://osf.io/v9ams/. Florent Cabric, Margrét V. Bjarnadóttir, Meng Ling, Guðbjörg Linda Rafnsdóttir, Petra Isenberg |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2021 | Document Domain Randomization for Deep Learning Document Layout ExtractionabstractWe present document domain randomization (DDR), the first successful transfer of convolutional neural networks (CNNs) trained only on graphically rendered pseudo-paper pages to real-world document segmentation. DDR renders pseudo-document pages by modeling randomized textual and non-textual contents of interest, with user-defined layout and font styles to support joint learning of fine-grained classes. We demonstrate competitive results using our DDR approach to extract nine document classes from the benchmark CS-150 and papers published in two domains, namely annual meetings of Association for Computational Linguistics (ACL) and IEEE Visualization (VIS). We compare DDR to conditions of style mismatch, fewer or more noisy samples that are more easily obtained in the real world. We show that high-fidelity semantic information is not necessary to label semantic classes but style mismatch between train and test can lower model accuracy. Using smaller training samples had a slightly detrimental effect. Finally, network models still achieved high test accuracy when correct labels are diluted towards confusing labels; this behavior hold across several classes. Meng Ling, Jian Chen 0006, Torsten Möller, Petra Isenberg, Tobias Isenberg 0001, Michael Sedlmair, Robert S. Laramee, Han-Wei Shen, Jian Wu 0006, C. Lee Giles |
ICDAR (1) | 1 |
| 2021 | VIS30K: A Collection of Figures and Tables From IEEE Visualization Conference PublicationsabstractWe present the VIS30K dataset, a collection of 29,689 images that represents 30 years of figures and tables from each track of the IEEE Visualization conference series (Vis, SciVis, InfoVis, VAST). VIS30K's comprehensive coverage of the scientific literature in visualization not only reflects the progress of the field but also enables researchers to study the evolution of the state-of-the-art and to find relevant work based on graphical content. We describe the dataset and our semi-automatic collection process, which couples convolutional neural networks (CNN) with curation. Extracting figures and tables semi-automatically allows us to verify that no images are overlooked or extracted erroneously. To improve quality further, we engaged in a peer-search process for high-quality figures from early IEEE Visualization papers. With the resulting data, we also contribute VISImageNavigator (VIN, visimagenavigator.github.io), a web-based tool that facilitates searching and exploring VIS30K by author names, paper keywords, title and abstract, and years. Jian Chen 0006, Meng Ling, Rui Li 0067, Petra Isenberg, Tobias Isenberg 0001, Michael Sedlmair, Torsten Möller, Robert S. Laramee, Han-Wei Shen, Katharina Wünsche |
IEEE Trans. Vis. Comput. Graph. | 2 |