EDBT 2026 Demo / reviewers in the wild / expert
Shanchao Liang
dblp:387/3037
· DBLP profile ↗
3ranked-venue papers
2as first author
3since 2021 · last 2025
0000-0002-7707-6551ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Information extraction and text analysis · 40% Image recognition and object detection · 20% Vision and language · 20% | |
| Software engineering, system software, and programming languages
3 papers |
Program synthesis and code generation · 85% Requirements engineering and software design · 8% Software testing · 8% |
Topics — the 8 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Vision and language
cross-modal alignment |
0.9 | 1 | 2025 | WAFFLE: Fine-tuning Multi-Modal Model for Automated Front-End Development · ACL (1) 2025 |
Natural language and speech › Information extraction and text analysis
document understanding |
0.9 | 1 | 2025 | LATTE: Improving Latex Recognition for Tables and Formulae with Iterative Refinement · AAAI 2025 |
Computer vision › Image recognition and object detection › text recognition › optical character recognition
mathematical formula recognition |
0.9 | 1 | 2025 | LATTE: Improving Latex Recognition for Tables and Formulae with Iterative Refinement · AAAI 2025 |
Natural language and speech › Information extraction and text analysis › document understanding
table recognition |
0.9 | 1 | 2025 | LATTE: Improving Latex Recognition for Tables and Formulae with Iterative Refinement · AAAI 2025 |
Program synthesis and code generation › code generation evaluation
code generation benchmark |
0.9 | 1 | 2025 | Can Language Models Replace Programmers for Coding? REPOCOD Says 'Not Yet' · ACL (1) 2025 |
Program synthesis and code generation › code generation with language models
repository-level code generation |
0.9 | 1 | 2025 | Can Language Models Replace Programmers for Coding? REPOCOD Says 'Not Yet' · ACL (1) 2025 |
Program synthesis and code generation
code generation with language models |
0.3 | 1 | 2025 | LATTE: Improving Latex Recognition for Tables and Formulae with Iterative Refinement · AAAI 2025 |
Software testing › test suite evaluation
test case quality assessment |
0.3 | 1 | 2025 | Can Language Models Replace Programmers for Coding? REPOCOD Says 'Not Yet' · ACL (1) 2025 |
Methods — techniques the papers use, named apart from their topics
vision-language model · 1.7structure-aware attention · 1.7retrieval-augmented generation · 1.7pass@1 evaluation · 1.7large language model · 1.7iterative refinement · 1.7fault localization · 1.7contrastive fine-tuning · 1.7benchmark construction · 1.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | LATTE: Improving Latex Recognition for Tables and Formulae with Iterative RefinementabstractPortable Document Format (PDF) files are dominantly used for storing and disseminating scientific research, legal documents, and tax information. LaTeX is a popular application for creating PDF documents. Despite its advantages, LaTeX is not WYSWYG---what you see is what you get, i.e., the LaTeX source and rendered PDF images look drastically different, especially for formulae and tables. This gap makes it hard to modify or export LaTeX sources for formulae and tables from PDF images, and existing work is still limited. First, prior work generates LaTeX sources in a single iteration and struggles with complex LaTeX formulae. Second, existing work mainly recognizes and extracts LaTeX sources for formulae; and is incapable or ineffective for tables. This paper proposes LATTE, the first iterative refinement framework for LaTeX recognition. Specifically, we propose delta-view as feedback, which compares and pinpoints the differences between a pair of rendered images of the extracted LaTeX source and the expected correct image. Such delta-view feedback enables our fault localization model to localize the faulty parts of the incorrect recognition more accurately and enables our LaTeX refinement model to repair the incorrect extraction more accurately. LATTE improves the LaTeX source extraction accuracy of both LaTeX formulae and tables, outperforming existing techniques as well as GPT-4V by at least 7.07% of exact match, with a success refinement rate of 46.08% (formula) and 25.51% (table). Nan Jiang 0012, Shanchao Liang, Chengxiao Wang, Jiannan Wang 0002, Lin Tan 0001 |
AAAI | 2 |
| 2025 | WAFFLE: Fine-tuning Multi-Modal Model for Automated Front-End DevelopmentabstractWeb development involves turning UI designs into functional webpages, which can be difficult for both beginners and experienced developers due to the complexity of HTML's hierarchical structures and styles.While Large Language Models (LLMs) have shown promise in generating source code, two major challenges persist in UI-to-HTML code generation: (1) effectively representing HTML's hierarchical structure for LLMs, and (2) bridging the gap between the visual nature of UI designs and the text-based format of HTML code.To tackle these challenges, we introduce WAFFLE, a new fine-tuning strategy that uses a structure-aware attention mechanism to improve LLMs' understanding of HTML's structure and a contrastive fine-tuning approach to align LLMs' understanding of UI images and HTML code.Models fine-tuned with WAFFLE show up to 9.00 pp (absolute percentage point) higher HTML match, 0.0982 higher CW-SSIM, 32.99 higher CLIP, and 27.12 pp higher LLEM on our new benchmark WebSight-Test and an existing benchmark Design2Code, outperforming current fine-tuning methods. Shanchao Liang, Nan Jiang 0012, Shangshu Qian, Lin Tan 0001 |
ACL (1) | 1 |
| 2025 | Can Language Models Replace Programmers for Coding? REPOCOD Says 'Not Yet'abstractRecently, a number of repository-level code generation benchmarks-such as CoderEval, DevEval, RepoEval, RepoBench, and Long-Code-Arena-have emerged to evaluate the capabilities of large language models (LLMs) beyond standalone benchmarks like HumanEval and MBPP.Thus, a natural question is, would LLMs have similar performance in real world coding tasks as their performance in these benchmarks?Unfortunately, one cannot answer this question, since these benchmarks consist of short completions, synthetic examples, or focus on limited scale repositories, failing to represent real-world coding tasks.To address these challenges, we create RE-POCOD, a Python code-generation benchmark containing complex tasks with realistic dependencies in real-world large projects and appropriate metrics for evaluating source code.It includes 980 whole-function generation tasks from 11 popular projects, 50.8 % of which require repository-level context.REPOCOD includes 314 developer-written test cases per instance for better evaluation.We evaluate ten LLMs on REPOCOD and find that none achieves more than 30% pass@1 on REPOCOD, indicating the necessity of building stronger LLMs that can help developers in real-world software development.In addition, we found that retrieval-augmented generation achieves better results than using target function dependencies as context. Shanchao Liang, Nan Jiang 0012, Yiran Hu, Lin Tan 0001 |
ACL (1) | 1 |