VLDB 2026 Research / reviewers in the wild / expert
Yinghao Tang
dblp:344/2547
· DBLP profile ↗
6ranked-venue papers
2as first author
6since 2021 · last 2026
—ORCID · unresolved
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer graphics and multimedia
2 papers |
Multimedia analysis and retrieval · 54% Visualization and visual analytics · 46% | |
| Artificial intelligence
2 papers |
Efficient and distributed learning · 84% Generative modeling · 16% | |
| Databases, data mining, and information retrieval
1 paper |
Machine learning and data management · 50% Data integration and cleaning · 50% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Cloud and datacenter computing · 100% |
Topics — the 6 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Data integration and cleaning
data preprocessing |
0.9 | 1 | 2025 | DataLab: A Unified Platform for LLM-Powered Business Intelligence · ICDE 2025 |
Visualization and visual analytics › visualization generation
automated visualization generation |
0.9 | 1 | 2025 | DataLab: A Unified Platform for LLM-Powered Business Intelligence · ICDE 2025 |
Machine learning › Efficient and distributed learning
distributed training |
0.8 | 1 | 2024 | DLRover-RM: Resource Optimization for Deep Recommendation Models Training in the cloud · Proc. VLDB Endow. 2024 |
Machine learning › Efficient and distributed learning › distributed training
recommendation model training |
0.8 | 1 | 2024 | DLRover-RM: Resource Optimization for Deep Recommendation Models Training in the cloud · Proc. VLDB Endow. 2024 |
Cloud and datacenter computing › cluster resource management and scheduling
cluster resource management |
0.8 | 1 | 2024 | DLRover-RM: Resource Optimization for Deep Recommendation Models Training in the cloud · Proc. VLDB Endow. 2024 |
Machine learning › Generative modeling › diffusion model
text-to-image generation |
0.3 | 1 | 2026 | IGenBench: Benchmarking the Reliability of Text-to-Infographic Generation · ACL (1) 2026 |
Methods — techniques the papers use, named apart from their topics
benchmark construction · 2.0large language model · 1.7inter-agent communication · 1.7domain knowledge incorporation · 1.7agent framework · 1.7resource-performance models · 1.5heuristic resource allocation · 1.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | IGenBench: Benchmarking the Reliability of Text-to-Infographic GenerationabstractYinghao Tang, Xueding Liu, Boyuan Zhang, Tingfeng Lan, Yupeng Xie, Jiale Lao, Yiyao Wang, Haoxuan Li, Tingting Gao, Bo Pan, Luoxuan Weng, Xiuqi Huang, Minfeng Zhu, Yingchaojie Feng, Yuyu Luo, Wei Chen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yinghao Tang, Xueding Liu, Tingfeng Lan, Jiale Lao, Yiyao Wang, Tingting Gao, Bo Pan 0004, Luoxuan Weng, Xiuqi Huang, Minfeng Zhu 0001, Yingchaojie Feng, Yuyu Luo, Wei Chen 0001 |
ACL (1) | 1 |
| 2026 | HyperPMT: MHC-Peptide-TCR Binding Prediction via UniGAT-Based Hypergraph Neural Network
Xinhong Wu, Yinghao Tang |
ICIC (30) | 2 |
| 2026 | Exploring Multimodal Prompt for Visualization Authoring With Large Language ModelsabstractRecent advances in large language models (LLMs) have shown great potential in automating the process of visualization authoring through simple natural language utterances. However, instructing LLMs using natural language is limited in precision and expressiveness for conveying visualization intent, leading to misinterpretation and time-consuming iterations. To address these limitations, we conduct an empirical study to understand how LLMs interpret ambiguous or incomplete text prompts in the context of visualization authoring, and the conditions making LLMs misinterpret user intent. Informed by the findings, we introduce visual prompts as a complementary input modality to text prompts, which help clarify user intent and improve LLMs' interpretation abilities. To explore the potential of multimodal prompting in visualization authoring, we design VisPilot, which enables users to easily create visualizations using multimodal prompts, including text, sketches, and direct manipulations on existing visualizations. We evaluate VisPilot through a controlled user study and an expert evaluation. The results suggest that multimodal prompts facilitate users in communicating spatial constraints, local references, and design preferences while maintaining comparable task efficiency to text-only prompting. We further discuss when text, visual, and hybrid prompts are beneficial for visualization authoring, and summarize design implications for future human-AI authoring systems. Zhen Wen 0001, Luoxuan Weng, Yinghao Tang, Runjin Zhang, Bo Pan 0004, Minfeng Zhu 0001, Wei Chen 0001 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2025 | DataLab: A Unified Platform for LLM-Powered Business IntelligenceabstractBusiness intelligence (BI) transforms large volumes of data within modern organizations into actionable insights for informed decision-making. Recently, large language model (LLM)-based agents have streamlined the BI workflow by automatically performing task planning, reasoning, and actions in executable environments based on natural language (NL) queries. However, existing approaches primarily focus on individual BI tasks such as NL2SQL and NL2VIS. The fragmentation of tasks across different data roles and tools lead to inefficiencies and potential errors due to the iterative and collaborative nature of BI. In this paper, we introduce DataLab, a unified BI platform that integrates a one-stop LLM-based agent framework with an augmented computational notebook interface. DataLab supports various BI tasks for different data roles in data preparation, analysis, and visualization by seamlessly combining LLM assistance with user customization within a single environment. To achieve this unification, we design a domain knowledge incorporation module tailored for enterprise-specific BI tasks, an inter-agent communication mechanism to facilitate information sharing across the BI workflow, and a cell-based context management strategy to enhance context utilization efficiency in BI notebooks. Extensive experiments demonstrate that DataLab achieves state-of-the-art performance on various BI tasks across popular research benchmarks. Moreover, DataLab maintains high effectiveness and efficiency on real-world datasets from Tencent, achieving up to a 58.58% increase in accuracy and a 61.65 % reduction in token cost on enterprise-specific BI tasks. Luoxuan Weng, Yinghao Tang, Yingchaojie Feng, Zhuo Chang, Ruiqin Chen, Haozhe Feng, Chen Hou, Danqing Huang, Yang Li 0106, Huaming Rao, Canshi Wei, Xiuqi Huang, Minfeng Zhu 0001, Yuxin Ma 0001, Bin Cui 0001, Peng Chen 0021, Wei Chen 0001 |
ICDE | 2 |
| 2024 | Black-Box Adversarial Attack Against Transformer-Based Object Detection Models in Vehicular Networks
Yinghao Tang, Jinbiao Lu, Xingkai Kang, Hongchen Guo |
ICA3PP (5) | 1 |
| 2024 | DLRover-RM: Resource Optimization for Deep Recommendation Models Training in the cloudabstractDeep learning recommendation models (DLRM) rely on large embedding tables to manage categorical sparse features. Expanding such embedding tables can significantly enhance model performance, but at the cost of increased GPU/CPU/memory usage. Meanwhile, tech companies have built extensive cloud-based services to accelerate training DLRM models at scale. In this paper, we conduct a deep investigation of the DLRM training platforms at AntGroup and reveal two critical challenges: low resource utilization due to suboptimal configurations by users and the tendency to encounter abnormalities due to an unstable cloud environment. To overcome them, we introduce DLRover, an elastic training framework for DLRMs designed to increase resource utilization and handle the instability of a cloud environment. DLRover develops a resource-performance model by considering the unique characteristics of DLRMs and a three-stage heuristic strategy to automatically allocate and dynamically adjust resources for DLRM training jobs for higher resource utilization. Further, DLRover develops multiple mechanisms to ensure efficient and reliable execution of DLRM training jobs. Our extensive evaluation shows that DLRover reduces job completion times by 31%, increases the job completion rate by 6%, enhances CPU usage by 15%, and improves memory utilization by 20%, compared to state-of-the-art resource scheduling frameworks. DLRover has been widely deployed at AntGroup and processes thousands of DLRM training jobs on a daily basis. DLRover is open-sourced and has been adopted by 10+ companies. Qinlong Wang, Tingfeng Lan, Yinghao Tang, Bo Sang, Ziling Huang, Yiheng Du, Jian Sha, Hui Lu 0001, Yuanchun Zhou, Ke Zhang 0048, MingJie Tang |
Proc. VLDB Endow. | 3 |