VLDB 2026 Research / reviewers in the wild / expert
Dongke Hu
dblp:433/8980
· DBLP profile ↗
1ranked-venue papers
0as first author
1since 2021 · last 2026
0009-0007-0587-1154ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Representation and self-supervised learning · 44% Information extraction and text analysis · 44% Language models and text generation · 13% | |
| Databases, data mining, and information retrieval
1 paper |
Data integration and cleaning · 100% |
Topics — the 4 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Representation and self-supervised learning › pre-training
pretraining data |
1.0 | 1 | 2026 | HiParse: A Hierarchical Framework for Optimizing Web Content Extraction to Enhance LLM Pre-training Data · KDD (1) 2026 |
Natural language and speech › Information extraction and text analysis
web information extraction |
1.0 | 1 | 2026 | HiParse: A Hierarchical Framework for Optimizing Web Content Extraction to Enhance LLM Pre-training Data · KDD (1) 2026 |
Data integration and cleaning › data extraction
web data extraction |
1.0 | 1 | 2026 | HiParse: A Hierarchical Framework for Optimizing Web Content Extraction to Enhance LLM Pre-training Data · KDD (1) 2026 |
Natural language and speech › Language models and text generation › large language model training
pretraining data quality |
0.3 | 1 | 2026 | HiParse: A Hierarchical Framework for Optimizing Web Content Extraction to Enhance LLM Pre-training Data · KDD (1) 2026 |
Methods — techniques the papers use, named apart from their topics
hybrid parser · 2.0hierarchical parsing · 2.0decision engine · 2.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HiParse: A Hierarchical Framework for Optimizing Web Content Extraction to Enhance LLM Pre-training DataabstractIn the industrial pre-training of large language models (LLMs), large-scale web content extraction faces a critical trade-off: rule-based methods sacrifice quality for efficiency, while model-based methods do the opposite. As a result, existing solutions deployed at scale often lead to significant data quality degradation and corpus loss. To address this challenge, which we faced directly at Ant Group, this paper introduces HiParse, a hierarchical adaptive web parsing framework deployed in production. The core of HiParse is to abstract this multi-dimensional challenge into a ''PESO'' (Precision, Efficiency, Scope, Overhead) evaluation model, quantifying parser capabilities and parsing requirements from a ''supply-demand'' perspective. The framework integrates a hybrid parser system and uses a lightweight decision engine to dynamically route different web pages to the optimal parsing strategy. Comprehensive experiments show that HiParse not only achieves superior parsing quality (9.03/10 score) but its direct output corpus, with only 30 billion tokens, also improved the average performance of a downstream LLM on 11 benchmarks by 7.6%. Furthermore, the framework has proven its high throughput, cost-effectiveness, and alignment with business parsing demands under real industrial loads. HiParse provides a high-fidelity, scalable, and cost-effective data solution for enterprise-level LLM pre-training. All the relevant models in the paper will be open source soon. Yuzhuo Fu, Zhuyan Zhou, Binwei Zeng, Dongke Hu, Xiangchun Wang, Wang Hong |
KDD (1) | 6 |