Dongke Hu

dblp:433/8980 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2026
0009-0007-0587-1154ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Representation and self-supervised learning · 44% Information extraction and text analysis · 44% Language models and text generation · 13%
Databases, data mining, and information retrieval
1 paper
Data integration and cleaning · 100%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Representation and self-supervised learning › pre-training
pretraining data
1.012026
HiParse: A Hierarchical Framework for Optimizing Web Content Extraction to Enhance LLM Pre-training Data · KDD (1) 2026
Natural language and speech › Information extraction and text analysis
web information extraction
1.012026
HiParse: A Hierarchical Framework for Optimizing Web Content Extraction to Enhance LLM Pre-training Data · KDD (1) 2026
Data integration and cleaning › data extraction
web data extraction
1.012026
HiParse: A Hierarchical Framework for Optimizing Web Content Extraction to Enhance LLM Pre-training Data · KDD (1) 2026
Natural language and speech › Language models and text generation › large language model training
pretraining data quality
0.312026
HiParse: A Hierarchical Framework for Optimizing Web Content Extraction to Enhance LLM Pre-training Data · KDD (1) 2026

Methods — techniques the papers use, named apart from their topics

hybrid parser · 2.0hierarchical parsing · 2.0decision engine · 2.0
YearPublicationVenuePosition
2026 HiParse: A Hierarchical Framework for Optimizing Web Content Extraction to Enhance LLM Pre-training Data
abstract
In the industrial pre-training of large language models (LLMs), large-scale web content extraction faces a critical trade-off: rule-based methods sacrifice quality for efficiency, while model-based methods do the opposite. As a result, existing solutions deployed at scale often lead to significant data quality degradation and corpus loss. To address this challenge, which we faced directly at Ant Group, this paper introduces HiParse, a hierarchical adaptive web parsing framework deployed in production. The core of HiParse is to abstract this multi-dimensional challenge into a ''PESO'' (Precision, Efficiency, Scope, Overhead) evaluation model, quantifying parser capabilities and parsing requirements from a ''supply-demand'' perspective. The framework integrates a hybrid parser system and uses a lightweight decision engine to dynamically route different web pages to the optimal parsing strategy. Comprehensive experiments show that HiParse not only achieves superior parsing quality (9.03/10 score) but its direct output corpus, with only 30 billion tokens, also improved the average performance of a downstream LLM on 11 benchmarks by 7.6%. Furthermore, the framework has proven its high throughput, cost-effectiveness, and alignment with business parsing demands under real industrial loads. HiParse provides a high-fidelity, scalable, and cost-effective data solution for enterprise-level LLM pre-training. All the relevant models in the paper will be open source soon.
Yuzhuo Fu, Zhuyan Zhou, Binwei Zeng, Dongke Hu, Xiangchun Wang, Wang Hong
KDD (1)6