Beicheng Xu

dblp:325/1751 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
5since 2021 · last 2026
0000-0003-4178-2451ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Efficient and distributed learning · 25% Optimization for machine learning · 25% Kernel, tree and ensemble methods · 25%
Databases, data mining, and information retrieval
2 papers
Data mining · 50% Web and social media mining · 34% Data integration and cleaning · 16%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Cloud and datacenter computing · 67% Performance modeling and evaluation · 33%

Topics — the 14 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Web and social media mining › information diffusion
diffusion network inference
1.222023
Multi-aspect Diffusion Network Inference · WWW 2023
Reconstructing Diffusion Networks from Incomplete Data · IJCAI 2022
Machine learning › Efficient and distributed learning
automated machine learning
1.012026
PSEO: Optimizing Post-hoc Stacking Ensemble Through Hyperparameter Tuning · AAAI 2026
Machine learning › Optimization for machine learning › hyperparameter optimization
combined algorithm selection and hyperparameter optimization
1.012026
PSEO: Optimizing Post-hoc Stacking Ensemble Through Hyperparameter Tuning · AAAI 2026
Machine learning › Kernel, tree and ensemble methods
ensemble learning
1.012026
PSEO: Optimizing Post-hoc Stacking Ensemble Through Hyperparameter Tuning · AAAI 2026
Cloud and datacenter computing
cluster resource management and scheduling
0.912025
A-Tune-Online: Efficient and QoS-Aware Online Configuration Tuning for Dynamic Workloads · ICDE 2025
Cloud and datacenter computing › configuration tuning
online tuning
0.912025
A-Tune-Online: Efficient and QoS-Aware Online Configuration Tuning for Dynamic Workloads · ICDE 2025
Performance modeling and evaluation
workload characterization
0.912025
A-Tune-Online: Efficient and QoS-Aware Online Configuration Tuning for Dynamic Workloads · ICDE 2025
Machine learning › Graph learning › graph diffusion › information diffusion
diffusion network inference
0.712023
Multi-aspect Diffusion Network Inference · WWW 2023
Data mining
network inference
0.712023
Multi-aspect Diffusion Network Inference · WWW 2023
Data mining
anomaly detection
0.612022
Reconstructing Diffusion Networks from Incomplete Data · IJCAI 2022
Data integration and cleaning › missing data
missing value imputation
0.612022
Reconstructing Diffusion Networks from Incomplete Data · IJCAI 2022
Data mining
network analysis
0.612022
Reconstructing Diffusion Networks from Incomplete Data · IJCAI 2022
Machine learning › Generative modeling › generative model
probabilistic generative model
0.212023
Multi-aspect Diffusion Network Inference · WWW 2023
Machine learning › Probabilistic and Bayesian machine learning
probabilistic inference
0.212022
Reconstructing Diffusion Networks from Incomplete Data · IJCAI 2022

Methods — techniques the papers use, named apart from their topics

probabilistic generative model · 1.3posterior inference · 1.3expectation-maximization · 1.1correlation analysis · 1.1hyperparameter tuning · 1.0binary quadratic programming · 1.0warm-start · 0.9lower confidence bound · 0.9bayesian optimization · 0.9
YearPublicationVenuePosition
2026 PSEO: Optimizing Post-hoc Stacking Ensemble Through Hyperparameter Tuning
abstract
The Combined Algorithm Selection and Hyperparameter Optimization (CASH) problem is fundamental in Automated Machine Learning (AutoML). Inspired by the success of ensemble learning, recent AutoML systems construct post-hoc ensembles for final predictions rather than relying on the best single model. However, while most CASH methods conduct extensive searches for the optimal single model, they typically employ fixed strategies during the ensemble phase that fail to adapt to specific task characteristics. To tackle this issue, we propose PSEO, a framework for post-hoc stacking ensemble optimization. First, we conduct base model selection through binary quadratic programming, with a trade-off between diversity and performance. Furthermore, we introduce two mechanisms to fully realize the potential of multi-layer stacking. Finally, PSEO builds a hyperparameter space and searches for the optimal post-hoc ensemble strategy within it. Empirical results on 80 public datasets show that PSEO achieves the best average test rank (2.96) among 16 methods, including post-hoc designs in recent AutoML systems and state-of-the-art ensemble learning methods.
Beicheng Xu, Keyao Ding, Yupeng Lu, Bin Cui 0001
AAAI1
2025 A-Tune-Online: Efficient and QoS-Aware Online Configuration Tuning for Dynamic Workloads
abstract
Automatic configuration tuning of online services with dynamic workloads has attracted increasing interest. Effective online tuning ensures configurations adapt to workload changes over time to maintain optimal online service performance. To be practical, online tuning must satisfy the dynamicity, efficiency, and Quality of Service (QoS) requirements. However, existing online tuning approaches fail to meet these requirements due to the inability to eliminate negative effects from historical observations. In this paper, we propose A-Tune-Online, an online configuration tuning system that tackles dynamic workloads, delivering superior tuning efficiency, and QoS guarantee simultaneously to a wide range of online scenarios. We identify that restarting the optimization based on explicit workload shift detection is necessary and critical to eliminate negative historical observations. First, to invoke optimization restarts appropriately, we design a multi-stage multi-indicator detection strategy based on heuristic rules and configuration replays. Then, to avoid initial efficiency drop after re-optimization, A-Tune-Online utilizes a similarity-based dual warm start scheme that transfers knowledge from similar historical workloads effectively. Finally, to prevent transient performance degradation from violating QoS guarantee after optimization restart, we leverage lower confidence bound to construct a safety region where each configuration is expected to perform better than the QoS requirement. Empirical study on five tuning scenarios showcases the superiority of A-Tune-Online compared with state-of-art tuning systems. A-Tune-Online achieves an average speedup of 2.90x and 1.72x compared with OnlineTune and DDPG+, respectively. We provide a version of our system in https://github.com/PKU-DAIR/A-Tune-Online.
Yu Shen 0003, Beicheng Xu, Yupeng Lu, Huaijun Jiang, Zhipeng Xie, Senbo Fu, Nan Zhang 0004, Yuxin Ren 0001, Ning Jia 0004, Xinwei Hu, Bin Cui 0001
ICDE2
2024 OpenBox: A Python Toolkit for Generalized Black-box Optimization
abstract
Black-box optimization (BBO) has a broad range of applications, including automatic machine learning, experimental design, and database knob tuning. However, users still face challenges when applying BBO methods to their problems at hand with existing software packages in terms of applicability, performance, and efficiency. This paper presents OpenBox, an open-source BBO toolkit with improved usability. It implements user-friendly interfaces and visualization for users to define and manage their tasks. The modular design behind OpenBox facilitates its flexible deployment in existing systems. Experimental results demonstrate the effectiveness and efficiency of OpenBox over existing systems. The source code of OpenBox is available at https://github.com/PKU-DAIR/open-box.
Huaijun Jiang, Yu Shen 0003, Yang Li 0106, Beicheng Xu, Sixian Du, Wentao Zhang 0001, Ce Zhang 0001, Bin Cui 0001
J. Mach. Learn. Res.4
2023 Multi-aspect Diffusion Network Inference
abstract
To learn influence relationships between nodes in a diffusion network, most existing approaches resort to precise timestamps of historical node infections. The target network is customarily assumed as an one-aspect diffusion network, with homogeneous influence relationships. Nonetheless, tracing node infection timestamps is often infeasible due to high cost, and the type of influence relationships may be heterogeneous because of the diversity of propagation media. In this work, we study how to infer a multi-aspect diffusion network with heterogeneous influence relationships, using only node infection statuses that are more readily accessible in practice. Equipped with a probabilistic generative model, we iteratively conduct a posteriori, quantitative analysis on historical diffusion results of the network, and infer the structure and strengths of homogeneous influence relationships in each aspect. Extensive experiments on both synthetic and real-world networks are conducted, and the results verify the effectiveness and efficiency of our approach.
Hao Huang 0001, Keqi Han, Beicheng Xu, Ting Gan
WWW3
2022 Reconstructing Diffusion Networks from Incomplete Data
abstract
To reconstruct the topology of a diffusion network, existing approaches customarily demand not only eventual infection statuses of nodes, but also the exact times when infections occur. In real-world settings, such as the spread of epidemics, tracing the exact infection times is often infeasible; even obtaining the eventual infection statuses of all nodes is a challenging task. In this work, we study topology reconstruction of a diffusion network with incomplete observations of the node infection statuses. To this end, we iteratively infer the network topology based on observed infection statuses and estimated values for unobserved infection statuses by investigating the correlation of node infections, and learn the most probable probabilities of the infection propagations among nodes w.r.t. current inferred topology, as well as the corresponding probability distribution of each unobserved infection status, which in turn helps update the estimate of unobserved data. Extensive experimental results on both synthetic and real-world networks verify the effectiveness and efficiency of our approach.
Hao Huang 0001, Keqi Han, Beicheng Xu, Ting Gan
IJCAI3