Boxuan Chen

dblp:320/7899 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
4since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 HourglassSketch: An Efficient and Scalable Framework for Graph Stream Summarization
abstract
Graph stream is a special kind of data stream, where every item coming in sequence represents an edge in a dynamic graph. Graph stream has wide application in many fields, including cyber security, social networks and financial fraud detection. In this paper, we propose HourglassSketch, a two-stage data structure, for high-accuracy graph stream summarization. In Stage 1, HourglassSketch uses a CocoSketch to accurately record a partial collection of large-weight edges. In Stage 2, HourglassSketch integrates a TowerSketch with a TCMSketch to approximately record the statistics of most small-weight edges. In addition, we propose a key technique named Error Funnel to further reduce its error margin. Theoretical analysis and experimental results demonstrate that HourglassSketch supports various kinds of query operation and adapts well to graph stream storage. HourglassSketch achieves up to 100x smaller error and 2.7x higher speed than prior work. We also explore the versatility of HourglassSketch as a hardware-friendly framework by implementing it on FPGA and P4 platforms. We have released our codes on GitHub.
Jiarui Guo, Boxuan Chen, Kaicheng Yang 0001, Tong Yang 0003, Zirui Liu 0002, Qiuheng Yin, Yuhan Wu 0001, Bin Cui 0001, Xi Peng 0006, Renhai Chen, Gong Zhang 0001
ICDE2
2025 Retrieval Model for Surface Dead Fuel Moisture Data Based on Microwave Soil Moisture Content
abstract
Wildfires are natural disasters that pose substantial threats to the environment. The accurate prediction of wildfire risk levels and timely implementation of effective mitigation measures are critical for wildfire prevention and ecological security maintenance. Fuel moisture content is an important factor affecting the spread and intensity of wildfires; however, there is currently a lack of large-scale data on surface dead fuel moisture content (DFMC). Therefore, in this study, we aimed to develop a more accurate method for retrieving DFMC by integrating multi-source satellite remote sensing data with machine learning algorithms. Evaluation of six retrieval models (extreme gradient boosting [XGB], linear regression [LR], generalized additive model [GAM], random forest [RF], convolutional neural network [CNN], and long short-term memory [LSTM]) confirmed the adaptability of XGB to four forest types, achieving an average R² value of 0.78. In addition, we used SHAP analysis to assess the importance of the model's influencing factors and identified the significance of soil moisture content (SMC) and evapotranspiration. Furthermore, comparative analysis of four input parameter combinations confirmed the pivotal role of SMC, with the integrated use of ERA5-Land reanalysis data and SMC data achieving optimal model performance while balancing large-scale applicability with accuracy. Additionally, a daily DFMC dataset for the Greater Khingan Mountains was constructed by integrating SMAP soil moisture data, ERA5 meteorological data, and the XGB model. In this study, we innovatively combined microwave remote sensing data with machine learning to provide a robust methodological framework for DFMC data retrieval via microwave remote sensing. Our results establish a foundation for enhancing the precision of forest fire risk early warning systems.
Yuanting Gao, Boxuan Chen, Jiale Fan, Tongxin Hu
IEEE Trans. Geosci. Remote. Sens.3
2025 CAFE+: Towards Compact, Adaptive, and Fast Embedding for Large-scale Online Recommendation Models
abstract
The growing memory demands of embedding tables in Deep Learning Recommendation Models (DLRMs) pose great challenges for model training and deployment. Existing embedding compression solutions cannot simultaneously achieve memory efficiency, low latency, and adaptability to dynamic data distribution. This article presents CAFE+, a Compact, Adaptive, and Fast Embedding compression framework that meets the above requirements. The design philosophy of CAFE+ is to dynamically allocate more memory to important features and less to unimportant ones. We assign unique embedding to important feature and allow multiple unimportant features sharing one embedding. We propose a fast and lightweight feature monitor, to real-time capture feature importance and report important features. We theoretically analyze the accuracy of our feature monitor and prove the superiority of CAFE+ from the aspect of model convergence. Extensive experiments show CAFE+ outperforms existing embedding compression methods, yielding \(3.94\%\) and \(3.94\%\) superior testing AUC on Criteo Kaggle dataset and CriteoTB dataset at a compression ratio of \(10{,}000\times\) . Building on our conference version [ 114 ], this journal version introduces several novel designs (implicit importance attenuation, adaptive threshold adjustment, and ColdSifter) that enable CAFE+ to more effectively adapt to long-term online learning and achieve better model quality. All codes are available at GitHub [ 112 ].
Zirui Liu 0002, Hailin Zhang 0004, Boxuan Chen, Zihan Jiang 0004, Yikai Zhao 0001, Yangyu Tao, Tong Yang 0003, Bin Cui 0001
ACM Trans. Inf. Syst.3
2024 CAFE: Towards Compact, Adaptive, and Fast Embedding for Large-scale Recommendation Models
abstract
Recently, the growing memory demands of embedding tables in Deep Learning Recommendation Models (DLRMs) pose great challenges for model training and deployment. Existing embedding compression solutions cannot simultaneously meet three key design requirements: memory efficiency, low latency, and adaptability to dynamic data distribution. This paper presents CAFE, a Compact, Adaptive, and Fast Embedding compression framework that addresses the above requirements. The design philosophy of CAFE is to dynamically allocate more memory resources to important features (called hot features), and allocate less memory to unimportant ones. In CAFE, we propose a fast and lightweight sketch data structure, named HotSketch, to capture feature importance and report hot features in real time. For each reported hot feature, we assign it a unique embedding. For the non-hot features, we allow multiple features to share one embedding by using hash embedding technique. Guided by our design philosophy, we further propose a multi-level hash embedding framework to optimize the embedding tables of non-hot features. We theoretically analyze the accuracy of HotSketch, and analyze the model convergence against deviation. Extensive experiments show that CAFE significantly outperforms existing embedding compression methods, yielding 3.92% and 3.68% superior testing AUC on Criteo Kaggle dataset and CriteoTB dataset at a compression ratio of 10000x. The source codes of CAFE are available at GitHub.
Hailin Zhang 0004, Zirui Liu 0002, Boxuan Chen, Yikai Zhao 0001, Tong Yang 0003, Bin Cui 0001
Proc. ACM Manag. Data3