Shen-Huan Lyu

dblp:255/7033 · DBLP profile ↗
← Back
20ranked-venue papers
6as first author
19since 2021 · last 2026
0000-0002-0173-8408ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 4 first-author · 12 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 5 since 2021Computer networks · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 A semi-supervised deep forest framework based on margin distribution optimization for tabular data
Shen-Huan Lyu, Jia-Le Xu, Yi-Xiao He, Yanyan Wang 0001, Qingfu Zhang 0001
Inf. Sci.1
2026 Improving multi-label contrastive learning by leveraging label distribution
Shen-Huan Lyu, Tian-Shuang Wu, Yanyan Wang 0001, Bin Tang 0002
Pattern Recognit.2
2026 Enhance and reuse: A dual-mechanism approach to boost deep forest for label distribution learning
Jia-Le Xu, Shen-Huan Lyu, Yu-Nian Wang, Zhihao Qu, Bin Tang 0002
Pattern Recognit.2
2026 Compressing model with few class-imbalance samples: An out-of-distribution expedition
Tian-Shuang Wu, Shen-Huan Lyu, Yanyan Wang 0001, Zhihao Qu
Pattern Recognit. Lett.2
2026 Interpreting Deep Forest through Feature Contribution and MDI Feature Importance
abstract
Deep forest is a non-differentiable deep model that has achieved impressive empirical success across a wide variety of applications, especially on categorical/symbolic or mixed modeling tasks. Many of the application fields prefer explainable models, such as random forests with feature contributions that can provide a local explanation for each prediction, and Mean Decrease Impurity (MDI) that can provide global feature importance. However, deep forest, as a cascade of random forests, possesses interpretability only at the first layer. From the second layer on, many of the tree splits occur on the new features generated by the previous layer, which makes existing explaining tools for random forests inapplicable. To disclose the impact of the original features in the deep layers, we design a calculation method with an estimation step followed by a calibration step for each layer, and propose our feature contribution and MDI feature importance calculation tools for deep forest. Experimental results on both simulated data and real-world data verify the effectiveness of our methods.
Yi-Xiao He, Shen-Huan Lyu, Yuan Jiang 0001
ACM Trans. Knowl. Discov. Data2
2026 Time-Efficient Identifying Key Tag Distribution in Large-Scale RFID Systems
abstract
With the proliferation of RFID-enabled applications, large-scale RFID systems often require multiple readers to ensure full coverage of numerous tags. In such systems, we sometimes pay more attention to a subset of tags instead of all, which are called key tags. This paper studies an under-investigated problemkey tag distribution identification, which aims to identify which key tags are beneath which readers. This is crucial for efficiently managing specific items of interest, which can quickly pinpoint key tags and help RFID readers covering these tags collaborate to improve the tag inventory efficiency. We propose a protocol called Kadept that identifies the key tag distribution by designing a sophisticated Cuckoo filter that teases out key tags as well as assigns each of them a singleton slot for response. With this design, a great number of trivial (non-key) tags will keep silent and free up bandwidth resources for key tags, and each key tag is sorted in a collision-free way and can be identified with only 1-bit response, which significantly improves the time efficiency. To enhance the scalability and efficiency of Kadept for high key tag proportions, we propose E-Kadept protocol, which accelerates the identification process by designing an incremental Cuckoo filter that reduces false positives and improves space efficiency. We theoretically analyze how to optimize protocol parameters of Kadept and E-Kadept, and conduct extensive simulations under different tag distribution scenarios. Compared with the state-of-the-art, E-Kadept can improve the time efficiency by a factor of 1.75×, when the ratio of key tags to all tags is 0.3.
Yanyan Wang 0001, Jia Liu 0008, Zhihao Qu, Shen-Huan Lyu, Bin Tang 0002
IEEE Trans. Mob. Comput.4
2025 Offline Model-Based Optimization by Learning to Rank
abstract
Offline model-based optimization (MBO) aims to identify a design that maximizes a black-box function using only a fixed, pre-collected dataset of designs and their corresponding scores. This problem has garnered significant attention from both scientific and industrial domains. A common approach in offline MBO is to train a regression-based surrogate model by minimizing mean squared error (MSE) and then find the best design within this surrogate model by different optimizers (e.g., gradient ascent). However, a critical challenge is the risk of out-of-distribution errors, i.e., the surrogate model may typically overestimate the scores and mislead the optimizers into suboptimal regions. Prior works have attempted to address this issue in various ways, such as using regularization techniques and ensemble learning to enhance the robustness of the model, but it still remains. In this paper, we argue that regression models trained with MSE are not well-aligned with the primary goal of offline MBO, which is to \textit{select} promising designs rather than to predict their scores precisely. Notably, if a surrogate model can maintain the order of candidate designs based on their relative score relationships, it can produce the best designs even without precise predictions. To validate it, we conduct experiments to compare the relationship between the quality of the final designs and MSE, finding that the correlation is really very weak. In contrast, a metric that measures order-maintaining quality shows a significantly stronger correlation. Based on this observation, we propose learning a ranking-based model that leverages learning to rank techniques to prioritize promising designs based on their relative scores. We show that the generalization error on ranking loss can be well bounded. Empirical results across diverse tasks demonstrate the superior performance of our proposed ranking-based method than twenty existing methods. Our implementation is available at \url{https://github.com/lamda-bbo/Offline-RaM}.
Rong-Xi Tan, Ke Xue 0001, Shen-Huan Lyu, Haopu Shang, Yaoyuan Wang, Sheng Fu, Chao Qian 0001
ICLR3
2025 Multi-Range Query in Commodity RFID Systems
abstract
Range Query (RQ) is to check whether there are any RFID tags with data beyond a given range. With about 46 billion RFID tags sold worldwide in 2023, time-efficient RQ becomes increasingly important for practical use, which can help users quickly pinpoint the target tags (if any) and give an early warning (e.g., fire alarm) to them for taking urgent actions and reducing the potential risk. However, existing work can deal with only a single range rather than multiple ranges that are very common in real-world applications. For example, foods in the refrigerator and the freezer have different temperature ranges for safe storing; treating them as one would probably give rise to query errors. In this paper, we study an under-investigated problem called multi-range query, which aims to achieve RQ in an RFID system with multiple query ranges. We propose a tailored protocol called anomalous tag identification (ATI) that quickly separates target tags from others and avoids querying all tags for saving communication overhead. In ATI, we design a fixedlength encoding vector together with standards-compliant select commands to deal with different ranges individually, without the need for any hardware modification. We implement the proposed protocols in commodity RFID systems. Experimental results show that ATI is superior to the baseline under different parameters, in terms of the time efficiency and space efficiency.
Yanyan Wang 0001, Jia Liu 0008, Zhihao Qu, Shen-Huan Lyu, Bin Tang 0002
IWQoS4
2025 Enhance learning efficiency of oblique decision tree via feature concatenation
Shen-Huan Lyu, Yi-Xiao He, Yanyan Wang 0001, Zhihao Qu, Bin Tang 0002
Inf. Sci.1
2024 The Role of Depth, Width, and Tree Size in Expressiveness of Deep Forest
abstract
Random forests are classical ensemble algorithms that construct multiple randomized decision trees and aggregate their predictions using naive averaging. Zhou and Feng [51] further propose a deep forest algorithm with multi-layer forests, which outperforms random forests in various tasks. The performance of deep forests is related to three hyperparameters in practice: depth, width, and tree size, but little has been known about its theoretical explanation. This work provides the first upper and lower bounds on the approximation complexity of deep forests concerning the three hyperparameters. Our results confirm the distinctive role of depth, which can exponentially enhance the expressiveness of deep forests compared with width and tree size. Experiments validate these theoretical findings. The detailed proof and code are available in the full version [31].
Shen-Huan Lyu, Jin-Hui Wu, Qin-Cheng Zheng
ECAI1
2024 Mask-Encoded Sparsification: Mitigating Biased Gradients in Communication-Efficient Split Learning
abstract
This paper introduces a novel framework designed to achieve a high compression ratio in Split Learning (SL) scenarios where resource-constrained devices are involved in large-scale model training. Our investigations demonstrate that compressing feature maps within SL leads to biased gradients that can negatively impact the convergence rates and diminish the generalization capabilities of the resulting models. Our theoretical analysis provides insights into how compression errors critically hinder SL performance, which previous methodologies underestimate. To address these challenges, we employ a narrow bit-width encoded mask to compensate for the sparsification error without increasing the order of time complexity. Supported by rigorous theoretical analysis, our framework significantly reduces compression errors and accelerates the convergence. Extensive experiments also verify that our method outperforms existing solutions regarding training efficiency and communication complexity. Our code can be found at https://github.com/BinaryMus/MaskSparsification.
Zhihao Qu, Shen-Huan Lyu, Miao Cai 0001
ECAI3
2024 Confidence-aware Contrastive Learning for Selective Classification
abstract
Selective classification enables models to make predictions only when they are sufficiently confident, aiming to enhance safety and reliability, which is important in high-stakes scenarios. Previous methods mainly use deep neural networks and focus on modifying the architecture of classification layers to enable the model to estimate the confidence of its prediction. This work provides a generalization bound for selective classification, disclosing that optimizing feature layers helps improve the performance of selective classification. Inspired by this theory, we propose to explicitly improve the selective classification model at the feature level for the first time, leading to a novel Confidence-aware Contrastive Learning method for Selective Classification, CCL-SC, which similarizes the features of homogeneous instances and differentiates the features of heterogeneous instances, with the strength controlled by the model's confidence. The experimental results on typical datasets, i.e., CIFAR-10, CIFAR-100, CelebA, and ImageNet, show that CCL-SC achieves significantly lower selective risk than state-of-the-art methods, across almost all coverage degrees. Moreover, it can be combined with existing methods to bring further improvement.
Yu-Chang Wu, Shen-Huan Lyu, Haopu Shang, Chao Qian 0001
ICML2
2024 Identifying Key Tag Distribution in Large-Scale RFID Systems
abstract
With the proliferation of RFID-enabled applications, multiple readers are required for the complete coverage of numerous tags in a large-scale RFID system. In this scenario, we sometimes pay more attention to a subset of tags instead of all, which are referred to as key tags. In this paper, we study an under-investigated problem key tag distribution identification, which aims to identify which key tags are beneath which readers. This is crucial for efficiently managing specific items of interest, which can quickly pinpoint key tags and help RFID readers covering these tags collaborate to improve the tag inventory efficiency. Since key tags typically make up a small part of all tags, it is time consuming to deal with all tags in the traditional way. We propose a protocol called Kadept that identifies the key tag distribution by using a sophisticatedly designed filter that teases out key tags as well as assigns each of them a singleton slot for response. With this design, a great number of trivial (non-key) tags will keep silent and free up bandwidth resources for key tags, and each key tag is sorted in a collision-free way and can be identified with only 1-bit response, which significantly improves the time efficiency. We theoretically analyze how can we optimize protocol parameters of Kadept and conduct extensive simulations under different tag distribution scenarios. Compared with the state-of-the-art, Kadept can improve the time efficiency by a factor of 3.7×, when the ratio of key tags to all tags is 0.1.
Yanyan Wang 0001, Jia Liu 0008, Shen-Huan Lyu, Zhihao Qu, Bin Tang 0002
IWQoS3
2024 Personalized Federated Learning with Feature Alignment via Knowledge Distillation
Guangfei Qi, Zhihao Qu, Shen-Huan Lyu, Ninghui Jia
PRICAI (2)3
2024 Multi-class imbalance problem: A multi-objective solution
Yi-Xiao He, Dan-Xuan Liu, Shen-Huan Lyu, Chao Qian 0001, Zhi-Hua Zhou
Inf. Sci.3
2023 On the Consistency Rate of Decision Tree Learning Algorithms
abstract
Decision tree learning algorithms such as CART are generally based on heuristics that maximizes the purity gain greedily. Though these algorithms are practically successful, theoretical properties such as consistency are far from clear. In this paper, we discover that the most serious obstacle encumbering consistency analysis for decision tree learning algorithms lies in the fact that the worst-case purity gain, i.e., the core heuristics for tree splitting, can be zero. Based on this recognition, we present a new algorithm, named Grid Classification And Regression Tree (GridCART), with a provable consistency rate $\mathcal{O}(n^{-1 / (d + 2)})$, which is the first consistency rate proved for heuristic tree learning algorithms.
Qin-Cheng Zheng, Shen-Huan Lyu, Shao-Qun Zhang, Yuan Jiang 0001, Zhi-Hua Zhou
AISTATS2
2022 Depth is More Powerful than Width with Prediction Concatenation in Deep Forest
abstract
Random Forest (RF) is an ensemble learning algorithm proposed by \citet{breiman2001random} that constructs a large number of randomized decision trees individually and aggregates their predictions by naive averaging. \citet{zhou2019deep} further propose Deep Forest (DF) algorithm with multi-layer feature transformation, which significantly outperforms random forest in various application fields. The prediction concatenation (PreConc) operation is crucial for the multi-layer feature transformation in deep forest, though little has been known about its theoretical property. In this paper, we analyze the influence of Preconc on the consistency of deep forest. Especially when the individual tree is inconsistent (as in practice, the individual tree is often set to be fully grown, i.e., there is only one sample at each leaf node), we find that the convergence rate of two-layer DF \textit{w.r.t.} the number of trees $M$ can reach $\mathcal{O}(1/M^2)$ under some mild conditions, while the convergence rate of RF is $\mathcal{O}(1/M)$. Therefore, with the help of PreConc, DF with deeper layer will be more powerful than the shallower layer. Experiments confirm theoretical advantages.
Shen-Huan Lyu, Yi-Xiao He, Zhi-Hua Zhou
NeurIPS1
2022 Improving generalization of deep neural networks by leveraging margin distribution
Shen-Huan Lyu, Lu Wang 0031, Zhi-Hua Zhou
Neural Networks1
2021 Improving Deep Forest by Exploiting High-order Interactions
abstract
Recent studies on deep forests have shown that deep learning frameworks can be built on non-differentiable modules without a backpropagation training process. However, the feature representations of deep forests only consist of predicted class probabilities. The information these class probabilities deliver is very limited and lacks diversity, especially when the number of output labels is far less than the number of input features. Besides, the prediction-based representations require us to save multiple layers of random forests to use them during testing, which is high-memory and high-time cost. In this paper, we propose a novel deep forest model that utilizes high-order interactions of input features to generate more informative and diverse feature representations. Specifically, we design a generalized version of Random Intersection Trees (gRIT) to discover stable high-order interactions and apply Activated Linear Combination (ALC) to transform them into hierarchical distributed representations. These interaction-based representations obviate the need to store random forests in the front layers, thus greatly improving the computational efficiency. Our experiments show that our method achieves highly competitive predictive performance with significantly reduced time and memory cost.
Yi-He Chen, Shen-Huan Lyu, Yuan Jiang 0001
ICDM2
2019 A Refined Margin Distribution Analysis for Forest Representation Learning
abstract
In this paper, we formulate the forest representation learning approach called \textsc{CasDF} as an additive model which boosts the augmented feature instead of the prediction. We substantially improve the upper bound of the generalization gap from $\mathcal{O}(\sqrt{\ln m/m})$ to $\mathcal{O}(\ln m/m)$, while the margin ratio of the margin standard deviation to the margin mean is sufficiently small. This tighter upper bound inspires us to optimize the ratio. Therefore, we design a margin distribution reweighting approach for deep forest to achieve a small margin ratio by boosting the augmented feature. Experiments confirm the correlation between the margin distribution and generalization performance. We remark that this study offers a novel understanding of \textsc{CasDF} from the perspective of the margin theory and further guides the layer-by-layer forest representation learning.
Shen-Huan Lyu, Zhi-Hua Zhou
NeurIPS1