VLDB 2026 Research / reviewers in the wild / expert
Jihong Li
dblp:01/4092
· DBLP profile ↗
24ranked-venue papers
2as first author
11since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 5 since 2021Computer networks · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021Systems, architecture and hardware · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Mutual Coupling-Aware 3D Non-Stationary Channel Modeling for TRIS Transceiver Systems
Kaiyang Ma, Lixiang Lian, Jihong Li, Haris Pervaiz, Guhan Zheng, Shunqing Zhang, Marco Di Renzo |
ICC | 3 |
| 2026 | An Alternating Directional Dual-RBF Approach for Joint Multi-BSs and Multi-RISs Deployment
Tao Yu 0008, Shunqing Zhang, Jihong Li, Kaixuan Huang, Wen Chen 0001, Qingqing Wu 0001 |
ICC | 4 |
| 2025 | Bayesian Model Comparison Based on Cross-Validated Estimation of F1 Measure
Yan Xue, Xuefei Cao, Xingli Yang, Jihong Li |
PRCV (1) | 5 |
| 2025 | Model Evaluation with Precision, Recall, and F1 Measure Based on Block-regularized m×2 Cross Validation for Text Corpus
Yan Xue, Xuefei Cao, Xingli Yang, Jihong Li |
PRICAI (4) | 5 |
| 2025 | A Unified QoS-Aware Multiplexing Framework for Next-Generation Immersive Communication With Legacy Wireless ApplicationsabstractImmersive communication, including emerging augmented reality, virtual reality, and holographic telepresence, has been identified as a key service for enabling next-generation wireless applications. To align with legacy wireless applications, such as enhanced mobile broadband or ultra-reliable low-latency communication, network slicing has been widely adopted. However, attempting to statistically isolate the above types of wireless applications through different network slices may lead to throughput degradation and increased queue backlog. To address these challenges, we establish a unified QoS-aware framework that supports immersive communication and legacy wireless applications simultaneously. Based on the Lyapunov drift theorem, we transform the original long-term throughput maximization problem into an equivalent short-term throughput maximization weighted by virtual queue length. Moreover, to cope with the challenges introduced by the interaction between large-timescale network slicing and short-timescale resource allocation, we propose an adaptive adversarial slicing (Ad2S) scheme for networks with invarying channel statistics. To track the network channel variations, we also propose a measurement extrapolation-Kalman filter (ME-KF)-based method and refine our scheme into Ad2S-non-stationary refinement (Ad2S-NR). Through extended numerical examples, we demonstrate that our proposed schemes achieve 3.86 Mbps throughput improvement and 63.96% latency reduction with 24.36% convergence time reduction. Within our framework, the trade-off between total throughput and user service experience can be achieved by tuning systematic parameters. Jihong Li, Shunqing Zhang, Tao Yu 0008, Guangjin Pan, Kaixuan Huang, Xiaojing Chen 0001, Yanzan Sun, Junyu Liu, Jiandong Li 0001, Derrick Wing Kwan Ng |
IEEE Internet Things J. | 1 |
| 2025 | Deep CNN Feature Resampling and Ensemble Based on Cross Validation for Image ClassificationabstractDeep convolutional neural networks (CNNs) such as AlexNet, VGGNet, ResNet, EfficientNet, and MobileNet have been extensively employed in image classification tasks. A common solution is directly feeding deep CNN features extracted from a deep network into a classification function. However, this solution may easily result in poor accuracy and robustness due to the single experimental result. One alternative is utilizing an ensemble of multiple deep networks. And this would bring very expensive, even unacceptable computational complexity. Thus, we propose a new deep CNN feature ensemble frame based on multiple cross validation resampling results of the single feature layer to cope with the above two issues. Theoretically, the proposed method is proved that having a smaller error rate than the single feature layer method and the same Rademacher complexity as the single feature layer method. Moreover, extensive experiments on several challenging image classification databases demonstrate the superiority of the proposed method. Yu Wang 0045, Xingli Yang, Jihong Li |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | Large-Scale Network Lifetime Inference Based on Universal Scaling FunctionabstractReliability evaluation of complex network is one of main topics in complex engineering systems, especially for Internet of Things (IoT). The reliability of IoT partially depends on its large-scale network. Especially, the lifetime distribution of large-scale network is critical for its health management. However, the large scale of the network usually leads to an expensive simulation time cost. Instead of direct simulation, we propose a method to infer the large-scale network lifetime using small-scale networks with the universal scaling function. We first find the scaling relationships between network lifetime and network size in a network model with failure coupling for two-dimensional square lattice network and Cayley tree network. Network lifetime with different size can be described by one universal scaling function. Then we perform theoretical analysis to derive the scaling relationships. Finally we apply these scaling relationships to wireless sensor network with coupled failures in more realistic situation as case study. From the simulation results of smaller-scale networks, we can infer the lifetime distribution of large-scale networks based on universal scaling function. The computation time and accuracy are compared with standard Monte Carlo simulation which shows that our method is faster and accurate. Our research shows that the proposed method using universal scaling functions can help us to infer the lifetime properties of large-scale networks with low computational cost. Our method can help fast reliability evaluations of large-scale complex networks with high accuracy. Shaobo Sui, Rui Peng 0001, Jihong Li, Mingyang Bai, Daqing Li |
IEEE Internet Things J. | 5 |
| 2024 | A Second-Order Noise Shaping SAR ADC With Parallel Multiresidual IntegratorabstractThis brief proposes a parallel multiresidual (PMR) integrator to enhance the noise-shaping (NS) effect for successive approximation register (SAR) analog-to-digital converter (ADC). The PMR employs passive integrators in parallel to simultaneously integrate the average result of the multiple sequential residual voltages. The proposed PMR technique provides an alternative scheme to enhance the NS rather than increasing the order of the integrator to suppress the instability and power. A prototype 7-bit second-order NS-SAR ADC is designed and simulated in a 130-nm CMOS process. PMR increases the effective number of bits (ENOBs) to 10.6 bit, which enhances the NS effect of 3.6 bit. It achieves a peak signal-to-noise and distortion ratio (SNDR) of 65.84 dB over a bandwidth of 1.3 kHz at the oversampling ratio (OSR) of 16. Longbin Zhu, Zhengtao Zhu, Risheng Su, Jianan Zheng, Siyuan Xie, Jihong Li, Fanyi Meng 0002, Zhijun Zhou, Keping Wang |
IEEE Trans. Very Large Scale Integr. Syst. | 8 |
| 2023 | We Need to Talk About Reproducibility in NLP Model ComparisonabstractNLPers frequently face reproducibility crisis in a comparison of various models of a realworld NLP task.Many studies have empirically showed that the standard splits tend to produce low reproducible and unreliable conclusions, and they attempted to improve the splits by using more random repetitions.However, the improvement on the reproducibility in a comparison of NLP models is limited attributed to a lack of investigation on the relationship between the reproducibility and the estimator induced by a splitting strategy.In this paper, we formulate the reproducibility in a model comparison into a probabilistic function with regard to a conclusion.Furthermore, we theoretically illustrate that the reproducibility is qualitatively dominated by the signal-tonoise ratio (SNR) of a model performance estimator obtained on a corpus splitting strategy.Specifically, a higher value of the SNR of an estimator probably indicates a better reproducibility.On the basis of the theoretical motivations, we develop a novel mixture estimator of the performance of an NLP model with a regularized corpus splitting strategy based on a blocked 3 × 2 cross-validation.We conduct numerical experiments on multiple NLP tasks to show that the proposed estimator achieves a high SNR, and it substantially increases the reproducibility.Therefore, we recommend the NLP practitioners to use the proposed method to compare NLP models instead of the methods based on the widely-used standard splits and the random splits with multiple repetitions. Yan Xue, Xuefei Cao, Xingli Yang, Yu Wang 0092, Ruibo Wang, Jihong Li |
EMNLP | 6 |
| 2023 | An Improved Cross-Validated Adversarial Validation Method
Zhengjiang Liu, Yan Xue, Ruibo Wang, Xuefei Cao, Jihong Li |
KSEM (1) | 6 |
| 2023 | Ensemble Feature Selection With Block-Regularized m × 2 Cross-ValidationabstractEnsemble feature selection (EFS) has attracted significant interest in the literature due to its great potential in reducing the discovery rate of noise features and stabilizing the feature selection results. In view of the superior performance of block-regularized m × 2 cross-validation on generalization performance and algorithm comparison, a novel EFS technology based on block-regularized m × 2 cross-validation is proposed in this study. Contrary to the traditional ensemble learning with a binomial distribution, the distribution of feature selection frequency in the proposed technique is approximated by a beta distribution more accurately. Furthermore, theoretical analysis of the proposed technique shows that it yields a higher selection probability for important features, lower selected risk for noise features, more true positives, and fewer false positives. Finally, the above conclusions are verified by the simulated and real data experiments. Xingli Yang, Yu Wang 0045, Ruibo Wang, Jihong Li |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2020 | Research on Creativity in the Age of IntelligenceabstractHow to meet the age of intelligence and train the innovative talents needed by the age of intelligence has become a research topic of the education circle. The characteristics of space fusion, man-machine fusion, knowledge explosion and focusing human nature in the age of intelligence put forward new requirements for innovative talents. Creative personality, critical thinking ability, digital learning ability, computational thinking, design thinking, and man-machine collaboration are essential components of creativity in the age of intelligence. Jihong Li, Huaibo Wang, Rongxia Zhuang, Ting-Wen Chang, Ronghuai Huang |
ICALT | 1 |
| 2020 | Research on differences of learning behaviors between students with different achievements in online learning environmentabstractIn recent years, the high dropout rate and low completion rate brought by MOOCs have attracted the attention of the researchers. To a certain extent, this study is providing a new perspective to reduce the problem of high school dropouts and low completion rates. A course of 44365 registered students in the spring of 2016 in the XuetangX platform was used as the analysis object. Lag Sequential Analysis (LSA) was used to analyze students with different levels of academic achievement to extract behavior sequence, find the regularity of the behavior patterns, and analyze the differences in behavior patterns. Huaibo Wang, Jihong Li, Ting-Wen Chang |
ICALT | 2 |
| 2020 | Tuning Parameter Selection Based on Blocked 3˟ 2 Cross-Validation for High-Dimensional Linear Regression Model
Xingli Yang, Yu Wang 0092, Ruibo Wang, Jihong Li |
Neural Process. Lett. | 5 |
| 2019 | Bayes Test of Precision, Recall, and F1 Measure for Comparison of Two Natural Language Processing ModelsabstractDirect comparison on point estimation of the precision (P), recall (R), and F1 measure of two natural language processing (NLP) models on a common test corpus is unreasonable and results in less replicable conclusions due to a lack of a statistical test. However, the existing t-tests in cross-validation (CV) for model comparison are inappropriate because the distributions of P, R, F1 are skewed and an interval estimation of P, R, and F1 based on a t-test may exceed [0,1]. In this study, we propose to use a block-regularized 3×2 CV (3×2 BCV) in model comparison because it could regularize the difference in certain frequency distributions over linguistic units between training and validation sets and yield stable estimators of P, R, and F1. On the basis of the 3×2 BCV, we calibrate the posterior distributions of P, R, and F1 and derive an accurate interval estimation of P, R, and F1. Furthermore, we formulate the comparison into a hypothesis testing problem and propose a novel Bayes test. The test could directly compute the probabilities of the hypotheses on the basis of the posterior distributions and provide more informative decisions than the existing significance t-tests. Three experiments with regard to NLP chunking tasks are conducted, and the results illustrate the validity of the Bayes test. Ruibo Wang, Jihong Li |
ACL (1) | 2 |
| 2019 | Block-regularized repeated learning-testing for estimating generalization error
Ruibo Wang, Jihong Li, Xingli Yang |
Inf. Sci. | 2 |
| 2019 | Calibrating GloVe model on the principle of Zipf's law
Xuefei Cao, Jihong Li, Ruibo Wang, Yu Wang 0092, Qian Niu, Junfeng Shi |
Pattern Recognit. Lett. | 2 |
| 2017 | Comparison of DC fault handling strategies for hybrid HVDC systemabstractThe hybrid high voltage direct current (Hybrid HVDC) system based on line commutated converter (LCC) and modular multilevel converter (MMC) is a feasible solution for long distance bulk power transmission. In order to handle the dc-side short-circuit fault, two kinds of methods can be adopted. The first one is to replace half-bridge sub-modules (HBSMs) with full-bridge SMs or clamp-double SMs which have the capability of dc fault clearance, and the second one is to arrange diodes at the dc port of MMC. In this paper, detailed processes of the dc fault solutions are described and discussed. The typical testing system is built in PSCAD/EMTDC, and the system performances under different strategies are compared. In addition, the non-block dc fault handling strategy based on full-bridge SMs is analyzed and tested. The results show that this kind of solution is not suitable for the bipolar HVDC system and the fault current cannot be blocked thoroughly. Jihong Li, Peng Qiu, Huangqing Xiao, Gaoren Liu, Zheng Xu 0011 |
IECON | 2 |
| 2017 | Block-Regularized m × 2 Cross-Validated Estimator of the Generalization ErrorabstractA cross-validation method based on [Formula: see text] replications of two-fold cross validation is called an [Formula: see text] cross validation. An [Formula: see text] cross validation is used in estimating the generalization error and comparing of algorithms' performance in machine learning. However, the variance of the estimator of the generalization error in [Formula: see text] cross validation is easily affected by random partitions. Poor data partitioning may cause a large fluctuation in the number of overlapping samples between any two training (test) sets in [Formula: see text] cross validation. This fluctuation results in a large variance in the [Formula: see text] cross-validated estimator. The influence of the random partitions on variance becomes serious as [Formula: see text] increases. Thus, in this study, the partitions with a restricted number of overlapping samples between any two training (test) sets are defined as a block-regularized partition set. The corresponding cross validation is called block-regularized [Formula: see text] cross validation ([Formula: see text] BCV). It can effectively reduce the influence of random partitions. We prove that the variance of the [Formula: see text] BCV estimator of the generalization error is smaller than the variance of [Formula: see text] cross-validated estimator and reaches the minimum in a special situation. An analytical expression of the variance can also be derived in this special situation. This conclusion is validated through simulation experiments. Furthermore, a practical construction method of [Formula: see text] BCV by a two-level orthogonal array is provided. Finally, a conservative estimator is proposed for the variance of estimator of the generalization error. Ruibo Wang, Yu Wang 0092, Jihong Li, Xingli Yang |
Neural Comput. | 3 |
| 2017 | Choosing Between Two Classification Learning Algorithms Based on Calibrated Balanced 5 × 2 Cross-Validated F-Test
Yu Wang 0092, Jihong Li |
Neural Process. Lett. | 2 |
| 2016 | Credible Intervals for Precision and Recall Based on a K-Fold Cross-Validated Beta DistributionabstractIn typical machine learning applications such as information retrieval, precision and recall are two commonly used measures for assessing an algorithm's performance. Symmetrical confidence intervals based on K-fold cross-validated t distributions are widely used for the inference of precision and recall measures. As we confirmed through simulated experiments, however, these confidence intervals often exhibit lower degrees of confidence, which may easily lead to liberal inference results. Thus, it is crucial to construct faithful confidence (credible) intervals for precision and recall with a high degree of confidence and a short interval length. In this study, we propose two posterior credible intervals for precision and recall based on K-fold cross-validated beta distributions. The first credible interval for precision (or recall) is constructed based on the beta posterior distribution inferred by all K data sets corresponding to K confusion matrices from a K-fold cross-validation. Second, considering that each data set corresponding to a confusion matrix from a K-fold cross-validation can be used to infer a beta posterior distribution of precision (or recall), the second proposed credible interval for precision (or recall) is constructed based on the average of K beta posterior distributions. Experimental results on simulated and real data sets demonstrate that the first credible interval proposed in this study almost always resulted in degrees of confidence greater than 95%. With an acceptable degree of confidence, both of our two proposed credible intervals have shorter interval lengths than those based on a corrected K-fold cross-validated t distribution. Meanwhile, the average ranks of these two credible intervals are superior to that of the confidence interval based on a K-fold cross-validated t distribution for the degree of confidence and are superior to that of the confidence interval based on a corrected K-fold cross-validated t distribution for the interval length in all 27 cases of simulated and real data experiments. However, the confidence intervals based on the K-fold and corrected K-fold cross-validated t distributions are in the two extremes. Thus, when focusing on the reliability of the inference for precision and recall, the proposed methods are preferable, especially for the first credible interval. Yu Wang 0092, Jihong Li |
Neural Comput. | 2 |
| 2015 | Measure for data partitioning in m × 2 cross-validation
Yu Wang 0092, Jihong Li |
Pattern Recognit. Lett. | 2 |
| 2015 | Confidence Interval for F1 Measure of Algorithm Performance Based on Blocked 3× 2 Cross-ValidationabstractIn studies on the application of machine learning such as Information Retrieval (IR), the focus is typically on the estimation of the F1measure of algorithm performance. Approximate symmetrical confidence intervals constructed by the F1value based on cross-validated L distribution are commonly used in the literature. However, theoretical analysis on the distribution of F1values shows that such distribution is actually non-symmetrical. Thus, simply using symmetrical distribution to approximate non-symmetrical distribution may be inappropriate and may result in a low degree of confidence and long interval length for the confidence interval. In the present study, a non-symmetrical confidence interval of the F1measure based on Beta prime distribution is constructed by using the F1value computed based on the average confusion matrix of a blocked 3 x 2 cross-validation. Experimental results show that in most cases, our method has high degrees of confidence. With an acceptable degree of confidence, our method has a shorter interval length than the approximate symmetrical confidence intervals based on the blocked 3 x 2 and 5 x 2 cross-validated L distributions. The approximate symmetrical confidence interval based on the 10-fold cross-validated L distribution has the shortest interval length of the four confidence intervals but with low degrees of confidence in all cases. Taking these two factors into consideration, our method is recommended. Yu Wang 0092, Jihong Li, Ruibo Wang, Xingli Yang |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2014 | Blocked 3×2 Cross-Validated t-Test for Comparing Supervised Classification Learning AlgorithmsabstractIn the research of machine learning algorithms for classification tasks, the comparison of the performances of algorithms is extremely important, and a statistical test of significance for generalization error is often used to perform it in the machine learning literature. In view of the randomness of partitions in cross-validation, a new blocked 3×2 cross-validation is proposed to estimate generalization error in this letter. We then conduct an analysis of variance of the blocked 3×2 cross-validated estimator. A relatively conservative variance estimator that considers the correlation between any two two-fold cross-validations, and was previously neglected in 5×2 cross-validated t and F-tests is put forward. A corresponding test using this variance estimator is presented to compare the performances of algorithms. Simulated results show that the performance of our test is comparable with that of 5×2 cross-validated tests but with less computation complexity. Yu Wang 0092, Ruibo Wang, Huichen Jia, Jihong Li |
Neural Comput. | 4 |