Haipeng Shen

dblp:11/3852 · DBLP profile ↗
← Back
9ranked-venue papers
1as first author
2since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Probabilistic and Bayesian machine learning · 30% Kernel, tree and ensemble methods · 30% Learning theory · 30%
Theoretical computer science
1 paper
Information theory · 50% Mathematical optimization · 50%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Bioinformatics and computational biology · 100%

Topics — the 9 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › regression
functional regression
0.912025
Optimal Functional Bilinear Regression with Two-dimensional Functional Covariates via Reproducing Kernel Hilbert Space · J. Mach. Learn. Res. 2025
Machine learning › Learning theory › statistical estimation
minimax estimation
0.912025
Optimal Functional Bilinear Regression with Two-dimensional Functional Covariates via Reproducing Kernel Hilbert Space · J. Mach. Learn. Res. 2025
Machine learning › Kernel, tree and ensemble methods › kernel methods
reproducing kernel hilbert space
0.912025
Optimal Functional Bilinear Regression with Two-dimensional Functional Covariates via Reproducing Kernel Hilbert Space · J. Mach. Learn. Res. 2025
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction
principal component analysis
0.212016
A General Framework for Consistency of Principal Component Analysis · J. Mach. Learn. Res. 2016
Information theory
asymptotic analysis
0.212016
A General Framework for Consistency of Principal Component Analysis · J. Mach. Learn. Res. 2016
Mathematical optimization › high-dimensional statistics
spiked covariance model
0.212016
A General Framework for Consistency of Principal Component Analysis · J. Mach. Learn. Res. 2016
Bioinformatics and computational biology › gene expression analysis
microRNA expression analysis
0.212013
Poisson factor models with applications to non-normalized microRNA profiling · Bioinform. 2013
Bioinformatics and computational biology › biostatistics › statistical bioinformatics
statistical genomics
0.212013
Poisson factor models with applications to non-normalized microRNA profiling · Bioinform. 2013
Bioinformatics and computational biology
transcriptomics
0.212013
Poisson factor models with applications to non-normalized microRNA profiling · Bioinform. 2013

Methods — techniques the papers use, named apart from their topics

three-term penalty · 0.9generalized cross-validation · 0.9RKHS · 0.9random matrix theory · 0.5asymptotic framework · 0.5singular value decomposition · 0.2poisson factor model · 0.2normalization · 0.2
YearPublicationVenuePosition
2025 Optimal Functional Bilinear Regression with Two-dimensional Functional Covariates via Reproducing Kernel Hilbert Space
abstract
Traditional functional linear regression usually takes a one-dimensional functional predictor as input and estimates the continuous coefficient function. Modern applications often generate two-dimensional covariates, which become matrices when observed at grid points. To avoid the inefficiency of the classical method involving estimation of a two-dimensional coefficient function, we propose a functional bilinear regression model, and introduce an innovative three-term penalty to impose roughness penalty in the estimation. The proposed estimator exhibits minimax optimal property for prediction under the framework of reproducing kernel Hilbert space. An iterative generalized cross-validation approach is developed to choose tuning parameters, which significantly improves the computational efficiency over the traditional cross-validation approach. The statistical and computational advantages of the proposed method over existing methods are further demonstrated via simulated experiments, the Canadian weather data, and a biochemical long-range infrared light detection and ranging data.
Jianlong Shao, Haipeng Shen, Hongtu Zhu
J. Mach. Learn. Res.3
2024 A clinically actionable and explainable real-time risk assessment framework for stroke-associated pneumonia
Lutao Dai, Hao Li 0030, Xingquan Zhao, Zixiao Li, Haipeng Shen
Artif. Intell. Medicine9
2016 Efficient dimension reduction for high-dimensional matrix-valued data
Dong Wang 0010, Haipeng Shen, Young Truong
Neurocomputing2
2016 A General Framework for Consistency of Principal Component Analysis
abstract
A general asymptotic framework is developed for studying consistency properties of principal component analysis (PCA). Our framework includes several previously studied domains of asymptotics as special cases and allows one to investigate interesting connections and transitions among the various domains. More importantly, it enables us to investigate asymptotic scenarios that have not been considered before, and gain new insights into the consistency, subspace consistency and strong inconsistency regions of PCA and the boundaries among them. We also establish the corresponding convergence rate within each region. Under general spike covariance models, the dimension (or number of variables) discourages the consistency of PCA, while the sample size and spike information (the relative size of the population eigenvalues) encourage PCA consistency. Our framework nicely illustrates the relationship among these three types of information in terms of dimension, sample size and spike size, and rigorously characterizes how their relationships affect PCA consistency.
Dan Shen 0002, Haipeng Shen, J. S. Marron
J. Mach. Learn. Res.2
2013 Poisson factor models with applications to non-normalized microRNA profiling
abstract
MOTIVATION: Next-generation (NextGen) sequencing is becoming increasingly popular as an alternative for transcriptional profiling, as is the case for micro RNAs (miRNA) profiling and classification. miRNAs are a new class of molecules that are regulated in response to differentiation, tumorigenesis or infection. Our primary motivating application is to identify different viral infections based on the induced change in the host miRNA profile. Statistical challenges are encountered because of special features of NextGen sequencing data: the data are read counts that are extremely skewed and non-negative; the total number of reads varies dramatically across samples that require appropriate normalization. Statistical tools developed for microarray expression data, such as principal component analysis, are sub-optimal for analyzing NextGen sequencing data. RESULTS: We propose a family of Poisson factor models that explicitly takes into account the count nature of sequencing data and automatically incorporates sample normalization through the use of offsets. We develop an efficient algorithm for estimating the Poisson factor model, entitled Poisson Singular Value Decomposition with Offset (PSVDOS). The method is shown to outperform several other normalization and dimension reduction methods in a simulation study. Through analysis of an miRNA profiling experiment, we further illustrate that our model achieves insightful dimension reduction of the miRNA profiles of 18 samples: the extracted factors lead to more accurate and meaningful clustering of the cell lines. AVAILABILITY: The PSVDOS software is available on request.
Seonjoo Lee, Pauline E. Chugh, Haipeng Shen, R. Eberle, Dirk P. Dittmer
Bioinform.3
2007 On scalable measurement-driven modeling of traffic demand in large WLANs
abstract
Models of traffic demand are fundamental inputs to the design and engineering of data networks. In this paper we address this requirement in the context of large-scale wireless infrastructures using real measurement data from the University of North Carolina (UNC) wireless campus network. Our modeling effort focuses on capturing the demand variation in both the spatial and temporal domain in a way that scales well with the size of the wireless network. The network traffic dynamics are studied over two different week-long monitoring periods at various levels of spatial aggregation, from individual buildings to the whole network. We model traffic workload in terms of wireless sessions and network flows and find several modeling elements that are reusable in both temporal and spatial dimensions. The same set of parametric distributions for the session-and flow-related traffic variables capture the network traffic demand in both monitoring periods. Even more interestingly, these same distributions can characterize traffic dynamics at finer spatial scales, such as a single building or a group of buildings. We use our models to generate synthetic traffic and compare with trace data. The comparison clearly illustrates the trade-off between model scalability and reusability, on the one hand, and accuracy in capturing local-scale traffic dynamics on the other. Our main contribution is a novel behavioral approach for traffic demand modeling in large wireless networks that features high flexibility in the exploitation of the spatial and temporal resolution available in data traces.
Merkourios Karaliopoulos, Maria Papadopouli, Elias Raftopoulos, Haipeng Shen
LANMAN4
2007 Robust estimation of the self-similarity parameter in network traffic using wavelet transform
Haipeng Shen, Zhengyuan Zhu, Thomas C. M. Lee
Signal Process.1
2005 Modeling client arrivals at access points in wireless campus-wide networks
abstract
Our goal is to model the arrival of wireless clients at the access points (APs) in a production 802.11 infrastructure. Such models are critical for benchmarks, simulation studies, design of capacity planning and resource allocation, and the administration and support of wireless infrastructures. Our contributions include a novel methodology for modeling the arrival processes of clients at APs and the use of a powerful visualization tool for finding detailed interior features and quantile plots with simulation envelope for goodness-of-fit test. Time-varying Poisson processes can model well the arrival processes of clients at APs. We validate these results by modeling the visit arrivals at different time intervals and APs. Furthermore, we propose a clustering of the APs based on their visit arrival and functionality of the area in which these APs are located.
Maria Papadopouli, Haipeng Shen, Manolis Spanakis
LANMAN2
2005 Short-Term Traffic Forecasting in a Campus-Wide Wireless Network
abstract
Our goal is to characterize the traffic load in an IEEE802.11 infrastructure. This can be beneficial in many domains, including coverage planning, resource reservation, network monitoring for anomaly detection, and producing more accurate simulation models. The key issue that drives this study is traffic forecasting at each wireless access point (AP) in an hourly timescale. We conducted an extensive measurement study of wireless users on a major university campus using the IEEE802.11 wireless infrastructure. We propose several traffic models that take into account the periodicity and recent traffic history for each AP and present a time-series forecasting methodology. Finally, we build and evaluate these forecasting algorithms and discuss our findings.
Maria Papadopouli, Haipeng Shen, Elias Raftopoulos, Manolis Ploumidis, Félix Hernández-Campos
PIMRC2