Haojie Ren

dblp:163/7656 · DBLP profile ↗
← Back
13ranked-venue papers
3as first author
13since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 7 since 2021Computer networks · 3 · 2 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 TransLLM: A Unified Multi-Task Large Language Model for Urban Transportation via Learnable Prompting
abstract
Jiaming Leng, Yunying Bi, Chuan Qin, Zhenya Huang, Bing Yin, Haojie Ren, Yanyong Zhang, Chao Wang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Jiaming Leng, Yunying Bi, Chuan Qin 0002, Zhenya Huang, Haojie Ren, Yanyong Zhang, Chao Wang 0086
ACL (1)6
2026 Singular Value Decomposition-based lightweight LSTM for time series forecasting
Hao Ren 0005, Haojie Ren, Xiaojun Liang, Chunhua Yang 0001, Weihua Gui 0001
Future Gener. Comput. Syst.4
2025 Conformal Prediction with Cellwise Outliers: A Detect-then-Impute Approach
abstract
Conformal prediction is a powerful tool for constructing prediction intervals for black-box models, providing a finite sample coverage guarantee for exchangeable data. However, this exchangeability is compromised when some entries of the test feature are contaminated, such as in the case of cellwise outliers. To address this issue, this paper introduces a novel framework called *detect-then-impute conformal prediction*. This framework first employs an outlier detection procedure on the test feature and then utilizes an imputation method to fill in those cells identified as outliers. To quantify the uncertainty in the processed test feature, we adaptively apply the detection and imputation procedures to the calibration set, thereby constructing exchangeable features for the conformal prediction interval of the test label. We develop two practical algorithms, $\texttt{PDI-CP}$ and $\texttt{JDI-CP}$, and provide a distribution-free coverage analysis under some commonly used detection and imputation procedures. Notably, $\texttt{JDI-CP}$ achieves a finite sample $1-2\alpha$ coverage guarantee. Numerical experiments on both synthetic and real datasets demonstrate that our proposed algorithms exhibit robust coverage properties and comparable efficiency to the oracle baseline.
Yajie Bao, Haojie Ren, Zhaojun Wang, Changliang Zou
ICML3
2025 e-GAI: e-value-based Generalized α-Investing for Online False Discovery Rate Control
abstract
Online multiple hypothesis testing has attracted a lot of attention in many applications, e.g., anomaly status detection and stock market price monitoring. The state-of-the-art generalized $\alpha$-investing (GAI) algorithms can control online false discovery rate (FDR) on p-values only under specific dependence structures, a situation that rarely occurs in practice. The e-LOND algorithm (Xu & Ramdas, 2024) utilizes e-values to achieve online FDR control under arbitrary dependence but suffers from a significant loss in power as testing levels are derived from pre-specified descent sequences. To address these limitations, we propose a novel framework on valid e-values named e-GAI. The proposed e-GAI can ensure provable online FDR control under more general dependency conditions while improving the power by dynamically allocating the testing levels. These testing levels are updated not only by relying on both the number of previous rejections and the prior costs, but also, differing from the GAI framework, by assigning less $\alpha$-wealth for each rejection from a risk aversion perspective. Within the e-GAI framework, we introduce two new online FDR procedures, e-LORD and e-SAFFRON, and provide strategies for the long-term performance to address the issue of $\alpha$-death, a common phenomenon within the GAI framework. Furthermore, we demonstrate that e-GAI can be generalized to conditionally super-uniform p-values. Both simulated and real data experiments demonstrate the advantages of both e-LORD and e-SAFFRON in FDR control and power.
Zijian Wei, Haojie Ren, Changliang Zou
ICML3
2025 Ghost Points Matter: Far-Range Vehicle Detection with a Single mmWave Radar in Tunnel
abstract
Vehicle detection in tunnels is crucial for traffic monitoring and accident response, yet remains underexplored. In this paper, we develop mmTunnel, a millimeter-wave radar system that achieves far-range vehicle detection in tunnels. The main challenge here is coping with ghost points caused by multi-path reflections, which lead to severe localization errors and false alarms. Instead of merely removing ghost points, we propose correcting them to true vehicle positions by recovering their signal reflection paths, thus reserving more data points and improving detection performance, even in occlusion scenarios. However, recovering complex 3D reflection paths from limited 2D radar points is highly challenging. To address this problem, we develop a multi-path ray tracing algorithm that leverages the ground plane constraint and identifies the most probable reflection path based on signal path loss and spatial distance. We also introduce a curve-to-plane segmentation method to simplify tunnel surface modeling such that we can significantly reduce the computational delay and achieve real-time processing.
Chenming He, Chengzhen Meng, Xiaoran Fan, Dequan Wang, Haojie Ren, Jianmin Ji, Yanyong Zhang
MobiCom6
2025 UniSense: Spatial-Uncertainty-Aware Collaborative Sensing for Autonomous Driving
abstract
Vehicle-to-vehicle collaborative perception faces fundamental deployment barriers: raw LiDAR data sharing requires over 300 Mbps per vehicle - far exceeding V2X network capacities, while network delays of 80-200ms create dangerous temporal misalignments at highway speeds. We present UniSense, a distributed collaborative perception system that enables efficient and reliable multi-vehicle perception through uncertainty-driven sensor data exchange. Instead of sharing raw sensor data, vehicles exchange compact uncertainty maps that identify regions requiring additional perceptual information. Our key innovations include: (1) a lightweight uncertainty quantification pipeline that runs in real-time on automotive hardware, identifying perception-critical regions while reducing bandwidth requirements by more than 10×, (2) a bandwidth-aware protocol that dynamically adapts data sharing based on network conditions and perception uncertainty, and (3) a selective motion compensation scheme that maintains temporal consistency. We evaluate UniSense through a year-long deployment with 16 roadside LiDAR nodes and autonomous vehicles across our campus. Our experimental results show that UniSense extends reliable perception range from local 80m to 140m, improving accuracy by 1.33× on average, up to 1.73×, over the state-of-the-art baselines, under communication constraints. The code and dataset are available at https://github.com/LetStarFly/UniSense.
Haojie Ren, Wuyang Zhang, Shuyao Shi, Yanyong Zhang
MobiSys1
2025 CAP: A General Algorithm for Online Selective Conformal Prediction with FCR Control
abstract
This paper studies the problem of post-selection predictive inference in an online fashion. To avoid devoting resources to unimportant units, a preliminary selection of the current individual before reporting its prediction interval is common and meaningful in online predictive tasks. As a result of the temporal multiplicity introduced by the online selection process, it becomes essential to control the real-time false coverage-statement rate (FCR) which measures the overall miscoverage level. We develop a general framework named CAP (Calibration after Adaptive Pick), which performs an adaptive pick rule on historical data to construct a calibration set if the current individual is selected. This is followed by the output of a conformal prediction interval for the unobserved label. We present a series of tractable procedures for the construction of the calibration set for various popular online selection rules. It has been demonstrated that CAP achieves an exact selection-conditional coverage guarantee in the finite-sample and distribution-free regimes. In order to address the issue of the distribution shift in online data, we also embed CAP into some recent dynamic conformal prediction algorithms and prove that the proposed method can deliver long-run FCR control. Numerical results on both synthetic and real data corroborate that CAP can effectively control FCR around the target level and yield narrower prediction intervals than existing baselines across various settings.
Yajie Bao, Yuyang Huo, Haojie Ren, Changliang Zou
J. Mach. Learn. Res.3
2025 Online Multiple Changepoint Detection With False Discovery Rate Control
abstract
Technological advances have led to the emergence of an increasing number of applications requiring the analysis of datastreams, that are characterized by an indefinitely long and time-evolving sequence, particularly in the healthcare domain. In such applications, the status of a stream can alternate, possibly many times, between a regular status and an irregular status. Consequently, it is necessary to develop statistical methodologies that constantly detect multiple changepoints in an online manner. While we may employ conventional methods of sequential change detection to trigger signals after the change occurs, no online procedure is available to quantify the uncertainty of the detected changes. In this work, we fill this gap by framing online multiple changepoint detection into an online multiple testing problem and proposing a new framework to test the null hypothesis that there is no change between neighboring signalled points. To obtain valid p-values for online multiple testing, we propose a data-fission-based procedure that is a simple yet effective way of dealing with the post-detection uncertainty quantification. It is shown that popular online false discovery rate control methods with those p-values can achieve finite-sample false discovery rate control. We evaluate the proposed method in simulation studies. The method is applied to health monitoring dataset, alleviating the false alarm issue in online data analysis.
Haoyu Geng, Haojie Ren, Zhaojun Wang, Changliang Zou
IEEE Trans. Inf. Theory3
2024 ByMI: Byzantine Machine Identification with False Discovery Rate Control
abstract
Various robust estimation methods or algorithms have been proposed to hedge against Byzantine failures in distributed learning. However, there is a lack of systematic approaches to provide theoretical guarantees of significance in detecting those Byzantine machines. In this paper, we develop a general detection procedure, ByMI, via error rate control to address this issue, which is applicable to many robust learning problems. The key idea is to apply the sample-splitting strategy on each worker machine to construct a score statistic integrated with a general robust estimation and then to utilize the symmetry property of those scores to derive a data-driven threshold. The proposed method is dimension insensitive and p-value free with the help of the symmetry property and can achieve false discovery rate control under mild conditions. Numerical experiments on both synthetic and real data validate the theoretical results and demonstrate the effectiveness of our proposed method on Byzantine machine identification.
Chengde Qian, Haojie Ren, Changliang Zou
ICML3
2024 PhD Forum Abstract: Cooperative Perception System with Roadside Assistance
abstract
In this paper, we mainly focus on cooperative perception systems for vehicle-road coordination. Specifically, this paper encompasses two main aspects: 1) discussing the spatio-temporal synchronization issues among roadside multiple LiDARs. In this part, we design a method to synchronize the spatio-temporal data among multiple LiDARs by matching trajectory points between them; 2) designing a cooperative perception system based on uncertainty. In this part, we design a scheme to reduce the communication volume of cooperative perception by lowering the communication frequency.
Haojie Ren
IPSN1
2024 Real-Time Selection Under General Constraints via Predictive Inference
abstract
Real-time decision-making gets more attention in the big data era. Here, we consider the problem of sample selection in the online setting, where one encounters a possibly infinite sequence of individuals collected over time with covariate information available. The goal is to select samples of interest that are characterized by their unobserved responses until the user-specified stopping time. We derive a new decision rule that enables us to find more preferable samples that meet practical requirements by simultaneously controlling two types of general constraints: individual and interactive constraints, which include the widely utilized False Selection Rate (FSR), cost limitations, and diversity of selected samples. The key elements of our approach involve quantifying the uncertainty of response predictions via predictive inference and addressing individual and interactive constraints in a sequential manner. Theoretical and numerical results demonstrate the effectiveness of the proposed method in controlling both individual and interactive constraints.
Yuyang Huo, Haojie Ren, Changliang Zou
NeurIPS3
2023 TrajMatch: Toward Automatic Spatio-Temporal Calibration for Roadside LiDARs Through Trajectory Matching
abstract
Recently, deploying sensors such as LiDARs on the roadside to monitor the passing traffic and assist autonomous vehicle perception has become popular. However, unlike autonomous vehicle systems, roadside sensor systems involve sensors from different subsystems, resulting in a lack of synchronization in both time and space between the sensors. Calibration is a critical technology that enables the central server to fuse data generated by different location infrastructures, which vastly improves sensing range and detection robustness. Regrettably, existing calibration algorithms frequently assume that LiDARs have significant overlap or that temporal calibration has already been achieved. However, since these assumptions do not always hold in real-world scenarios, the calibration results obtained from existing algorithms are frequently unsatisfactory. In this paper, we propose TrajMatch - the first system that can automatically calibrate roadside LiDARs in both time and space. The main idea is to automatically calibrate the sensors based on the result of the detection/tracking task, rather than relying on extracting special features. Furthermore, we propose a novel mechanism for evaluating calibration parameters that align with our algorithm, and we demonstrate its effectiveness through experiments. This mechanism can also guide parameter iterations for multiple calibrations, further enhancing the accuracy and efficiency of our calibration method. Finally, to evaluate the performance of TrajMatch, we collected two datasets, one simulated dataset LiDARnet-sim 1.0 and one real-world dataset. The experimental results show that TrajMatch can achieve a spatial calibration error of less than$10cm$and a temporal calibration error of less than$1.5ms$.
Haojie Ren, Sha Zhang 0002, Sugang Li, Yao Li 0016, Xinchen Li, Jianmin Ji, Yu Zhang 0086, Yanyong Zhang
IEEE Trans. Intell. Transp. Syst.1
2022 AutoMS: Automatic Model Selection for Novelty Detection with Error Rate Control
abstract
Given an unsupervised novelty detection task on a new dataset, how can we automatically select a ''best'' detection model while simultaneously controlling the error rate of the best model? For novelty detection analysis, numerous detectors have been proposed to detect outliers on a new unseen dataset based on a score function trained on available clean data. However, due to the absence of labeled data for model evaluation and comparison, there is a lack of systematic approaches that are able to select a ''best'' model/detector (i.e., the algorithm as well as its hyperparameters) and achieve certain error rate control simultaneously. In this paper, we introduce a unified data-driven procedure to address this issue. The key idea is to maximize the number of detected outliers while controlling the false discovery rate (FDR) with the help of Jackknife prediction. We establish non-asymptotic bounds for the false discovery proportions and show that the proposed procedure yields valid FDR control under some mild conditions. Numerical experiments on both synthetic and real data validate the theoretical results and demonstrate the effectiveness of our proposed AutoMS method. The code is available at https://github.com/ZhangYifan1996/AutoMS.
Haojie Ren, Changliang Zou, Dejing Dou
NeurIPS3