VLDB 2026 Research / reviewers in the wild / expert
Jiarui Cai
dblp:250/5594
· DBLP profile ↗
12ranked-venue papers
4as first author
11since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Joint Spatiotemporal-Frequency-Aware Feature Fusion for Vehicle Trajectory PredictionabstractTo guarantee rational decision-making and safety of intelligent transportation systems and autonomous driving, existing vehicle trajectory prediction (VTP) methods extract spatial and temporal features from complex traffic environments to achieve accurate forecast results. However, they generally do not use the frequency-domain information inherently embedded in vehicle trajectory data, resulting in lower prediction accuracy. To solve this problem, we propose a joint spatio-temporal-frequency-aware feature fusion (STFA-FF) method for VTP. Firstly, an intention recognition network is proposed to integrate spatial features and temporal features to infer driving intentions with high precision. Secondly, to fully utilize frequency-domain features, we present a multi-scale frequency-domain feature extraction (MSFDFE) module to map vehicle trajectory data into the frequency domain, incorporate the high-frequency attenuation mask to suppress high-frequency noise, and deeply integrate short-term variations with long-term trends. Additionally, a frequency-domain channel selection (FDCS) module is proposed to dynamically select key frequency channels related to driving modes. Furthermore, a multi-domain feature fusion prediction network is proposed to process the spatial, temporal and frequency-domain features to generate the final trajectory prediction results. Finally, experimental results demonstrate that the proposed method significantly outperforms mainstream approaches in prediction accuracy and robustness. Jiarui Cai, Kai Liu 0005, Yining Yue, Kaiquan Cai, Yanbo Zhu, Jiaqin Wang |
IEEE Internet Things J. | 1 |
| 2026 | InsQABench: Benchmarking Chinese insurance domain question answering with large language modelsabstractWe present InsQABench-the first comprehensive benchmark for evaluating LLMs’ capabilities in Chinese insurance QA. InsQABench comprises 95K carefully curated QA pairs derived from real-world insurance documents, covering 3 distinct tasks, 44 question types, and 55 specialized insurance topics. Our experiments evaluated and reported the performance of mainstream LLMs under both fine-tuned and zero-shot settings, demonstrating that fine-tuning on InsQABench can significantly improve model performance. We also introduced two frameworks that further enhanced task-specific performance, achieving 4.91% and 5.11% enhancement in accuracy over the next best-performing model. Binbin Lin 0002, Jiarui Cai, Xiaojin Zhang 0002, Zhongyu Wei, Wei Chen 0088 |
Inf. Process. Manag. | 5 |
| 2026 | A multi-factor decoupling repeat aware network for session-based recommendation
Huiying Wang, Qifeng Zhou, Jiarui Cai, Yihui Qiu |
Multim. Syst. | 3 |
| 2024 | Hyperbolic Learning with Synthetic Captions for Open-World DetectionabstractOpen-world detection poses significant challenges, as it requires the detection of any object using either object class labels or free-form texts. Existing related works often use large-scale manual annotated caption datasets for training, which are extremely expensive to collect. Instead, we propose to transfer knowledge from vision-language models (VLMs) to enrich the open-vocabulary descriptions au-tomatically. Specifically, we bootstrap dense synthetic captions using pretrained VLMs to provide rich descriptions on different regions in images, and incorporate these captions to train a novel detector that generalizes to novel concepts. To mitigate the noise caused by hallucination in syn-thetic captions, we also propose a novel hyperbolic vision-language learning approach to impose a hierarchy between visual and caption embeddings. We call our detector “Hy-perLearner”. We conduct extensive experiments on a wide variety of open-world detection benchmarks (COCO, LVIS, Object Detection in the Wild, RefCoCo) and our results show that our model consistently outperforms existing state-of-the-art methods, such as GLIP, GLIPv2 and Grounding DINO, when using the same backbone. Fanjie Kong, Yanbei Chen, Jiarui Cai, Davide Modolo |
CVPR | 3 |
| 2024 | Mitigating Bias of Deep Neural Networks for Trustworthy Traffic Perception in Autonomous SystemsabstractWith the rapid advancement of deep learning technology, feature extraction backbones that are effectively trained have found increasing use in various traffic perception tasks, such as vehicle recognition and roadway user detection and classification. However, given the naturally imbalanced distribution of objects in the real world, deep learning networks can inadvertently act as bias amplifiers, leading to unfair detection and classification outcomes. Addressing and quantifying this bias in traffic applications has thus become a pressing challenge. In response, this research introduces the first comprehensive traffic imbalance object recognition dataset tailored for autonomous vehicles, called the Autonomous-vehicle Long-tail Image Dataset (ALIDA). This dataset reflects real-world sample distribution and includes four categories—motorized users, non-motorized users, roadway facilities, and traffic signs—spanning 87 classes and totaling 37,558 images. Our experimental results confirm that these backbones may struggle to accurately recognize less common objects with limited training data, such as children and wheelchair users. To mitigate such biases and improve traffic perception equality, we introduce a DEbiased Traffic Object Recognition (DETOR) scheme. This scheme leverages both few-shot and representation learning techniques. Employing DETOR, the residual neural network achieved a 290% increase in accuracy for recognizing minority classes, such as children, motorcyclists, deer, and bears. This not only enhances the effectiveness but also significantly improves the fairness and scalability of traffic perception using deep neural networks. Hao (Frank) Yang, Yang Zhao 0013, Jiarui Cai, Meixin Zhu, Jenq-Neng Hwang, Yiran Chen 0001 |
IV | 3 |
| 2022 | LUNA: Localizing Unfamiliarity Near Acquaintance for Open-Set Long-Tailed RecognitionabstractThe predefined artificially-balanced training classes in object recognition have limited capability in modeling real-world scenarios where objects are imbalanced-distributed with unknown classes. In this paper, we discuss a promising solution to the Open-set Long-Tailed Recognition (OLTR) task utilizing metric learning. Firstly, we propose a distribution-sensitive loss, which weighs more on the tail classes to decrease the intra-class distance in the feature space. Building upon these concentrated feature clusters, a local-density-based metric is introduced, called Localizing Unfamiliarity Near Acquaintance (LUNA), to measure the novelty of a testing sample. LUNA is flexible with different cluster sizes and is reliable on the cluster boundary by considering neighbors of different properties. Moreover, contrary to most of the existing works that alleviate the open-set detection as a simple binary decision, LUNA is a quantitative measurement with interpretable meanings. Our proposed method exceeds the state-of-the-art algorithm by 4-6% in the closed-set recognition accuracy and 4% in F-measure under the open-set on the public benchmark datasets, including our own newly introduced fine-grained OLTR dataset about marine species (MS-LT), which is the first naturally-distributed OLTR dataset revealing the genuine genetic relationships of the classes. Jiarui Cai, Yizhou Wang 0005, Hung-Min Hsu, Jenq-Neng Hwang, Kelsey Magrane, Craig S. Rose |
AAAI | 1 |
| 2022 | MeMOT: Multi-Object Tracking with MemoryabstractWe propose an online tracking algorithm that performs the object detection and data association under a common framework, capable of linking objects after a long time span. This is realized by preserving a large spatio-temporal memory to store the identity embeddings of the tracked objects, and by adaptively referencing and aggregating useful information from the memory as needed. Our model, called MeMOT, consists of three main modules that are all Transformer-based: 1) Hypothesis Generation that produce object proposals in the current video frame; 2) Memory Encoding that extracts the core information from the memory for each tracked object; and 3) Memory Decoding that solves the object detection and data association tasks simultaneously for multi-object tracking. When evaluated on widely adopted MOT benchmark datasets, MeMOT observes very competitive performance. Jiarui Cai, Yuanjun Xiong, Wei Xia 0009, Zhuowen Tu, Stefano Soatto |
CVPR | 1 |
| 2022 | Traffic-Informed Multi-Camera Sensing (TIMS) System Based on Vehicle Re-IdentificationabstractSurveillance cameras are widely deployed traffic sensors, due to their affordable prices and being able to capture rich information. However, current surveillance systems have not been fully exploited: these cameras are isolated and can only extract information from their own fixed views. To enable a collaborative sensing system, we propose a novel framework called Traffic-Informed Multi-camera Sensing (TIMS) system for network-level traffic information estimation. By pushing multi-camera Re-IDentification (ReID) workflow towards network-wide traffic information extraction, TIMS system integrates a customized metric-learning vision-based vehicle ReID method (TIM-ReID) and establishing a traffic-informed workflow. To integrate the traffic network connection information along with visual and vehicle attributes features, the road network is extracted as a weighted graph through the Spatial-temporal Camera Graph Inference Model (StCGIM) and serves for matching and re-ranking ReID candidates. Moreover, an Accuracy Model (AAM) is designed to provide accurate, reliable and comprehensive traffic information estimation, including both the values and distribution of parameters under a high penetration rate. In experiments based on real-world multi-camera datasets captured in the city of Seattle, the customized TIM-ReID outperforms existing state-of-the-art methods, and delivers accurate cross-camera information estimation, whose value error is less than 8% and the Kullback-Leibler (KL) distance between the estimated and real distribution is less than 3.42 among all the evaluated camera pairs. TIMS system empowers cameras to work collaboratively through an interactive brain, and provides users with valuable and comprehensive traffic information. Hao (Frank) Yang, Jiarui Cai, Meixin Zhu, Yinhai Wang |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2021 | ACE: Ally Complementary Experts for Solving Long-Tailed Recognition in One-ShotabstractOne-stage long-tailed recognition methods improve the overall performance in a "seesaw" manner, i.e., either sacrifice the head’s accuracy for better tail classification or elevate the head’s accuracy even higher but ignore the tail. Existing algorithms bypass such trade-off by a multi-stage training process: pre-training on imbalanced set and fine-tuning on balanced set. Though achieving promising performance, not only are they sensitive to the generalizability of the pre-trained model, but also not easily integrated into other computer vision tasks like detection and segmentation, where pre-training of classifiers solely is not applicable. In this paper, we propose a one-stage long-tailed recognition scheme, ally complementary experts (ACE), where the expert is the most knowledgeable specialist in a sub-set that dominates its training, and is complementary to other experts in the less-seen categories without being disturbed by what it has never seen. We design a distribution-adaptive optimizer to adjust the learning pace of each expert to avoid over-fitting. Without special bells and whistles, the vanilla ACE outperforms the current one-stage SOTA method by 3 ~ 10% on CIFAR10-LT, CIFAR100-LT, ImageNet-LT and iNaturalist datasets. It is also shown to be the first one to break the "seesaw" trade-off by improving the accuracy of the majority and minority categories simultaneously in only one stage. Code and trained models are at https://github.com/jrcai/ACE. Jiarui Cai, Yizhou Wang 0005, Jenq-Neng Hwang |
ICCV | 1 |
| 2021 | ROD2021 Challenge: A Summary for Radar Object Detection Challenge for Autonomous Driving ApplicationsabstractThe Radar Object Detection 2021 (ROD2021) Challenge, held in the ACM International Conference on Multimedia Retrieval (ICMR) 2021, has been introduced to detect and classify objects purely using an FMCW radar for autonomous driving applications. As a robust sensor to all-weather conditions, radar has rich information hidden in the radio frequencies, which can potentially achieve object detection and classification. This insight will provide a new object perception solution for an autonomous vehicle even in adverse driving scenarios. The ROD2021 Challenge is the first public benchmark focusing on this topic, which attracts great attention and participation. There are more than 260 participants among 37 teams from more than 10 countries with different academic and industrial affiliations, contributing about 300 submissions in the first phase and 400 submissions in the second phase. The final performance is evaluated by average precision (AP). Results add strong value and a better understanding of the radar object detection task for the autonomous vehicle community. Yizhou Wang 0005, Jenq-Neng Hwang, Gaoang Wang, Hui Liu 0011, Kwang-Ju Kim, Hung-Min Hsu, Jiarui Cai, Haotian Zhang 0005, Zhongyu Jiang, Renshu Gu |
ICMR | 7 |
| 2021 | Multi-Target Multi-Camera Tracking of Vehicles Using Metadata-Aided Re-ID and Trajectory-Based Camera Link ModelabstractIn this paper, we propose a novel framework for multi-target multi-camera tracking (MTMCT) of vehicles based on metadata-aided re-identification (MA-ReID) and the trajectory-based camera link model (TCLM). Given a video sequence and the corresponding frame-by-frame vehicle detections, we first address the isolated tracklets issue from single camera tracking (SCT) by the proposed traffic-aware single-camera tracking (TSCT). Then, after automatically constructing the TCLM, we solve MTMCT by the MA-ReID. The TCLM is generated from camera topological configuration to obtain the spatial and temporal information to improve the performance of MTMCT by reducing the candidate search of ReID. We also use the temporal attention model to create more discriminative embeddings of trajectories from each camera to achieve robust distance measures for vehicle ReID. Moreover, we train a metadata classifier for MTMCT to obtain the metadata feature, which is concatenated with the temporal attention based embeddings. Finally, the TCLM and hierarchical clustering are jointly applied for global ID assignment. The proposed method is evaluated on the CityFlow dataset, achieving IDF1 76.77%, which outperforms the state-of-the-art MTMCT methods. Hung-Min Hsu, Jiarui Cai, Yizhou Wang 0005, Jenq-Neng Hwang, Kwang-Ju Kim |
IEEE Trans. Image Process. | 2 |
| 2020 | Where is my infusion pump? Harnessing network dynamics for improved hospital equipment fleet managementabstractOBJECTIVE: Timely availability of intravenous infusion pumps is critical for high-quality care delivery. Pumps are shared among hospital units, often without central management of their distribution. This study seeks to characterize unit-to-unit pump sharing and its impact on shortages, and to evaluate a system-control tool that balances inventory across all care areas, enabling increased availability of pumps. MATERIALS AND METHODS: A retrospective study of 3832 pumps moving in a network of 5292 radiofrequency and infrared sensors from January to November 2017 at The Johns Hopkins Hospital in Baltimore, Maryland. We used network analysis to determine whether pump inventory in one unit was associated with inventory fluctuations in others. We used a quasi-experimental design and segmented regressions to evaluate the effect of the system-control tool on enabling safe inventory levels in all care areas. RESULTS: We found 93 care areas connected through 67,111 pump transactions and 4 discernible clusters of pump sharing. Up to 17% (95% confidence interval, 7%-27%) of a unit's pump inventory was explained by the inventory of other units within its cluster. The network analysis supported design and deployment of a hospital-wide inventory balancing system, which resulted in a 44% (95% confidence interval, 36%-53%) increase in the number of care areas above safe inventory levels. CONCLUSIONS: Network phenomena are essential inputs to hospital equipment fleet management. Consequently, benefits of improved inventory management in strategic unit(s) are capable of spreading safer inventory levels throughout the hospital. Diego A. Martinez, Jiarui Cai, Jimi Oke, Andrew S. Jarrell, Felipe Feijoo, Jeffrey Appelbaum, Eili Y. Klein, Sean Barnes, Scott R. Levin |
J. Am. Medical Informatics Assoc. | 2 |