Jiajing Chen

dblp:148/5615 · DBLP profile ↗
← Back
20ranked-venue papers
7as first author
15since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 5 first-author · 7 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Computer networks · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Intelligent Reinforcement-Learning Routing Protocol With Integrated Power Control for Underwater Acoustic Sensor Networks
Jianmin Yang, Jiajing Chen, Tongwei Zhang, Guangjie Han
IEEE Internet Things J.4
2025 Learning-based robust direction-of-arrival estimation with array imperfections
Jiajing Chen, Yixin Jiang, Qingjiang Shi, Xiang Cheng 0001, Xuesong Cai
Signal Process.1
2025 A Scheme of Gradient-Based Iterative Identification With Separable Synchronous Variable-Innovation for Exp-ARX Systems With Arbitrary Data Loss
abstract
This article addresses the parameter identification of the Exp-ARX system with arbitrary data loss. To effectively tackle the challenges associated with data loss to enhance data utilization and identification accuracy, a novel iterative method for constructing output predictions is proposed, which operates in multiple steps starting from the most recently previous output value. To reduce the computational complexity caused by numerous characteristic parameters and further improve the accuracy of parameter estimation, an iterative algorithm based on separable synchronization is proposed by using the hierarchical principle. The key problem of this algorithm is to transform the original system into two subsystems and estimate their parameters separately. For each subsystem, a gradient-based iterative algorithm with variable innovation is introduced to adjust the length of innovation data adaptively. The performance of the proposed algorithm has been verified through simulation and applications, demonstrating significant improvements in handling data loss and parameter estimation accuracy.
Ya Gu, Jiajing Chen, Yonghong Tan 0001
IEEE Trans. Ind. Informatics2
2024 Letting 3D Guide the Way: 3D Guided 2D Few-Shot Image Classification
abstract
Existing few-shot image classification networks aim to perform prediction on images belonging to classes that were not seen during training, with only a few labeled images, which are randomly picked from the same image pool as the support set. However, this traditional approach has two main issues: (i) in real-world applications, since support images are randomly picked, the angle they were captured from can be very different from that of the query image, causing the images to look very different and making it hard to match them; (ii) since support and query images, for both training and testing, are sampled from the same image pool, models can overfit the dataset, especially if the image pool contains images with similar color, texture or view angle. Thus, good performance on a dataset does not reflect a model’s real ability. To address these issues, we propose a novel few-shot learning approach referred to as the 3D guided 2D (3DG2D) few-shot image classification. In our proposed approach, the queries are 2D images, and the support set is composed of 3D mesh data, providing different views of an object, in contrast to randomly picked images providing a single view. From each 3D mesh, 14 projection images are generated from different angles. Thus, these projections have significant variance among themselves. To address this challenge, we also propose the Angle Inference Module (AIM), which is used to infer the view angle of a query image so that more attention is given to projection images corresponding to the same view angle as the query image to achieve better prediction performance. We perform experiments on ModelNet40, Toys4K and ShapeNet datasets with 4-fold cross validation, and show that our 3DG2D few-shot classification approach consistently outperforms the state-of-the-art baselines.
Jiajing Chen, Minmin Yang, Senem Velipasalar
WACV1
2024 Rethinking the Evaluation of Driver Behavior Analysis Approaches
abstract
Crashes caused by distracted driving result in more than 3000 deaths every year in the U.S. Distracted driver behavior detection is instrumental for driver assist systems. Researchers have focused on autonomously detecting distracted driver behavior so that drivers can be alerted in time to reduce the risk of crashes. Despite the large number of approaches presented in the literature, there are still issues related to proper performance evaluation, reproducibility and lack of or very slow adoption of these approaches by the transportation industry. Most existing approaches either do not provide documented and usable codes or use private datasets, or do not present the experiment details, such as data split, sometimes resulting in inflated accuracy numbers. Moreover, these factors also make many results not reproducible. In addition, the performance metrics should be chosen carefully to measure various aspects of different methods, including their generalizability, and action localization ability in time. In this work, we perform a commensurate comparison of different state-of-the-art methods by using different data splits and performance metrics on the StateFarm distracted driving and AI CITY Challenge datasets. With the data split experiments, we highlight the importance of leave-N-driver-out cross validation, since these models should perform well in real-world testing with never-before-seen drivers. The results show the importance of data splitting and the performance metric for the comparison and evaluation of different methods, and their significant effects on the results.
Weiheng Chai, Jiyang Wang, Jiajing Chen, Senem Velipasalar, Anuj Sharma 0001
IEEE Trans. Intell. Transp. Syst.3
2024 Vision-Language Models Can Identify Distracted Driver Behavior From Naturalistic Videos
abstract
Recognizing the activities causing distraction in real-world driving scenarios is critical for ensuring the safety and reliability of both drivers and pedestrians on the roadways. Conventional computer vision techniques are typically data-intensive and require a large volume of annotated training data to detect and classify various distracted driving behaviors, thereby limiting their generalization ability, efficiency and scalability. We aim to develop a generalized framework that showcases robust performance with access to limited or no annotated training data. Recently, vision-language models have offered large-scale visual-textual pretraining that can be adapted to task-specific learning like distracted driving activity recognition. Vision-language pretraining models like CLIP have shown significant promise in learning natural language-guided visual representations. This paper proposes a CLIP-based driver activity recognition approach that identifies driver distraction from naturalistic driving images and videos. CLIP’s vision embedding offers zero-shot transfer and task-based finetuning, which can classify distracted activities from naturalistic driving video. Our results show that this framework offers state-of-the-art performance on zero-shot transfer, finetuning and video-based models for predicting the driver’s state on four public datasets. We propose frame-based and video-based frameworks developed on top of the CLIP’s visual representation for distracted driving detection and classification tasks and report the results. Our code is available at https://github.com/zahid-isu/DriveCLIP
Md. Zahid Hasan, Jiajing Chen, Jiyang Wang, Mohammed Shaiqur Rahman, Ameya Joshi, Senem Velipasalar, Chinmay Hegde, Anuj Sharma 0001, Soumik Sarkar
IEEE Trans. Intell. Transp. Syst.2
2023 ViewNet: A Novel Projection-Based Backbone with View Pooling for Few-shot Point Cloud Classification
abstract
Although different approaches have been proposed for 3D point cloud-related tasks, few-shot learning (FSL) of 3D point clouds still remains under-explored. In FSL, un-like traditional supervised learning, the classes of training and test data do not overlap, and a model needs to rec-ognize unseen classes from only a few samples. Existing FSL methods for 3D point clouds employ point-based models as their backbone. Yet, based on our extensive experiments and analysis, we first show that using a point-based backbone is not the most suitable FSL approach, since (i) a large number of points' features are discarded by the max pooling operation used in 3D point-based backbones, decreasing the ability of representing shape information; (ii) point-based backbones are sensitive to occlusion. To address these issues, we propose employing a projection-and 2D Convolutional Neural Network-based backbone, referred to as the ViewNet, for FSL from 3D point clouds. Our approach first projects a 3D point cloud onto six different views to alleviate the issue of missing points. Also, to generate more descriptive and distinguishing features, we propose View Pooling, which combines different projected plane combinations into five groups and performs max-pooling on each of them. The experiments performed on the ModelNet40, ScanObjectNN and ModelNet40-C datasets, with cross validation, show that our method consistently outperforms the state-of-the-art baselines. Moreover, compared to traditional image classification backbones, such as ResNet, the proposed ViewNet can extract more distinguishing features from multiple views of a point cloud. We also show that ViewNet can be used as a backbone with different FSL heads and provides improved performance compared to traditionally used backbones.
Jiajing Chen, Minmin Yang, Senem Velipasalar
CVPR1
2023 ToThePoint: Efficient Contrastive Learning of 3D Point Clouds via Recycling
abstract
Recent years have witnessed significant developments in point cloud processing, including classification and segmentation. However, supervised learning approaches need a lot of well-labeled data for training, and annotation is labor-and time-intensive. Self-supervised learning, on the other hand, uses unlabeled data, and pretrains a back-bone with a pretext task to extract latent representations to be used with the downstream tasks. Compared to 2D images, self-supervised learning of 3D point clouds is under-explored. Existing models, for self-supervised learning of 3D point clouds, rely on a large number of data samples, and require significant amount of computational re-sources and training time. To address this issue, we propose a novel contrastive learning approach, referred to as To ThePoint. Different from traditional contrastive learning methods, which maximize agreement between features obtained from a pair of point clouds formed only with different types of augmentation, ToThePoint also maximizes the agreement between the permutation invariant features and features discarded after max pooling. We first perform self-supervised learning on the ShapeNet dataset, and then evaluate the performance of the network on different downstream tasks. In the downstream task experiments, performed on the ModelNet40, ModelNet40C, ScanobjectNN and ShapeNet-Part datasets, our proposed ToThe-Point achieves competitive, if not better results compared to the state-of-the-art baselines, and does so with significantly less training time (200 times faster than baselines).
Xinglin Li, Jiajing Chen, Jinhui Ouyang, Hanhui Deng, Senem Velipasalar, Di Wu 0002
CVPR2
2023 Cross-Modality Feature Fusion Network for Few-Shot 3D Point Cloud Classification
abstract
Recent years have witnessed significant progress in the field of few-shot image classification while few-shot 3D point cloud classification still remains under-explored. Real-world 3D point cloud data often suffers from occlusions, noise and deformation, which make the few-shot 3D point cloud classification even more challenging. In this paper, we propose a cross-modality feature fusion network, for few-shot 3D point cloud classification, which aims to recognize an object given only a few labeled samples, and provides better performance even with point cloud data with missing points. More specifically, we train two models in parallel. One is a projection-based model with ResNet18 as the backbone and the other one is a point-based model with a DGCNN backbone. Moreover, we design a Support-Query Mutual Attention (sqMA) module to fully exploit the correlation between support and query features. Extensive experiments on three datasets, namely ModelNet40, ModelNet40-C and ScanObjectNN, show the effectiveness of our method, and its robustness to missing points. Our proposed method outperforms different state-of-the-art baselines on all datasets. The margin of improvement is even larger on the ScanObjectNN dataset, which is collected from real-world scenes and is more challenging with objects having missing points.
Minmin Yang, Jiajing Chen, Senem Velipasalar
WACV2
2023 Gender and country biases in Wikipedia citations to scholarly publications
abstract
Abstract Ensuring Wikipedia cites scholarly publications based on quality and relevancy without biases is critical to credible and fair knowledge dissemination. We investigate gender‐ and country‐based biases in Wikipedia citation practices using linked data from the Web of Science and a Wikipedia citation dataset. Using coarsened exact matching, we show that publications by women are cited less by Wikipedia than expected, and publications by women are less likely to be cited than those by men. Scholarly publications by authors affiliated with non‐Anglosphere countries are also disadvantaged in getting cited by Wikipedia, compared with those by authors affiliated with Anglosphere countries. The level of gender‐ or country‐based inequalities varies by research field, and the gender‐country intersectional bias is prominent in math‐intensive STEM fields. To ensure the credibility and equality of knowledge presentation, Wikipedia should consider strategies and guidelines to cite scholarly publications independent of the gender and country of authors.
Jiajing Chen, Erjia Yan, Chaoqun Ni
J. Assoc. Inf. Sci. Technol.2
2023 Driver Head Pose Detection From Naturalistic Driving Data
abstract
Driver behavior analysis plays an important role in driver assistance systems. A driver’s face and head pose hold the key towards understanding whether the driver’s attention and concentration are on the road while driving. Naturalistic driving studies (NDS) allow observing drivers in real-time under naturalistic traffic conditions. Yet, data collected in NDS often comprise low-resolution videos usually with more challenging camera positions compared to controlled studies. For instance, when the camera is not directly facing the driver, classifying head pose becomes more challenging, since the variation between different classes becomes much smaller. In this paper, we propose three different approaches to classify a driver’s head pose from naturalistic videos, which were captured by a camera providing a side view, instead of directly facing the driver. These approaches employ a sequence of five key points on the driver’s face. We compare these three proposed approaches with each other as well as with three different baselines by using leave-one-driver-out cross-validation on nine different drivers. Results show that our proposed method employing a Bidirectional Gated Recurrent Unit (BiGRU) outperforms the best performing baseline by 11% in terms of overall accuracy.
Weiheng Chai, Jiajing Chen, Jiyang Wang, Senem Velipasalar, Archana Venkatachalapathy, Yaw Adu-Gyamfi, Jennifer Merickel, Anuj Sharma 0001
IEEE Trans. Intell. Transp. Syst.2
2022 Why Discard if You can Recycle?: A Recycling Max Pooling Module for 3D Point Cloud Analysis
abstract
In recent years, most 3D point cloud analysis models have focused on developing either new network architectures or more efficient modules for aggregating point features from a local neighborhood. Regardless of the network architecture or the methodology used for improved feature learning, these models share one thing, which is the use of max-pooling in the end to obtain permutation invariant features. We first show that this traditional approach causes only a fraction of 3D points contribute to the permutation-invariant features, and discards the rest of the points. In order to address this issue and improve the performance of any baseline 3D point classification or segmentation model, we propose a new module, referred to as the Recycling Max-Pooling (RMP) module, to recycle and utilize the features of some of the discarded points. We incorporate a refinement loss that uses the recycled features to refine the prediction loss obtained from the features kept by traditional max-pooling. To the best of our knowledge, this is the first work that explores recycling of still useful points that are traditionally discarded by max-pooling. We demonstrate the effectiveness of the proposed RMP module by incorporating it into several milestone baselines and state-of-the-art networks for point cloud classification and indoor semantic segmentation tasks. We show that RPM, without any bells and whistles, consistently improves the performance of all the tested networks by using the same base network implementation and hyper-parameters. The code is provided in the supplementary material.
Jiajing Chen, Burak Kakillioglu, Huantao Ren, Senem Velipasalar
CVPR1
2022 Gaitpoint: A Gait Recognition Network Based on Point Cloud Analysis
abstract
We propose a novel gait recognition method that combines convolutional features with features of human pose key points obtained by a point cloud analysis model. Currently, most state-of-the-art works on gait recognition rely on only images and are purely based on convolutional neural networks. Most of these methods are very sensitive to small variations in the appearance of a walking person. For instance, if a person wears a coat or carries a bag, the accuracy of these methods may drop significantly. To address this problem, we propose to treat a sequence of human key points as a point cloud and combine human key point features and convolution feature map for final prediction. The experimental results show the promise of this approach, which outperforms three state-oft-he-art baselines in all walking scenarios, including the ones involving heavy clothing or carried items.
Jiajing Chen, Huantao Ren, Frank Sicong Chen, Senem Velipasalar, Vir V. Phoha
ICIP1
2022 Background-Aware 3-D Point Cloud Segmentation With Dynamic Point Feature Aggregation
abstract
With the proliferation of Lidar sensors and 3D vision cameras, 3D point cloud analysis has attracted significant attention in recent years. After the success of the pioneer work PointNet, deep learning-based methods have been increasingly applied to various tasks, including 3D point cloud segmentation and 3D object classification. In this paper, we propose a novel 3D point cloud learning network, referred to as Dynamic Point Feature Aggregation Network (DPFA-Net), by selectively performing the neighborhood feature aggregation with dynamic pooling and an attention mechanism. DPFA-Net has two variants for semantic segmentation and classification of 3D point clouds. As the core module of the DPFA-Net, we propose a Feature Aggregation layer, in which features of the dynamic neighborhood of each point are aggregated via a self-attention mechanism. In contrast to other segmentation models, which aggregate features from fixed neighborhoods, our approach can aggregate features from different neighbors in different layers providing a more selective and broader view to the query points, and focusing more on the relevant features in a local neighborhood. In addition, to further improve the performance of the proposed semantic segmentation model, we present two novel approaches, namely Two-Stage BF-Net and BF-Regularization to exploit the background-foreground information. Experimental results show that the proposed DPFA-Net achieves the state-of-the-art overall accuracy score for semantic segmentation on the S3DIS dataset, and provides a consistently satisfactory performance across different tasks of semantic segmentation, part segmentation, and 3D object classification. It is also computationally more efficient compared to other methods.
Jiajing Chen, Burak Kakillioglu, Senem Velipasalar
IEEE Trans. Geosci. Remote. Sens.1
2021 Feasibility of capturing real-world data from health information technology systems at multiple centers to assess cardiac ablation device outcomes: A fit-for-purpose informatics analysis report
abstract
OBJECTIVE: The study sought to conduct an informatics analysis on the National Evaluation System for Health Technology Coordinating Center test case of cardiac ablation catheters and to demonstrate the role of informatics approaches in the feasibility assessment of capturing real-world data using unique device identifiers (UDIs) that are fit for purpose for label extensions for 2 cardiac ablation catheters from the electronic health records and other health information technology systems in a multicenter evaluation. MATERIALS AND METHODS: We focused on data capture and transformation and data quality maturity model specified in the National Evaluation System for Health Technology Coordinating Center data quality framework. The informatics analysis included 4 elements: the use of UDIs for identifying device exposure data, the use of standardized codes for defining computable phenotypes, the use of natural language processing for capturing unstructured data elements from clinical data systems, and the use of common data models for standardizing data collection and analyses. RESULTS: We found that, with the UDI implementation at 3 health systems, the target device exposure data could be effectively identified, particularly for brand-specific devices. Computable phenotypes for study outcomes could be defined using codes; however, ablation registries, natural language processing tools, and chart reviews were required for validating data quality of the phenotypes. The common data model implementation status varied across sites. The maturity level of the key informatics technologies was highly aligned with the data quality maturity model. CONCLUSIONS: We demonstrated that the informatics approaches can be feasibly used to capture safety and effectiveness outcomes in real-world data for use in medical device studies supporting label extensions.
Guoqian Jiang, Sanket S. Dhruva, Jiajing Chen, Wade L. Schulz, Amit A. Doshi, Peter A. Noseworthy, Yue Yu 0012, Hobart Patrick Young, Eric Brandt, Keondae R. Ervin, Nilay D. Shah, Joseph S. Ross, Paul Coplan, Joseph P. Drozda
J. Am. Medical Informatics Assoc.3
2017 Performance Evaluation of a Hybrid-Beamforming Sounder for 26 GHz Channel Measurements
abstract
Both the conventional time-division-multiplexing channel sounder and the horn-antenna-rotating millimeter-wave channel sounder suffer from a serious problem that significant time consumptions are inevitable when sounding a spatial channel in 3- dimensions. In this paper, a novel channel sounder utilizing Hybrid BeamForming (HBF) techniques is presented which can complete a 3D channel measurement in microseconds, facilitating wideband MIMO channel sounding in time-variant cases, such as vehicular scenarios. The sounder is equipped with 64 transmitter (Tx) antennas and 64 receiver (Rx) antennas. The Tx makes use of digital beamforming (DBF) to compose 8 directive radiation patterns by superimposing multiple weighted patterns each generated with 8 antennas through analog beamforming. The Rx receives signals through 8 parallel front-end chains each connected to an 8- element antenna array. A 26-GHz channel measurement campaign was conducted by using the sounder in a line-of-sight indoor environment. The results demonstrate the applicability of the sounder in measuring millimeter-wave channels. Moreover, special notes are given on the appropriate usage of parameter estimation algorithms to processing the HBF channel sounding data.
Xuefeng Yin, Xi Chu, Jiajing Chen, Zhimeng Zhong
VTC Fall3
2016 Measurement-based massive MIMO channel modeling in 13-17 GHz for indoor hall scenarios
abstract
In this contribution, a recently conducted measurement campaign is introduced for investigating the characteristics of propagation channels for massive multiple-input multiple-output (MIMO) scenarios in a lecture hall environment. The channel responses for waves of higher frequency band ranging from 13 GHz to 17 GHz was measured with a vertically standing virtual two-dimensional (2-D) 20 × 20 = 400-element planar antenna array at the receiver (Rx) side, and an omni-directional antenna at the transmitter (Tx) side. Measurements were performed for four different locations of the Tx antenna in line-of-sight (LoS) scenarios. The variation of channel characteristics such as the narrowband channel gain, the K-factor and the composite delay spread, across the 2-D array aperture is investigated. The results are important for generating realistic channel realizations for the designing and performance evaluation of the algorithms for massive MIMO communication in the context of the fifth generation (5G) wireless networks.
Jiajing Chen, Xuefeng Yin, Stephen Wang 0001
ICC1
2015 Empirical Geometry-Based Random-Cluster Model for High-Speed-Train Channels in UMTS Networks
abstract
In this paper, a recently conducted measurement campaign for high-speed-train (HST) channels is introduced, where the downlink signals of an in-service Universal Mobile Terrestrial System (UMTS) deployed along an HST railway between Beijing and Shanghai were acquired. The channel impulse responses (CIRs) are extracted from the data received in the common pilot channels (CPICHs). Within 1318 km, 144 base stations (BSs) were detected. Multipath components (MPCs) estimated from the CIRs are clustered and associated across the time slots. The results show that, limited by the sounding bandwidth of 3.84 MHz, most of the channels contain a single line-of-sight (LoS) cluster, and the rest consists of several LoS clusters due to distributed antennas, leaking cable, or neighboring BSs sharing the same CPICH. A new geometry-based random-cluster model is established for the clusters' behavior in delay and Doppler domains. Different from conventional models, the time-evolving behaviors of clusters are characterized by random geometrical parameters, i.e., the relative position of BS to railway, and the train speed. The distributions of these parameters, and the per-cluster path loss, shadowing, delay, and Doppler spreads, are extracted from the measurement data.
Xuefeng Yin, Xuesong Cai, Xiang Cheng 0001, Jiajing Chen
IEEE Trans. Intell. Transp. Syst.4
2013 Crossing the Early Adopter Chasm for HIE
Cynthia LeRouge, Brandon L. Rahn, Jiajing Chen, Monica C. Tremblay, Kelvin Hanson
AMIA3
2012 Empirical models of cross-correlation for small-scale fading in co-existing channels
abstract
In this paper, stochastic modeling of the cross-correlation of small-scale fading (SSF) is performed for co-existing propagation channels based on measurement data collected using the relay-Band-Exploration-and-Channel-Sounder (rBECS) systems in urban areas, Korea. We define six geometrical parameters which are considered as variables in the cross-correlation models proposed. These parameters characterize the geographic properties of a three-node co-operative relay system that consists of a base station, a relay station and a mobile station. The sensitivity of the distribution of SSF cross-correlation coefficients with respect to the proposed variables is assessed. Three of them are identified as sensible variables for modeling.
Xuefeng Yin, Jinyi Liang, Jiajing Chen, Jae Joon Park, Myung Don Kim, Hyun Kyu Chung
APCC3