Jiyang Wang

dblp:153/3400 · DBLP profile ↗
← Back
16ranked-venue papers
5as first author
13since 2021 · last 2027
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 7 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2027 Knowledge-data co-driven hierarchical traction control of distributed-drive electric vehicles without explicit online road-friction estimation
Junchen Lin, Jiyang Wang
Expert Syst. Appl.6
2025 DualMixer: An Efficient and Robust Model for Multivariate Time Series Forecasting
abstract
Time series forecasting plays a vital role in various fields such as energy forecasting and transportation planning. Although Transformer-based models have made remarkable progress in time series forecasting, they still face bottlenecks in efficiency and performance robustness when there are multivariate correlations, multi-period patterns and context anomalies in the series. In order to solve the above problems, we propose a DualMixer model composed of the SeriesMixer and StampMixer modules to efficiently utilize the historical observation series by SeriesMixer and fully utilize timestamps of future series by StampMixer. Concretely, SeriesMixer first efficiently captures multivariate correlations and mixes correlation features between each variate by Kernel-Attention we designed, then learns nonlinear representations between temporal dimension by the feed-forward network, and finally decodes the representations by linear projection to make preliminary forecasting for future series. StampMixer mixes the features of future timestamps by MLP to extract the global information of future series. Finally, DualMixer aggregates the outputs of SeriesMixer and StampMixer to achieve more accurate prediction of future series. Consequently, DualMixer is able to achieve optimal performance in both long-term and short-term multivariate time series forecasting tasks with excellent efficiency and robustness. Code is available at this repository: https://github.com/Truffle-zzl/DualMixer.
Zhenlong Zou, Jiyang Wang
IJCNN2
2025 Hybrid wind speed optimization forecasting system based on linear and nonlinear deep neural network structure and data preprocessing fusion
Jiyang Wang, Jifeng Che, Zhiwu Li 0001, Jialu Gao, Linyue Zhang
Future Gener. Comput. Syst.1
2024 S2E: Towards an End-to-End Entity Resolution Solution from Acoustic Signal
abstract
Traditional cascading Entity Resolution (ER) pipeline suffers from propagated errors from upstream tasks. We address this issue by formulating a new end-to-end (E2E) ER problem, Signal-to-Entity (S2E), resolving query entity mentions to actionable entities in textual catalogs directly from audio queries instead of audio transcriptions in raw or parsed format. Additionally, we extend the E2E Spoken Language Understanding framework by introducing a novel dimension to ER research. We adapt three public datasets for the S2E task, and propose a novel solution, which aligns the multimodal signals via an effective retrieval co-attention mechanism and refined multimodal objectives. Despite 42% smaller in terms of the total model size, the proposed design outperforms the cascading baseline by 2.6%, 47.0%, and 73.3% across the three datasets respectively with different acoustic conditions.
Kangrui Ruan, Jiyang Wang, Helian Feng, Ali Kebarighotbi
ICASSP3
2024 SimpliMix: A Simplified Manifold Mixup for Few-shot Point Cloud Classification
abstract
Few-shot learning often assumes that base classes are abundant and diverse with plentiful well-labeled samples for each class. This ensures that models can generalize effectively from a small amount of data by leveraging prior knowledge learned from base classes. This assumption holds for 2D few-shot learning since the benchmark datasets are large and diverse. However, 3D point cloud few-shot benchmarks are low in magnitude and diversity. We conduct experiments and show that many existing methods overlook this issue and suffer from overfitting on base classes, which hinders generalization ability and test performance. To alleviate the overfitting issue, we propose a simplified manifold mixup, referred to as the SimpliMix, which mixes hidden representations and forces the models to learn more generalized features. We incorporate SimpliMix into existing prototype-based models, perform experiments on ModelNet40-FS, ModelNet40-C-FS and ScanObjectNN-FS datasets, and improve the models by a significant margin. We further conduct cross-domain few-shot classification experiments and show that networks with SimpliMix learn more generalized and transferable features and achieve better performance. The code is available at https://github.com/LexieYang/SimpliMix
Minmin Yang, Weiheng Chai, Jiyang Wang, Senem Velipasalar
WACV3
2024 Rethinking the Evaluation of Driver Behavior Analysis Approaches
abstract
Crashes caused by distracted driving result in more than 3000 deaths every year in the U.S. Distracted driver behavior detection is instrumental for driver assist systems. Researchers have focused on autonomously detecting distracted driver behavior so that drivers can be alerted in time to reduce the risk of crashes. Despite the large number of approaches presented in the literature, there are still issues related to proper performance evaluation, reproducibility and lack of or very slow adoption of these approaches by the transportation industry. Most existing approaches either do not provide documented and usable codes or use private datasets, or do not present the experiment details, such as data split, sometimes resulting in inflated accuracy numbers. Moreover, these factors also make many results not reproducible. In addition, the performance metrics should be chosen carefully to measure various aspects of different methods, including their generalizability, and action localization ability in time. In this work, we perform a commensurate comparison of different state-of-the-art methods by using different data splits and performance metrics on the StateFarm distracted driving and AI CITY Challenge datasets. With the data split experiments, we highlight the importance of leave-N-driver-out cross validation, since these models should perform well in real-world testing with never-before-seen drivers. The results show the importance of data splitting and the performance metric for the comparison and evaluation of different methods, and their significant effects on the results.
Weiheng Chai, Jiyang Wang, Jiajing Chen, Senem Velipasalar, Anuj Sharma 0001
IEEE Trans. Intell. Transp. Syst.2
2024 Vision-Language Models Can Identify Distracted Driver Behavior From Naturalistic Videos
abstract
Recognizing the activities causing distraction in real-world driving scenarios is critical for ensuring the safety and reliability of both drivers and pedestrians on the roadways. Conventional computer vision techniques are typically data-intensive and require a large volume of annotated training data to detect and classify various distracted driving behaviors, thereby limiting their generalization ability, efficiency and scalability. We aim to develop a generalized framework that showcases robust performance with access to limited or no annotated training data. Recently, vision-language models have offered large-scale visual-textual pretraining that can be adapted to task-specific learning like distracted driving activity recognition. Vision-language pretraining models like CLIP have shown significant promise in learning natural language-guided visual representations. This paper proposes a CLIP-based driver activity recognition approach that identifies driver distraction from naturalistic driving images and videos. CLIP’s vision embedding offers zero-shot transfer and task-based finetuning, which can classify distracted activities from naturalistic driving video. Our results show that this framework offers state-of-the-art performance on zero-shot transfer, finetuning and video-based models for predicting the driver’s state on four public datasets. We propose frame-based and video-based frameworks developed on top of the CLIP’s visual representation for distracted driving detection and classification tasks and report the results. Our code is available at https://github.com/zahid-isu/DriveCLIP
Md. Zahid Hasan, Jiajing Chen, Jiyang Wang, Mohammed Shaiqur Rahman, Ameya Joshi, Senem Velipasalar, Chinmay Hegde, Anuj Sharma 0001, Soumik Sarkar
IEEE Trans. Intell. Transp. Syst.3
2023 SimTDE: Simple Transformer Distillation for Sentence Embeddings
abstract
In this paper we introduce SimTDE, a simple knowledge distillation framework to compress sentence embeddings transformer models with minimal performance loss and significant size and latency reduction. SimTDE effectively distills large and small transformers via a compact token embedding block and a shallow encoding block, connected with a projection layer, relaxing dimension match requirement. SimTDE simplifies distillation loss to focus only on token embedding and sentence embedding. We evaluate on standard semantic textual similarity (STS) tasks and entity resolution (ER) tasks. It achieves 99.94% of the state-of-the-art (SOTA) SimCSE-Bert-Base performance with 3 times size reduction and 96.99% SOTA performance with 12 times size reduction on STS tasks. It also achieves 99.57% of teacher's performance on multi-lingual ER data with a tiny transformer student model of 1.4M parameters and 5.7MB size. Moreover, compared to other distilled transformers SimTDE is 2 times faster at inference given similar size and still 1.17 times faster than a model 33% smaller (e.g. MiniLM). The easy-to-adopt framework, strong accuracy and low latency of SimTDE can widely enable runtime deployment of SOTA sentence embeddings.
Jian Xie 0002, Jiyang Wang, Zimeng Qiu, Ali Kebarighotbi, Farhad Ghassemi
SIGIR3
2023 Wind speed interval prediction based on multidimensional time series of Convolutional Neural Networks
Jiyang Wang, Zhiwu Li 0001
Eng. Appl. Artif. Intell.1
2023 Driver Head Pose Detection From Naturalistic Driving Data
abstract
Driver behavior analysis plays an important role in driver assistance systems. A driver’s face and head pose hold the key towards understanding whether the driver’s attention and concentration are on the road while driving. Naturalistic driving studies (NDS) allow observing drivers in real-time under naturalistic traffic conditions. Yet, data collected in NDS often comprise low-resolution videos usually with more challenging camera positions compared to controlled studies. For instance, when the camera is not directly facing the driver, classifying head pose becomes more challenging, since the variation between different classes becomes much smaller. In this paper, we propose three different approaches to classify a driver’s head pose from naturalistic videos, which were captured by a camera providing a side view, instead of directly facing the driver. These approaches employ a sequence of five key points on the driver’s face. We compare these three proposed approaches with each other as well as with three different baselines by using leave-one-driver-out cross-validation on nine different drivers. Results show that our proposed method employing a Bidirectional Gated Recurrent Unit (BiGRU) outperforms the best performing baseline by 11% in terms of overall accuracy.
Weiheng Chai, Jiajing Chen, Jiyang Wang, Senem Velipasalar, Archana Venkatachalapathy, Yaw Adu-Gyamfi, Jennifer Merickel, Anuj Sharma 0001
IEEE Trans. Intell. Transp. Syst.3
2022 Capsule network-based semantic segmentation model for thermal anomaly identification on building envelopes
Chenbin Pan, Jiyang Wang, Weiheng Chai, Burak Kakillioglu, Yasser El Masri, Eleanna Panagoulia, Norhan Bayomi, John E. Fernandez, Tarek Rakha, Senem Velipasalar
Adv. Eng. Informatics2
2022 Taking a Deeper Look at the Brain: Predicting Visual Perceptual and Working Memory Load From High-Density fNIRS Data
abstract
Predicting workload using physiological sensors has taken on a diffuse set of methods in recent years. However, the majority of these methods train models on small datasets, with small numbers of channel locations on the brain, limiting a model's ability to transfer across participants, tasks, or experimental sessions. In this paper, we introduce a new method of modeling a large, cross-participant and cross-session set of high density functional near infrared spectroscopy (fNIRS) data by using an approach grounded in cognitive load theory and employing a Bi-Directional Gated Recurrent Unit (BiGRU) incorporating attention mechanism and self-supervised label augmentation (SLA). We show that our proposed CNN-BiGRU-SLA model can learn and classify different levels of working memory load (WML) and visual processing load (VPL) across participants. Importantly, we leverage a multi-label classification scheme, where our models are trained to predict simultaneously occurring levels of WML and VPL. We evaluate our model using leave-one-participant-out (LOOCV) as well as 10-fold cross validation. Using LOOCV, for binary classification (off/on), we reached an F1-score of 0.9179 for WML and 0.8907 for VPL across 22 participants (each participant did 2 sessions). For multi-level (off, low, high) classification, we reached an F1-score of 0.7972 for WML and 0.7968 for VPL. Using 10-fold cross validation, for multi-level classification, we reached an F1-score of 0.7742 for WML and 0.7741 for VPL.
Jiyang Wang, Trevor Grant, Senem Velipasalar, Baocheng Geng, Leanne M. Hirshfield
IEEE J. Biomed. Health Informatics1
2022 A Survey on Driver Behavior Analysis From In-Vehicle Cameras
abstract
Distracted or drowsy driving is unsafe driving behavior responsible for thousands of crashes every year. Studying driver behavior has challenges associated with observing drivers in their natural environment. The naturalistic driving study (NDS) has become the most sought-after approach, since it eliminates the bias of a controlled setup, allowing researchers to understand drivers’ behavior in real-world scenarios. Video recordings collected in NDS research are incredibly insightful in identifying driver errors. Computer vision techniques have been used to autonomously analyze video data and classify drivers’ behavior. While computer vision scientists focus on image analytics, NDS researchers are interested in the factors impacting driver behavior. This survey paper makes a concerted effort to serve both communities by comprehensively reviewing studies, describing their data collection, computer vision techniques implemented, and performance in classifying driver behavior. The scope is limited to studies employing at least one camera observing the driver inside a vehicle. Based on their objective, papers have been classified as detecting low-level (e.g. head orientation) or high-level (e.g. distraction detection) driver information. Papers have been further classified based on the datasets they employ. In addition to twelve public datasets, many private datasets have also been identified, and their data collection design is discussed to highlight any impact on model performance. Across each task, algorithms employed and their performance are discussed to establish a baseline. A comparison of different frameworks for NDS video data analytics throws light on the existing gaps in the state-of-the-art that can be addressed by future computer vision research.
Jiyang Wang, Weiheng Chai, Archana Venkatachalapathy, Kai Liang Tan, Arya Haghighat, Senem Velipasalar, Yaw Adu-Gyamfi, Anuj Sharma 0001
IEEE Trans. Intell. Transp. Syst.1
2019 Research on combined model based on multi-objective optimization and application in time series forecast
Shenghui Zhang, Jiyang Wang, Zhen-hai Guo
Soft Comput.2
2004 Optical Ethernet: making Ethernet carrier class for professional services
abstract
The existing overlaid data network architecture deployed by most service providers has shown its shortcomings in supporting professional business services due to its complexity and high cost. This paper introduces a new optical transport technology that is based on Ethernet but integrates all the required features to consolidate multiple layers below the IP layer in the existing architecture into one, thus simplifying the architecture and significantly reducing the cost both in network buildout and in network operations. The paper provides the technical details on how the limitations of traditional Ethernet are overcome and how Ethernet becomes carrier class. It also introduces the new services enabled by carrier-class Ethernet and analyzes the impact of it on Internet evolution.
Jiyang Wang
Proc. IEEE1
2001 Joint frequency, 2-D AOA and polarization estimation in broad-band
Jiyang Wang, Tiangi Chen
Sci. China Ser. F Inf. Sci.2