Jiajia Li 0004

dblp:89/9032-4 · DBLP profile ↗
← Back
14ranked-venue papers
7as first author
14since 2021 · last 2026
0000-0001-8158-3970ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 7 · 4 first-author · 7 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A weakly supervised framework for CT-based infectious pancreatic necrosis prediction with CAM-guided lesion localization
Jiajia Li 0004, Yuechen Liu
Vis. Comput.2
2025 A Progressive Local Variance-guided Strategy for Improving Data Augmentation Reliability
abstract
Recently, CutMix-based augmentation has emerged as a promising strategy for providing regularization to deep neural networks. However, the randomness in cropping may result in uninformative or non-representative regions being selected, resulting in a synthesized image without the desired features. To address these issues, we propose a simple, flexible, and effective augmentation strategy called Progressive LOcal Variance-guided Mix (PlovMix). PlovMix indicates the effective regions based on the information density distribution of the image, which maintains the consistency between synthetic images and the corresponding labels, and further improves the reliability of the augmented data. Additionally, our method generates idiosyncratic shape-free mask for image, which helps the network learn more appropriate feature distributions from the diverse synthetic data. Finally, Experimental results demonstrate PlovMix significantly improves the generalization performance of popular deep networks on various datasets, such as CIFAR-10, CIFAR-100, Tiny ImageNet, and FGVC-Aircraft.
Zheyuan Wang, Dezhi Wu, Haoran Liao, Jiajia Li 0004
ICASSP7
2025 MHFNet: A Multimodal Hybrid-Embedding Fusion Network for Automatic Sleep Staging
abstract
Scoring sleep stages is essential for evaluating the status of sleep continuity and comprehending its structure. Despite previous attempts, automating sleep scoring remains challenging. First, most existing works did not fuse local and global temporal information. Second, the correlation for special waves in different signals is rarely used in sleep staging modeling. Third, the logic of scoring rules based on adjacent epochs is not considered in developing sleep staging models. This paper introduces a multimodal hybrid-embedding fusion network (MHFNet), which aims to tackle these challenges in automating sleep stage scoring. MHFNet comprises multi-stream Xception blocks to extract wave characteristics, a hybrid time-embedding module to combine local and global temporal information, a dual-path gate transformer to fuse and enhance attention features, and a refined output header to reconstruct sleep scoring. We perform experiments using three publicly available datasets (SleepEDF-ST, SleepEDF-SC, and SHHS). Experimental results indicate the superiority of MHFNet over baseline approaches in cross-validation. Moreover, at the individual level, MHFNet yielded an average $R^{2}$ score improvement of 9$\%$ in the testing dataset compared to state-of-the-art models, paving the way for its applications in real-world sleep medicine.
Ruhan Liu, Jiajia Li 0004, Bin Sheng 0001, David Dagan Feng, Ping Zhang 0016
IEEE J. Biomed. Health Informatics2
2024 SSM-Net: Semi-supervised multi-task network for joint lesion segmentation and classification from pancreatic EUS images
Jiajia Li 0004, Lei Zhu 0003, Ping Zhang 0016, Ruhan Liu, Bin Sheng 0001
Artif. Intell. Medicine1
2024 Saliency-Aware Dual Embedded Attention Network for Multivariate Time-Series Forecasting in Information Technology Operations
abstract
In the field of artificial intelligence for information technology operations, operational data are often modeled as aperiodic multivariate time series, which contain rich multidimensional and nonlinear patterns. However, the existing approaches are unable to effectively acquire knowledge and recognize patterns due to their reliance on processing and modeling periodic patterns. To address this issue, this article proposes a novel deep-saliency-aware dual embedded attention network for aperiodic multivariate time-series forecasting. Our network consists of three main components: 1) a convolutional-neural-network- and transformer-based component for saliency representation of the aperiodic patterns; 2) a lightweight recurrent neural network component for capturing long-term dependence features; and 3) an attention mechanism for fusing the latent representations from the former components. Extensive empirical studies are conducted on a real-world dataset and five other public datasets to evaluate the proposed network against four state-of-the-art models. The results show that our method achieves impressive high performance on most evaluation metrics. Furthermore, the data and code used in this study are publicly available, which can facilitate progress in the community.
Jiajia Li 0004, Feng Tan 0002, Pengwei Hu 0001, Xin Luo 0001
IEEE Trans. Ind. Informatics1
2024 DSMT-Net: Dual Self-Supervised Multi-Operator Transformation for Multi-Source Endoscopic Ultrasound Diagnosis
abstract
Pancreatic cancer has the worst prognosis of all cancers. The clinical application of endoscopic ultrasound (EUS) for the assessment of pancreatic cancer risk and of deep learning for the classification of EUS images have been hindered by inter-grader variability and labeling capability. One of the key reasons for these difficulties is that EUS images are obtained from multiple sources with varying resolutions, effective regions, and interference signals, making the distribution of the data highly variable and negatively impacting the performance of deep learning models. Additionally, manual labeling of images is time-consuming and requires significant effort, leading to the desire to effectively utilize a large amount of unlabeled data for network training. To address these challenges, this study proposes the Dual Self-supervised Multi-Operator Transformation Network (DSMT-Net) for multi-source EUS diagnosis. The DSMT-Net includes a multi-operator transformation approach to standardize the extraction of regions of interest in EUS images and eliminate irrelevant pixels. Furthermore, a transformer-based dual self-supervised network is designed to integrate unlabeled EUS images for pre-training the representation model, which can be transferred to supervised tasks such as classification, detection, and segmentation. A large-scale EUS-based pancreas image dataset (LEPset) has been collected, including 3,500 pathologically proven labeled EUS images (from pancreatic and non-pancreatic cancers) and 8,000 unlabeled EUS images for model development. The self-supervised method has also been applied to breast cancer diagnosis and was compared to state-of-the-art deep learning models on both datasets. The results demonstrate that the DSMT-Net significantly improves the accuracy of pancreatic and breast cancer diagnosis.
Jiajia Li 0004, Lei Zhu 0003, Ruhan Liu, Dinggang Shen, Bin Sheng 0001
IEEE Trans. Medical Imaging1
2023 Making the Implicit Explicit: Depression Detection in Web across Posted Texts and Images
abstract
The utilization of web social media for depression detection has been proven effective in recent years since the multimedia signal on web can reflect users’ emotions, feelings, and personality traits in advance. However, most earlier studies simply used users’ submitted words or user profiles to predict depression risk. The implicit information accessible in users’ posted images, which can be effective in depression detection, still remains unexplored. In this paper, an implicit and explicit multi-modal feature fusion (IEMFF) model is proposed for depression detection. We successfully make the implicit information inherent in users’ posted images explicit and further incorporate such explicit features with the textual features directly extracted from user-posted texts. A multi-modal feature fusion approach is applied for depression detection. Extensive experiments have been conducted on public Twitter datasets. Experimental results show that our approach has achieved state-of-the-art performance for depression detection.
Pengwei Hu 0001, Chenhao Lin, Jiajia Li 0004, Feng Tan 0002, Xue Han 0018, Xi Zhou 0007, Lun Hu
BIBM3
2023 MLGL: Model-free Lesion Generation and Learning for Diabetic Retinopathy Diagnosis
abstract
The approaches based on deep learning have achieved remarkable success in diabetic retinopathy detection. Due to the accountability in medical diagnosis, the interpretability of computer-aided diagnosis has recently been investigated. However, few existing approaches make full use of the explainable evidence to improve the diagnosis accuracy. In this paper, we propose a Model-free Lesion Generation and Learning (MLGL) framework to study the interpretability of diabetic retinopathy detection. We first generate visual explanations for diabetic retinopathy diagnosis using the proposed Gated Multi-layer Saliency Map (GMSM) module, which locates the accurate region of lesions by combining multi-layer heatmaps. Then we use the GMSM to extract the lesion patches and conduct the adaptive lesion transfer, iteratively generating new retinal fundus images with lesions. Especially, in this process, no additional generative models are trained. Finally, we merge the generated and original retinal fundus images for the model's training to learn robust lesion features. Overall, our method provides accurate explainable evidence and further addresses the data imbalance problem in diabetic retinopathy detection. The experimental results on four public datasets demonstrate the efficiency of our approach.
Jiajia Li 0004, Chenhao Lin, Feng Tan 0002, Lun Hu, Pengwei Hu 0001
BIBM1
2023 PatternRCA: A Pattern-Aware Root Cause Analysis Framework for Multi-Dimensional Time Series
abstract
Root cause analysis for multi-dimensional time series from large scale micro-service scenarios aims at identifying the set of anomaly attributes by monitoring operational metrics. The online metrics provide a general indication to investigate these attributes' inter-dependencies and can guide the overall exploration process. However, the problem space for the root cause localization still remains largely challenging due to the combinatorial explosion of possible attribute combinations. This leads researchers and practitioners to (a) assume some prior distributions on the data set; (b) assume some data patterns on the attribute combinations; (c) perform pruning techniques to reduce the search space. Furthermore, state-of-the-art root cause analysis methods are often tied to one or more of these assumptions, which makes it difficult to be robust to general scenarios. In this paper, we conclude the heterogeneity in the data patterns by analyzing several open and industrial datasets. A uniform analytical framework, PatternRCA, is proposed such that it can be aware of the patterns in the metrics while avoiding explicit assumptions about them. We design an offline learning procedure that enables the framework to detect existing data patterns, which then can guide it to do fine-grain exploration in online metrics. Our extensive evaluation results show that PatternRCA outperforms state-of-the-art models with better benchmark results in public datasets. Meanwhile, it can scale to complex root cause analysis tasks on datasets with hybrid patterns in production environments.
Fulong Tian, Peijiao Xue, Jiajia Li 0004, Feng Tan 0002, Hongyang Chen 0001, Linghe Kong
ICDM6
2023 TransOrga: End-To-End Multi-modal Transformer-Based Organoid Segmentation
Jiajia Li 0004, Zhu-Hong You, Lun Hu, Pengwei Hu 0001, Feng Tan 0002
ICIC (3)2
2023 Artificial intelligence accelerates multi-modal biomedical process: A Survey
Jiajia Li 0004, Xue Han 0018, Feng Tan 0002, Xi Zhou 0007, Lun Hu, Pengwei Hu 0001
Neurocomputing1
2022 CDX-NET: Cross-Domain Multi-Feature Fusion Modeling Via Deep Neural Networks for Multivariate Time Series Forecasting in AIOps
abstract
In the application of Artificial Intelligence for IT Operations (AIOps), monitoring data are usually modeled as MTS (Multivariate Time Series). The prediction of MTS has been widely studied and various models, including statistic algorithms and deep learning networks, have been proposed, which attempt to capture the multi-dimensional and non-linear features. To this end, this paper focuses on one important type of time series: aperiodic MTS. Our solution introduces a deep neural network named CDX-Net to describe and analyze aperiodic MTS from both temporal and spectral domains. We also propose the integration of the convolution neural network (CNN), recurrent neural network (RNN) and attention mechanism into the predictive model. The introduction of these modules can effectively improve the feature extraction and feature fusion procedures. We conduct performance evaluation on a real-world dataset from an AIOps application and the correlation between the predicted result and ground-truth is found to be significant. The proposed model is compared with several state-of-the-art baseline methods. Empirical results show that our model achieves better performance in most evaluation metrics while others can perform better under some particular settings.
Jiajia Li 0004
ICASSP1
2022 B-AT-KD: Binary attention map knowledge distillation
Jiajia Li 0004, Huiyong Chu, Zichen Zhang 0012, Feng Tan 0002, Pengwei Hu 0001
Neurocomputing3
2022 Automatic Detection and Classification System of Domestic Waste via Multimodel Cascaded Convolutional Neural Network
abstract
Domestic waste classification was incorporated into legal provisions recently in China. However, relying on manpower to detect and classify domestic waste is highly inefficient. To that end, in this article, we propose a multimodel cascaded convolutional neural network (MCCNN) for domestic waste image detection and classification. MCCNN combined three subnetworks (DSSD, YOLOv4, and Faster-RCNN) to obtain the detections. Moreover, to suppress the false-positive predicts, we utilized a classification model cascaded with the detection part to judge whether the detection results are correct. To train and evaluate MCCNN, we designed a large-scale waste image dataset (LSWID), containing 30 000 domestic waste multilabeled images with 52 categories. To the best of our knowledge, the LSWID is the largest dataset on domestic waste images. Furthermore, a smart trash can is designed and applied to a Shanghai community, which helped to make waste recycling more efficient. Experimental results showed a state-of-the-art performance, with an average improvement of 10% in detection precision.
Jiajia Li 0004, Jie Chen 0097, Bin Sheng 0001, Ping Li 0016, Po Yang 0001, David Dagan Feng, Jun Qi 0001
IEEE Trans. Ind. Informatics1