EDBT 2026 Demo / reviewers in the wild / expert
Jinglin Zhang 0001
dblp:83/2890-1
· DBLP profile ↗
30ranked-venue papers
1as first author
27since 2021 · last 2026
0000-0003-1618-8493ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 1 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Databases, data management, data science and information retrieval · 4 · 3 since 2021Computer networks · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | KFTD: Koopman-Fourier Time-Differentiable Network for Continuous Ocean Spatiotemporal ForecastingabstractAccurate oceanic forecasting is critical for climate monitoring and disaster early-warning. However, ocean spatiotemporal forecasting encounters the double challenges of modeling complex dynamical systems and ensuring computational efficiency. We present Koopman–Fourier Time-Differentiable (KFTD) Network, a time-continuous two-stage paradigm that decouples interpolation from prediction to achieve efficient and scalable spatiotemporal modeling. We map complex nonlinear dynamics into the Koopman linear space and exploit Fourier analysis to enable continuous-time interpolation at arbitrary sub-steps. A lightweight residual network consumes the high-fidelity intermediate states to yield the final forecast. Unlike diffusion models, KFTD eliminates multi-step noise sampling and directly evolves the system in continuous time, yielding a 4× computational speed-up. We further introduce a D-PP Loss that supports arbitrary PDE constraints in an end-to-end manner, breaking the physical-consistency bottleneck of pure data-driven approaches. Empirical results on four ocean datasets confirm that our continuous-time framework reduces MSE by an average of 5.6% (up to 12.7% for SST) and improves efficiency over MCVD by 76.25%. Qinghui Chen, Hailong Liu 0007, Jinglin Zhang 0001, Cong Bai |
KDD (1) | 4 |
| 2026 | KAN-FIF: Spline-Parameterized Lightweight Physics-based Tropical Cyclone Estimation on Meteorological Satellite
Jiakang Shen, Qinghui Chen, Runtong Wang, Chenrui Xu, Jinglin Zhang 0001, Cong Bai, Feng Zhang 0041 |
KDD (1) | 5 |
| 2026 | 3D-MolGL: A multimodal framework for integrating 3D molecular graphs into language models
Huizhi Li, Dagang Li 0001, Jinglin Zhang 0001, Yuhui Zheng, Cong Bai |
Expert Syst. Appl. | 3 |
| 2026 | Unification of Closed-Open Industrial Detection Scenarios: New Large-Scale Benchmarks, Challenges and BaselinesabstractLarge-scale Visual-Language Models (LVLMs) have achieved remarkable success in natural visual tasks, yet their application to industrial defect detection remains challenging due to two fundamental limitations: (i) the scarcity of large-scale industrial datasets that cover diverse defect categories across multiple domains, and (ii) the reliance on manual prompts (points, boxes, masks) that introduce subjective noise and lack text-visual interaction for fine-grained understanding. To address these challenges, we introduce a Large-Scale Multi-Modal Industrial Open-Closed benchmark (MMIOC-1 M) containing over one million samples across 14 super-categories, 29 industrial scenes, and 351 defect subcategories. To our knowledge, MMIOC-1 M is the first unified largest benchmark supporting both open-vocabulary and closed-set industrial detection, providing valuable pre-training data for LVLMs in industrial scenarios. Furthermore, we propose a Refined Text-Visual Prompt Network (RTVPNet) that incorporates three key innovations: (1) an expert-assisted domain projection mechanism that enables rapid adaptation of general vision models to industrial domains, (2) an energy-based sparse sampling strategy that automatically generates refined visual prompts without manual intervention, and (3) a bidirectional text-visual interaction module that enhances cross-modal semantic alignment and understanding. Extensive experiments demonstrate that RTVPNet achieves state-of-the-art performance on MMIOC-1 M, LVIS, and COCO benchmarks while maintaining computational efficiency. Jinglin Zhang 0001, Qinghui Chen, Gang Li 0005, Da Chen 0002, Shuainan Jing, Dagang Li 0001, Cong Liu 0012, Cong Bai, Shengyong Chen |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2026 | CWRNN-INVR: A Coupled WarpRNN Based Implicit Neural Video RepresentationabstractImplicit Neural Video Representation (INVR) has emerged as a novel approach for video representation and compression, using learnable grids and neural networks. Existing methods focus on developing new grid structures efficient for latent representation and neural network architectures with large representation capability, lacking the study on their roles in video representation. In this paper, the difference between INVR based on neural network and INVR based on grid is first investigated from the perspective of video information composition to specify their own advantages, i.e., neural network for general structure while grid for specific detail. Accordingly, an INVR based on mixed neural network and residual grid framework is proposed, where the neural network is used to represent the regular and structured information and the residual grid is used to represent the remaining irregular information in a video. A Coupled WarpRNN-based multi-scale motion representation and compensation module is specifically designed to explicitly represent the regular and structured information, thus terming our method as CWRNN-INVR. For the irregular information, a mixed residual grid is learned where the irregular appearance and motion information are represented together. The mixed residual grid can be combined with the coupled WarpRNN in a way that allows for network reuse. Experiments show that our method achieves the best reconstruction results compared with the existing methods, with an average PSNR of 33.73 dB on the UVG dataset under the 3M model and outperforms existing INVR methods in other downstream tasks. The code can be found athttps://github.com/yiyang-sdu/CWRNN-INVR.git. Yanbo Gao, Shuai Li 0005, Jinglin Zhang 0001, Hui Yuan 0001, Mao Ye 0001, Xingyu Gao 0001 |
IEEE Trans. Multim. | 5 |
| 2026 | Detecting Root Causes for Process Performance Anomalies Using Causal InferenceabstractProcess execution time is a key performance indicator for evaluating bottlenecks in business processes. Cases and activities that exceed the specified time constraints can be seen as anomalies, affecting process performance and leading to risks such as delays and customer complaints. Identifying the root causes of these anomalies can help formulate effective intervention measures. However, this task is inherently complex, and conducting incomplete or inaccurate analysis can result in misguided interventions that inadvertently exacerbate process inefficiencies. To address these challenges, this paper proposes a traceability-based root cause analysis approach for process performance anomalies using causal inference. Specifically, the approach begins by extracting hidden contextual information from the event log to enrich the pool of potential causal factors. Then formulates causal hypotheses linking these factors to observed performance anomalies (at both the case and activity level) and establishes potential causal relations through a traceability mechanism. A meta-learning based causal inference approach is used to estimate the strength of causal effects. The proposed approach is evaluated against a state-of-the-art approach using four synthetic event logs with known root causes and nine public real-life event logs. Experimental results demonstrate that the proposed approach delivers accurate insights into the root causes of process performance anomalies in synthetic event logs, while maintaining high efficiency in the comprehensive analysis of potential causal factors. Cong Liu 0012, Qingtian Zeng, Youxi Wu, Jinglin Zhang 0001, Xixi Lu 0001, Long Cheng 0003 |
IEEE Trans. Serv. Comput. | 5 |
| 2025 | Dust-Mamba: An Efficient Dust Storm Detection Network with Multiple Data SourcesabstractAccurate detection of dust storms is challenging due to complex meteorological interactions. With the development of deep learning, deep neural networks have been increasingly applied to dust storm detection, offering better learning and generalization capabilities compared to traditional physical modeling. However, existing methods face some limitations, leading to performance bottlenecks in dust storm detection. From the task perspective, existing research focuses on occurrence detection while neglecting intensity detection. From the data perspective, existing research fails to explore the utilization of multi-source data. From the model perspective, most models are built on convolutional neural networks, which have an inherent limitation in capturing long-range dependencies. To address the challenges mentioned, this study proposes Dust-Mamba. To the best of our knowledge, this study is the first attempt to accomplish both the occurrence and intensity detection of dust storms with advanced deep learning technology. In Dust-Mamba, multi-source data is introduced to provide a comprehensive perspective, Mamba and attention are applied to boost feature selection while maintaining long-range modeling capability. Additionally, this study proposes Structure Sharing Transfer Learning Strategies for intensity detection, which further enhances the performance of Dust-Mamba with minimal time cost. As shown by experiments, Dust-Mamba achieves Dice scores of 0.963 for occurrence detection and 0.560 for intensity detection, surpassing several baseline models. In conclusion, this study offers valuable baselines for dust storm detection, with significant reference value and promising application potential. Cong Bai, Zhonghao Lin, Jinglin Zhang 0001, Shengyong Chen |
AAAI | 3 |
| 2025 | MetricGrids: Arbitrary Nonlinear Approximation with Elementary Metric Grids based Implicit Neural RepresentationabstractThis paper presents MetricGrids, a novel grid-based neural representation that combines elementary metric grids in various metric spaces to approximate complex nonlinear signals. While grid-based representations are widely adopted for their efficiency and scalability, the existing feature grids with linear indexing for continuous-space points can only provide degenerate linear latent space representations, and such representations cannot be adequately compensated to represent complex nonlinear signals by the following compact decoder. To address this problem while keeping the simplicity of a regular grid structure, our approach builds upon the standard grid-based paradigm by constructing multiple elementary metric grids as high-order terms to approximate complex nonlinearities, following the Taylor expansion principle. Furthermore, we enhance model compactness with hash encoding based on different sparsities of the grids to prevent detrimental hash collisions, and a high-order extrapolation decoder to reduce explicit grid storage requirements. experimental results on both 2D and 3D reconstructions demonstrate the superior fitting and rendering accuracy of the proposed method across diverse signal types, validating its robustness and generalizability. Code is available at https://github.com/wangshu31/MetricGrids. Yanbo Gao, Shuai Li 0005, Chong Lv, Chuankun Li, Hui Yuan 0001, Jinglin Zhang 0001 |
CVPR | 8 |
| 2025 | Causal Discovery from Shifted Multiple Environments
Dezhi Yang, Guoxian Yu, Jun Wang 0035, Jinglin Zhang 0001, Carlotta Domeniconi |
KDD (1) | 4 |
| 2025 | Dual-path aggregation transformer network for super-resolution with images occlusions and variability
Qinghui Chen, Lunqian Wang, Xinghua Wang 0008, Bo Xia, Hao Ding 0014, Jinglin Zhang 0001 |
Eng. Appl. Artif. Intell. | 8 |
| 2025 | Multi-Domain Adversarial Variational Bayesian Inference for Domain GeneralizationabstractDomain generalization aims to learn common knowledge from multiple observed source domains and transfer it to unseen target domains, e.g. the object recognition in varieties of visual environments. Traditional domain generalization methods aim to learn the feature representation of the raw data with its distribution invariant across domains. This relies on the assumption that the two posterior distributions (the distributions of the label given the feature distribution and given the raw data) are stable in different domains. However, this does not always hold in many practical situations. In this paper, we relax the above assumption by permitting the posterior distribution of the label given the raw data changes in difference domains, and thus focuses on a more realistic learning problem that infers the conditional domain-invariant feature representation. Specifically, a multi-domain adversarial variational Bayesian inference approach is proposed to minimize the inter-domain discrepancy of the conditional distributions of the feature given the label. Besides, it is imposed by the constraints from the adversarial learning and feedback mechanism to enhance the condition invariant feature representation. The extensive experiments on two datasets demonstrate the effectiveness of our approach, as well as the state-of-the-art performance comparing with thirteen methods. Zhifan Gao, Saidi Guo, Chenchu Xu, Jinglin Zhang 0001, Mingming Gong, Javier Del Ser, Shuo Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | SPMNet: A Siamese Pyramid Mamba Network for Very-High-Resolution Remote Sensing Change DetectionabstractVery-High-Resolution (VHR) remote sensing images are characterized by extremely high spatial resolution, incorporating higher pixel density and larger image sizes, which pose challenges for existing methods to extract complex texture features. Furthermore, due to the wide-area and high-resolution imaging strategy, VHR change detection images suffer from a severe imbalance between change pixels and non-change pixels, increasing the difficulty of handling change detection tasks. To address these challenges, we introduced the Omnidirectional Selective Scan Module (OSSM), which has the capability to process long sequences. By integrating it with the lightweight Siamese Feature Pyramid Network (SFPN), we designed a hybrid CNN-Mamba backbone, referred to as SPMamba. This backbone captures both global and local information within bitemporal feature maps at each stage, enhancing the precision of texture feature extraction. Additionally, to integrate the semantic features from each branch in SPMamba and reduce noise interference from non-target change areas, we developed a Hybrid Fusion Module (HFM). The HFM consists of two fusion modules: the High-Low Channel Fusion Module (HLM) and the Bilateral Channel Fusion Module (BCM), which facilitates both feature-level and channel-level integration, enhancing the sensitivity of the model to subtle changes. Extensive experimental results demonstrate that SPMNet achieves the highest F1-score of 91.80%, 90.99%, and 96.04% on the WHU-CD, LEVIR-CD, and CDD-CD datasets, respectively, outperforming eleven state-of- the-art methods. Moreover, the effects of varying image sizes on model training are thoroughly analyzed. Jinze Song, Yunlong Ji, Wenyin Zhang, Jinglin Zhang 0001, Xing Wang 0002 |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2025 | S2DBFT: Spectral-Spatial Dual-Branch Fusion Transformer for Hyperspectral Image ClassificationabstractConvolutional neural networks (CNNs) and Transformer-based models have achieved remarkable success in hyperspectral image (HSI) classification tasks due to their outstanding ability to extract spatial and spectral features. However, most existing methods process spatial and spectral features separately, making it difficult to effectively learn their interactive features. To address this issue, we propose a spectral-spatial dual-branch fusion Transformer (S2DBFT) for HSI classification. Initially, we construct a spectral feature extraction module (SPEEM) and a spatial feature extraction module (SPAEM) to extract low-level features. These two modules consist of a one-dimensional convolution layer and a two-dimensional convolution layer, respectively, performing shallow extraction of spectral and spatial features. Next, the two feature sets obtained are fused through a weighted fusion process. Additionally, we design a multi-head spectral-spatial self-attention (MHS3A) mechanism to enhance the interactive fusion of spectral and spatial features. Upon completion of feature fusion, a linear layer is used to obtain the sample labels. Extensive experiments on four HSI datasets demonstrate the effectiveness of the proposed S2DBFT, compared to existing state-of-the-art methods. In terms of performance evaluation, the overall accuracy and average accuracy indicate the superiority and generalizability of S2DBFT. Meng Huang 0003, Ming Li 0026, Jian Zhang 0082, Shandong Wang, Jinglin Zhang 0001, Heng Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2025 | Representation Learning Based on Co-Evolutionary Combined With Probability Distribution Optimization for Precise Defect LocationabstractVisual defect detection methods based on representation learning play an important role in industrial scenarios. Defect detection technology based on representation learning has made significant progress. However, existing defect detection methods still face three challenges: first, the extreme scarcity of industrial defect samples makes training difficult. Second, due to the characteristics of industrial defects, such as blur and background interference, it is challenging to obtain fuzzy defect separation edges and context information. Third, industrial defects cannot obtain accurate positioning information. This article proposes feature co-evolution interaction architecture (CIA) and glass container defect dataset to address the above challenges. Specifically, the contributions of this article are as follows: first, this article designs a glass container image acquisition system that combines RGB and polarization information to create a glass container defect dataset containing more than 60000 samples to alleviate the sample scarcity problem in industrial scenarios. Subsequently, this article designs the CIA. CIA optimizes the probability distribution of features through the co-evolution of edge and context features, thereby improving detection accuracy in blurred defects and noisy environments. Finally, this article proposes a novel inforced IoU loss (IIoU loss), which can obtain more accurate position information by being aware of the scale changes of the predicted box. Defect detection experiments in three mainstream industrial manufacturing categories (Northeastern University (NEU)-Det, glass containers, wood) show that CIA only uses 22.5 GFLOPs, and mean average precision (mAP) (NEU-Det: 88.74%, glass containers: 95.38%, wood: 68.42%) outperforms state-of-the-art methods. Jinglin Zhang 0001, Qinghui Chen, Gang Li 0005, Shijiao Ding, Maomao Xiong, Shengyong Chen |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2025 | Boosting Pseudo-Labeling With Curriculum Self-Reflection for Attributed Graph ClusteringabstractAttributed graph clustering is an unsupervised learning task that aims to partition various nodes of a graph into distinct groups. Existing approaches focus on devising diverse pretext tasks to obtain suitable supervised information for representation learning, among which the predictive methods show great potential. However, these methods 1) generate auxiliary task bias toward the clustering target and 2) introduce label noise due to static thresholds. To address this issue, we propose a new self-supervised learning method, namely, pseudo-labeling with curriculum self-reflection (PLCSR), that learns reliable pseudo-labels by mining its information to achieve progressive processing of nodes in a self-reflection manner. First, a self-auxiliary encoder is constructed using the exponential moving average (EMA) of the original encoder's parameters to replace the auxiliary tasks, which provides an additional perspective of finding highly confident pseudo-labels. Second, a curriculum selection strategy using dynamic thresholds is designed to take full advantage of graph nodes more accurately. Besides simple nodes with high confidence at the initial stage, nodes that yield consistent predictions from both encoders are then assigned pseudo-labels to avoid the under-learning problem. For the rest difficult nodes that are highly uncertain, we abstain from making judgments to minimize their adverse impact on the model. Extensive experiments have shown that PLCSR significantly outperforms the state-of-the-art predictive method CDRS, achieving more than 6% improvements in terms of clustering accuracy. The code is available at: https://github.com/Jillian555/PLCSR. Pengfei Zhu 0001, Yu Wang 0106, Bin Xiao 0002, Jinglin Zhang 0001, Wanyu Lin, Qinghua Hu |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | Federated Causality Learning with Explainable Adaptive OptimizationabstractDiscovering the causality from observational data is a crucial task in various scientific domains. With increasing awareness of privacy, data are not allowed to be exposed, and it is very hard to learn causal graphs from dispersed data, since these data may have different distributions. In this paper, we propose a federated causal discovery strategy (FedCausal) to learn the unified global causal graph from decentralized heterogeneous data. We design a global optimization formula to naturally aggregate the causal graphs from client data and constrain the acyclicity of the global graph without exposing local data. Unlike other federated causal learning algorithms, FedCausal unifies the local and global optimizations into a complete directed acyclic graph (DAG) learning process with a flexible optimization objective. We prove that this optimization objective has a high interpretability and can adaptively handle homogeneous and heterogeneous data. Experimental results on synthetic and real datasets show that FedCausal can effectively deal with non-independently and identically distributed (non-iid) data and has a superior performance. Dezhi Yang, Xintong He, Jun Wang 0035, Guoxian Yu, Carlotta Domeniconi, Jinglin Zhang 0001 |
AAAI | 6 |
| 2024 | DACA: A domain adaptive fault diagnosis approach with class-aware based on cross-domain extreme imbalance data
Yuanjiang Li, Yang Yu 0005, Runze Mao, Linchang Ye, Ruochen Liu 0005, Tao Lang, Jinglin Zhang 0001 |
Expert Syst. Appl. | 9 |
| 2024 | Graph Convolutional Network Discrete Hashing for Cross-Modal RetrievalabstractWith the rapid development of deep neural networks, cross-modal hashing has made great progress. However, the information of different types of data is asymmetrical, that is to say, if the resolution of an image is high enough, it can reproduce almost 100% of the real-world scenes. However, text usually carries personal emotion and it is not objective enough, so we generally think that the information of image will be much richer than text. Although most of the existing methods unify the semantic feature extraction and hash function learning modules for end-to-end learning, they ignore this issue and do not use information-rich modalities to support information-poor modalities, leading to suboptimal results, although they unify the semantic feature extraction and hash function learning modules for end-to-end learning. Furthermore, previous methods learn hash functions in a relaxed way that causes nontrivial quantization losses. To address these issues, we propose a new method called graph convolutional network (GCN) discrete hashing. This method uses a GCN to bridge the information gap between different types of data. The GCN can represent each label as word embedding, with the embedding regarded as a set of interdependent object classifiers. From these classifiers, we can obtain predicted labels to enhance feature representations across modalities. In addition, we use an efficient discrete optimization strategy to learn the discrete binary codes without relaxation. Extensive experiments conducted on three commonly used datasets demonstrate that our proposed method graph convolutional network-based discrete hashing (GCDH) outperforms the current state-of-the-art cross-modal hashing methods. Cong Bai, Jinglin Zhang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2023 | Long-Tail Cross Modal HashingabstractExisting Cross Modal Hashing (CMH) methods are mainly designed for balanced data, while imbalanced data with long-tail distribution is more general in real-world. Several long-tail hashing methods have been proposed but they can not adapt for multi-modal data, due to the complex interplay between labels and individuality and commonality information of multi-modal data. Furthermore, CMH methods mostly mine the commonality of multi-modal data to learn hash codes, which may override tail labels encoded by the individuality of respective modalities. In this paper, we propose LtCMH (Long-tail CMH) to handle imbalanced multi-modal data. LtCMH firstly adopts auto-encoders to mine the individuality and commonality of different modalities by minimizing the dependency between the individuality of respective modalities and by enhancing the commonality of these modalities. Then it dynamically combines the individuality and commonality with direct features extracted from respective modalities to create meta features that enrich the representation of tail labels, and binaries meta features to generate hash codes. LtCMH significantly outperforms state-of-the-art baselines on long-tail datasets and holds a better (or comparable) performance on datasets with balanced labels. Zijun Gao, Jun Wang 0035, Guoxian Yu, Zhongmin Yan, Carlotta Domeniconi, Jinglin Zhang 0001 |
AAAI | 6 |
| 2023 | Incentive-Boosted Federated CrowdsourcingabstractCrowdsourcing is a favorable computing paradigm for processing computer-hard tasks by harnessing human intelligence. However, generic crowdsourcing systems may lead to privacy-leakage through the sharing of worker data. To tackle this problem, we propose a novel approach, called iFedCrowd (incentive-boosted Federated Crowdsourcing), to manage the privacy and quality of crowdsourcing projects. iFedCrowd allows participants to locally process sensitive data and only upload encrypted training models, and then aggregates the model parameters to build a shared server model to protect data privacy. To motivate workers to build a high-quality global model in an efficacy way, we introduce an incentive mechanism that encourages workers to constantly collect fresh data to train accurate client models and boosts the global model training. We model the incentive-based interaction between the crowdsourcing platform and participating workers as a Stackelberg game, in which each side maximizes its own profit. We derive the Nash Equilibrium of the game to find the optimal solutions for the two sides. Experimental results confirm that iFedCrowd can complete secure crowdsourcing projects with high quality and efficiency. Xiangping Kang, Guoxian Yu, Jun Wang 0035, Wei Guo 0017, Carlotta Domeniconi, Jinglin Zhang 0001 |
AAAI | 6 |
| 2023 | Intelligent Internet of Things in Mammography Screening Using Multicenter Transformation Between Unified CapsulesabstractMammography screening is one of the important applications for the intelligent Internet of Things (IoT). Due to the efficient and personalized cyber-medicine system, early diagnosis can successfully reduce the breast cancer mortality rate by AI-driven healthcare. However, it is a huge challenge to extend the conventional single-center into the multicenter mammography screening, thus improving the effectiveness and robustness of intelligent IoT-based devices. To address this problem, we utilize multicenter mammograms by the modified capsule neural network and propose a novel framework called multicenter transformation between unified capsules (MLT-UniCaps) in this article. The proposed MLT-UniCaps is composed of Attentional Pose Embedding, Dynamic Source Capsule Traversal, and Adaptive Target Capsule Fusion to realize an intelligent remote assistant diagnosis. Attentional Pose Embedding extracts feature vectors via variations in position, orientation, scale, and lighting as the poses through an adversarial convolutional neural network with an attention-based layer. Based on the pose presentation, Dynamic Source Capsule Traversal deploys a dynamic routing mechanism between neurons to build a source cancer classifier for single-center mammography screening. Using the source cancer classifier, Adaptive Target Capsule Fusion integrates various centers of mammograms as the universal cancer detectors and optimizes heterogeneous distribution among them by the transformation-likelihood maximization. Owing to the three components, MLT-UniCaps effectively improves the results of single-center mammography screening and works in the multicenter breast cancer diagnosis. By comprehensive experiments on 58 965 samples, the proposed MLT-UniCaps obtains 90.1% of overall classification accuracy on single-center trials and 73.8% of overall F1 score on multicenter trials. All the experimental results illustrated that our MLT-UniCaps, an intelligent IoT-based clinical tool, inures the benefit of mammography screening. Xuegang Hu, Jinglin Zhang 0001, Chenchu Xu, Zhifan Gao |
IEEE Internet Things J. | 3 |
| 2023 | Mine Diversified Contents of Multispectral Cloud Images Along With Geographical Information for Multilabel ClassificationabstractMultispectral multilabel cloud image classification (MSMLCIC) aims to predict a set of labels presented in a multispectral (MS) cloud image, which usually contains more than one cloud type or weather system. However, the exploration of diversified contents reflected by multiple bands of MS image is limited and the consideration of geographical information (time and location information) is insufficient. To cope with the abovementioned problems, this work proposes the multispectral cloud image multilabel classifier with group feature extractor and geo-queries (MS-GoGo). With a group feature extractor, different bands of MS images are processed separately according to the content they reflected, and a group of distinctive yet complementary image features are generated. Geo-queries are responsible for implicitly embedding different labels with time and location information to probe the corresponding similar semantic ingredients. Due to the coarse classification of the existing dataset, a new dataset named LSCIDMR-V2 is generated with fine-grained cloud-type annotation and multichannel data. The experiment shows that, using the group feature extractor and geo-queries, the popular used metric subset accuracy is improved from 40.06 to 42.87 and 44.35, respectively. The proposed method achieves the mean average precision of 82.40, outperforming state-of-the-art methods. Dongxiaoyuan Zhao, Jinglin Zhang 0001, Cong Bai |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | An Effective Federated Learning Verification Strategy and Its Applications for Fault Diagnosis in Industrial IoT SystemsabstractDue to the diverse equipment and uneven load distribution in industrial environments, data regarding faults are often unbalanced. Moreover, data and models from clients may become contaminated or damaged, affecting diagnostic performance. To overcome these problems, this study proposes a stacking model for diagnosing interturn short circuit (ITSC) faults in permanent magnet synchronous motors (PMSMs). Federated learning (FL) is used to train the model to increase data security and overcome data islanding in distributed scenarios. Moreover, an improved verification strategy was adopted to select appropriate client models in each round to update the FL global model. We created a secondary server-side data set to validate the client weightings. The data set contains clean sample data for all ITSC fault categories. By calculating the fault diagnosis accuracy of the global model on the auxiliary data set, the model eliminates low-quality clients with uneven fault distributions. The improved particle swarm optimization (PSO) is used to optimize the weight coefficients of clients involved in aggregation, improving the robustness of the aggregation strategy under a joint learning system. In evaluation experiments, compared with the federated average (FedAvg) model, the proposed dynamic verification model exhibited the better diagnostic accuracy in situations of data imbalance, incurred lower communication costs, and prevented local oscillations in the model. Yuanjiang Li, Kai Zhu 0005, Cong Bai, Jinglin Zhang 0001 |
IEEE Internet Things J. | 5 |
| 2022 | Rainformer: Features Extraction Balanced Network for Radar-Based Precipitation NowcastingabstractPrecipitation nowcasting is one of the fundamental challenges in natural hazard research. High-intensity rainfall, especially the rainstorm, will lead to the enormous loss of people’s property. Existing methods usually utilize convolution operation to extract rainfall features and increase the network depth to expand the receptive field to obtain fake global features. Although this scheme is simple, only local rainfall features can be extracted leading to insensitivity to high-intensity rainfall. This letter proposes a novel precipitation nowcasting framework named Rainformer, in which, two practical components are proposed: the global features extraction unit and the gate fusion unit (GFU). The former provides robust global features learning ability depending on the window-based multi-head self-attention (W-MSA) mechanism, while the latter provides a balanced fusion of local and global features. Rainformer has a simple yet efficient architecture and significantly improves the accuracy of rainfall prediction, especially on high-intensity rainfall. It offers a potential solution for real-world applications. The experimental results show that Rainformer outperforms seven state of the arts methods on the benchmark database and provides more insights into the high-intensity rainfall prediction task. Cong Bai, Jinglin Zhang 0001, Shengyong Chen |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2022 | Multi-granularity episodic contrastive learning for few-shot learning
Pengfei Zhu 0001, Yu Wang 0106, Jinglin Zhang 0001 |
Pattern Recognit. | 4 |
| 2022 | LSCIDMR: Large-Scale Satellite Cloud Image Database for Meteorological ResearchabstractPeople can infer the weather from clouds. Various weather phenomena are linked inextricably to clouds, which can be observed by meteorological satellites. Thus, cloud images obtained by meteorological satellites can be used to identify different weather phenomena to provide meteorological status and future projections. How to classify and recognize cloud images automatically, especially with deep learning, is an interesting topic. Generally speaking, large-scale training data are essential for deep learning. However, there is no such cloud images database to date. Thus, we propose a large-scale cloud image database for meteorological research (LSCIDMR). To the best of our knowledge, it is the first publicly available satellite cloud image benchmark database for meteorological research, in which weather systems are linked directly with the cloud images. LSCIDMR contains 104 390 high-resolution images, covering 11 classes with two different annotation methods: 1) single-label annotation and 2) multiple-label annotation, called LSCIDMR-S and LSCIDMR-M, respectively. The labels are annotated manually, and we obtain a total of 414 221 multiple labels and 40 625 single labels. Several representative deep learning methods are evaluated on the proposed LSCIDMR, and the results can serve as useful baselines for future research. Furthermore, experimental results demonstrate that it is possible to learn effective deep learning models from a sufficiently large image database for the cloud image classification. Cong Bai, Minjing Zhang, Jinglin Zhang 0001, Jianwei Zheng 0001, Shengyong Chen |
IEEE Trans. Cybern. | 3 |
| 2022 | Automated CCA-MWF Algorithm for Unsupervised Identification and Removal of EOG Artifacts From EEGabstractAffective brain computer interface (ABCI) enables machines to perceive, understand, express and respond to people's emotions. Therefore, it is expected to play an important role in emotional care and mental disorder detection. EEG signals are most frequently adopted as the physiology measurement in ABCI applications. Eye blinking and movements introduce lots of artifacts into raw EEG data, which seriously affect the quality of EEG signal and the subsequent emotional EEG feature engineering and recognition. In this paper, we propose a fully automatic and unsupervised ocular artifact identification and removal algorithm named automated canonical correlation analysis (CCA)-multi-channel wiener filter (MWF) (ACCAMWF). Firstly, spatial distribution entropy (SDE) and spectral entropy (SE) are computed to automatically annotate artifact segments. Then, CCA algorithm is used to extract neural signal from artifact contaminated data to further supplement the clean EEG data. Finally, MWF is trained to remove ocular artifacts from multiple channel EEG data adaptively. Extensive experiments have been carried out on semi-simulated EEG/EOG dataset and real eye blinking-contaminated EEG dataset to verify the effectiveness of our method when compared to two state-of-the-art algorithms. The results clearly demonstrate that ACCAMWF is a promising solution for removing EOG artifacts from emotional EEG data. Minmin Miao, Baoguo Xu, Jinglin Zhang 0001, Joel J. P. C. Rodrigues, Victor Hugo C. de Albuquerque |
IEEE J. Biomed. Health Informatics | 4 |
| 2020 | Deep Adversarial Discrete Hashing for Cross-Modal RetrievalabstractCross-modal hashing has received widespread attentions on cross-modal retrieval task due to its superior retrieval efficiency and low storage cost. However, most existing cross-modal hashing methods learn binary codes directly from multimedia data, which cannot fully utilize the semantic knowledge of the data. Furthermore, they cannot learn the ranking based similarity relevance of data points with multi-label. And they usually use a relax constraint of hash code which causes non-negligible quantization loss in the optimization. In this paper, a hashing method called Deep Adversarial Discrete Hashing (DADH) is proposed to address these issues for cross-modal retrieval. The proposed method uses adversarial training to learn features across modalities and ensure the distribution consistency of feature representations across modalities. We also introduce a weighted cosine triplet constraint which can make full use of semantic knowledge from the multi-label to ensure the precise ranking relevance of item pairs. In addition, we use a discrete hashing strategy to learn the discrete binary codes without relaxation, by which the semantic knowledge from label in the hash codes can be preserved while the quantization loss can be minimized. Ablation experiments and comparison experiments on two cross-modal databases show that the proposed DADH improves the performance and outperforms several state-of-the-art hashing methods for cross-modal retrieval. Cong Bai, Jinglin Zhang 0001, Shengyong Chen |
ICMR | 4 |
| 2019 | Supervised learning based discrete hashing for image retrieval
Cong Bai, Jinglin Zhang 0001, Zhi Liu 0003, Shengyong Chen |
Pattern Recognit. | 3 |
| 2015 | K-means based histogram using multiresolution feature vectors for color texture database retrieval
Cong Bai, Jinglin Zhang 0001, Zhi Liu 0003, Wanlei Zhao |
Multim. Tools Appl. | 2 |