Lei Liu 0029

dblp:21/2715-29 · DBLP profile ↗
← Back
15ranked-venue papers
1as first author
14since 2021 · last 2026
0000-0002-0625-6248ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 9 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
YearPublicationVenuePosition
2026 TDFormer: A novel triple decoupled transformer for accurate multi-step wind power forecasting
Lei Liu 0029, Qiuju Chen, Bin Li 0025
Eng. Appl. Artif. Intell.1
2026 Curriculum trustworthy multi-modal learning
abstract
Trustworthy multi-modal learning reliably integrates multiple data sources. However, current methods often face a challenge, which is the inherent non-convex nature of deep neural networks. It leads to their susceptibility to local minima, ultimately resulting in a reduced capacity for generalization. To address this issue, we first create a theoretical framework, which extends the application of curriculum learning in multi-modal scenarios. Secondly, we propose a novel curriculum termed the Dynamic SRM Curriculum (DSRMC). It consists of two modules: a scoring function and a training schedule. The scoring function sorts samples from simple to complex. The training scheduler aims to manage the quantity of samples supplied at each round during training. DSRMC facilitates positioning the learned model in a flatter region of the loss landscape, thereby enhancing its overall generalization ability. Building on DSRMC, we eventually propose an innovative method termed as Curriculum Trustworthy Multi-modal Learning (CTML). It applies DSRMC in multi-modal learning application scenarios. Extensive experiments conducted on three open datasets show that the proposed CTML outperforms state-of-the-art methods, with a maximum improvement of 6.7% in macro F1 score. Our code and dataset are publicly available on https://github.com/HackerHyper/DSRMC.git .
Xin Zou 0001, Jun Sun 0014, Lingfang Zeng, Linqing Feng, Lei Liu 0029, Chang Tang
Expert Syst. Appl.7
2026 Large language models are good attackers: Efficient and stealthy textual backdoor attacks
Ziqiang Li 0001, Yueqi Zeng, Lei Liu 0029, Zhangjie Fu 0001, Bin Li 0025
Pattern Recognit.4
2026 Generative Diffusion Contrastive Network for Multi-View Clustering
abstract
In recent years, Multi-View Clustering (MVC) has been significantly advanced under the influence of deep learning. By integrating heterogeneous data from multiple views, MVC enhances clustering analysis, making multi-view fusion critical to clustering performance. However, multi-view fusion remains challenged by low-quality data, primarily stemming from two reasons: 1) Certain views are contaminated by noisy data. 2) Some views suffer from missing data. This paper proposes a novel Stochastic Generative Diffusion Fusion (SGDF) method to address this problem. SGDF leverages a multiple generative mechanism for the multi-view feature of each sample. It exhibits robustness against low-quality data. Building on SGDF, we further present the Generative Diffusion Contrastive Network (GDCN). Extensive experiments show that GDCN achieves the state-of-the-art results in deep MVC tasks. The source code is publicly available athttps://github.com/HackerHyper/GDCN.
Xin Zou 0001, Lei Liu 0029, Chang Tang, Li-Rong Dai 0001
IEEE Signal Process. Lett.4
2025 VQLTI: Long-Term Tropical Cyclone Intensity Forecasting with Physical Constraints
abstract
Tropical cyclone (TC) intensity forecasting is crucial for early disaster warning and emergency decision-making. Numerous researchers have explored deep-learning methods to address computational and post-processing issues in operational forecasting. Regrettably, they exhibit subpar long-term forecasting capabilities. We use two strategies to enhance long-term forecasting. (1) By enhancing the matching between TC intensity and spatial information, we can improve long-term forecasting performance. (2) Incorporating physical knowledge and physical constraints can help mitigate the accumulation of forecasting errors. To achieve the above strategies, we propose the VQLTI framework. VQLTI transfers the TC intensity information to a discrete latent space while retaining the spatial information differences, using large-scale spatial meteorological data as conditions. Furthermore, we leverage the forecast from the weather prediction model FengWu to provide additional physical knowledge for VQLTI. Additionally, we calculate the potential intensity (PI) to impose physical constraints on the latent variables. In the global long-term TC intensity forecasting, VQLTI achieves state-of-the-art results for the 24h to 120h, with the MSW (Maximum Sustained Wind) forecast error reduced by 35.65%-42.51% compared to ECMWF-IFS.
Lei Liu 0029, Tao Han 0002, Bin Li 0025, Lei Bai 0001
AAAI2
2025 Global Tropical Cyclone Intensity Forecasting with Multi-modal Multi-scale Causal Autoregressive Model
abstract
Accurate forecasting of tropical cyclone (TC) intensity is crucial for formulating disaster risk reduction strategies. Current methods predominantly rely on limited spatiotemporal information from ERA5 data and neglect the causal relationships between these physical variables, failing to fully capture the spatial and temporal patterns required for intensity forecasting. To address this issue, we propose a Multi-modal multi-Scale Causal AutoRegressive model (MSCAR), which is the first model that combines causal relationships with large-scale multimodal data for global TC intensity autoregressive forecasting. Furthermore, given the current absence of a TC dataset that offers a wide range of spatial variables, we present the Satellite and ERA5-based Tropical Cyclone Dataset (SETCD), which stands as the longest and most comprehensive global dataset related to TCs. Experiments on the dataset show that MSCAR outperforms the state-of-the-art methods, achieving maximum reductions in global and regional forecast errors of 9.52% and 6.74%, respectively. The code and dataset are publicly available at https://github.com/1457756434/MSCAR.git.
Lei Liu 0029, Tao Han 0002, Bin Li 0025, Lei Bai 0001
ICASSP3
2025 Trusted Mamba Contrastive Network for Multi-View Clustering
abstract
Multi-view clustering can partition data samples into their categories by learning a consensus representation in an unsupervised way and has received more and more attention in recent years. However, there is an untrusted fusion problem. The reasons for this problem are as follows: 1) The current methods ignore the presence of noise or redundant information in the view; 2) The similarity of contrastive learning comes from the same sample rather than the same cluster in deep multi-view clustering. It causes multi-view fusion in the wrong direction. This paper proposes a novel multi-view clustering network to address this problem, termed as Trusted Mamba Contrastive Network (TMCN). Specifically, we present a new Trusted Mamba Fusion Network (TMFN), which achieves a trusted fusion of multi-view data through a selective mechanism. Moreover, we align the fused representation and the view-specific representation using the Average-similarity Contrastive Learning (AsCL) module. AsCL increases the similarity of view presentation from the same cluster, not merely from the same sample. Extensive experiments show that the proposed method achieves state-of-the-art results in deep multi-view clustering tasks. The source code is available at https://github.com/HackerHyper/TMCN.
Xin Zou 0001, Lei Liu 0029, Zhangmin Huang, Chang Tang, Li-Rong Dai 0001
ICASSP3
2025 Dynamic SRM Curriculum for Trustworthy Multi-modal Classification
abstract
Trustworthy multi-modal learning integrates multiple sources of data reliably. However, the current methods still focus on performance improvement by developing deep multi-modal networks. These approaches frequently encounter challenges due to the inherent non-convex nature of deep neural networks and their vulnerability to local minima, ultimately leading to a diminished ability for generalization. To address this problem, we present a novel curriculum termed the Dynamic SRM Curriculum (DSRMC). Within DSRMC, the deep trustworthy multi-modal networks undergo training with data provided sequentially, progressing from simple to complex samples. This training strategy mimics the human learning process, commencing with fundamental concepts and gradually advancing to tackle more complex and abstract ideas. Building upon DSRMC, we propose an innovative Curriculum Trustworthy Multi-modal Learning (CTML) method. CTML makes it easier to place the learned model in a flatter area, which improves its overall ability for generalization. Comprehensive experiments on three public datasets demonstrate that the proposed CTML performs better than state-of-the-art methods, achieving a maximum improvement of 6.7% on macroF1.
Cui Yu, Xin Zou 0001, Zhangmin Huang, Chenshu Hu, Jun Sun 0014, Bo Lyu, Lei Liu 0029, Chang Tang, Li-Rong Dai 0001
ICASSP8
2025 CLIP Multi-modal Hashing for Multimedia Retrieval
Mingkai Sheng, Zhangmin Huang, Jingfei Chang, Jinling Jiang, Lei Liu 0029
MMM (1)8
2025 Chiron-o1: Igniting Multimodal Large Language Models towards Generalizable Medical Reasoning via Mentor-Intern Collaborative Search
abstract
Multimodal large language models (MLLMs) have begun to demonstrate robust reasoning capabilities on general tasks, yet their application in the medical domain remains in its early stages. Constructing chain-of-thought (CoT) training data is essential for bolstering the reasoning abilities of medical MLLMs. However, existing approaches exhibit a deficiency in offering a comprehensive framework for searching and evaluating effective reasoning paths towards critical diagnosis. To address this challenge, we propose Mentor-Intern Collaborative Search (MICS), a novel reasoning-path searching scheme to generate rigorous and effective medical CoT data. MICS first leverages mentor models to initialize the reasoning, one step at a time, then prompts each intern model to continue the thinking along those initiated paths, and finally selects the optimal reasoning path according to the overall reasoning performance of multiple intern models. The reasoning performance is determined by an MICS-Score, which assesses the quality of generated reasoning paths. Eventually, we construct MMRP, a multi-task medical reasoning dataset with ranked difficulty, and Chiron-o1, a new medical MLLM devised via a curriculum learning strategy, with robust visual question-answering and generalizable reasoning capabilities. Extensive experiments demonstrate that Chiron-o1, trained on our CoT dataset constructed using MICS, achieves state-of-the-art performance across a list of medical visual question answering and reasoning benchmarks. Codes are available at https://github.com/Yankai96/Chiron-o1
Yankai Jiang 0003, Wenjie Lou, Lilong Wang, Mianxin Liu, Lei Liu 0029, Xiaosong Wang 0001
NeurIPS8
2025 Adaptive Confidence Multi-View Learning
abstract
Multi-view hashing is a crucial technology for multimedia retrieval because it transforms heterogeneous data from many viewpoints into binary hash codes. However, the existing approaches focus mostly on the complementarity among multiple views while being without confidence fusion. Furthermore, redundant noise is present in the single-view data in real-world application contexts. We present an innovativeAdaptive Confidence Multi-View Learning(ACMVL) method to perform confidence fusion and remove extraneous noise. Initially, a confidence network is constructed to eliminate noise data and extract useful information from various single-view features. Moreover, an adaptive confidence multi-view network is utilized to quantify the confidence of each view and further fuse multiple view features using a weighted summation. Here, we propose anAutomatic View Confidence Metric(AVCM) as a score for evaluating the confidence of views. Finally, to improve the semantic representation of the fused feature, a dilation network is created. Based on ACMVL, we introduce a novelAdaptive Confidence Multi-View Hashing(ACMVH) method. To our knowledge, we are the pioneers in using confidence learning for multimedia retrieval. Comprehensive experiments on three publicly available datasets demonstrate that our ACMVH outperforms the state-of-the-art methods (maximum improvement of 3.24% on mAP).
Lei Liu 0029, Chang Tang, Li-Rong Dai 0001
IEEE Trans. Multim.2
2024 Adaptive Confidence Multi-View Hashing for Multimedia Retrieval
abstract
The multi-view hash method converts heterogeneous data from multiple views into binary hash codes, which is one of the critical technologies in multimedia retrieval. However, the current methods mainly explore the complementarity among multiple views while lacking confidence in learning and fusion. Moreover, in practical application scenarios, the single-view data contains redundant noise. To conduct confidence learning and eliminate unnecessary noise, we propose a novel Adaptive Confidence Multi-View Hashing (ACMVH) method. First, a confidence network is developed to extract useful information from various single-view features and remove noise information. Furthermore, an adaptive confidence multi-view network is employed to measure the confidence of each view and then fuse multi-view features through a weighted summation. Lastly, a dilation network is designed to further enhance the feature representation of the fused features. To the best of our knowledge, we pioneer the application of confidence learning into the field of multimedia retrieval. Extensive experiments on two public datasets show that the proposed ACMVH performs better than state-of-the-art methods (maximum increase of 3.24%). The source code is available at https://github.com/HackerHyper/ACMVH.
Zhangmin Huang, Lei Liu 0029, Lingfang Zeng
ICASSP5
2024 Boosted Curriculum Multi-View Hashing for Multimedia Retrieval
abstract
The multi-view hash method plays a pivotal role in multimedia retrieval, transforming diverse data from multiple perspectives into binary hash codes. While existing methods primarily emphasize complementarity across multiple views, they often face challenges associated with the non-convex nature of deep neural networks, ultimately causing a decrease in generalization ability. To overcome this limitation, we propose a novel curriculum calledAutomatic Multiple Loss Curriculum(AMLC). In AMLC, the deep multi-view hashing network undergoes training with data presented sequentially, progressing from simple to complex samples. This training strategy mirrors the human learning process, commencing with fundamental concepts and progressively advancing to tackle more intricate and abstract ideas. Building upon AMLC, we propose theBoosted Curriculum Multi-View Hashing(BCMVH) method. BCMVH facilitates the positioning of the learned model in a more flat region, enhancing its overall generalization capability. Extensive experiments conducted on three public datasets demonstrate that the proposed BCMVH outperforms state-of-the-art methods, achieving a maximum improvement of 3.17% in terms of mean Average Precision.
Zhangmin Huang, Lei Liu 0029, Chang Tang, Li-Rong Dai 0001
IEEE Signal Process. Lett.3
2022 Prediction of Vertical Profile of NO₂ Using Deep Multimodal Fusion Network Based on the Ground-Based 3-D Remote Sensing
abstract
The vertical distribution profiles of NO2are essential for understanding the mechanisms, detecting near-surface emissions, and tracking pollutant transportation at high altitude. However, most of the published NO2studies are based on the surface 2-D measurements. The ground-based 3-D remote-sensing stations were recently built to measure vertical distribution profiles of NO2. However, the stations were spatially sparse due to the high cost and could not make the measurements without sunlight. In this study, we first developed a multimodel fusion network (MF-net) based on the sparse vertical observations from the Jing-Jin-Ji region. We achieved the 3-D profile prediction of NO2in the range of 39.005–41.405N and 115.005–117.905E with 24-h coverage. The MF-net significantly surpassed the conventional WRF-CHEM model and provided a more accurate evaluation of the NO2transmission between Beijing and the neighboring cities. Besides, the MF-net covers the monitoring of NO2to the whole study area and extends the monitoring time to the entire day (24 h), making it serviceable for continuous spatial-temporal estimation of NO2and its transmission in pollution events. The MF-net provides more robust data support to formulate reasonable and effective pollution prevention and control measures.
Shulin Zhang, Bo Li 0005, Lei Liu 0029, Qihou Hu, Yizhi Zhu, Mingzhai Sun, Cheng Liu 0005
IEEE Trans. Geosci. Remote. Sens.3
2020 Dehaze of Cataractous Retinal Images Using an Unpaired Generative Adversarial Network
abstract
Cataracts are the leading cause of visual impairment worldwide. Examination of the retina through cataracts using a fundus camera is challenging and error-prone due to degraded image quality. We sought to develop an algorithm to dehaze such images to support diagnosis by either ophthalmologists or computer-aided diagnosis systems. Based on the generative adversarial network (GAN) concept, we designed two neural networks: CataractSimGAN and CataractDehazeNet. CataractSimGAN was intended for the synthesis of cataract-like images through unpaired clear retinal images and cataract images. CataractDehazeNet was trained using pairs of synthesized cataract-like images and the corresponding clear images through supervised learning. With two networks trained independently, the number of hyper-parameters was reduced, leading to better performance. We collected 400 retinal images without cataracts and 400 hazy images from cataract patients as the training dataset. Fifty cataract images and the corresponding clear images from the same patients after surgery comprised the test dataset. The clear images after surgery were used for reference to evaluate the performance of our method. CataractDehazeNet was able to enhance the degraded image from cataract patients substantially and to visualize blood vessels and the optic disc, while actively suppressing the artifacts common in application of similar methods. Thus, we developed an algorithm to improve the quality of the retinal images acquired from cataract patients. We achieved high structure similarity and fidelity between processed images and images from the same patients after cataract surgery.
Yuhao Luo 0001, Kun Chen 0005, Lei Liu 0029, Jianbo Mao, Genjie Ke, Mingzhai Sun
IEEE J. Biomed. Health Informatics3