Wei Liu 0123

dblp:49/3283-123 · DBLP profile ↗
← Back
20ranked-venue papers
5as first author
17since 2021 · last 2026
0000-0002-2343-8254ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 2 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 7 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Security and privacy · 1
YearPublicationVenuePosition
2026 EKT-ML: An efficient knowledge tracing model with multi-task learning
Wei Liu 0123, Bo Yang 0011, Haotian Su, Yaowei Wang 0001, Qing Li 0001
Expert Syst. Appl.1
2026 Bidirectional multi-scale masked transformer for single image deraining
Miao Qi, Wei Liu 0123, Ziqiang Huang, Hangyu Nie
Multim. Syst.2
2026 No-reference dehazed image quality assessment via perception-driven interactive feature representation learning
Hangyu Nie, Ziqiang Huang, Miao Qi, Junjun Jiang, Jiayi Ma 0001, Wei Liu 0123
Pattern Recognit.6
2026 Self-Iteration Image Haze Removal Using a Deep Curve-Dehazing Model
abstract
This paper proposes a novel dehazing method termed Haze-Restoration Curve Model (HRCM), which transforms the single-image dehazing task into a specific curve estimation problem, achieving haze removal through an intuitive and simple nonlinear curve mapping. Unlike methods based on Atmospheric Scattering Model (ASM), HRCM does not require the computation of complex physical parameters. Instead, it estimates two intuitive curvature adjustment coefficients. Moreover, compared to recent end-to-end dehazing methods, HRCM circumvents the challenging modeling of static mapping functions, thereby improving the generalization ability and dehazing performance of the model. All of these are attributed to a meticulously designed dehazing curve, which first reversing the hazy image to highlight obscured regions, and then specifies a set of high-order functions to remap hazy pixels for image restoration. Moreover, to estimate the curve parameters, we designed a dual-branch Deep Dehaze Curve Estimation Network(DDCEN), which consists of the Residual Swin Transformer Block(RTSB) and the Large kernel convolutional Attention Block(LAB). Specifically, RTSB captures the global fog density distribution features of foggy images by introducing window self-attention and shifted window mechanisms, providing support for global semantic information for subsequent parameter estimation. LAB captures local multi-scale features by constructing a large receptive field, and uses the attention mechanism of feature pooling in horizontal and vertical directions to focus on detail regions, refining the local details of the parameter map. Extensive experiments on synthetic and real-world hazy image datasets demonstrate that the proposed approach achieves superior performance in terms of quantitative accuracy and subjective visual quality compared to the current state-of-the-art methods. The source code of our HRCM is available at https://github.com/larrylanrui/HRCM.
Wei Liu 0123, Rui Nan, Jiayi Ma 0001, Xin Chen 0003, Guoping Qiu
IEEE Trans. Circuits Syst. Video Technol.1
2026 MolReFlect: Toward In-Context Fine-Grained Alignments Between Molecules and Texts
abstract
Molecule discovery is a pivotal research field, impacting everything from medicine to materials. Recently, Large Language Models (LLMs) have been widely adopted in molecular understanding and generation, serving as a bridge between the molecular space and the natural language space, yet the alignment between molecules and their corresponding captions remains a significant challenge. Previous endeavors typically treat molecules as monolithic inputs, lacking an intermediate reasoning process and sacrificing explainability. In this work, we define fine-grained alignments as the precise correspondence between a molecule's sub-structures and the textual phrases that explain their properties. These alignments are crucial for LLMs to understand molecules in a more accurate and explainable manner. Normally, such fine-grained alignments require expert annotation, which is both costly and time-consuming. To allow LLMs to automatically label and learn the fine-grained alignments, we propose MolReFlect, a novel teacher-student framework, where a teacher LLM first generates and refines mappings between caption phrases and SMILES substructures and then explicitly teaches these detailed alignments to a student LLM. Experimental results demonstrate that MolReFlect enables LLMs to significantly outperform previous baselines, achieving the state-of-the-art performance in the molecule-caption translation task. Our codes are available via: https://github.com/phenixace/MolReFlect.
Jiatong Li 0003, Wei Liu 0123, Jingdi Lei, Di Zhang 0026, Wenqi Fan, Dongzhan Zhou, Qing Li 0001
IEEE Trans. Knowl. Data Eng.3
2025 ChemVLM: Exploring the Power of Multimodal Large Language Models in Chemistry Area
abstract
Large Language Models (LLMs) have achieved remarkable success and have been applied across various scientific fields, including chemistry. However, many chemical tasks require the processing of visual information, which cannot be successfully handled by existing chemical LLMs. This brings a growing need for models capable of integrating multimodal information in the chemical domain. In this paper, we introduce ChemVLM, an open-source chemical multimodal large language model specifically designed for chemical applications. ChemVLM is trained on a carefully curated bilingual multimodal dataset that enhances its ability to understand both textual and visual chemical information, including molecular structures, reactions, and chemistry examination questions. We develop three datasets for comprehensive evaluation, tailored to Chemical Optical Character Recognition (OCR), Multimodal Chemical Reasoning (MMCR), and Multimodal Molecule Understanding tasks. We benchmark ChemVLM against a range of open-source and proprietary multimodal large language models on various tasks. Experimental results demonstrate that ChemVLM achieves competitive performance across all evaluated tasks.
Junxian Li 0001, Di Zhang 0026, Xunzhi Wang, Zeying Hao, Jingdi Lei, Cai Zhou, Wei Liu 0123, Yaotian Yang, Xinrui Xiong, Weiyun Wang, Zhe Chen 0013, Wenhai Wang, Wei Li 0076, Mao Su, Shufei Zhang, Wanli Ouyang, Dongzhan Zhou
AAAI8
2025 FAGCL: frequency-based augmentation graph contrastive learning for recommendation
Bo Yang 0011, Zimu Li, Wei Liu 0123
Appl. Intell.4
2025 DMAM: Difficulty-enhanced multi-view attention-based model for knowledge tracing
Bo Yang 0011, Wei Liu 0123
Knowl. Based Syst.3
2025 A simple yet effective difficulty-aware bucketed fine-tuning strategy for LLM-based recommendation
Qianyang Zhu, Bo Yang 0011, Wei Liu 0123, Jiajin Wu
Knowl. Based Syst.3
2025 Large Language Models are in-Context Molecule Learners
abstract
Large Language Models (LLMs) have demonstrated exceptional performance in biochemical tasks, especially the molecule caption translation task, which aims to bridge the gap between molecules and natural language texts. However, previous methods in adapting LLMs to the molecule-caption translation task required extra domain-specific pre-training stages, suffered weak alignment between molecular and textual spaces, or imposed stringent demands on the scale of LLMs. To resolve the challenges, we propose In-Context Molecule Adaptation (ICMA), as a new paradigm allowing LLMs to learn the molecule-text alignment from context examples via In-Context Molecule Tuning. Specifically, ICMA incorporates the following three stages: Hybrid Context Retrieval, Post-retrieval Re-ranking, and In-context Molecule Tuning. Initially, Hybrid Context Retrieval utilizes BM25 Caption Retrieval and Molecule Graph Retrieval to retrieve similar informative context examples. Additionally, Post-retrieval Re-ranking is composed of Sequence Reversal and Random Walk selection to further improve the quality of retrieval results. Finally, In-Context Molecule Tuning unlocks the in-context learning and reasoning capability of LLMs with the retrieved examples and adapts the parameters of LLMs for better alignment between molecules and texts. Experimental results demonstrate that ICMA can empower LLMs to achieve state-of-the-art or comparable performance without extra training corpora and intricate structures, showing that LLMs are inherently in-context molecule learners.
Jiatong Li 0003, Wei Liu 0123, Zhihao Ding, Wenqi Fan, Qing Li 0001
IEEE Trans. Knowl. Data Eng.2
2024 MCL4SRec: A Sequential Recommendation Model with Multi-level Contrastive Learning
abstract
Sequential recommendation (SR) plays an important role across various platforms, aiming to predict users’ next items of interest based on their historical interaction sequences. Recent SR studies have employed deep learning techniques, such as Recurrent Neural Networks and Self-Attention (SA) mechanism, demonstrating promising results. Inspired by the emergence of contrastive learning methods, some SR models have utilized contrastive learning to improve the accuracy of recommendations. However, existing SR models employing contrastive learning primarily construct positive and negative sample pairs only from user interaction sequences, i.e., through sequence-level contrastive learning. In our research, we argue that there also exists semantic similarities between items, which can be used to conduct the item-level constructive learning, resulting in better recommendation accuracy. In this paper, we propose MCL4SRec, an SA-based SR model that combines sequence-level and item-level contrastive learning to enhance recommendation accuracy. In our proposed MCL4SRec, the item-level contrastive learning module utilizes items’ category information to construct positive and negative sample pairs, capturing semantic similarities and differences between items. Additionally, in MCL4SRec, we propose to use more side information such as category and brand to further improve the accuracy of recommendations. We conduct extensive experiments on three widely-used datasets to evaluate the proposed MCL4SRec. Experimental results indicate that the average improvements compared with the recent well-known baselines range from ${7. 7 3 \%}$ to ${1 6. 1 8 \%}$ in HR and NDCG, demonstrating the effectiveness of MCL4SRec for SR tasks.
Zhuohan Hu, Bo Yang 0011, Jialiang Lin 0004, Jiajin Wu, Wei Liu 0123
FUSION5
2024 AFBench: A Large-scale Benchmark for Airfoil Design
abstract
Data-driven generative models have emerged as promising approaches towards achieving efficient mechanical inverse design. However, due to prohibitively high cost in time and money, there is still lack of open-source and large-scale benchmarks in this field. It is mainly the case for airfoil inverse design, which requires to generate and edit diverse geometric-qualified and aerodynamic-qualified airfoils following the multimodal instructions, \emph{i.e.,} dragging points and physical parameters. This paper presents the open-source endeavors in airfoil inverse design, \emph{AFBench}, including a large-scale dataset with 200 thousand airfoils and high-quality aerodynamic and geometric labels, two novel and practical airfoil inverse design tasks, \emph{i.e.,} conditional generation on multimodal physical parameters, controllable editing, and comprehensive metrics to evaluate various existing airfoil inverse design methods. Our aim is to establish \emph{AFBench} as an ecosystem for training and evaluating airfoil inverse design methods, with a specific focus on data-driven controllable inverse design models by multimodal instructions capable of bridging the gap between ideas and execution, the academic research and industrial applications. We have provided baseline models, comprehensive experimental observations, and analysis to accelerate future research. Our baseline model is trained on an RTX 3090 GPU within 16 hours. The codebase, datasets and benchmarks will be available at \url{https://hitcslj.github.io/afbench/}.
Jian Liu 0036, Hairun Xie, Wei Liu 0123, Wanli Ouyang, Junjun Jiang, Xianming Liu 0005, Shixiang Tang
NeurIPS6
2024 Empowering and Assessing the Utility of Large Language Models in Crop Science
abstract
Large language models (LLMs) have demonstrated remarkable efficacy across knowledge-intensive tasks. Nevertheless, their untapped potential in crop science presents an opportunity for advancement. To narrow this gap, we introduce CROP, which includes a novel instruction tuning dataset specifically designed to enhance LLMs’ professional capabilities in the crop science sector, along with a benchmark that serves as a comprehensive evaluation of LLMs’ understanding of the domain knowledge. The CROP dataset is curated through a task-oriented and LLM-human integrated pipeline, comprising 210,038 single-turn and 1,871 multi-turn dialogues related to crop science scenarios. The CROP benchmark includes 5,045 multiple-choice questions covering three difficulty levels. Our experiments based on the CROP benchmark demonstrate notable enhancements in crop science-related tasks when LLMs are fine-tuned with the CROP dataset. To the best of our knowledge, CROP dataset is the first-ever instruction tuning dataset in the crop science domain. We anticipate that CROP will accelerate the adoption of LLMs in the domain of crop science, ultimately contributing to global food production.
Renqi Chen, Wei Liu 0123, Zhonghang Yuan, Xinzhe Zheng 0001, Zhefan Wang 0002, Hang Yan 0001, Han-Sen Zhong, Xiqing Wang, Wanli Ouyang, Nanqing Dong
NeurIPS4
2022 Deep representation learning for face hallucination
Tao Lu 0001, Yu Wang 0140, Ruobo Xu, Wei Liu 0123, Wenhua Fang, Yanduo Zhang
Multim. Tools Appl.4
2021 Face Hallucination via Split-Attention in Split-Attention Network
abstract
Recently, convolutional neural networks (CNNs) have been widely employed to promote the face hallucination due to the ability to predict high-frequency details from a large number of samples. However, most of them fail to take into account the overall facial profile and fine texture details simultaneously, resulting in reduced naturalness and fidelity of the reconstructed face, and further impairing the performance of downstream tasks (e.g., face detection, facial recognition). To tackle this issue, we propose a novel external-internal split attention group (ESAG), which encompasses two paths responsible for facial structure information and facial texture details, respectively. By fusing the features from these two paths, the consistency of facial structure and the fidelity of facial details are strengthened at the same time. Then, we propose a split-attention in split-attention network (SISN) to reconstruct photorealistic high-resolution facial images by cascading several ESAGs. Experimental results on face hallucination and face recognition unveil that the proposed method not only significantly improves the clarity of hallucinated faces, but also encourages the subsequent face recognition performance substantially. Codes have been released at https://github.com/mdswyz/SISN-Face-Hallucination.
Tao Lu 0001, Yuanzhi Wang, Yanduo Zhang, Yu Wang 0140, Wei Liu 0123, Zhongyuan Wang 0001, Junjun Jiang
ACM Multimedia5
2021 Single Image Super-Resolution via Multi-Scale Information Polymerization Network
abstract
Recently, the performances of deep convolution neural networks (CNNs)-based single-image super-resolution (SISR) have been significantly improved. However, most of the existing CNN-based SISR methods mainly focus on wider or deeper networks and ignore the potential relationship between multi-scale features, leading to the limited representation ability of the reconstructed network. To address this problem, we propose a new multi-scale information polymerization network (MIPN). Specifically, we propose a multi-scale information polymerization block (MIPB), which uses convolution layers of different convolution kernel sizes to extract multi-scale image features, and effectively polymerizate the extracted features together to obtain fine image features. Moreover, we also propose a shallow residual block in MIPB. Compared with the traditional convolution layer, this proposed block can effectively extract image features without increasing the number of parameters. Extensive experiments show that the proposed method performs better than several state-of-the-art methods in quantitative and visual quality indicators.
Tao Lu 0001, Yu Wang 0140, Jiaming Wang 0001, Wei Liu 0123, Yanduo Zhang
IEEE Signal Process. Lett.4
2021 Image Defogging Quality Assessment: Real-World Database and Method
abstract
Fog removal from an image is an active research topic in computer vision. However, current literature is weak in the following two areas which in many ways are hindering progress for developing defogging algorithms. First, there is no true real-world and naturally occurring foggy image datasets suitable for developing defogging models. Second, there is no suitable mathematically simple and easy to use image quality assessment (IQA) methods for evaluating the visual quality of defogged images. We address these two aspects in this paper. We first introduce a new foggy image dataset called multiple real-world foggy image dataset (MRFID). MRFID contains foggy and clear images of 200 outdoor scenes. For each scene, one clear image and 4 foggy images of different densities defined as slightly foggy, moderately foggy, highly foggy, and extremely foggy, are manually selected from images taken from these scenes over the course of one calendar year. We then process the foggy images of MRFID using 16 defogging methods to obtain 12,800 defogged images (DFIs) and perform a comprehensive subjective evaluation of the visual quality of the DFIs. Through collecting the mean opinion score (MOS) of 120 subjects and evaluating a variety of fog-relevant image features, we have developed a new Fog-relevant Feature based SIMilarity index (FRFSIM) for assessing the visual quality of DFIs. We present extensive experimental results to show that our new visual quality assessment measure, the FRFSIM, is more consistent with the MOS than other IQA methods and is therefore more suitable for evaluating defogged images than other state-of-the-art IQA methods. Our dataset and relevant code are available at http://www.vistalab.ac.cn/MRFID-for-defogging/.
Wei Liu 0123, Fei Zhou 0001, Tao Lu 0001, Jiang Duan, Guoping Qiu
IEEE Trans. Image Process.1
2020 End-to-End Single Image Fog Removal Using Enhanced Cycle Consistent Adversarial Networks
abstract
Single image defogging is a classical and challenging problem in computer vision. Existing methods towards this problem mainly include handcrafted priors based methods that rely on the use of the atmospheric degradation model and learning-based approaches that require paired fog-fogfree training example images. In practice, however, prior-based methods are prone to failure due to their own limitations and paired training data are extremely difficult to acquire. Moreover, there are few studies on the unpaired trainable defogging network in this field. Thus, inspired by the principle of CycleGAN network, we have developed an end-to-end learning system that uses unpaired fog and fogfree training images, adversarial discriminators and cycle consistency losses to automatically construct a fog removal system. Similar to CycleGAN, our system has two transformation paths; one maps fog images to a fogfree image domain and the other maps fogfree images to a fog image domain. Instead of one stage mapping, our system uses a two stage mapping strategy in each transformation path to enhance the effectiveness of fog removal. Furthermore, we make explicit use of prior knowledge in the networks by embedding the atmospheric degradation principle and a sky prior for mapping fogfree images to the fog images domain. In addition, we also contribute the first real world nature fog-fogfree image dataset for defogging research. Our multiple real fog images dataset (MRFID) contains images of 200 natural outdoor scenes. For each scene, there is one clear image and corresponding four foggy images of different fog densities manually selected from a sequence of images taken by a fixed camera over the course of one year. Qualitative and quantitative comparison against several state-of-the-art methods on both synthetic and real world images demonstrate that our approach is effective and performs favorably for recovering a clear image from a foggy image.
Wei Liu 0123, Xianxu Hou, Jiang Duan, Guoping Qiu
IEEE Trans. Image Process.1
2019 A physics based generative adversarial network for single image defogging
Wei Liu 0123, Rongguo Yao, Guoping Qiu
Image Vis. Comput.1
2009 Design Optimization under Aleatory and Epistemic Uncertainties
abstract
Deterministic design optimization (DO) may lead to unreliable designs due to not taking into consideration uncertainties. In recent years, DO under uncertainty has attracted more and more attention. In DO practice, both aleatory and epistemic uncertainties may exist, furthermore, the two types of uncertainties may affect both the objective function and the constrains of the optimization problem. In this paper, we develop a DO model which could cater for the above-mentioned situation. We also propose a new algorithm, sequential optimization and reliability and possibility assessment (SORPA), to solve the developed DO model. Numerical example is given, and the results obtained demonstrate the feasibility of the developed DO model and the efficiency of the proposed algorithm.
Bo Yang 0011, Wei Liu 0123
DASC3