Jianyu Liu

dblp:192/9939 · DBLP profile ↗
← Back
15ranked-venue papers
6as first author
12since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 3 first-author · 7 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 MOON: Generative MLLM-based Multimodal Representation Learning for E-commerce Product Understanding
abstract
With the rapid advancement of e-commerce, exploring general representations rather than task-specific ones has attracted increasing research attention. For product understanding, although existing discriminative dual-flow architectures drive progress in this field, they inherently struggle to model the many-to-one alignment between multiple images and texts of products. Therefore, we argue that generative Multimodal Large Language Models (MLLMs) hold significant potential for improving product representation learning. Nevertheless, achieving this goal still remains non-trivial due to several key challenges: the lack of multimodal and aspect-aware modeling modules in typical LLMs; the common presence of background noise in product images; and the absence of a standard benchmark for evaluation. To address these issues, we propose the first generative MLLM-based model named MOON for product representation learning. Our method (1) employs a guided Mixture-of-Experts (MoE) module for targeted modeling of multimodal and aspect-specific product content; (2) effectively detects core semantic regions in product images to mitigate the distraction and interference caused by background noise; and (3) introduces the specialized negative sampling strategy to increase the difficulty and diversity of negative samples. In addition, we release a large-scale multimodal benchmark MBE for various product understanding tasks. Experimentally, our model demonstrates competitive zero-shot performance on both our benchmark and the public dataset, showcasing strong generalization across various downstream tasks, including cross-modal retrieval, product classification, and attribute prediction. Furthermore, the case study and visualization illustrate the effectiveness of MOON for product understanding.
Daoze Zhang, Chenghan Fu, Zhanheng Nie, Jianyu Liu, Wanxian Guan, Pengjie Wang 0002, Jian Xu 0015, Bo Zheng 0007
WSDM4
2026 ICCPNet: inter-layer coupling channel pruning network based on elastic scaling scoring for complex traffic scene
Cuijin Li, Jianyu Liu, Zhong Qu
Appl. Intell.2
2026 Fourier-Taylor progressive knowledge distillation for RGB+IR object detection in complex environments
Cuijin Li, Jianyu Liu
Expert Syst. Appl.2
2025 From Superficial to Deep: Integrating External Knowledge for Follow-up Question Generation Using Knowledge Graph and LLM
abstract
In a conversational system, dynamically generating follow-up questions based on context can help users explore information and provide a better user experience. Humans are usually able to ask questions that involve some general life knowledge and demonstrate higher order cognitive skills. However, the questions generated by existing methods are often limited to shallow contextual questions that are uninspiring and have a large gap to the human level. In this paper, we propose a three-stage external knowledge-enhanced follow-up question generation method, which generates questions by identifying contextual topics, constructing a knowledge graph (KG) online, and finally combining these with a large language model to generate the final question. The model generates information-rich and exploratory follow-up questions by introducing external common sense knowledge and performing a knowledge fusion operation. Experiments show that compared to baseline models, our method generates questions that are more informative and closer to human questioning levels while maintaining contextual relevance.
Jianyu Liu, Yi Huang 0017, Junlan Feng, Guilin Qi
COLING1
2025 Metal Artifact Reduction Methods Using Deep Generative Models for Cultural Relics X-ray CT Images
abstract
Computerized tomography (CT) provides non-invasive visualization of internal structural information without losing any detail. It has proven to be very useful in protecting cultural relics. However, metal cultural relics are frequently accompanied by destructive metal artifacts in x-ray CT images, making it impossible for traditional methods to obtain detailed information from the cultural relics. In recent years, deep generative-based models have demonstrated great promise for solving such problems. However, due to the complicated structure and diverse materials of cultural relics, it is difficult to accurately restore the highly heterogeneous details of cultural relics in practical applications. As a result, the primary focus of this study is on the removal of metal items from cultural relics using deep generative models and achieves effective restoration of highly heterogeneous details by combining mask-guided strategies. Specifically, we first collaborated with the Palace Museum to build a cultural relic CT dataset and specifically divided artifacts into two categories: sharp edge smoothing and edge distortion according to their complexity. Second, many deep generative networks include CycleGAN, CSGAN, MUNIT, DeblurGAN, and DRIT were trained. Additionally, segmentation masks were blended to create artifact-free images. Finally, the performance was enhanced via dynamic weight adjustment. The effectiveness has been qualitatively and quantitatively validated on the cultural relic CT dataset. PSNR and SSIM metrics confirm the model’s ability to restore fine details, providing reliable support for future cultural relic protection.
Daxin Peng, Jianqiang Li 0002, Liang Qu, Jianyu Liu, Zhenbin Xie, Junyu Zhao, Qixin Chen, Wenyi Liang
COMPSAC6
2025 DREAM: Disentangling Risks to Enhance Safety Alignment in Multimodal Large Language Models
abstract
Jianyu Liu, Hangyu Guo, Ranjie Duan, Xingyuan Bu, Yancheng He, Shilong Li, Hui Huang, Jiaheng Liu, Yucheng Wang, Chenchen Jing, Xingwei Qu, Xiao Zhang, Pei Wang, Yanan Wu, Jihao Gu, Yangguang Li, Jianke Zhu. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Jianyu Liu, Hangyu Guo, Ranjie Duan, Xingyuan Bu, Yancheng He, Hui Huang 0021, Chenchen Jing, Xingwei Qu, Jihao Gu, Yangguang Li 0001, Jianke Zhu
NAACL (Long Papers)1
2025 Remaining useful life prediction with uncertainty quantification for rotating machinery: A method based on explainable variational deep gaussian process
Shuo Cui, Wan Qiao, Jianyu Liu, Guoxin Wu
Neurocomputing4
2025 Two peak-finding algorithms for two-dimensional unimodal symmetric signals based on mirroring and interpolating
Bin Wan, Jianyu Liu, Zhenfeng Chen
J. Supercomput.5
2024 PRIMO: Progressive Induction for Multi-hop Open Rule Generation
abstract
Open rules refer to the implication from premise atoms to hypothesis atoms, which captures various relationships between instances in the real world. Injecting open rule knowledge into the machine helps to improve the performance of downstream tasks such as dialogue and relation extraction. Existing approaches focus on single-hop open rule generation, ignoring scenarios involving multiple hops, leading to logical inconsistencies between premise and hypothesis atoms, as well as semantic duplication of generated rule atoms. To address these issues, we propose a progressive multi-stage open rule generation method called PRIMO. We introduce ontology information during the rule generation stage to reduce ambiguity and improve rule accuracy. PRIMO constructs a multi-stage structure consisting of generation, extraction, and rank modules to fully leverage the latent knowledge within the language model across multiple dimensions. Furthermore, we employ reinforcement learning from human feedback to further optimize model, enhancing the model’s understanding of commonsense knowledge. Experimental results demonstrate that compared to baseline models, PRIMO significantly enhances rule quality and diversity while reducing the repetition rate of rule atoms.
Jianyu Liu, Guilin Qi
LREC/COLING1
2024 Application of SNNS Model Based On Multi-Dimensional Attention In Drone Radio Frequency Signal Classification
abstract
Spiking Neural Networks (SNNs) are attracting attention due to their energy efficiency and importance in neuromorphic computing. Therefore, we propose an SNN-based method for classifying drone RF signals in complex electromagnetic environments. Specifically, we designed a new SNNs model called Spiking-EfficientNet based on EfficientNetV2 and improved its performance with a multidimensional attention mechanism. Experimental results demonstrate that Spiking-EfficientNet achieved classification accuracy of 99.13% and 96.02% on the ZK RF and DroneDetectV2 datasets. Importantly, Spiking-EfficientNet not only outperforms traditional Artificial Neural Networks (ANNs) in performance, but also exhibits significantly lower energy consumption. The energy consumption is only 20.1% of EfficientNetV2, 2.56% of VGG11, 10.71% of ResNet18, and 61.15% of MobileNetV2. This study demonstrates the significant potential of SNNs in drone RF signal classification and provides a low-power solution.
Zheng Si, Jianyu Liu, Yinhao Zhou
ICASSP3
2024 Difficulty-controllable question generation over knowledge graphs: A counterfactual reasoning approach
Jianyu Liu, Zeyi Miao, Qizhi Min
Inf. Process. Manag.2
2023 A two-stage anomaly detection framework: Towards low omission rate in industrial vision applications
Jianyu Liu, Zhouwang Yang, Yanzhi Song
Adv. Eng. Informatics1
2019 Clinical Knowledge Graph Embedding Representation Bridging the Gap between Electronic Health Records and Prediction Models
abstract
Learning knowledge embedding representation is an increasingly important technology. However, the choice of hyperparameters is seldom justified and usually relies on exhaustive search. Understanding the effect of hyperparameter combinations on embedding quality is crucial to avoid the inefficient process and enhance practicality of embedding representation along subsequent machine learning applications. This work focuses on translational embedding models for multi-relational categorized data in the clinical domain. We trained and evaluated models with different combinations of hyperparameters on two clinical datasets. We contrasted the results by comparing metric distributions and fitting a random forest regression model. Classifiers were trained to assess embedding representation quality. Finally, clustering was tested as a validation protocol. We observed consistent patterns of hyperparameter preference and identified those that achieved better results respectively. However, results show different patterns regarding link prediction, which is taken as strong evidence that traditional evaluation protocol used for open-domain data does not necessarily lead to the best embedding representation for categorized data.
Matthew Wai Heng Chung, Jianyu Liu, Hegler Tissot
ICMLA2
2018 Application of Multi-Source Data on Structural Framework Study in the Western Beishan Orogenic Belt, Northwest China
abstract
The remote sensing data have respective advantages in application of structural extraction and research. The Landsat 8 OLI image can be used to rapidly delineate the large-scale lineaments due to its large swath width; the ASTER data can be used to delineate mid-scale structures shown by the boundaries of different rocks on the basis of more SWIR bands; the high spatial resolution Worldview-2 images can be used for detailed visual interpretation of small-scale structures. In this study, we applied the OLI, ASTER and Worldview-2 data to the research of the structural framework in western Beishan orogenic belt and provide some new evidence for its tectonic evolution. Our results show that the faults in the study area constitute an imbricate fan system. It proved the ability of integrated utilization of OLI, ASTER and Worldview-2 data in the regional structural framework research.
Jianyu Liu, Genhou Wang
IGARSS1
2018 A New Method for Lithological Dscrimination and Mapping by Using Aster Data in Dong Co Area, Northern Tibet
abstract
A new method is proposed to discriminate various lithological units, including dunites, harzburgites, cumulate gabbros, serpentinites and basalts of the Dong Co ophiolites and sedimentary rocks and volcanic rocks, in Dong Co area by using Advanced Space borne Thermal Emission and Reflection Radiometer (ASTER) data. In this paper, firstly, based on the ASTER image spectra of the known lithological units, a subset of bands are selected to enhance the differences between each unit and eliminate interference information by principle component analysis (PCA). Then, the RGB image of PC1, PC2 and PC3 effectively discriminates various rock units and also contains the texture information of sedimentary rocks, which is key evidence for lithological discrimination. At last, a new lithological map of Dong Co area was proposed, which is consistent with the geological map. It is concluded that this method is rapid and cost effective to map ophiolites in exposed areas in Tibet where it is difficult to do conventional fieldwork.
Jianyu Liu, Genhou Wang, Limin Jia 0003
IGARSS1