Liming Xu

dblp:74/8609 · DBLP profile ↗
← Back
44ranked-venue papers
18as first author
39since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 20 · 5 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 7 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 3 first-author · 7 since 2021Computer networks · 3 · 3 first-author · 3 since 2021Software engineering, systems software and programming languages · 3 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 EmoMAS: Emotion-Aware Multi-Agent System for High-Stakes Edge-Deployable Negotiation with Bayesian Orchestration
abstract
Large language models (LLMs) have been widely used for automated negotiation, but their high computational cost and privacy risks limit deployment in privacy-sensitive, on-device settings such as mobile assistants or rescue robots. Small language models (SLMs) offer a viable alternative, yet struggle with the complex emotional dynamics of high-stakes negotiation. We introduce EmoMAS, a Bayesian multi-agent framework that transforms emotional decision-making from reactive to strategic. EmoMAS leverages a Bayesian orchestrator to coordinate three specialized agents: game-theoretic, reinforcement learning, and psychological coherence models. The system fuses their real-time insights to optimize emotional state transitions while continuously updating agent reliability based on negotiation feedback. This mixture-of-agents architecture enables online strategy learning without pre-training. We further introduce four high-stakes, edge-deployable negotiation benchmarks across debt, healthcare, emergency response, and educational domains. Through extensive agent-to-agent simulations across all benchmarks, both SLMs and LLMs equipped with EmoMAS consistently surpass all baseline models in negotiation performance while balancing ethical behavior. These results show that strategic emotional intelligence is a key driver of negotiation success. By treating emotional expression as a strategic variable within a Bayesian multi-agent optimization framework, EmoMAS establishes a new paradigm for effective, private, and adaptive negotiation AI suitable for high-stakes edge deployment. The code is available at https://github.com/Yunbo-max/EmoMAS.
Yunbo Long, Liming Xu
ACL (1)3
2026 ShiftingNet : Lightweight Crop Leaf Disease Classification Model With Channel-Wise Feature Shifting
abstract
ABSTRACT This paper proposes ShiftingNet, a lightweight classification model based on the improved EfficientNetV2 architecture, which embeds the Channel‐wise Feature Shifting (CFS) operation. Designed as an efficient diagnostic component for agricultural expert systems, ShiftingNet aims to automate classification performance for diverse crop leaf diseases under challenging agricultural conditions. To address the limitations in balancing global and local feature representations and cross‐environment generalisation, we design two novel modules: Channel‐wise Feature Shifting Convolution (CFSConv) and Fused Channel‐wise Feature Shifting Convolution (Fused‐CFSConv). These modules integrate the CFS residual connection and DropPath regularisation into the original MBConv and Fused‐MBConv, while introducing the Squeeze‐and‐Excitation (SE) and Coordinate Attention (CA) mechanisms, respectively. We construct a multi‐source corn dataset, MixCorn, which fuses PlantVillage laboratory images, PlantDoc network images, and CD&S field samples, covering different illumination conditions, backgrounds, and disease scales. Experiments show that ShiftingNet achieves classification accuracies of 99.84% and 99.08% on the PlantVillage and MixCorn datasets, respectively, with only 9.92 M parameters. This demonstrates advantages in knowledge acquisition efficiency and computational cost, providing a theoretical foundation for constructing resource‐constrained mobile expert systems. Robustness evaluation under common perturbations and Grad‐CAM‐based interpretability analysis further validate the model's reliability for automated decision‐making. Ablation studies further confirm that the CFS operation improves feature representation and classification performance.
Dongen Guo, Linbo Han, Penghua Yan, Ziqi Jia, Liming Xu
Expert Syst. J. Knowl. Eng.6
2026 Efficient and privacy-preserved link prediction via condensed graphs
abstract
Link prediction plays a vital role in uncovering hidden relationships within complex networks, enabling applications such as identifying potential customers and products. However, this task faces critical challenges, including growing concerns over data privacy and the substantial computational and storage costs associated with large-scale networks. Condensed graphs, which are significantly smaller yet retain essential structural information, have emerged as a promising solution for preserving data utility while enhancing privacy. Despite this potential, existing methods like HyDRO rely on random node selection strategies designed primarily for node classification, overlooking connectivity patterns crucial for effective link prediction. Moreover, these methods lack rigorous evaluations of privacy risks associated with nodes and links in the condensed graphs—an essential consideration for sensitive real-world networks, where privacy concerns far exceed those of public datasets such as citation graphs. To address these limitations, we introduce HyDRO + , a novel graph condensation method guided by algebraic Jaccard similarity. By leveraging local connectivity patterns, HyDRO + generates structurally-aware condensed graphs that preserve link information more effectively. We further introduce a comprehensive evaluation framework that rigorously assesses both node- and link-level privacy leakage. Extensive experiments on four real-world networks demonstrate that HyDRO + consistently outperforms state-of-the-art methods—and even the original networks—in balancing link prediction accuracy and privacy preservation. Notably, it achieves nearly 20 × faster training and reduces storage requirements by a factor of 452 on the Computers dataset. This work represents the first attempt to use condensed graphs for privacy-preserving link prediction in real-world complex networks, offering a practical and scalable solution for secure information sharing in large-scale networks.
Yunbo Long, Liming Xu, Alexandra Brintrup
Expert Syst. Appl.2
2026 DENAS: Differential evolution neural architecture search for prediction of diabetic retinopathy
Hangjiang Liu, Liming Xu, Jie Shao 0001, Weisheng Li 0001
Neurocomputing3
2026 Multi-scale dense network with pyramid pooling for glaucoma screening
Yingxiu Jin, Liming Xu
J. Vis. Commun. Image Represent.5
2026 SODAS: Second-order optimization differential architecture search for diabetic retinopathy prediction
Liming Xu, Jie Shao 0001, Jiancheng Lv 0001, Weisheng Li 0001
Neural Networks2
2026 C-GAN: Medical Image Steganography Based on Convergent GANs With Localization
abstract
Image steganography aims to hide secret message into cover image in an imperceptible and undetectable way, and only allows the informed receivers to decode stego image. Generative adversarial nets have been proved to be promising against other generative models, and some recently propose to use GANs to hide secret message to reach image steganography. However, it is still facing low embedding capacity, high detectability and poor convergence. To hand the pitfalls, we propose a novel medical image steganography method with convergent divergence measurement to achieve large capacity and undetectable hiding. Specifically, generator, extractor and discriminator are jointed into end-to-end framework where generator yields visually and detectably indistinguishable steganography image from which extractor recovers diagnose report while discriminator tries to distinguish the steganographic and original images. Then, we design Zero-centered Wasserstein distance to achieve controllable and stable training. It can be proved that the proposed method with the defined Zero-centered Wasserstein can converge to a local equilibrium with finite discriminator updates per generator updates. Besides, local regularization which can be also proved to be effective for accelerating convergence is imposed on generator to improve embedding capacity and achieve homeomorphism manifold mapping in low-dimension latent space. Extensive experiments on Open-I, LGK and COV-CTR medical dataset show that the proposed method outperforms recent state-of-the-art methods in capacity, detectability and convergence rate.
Liming Xu, Bochuan Zheng, Weisheng Li 0001
IEEE Trans. Dependable Secur. Comput.1
2025 Multi-modal semantic feature alignment medical cross-modal hashing
Qinghai Liu, Qianlin Wu, Lun Tang, Liming Xu, Qianbin Chen
Eng. Appl. Artif. Intell.4
2025 State estimation of Lithium-ion Batteries with state space model
Zihao Lv, Yu Xue 0003, Liming Xu
Eng. Appl. Artif. Intell.6
2025 Modal disentangled generative adversarial networks for bidirectional magnetic resonance image synthesis
Liming Xu, Yanrong Lei
Eng. Appl. Artif. Intell.1
2025 A survey on deep learning-based algorithms for the traveling salesman problem
abstract
Abstract This paper presents an overview of deep learning (DL)-based algorithms designed for solving the traveling salesman problem (TSP), categorizing them into four categories: end-to-end construction algorithms, end-to-end improvement algorithms, direct hybrid algorithms, and large language model (LLM)-based hybrid algorithms. We introduce the principles and methodologies of these algorithms, outlining their strengths and limitations through experimental comparisons. End-to-end construction algorithms employ neural networks to generate solutions from scratch, demonstrating rapid solving speed but often yielding subpar solutions. Conversely, end-to-end improvement algorithms iteratively refine initial solutions, achieving higher-quality outcomes but necessitating longer computation times. Direct hybrid algorithms directly integrate deep learning with heuristic algorithms, showcasing robust solving performance and generalization capability. LLM-based hybrid algorithms leverage LLMs to autonomously generate and refine heuristics, showing promising performance despite being in early developmental stages. In the future, further integration of deep learning techniques, particularly LLMs, with heuristic algorithms and advancements in interpretability and generalization will be pivotal trends in TSP algorithm design. These endeavors aim to tackle larger and more complex real-world instances while enhancing algorithm reliability and practicality. This paper offers insights into the evolving landscape of DL-based TSP solving algorithms and provides a perspective for future research directions.
Jingyan Sui, Shizhe Ding, Xulin Huang, Boyang Xia, Zhenxin Ding, Liming Xu, Haicang Zhang, Chungong Yu, Dongbo Bu
Frontiers Comput. Sci.8
2025 Graph convolutional lifelong medical cross-modal hashing
Dengping Zhao, Liming Xu
Neurocomputing3
2025 Dual-modality visual feature flow for medical report generation
abstract
Medical report generation, a cross-modal task of generating medical text information, aiming to provide professional descriptions of medical images in clinical language. Despite some methods have made progress, there are still some limitations, including insufficient focus on lesion areas, omission of internal edge features, and difficulty in aligning cross-modal data. To address these issues, we propose Dual-Modality Visual Feature Flow (DMVF) for medical report generation. Firstly, we introduce region-level features based on grid-level features to enhance the method's ability to identify lesions and key areas. Then, we enhance two types of feature flows based on their attributes to prevent the loss of key information, respectively. Finally, we align visual mappings from different visual feature with report textual embeddings through a feature fusion module to perform cross-modal learning. Extensive experiments conducted on four benchmark datasets demonstrate that our approach outperforms the state-of-the-art methods in both natural language generation and clinical efficacy metrics.
Quan Tang 0006, Liming Xu, Yongheng Wang, Bochuan Zheng, Jiancheng Lv 0001, Weisheng Li 0001
Medical Image Anal.2
2025 Pre- to post-contrast medical image synthesis with outline-guide accelerate diffusion model
Xueying Fan, Liming Xu, Bochuan Zheng
Neural Networks2
2025 Evolutionary architecture search for generative adversarial networks using an aging mechanism-based strategy
Wenxing Man, Liming Xu
Neural Networks2
2025 Deep Disease Label-guided Graph Convolutional Network for Medical Report Generation
abstract
Medical report generation which extracts pathological information within medical images and subsequently produces diagnostic text autonomously aims to alleviate the workload of medical experts and offers auxiliary support in diagnoses. Despite some preliminary progress have been made, several limitations still persist, including lack of specificity in extracted visual features, insufficient consideration of cross-modal alignment and extensive preparatory work required for prior knowledge. To address these issues, we, in this article, propose a novel deep label-guided graph convolutional network for medical report generation which utilizes disease label to guide to extract pathological information from medical images. To be specific, we first construct graph convolutional network to guide the model to extract the specific visual features based on disease labels, which allowing us to selectively extract disease specificity information resided in medical images. Then, we develop cross-modal alignment module to guide the alignment across medical image, diagnose report and disease label, which enables more accurate generation with more precise description. Besides, we build pre-constructed relational matrix to guide report generation model to learn the relationship between visual features and disease types with minimal additional workload to further reduce intensive workload. Extensive experiments on three benchmark datasets, i.e., IU X-ray, MIMIC-CXR, and COV-CTR, demonstrate that the proposed method outperforms the recent state-of-the-art medical report generation methods. Ours shows a 9.2% improvement in BLEU-4 score on the IU X-ray dataset, and both BLEU-4 and CIDEr scores improve by 6.31% on the MIMIC-CXR dataset. Additionally, the results show that it can be easily to applied and extended to medical image report generation with different modalities.
Liming Xu, Yongheng Wang, Quan Tang 0006, Jiancheng Lv 0001
ACM Trans. Knowl. Discov. Data1
2025 Deep Rib Fracture Instance Segmentation and Classification From CT on the RibFrac Challenge
abstract
Rib fractures are a common and potentially severe injury that can be challenging and labor-intensive to detect in CT scans. While there have been efforts to address this field, the lack of large-scale annotated datasets and evaluation benchmarks has hindered the development and validation of deep learning algorithms. To address this issue, the RibFrac Challenge was introduced, providing a benchmark dataset of over 5,000 rib fractures from 660 CT scans, with voxel-level instance mask annotations and diagnosis labels for four clinical categories (buckle, nondisplaced, displaced, or segmental). The challenge includes two tracks: a detection (instance segmentation) track evaluated by an FROC-style metric and a classification track evaluated by an F1-style metric. During the MICCAI 2020 challenge period, 243 results were evaluated, and seven teams were invited to participate in the challenge summary. The analysis revealed that several top rib fracture detection solutions achieved performance comparable or even better than human experts. Nevertheless, the current rib fracture classification solutions are hardly clinically applicable, which can be an interesting area in the future. As an active benchmark and research resource, the data and online evaluation of the RibFrac Challenge are available at the challenge website (https://ribfrac.grand-challenge.org/). In addition, we further analyzed the impact of two post-challenge advancements-large-scale pretraining and rib segmentation-based on our internal baseline for rib fracture detection. These findings lay a foundation for future research and development in AI-assisted rib fracture diagnosis.
Jiancheng Yang, Kaiming Kuang, Donglai Wei 0001, Shixuan Gu, Jianying Liu, Zhizhong Chai, Yongjie Xiao, Hao Chen 0011, Liming Xu, Bang Du, Xiangyi Yan, Hao Tang 0010, Adam M. Alessio, Gregory Holste, Jianye He, Lixuan Che, Hanspeter Pfister, Ming Li 0005, Bingbing Ni
IEEE Trans. Medical Imaging13
2025 Multi-scale Consistency Deep Lifelong Cross-modal Hashing
abstract
Deep cross-modal hashing methods provide effective and efficient solutions for large-scale cross-modal retrieval. However, existing cross-modal hashing methods fail to capture the dynamic changes of real-world data, and suffer from serious performance degradation when retrieving streaming data. In this paper, we propose a novel hashing method to achieve accurate cross-modal retrieval under continuous and streaming scenarios. Specifically, regularization-based lifelong learning module is introduced to balance plasticity for learning new knowledge and stability for maintaining old knowledge, and update incremental hash codes without retraining cumulative data. Then, multi-scale consistency network which employs multi-scale feature fusion module to extract fine-grained features among multi-scale modalities is introduced to learn multi-level semantic representations with consistency. Additionally, modality alignment with variational information bottleneck is designed to remove irrelevant information and obtain unified representation, which can be proved to be effective to yield high-quality hash code with new and old knowledge. Extensive experiments show that ours gains the advanced performance and the better adaptability to continuous and streaming environments.
Liming Xu, Jie Shao 0001, Weisheng Li 0001
ACM Trans. Multim. Comput. Commun. Appl.1
2025 Deep Differential Lifelong Cross-modal Hashing for Stream Medical Data Retrieval
abstract
With the explosive growth of stream medical multi-modal data, it is significant to develop an efficient cross-modal retrieval algorithm to achieve effective medical data search. Within it, deep cross-modal hashing which maps cross-modal data into low-dimensional Hamming space where similarity in high-dimension space is preserved has made much progress. However, most of deep cross-modal hashing algorithms are usually facing disability of adapting to dynamic stream medical data, non-differentiable optimization, and unaligned semantic across modalities. To address these, we, in this article, propose a novel deep differential lifelong cross-modal hashing method for large-scale stream medical data retrieval. Specifically, we first design lifelong learning module to keep the learned hash code of base data unchanged and directly learn hash code of incremental data with new categories to achieve continuous retrieval of stream medical data, which effectively mitigates catastrophic forgetting, as well as significantly reduces training time and computation resource. Then, we introduce differential cross-modal hashing module to generate discriminative binary hash codes, which yields continuous and differentiable optimization and improves accuracy. Besides, we design semantic alignment module which embeds intra-modal and inter-modal losses to maintain the semantic similarity and dis-similarity among stream medical data across modalities. Extensive experiments on benchmark medical datasets show that our proposed method can retrieve dynamic stream medical cross-modal data effectively and obtain higher retrieval performance comparing with recent state-of-the-art approaches.
Liming Xu, Dengping Zhao, Bochuan Zheng
ACM Trans. Multim. Comput. Commun. Appl.1
2024 Uncovering the Metaverse within Everyday Environments: A Coarse-to-Fine ApproachBehaviors
abstract
The recent release of the Apple Vision Pro has reignited interest in the metaverse, showcasing the intensified efforts of technology giants in developing platforms and devices to facilitate its growth. As the metaverse continues to proliferate, it is foreseeable that everyday environments will become increasingly saturated with its presence. Consequently, uncovering links to these metaverse items will be a crucial first step to interacting with this new augmented world. In this paper, we address the problem of establishing connections with virtual worlds within everyday environments, especially those that are not readily discernible through direct visual inspection. We introduce a vision-based approach leveraging Artcode visual markers to uncover hidden metaverse links embedded in our ambient surroundings. This approach progressively localises the access points to the metaverse, transitioning from coarse to fine localisation, thus facilitating an exploratory interaction process. Detailed experiments are conducted to study the performance of the proposed approach, demonstrating its effectiveness in Artcode localisation and enabling new interaction opportunities.
Liming Xu, Dave Towey, Andrew P. French, Steve Benford
COMPSAC1
2024 DIFNet: Dual-Domain Information Fusion Network for Image Denoising
Zedong Wu, Wenxu Shi, Liming Xu, Zicheng Ding, Bochuan Zheng
PRCV (8)3
2024 Deep Lifelong Cross-Modal Hashing
abstract
Hashing methods have made significant progress in cross-modal retrieval tasks with fast query speed and low storage cost. Among them, deep learning-based hashing achieves better performance on large-scale data due to its excellent extraction and representation ability for nonlinear heterogeneous features. However, there are still two main challenges in catastrophic forgetting when data with new categories arrive continuously, and time-consuming for non-continuous hashing retrieval to retrain for updating. To this end, we, in this paper, propose a novel deep lifelong cross-modal hashing to achieve lifelong hashing retrieval instead of re-training hash function repeatedly when new data arrive. Specifically, we design lifelong learning strategy to update hash functions by directly training the incremental data instead of retraining new hash functions using all the accumulated data, which significantly reduce training time. Then, we propose lifelong hashing loss to enable original hash codes participate in lifelong learning but remain invariant, and further preserve the similarity and dis-similarity among original and incremental hash codes to maintain performance. Additionally, considering distribution heterogeneity when new data arriving continuously, we introduce enhanced-semantic similarity to supervise hash learning, and it has been proven that the similarity improves performance with detailed analysis. Experimental results on benchmark datasets show that our proposed method achieves comparative performance comparing with recent state-of-the-art cross-modal hashing methods, and it yields substantial average increments over 20% in retrieval accuracy and almost reduces over 80% training time when new data arrives continuously.
Liming Xu, Bochuan Zheng, Weisheng Li 0001, Jiancheng Lv 0001
IEEE Trans. Circuits Syst. Video Technol.1
2024 CGFTrans: Cross-Modal Global Feature Fusion Transformer for Medical Report Generation
abstract
Medical report generation, as a cross-modal automatic text generation task, can be highly significant both in research and clinical fields. The core is to generate diagnosis reports in clinical language from medical images. However, several limitations persist, including a lack of global information, inadequate cross-modal fusion capabilities, and high computational demands. To address these issues, we propose cross-modal global feature fusion Transformer (CGFTrans) to extract global information meanwhile reduce computational strain. Firstly, we introduce mesh recurrent network to capture inter-layer information at different levels to address the absence of global features. Then, we design feature fusion decoder and define 'mid-fusion' strategy to separately fuse visual and global features with medical report embeddings, which enhances the ability of the cross-modal joint learning. Finally, we integrate shifted window attention into Transformer encoder to alleviate computational pressure and capture pathological information at multiple scales. Extensive experiments conducted on three datasets demonstrate that the proposed method achieves average increments of 2.9%, 1.5%, and 0.7% in terms of the BLEU-1, METEOR and ROUGE-L metrics, respectively. Besides, it achieves average increments -22.4% and 17.3% training time and images throughput, respectively.
Liming Xu, Quan Tang 0006, Bochuan Zheng, Jiancheng Lv 0001, Weisheng Li 0001
IEEE J. Biomed. Health Informatics1
2023 A Study of Chinese Medicine Entity Recognition Method by Fusing Multi-Features and Pointer Networks
abstract
The recognition of named entities in Traditional Chinese medicine (TCM) is a difficult task in the field of medical information extraction, which often contains a large number of domain nouns and specialized terms with high semantic complexity and unclear entity boundaries and multiple meanings among some entities. In order to effectively solve the recognition problem of named entities in TCM, and to address the phenomenon of underutilized semantic information in entity recognition tasks, an entity recognition method incorporating Chinese character multi-features and SPAN pointer networks is proposed to obtain character feature vectors of data using the powerful characterization information of the pre-training model BERT, connect the character vectors with lexical and radical feature embeddings, and obtain the long-range textual context through BiGRU and Attention layer to obtain the contextual information of long-range text, and finally use SPAN pointer network to achieve the start boundary determination of the entity and complete the extraction of the entity. In addition, adversarial training and focal loss are added to reduce the risk of overfitting and enhance the generalization ability and robustness of the model. The final experiments show that this method has superior performance in dealing with the named entity recognition problem of TCM.
Zihao Lv, Liming Xu
SMC3
2023 Deep image captioning: A review of methods, trends and future challenges
Liming Xu, Quan Tang 0006, Jiancheng Lv 0001, Bochuan Zheng, Weisheng Li 0001
Neurocomputing1
2023 BH2I-GAN: Bidirectional Hash_code-to-Image Translation using Multi-Generative Multi-Adversarial Nets
Liming Xu, Weisheng Li 0001, Yicai Xie
Pattern Recognit.1
2023 A Robust Shape-Aware Rib Fracture Detection and Segmentation Framework With Contrastive Learning
abstract
The rib fracture is a common type of thoracic skeletal trauma, and its inspections using computed tomography (CT) scans are critical for clinical evaluation and treatment planning. However, it is often challenging for radiologists to quickly and accurately detect rib fractures due to tiny objects and blurriness in large 3D CT images. Previous diagnoses for automatic rib fracture mostly relied on deep learning (DL)-based object detection, which highly depends on label quality and quantity. Moreover, general object detection methods did not take into consideration the typically elongated and oblique shapes of ribs in 3D volumes. To address these issues, we propose a shape-aware method based on DL called SA-FracNet for rib fracture detection and segmentation. First, we design a pixel-level pretext task founded on contrastive learning on massive unlabeled CT images. Second, we train the fine-tuned rib fracture detection model based on the pre-trained weights. Third, we develop a fracture shape-aware multi-task segmentation network to delineate the fracture based on the detection result. Experiments demonstrate that our proposed SA-FracNet achieves state-of-the-art rib fracture detection and segmentation performance on the public RibFrac dataset, with a detection sensitivity of 0.926 and segmentation Dice of 0.754. Test on a private dataset also validates the robustness and generalization of our SA-FracNet.
Zheng Cao 0005, Liming Xu, Danny Ziyi Chen, Honghao Gao, Jian Wu 0001
IEEE Trans. Multim.2
2023 MFGAN: Multi-modal Feature-fusion for CT Metal Artifact Reduction Using GANs
abstract
Due to the existence of metallic implants in certain patients, the Computed Tomography (CT) images from these patients are often corrupted by undesirable metal artifacts, which causes severe problem of metal artifact. Although many methods have been proposed to reduce metal artifact, reduction is still challenging and inadequate. Some reduced results are suffering from symptom variance, second artifact, and poor subjective evaluation. To address these, we propose a novel method based on generative adversarial nets (GANs) to reduce metal artifacts. Specifically, we firstly encode interactive information (text) and imaging CT (image) to yield multi-modal feature-fusion representation, which overcomes representative ability limitation of single-modal CT images. The incorporation of interaction information constrains feature generation, which ensures symptom consistency between corrected and target CT. Then, we design an enhancement network to avoid second artifact and enhance edge as well as suppress noise. Besides, three radiology physicians are invited to evaluate the corrected CT image. Experiments show that our method gains significant improvement over other methods. Objectively, ours achieves an average increment of 7.44% PSNR and 6.12% SSIM on two medical image datasets. Subjectively, ours outperforms others in comparison in term of sharpness, resolution, invariance, and acceptability.
Liming Xu, Weisheng Li 0001, Bochuan Zheng
ACM Trans. Multim. Comput. Commun. Appl.1
2022 Connecting Everyday Objects with the Metaverse: A Unified Recognition Framework
abstract
The recent Facebook rebranding to Meta has drawn renewed attention to the metaverse. Technology giants, amongst others, are increasingly embracing the vision and opportunities of a hybrid social experience that mixes physical and virtual interactions. As the metaverse gains in traction, it is expected that everyday objects may soon connect more closely with virtual elements. However, discovering this “hidden” virtual world will be a crucial first step to interacting with it in this new augmented world. In this paper, we address the problem of connecting phys-ical objects with their virtual counterparts, especially through connections built upon visual markers. We propose a unified recognition framework that guides approaches to the metaverse access points. We illustrate the use of the framework through experimental studies under different conditions, in which an interactive and visually attractive decoration pattern, an Artcode, is used as the approach to enable the connection. This paper will be of interest to, amongst others, researchers working in Interaction Design or Augmented Reality who are seeking techniques or guidelines for augmenting physical objects in an unobtrusive, complementary manner.
Liming Xu, Dave Towey, Andrew P. French, Steve Benford
COMPSAC1
2022 Multiple deep neural networks with multiple labels for cross-modal hashing retrieval
Yicai Xie, Tinghua Wang, Liming Xu, Dingjie Wang
Eng. Appl. Artif. Intell.4
2022 CP-GAN: Meet the high requirements of diagnose report to medical image by content preservation
abstract
Abstract Medical image generation from diagnostic report has important research significance for medical‐aided diagnosis. This research can improve the diagnosis speed and accuracy of doctors and effectively save the storage resources of hospitals. The research has made significant progress in the field of natural images, but rarely used in medical images. Medical images have higher requirements for image quality. In this paper, a method based on attention mechanism and content preservation loss to improve image quality is proposed. The model consists of three stages, and each stage generates feature map of different scale combined with attention features as the input of the next stage to optimise semantic consistency between image and text. The last stage will generate an image, whose size is 256 × 256. And the content preservation loss can optimise the similarity between the generated image and the real image by low‐level and high‐level features to meet the high requirements of medical image for texture details. The content preservation loss consists of MSE loss, VGG loss and TV loss. The experiments on two datasets prove that the method can achieve excellent results.
Zhengyi Huang, Liming Xu, Yicai Xie
IET Image Process.3
2022 Deep online cross-modal hashing by a co-training mechanism
Yicai Xie, Tinghua Wang, Yun Yi, Liming Xu
Knowl. Based Syst.5
2022 Multi-Manifold Deep Discriminative Cross-Modal Hashing for Medical Image Retrieval
abstract
Benefitting from the low storage cost and high retrieval efficiency, hash learning has become a widely used retrieval technology to approximate nearest neighbors. Within it, the cross-modal medical hashing has attracted an increasing attention in facilitating efficiently clinical decision. However, there are still two main challenges in weak multi-manifold structure perseveration across multiple modalities and weak discriminability of hash code. Specifically, existing cross-modal hashing methods focus on pairwise relations within two modalities, and ignore underlying multi-manifold structures across over 2 modalities. Then, there is little consideration about discriminability, i.e., any pair of hash codes should be different. In this paper, we propose a novel hashing method named multi-manifold deep discriminative cross-modal hashing (MDDCH) for large-scale medical image retrieval. The key point is multi-modal manifold similarity which integrates multiple sub-manifolds defined on heterogeneous data to preserve correlation among instances, and it can be measured by three-step connection on corresponding hetero-manifold. Then, we propose discriminative item to make each hash code encoded by hash functions be different, which improves discriminative performance of hash code. Besides, we introduce Gaussian-binary Restricted Boltzmann Machine to directly output hash codes without using any continuous relaxation. Experiments on three benchmark datasets (AIBL, Brain and SPLP) show that our proposed MDDCH achieves comparative performance to recent state-of-the-art hashing methods. Additionally, diagnostic evaluation from professional physicians shows that all the retrieved medical images describe the same object and illness as the queried image.
Liming Xu, Bochuan Zheng, Weisheng Li 0001
IEEE Trans. Image Process.1
2022 IDHashGAN: Deep Hashing With Generative Adversarial Nets for Incomplete Data Retrieval
abstract
Benefiting from low storage costs and high retrieval efficiency, hash learning has been a widely adopted technology for approximating nearest neighbor in large-scale data retrieval. Deep learning to hash greatly improves image retrieval performance by integrating feature learning and hash coding into an end-to-end framework. However, subject to application scope, most existing deep hashing methods only apply to retrieval of complete data and have undesirable results when retrieving incomplete but valuable data. In this paper we propose IDHashGAN, a novel deep hashing model with generative adversarial networks to retrieve incomplete data, in which feature restoration, feature learning and hash coding are integrated into an unified end-to-end framework. The proposed model consists of four key components: (1) reconstructive and generative loss are used to generate continuous feature of incomplete data in generative network; (2) supervised manifold similarity is proposed to improve retrieval accuracy and obtain good user acceptance; (3) adversarial and classified loss are designed to distinguish authenticity and similarity in discriminative network; and (4) encoding and quantization loss are adopted to preserve similarity and control hash quality. Extensive experiments on benchmark datasets show that IDHashGAN is competitive on complete dataset and yields substantial boosts of 70% on incomplete datasets compared to state-of-the-art hashing methods.
Liming Xu, Weisheng Li 0001, Ling Bai
IEEE Trans. Multim.1
2021 Learning 3-opt heuristics for traveling salesman problem via deep reinforcement learning
abstract
Traveling salesman problem (TSP) is a classical combinatorial optimization problem. As it represents a large number of important practical problems, it has received extensive studies and a great variety of algorithms have been proposed to solve it, including exact and heuristic algorithms. The success of heuristic algorithms relies heavily on the design of powerful heuristic rules, and most of the existing heuristic rules were manually designed by experienced experts to model their insights and observations on TSP instances and solutions. Recent studies have shown an alternative promising design strategy that directly learns heuristic rules from TSP instances without any manual interference. Here, we report an iterative improvement approach (called Neural-3-OPT) that solves TSP through automatically learning effective 3-opt heuristics via deep reinforcement learning. In the proposed approach, we adopt a pointer network to select 3 links from the current tour,and a feature-wise linear modulation network to select an appropriate way to reconnect the segments after removing the selected 3 links. We demonstrate that our approach achieves state-of-the-art performance on both real TSP instances and randomly-generated instances than, to the best of our knowledge, the existing neural network-based approaches.
Jingyan Sui, Shizhe Ding, Liming Xu, Dongbo Bu
ACML4
2021 Cross-modal variable-length hashing based on hierarchy
abstract
Due to the emergence of the era of big data, cross-modal learning have been applied to many research fields. As an efficient retrieval method, hash learning is widely used frequently in many cross-modal retrieval scenarios. However, most of existing hashing methods use fixed-length hash codes, which increase the computational costs for large-size datasets. Furthermore, learning hash functions is an NP hard problem. To address these problems, we initially propose a novel method named Cross-modal Variable-length Hashing Based on Hierarchy (CVHH), which can learn the hash functions more accurately to improve retrieval performance, and also reduce the computational costs and training time. The main contributions of CVHH are: (1) We propose a variable-length hashing algorithm to improve the algorithm performance; (2) We apply the hierarchical architecture to effectively reduce the computational costs and training time. To validate the effectiveness of CVHH, our extensive experimental results show the superior performance compared with recent state-of-the-art cross-modal methods on three benchmark datasets, WIKI, NUS-WIDE and MIRFlickr.
Xiaojun Qi 0002, Yicai Xie, Liming Xu
Intell. Data Anal.5
2021 Remote sensing image super-resolution using cascade generative adversarial nets
Dongen Guo, Liming Xu, Weisheng Li 0001, Xiaobo Luo
Neurocomputing3
2021 DSAGAN: A generative adversarial network based on dual-stream attention mechanism for anatomical and functional image fusion
Jun Fu 0004, Weisheng Li 0001, Jiao Du, Liming Xu
Inf. Sci.4
2021 Using metamorphic relations to verify and enhance Artcode classification
Liming Xu, Dave Towey, Andrew P. French, Steve Benford, Zhiquan Zhou 0001, Tsong Yueh Chen
J. Syst. Softw.1
2020 Multi-granularity generative adversarial nets with reconstructive sampling for image inpainting
Liming Xu, Weisheng Li 0001
Neurocomputing1
2020 BPGAN: Bidirectional CT-to-MRI prediction using multi-generative multi-adversarial nets with spectral normalization and localization
Liming Xu, Weisheng Li 0001, Jianbo Lei
Neural Networks1
2018 Recognition method for apple fruit based on SUSAN and PCNN
Liming Xu, Jidong Lv
Multim. Tools Appl.1
2016 Accountable Artefacts: The Case of the Carolan Guitar
abstract
We explore how physical artefacts can be connected to digital records of where they have been, who they have encountered and what has happened to them, and how this can enhance their meaning and utility. We describe how a travelling technology probe in the form of an augmented acoustic guitar engaged users in a design conversation as it visited homes, studios, gigs, workshops and lessons, and how this revealed the diversity and utility of its digital record. We describe how this record was captured and flexibly mapped to the physical guitar and proxy artefacts. We contribute a conceptual framework for accountable artefacts that articulates how multiple and complex mappings between physical artefacts and their digital records may be created, appropriated, shared and interrogated to deliver accounts of provenance and use as well as methodological reflections on technology probes.
Steve Benford, Adrian Hazzard, Alan Chamberlain, Kevin Glover, Christopher Greenhalgh, Liming Xu, Michaela Hoare, Dimitrios Paris Darzentas
CHI6
2015 A land cover adaptive topographic correction and evaluation method for remote sensing data
abstract
Most of the empirical topographic correction methods are based on the universal assumptions of the relationship between radiations and solar incident angles. The correction accuracy is hardly to be accessed quantitatively. This paper introduces a land cover adaptive C (LCAC) method for topographic correction, and verifies its advantage quantitatively. Experiments on synthetic and real remote sensing data are performed. The synthetic data derives from the SMARTS2 model, which simulates the surface reflectance and the atmospheric conditions. The global land cover map produced by the National Geomatics Center of China is taken as the auxiliary data. The LCAC outperforms the traditional C method in experiments on both synthetic and real remote sensing data by visual and quantitative assessments.
Huifang Li 0001, Liming Xu, Huanfeng Shen, Wei Li 0318, Liqin Cao
IGARSS2