Ming Xia 0005

dblp:54/344-5 · DBLP profile ↗
← Back
18ranked-venue papers
7as first author
15since 2021 · last 2026
0000-0002-8094-2203ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 8 · 6 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Enhancing the Interpretation of Skin Lesion Diagnosis: Concept Adaptive Fine-Tuning of Vision-Language Models
abstract
Significant progress has been made in applying deep learning for the automatic diagnosis of skin lesions. However, most models remain unexplainable, which severely hinders their application in clinical settings. Concept-based ante-hoc interpretable models have the potential to clarify the decision-making process of diagnosis by learning high-level, human-understandable concepts, while they can only provide numerical values of conceptual contributions. Pre-trained Vision-Language Models (VLMs) can learn rich vision-language correlations from large-scale image-text pairs. Fine-tuning pre-trained VLMs for specific downstream tasks is an effective way to reduce data requirements. Nevertheless, when there is a substantial disparity between the pre-trained model and the target task, existing tuning methods frequently struggle to generalize, necessitating substantial training data to fully adapt VLMs to specialized medical tasks. In this work, we propose a concept adaptive fine-tuning (CptAFT) method based on the pre-trained VLM, BiomedCLIP, to develop a concept-based multi-modal interpretable skin lesion diagnosis model. By incorporating medical texts, such as reports and conceptual terms, our model can recognize fine-grained features and provide robust, natural language-driven interpretability. Moreover, our concept-adaptive method that reconstructs images using concept logits and imposes a consistency loss with the original image, enabling the VLM to quickly adapt to the task with a small amount of training data. Extensive experimental results demonstrate that our approach outperforms state-of-the-art closed box and interpretable models in both classification performance and medically relevant interpretability. In particular, after fine-tuning with a small amount of data, our model outperforms MONET, a model trained on the large Skin Disease Image-Report dataset, by 8.28% in concept recognition ability, demonstrating the interpretability of our model.
Yating Zhu, Xiaoyan Wang 0007, Ming Xia 0005, Pan Mu, Haigen Hu, Xiaoqin Zhang 0002
IEEE J. Biomed. Health Informatics4
2026 Toward Ubiquitous Vision-Augmented Intelligent Communications: A Generalized Loosely Coupled Approach
abstract
Existing studies have shown that merging sensory data, especially visual signals, into Radio Frequency (RF) networks can enhance their perception and adaptability to hidden disruptive environmental factors (e.g., blockages). A majority of these methods assume tight integration of the RF and non-RF components within a wireless system, making them rely heavily on prior network deployment knowledge and applicable to only specific deployments. Aiming at ubiquitous physically assisted intelligent communications, this article presents a new LOOSELY COUPLED approach that incorporates the videos from widely installed external surveillance cameras to aid the operation of existing networks. To enable adaptive integration across individual sensing and communication systems in various unknown deployments, we devise a novelobject-centric hierarchical learningscheme that can combine decoupled visual and RF signals to autonomously discover the compositional structure of the blockage scene and analyze its spatiotemporal development for inferring the future blockage condition changes. We demonstrate the application of these structural representations in building a generalized Vision-Perceptive Link Quality Predictor (VP-LQP). To support the training and testing of VP-LQP, we construct a new real-world multi-scenario dataset, and the experimental results validate the superiority of our design.
Ming Xia 0005, Ziyang Lin, Jiaquan Jin, Yu Hen Hu, Zhen Cheng 0001, Kaikai Chi
IEEE Trans. Wirel. Commun.1
2025 LGCE-Net: a local and global contextual encoding network for effective and efficient medical image segmentation
Yating Zhu, Meifang Peng, Xiaoyan Wang 0007, Ming Xia 0005, Xiaoting Shen, Weiwei Jiang 0005
Appl. Intell.5
2025 Lightweight Multi-Stage Aggregation Transformer for robust medical image segmentation
Xiaoyan Wang 0007, Yating Zhu, Dongyan Guo, Pan Mu, Ming Xia 0005, Cong Bai, Zhongzhao Teng, Shengyong Chen
Medical Image Anal.7
2025 Visualizing the Smart Environment in AR: An Approach Based on Visual Geometry Matching
abstract
This article presentsInsight, an AR system for visualizing the IoT-enabled smart environment without relying on the unique appearances, barcodes, world coordinates, or wireless signals of IoT infrastructures. The system analyzes the camera video and motion data taken by mobile AR equipment to extract the self and cross visual geometries describing the poses and geographic distribution of nearby IoT devices. To recognize IoT devices using the extracted geometries,Insightoperates in two phases. At deployment time, it learns pairwise mappings from the visual geometries to the corresponding device identities. After that, it leverages the geometries scanned at run time to look for a partial assignment to the recorded geometries, allowing it to automatically recognize the IoT devices in AR view. As such, our system turns the IoT device recognition task into a geometry matching problem, which is further formalized as to perform Subset, Incomplete, and Duplicated Point Cloud Registration (SID-PCR) in this work. We design a deep neural network paying specific edge- and spectral-wise graph attention to solve SID-PCR, and implement a prototype that adaptively requests visual geometry scan and registration operations for accurate recognition. The performance ofInsightis validated using both synthetic data and a real-world testbed.
Ming Xia 0005, Min Huang 0022, Qiuqi Pan, Xiaoyan Wang 0007, Kaikai Chi
IEEE Trans. Mob. Comput.1
2024 A Multi-Scale Grid Attention Based Network for Blood Vessel Extraction from Retinal Fundus Images
abstract
Accurate blood vessel extraction in medical images is a key step in the diagnose of vascular-related diseases and can also provide important guiding information for surgery. However, blood vessels are characterized by elongated and tortuous, irregular shapes, making the task of vessel segmentation challenging. Multi-scale contextual features can help to better understand the structure and morphology of an image, and attention mechanisms plays a quite essential role in the abstraction of contextual features. In this paper, we propose a multi-scale dual-branching blood vessel image segmentation method based on the grid attention mechanism. In the encoder, a dual combination of CNN and grid attention is used for multi-scale contextual information extraction, and then channel attention is utilized to fusion the feature representations extracted in a weighted manner, which can be better adapted to complicated vascular structures. In the decoder, each feature map of the encoder layer is up-sampled by sub-pixel convolution, and subsequently, the high-resolution feature maps obtained are concatenated to generate the final segmentation mask by further convolution operations. This design can fully utilize the multi-scale feature information extracted by the encoder while preserving the spatial structure and details of the image. We conducted experiments on two public retinal fundus image datasets and the results demonstrate that our approach outperforms CNN-based, attention-based and state-of-the-art CNN-Attention combined models.
Xiaoyan Wang 0007, Meifang Peng, Yuanhao Zheng, Jianhao Yu, Yating Zhu, Ming Xia 0005
SMC6
2024 Optimization of Detection Interval for Mobile Multiuser Molecular Communication With Anomalous Diffusion in Internet of Nano Things
abstract
The Internet of Nano Things (IoNT) as an innovative technology which integrates nanotechnology and communication technology involves the use of nanomachines to create systems that can interact with each other. IoNT provides possibilities for applications in many fields, i.e., drug delivery, healthcare, and environment monitoring. It promotes the development of molecular communication (MC) and nanonetworks. In this article, a mobile multiuser MC system where multiple transmitters and one receiver move in the anomalous diffusion channel is studied. First, the molecular division multiple access technique where multiple transmitters convey information by delivering different types of information molecules is employed. The mathematical expression of the average bit error rate (BER) of this MC system is derived. Then we formulate a multiobjective optimization problem with the constraint that the detection interval has lower and upper bounds in order to achieve minimum average BER of this system. Furthermore, we use the alternative search algorithm with gradient projection to solve this optimization problem and obtain the optimal detection interval. Finally, numerical results show this algorithm has good convergence and the optimization results can be approximate to the optimal solutions obtained by the exhaustive search.
Zhen Cheng 0001, Ming Xia 0005, Kaikai Chi
IEEE Internet Things J.4
2023 Multi-Stage Aggregation Transformer for Medical Image Segmentation
abstract
Capturing rich multi-scale features is essential for resolving complex variations in medical image segmentation. In this paper, we explore how to fully utilize the advantages of Convolutional neural networks (CNN) and Transformer, and propose a novel multi-stage aggregation architecture named MA-Transformer for accurate segmentation of medical images with large variations and blurs. Specifically, an encoder module is introduced in each stage, which is a dual-branch structure parallelly combining Transformers and convolutions. By such design, the self-attention can provide a global context for CNN to extract multi-resolution complementary features stage by stage, thus the feature representations are gradually enhanced with local details and contextual information. Multi-scale semantic features are then combined with skip connections in the decoder to produce the final result. Extensive experiments on public medical imaging datasets demonstrate our superior segmentation performance, compared to the state-of-the-art CNN-based, Transformer-based approaches and CNN-Transformer combined approaches. Code will be made publicly available.
Xiaoyan Wang 0007, Minghan Shao, Dongyan Guo, Ming Xia 0005, Cong Bai
ICASSP6
2023 Toward Sustainable Transportation: Robust Lane-Change Monitoring With a Single Back View Cabin Camera
abstract
The risk of death and injury from traffic crashes has been universally recognized as one of the most serious threats to sustainable development. Among all the factors in traffic crashes, aggressive driving in which the driver usually makes excessive lane changes to overtake other vehicles is prevalent. As such, monitoring lane changes and providing real-time warnings is beneficial for improving transportation sustainability. This article presents BackWatch, a novel vehicle-mounted sensing system that uses a back view cabin camera monitoring the steering wheel rotations to track lane-change events. BackWatch consists of an encoder network to extract essential visual features of steering wheel rotations, and an inference network incorporating the visual and GPS speed features to recognize the resulting lane changes. Our system does not rely on precise coordinate alignment between the monitoring device and the vehicle, nor the wearables worn by the driver, and is robust against different drivers, vehicles, driving speeds, and environmental settings. We evaluate the system based on 16 hours of real-world on-road driving data collected from three pairs of cars and drivers under different traffic and environmental conditions. The results show that BackWatch achieves 0.952 of precision and 0.981 of recall on the detection of lane changes.
Ming Xia 0005, Linghao Ying, Kaikai Chi, Keping Yu
IEEE Trans. Intell. Transp. Syst.1
2023 Physical-Assisted Routing for Proactive Avoidance of Nomadic Obstacles in IoT
abstract
With the broadening of the radio spectrum to higher frequency bands, wireless links are more prone to blockages by nomadic obstacles. However, existing routing schemes mostly follow the network-oriented design principle, which makes it difficult to react quickly to sudden obstruction. This article proposes PAR, a novel physical-assisted routing scheme for the Internet of Things. PAR takes a physical-oriented viewpoint attempting to mitigate unexpected link blockages by leveraging sensor observations of obstacles at individual nodes. It analyzes the signatures of sensor measurements to directly estimate the positions and sizes of obstacles and to proactively infer remedial routing decisions even before the transmission failure occurs. During the network deployment time, PAR lets each node learn and calibrate the geographic distribution of neighbors with respect to its sensor measurements by exchanging the sensor observations of a mobile guidance obstacle. During the runtime, an obstacle bypassing algorithm is then developed to find the shortest detour routes by comparing the local sensor measurements with the calibrated positions of neighbors to immediately resume data forwarding. We evaluate the efficacy of PAR using simulation and a real-world testbed. It is observed that PAR significantly improves the success rate while reducing redundant hops and routing control overhead.
Ming Xia 0005, Jiaquan Jin, Biqian Liu, Yu Hen Hu, Xiaoyan Wang 0007, Kaikai Chi
ACM Trans. Sens. Networks1
2022 Automatic and accurate segmentation of peripherally inserted central catheter (PICC) from chest X-rays using multi-stage attention-guided learning
Xiaoyan Wang 0007, Ye Sheng, Chenglu Zhu, Cong Bai, Ming Xia 0005, Zhanpeng Shao, Ruiyi Zhao, Zhenjie Liu
Neurocomputing7
2022 SSA-Net: Spatial self-attention network for COVID-19 pneumonia infection segmentation with semi-supervised few-shot learning
Xiaoyan Wang 0007, Yiwen Yuan, Dongyan Guo, Ming Xia 0005, Zhenhua Wang 0003, Cong Bai, Shengyong Chen
Medical Image Anal.6
2021 Multi-scale Hierarchical Transformer structure for 3D medical image segmentation
abstract
Transformers have demonstrated great potential in computer vision tasks which is benefited from its ability of long-range dependency modeling. However, its core material multi-head self-attention (MHSA) has high computational and spacial complexity. Besides, tokens embedded with 1D position sequence are unable to represent inductive bias of locality and neighbor contextual information, which are essential for high-resolution downstream vision tasks. Thus, transformer is hard to use in image segmentation tasks, especially in 3D medical images. In this paper, we propose a novel multi-scale hierarchical framework (MSHT) that efficiently incorporates the local modeling ability of CNN and long-range modeling ability of transformer for accurate 3D medical image segmentation. Unlike many prior Transformer-based solutions, the proposed MSHT first adopts parallel CNN and transformer as encoder block to extract the global and local feature representations. As the core component for our MSHT, a position embedding method for 3D medical image (3D IPE) is proposed to generate relational and absolute position encoding for tokens fed in transformer block. We conduct an extensive evaluation on both Lits2017 dataset and Kits2019 dataset. The results indicate that our MSHT leads to a substantial performance over other CNN-based methods on 3D image segmentation.
Xiaoyan Wang 0007, Bangze Zhang, Cong Bai, Ming Xia 0005, Peiliang Sun
BIBM6
2021 INSIGHT: An AR-Enabled User Interface for Vision-Based Markerless Interaction with IoT Nodes
abstract
Real-world, long-running Internet of Things (IoT) requires intense user-node interaction in the deployment, network operation, and maintenance stages. The rapidly increasing number of IoT nodes urges a new user-friendly interface to reduce the interaction complexity. This paper presents INSIGHT, an AR-enabled user interface for IoT nodes allowing the users to directly grasp perceptual node information from the surrounding videos shot by mobile AR devices such as smart glasses and phones. INSIGHT incorporates camera, magnetic, and IMU sensor data from an AR device to detect, locate, and recognize IoT nodes markerlessly. An intuitive user interface will then be overlaid on each recognized node for fetching/feeding information from/to the node. We have validated the performance of INSIGHT in terms of localization accuracy and the capability to distinguish nearby nodes. The results show that INSIGHT works well even in densely deployed networks.
Ming Xia 0005, Xiaoyan Wang 0007, Zhen Cheng 0001, Kaikai Chi
WCNC1
2021 PaI-Net: A modified U-Net of reducing semantic gap for surgical instrument segmentation
abstract
Abstract Tracking the instruments in a surgical scene is an essential task in minimally invasive surgery. However, due to the unpredictability of scenes, automatically segmenting the instruments is very challenging. In this paper, a novel method named parallel inception network (PaI‐Net) is proposed, in which an attention parallel module (APM) and an output fusion module (OFM) are integrated with U‐Net to improve the segmentation ability. Specially, APM utilizes multi‐scale convolution kernels and global average pooling operations to extract semantic information and global context information of different scales, while OFM combines the feature maps of the decoder part to aggregate the abundant boundary information of shallow layers and the rich semantic information of deep layers together, which achieve a significant improvement in generating segmentation masks. Finally, the evaluation of proposed method on robotic instruments segmentation task from Medical Image Computing and Computer Assisted Intervention Society (MICCAI) and retinal image segmentation task from International Symposium on Biomedical Imaging (ISBI) show that our model has achieved advanced performance on multi‐scale semantic segmentation and is superior to the current state‐of‐the‐art models.
Xiaoyan Wang 0007, Xingyu Zhong, Cong Bai, Ruiyi Zhao, Ming Xia 0005
IET Image Process.7
2020 PACE: Physically-Assisted Channel Estimation
abstract
Radio link quality is highly influenced by changes in the physical environment. To sustain reliable and efficient data delivery, link quality estimation is essential for Cyber-Physical Systems (CPSs) or Internet of Things (IoT). Network-based link quality estimation methods estimate the link quality by monitoring data transmissions. In a dynamic environment, the accuracy of link quality so estimated may become degraded because the accuracy must be balanced against the overhead of data transmissions. In this work, we propose to incorporate sensor readings available in a CPS/IoT system to augment existing link quality estimation. We call this a Physical-Assisted Channel Estimator (PACE). By analyzing sensor readings that are highly correlated to the link quality, PACE may detect the change of link quality in real-time. Evaluation conducted on a real intelligent parking system shows that compared to existing network-based methods, PACE reacts to persistent disturbances much more quickly without sacrificing robustness to transient fluctuations, and achieves higher accuracy even under a low data transmission rate. With PACE, the data delivery performance of routing protocols can be significantly improved. We expect PACE to be the first milestone towards Physical-Assisted Cyber Systems (PACSs) for fulfilling the vision of environment-aware computing and communication.
Ming Xia 0005, Biqian Liu, Yu Hen Hu, Kaikai Chi, Xiaoyan Wang 0007, Jiajia Liu 0001
IEEE Trans. Wirel. Commun.1
2019 Mode-oriented hybrid programming of sensor network nodes for supporting rapid and flexible utility assembly
Ming Xia 0005, Kaikai Chi, Xiaoyan Wang 0007, Zhen Cheng 0001
Comput. Networks1
2017 Minimization of Transmission Completion Time in Wireless Powered Communication Networks
abstract
Recently, the newly emerging wireless powered communication network (WPCN) has drawn significant interests, where network nodes are powered by the energy harvested from the radio-frequency (RF) signal. This paper studies the WPCN where one hybrid sink (H-sink) coordinates the wireless energy/information transmissions to/from a set of one-hop nodes powered by the harvested RF energy only. The transmission completion time (TCT) minimization for the uplink (UL) transmissions of a given number of bits per node is considered. First, we prove that the harvest-then-transmit (HTT) transmission strategy is one of the transmission strategies able to achieve the minimal TCT, where all nodes first harvest the RF energy broadcast by the H-sink in the downlink and then send their independent information to the H-sink in the UL by time-division multiple access. Then for the HTT transmission, we prove that in order to achieve the minimal TCT, each node must transmit with constant power and consume all available energy, which helps to simplify the considered TCT minimization problem to be the optimization of time allocated for the H-sink's wireless energy transfer and the nodes' wireless information transmissions, and we formulate the optimal time allocation problem as a nonlinear optimization problem. Finally, we prove that it is a convex optimization problem. Due to the inexistence of explicit closed-form expressions of optimal time allocations to minimize TCT, one efficient algorithm is presented to obtain the optimal time allocations. Simulation results show that, compared with the available transmission strategies, the designed TCT-minimized transmission achieves a significantly smaller TCT.
Kaikai Chi, Yihua Zhu 0001, Yanjun Li 0004, Liang Huang 0006, Ming Xia 0005
IEEE Internet Things J.5