VLDB 2026 Research / reviewers in the wild / expert
Shouxiang Ni
dblp:299/4220
· DBLP profile ↗
8ranked-venue papers
2as first author
8since 2021 · last 2026
0009-0006-9416-0102ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 5 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Semi-supervised image segmentation via selective self-ensembling and boundary uncertainty suppressionabstractImage segmentation plays a key role in many image-guided clinical applications and deep learning technology has proven effective for this task when sufficient labeled images are available. However, it is very time-consuming and labor-intensive to obtain adequate pixel-level image labels. To alleviate the scarcity of labeled images, we propose a novel semi-supervised segmentation method based on the available uncertainty-aware mean teacher (UAMT) framework by introducing two different strategies, i.e., selective self-ensembling (SSE) and boundary uncertainty suppression (BUS). The SSE dynamically selects multiple best student models across different training steps to update the teacher model's weights, while the BUS reduces boundary segmentation errors and improves the quantitative potential of loss functions through a unique uncertainty estimation function. With the two strategies, our proposed method was able to obtain promising segmentation performance with limited labeled images and abundant unlabeled ones. We trained and validated our proposed method by segmenting multiple objects from three public datasets (i.e., PROMISE, REFUGE, and RETA). Extensive experiments showed that our proposed method achieved better segmentation performance than the UAMT, along with the average Dice score (DSC) of 0.7990 for three different objects, and can compete with several existing semi-supervised methods (i.e., HCMT, SASSNet, and DTC). • A novel semi-supervised learning method was developed for accurate image segmentation. • A performance-driven exponential moving average (pEMA) was proposed for semi-supervised learning. • A unique exponential function was proposed for uncertainty estimation. • Extensive segmentation experiments showed the advantage of the developed method. Xiaoguo Yang, Yabo Wu, Shouxiang Ni, Hongmei He, Chuoying Tan, Wencan Wu, Quanyong Yi |
Eng. Appl. Artif. Intell. | 3 |
| 2026 | Low-Latency Versus High-Precision: LAM-Based Cross-Modal Semantic Communication for Emergency ResponseabstractCross-modal semantic communication plays a crucial role in emergency response systems. However, there are still two major challenges in practical applications: insufficient semantic compression due to the computing constraints of the source device, and poor signal generation at the sink device caused by harsh communication environment. To this end, this work fully leverages multi-modal interactive large AI model (LAM) distillation to enhance sparse feature representation for dynamic residual compression, while utilizing LAM adversarial feature mapping to repair the feature impairment for accurate signal generation, thereby achieving low latency and high reliable communication. Specifically, we first propose a scalable cross-modal semantic communication framework, which constructs a multi-modal semantic knowledge base (MSKB) by semantic similarity to support cross-modal semantic codec. On this basis, residual-guided cross-modal semantic encoding (RCSE) is designed, which employs video-infrared bidirectional knowledge distillation to lightweight LAM for feature extraction, and initially compresses features by inter-modal correlations. Further, according to the link state, the similar features in MSKB are dynamically stripped and the residual features are compressed for low-latency transmission. Additionally, an adversarial mapping-based cross-modal semantic decoding (AMCSD) is developed, which leverages adversarial mutual information (MI) to map features in LAM to impaired features, and discriminate semantic fidelity to enhance the mapping robustness, ensuring accurate signal generation. Finally, an emergency simulation platform is constructed. Experiments demonstrate that our scheme reduces the transmission latency by 27.64% and improves the signal generation precision by 14.06%. Dan Wu 0001, Shouxiang Ni, Liang Zhou 0002 |
IEEE J. Sel. Areas Commun. | 3 |
| 2026 | Cross-Modal Private and Covert Communication With Constraints of Computation and BandwidthabstractWith the popularity of unmanned inspection and telemedicine applications, edge source devices need to transmit multimodal data such as video, text, infrared, etc. in real-time in a computing power and bandwidth-constrained environment, whereas traditional single-modal compression and steganography methods struggle to balance efficiency and security. To address the current lack of end-to-end efficient methods that jointly consider multimodal coding and steganography, this paper is the first to propose a cross-modal private and covert communication scheme with constraints of computation and bandwidth, unifying cross-modal fusion, invertible steganography, and hyperprior image compression into a single model, and introducing a multi-stage joint training strategy. First, a lightweight composition-steganography-compression pipeline is designed at the edge source device, which composes visible and infrared signals into a semantically enhanced image, generates a natural stego image through invertible steganography, and produces a compact bit-stream through deep compression, achieving a 51.5% reduction in bitrate and a 48.3% reduction in encoding latency. Second, a multi-stage decoding-cross-modal reconstruction pipeline is built at the sink device to sequentially complete bitstream decompression, inverse steganography, and cross-modal reconstruction, ultimately outputting visible images (PSNR 30.37 dB) and infrared images (PSNR 31.58 dB), with a 40.9% reduction in decoding latency. Finally, a three-stage joint training further enhances PSNR and saves an additional 6% bitrate. Experimental results validate the efficiency, security, and robustness of the proposed method in resource-constrained environments. Shouxiang Ni, Jinmin Gu, Xinbiao Yi, Dan Wu 0001, Anjie Jiang, Liang Zhou 0002 |
IEEE J. Sel. Areas Commun. | 1 |
| 2023 | Cloud-Edge-Client Collaborative Learning in Digital Twin Empowered Mobile NetworksabstractDigital twin (DT) has emerged as a key enabler for the intelligent-oriented evolution of mobile networks. With the rise of privacy concerns for enabling intelligent applications in DT-empowered mobile networks (DTMNs), federated learning has garnered wide attention due to its potential on breaking down data silos. However, the data privacy of federated learning is greatly threatened by emerging gradient leakage attacks, and the need for frequent knowledge exchange limits its training efficiency over resource-constrained DTMNs. To circumvent such dilemmas, this work first proposes a privacy-enhanced federated learning framework based on cloud-edge-client collaborations. Particularly, model splitting between clients and edge servers makes gradient leakage attacks computationally prohibitive, and cloud-side partial model aggregation provides hierarchical data utility. To improve the training efficiency of the proposed learning framework, we further establish its communication and computation cost models, and develop a DT-assisted multi-agent deep reinforcement learning-based resource scheduler for joint client association and channel assignment. Finally, as a case study of intelligent applications in DTMNs, a human-robot collaborative nursing task is designed to evaluate the practical performance of our proposed scheduler. Experimental results show its superiority in saving training costs and preserving learning accuracy. Lindong Zhao, Shouxiang Ni, Dan Wu 0001, Liang Zhou 0002 |
IEEE J. Sel. Areas Commun. | 2 |
| 2022 | Haptic Signal Reconstruction in eHealth Internet of ThingsabstractWith the haptic technology continuously enlarging the eHealth Industry Internet of Things (IIoT) ecosystem, haptic perception service which requires effective haptic signal reconstruction for immersive experience has become an indispensable function. However, the majority of existing haptic signal reconstruction methods are generally inefficient because of undergoing extremely complex operations or inefficient feature representations. To resolve this dilemma, this article proposes a long short-term memory-based force reconstruction network (LSTM-FRN) by designing a novel sparse attention module for low-latency reconstruction and a novel metric learning-based constraint for high-precision reconstruction, yielding an excellent tradeoff between the computational complexity and feature representation. To train our network, we construct a large-scale data set of synchronous needle motion signals and haptic signals in acupuncture needle insertion. Finally, we build an interactive needle insertion training system (HapAR-NITS) by integrating augmented reality (AR), the LSTM-FRN-based haptic reconstruction as well as a skill assessment subsystem. Comprehensive experiments demonstrate that the proposed multiple technologies enable our HapAR-NITS to achieve satisfying immersive experience and manipulation effects. Ang Li 0012, Shouxiang Ni, Liang Zhou 0002 |
IEEE Internet Things J. | 3 |
| 2022 | Edge-Based Cross-Modal Communications for Remote HealthcareabstractMedical robots with audio-video-haptic streams, as indispensable devices for remote healthcare, are playing ever-increasing roles in mitigating the spread of infectious diseases. However, existing medical robots are far from precise and efficient because of the following two technical challenges, including i) how to ensure the haptic fidelity for precise manipulation, and ii) how to alleviate the impact of haptic streams on the quality of visual navigation for efficient operation. To this end, this work explores the benefits of edge-based cross-modal communications (CMCs), which take full advantage of potential correlations among different modalities’ streams, to realize high reliability and throughput. Specifically, to compensate for the reliability loss caused by wireless transmission, a semantic-aided cross-modal reconstruction framework is firstly designed at edge nodes for high haptic fidelity. Then, a user experience-driven stream scheduling strategy is developed to enhance the visual quality by fully leveraging edge computing and network slicing. In particular, different from traditionally interrupting audio/video stream transmission to prioritize haptic streams, we jointly schedule resources to different modalities’ streams via estimating haptic arrival time. Finally, as a classical case study, we independently construct a remote throat swab sampling platform based on CMCs to evaluate practical performance, and numerical results indicate the significant improvements in terms of various metrics. Shouxiang Ni, Dan Wu 0001, Liang Zhou 0002 |
IEEE J. Sel. Areas Commun. | 2 |
| 2021 | A Virtual and Non-contact Bounce System Based on Ultrasonic Haptic TechnologyabstractHaptics has become an important research direction at the present stage. It is an inevitable trend to integrate haptic feedback into human-computer interaction technology in the future. However, traditional contact haptic feedback technology has disadvantages such as fixed equipment and limited space. In order to overcome this difficulty, this paper mainly studies the non-contact haptic feedback, and proposes a virtual bounce system based on ultrasonic haptic. First of all, we analyze the principle of the main modules of the system to provide a theoretical basis for the system design. Secondly, in the design, we combine tactile stimuli with visual stimuli to enhance the visualization and comfort of users during operation. Then, we analyze the performance of the system through experiment. The results show that the virtual bounce system proposed in this paper can free users from the shackles of devices, improve the interaction comfort and immersive experience of users. Gaolong Jiang, Shouxiang Ni |
IWCMC | 4 |
| 2021 | A Solution of Human-Computer Remote Interaction with Tactile FeedbackabstractIn human-computer interaction, we need to use human's multiple sensory channels and motion channels to improve its naturalness and efficiency. These channels include voice, handwriting, posture, vision, expression, etc. Visual-based human-computer interaction systems are often mentioned in the literature. The system uses a camera to monitor and track the operator's posture, and sets corresponding commands to control the movement of the robot. On top of that, we add tactile feedback to allow the robot to accomplish more tasks. In the vision system, first we obtain color images and depth images through a depth camera. Secondly, we estimate the human posture through OpenPose. Finally, we map the motion of the human arm to the motion of the robotic arm. In the tactile system, first we control the movement of the magic hand through the exoskeleton. Secondly, we obtain the force values of the fingers through the magic hand. Finally, we provide force feedback to each finger through the exoskeleton. Shouxiang Ni |
IWCMC | 1 |