VLDB 2026 Research / reviewers in the wild / expert
Ang Li 0012
dblp:33/2805-12
· DBLP profile ↗
9ranked-venue papers
4as first author
9since 2021 · last 2026
0000-0001-7149-3250ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 5 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MobileFold: Linear-complexity global modeling via spatial fold convolution
Ang Li 0012 |
Pattern Recognit. Lett. | 1 |
| 2026 | Towards General Cross-Modal Visual Coding for Emergency CommunicationsabstractMulti-modal visual signals are prevalent in emergency communications. To ensure high reliability of signal transmission under bandwidth constraints, it is crucial to compress redundant information both within and between modalities as much as possible, and ensure the fidelity of the reconstructed signals. Most existing studies depend exclusively on single-modal coding schemes and fail to effectively leverage the semantic correlations between modalities. In this paper, we introduce an end-to-end general cross-modal visual coding scheme, namely CMVC, which aims to jointly compress multi-modal visual signals (such as visible and infrared signals). First, we propose a cross-modal asynchronous entropy module that extracts common features using a cross-attention mechanism. Additionally, we enhance the accuracy of common features extraction by maximizing mutual information loss. This module further compresses multi-modal visual signals by compressing only the residual features between modalities. Second, we propose a cascaded enhancement module based on cross-modal Mamba that fuses complementary information to enhance the reconstruction quality of multi-modal visual signals. Finally, extensive experimental results demonstrate that our scheme significantly outperforms other advanced methods on visible-infrared datasets. Even at low bitrates, multi-modal visual signals can still achieve excellent reconstruction quality. Additionally, our scheme exhibits outstanding compression and reconstruction performance when applied to visible-depth signals, effectively demonstrating its robustness and generalizability. Lindong Zhao, Ang Li 0012, Bin Kang, Dan Wu 0001, Liang Zhou 0002 |
IEEE Trans. Multim. | 4 |
| 2025 | The Best of Both Worlds: Task-Oriented Cross-Modal Semantic TransmissionabstractTraditional communication faces significant challenges in multimodal scenarios, including surging network capacity demands and the neglect of semantic value. Although semantic communication achieves data compression through semantic feature extraction and refinement, existing methods have drawbacks such as inflexible compression, high computational complexity, and the separation of feature extraction and refinement from transmission scheduling, making it difficult to trade-off semantic integrity and transmission efficiency. To this end, this paper proposes a task-oriented cross-modal semantic transmission scheme, which is based on the task requirements, dynamically adjusts the feature fusion strength and utilizes semantic correlations to enhance the task-related features, and then evaluates the feature priority and selects the task-critical features for transmission, realizing the best of both worlds. Specifically, we design a task feedback-based cross-modal feature fusion method, which establishes a mapping between computing state and feature fusion level, and dynamically optimizes the fusion weight decomposed by the fusion level through task loss. On this basis, the cross-modal features are aligned and complemented using semantic correlation to refine the task-relevant features. Further, we propose a feature priority-based multi-mode semantic transmission method, which determines the feature priority by a task-response-based feature importance assessment model. Accordingly, a reinforcement learning (RL)-based dual-modal feature selection strategy is designed to select task-critical features for reliable transmission by coupling transmission performance with task requirements. Additionally, simulation results show that compared with the baseline, our method improves task accuracy by 10.6% and reduces transmission latency by 7.2% on average. Dan Wu 0001, Ang Li 0012, Liang Zhou 0002 |
GLOBECOM | 4 |
| 2024 | Toward Low-Latency Cross-Modal Communication: A Flexible Prediction SchemeabstractTo ensure the users’ immersive experience in cross-modal communication, overcoming the end-to-end (E2E) latency through prediction has attracted attention and shown its superiority. However, existing prediction schemes encounter formidable challenges in the presence of multi-modal signals, primarily to adapt and satisfy the prediction requirements of diverse multi-modal services, as well as to fully exploit and effectively utilize the correlation features of multi-modal signals for precise prediction. To this end, this work presents a flexible prediction scheme for low-latency cross-modal communication. Specifically, we first propose an adaptive prediction-aware cross-modal communication framework, which reduces the delay by predicting and transmitting the future multi-modal signals in advance, and flexibly adjusts the prediction horizon to satisfy the prediction accuracy of different multi-modal services. Next, we design an information gain-assisted graph attention (IGGA) method for cross-modal signal prediction, which leverages the graph attention block to extract the intra-modal, inter-modal spatial and temporal correlation features, and effectively optimize and utilize these features with the information gain (IG), thereby facilitating precise cross-modal signal prediction. Finally, numerical experiments conducted on a self-built dataset, a public dataset, and a multi-modal acupuncture platform demonstrate the superiority of the proposed scheme in low-latency cross-modal communication. Ang Li 0012, Dan Wu 0001, Liang Zhou 0002, Yi Qian 0001 |
IEEE Trans. Mob. Comput. | 3 |
| 2024 | Toward General Cross-Modal Signal Reconstruction for Robotic TeleoperationabstractThe multi-modal robotic teleoperation, as an important application in human-computer interaction (HCI), is playing a significant role in various domains such as industry, healthcare, and education. However, existing robotic teleoperation systems face significant challenges with multi-modal signals, primarily in designing a cross-modal communication architecture that caters to diverse modal requirements and ensuring high-quality cross-modal signal reconstruction even in poor network conditions. To this end, this work proposes a general cross-modal signal reconstruction scheme by taking full advantage of the correlation among different modality signals. Specifically, we first propose a scalable cross-modal communication architecture that meets the diverse needs of various modality signals using multi-modal encoding and multi-directional decoding, eliminating the need for a specialized feature extraction model. Next, we design a masked auto-encoder with discriminator assistance (MAE-D) cross-modal signal reconstruction method, which leverages the idea of generative confrontation by combining the codec for signal reconstruction with the discriminator responsible for assessing the authenticity of the reconstructed signal to achieve accurate and efficient cross-modal signal reconstruction. Finally, numerical experiments conducted on our self-built multi-modal dataset, a public dataset, and a teleoperation simulation platform demonstrate that the proposed scheme offers significant advantages in cross-modal signal reconstruction. Ang Li 0012, Dan Wu 0001, Liang Zhou 0002 |
IEEE Trans. Multim. | 2 |
| 2023 | Cross-Scale Haptic Object Recognition for Intelligent RobotabstractHaptic technology enables robots to touch and understand the interactions between objects in the reality. Advanced haptic sensing systems can not only collect pressure, temperature and stiffness of touched objects, but also avoid destructive operations, and assist in navigation and posture control for robots. In order to smoothly interact with different types of objects, in the haptic system, it is necessary to develop haptic object recognition methods for effective haptic perception capability. However, compared to RGB images, haptic images collected by opticallybased haptic sensors are similar in appearance, which makes traditional convolutional neural networks (e.g.,ResNet, VGG, etc.) ineffective. Therefore, in this paper, we are inspired by popular attention mechanism and multi-scale strategies, and propose a cross-scale attention based haptic object recognition network for object-robot interaction. In particular, On the one hand, we design a cross-scale attention module in convolutional neural networks to acquire spatial contextual feature. On the other hand, we design a learnable bilinear fusion strategy to integrate above spatial contextual feature with original haptic feature, so as to effectively discriminate haptic images. Experimental results on ViTac dataset have shown the effectiveness of our approach. Ang Li 0012, Xin Wei 0001 |
IWCMC | 1 |
| 2023 | Global Information-Assisted Fine-Grained Visual Categorization in Internet of ThingsabstractIn fine-grained visual categorization (FGVC), most part-based frameworks do not work effectively in some extremely challenging scenarios such as partial occlusion. This limitation is due to the heavy disorder of local features extracted from such occluded targets. To address this issue, we propose a global information-assisted network (GIAN), where auxiliary global information can search the useful elements of local information and integrate with them for an efficient unified feature representation. In particular, in order to acquire the global information, we design a global attention-concentrated convolutional neural network (GAC-CNN) by extending a convolutional neural network with a nonlocal GCN module. Then, the unified feature representation is produced by two strategies. On the one hand, a global–local aggregation strategy is developed to selectively integrate global features with local features through consistency evaluation and reweighting method. On the other hand, an alternative knowledge distillation strategy is developed to help generate more powerful global and local features. Two strategies collaboratively make the unified features more robust and more discriminative than traditional part-based features. Experimental results show that the proposed GIAN can achieve accuracies of 92.8%, 93.8%, and 95.7% on CUB-200-2011, FGVC Aircraft, and Stanford Cars, respectively. Ang Li 0012, Bin Kang, Dan Wu 0001, Liang Zhou 0002 |
IEEE Internet Things J. | 1 |
| 2022 | Haptic Signal Reconstruction in eHealth Internet of ThingsabstractWith the haptic technology continuously enlarging the eHealth Industry Internet of Things (IIoT) ecosystem, haptic perception service which requires effective haptic signal reconstruction for immersive experience has become an indispensable function. However, the majority of existing haptic signal reconstruction methods are generally inefficient because of undergoing extremely complex operations or inefficient feature representations. To resolve this dilemma, this article proposes a long short-term memory-based force reconstruction network (LSTM-FRN) by designing a novel sparse attention module for low-latency reconstruction and a novel metric learning-based constraint for high-precision reconstruction, yielding an excellent tradeoff between the computational complexity and feature representation. To train our network, we construct a large-scale data set of synchronous needle motion signals and haptic signals in acupuncture needle insertion. Finally, we build an interactive needle insertion training system (HapAR-NITS) by integrating augmented reality (AR), the LSTM-FRN-based haptic reconstruction as well as a skill assessment subsystem. Comprehensive experiments demonstrate that the proposed multiple technologies enable our HapAR-NITS to achieve satisfying immersive experience and manipulation effects. Ang Li 0012, Shouxiang Ni, Liang Zhou 0002 |
IEEE Internet Things J. | 1 |
| 2021 | Understanding Digital Forensic Characteristics of Smart Speaker EcosystemsabstractWith a built-in intelligent personal voice assistant providing Q&A services, smart speaker ecosystems combine multiple compatible components, including the internet of things (IoT) technology, mobile devices, and cloud computing. However, as it is closely related to people's daily lives, security and privacy issues have gained worldwide attention. Components in the ecosystem are interconnected and chained together to enable the ecosystem to perform increasingly diverse operations. By collecting meaningful data from smart speaker ecosystems, we can reconstruct user behavior and provide a holistic explanation for finding the root cause of an observable symptom. This highlights the need for digital forensic research to enhance the security and privacy of smart speaker ecosystems. In this paper, we first discuss the digital forensic characteristics of a smart speaker ecosystem. Then, we propose a proof-of-concept digital forensic tool based on data provenance, that supports the identification, acquisition, and analysis of client-side artifacts from local devices. Ang Li 0012, Xiao Fu 0005, Bin Luo 0003, Xiaojiang Du, Mohsen Guizani |
GLOBECOM | 2 |