Tao Ren 0001

dblp:61/11243-1 · DBLP profile ↗
← Back
41ranked-venue papers
8as first author
40since 2021 · last 2026
0000-0003-0408-9447ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 14 · 6 first-author · 14 since 2021Artificial intelligence and machine learning · 11 · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 11 since 2021Systems, architecture and hardware · 6 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 From Scene to Object: Enhancing Open-Vocabulary Object Detection via Foreground-Background Context Reasoning
abstract
Open-Vocabulary Object Detection (OVOD) aims to detect both known and novel categories in complex visual scenes, surpassing the limitations of conventional closed-set detectors. Recent advances in vision-language models (VLMs) like CLIP have enabled zero-shot recognition by aligning visual features with large-scale textual embeddings. However, current OVOD approaches often fall short by overlooking critical contextual and semantic cues necessary for discovering a broader range of novel objects. To address this, we propose BFDet, a scene-to-object reasoning framework that leverages the complementary strengths of Large Language Models (LLMs) and VLMs. BFDet introduces a novel scene-to-object reasoning mechanism grounded in foreground-background context interaction. It first uses high-confidence objects to infer the scene-level background. This scene background then guides the discovery of foreground objects by prompting an LLM to generate scene-sensitive novel object candidates. These candidates are subsequently verified through cross-modal alignment and used as high-quality pseudo-labels to enrich detector training. Designed as a plug-and-play module, BFDet integrates seamlessly into existing detection pipelines and consistently improves performance on novel categories across COCO and LVIS benchmarks.
Yanqi Li, Jianwei Niu 0002, Ningbo Gu, Tao Ren 0001
AAAI4
2026 DoKnowAD: Calibrating Normal Representations with Refined Domain Knowledge to Enhance Time Series Anomaly Detection
abstract
Time series anomaly detection (TSAD) is critical in various real-world applications. Due to the high cost of manual annotation, unsupervised methods are commonly employed to distinguish abnormal patterns from normal ones based on data or representation characteristics. However, the limited coverage of a single dataset often leads to misclassifying test-time normal patterns that deviate from the training distribution as anomalies. In view of this, we propose to introduce domain knowledge from auxiliary datasets (AuxSets) to enhance domain-level normality understanding in the target dataset (TargetSet). However, through in-depth analysis on the representation space of the TargetSet after incorporating AuxSets, we find that consistent knowledge about normality from homogeneous AuxSets do little help to TargetSet, while diverse knowledge from heterogeneous AuxSets can bring semantic confusion of normality for TargetSet, both of which can degrade TargetSet detection performance. To address the issue, we design DoKnowAD, a framework that introduces a Representation HyperVolume Estimation metric to identify helpful heterogeneous AuxSets, and further adopts contrastive learning to enforce loose coupling between datasets and high cohesion within single dataset to calibrate the TargetSet’s representation space, thus mitigating knowledge confusion. Extensive experiments on five popular datasets across different domains demonstrate that DoKnowAD consistently outperforms existing TSAD baselines in various metrics.
Shiwang Xing, Jianwei Niu 0002, Tao Ren 0001
AAAI3
2026 Many Minds, One Path: LLM-Augmented Consensus Decision for Distributed Control in Multi-Agent Collaborative Stable Scenarios
Zhuohao Yu 0002, Zhe Li 0025, Tao Ren 0001, Chenxue Wang, Junjie Wang 0001, Qing Wang 0001
AAAI3
2026 EdgeFormer: Latency-Aware Collaborative Multi-Head Attention of Transformer Inference in Edge Networks
abstract
Recent breakthroughs in Transformer-based large models, have driven widespread tasks, yet their reliance on centralized cloud deployment raises significant privacy risks due to sensitive data exposure.While edgebased collaborative inference offers a privacypreserving alternative, existing methods face critical limitations: static model partitioning cannot adapt to dynamic edge resource fluctuations, and rigid multi-head attention handling overlooks semantic-critical prioritization and parallelism.We propose EdgeFormer, a latency-aware framework for distributed Transformer inference in resource-constrained edge networks.EdgeFormer dynamically allocates model blocks across devices via efficiencystorage trade-off optimization and introduces collaborative Multi-Head Attention (cMHA), which distributes semantic-critical attention heads across devices while pruning redundant ones under real-time constraints.We further develop LiScore, a composite metric integrating attention diversity and latency costs, alongside a similarity-based retrieval method to reduce recomputation overhead.Extensive experiments demonstrate that EdgeFormer achieves up to 2.01× inference acceleration over state-of-theart baselines with ≤1.06% accuracy loss, maintaining robustness under varying edge conditions.
Jianwei Niu 0002, Bin Dai 0009, Tao Ren 0001
ACL (1)4
2026 MM2SQL: A Benchmark and Method for Visually-Grounded SQL Generation
Shengze Shi, Tao Ren 0001, Tingrui Yang, Jun Hu 0015
ICDE2
2026 SemCache: Semantic-Aware Cache Sharing for Efficient Multi-User LoRA-Adapted LLM Inference at the Edge
Tao Ren 0001, Zheyuan Hu 0001, Jianwei Niu 0002
INFOCOM1
2026 Benefit From Noise: Detecting Time-Series Anomaly by Distinguishing Prior and Posterior Noises
abstract
With the rapid development of digital technologies, a large range of real-world systems, spanning from cloud servers, IoT devices, to industrial control systems, continuously generate vast amounts of time series data. Time series anomaly detection (AD) plays a crucial role in maintaining system stability by identifying unusual patterns from normal distributions, with the primary challenge lies in learning effective anomaly-discriminative representations. Recently, diffusion models have been applied to time series AD due to their strong representational capabilities. However, existing diffusion-based methods typically rely on reconstruction errors, which not only fail to fully exploit the representational potential of diffusion models but also be computationally intensive. To address these limitations, through experimental observation and theoretical analysis, we show thatspecific regions of the diffusion noises exhibit stronger representation capabilitiesfor normal patterns, which can be leveraged to enhance AD performance and reduce computational costs. Building on these insights, we propose NoiseAD, a diffusion noise-guided anomaly detection method incorporating an optimal noise steps selection approach to identify diffusion steps with higher resolution. Extensive experiments on diverse benchmarks demonstrate the superiority of NoiseAD over state-of-the-art methods, further substantiated by insightful visualizations. Code could be available athttps://github.com/shiwang-Xing/NoiseAD.
Shiwang Xing, Jianwei Niu 0002, Tao Ren 0001, Joel J. P. C. Rodrigues
IEEE Trans. Knowl. Data Eng.3
2025 VisQ2SQL: Towards SQL-Driven Data Visualization via LLMs-Grounded Preference Learning
abstract
Text-to-Visualization (Text2Vis) aims to democratize data insights for non-expert users by transforming natural language query (NLQ) into visualization specification (VS). In view of the high dependence of rule-based methods on predefined VS templates and poor NLQ understanding ability of small data-driven methods, recent works resort to leveraging pre-trained LLMs to perform NLQ-understanding and VS-generating in Text2Vis tasks via prompt-guided in-context learning. However, existing LLM-based methods still fall short of satisfactory end-to-end Text2Vis performances primarily owing to the limited ability of pre-trained LLMs in directly retrieving and operating NLQ-intended tabular data. Inspired by the SQL generating ability born with latest LLMs, this paper proposes the idea of harnessing LLMs for SQL-driven visualization data retrieval and operation. Nonetheless, there remains a nonneglectable gap between data visualization queries in Text2Vis tasks and SQL retrieving queries in LLM corpus. To fill the gap, this paper proposes a Visualization Query to SQL (VisQ2SQL) framework to obtain NLQ-intended data, primarily by fine-tuning LLMs through preference learning data-retrieval SQLs induced from VS and those generated by LLMs. We conduct extensive experiments to demonstrate the superiority of VisQ2SQL over SOTA methods, and various ablation studies to verify the efficacy of VisQ2SQL.
Shengze Shi, Tao Ren 0001, Jun Hu 0015
ICASSP2
2025 Benefit from Seen: Enhancing Open-Vocabulary Object Detection by Bridging Visual and Textual Co-Occurrence Knowledge
Yanqi Li, Jianwei Niu 0002, Tao Ren 0001
ICCV3
2025 Enabling Communication-efficient and Robust Federated Learning over Packet Lossy Networks via Random Interleaved Vector Quantization
abstract
In packet erasure networks, federated learning (FL) typically suffers more prohibitive communication overhead from massive retransmissions of high-dimensional gradients. As a result, recent studies are dedicated to developing retransmission-free gradient compression techniques with erasure resilience. Nonetheless, two limitations remain unsolved: existing works neither explore why packet erasure degrades the performance of FL nor exploit the spatial correlations among gradient entries for better compression. In this paper, we investigate FL performance degradation via analyzing model updating deviation and find that the deviation is exacerbated by dependencies among lost gradient entries. On top of this observation, we propose FedRIVQ, a communication-efficient and robust FL framework taking a customized compressor termed random interleaved vector quantization (VQ). FedRIVQ leverages the spatial correlations among gradient entries with VQ and randomly interleaves these entries prior to VQ to eliminate their dependencies. These innovations allow all gradient entries to share an identical erasure probability, thereby packet erasure is equivalent to random erasure, which significantly improves both communication efficiency and the robustness of FL. Theoretical analysis and experimental results consistently demonstrate the effectiveness of our designs.
Yixuan Guan 0001, Jianwei Niu 0002, Tao Ren 0001, Xuefeng Liu 0001
ICME3
2025 MAIF: Efficient Multi-agent Communication via Intention Filter
abstract
Effective information interaction can enhance the coordination capabilities in collaborative multi-agent reinforcement learning (MARL). A popular communication scheme is the exchange of agents’ intention information, which typically involves broadcasting agents’ intention information to all other agents. This not only increases the communication overhead of the entire system but also interferes with the decision-making of agents to some extent due to the reception of irrelevant intention information from other agents. In this paper, we propose the Multi-Agent Intention Filter (MAIF), which filters out intentions irrelevant to the agent, allowing the agent to focus on the intentions of agents with whom it is more likely to collaborate, thus promoting effective cooperation among agents. Specifically, our method first uses a State Simulator to coordinate the joint intentions of agents based on their current joint observations and predict target states, providing prior knowledge for subsequent intention filtering. Then, we calculate the causal effect of agents’ intentions on target states and filter out intentions that are irrelevant to the agent, thereby promoting effective cooperation. Each agent will use the filtered intention information to assist in decision-making. Experimental results show that our method outperforms strong baselines in multiple cooperative MARL tasks under various task settings.
Chengcheng Wu, Jianwei Niu 0002, Tao Ren 0001
IJCNN3
2025 ExplabOff: Towards Explorative and Collaborative Task Offloading via Mutual Information-Enhanced MARL
Tao Ren 0001, Zheyuan Hu 0001, Jianwei Niu 0002
INFOCOM1
2025 Closing the Feedback Loop in Text2Vis: Refining Visualization with Vision-Language Models
abstract
Text-to-Visualization (Text2Vis) generates data visualizations directly from natural language queries, democratizing access to data insights. Early Text2Vis efforts, primarily relying on rule-based systems and machine learning models, struggled to handle semantically intricate queries. The advent of large language models (LLMs) allows for better generalization in generating visualization code. However, LLM-based approaches have mainly focused on textual or code-level optimizations, neglecting the potential benefits of assessing and improving visualized charts. Hence, we propose Visualization Refinement (VisRef), a novel framework based on vision-language models (VLMs) to enhance Text2Vis outputs. (1) Knowledge Extraction -- VisRef extracts visualization assessment knowledge through a hierarchical contrastive prompt and multi-granularity quality assessment framework by comparing superior ground-truth charts with inferior Text2Vis outputs; and (2) VLM Fine-Tuning -- This knowledge is used to fine-tune a VLM through a two-stage approach, including warm-up and iterative preference alignment phases, to judge visualization quality and provide code-level refinement suggestions. Experimental results demonstrate that VisRef significantly outperforms state-of-the-art approaches, including LLM-based and VLM-prompted, and exhibits strong orthogonal compatibility with existing approaches.
Shengze Shi, Tao Ren 0001, Guoliang Zhu, Guan Dong Feng, Jun Hu 0015
ACM Multimedia2
2025 WFF: Wavelet-based Information Fusion for Multimodal Knowledge Graph Link Prediction
abstract
Multimodal link prediction on multimodal knowledge graphs is an inference task aimed at finding missing triples, which seeks to improve prediction accuracy by leveraging a wide range of information. However, current multimodal knowledge graph link prediction methods are primarily designed in the spatial domain, necessitating ever-growing complexity in fusion strategies. In addition, most of them focus on only three modalities (text, image, and structure). Forward-looking approaches, however, should accommodate a broader array of modalities. Motivated by the operational simplification enabled by transforming features into the frequency (or time-frequency) domain, we propose a wavelet-transform-based multimodal link prediction method, WFF, which offers high modal extensibility and low fusion complexity. Specifically, for unimodal information, we designed a Unimodal Time-Frequency Knowledge Enhancement module, UTFKE, which extracts time-frequency features via discrete wavelet transform and enhances information quality through adaptive filtering. To address the challenge of multimodal fusion, we devised a Multimodal Time-Frequency Knowledge Fusion module MTFKF that supports high modal extensibility and enables effective, efficient integration. Extensive experiments on multiple well-known datasets demonstrate that WFF outperforms strong baselines and achieves state-of-the-art performance. In addition, WFF extends modality to audio and video, further validating the model's effectiveness. Our code is available at https://github.com/xxd12315/WFF.
Xiaodi Xu, Ye Wang 0021, Tao Ren 0001, Tian Qiao
ACM Multimedia4
2025 SITOff: Enabling Size-Insensitive Task Offloading in D2D-Assisted Mobile Edge Computing
abstract
Mobile edge computing (MEC), along with device-to-device (D2D) assisted MEC (D-MEC), are promising technologies that could improve the quality-of-experience for mobile devices (MDs) by offloading their tasks to edge servers or nearby idle MDs. There is a popular trend to develop distributed task offloading algorithms using multi-agent reinforcement learning (MARL), whose adoption of central critics during training makes the offloading still size-sensitive. Therefore, this paper proposes a Size-Insensitive Task Offloading (SITOff) algorithm for D-MEC based on fully-distributed offloading without maintaining any central venue. Specifically, taking advantage of the inherent graph-like structure of D-MEC, SITOff adopts graphs to represent MDs’ states and relationships and form each MD's local knowledge about D-MEC through graph computation. Furthermore, considering the limitation of local knowledge in performing whole performance-oriented offloading, each MD utilizes D2D-transmitting to exchange knowledge with its neighbors and form a comprehensive knowledge about D-MEC to enhance the coordination of distributed offloading. Additionally, regarding the different impacts of neighbors’ knowledge, each MD leverages attention mechanisms to selectively learn its neighbors’ knowledge during knowledge-exchange. Extensive experimental results show the superiority of SITOff over state-of-the-art MARL-based offloading algorithms in D-MEC with various MDs, and the easy collaboration of SITOff with curriculum-learning for large-scale D-MEC offloading.
Zheyuan Hu 0001, Jianwei Niu 0002, Tao Ren 0001, Xuefeng Liu 0001, Mohsen Guizani
IEEE Trans. Mob. Comput.3
2024 Integrating Structure and Text for Enhancing Hyper-relational Knowledge Graph Representation via Structure Soft Prompt Tuning
abstract
Different from traditional knowledge graphs, where facts are usually represented as (subject, relation, object), hyper-relational knowledge graphs (HKGs) allow facts to be associated with additional relation-entity pairs to constrain the validity of facts. HKGs contain a substantial amount of textual information, which plays a crucial role in enriching representations. However, existing HKG embedding methods mainly rely on structural information but overlook textual information in HKGs, which are less effective in representing entities with limited structural information. To address this issue, the paper proposes HIST (Hyper-relational Knowledge Graph Encoder Integrating Structure and Text), which incorporates textual information and structural information in HKGs to enhance representations of entities and relations. HIST adopts the graph convolutional network to extract structural information and utilizes it to generate the Structure Soft Prompt. During the Structure Soft Prompt Tuning process, the textual information and structural information are fully integrated to generate more comprehensive representations. Additionally, an effective contrastive learning method for HKG embedding is formulated to improve the efficiency of negative sampling. Experimental results show that HIST achieves state-of-the-art performance on several public datasets. Our code is available at https://github.com/QieFangBaiLuQingYaJian/HIST.
Hui Wang 0117, Xiaodi Xu, Ye Wang 0021, Tao Ren 0001
CIKM6
2024 FedMDC: Enabling Communication-Efficient Federated Learning over Packet Lossy Networks via Multiple Description Coding
abstract
Federated learning (FL) generally suffers significant communication overhead from high-traffic gradient synchronization. The majority of existing studies on this problem aim at compressing gradients under the premise of reliable transmission. While transmission reliability can be ensured via TCP by default, the notably increased latency and retransmitted packets are prohibitive for most clients in FL. To tackle this issue, we propose FedMDC, a retransmission-free compression framework for FL over packet lossy networks. Given clients’ limited resources, FedMDC adopts multiple description coding to encode gradients into redundant descriptions for erasure resilience simply through multiplying an overcomplete matrix; and then quantizes these descriptions for compression. To further reduce quantization distortion and computational overhead, a reduced decoding algorithm is developed by decoding the aggregation of all clients’ encodings in conjunction with a customized dither quantization design. Besides, FedMDC explicitly supports adaptive bitrates subject to clients’ heterogeneous communication budgets, which maximize resource utilization to facilitate distortion reduction and accelerate model convergence. Theoretical analysis and experimental results both demonstrate the effectiveness of our scheme.
Yixuan Guan 0001, Xuefeng Liu 0001, Tao Ren 0001, Jianwei Niu 0002
ICME3
2024 NID-SLAM: Neural Implicit Representation-based RGB-D SLAM In Dynamic Environments
abstract
Neural implicit representations have been explored to enhance visual SLAM algorithms, especially in providing high-fidelity dense map. Existing methods operate robustly in static scenes but struggle with the disruption caused by moving objects. In this paper we present NID-SLAM, which significantly improves the performance of neural SLAM in dynamic environments. We propose a new approach to enhance inaccurate regions in semantic masks, particularly in marginal areas. Utilizing the geometric information present in depth images, this method enables accurate removal of dynamic objects, thereby reducing the probability of camera drift. Additionally, we introduce a keyframe selection strategy for dynamic scenes, which enhances camera tracking robustness against large-scale objects and improves the efficiency of mapping. Experiments on publicly available RGB-D datasets demonstrate that our method outperforms competitive neural SLAM approaches in tracking accuracy and mapping quality in dynamic environments.
Jianwei Niu 0002, Qingfeng Li 0004, Tao Ren 0001, Chen Chen 0141
ICME4
2024 M3OFF: Module-Compositional Model-Free Computation Offloading in Multi-Environment MEC
abstract
Computation offloading is one of the key issues in mobile edge computing (MEC) that alleviates the tension between user equipment's limited capabilities and mobile application's high requirements. To achieve model-free computation offloading when reliable MEC dynamics are unavailable, deep reinforcement learning (DRL) has become a popular methodology. However, most existing DRL-based offloading approaches are developed for a single MEC environment, with invariant system bandwidth, edge capability, task types, etc., while realistic MEC scenarios tend to be of high diversity. Unfortunately, in multi-MEC environments, DRL-based offloading faces at least two challenges, learning inefficiency and interference of offloading experiences. To address the challenges, we propose a DRL-based Multi-environmental Module-compositional Modelfree computation OFFloading (M3OFF) framework. M3OFF generates offloading policies using module composition instead of a single DRL network so that learning efficiency could be improved by reusing the same modules and learning interference could be reduced by composing different modules. Furthermore, we design multiple module composition-specific training methods for M3OFF, including alternate modules-and-composer updates to improve training stability, loss-regularization to avoid module degeneration, and module-dropout to mitigate overfitting. Extensive experimental results on both simulation and testbed demonstrate that M3OFF outperforms the performances of most state-of-the-arts in multi-MEC and reaches close to single-MEC.
Tao Ren 0001, Zheyuan Hu 0001, Jianwei Niu 0002, Weikun Feng, Hang He
INFOCOM1
2024 FedTC: Enabling Communication-Efficient Federated Learning via Transform Coding
abstract
Federated learning (FL) enables distributed training via periodically synchronizing model updates among participants. Communication overhead becomes a dominant constraint of FL since participating clients usually suffer from limited bandwidth. To tackle this issue, top-k based gradient compression techniques are broadly explored in FL context, manifesting powerful capabilities in reducing gradient volumes via picking significant entries. However, previous studies are primarily conducted on the raw gradients where massive spatial redundancies exist and positions of non-zero (top-k) entries vary greatly between gradients, which both impede the achievement of deeper compressions. Top-k may also degrade the performance of trained models due to biased gradient estimations. Targeting the above issues, we propose FedTC, a novel transform coding based compression framework. FedTC transforms gradients into a new domain with more compact energy distributions, which facilitates reducing spatial redundancies and biases in subsequent sparsification. Furthermore, non-zero entries across clients from different rounds become highly aligned in the transform domain, motivating us to partition the gradients into smaller entry blocks with various alignment levels to better exploit these alignments. Lastly, positions and values of non-zero entries are independently compressed in a block-wise manner with our customized designs, through which a higher compression ratio is achieved. Theoretical analysis and extensive experiments consistently demonstrate the effectiveness of our approach.
Yixuan Guan 0001, Xuefeng Liu 0001, Jianwei Niu 0002, Tao Ren 0001
INFOCOM4
2024 Achieving Fast Environment Adaptation of DRL-Based Computation Offloading in Mobile Edge Computing
abstract
One of the key issues in mobile edge computing (MEC) is computation offloading, most policies of which are developed based on mathematical programming (MP). Due to the high computational complexity of iterative programming in MP-based policies, recent years have seen a popular trend to develop offloading policies based on deep reinforcement learning (DRL). However, on account of the poor generalization ability of DRL models in MEC environments with different network sizes and settings, it is difficult to directly apply DRL-based offloading policies in unseen MEC environments. Motivated by this, we propose a DRL-based environment-adaptive offloading framework (DEAT), including a size-adaptive scheme (SIED) and setting-adaptive component (SEAL). SIED leverages the idea of ‘time division multiplexing’ to adapt to varying MEC network sizes and order-unaware feature extraction to mitigate impacts of different size-changing orders. SEAL adopts system dynamics embedding and offloading policy embedding, which guide the finding of the closest pre-training MEC environment and offloading policy, respectively, to achieve fast setting-adaptation with only few exploring interactions in unseen MEC environments. Extensive experiments are conducted via both simulation and testbed to demonstrate the adaptation performance advantages of DEAT in unseen MEC environments compared to the state-of-the-art offloading approaches.
Zheyuan Hu 0001, Jianwei Niu 0002, Tao Ren 0001, Mohsen Guizani
IEEE Trans. Mob. Comput.3
2023 Joint Optimization of System Bandwidth and Transmitting Power in Space-Air-Ground Integrated Mobile Edge Computing
Yuan Qiu 0006, Jianwei Niu 0002, Tao Ren 0001, Xinzhong Zhu, Kuntuo Zhu
ICA3PP (6)5
2023 TransOff: Towards Fast Transferable Computation Offloading in MEC via Embedded Reinforcement Learning
abstract
Mobile edge computing (MEC) has been proposed as a promising paradigm to provide mobile devices with both satisfactory computing capacity and task latency. One key issue in MEC is computation offloading (CompOff), which has attracted numerous research interests. Most existing CompOff approaches are developed based on iterative programming (IterProg), that calculates a CompOff action based on system dynamics each time mobile tasks arrive. Due to the heavy dependency of IterProg on reliable system dynamics, as well as the online computational burden, recent years have seen a popular trend to develop CompOff approaches based on deep reinforcement learning (DRL), which could generate real-time model-free CompOff actions. However, due to the intrinsic poor generalization of DRL, it is hard to directly apply DRL-based policies in new MEC environments, and long-time fine-tuning is often required. To address the challenge, this paper proposes a fast transferable CompOff framework (named TransOff), based on the idea of embedded reinforcement learning. Specifically, TransOff is composed of multiple primitive CompOff policies (pCOPs) and a multiplicative composition function (MCF). The pCOPs and MCF are pre-trained in a diverse variety of MEC environments. When encountering new MEC environments, pCOPs are kept fixed to prevent catastrophic forgetting of pre-trained CompOff skills, while only MCF is fine-tuned to produce new compositions of pCOPs to achieve fast transfer. We conduct extensive experiments via both numerical simulation and real testbed, indicating the fast transfer ability of TransOff compared to the state-of-the-art DRL-based and meta learning-based CompOff approaches.
Zheyuan Hu 0001, Jianwei Niu 0002, Tao Ren 0001
ICDCS3
2023 GCFormer: Granger Causality based Attention Mechanism for Multivariate Time Series Anomaly Detection
abstract
Multivariate time series anomaly detection, crucial for ensuring the safety of real-world systems, primarily focuses on extracting characteristics from time series under normal condition, and identifying potential anomalies throughout the evaluation process. Recent studies have achieved fruitful progress through mining the spatio-temporal relationships from multivariate time series, however, these approaches mostly neglect the latency among series which could lead to higher false alarm. Granger causality presents a promising solution to extract these inherent time-lagged relationships. Nonetheless, the intricate and dynamic relationships among numerous time series in real-world systems surpass the ability of linear Granger causality. To address this, we extend the linear Granger causality and propose the Granger Causal Former (GCFormer), a novel approach that leverages attention mechanisms to learn the inherent causal spatio-temporal relationships between historical and current timestamps across multiple time series. Specifically, GCFormer develops a Spatio-Mask (SM) to select the top-k most relevant series and a Temporal-Mask (TM) to concentrate attention on more recent historical timestamps. Moreover, to mitigate overfitting and ensure a smooth training process, GCFormer introduces an adjust top-k method and a TM penalty term. We evaluated GCFormer on four real-world benchmark datasets, demonstrating its superior performance over state-of-the-art approaches. Further analysis and a case study highlight the model’s novelty and interpretability.
Shiwang Xing, Jianwei Niu 0002, Tao Ren 0001
ICDM3
2023 Enabling Communication-Efficient Federated Learning via Distributed Compressed Sensing
abstract
Federated learning (FL) trains a shared global model by periodically aggregating gradients from local devices. Communication overhead becomes a principal bottleneck in FL since participating devices usually suffer from limited bandwidth and unreliable connections in uplink transmission. To address this problem, the gradient compression methods based on compressed sensing (CS) theory have been put forward recently. However, most existing CS-based works compress gradients independently, ignoring the gradient correlations between participants or adjacent communication rounds, which constrains the achievement of higher compression rates. In view of the above observation, we propose a novel gradient compression scheme named FedDCS, guided by distributed compressed sensing (DCS) theory. Following the design philosophy of separate encoding and joint decoding in DCS, FedDCS compresses gradients for participants in each round separately while reconstructing them at the central server jointly via fully exploiting correlated gradients from the previous round, which are known as side information (SI). Benefiting from this design, reconstruction performance is significantly improved with fewer decoding errors also iterations under the identical compression rate, and the total uploading bits to achieve model convergence are considerably reduced. Theoretical analysis and extensive experiments conducted on MNIST and Fashion-MNIST both verify the effectiveness of our approach.
Yixuan Guan 0001, Xuefeng Liu 0001, Tao Ren 0001, Jianwei Niu 0002
INFOCOM3
2023 FEAT: Towards Fast Environment-Adaptive Task Offloading and Power Allocation in MEC
Tao Ren 0001, Zheyuan Hu 0001, Hang He, Jianwei Niu 0002, Xuefeng Liu 0001
INFOCOM1
2023 An enhanced data-driven framework for early kick detection based on imbalanced multivariate time series classification
Shiwang Xing, Jianwei Niu 0002, Haige Wang, Tao Ren 0001, Xiaoyan Shi
Neural Comput. Appl.4
2022 Anatomical Landmarks Annotation on 2D Lateral Cephalograms with Channel Attention
abstract
Cephalometric tracing is widely used in orthodontic diagnosis and treatment planning. Since manual landmark lo-calization suffers from severe inter-observer and intra-observer inconsistency, a large number of efforts have been made by researchers to develop automatic localization methods. However, most of the existing methods are developed based on rules which sample uniformly from origin images rather than with the highest density in a focal point and ignore intermediate layers' results in their networks or their outputs' channels. To address the issue, this paper proposes a deep learning model based on multi-scale and multi-channel attention to identify landmarks. The channel attention network is first trained by multi-scale image patches cropped from 100 Cephalograms, and then enhanced by cross-layer connections to extract high-level features, finally involved in the collaboration with a three-layer MLP module to accurately locate coordinates. We conduct extensive evaluation on a real cephalometric X-ray data-set and a Non-public dataset, both achieve promising performance improvements especially in terms of high-precision detection.
Dongfeng Du, Tao Ren 0001, Chen Chen 0141, Yiran Jiang, Guangying Song, Qingfeng Li 0004, Jianwei Niu 0002
CCGRID2
2022 Multi-scale Fusion and Global Semantic Encoding for Affordance Detection
abstract
Affordance detection is of great importance in robot operational tasks, due to its capability of helping robots effectively interact with objects. Many affordance detectors have been proposed, primarily based on two-stage object detection, significantly suffering from the slow detection speed. Hence, recent years have saw the popularity of one-stage affordance detectors based on encoder-decoder structures that adopt dilated convolutions to extract high-resolution feature maps. However, dilated convolutions on high resolution features tend to be computation and memory-intensive, greatly limiting the practicality of one-stage detectors. To address the issue, this paper proposes a novel convolution neural network (CNN) based encoder-decoder architecture, without the need of adopting dilated convolution. A repeated multi-scale feature-map-fusion network is introduced to produce high-resolution features, effectively improving the feature representation performance of the model. Besides, a semantic encode module is embedded to capture global semantic information and enhance category-relevant feature maps. Extensive experiments show that the proposed framework outperforms the start-of-art methods with only 1/2 of the computational cost, while maintaining the inference at the speed of 26ms per image, indicating the promising affordance-detection performance of our network on IIT-AFF dataset and UMD dataset.
Huiyong Li 0005, Tao Ren 0001, Yuanbo Dou, Qingfeng Li 0004
IJCNN3
2022 pFedGF: Enabling Personalized Federated Learning via Gradient Fusion
abstract
Data heterogeneity is one of the main challenges faced by federated learning (FL). Unlike traditional FL methods (e.g. FedAvg) which train a global model for all clients, personalized federated learning (PFL) can address the above problem by training a personalized model for each client. Current mainstream PFL researches first obtain a global model through collaborative training among all clients and then fine-tune the global model on each client's local data to obtain personalized models. However, this two-staged approach has a drawback: when the heterogeneity of different clients is large, the obtained final global model can deviate from the distributions of all clients, and therefore is not a good starting point for updating personalized models. In this paper, we propose pFedGF, a new PFL method based on gradient fusion. Different from traditional two-staged PFL, in each round of pFedGF, each client maintains two gradients simultaneously, a global gradient to capture information from all clients, and a local gradient that reflects the specific distribution of each client. The two gradients are fused to obtain the updated direction of the personalized model for each client. We carried out experiments on MNIST, FMNIST, and CIFAR-10 datasets. The results demonstrate that in the presence of data heterogeneity, pFedGF outperforms other PFL methods.
Xinghao Wu, Jianwei Niu 0002, Xuefeng Liu 0001, Tao Ren 0001, Zhangmin Huang, Zhetao Li
IPDPS4
2022 Deep Reinforcement Learning Based Computation Offloading in Heterogeneous MEC Assisted by Ground Vehicles and Unmanned Aerial Vehicles
Hang He, Tao Ren 0001, Dong Liu 0008, Jianwei Niu 0002
WASA (3)2
2022 Meta-MADDPG: Achieving Transfer-Enhanced MEC Scheduling via Meta Reinforcement Learning
Tao Ren 0001, Dong Liu 0008, Jianwei Niu 0002
WASA (3)2
2022 Toward Mobility-Aware Computation Offloading and Resource Allocation in End-Edge-Cloud Orchestrated Computing
abstract
Mobile devices (MDs) have undergone a booming development, yet are still capacity limited in computation and energy resources and thus could face troubles when serving computation-intensive and delay-sensitive applications. Mobile-edge computing (MEC) has been proposed to accommodate MDs with both satisfactory latency and acceptable resources, by offloading MDs’ tasks to near-deployed edge servers (ESs). Whereas, offloading tasks solely to ESs are difficult to meet distinct requirements of various applications, which leads to the emergence of end–edge–cloud orchestrated computing (EECOC). However, studies on EECOC are still insufficient, most existing of which do not consider MDs’ dynamical movements partly due to the intractability of associating optimal ESs with moving MDs. To address the issue, a novel deep reinforcement learning (DRL)-based mobility-aware (MA) EECOC scheduling approach is proposed in this article. With the goal of minimizing maximal task latency, we first formulate and transform the optimization problem into a Markov decision problem (MDP). Then, we enable DRL with elaborately designed reward functions and integrate it with NoisyNet to obtain near-optimal solutions. Furthermore, a MA component based on ConvLSTM is developed to extract MDs’ temporal–spatial distribution features and predict their movements, which are further utilized to facilitate the decision making of computation-offloading and resource-allocation actions. Extensive experimental results indicate the promising performance improvements of our approach against the state-of-the-art approaches in various scenarios.
Bin Dai 0009, Jianwei Niu 0002, Tao Ren 0001, Mohammed Atiquzzaman
IEEE Internet Things J.3
2022 Enabling Efficient Scheduling in Large-Scale UAV-Assisted Mobile-Edge Computing via Hierarchical Reinforcement Learning
abstract
Due to the high maneuverability and flexibility, unmanned aerial vehicles (UAVs) have been considered as a promising paradigm to assist mobile edge computing (MEC) in many scenarios including disaster rescue and field operation. Most existing research focuses on the study of trajectory and computation-offloading scheduling for UAV-assisted MEC in stationary environments, and could face challenges in dynamic environments where the locations of UAVs and mobile devices (MDs) vary significantly. Some latest research attempts to develop scheduling policies for dynamic environments by means of reinforcement learning (RL). However, as these need to explore in high-dimensional state and action space, they may fail to cover in large-scale networks where multiple UAVs serve numerous MDs. To address this challenge, we leverage the idea of “divide-and-conquer” and propose HT3O, a scalable scheduling approach for large-scale UAV-assisted MEC. First, HT3O is built with neural networks via deep RL to obtain real-time scheduling policies for MEC in dynamic environments. More importantly, to make HT3O more scalable, we decompose the scheduling problem into two-layered subproblems and optimize them alternately via hierarchical RL. This not only substantially reduces the complexity of each subproblem, but also improves the convergence efficiency. Experimental results show that HT3O can achieve promising performance improvements over state-of-the-art approaches.
Tao Ren 0001, Jianwei Niu 0002, Bin Dai 0009, Xuefeng Liu 0001, Zheyuan Hu 0001, Mingliang Xu 0001, Mohsen Guizani
IEEE Internet Things J.1
2022 Enhancing generalization of computation offloading policies in novel mobile edge computing environments by exploiting experience utility
Tao Ren 0001, Jianwei Niu 0002, Yuan Qiu 0006
J. Syst. Archit.1
2022 An Efficient Online Computation Offloading Approach for Large-Scale Mobile Edge Computing via Deep Reinforcement Learning
abstract
Mobile edge computing (MEC) has been envisioned as a promising paradigm that could effectively enhance the computational capacity of wireless user devices (WUDs) and quality of experience of mobile applications. One of the most crucial issues of MEC is computation offloading, which decides how to offload WUDs’ tasks to edge severs for further intensive computation. Conventional mathematical programming-based offloading approaches could face troubles in dynamic MEC environments due to the time-varying channel conditions (caused primarily by WUD mobility). To address the problem, reinforcement learning (RL) based offloading approaches have been proposed, which develop offloading policies by mapping MEC states to offloading actions. However, these approaches could fail to converge in large-scale MEC due to the exponentially-growing state and action spaces. In this article, we propose a novel online computation offloading approach that could effectively reduce task latency and energy consumption in dynamic MEC with large-scale WUDs. First, a RL-based computation offloading and energy transmission algorithm is proposed to accelerate the learning process. Then, a joint optimization method is adopted to develop the allocating algorithm, which obtains near-optimal solutions for energy and computation resources allocation. Simulation results show that the proposed approach can converge efficiently and achieve significant performance improvements over baseline approaches.
Zheyuan Hu 0001, Jianwei Niu 0002, Tao Ren 0001, Bin Dai 0009, Qingfeng Li 0004, Mingliang Xu 0001, Sajal K. Das 0001
IEEE Trans. Serv. Comput.3
2021 ArtCoder: An End-to-End Method for Generating Scanning-Robust Stylized QR Codes
abstract
Quick Response (QR) code is one of the most worldwide used two-dimensional codes. Traditional QR codes appear as random collections of black-and-white modules that lack visual semantics and aesthetic elements, which inspires the recent works to beautify the appearances of QR codes. However, these works adopt fixed generation algorithms and therefore can only generate QR codes with a pre-defined style. In this paper, combining the Neural Style Transfer technique, we propose a novel end-to-end method, named ArtCoder, to generate the stylized QR codes that are personalized, diverse, attractive, and scanning-robust. To guarantee that the generated stylized QR codes are still scanning-robust, we propose a Sampling-Simulation layer, a module-based code loss, and a competition mechanism. The experimental results show that our stylized QR codes have high-quality in both the visual effect and the scanning-robustness, and they are able to support the real-world application.
Hao Su 0001, Jianwei Niu 0002, Xuefeng Liu 0001, Qingfeng Li 0004, Ji Wan, Mingliang Xu 0001, Tao Ren 0001
CVPR7
2021 Distributed Task Offloading based on Multi-Agent Deep Reinforcement Learning
abstract
Recent years have witnessed the increasing popularity of mobile applications, e.g., virtual reality, unmanned driving, which are generally computation-intensive and latency-sensitive, posing a major challenge for resource-limited user equipment (UE). Mobile edge computing (MEC) has been proposed as a promising approach to alleviate the problem, by offloading mobile tasks to the edge server (ES) deployed in close proximity to UE. However, most existing task offloading algorithms are primarily based on centralized scheduling, which could suffer from the ‘curse of dimensionality’ in large MEC environments. To address this issue, this paper proposes a fully distributed task offloading approach based on multi-agent deep reinforcement learning, whose critic and actor neural networks are trained under the assistance of global and local network states, respectively. In addition, we design a model parameter aggregation mechanism, along with a normalized fine-tuned reward function, to further improve the learning efficiency of the training process. Simulation results show that our proposed approach could achieve substantial performance improvements over baseline approaches.
Shucheng Hu, Tao Ren 0001, Jianwei Niu 0002, Zheyuan Hu 0001, Guoliang Xing
MSN2
2021 An application of multi-objective reinforcement learning for efficient model-free control of canals deployed with IoT networks
Tao Ren 0001, Jianwei Niu 0002, Jiahe Cui, Zhenchao Ouyang, Xuefeng Liu 0001
J. Netw. Comput. Appl.1
2021 An Efficient Model-Free Approach for Controlling Large-Scale Canals via Hierarchical Reinforcement Learning
abstract
Large-scale canals with cascaded pools are constructed wordwide to divert water from rich to arid areas to mitigate water shortages. Efficient control of canals is essential to improve water-diversion performance. Numerous model-based approaches have been proposed and made great progress for canal control. However, when the predictive model is unavailable or unpromising for long time step predictions, model-free approaches could be considered as a possible way to achieve efficient control. Since most existing model-free approaches are focused on control of small canals or reservoirs, this article proposes a new control approach named policy and action reinforcement learning (PARL) for large-scale canals. We leverage the idea of “divide and conquer” to decompose the control task of large-scale canals into policy learning and action learning subtasks, and develop PARL by means of hierarchical reinforcement learning. Extensive experiments are conducted via numerical simulation on the case study of Chinese South to North Water Transfer Project, and experimental results show that PARL can achieve desirable performance improvements over other model-free learning approaches.
Tao Ren 0001, Jianwei Niu 0002, Xuefeng Liu 0001, Jiyan Wu, Xiaohui Lei
IEEE Trans. Ind. Informatics1
2020 MBBNet: An edge IoT computing-based traffic light detection solution for autonomous bus
Zhenchao Ouyang, Jianwei Niu 0002, Tao Ren 0001, Yanqi Li, Jiahe Cui, Jiyan Wu
J. Syst. Archit.3