Xiongwei Wu

dblp:172/1093 · DBLP profile ↗
← Back
24ranked-venue papers
15as first author
10since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 4 first-author · 5 since 2021Computer networks · 9 · 6 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 6 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2025 Heterogeneous multi-modal graph network for arterial travel time prediction
Hangyu He, Mengyun Xu, Xiongwei Wu
Appl. Intell.4
2025 A Macroscopic-Fundamental-Function-Aided Neural Networks for Traffic Flows Prediction From Mobile Signaling Data
abstract
The widespread deployment of the Internet of Things (IoT) allows for the utilization of mobile signaling data (MSD) in cellular networks to perceive the underlying traffic states on road networks. Nevertheless, the application of MSD for traffic flow prediction can be impeded by the limited positioning accuracy of MSD and stringent privacy policies. To surmount these obstacles, we propose a traffic flow prediction method that integrates an estimation process using a macroscopic fundamental function (M) with a positional factor (P), combined with the sequence-to-sequence (S2S) neural networks model (MPS2S). First, we estimate traffic flow using a macroscopic fundamental function (M). In this approach, we introduce a positional factor (P) that fuses road inflows and outflows to capture the relationship between cellular network and road network. Next, we employ an S2S neural network model, which considers the spatial and temporal aggregation information to infer future traffic flows. Specifically, an encoder estimates traffic parameters from MSD, a decoder infers multistep traffic flows, and a coefficients estimation method introduces the macroscopic fundamental function into the neural network. To validate the prediction performance of our proposed model, we collect video data from roadside cameras to measure ground-truth values. In addition, we analyze the impact of the positional factor on our estimation method and the impact of spatial and temporal aggregation on our prediction model in real-world freeway scenarios. The results indicate that MSD can link the potential traffic information, and our method contributes to better traffic flow prediction.
Xiongwei Wu, Mengyun Xu, Chunping Li, Xuesong Wu 0004, Li Li 0082
IEEE Internet Things J.2
2024 OVFoodSeg: Elevating Open-Vocabulary Food Image Segmentation via Image-Informed Textual Representation
abstract
In the realm of food computing, segmenting ingredients from images poses substantial challenges due to the large intra-class variance among the same ingredients, the emergence of new ingredients, and the high annotation costs as-sociated with large food segmentation datasets. Existing approaches primarily utilize a closed-vocabulary and static text embeddings setting. These methods often fall short in effectively handling the ingredients, particularly new and diverse ones. In response to these limitations, we introduce OVFoodSeg, a framework that adopts an open-vocabulary setting and enhances text embeddings with visual context. By integrating vision-language models (VLMs), our approach enriches text embedding with image-specific infor-mation through two innovative modules, e.g., an image-to-text learner FoodLearner and an Image-Informed Text Encoder. The training process of OVFoodSeg is divided into two stages: the pre-training of FoodLearner and the sub-sequent learning phase for segmentation. The pre-training phase equips FoodLearner with the capability to align visual information with corresponding textual representations that are specifically related to food, while the second phase adapts both the FoodLearner and the Image-Informed Text Encoder for the segmentation task. By addressing the de-ficiencies of previous models, OVFoodSeg demonstrates a significant improvement, achieving an 4.9% increase in mean Intersection over Union (mIoU) on the FoodSeg103 dataset, setting a new milestone for food image segmentation.
Xiongwei Wu, Sicheng Yu, Ee-Peng Lim, Chong-Wah Ngo
CVPR1
2023 A novel cooperative path planning method based on UCR-FCE and behavior regulation for large-scale multi-robot system
Wei Tang 0003, Jingxi Zhang, Xiongwei Wu
Appl. Intell.5
2022 Class Re-Activation Maps for Weakly-Supervised Semantic Segmentation
abstract
Extracting class activation maps (CAM) is arguably the most standard step of generating pseudo masks for weakly-supervised semantic segmentation (WSSS). Yet, we find that the crux of the unsatisfactory pseudo masks is the binary crossentropy loss (BCE) widely used in CAM. Specifically, due to the sum-over-class pooling nature of BCE, each pixel in CAM may be responsive to multiple classes co-occurring in the same receptive field. As a result, given a class, its hot CAM pixels may wrongly invade the area belonging to other classes, or the non-hot ones may be actually a part of the class. To this end, we introduce an embarrassingly simple yet surprisingly effective method: Reactivating the converged CAM with BCE by using softmax crossentropy loss (SCE), dubbed ReCAM. Given an image, we use CAM to extract the feature pixels of each single class, and use them with the class label to learn another fully-connected layer (after the backbone) with SCE. Once converged, we extract ReCAM in the same way as in CAM. Thanks to the contrastive nature of SCE, the pixel response is disentangled into different classes and hence less mask ambiguity is expected. The evaluation on both PASCAL VOC and MS COCO shows that ReCAM not only generates high-quality masks, but also supports plug-and-play in any CAM variant with little overhead. Our code is public at https://github.com/zhaozhengChenIReCAM.
Zhaozheng Chen, Xiongwei Wu, Xian-Sheng Hua 0001, Hanwang Zhang, Qianru Sun
CVPR3
2022 A Multi-Attention Tensor Completion Network for Spatiotemporal Traffic Data Imputation
abstract
The widespread deployment of road sensors in the Internet of Things (IoT) allows for fine-grained data integration, which is a fundamental demand for data-driven applications. Sensing data with inevitable missing and substantial anomalies are unavoidable, due to unstable network communication, faulty sensors, etc. Recent tensor completion studies have demonstrated the superiority of deep learning in imputation tasks by precisely capturing the intricate spatiotemporal dependencies/correlations. However, ignoring the significance of initial interpolation in these methods results in unstable performance, especially for complicated missing scenarios across large-scale data. Additionally, the existing interpolation methods utilize recursive signal propagation along spatiotemporal dimensions, which produce noise accumulation where the dependencies are uncorrelated. In this study, we design a multiattention tensor completion network (MATCN) for modeling multidimensional representation in the presence of missing entries. MATCN sparsely sampled historical fragments and utilized a gated diffusion convolution layer to generate the initial schemes, which mitigate the exposure bias existing in previous traffic imputation models. In addition, we develop a spatial signal propagation module and a temporal self-attention module as the basic stack block of deep networks, which executes representation aggregation and dynamic dependencies extraction at the spatiotemporal level. This architecture empowers MATCN with progressive completion capacities for complex data missing scenarios. Numerical experiments on four real-world traffic data sets with various missing scenarios demonstrate the superiority of MATCN over multiple state-of-the-art imputation baselines.
Xuesong Wu 0004, Mengyun Xu, Xiongwei Wu
IEEE Internet Things J.4
2021 PolarNet: Learning to Optimize Polar Keypoints for Keypoint Based Object Detection
Xiongwei Wu, Doyen Sahoo, Steven C. H. Hoi
ICLR1
2021 A Large-Scale Benchmark for Food Image Segmentation
abstract
Food image segmentation is a critical and indispensible task for developing health-related applications such as estimating food calories and nutrients. Existing food image segmentation models are underperforming due to two reasons: (1) there is a lack of high quality food image datasets with fine-grained ingredient labels and pixel-wise location masks---the existing datasets either carry coarse ingredient labels or are small in size; and (2) the complex appearance of food makes it difficult to localize and recognize ingredients in food images, e.g., the ingredients may overlap one another in the same image, and the identical ingredient may appear distinctly in different food images.
Xiongwei Wu, Ying Liu 0004, Ee-Peng Lim, Steven C. H. Hoi, Qianru Sun
ACM Multimedia1
2021 Caching Transient Content for IoT Sensing: Multi-Agent Soft Actor-Critic
abstract
Edge nodes (ENs) in Internet of Things commonly serve as gateways to cache sensing data while providing accessing services for data consumers. This paper considers multiple ENs that cache sensing data under the coordination of the cloud. Particularly, each EN can fetch content generated by sensors within its coverage, which can be uploaded to the cloud via fronthaul and then be delivered to other ENs beyond the communication range. However, sensing data are usually transient with time whereas frequent cache updates could lead to considerable energy consumption at sensors and fronthaul traffic loads. Therefore, we adopt Age of Information to evaluate data freshness and investigate intelligent caching policies to preserve data freshness while reducing cache update costs. Specifically, we model the cache update problem as a cooperative multi-agent Markov decision process with the goal of minimizing the long-term average weighted cost. To efficiently handle the exponentially large number of actions, we devise a novel reinforcement learning approach, which is a discrete multi-agent variant of soft actor-critic (SAC). Furthermore, we generalize the proposed approach into a decentralized control, where each EN can make decisions based on local observations only. Simulation results demonstrate the superior performance of the proposed SAC-based caching schemes.
Xiongwei Wu, Xiuhua Li 0001, Jun Li 0004, Pak-Chung Ching, Victor C. M. Leung, H. Vincent Poor
IEEE Trans. Commun.1
2021 Multi-Agent Reinforcement Learning for Cooperative Coded Caching via Homotopy Optimization
abstract
Introducing cooperative coded caching into small cell networks is a promising approach to reducing traffic loads. By encoding content via maximum distance separable (MDS) codes, coded fragments can be collectively cached at small-cell base stations (SBSs) to enhance caching efficiency. However, content popularity is usually time-varying and unknown in practice. As a result, cached content is anticipated to be intelligently updated by taking into account limited caching storage and interactive impacts among SBSs. In response to these challenges, we propose a multi-agent deep reinforcement learning (DRL) framework to intelligently update cached content in dynamic environments. With the goal of minimizing long-term expected fronthaul traffic loads, we first model dynamic coded caching as a cooperative multi-agent Markov decision process. Owing to the use of MDS coding, the resulting decision-making falls into a class of constrained reinforcement learning problems with continuous decision variables. To deal with this difficulty, we custom-build a novel DRL algorithm by embedding homotopy optimization into a deep deterministic policy gradient formalism. Next, to empower the caching framework with an effective trade-off between complexity and performance, we propose centralized, and partially and fully decentralized caching controls by applying the derived DRL approach. Simulation results demonstrate the superior performance of the proposed multi-agent framework.
Xiongwei Wu, Jun Li 0004, Ming Xiao 0001, Pak-Chung Ching, H. Vincent Poor
IEEE Trans. Wirel. Commun.1
2020 Deep Reinforcement Learning for IoT Networks: Age of Information and Energy Cost Tradeoff
abstract
In most Internet of Things (IoT) networks, edge nodes are commonly used as to relays to cache sensing data generated by IoT sensors as well as provide communication services for data consumers. However, a critical issue of IoT sensing is that data are usually transient, which necessitates temporal updates of caching content items while frequent cache updates could lead to considerable energy cost and challenge the lifetime of IoT sensors. To address this issue, we adopt the Age of Information (AoI) to quantity data freshness and propose an online cache update scheme to obtain an effective tradeoff between the average AoI and energy cost. Specifically, we first develop a characterization of transmission energy consumption at IoT sensors by incorporating a successful transmission condition. Then, we model cache updating as a Markov decision process to minimize average weighted cost with judicious definitions of state, action, and reward. Since user preference towards content items is usually unknown and often temporally evolving, we therefore develop a deep reinforcement learning (DRL) algorithm to enable intelligent cache updates. Through trial-and-error explorations, an effective caching policy can be learned without requiring exact knowledge of content popularity. Simulation results demonstrate the superiority of the proposed framework.
Xiongwei Wu, Xiuhua Li 0001, Jun Li 0004, Pak-Chung Ching, H. Vincent Poor
GLOBECOM1
2020 Latency-Minimized Design of secure transmissions in UAV-Aided Communications
abstract
Unmanned aerial vehicles (UAVs) can be utilized as aerial base stations to provide communication service for remote mobile users due to their high mobility and flexible deployment. However, the line-of-sight (LoS) wireless links are vulnerable to be intercepted by the eavesdropper (Eve), which presents a major challenge for UAV-aided communications. In this paper, we propose a latency-minimized transmission scheme for satisfying legitimate users' (LUs') content requests securely against Eve. By leveraging physical-layer security (PLS) techniques, we formulate a transmission latency minimization problem by jointly optimizing the UAV trajectory and user association. The resulting problem is a mixed-integer nonlinear program (MINLP), which is known to be NP hard. Furthermore, the dimension of optimization variables is indeterminate, which again makes our problem very challenging. To efficiently address this, we utilize bisection to search for the minimum transmission delay and introduce a variational penalty method to address the associated subproblem via an inexact block coordinate descent approach. Moreover, we present a characterization for the optimal solution. Simulation results are provided to demonstrate the superior performance of the proposed design.
Xiongwei Wu, Qiang Li 0017, Yawei Lu, H. Vincent Poor, Victor C. M. Leung, Pak-Chung Ching
ICASSP1
2020 Task Offloading for Automatic Speech Recognition in Edge-Cloud Computing Based Mobile Networks
abstract
Explosively increasing multimedia services and applications, e.g., automatic speech recognition (ASR), have aggravated the burden on the cloud server in mobile networks. To address the challenge, mobile edge computing has emerged for partially alleviating the workload of the cloud server and enhancing the quality of service of mobile users. In this paper, we aim to employ the technique of edge-cloud computing to accelerate the processing of ASR tasks generated by users in mobile networks. Particularly, we deploy a convolutional neural network based encoder in each edge server to extract features of the audio data. Based on certain network constraints (i.e., user association and edge servers’ storage/computing capacity), we propose a low-complexity and distributed iterative greedy method to address the formulated nonlinear mixed-integer nonconvex optimization problem. Simulation results demonstrate the effectiveness of the proposed scheme on reducing the total delay in the network.
Shitong Cheng, Zhenghui Xu, Xiuhua Li 0001, Xiongwei Wu, Qilin Fan, Xiaofei Wang 0001, Victor C. M. Leung
ISCC4
2020 Meta-RCNN: Meta Learning for Few-Shot Object Detection
abstract
Despite significant advances in deep learning based object detection in recent years, training effective detectors in a small data regime remains an open challenge. This is very important since labelling training data for object detection is often very expensive and time-consuming. In this paper, we investigate the problem of few-shot object detection, where a detector has access to only limited amounts of annotated data. Based on the meta-learning principle, we propose a new meta-learning framework for object detection named "Meta-RCNN", which learns the ability to perform few-shot detection via meta-learning. Specifically, Meta-RCNN learns an object detector in an episodic learning paradigm on the (meta) training data. This learning scheme helps acquire a prior which enables Meta-RCNN to do few-shot detection on novel tasks. Built on top of the popular Faster RCNN detector, in Meta-RCNN, both the Region Proposal Network (RPN) and the object classification branch are meta-learned. The meta-trained RPN learns to provide class-specific proposals, while the object classifier learns to do few-shot classification. The novel loss objectives and learning strategy of Meta-RCNN can be trained in an end-to-end manner. We demonstrate the effectiveness of Meta-RCNN in few-shot detection on three datasets (Pascal-VOC, ImageNet-LOC and MSCOCO) with promising results.
Xiongwei Wu, Doyen Sahoo, Steven C. H. Hoi
ACM Multimedia1
2020 Recent advances in deep learning for object detection
Xiongwei Wu, Doyen Sahoo, Steven C. H. Hoi
Neurocomputing1
2020 Single-shot bidirectional pyramid networks for high-quality object detection
Xiongwei Wu, Doyen Sahoo, Daoxin Zhang, Jianke Zhu, Steven C. H. Hoi
Neurocomputing1
2020 Feature agglomeration networks for single stage face detection
Xiongwei Wu, Steven C. H. Hoi, Jianke Zhu
Neurocomputing2
2020 Joint Long-Term Cache Updating and Short-Term Content Delivery in Cloud-Based Small Cell Networks
abstract
Explosive growth of mobile data demand may impose a heavy traffic burden on fronthaul links of cloud-based small cell networks (C-SCNs), which deteriorates users' quality of service (QoS) and requires substantial power consumption. This paper proposes an efficient maximum distance separable (MDS) coded caching framework for a cache-enabled C-SCNs, aiming at reducing long-term power consumption while satisfying users' QoS requirements in short-term transmissions. To achieve this goal, the cache resource in small-cell base stations (SBSs) needs to be reasonably updated by taking into account users' content preferences, SBS collaboration, and characteristics of wireless links. Specifically, without assuming any prior knowledge of content popularity, we formulate a mixed timescale problem to jointly optimize cache updating, multicast beamformers in fronthaul and edge links, and SBS clustering. Nevertheless, this problem is anti-causal because an optimal cache updating policy depends on future content requests and channel state information. To handle it, by properly leveraging historical observations, we propose a two-stage updating scheme by using Frobenius-Norm penalty and inexact block coordinate descent method. Furthermore, we derive a learning-based design, which can obtain effective trade-off between accuracy and computational complexity. Simulation results demonstrate the effectiveness of the proposed two-stage framework.
Xiongwei Wu, Qiang Li 0017, Xiuhua Li 0001, Victor C. M. Leung, Pak-Chung Ching
IEEE Trans. Commun.1
2019 Latency Driven Fronthaul Bandwidth Allocation and Cooperative Beamforming for Cache-enabled Cloud-based Small Cell Networks
abstract
This paper considers content delivery of the cache-enabled small cell networks (C-SCNs), where users with the same request form a multicast group and are served by a cluster of small-cell base stations (SBSs) under the coordination of the central processor. The performance of such a coordination is severely limited by the fronthaul link, which may be saturated and degrade quality of service (QoS). To improve user QoS, we propose a latency driven scheme by jointly optimizing fronthaul bandwidth allocation, multicast beamforming, and BS clustering. Accordingly, with min-max fairness among multicast groups, a latency minimization problem is formulated under the constraints of fronthaul bandwidth and transmission power. The resultant problem is a mixed-integer nonlinear program, which is NP-hard. To address such a complex problem, a quadratic penalty-based algorithm is proposed by using a reformulation of binary constraint. Meanwhile, we present the necessary condition for an optimal solution, which shows that fronthaul bandwidth allocation is inherently adaptive to cached contents and patterns of BS cooperation. Finally, simulation results demonstrate that the proposed scheme can effectively reduce latency under different caching strategies.
Xiongwei Wu, Xiuhua Li 0001, Qiang Li 0017, Victor C. M. Leung, Pak-Chung Ching
ICASSP1
2019 Joint Long-Term Cache Allocation and Short-Term Content Delivery in Green Cloud Small Cell Networks
abstract
Recent years have witnessed an exponential growth of mobile data traffic, which may lead to a serious traffic burn on the wireless networks and considerable power consumption. Network densification and edge caching are effective approaches to addressing these challenges. In this study, we investigate joint long-term cache allocation and short-term content delivery in cloud small cell networks (C-SCNs), where multiple small-cell BSs (SBSs) are connected to the central processor via fronthaul and can store popular contents so as to reduce the duplicated transmissions in networks. Accordingly, a long-term power minimization problem is formulated by jointly optimizing multicast beamforming, BS clustering, and cache allocation under quality of service (QoS) and storage constraints. The resultant mixed timescale design problem is an anticausal problem because the optimal cache allocation depends on the future file requests. To handle it, a two-stage optimization scheme is proposed by utilizing historical knowledge of users' requests and channel state information. Specifically, the online content delivery design is tackled with a penalty-based approach, and the periodic cache updating is optimized with a distributed alternating method. Simulation results indicate that the proposed scheme significantly outperforms conventional schemes and performs extremely close to a genie-aided lower bound in the low caching region.
Xiongwei Wu, Qiang Li 0017, Xiuhua Li 0001, Victor C. M. Leung, Pak-Chung Ching
ICC1
2019 FoodAI: Food Image Recognition via Deep Learning for Smart Food Logging
abstract
An important aspect of health monitoring is effective logging of food consumption. This can help management of diet-related diseases like obesity, diabetes, and even cardiovascular diseases. Moreover, food logging can help fitness enthusiasts, and people who wanting to achieve a target weight. However, food-logging is cumbersome, and requires not only taking additional effort to note down the food item consumed regularly, but also sufficient knowledge of the food item consumed (which is difficult due to the availability of a wide variety of cuisines). With increasing reliance on smart devices, we exploit the convenience offered through the use of smart phones and propose a smart-food logging system: FoodAI, which offers state-of-the-art deep-learning based image recognition capabilities. FoodAI has been developed in Singapore and is particularly focused on food items commonly consumed in Singapore. FoodAI models were trained on a corpus of 400,000 food images from 756 different classes.
Doyen Sahoo, Hao Wang 0094, Shu Ke, Xiongwei Wu, Hung Le 0003, Palakorn Achananuparp, Ee-Peng Lim, Steven C. H. Hoi
KDD4
2019 Joint Fronthaul Multicast and Cooperative Beamforming for Cache-Enabled Cloud-Based Small Cell Networks: An MDS Codes-Aided Approach
abstract
The performance of cloud-based small cell networks (C-SCNs) relies highly on a capacity-limited fronthaul, which degrade quality of service when it is saturated. Coded caching is a promising approach to addressing these challenges, as it provides abundant opportunities for fronthaul multicast and cooperative transmissions. This paper investigates cache-enabled C-SCNs, in which small-cell base stations (SBSs) are connected to the central processor via fronthaul, and can prefetch popular contents by applying maximum distance separable (MDS) codes. To fully capture the benefits of fronthaul multicast and cooperative transmissions, an MDS codes-aided transmission scheme is first proposed. We formulate the problem to minimize the content delivery latency by jointly optimizing fronthaul bandwidth allocation, SBS clustering, and beamforming. To efficiently solve the resulting nonlinear integer programming problem, we propose a penalty-based design by leveraging variational reformulations of binary constraints. To improve the solution of the penalty-based design, a greedy SBS clustering design is also developed. Furthermore, closed-form characterization of the optimal solution is obtained, through which the benefits of MDS codes can be quantified. The simulation results are given to demonstrate the significant benefits of the proposed MDS codes-aided transmission scheme.
Xiongwei Wu, Qiang Li 0017, Victor C. M. Leung, Pak-Chung Ching
IEEE Trans. Wirel. Commun.1
2018 Content Delivery Design for Cache-Aided Cloud Radio Access Network to Achieve Low Latency
abstract
In this paper, we examine the content delivery design for a cache-aided cloud radio access network (CA-CRAN), where users are served by multiple base stations (BSs) that are connected to cloud processor via fronthaul link. We propose a unified framework for cooperative delivery, which aims to minimize the total latency in the network. With fairness among users and physical-layer transmission, beamformers and content assignment are jointly optimized to fully exploit the benefits of caching. To address the resulting mixed binary nonconvex problem, a successive convex approximation (SCA)-based algorithm is derived with low complexity. Through simulations, the proposed design reduces latency significantly compared with existing work.
Xiongwei Wu, Pak-Chung Ching
ICASSP1
2018 Three-User Mimo Broadcast Channel with Delayed Csit: A Higher Achievable DoF
abstract
Degrees of freedom (DoF) of the three-user multiple-input multiple-output (MIMO) broadcast channel (BC) with delayed CSIT was derived for most antenna configurations except for the case of , where transmitter has M antennas and each receiver has N antennas. In this paper, for that problem, we propose an effective scheme for acquiring a higher achievable DoF than the value via existing methods. In the initial transmission phase, we transmit more data symbols than the amount that the receivers can instantaneously decode. Then, we generate auxiliary symbols for decoding the data symbols. Specifically, our scheme introduces an integrated design for the generation of auxiliary symbols. As a result, a higher achievable DoF, i.e., [12MN/(7M+2N)], can be achieved for specific antenna configurations, where .2N <; M <; 2.5N.
Tong Zhang 0026, Xiongwei Wu, Yinfei Xu, Yao Ge 0001, Pak-Chung Ching
ICASSP2