VLDB 2026 Research / reviewers in the wild / expert
Xianbin Cao 0001
dblp:22/3485 · also Xian-Bin Cao 0001
· DBLP profile ↗
139ranked-venue papers
18as first author
65since 2021 · last 2026
0000-0002-5042-7884ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 50 · 7 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 45 · 5 first-author · 18 since 2021Computer networks · 35 · 2 first-author · 25 since 2021Applied, interdisciplinary, general and emerging computing · 17 · 2 first-author · 8 since 2021Security and privacy · 6 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 4Databases, data management, data science and information retrieval · 3 · 1 first-authorSystems, architecture and hardware · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Vision-Augmented LLM for Communication Beam Steering Compensation
Dingyi Lu, Peng Yang 0009, Zehui Xiong, Xianbin Cao 0001, Tony Q. S. Quek |
WCNC | 5 |
| 2026 | Noise-Robust tiny object localization with flows
Huixin Sun, Linlin Yang 0001, Ronyu Chen, Kerui Gu, Baochang Zhang 0001, Angela Yao, Xianbin Cao 0001 |
Pattern Recognit. | 7 |
| 2026 | Industrial Scene Gas Leakage Detection: A Cross-Attention Based Multimodal Feature Difference Network and a New BenchmarkabstractIndustrial gas leakage detection is critically important for safety and environmental protection. While infrared imaging enables detection of invisible gases, two challenges remain: existing datasets lack realistic industrial scenarios, and current methods struggle to distinguish gas plumes from background interferences or segment discontinuous gas distributions. This paper introduces a benchmark comprising an Industrial RGB-Thermal Dataset (IRTD) with gas emission and leakage data from laboratory and industrial sites. A VLM-assisted RGBThermal detection framework with a Cross-Attention based Feature Difference (CAFD) module is designed to enhance gasspecific feature differentiation by computing inter-modal feature discrepancies. Evaluations on public datasets and IRTD demonstrate state-of-the-art results. Linlin Yang 0001, Xingyu Guo, Sheng Xu 0007, Xianbin Cao 0001, Baochang Zhang 0001 |
IEEE Signal Process. Lett. | 6 |
| 2026 | Adaptive Subarray Segmentation: A New Paradigm of Spatial Non-Stationary Near-Field Channel Estimation for XL-MIMO SystemsabstractTo address the complexities of spatial non-stationary (SnS) effects and spherical wave propagation in near-field channel estimation (CE) for extremely large-scale multiple-input multiple-output (XL-MIMO) systems, this paper proposes an SnS-aware CE framework based on adaptive subarray partitioning. We first investigate spherical wave propagation and various SnS characteristics and construct an SnS near-field channel model for XL-MIMO systems. Due to the limitations of uniform subarray patterns in capturing SnS, we analyze the adverse effects of the non-ideal array segmentation (over- and under-segmentation) on CE accuracy. To counter these issues, we develop a dynamic hybrid beamforming-assisted power-based subarray segmentation paradigm (DHBF-PSSP), which integrates power measurements with a dynamic hybrid beamforming structure to enable joint subarray partitioning and decoupling. A power-adaptive subarray segmentation (PASS) algorithm leverages the statistical properties of power profiles, while subarray decoupling is achieved via a subarray segmentation-based sampling method (SS-SM) under radio frequency (RF) chain constraints. For subarray CE, we propose a subarray segmentation-based assorted block sparse Bayesian learning algorithm under the multiple measurement vectors framework (SS-ABSBL-MMV). This algorithm exploits angular-domain block sparsity under a discrete Fourier transform (DFT) codebook and inter-subcarrier structured sparsity. Simulation results confirm that the proposed framework outperforms existing methods in CE performance. Shuhang Yang, Puguang An, Peng Yang 0009, Xianbin Cao 0001, Dapeng Oliver Wu, Tony Q. S. Quek |
IEEE Trans. Commun. | 4 |
| 2026 | Toward Semantic-Aware Aerial Video Anomaly Detection by Exploiting Multimodal Large Language ModelabstractDrones have become increasingly widely applied in surveillance systems due to their mobility, making aerial video anomaly detection methods more crucial. Anomalies in aerial videos often present as semantic conflicts, such as the presence of unexpected objects or unusual behaviors that do not align with the context. Previous approaches often relied on manually crafted knowledge graphs to detect such conflicts, which suffer from poor scalability. Recently, owing to their sufficient alignment training, multimodal large language models (MLLMs) have emerged as a generalized solution for semantic understanding. However, the direct application of MLLMs does not yield satisfactory anomaly detection performance in aerial videos. First, aerial videos often manifest platform-induced pseudo-motion, which obscures the true motion of objects and exacerbates detection errors. Second, without sufficient labeled data for fine-tuning, generic MLLMs often lack scene-level semantic guidance to reliably distinguish abnormal events that include contextually inappropriate behaviors. To address these challenges, we propose SemAero, an MLLM-based framework to address these challenges by: 1) designing an ego-motion reduction module to enhance model perception on object movement, 2) generating scene-specific prompts adaptively with step-by-step guidance for reasonable output, and 3) refining scores with dual-stream consistent feature for better domain-specific anomaly detection. Evaluated across 8 diverse aerial scenes and 73 sub-datasets, SemAero achieves a 3.09% improvement in AUC-ROC over the second-best model, demonstrating its ability in aerial video anomaly detection. Ruoheng Li, Xuhui Liu, Yutao Hu 0002, Xianbin Cao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | SpectraBayes: Exploring Traffic Series Reconstruction in Frequency Domain for Anomaly DetectionabstractTimely detection of anomalies in traffic systems is crucial for mitigating risks and economic losses. Current time series anomaly detection methods often use reconstruction errors to identify anomalies, and their accuracy depends on how well they can reconstruct normal patterns from the original series. However, traffic data typically exhibit high noise, spatiotemporal heterogeneity, and uncertain inter-series correlations, complicating the learning of normal patterns. To address these challenges, we introduce SpectraBayes, which explores the reconstruction of density and volume series in the frequency domain for anomaly detection. First, we transform the series into the frequency domain and apply a low-pass filter to remove noise. Then, we embed periodic information into the frequency-domain representation through phase shifts to enhance the temporal awareness. Additionally, we model the inter-series correlations between density and volume resiliently using cross-spectrum probabilistic modeling. Optimized by maximizing the Evidence Lower Bound (ELBO), SpectraBayes ensures robust reconstruction while avoiding overfitting against the uncertain data. SpectraBayes outperforms 21 existing anomaly detection models on traffic series anomaly detection tasks, achieving mean improvements of 2.71% across three metrics over the second-best model. Furthermore, it is lightweight and maintains robust performance under varying noise levels. Ruoheng Li, Diyin Tang, Fei Wang 0014, Xianbin Cao 0001 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2025 | SET: Spectral Enhancement for Tiny Object DetectionabstractDeep learning has significantly advanced the object detection field. However, tiny object detection (TOD) remains a challenging problem. We provide a new analysis method to examine the TOD challenge through occlusion-based attribution analysis in the frequency domain. We observe that tiny objects become less distinct after feature encoding and can benefit from the removal of high-frequency information. In this paper, we propose a novel approach named Spectral Enhancement for Tiny object detection (SET), which amplifies the frequency signatures of tiny objects in a heterogeneous architecture. SET includes two modules. The Hierarchical Background Smoothing (HBS) module suppresses high-frequency noise in the background through adaptive smoothing operations. The Adversarial Perturbation Injection (API) module leverages adversarial perturbations to increase feature saliency in critical regions and prompt the refinement of object features during training. Extensive experiments on four datasets demonstrate the effectiveness of our method. Especially, SET boosts the prior art RFLA by 3.2% AP on the AI-TOD dataset. Huixin Sun, Runqi Wang, Yanjing Li, Linlin Yang 0001, Shaohui Lin, Xianbin Cao 0001, Baochang Zhang 0001 |
CVPR | 6 |
| 2025 | Uncertainty-Aware Gradient Stabilization for Small Object Detection
Huixin Sun, Yanjing Li, Linlin Yang 0001, Xianbin Cao 0001, Baochang Zhang 0001 |
ICCV | 4 |
| 2025 | Prompt as Knowledge Bank: Boost Vision-language model via Structural Representation for zero-shot medical detectionabstractZero-shot medical detection can further improve detection performance without relying on annotated medical images even upon the fine-tuned model, showing great clinical value. Recent studies leverage grounded vision-language models (GLIP) to achieve this by using detailed disease descriptions as prompts for the target disease name during the inference phase.
However, these methods typically treat prompts as equivalent context to the target name, making it difficult to assign specific disease knowledge based on visual information, leading to a coarse alignment between images and target descriptions. In this paper, we propose StructuralGLIP, which introduces an auxiliary branch to encode prompts into a latent knowledge bank layer-by-layer, enabling more context-aware and fine-grained alignment. Specifically, in each layer, we select highly similar features from both the image representation and the knowledge bank, forming structural representations that capture nuanced relationships between image patches and target descriptions. These features are then fused across modalities to further enhance detection performance.
Extensive experiments demonstrate that StructuralGLIP achieves a +4.1\% AP improvement over prior state-of-the-art methods across seven zero-shot medical detection benchmarks, and consistently improves fine-tuned models by +3.2\% AP on endoscopy image datasets. Yuguang Yang 0007, Tongfei Chen, Linlin Yang 0001, Chunyu Xie, Dawei Leng, Xianbin Cao 0001, Baochang Zhang 0001 |
ICLR | 7 |
| 2025 | Efficient Low-Bit Quantization with Adaptive Scales for Multi-Task Co-TrainingabstractCo-training can achieve parameter-efficient multi-task models but remains unexplored for quantization-aware training. Our investigation shows that directly introducing co-training into existing quantization-aware training (QAT) methods results in significant performance degradation. Our experimental study identifies that the primary issue with existing QAT methods stems from the inadequate activation quantization scales for the co-training framework. To address this issue, we propose Task-Specific Scales Quantization for Multi-Task Co-Training (TSQ-MTC) to tackle mismatched quantization scales. Specifically, a task-specific learnable multi-scale activation quantizer (TLMAQ) is incorporated to enrich the representational ability of shared features for different tasks. Additionally, we find that in the deeper layers of the Transformer model, the quantized network suffers from information distortion within the attention quantizer. A structure-based layer-by-layer distillation (SLLD) is then introduced to ensure that the quantized features effectively preserve the information from their full-precision counterparts. Our extensive experiments in two co-training scenarios demonstrate the effectiveness and versatility of TSQ-MTC. In particular, we successfully achieve a 4-bit quantized low-level visual foundation model based on IPT, which attains a PSNR comparable to the full-precision model while offering a $7.99\times$ compression ratio in the $\times4$ super-resolution task on the Set5 benchmark. Linlin Yang 0001, Yanjing Li, Guodong Guo, Xianbin Cao 0001, Baochang Zhang 0001 |
ICLR | 6 |
| 2025 | Hyperchaos and HVS-Adaptive Video Watermarking EmbeddingabstractWatermarking for video playback authorization faces the classic challenge of balancing imperceptibility and robustness, while also maintaining resilience against statistical attacks. This paper introduces a novel scheme that integrates hyperchaos, a human visual system (HVS) model, and asymmetric modulation to address these challenges. First, a four-dimensional hyperchaotic system is constructed to achieve triple dynamic randomization of the watermark information, embedding locations, and embedding strength, thereby enhancing security. Guided by an HVS-based just noticeable distortion (JND) model, a spatio-temporally adaptive embedding strength is then derived, maximizing robustness under strict imperceptibility constraints. Furthermore, a blind extraction mechanism using coefficient-relation modulation is designed, inherently improving resilience against common video processing and malicious attacks. Collectively, these strategies unify imperceptibility, robustness, and security. The experimental results confirm the algorithm’s superior performance against benchmarks. It exhibits stronger resistance to statistical analysis, with a mean Kullback-Leibler (KL) divergence of only 0.003, and enhanced watermark robustness, shown by a 61.8% improvement in normalized correlation (NC). For imperceptibility, it achieves a gain in the peak signal-to-noise ratio (PSNR) over 2.2 dB. Kesong Wu, Maowei Li, Peng Yang 0009, Jiangtian Nie, Xianbin Cao 0001 |
TrustCom | 6 |
| 2025 | Multi-UAV Trajectory Generation for Fresh Data Collection: A Diffusion-based Reinforcement Learning ApproachabstractThis paper investigates the trajectory generation problem for multi-unmanned aerial vehicle (UAV)-enabled uplink data collection. Specifically, we minimize the age-of-information (AoI) and maximize the coverage as well as the amount of collected data by planning the multi- UAV trajectory considering the energy consumption and collisions constraints. Motivated by diffusion models' exceptional generative capabilities, we propose a multi-UAV trajectory generation (MUTG) solution based on soft actor-critic and diffusion to solve the optimization problem. A diffusion model-based predictor is designed to obtain the action policy, where a hierarchical graph-transformer network is developed to extract entities' interactive information as a conditional guide for the diffusion. Numerical results verify the effectiveness and superiority compared with benchmark schemes in terms of average AoI, user coverage and data collection ratio. Ziping Yu, Meng Xiao 0002, Zhongliang Zhao, Xianbin Cao 0001, Yang Liu 0003, Tony Q. S. Quek |
WCNC | 5 |
| 2025 | M3DP: Optimizing 2D vision tasks with minimal 3D object information
Yanjing Li, Linlin Yang 0001, Xinkai Liang, Xianbin Cao 0001, Qi Wang 0009, Baochang Zhang 0001 |
Neurocomputing | 5 |
| 2025 | DNN Task Assignment in UAV Networks: A Generative AI Enhanced Multiagent Reinforcement Learning Approachabstractuncrewed aerial vehicles (UAVs) offer high mobility and flexible deployment capabilities, making them ideal for Internet of Things (IoT) applications. However, the substantial amount of data generated by various applications within the existing low-altitude network requires processing through deep neural networks (DNN) on UAVs, which is challenging due to their limited computational resources. To address this issue, we propose a two-stage optimization method for flight path planning and task allocation based on a mother-child UAV swarm system. In the first stage, we employ a greedy algorithm to solve the path planning problem by considering the task size of the target area to be inspected and the shortest flight path as constraints. The goal is to minimize both the flight path of the UAV and the overall cost of the system. In the second stage, we introduce a novel DNN task assignment algorithm that combines multiagent deep deterministic policy gradient (MADDPG) and generative diffusion models (GDMs), named GDM-MADDPG. This algorithm takes advantage of the reverse denoising process of GDM to replace the actor network in MADDPG. It enables UAVs to generate specific DNN task assignment actions based on agents’ observations in a dynamic environment, thereby improving the efficiency of task assignment and overall system performance. The simulation results demonstrate that our algorithm outperforms the benchmarks in terms of path planning, Age of Information (AoI), task completion rate, and system utility, demonstrating its effectiveness. Qian Chen 0019, Wenjie Weng, Binhan Liao, Jiacheng Wang 0001, Xianbin Cao 0001, Xiaohuan Li 0001 |
IEEE Internet Things J. | 6 |
| 2025 | Engineering and technology for low-altitude economy infrastructure
Harry Shum, Xianbin Cao 0001, Mark Hansen |
Frontiers Inf. Technol. Electron. Eng. | 3 |
| 2025 | Coevolutionary genetic programming for large-scale dynamic multi-aircraft task allocationabstractMulti-aircraft task allocation (MATA) plays a vital role in improving mission efficiency under dynamic conditions. This paper proposes a novel coevolutionary genetic programming (CoGP) framework that automatically designs high-performance reactive heuristics for dynamic MATA problems. Unlike conventional single-tree genetic programming (GP) methods, CoGP jointly develops two interacting populations, i.e., task prioritizing heuristics and aircraft selection heuristics, to explicitly model the coupling between these two interdependent decision phases. A comprehensive terminal set is constructed to represent the dynamic states of aircraft and tasks, whereas a low-level heuristic template translates developed trees into executable allocation strategies. Extensive experiments on public benchmark instances simulating post-disaster emergency delivery demonstrate that CoGP achieves superior performance compared with state-of-the-art GP and heuristic methods, exhibiting strong adaptability, scalability, and real-time responsiveness in complex and dynamic rescue environments. Ce Yu, Xianbin Cao 0001, Wenbo Du 0001 |
Frontiers Inf. Technol. Electron. Eng. | 2 |
| 2025 | Connection Performance Modeling and Analysis of a Radiosonde Network in a TyphoonabstractThis paper is concerned with the theoretical modeling and analysis of uplink connection performance of a radiosonde network deployed in a typhoon. Similar to existing works, the stochastic geometry theory is leveraged to derive the expression of the uplink connection probability (CP) of a radiosonde. Nevertheless, existing works assume that network nodes are spherically or uniformly distributed. Different from the existing works, this paper investigates two particular motion patterns of radiosondes in a typhoon, which significantly challenges the theoretical analysis. According to their particular motion patterns, this paper first separately models the distributions of horizontal and vertical distances from a radiosonde to its receiver. Secondly, this paper derives the closed-form expressions of cumulative distribution function (CDF) and probability density function (PDF) of a radiosonde’s three-dimensional (3D) propagation distance to its receiver. Thirdly, this paper derives the analytical expression of the uplink CP for any radiosonde in the network. Finally, extensive numerical simulations are conducted to validate the theoretical analysis, and the influence of various network design parameters is comprehensively discussed. Simulation results show that when the signal-to-interference-noise ratio (SINR) threshold is below -35 dB, and the density of radiosondes remains under 0.01/km3, the uplink CP approaches 26%, 39%, and 50% in three patterns. Hanyi Liu, Xianbin Cao 0001, Peng Yang 0009, Zehui Xiong, Tony Q. S. Quek, Dapeng Oliver Wu |
IEEE Trans. Commun. | 2 |
| 2025 | EMOR: Energy-Efficient Mixture Opportunistic Routing Based on Reinforcement Learning for Lunar Surface Ad-Hoc NetworksabstractThe lunar surface ad-hoc network is a critical component of the international lunar research station and an extension of the earth-moon communication networks. Its high reliability and low delay are essential for ensuring the safety of the lunar station and improving the efficiency of node collaboration. However, due to the lack of large-scale grid infrastructures, the network must operate autonomously for long periods under strong energy constraints. We propose EMOR, a cross-layer routing protocol, which aims to achieve sustainable high reliability and low latency while balancing energy recovery and consumption. EMOR improves reliability through the “parallel” forwarding feature of opportunistic routing and reduces delay through a mixture of table-based and timer-based routing mechanisms. Moreover, EMOR uses reinforcement learning to analyze the environment and calculate the weights of energy and progress to guide the emphasis on multi-metrics routing. To balance energy consumption and recovery, EMOR introduces a dynamic duty cycle in the MAC layer. Compared to table-based routing and the latest opportunistic routing, EMOR maintains the optimal end-to-end delay in the order of 1ms while improving the packet delivery ratio 6% to 21% higher than other protocols. Moreover, the network lifetime using EMOR is extended by 75.5% to 242%. Zhiyuan Qu, Zhongliang Zhao, Xianbin Cao 0001, Yang Liu 0003, Tony Q. S. Quek |
IEEE Trans. Commun. | 4 |
| 2025 | Diffusion Self-Distillation for Remote Sensing Scene ClassificationabstractRemote sensing scene classification, a fundamental task in remote image analysis, has obtained rapid progress due to the powerful capabilities of Convolutional Neural Networks (CNNs). Achieving precise classification performance heavily relies on the feature extraction capacity of the network. However, due to the large variation and severe distortion within the images, extracting robust feature representations is necessary but challenging. Self-distillation could enhance the shallow layers by providing stronger gradients and more accurate supervision from deeper layers, thereby promoting the extraction of spatially detailed features. Nonetheless, due to the limited capacity of shallow layers to learn truly valuable knowledge, shallow layer features can be viewed as the noisy version of deep layer features and contain more disruptive factors, which significantly impedes the effectiveness of self-distillation. To address this issue, in this paper, we establish the Diffusion Self-Distillation Network (DSDNet), which incorporates the conditional diffusion denoising model into the self-distillation framework. Specifically, DSDNet filters noise from shallow features through the diffusion denoising process, enabling more precise and accurate distillation between the refined student features and the teacher features. Extensive experiments on four challenging remote sensing datasets emonstrate that the proposed DSDNet achieves significant performance improvements over various backbone networks with negligible increases in parameters, delivering state-of-the-art classification performance. Our code and dataset are available on https://github.com/toggle1995/DSDNet. Yutao Hu 0002, Lei Zhang 0001, Xiaoyan Luo, Xianbin Cao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | Hierarchical Self-Distilled Feature Learning for Fine-Grained Visual CategorizationabstractFine-grained visual categorization (FGVC) relies on hierarchical features extracted by deep convolutional neural networks (CNNs) to recognize closely alike objects. Particularly, shallow layer features containing rich spatial details are vital for specifying subtle differences between objects but are usually inadequately optimized due to gradient vanishing during backpropagation. In this article, hierarchical self-distillation (HSD) is introduced to generate well-optimized CNNs features for accurate fine-grained categorization. HSD inherits from the widely applied deep supervision and implements multiple intermediate losses for reinforced gradients. Besides that, we observe that the hard (one-hot) labels adopted for intermediate supervision hurt the performance of FGVC by enforcing overstrict supervision. As a solution, HSD seeks self-distillation where soft predictions generated by deeper layers of the network are hierarchically exploited to supervise shallow parts. Moreover, self-information entropy loss (SIELoss) is designed in HSD to adaptively soften intermediate predictions and facilitate better convergence. In addition, the gradient detached fusion (GDF) module is incorporated to produce an ensemble result with multiscale features via effective feature fusion. Extensive experiments on four challenging fine-grained datasets show that, with neglectable parameter increase, the proposed HSD framework and the GDF module both bring significant performance gains over different backbones, which also achieves state-of-the-art classification performance. Yutao Hu 0002, Xuhui Liu, Xiaoyan Luo, Yao Hu 0002, Xianbin Cao 0001, Baochang Zhang 0001, Jun Zhang 0007 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2024 | Bi-ViT: Pushing the Limit of Vision Transformer QuantizationabstractVision transformers (ViTs) quantization offers a promising prospect to facilitate deploying large pre-trained networks on resource-limited devices. Fully-binarized ViTs (Bi-ViT) that pushes the quantization of ViTs to its limit remain largely unexplored and a very challenging task yet, due to their unacceptable performance. Through extensive empirical analyses, we identify the severe drop in ViT binarization is caused by attention distortion in self-attention, which technically stems from the gradient vanishing and ranking disorder. To address these issues, we first introduce a learnable scaling factor to reactivate the vanished gradients and illustrate its effectiveness through theoretical and experimental analyses. We then propose a ranking-aware distillation method to rectify the disordered ranking in a teacher-student framework. Bi-ViT achieves significant improvements over popular DeiT and Swin backbones in terms of Top-1 accuracy and FLOPs. For example, with DeiT-Tiny and Swin-Tiny, our method significantly outperforms baselines by 22.1% and 21.4% respectively, while 61.5x and 56.1x theoretical acceleration in terms of FLOPs compared with real-valued counterparts on ImageNet. Our codes and models are attached on https://github.com/YanjingLi0202/Bi-ViT/ . Yanjing Li, Sheng Xu 0007, Mingbao Lin, Xianbin Cao 0001, Chuanjian Liu, Baochang Zhang 0001 |
AAAI | 4 |
| 2024 | Swarm-RE: Hierarchical Opportunistic Routing and Fast Terrain Exploration in Planetary SurfaceabstractWireless sensor networks (WSNs) can be applied to planetary surface exploration due to the advantages of large coverage areas, low cost, and all-time monitoring. However, obtaining sensor data stably and efficiently is difficult due to energy constraints and environmental interference. In this paper, we propose Swarm-Re, including hierarchical heterogeneous opportunistic routing (HHOR) for cooperative air-ground network communication and fast terrain exploration algorithm to collect information efficiently. The sensor nodes are clustered by the DBSCAN algorithm based on the estimated SNR and network connectivity. For cross-cluster communication, HHOR deploys UAV nodes to achieve obstacle crossing. Through the experiment of Swarm-RE, HHOR consumes the least energy and accomplishes a high data delivery rate compared with representative routing protocols. The HHOR protocol shows the lowest expected end-to-end delay and highest channel utilization in both intra-cluster and cross-cluster communications. Relying on HHOR and UAV deployment, the Fast exploration algorithm can improve the efficiency with an average of 11s per Number of UAV. The result shows that the Swarm-RE provides efficient routine and exploration ability in planetary surface. Ziping Yu, Jingxuan Chen, Zhongliang Zhao, Xianbin Cao 0001 |
ICC | 5 |
| 2024 | ROI-Aware Dynamic Network Quantization for Neural Video Compression
Baochang Zhang 0001, Xianbin Cao 0001 |
ICPR (5) | 3 |
| 2024 | MAFormer: A transformer network with multi-scale attention fusion for visual recognition
Huixin Sun, Baochang Zhang 0001, Xianbin Cao 0001, Errui Ding, Shumin Han |
Neurocomputing | 7 |
| 2024 | Enhancing AIoT Device Association With Task Offloading in Aerial MEC NetworksabstractUnmanned aerial vehicles (UAVs) have emerged as a promising solution for enhancing mobile-edge computing (MEC) networks. However, the integration of UAVs into MEC networks poses unique challenges, such as the presence of dynamic devices and complex resource allocation. This research investigates the problem of task offloading in a distributed MEC network with multiple ground and aerial base stations (UAV base stations). With a focus on the cost-sensitive nature of Internet of Things Devices (IoTDs), our objective is to maximize the Quality of Experience (QoE) in terms of average task response time and cache queue length in IoTDs by jointly optimizing device association, offloading decision, and UAV trajectory planning. To address the combinatorial and nonconvex nature of the problem, we propose an artificial intelligence (AI)-based optimization scheme. First, the association between IoTDs and stations is determined using a recursive selection and replacement transmission-rate-based (RSRT) algorithm. Subsequently, the offloading problem is formulated as a 0-1 Backpack Problem with variable value, for which we present a backtracking task offloading (BTO) algorithm. Additionally, we employ a multiagent deep deterministic policy gradient (MADDPG) approach to determine the trajectory planning of UAVs. Numerical results demonstrate the effectiveness of the proposed scheme in terms of reduction in average response time, and cache queue length in IoTDs within the MEC system when compared to benchmark schemes. Jingxuan Chen, Peng Yang 0009, Siqiao Ren, Zhongliang Zhao, Xianbin Cao 0001, Dapeng Oliver Wu |
IEEE Internet Things J. | 5 |
| 2024 | Learning Foreground Information Bottleneck for few-shot semantic segmentation
Yutao Hu 0002, Xiaoyan Luo, Jungong Han, Xianbin Cao 0001, Jun Zhang 0007 |
Pattern Recognit. | 5 |
| 2024 | A Spatial-Temporal Approach for Multi-Airport Traffic Flow Prediction Through Causality GraphsabstractAccurate airport traffic flow estimation is crucial for the secure and orderly operation of the aviation system. Recent advances in machine learning have achieved promising prediction results in the single-airport scenario. However, these works overlook the variational spatial interactions hidden among airports and show limited performances on the traffic flow prediction task for the aviation system which is composed of several airports. In this paper, we consider the multi-airport scenario and propose a novel spatio-temporal hybrid deep learning model to efficiently capture spatial correlations as well as temporal dependencies in a parallelized way. Specifically, we introduce the causal inference among airports to model their interactions and thus construct adaptive causality graphs in a data-driven manner to address the heterogeneity of airports. Furthermore, given that multi-source features are not applicable for all airports, a feature mask module is designated to adaptively select the features in spatial information mining. Extensive experiments are conducted on the real data of top-30 busiest airports in China. The results show that our spatio-temporal deep learning approach is superior to state-of-the-art methodologies and the improvement ratio is up to 4.7% against benchmarks. Ablation studies emphasize the power of the proposed adaptive causality graph and the feature mask module. All of these prove the effectiveness of the proposed methodology. Wenbo Du 0001, Shenwen Chen, Zhishuai Li, Xianbin Cao 0001 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2024 | Energy-Efficient URLLC Service Provision via a Near-Space Information NetworkabstractThe integration of a near-space information network (NSIN) with the reconfigurable intelligent surface (RIS) is envisioned to significantly enhance the communication performance of future wireless communication systems by proactively altering wireless channels. This paper investigates the problem of deploying a RIS-integrated NSIN to provide energy-efficient, ultra-reliable and low-latency communications (URLLC) services. We mathematically formulate this problem as a resource optimization problem, aiming to maximize the effective throughput and minimize the system power consumption, subject to URLLC and physical resource constraints. The formulated problem is challenging in terms of accurate channel estimation, RIS phase alignment, and effective solution design. We propose a joint resource allocation algorithm to handle these challenges. In this algorithm, we develop an accurate channel estimation approach by exploring message passing and optimize phase shifts of RIS reflecting elements to further increase the channel gain. Besides, we derive an analysis-friendly expression of decoding error probability and decompose the problem into two-layered optimization problems by analyzing the monotonicity, which makes the formulated problem analytically tractable. Extensive simulations have been conducted to verify the performance of the proposed algorithm. Simulation results show that the proposed algorithm can achieve outstanding channel estimation performance and is more energy-efficient than diverse benchmark algorithms. Puguang An, Peng Yang 0009, Xianbin Cao 0001, Kun Guo 0002, Yue Gao 0001, Tony Q. S. Quek |
IEEE Trans. Wirel. Commun. | 3 |
| 2024 | Joint 3D Deployment and Beamforming for RSMA-Enabled UAV Base Station With Geographic InformationabstractThis paper studies the joint three-dimensional (3D) deployment and beamforming problem for a rate-splitting multiple access (RSMA)-enabled unmanned aerial vehicle base station (UBS) assisted by geographic information. Specifically, we maximize the minimum achievable rate among users by optimizing the beamforming, rate allocation and UBS deployment considering the power and building blockages constraints. To solve the intractable problem, an alternating optimization scheme is proposed. In particular, we first split the formulated problem into three sub-problems of deployment region modeling, joint beamforming and rate allocation, and 3D UBS deployment. For the first sub-problem, we define the allowable deployment region with geographic information with the aim of ensuring line-of-sight connections between the UBS and users. The feasible region is expressed as tractable constraints via the Big-M method and penalty function method. For the other sub-problems, semi-definite programming and successive convex approximation are employed to design the joint beamforming and rate allocation, and UBS deployment, respectively. These two sub-problems are optimized iteratively until convergence. Finally, numerical results validate the superiority of our proposed solution in comparison with the benchmark schemes with regard to the minimum achievable rate. Meng Xiao 0002, Huanxi Cui, Zhongliang Zhao, Xianbin Cao 0001, Dapeng Oliver Wu |
IEEE Trans. Wirel. Commun. | 4 |
| 2023 | Representation Disparity-aware Distillation for 3D Object DetectionabstractIn this paper, we focus on developing knowledge distillation (KD) for compact 3D detectors. We observe that off-the-shelf KD methods manifest their efficacy only when the teacher model and student counterpart share similar intermediate feature representations. This might explain why they are less effective in building extreme-compact 3D detectors where significant representation disparity arises due primarily to the intrinsic sparsity and irregularity in 3D point clouds. This paper presents a novel representation disparity-aware distillation (RDD) method to address the representation disparity issue and reduce performance gap between compact students and over-parameterized teachers. This is accomplished by building our RDD from an innovative perspective of information bottleneck (IB), which can effectively minimize the disparity of proposal region pairs from student and teacher in features and logits. Extensive experiments are performed to demonstrate the superiority of our RDD over existing KD methods. For example, our RDD increases mAP of CP-Voxel-S to 57.1% on nuScenes dataset, which even surpasses teacher performance while taking up only 42% FLOPs. Yanjing Li, Sheng Xu 0007, Mingbao Lin, Jihao Yin, Baochang Zhang 0001, Xianbin Cao 0001 |
ICCV | 6 |
| 2023 | Q-DM: An Efficient Low-bit Quantized Diffusion ModelabstractDenoising diffusion generative models are capable of generating high-quality data, but suffers from the computation-costly generation process, due to a iterative noise estimation using full-precision networks. As an intuitive solution, quantization can significantly reduce the computational and memory consumption by low-bit parameters and operations. However, low-bit noise estimation networks in diffusion models (DMs) remain unexplored yet and perform much worse than the full-precision counterparts as observed in our experimental studies. In this paper, we first identify that the bottlenecks of low-bit quantized DMs come from a large distribution oscillation on activations and accumulated quantization error caused by the multi-step denoising process. To address these issues, we first develop a Timestep-aware Quantization (TaQ) method and a Noise-estimating Mimicking (NeM) scheme for low-bit quantized DMs (Q-DM) to effectively eliminate such oscillation and accumulated error respectively, leading to well-performed low-bit DMs. In this way, we propose an efficient Q-DM to calculate low-bit DMs by considering both training and inference process in the same framework. We evaluate our methods on popular DDPM and DDIM models. Extensive experimental results show that our method achieves a much better performance than the prior arts. For example, the 4-bit Q-DM theoretically accelerates the 1000-step DDPM by 7.8x and achieves a FID score of 5.17, on the unconditional CIFAR-10 dataset. Yanjing Li, Sheng Xu 0007, Xianbin Cao 0001, Baochang Zhang 0001 |
NeurIPS | 3 |
| 2023 | AOR: Adaptive opportunistic routing based on reinforcement learning for planetary surface exploration
Ziping Yu, Zhongliang Zhao, Xianbin Cao 0001 |
Comput. Commun. | 4 |
| 2023 | DCP-NAS: Discrepant Child-Parent Neural Architecture Search for 1-bit CNNs
Yanjing Li, Sheng Xu 0007, Xianbin Cao 0001, Lian Zhuo, Baochang Zhang 0001, Tian Wang 0002, Guodong Guo |
Int. J. Comput. Vis. | 3 |
| 2023 | Cooperative path planning optimization for multiple UAVs with communication constraints
Xianbin Cao 0001, Wenbo Du 0001 |
Knowl. Based Syst. | 2 |
| 2023 | Deep Reinforcement Learning Based Resource Allocation in Multi-UAV-Aided MEC NetworksabstractResource allocation for mobile edge computing (MEC) in unmanned aerial vehicle (UAV) networks has been a popular research issue. Different from existing works, this paper considers a multi-UAV-aided uplink communication scenario and investigates a resource allocation problem of minimizing the total system latency and the energy consumption, subject to constraints on transmit power of mobile users (MUs), system latency caused by transmission and computation. The problem is confirmed to be a challenging time-series mixed-integer non-convex programming problem, and we propose a joint UAV Movement control, MU Association and MU Power control (UMAP) algorithm to solve it effectively, where three sub-problems are optimized iteratively. Specifically, UAV movement and MU association are optimized utilizing deep reinforcement learning (DRL) to decrease the energy consumption and system latency. Next, a closed-form solution of the MU transmit power is derived. Finally, simulation results show that the UMAP algorithm can significantly decrease the system latency and energy consumption and increase the coverage rate compared with benchmark algorithms. Jingxuan Chen, Xianbin Cao 0001, Peng Yang 0009, Meng Xiao 0002, Siqiao Ren, Zhongliang Zhao, Dapeng Oliver Wu |
IEEE Trans. Commun. | 2 |
| 2023 | Network Topology Inference Based on Timing Meta-DataabstractA set of low-cost sensors is deployed to infer the network topology of a self-organizing wireless network. The sensors operate in a non-invasive fashion, extracting only the timings of data packets and acknowledgment (ACK) packets from all nodes in a network. The meta-data also reports the source node of each packet, but not the destination nodes or the contents of the packets. A central processor collects the meta-data from the sensors, and the goal is for the processor to infer the network topology based solely on such information. Prior work leveraged causality metrics to identify which links are active. If the data timings and ACK timings of two nodes– say node 1 and node 2, respectively– are causally related, this may be taken as evidence that node 1 is communicating to node 2 (which sends back ACK packets to node 1). This paper starts with the observation that packet losses can weaken the causality relationship between data and ACK timing streams. To obviate this problem, a new Expectation Maximization (EM)-based algorithm is introduced– EM-causality discovery algorithm (EM-CDA)– which treats packet losses as latent variables. EM-CDA iterates between the estimation of packet losses and the evaluation of causality metrics. The method is validated through extensive experiments in wireless sensor networks on the NS-3 simulation platform. Wenbo Du 0001, Tao Tan 0006, Haijun Zhang 0001, Xianbin Cao 0001, Osvaldo Simeone |
IEEE Trans. Commun. | 4 |
| 2023 | Boosting Variational Inference With Margin Learning for Few-Shot Scene-Adaptive Anomaly DetectionabstractAnomaly detection in surveillance videos aims to identify frames where abnormal events happen. Existing approaches assume that the training and testing videos are from the same scene, exhibiting poor generalization performance when encountering an unseen scene. In this paper, we propose a Variational Anomaly Detection Network (VADNet), which is characterized by its high scene-adaptation - it can identify abnormal events in a new scene only via referring to a few normal samples without fine-tuning. Our model embodies two major innovations. First, a novel Variational Normal Inference (VNI) module is proposed to formulate image reconstruction in a conditional variational auto-encoder (CVAE) framework, which learns a probabilistic decision model instead of a traditional deterministic one. Secondly, a Margin Learning Embedding (MLE) module is leveraged to boost the variational inference and aid in distinguishing normal events. We theoretically demonstrate that minimizing the triplet loss in MLE module facilitates maximizing the evidence lower bound (ELBO) of CVAE, which promotes the convergence of VNI. By incorporating variational inference with margin learning, VADNet becomes much more generative that is able to handle the uncertainty caused by the changed scene and limited reference data. Extensive experiments on several datasets demonstrate that the proposed VADNet can adapt to a new scene effectively without fine-tuning and achieve remarkable performance, which outperforms other methods significantly and establishes new state-of-the-art in the case of few-shot scene-adaptive anomaly detection. We believe our method is closer to real-world application due to its strong generalization ability. All codes are released inhttps://github.com/huangxx156/VADNet. Yutao Hu 0002, Xiaoyan Luo, Jungong Han, Baochang Zhang 0001, Xianbin Cao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2023 | Airport Capacity Prediction With Multisource Features: A Temporal Deep Learning ApproachabstractAccurate airport capacity estimation is crucial for the secure and orderly operation of the aviation system. However, such estimation is a non-trivial task as capacity depends on various meteorological and operational features. The complex coupling characteristics among these multi-source features have proved to be challenging for most of the traditional regression models. Recently, enhanced by its excellent ability to mine nonlinear relationships, the machine learning methods trigger widely applications. However, due to the imbalance of features scatter and the neglect of temporal dependences in aviation systems, existing machine learning methods for airport capacity prediction still have room for improvement. In light of these, this paper presents a novel airport capacity prediction method based on the multi-channel fusion Transformer model (MF-Transformer). Besides the commonly used aviation features, we unprecedentedly harness the power of the high-dimensional meteorological feature for accurate prediction. As to the model, we construct a multi-channel feature fusion structure, which includes a three-channel network for multi-source features extraction and an attention-based feature fusion module between channels. In each channel, the Transformer-based model is utilized to capture the temporal dependences of features. We conduct experiments on the capacity prediction tasks of the Beijing Capital International Airport which is the largest airport in China and verify that the proposed MF-Transformer outperforms benchmarks under different prediction horizons. Wenbo Du 0001, Shenwen Chen, Zhishuai Li, Xianbin Cao 0001 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2023 | Latent Domain Generation for Unsupervised Domain Adaptation Object CountingabstractUnsupervised cross-domain crowd counting has recently received great attention in computer vision, which generalizes the model from the source domain to the unlabeled target domain. However, it is an extremely challenging task because only unlabeled data is available from the target domain and the domain gap between two domains is implicit in crowd counting. In this paper, we propose a latent domain generation method to improve the generalization ability of unsupervised domain adaptation crowd counting by generating a latent domain. To this end, we propose a domain generator with random perturbations to learn a new latent distribution derived from the original source distribution. The latent domain generator can extract target information sampled in its stochastic latent representation, which preserves the original target information and enhances the variational ability. Meanwhile, to ensure that the generated latent domain is consistent with the source domain in counting performance, we introduce a consistency loss to encourage similar output from latent and source domains. Moreover, to enhance the adaptation ability of the generated latent domain, we apply the adversarial loss to achieve alignment between the latent and target domains. The domain generator with the adversarial loss and consistency loss ensures that the generated domain is aligned to the target while also improving the robustness of the original source domain model. The experiment indicates that our framework can effortlessly extend to scenarios with different objects (crowd, cars). The experiments also demonstrate the effectiveness of our method on unsupervised realistic-to-realistic crowd counting problems. Yandan Yang, Jun Xu 0019, Xianbin Cao 0001, Xiantong Zhen, Ling Shao 0001 |
IEEE Trans. Multim. | 4 |
| 2022 | IDa-Det: An Information Discrepancy-Aware Distillation for 1-Bit Detectors
Sheng Xu 0007, Yanjing Li, Bohan Zeng, Teli Ma, Baochang Zhang 0001, Xianbin Cao 0001, Peng Gao 0007, Jinhu Lü 0001 |
ECCV (11) | 6 |
| 2022 | Q-ViT: Accurate and Fully Quantized Low-bit Vision TransformerabstractThe large pre-trained vision transformers (ViTs) have demonstrated remarkable performance on various visual tasks, but suffer from expensive computational and memory cost problems when deployed on resource-constrained devices. Among the powerful compression approaches, quantization extremely reduces the computation and memory consumption by low-bit parameters and bit-wise operations. However, low-bit ViTs remain largely unexplored and usually suffer from a significant performance drop compared with the real-valued counterparts. In this work, through extensive empirical analysis, we first identify the bottleneck for severe performance drop comes from the information distortion of the low-bit quantized self-attention map. We then develop an information rectification module (IRM) and a distribution guided distillation (DGD) scheme for fully quantized vision transformers (Q-ViT) to effectively eliminate such distortion, leading to a fully quantized ViTs. We evaluate our methods on popular DeiT and Swin backbones. Extensive experimental results show that our method achieves a much better performance than the prior arts. For example, our Q-ViT can theoretically accelerates the ViT-S by 6.14x and achieves about 80.9% Top-1 accuracy, even surpassing the full-precision counterpart by 1.0% on ImageNet dataset. Our codes and models are attached on https://github.com/YanjingLi0202/Q-ViT Yanjing Li, Sheng Xu 0007, Baochang Zhang 0001, Xianbin Cao 0001, Peng Gao 0007, Guodong Guo |
NeurIPS | 4 |
| 2022 | Attentive encoder-decoder networks for crowd counting
Xuhui Liu, Yutao Hu 0002, Baochang Zhang 0001, Xiantong Zhen, Xiaoyan Luo, Xianbin Cao 0001 |
Neurocomputing | 6 |
| 2022 | Trajectory Design for UAV-Based Internet of Things Data Collection: A Deep Reinforcement Learning ApproachabstractIn this article, we investigate an unmanned aerial vehicle (UAV)-assisted Internet of Things (IoT) system in a sophisticated 3-D environment, where the UAV’s trajectory is optimized to efficiently collect data from multiple IoT ground nodes. Unlike existing approaches focusing only on a simplified 2-D scenario and the availability of perfect channel state information (CSI), this article considers a practical 3-D urban environment with imperfect CSI, where the UAV’s trajectory is designed to minimize data collection completion time subject to practical throughput and flight movement constraints. Specifically, inspired by the state-of-the-art deep reinforcement learning approaches, we leverage the twin-delayed deep deterministic policy gradient (TD3) to design the UAV’s trajectory and we present a TD3-based trajectory design for completion time minimization (TD3-TDCTM) algorithm. In particular, we set an additional information, i.e., the merged pheromone, to represent the state information of the UAV and environment as a reference of reward which facilitates the algorithm design. By taking the service statuses of the IoT nodes, the UAV’s position, and the merged pheromone as input, the proposed algorithm can continuously and adaptively learn how to adjust the UAV’s movement strategy. By interacting with the external environment in the corresponding Markov decision process, the proposed algorithm can achieve a near-optimal navigation strategy. Our simulation results show the superiority of the proposed TD3-TDCTM algorithm over three conventional nonlearning-based baseline methods. Yang Wang 0154, Zhen Gao 0001, Jun Zhang 0007, Xianbin Cao 0001, Dezhi Zheng, Yue Gao 0001, Derrick Wing Kwan Ng, Marco Di Renzo |
IEEE Internet Things J. | 4 |
| 2022 | A Hierarchical Incentive Design Toward Motivating Participation in Coded Federated LearningabstractFederated Learning (FL) is a privacy-preserving collaborative learning approach that trains artificial intelligence (AI) models without revealing local datasets of the FL workers. While FL ensures the privacy of the FL workers, its performance is limited by several bottlenecks, which become significant given the increasing amounts of data generated and the size of the FL network. One of the main challenges is the straggler effects where the significant computation delays are caused by the slow FL workers. As such, Coded Federated Learning (CFL), which leverages coding techniques to introduce redundant computations to the FL server, has been proposed to reduce the computation latency. In CFL, the FL server helps to compute a subset of the partial gradients based on the composite parity data and aggregates the computed partial gradients with those received from the FL workers. In order to implement the coding schemes over the FL network, incentive mechanisms are important to allocate the resources of the FL workers and data owners efficiently in order to complete the CFL training tasks. In this paper, we consider a two-level incentive mechanism design problem. In the lower level, the data owners are allowed to support the FL training tasks of the FL workers by contributing their data. To model the dynamics of the selection of FL workers by the data owners, an evolutionary game is adopted to achieve an equilibrium solution. In the upper level, a deep learning based auction is proposed to model the competition among the model owners. Jer Shyuan Ng, Wei Yang Bryan Lim, Zehui Xiong, Xianbin Cao 0001, Dusit Niyato, Cyril Leung, Dong In Kim 0001 |
IEEE J. Sel. Areas Commun. | 4 |
| 2022 | Cross-Domain Attention Network for Unsupervised Domain Adaptation Crowd CountingabstractUnsupervised domain adaptation crowd counting (UDACC) has been studied with practical research utility by getting rid of the labeling burden on large-scale dense crowds in the target domain. Current methods generalize well within the specific domain gap by directly aligning domain distributions or translating synthetic data to realistic images. However, it is difficult to define domain gaps among complex real-world datasets, in which the images vary greatly in style, density level and/or content. To tackle this problem, in this paper, we propose a Cross-Domain Attention Network (CDANet), which can effectively generalize the model to the unlabeled domain on both unsupervised synthetic-to-realistic and realistic-to-realistic crowd counting. Specifically, we propose a Cross-Domain Attention Module (CDAM) to learn domain-related information between the source and target domain, which extracts relations in cross-domain attentive information, thus enhancing crowd-informative features. Moreover, to make our CDAM invariant to domain shifts, we introduce a consistency penalty to ensure that the attention maps are consistent before and after the domain shifting. Thus our CDANet can pay attention to the shared counting information across domains, while remaining its invariant ability during domain adaptation. Extensive experiments on several common benchmarks for UDACC demonstrate that our CDANet gets competitive results on both unsupervised synthetic-to-realistic and realistic-to-realistic UDACC tasks. Jun Xu 0019, Xiaoyan Luo, Xianbin Cao 0001, Xiantong Zhen |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | Variational Self-Distillation for Remote Sensing Scene ClassificationabstractSupported by deep learning techniques, remote sensing scene classification, a fundamental task in remote image analysis, has recently obtained remarkable progress. However, due to the severe uncertainty and perturbation within an image, it is still a challenging task and remains many unsolved problems. In this paper, we note that regular one-hot labels cannot precisely describe remote sensing images, and they fail to provide enough information for supervision and limiting the discriminative feature learning of the network. To solve this problem, we propose a Variational Self-Distillation Network (VSDNet), in which the class entanglement information from the prediction vector acts as the supplement to the category information. Then, the exploited information is hierarchically distilled from the deep layers into the shallow parts via a Variational Knowledge Transfer (VKT) module. Notably, the VKT module performs knowledge distillation in a probabilistic way through variational estimation, which enables end-to-end optimization for mutual information and promotes robustness to uncertainty within the image. Extensive experiments on four challenging remote sensing datasets demonstrate that, with a negligible parameter increase, the proposed VSDNet brings a significant performance improvement over different backbone networks and delivers state-of-the-art results. Yutao Hu 0002, Xiaoyan Luo, Jungong Han, Xianbin Cao 0001, Jun Zhang 0007 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | A Deep Unsupervised Learning Approach for Airspace Complexity EvaluationabstractAirspace complexity is a critical metric in current Air Traffic Management systems for indicating the security degree of airspace operations. Airspace complexity can be affected by many coupling factors in a complicated and nonlinear way, making it extremely difficult to be evaluated. In recent years, machine learning has been proved as a promising approach and achieved significant results in evaluating airspace complexity. However, existing machine learning based approaches require a large number of airspace operational data labeled by experts. Due to the high cost in labeling the operational data and the dynamical nature of the airspace operating environment, such data are often limited and may not be suitable for the changing airspace situation. In light of these, we propose a novel unsupervised learning approach for airspace complexity evaluation based on a deep neural network trained by unlabeled samples. We introduce a new loss function to better address the characteristics pertaining to airspace complexity data, including dimension coupling, category imbalance, and overlapped boundaries. Due to these characteristics, the generalization ability of existing unsupervised models is adversely impacted. The proposed approach is validated through extensive experiments based on the real-world data of six sectors in Southwestern China airspace. Experimental results show that our deep unsupervised model outperforms the state-of-the-art methods in terms of airspace complexity evaluation accuracy. Biyue Li, Wenbo Du 0001, Yu Zhang 0087, Jun Chen 0009, Ke Tang 0001, Xianbin Cao 0001 |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2022 | Reputation-Aware Hedonic Coalition Formation for Efficient Serverless Hierarchical Federated LearningabstractAmid growing concerns on data privacy, Federated Learning (FL) has emerged as a promising privacy preserving distributed machine learning paradigm. Given that the FL network is expected to be implemented at scale, several studies have proposed system architectures towards improving the network scalability and efficiency. Specifically, the Hierarchical FL (HFL) network utilizes cluster heads, e.g., base stations, for the intermediate aggregation and relay of model parameters. Serverless FL is also proposed recently, in which the data owners, i.e., workers, exchange the local model parameters among a neighborhood of workers. This decentralized approach reduces the risk of a single point of failure but inevitably incurs significant communication overheads. To achieve the best of both worlds, we propose the Serverless Hierarchical Federated Learning (SHFL) framework in this paper. The SHFL framework adopts a two-layer system architecture. In the lower layer, the FL workers are grouped into clusters under cluster heads. In the upper layer, the cluster heads exchange the intermediate parameters with their one-hop neighbors without the aid of a central server. To improve the sustainable efficiency of the FL system while taking into account the incentive design for workers marginal contributions in the system, we propose the reputation-aware hedonic coalition formation game in this paper. Specifically, the workers are rewarded for their marginal contribution to the cluster, whereas the reputation opinions of each cluster head is updated in a decentralized manner, thereby deterring malicious behaviors by the cluster head. This improves the performance of the network since cluster heads with higher reputation scores are more reliable in relaying the intermediate model parameters. The simulation results show that our proposed hedonic coalition formation algorithm converges to a Nash-stable partition and improves the network efficiency. Jer Shyuan Ng, Wei Yang Bryan Lim, Zehui Xiong, Xianbin Cao 0001, Jiangming Jin, Dusit Niyato, Cyril Leung, Chunyan Miao |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2022 | Feeling of Presence Maximization: mmWave-Enabled Virtual Reality Meets Deep Reinforcement LearningabstractThis paper investigates the problem of providing ultra-reliable and power-efficient virtual reality (VR) experiences for wireless mobile users. To ensure reliable ultra-high-definition (UHD) video frame delivery to mobile users and enhance their immersive visual experiences, a coordinated multipoint (CoMP) transmission technique and millimeter wave (mmWave) communications are exploited. Owing to user movement and time-varying wireless channels, the wireless VR experience enhancement problem is formulated as a sequence-dependent and mixed-integer problem with a goal of maximizing users’ feeling of presence (FoP) in the virtual world, subject to power consumption constraints on access points (APs) and users’ head-mounted displays (HMDs). The problem, however, is hard to be directly solved due to the lack of users’ accurate tracking information and the sequence-dependent and mixed-integer characteristics. To overcome this challenge, we develop a parallel echo state network (ESN) learning method to predict users’ tracking information by training fresh and historical tracking samples separately collected by APs. With the learnt results, we propose a deep reinforcement learning (DRL) based optimization algorithm to solve the formulated problem. In this algorithm, we implement deep neural networks (DNNs) as a scalable solution to produce integer decision variables and solve a continuous power control problem to criticize the integer decision variables. Finally, the performance of the proposed algorithm is compared with various benchmark algorithms, and the impact of different design parameters is also discussed. Simulation results demonstrate that the proposed algorithm is more 4.14% power-efficient than the benchmark algorithms. Peng Yang 0009, Tony Q. S. Quek, Jingxuan Chen, Chaoqun You, Xianbin Cao 0001 |
IEEE Trans. Wirel. Commun. | 5 |
| 2021 | Variational Prototype Inference for Few-Shot Semantic SegmentationabstractIn this paper, we propose variational prototype inference to address few-shot semantic segmentation in a probabilistic framework. A probabilistic latent variable model infers the distribution of the prototype that is treated as the latent variable. We formulate the optimization as a variational inference problem, which is established with an amortized inference network based on an auto-encoder architecture. The probabilistic modeling of the prototype enhances its generalization ability to handle the inherent uncertainty caused by limited data and the huge intra-class variations of objects. Moreover, it offers a principled way to incorporate the prototype extracted from support images into the prediction of the segmentation maps for query images. We conduct extensive experimental evaluations on three benchmark datasets. Ablation studies show the effectiveness of variational prototype inference for few-shot semantic segmentation by probabilistic modeling. On all three benchmarks, our proposal achieves high segmentation accuracy and surpasses previous methods by considerable margins. Yandan Yang, Xianbin Cao 0001, Xiantong Zhen, Cees Snoek, Ling Shao 0001 |
WACV | 3 |
| 2021 | IncreACO: Incrementally Learned Automatic Check-out with Photorealistic Exemplar AugmentationabstractAutomatic check-out (ACO) emerges as an integral component in recent self-service retailing stores, which aims at automatically detecting and counting the randomly placed products upon a check-out platform. Existing data-driven counting works still have difficulties in generalizing to real-world retail product counting scenarios, since (1) real check-out images are hard to collect or cover all products and their possible layouts, (2) rapid updating of the product list leads to frequent and tedious re-training of the counting models. To overcome these obstacles, we contribute a practical automatic check-out framework tailored to real-world retail product counting scenarios, consisting of a photorealistic exemplar augmentation to generate physically reliable and photorealistic check-out images from canonical exemplars scanned for each product and an incremental learning strategy to match the updating nature of the ACO system with much fewer training effort. Through comprehensive studies, we show that the proposed IncreACO serves as an effective framework on the recent Retail Product Checkout (RPC) dataset, where the proposed photorealistic exemplar augmentation remarkably improves the counting performance against the state-of-the-art methods (77.15% v.s. 72.83% in counting accuracy), whilst the proposed incremental learning framework consistently extends the counting performance to new categories. Yandan Yang, Lu Sheng, Dong Xu 0001, Xianbin Cao 0001 |
WACV | 6 |
| 2021 | Long-range Attention Network for Multi-View StereoabstractLearning-based multi-view stereo (MVS) has recently gained great popularity, which can efficiently infer depth map and reconstruct fine-grained scene geometry. Previous methods calculate the variance of the corresponding pixel pairs to determine whether they are matched mostly based on the pixel-wise measure, which fails to consider the interdependence among pixels and is ineffective on the matching of texture-less or occluded regions. These false matching problems challenge MVS and result in its most failure cases. To address the issues, we introduce a Long-range Attention Network (LANet) to selectively aggregate reference features to each position to capture the long-range interdependence across the entire space. As a result, similar features relate to each other regardless of their distance, propagating more guiding information for the effective match. Furthermore, we introduce a new loss to supervise the intermediate probability volume by constraining its distribution reasonably centered at the true depth. Extensive experiments on large-scale DTU dataset demonstrate that the proposed LANet achieves the new state-of-the-art performance, outperforming previous methods by a large margin. Our method is generic and also achieves comparable results on outdoor Tanks and Temples dataset without any fine-tuning, which validates our method's generalization ability. Yutao Hu 0002, Xianbin Cao 0001, Baochang Zhang 0001 |
WACV | 4 |
| 2021 | Energy-Efficient Resource Allocation in a Multi-UAV-Aided NOMA NetworkabstractThis paper is concerned with the resource allocation in a multi-unmanned aerial vehicle (UAV)-aided network for providing enhanced mobile broadband (eMBB) services for user equipments. Different from most of the existing network resource allocation approaches, we investigate a joint non-orthogonal user association, subchannel allocation and power control problem. The objective of the problem is to maximize the network energy efficiency under the constraints on user equipments' quality of service, UAVs' network capacity and power consumption. We formulate the energy efficiency maximization problem as a challenging mixed-integer non-convex programming problem. To alleviate this problem, we first decompose the original problem into two subproblems, namely, an integer non-linear user association and subchannel allocation subproblem and a non-convex power control subproblem. We then design a two-stage approximation strategy to handle the non-linearity of the user association and subchannel allocation subproblem and exploit a successive convex approximation approach to tackle the non-convexity of the power control subproblem. Based on the derived results, we develop an iterative algorithm with provable convergence to mitigate the original problem. Simulation results show that our proposed framework can improve energy efficiency compared with several benchmark algorithms. Xing Xi, Xianbin Cao 0001, Peng Yang 0009, Jingxuan Chen, Dapeng Oliver Wu |
WCNC | 2 |
| 2021 | RAN Slicing for Massive IoT and Bursty URLLC Service Multiplexing: Analysis and OptimizationabstractFuture wireless networks are envisioned to serve massive Internet of Things (mIoT) via some radio access technologies, where the random access channel (RACH) procedure should be exploited for IoT devices to access the networks. However, the theoretical analysis of the RACH procedure for massive IoT devices is challenging. To address this challenge, we first correlate the RACH request of an IoT device with the status of its maintained queue and analyze the evolution of the queue status by the probability theory. Based on the analysis result, we then derive the closed-form expression of the random access (RA) success probability, which is a significant indicator characterizing the RACH procedure of the device by the stochastic geometry theory. Besides, considering the agreement on converging different services onto a shared infrastructure, we investigate the radio access network (RAN) slicing for mIoT and bursty ultrareliable and low-latency communication (URLLC) service multiplexing. Specifically, we formulate the RAN slicing problem as an optimization one to maximize the total RA success probabilities of all IoT devices and provide URLLC services for URLLC devices in an energy-efficient way. A slice resource optimization (SRO) algorithm, exploiting relaxation and approximation with provable tightness and error bound, is then proposed to mitigate the optimization problem. Simulation results demonstrate that the proposed SRO algorithm can effectively implement the service multiplexing of mIoT and bursty URLLC traffic. Peng Yang 0009, Xing Xi, Tony Q. S. Quek, Jingxuan Chen, Xianbin Cao 0001, Dapeng Oliver Wu |
IEEE Internet Things J. | 5 |
| 2021 | Fast Pseudospectrum Estimation for Automotive Massive MIMO RadarabstractSubspace methods, e.g., multiple signal classification algorithm (MUSIC), show great promise to high-resolution environment sensing in the 6G-enabled mobile Internet of Things (IoT), e.g., the emerging unmanned systems. Existing schemes, aiming to simplify the computational 1-D search of the MUSIC pseudospectrum, unfortunately have still an unaffordable complexity or the compromised accuracy, especially when the millimeter-wave massive multiple-input–multiple-output (MIMO) radar is considered. In this work, we address the fast and accurate estimation of the high-resolution pseudospectrum in massive MIMO radars. To enable real-time automotive sensing, we first formulate this computational procedure as one matrix product problem, which is then solved by leveraging randomized matrix sketching techniques. To be specific, we compute the large matrix productapproximatelyby the product of two small matrices abstracted via random sampling. To minimize the approximation error, we further design another sampling, pruning, and recomputing (SaPRe) algorithm, which refines the approximated results and thus attains the exact pseudospectrum. Finally, the theoretical analysis and numerical simulations are provided to validate the proposed methods. Our fast approaches dramatically reduce the time complexity and simultaneously attain the accurate Direction-of-Arrival (DoA) estimation, which have the great potential to real time and high-resolution automotive sensing with massive MIMO radars. Bin Li 0002, Shusen Wang, Zhiyong Feng 0001, Jun Zhang 0007, Xianbin Cao 0001, Chenglin Zhao |
IEEE Internet Things J. | 5 |
| 2021 | Proactive UAV Network Slicing for URLLC and Mobile Broadband Service MultiplexingabstractThe unmanned aerial vehicle (UAV) network that is convinced as a significant component of 5G and emerging 6G wireless networks is desired to accommodate multiple types of service requirements simultaneously. However, how to converge different types of services onto a common UAV network without deploying an individual network solution for each type of service is challenging. We tackle this challenge in this paper through slicing the UAV network, i.e., creating logical UAV networks customized for specific requirements. To this end, we formulate the UAV network slicing problem as a sequential decision problem to provide mobile broadband (MBB) services for ground mobile users while satisfying ultra-reliable and low-latency requirements of UAV control and non-payload signal delivery. This problem, however, is difficult to be directly solved mainly due to the sequence-dependent characteristic and the lack of accurate location information of mobile users and accurate and tractable channel gain models in practice. To overcome these difficulties, we propose a novel solution approach based on learning and optimization methods. Particularly, we develop a distributed learning method to predict mobile users’ locations, where partial user location information stored on each UAV is utilized to train user location prediction networks. To achieve accurate channel gain models, we design deep neural networks (DNNs) that are trained by signal measurements at each UAV. To cope with the challenging sequence-dependent characteristic of the problem, we develop a Lyapunov-based optimization framework with provable performance guarantees to decompose the original problem into a sequence of separate optimization subproblems based on the learned results. Finally, an iterative optimization scheme joint with a successive convex approximation technique is exploited to solve these subproblems. Simulation results demonstrate the accuracy of the learning methods as well as the effectiveness of the Lyapunov-based optimization framework. Peng Yang 0009, Xing Xi, Kun Guo 0002, Tony Q. S. Quek, Jingxuan Chen, Xianbin Cao 0001 |
IEEE J. Sel. Areas Commun. | 6 |
| 2021 | A semi-supervised zero-shot image classification method based on soft-target
Zhong Ji, Qiang Wang 0056, Biying Cui, Yanwei Pang, Xianbin Cao 0001, Xuelong Li 0001 |
Neural Networks | 5 |
| 2021 | Multi-source off-grid DOA estimation with single snapshot using non-uniform linear arrays
Xianbin Cao 0001, Xiangrong Wang 0001, Maria Greco 0001, Fulvio Gini |
Signal Process. | 2 |
| 2021 | Network Resource Allocation for eMBB Payload and URLLC Control Information Communication Multiplexing in a Multi-UAV Relay NetworkabstractUnmanned aerial vehicle (UAV) relay networks are convinced to be a significant complement to terrestrial infrastructures to provide robust network capacity. However, most of the existing works either considered enhanced mobile broadband (eMBB) payload communication or ultra-reliable and low latency communications (URLLC) control information communication. In this paper, we investigate resource allocation for the eMBB payload and URLLC control information communication multiplexing in a multi-UAV relay network. We firstly propose a multi-UAV relay model comprehensively considering path loss, small-scale channel fading and different quality of service requirements of eMBB and URLLC communications. Then we formulate the multiplexing problem as a joint user association, bandwidth and transmit power optimization problem to improve total transmission data rate and reduce power consumption. The solution of this problem is challenging due to different capacity characteristics of eMBB and URLLC communications, the coupling of continuous variables and integer variables, and the non-convexity. To mitigate these challenges, we equivalently decompose the original optimization problem into a URLLC problem and an eMBB problem. For the URLLC problem, we derive closed-form expressions of the optimal bandwidth and transmit power. For the eMBB problem, we develop an iterative solution framework of alternatively optimizing user association, bandwidth and transmit power. Xing Xi, Xianbin Cao 0001, Peng Yang 0009, Jingxuan Chen, Tony Q. S. Quek, Dapeng Oliver Wu |
IEEE Trans. Commun. | 2 |
| 2021 | How Should I Orchestrate Resources of My Slices for Bursty URLLC Service Provision?abstractFuture wireless networks are convinced to provide flexible and cost-efficient services via exploiting network slicing techniques. However, it is challenging to configure slicing systems for bursty ultra-reliable and low latency communications (URLLC) service provision due to its stringent requirements on low packet blocking probability and low codeword decoding error probability. In this paper, we propose to orchestrate network resources for a slicing system to guarantee more reliable bursty URLLC transmission. We re-cut physical resource blocks and derive the minimum upper bound of bandwidth for URLLC transmission with a low packet blocking probability. We correlate coordinated multipoint beamforming with channel uses and derive the minimum upper bound of channel uses for URLLC transmission with a low codeword decoding error probability. Considering the agreement on converging diverse services onto shared infrastructures, we further investigate the network slicing for URLLC and enhanced mobile broadband (eMBB) service multiplexing. Particularly, we formulate the service multiplexing as an optimization problem, which is challenging to be mitigated due to requirements of future channel information and of tackling a two timescale issue. To address the challenges, we develop a resource optimization algorithm based on a sample average approximate technique and a distributed optimization method with provable performance guarantees. Peng Yang 0009, Xing Xi, Tony Q. S. Quek, Jingxuan Chen, Xianbin Cao 0001, Dapeng Oliver Wu |
IEEE Trans. Commun. | 5 |
| 2021 | Attentional Kernel Encoding Networks for Fine-Grained Visual CategorizationabstractFine-grained visual categorization aims to recognize objects from different sub-ordinate categories, which is a challenging task due to subtle visual differences between images. It is highly desired to identify discriminative regions while achieving highly non-linear compact representation for fine-grained visual categorization. However, existing methods either rely on manually defined part-based annotations to indicate the distinctive regions or operate on longitudinal vectors to capture the non-linear information, which may lose important spatial layout information. In this paper, we propose the Attentional Kernel Encoding Networks (AKEN) for fine-grained visual categorization. Specifically, the AKEN aggregates feature maps from the last convolutional layer of ConvNets to obtain a holistic feature representation. By Fourier embedding, it encodes features from both the longitudinal and transverse directions, which largely retains the spatial layout information. Moreover, we incorporate a Cascaded Attention (Cas-Attention) module to highlight local regions that distinguish among subordinate categories, enabling the AKEN to extract the most discriminative features. Working in conjunction with the attention mechanism, the proposed AKEN combines the strengths of ConvNets and kernels for non-linear feature learning, which can establish discriminative and descriptive feature representations for fine-grained image categorization. Experiments on three benchmark datasets show that the proposed AKEN delivers highly competitive performance, surpassing most existed methods and achieving state-of-the-art results. Yutao Hu 0002, Yandan Yang, Jun Zhang 0007, Xianbin Cao 0001, Xiantong Zhen |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2021 | Fast Nearest Subspace Search via Random Angular HashingabstractSubspaces frequently offer powerful representation in many tasks including recognition, retrieval, and optimization. In these tasks, the nearest subspaces (i.e., subspace-to-subspace search) often inevitably arise. Several studies in the literature have attempted to address this hard problem using techniques such as locality-sensitive hashing. Unfortunately, these subspace hashing methods are severely affected by poor scaling, with consequently high computational cost or unsatisfying accuracy, when the subspaces originally distribute with arbitrary dimensions. Accordingly, in this paper, we propose random angular hashing, a new and efficient type of locality-sensitive hashing, for linear subspaces of arbitrary dimension. The method we proposed preserves the angular distances among subspaces by randomly projecting their orthonormal basis and then encoding them with binary codes, meanwhile not only achieving fast computation but also maintaining a powerful collision probability. Moreover, its flexibility to easily get a balance between efficiency and accuracy in terms of performance. The extensive experimental results on tasks of face recognition, video de-duplication, and gesture recognition demonstrate that the proposed approach performs better than the state-of-the-art methods heavily, in terms of both accuracy and efficiency (up to 16× speedup). Yi Xu 0013, Xianglong Liu 0001, Binshuai Wang, Renshuai Tao, Ke Xia, Xianbin Cao 0001 |
IEEE Trans. Multim. | 6 |
| 2021 | Predictive UAV Base Station Deployment and Service Offloading With Distributed Edge LearningabstractIn modern networks, edge computing will be responsible for processing and learning from the critical network- and user-generated data, such as wireless link usage, mobility information, application requests, and many others. The presence of Artificial Intelligence-based (AI) applications at the edge of the network will enable the network to predict necessary user behavior and its impact on network infrastructure, such as base station overloading. One of the main strategies for offloading users and base stations is to deploy UAV base stations, or flying base stations, which can dynamically provide service and connectivity. In this article, we introduce a framework for distributed learning over Multi-access Edge Computing (MEC), which manages data applications in a fully distributed setting across edge servers, thus reducing the cost of collecting user information in a centralized server. We couple the proposed distributed learning with a novel similarity metric for user trajectories, which can aggregate neural network models with similar costs as other model aggregation techniques. However, the aggregation technique can achieve much higher accuracy. Furthermore, we apply the proposed distributed learning scheme to manage and deploy flying base stations to areas that experience high demand or poor user connectivity, thus optimizing connectivity in terms of user satisfaction, delay, and network throughput. Zhongliang Zhao, Lucas Pacheco, Hugo Santos, Antonio Di Maio, Denis do Rosário, Eduardo Cerqueira, Torsten Braun, Xianbin Cao 0001 |
IEEE Trans. Netw. Serv. Manag. | 9 |
| 2021 | Alignment Enhancement Network for Fine-grained Visual CategorizationabstractFine-grained visual categorization (FGVC) aims to automatically recognize objects from different sub-ordinate categories. Despite attracting considerable attention from both academia and industry, it remains a challenging task due to subtle visual differences among different classes. Cross-layer feature aggregation and cross-image pairwise learning become prevailing in improving the performance of FGVC by extracting discriminative class-specific features. However, they are still inefficient to fully use the cross-layer information based on the simple aggregation strategy, while existing pairwise learning methods also fail to explore long-range interactions between different images. To address these problems, we propose a novel Alignment Enhancement Network (AENet), including two-level alignments, Cross-layer Alignment (CLA) and Cross-image Alignment (CIA). The CLA module exploits the cross-layer relationship between low-level spatial information and high-level semantic information, which contributes to cross-layer feature aggregation to improve the capacity of feature representation for input images. The new CIA module is further introduced to produce the aligned feature map, which can enhance the relevant information as well as suppress the irrelevant information across the whole spatial region. Our method is based on an underlying assumption that the aligned feature map should be closer to the inputs of CIA when they belong to the same category. Accordingly, we establish Semantic Affinity Loss to supervise the feature alignment within each CIA block. Experimental results on four challenging datasets show that the proposed AENet achieves the state-of-the-art results over prior arts. Yutao Hu 0002, Xuhui Liu, Baochang Zhang 0001, Jungong Han, Xianbin Cao 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2021 | Multicast eMBB and Bursty URLLC Service Multiplexing in a CoMP-Enabled RANabstractThis paper is concerned with slicing a radio access network (RAN) for simultaneously serving two 5G-and-Beyond typical use cases, i.e., enhanced mobile broadband (eMBB) and ultra-reliable and low-latency communications (URLLC). Although many researches have been conducted to tackle this issue, few of them have considered the impact of bursty URLLC. The bursty characteristic of URLLC traffic may significantly increase the difficulty of RAN slicing in terms of ensuring an ultra-low packet blocking probability. To reduce the probability, we re-visit the structure of physical resource blocks orchestrated for URLLC traffic based on theoretical results. Meanwhile, we formulate the problem of slicing a RAN enabling coordinated multi-point (CoMP) transmissions for multicast eMBB and bursty URLLC service multiplexing as a multi-timescale optimization problem aiming at maximizing eMBB and URLLC slice utilities, subject to physical resource constraints. To mitigate this problem, we transform it into multiple single timescale problems by exploring sample average approximations. An iterative algorithm with provable performance guarantees is developed to obtain solutions to these single timescale problems and aggregate obtained solutions into those of the multi-timescale problem. We also design a CoMP-enabled RAN slicing system prototype and compare the iterative algorithm with the state-of-the-art algorithm to verify its effectiveness. Peng Yang 0009, Xing Xi, Yaru Fu, Tony Q. S. Quek, Xianbin Cao 0001, Dapeng Oliver Wu |
IEEE Trans. Wirel. Commun. | 5 |
| 2020 | NAS-Count: Counting-by-Density with Neural Architecture Search
Yutao Hu 0002, Xuhui Liu, Baochang Zhang 0001, Jungong Han, Xianbin Cao 0001, David S. Doermann |
ECCV (22) | 6 |
| 2020 | Few-Shot Semantic Segmentation with Democratic Attention Networks
Yutao Hu 0002, Yandan Yang, Xianbin Cao 0001, Xiantong Zhen |
ECCV (13) | 5 |
| 2020 | Repeatedly Energy-Efficient and Fair Service Coverage: UAV SlicingabstractUnmanned aerial vehicle (UAV) networks are convinced as a significant part of 5G and emerging 6G wireless networks. UAV slicing is a promising proposal of converging different services onto a common UAV network without deploying individual network solution for each type of service. This paper is concerned with UAV slicing for providing energy-efficient and fair service coverage for enhanced mobile broad-band (eMBB) users (UEs). Aiming at physically configuring UAV slices, the UAV slicing problem is formulated as a time-dependent mixed-integer-non-convex programming problem with a goal of maximizing all UEs' data rates while minimizing UAVs' total transmit power. To mitigate this challenging problem, we first decompose the original problem into two time-dependent subproblems using a Lyapunov approach. We then derive the procedure of tackling the non-convexity and the mixed-integer property of the subproblems by exploring a successive convex approximate (SCA) method and an alternative optimization scheme, respectively. Based on the derived results, we develop an algorithm with provable performance guarantees to mitigate the two subproblems repeatedly. Peng Yang 0009, Xing Xi, Tony Q. S. Quek, Jingxuan Chen, Xianbin Cao 0001, Dapeng Oliver Wu |
GLOBECOM | 5 |
| 2020 | You Only Need The Image: Unsupervised Few-Shot Semantic Segmentation With Co-Guidance NetworkabstractFew-shot semantic segmentation has recently attracted attention for its ability to segment unseen-class images with only a few annotated support samples. Yet existing methods not only need to be trained with a large scale of pixel-level annotations on certain seen classes, but also require a few annotated support image-mask pairs for the guidance of segmentation on each unseen class. In this paper, we propose the Co-guidance Network (CGNet) for unsupervised few-shot segmentation, which eliminates requirements of annotation on both seen and unseen classes. Specifically, CGNet segments unseen-class images with only unlabeled support images by the newly designed co-guidance mechanism. Moreover, CGNet is trained on seen classes by a novel co-existence recognition loss, which further removes the need of pixel-level annotations. Extensive experiments on the PASCAL -5idataset show that the unsupervised CGNet performs comparably with the state-of-the-art fully-supervised few-shot methods, while largely alleviating annotation requirement. Yandan Yang, Xianbin Cao 0001, Xiantong Zhen |
ICIP | 4 |
| 2020 | Model-Agnostic Metric for Zero-Shot LearningabstractZero-shot Learning (ZSL) aims to learn a classifier to recognize unseen categories without training samples. Most ZSL works based on embedding models handle the visual space and the semantic space through a common metric space and then apply a simple nearest neighbor search which directly leads to the hubness problem, one of the main challenges of ZSL. Contrary to recent works, whose conclusions about hubs are drawn based on Euclidean and specific models like ridge regression, we adopt cosine metric and for the first time prove cosine is model-agnostic to alleviate the hubness problem in ZSL. Assuming that the normalized mapped semantic vectors follow a uniform distribution, we provide theoretical analysis which demonstrates that hubs can be better reduced with a higher-dimensional cosine metric space. Moreover, we introduce a diversity-based regularizer with the cosine metric which underpins the assumption about the uniform distribution and further improves the model's discriminative ability. Extensive experiments on five benchmarks and large-scale Imagenet dataset show that our method can improve the performance, surpassing previous embedding methods by large margins. Qiang Qiu 0001, Xiantong Zhen, Xianbin Cao 0001 |
WACV | 6 |
| 2020 | Millimeter-Wave Full-Duplex UAV Relay: Joint Positioning, Beamforming, and Power ControlabstractIn this paper, a full-duplex unmanned aerial vehicle (FD-UAV) relay is employed to increase the communication capacity of millimeter-wave (mmWave) networks. Large antenna arrays are equipped at the source node (SN), destination node (DN), and FD-UAV relay to overcome the high path loss of mmWave channels and to help mitigate the self-interference at the FD-UAV relay. Specifically, we formulate a problem for maximization of the achievable rate from the SN to the DN, where the UAV position, analog beamforming, and power control are jointly optimized. Since the problem is highly non-convex and involves high-dimensional, highly coupled variable vectors, we first obtain the conditional optimal position of the FD-UAV relay for maximization of an approximate upper bound on the achievable rate in closed form, under the assumption of a line-of-sight (LoS) environment and ideal beamforming. Then, the UAV is deployed to the position which is closest to the conditional optimal position and yields LoS paths for both air-to-ground links. Subsequently, we propose an alternating interference suppression (AIS) algorithm for the joint design of the beamforming vectors and the power control variables. In each iteration, the beamforming vectors are optimized for maximization of the beamforming gains of the target signals and the successive reduction of the interference, where the optimal power control variables are obtained in closed form. Our simulation results confirm the superiority of the proposed positioning, beamforming, and power control method compared to three benchmark schemes. Furthermore, our results show that the proposed solution closely approaches a performance upper bound for mmWave FD-UAV systems. Lipeng Zhu 0001, Jun Zhang 0007, Zhenyu Xiao, Xianbin Cao 0001, Xiang-Gen Xia 0001, Robert Schober |
IEEE J. Sel. Areas Commun. | 4 |
| 2020 | Multi-scale Supervised Attentive Encoder-Decoder Network for Crowd CountingabstractCrowd counting is a popular topic with widespread applications. Currently, the biggest challenge to crowd counting is large-scale variation in objects. In this article, we focus on overcoming this challenge by proposing a novel Attentive Encoder-Decoder Network (AEDN), which is supervised on multiple feature scales to conduct crowd counting via density estimation. This work has three main contributions. First, we augment the traditional encoder-decoder architecture with our proposed residual attention blocks, which, beyond skip-connected encoded features, further extend the decoded features with attentive features. AEDN is better at establishing long-range dependencies between the encoder and decoder, therefore promoting more effective fusion of multi-scale features for handling scale-variations. Second, we design a new KL-divergence-based distribution loss to supervise the scale-aware structural differences between two density maps, which complements the pixel-isolated MSE loss and better optimizes AEDN to generate high-quality density maps. Third, we adopt a multi-scale supervision scheme, such that multiple KL divergences and MSE losses are deployed at all decoding stages, providing more thorough supervisions for different feature scales. Extensive experimental results on four public datasets, including ShanghaiTech Part A, ShanghaiTech Part B, UCF-CC-50, and UCF-QNRF, reveal the superiority and efficacy of the proposed method, which outperforms most state-of-the-art competitors. Baochang Zhang 0001, Xianbin Cao 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2019 | Projection Convolutional Neural Networks for 1-bit CNNs via Discrete Back PropagationabstractThe advancement of deep convolutional neural networks (DCNNs) has driven significant improvement in the accuracy of recognition systems for many computer vision tasks. However, their practical applications are often restricted in resource-constrained environments. In this paper, we introduce projection convolutional neural networks (PCNNs) with a discrete back propagation via projection (DBPP) to improve the performance of binarized neural networks (BNNs). The contributions of our paper include: 1) for the first time, the projection function is exploited to efficiently solve the discrete back propagation problem, which leads to a new highly compressed CNNs (termed PCNNs); 2) by exploiting multiple projections, we learn a set of diverse quantized kernels that compress the full-precision kernels in a more efficient way than those proposed previously; 3) PCNNs achieve the best classification performance compared to other state-ofthe-art BNNs on the ImageNet and CIFAR datasets. Jiaxin Gu, Ce Li 0002, Baochang Zhang 0001, Jungong Han, Xianbin Cao 0001, Jianzhuang Liu, David S. Doermann |
AAAI | 5 |
| 2019 | Attentive Temporal Pyramid Network for Dynamic Scene ClassificationabstractDynamic scene classification is an important yet challenging problem especially with the presence of defected or irrelevant frames due to unconstrained imaging conditions such as illumination, camera motion and irrelevant background. In this paper, we propose the attentive temporal pyramid network (ATP-Net) to establish effective representations of dynamic scenes by extracting and aggregating the most informative and discriminative features. The proposed ATP-Net detects informative features of frames that contain the most relevant information to scenes by a temporal pyramid structure with the incorporated attention mechanism. These frame features are effectively fused by a newly designed kernel aggregation layer based on kernel approximation into a discriminative holistic representations of dynamic scenes. The proposed ATP-Net leverages the strength of attention mechanism to select the most relevant frame features and the ability of kernels to achieve optimal feature fusion for discriminative representations of dynamic scenes. Extensive experiments and comparisons are conducted on three benchmark datasets and the results show our superiority over the state-of-the-art methods on all these three benchmark datasets. Yuanjun Huang, Xianbin Cao 0001, Xiantong Zhen, Jungong Han |
AAAI | 2 |
| 2019 | Crowd Counting and Density Estimation by Trellis Encoder-Decoder NetworksabstractCrowd counting has recently attracted increasing interest in computer vision but remains a challenging problem. In this paper, we propose a trellis encoder-decoder network (TEDnet) for crowd counting, which focuses on generating high-quality density estimation maps. The major contributions are four-fold. First, we develop a new trellis architecture that incorporates multiple decoding paths to hierarchically aggregate features at different encoding stages, which improves the representative capability of convolutional features for large variations in objects. Second, we employ dense skip connections interleaved across paths to facilitate sufficient multi-scale feature fusions, which also helps TEDnet to absorb the supervision information. Third, we propose a new combinatorial loss to enforce similarities in local coherence and spatial correlation between maps. By distributedly imposing this combinatorial loss on intermediate outputs, TEDnet can improve the back-propagation process and alleviate the gradient vanishing problem. Finally, on four widely-used benchmarks, our TEDnet achieves the best overall performance in terms of both density map quality and counting accuracy, with an improvement up to 14% in MAE metric. These results validate the effectiveness of TEDnet for crowd counting. Zehao Xiao, Baochang Zhang 0001, Xiantong Zhen, Xianbin Cao 0001, David S. Doermann, Ling Shao 0001 |
CVPR | 5 |
| 2019 | Relational Attention Network for Crowd CountingabstractCrowd counting is receiving rapidly growing research interests due to its potential application value in numerous real-world scenarios. However, due to various challenges such as occlusion, insufficient resolution and dynamic backgrounds, crowd counting remains an unsolved problem in computer vision. Density estimation is a popular strategy for crowd counting, where conventional density estimation methods perform pixel-wise regression without explicitly accounting the interdependence of pixels. As a result, independent pixel-wise predictions can be noisy and inconsistent. In order to address such an issue, we propose a Relational Attention Network (RANet) with a self-attention mechanism for capturing interdependence of pixels. The RANet enhances the self-attention mechanism by accounting both short-range and long-range interdependence of pixels, where we respectively denote these implementations as local self-attention (LSA) and global self-attention (GSA). We further introduce a relation module to fuse LSA and GSA to achieve more informative aggregated feature representations. We conduct extensive experiments on four public datasets, including ShanghaiTech A, ShanghaiTech B, UCF-CC-50 and UCF-QNRF. Experimental results on all datasets suggest RANet consistently reduces estimation errors and surpasses the state-of-the-art approaches by large margins. Zehao Xiao, Fan Zhu 0001, Xiantong Zhen, Xianbin Cao 0001, Ling Shao 0001 |
ICCV | 6 |
| 2019 | Attentional Neural Fields for Crowd CountingabstractCrowd counting has recently generated huge popularity in computer vision, and is extremely challenging due to the huge scale variations of objects. In this paper, we propose the Attentional Neural Field (ANF) for crowd counting via density estimation. Within the encoder-decoder network, we introduce conditional random fields (CRFs) to aggregate multi-scale features, which can build more informative representations. To better model pair-wise potentials in CRFs, we incorperate non-local attention mechanism implemented as inter- and intra-layer attentions to expand the receptive field to the entire image respectively within the same layer and across different layers, which captures long-range dependencies to conquer huge scale variations. The CRFs coupled with the attention mechanism are seamlessly integrated into the encoder-decoder network, establishing an ANF that can be optimized end-to-end by back propagation. We conduct extensive experiments on four public datasets, including ShanghaiTech, WorldEXPO 10, UCF-CC-50 and UCF-QNRF. The results show that our ANF achieves high counting performance, surpassing most previous methods. Fan Zhu 0001, Xiantong Zhen, Xianbin Cao 0001, Ling Shao 0001 |
ICCV | 6 |
| 2019 | Model-Free Tracking With Deep Appearance and Motion Features IntegrationabstractBeing able to track an anonymous object, a model-free tracker is comprehensively applicable regardless of the target type. However, designing such a generalized framework is challenged by the lack of object-oriented prior information. As one solution, a real-time model-free object tracking approach is designed in this work relying on Convolutional Neural Networks (CNNs). To overcome the object-centric information scarcity, both appearance and motion features are deeply integrated by the proposed AMNet, which is an end-to-end offline trained two-stream network. Between the two parallel streams, the ANet investigates appearance features with a multi-scale Siamese atrous CNN, enabling the tracking-by-matching strategy. The MNet achieves deep motion detection to localize anonymous moving objects by processing generic motion features. The final tracking result at each frame is generated by fusing the output response maps from both sub-networks. The proposed AMNet reports leading performance on both OTB and VOT benchmark datasets with favorable real-time processing speed. Peizhao Li, Xiantong Zhen, Xianbin Cao 0001 |
WACV | 4 |
| 2019 | Multi-Scale Aggregation Network for Direct Face AlignmentabstractFace alignment has been extensively researched in computer vision while remaining a challenging task. Direct face alignment based on convolutional neural networks (CNN) without relying on cascaded regression has recently emerged and achieved promising performance. In this paper, we propose a multi-scale aggregation network (MAN) for direct face alignment by aggregating features from intermediate layers of a CNN. Specifically, MAN adopts a new convolutional architecture to aggregate features at all scales in different semantic levels, which establishes highly informative facial representations for accurate alignment. Moreover, we introduce the attention mechanism into the network, which drives it to focus on the spatial regions closely related to facial landmarks for further improved performance. Our MAN achieves a general end-to-end learning architecture for multi-scale feature aggregation, which, coupled with spatial attention mechanism, is well-suited for direct face alignment. Extensive experiments conducted on four benchmark datasets, including AFLW, 300W, CelebA and 300VW, show that MAN consistently produces high performance and surpasses several state-of-the-art methods in most cases. Peizhao Li, Xiantong Zhen, Xianbin Cao 0001 |
WACV | 5 |
| 2019 | Heterogeneous pigeon-inspired optimization
Zhuxi Zhang, Jun Chen 0009, Wenbo Du 0001, Xianbin Cao 0001 |
Sci. China Inf. Sci. | 7 |
| 2019 | Object detection and tracking under Complex environment using deep learning-based LPMabstractObject detection and tracking under complex environment are challenging because of the disturbances induced by background clutter, illumination changes, occlusions and other factors. The bulk of traditional algorithms basically rely on hand‐crafted features, which are not sufficiently robust to a complex environment. Moreover, the processes of detection and tracking are separated, which leads to the overall efficiency not high. In this study, a novel local probability model (LPM)‐based mean shift (MS) algorithm is proposed to integrate object detection and tracking. The main contributions include: (i) a new framework based on the combination of LPM and MS is established for the integration of object tracking and detection. (ii) For object detection, the training and prediction of LPM are built by stacked denoising autoencoders based deep learning. (iii) For object tracking, an MS tracking algorithm leveraging LPM is modified to improve the tracking efficiency under a complex environment. Experimental results demonstrate that the proposed method is superior to the colour histograms based MS and histograms of oriented gradients based MS in terms of robustness and tracking accuracy. Yundong Li, Qichen Zhou, Xianbin Cao 0001, Zhifeng Xiao |
IET Comput. Vis. | 5 |
| 2019 | ASiam: adaptive Siamese regression tracking with adversarial template generation and motion-based failure recoveryabstractObject tracking is challenged by the varying appearances of targets and the real‐time requirement. Siamese regression trackers, being one of the most popular tracking paradigms, excel in efficiency but suffer at adaptability to cope with appearance variations. To improve their adaptability, the authors propose a new adaptive Siamese (ASiam) tracker, which integrates a novel adversarial template generation module and a motion‐based failure recovery module. The template generation module exploits the temporal coherence and evolution of target appearance variations encoded in preceding tracklets and then generates an adaptive target template online which approximates the varying target in the coming frame. This generation module is optimised via adversarial learning to achieve accurate appearance prediction and sharp template quality. The generated template, together with a search region, are fed into a Siamese tracking backbone to compute an appearance response map via dense similarity computation in a sliding‐window way. At frames where the Siamese tracking fails, the failure recovery module is invoked to perform deep frame differencing motion detection to provide a motion response map. By fusing different response maps, the drifted tracker can be re‐calibrated. Extensive experiments on the OTB2013, OTB2015, and VOT2016 datasets prove the accuracy and efficiency of the proposed tracker. Zehao Xiao, Baochang Zhang 0001, Xianbin Cao 0001 |
IET Image Process. | 4 |
| 2019 | Attentional Information Fusion Networks for Cross-Scene Power Line DetectionabstractThe power line is one of the most hazardous obstacles for low-altitude aircrafts. As aircrafts usually encounter scenes like never before during the flight, cross-scene power line detection is the key for their flight safety. However, compared to regular object detection tasks, cross-scene power line detection is extremely challenging due to its weak visual appearance and widespread existence. In this letter, we propose a cross-scene power line detection method based on attentional information fusion networks. Specifically, we construct a fully convolutional network with attention and information fusion mechanism for cross-scene detection. The two main modules make full use of the semantic and location information, which enables the model to focus more on power lines rather than the unexpected scenes. To the best of author knowledge, our method establishes the first end-to-end convolutional architecture for pixelwise power line detection. Experimental results have shown that our method outperforms previous methods by large margins for cross-scene power line detection. Yan Li 0054, Zehao Xiao, Xiantong Zhen, Xianbin Cao 0001 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2019 | Joint Tx-Rx Beamforming and Power Allocation for 5G Millimeter-Wave Non-Orthogonal Multiple Access NetworksabstractIn this paper, we investigate the combination of non-orthogonal multiple access and millimeter-wave communications (mmWave-NOMA). A downlink cellular system is considered, where an analog phased array is equipped at both the base station and users. A joint Tx-Rx beamforming and power allocation problem is formulated to maximize the achievable sum rate (ASR) subject to a minimum rate constraint for each user. As the problem is non-convex, we propose a sub-optimal solution with three stages. In the first stage, the optimal power allocation with a closed form is obtained for an arbitrary fixed Tx-Rx beamforming. In the second stage, the optimal Rx beamforming with a closed form is designed for an arbitrary fixed Tx beamforming. In the third stage, the original joint Tx-Rx beamforming and power allocation problem is reduced to a Tx beamforming problem by using the previous results, and a boundary-compressed particle swarm optimization (BC-PSO) algorithm is proposed to obtain a sub-optimal solution. Extensive performance evaluations are conducted to verify the rational of the proposed solution, and the results show that the proposed sub-optimal solution can achieve a significantly better performance in terms of ASR compared with those of the state-of-the-art schemes and the conventional mmWave orthogonal multiple access (mmWave-OMA) system. Lipeng Zhu 0001, Jun Zhang 0007, Zhenyu Xiao, Xianbin Cao 0001, Dapeng Oliver Wu, Xiang-Gen Xia 0001 |
IEEE Trans. Commun. | 4 |
| 2019 | Long-Short-Term Features for Dynamic Scene ClassificationabstractDynamic scene classification has been extensively studied in computer vision due to its widespread applications. The key to dynamic scene classification lies in jointly characterizing spatial appearance and temporal dynamics to achieve informative representation, which remains an outstanding task in the literature. In this paper, we propose a unified framework to extract spatial and temporal features for dynamic scene representation. More specifically, we deploy two variants of deep convolutional neural networks to encode spatial appearance and short-term dynamics into short-term deep features (STDF). Based on STDF, we propose using the autoregressive moving average model to extract long-term frequency features (LTFF). By combining STDF and LTFF, we establish the long-short-term feature (LSTF) representations of dynamic scenes. The LSTF characterizes both spatial and temporal patterns of dynamic scenes for comprehensive and information representation that enables more accurate classification. Extensive experiments on three-dynamic scene classification benchmarks have shown that the proposed LSTF achieves high performance and substantially surpasses the state-of-the-art methods. Yuanjun Huang, Xianbin Cao 0001, Qi Wang 0009, Baochang Zhang 0001, Xiantong Zhen, Xuelong Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2019 | Glance and Stare: Trapping Flying Birds in Aerial Videos by Adaptive Deep Spatio-Temporal FeaturesabstractFlying bird detection has recently attracted increasing attention in computer vision, which becomes an urgent task with the opening up of the low-altitude airspace. However, compared to conventional object detection tasks, it is much more challenging to trap flying birds in aerial videos due to small target sizes, complex backgrounds of great variations and disturbances of bird-like objects. In this paper, we propose a unified framework termed glance-and-stare detection (GSD) to trap flying birds in aerial videos. The GSD is inspired by the fact that human beings first glance at the whole image and then stare at the areas where the suspected object is most likely to appear until the confirmation is obtained. Specifically, we propose the zooming-in algorithm to generate region proposals for accurate localization of flying birds; to represent region proposal sequences of different lengths, we propose adaptive deep spatio-temporal features by leveraging the strength of 3D convolutional neural networks, based on which classification is conducted to achieve final detection. In contrast to conventional methods, the GSD enables localization and classification to be conducted jointly in an alternating iterative way, which mutually enhances each other to improve their performance. In order to validate the proposed GSD algorithm, we build flying bird data sets including images and videos, which provide new benchmarks for evaluation of flying bird detection systems. Experiments on the data sets demonstrate that the GSD can achieve high detection accuracy and largely outperform the state-of-the-art detection methods. Shuman Tian, Xianbin Cao 0001, Yan Li 0054, Xiantong Zhen, Baochang Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2019 | Deep Ensemble Machine for Video ClassificationabstractVideo classification has been extensively researched in computer vision due to its wide spread applications. However, it remains an outstanding task because of the great challenges in effective spatial-temporal feature extraction and efficient classification with high-dimensional video representations. To address these challenges, in this paper, we propose an end-to-end learning framework called deep ensemble machine (DEM) for video classification. Specifically, to establish effective spatio-temporal features, we propose using two deep convolutional neural networks (CNNs), i.e., vision and graphics group and C3-D to extract heterogeneous spatial and temporal features for complementary representations. To achieve efficient classification, we propose ensemble learning based on random projections aiming to transform high-dimensional features into a set of lower dimensional compact features in subspaces; an ensemble of classifiers is trained on the subspaces and combined with a weighting layer during the backpropagation. To further enhance the performance, we introduce rectified linear encoding (RLE) inspired from error-correcting output coding to encode the initial outputs of classifiers, followed by a softmax layer to produce the final classification results. DEM combines the strengths of deep CNNs and ensemble learning, which establishes a new end-to-end learning architecture for more accurate and efficient video classification. We show the great effectiveness of DEM by extensive experiments on four data sets for diverse video classification tasks including action recognition and dynamic scene classification. Results have shown that DEM achieves high performance on all tasks with an improvement of up to 13% on CIFAR10 data set over the baseline model. Jiewan Zheng, Xianbin Cao 0001, Baochang Zhang 0001, Xiantong Zhen, Xiangbo Su |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2019 | Millimeter-Wave NOMA With User Grouping, Power Allocation and Hybrid BeamformingabstractThis paper investigates the application of non-orthogonal multiple access in millimeter-Wave communications (mmWave-NOMA). Particularly, we consider downlink transmission with a hybrid beamforming structure. A user grouping algorithm is first proposed according to the channel correlations of the users. Whereafter, a joint hybrid beamforming and power allocation problem is formulated to maximize the achievable sum rate, subject to a minimum rate constraint for each user. To solve this non-convex problem with high-dimensional variables, we first obtain the solution of power allocation under arbitrary fixed hybrid beamforming, which is divided into intra-group power allocation and inter-group power allocation. Then, given arbitrary fixed analog beamforming, we utilize the approximate zero-forcing method to design the digital beamforming to minimize the inter-group interference. Finally, the analog beamforming problem with the constant-modulus constraint is solved with a proposed boundary-compressed particle swarm optimization algorithm. The simulation results show that the proposed joint approach, including user grouping, hybrid beamforming and power allocation, outperforms the state-of-the-art schemes and the conventional mmWave orthogonal multiple access system in terms of achievable sum rate, and energy efficiency. Lipeng Zhu 0001, Jun Zhang 0007, Zhenyu Xiao, Xianbin Cao 0001, Dapeng Oliver Wu, Xiang-Gen Xia 0001 |
IEEE Trans. Wirel. Commun. | 4 |
| 2018 | Deep Collaborative Tracking Networks
Xiantong Zhen, Baochang Zhang 0001, Xianbin Cao 0001 |
BMVC | 5 |
| 2018 | In Defense of Single-column Networks for Crowd Counting
Ze Wang 0008, Zehao Xiao, Qiang Qiu 0001, Xiantong Zhen, Xianbin Cao 0001 |
BMVC | 6 |
| 2018 | Attentional Alignment Networks
Baochang Zhang 0001, Xiantong Zhen, Xianbin Cao 0001 |
BMVC | 6 |
| 2018 | Modulated Convolutional NetworksabstractDespite great effectiveness of very deep and wide Convolutional Neural Networks (CNNs) in various computer vision tasks, the significant cost in terms of storage requirement of such networks impedes the deployment on computationally limited devices. In this paper, we propose new modulated convolutional networks (MCNs) to improve the portability of CNNs via binarized filters. In MCNs, we propose a new loss function which considers the filter loss, center loss and softmax loss in an end-to-end framework. We first introduce modulation filters (M-Filters) to recover the unbinarized filters, which leads to a new architecture to calculate the network model. The convolution operation is further approximated by considering intra-class compactness in the loss function. As a result, our MCNs can reduce the size of required storage space of convolutional filters by a factor of 32, in contrast to the full-precision model, while achieving much better performances than state-of-the-art binarized models. Most importantly, MCNs achieve a comparable performance to the full-precision Resnets and WideResnets. The code will be available publicly soon. Baochang Zhang 0001, Ce Li 0002, Rongrong Ji, Jungong Han, Xianbin Cao 0001, Jianzhuang Liu |
CVPR | 6 |
| 2018 | Offline and Online Search: UAV Multiobjective Path Planning Under Dynamic Urban EnvironmentabstractThis paper is concerned with path planning for unmanned aerial vehicles (UAVs) flying through low altitude urban environment. Although many different path planning algorithms have been proposed to find optimal or near-optimal collision-free paths for UAVs, most of them either do not consider dynamic obstacle avoidance or do not incorporate multiple objectives. In this paper, we propose a multiobjective path planning (MOPP) framework to explore a suitable path for a UAV operating in a dynamic urban environment, where safety level is considered in the proposed framework to guarantee the safety of UAV in addition to travel time. To this aim, two types of safety index maps (SIMs) are developed first to capture static obstacles in the geography map and unexpected obstacles that are unavailable in the geography map. Then an MOPP method is proposed by jointly using offline and online search, where the offline search is based on the static SIM and helps shorten the travel time and avoid static obstacles, while the online search is based on the dynamic SIM of unexpected obstacles and helps bypass unexpected obstacles quickly. Extensive experimental results verify the effectiveness of the proposed framework under the dynamic urban environment. Zhenyu Xiao, Xianbin Cao 0001, Xing Xi, Peng Yang 0009, Dapeng Oliver Wu |
IEEE Internet Things J. | 3 |
| 2018 | Guest Editorial Airborne Communication NetworksabstractWelcome to the IEEE JSAC special issue onAirborne Communication Networks. The goal of this special issue is to disseminate the contributions in the field of airborne communication networks. Xianbin Cao 0001, Seong-Lyun Kim, Katia Obraczka, Cheng-Xiang Wang 0001, Dapeng Oliver Wu, Halim Yanikomeroglu |
IEEE J. Sel. Areas Commun. | 1 |
| 2018 | Airborne Communication Networks: A SurveyabstractOwing to the explosive growth of requirements of rapid emergency communication response and accurate observation services, airborne communication networks (ACNs) have received much attention from both industry and academia. ACNs are subject to heterogeneous networks that are engineered to utilize satellites, high-altitude platforms (HAPs), and low-altitude platforms (LAPs) to build communication access platforms. Compared to terrestrial wireless networks, ACNs are characterized by frequently changed network topologies and more vulnerable communication connections. Furthermore, ACNs have the demand for the seamless integration of heterogeneous networks such that the network quality-of-service (QoS) can be improved. Thus, designing mechanisms and protocols for ACNs poses many challenges. To solve these challenges, extensive research has been conducted. The objective of this special issue is to disseminate the contributions in the field of ACNs. To present this special issue with the necessary background and offer an overall view of this field, three key areas of ACNs are covered. Specifically, this paper covers LAP-based communication networks, HAP-based communication networks, and integrated ACNs. For each area, this paper addresses the particular issues and reviews major mechanisms. This paper also points out future research directions and challenges. Xianbin Cao 0001, Peng Yang 0009, Mohamed Alzenad, Xing Xi, Dapeng Oliver Wu, Halim Yanikomeroglu |
IEEE J. Sel. Areas Commun. | 1 |
| 2018 | Feature Adaptation and Augmentation for Cross-Scene Hyperspectral Image ClassificationabstractCross-scene hyperspectral image (HSI) classification has recently become increasingly popular due to its crucial use in various applications. It poses great challenges to existing domain adaptation methods because of the data set shift, that is, two scenes exhibit huge distribution discrepancy. To tackle this problem, we propose a new domain adaptation method called hyperspectral feature adaptation and augmentation (HFAA) for cross-scene HSI classification. The proposed HFAA method learns a common subspace by introducing two different projection matrices to extract the transferable knowledge from the source domain to the target domain. To further enhance the common subspace representation, we propose to augment it by the feature selection strategy. HFAA can make full use of the original features from both source and target domains, and increase the similarity of the samples with the same label from the two domains. Our proposed HFAA method achieves compact but discriminative feature representations, which make it well suited for data sets with a large number of classes and huge interclass ambiguity. Experimental results on the Earth Observing 1 hyperspectral data set show that HFAA can produce state-of-the-art performance and surpass previous methods. Xianbin Cao 0001, Yan Li 0054, Dong Xu 0001 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2018 | Correlation-Based Tracking of Multiple Targets With Hierarchical Layered StructureabstractVisual target tracking is one of the most important research areas in the field of computer vision. Within this realm, multiple targets tracking (MTT) under complicated scene stands out for its great availability in real life applications, such as urban traffic surveillance and sports video analysis. However, in MTT, main difficulties arise from large variation in target saliency and significant motion heterogeneity, which may result in the failure of tracking weak targets. To tackle this challenge, a novel hierarchical layered tracking structure is proposed to perform tracking sequentially layer-by-layer. Upon this layered structure, we establish an intertarget mutual assistance mechanism on basis of intertarget correlation exploited among targets. The tracking results of a subset of targets can be utilized as additional prior information for tracking other targets. Specifically, a nonlinear motion model as well as a target interaction model basing on the intertarget correlation are proposed to effectively estimate the possible target region-of-interest to facilitate the prediction-based tracking. Moreover, the concept of motion entropy is introduced to quantitatively measure the degree of motion heterogeneity within the tracking scene for layer construction. Compared to other existing methods, extensive experiments demonstrated that the proposed method is capable of achieving higher tracking performance in complicated scenes, where targets are characterized with great heterogeneity. Xianbin Cao 0001, Pingkun Yan |
IEEE Trans. Cybern. | 1 |
| 2018 | Online Multi-Object Tracking Using Hierarchical Constraints for Complex ScenariosabstractOnline multi-object tracking (MOT) in an intelligent vehicle platform aims at locating the surrounding objects in real time, which remains far from being solved in complex scenarios, due to various motion patterns of tracked objects and severe occlusions caused by cluttered background or other objects. In this paper, we establish a unified online MOT framework for complex scenarios that employs a hierarchical model to improve the solution of data association, termed hierarchical MOT (HMOT). Incorporating the multiple Gaussians uncertainty theory into the individual motion model for each target followed by imposing interaction constraint to re-associate the tracklets with lower confidence leads our algorithm to achieve accurate multi-object tracking. With such a model, individual objects are not only more precisely associated across frames, but also dynamically constrained with each other in a global manner. Experiments on challenging data sets verify the performance of the proposed HMOT approach over the other state-of-the-art MOT methods. Xianbin Cao 0001, Yan Li 0054, Baochang Zhang 0001 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2018 | Joint Power Control and Beamforming for Uplink Non-Orthogonal Multiple Access in 5G Millimeter-Wave CommunicationsabstractIn this paper, we investigate the combination of two key enabling technologies for the fifth generation wireless mobile communication, namely millimeter-wave (mm-wave) communications and non-orthogonal multiple access (NOMA). In particular, we consider a typical two-user uplink mm-wave-NOMA system, where the base station equips an analog beamforming structure with a single radio-frequency chain and serves two NOMA users. An optimization problem is formulated to maximize the achievable sum rate of the two users while ensuring a minimal rate constraint for each user. The problem turns to be a joint power control and beamforming problem, i.e., we need to find the beamforming vectors to steer to the two users simultaneously subject to an analog beamforming structure, and meanwhile control appropriate power on them. As direct search for the optimal solution of the non-convex problem is too complicated, we propose decomposing the original problem into two sub-problems that are relatively easy to solve: one is a power control and beam gain allocation problem, and the other is an analog beamforming problem under a constant-modulus constraint. The rationale of the proposed solution is verified by extensive simulations, and the performance evaluation results show that the proposed sub-optimal solution achieves a close-to-bound uplink sum-rate performance. Lipeng Zhu 0001, Jun Zhang 0007, Zhenyu Xiao, Xianbin Cao 0001, Dapeng Oliver Wu, Xiang-Gen Xia 0001 |
IEEE Trans. Wirel. Commun. | 4 |
| 2017 | Optimum array configurations of maximum output SNR for quiescent beamformingabstractIn this paper, we consider optimum array configurations for multiple satellite signals in interference-free environment. The two measures of maximum output signal-to-noise ratio (SNR) and equal gains towards all sources incident on the array are considered for the array design. As it is computationally exhaustive to enumerate all configurations and implement eigenvalue decomposition to compare respective maximum eigenvalues, we resort to the relaxation of maximizing the lower bound of the output SNR. Subsequently, an iterative linear fractional programming method is proposed to maximize the spectral norm of the source covariance matrix. Simulation examples confirm that the array configuration plays a vital role in determining the array processing performance in interference-free scenarios. The selected optimum subarrays achieve maximum performance preservations with a dramatically reduced cost. Xiangrong Wang 0001, Moeness G. Amin, Xianbin Cao 0001 |
ICASSP | 3 |
| 2017 | Routing protocol design for drone-cell communication networksabstractThis paper is concerned with the design of routing protocol capable of congestion mitigation for drone-cells communication networks where drone-cells remain stationary in the sky as relays. All of the (distance or hop-count based) existing routing protocols can perform well when the network is lightly loaded. Once the network is heavily loaded, a large number of packets might be backlogged in queues of network nodes since these protocols can not be aware of the network congestion condition. In this paper, we propose a queuing delay and transmission delay based routing protocol (QDTD) to relieve the network congestion caused by heavily loaded traffic. First, QDTD designs a novel ForWard-Back (FWB) queue architecture that significantly reduces the number of queues maintained at each network node. Second, both queuing delay and transmission delay are leveraged as a routing metric to enhance the performance of QDTD. Experimental results show that QDTD can effectively relieve the network congestion and reduce the overall network delay and achieve high throughput. Peng Yang 0009, Xianbin Cao 0001, Zhenyu Xiao, Xing Xi, Dapeng Oliver Wu |
ICC | 2 |
| 2017 | An evolutionary approach for dynamic single-runway arrival sequencing and scheduling problem
Xiao-Peng Ji, Xianbin Cao 0001, Wenbo Du 0001, Ke Tang 0001 |
Soft Comput. | 2 |
| 2017 | Proactive Drone-Cell Deployment: Overload Relief for a Cellular Network Under Flash Crowd TrafficabstractThis paper is concerned with providing radio access network (RAN) elements (supply) for flash crowd traffic demands. The concept of multi-tier cells [heterogeneous networks (HetNets)] has been introduced in 5G network proposals to alleviate the erratic supply–demand mismatch. However, since the locations of the RAN elements are determined mainly based on the long-term traffic behavior in 5G networks, even the HetNet architecture will have difficulty in coping up with the cell overload induced by flash crowd traffic. In this paper, we propose a proactive drone-cell deployment framework to alleviate overload conditions caused by flash crowd traffic in 5G networks. First, a hybrid distribution and three kinds of flash crowd traffic are developed in this framework. Second, we propose a prediction scheme and an operation control scheme to solve the deployment problem of drone cells according to the information collected from the sensor network. Third, the software-defined networking technology is employed to seamlessly integrate and disintegrate drone cells by reconfiguring the network. Our experimental results have shown that the proposed framework can effectively address the overload caused by flash crowd traffic. Peng Yang 0009, Xianbin Cao 0001, Zhenyu Xiao, Xing Xi, Dapeng Oliver Wu |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2017 | Supervised Local Descriptor Learning for Human Action RecognitionabstractLocal features have been widely used in computer vision tasks, e.g., human action recognition, but it tends to be an extremely challenging task to deal with large-scale local features of high dimensionality with redundant information. In this paper, we propose a novel fully supervised local descriptor learning algorithm called discriminative embedding method based on the image-to-class distance (I2CDDE) to learn compact but highly discriminative local feature descriptors for more accurate and efficient action recognition. By leveraging the advantages of the I2C distance, the proposed I2CDDE incorporates class labels to enable fully supervised learning of local feature descriptors, which achieves highly discriminative but compact local descriptors. The objective of our I2CDDE is to minimize the I2C distances from samples to their corresponding classes while maximizing the I2C distances to the other classes in the low-dimensional space. To further improve the performance, we propose incorporating a manifold regularization based on the graph Laplacian into the objective function, which can enhance the smoothness of the embedding by extracting the local intrinsic geometrical structure. The proposed I2CDDE for the first time achieves fully supervised learning of local feature descriptors. It significantly improves the performance of I2C-based methods by increasing the discriminative ability of local features while greatly reducing the computational burden by dimensionality reduction to handle large-scale data. We apply the proposed I2CDDE algorithm to human action recognition on four widely used benchmark datasets. The results have shown that I2CDDE can significantly improve I2C-based classifiers and achieves state-of-the-art performance. Xiantong Zhen, Feng Zheng 0001, Ling Shao 0001, Xianbin Cao 0001, Dan Xu 0002 |
IEEE Trans. Multim. | 4 |
| 2017 | Output Constraint Transfer for Kernelized Correlation Filter in TrackingabstractThe kernelized correlation filter (KCF) is one of the state-of-the-art object trackers. However, it does not reasonably model the distribution of correlation response during tracking process, which might cause the drifting problem, especially when targets undergo significant appearance changes due to occlusion, camera shaking, and/or deformation. In this paper, we propose an output constraint transfer (OCT) method that by modeling the distribution of correlation response in a Bayesian optimization framework is able to mitigate the drifting problem. OCT builds upon the reasonable assumption that the correlation response to the target image follows a Gaussian distribution, which we exploit to select training samples and reduce model uncertainty. OCT is rooted in a new theory which transfers data distribution to a constraint of the optimized variable, leading to an efficient framework to calculate correlation filters. Extensive experiments on a commonly used tracking benchmark show that the proposed method significantly improves KCF, and achieves better performance than other state-of-the-art trackers. To encourage further developments, the source code is made available. Baochang Zhang 0001, Xianbin Cao 0001, Qixiang Ye, Chen Chen 0001, LinLin Shen, Alessandro Perina, Rongrong Ji |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2015 | Path Planning for Single Unmanned Aerial Vehicle by Separately Evolving WaypointsabstractEvolutionary algorithm-based unmanned aerial vehicle (UAV) path planners have been extensively studied for their effectiveness and flexibility. However, they still suffer from a drawback that the high-quality waypoints in previous candidate paths can hardly be exploited for further evolution, since they regard all the waypoints of a path as an integrated individual. Due to this drawback, the previous planners usually fail when encountering lots of obstacles. In this paper, a new idea of separately evaluating and evolving waypoints is presented to solve this problem. Concretely, the original objective and constraint functions of UAVs path planning are decomposed into a set of new evaluation functions, with which waypoints on a path can be evaluated separately. The new evaluation functions allow waypoints on a path to be evolved separately and, thus, high-quality waypoints can be better exploited. On this basis, the waypoints are encoded in a rotated coordinate system with an external restriction and evolved with JADE, a state-of-the-art variant of the differential evolution algorithm. To test the capabilities of the new planner on planning obstacle-free paths, five scenarios with increasing numbers of obstacles are constructed. Three existing planners and four variants of the proposed planner are compared to assess the effectiveness and efficiency of the proposed planner. The results demonstrate the superiority of the proposed planner and the idea of separate evolution. Peng Yang 0008, Ke Tang 0001, José Antonio Lozano 0001, Xianbin Cao 0001 |
IEEE Trans. Robotics | 4 |
| 2014 | Hierarchical incorporation of shape and shape dynamics for flying bird detection
Jun Zhang 0007, Qunyu Xu, Xianbin Cao 0001, Pingkun Yan, Xuelong Li 0001 |
Neurocomputing | 3 |
| 2014 | Ego motion guided particle filter for vehicle tracking in airborne videos
Xianbin Cao 0001, Changcheng Gao, Jinhe Lan, Yuan Yuan 0001, Pingkun Yan |
Neurocomputing | 1 |
| 2014 | Robust object tracking using least absolute deviation
Fuxiang Wang, Xianbin Cao 0001, Jun Zhang 0007 |
Image Vis. Comput. | 3 |
| 2013 | Tracking vehicles as groups in airborne videos
Xianbin Cao 0001, Zhengrong Shi, Pingkun Yan, Xuelong Li 0001 |
Neurocomputing | 1 |
| 2013 | Pedestrian detection in unseen scenes by dynamically updating visual words
Xianbin Cao 0001, Bo Ning 0003, Yuan Yuan 0001, Pingkun Yan |
Neurocomputing | 1 |
| 2013 | Transfer learning for pedestrian detection
Xianbin Cao 0001, Zhong Wang 0008, Pingkun Yan, Xuelong Li 0001 |
Neurocomputing | 1 |
| 2013 | Special issue on image feature detection and description
Yanwei Pang, Xianbin Cao 0001, Lei Zhang 0001, Amir Hussein |
Neurocomputing | 2 |
| 2012 | Vehicle detection and tracking in airborne videos by multi-motion layer analysis
Xianbin Cao 0001, Jinhe Lan, Pingkun Yan, Xuelong Li 0001 |
Mach. Vis. Appl. | 1 |
| 2012 | Visual Attention Accelerated Vehicle Detection in Low-Altitude Airborne Video of Urban EnvironmentabstractOne of the primary goals of the airborne vehicle detection system is to reduce the risks of incident collisions and to relieve traffic jam caused by the increasing number of vehicles. Different from the stationary systems, which are usually fixed on buildings, the airborne systems in unmanned aircrafts or satellites take the advantages of wider view angle and higher mobility. However, detecting vehicles in airborne videos is a challenging task because of the scene complexity and platform movement. The direct application of the traditional image processing techniques to the problem may result in low detection rate or cannot meet the requirements of real-time applications. To address these problems, a new and efficient method composed by two stages, attention focus extraction and vehicle classification is proposed in this paper. Our work makes two key contributions. The first is the introduction of a new attention focus extraction algorithm, which can quickly detect the candidate vehicle regions to make the algorithm focus on much smaller regions for faster computation. The second contribution is a simple and efficient classification process, which is built using the AdaBoost learning algorithm. The classification process, which is a hierarchical structure, is designed to obtain a lower false alarm rate by looking for vehicles in the candidate regions. Experimental results demonstrate that, compared with other representative algorithms, our method can obtain better performance in terms of higher detection rate and lower false positive rate, while meeting the needs of real-time application. Xianbin Cao 0001, Renjun Lin, Pingkun Yan, Xuelong Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2012 | Selecting Key Poses on Manifold for Pairwise Action RecognitionabstractIn action recognition, bag of visual words based approaches have been shown to be successful, for which the quality of codebook is critical. In a large vocabulary of poses (visual words), some key poses play a more decisive role than others in the codebook. This paper proposes a novel approach for key poses selection, which models the descriptor space utilizing a manifold learning technique to recover the geometric structure of the descriptors on a lower dimensional manifold. A PageRank-based centrality measure is developed to select key poses according to the recovered geometric structure. In each step, a key pose is selected from the manifold and the remaining model is modified to maximize the discriminative power of selected codebook. With the obtained codebook, each action can be represented with a histogram of the key poses. To solve the ambiguity between some action classes, a pairwise subdivision is executed to select discriminative codebooks for further recognition. Experiments on benchmark datasets showed that our method is able to obtain better performance compared with other state-of-the-art methods. Xianbin Cao 0001, Bo Ning 0003, Pingkun Yan, Xuelong Li 0001 |
IEEE Trans. Ind. Informatics | 1 |
| 2012 | Detection of Sudden Pedestrian Crossings for Driving Assistance SystemsabstractIn this paper, we study the problem of detecting sudden pedestrian crossings to assist drivers in avoiding accidents. This application has two major requirements: to detect crossing pedestrians as early as possible just as they enter the view of the car-mounted camera and to maintain a false alarm rate as low as possible for practical purposes. Although many current sliding-window-based approaches using various features and classification algorithms have been proposed for image-/video-based pedestrian detection, their performance in terms of accuracy and processing speed falls far short of practical application requirements. To address this problem, we propose a three-level coarse-to-fine video-based framework that detects partially visible pedestrians just as they enter the camera view, with low false alarm rate and high speed. The framework is tested on a new collection of high-resolution videos captured from a moving vehicle and yields a performance better than that of state-of-the-art pedestrian detection while running at a frame rate of 55 fps. Yanwu Xu 0001, Dong Xu 0001, Stephen Lin 0001, Tony X. Han, Xianbin Cao 0001, Xuelong Li 0001 |
IEEE Trans. Syst. Man Cybern. Part B | 5 |
| 2011 | Accelerating Vehicle Detection in Low-Altitude Airborne Urban VideoabstractThe limitation of the existing methods of traffic data collection is that they rely on techniques that are strictly local in nature. The airborne system in unmanned aircrafts provides the advantages of wider view angle and higher mobility. However, detecting vehicles in airborne videos is a challenging task because of the scene complexity and platform movement. Most of the techniques used in stationary platforms cannot perform well in this situation. A new and efficient method based on Bayes model is proposed in this paper. This method can be divided into two stages, attention focus extraction and vehicle classification. Experimental results demonstrated that, compared with other representative algorithms, our method obtained better performance with higher detection rate, lower false positive rate and faster detection speed. Xianbin Cao 0001, Renjun Lin, Pingkun Yan, Xuelong Li 0001 |
ICIG | 1 |
| 2011 | KLT Feature Based Vehicle Detection and Tracking in Airborne VideosabstractAirborne vehicle detection and tracking systems equipped on unmanned aerial vehicles (UAVs) are difficult to develop because of factors like UAV motion, scene complexity and so on. In this paper, we propose a new framework of multi-motion layer analysis to detect and track moving vehicles in airborne platform. Moving vehicles are firstly detected by registration and temporal differencing to establish motion layers. After motion layers are constructed, they are maintained over time for tracking vehicles. All vehicles are tracked by maintaining their corresponding motion layers. Our experimental results showed that compared with other previous algorithms, our method can achieve better results in terms of detection and tracking performance. Xianbin Cao 0001, Jinhe Lan, Pingkun Yan, Xuelong Li 0001 |
ICIG | 1 |
| 2011 | Linear SVM classification using boosting HOG features for vehicle detection in low-altitude airborne videosabstractVisual surveillance from low-altitude airborne platforms has been widely addressed in recent years. Moving vehicle detection is an important component of such a system, which is a very challenging task due to illumination variance and scene complexity. Therefore, a boosting Histogram Orientation Gradients (boosting HOG) feature is proposed in this paper. This feature is not sensitive to illumination change and shows better performance in characterizing object shape and appearance. Each of the boosting HOG feature is an output of an adaboost classifier, which is trained using all bins upon a cell in traditional HOG features. All boosting HOG features are combined to establish the final feature vector to train a linear SVM classifier for vehicle classification. Compared with classical approaches, the proposed method achieved better performance in higher detection rate, lower false positive rate and faster detection speed. Xianbin Cao 0001, Changxia Wu, Pingkun Yan, Xuelong Li 0001 |
ICIP | 1 |
| 2011 | Cooperative Co-evolution with Weighted Random Grouping for Large-Scale Crossing Waypoints Locating in Air Route NetworkabstractThe large-scale Crossing Waypoints Location Problem (CWLP) is a crucial problem in the design of Air Route Network (ARN). CWLP is fully non-separable and non-differentiable, and thus traditional algorithms can hardly deal with it. This paper proposes an algorithm named Cooperative Co-evolution with Weighted Random Grouping (CCWR) to tackle it. CCWR employs the weighted random (WR) grouping strategy, which is specifically designed for CWLP, to divide the large-scale Crossing Waypoints (CWs) into small sub-groups and an Evolutionary Algorithm (EA) to solve the smaller scale CWs location problem in each sub-group. Experiments on the database of the ARN in China have been carried out to evaluate the performance of CCWR. The results showed that CCWR is superior to a number of state-of-the-art algorithms, and the advanced performance of CCWR is mainly due to the WR grouping strategy. Mingming Xiao, Jun Zhang 0007, Kaiquan Cai, Xianbin Cao 0001, Ke Tang 0001 |
ICTAI | 4 |
| 2011 | Rapid pedestrian detection in unseen scenes
Xianbin Cao 0001, Zhong Wang 0008, Pingkun Yan, Xuelong Li 0001 |
Neurocomputing | 1 |
| 2011 | Vehicle Detection and Motion Analysis in Low-Altitude Airborne Video Under Urban EnvironmentabstractVisual surveillance from low-altitude airborne platforms plays a key role in urban traffic surveillance. Moving vehicle detection and motion analysis are very important for such a system. However, illumination variance, scene complexity, and platform motion make the tasks very challenging. In addition, the used algorithms have to be computationally efficient in order to be used on a real-time platform. To deal with these problems, a new framework for vehicle detection and motion analysis from low-altitude airborne videos is proposed. Our paper has two major contributions. First, to speed up feature extraction and to retain additional global features in different scales for higher classification accuracy, a boosting light and pyramid sampling histogram of oriented gradients feature extraction method is proposed. Second, to efficiently correlate vehicles across different frames for vehicle motion trajectories computation, a spatio-temporal appearance-related similarity measure is proposed. Compared to other representative existing methods, our experimental results showed that the proposed method is able to achieve better performance with higher detection rate, lower false positive rate, and faster detection speed. Xianbin Cao 0001, Changxia Wu, Jinhe Lan, Pingkun Yan, Xuelong Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2011 | An Efficient Tree Classifier Ensemble-Based Approach for Pedestrian DetectionabstractClassification-based pedestrian detection systems (PDSs) are currently a hot research topic in the field of intelligent transportation. A PDS detects pedestrians in real time on moving vehicles. A practical PDS demands not only high detection accuracy but also high detection speed. However, most of the existing classification-based approaches mainly seek for high detection accuracy, while the detection speed is not purposely optimized for practical application. At the same time, the performance, particularly the speed, is primarily tuned based on experiments without theoretical foundations, leading to a long training procedure. This paper starts with measuring and optimizing detection speed, and then a practical classification-based pedestrian detection solution with high detection speed and training speed is described. First, an extended classification/detection speed metric, named feature-per-object (fpo), is proposed to measure the detection speed independently from execution. Then, an fpo minimization model with accuracy constraints is formulated based on a tree classifier ensemble, where the minimum fpo can guarantee the highest detection speed. Finally, the minimization problem is solved efficiently by using nonlinear fitting based on radial basis function neural networks. In addition, the optimal solution is directly used to instruct classifier training; thus, the training speed could be accelerated greatly. Therefore, a rapid and accurate classification-based detection technique is proposed for the PDS. Experimental results on urban traffic videos show that the proposed method has a high detection speed with an acceptable detection rate and a false-alarm rate for onboard detection; moreover, the training procedure is also very fast. Yanwu Xu 0001, Xianbin Cao 0001, Hong Qiao |
IEEE Trans. Syst. Man Cybern. Part B | 2 |
| 2010 | Rapid classification based pedestrian detection in changing scenesabstractHow to adapt to changing scenes in pedestrian detection is a difficult problem in visual monitoring. This paper proposed a pedestrian detection method in changing scenes. Response to the requirements of high detection speed and high detection rate of pedestrian detection method in changing scenes, this paper mainly consists of two parts: (1) proposing a general ternary classification framework. It is based on cascade classification framework and each stage is a ternary detection pattern, that is, through comparing stage threshold to exclude current pedestrians or non-pedestrians object and objects which is difficult determine will enter the next layer filtering. Such detection framework is faster than traditional method and is suitable for real time pedestrian detection system. (2) Considering the above mentioned detection framework relies on thresholds, the parameters of cascade classifier which trained in old scene require adaptive adjustment in a new scenario. We design a pedestrian method in changing scenes, using a small amount of data in new scene to assist the old scene classifier, taking cross entropy method to quickly optimizing these parameters combination so that the optimized classifier can be better adapt to pedestrian detection in changing scenes. The new classifier can receive high detection rate and high detection speed. Taking AHHF dataset as an old scene and NICTA dataset as the new scene, experiments show that the proposed method can apply to pedestrian detection in new scene and obtain good results. Zhong Wang 0008, Xianbin Cao 0001 |
SMC | 2 |
| 2009 | Associated evolution of a support vector machine-based classifier for pedestrian detection
Xianbin Cao 0001, Yanwu Xu 0001, Hong Qiao |
Inf. Sci. | 1 |
| 2009 | Mr-SDM: a novel statistical deformable model for object deformation
Qizhen He, Horace Ho-Shing Ip, Jun Feng 0003, Xianbin Cao 0001 |
Vis. Comput. | 4 |
| 2008 | Multiobjective evolutionary algorithm with constraint handling for aircraft landing schedulingabstractAircraft landing scheduling is a multiobjective optimization problem with lots of constraints, which is difficult to be dealt with by traditional multiobjective evolutionary algorithms with general constraint handling strategies such as constraint-dominate definition. In this paper we pertinently designed an effective constraint handling method, and then presented a multiobjective evolutionary algorithm using the constraint handing method to solve the aircraft landing scheduling problem. Experiments show that our method is able to locate the feasible region in the search space, obtain the jagged Pareto front, and thereby provide efficient schedule for aircraft landing. Yuanping Guo, Xianbin Cao 0001, Jun Zhang 0007 |
IEEE Congress on Evolutionary Computation | 2 |
| 2008 | A multi-objective evolutionary approach to aircraft landing scheduling problemsabstractScheduling aircraft landings has been a complex and challenging problem in air traffic control for long time. In this paper, we propose to solve the aircraft landing scheduling problem (ALSP) using multi-objective evolutionary algorithms (MOEAs). Specifically, we consider simultaneously minimizing the total scheduled time of arrival and the total cost, and formulate the ALSP as a 2-objective optimization problem. A MOEA named Multi-Objective Neighborhood Search Differential Evolution (MONSDE) is applied to solve the 2-objective ALSP. Besides, a ranking scheme named non-dominated average ranking is also proposed to determine the optimal landing sequence. Advantages of our approaches are demonstrated on two example scenarios. Ke Tang 0001, Zai Wang, Xianbin Cao 0001, Jun Zhang 0007 |
IEEE Congress on Evolutionary Computation | 3 |
| 2008 | A Low-Cost Pedestrian-Detection System With a Single Optical CameraabstractThe ultimate purpose of a pedestrian-detection system (PDS) is to reduce pedestrian-vehicle-related injury. Most such systems tend to adopt expensive sensors, such as infrared devices, in expectation of better performance. In comparison, a low-cost optical-camera-based system has much potential practical value, including a greater detection range, and can easily be trained to detect other objects. However, such low-cost systems are difficult to design (e.g., little original information can be collected, and the scene is very complex). To address these problems, an effective and reliable classifier is needed. The classifier should have a proper structure, its features need to be well selected, and a large number of high-quality samples are necessary for training. In this paper, we present a low-cost PDS which only uses a single optical camera. We design a cascade classifier to achieve an effective and reliable detection. First, our system scans two sequential frames at each zoom scale with a sliding window. Second, with each window, both appearance and motion features are extracted. A well-trained cascade classifier, combining statistical learning with a decomposed support-vector-machine classifier, then determines whether the window contains a human body. At the same time, to provide as much information as possible about the pedestrian, a small-scale weighted template tree trained by a coevolutionary algorithm is adopted to identify each pedestrian's direction, and the distance of each from the vehicle is also provided using an estimation algorithm. During the training procedure, we select key features by using the AdaBoost algorithm and a large number of high-quality samples. Experimental results demonstrate that the system is suitable for pedestrian detection in city traffic: The detection speed is more than 10 ft/s, the detection rate reaches 80%, and the false positive rate is no more than 0.30/00. Xianbin Cao 0001, Hong Qiao, John A. Keane |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2007 | How Contents Influence Clustering Features in the WebabstractIn World Wide Web, contents of web documents play important roles in the evolution process because of their effects on linking preference. A majority of topological properties are content-related, and among them the clustering features are sensitive to contents of Web documents. In this paper, we first observe the impacts of content similarity on web links by introducing a metric called Linkage Probability. Then we investigate how contents influence the formation mechanism of the most basic cluster, triangle, with a metric named Triangularization Probability. Experimental results indicate that content similarity has a positive function in the process of cluster formation in theWeb. Theoretical analysis predicts the contents influence on the clustering features in the Web very well. Xueqi Cheng 0001, Fuxin Ren, Xianbin Cao 0001 |
Web Intelligence | 3 |
| 2007 | Modeling the Evolution of Web using Vertex Content SimilarityabstractIn the evolution process of World Wide Web, contents of web pages play important roles because of their direct effect on linking preference. In this paper, we propose a model which combines vertex connectivity and content similarity in a proportional manner. Analytical solutions indicate that our model exhibits a power-law degree distribution with variable exponent determined by the weight of content similarity. Distribution of content similarity on connected vertex pairs shows content similar web pages trend to be linked together. Simulation results show our model yields remarkably agreements of both degree and content similarity distributions with real network. Xianbin Cao 0001, Yuanping Guo, Xueqi Cheng 0001 |
Web Intelligence | 2 |
| 2006 | An Evolutionary Support Vector Machines Classifier for Pedestrian DetectionabstractIn a pedestrian detection system, a classifier is usually designed to recognize whether a candidate is a pedestrian. Support vector machines (SVM) has become a primary technique to train a classifier for pedestrian detection. However, it is hard to give the best training model which has a tremendous effect to the performance of a SVM classifier. In this paper, we design special code/decode scheme and evaluation function for a training model firstly; and then use genetic algorithm to optimize key parameters which represent the SVM training model. Therefore a most suitable SVM classifier can be obtained for pedestrian detection. Experiments have been carried out in a single camera based pedestrian detection system. The results show that the evolutionary SVM classifier has a better detection rate; moreover, RBF kernel is more suitable than polynomial kernel when chosen in an evolutionary SVM classifier for pedestrian detection Xianbin Cao 0001, Yanwu Xu 0001, Hong Qiao |
IROS | 2 |
| 2006 | A Multiclass Classifier to Detect Pedestrians and Acquire Their Moving Styles
Xianbin Cao 0001, Hong Qiao, Fei-Yue Wang 0001 |
ISI | 2 |
| 2006 | Fast Pedestrian Detection Using Color Information
Yanwu Xu 0001, Xianbin Cao 0001, Hong Qiao, Fei-Yue Wang 0001 |
ISI | 2 |
| 2005 | Application of Cooperative Co-evolution in Pedestrian Detection Systems
Xianbin Cao 0001, Hong Qiao, Fei-Yue Wang 0001 |
ISI | 1 |
| 2005 | Application of a Decomposed Support Vector Machine Algorithm in Pedestrian Detection from a Moving Vehicle
Hong Qiao, Fei-Yue Wang 0001, Xianbin Cao 0001 |
ISI | 3 |
| 2002 | An immune genetic algorithm based on immune regulationabstractUsing immune regulation mechanisms that include density regulation and network regulation, this paper proposes a novel immune genetic algorithm. Its core idea is that all individuals compose an antibody network, and it utilizes a density regulation mechanism to adjust individual diversity at an individual level and network regulation mechanism to achieve dynamic balance between individual diversity and population convergence. The dynamic regulative ability of this algorithm is analyzed and the approach to choosing the parameters is also given. As a novel adaptive resolving algorithm, it can be used to solve many complex optimization problems. This paper discusses solution of the frequency assignment problem and analyzes the parameters' influence upon the performance of this algorithm. The experimental results prove that this algorithm has good performance and can properly maintain the balance between individual diversity and population convergence. Wenjian Luo, Xianbin Cao 0001, Xufa Wang |
IEEE Congress on Evolutionary Computation | 2 |
| 2001 | NIDS Research Based on Artificial Immunology
Wenjian Luo, Xianbin Cao 0001, Xufa Wang |
ICICS | 2 |