EDBT 2026 Demo / reviewers in the wild / expert
Shengjie Zhao 0001
dblp:47/5183-1
· DBLP profile ↗
155ranked-venue papers
12as first author
101since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 52 · 5 first-author · 30 since 2021Graphics, computer vision, multimedia, augmented reality and games · 49 · 3 first-author · 29 since 2021Artificial intelligence and machine learning · 29 · 27 since 2021Applied, interdisciplinary, general and emerging computing · 20 · 19 since 2021Human-computer interaction and ubiquitous computing · 9 · 9 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 3 since 2021Systems, architecture and hardware · 1Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DEPO: Dual-Efficiency Preference Optimization for LLM AgentsabstractRecent advances in large language models (LLMs) have greatly improved their reasoning and decision-making abilities when deployed as agents. Richer reasoning, however, often comes at the cost of longer chain of thought (CoT), hampering interaction efficiency in real-world scenarios. Nevertheless, there still lacks systematic definition of LLM‑Agent efficiency, hindering targeted improvements. To this end, we introduce dual‑efficiency, comprising (i) step-level efficiency, which minimizes tokens per step, and (ii) trajectory-level efficiency, which minimizes the number of steps to complete a task. Building on this definition, we propose DEPO, a dual-efficiency preference‑based optimization method that jointly rewards succinct responses and fewer action steps. Experiments on WebShop and BabyAI show that DEPO cuts token usage by up to 60.9% and steps by up to 26.9%, while achieving up to a 29.3% improvement in task performance. DEPO also generalizes to three out-of-domain math benchmarks and retains its efficiency gains when trained on only 25% of the data. Mengshi Zhao, Yuying Zhao, Beier Zhu, Hanwang Zhang, Shengjie Zhao 0001, Chaochao Lu |
AAAI | 7 |
| 2026 | Uncovering Pretraining Code in LLMs: A Syntax-Aware Attribution ApproachabstractAs large language models (LLMs) become increasingly capable, concerns over the unauthorized use of copyrighted and licensed content in their training data have grown, especially in the context of code. Open-source code, often protected by open source licenses (e.g, GPL), poses legal and ethical challenges when used in pretraining. Detecting whether specific code samples were included in LLM training data is thus critical for transparency, accountability, and copyright compliance. We propose SynPrune, a syntax-pruned membership inference attack method tailored for code. Unlike prior MIA approaches that treat code as plain text, SynPrune leverages the structured and rule-governed nature of programming languages. Specifically, it identifies and excludes consequent tokens that are syntactically required and not reflective of authorship, from attribution when computing membership scores. Experimental results show that SynPrune consistently outperforms the state-of-the-arts. Our method is also robust across varying function lengths and syntax categories. Yuanheng Li, Zhuoyang Chen, Xiaoyun Liu, Mingwei Liu 0002, Yang Shi 0002, Kaifeng Huang 0001, Shengjie Zhao 0001 |
AAAI | 8 |
| 2026 | On the Localization Probability of RIS-Assisted Systems: A Stochastic Geometry Perspective
Junqi Guo, Junyuan Wang 0001, Shengjie Zhao 0001 |
ICC | 5 |
| 2026 | Scaling Multimodal Retrieval and Generation for Long Documents through Visual Tiling and Context CompressionabstractMultimodal retrieval-augmented generation (MRAG) provides a powerful paradigm for long-document reasoning by integrating retrieval with generative modeling. However, scaling MRAG to multi-page documents remains challenging due to noisy cross-modal retrieval and the quadratic computational cost of long-context inference. Existing multimodal large language models (MLLMs) lack mechanisms to structure visual representations and efficiently manage contextual memory, limiting their scalability and generalization. We propose LMDocRag, a principled framework for efficient multimodal retrieval and generation over long documents. Our approach is based on the insight that document understanding benefits from structured visual decomposition and representation-level compression. Specifically, we introduce a document tiling strategy that transforms long document images into semantically localized visual units, enabling unified hybrid retrieval in a structured embedding space. Building on this decomposition, we develop a layout-guided visual token sparsification method that learns to preserve structurally salient regions while suppressing redundant visual representations, and a KV-cache compression scheme that reduces autoregressive memory growth by selectively retaining informative contextual states. Extensive experiments on multimodal long-document benchmarks demonstrate that LMDocRag improves question-answering accuracy while substantially reducing visual token count and inference complexity. Weichao Chen 0001, Shengjie Zhao 0001 |
ICMR | 3 |
| 2026 | GMKAT: Geospatial Multipath Kolmogorov-Arnold Transformer for flood susceptibility mapping
Minzhen Cao, Hao Deng 0002, Hongwei Dai, Shengjie Zhao 0001 |
Appl. Intell. | 4 |
| 2026 | A personalized active learning strategy with enhanced user satisfaction for recommender systems
Siwei Qian, Jie Wang 0016, Shengjie Zhao 0001 |
Expert Syst. Appl. | 3 |
| 2026 | RSUTrajRec: Multi-granularity trajectory recovery based on roadside units sensing
Xianjing Wu, Xutao Chu, Jianyu Wang 0003, Shengjie Zhao 0001 |
Expert Syst. Appl. | 4 |
| 2026 | SETFusion: A semantic transformer for infrared and visible image fusion
Wei Tang 0018, Fazhi He, Lin Zhang 0014, Shengjie Zhao 0001 |
Pattern Recognit. | 4 |
| 2026 | CCTformer: Calibrated Context-Aware Transformer for Correspondence PruningabstractCorrespondence pruning is a core problem in computer vision, aiming to distinguish correct matches from a large set of initial correspondences. Recent Transformer-based methods, such as VSFormer, have shown remarkable performance by jointly modeling visual appearance and spatial relationships. Nevertheless, these methods often depend on abstract, scene-level visual representations, which are insufficiently precise for validating individual correspondences. In addition, their context aggregation modules often struggle to capture multi-granularity corresponding features, particularly under complex geometric variations. To overcome these limitations, we propose the Calibrated Context-aware Transformer with two key components. The Calibrated Visual Cues Extractor samples features directly at putative match locations to provide well-aligned, correspondence-specific visual cues. The Dual-Perspective Graph Transformer replaces conventional graph attention to better model both node-level and edge-level interactions. Together, these modules enable more robust local-global context fusion. Extensive experiments on YFCC100 M and SUN3D show that CCTformer consistently outperforms the baseline and other state-of-the-art methods. Shengjie Zhao 0001, Yizhang Liu |
IEEE Signal Process. Lett. | 2 |
| 2026 | HT-DRIVE: Heterogeneous-Temporal GNN-Aided Multi-Agent Reinforcement Learning for Resource Allocation in Platoon-Based V2X NetworksabstractPlatoon-based Vehicle-to-Everything (V2X) is a promising communication paradigm for intelligent vehicular networks. In such networks, each platoon leader (PL) engages in a high-throughput vehicle-to-infrastructure (V2I) link to deliver entertainment data, while maintaining low-latency and high-reliability vehicle-to-vehicle (V2V) links with multiple platoon members (PMs) to convey safety-critical information, thus posing conflicting performance requirements. Besides, the temporal variations and co-channel interferences further complicate the problem. To address these issues, we propose a heterogeneity-aware and dynamics-adaptive resource allocation framework, termed HT-DRIVE, which integrates an advanced heterogeneous temporal graph neural network (HTGNN) with a tailored multi-agent reinforcement learning (MARL) mechanism. Specifically, we design an HTGNN module to capture heterogeneous and temporal patterns within the platoon-based V2X network, generating structured embeddings that serve as state inputs to a modified multi-agent twin delayed deep deterministic policy gradient (MATD3) algorithm for downstream decision-making. To facilitate end-to-end optimization, several enhancements are also incorporated, including a hybrid action design to jointly handle discrete subchannel selection and continuous power control, a dual-critics architecture with reduced computational complexity, and customized loss functions for collaborative training. Extensive simulations demonstrate that HT-DRIVE significantly outperforms MATD3, multi-agent deep deterministic policy gradient (MADDPG) and heuristic methods under various settings, validating its effectiveness and robustness. Yanming Huang, Fengxia Han, Shengjie Zhao 0001 |
IEEE Trans. Commun. | 3 |
| 2026 | Improving Hate Speech Detection via Robust Knowledge DistillationabstractHate speech proliferating on social media disrupts the harmony of the Internet, making its detection a challenging task in natural language processing. Despite the recent advances in hate speech detection based on pre-trained language models (PLMs), their large parameter scale limits their applicability on resource-constrained devices, and the substantial noise and data imbalance in online speech severely weaken the model’s ability to extract valid information. To address these issues, we propose a novel robust knowledge distillation framework for improving hate speech detection. This framework aims to leverage the knowledge of a complex teacher model to supervise the training of a compact student model. Specifically, we perform linguistically motivated denoising to mitigate the impact of non-semantic noise on model training, and then transfer the fine-tuned teacher’s knowledge to the student for model compression. To handle data imbalance, we incorporate focal loss into knowledge distillation to enhance the model’s discriminative capability on hard samples. Moreover, we design a dual-round data augmentation strategy to improve model robustness against noisy and perturbations. Experiments on the SE and DV hate speech datasets reveal that our framework boosts the lightweight model’s F1-score by up to 4.3% over the state-of-the-art baseline, while reducing inference time by nearly 40% and resource consumption by 35%–50%. Yongjie Gui, Hao Deng 0002, Shengjie Zhao 0001 |
IEEE Trans. Comput. Soc. Syst. | 3 |
| 2026 | MSFormer: Multi-Scale Transformer With Hierarchical and Local Awareness for Traffic Flow Prediction
Shilong Dong, Shengjie Zhao 0001, Jiafeng Huang, Kenan Ye, Wenzhen Jia |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2026 | MPVFedLoc: Enabling Pervasive Indoor Localization Through Multi-Perspective Views and Distributed Federated LearningabstractWith camera-equipped phones becoming essential items carried by people, utilizing multi-perspective views (MPV) captured from the surrounding environment has emerged as a promising approach to achieve pervasive localization in indoor environments. This MPV-based localization requires extensive data, necessitating the use of crowdsourcing for data collection and training. In this distributed process, multiple clients can process and share results, raising concerns about privacy breaches. To address this challenge, this paper introduces federated learning (FL) into MPV-based localization, resulting in theMPVFedLocalgorithm, which facilitates distributed learning without exchanging raw local data. However, FL-based methods often experience reduced accuracy due to data heterogeneity. To overcome this, we propose a model self-supervised federated learning framework withinMPVFedLoc. This framework integrates self-supervised learning at the model level and incorporates a model self-supervised loss into the local training objective to mitigate the bias between the global and local models. To evaluate the performance of MPV-based localization, we construct a benchmark dataset named TJF_Building and conduct extensive experiments. Additionally, we compare the performance ofMPVFedLocwith three state-of-the-art FL methods in both homogeneous and heterogeneous settings. Experimental results demonstrate the robustness and effectiveness ofMPVFedLoc, particularly in handling heterogeneous scenarios. Further experiments on two typical image classification datasets also highlight the potential ofMPVFedLocfor diverse tasks. Junyuan Wang 0001, Feng Yan 0004, Shengjie Zhao 0001 |
IEEE Trans. Mob. Comput. | 5 |
| 2026 | Learning Local Semantic Signals and Inter-Class Discrepancy for Weakly Supervised Video Anomaly DetectionabstractWeakly supervised video anomaly detection (WSVAD) aims to retrieve temporal intervals containing anomalous events within untrimmed videos, leveraging video-level annotations. The existing methodologies exhibit unsatisfactory performance, primarily attributed to the absence of frame-level annotations. To alleviate this issue, we argue that representations exhibiting high similarity within local regions can offer dependable semantic knowledge. With this insight, we introduce a novel pseudo-label optimization mechanism grounded in the similarity of local features, specifically designed to direct framelevel anomaly predictions. Besides, we have noticed that the majority of videos utilized for anomaly detection originate from surveillance scenarios and are captured by stationary cameras. Consequently, the frames within a given video share consistent background properties. Motivated by this observation, we propose an intra-frame background-foreground separation strategy for each frame to extract discriminative visual representations. Given the considerable similarity in the backgrounds of most frames, the distinctions among the features of diverse frames become subtly indistinct. As a result, to ensure the rationality of predictions, we encourage maximizing the inter-class variance between normal and abnormal frames. Extensive experiments and ablation studies, encompassing both coarse-grained and fine-grained, have been conducted on XD-Violence, UCF-Crime, and ShanghaiTech benchmark datasets. The results demonstrate that our proposed method achieves substantial improvements, outperforming the current state-of-the-art approaches. Yu Wang 0174, Shengjie Zhao 0001, Jianyu Wang 0003, Xutao Chu |
IEEE Trans. Multim. | 2 |
| 2026 | Targeted Mining of Time-Interval Related PatternsabstractCompared with frequent pattern mining, sequential pattern mining emphasizes the temporal aspect and finds broad applications across various fields. However, numerous studies treat temporal events as single time points, neglecting their durations. Time-interval-related pattern (TIRP) mining is introduced to address this issue and has been applied to healthcare analytics, stock prediction, etc. Typically, mining all patterns is not only computationally challenging for accurate forecasting but also resource-intensive in terms of time and memory. Targeting the extraction of TIRPs based on specific criteria can improve data analysis efficiency and better align with customer pReferences. Therefore, this article proposes a novel algorithm called TaTIRP to discover targeted TIRP. In addition, we develop multiple pruning strategies to eliminate redundant extension operations, thereby enhancing performance on large-scale datasets. Finally, we conduct experiments on various real-world and synthetic datasets to validate the accuracy and efficiency of the proposed algorithm. Shuang Liang 0001, Wensheng Gan, Philip S. Yu, Shengjie Zhao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2026 | Robust Traffic Forecasting With Disentangled Spatiotemporal Graph Neural NetworksabstractTraffic prediction is a cornerstone of intelligent transportation systems (ITSs). The effectiveness of existing spatiotemporal graph neural networks (STGNNs) heavily relies on the independent identically distributed (i.i.d.) assumption of traffic data, which is frequently violated in practice because of distribution shifts owing to exogenous factors. While learning features that remain stable across all environments is promising for modeling robust frameworks, the fundamental challenge involves the decomposition of invariant features from the dynamic nature of spatiotemporal dependencies. In this article, we propose the disentangled spatiotemporal (DIST) graph neural networks, a novel framework for robust traffic forecasting considering distribution shifts. In DIST, latent invariant variables are explicitly decoupled from dynamically evolving spatiotemporal dependencies, enabling the learning of topology-agnostic representations resilient to distribution shifts. Specifically, we formulate a causality-driven learning objective that guides the separation of invariant variables from various exogenous factors. We then propose a spatiotemporal graph modeling module that can adaptively capture spatiotemporal dependencies in evolving traffic systems. Furthermore, we present a graph perturbation module to simulate topology variations during training, thereby encouraging the model to identify perturbation-sensitive dependencies and infer invariant and variant features for prediction and intervention tasks. The prediction risk and its variance on multiple interventional distributions are minimized in our learning strategy, allowing the model to identify invariant features, thus improving its robustness. The results of comprehensive real-world experiments demonstrate the superiority of our approach. The source code is available: https://github.com/tingwang25/DIST. Ting Wang 0019, Rui Luo 0002, Daqian Shi, Hao Deng 0002, Shengjie Zhao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2026 | MECI: A Multi-Model Motion Capture-Free Event Dataset Featuring Large-Scale Challenging Indoor EnvironmentsabstractRecently, several event-based datasets have emerged to foster the application of the new event camera to classic vision tasks like Simultaneous Localization and Mapping (SLAM). However, current indoor benchmark datasets depend on the expensive motion capture system to obtain ground-truth trajectories, restricting data acquisition to small object-centric scenes or single-room environments due to infrastructure costs and spatial limitations. Furthermore, these datasets lack sensor diversity, relying solely on a single event camera model that hinders practical cross-device generalization. To address the above limitations, we propose MECI, the first Multi-model Event dataset targeting Challenging Indoor environments, especially including large-scale scenes with complete trajectory ground-truth provided. Specifically, MECI includes 38 posed sequences of visual-inertial-event data from three different models of event cameras with varying resolutions and data frequencies. Apart from typical challenging factors such as illumination changes, motion blur, and dynamics, these sequences involve large-scale indoor scenes (across rooms and floors) with each room occupying approximately 60 m \({}^{2}\) and a maximum trajectory length of 343.8 m. We innovatively leverage ETS (Electronic Total Station) measurements and AprilTag markers to provide complete 6-DoF trajectories with low cost and minimal environmental modifications. Extensive experiments show that our MECI is an effective yet challenging benchmark dataset not only for visual SLAM, but also for event-image reconstruction. Our project page is https://cslinzhang.github.io/MECI_dataset/ . Yang Chen 0037, Lin Zhang 0014, Shengjie Zhao 0001, Yicong Zhou |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2026 | CaneSpeaker: An LLM-Assisted Speaker for Generating Human-Like Navigation InstructionsabstractNavigation instruction generation aims to address data scarcity in Vision-and-Language Navigation (VLN) by generating navigation instructions for unannotated routes from data sources like simulators or online data. However, existing methods usually suffer from high reliance on panoramic views, poor cross-task generalization ability, and limited availability of training data. To address these challenges, we propose a novel speaker, CaneSpeaker, to generate human-like instructions from front-facing images for a variety of VLN tasks. First, to mitigate the limited amount of speaker training data, we propose an Large Language Model (LLM)-based instruction augmentation method, LLM-IA, that utilizes an off-the-shelf LLM to create augmented instructions for training by distilling and reformulating existing instructions. This method allows us to collect an instruction-augmented dataset with human-level accuracy for speaker training, namely Rx2R. Second, to eliminate the dependency on panoramic views, we propose a novel Vision-Language Model (VLM)-based speaker architecture, VL-Sp. By leveraging the advanced reasoning capabilities of a pre-trained VLM, CaneSpeaker can effectively generate high-quality instructions directly from front-facing images without relying on panoramic views. Also, the prompt-based characteristic of the VLM allows us to devise a unified input representation to enable the processing of multiple VLN tasks, thus further addressing the problem of data scarcity by combining multiple datasets from different VLN tasks. Finally, we utilize CaneSpeaker to synthesize a large-scale augmented dataset, CANE, from unannotated routes in the Matterport3D Simulator. Comprehensive experiments demonstrate that CaneSpeaker generates precise instructions with diverse expressions across various VLN tasks, and the VLN agent trained on our datasets obviously outperforms its counterparts. The source codes and datasets are available at https://github.com/zheng19845/CaneSpeaker . Yuanyu Zheng, Lin Zhang 0014, Yunda Sun, Ying Shen 0005, Shengjie Zhao 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2026 | Learning from Rendering: Realistic and Controllable Extreme Rainy Image Synthesis for Autonomous Driving SimulationabstractAutonomous driving simulators provide an effective and low-cost alternative for evaluating or enhancing visual perception models. However, the reliability of evaluation depends on the diversity and realism of the generated scenes. Extreme weather conditions, particularly extreme rainfalls, are rare and costly to capture in real-world settings. While simulated environments can help address this limitation, existing rainy image synthesizers often suffer from poor controllability over illumination and limited realism, which significantly undermines the effectiveness of the model evaluation. To that end, we propose a learning-from-rendering rainy image synthesizer, which combines the benefits of the realism of rendering-based methods and the controllability of learning-based methods. To validate the effectiveness and generalizability of our extreme rainy image synthesizer on the semantic segmentation task, a continuous set of pixel-accurately labeled extreme rainy images is necessary. By integrating the proposed synthesizer with the CARLA driving simulator, we develop CARLARain—an extreme rainy street scene simulator which can obtain paired rainy-clean images and labels under complex illumination conditions. Qualitative and quantitative experiments validate that CARLARain can effectively improve the accuracy of semantic segmentation models in extreme rainy scenes, with the models’ accuracy (mIoU) improved by \(5{-}8\%\) on the synthetic dataset and significantly enhanced in real extreme rainy scenarios under complex illuminations. Our source code and datasets are available at https://kb824999404.github.io/HRIG/ . Kaibin Zhou, Kaifeng Huang 0001, Hao Deng 0002, Zelin Tao, Ziniu Liu, Lin Zhang 0014, Shengjie Zhao 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 7 |
| 2025 | DFFL: Federated Learning with Data FairnessabstractFederated Learning (FL) is a promising solution for distributed cooperative training in various scenarios. FL empowers data-rich devices to contribute to a comprehensive model, but it suffers from the data fairness problem, where local datasets are not fairly accounted for in the global model. This problem is especially devastating in the scenario of heterogeneous devices with non-lID data distributions. To resolve this issue, we propose Federated Learning with Data Fairness (DFFL), which incorporates model factorization to adapt local models of different sizes to devices of various computation capabilities, and a novel DNM metric for client selection to boost its performance on devices with minor data distributions. Experiments show that, the proposed DFFL outperforms existing FL schemes, achieving higher accuracy with much lower (82 % less local parameters on average) computation load, while providing a more robust assurance of data fairness among devices. Yixun Gu, Jie Wang 0016, Shengjie Zhao 0001 |
CSCWD | 4 |
| 2025 | Self-Regeneration: Step-Level Regeneration Improves the LLM's Causal Reasoning AbilityabstractLarge Language Models (LLMs) have performed remarkably on different tasks. The discovery of self-evolution approaches reveals that LLMs can think like humans which improves their reasoning abilities without external assistance. In this paper, we introduce a step-level framework called Self-Regeneration that can assist LLMs in deciding the consistency of each step by regenerating and giving out the confidence of responses. We select the casual reasoning domain and test Self-Regeneration on some popular LLMs. The result shows our approach improves LLMs' reasoning ability and increases the accuracy of responses. Wenlin Jiang, Shengjie Zhao 0001, Hongwei Dai |
CSCWD | 2 |
| 2025 | GLFMamba-U: Global-Local Fused Mamba-Unet
Ziniu Liu, Fengxia Han, Daqiang Zhang 0001, Mingqing Liu 0002, Hao Deng 0002, Shengjie Zhao 0001 |
ICANN (1) | 8 |
| 2025 | Classifier Recalibration for Human-Object Interaction Detection
Shuwei Yan, Shuang Liang 0001, Kenan Ye, Baihua Liu, Chi Xie 0001, Shengjie Zhao 0001 |
ICIC (6) | 6 |
| 2025 | ESTJ: Efficient Semantic Segmentation via Token Joint MergingabstractVision Transformers (ViTs) leverage the attention mechanism for feature extraction but often suffer from high computational costs. To address this issue, prior works have introduced token reduction methods involving fixed-window local merging and global Bipartite Matching. However, these methods face significant challenges, such as insufficient merging due to fixed-size local windows and incorrect merging of informative tokens in global merging. To overcome these limitations, we propose Efficient Semantic Segmentation via Token Joint Merging (ESTJ) for ViT-based semantic segmentation networks. Specifically, ESTJ merges tokens using two strategies: Hierarchical Condition Pooling (HCP), which employs hierarchical local windows to effectively select sufficient tokens, and Protected Bipartite Matching (PBM), designed to preserve informative tokens using average similarity between a token and all other tokens. Experimental results demonstrate that ESTJ improves throughput by 75%, reduces GFLOPs by 40%, and enhances mIoU by up to 1.1%. Moreover, ESTJ can adjust the merging threshold during inference to adapt to scenarios that prioritize efficiency or accuracy. Compared to existing methods, ESTJ achieves a better balance between computational efficiency and segmentation accuracy. Ziniu Liu, Mingqing Liu 0002, Fengxia Han, Xingtong Liu, Hao Deng 0002, Shengjie Zhao 0001 |
ICME | 8 |
| 2025 | SAVE-GSL: Scalable and Expressive Graph Structure Learning for Large GraphsabstractGraph structure learning (GSL) has emerged as a promising approach for optimizing graph structures to enhance downstream task performance. However, the quadratic complexity of GSL renders it impractical for large graphs, such as social networks. While several attempts have been made to mitigate the scalability issue of GSL, they struggle to ensure efficiency and expressiveness simultaneously. In this paper, we propose a novel scalable and expressive graph structure learning framework (SAVE-GSL), that models all-pair interactions with linear complexity. Specifically, we design cluster and expander sparse patterns for efficient expressiveness enhancement, and adaptively fuse them to generate the optimized graph. Moreover, leveraging these sparse patterns, we theoretically prove that SAVE-GSL is an efficient universal approximator of permutation-equivariant functions, providing a formal justification for its superior expressiveness. Extensive experiments demonstrate that SAVE-GSL outperforms state-of-the-art schemes in both efficiency and accuracy on social network datasets and other graph datasets of various scales. Manxin Xu, Shengjie Zhao 0001, Jin Zeng 0004, Weichao Chen 0001, Shilong Dong |
ICME | 2 |
| 2025 | Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level ComputersabstractContrastive Language-Image Pre-training (CLIP) has attracted a surge of attention for its superior zero-shot performance and excellent transferability to downstream tasks. However, training such large-scale models usually requires substantial computation and storage, which poses barriers for users with consumer-level computers. Motivated by this observation, in this paper we investigate how to achieve competitive performance on a single Nvidia RTX3090 GPU and with one terabyte of storage for the dataset. On one hand, we simplify the transformer block structure and combine Weight Inheritance with multi-stage Knowledge Distillation (WIKD), thereby reducing the number of parameters and improving the inference speed during training as well as deployment. On the other hand, confronted with the convergence challenge posed by a limited dataset, we generate synthetic captions for each sample as data augmentation, and devise a novel Pair Matching (PM) loss to fully exploit the distinction among positive and negative image-text pairs. Extensive experiments demonstrate that our model can achieve a new state-of-the-art datascale-parameter-accuracy tradeoff, which could further popularize the CLIP model and enable its deployment in consumer devices. Shengjie Zhao 0001, Weichao Chen 0001, Deniz Gündüz |
IJCNN | 2 |
| 2025 | MF2former: Multi-Feature Fusion Transformer for Traffic Flow PredictionabstractIn urban planning, traffic flow prediction, a core component of Intelligent Transportation Systems, has made significant progress with the development of deep learning. The key problem of traffic flow prediction lies in capturing the complex spatio-temporal correlations in traffic flow. In recent years, more and more research has tended to apply Transformer-based models to solve this problem. However, Transformer-based models have two major limitations for traffic flow prediction: i) Most methods only focus on extracting data features within the attention head, while ignoring the correlation between these heads, making it difficult to integrate the multi-features of traffic data; ii) Most methods do not recognize the unique impact of nodes that serve as pivotal traffic hubs in the traffic networks, which cause the Transformer to excessively focus on the influence of non-pivotal nodes. In this study, we propose a novel Transformer-based model, the Multi-Feature Fusion Transformer (MF2former), aimed at addressing the above limitations of traffic flow prediction. MF2former incorporates the Augmented Synergistic Transformer Module to achieve a comprehensive multi-feature fusion of traffic data by enhancing the information capacity within self-attention heads and performing information fusion between these heads. Additionally, our model incorporates the Pivotal Node Module, which extracts pivotal nodes from all nodes and masks the global receptive field of the Transformer to enhance its focus on pivotal nodes in the traffic network. Our model is evaluated using two real-world traffic datasets, demonstrating superior performance compared to existing methods. This study provides a stable framework for accurate traffic flow prediction, offering valuable insights for urban planners and commuters. Shengjie Zhao 0001, Shilong Dong, Jiafeng Huang, Yuhang Wan, Wenzhen Jia, Hao Deng 0002 |
IJCNN | 2 |
| 2025 | ShotBench: Expert-Level Cinematic Understanding in Vision-Language ModelsabstractRecent Vision-Language Models (VLMs) have shown strong performance in general-purpose visual understanding and reasoning, but their ability to comprehend the visual grammar of movie shots remains underexplored and insufficiently evaluated. To bridge this gap, we present \textbf{ShotBench}, a dedicated benchmark for assessing VLMs’ understanding of cinematic language. ShotBench includes 3,049 still images and 500 video clips drawn from more than 200 films, with each sample annotated by trained annotators or curated from professional cinematography resources, resulting in 3,608 high-quality question-answer pairs. We conduct a comprehensive evaluation of over 20 state-of-the-art VLMs across eight core cinematography dimensions. Our analysis reveals clear limitations in fine-grained perception and cinematic reasoning of current VLMs. To improve VLMs capability in cinematography understanding, we construct a large-scale multimodal dataset, named ShotQA, which contains about 70k Question-Answer pairs derived from movie shots.
Besides, we propose ShotVL and train this VLM model with a two-stage training strategy, integrating both supervised fine-tuning and Group Relative Policy Optimization (GRPO). Experimental results demonstrate that our model achieves substantial improvements, surpassing all existing strongest open-source and proprietary models evaluated on ShotBench, establishing a new state-of-the-art performance. Jingwen He, Dian Zheng, Yuhao Dong, Fan Zhang 0045, Yinan He, Weichao Chen 0001, Yu Qiao 0001, Wanli Ouyang, Shengjie Zhao 0001, Ziwei Liu 0002 |
NeurIPS | 12 |
| 2025 | Point4Bit: Post Training 4-bit Quantization for Point Cloud 3D DetectionabstractVoxel-based 3D object detectors have achieved remarkable performance in point cloud perception, yet their high computational and memory demands pose significant challenges for deployment on resource-constrained edge devices. Post-training quantization (PTQ) provides a practical means to compress models and accelerate inference; however, existing PTQ methods for point cloud detection are typically limited to INT8 and lack support for lower-bit formats such as INT4, which restricts their deployment potential. In this paper, we present Point4bit, the first general 4-bit PTQ framework tailored for voxel-based 3D object detectors. To tackle challenges in low-bit quantization, we propose two key techniques: (1) Foreground-aware Piecewise Activation Quantization (FA-PAQ), which leverages foreground structural cues to improve the quantization of sparse activations; and (2) Gradient-guided Key Weight Quantization (G-KWQ), which preserves task-critical weights through gradient-based analysis to reduce quantization-induced degradation. Extensive experiments demonstrate that Point4bit achieves INT4 quantization with minimal accuracy loss with less than 1.5\% accuracy drop. Moreover, we validate its generalization ability on point cloud classification and segmentation tasks, demonstrating broad applicability. Our method further advances the bit-width limitation of point cloud quantization to 4 bits, demonstrating strong potential for efficient deployment on resource-constrained edge devices. Jianyu Wang 0003, Yu Wang 0174, Shengjie Zhao 0001, Sifan Zhou |
NeurIPS | 3 |
| 2025 | TriVSS-Net: Visual, Spatial, and Semantic Fusion Transformer for Two-View Correspondence Learning
Yizhang Liu, Shengjie Zhao 0001 |
PRCV (2) | 4 |
| 2025 | HC-RL-MPC: Epidemic Multi-scale Hierarchical Control Framework Based on Reinforcement Learning and Model Predictive ControlabstractHuman mobility restrictions are considered effective measures for epidemic mitigation but require adaptability to complex transmission environments and consistency across spatial scales. Existing mainstream approaches, including model predictive control (MPC) and reinforcement learning (RL), still face limitations. MPC relies on precise mathematical models and struggles to handle large-scale dynamic scenarios. Meanwhile, RL suffers from the curse of dimensionality, making fine-grained control challenging. In this work, we propose a novel epidemic multi-scale Hierarchical Control framework based on Reinforcement Learning and Model Predictive Control (HC-RL-MPC). At the high level, RL generates regional mobility restriction quota to balance medical and socioeconomic. At the low level, MPC clusters optimize the inter-subregion mobility quota allocation under high-level constraints. Furthermore, to ensure cross-scale policy consistency, we introduce a Policy Inverse Guidance (PIG) mechanism, where low-level modules provide gradient feedback and real-time results to high-level policy. Experimental results demonstrate that HC-RL-MPC generates macro-micro consistent mobility restriction strategies. Compared with existing models, it significantly reduces training complexity and improves cumulative rewards, providing an efficient solution for multi-scale dynamic decision-making. Xueting Luo, Hao Deng 0002, Shengjie Zhao 0001 |
SMC | 4 |
| 2025 | 3DGCformer: 3-Dimensional Graph Convolutional transformer for multi-step origin-destination matrix forecasting
Yiou Huang, Hao Deng 0002, Shengjie Zhao 0001 |
Appl. Intell. | 3 |
| 2025 | H2-MARL: Multi-agent reinforcement learning for Pareto optimality in hospital capacity strain and human mobility during epidemic
Xueting Luo, Hao Deng 0002, Jihong Yang, Huanhuan Guo, Mingqing Liu 0002, Jiming Wei, Shengjie Zhao 0001 |
Expert Syst. Appl. | 9 |
| 2025 | Deep Unrolled Weighted Graph Laplacian Regularization for Depth Completion
Jin Zeng 0004, Qingpeng Zhu, Tongxuan Tian, Wenxiu Sun, Lin Zhang 0014, Shengjie Zhao 0001 |
Int. J. Comput. Vis. | 6 |
| 2025 | PVBF: A framework for mitigating parameter variation imbalance in online continual learning
Zelin Tao, Hao Deng 0002, Mingqing Liu 0002, Lijun Zhang 0005, Shengjie Zhao 0001 |
Neural Networks | 5 |
| 2025 | Online indoor visual odometry with semantic assistance under implicit epipolar constraints
Yang Chen 0037, Lin Zhang 0014, Shengjie Zhao 0001, Yicong Zhou |
Pattern Recognit. | 3 |
| 2025 | TransMatch: Transformer-based correspondence pruning via local and global consensus
Yizhang Liu, Shengjie Zhao 0001 |
Pattern Recognit. | 3 |
| 2025 | I-DACS: Always Maintaining Consistency Between Poses and the Field for Radiance Field Construction Without Pose PriorabstractThe radiance field, emerging as a novel 3D scene representation, has found widespread application across diverse fields. Standard radiance field construction approaches rely on the ground-truth poses of key-frames, while building the field without pose prior remains a formidable challenge. Recent advancements have made strides in mitigating this challenge, albeit to a limited extent, by jointly optimizing poses and the radiance field. However, in these schemes, the consistency between the radiance field and poses is achieved completely by training. Once the poses of key-frames undergo changes, long-term training is required to readjust the field to fit them. To address such a limitation, we propose a new solution for radiance field construction without pose prior, namely I-DACS (Incremental radiance field construction with Direction-Aware Color Sampling). Diverging from most of the existing global optimization solutions, we choose to incrementally solve the poses and construct a radiance field within a sliding-window framework. The poses are unequivocally retrieved from the radiance field, devoid of any constraints and accompanying noise from other observation models, so as to achieve the consistency of poses to the field. Besides, in the radiance field, the color information is much higher-frequency and more time-consuming to learn compared with the density. To accelerate training, we isolate the color information to a distinct color field, and construct the color field based on an innovative direction-aware color sampling strategy, by which the color field can be derived directly from images without training. The color field obtained in this way is always consistent with the poses, and intricate details of training images can be retained to the utmost extent. Extensive experimental results evidently showcase both the remarkable training speed and the outstanding performance in rendering quality and localization accuracy achieved by I-DACS. To make our results reproducible, the source code has been released athttps://cslinzhang.github.io/I-DACS-MainPage/. Tianjun Zhang, Lin Zhang 0014, Shengjie Zhao 0001, Yicong Zhou |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | A Task-Oriented Spatial Graph Structure Learning Method for Traffic ForecastingabstractTraffic forecasting is the foundation of intelligent transportation systems (ITS). In recent, graph neural networks (GNNs) have successfully captured spatial-temporal dependencies to forecast traffic conditions by transforming traffic data in the graph domain. Nevertheless, the existing methods focus only on learning informative graph representations and fail to model informative graph structures, which hinders the capture of dynamic spatial-temporal dependencies caused by dynamic factors such as weather, accidents, and special events. In this paper, we propose a novel task-oriented Spatial Graph Structure Learning (SGSL) method, which aims to capture dynamic dependencies by jointly learning graph structures and graph representations. Compared to methods that use spectral graph representations, we exploit a learnable spatial graph to effectively model dynamic dependencies in traffic data. Moreover, we directly define graph convolutions on spatial relations to specify different edge weights when aggregating the information of spatial neighbours. Thus, the graph structure alterations, i.e., the relation changes, and the time-varying weights of relations can be encapsulated, thereby effectively representing dynamic dependencies. The gradient descent strategy is introduced to periodically learn a spatial graph through joint optimization with a newly designed deep graph learning model named GAT-nLSTM. In this manner, the intrinsic behaviours of nodes are learned to capture correlations across periods. Notably, the optimization process is performed under the traffic forecasting constraint to ensure that the learned spatial graph is specific to this task. Compared with those of state-of-the-art baselines, the experimental results obtained on real-world traffic datasets show significant improvement, which verifies the superiority of the proposed SGSL. Ting Wang 0019, Shengjie Zhao 0001, Wenzhen Jia, Daqian Shi |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2025 | Attentive Radiate Graph for Pedestrian Trajectory Prediction in Disconnected ManifoldsabstractPedestrian trajectory prediction grapples with the demanding feat of modeling complex interactions and learning multimodal distribution to navigate different human-centric environments. Despite superior performance in reducing distance-based metrics, recent works tend to predict out-of-distribution trajectories, as the distribution of forthcoming paths comprises a blend of various manifolds that may be disconnected. These unrealistic trajectories can potentially jeopardize the safety of traffic participants and result in significant damage. To meet these challenges, we propose DMPred, a graph-based generator adversarial network that generates realistic multimodal trajectory predictions by better modeling the social interactions of pedestrians across different scenes in disconnected manifolds. The core of DMPred is an attentive radiate graph sequence constructed by considering the localized influence radiating from pedestrian movements, which is followed by a spatiotemporal extractor that stores and reuses potentially forgotten neighboring pedestrian information to allow for better extraction of complex interactions. Additionally, a collection of generators is utilized for forecasting, which incorporates spectral clustering on trajectories during the prior learning process of multiple generators to help reduce model redundancy and enhance flexibility for various prediction scenarios. Through extensive experiments on multiple real-world and simulation datasets, we demonstrate that DMPred obtains highly competitive results with efficacy in predicting realistic multimodal trajectories. Peiyuan Zhu 0001, Shengjie Zhao 0001, Hao Deng 0002, Fengxia Han |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2025 | HT-FL: Hybrid Training Federated Learning for Heterogeneous Edge-Based IoT NetworksabstractWith the continuous rolling-out of edge computing, Federated Learning (FL) has become a promising solution for intelligent Internet-of-things (IoT). In addition to resource constraints, deploying FL schemes in IoT networks is greatly challenged byheterogeneityin multiple dimensions. While heterogeneity in data distribution and computation capability has been extensively studied, the impact of distinct, even hybrid training paradigms on FL performances remains largely unknown. To answer this open question in the IoT context, we propose aHybrid-Training Federated Learning(HT-FL) algorithm for the power-constrained IoT networks, incorporating both sequential and parallel training that naturally adapts to various sub-network topologies, while greatly reducing the energy consumption during the training stage. We demonstrate through analysis that the convergence of HT-FL is theoretically guaranteed, achieving$O (\frac{1}{\sqrt{K}})$for carefully chosen learning rates. Experiments on multiple datasets show that, the proposed HT-FL outperforms existing FL schemes on multiple training tasks under various data distribution settings, while reducing an average of 20% energy consumption. In a more practical sense, a self-adaptive parameter-tuning strategy is also designed for HT-FL deployment, which can be easily extended to other multi-layer FL schemes in complex application scenarios. Yixun Gu, Jie Wang 0016, Shengjie Zhao 0001 |
IEEE Trans. Mob. Comput. | 3 |
| 2025 | ATM-NeRF: Accelerating Training for NeRF Rendering on Mobile Devices via Geometric RegularizationabstractRecently, an increasing number of researchers have been dedicated to transferring the impressive novel view synthesis capability of Neural Radiance Fields (NeRF) to resource-constrained mobile devices. One common solution is to pre-train NeRF and bake it into textured meshes which are well supported by mobile graphics hardware. However, the training process of existing methods often requires several hours even with multiple high-end NVIDIA V100 GPUs. The underlying reason is that these schemes mainly rely on photometric rendering loss, neglecting the geometric relationship between the pre-trained NeRF and the baked results. Standing on this point, we presentATM-NeRF(AcceleratingTraining forMobile rendering based onNeRF), which is the first to apply effective geometric regularization constraints during both the pre-training and the baking training stages for faster convergence. Specifically, in the initial NeRF pre-training stage, we enforce consistency of the multi-resolution density grids representing the scene geometry to mitigate the shape-radiance ambiguity problem to some extent, achieving a coarse mesh with smoothness. In the second stage, we utilize the positions and geometric features of 3D points projected from the pre-trained posed depths to provide geometric supervision for joint refinement of geometry and appearance of the coarse mesh. As a result, our ATM-NeRF achieves comparable rendering quality to MobileNeRF with a training speed that is about$30\times \sim 70\times$faster while maintaining finer structure details of the exported mesh. Yang Chen 0037, Lin Zhang 0014, Shengjie Zhao 0001, Yicong Zhou |
IEEE Trans. Multim. | 3 |
| 2025 | SQL-Net: Semantic Query Learning for Point-Supervised Temporal Action LocalizationabstractPoint-supervised Temporal Action Localization (PS-TAL) detects temporal intervals of actions in untrimmed videos with a label-efficient paradigm. However, most existing methods fail to learn action completeness without instance-level annotations, resulting in fragmentary region predictions. In fact, the semantic information of snippets is crucial for detecting complete actions, meaning that snippets with similar representations should be considered as the same action category. To address this issue, we propose a novel representation refinement framework with a semantic query mechanism to enhance the discriminability of snippet-level features. Concretely, we set a group of learnable queries, each representing a specific action category, and dynamically update them based on the video context. With the assistance of these queries, we expect to search for the optimal action sequence that agrees with their semantics. Besides, we leverage some reliable proposals as pseudo labels and design a refinement and completeness module to refine temporal boundaries further, so that the completeness of action instances is captured. Finally, we demonstrate the superiority of the proposed method over existing state-of-the-art approaches on THUMOS14 and ActivityNet13 benchmarks. Notably, thanks to completeness learning, our algorithm achieves significant improvements under more stringent evaluation metrics. Yu Wang 0174, Shengjie Zhao 0001 |
IEEE Trans. Multim. | 2 |
| 2025 | Skeleton-Aware Graph-Based Adversarial Networks for Human Pose Estimation from Sparse IMUsabstractRecently, sparse-inertial human pose estimation (SI-HPE) with only a few IMUs has shown great potential in various fields. The most advanced work in this area achieved fairish results using only six IMUs. However, there are still two major issues that remain to be addressed. First, existing methods typically treat SI-HPE as a temporal sequential learning problem and often ignore the important spatial prior of skeletal topology. Second, there are far more synthetic data in their training data than real data, and the data distribution of synthetic data and real data is quite different, which makes it difficult for the model to be applied to more diverse real data. To address these issues, we propose “Graph-based Adversarial Inertial Poser (GAIP),” which tracks body movements using sparse data from six IMUs. To make full use of the spatial prior, we design a multi-stage pose regressor with graph convolution to explicitly learn the skeletal topology. A joint position loss is also introduced to implicitly mine spatial information. To enhance the generalization ability, we propose supervising the pose regression with an adversarial loss from a discriminator, bringing the ability of adversarial networks to learn implicit constraints into full play. Additionally, we construct a real dataset that includes hip support movements and a synthetic dataset containing various motion categories to enrich the diversity of inertial data for SI-HPE. Extensive experiments demonstrate that GAIP produces results with more precise limb movement amplitudes and relative joint positions, accompanied by smaller joint angle and position errors compared to state-of-the-art counterparts. The datasets and codes are publicly available at https://cslinzhang.github.io/GAIP/ . Kaixin Chen 0003, Lin Zhang 0014, Zhong Wang 0009, Shengjie Zhao 0001, Yicong Zhou |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2025 | Towards a Robust Visual-Inertial-Surround-View SLAM System for Autonomous Indoor ParkingabstractAn autonomous parking system is a low-speed unmanned driving system applied in indoor parking environments. Real-time and high-precision vehicle localization and map construction of the environment are two core functional modules of the system. Camera and IMU (Inertial Measurement Unit) sensors provide complementary data to create a Visual-Inertial Simultaneous Localization and Mapping (VI-SLAM) system. However, existing SLAM systems face challenges in complex parking environments. Moreover, limitations inherent in VI-SLAM systems further compromise their perception accuracy, affecting both localization and optimization. This article addresses the shortcomings of current VI-SLAM systems by proposing the RVIS SLAM system. This robust semantic SLAM system integrates data from three sensors: a front-view camera, an IMU, and a surround-view system. To ensure localization accuracy, the system utilizes metric information from common semantic objects on the ground. These objects include parking-slots, speed bumps, and parking-slot numbers captured in surround-view images to build scale-aware constraints. These constraints refine the initial scale of the SLAM system, which is often compromised under low IMU excitation conditions. Additionally, in optimization, SLAM systems ideally assume that the front-end produces optimization graphs without data association outliers. However, in real-world indoor parking environments, sensor noise and vehicle vibrations make this assumption unrealistic. To mitigate the adverse effects of outliers in SLAM systems, this article proposes a robust surround-view semantic data association strategy. This strategy quantifies the uncertainty of surround-view semantic landmarks for the first time, ensuring reliable localization and mapping in challenging environments. Extensive experiments in typical indoor parking environments validate the effectiveness and efficiency of the proposed RVIS SLAM system. Xuan Shao, Lin Zhang 0014, Tianjun Zhang, Shengjie Zhao 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2024 | CASTNet: Convolution Augmented Graph Sampling Transformer Network for Traffic Flow ForecastingabstractAccurate and efficient traffic flow forecasting plays an essential role in urban management which is conductive to traffic safety and travel experience. Despite the tremendous progress of attention-based traffic flow forecasting methods, complicated spatial and temporal correlations of traffic data pose difficulty for the network in simultaneously capturing the long-range contexts and the local feature details, leading to limited performance. To this end, we propose Convolution Augmented Graph Sampling Transformer Network (CASTNet) for accurate traffic flow prediction. Specifically, we design a hybrid graph transformer module to utilize the graph convolutions to perceive local structural details, which are incorporated into the transformer to augment the attention computation for more accurate global dependency modeling. Moreover, we develop a sampling-based attention layer based on the approximated sampling strategy to further optimize the graph structure, which is implemented with linear complexity. Finally, the spatial features from the graph transformer module are fused with the temporal features from the sample convolution and interaction network. Experimental results on three real-world traffic flow datasets demonstrate the superiority of the proposed CASTNet over existing schemes. Shengjie Zhao 0001, Jin Zeng 0004, Shilong Dong, Geyunqian Zu |
CSCWD | 2 |
| 2024 | KEWS: A KPIs-Based Evaluation Framework of Workload Simulation On Microservice SystemabstractSimulating the workload is an essential procedure in microservice systems as it helps augment realistic workloads whilst safeguarding user privacy. The efficacy of such simulation depends on its dynamic assessment. The straightforward and most efficient approach to this is comparing the original workload with the simulated one using Key Performance Indicators (KPIs), which capture the state of the system. Nonetheless, due to the extensive volume and complexity of KPIs, fully evaluating them is not feasible, and measuring their similarity poses a significant challenge. This paper introduces a similarity metric algorithm for KPIs, the Extended Shape-Based Distance (ESBD), which gauges similarity in both shape and intensity. Additionally, we propose a KPI-based Evaluation Framework for Workload Simulations (KEWS), comprising three modules: preprocessing, compression, and evaluation. These methodologies effectively counteract the adverse effects of KPIs’ characteristics and offer a holistic evaluation. Experimental results substantiate the effectiveness of both ESBD and KEWS. Pengsheng Li, Qingfeng Du, Shengjie Zhao 0001 |
CSCWD | 3 |
| 2024 | Advancing Root Cause Analysis in Cloud-native System with Knowledge Graph Path Embedding TranslationabstractCloud computing technologies, including cloud-native and containerization, have gained prominence in recent years, attributed to their exceptional scalability, enhanced resource utilization, and expedited deployment capabilities. However, their inherent complexity and the intricate interplay of internal components heighten the risk of sporadic and unforeseen anomalies. To address these challenges, Root Cause Analysis (RCA) is employed to accurately identify problematic services (pods) and mine the precise faults behind observed anomalies. Tailored to the limitations of conventional RCA algorithms, we propose a novel approach that jointly models operation entities and their relationships as learnable embeddings. Additionally, this method integrates fault propagation information to further improve RCA accuracy. Our evaluation involves developing a prototype within the Kubernetes cloud-native system. Extensive experimental results validate the efficacy of our approach. Pengsheng Li, Qingfeng Du, Shengjie Zhao 0001, Pei Fang |
CSCWD | 3 |
| 2024 | BEAVP: A Bidirectional Enhanced Adversarial Model for Video PredictionabstractPredicting future frames in videos is crucial for motion understanding and behavior analysis. However, despite significant advancements, existing stochastic methods have insufficient utilization of motion patterns, leading to blurry motion in long-term predictions. Most of the previous work also lacks constraints to effectively address the unconstrained nature of spacetime-varying motion. In this paper, we propose a stochastic video prediction model based on coupled GANs. The pair of GANs could model motion trends based on adjacent frames organized in sequential and reverse orders, respectively. We assume a common latent space assumption and build bridges between forward prediction and backward prediction by leveraging the constraints of weight-sharing and cycle- consistency. Specifically, we propose to learn a joint distribution with adjacent frames in opposite orders drawn from the marginal distributions and enhance forward prediction with an in-depth exploration of motion patterns. Through experiments on several challenging datasets that include spacetime-varying human motion, we show that our model surpasses the performance of state-of-the-art models, thus validating the effectiveness of our proposed approach. Peiyuan Zhu 0001, Shengjie Zhao 0001, Fengxia Han, Hao Deng 0002 |
FG | 2 |
| 2024 | SEA-GNN: Sequence Extension Augmented Graph Neural Network for Sequential RecommendationabstractSequential recommendation aims to anticipate the next preference of users by examining their recent interactions. Recently, graph neural networks (GNNs) have been widely utilized in sequential recommendation, but existing schemes focus on interactions within individual sequences and tend to connect irrelevant items in case of insufficient historical data. In this work, we propose Sequence Extension Augmented GNN (SEA-GNN) which augments the node representation learning with inter-sequence global context aggregation while maintaining intra-sequence local preference. Specifically, we augment the graph construction with sequence extension that diversifies the item connections to exploit global context and robustify the node representation against data insufficiency. Meanwhile, we extract local preference based on the intra-sequence user-item graph to enhance the node representation with user-specific interest. Experimental results demonstrate the superiority of the proposed algorithm compared with existing schemes in recommendation performance. Geyunqian Zu, Shengjie Zhao 0001, Jin Zeng 0004, Shilong Dong |
ICASSP | 2 |
| 2024 | Optimal Power Allocation for Location Privacy Security in Wireless LocalizationabstractThe prevalence of location-based services has made positional information indispensable to everyday life, which has sparked growing concerns about the security of location privacy. In this paper, we propose to protect location privacy in wireless lo-calization from the perspective of power allocation. A closed-form expression of location secrecy metric (LSM) is first established to quantify the degree of risk that the location can be inferred by the eavesdropper. Then, a transmit power optimization problem constrained by the LSM lower bound is formulated. Through problem transformation and fractional programming, the optimal power allocation is ultimately obtained. Simulation results show that compared with the other existing strategies, the proposed power allocation strategy can effectively protect location privacy at the cost of smaller nosltioning accuracy loss. Yuzhuo Dai, Jie Wang 0016, Junyuan Wang 0001, Shengjie Zhao 0001 |
ICC | 5 |
| 2024 | Controllable Rain Image Generation: Balance Between Diversity and Controllability
Kaibin Zhou, Shengjie Zhao 0001, Hao Deng 0002 |
ICIC (6) | 2 |
| 2024 | Weakly-Supervised Action Localization by Hierarchical Attention Mechanism with Multi-Scale Fusion StrategiesabstractWeakly-supervised temporal action localization focuses on locating action intervals when merely video-level supervised signals are available. Conventional methods mostly rely on the attention framework, which generates a set of scores indicating the confidence that the video snippet belongs to the foreground, the background, and the context, respectively. However, such methods fail to consider the structural properties of snippet-level features when generating attention scores, and these structural properties are critical for capturing contextual information in temporal tasks. To this end, we propose a hierarchical attention generation mechanism with multi-scale fusion strategies to model such structural information. Besides, to resolve action-context confusion issues that are quite intractable in weakly-supervised action localization tasks, metric learning is further introduced into our framework to suppress context features from approaching action features, while encouraging them to be close to background features. Finally, our model is evaluated on THUMOS14 and ActivityNet1.3 benchmarks, and the results demonstrate that the proposed approach achieves desirable performance. Yu Wang 0174, Shengjie Zhao 0001 |
ICME | 2 |
| 2024 | MGSTA : Meta Learning Based Graph Convolutional Stacked Temporal Attention Neural Network for Traffic Flow ForecastingabstractRecently, numerous deep learning-based methodologies have been applied to the domain of traffic flow forecasting, showcasing a significant enhancement in predictive accuracy when compared to the conventional statistical approaches. However, a critical reexamination of traffic flow prediction from the perspective of regional traffic diversity reveals that prevailing spatio-temporal traffic networks tend to neglect the heterogeneity of urban traffic patterns. The spatio-temporal correlations within urban regions intricately links to regional characteristic data (such as points of interests,i.e., POIs and road network density), which in turn influence the variations in traffic flow across diverse urban areas. In light of this, we propose a meta-learning-based graph convolutional stacked temporal attention network (MGSTA). Different from the mainstream sptio-temporal traffic network that iteratively optimize the model parameters by gradient descent after their generation. The parameters of the proposed spatio-temporal module are generated by a series of meta-learning layers, which are specifically designed for learning these meta characteristic data. Experiments on multiple datasets have proved the superior performance of the proposed model and the substantial improvements brought by this new idea to the spatio-temporal feature learning capabilities of the traffic prediction model. Shengjie Zhao 0001, Yushan Feng, Fengxia Han |
IJCNN | 2 |
| 2024 | RoSe: Rotation-Invariant Sequence-Aware Consensus for Robust Correspondence PruningabstractCorrespondence pruning has recently drawn considerable attention as a crucial step in image matching. Existing methods typically achieve this by constructing neighborhoods for each feature point and imposing neighborhood consistency. However, the nearest-neighbor matching strategy often results in numerous many-to-one correspondences, thereby reducing the reliability of neighborhood information. Furthermore, the smoothness constraint fails in cases of large-scale rotations, leading to misjudgments. To address the above issues, this paper proposes a novel robust correspondence pruning method termed RoSe, which is based on rotation-invariant sequence-aware consensus. We formulate the correspondence pruning problem as a mathematical optimization problem and derive a closed-form solution. Specifically, we devise a rectified local neighborhood construction strategy that effectively enlarges the distribution between inliers and outliers. Meanwhile, to accommodate large-scale rotation, we propose a relative sequence-aware consistency as an alternative to existing smoothness constraints, which can better characterize the topological structure of inliers. Experimental results on image matching and registration tasks demonstrate the effectiveness of our method. Robustness analysis involving diverse feature descriptors and varying rotation degrees further showcases the efficacy of our method. Yizhang Liu, Shengjie Zhao 0001 |
ACM Multimedia | 4 |
| 2024 | Depth Camera-LiDAR Fusion Based UAV Autonomous Navigation for Parcel Delivery in Complex Urban EnvironmentabstractUnmanned aerial vehicle (UAV) delivery in urban environments demonstrates significant growth potential due to its efficiency and eco-friendliness. The challenges of autonomous navigation and obstacle avoidance in complex environments are crucial study aspects of UAV delivery. This paper proposes a depth camera-LiDAR fusion based UAV autonomous navigation protocol, enabling autonomous navigation and obstacle avoidance for UAV parcel delivery in complex urban environments. Experiments are conducted in a highly realistic simulation environment to validate performance of the proposed method. The results indicate that the proposed method can efficiently and accurately achieve autonomous navigation and obstacle avoidance with high robustness. Chengdao Chi, Bing Li 0025, Hao Deng 0002, Shengjie Zhao 0001 |
SMC | 4 |
| 2024 | Beamforming Design With Partial Group Successive Interference Cancelation for ISAC SystemsabstractIntegrated sensing and communication (ISAC) is an emerging paradigm in the sixth-generation mobile communication systems (6G) to address the spectrum scarcity and realize the vision of the Internet of Everything (IoE). In this article, we consider a multiple-input-multiple-output (MIMO) ISAC system where the dual-functional radar-communication (DFRC) base station (BS) detects the targets and communicates with multiple downlink users. To meliorate the severe interference management and improve the transmission performance of the ISAC system, we identify a specific transceiver design that splits each independent message into multiple layers at the transmitter and employs a partial group successive interference cancelation scheme at the receivers. To coordinate the communication and radar performance, we formulate an optimization problem to approximate the beamformers to the desired radar beampattern subject to the achievable rate regions. Since, the formulated problem is nonconvex and NP-hard, we propose an iterative algorithm based on the semi-definite programming relaxation, which optimizes the beamformers and rate vectors alternatively to yield near-optimal solutions. Numerical results demonstrate the superior performance of the proposed transceiver design in improving the achievable transmission rate and obtaining better interference management in the ISAC system. Mengqiu Chai, Shengjie Zhao 0001, Fengxia Han, Yuan Liu 0030 |
IEEE Internet Things J. | 2 |
| 2024 | Joint Sensing and Power Transfer via Distributed Coupled-Cavity LasersabstractPositioning and power transfer are crucial demands in the existing Internet of Things networks, where intracavity laser-based systems are proposed as a potential alternative for providing sufficient wireless power and high-accuracy positioning simultaneously. However, existing intracavity laser-based systems still face challenges in improving power transfer efficiency and system Field of View (FoV). This article proposes a joint sensing and power transfer (JSPT) system based on distributed coupled-cavity laser (DCCL) design. Power efficiency and FoV are enhanced owing to the external-cavity feedback in DCCL. The angle of arrival estimation based on the intrinsic self-alignment feature and polarization self-modulation ranging based on the self-mixing feature are achieved in the DCCL-based JSPT system. Moreover, we build analytical models for revealing the principle of DCCL and verifying the system performance relying on diffraction propagation-based beam field simulation and rate equation-based gain simulation. Numerical results demonstrate that DCCL-based JSPT achieves 4-W charging power and mm-level 3-D positioning within an FoV of ±40° and 2-m vertical distances, showing its capability for Internet of Things applications. Hao Deng 0002, Shengjie Zhao 0001, Mingqing Liu 0002, Qingwen Liu 0001 |
IEEE Internet Things J. | 2 |
| 2024 | MPEG: A Multi-Perspective Enhanced Graph Attention Network for Causal Emotion Entailment in ConversationsabstractEmotion causes constitute a pivotal component in the comprehension of emotional conversations. Recently, a new task named Causal Emotion Entailment (CEE) has been proposed to identify the causal utterances for the target emotional utterance in a conversation. Although researchers have achieved some progress in solving this problem, they failed to adequately incorporate speaker characteristics and overlooked the effects of temporal relations in conversation structures. To fill such a research gap to some extent, we propose a novel causal emotion entailment framework, namely MPEG (Multi-Perspective Enhanced Graph attention network). The training of MPEG consists of three stages. Firstly, we utilize a speaker-aware pre-trained model and two attention mechanisms to obtain the utterance representations that incorporate local contexts as well as the speaker and emotional information. Then, these representations are fed into a graph attention network to model the conversation structures and emotional dynamics from both local and global perspectives. Finally, a fully-connected network is implemented to predict the relationships between emotional utterances and causal utterances. Experimental results show that MPEG achieves state-of-the-art performance. The source code is available athttps://github.com/slptongji/MPEG. Ying Shen 0005, Xuri Chen, Lin Zhang 0014, Shengjie Zhao 0001 |
IEEE Trans. Affect. Comput. | 5 |
| 2024 | Unsupervised BLSTM-Based Electricity Theft Detection with Training Data ContaminatedabstractElectricity theft can cause economic damage and even increase the risk of outage. Recently, many methods have implemented electricity theft detection on smart meter data. However, how to conduct detection on the dataset without any label still remains challenging. In this article, we propose a novel unsupervised two-stage approach under the assumption that the training set is contaminated by attacks. Specifically, the method consists of two stages: (1) a Gaussian mixture model is employed to cluster consumption patterns with respect to different habits of electricity usage, and with the goal of improving the accuracy of the model in the posterior stage; (2) an attention-based bidirectional long short-term memory encoder-decoder scheme is employed to improve the robustness against the non-malicious changes in usage patterns leveraging the process of encoding and decoding. Quantifying the similarity of consumption patterns and reconstruction errors, the anomaly score is defined to improve detection performance. Experiments on a real dataset show that the proposed method outperforms the state-of-the-art unsupervised detectors. Qiushi Liang, Shengjie Zhao 0001, Jiangfan Zhang, Hao Deng 0002 |
ACM Trans. Cyber Phys. Syst. | 2 |
| 2024 | TriKF: Triple-Perspective Knowledge Fusion Network for Empathetic Question GenerationabstractQuestioning is one of the essential tactics for demonstrating empathy in social dialogues. Effective questioning can guide individuals to express their experiences, feelings, and thoughts, aiming to establish emotional connections and deepen interpersonal understanding. However, how to generate empathetic questions in emotional support conversations remains an unresolved issue. To fill this research gap to some extent, we propose an empathetic question generation (QG) framework called triple-perspective knowledge fusion (TriKF), which incorporates external knowledge from the perspectives of events, cognition, and affection to comprehensively understand the dialogue context. Specifically, this framework acquires commonsense knowledge from these three perspectives and integrates them into the dialogue context to enrich the contextual information. To the best of our knowledge, this is the first method proposed for empathetic QG. Additionally, we construct an empathetic question dataset, namely EQ-EMAC. This dataset comprises 4213 dialogues with single user inputs and multiple empathetic question responses, which can be utilized to assess the effectiveness and generalization capability of empathetic QG models. Experimental results have demonstrated the effectiveness of TriKF on the task of empathetic QG compared with seven baseline models. Ying Shen 0005, Xuri Chen, Lin Zhang 0014, Shengjie Zhao 0001 |
IEEE Trans. Comput. Soc. Syst. | 5 |
| 2024 | The Hierarchical Clustering of Human Mobility BehaviorsabstractHuman mobility flow prediction forecasts the number of passengers coming into (inflows) or leaving from (outflows) every region of a city. It is crucial for many applications such as mobile marketing or route optimization. People have proposed deep learning methods to predict human mobility flows. However, existing methods neglect the hierarchical nature of human mobility behaviors. Each of us is unique, and we live beyond our neighborhoods. In the process of cross-regional activities, humans migrate in the hierarchical structure of buildings, neighborhoods, regions, cities, and countries. In this article, we propose a predictive framework, hierarchical fuzzy C-means (Hierarchical-FCM)-residual networks (ResNets), to capture the hierarchical structure of human mobility for prediction. First, we offer a Hierarchical-FCM clustering algorithm that is trained to learn the relationship between proximity and road network hierarchically on large human mobility data. Second, we design a fusion model to incorporate the knowledge learned from hierarchical clusters into deep ResNets to improve prediction accuracy. We compare our framework with 25 existing state-of-the-art models, ranging from traditional time-series and machine learning predictors, such as ARIMA and RNN, to the latest deep learning methods designed for human mobility prediction, such as ST-ResNets and DeepST. Our framework outperforms all existing models, by a margin of 2%–68% in terms of prediction accuracy, showing its effectiveness. We are among the first to incorporate the hierarchical structure of human mobility into location clustering for human mobility flow prediction. We empirically show that incorporating such knowledge significantly improves prediction performance. Wenzhen Jia, Kai Zhao 0011, Shengjie Zhao 0001 |
IEEE Trans. Comput. Soc. Syst. | 3 |
| 2024 | Global Localization in Large-Scale Point Clouds via Roll-Pitch-Yaw Invariant Place Recognition and Low-Overlap Global RegistrationabstractFor autonomous ground vehicles, global localization with 3D LiDAR is an indispensable part of tasks such as navigation. Usually, global localization using LiDAR is subdivided into two sub-problems, place recognition and global registration. For place recognition, the recent emerging schemes based on deep learning either rely on 3D convolution with high complexity or need to learn features from various forward perspectives. To mitigate this, we propose a model with roll-pitch-yaw invariance that represents point clouds as probabilistic voxels and generates occupancy grids from a bird’s-eye view, fulfilling robust place recognition by learning aggregated embeddings from a fixed perspective. For low-overlap global registration, the traditional handcraft feature-based methods are mostly limited to dense object-level point clouds, while the state-of-the-art learning-based approaches often rely on complex 3D convolution and additional feature association learning. To fill this gap to some extent, we propose to estimate the relative roll-pitch angles and vertical translation by fitting and aligning the ground plane of the point clouds and to determine the horizontal translations and yaw angle by matching their projected occupancy grids. Extensive experiments corroborate the superior recall and generalization ability of our place recognition model, as well as the advanced success rate and accuracy of our 3D registration approach. Especially in the recognition and registration of hard samples, our results far exceed those of our counterparts by large margins. To ensure full reproducibility, the relevant codes and data are made available online. Zhong Wang 0009, Lin Zhang 0014, Shengjie Zhao 0001, Yicong Zhou |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Ct-LVI: A Framework Toward Continuous-Time Laser-Visual-Inertial Odometry and MappingabstractOwing to the inherent complementarity among LiDAR, camera, and IMU, a growing effort has been paid to laser-visual-inertial SLAM recently. The existing approaches, however, are limited in two aspects. First, at the front-end, they usually employ a discrete-time representation that requires high-precision hardware/software synchronization and are based on geometric laser features, leading to low robustness and scalability. Second, at the backend, visual loop constraints suffer from scale ambiguity and the sparseness of the point cloud deteriorates the scan-to-scan loop detection. To solve these problems, for the front-end, we propose a continuous-time laser-visual-inertial odometry which formulates the carrier trajectory in continuous time, organizes point clouds in probabilistic submaps, and jointly optimizes the loss terms of laser anchors, visual reprojections, and IMU readings, achieving accurate pose estimation even with fast motion or in unstructured scenes where it is difficult to extract meaningful geometric features. At the backend, we propose building 5-DoF laser constraints by matching projected 2D submaps and 6-DoF visual constraints via laser-aided visual relocalization, ensuring mapping consistency in large-scale scenes. Results show that our framework achieves high-precision estimation and is more robust than its counterparts when the carrier works in large scenes or with fast motion. The relevant codes and data are open-sourced at https://cslinzhang.github.io/Ct-LVI/Ct-LVI.html. Zhong Wang 0009, Lin Zhang 0014, Shengjie Zhao 0001, Yicong Zhou |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Adaptive Guided Convolution Generated With Spatial Relationships for Point Clouds AnalysisabstractRepresenting valuable semantics in irregular 3-D point clouds with spatial-variant relationships is challenging. Current shared/weighted point convolution methods have limited efficiency and struggle to represent distinctive weight matrices precisely. In thisarticle, we present guided convolution (GuidedConv) operation, a plug-and-play point convolution operator for 3-D point cloud analysis. Compared to existing point convolution methods, GuidedConv implements a guided-share method for the weight matrices, which dynamically assigns corresponding weight matrices to points depending on their local spatial relationships without increasing the kernels in a brute-force way or equipping sophisticated auxiliary networks. The key of GuidedConv is to dynamically assigns weight matrices generated by kernel base to the points according to a guided mask, which is learned adaptively from the spatial relationships of 3-D point clouds in a data-driven manner. Guided masks establish a one-of-a-kind correspondence between the weight matrices and the points. In this way, GuidedConv possesses powerful capabilities to characterize different semantics and is more efficient and precise than current point convolution operators. Extensive qualitative and quantitative evaluations clearly show integrating GuidedConv into several relatively typical networks without modifying configurations to achieve on-par or better performances than state-of-the-art point cloud classification and segmentation methods on several benchmark datasets. Thorough ablation studies and visualizations demonstrate the flexibility and effectiveness of GuidedConv. Xutao Chu, Shengjie Zhao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | A Survey on Knowledge Graph Related Research in Smart City DomainabstractKnowledge graph employs the specific graph structure to store knowledge in the form of entities, relations, attributes, and so forth, which can effectively represent correlations among data and has been applied in many fields, including search engine optimization, intelligent question answering, and recommendation systems. In this article, we mainly focus on the research and application of the domain-specific knowledge graph in the field of the smart city, which has not been fully paid attention to. Currently, the major problem faced by the smart city lies in data mining and proper application. On the one hand, data are usually stored by government management departments, which creates challenges such as high data storing overhead and inefficient data usage. On the other hand, data cannot be coordinated and collaborated between different city management systems, because data silos exist. By constructing the corresponding knowledge graph, the data of urban traffic, services, and public resources are integrated to provide help for city builders and managers to make important decisions. Therefore, we will review the related literature on the knowledge graph existing in the smart city domain to expore reasearch scopes. Specifically, we will analyze and summarize knowledge graph construction research in the field of smart cities from four perspectives, i.e., smart city ontology, urban data processing, urban knowledge graph construction, and their application. Finally, the research limitations and prospects of the urban knowledge graph are provided. Zhu Wang 0016, Fengxia Han, Shengjie Zhao 0001 |
ACM Trans. Knowl. Discov. Data | 3 |
| 2024 | Human Mobility Prediction Based on Trend Iteration of Spectral ClusteringabstractHuman mobility prediction is crucial for epidemic control, urban planning, and traffic forecasting systems. We observe urban traffic flow prediction has a hierarchical structure, in which human mobility prediction should consider not only the spatial and the temporal relationships, but also the high-level mobility trend between individuals and regions. In this paper, we propose a human mobility clustering algorithm based on trend iteration of spectral clustering (TISC) to incorporate the high-level human mobility trend between individuals and regions. We integrate our TISC clustering algorithm with two existing urban traffic flow predictive models: namely, deep spatio-temporal residual network (ST-ResNet) and deep spatio-temporal 3D network (ST-3DNet). By adapting our TISC clustering algorithm, the prediction accuracy of both algorithms has been improved significantly (30.96$\%$for ST-ResNet and 24.66$\%$for ST-3DNet). We also compare the TISC-based predictive framework with 26 state-of-the-art human mobility prediction algorithms. We observe that our TISC algorithm considerably outperforms all 26 methods, reducing the predictive error from 6.93% to 69.55$\%$. Wenzhen Jia, Shengjie Zhao 0001, Kai Zhao 0011 |
IEEE Trans. Mob. Comput. | 2 |
| 2024 | Greta: Towards a General Roadside Unit Deployment FrameworkabstractAs an essential component, roadside units (RSUs) play an indispensable role in realizing Vehicle-to-Everything (V2X) by seamlessly connecting various intelligent devices and vehicles. To facilitate the construction of V2X, much research has been done in designing effective RSU deployment strategies. However, most of these efforts are largely limited by design utility and deployment scalability. To address the limitations of previous works, this paper proposes a general RSU deployment framework,Greta, which can evaluate candidate deployment sites from different perspectives with rich input data, and satisfy different requirements on optimization metrics. To this end, we model the general RSU deployment problem as a customized reinforcement learning (RL) problem that intelligently explores the deployment environment to find a good deployment strategy. Specifically, we design an effective data profiling network to extract features from multi-modality input data. These extracted features are gradually weighted, fused, and encoded as part of the state representation of the RL model. We further design new reward functions considering various deployment metrics and propose an action space pruning scheme to speed up model training. We implement a prototype system ofGretaand extensively evaluate its performance using real-world data. The results showGretaachieves remarkable performance gains compared to recent RSU deployment methods. Xianjing Wu, Zhidan Liu 0001, Zhenjiang Li 0001, Shengjie Zhao 0001 |
IEEE Trans. Mob. Comput. | 4 |
| 2024 | Action-Semantic Consistent Knowledge for Weakly-Supervised Action LocalizationabstractWeakly-supervised temporal action localization aims to detect temporal intervals of actions in arbitrarily long untrimmed videos with only video-level annotations. Owing to label sparsity, learning action consistency is intractable. In this paper, we assume that frames with similar representations in a given video should be considered as the same action. To this end, we develop a query-based contrastive learning paradigm to ensure action-semantic consistency. This mechanism encourages normalized embeddings with the same class to be pulled closer together, while embeddings from different classes are repelled apart. Besides, we design a two-branch framework, consisting of a class-aware branch and a class-agnostic branch, to learn salient features and fine-grained clues respectively. To further guarantee the action-semantic consistency of the two branches, unlike previous methods that handle each branch independently, we model the relationship between the two branches to avoid unreasonable predictions. Finally, the proposed model demonstrates superior performance over existing methods on the publicly available THUMOS-14 and ActivityNet-1.3 datasets. Substantial experiments and ablation studies also demonstrate the effectiveness of our model. Yu Wang 0174, Shengjie Zhao 0001 |
IEEE Trans. Multim. | 2 |
| 2024 | An Underwater Organism Image Dataset and a Lightweight Module Designed for Object Detection NetworksabstractLong-term monitoring and recognition of underwater organism objects are of great significance in marine ecology, fisheries science and many other disciplines. Traditional techniques in this field, including manual fishing-based ones and sonar-based ones, are usually flawed. Specifically, the method based on manual fishing is time-consuming and unsuitable for scientific researches, while the sonar-based one, has the defects of low acoustic image accuracy and large echo errors. In recent years, the rapid development of deep learning and its excellent performance in computer vision tasks make vision-based solutions feasible. However, the researches in this area are still relatively insufficient in mainly two aspects. First, to our knowledge, there is still a lack of large-scale datasets of underwater organism images with accurate annotations. Second, in consideration of the limitation on hardware resources of underwater devices, an underwater organism detection algorithm that is both accurate and lightweight enough to be able to infer in real time is still lacking. As an attempt to fill in the aforementioned research gaps to some extent, we established the Multiple Kinds of Underwater Organisms (MKUO) dataset with accurate bounding box annotations of taxonomic information, which consists of 10,043 annotated images, covering eighty-four underwater organism categories. Based on our benchmark dataset, we evaluated a series of existing object detection algorithms to obtain their accuracy and complexity indicators as the baseline for future reference. In addition, we also propose a novel lightweight module, namely Sparse Ghost Module, designed especially for object detection networks. By substituting the standard convolution with our proposed one, the network complexity can be significantly reduced and the inference speed can be greatly improved without obvious detection accuracy loss. To make our results reproducible, the dataset and the source code are available online at https://cslinzhang.github.io/MKUO-and-Sparse-Ghost-Module/ . Jiafeng Huang, Tianjun Zhang, Shengjie Zhao 0001, Lin Zhang 0014, Yicong Zhou |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2024 | I2P Registration by Learning the Underlying Alignment Feature Space from Pixel-to-Point SimilaritiesabstractEstimating the relative pose between a camera and a LiDAR holds paramount importance in facilitating complex task execution within multi-agent systems. Nonetheless, current methodologies encounter two primary limitations. First, amid the cross-modal feature extraction, they typically employ separate modal branches to extract cross-modal features from images and point clouds. This approach results in the feature spaces of images and point clouds being misaligned, thereby reducing the robustness of establishing correspondences. Second, due to the scale differences between images and point clouds, one-to-many pixel-point correspondences are inevitably encountered, which will mislead the pose optimization. To address these challenges, we propose a framework named I mage-to- P oint cloud registration by learning the underlying alignment feature space from P ixel-to- P oint SIM imilarities (I2P \({}_{\mathbf{ppsim}}\) ) . Central to \(\text{I2P}_{\text{ppsim}}\) is a Shared Feature Alignment Module (SFAM). It is designed under on a coarse-to-fine architecture and uses a weight-sharing network to construct an alignment feature space. Benefiting from SFAM, \(\text{I2P}_{\text{ppsim}}\) can effectively identify the co-view regions between images and point clouds and establish high-reliability 2D-3D correspondences. Moreover, to mitigate the one-to-many correspondence issue, we introduce a similarity maximization strategy termed point-max. This strategy effectively filters out outliers, thereby establishing accurate 2D-3D correspondences. To evaluate the efficacy of our framework, we conduct extensive experiments on KITTI Odometry and Oxford Robotcar. The results corroborate the effectiveness of our framework in improving image-to-point cloud registration. To make our results reproducible, the source codes have been released at https://cslinzhang.github.io/I2P Yunda Sun, Lin Zhang 0014, Zhong Wang 0009, Yang Chen 0037, Shengjie Zhao 0001, Yicong Zhou |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2023 | FedWNS: Data Distribution-Wise Node Selection in Federated Learning via Reinforcement LearningabstractTo deal with the discrepancy between global and local objectives in the federated learning invoked by the non-independent, identically distributed (non-IID) data and mitigate the impact of catastrophic forgetting in the training phase, we propose a federated learning framework with data distribution-wise reinforcement learning to perform node selection to accelerate the convergence process and alleviate the accuracy degradation. In this framework, the agent on the central server observes the number of samples every node owns, the derived distribution information of every dataset, and the current local and global accuracy. Then infer the selected node-set to participate in the current federated learning round through policy network in reinforcement learning. Finally, we conduct simulations with publicly data sets. Simulation results indicate that our FedWNS outperforms the existing FedAvg and CSFedAvg on the testing accuracy and the communication rounds to reach target accuracy under different settings. Chengwu Tu, Shengjie Zhao 0001, Hao Deng 0002 |
CSCWD | 2 |
| 2023 | SVFNeXt: Sparse Voxel Fusion for LiDAR-Based 3D Object Detection
Deze Zhao, Shengjie Zhao 0001, Shuang Liang 0001 |
PRICAI (3) | 2 |
| 2023 | Joint UAV Trajectory Scheduling and Network Routing for FANETs: A Reinforcement Learning ApproachabstractFlying Ad-hoc Networks (FANETs) consisting of multiple flexible unmanned aerial vehicles (UAVs) have gained ever-increasing attention due to their advantages in flexible deployment and enhanced data link connectivity. However, existing FANET schemes do not consider the mutual influence between data routing and UAV trajectory scheduling, resulting in limited network performance especially when the number of UAVs in FANET is large. This paper proposes a joint optimization framework for routing and UAV trajectory scheduling in FANET to enhance communication reliability and efficiency. Unlike previous studies that consider routing and trajectory scheduling separately, this framework formulates the routing problem as a Markov Decision Process (MDP) and employs trajectory schedule to guide the state transition probability, addressing the limitation of routing algorithms that passively perceive network topology. In addition, we exploit reinforcement learning to iteratively determine the optimal routing and trajectory scheduling strategy with the support of network topology prediction. Simulation results demonstrate that our proposed approach can significantly reduce latency, communication overhead, and packet loss in FANET. Junyi Gan, Bing Li 0025, Shengjie Zhao 0001 |
SMC | 3 |
| 2023 | Principal graph embedding convolutional recurrent network for traffic flow prediction
Yang Han 0007, Shengjie Zhao 0001, Hao Deng 0002, Wenzhen Jia |
Appl. Intell. | 2 |
| 2023 | Joint task assignment and resource allocation in VFC based on mobility prediction information
Xianjing Wu, Shengjie Zhao 0001, Hao Deng 0002 |
Comput. Commun. | 2 |
| 2023 | An Automatic Malaria Disease Diagnosis Framework Integrating Blockchain-Enabled Cloud-Edge Computing and Deep LearningabstractMalaria is a life-threatening disease, which mainly occurs in developing countries and regions with poor sanitary conditions. Early diagnosis of malaria will effectively decrease the death rate. In this article, we develop an automatic malaria disease diagnosis framework integrating blockchain-enabled cloud–edge computing and deep learning. The diagnosis task is divided into malaria parasite segmentation from blood smear images and classification of parasite species and stages. To meet the massive demand for deep learning training, we design a diagnosis pipeline that is deployed in a cloud–edge paradigm to utilize both local and remote resources. At edge nodes, preprocessed data sets are classified by U-Net in a supervised approach to generate coarse probability maps. Then, the normalized images and generated probability maps are uploaded to the cloud server. At the cloud, the uploaded probability maps are used to weakly supervise the stacked dilated U-Net (SDU-Net) to segment infected cells. Further classifications of malaria parasites species and stages are conducted by a pretrained MobileNet V1. The blockchain technology is adopted during the data transmission process. The diagnosis results will be sent back to the original local hospital immediately through the cloud. Our framework improves the diagnosis accuracy and eases the burden of deep learning training. Evaluation on real data collection MP-IDB demonstrated the effectiveness of our method. Shengjie Zhao 0001, Chenxi Huang 0001 |
IEEE Internet Things J. | 2 |
| 2023 | LWS: A framework for log-based workload simulation in session-based SUT
Yongqi Han 0001, Qingfeng Du, Jincheng Xu, Shengjie Zhao 0001, Zhekang Chen, Kanglin Yin, Dan Pei |
J. Syst. Softw. | 4 |
| 2023 | Compressive Spectral Imaging via Misalignment Induced Equivalent Grayscale Coded ApertureabstractCoded aperture snapshot spectral imager (CASSI) senses the spectral information of a 2-D scene and captures a set of coded measurement data that can be used to reconstruct the 3-D spatio-spectral datacube of the input scene by compressive sensing algorithms. The coded aperture (CA) in CASSI plays a crucial role in modulating the spatial information. The pixels in CA are typically square, switched binary ON–OFF, and aligned with the pixels of focal plane array (FPA). Instead of this binary modulation, this letter explores a simple yet effective approach to enabling an equivalent grayscale modulation, which can increase the sensing degree of freedom in CASSI systems. In particular, we deliberately introduce misalignment between the CA pixels and the FPA pixels, such that the spatial modulation of one FPA pixel is determined by four adjacent CA pixels instead of one. Numerical experiments show that the proposed equivalent grayscale modulation induced by misalignment can significantly improve the CASSI reconstruction when compared with current methods, whether a random CA or an optimal blue noise CA is used. More importantly, it does not incur in any cost to the CASSI system. Tong Zhang 0001, Shengjie Zhao 0001, Andres Ramirez-Jaime, Qile Zhao, Gonzalo R. Arce |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2023 | An Embedding-Driven Multi-Hop Spatio-Temporal Attention Network for Traffic PredictionabstractTraffic prediction is an important part of modern intelligent transportation systems (ITS), which helps transportation management and city planning. However, it is a very challenging task for modeling complex spatio-temporal dependencies, since the traffic data belongs to highly periodic multivariate time series which makes it hard to model accurate spatial dependencies only from time series and observed geolocation information of road segments. The existing research mainly focuses on finding ways of capturing dynamic spatial dependencies of road segments while neglecting the importance of periodicity, and few studies have explored a pure embedding-driven method that is robust to corrupted data to model periodicity. In this paper, we propose an embedding-driven multi-hop spatio-temporal attention network for traffic prediction (PIANOFORTE), which mainly focuses on leveraging the multi-scale periodicity of traffic data. Specifically, the proposed network applies a designed Fourier-series-based embedding, to capture the periodicity, which is more in line with real-world facts. Driven by the designed embedding, both local and global temporal dependencies are modeled properly by combining the attention-based methods and the convolution-based methods. Besides, we implement a trial that can hardly be seen in the existing traffic prediction works to combine the graph self-attention mechanism with a multi-hop diffusion process to explore the large-scale structural information on a designed set of graphs. Experiments on two real-world traffic datasets which contains traffic speed data for months show the effectiveness of our proposed methods. The experiments also suggest the methods can provide stable reasonable and smooth predictions for completely corrupted data. Shengjie Zhao 0001, Fengxia Han |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2023 | SLAM for Indoor Parking: A Comprehensive Benchmark Dataset and a Tightly Coupled Semantic FrameworkabstractFor the task of autonomous indoor parking, various Visual-Inertial Simultaneous Localization And Mapping (SLAM) systems are expected to achieve comparable results with the benefit of complementary effects of visual cameras and the Inertial Measurement Units. To compare these competing SLAM systems, it is necessary to have publicly available datasets, offering an objective way to demonstrate the pros/cons of each SLAM system. However, the availability of such high-quality datasets is surprisingly limited due to the profound challenge of the groundtruth trajectory acquisition in the Global Positioning Satellite denied indoor parking environments. In this article, we establish BeVIS, a large-scale Be nchmark dataset with V isual (front-view), I nertial and S urround-view sensors for evaluating the performance of SLAM systems developed for autonomous indoor parking, which is the first of its kind where both the raw data and the groundtruth trajectories are available. In BeVIS, the groundtruth trajectories are obtained by tracking artificial landmarks scattered in the indoor parking environments, whose coordinates are recorded in a surveying manner with a high-precision Electronic Total Station. Moreover, the groundtruth trajectories are comprehensively evaluated in terms of two respects, the reprojection error and the pose volatility, respectively. Apart from BeVIS, we propose a novel tightly coupled semantic SLAM framework, namely VIS SLAM -2, leveraging V isual (front-view), I nertial, and S urround-view sensor modalities, specially for the task of autonomous indoor parking. It is the first work attempting to provide a general form to model various semantic objects on the ground. Experiments on BeVIS demonstrate the effectiveness of the proposed VIS SLAM -2. Our benchmark dataset BeVIS is publicly available at https://shaoxuan92.github.io/BeVIS . Xuan Shao, Ying Shen 0005, Lin Zhang 0014, Shengjie Zhao 0001, Dandan Zhu 0001, Yicong Zhou |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2023 | Temporal Dropout for Weakly Supervised Action LocalizationabstractWeakly supervised action localization is a challenging problem in video understanding and action recognition. Existing models usually formulate the training process as direct classification using video-level supervision. They tend to only locate the most discriminative parts of action instances and produce temporally incomplete detection results. A natural solution for this problem, the adversarial erasing strategy, is to remove such parts from training so that models can attend to complementary parts. Previous works do it in an offline and heuristic way. They adopt a multi-stage pipeline, where discriminative regions are determined and erased under the guidance of detection results from last stage. Such a pipeline can be both ineffective and inefficient, possibly hindering the overall performance. On the contrary, we combine adversarial erasing with dropout mechanism and propose a Temporal Dropout Module that learns where to remove in a data-driven and online manner. This plug-and-play module is trained without iterative stages, which not only simplifies the pipeline but also makes the regularization during training easier and more adaptive. Experiments show that the proposed method outperforms previous erasing-based methods by a large margin. More importantly, it achieves universal improvement when plugged into various direct classification methods and obtains state-of-the-art performance. Chi Xie 0001, Zikun Zhuang, Shengjie Zhao 0001, Shuang Liang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2022 | Multi-mode Light: Learning Special Collaboration Patterns for Traffic Signal Control
Shengjie Zhao 0001, Hao Deng 0002 |
ICANN (2) | 2 |
| 2022 | A Graph Convolutional Stacked Temporal Attention Neural Network for Traffic Flow ForecastingabstractAs the foundation of route planning and applications for intelligent transportation systems, accurate spatio-temporal traffic forecasting plays an essential role in improving both road utilization and traffic safety. Recently, the graph convolution network (GCN), recurrent neural network (RNN) and many other deep-learning based methods have been adopted for traffic flow forecasting and performed much better than the conventional statistical approaches. However, some node information may be lost during the propagation in graph convolutional layers, and the existing methods are insufficient to model the temporal dependencies especially for long-range sequences. In order to address these deficiencies, we innovatively come up with a graph convolutional stacked temporal attention neural network (GSTA), which can simultaneously extract the spatial and temporal features to forecast the traffic flow with higher accuracy. Specifically, our proposed framework uses a mix-hop GCN to better capture the spatial dependencies by preserving more useful information compared with the traditional GCN. Moreover, to identify the relations among traffic flow data over different time steps, we adopt an attention mechanism and introduce the temporal feature through embedding technology to capture the temporal regularity. We evaluate the proposed GSTA on two real-world traffic datasets, and the experimental results demonstrate the performance of our proposed model is significantly superior to several existing methods. Yushan Feng, Fengxia Han, Shengjie Zhao 0001 |
IJCNN | 3 |
| 2022 | Federated Semi-Supervised Learning Through a Combination of Self and Cross Model EnsemblingabstractMedical image segmentation (MIS) plays a vital role in modern computer-aided diagnosis systems. Deep learning technology has achieved promising results in MIS in recent years. However, deep learning (DL) is a data-hungry technology and in the domain of healthcare, a single hospital usually cannot afford to collect adequate medical images to train a robust DL model. Moreover, a medical image usually contains sensitive data related to patients' privacy, making it's infeasible to build a larger dataset by collecting images from different hospitals. Federated learning (FL), a recently proposed privacy-protecting collaborative paradigm aiming at allowing different data owners to collaboratively train a model without exposing raw data, seems to be a proper solution to these problems. Some recent works have verified the feasibility of applying FL to MIS but most of these works only confine to fully supervised scenarios. Unfortunately, most hospitals in realistic usually cannot provide fully labeled data due to lack of labor. In this paper, we study a challenging but more practical problem in which each hospital can only provide a few labeled data combined with some other unlabeled data. To effectively handle such a problem, we propose a novel and robust federated semi-supervised learning (FSSL) framework, which improves over the mean teacher mechanism with a cross-clients ensemble module and a model-wise self-ensembling module. We evaluate our method on two public medical image datasets and the results show that, in the challenging FSSL scenario, our method can effectively leverage unlabeled data to boost the model performance by a considerable margin. Notably, our method also outperforms other existing FSSL approaches designed for MIS. Tingjie Wen, Shengjie Zhao 0001, Rongqing Zhang 0001 |
IJCNN | 2 |
| 2022 | Joint Transmit Power and Trajectory Optimization for Two-Way Multihop UAV Relaying NetworksabstractUnmanned aerial vehicle (UAV) has been more and more widely used in military and civilian, with its unique advantages of flexibility, convenience, and wide coverage. As flying stations, UAVs can quickly set up relay communication links for different missions, to enhance the receiving signal power, increase the system capacity, and expand the communication coverage. In this article, we investigate a two-way multihop UAV relaying network, where there are two ground users as sources and multiple UAVs as relays to help the two ground sources exchange information. For the purpose of enhancing the efficiency of the investigated UAV-assisted relaying, we come up with a productive two-way multihop UAV relaying pattern, which can achieve a data rate of$({1}/{2})$data packets per time slot with the decode-and-forward protocol. Then, we further formulate a joint transmit power and trajectory optimization problem for the UAVs in this two-way multihop relaying scenario. The formulated problem is nonconvex which makes it difficult to solve directly; hence, we propose an iterative algorithm to obtain an approximate optimal solution based on block coordinate descent and successive convex optimization techniques. Numerical results demonstrate that our proposed two-way multihop UAV relaying network achieves significant throughput gains compared with other benchmark schemes. Bing Li 0025, Shengjie Zhao 0001, Rongqing Zhang 0001, Liuqing Yang 0001 |
IEEE Internet Things J. | 2 |
| 2022 | Hyper-clustering enhanced spatio-temporal deep learning for traffic and demand prediction in bike-sharing systems
Shengjie Zhao 0001, Kai Zhao 0011, Yusen Xia, Wenzhen Jia |
Inf. Sci. | 1 |
| 2022 | Rectified Neighborhood Construction for Robust Feature Matching With Heavy OutliersabstractThis letter is concerned with constructing reliable neighborhoods for the local consistency-based feature matching methods. To alleviate the impact of outliers on neighborhood construction, we propose a rectified neighborhood construction strategy (RNC), which can effectively enlarge the distribution between inliers and outliers. Besides, we also integrate an adaptive parameter estimation into the aforementioned rectified strategy, and it can contribute to determining a reasonable parameter for the rectified strategy. Finally, the experimental results on two representative remote sensing image data sets show that the proposed method can achieve satisfactory feature matching results compared with some state-of-the-arts. Yizhang Liu, Brian Nlong Zhao, Shengjie Zhao 0001 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2022 | Progressive Motion Coherence for Remote Sensing Image MatchingabstractIn this article, we present a feature-based remote sensing (RS) image matching method termed progressive motion coherence (PMC). We formulate the matching problem into a mathematical model and derive a closed-form solution. The objective function is only based on two novel coherence constraints, namely, efficient neighborhood element coherence and relative order-aware motion coherence, and hence, it is general enough and can be applied to RS image matching with different image types and degradations. The efficient neighborhood element coherence uses the Jaccard distance to measure the dissimilarity of two neighborhoods, which are lists composed of$k$nearest neighbors of feature points. To prevent overpenalization on the outliers, we combine it with an exponential function, which is simple yet efficient. The relative order-aware motion coherence is an alternative to motion smoothness, which is based on the observation that the relative order of neighboring matches for inliers in a small region can be well preserved, while for outliers, the relative order changes greatly. The above two coherences are robust to large rotation changes and low ratio inliers. Extensive experiments on five RS image datasets compared with seven state of the arts demonstrate that our PMC is more efficient and robust than the competitors. Yizhang Liu, Brian Nlong Zhao, Shengjie Zhao 0001, Lin Zhang 0014 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Multi-Vehicle Collaborative Learning for Trajectory Prediction With Spatio-Temporal Tensor FusionabstractAccurate behavior prediction of other vehicles in the surroundings is critical for intelligent transportation systems. Common practices to reason about the future trajectory are through their historical paths. However, the impact of traffic context is ignored, which means the beneficial environment information is deserted. Although a few methods are proposed to exploit the surrounding vehicle information, they simply model the influence according to spatial relations without considering the temporal information among them. In this paper, a novel multi-vehicle collaborative learning with spatio-temporal tensor fusion model for vehicle trajectory prediction is proposed, which introduces a novel auto-encoder social convolution mechanism and a fancy recurrent social mechanism to model spatial and temporal information among multiple vehicles, respectively. Furthermore, the generative adversarial network is incorporated into our framework to handle the inherent multi-modal characteristics of the agent motion behavior. Finally, we evaluate the proposed multi-vehicle collaborative learning model on NGSIM US-101 and I-80 benchmark datasets. Experimental results demonstrate that the proposed approach outperforms the state-of-the-art for vehicle trajectory prediction. Additionally, we also present qualitative analyses of the multi-modal vehicle trajectory generation and the impacts of surrounding vehicles on trajectory prediction under various circumstances. Yu Wang 0174, Shengjie Zhao 0001, Rongqing Zhang 0001, Xiang Cheng 0001, Liuqing Yang 0001 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2022 | Online Correction of Camera Poses for the Surround-view System: A Sparse Direct ApproachabstractThe surround-view module is an indispensable component of a modern advanced driving assistance system. By calibrating the intrinsics and extrinsics of the surround-view cameras accurately, a top-down surround-view can be generated from raw fisheye images. However, poses of these cameras sometimes may change. At present, how to correct poses of cameras in a surround-view system online without re-calibration is still an open issue. To settle this problem, we introduce the sparse direct framework and propose a novel optimization scheme of a cascade structure. This scheme is actually composed of two levels of optimization and two corresponding photometric error based models are proposed. The model for the first-level optimization is called the ground model, as its photometric errors are measured on the ground plane. For the second level of the optimization, it’s based on the so-called ground-camera model, in which photometric errors are computed on the imaging planes. With these models, the pose correction task is formulated as a nonlinear least-squares problem to minimize photometric errors in overlapping regions of adjacent bird’s-eye-view images. With a cascade structure of these two levels of optimization, an appropriate balance between the speed and the accuracy can be achieved. Experiments show that our method can effectively eliminate the misalignment caused by cameras’ moderate pose changes in the surround-view system. Source code and test cases are available online at https://cslinzhang.github.io/CamPoseCorrection/ . Tianjun Zhang, Hao Deng 0002, Lin Zhang 0014, Shengjie Zhao 0001, Xiao Liu 0030, Yicong Zhou |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2022 | Flexible and Reliable Multiuser SWIPT IoT Network Enhanced by UAV-Mounted Intelligent Reflecting SurfaceabstractIntelligent reflecting surface (IRS) cooperated with the simultaneous wireless information and power transfer (SWIPT) can reinforce the desired signal and deal with the energy supply problem effectively. By leveraging the on-demand mobility of unmanned aerial vehicles (UAVs), the IRS cooperated SWIPT can be deployed in more flexible and reliable scenarios. In this article, we investigate the UAV-mounted IRS-assisted SWIPT for Internet of Things (IoT) networks. In particular, a UAV-mounted IRS is deployed to assist the information transmission and power transfer from the access point to several IoT devices simultaneously. Taking full advantage of the UAV-mounted IRS in attending multiple IoT devices flexibly, a time division multiple access (TDMA)-based scheduling protocol is proposed to serve different IoT devices alternatively during the UAV flying along an optimized trajectory with the information and power transfer executed. Then, an optimization problem of maximizing the minimum average achievable rate of multiple devices is formulated with the specific energy harvesting requirement guaranteed. To solve the nonconvex problem, we leverage the successive convex approximation and block coordinate descent methods to develop an iterative algorithm. Simulation results demonstrate that with the help of the more flexible and reliable UAV-mounted IRS, the minimum achievable rate of the IoT network can be significantly improved. Yuan Liu 0030, Fengxia Han, Shengjie Zhao 0001 |
IEEE Trans. Reliab. | 3 |
| 2021 | On improving the cooperative localization performance for IoT WSNs
Feng Yan 0004, Shengjie Zhao 0001, Song Xing, Lianfeng Shen |
Ad Hoc Networks | 3 |
| 2021 | A survey on unmanned aerial vehicle relaying networksabstractAbstract With the explosive growth of data communications, existing infrastructure networks are under ever‐increasing pressure. Due to the advantages of fully controllable mobility, rapid deployment, and low cost, the unmanned aerial vehicles (UAVs) have attracted much attentions from both industry and academia in recent years, and it has become an inevitable trend to employ UAVs to enhance the network performance in different environments. As an important paradigm of UAV‐assisted communications, UAV relaying communications has been regarded as a promising solution in enhancing connectivity and improving transmission rate. This paper for the first time comprehensively summarizes UAV relaying communications and its application scenarios, including single UAV relaying networks, multi‐user UAV relaying networks, multi‐hop UAV relaying networks, as well as Internet of UAVs, and deeply analyzes the key technologies and challenges to be solved under this topic. Furthermore, the state‐of‐the‐art researches and opportunities of UAV relaying communications are discussed in detail. Bing Li 0025, Shengjie Zhao 0001, Ruiqin Miao, Rongqing Zhang 0001 |
IET Commun. | 2 |
| 2021 | An efficient multi-sensor fusion and tracking protocol in a vehicle-road collaborative systemabstractAbstract Nowadays, driving safety has become an important topic in the field of intelligent driving. The key step in the intelligent transportation system (ITS) is to detect and track the vehicles on the road through various sensors equipped on the vehicle. While putting the sensors aside the road can have benefits in the vehicle‐road collaborative scenario. In this paper, an efficient and applicable multi‐sensor fusion protocol for a vehicle‐road collaborative system is proposed where roadside units (RSUs) are equipped with cameras and radars. To fully use the history information and the multi‐source heterogeneous sensing data to improve the reliability of multi‐sensor fusion, a Bayesian‐based two‐layer fusion scheme is further proposed. Moreover, the efficiency and robustness of the proposed scheme are evaluated in a real highway environment in Changsha, China. Zhehan Xu, Shengjie Zhao 0001, Rongqing Zhang 0001 |
IET Commun. | 2 |
| 2021 | UAV-Assisted Data Collection With Nonorthogonal Multiple AccessabstractUnmanned aerial vehicles (UAVs) facilitate information collection greatly in the Internet-of-Things (IoT) systems due to their superior flexibility and mobility. On the other hand, nonorthogonal multiple access (NOMA) is regarded as a promising technology to provide high spectral efficiency and support massive connectivity in fifth-generation networks. The integration of NOMA into UAV-assisted wireless networks shows great potential, but how to determine the user grouping and power allocation in NOMA according to the high mobility of UAV is challenging. In this article, we propose a general NOMA-enabled UAV-assisted data collection (NUDC) protocol to maximize the sum rate of a wireless sensor network (WSN), where the location of UAV, sensor grouping, and power control are jointly considered. Moreover, a joint signal-to-interference ratio (SIR) hypergraph-based grouping and power control (SHG-PC) NOMA scheme is provided to obtain the appropriate sensor grouping and the optimal power control solutions efficiently, in which the hypergraph and the greedy coloring algorithm are exploited to find out the optimized group relationships. Extensive simulation results demonstrate the efficiency of our proposed protocol. Weichao Chen 0001, Shengjie Zhao 0001, Rongqing Zhang 0001, Yi Chen 0013, Liuqing Yang 0001 |
IEEE Internet Things J. | 2 |
| 2021 | Full-Duplex UAV Relaying for Multiple User PairsabstractBased on the advantages of small size, lightweight, as well as flexible deployment and recycling, unmanned aerial vehicle (UAV) has been more and more widely used in military and civilian. As flying relays, UAVs can quickly set up relay communication links for different missions, to enhance the receiving signal power, increase the system capacity, and expand the communication coverage. In this article, we investigate full-duplex (FD) UAV relaying for multiple source-destination pairs. To fully exploit the flying flexibility of the UAV in serving multiple source-destination pairs, we propose a scheduling protocol that exploits time-division multiple access (TDMA) to serve different source-destination pairs in turns when flying along an optimized trajectory. Then, we further formulate a joint optimization problem of the TDMA-based user scheduling, the dynamic UAV trajectory, and the UAV transmit power to maximize the system throughput. The formulated problem is nonconvex that makes it difficult to solve directly, hence we propose an iterative algorithm to obtain an approximate optimal solution based on block coordinate descent and successive convex optimization techniques. Simulation results demonstrate that our proposed FD-based UAV relaying network achieves significant throughput gains compared with the half-duplex (HD) baseline, and the TDMA-based protocol outperforms the OFDMA-based ones with fixed UAV position/trajectory when the UAV helps relay information for multiple source-destination pairs. Bing Li 0025, Shengjie Zhao 0001, Rongqing Zhang 0001, Liuqing Yang 0001 |
IEEE Internet Things J. | 2 |
| 2021 | Distributed Soft Clustering Algorithm for IoT Based on Finite Time Average ConsensusabstractClustering is a common technique for statistical data analysis and it has been widely used in many fields. This article investigates data clustering over the Internet-of-Things (IoT) network. Facing the IoT network challenges, including data volume, communication latency, and information security, we here propose a distributed soft clustering algorithm for the IoT environments where each IoT node may have data from multiple clusters. Considering that the main task of soft clustering is to compute each cluster center in a weighted averaging fashion, our distributed clustering method resorts to an efficient finite-time average-consensus algorithm. Moreover, to make the distributed clustering algorithm more stable and be able to escape from some bad local optimum, we propose a distributed deterministic initialization method based on data variance partitioning. Experiments show that the proposed distributed soft clustering algorithm can offer the same performance as its centralized counterpart in terms of both convergence and clustering quality. Besides, unlike most clustering methods relying on probabilistic initialization, our algorithm could provide stable clustering quality which makes it more suitable for IoT networks. A real-world case study about the clustering analysis for distributed data sets collected by environmental monitoring stations is offered, which shows the potential of our algorithms in practical applications. Shengjie Zhao 0001, Qingjiang Shi |
IEEE Internet Things J. | 3 |
| 2021 | Generalized User Grouping in NOMA Based on Overlapping Coalition Formation GameabstractNon-orthogonal multiple access (NOMA) is regarded as a promising technology to provide high spectral efficiency and support massive connectivity in 5G systems. In most existing NOMA user grouping approaches, users are grouped into disjoint groups, which may lead to a waste of power resources within each NOMA group. Motivated by this, in this paper we propose a novel generalized user grouping (GuG) concept for NOMA from an overlapping perspective, which allows each user to participate in multiple groups but subject to individual maximum power constraint. In order to achieve effective GuG and maximize the system sum rate, we formulate a joint power control and GuG optimization problem. Then, we address this problem by exploiting the overlapping coalition formation (OCF) game framework, and we further propose an OCF-based algorithm in which each user can be self-organized into a desirable overlapping coalition structure. Simulation results verify the efficiency of GuG in NOMA systems and indicate that compared with traditional NOMA user grouping schemes, our proposed OCF-based GuG NOMA scheme achieves significant performance gains in terms of system sum rate. Weichao Chen 0001, Shengjie Zhao 0001, Rongqing Zhang 0001, Liuqing Yang 0001 |
IEEE J. Sel. Areas Commun. | 2 |
| 2021 | Efficient Selective Context Network for Accurate Object DetectionabstractSingle-stage detectors have gained great attention due to their high detection accuracy and real-time speed. To detect multi-scale objects, single-stage detectors make scale-aware predictions based on multiple pyramid layers. However, the insufficient context exploration in shallow pyramid layers leads to the detection accuracy of small objects being far from satisfactory. To tackle this problem, we propose a scheme to selectively extract multi-scale context with attention-adaptive weights. Specifically, we propose an efficient selective context network for accurate object detection. It incorporates an enhanced context module and a triple attention module. The enhanced context module consists of multi-branches to extract original-scale, small-scale, and large-scale contextual information. To make full use of this context and filter out noisy information, the triple attention module, which contains global-level, channel-level, and spatial-level attentions, is introduced to carry out selective context fusion. The two modules are easy to implement and can efficiently boost the accuracy of object detection. The performance of our method is validated on two benchmarks: PASCAL VOC and MS COCO. For a 512×512 input, our detector with VGG16 achieves competitive results (80.9 on the Pascal VOC 2012 test set in the case of single-scale inference without MS COCO pre-training). On the MS COCO test-dev set, our detector with ResNet101 outperforms RetinaNet500 by 2.5% AP in terms of overall performance and its speed is 48 milliseconds on a Titan XP GPU. As a result, ESCNet achieves a better trade-off between accuracy and speed. Jing Nie 0001, Yanwei Pang, Shengjie Zhao 0001, Jungong Han, Xuelong Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2021 | Generalized User Grouping in NOMA: An Overlapping PerspectiveabstractNon-orthogonal multiple access (NOMA) is regarded as a promising technology to provide high spectral efficiency and support massive connectivity in 5G systems. Traditionally, NOMA user grouping is non-overlapping, leading to a waste of power resources within each NOMA group. Motivated by this, in this paper we propose a novel generalized user grouping (GuG) concept for NOMA from an overlapping perspective, which allows each user to participate in multiple user groups but subject to individual maximum power constraint. In order to achieve effective GuG and maximize the system sum rate, we formulate a joint power control and GuG optimization problem. Then we further provide a machine learning-based GuG scheme to obtain the optimized feasible GuG and the optimal power control solutions efficiently, in which the established machine learning-based model is exploited to explore the relative relationships of channel gains of users and obtain several fixed grouping patterns via Merge operation. Simulation results verify the efficiency of GuG in NOMA systems and indicate that compared with traditional NOMA user grouping schemes, our proposed GuG scheme achieves significant performance gains in terms of system sum rate. Weichao Chen 0001, Shengjie Zhao 0001, Rongqing Zhang 0001, Hong Chen 0003, Liuqing Yang 0001 |
IEEE Trans. Wirel. Commun. | 2 |
| 2020 | Generalized User Grouping in NOMA Based on Overlapping Coalition Formation GameabstractNon-orthogonal multiple access (NOMA) is regarded as a promising technology to provide high spectral efficiency and support massive connectivity in 5G systems. In most existing NOMA user grouping approaches, users are grouped into disjoint groups, which may lead to a waste of power resources within each NOMA group. Motivated by this, in this paper we propose a novel generalized user grouping (GuG) concept for NOMA from an overlapping perspective, which allows each user to participate in multiple groups but subject to individual maximum power constraint. In order to achieve effective GuG and maximize the system sum rate, we formulate a joint power control and GuG optimization problem. Then, we address this problem by exploiting the overlapping coalition formation (OCF) game framework, and we further propose an OCF-based algorithm in which each user can be self-organized into a desirable overlapping coalition structure. Simulation results verify the efficiency of GuG in NOMA systems and show that our proposed OCF-based GuG NOMA scheme achieves significant performance gains in terms of system sum rate. Weichao Chen 0001, Shengjie Zhao 0001, Rongqing Zhang 0001, Yi Chen 0013, Liuqing Yang 0001 |
GLOBECOM | 2 |
| 2020 | Machine Learning-Based Generalized User Grouping in NOMAabstractNon-orthogonal multiple access (NOMA) provides high spectral efficiency and supports massive connectivity in 5G systems. Traditionally, NOMA user grouping is non-overlapping, leading to a waste of power resources within each NOMA group. Motivated by this, we propose a novel generalized user grouping (GuG) concept for NOMA from an overlapping perspective, which allows each user to participate in multiple user groups but subject to individual maximum power constraint. We formulate a joint power control and GuG optimization problem, and then provide a machine learning-based GuG scheme to obtain the optimized feasible GuG and the optimal power control solutions efficiently. Simulation results show significant performance gains in terms of system sum rate. Weichao Chen 0001, Shengjie Zhao 0001, Rongqing Zhang 0001, Yi Chen 0013, Liuqing Yang 0001 |
GLOBECOM | 2 |
| 2020 | Learning-Based Massive BeamformingabstractDeveloping resource allocation algorithms with strong real-time and high efficiency has been an imperative topic in wireless networks. Conventional optimization-based iterative resource allocation algorithms often suffer from slow convergence, especially for massive multiple-input-multiple-output (MIMO) beamforming problems. This paper studies learningbased efficient massive beamforming methods for multi-user MIMO networks. The considered massive beamforming problem is challenging in two aspects. First, the beamforming matrix to be learned is quite high-dimensional in case with a massive number of antennas. Second, the objective is often time-varying and the solution space is not fixed due to some communication requirements. All these challenges make learning representation for massive beamforming an extremely difficult task. In this paper, by exploiting the structure of the most popular WMMSE beamforming solution, we propose convolutional massive beamforming neural networks (CMBNN) using both supervised and unsupervised learning schemes with particular design of network structure and input/output. Numerical results demonstrate the efficacy of the proposed CMBNN in terms of running time and system throughput. Shengjie Zhao 0001, Qingjiang Shi |
GLOBECOM | 2 |
| 2020 | Consensus-Based Distributed Clustering for IoTabstractClustering is a common technique for statistical data analysis and it has been widely used in many fields. When the data is collected via a distributed network or distributedly stored, data analysis algorithms have to be designed in a distributed fashion. This paper investigates data clustering with distributed data. Facing the distributed network challenges including data volume, communication latency, and information security, we here propose a distributed clustering algorithm where each IoT device may have data from multiple clusters. Considering that the main task of clustering is to compute each cluster center in a weighted averaging fashion, our distributed clustering method resorts to an efficient finite-time average-consensus algorithm. Experiments show that the proposed distributed clustering algorithm can offer the same convergence and clustering quality as its centralized counterpart but with less data traffic. Besides, experiments also show that our proposed algorithms outperforms the existing methods. Shengjie Zhao 0001, Qingjiang Shi |
ICASSP | 3 |
| 2020 | Towards Adaptive Semantic Segmentation By Progressive Feature RefinementabstractAs one of the fundamental tasks in computer vision, semantic segmentation plays an important role in real world applications. Although numerous deep learning models have made notable progress on several mainstream datasets with the rapid development of convolutional networks, they still encounter various challenges in practical scenarios. Unsupervised adaptive semantic segmentation aims to obtain a robust classifier trained with source domain data, which is able to maintain stable performance when deployed to a target domain with different data distribution. In this paper, we propose an innovative progressive feature refinement framework, along with domain adversarial learning to boost the transferability of segmentation networks. Specifically, we firstly align the multi-stage intermediate feature maps of source and target domain images, and then a domain classifier is adopted to discriminate the segmentation output. As a result, the segmentation models trained with source domain images can be transferred to a target domain without significant performance degradation. Experimental results verify the efficiency of our proposed method compared with state-of-the-art methods. Shengjie Zhao 0001, Rongqing Zhang 0001 |
ICIP | 2 |
| 2020 | A Study Of Parking-Slot Detection With The Aid Of Pixel-Level Domain AdaptationabstractThe self-parking system is an important component of self-driving vehicles. Such a system needs to detect and locate the parking-slots from surround-view images, and then guide the vehicle to the designated parking-slot. In the real world, the appearances and environmental conditions of parking-slots can be rich and varied. Thus, to train the parking-slot detection model, it is necessary to collect and label a huge quantity of surround-view images covering as many real cases as possible. Such a process is cumbersome and costly, and will be repeated whenever encountering an unseen parking condition that is quite different from the ones covered by existing training set. To this end, in this paper we propose an extensible pipeline, namely FakePS, to assist parking-slot detection model training by making use of synthetic data. Specifically, with FakePS, we can first build various simulated parking scenes and collect labeled surround-view images automatically. Besides, we resort to pixel-level domain adaptation strategies to enhance the realism of the synthetic images using unlabeled real images while preserving their label information. The efficacy of FakePS has been corroborated by experimental results. Lin Zhang 0014, Ying Shen 0005, Yong Ma 0005, Shengjie Zhao 0001, Yicong Zhou |
ICME | 5 |
| 2020 | Oecs: Towards Online Extrinsics Correction For The Surround-View SystemabstractA typical surround-view system consists of four fisheye cameras. By performing an offline calibration that determines both the intrinsics and extrinsics of the system, surround-view images can be synthesized at runtime. However, poses of calibrated cameras sometimes may change. In such a case, if cameras' extrinsics are not updated accordingly, observable geometric misalignment will appear in surround-views. Most existing solutions to this problem resort to re-calibration, which is quite cumbersome. Thus, how to correct cameras' extrinsics in an online manner without using re-calibration is still an open issue. In this paper, we attempt to propose a novel solution to this problem and the proposed solution is referred to as “Online Extrinsics Correction for the Surround-view system OECS for short. We first design a Bi-Camera error model, measuring the photometric discrepancy between two corresponding pixels on images captured by two adjacent cameras. Then, by minimizing the system's overall BiCamera error, cameras' extrinsics can be optimized and the optimization is conducted within a sparse direct framework. The efficacy and efficiency of OECS are validated by experiments. Data and source code used in this work are publicly available at https://z619850002.github.io/OECage/. Tianjun Zhang, Lin Zhang 0014, Ying Shen 0005, Yong Ma 0005, Shengjie Zhao 0001, Yicong Zhou |
ICME | 5 |
| 2020 | Zero-Shot Restoration of Underexposed Images via Robust Retinex DecompositionabstractUnderexposed images often suffer from serious quality degradation such as poor visibility and latent noise in the dark. Most previous methods for underexposed images restoration ignore the noise and amplify it during stretching contrast. We predict the noise explicitly to achieve the goal of denoising while restoring the underexposed image. Specifically, a novel three-branch convolution neural network, namely RRDNet (short for Robust Retinex Decomposition Network), is proposed to decompose the input image into three components, illumination, reflectance and noise. As an image-specific network, RRDNet doesn't need any prior image examples or prior training. Instead, the weights of RRDNet will be updated by a zero-shot scheme of iteratively minimizing a specially designed loss function. Such a loss function is devised to evaluate the current decomposition of the test image and guide noise estimation. Experiments demonstrate that RRDNet can achieve robust correction with overall naturalness and pleasing visual quality. To make the results reproducible, the source code has been made publicly available at https://aaaaangel.github.io/RRDNet-Homepage. Lin Zhang 0014, Ying Shen 0005, Yong Ma 0005, Shengjie Zhao 0001, Yicong Zhou |
ICME | 5 |
| 2020 | Cross-Domain Semantic Segmentation of Urban Scenes via Multi-Level Feature AlignmentabstractSemantic segmentation is an essential task in plenty of real-life applications such as virtual reality, video analysis, autonomous driving, etc. Recent advancements in fundamental vision-based tasks ranging from image classification to semantic segmentation have demonstrated deep learning-based models' high capability in learning complicated representation on large datasets. Nevertheless, manually labeling semantic segmentation dataset with pixel-level annotation is extremely labor-intensive. To address this problem, we propose a novel multi-level feature alignment framework for cross-domain semantic segmentation of urban scenes by exploiting generative adversarial networks. In the proposed multi-level feature alignment method, we first translate images from one domain to another one. Then the discriminative feature representations extracted by the deep neural network are concatenated, followed by domain adversarial learning to make the intermediate feature distribution of the target domain images close to those in the source domain. With these domain adaptation techniques, models trained with images in the source domain where the labels are easy to acquire can be deployed to the target domain where the labels are scarce. Experimental evaluations on various mainstream benchmarks confirm the effectiveness as well as robustness of our approach. Shengjie Zhao 0001, Rongqing Zhang 0001 |
ICPR | 2 |
| 2020 | Block-Diagonal Zero-Forcing Beamforming for Weighted Sum-Rate Maximization in Multi-User Massive MIMO SystemsabstractBeamforming is one of the most important transmission technologies to improve the quality of communication in cellular network systems. However, in massive multiple-input-multiple-output (MIMO) systems, conventional optimization-based iterative beamforming algorithms often suffer from high computational complexity and thus are not suitable for practical applications. This paper focuses on low complexity beamforming technique for multi-user massive MIMO systems. Specifically, a block-diagonal zero-forcing (BD-ZF) beamforming algorithm is proposed for achieving weighted sum-rate maximization. We show that the BD-ZF beamforming problems can be globally solved using water-filling algorithms. Extensive simulations demonstrate that the proposed BD-ZF beamforming method can offer better performance than the state-of-art low complexity beamforming technique ZF (even could sometimes coincide with the performance of the popular WMMSE algorithm) but with only an extra little bit computational overhead. Shengjie Zhao 0001, Qingjiang Shi |
ISCC | 2 |
| 2020 | UAV-Assisted Data Collection with Non-Orthogonal Multiple AccessabstractUnmanned aerial vehicles (UAVs) facilitate information collection greatly in Internet of Things (IoT) systems. On the other hand, non-orthogonal multiple access (NOMA) is regarded as a promising technology to provide high spectral efficiency and support massive connectivity in 5G networks. The integration of NOMA into UAV-assisted wireless networks shows great potential, but how to determine the user grouping and power allocation in NOMA according to the different locations of UAV is challenging. In this paper, we propose a general NOMA-enabled UAV-assisted data collection (NUDC) protocol to solve the formulated sum rate maximization problem such that the location of UAV, sensor grouping, and power control are jointly considered. Moreover, a joint signal-to-interference-ratio (SIR) hypergraph-based grouping and power control (SHG-PC) NOMA scheme is provided to obtain the appropriate sensor grouping and the optimal power control solutions efficiently. Extensive simulation results demonstrate the effectiveness of our proposed protocol. Weichao Chen 0001, Shengjie Zhao 0001, Rongqing Zhang 0001, Liuqing Yang 0001 |
WCNC | 2 |
| 2020 | Mobility Prediction-Based Joint Task Assignment and Resource Allocation in Vehicular Fog ComputingabstractMost recently, vehicular fog computing (VFC) has been regarded as a novel and promising architecture to effectively reduce the computation time of various vehicular application tasks in Internet of vehicles (IoV). However, the high mobility of vehicles makes the topology of vehicular networks change fast, and thus it is a big challenge to coordinate vehicles for VFC in such a highly mobile scenario. In this paper, we investigate the joint task assignment and resource allocation optimization problem by taking the mobility effect into consideration in vehicular fog computing. Specifically, we formulate the joint optimization problem from a Min-Max perspective in order to reduce the overall task latency. Then we decompose the nonconvex problem into two sub-problems, i.e., one to one matching and bandwidth resource allocation, respectively. In addition, considering the relatively stable moving patterns of a vehicle in a short period, we further introduce the mobility prediction to design a mobility prediction-based scheme to obtain a better solution. Simulation results verify the efficiency of our proposed mobility prediction-based scheme in reducing the overall task completion latency in VFC. Xianjing Wu, Shengjie Zhao 0001, Rongqing Zhang 0001, Liuqing Yang 0001 |
WCNC | 2 |
| 2020 | Time-sync comments denoising via graph convolutional and contextual encoding
Zhenyu Liao 0002, Yikun Xian, Chenxi Zhang 0001, Shengjie Zhao 0001 |
Pattern Recognit. Lett. | 5 |
| 2020 | High-Level Semantic Networks for Multi-Scale Object DetectionabstractTo better solve scale variance problem, deep multi-scale methods usually detect objects of different scales by different in-network layers. However, the semantic levels of features from different layers are usually inconsistent. In this paper, we propose a multi-branch and high-level semantic network by gradually splitting a base network into multiple different branches. As a result, the different branches have same depth and the output features of different branches have similarly high-level semantics. Due to the difference of receptive fields, the different branches are suitable to detect objects of different scales. Meanwhile, the multi-branch network does not introduce additional parameters by sharing the convolutional weights of different branches. To further improve detection performance, skip-layer connections are used to add context to the branch of relatively small receptive field, and dilated convolution is incorporated to enlarge the resolutions of output feature maps. When they are embedded into Faster RCNN architecture, the weighted scores of proposal generation network and proposal classification network are further proposed. Experiments on three pedestrian datasets (i.e., the KITTI dataset, the Caltech dataset, and the Citypersons dataset), one face dataset (i.e., the WIDER FACE dataset), and two general object datasets (i.e., the COCO benchmark and the PASCAL VOC dataset) demonstrate the effectiveness and generality of proposed method. On these datasets, our method achieves state-of-the-art performance. Jiale Cao, Yanwei Pang, Shengjie Zhao 0001, Xuelong Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2020 | Dehazing Evaluation: Real-World Benchmark Datasets, Criteria, and BaselinesabstractOn benchmark images, modern dehazing methods are able to achieve very comparable results whose differences are too subtle for people to qualitatively judge. Thus, it is imperative to adopt quantitative evaluation on a vast number of hazy images. However, existing quantitative evaluation schemes are not convincing due to a lack of appropriate datasets and poor correlations between metrics and human perceptions. In this work, we attempt to address these issues, and we make two contributions. First, we establish two benchmark datasets, i.e., the BEnchmark Dataset for Dehazing Evaluation (BeDDE) and the EXtension of the BeDDE (exBeDDE), which had been lacking for a long period of time. The BeDDE is used to evaluate dehazing methods via full reference image quality assessment (FR-IQA) metrics. It provides hazy images, clear references, haze level labels, and manually labeled masks that indicate the regions of interest (ROIs) in image pairs. The exBeDDE is used to assess the performance of dehazing evaluation metrics. It provides extra dehazed images and subjective scores from people. To the best of our knowledge, the BeDDE is the first dehazing dataset whose image pairs were collected in natural outdoor scenes without any simulation. Second, we provide a new insight that dehazing involves two separate aspects, i.e., visibility restoration and realness restoration, which should be evaluated independently; thus, to characterize them, we establish two criteria, i.e., the visibility index (VI) and the realness index (RI), respectively. The effectiveness of the criteria is verified through extensive experiments. Furthermore, 14 representative dehazing methods are evaluated as baselines using our criteria on BeDDE. Our datasets and relevant code are available at https://github.com/xiaofeng94/BeDDE-for-defogging. Shiyu Zhao 0001, Lin Zhang 0014, Shuaiyi Huang, Ying Shen 0005, Shengjie Zhao 0001 |
IEEE Trans. Image Process. | 5 |
| 2019 | Seamless 3D Surround View with a Novel Burger ModelabstractIn recent years, the 3D surround view (3D-SV) system has become a hot research topic in the field of Advanced Driver Assistance Systems (ADAS). It can be used to form a stereoscopic view of the surrounding 3D environment by using 4 car-mounted surround cameras, and users can switch the viewpoints for virtual observation conveniently. However, there are still many problems in how to stitch calibrated images to the panoramic view and how to project the panorama to the surround view. In this paper, we introduce the graph cut algorithm and multi-band blending to the panorama stitching phase. In addition, we design a new hamburger-shaped 3D geometric model to be the carrier of the panorama for texture mapping. Our 3D-SV system can make drivers have an immersive visual experience. Experimental results show that the 3D-SV generated by our method is less distorted and looks more natural than the other competitors. Lin Zhang 0014, Ying Shen 0005, Shengjie Zhao 0001 |
ICIP | 5 |
| 2019 | DMPR-PS: A Novel Approach for Parking-Slot Detection Using Directional Marking-Point RegressionabstractThe self-parking system plays an important role in autonomous driving, and one of its critical issues is parking-slot detection. Previous studies in this field are mostly based on off-the-shelf models designed for universal purposes, which have various limitations in solving specific problems. In this paper, we propose a parking-slot detection method using directional marking-point regression, namely DMPR-PS. Instead of utilizing multiple off-the-shelf models, DMPR-PS uses a novel CNN-based model specially designed for directional marking-point regression. Given a surround-view image I, the model predicts position, shape and orientation of each marking-point on I. From marking-points, parking-slots on I could be easily inferred using geometric rules. DMPR-PS outperforms state-of-the-art competitors on the benchmark dataset with a precision rate of 99.42% and a recall rate of 99.37%, while achieving a real-time detection speed of 12ms per frame on Nvidia Titan Xp. To make the results reproducible, the source code is available at https://github.com/Teoge/DMPR-PS. Lin Zhang 0014, Ying Shen 0005, Shengjie Zhao 0001, Yukai Yang |
ICME | 5 |
| 2019 | Revisit Surround-view Camera System CalibrationabstractThe surround-view system is an essential component of an advanced driver assistance system especially when the vehicle runs in tight parking space or on a narrow road. To ensure successful maneuvering, a panoramic bird's-eye image with no blind spots is necessarily called for. Hence, a typical surround-view system consists of several cameras mounted around the vehicle capturing images from a top-down viewpoint, and an accurate extrinsic calibration for such system is prerequisite for providing a seamless surround-view image. To achieve this goal, this paper presents a novel extrinsic calibration pipeline which is both easy-to-use and reliable to operate on multiple cameras. Instead of taking the vehicle to a fixed position in a specific calibration site, a single chessboard is the only demand. We adopt a novel refinement procedure that jointly optimizes camera poses in a closed-loop manner. The effectiveness and efficiency of the proposed pipeline to calibrate a surround-view camera system has been corroborated by experiments. Xuan Shao, Xiao Liu 0030, Lin Zhang 0014, Shengjie Zhao 0001, Ying Shen 0005, Yukai Yang |
ICME | 4 |
| 2019 | Pay By Showing Your Palm: A Study of Palmprint Verification on Mobile PlatformsabstractWith the fast development of smart mobile devices, mobile phones have gradually become an indispensable part of people's lives. Many biometric technologies based on mobile platforms have also developed rapidly, such as face verification and fingerprint recognition. However, the great potential of palmprint has been neglected. In this paper, we conducted a thorough study of palmprint verification on mobile devices for the first time. Firstly, we established an annotated, palmprint dataset named MPD, which was collected by multi-brands phones in two different sessions. As the largest dataset in this field, MPD contains 16,000 palm images from 200 subjects. Secondly, we built a DCNN-based palmprint verification system named DeepMPV for mobile platforms. The efficiency and performance of our system have been corroborated on our collected dataset. The labelled dataset and the source code are publicly available at https://cslinzhang.github.io/deepmpv/. Lin Zhang 0014, Xiao Liu 0030, Shengjie Zhao 0001, Ying Shen 0005, Yukai Yang |
ICME | 4 |
| 2019 | Evaluation of Defogging: A Real-World Benchmark Dataset, A New Criterion and BaselinesabstractModern defogging methods are able to achieve very comparable results whose differences are too subtle for people to qualitatively judge. On the other hand, existing quantitative evaluation methods are also not convincing due to a lack of proper datasets. In this work, we attempt to address these issues and establish a long-term lacking benchmark dataset, namely BeDDE (BEnchmark Dataset for Defogging Evaluation), for evaluating the performance of defogging algorithms. To our knowledge, BeDDE is the first real-world dataset comprising foggy images with their registered clear counterparts. Using BeDDE, we set up a new criterion for evaluating defogging methods where VSI, a full reference image quality assessment metric, is calculated and averaged on registered ROIs of all image pairs. The evaluation results of the proposed criterion correlate well with human judgements. 10 state-of-the-art defogging methods are evaluated as baselines on BeDDE. BeDDE is available online. Shiyu Zhao 0001, Lin Zhang 0014, Shuaiyi Huang, Ying Shen 0005, Shengjie Zhao 0001, Yukai Yang |
ICME | 5 |
| 2019 | Online Camera Pose Optimization for the Surround-view SystemabstractSurround-view system is an important information medium for drivers to monitor the driving environment. A typical surround-view system consists of four to six fish-eye cameras arranged around the vehicle. From these camera inputs, a top-down image of the ground around the vehicle, namely the surround-view image can be generated with well calibrated camera poses. Although existing surround-view system solutions can estimate camera poses accurately in off-line environment, how to correct the camera poses' change in online environment is still an open issue. In this paper, we propose a camera pose optimization method for surround-view system in online environment. Our method consists of two models: Ground Model and Ground-Camera Model, both of which correct the camera poses by minimizing photometric errors between ground projections of adjacent cameras. Experiments show that our method can effectively correct the geometric misalignment of the surround-view image caused by camera poses' change. Since our method is highly automated with low requirement of calibration site and manual operation, it has a wide range of applications and is convenient for the end-users. To make the results reproducible, the source code is publicly available at https://cslinzhang.github.io/CamPoseOpt/. Xiao Liu 0030, Lin Zhang 0014, Ying Shen 0005, Shaoming Zhang, Shengjie Zhao 0001 |
ACM Multimedia | 5 |
| 2019 | Zero-Shot Restoration of Back-lit Images Using Deep Internal LearningabstractHow to restore back-lit images still remains a challenging task. State-of-the-art methods in this field are based on supervised learning and thus they are usually restricted to specific training data. In this paper, we propose a "zero-shot" scheme for back-lit image restoration, which exploits the power of deep learning, but does not rely on any prior image examples or prior training. Specifically, we train a small image-specific CNN, namely ExCNet (short for Exposure Correction Network) at test time, to estimate the "S-curve" that best fits the test back-lit image. Once the S-curve is estimated, the test image can be then restored straightforwardly. ExCNet can adapt itself to different settings per image. This makes our approach widely applicable to different shooting scenes and kinds of back-lighting conditions. Statistical studies performed on 1512 real back-lit images demonstrate that our approach can outperform the competitors by a large margin. To the best of our knowledge, our scheme is the first unsupervised CNN-based back-lit image restoration method. To make the results reproducible, the source code is available at https://cslinzhang.github.io/ExCNet/. Lin Zhang 0014, Lijun Zhang 0005, Xiao Liu 0030, Ying Shen 0005, Shaoming Zhang, Shengjie Zhao 0001 |
ACM Multimedia | 6 |
| 2019 | Semantic Prior Guided Face InpaintingabstractFace inpainting is a sub-task of image inpainting designed to repair broken or occluded incomplete portraits. Due to the high complexity of face image details, inpainting on the face is more difficult. At present, face-related tasks often draw on excellent methods from face recognition and face detection, using multitasking to boost its effect. Therefore, this paper proposes to add the face prior knowledge to the existing advanced inpainting model, combined with perceptual loss and SSIM loss to improve the model repair efficiency. A new face inpainting process and algorithm is implemented, and the repair effect is improved. Xiaobo Zhou 0001, Shengjie Zhao 0001, Xiaoyan Zhang 0002 |
MMAsia | 3 |
| 2019 | Anomaly detection for cellular networks using big data analyticsabstractBroadband connectivity and mobile technology have been widely applied in the world. With these advanced technologies, the proliferation of smart devices and their applications by accessing mobile internet have come up with a giant leap forward, leading to the ever‐increasing scale and complexity of cellular networks. This presents imminent challenges to anomaly detection in cellular networks. In this study, the authors discuss challenges and current literature of anomaly detection for cellular networks to embrace the ‘big data’ era. First, they review the state‐of‐the‐art techniques in the area of anomaly detection in cellular networks. Then, the challenges are pinpointed for anomaly detection due to the cellular network big data. Finally, they introduce a big data analytic‐based anomaly detection method for cellular networks. Bing Li 0025, Shengjie Zhao 0001, Rongqing Zhang 0001, Qingjiang Shi, Kai Yang 0001 |
IET Commun. | 2 |
| 2019 | Channel-Aware D2D-Assisted Wireless Distributed Storage SystemsabstractDevice-to-device (D2D) communications and distributed storage are enabling technologies for the future Internet of Things systems. In this article, we consider power-efficient content delivery in a D2D-assisted wireless distributed storage system, where the partial downloading scheme is employed such that the content requester can download a portion of the stored content from neighboring storage devices. The basic idea is that by downloading a small amount of data from many storage devices, the power consumption is much lower than downloading the entire content from a single device. In designing such a system, we aim to minimize the total power consumption by properly allocating the channel and the amount of transmitted data for each storage device. Moreover, to account for the case that there are more devices than communication channels, we allow multiple devices to share the same channel by employing successive interference cancelation (SIC) decoding. The optimization problem is an integer program and by taking the alternative minimization approach, we decouple it into two subproblems: 1) the packet allocation subproblem, for which we provide the optimal solution and 2) the channel allocation subproblem, for which we provide the efficient suboptimal solution. Fengxia Han, Xiaodong Wang 0001, Shengjie Zhao 0001 |
IEEE Internet Things J. | 3 |
| 2018 | Design of Network Coding for Wireless Broadcast and Multicast With Optimal DecodersabstractThis paper considers the design of network coding schemes for reliable wireless broadcast and multicast transmissions, in which the same packet is broadcast to a group of receivers. Network coding across multiple broadcasted packets is employed to generate redundant packets for the broadcast retransmissions so that the lost packets can be recovered. It is assumed that optimal decoders are employed at the receivers and the focus is on the design of short block codes with small numbers of redundant bits. To this end, use if first made of the residual graph representation to calculate the error probability of the optimal decoder. Then two code design schemes are proposed to minimize the error probability, including a low-complexity deterministic greedy code design algorithm as well as a stochastic code construction algorithm inspired by the simulated annealing technique. Extensive simulation studies have been carried out to assess the performance of the proposed schemes. It is seen that for a given number of retransmissions, the proposed network coding schemes can considerably increase the average number of recovered packages per user at the receivers and thereby improve the spectral efficiency over traditional coding methods. Guosen Yue, Kai Yang 0001, Shengjie Zhao 0001, H. Vincent Poor |
IEEE Trans. Wirel. Commun. | 3 |
| 2017 | Decentralized Beamforming for Weighted Sum Energy Efficiency Maximization in MIMO SystemsabstractThis paper considers a joint transceiver design for the weighed sum energy efficiency maximization problem for downlink transmissions in a multi-cell multi-user multiple-input multiple-output (MIMO) system. To make the formulated problem be more valuable in practice, a more practical power consumption model is adopted in which part of the processing power is dependent on data rate. With per-user rate requirements and per-base station (BS) power constraints, the resulting optimization problem is non-convex. A centralized solution is firstly proposed, which alternatively performs transmit beamforming optimization and receive beamforming optimization with the help of successive convex approximation (SCA) based iteration. Based on the centralized design that requires global channel state information (CSI), a more meaningful decentralized algorithm is further designed, which solves the problem with the requirement of only exchanging quite a small amount of information among BSs. Specifically, the key methodology used in the decentralized algorithm is to reformulate transmit beamforming optimization as an equivalent global consensus problem when receive beamforming is fixed as minimum mean square error (MMSE) receiver, which is solved effectively by exploiting the alternative direction method of multipliers (ADMM) technique. Numerical results demonstrate the effectiveness and superiority of our proposed algorithms, and also illustrate the impact of the rate-dependent power consumption on the energy efficiency. Fengxia Han, Shengjie Zhao 0001, Lu Zhang 0015, Kai Yang 0001 |
GLOBECOM | 2 |
| 2017 | Distributed Adaptive Range Extension Setting for Small Cells in Heterogeneous Cellular NetworkabstractFor heterogeneous network, which has been deployed by 4G systems in a small scale and is viewed as one pioneering technology for making cellular networks be evolved into 5G systems, more efficiently achieving load balance between Macrocells and small cells has attracted increasing attention. Specifically, when Macrocells and Picocells use the completely same frequency spectrum resource, range extension (RE) of Picocells is the key approach for load balancing. It has been observed that, when increasing the RE bias for augmenting the load balancing benefit, the link failure probability of downlink control channel will increase for user equipments (UEs) associated to Picocells because of inter-cell interference. Enhanced inter-cell interference cancellation (eICIC) techniques have been developed to solve this problem, where the time division multiplexing (TDM) based techniques have been mainly focused. However, due to reducing available time-domain resources, the TDM based eICIC techniques will cause inherent performance loss in Macrocells. In this paper, without exploiting any existing eICIC techniques, the aforementioned problem is solved via designing one adaptive per-Picocell RE biasing scheme. In this scheme, with the awareness of unacceptable interference, each Picocell wisely and autonomously adjusts its RE bias among a set of candidate values. Furthermore, this scheme is implemented in a distributed way, without the need of global optimization. Simulation results verify good performance of the proposed scheme, in terms of ICIC effect and load balance benefit. Lu Zhang 0015, Shengjie Zhao 0001, Peng Shang, Jimin Liu, Fengxia Han |
VTC Spring | 2 |
| 2016 | An efficient anonymous authentication protocol using batch operations for VANETs
Yawei Liu, Zongjian He, Shengjie Zhao 0001 |
Multim. Tools Appl. | 3 |
| 2016 | Identifying Region-Wide Functions Using Urban Taxicab TrajectoriesabstractWith the urban development and enlargement, various regions such as residential zones and administrative districts now appear as parts of cities. People exhibit different mobility patterns in each region, which is closely relevant to region-wide functions. In this article, we propose a scheme to discover region-wide functions using large-scale Shanghai taxicab trajectories that capture enormous traces for more than 13,000 taxicabs over a period of about 3 years. We investigate these taxicab trajectories and conduct an extensive preliminary study. Then, we divide the city into disjointed regions using Voronoi decomposition. By incorporating people's pick-up and drop-off information, we refine the Voronoi partitioning results to identify region-wide functional areas. Finally, we study people's movement frequency on weekdays and weekends for every kind of urban functional regions. We also look into human mobility within or across the identified urban functional regions. Experimental results show that human movement is bounded with the function of urban regions, and more than 90% of people visit neighboring (less than 20km travel distance) functional regions with high probability. Daqiang Zhang 0001, Jiafu Wan, Zongjian He, Shengjie Zhao 0001, Sang Oh Park |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2015 | An Efficient RFID Search Protocol Based On Clouds
Daqiang Zhang 0001, Yuming Qian, Jiafu Wan, Shengjie Zhao 0001 |
Mob. Networks Appl. | 4 |
| 2015 | NextMe: Localization Using Cellular Traces in Internet of ThingsabstractThe Internet of Things (IoT) opens up tremendous opportunities to location-based industrial applications that leverage both Internet-resident resources and phones' processing power and sensors to provide location information. Location-based service is one of the vital applications in commercial, economic, and public domains. In this paper, we propose a novel localization scheme called NextMe, which is based on cellular phone traces. We find that the mobile call patterns are strongly correlated with the co-locate patterns. We extract such correlation as social interplay from cellular calls, and use it for location prediction from temporal and spatial perspectives. NextMe consists of data preprocessing, call pattern recognition, and a hybrid predictor. To design the call pattern recognition module, we introduce the notions of critical calls and corresponding patterns. In addition, NextMe does not require that the cell tower addresses should be bounded with concrete coordinates, e.g., global positioning system (GPS) coordinates. We validate NextMe across MIT Reality Mining Dataset, involving 500 000 h of continuous behavior information and 112 508 cellular calls. Experimental results show that NextMe achieves fine-grained prediction accuracy at cell tower level in the forthcoming 1-6 h with 12% accuracy enhancement averagely from cellular calls. Daqiang Zhang 0001, Shengjie Zhao 0001, Laurence T. Yang, Min Chen 0003, Yunsheng Wang 0001, Huazhong Liu |
IEEE Trans. Ind. Informatics | 2 |
| 2014 | A novel Frequency Selective sounding scheme for TDD LTE-advanced systemsabstractIn this paper we propose a novel Frequency Selective SRS (FS-SRS) scheme to improve the quality of CSI. According to the CQIs (channel quality indications), which are periodically reported by each user, and the bandwidth requirement of each user, BSs dynamically schedule each user to send SRSs only on the requisite bandwidth with best CQIs instead of on full bandwidth or on specified sub-bandwidth as in the current LTE TDD systems. Simulation results show that the proposed FS-SRS scheme outperforms the current schemes in estimation accuracy for various frequency-selective and/or time-selective fading channels. In addition, the proposed FS-SRS scheme offers unique advantages over the current schemes. Firstly, it is robust to frequency selective channels. Secondly, it is hardly impacted by timing offset. Thirdly, it is adaptive to the change of the underlying scheduler. Fourthly, it can increase SRS capacity. Finally, it can be extended to multi-user MIMO (multiple-input multiple-output) systems, called extended FS-SRS scheme. Moreover, the proposed schemes are applicable to the uplink SRS design in LTE-Advanced systems. Shengjie Zhao 0001, Baolong Zhou, Daqiang Zhang 0001 |
GLOBECOM | 1 |
| 2012 | A novel CQI-assisted sounding design for TDD LTE-Advanced systemabstractIn the long term evolution (LTE) time division duplex (TDD) systems, base stations use uplink sounding reference signals (SRSs) to estimate downlink channel state information (CSI) for downlink beamforming transmissions. Because SRSs sent by up to 8 users are multiplexed on the same time-frequency resources, there exist mutual interferences among these SRSs over fading channels. Thus this scheme inevitably impacts the estimation accuracy of CSI and thereby degrades the beamforming performance. In this paper, we propose a novel sounding scheme to improve the quality of CSI, in which according to the CQIs (channel quality indications) periodically reported by each user and the bandwidth requirement of each user, BS dynamically schedules each user to send SRSs only on the requisite bandwidth with best CQIs instead of on full bandwidth or specified sub-bandwidth as in the current LTE TDD systems. Simulations show that the proposed scheme obtains much better estimation accuracy than the current scheme for various frequency-selective and/or time-selective fading channels, and is robust to frequency selective channels. Therefore the proposed scheme is applicable to the uplink SRS design in LTE-Advanced and other MIMO-OFDM TDD systems. Baolong Zhou, Ling-ge Jiang, Chen He 0001, Shengjie Zhao 0001 |
ICC | 4 |
| 2012 | A Frequency-Domain Sounding Scheme for LTE TDD Beamforming SystemsabstractIn the 3rdgeneration partnership project (3GPP) long term evolution (LTE) time division duplex (TDD) systems, base stations use uplink sounding reference signals (SRSs) to estimate downlink channel state information (CSI) for downlink beamforming transmissions. Because SRSs sent by up to 8 users are multiplexed on the same time-frequency resources, there exist mutual interferences among these SRSs over fading channels. Thus the current scheme inevitably impacts the estimation accuracy of CSI and thereby degrades the beamforming performance. Based on the characteristics of TDD systems, we propose a novel sounding scheme to improve the quality of CSI, in which according to the CQIs (channel quality indications) periodically reported by each user and the bandwidth requirement of each user, BS schedules each user to send SRSs only on the requisite bandwidth with best CQIs instead of on full bandwidth or specified sub-bandwidth as in the current LTE TDD systems. Simulations show that the proposed scheme outperforms the current scheme in estimation accuracy for various frequency-selective and/or time-selective fading channels, and is especially robust to frequency selective channels. The proposed scheme has been adopted as a candidate scheme for 3GPP LTE-Advanced. Baolong Zhou, Ling-ge Jiang, Shengjie Zhao 0001, Lu Zhang 0015, Chen He 0001, Zhining Jiang |
VTC Spring | 3 |
| 2012 | Sounding reference signal design for TDD LTE-Advanced systemabstractIn the 3rdgeneration partnership project (3GPP) long term evolution (LTE) time division duplex (TDD) systems, base stations use uplink sounding reference signals (SRSs) to estimate downlink channel state information (CSI) for downlink beamforming transmissions. Because SRSs sent by up to 8 users are multiplexed on the same time-frequency resources, there exist mutual interferences among these SRSs over fading channels. Thus the current scheme inevitably impacts the estimation accuracy of CSI and thereby degrades the beamforming performance. Based on the characteristics of TDD systems, we propose a novel sounding scheme to improve the quality of CSI, in which according to the CQIs (channel quality indications) periodically reported by each user and the bandwidth requirement of each user, BS dynamically schedules each user to send SRSs only on the requisite bandwidth with best CQIs instead of on full bandwidth or specified sub-bandwidth as in the current LTE TDD systems. Simulations show that the proposed scheme outperforms the current scheme in estimation accuracy for various frequency-selective and/or time-selective fading channels, and is especially robust to frequency selective channels. Therefore the proposed scheme is applicable to the uplink SRS design in LTE-Advanced and has been adopted as a candidate scheme for 3GPP LTE-Advanced. Baolong Zhou, Ling-ge Jiang, Shengjie Zhao 0001 |
WCNC | 3 |
| 2012 | Sounding reference signal pattern design for time division duplex multiple-input multiple-output and orthogonal frequency division multiplexing systemsabstractIn the long-term evolution (LTE) time division duplex (TDD) systems, base stations (BSs) use uplink sounding reference signals (SRSs) to estimate downlink channel state information (CSI) for downlink beamforming transmissions. In the current SRS pattern scheme, because SRSs of up to eight users are multiplexed on the same time-frequency resources, there exist mutual interferences among these SRSs over fading channels. Thus, the current scheme impacts the estimation accuracy of CSI and degrades the beamforming performance. Based on the characteristics of TDD systems, the authors propose a SRS pattern scheme, in which each user sends SRSs only on the allocated bandwidth instead of on full bandwidth as in the current LTE TDD systems, to eliminate the interferences among SRSs. Simulations show that the proposed scheme outperforms the current scheme in the estimation accuracy for various frequency-selective and time-selective fading channels. In addition, compared with the current SRS pattern scheme, the proposed SRS pattern scheme is robust to frequency-selective channels, and does not be impacted by the timing offset, and provides the flexibility for the choice of channel estimator at the BS side. Baolong Zhou, Ling-ge Jiang, Chen He 0001, Shengjie Zhao 0001 |
IET Commun. | 4 |
| 2012 | Practical robust uplink pilot time interval optimisation scheme for time-division duplex multiple-input-single-output beamforming systemabstractFocusing on a time-division duplex (TDD) multiple-input-single-output (MISO) beamforming system, this study investigates the robust uplink pilot time interval (UPTI) design to overcome the impact of delay and channel estimation error. In TDD beamforming systems, the base station estimates the downlink (DL) channel state information (CSI) by exploiting the received uplink (UL) pilots and the channel reciprocity. Then, utilising the estimated DL CSI, the beamforming vector in the DL transmission is generated. Owing to the constraints of the TDD frame structure and the UL pilot overhead, there inevitably exist delay and estimation error between the estimated and the actual DLCSI, which would degrade system performance. In this study, the upper bound is first derived on ergodic rate for TDD MISO beamforming systems with channel estimation error and delay. Then, with any one given setting of the normalised UL pilot overhead, the optimal robust UPTI is designed, which maximises the upper bound on ergodic rate in the worst case of the CSI delay (i.e. in the case of maximum delay). Simulation results validate that the designed optimal robust UPTI can not on I y maximise the upper bound on ergodic rate but also be adaptive to the variation of channel conditions very well. Baolong Zhou, Ling-ge Jiang, Chen He 0001, Shengjie Zhao 0001 |
IET Commun. | 4 |
| 2011 | Optimal Uplink Pilot Time Interval Design for TDD MISO Beamforming Systems with Channel Estimation Error and DelayabstractThis paper proposes an optimization scheme on uplink pilots time interval (UPTI) in terms of maximum average post-processing SNR (signal to noise ratio) for a time division duplex (TDD) multiple input single output (MISO) beamforming system with channel estimation error and delay. In TDD system, the base station estimates the channel state information (CSI) at transmitter based on uplink pilots and then uses it to generate the beamforming vector in the downlink transmission. Because of the constraints of the TDD frame structure and the uplink pilot overhead, there inevitably exists delay and channel estimation error between CSI estimation and its use. In this paper, we first derive average post-processing SNR for TDD MISO beamforming system with channel estimation error and delay. We then obtain the optimal UPTI, which maximizes average post-processing SNR, given the normalized pilot overhead (the number of pilot symbols per data symbol). The simulation results validate that the optimal UPTI not only maximizes the average post-processing SNR but also maximizes the system ergodic rate. Especially our research is valuable for the uplink sounding reference signal design in LTE-Advanced system. Baolong Zhou, Ling-ge Jiang, Lei Zhang 0015, Chen He 0001, Shengjie Zhao 0001 |
ICC | 5 |
| 2011 | Capacity of TDD MISO Beamforming Systems with Channel Estimation Error and DelayabstractThis paper examines the capacity of a time division duplex (TDD) multiple input single output (MISO) beamforming system with channel estimation error and delay over the time varying fading channel. In TDD system, the base station estimates the channel state information (CSI) at transmitter based on uplink pilots and then uses it to generate the beamforming vector in the downlink transmission. Because of the constraints of the TDD frame structure and the uplink pilot overhead, there inevitably exists delay and channel estimation error between CSI estimation and its use. In this paper, we first derive the capacity upper bound for TDD MISO beamforming system with channel estimation error and delay. We then obtain the optimal uplink pilot time interval (UPTI), which maximizes the capacity upper bound, given the normalized pilot overhead (the number of pilot symbols per data symbol). The simulation results validate that the optimal UPTI not only maximizes the capacity upper bound but also maximizes the system ergodic capacity. Especially our research is valuable for the uplink sounding reference signal design in LTE-Advanced system. Baolong Zhou, Ling-ge Jiang, Lei Zhang 0015, Chen He 0001, Shengjie Zhao 0001, Zhining Jiang |
VTC Spring | 5 |
| 2011 | On BER of TDD Multiuser MIMO System with Channel Estimation Error and DelayabstractIn Multiuser Multiple-Input Multiple-Output (MU-MIMO) downlink systems, the zero-forcing algorithm is a simple and effective technique for separating users and data streams of each user at the transmitter side, but its performance depends greatly on the accuracy of the available channel state information (CSI) at the transmitter. Due to the existence of the CSI delay and channel estimation error in practical systems, CSI is always imperfect, which degrades system performance significantly. In this paper, by using the correlation between the actual channel and the estimated one as well as the channel's time-correlation, we develop a novel bit error rate (BER) expression for M-QAM (quadrature amplitude modulation) signal in time division duplex (TDD) downlink MU-MIMO system with channel estimation error and CSI delay. We find that channel estimation error causes array gain loss while CSI delay causes diversity gain loss. Moreover, CSI delay causes more performance degradation than channel estimation error at high signal to noise ratio (SNR) for time varying channel. Especially our research is valuable for adaptive modulation and coding scheme and the optimization of MU-MIMO system. Numerical simulations show accurate agreement with the analytical expressions. Baolong Zhou, Ling-ge Jiang, Chen He 0001, Lei Zhang 0015, Shengjie Zhao 0001, Zhang Yi 0006, Lingfeng Lin |
VTC Spring | 5 |
| 2011 | An Optimal Scheme for TDD Beamforming Systems with Imperfect Channel State InformationabstractAn optimal scheme on uplink pilot time interval (UPTI) to maximize average post-processing SNR (signal to noise ratio) is proposed in order to overcome the impact of channel estimation error and delay on a time division duplex (TDD) multiple input single output (MISO) beamforming system. In TDD system, the base station estimates the channel state information (CSI) at transmitter based on uplink pilots and then uses it to generate the beamforming vector in the downlink transmission. Because of the constraints of the TDD frame structure and the uplink pilot overhead, there inevitably exists delay and channel estimation error between CSI estimation and its use. In this paper, we first derive average post-processing SNR for TDD MISO beamforming system with channel estimation error and delay. We then obtain the optimal UPTI, which maximizes average post-processing SNR, given the normalized pilot overhead (the number of pilot symbols per data symbol). The simulation results validate that the optimal UPTI not only maximizes the average post-processing SNR but also minimizes the BER. Especially our research is valuable for the uplink sounding reference signal design in LTE- Advanced system. Baolong Zhou, Ling-ge Jiang, Lei Zhang 0015, Chen He 0001, Shengjie Zhao 0001, Zhining Jiang |
VTC Spring | 5 |
| 2011 | Impact of imperfect channel state information on TDD downlink Multiuser MIMO systemabstractIn downlink Multiuser Multiple-Input Multiple-Output (MU-MIMO) systems, the zero-forcing (ZF) transmission is a simple and effective technique for separating users and data streams of each user at the transmitter side, but its performance depends greatly on the accuracy of the available channel state information (CSI) at the transmitter. Due to the existence of the CSI delay and channel estimation error in practical systems, CSI is always imperfect, which degrades system performance significantly. In this paper, by characterizing CSI inaccuracies due to channel delay and channel estimation error, we develop a novel bit error rate (BER) expression for M-QAM (quadrature amplitude modulation) signal in time division duplex (TDD) downlink MU-MIMO system where each user is equipped with one antenna. By simulation, we find that channel estimation error causes array gain loss while CSI delay causes diversity gain loss. Moreover, CSI delay causes more performance degradation than channel estimation error at high signal to noise ratio (SNR) for time varying channel. Especially our research is valuable for adaptive modulation and coding scheme and the optimization of MU-MIMO system. Numerical simulations show accurate agreement with the analytical expressions. Baolong Zhou, Ling-ge Jiang, Lei Zhang 0015, Chen He 0001, Shengjie Zhao 0001, Lingfeng Lin |
WCNC | 5 |
| 2010 | Optimal Resource Allocation for Video Delivery over MIMO OFDM Wireless SystemsabstractA scalable video delivery transmission framework over MIMO OFDM wireless channels by combining power allocation and antenna selection scheme has been proposed. The framework consisting of independently decodable layers (3D ESCOT), UEP channel coding structure, and antenna selection scheme provide more system error resilience even during deep fading period in wireless channels. A new algorithm is proposed to obtain the optimal power allocation and optimal transmission rate allocation among multiple layers, subject to constraints on the total transmission rate and the total power level. This proposed joint power allocation and antenna selection algorithm in MIMO OFDM system enables us to both overcome the challenge of the full CSI that was assumed in the existing approaches and utilize the estimated CSI from the receiver so as to achieve the optimal solution for a video bitstream in terms of its total expected distortion. Shengjie Zhao 0001, Luoning Gui |
GLOBECOM | 1 |
| 2010 | A distributed layered modulation for relay downlink cooperative transmissionabstractIn downlink relay systems, the conventional straightforward relaying scheme named decode-and-forward (DF) protocol is used at relay station (RS) to decode the data received from base station (BS) and forward the decoded data to the mobile station (MS) to improve the spatial diversity gain. However, conventional DF mode is spectral inefficient and not capacity optimal since the wireless channels in BS-to-RS, BS-to-MS, and RS-to-MS links are asymmetric, i.e., different channel qualities are obtained in different links in the practical scenario. In this paper, a distributed layered-modulation scheme is proposed to improve the DF mode performance in the presence of asymmetric links. With this proposed scheme, BS and/or RS transmit/re-transmit a layered modulated signal comprising a higher-order and a lower-order modulation components according to the different channel conditions in BS-to-RS and RS-to-MS links, respectively, and, therefore, the channel capacity of each link can be utilized efficiently. In addition, a two-layer demodulation is proposed at the destination by exploiting the two copies of transmitted packet from BS and RS, respectively. In order to further enhance the decoding performance, a soft combining method, in which the bit Log Likelihood Ratio (LLR) of the signals that received during the broadcast phase and the relaying phase are combined before turbo decoding, is used at MS receiver. Simulation results show that the proposed distributed layered modulation scheme obtains a significant performance gain in terms of block error rate (BLER) compared to the existing DF relaying scheme. Mingli You, Shengjie Zhao 0001, Lu Zhang 0015 |
PIMRC | 3 |
| 2010 | A sub-optimal interference-aware precoding scheme with other-cell interference for downlink multi-user MIMO channelabstractWe propose a linear precoding technique, called multiuser generalized eigenmode transmission (MGET), based on the leakage-based precoding technique. A signal-to-leakage-plus-noise (SLNR) precoding scheme is one approach for linear precoding in the multiple-input multiple-output broadcast channel that sends multiple data streams to different users in the same cell. Unfortunately, SLNR-based scheme neglects other-cell interference, which limits the performance of users at the edge of the cell. MGET addresses the shortcomings of previous SLNR-based beamformers by transmitting to each user on one or more eigenmodes chosen using a greedy algorithm as well as presenting an interference-aware enhancement to SLNR precoding scheme that uses a whitening filter for interference suppression at the receiver and a novel precoder using the interference-plus-noise covariance matrix for each user at the transmitter. We consider the typical sum-power constraint (SPC). Numerical results show that the proposed MGET technique outperforms previous linear techniques. Shengjie Zhao 0001, Mingli You, Luoning Gui |
PIMRC | 1 |
| 2010 | Scalable video transmission over MIMO OFDM wireless systems with antenna selectionabstractA framework of joint antenna selection and source coding strategy for scalable video transmission over wireless systems was proposed. Each layer of a bit-stream could be switched among multiple transmit antennas in order to achieve unequal error protection (UEP) for prioritized video layered bit-streams. At first wavelet-based 3D ESCOT codec generates layered bit streams that need prioritized delivery. Then we obtain the ordering of transmit antennas with channel strengths as partial CSI. Finally by the proposed joint antenna selection and layered source coding algorithm, we are able to transmit higher priority layers of a bit stream into antennas with higher channel strengths. Therefore we can achieve UEP for layered scalable video coding transmission over MIMO OFDMsystem. Simulation results show that the proposed framework improved the system performance significantly. The proposed joint antenna selection and layered source coding algorithm enables us to achieve UEP for layered video transmission over MIMO OFDM system with low complexity of algorithm at the transmit side as well as achieve different QoS at receive side. Shengjie Zhao 0001, Mingli You, Luoning Gui |
PIMRC | 1 |
| 2010 | An Interference-Aware Precoding Scheme with Other-Cell Interference for Downlink Multi-User MIMO ChannelabstractWe propose a linear precoding technique, called multiuser generalized eigenmode transmission (MGET), based on the leakage-based precoding technique. A signal-to-leakage-plus-noise (SLNR) precoding scheme is one approach for linear precoding in the multiple-input multiple-output broadcast channel that sends multiple data streams to different users in the same cell. Unfortunately, SLNR-based scheme neglects other-cell interference, which limits the performance of users at the edge of the cell. MGET addresses the shortcomings of previous SLNR-based beamformers by presenting an interference-aware enhancement to SLNR precoding scheme that uses a whitening filter for interference suppression at the receiver and a novel precoder using the interference-plus-noise covariance matrix for each user at the transmitter. Numerical results show that the proposed MGET technique outperforms the traditional block diagonalization (BD) linear technique in reasonable SINR regime. Shengjie Zhao 0001, Lu Zhang 0015, Luoning Gui |
VTC Fall | 1 |
| 2009 | Scalable video delivery over MIMO OFDM wireless systems using joint power allocation and antenna selectionabstractA scalable video delivery transmission framework over MIMO OFDM wireless channels by combining power allocation and antenna selection scheme has been proposed. The framework consisting of independently decodable layers (3D ESCOT), UEP channel coding structure, and antenna selection scheme provide more system error resilience even during deep fading period in wireless channels. A new algorithm is proposed to obtain the optimal power allocation and optimal transmission rate allocation among multiple layers, subject to constraints on the total transmission rate and the total power level. This proposed joint power allocation and antenna selection algorithm in MIMO OFDM system enables us to both overcome the challenge of the full CSI that was assumed in the existing approaches and utilize the estimated CSI from the receiver so as to achieve the optimal solution for a video bitstream in terms of its total expected distortion. Generally, as the wireless channel conditions change, the framework we proposed can scale the video streams and transport the scaled video streams to receivers with a smooth change of perceptual quality. Shengjie Zhao 0001, Mingli You, Luoning Gui |
MMSP | 1 |
| 2006 | Progressive Video Delivery over Wideband Wireless Channels Using Space-Time Differentially Coded OFDM SystemsabstractProgressive video delivery over wireless networks is very challenging due to the time-varying nature of wireless channels and limited power in the mobile devices. This paper proposes an end-to-end architecture for multilayer progressive video delivery over space-time differentially coded orthogonal frequency division multiplexing (STDC-OFDM) systems. An input video sequence is compressed by 3D-ESCOT into a layered bitstream. We input multiple layers of the bitstream in series to a STDC-OFDM channel. Different video source layers are protected by different error protection schemes in order to achieve unequal error protection. In progressive transmission, the reconstruction quality is important not only at the target transmission rate but also at the intermediate rates. So, the error protection strategy needs to optimize the average performance over the set of intermediate rates. We propose to use progressive joint source-channel coding to generate operational transmission distortion-rate (TD-R) functions and operational transmission distortion-power (TD-P) functions for multiple layers before forming the operational transmission distortion-power-rate (TD-PR) surfaces. Lagrange multipliers are then employed on the fly to obtain the optimal power allocation and optimal rate allocation among multiple layers, subject to constraints on the total transmission rate and the total power level. Progressive joint source-channel coding offers the scalability feature to handle bandwidth variations and changes in channel conditions. By extending the rate-distortion function in source coding to the TD-PR surface in joint source-channel coding, our work can use the "equal slope" argument to effectively solve the transmission rate allocation problem as well as the transmission power allocation problem for multilayer video transmission. Experiments show that our scheme achieves significant improvement over a nonoptimal system with the same total power level and total transmission rate. Shengjie Zhao 0001, Zixiang Xiong, Xiaodong Wang 0001, Jianping Hua |
IEEE Trans. Mob. Comput. | 1 |
| 2005 | Optimal Resource Allocation for Wireless Video over CDMA NetworksabstractWe present a multiple-channel video transmission scheme in wireless CDMA networks over multipath fading channels. We map an embedded video bitstream, which is encoded into multiple independently decodable layers by 3D-ESCOT video coding technique, to multiple CDMA channels. One video source layer is transmitted over one CDMA channel. Each video source layer is protected by a product channel code structure. A product channel code is obtained by the combination of a row code based on rate compatible punctured convolutional code (RCPC) with cyclic redundancy check (CRC) error detection and a source-channel column code, i.e., systematic rate-compatible Reed-Solomon (RS) style erasure code. For a given budget on the available bandwidth and total transmit power, the transmitter determines the optimal power allocations and the optimal transmission rates among multiple CDMA channels, as well as the optimal product channel code rate allocation, i.e., the optimal unequal Reed-Solomon code source/parity rate allocations and the optimal RCPC rate protection for each channel. In formulating such an optimization problem, we make use of results on the large-system CDMA performance for various multiuser receivers in multipath fading channels. The channel is modeled as the concatenation of wireless BER channel and a wireline packet erasure channel with a fixed packet loss probability. By solving the optimization problem, we obtain the optimal power level allocation and the optimal transmission rate allocation over multiple CDMA channels. For each CDMA channel, we also employ a fast joint source-channel coding algorithm to obtain the optimal product channel code structure. Simulation results show that the proposed framework allows the video quality to degrade gracefully as the fading worsens or the bandwidth decreases, and it offers improved video quality at the receiver. Shengjie Zhao 0001, Zixiang Xiong, Xiaodong Wang 0001 |
IEEE Trans. Mob. Comput. | 1 |
| 2003 | Optimal resource allocation for wireless video over CDMA networksabstractWe present a multiple-channel video transmission scheme in wireless CDMA networks over multipath fading channels. We map an embedded video bitstream, which is encoded into multiple independently decodable layers by 3D-ESCOT video coding techniques, to multiple CDMA channels. Each video source layer is protected by a product channel code structure. For a given budget on the available bandwidth and total transmit power, the transmitter determines the optimal power allocations and the optimal transmission rates among multiple CDMA channels, as well as the optimal product channel code rate allocation. We make use of results on the large-system CDMA performance for various multiuser receivers in multipath fading channels. Simulation results show that the proposed framework allows the video quality to degrade gracefully as the fading worsens or the bandwidth decreases, and it offers improved video quality at the receiver. Shengjie Zhao 0001, Zixiang Xiong, Xiaodong Wang 0001 |
ICME | 1 |
| 2002 | Joint error control and power allocation for video transmission over CDMA networks with multiuser detectionabstractError control and power allocation for transmitting wireless video over CDMA networks are considered in conjunction with multiuser detection. We map a layered video bitstream to several CDMA fading channels and inject multiple source/parity layers into each of these channels at the transmitter. At the receiver, we employ a linear minimum mean-square error (MMSE) multiuser detector in the uplink and two types of blind linear MMSE detectors in the downlink, for demodulating the received data. For given constraints on the available bandwidth and transmit power, the transmitter determines the optimal power allocation among different CDMA fading channels and the optimal number of source and parity packets to send that offer the best video quality. Simulation results show a performance gain of up to 1.5 dB with joint optimization over that with rate optimization only. Shengjie Zhao 0001, Zixiang Xiong, Xiaodong Wang 0001 |
ICC | 1 |
| 2002 | Joint error control and power allocation for video transmission over CDMA networks with multiuser detectionabstractError control and power allocation for transmitting wireless video over CDMA networks are considered in conjunction with multiuser detection. We map a layered video bitstream to several CDMA fading channels and inject multiple source/parity layers into each of these channels at the transmitter. At the receiver, we employ a linear minimum mean-square error (MMSE) multiuser detector in the uplink and two types of blind linear MMSE detectors, i.e., the direct-matrix-inversion blind detector and the subspace blind detector, in the downlink, for demodulating the received data. For given constraints on the available bandwidth and transmit power, the transmitter determines the optimal power allocation among different CDMA fading channels and the optimal number of source and parity packets to send that offer the best video quality. We formulate a combined optimization problem and give the optimal joint rate and power allocation for each of these three receivers. Simulation results show a performance gain of up to 3.5 dB with joint optimization over with rate optimization only. Shengjie Zhao 0001, Zixiang Xiong, Xiaodong Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |