EDBT 2026 Demo / reviewers in the wild / expert
Xiao-Ping Zhang 0002
dblp:157/0177 · also Xiao-Ping (Steven) Zhang, Xiaoping Zhang 0002
· DBLP profile ↗
286ranked-venue papers
8as first author
160since 2021 · last 2026
0000-0001-5241-0069ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 161 · 7 first-author · 70 since 2021Computer networks · 48 · 45 since 2021Artificial intelligence and machine learning · 38 · 28 since 2021Applied, interdisciplinary, general and emerging computing · 31 · 18 since 2021Systems, architecture and hardware · 17 · 1 first-author · 10 since 2021Databases, data management, data science and information retrieval · 6 · 5 since 2021Security and privacy · 2Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PSformer: Parameter-Efficient Transformer with Segment Shared Attention for Time Series Forecasting
Hongkang Zhang, Shao-Lun Huang, Danny Dongning Sun, Xiao-Ping Zhang 0002 |
ICPR (1) | 7 |
| 2026 | OMGRec: One-time Matching-based Generative Rerank with Permutation-level Modeling in E-commerce
Zhibo Xiao, Chuxin Chen, Chengyu Lai, Qijie Shen, Jiuning Lin, Dimin Wang, Xiao-Ping Zhang 0002 |
WWW | 9 |
| 2026 | HF-IDS: A Heuristic Factors-Based Semi-Supervised Model for Intrusion Detection Systems in IoT NetworksabstractSemi-supervised learning in intrusion detection systems (IDS) faces three major challenges: the scarcity of labeled samples, class imbalance, and distribution divergence between labeled and unlabeled data. Moreover, the third challenge may lead to an extreme situation where labeled data fail to cover all categories. To address these issues, we develop a heuristic factors based semi-supervised model named HF-IDS. Specifically, we employ symbolic regression, a novel feature engineering technique, to generate class-indicating factors. These factors guide a multi-level clustering for pseudo-label addition to some unlabeled data. For classification, we construct an Ensemble of Graph Neural Networks (E-GNNs). Different from common ensemble learning methods, each of the GNN classifiers is equipped with a unique graph filter constructed by the factors. Through these filters, we topologically reweight different classes to enhance the model’s effectiveness in handling with class imbalance problem. The model is evaluated in both multi-class and binary classification settings, with the binary case simulating the extreme missing-category scenario. Experiments on NSL-KDD and CICIDS-2017 datasets show that HF-IDS consistently outperforms state-of-the-art baselines in accuracy, precision, recall, and F1_score. Xiao-Ping Zhang 0002, Konstantinos N. Plataniotis |
IEEE Internet Things J. | 2 |
| 2026 | QUIDS: Quality-Informed Incentive-Driven Multiagent Dispatching System for Mobile CrowdsensingabstractThis paper addresses the challenges of achieving optimal quality of information (QoI) in non-dedicated vehicular mobile crowdsensing (NVMCS) system, where vehicles not originally designed for sensing are leveraged to collect real-time data as they traverse urban environments. These challenges are exacerbated by the interrelated issues of sensing coverage, sensing reliability, and the inherently dynamic nature of participating vehicles. To tackle these challenges, we propose QUIDS, a QUality-informed Incentive-driven multi-agent Dispatching System, which ensures high sensing coverage and sensing reliability under budget constraints in NVMCS systems. QUIDS improves QoI by introducing a novel metric, Aggregated Sensing Quality (ASQ), designed to quantitatively capture the concept of QoI by integrating both sensing coverage and sensing reliability. Moreover, we develop a Mutually Assisted Belief-aware Vehicle Dispatching algorithm that estimates sensing reliability and allocates monetary incentives under uncertain vehicle conditions, thereby further improving ASQ. Evaluation using real-world data collected from a deployed NVMCS system in a metropolitan area demonstrates the effectiveness of QUIDS. The ASQ metric shows a 38% improvement over non-dispatching scenarios and a 10% enhancement over state-of-the-art methods. Additionally, QUIDS reduces reconstruction map errors by 39–74% across various reconstruction algorithms, validating its efficacy in improving QoI within NVMCS systems. Addressing the often-overlooked issue of sensing reliability in existing studies, the QUIDS system leverages non-dedicated vehicles and incorporates a quality-informed incentive-driven dispatching system to jointly optimize sensing coverage and sensing reliability. This enables low-cost, high-quality, and scalable urban environmental monitoring without the need for dedicated sensing infrastructure, and makes the system applicable to diverse smart-city scenarios such as traffic monitoring and environmental sensing. Zuxin Li, Fanhang Man, Xuecheng Chen, Susu Xu, Fan Dang 0001, Chaopeng Hong, Yunhao Liu 0001, Xiao-Ping Zhang 0002, Xinlei Chen |
IEEE Internet Things J. | 9 |
| 2026 | An Adaptive Non-Linear Graph Filter in Semi-Supervised Graph Based ClassificationabstractOvercoming class imbalance is a critical challenge for graph-based semi-supervised classification methods. In this letter, we address this issue from the perspective of graph filtering and propose a novel adaptive graph filter. By introducing learnable thresholds into the adjacency matrix, the filter enables dynamic suppression of majority-class bias during label propagation through the incorporation of discontinuities. Additionally, we develop a modified Fruit Fly Optimization Algorithm (m-FOA) to optimize the filter's coefficients, which achieves lower loss and faster convergence compared to other heuristic algorithms. To evaluate the effectiveness of our approach, we conduct a Monte Carlo simulation on a real-world dataset. The results demonstrate that our method outperforms the baseline methods in both classification accuracy and efficiency when handling class imbalance. We note that the model's scalability to very large graphs is limited and the solving procedure can be time-consuming due to the dense construction of the filter. Xinchun Yu, Xiao-Ping Zhang 0002, Konstantinos N. Plataniotis |
IEEE Signal Process. Lett. | 3 |
| 2026 | UVCG: Leveraging Temporal Consistency for Video Protection Against Stable-Diffusion-Based Editing
Kaizhou Li, Xinchun Yu, Xiao-Ping Zhang 0002 |
IEEE Signal Process. Lett. | 5 |
| 2026 | WDMamba: When Wavelet Degradation Prior Meets Vision Mamba for Image DehazingabstractIn this paper, we reveal a novel haze-specific wavelet degradation prior observed through wavelet transform analysis, which shows that haze-related information predominantly resides in low-frequency components. Exploiting this insight, we propose a novel dehazing framework, WDMamba, which decomposes the image dehazing task into two sequential stages: low-frequency restoration followed by detail enhancement. This coarse-to-fine strategy enables WDMamba to effectively capture features specific to each stage of the dehazing process, resulting in high-quality restored images. Specifically, in the low-frequency restoration stage, we integrate Mamba blocks to reconstruct global structures with linear complexity, efficiently removing overall haze and producing a coarse restored image. Thereafter, the detail enhancement stage reinstates fine-grained information that may have been overlooked during the previous phase, culminating in the final dehazed output. Furthermore, to enhance detail retention and achieve more natural dehazing, we introduce a self-guided contrastive regularization during network training. By utilizing the coarse restored output as a hard negative example, our model learns more discriminative representations, substantially boosting the overall dehazing performance. Extensive evaluations on public dehazing benchmarks demonstrate that our method surpasses state-of-the-art approaches both qualitatively and quantitatively. Code is available at https://github.com/SunJ000/WDMamba. Heng Liu 0002, Yongzhen Wang 0001, Xiao-Ping Zhang 0002, Mingqiang Wei |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | Embodied Spatial Affordance: Spatial-Aware Affordance Learning for Embodied Navigation and ManipulationabstractEmbodied navigation and manipulation are fundamental capabilities for embodied agents operating in physical environments. A key challenge in this process is understanding the spatial context and the affordances of the environment, which involves recognizing how objects can be interacted with (object affordance) and identifying suitable locations for movement and object placement (free space affordance). While Vision-Language Models (VLMs) have shown promise in high-level task planning, their ability to translate reasoning into precise executable actions remains limited, particularly in image-based spatial understanding and precise affordance localization-a critical gap in image processing for robotics. To bridge this gap, we propose EspA, a novel image-to-keypoint model that leverages spatial-aware affordance learning to predict actionable affordances directly from 2D image inputs. Built on a hierarchical vision-language architecture, EspA jointly reasons about object affordances and free space affordances, enabling pixel-level localization of both types of interactions. Crucially, EspA translates language instructions into precise 2D affordance keypoints from observed images, which are then projected into 3D actionable coordinates using depth information. To support this unified affordance reasoning, we introduce the Embodied Spatial Affordance (ESA) dataset, which captures both object-centric interactions and free space contexts. By jointly modeling these affordances in a shared representation space, EspA overcomes the limitations of prior works that treat them independently. The dataset's fine-grained annotations enable our model to learn the intricate relationship between object functionality and spatial feasibility, significantly enhancing the spatial understanding in embodied tasks. Extensive experimental results demonstrate that EspA outperforms existing state-of-the-art Vision-Language Models (VLMs), both open-source and closed-source, in object and free space affordance prediction. Furthermore, it exhibits superior performance in real-world embodied navigation and manipulation experiments. Our work advances the field of image-based spatial reasoning by providing a scalable solution for translating high-level instructions into low-level actionable affordances. We believe this work paves the way for more robust and versatile embodied agents capable of effectively interacting with complex environments. The dataset, benchmark, and evaluation code will be publicly available to facilitate future research. Project website: https://embodied-spatial-affordance.github.io/. Xiaoshuai Hao, Yingbo Tang, Long Chen 0015, Wei Zhou 0021, Jungong Han, Wenbo Ding 0001, Xiao-Ping Zhang 0002 |
IEEE Trans. Image Process. | 8 |
| 2026 | STeP-Diff: Spatio-Temporal Physics-Informed Diffusion Models for Mobile Fine-Grained Pollution ForecastingabstractFine-grained air pollution forecasting is crucial for urban management and the development of healthy buildings. Deploying portable sensors on mobile platforms such as cars and buses offers a low-cost, easy-to-maintain, and wide-coverage data collection solution. However, due to the random and uncontrollable movement patterns of these non-dedicated mobile platforms, the resulting sensor data are often incomplete and temporally inconsistent. By exploring potential training patterns in the reverse process of diffusion models, we proposeSpatio-TemporalPhysics-InformedDiffusion Models (STeP-Diff). STeP-Diff leverages DeepONet to model the spatial sequence of measurements along with a PDE-informed diffusion model to forecast the spatio-temporal field from incomplete and time-varying data. Through a PDE-constrained regularization framework, the denoising process asymptotically converges to the convection-diffusion dynamics, ensuring that predictions are both grounded in real-world measurements and aligned with the fundamental physics governing pollution dispersion. To assess the performance of the system, we deployed 59 self-designed portable sensing devices in two cities, operating for 14 days to collect air pollution data. Compared to the second-best performing algorithm, our model achieved improvements of up to 89.12% in MAE, 82.30% in RMSE, and 25.00% in MAPE, with extensive evaluations demonstrating that STeP-Diff effectively captures the spatio-temporal dependencies in air pollution fields. Weijie Hong, Huandong Wang, Qiuhua Wang, Yali Song, Xiao-Ping Zhang 0002, Yong Li 0008, Xinlei Chen |
IEEE Trans. Knowl. Data Eng. | 7 |
| 2026 | Aerial Shepherds: Enabling Hierarchical Localization in Heterogeneous MAV SwarmsabstractA heterogeneous micro aerial vehicles (MAV) swarm consists of resource-intensive but expensive advanced MAVs (AMAVs) and resource-limited but cost-effective basic MAVs (BMAVs), offering opportunities in diverse fields. Accurate and real-time localization is crucial for MAV swarms, but current practices lack a low-cost, high-precision, and real-time solution, especially for lightweight BMAVs. We find an opportunity to accomplish the task by transforming AMAVs into mobile localization infrastructures for BMAVs. However, translating this insight into a practical system is challenging due to issues in estimating locations with diverse and unknown localization errors of BMAVs, and allocating resources of AMAVs considering interconnected influential factors. This work introduces TransformLoc, a new framework that transforms AMAVs into mobile localization infrastructures, specifically designed for low-cost and resource-constrained BMAVs. We design an error-aware joint location estimation model to perform intermittent joint estimation for BMAVs and introduce a similarity-instructed adaptive grouping-scheduling strategy to allocate resources of AMAVs dynamically. TransformLoc achieves a collaborative, adaptive, and cost-effective localization system suitable for large-scale heterogeneous MAV swarms. We implement and validate TransformLoc on industrial drones. Results show it outperforms all baselines by up to 68% in localization performance, improving navigation success rates by 60%. Extensive robustness and ablation experiments further highlight superiority of its design. Haoyang Wang 0012, Jingao Xu, Chenyu Zhao 0002, Yuhan Cheng, Xuecheng Chen, Chaopeng Hong, Xiao-Ping Zhang 0002, Yunhao Liu 0001, Xinlei Chen |
IEEE Trans. Mob. Comput. | 7 |
| 2026 | A Novel Integrated Sensing and Communication Scheme in UAVs-Enabled Vehicular Networks With MARL-Driven Adaptive ControlabstractIn this paper, we propose a novel integrated sensing and communication (ISAC) scheme tailored for UAVs-enabled vehicular networks, which leverages the information coverage capabilities of multiple UAVs and addresses critical challenges posed by multiple moving users. Unlike many traditional scheme, our scheme efficiently leverages ISAC signal echoes and real-time data uploads to provide communication services while achieving accurate sensing, thereby overcoming issues of resource waste and low operational efficiency. In the scheme, we aim to optimize both communication and sensing indicators, taking into account practical issues such as energy saving and collision avoidance for UAVs. However, the inherent complexity of multi-objective stochastic optimization in dynamic environments and limited communication resources render centralized UAV control inconvenient. To address the above challenges, we propose a novel multi-agent reinforcement learning (MARL) algorithm based on local information to realize the distributed adaptive control of motion decision, power selection, and channel allocation for UAVs. The algorithm combines random network distillation (RND) and dynamic data augmentation with multi-agent deep deterministic policy gradient (MADDPG) to encourage agents to explore effectively under sparse rewards and improve MADDPG's policy learning ability in finite data, thus approaching the global optimal solution. Experimental results demonstrate that the proposed algorithm can improve communication and sensing performance by more than 16.71% and 68.26% compared with other baselines and satisfy the set constraints. Furthermore, by adjusting hyperparameters, we can optimize the ISAC performance while achieving different energy savings levels for UAVs, proving that the designed scheme can reduce the waste of resources and improve the ISAC operation efficiency. Ziyuan Wang 0002, Xiao-Ping Zhang 0002, Wenbo Ding 0001, Yuhan Dong, Xinlei Chen |
IEEE Trans. Mob. Comput. | 2 |
| 2026 | An Ensemble MARL Approach for Heterogeneous UAV Swarm Target Search in 3D Space
Changxu Wei, Ziyuan Wang 0002, Yixian Zhang, Wenbo Ding 0001, Xiao-Ping Zhang 0002 |
IEEE Trans. Mob. Comput. | 6 |
| 2026 | Real-Scene Image Dehazing via Laplacian Pyramid-Based Conditional Diffusion ModelabstractRecent diffusion models have demonstrated exceptional efficacy across various image restoration tasks, but still suffer from time-consuming and substantial computational resource consumption. To address these challenges, we present LPCDiff, a novel Laplacian Pyramid-based Conditional Diffusion model designed for real-scene image dehazing. LPCDiff leverages the Laplacian pyramid decomposition to decouple the input image into two components: the low-resolution low-pass image and the high-frequency residuals. These components are subsequently reconstructed through a diffusion model and a well-designed high-frequency residual recovery module. With such a strategy, LPCDiff can substantially accelerate inference speed and reduce computational costs without sacrificing image fidelity. In addition, the framework empowers the model to capture intrinsic high-frequency details and low-frequency structural information within the image, resulting in sharper and more realistic haze-free outputs. Moreover, to extract more valuable information from the limited training data, we introduce a low-frequency refinement module to further enhance the intricate details of the final dehazed images. Through extensive experimentation, our method significantly outperforms 12 state-of-the-art approaches on three real-world and one synthetic image dehazing benchmarks. Code is available athttps://github.com/yz-wang/LPCDiff. Yongzhen Wang 0001, Heng Liu 0002, Xiao-Ping Zhang 0002, Mingqiang Wei |
IEEE Trans. Multim. | 4 |
| 2026 | SniffySquad: Patchiness-Aware Gas Source Localization with Multi-Robot CollaborationabstractGas source localization is pivotal for the rapid mitigation of gas leakage disasters, where mobile robots emerge as a promising solution. However, existing methods predominantly schedule robots’ movements based on reactive stimuli or simplified gas plume models. These approaches typically excel in idealized, simulated environments but fall short in real-world gas environments characterized by their patchy distribution. In this work, we introduce SniffySquad , a multi-robot olfaction-based system designed to address the inherent patchiness in gas source localization. SniffySquad incorporates a patchiness-aware active sensing approach that enhances the quality of data collection and estimation. Moreover, it features an innovative collaborative role adaptation strategy to boost the efficiency of source-seeking endeavors. Extensive evaluations demonstrate that our system achieves an increase in the success rate by \(20\%+\) and an improvement in path efficiency by \(30\%+\) , outperforming state-of-the-art gas source localization solutions. Yuhan Cheng, Xuecheng Chen, Haoyang Wang 0012, Jingao Xu, Chaopeng Hong, Susu Xu, Xiao-Ping Zhang 0002, Yunhao Liu 0001, Xinlei Chen |
ACM Trans. Sens. Networks | 8 |
| 2026 | Enhancing Communication Security in TDMA-Based Multi-User VLC With Obstructed Links: A Relay-Aided Cooperative Transmission and Jamming ApproachabstractIn this paper, we propose an innovative relay-aided physical layer security scheme leveraging cooperative transmission and jamming that operates effectively in time-division-multiple-access-based multi-user visible light communication (MuVLC) systems with obstructed communication links. First, to circumvent obstacles between the light source (LS) and legal receivers, we utilize a transmission relay (TR) to forward information signals from the LS to the receiver plane. Subsequently, to ensure communication security, we select a jamming relay to enlarge the channel capacity gap between eavesdroppers and legal receivers through cooperative jamming. Using stochastic geometry, we derive closed-form expressions for key performance metrics, including the cumulative distribution function of secrecy capacity, security outage probability, and their lower and upper bounds, both with and without relay cooperation. To further enhance system security, we introduce a disk-shaped security-protected zone around the TR. All analytical expressions are numerically verified through Monte Carlo simulations, serving as benchmarks to evaluate security performance. Numerical results confirm that our proposed scheme can ensure continuous signal reception and high-level security. Furthermore, while relay cooperation modestly enhances system security, the introduction of a security-protected zone around the TR yields substantial security improvement, which offers practical insights for designing secure time-division-multiple-access-based MuVLC where communication links are obstructed. Yuhan Dong, Xinchun Yu, Yongkang Ding, Jian Song 0004, Xiao-Ping Zhang 0002 |
IEEE Trans. Wirel. Commun. | 6 |
| 2025 | Multi-View 3D Human Pose Estimation with Weakly Synchronized ImagesabstractMulti-view 3D human pose estimation (MHPE) is an important research task in computer vision. To maintain consistency during the data collection, hardware synchronization devices are commonly used to connect cameras, ensuring that images from different views are captured simultaneously. However, synchronizing with extra devices has two apparent limitations: the hardware is i) usually expensive and ii) less flexible for deployment in outdoor open scenarios. Suppose the model can improve its tolerance for the time differences in multi-view image capture. In that case, the difficulty and cost of deployment will be greatly reduced, and MHPE will become more widespread. In this paper, we try to answer how to build a model that performs pose estimation directly using ''weakly synchronized images" from multiple views, where the captured images shift from each other within a frame. To this end, we introduce a new multi-view 3D human pose estimation task given weakly synchronized image inputs. Apart from existing well-synchronized datasets, we present the first weakly synchronized dataset comprising 800k images. Thereon, we propose SyncDiffPose, a novel model based on the diffusion method for pose estimation to denoise the error in such data. By combining simple synchronization strategies, e.g., the timer method, our approach can perform pose estimation without hardware calibration. Ruiwen Gu, Junliang Xing, Xinchun Yu, Xiao-Ping Zhang 0002 |
AAAI | 6 |
| 2025 | Structure-prior Informed Diffusion Model for Graph Source Localization with Limited DataabstractSource localization in graph information propagation is essential for mitigating network disruptions, including misinformation spread, cyber threats, and infrastructure failures. Existing deep generative approaches face significant challenges in real-world applications due to limited propagation data availability. We present SIDSL (Structure-prior Informed Diffusion model for Source Localization), a generative diffusion framework that leverages topology-aware priors to enable robust source localization with limited data. SIDSL addresses three key challenges: unknown propagation patterns through structure-based source estimations via graph label propagation, complex topology-propagation relationships via a propagation-enhanced conditional denoiser with GNN-parameterized label propagation module, and class imbalance through structure-prior biased diffusion initialization. By learning pattern-invariant features from synthetic data generated by established propagation models, SIDSL enables effective knowledge transfer to real-world scenarios. Experimental evaluation on four real-world datasets demonstrates superior performance with 7.5-13.3% F1 score improvements over baselines, including over 19% improvement in few-shot and 40% in zero-shot settings, validating the framework's effectiveness for practical source localization. Our code can be found here (https://github.com/tsinghua-fib-lab/SIDSL). Jingtao Ding, Xiaojun Liang, Yong Li 0025, Xiao-Ping Zhang 0002 |
CIKM | 5 |
| 2025 | ATP-LLaVA: Adaptive Token Pruning for Large Vision Language ModelsabstractLarge Vision Language Models (LVLMs) have achieved significant success across multi-modal tasks. However, the computational cost of processing long visual tokens can be prohibitively expensive on resource-limited devices. Previous methods have identified redundancy in visual tokens within the Large Language Model (LLM) decoder layers and have mitigated this by pruning tokens using a predefined or fixed ratio, thereby reducing computational over-head. Nonetheless, we observe that the impact of pruning ratio varies across different LLM layers and instances (image-prompt pairs). Therefore, it is essential to develop a layer-wise and instance-wise vision token pruning strategy to balance computational cost and model performance effectively. We propose ATP-LLaVA, a novel approach that adaptively determines instance-specific token pruning ratios for each LLM layer. Specifically, we introduce an Adaptive Token Pruning (ATP) module, which computes the importance score and pruning threshold based on input instance adaptively. The ATP module can be seamlessly integrated between any two LLM layers with negligible computational overhead. Additionally, we develop a Spatial Augmented Pruning (SAP) strategy that prunes visual tokens with both token redundancy and spatial modeling perspectives. Our approach reduces the average token count by 75% while maintaining performance, with only a minimal 1.9% degradation across seven widely used benchmarks. Xubing Ye, Yukang Gan, Yixiao Ge, Xiao-Ping Zhang 0002, Yansong Tang |
CVPR | 4 |
| 2025 | Transfer Risk Map: Mitigating Pixel-level Negative Transfer in Medical SegmentationabstractHow to mitigate negative transfer in transfer learning is a long-standing and challenging issue, especially in the application of medical image segmentation. Existing methods for reducing negative transfer focuses on classification or regression tasks, ignoring the non-uniform negative transfer risk in different image regions. In this work, we propose a simple yet effective weighted fine-tuning method that directs the model’s attention towards regions with significant transfer risk for medical semantic segmentation. Specifically, we compute a transferability-guided transfer risk map to quantify the transfer hardness for each pixel and the potential risks of negative transfer. During the fine-tuning phase, we introduce a map-weighted loss function, normalized with image foreground size to counter class imbalance. Extensive experiments on brain segmentation datasets show our method significantly improves the target task performance, with gains of 4.37% on FeTS2021 and 1.81% on iSeg-2019, avoiding negative transfer across modalities and tasks. Meanwhile, a 2.9% gain under a few-shot scenario validates the robustness of our approach. Shutong Duan, Yang Tan 0004, Yang Li 0104, Xiao-Ping Zhang 0002 |
ICASSP | 6 |
| 2025 | Reinforced Domain Selection for Continuous Domain AdaptationabstractContinuous Domain Adaptation (CDA) effectively bridges significant domain shifts by progressively adapting from the source domain through intermediate domains to the target domain. However, selecting intermediate domains without explicit metadata remains a substantial challenge that has not been extensively explored in existing studies. To tackle this issue, we propose a novel framework that combines reinforcement learning with feature disentanglement to conduct domain path selection in an unsupervised CDA setting. Our approach introduces an innovative unsupervised reward mechanism that leverages the distances between latent domain embeddings to facilitate the identification of optimal transfer paths. Furthermore, by disentangling features, our method facilitates the calculation of unsupervised rewards using domain-specific features and promotes domain adaptation by aligning domain-invariant features. This integrated strategy is designed to simultaneously optimize transfer paths and target task performance, enhancing the effectiveness of domain adaptation processes. Extensive empirical evaluations on datasets such as Rotated MNIST and ADNI demonstrate substantial improvements in prediction accuracy and domain selection efficiency, establishing our method’s superiority over traditional CDA approaches. Huaze Tang, Yanru Wu, Yang Li 0104, Xiao-Ping Zhang 0002 |
ICASSP | 5 |
| 2025 | GLST-GCN: Global-Local Spatio-Temporal Graph Convolutional Network for Skeleton-based Hand Motion PredictionabstractThis paper introduces a new task: skeleton-based hand motion sequence prediction, which can be applied to VR/AR systems and human-computer interaction systems to enhance user experience. To tackle this task, we performed a comprehensive analysis of hand movement patterns and propose a Global-Local Spatio-Temporal Graph Convolutional Network (GLST-GCN). Based on the local and global correlation characteristics of hand motions, the GLS block of GLST-GCN is proposed to extract spatial features through local branches and global modules. Additionally, we considered the variation in hand movement speeds, which leads to inconsistencies in temporal features across different time scales, and thus proposed the MGLT block to model both global and local temporal dependencies across multiple scales. We also conduct extensive experiments based on the remaked Bighand2.2M and FPHA datasets, and numerical results show that our proposed method offers state-of-the-art performance. Wenrui Yang, Xinchun Yu, Xiao-Ping Zhang 0002 |
ICASSP | 3 |
| 2025 | Mean-Field Aided QMIX: A Scalable and Flexible Q-Learning Approach for Large-Scale Agent GroupsabstractValue decomposition methods are effective for multi-agent reinforcement learning (MARL), with QMIX being one of the most advanced. However, it struggles with scalability and flexibility in large-scale agent systems. The introduction of mean-field theory into MARL provides a potential solution to both scalability and flexibility challenges. In this paper, we propose Mean-Field Aided QMIX (MF-QMIX), a novel algorithm that addresses these challenges by incorporating mean-field theory into the QMIX framework. MF-QMIX significantly reduces the computational complexity from O(n) in QMIX to O(1), making it highly scalable and suitable for large-scale environments with many agents. Additionally, MF-QMIX is flexible, allowing the trained model to be applied across various group sizes without retraining, thus accommodating dynamic changes in the number of agents. Our extensive experiments demonstrate that MF-QMIX outperforms existing methods, both in computational efficiency and adaptability, across cooperative and competitive scenarios. This establishes MF-QMIX as a scalable and flexible solution for large-scale MARL problems. Huaze Tang, Xiao-Ping Zhang 0002, Wenbo Ding 0001 |
ICASSP | 3 |
| 2025 | Discern-XR: An Online Classifier for Metaverse Network TrafficabstractIn this paper, we design an exclusive Metaverse network traffic classifier, named Discern-XR, to help Internet service providers (ISP) and router manufacturers enhance the quality of Metaverse services. Leveraging segmented learning, the Frame Vector Representation (FVR) algorithm and Frame Identification Algorithm (FIA) are proposed to extract critical frame-related statistics from raw network data having only four application-level features. A novel Augmentation, Aggregation, and Retention Online Training (A2R-OT) algorithm is proposed to find an accurate classification model through online training methodology. In addition, we contribute to the real-world Metaverse dataset comprising virtual reality (VR) games, VR video, VR chat, augmented reality (AR), and mixed reality (MR) traffic, providing a comprehensive benchmark. Discern-XR outperforms state-of-the-art classifiers by 7 % while improving training efficiency and reducing false-negative rates. Our work advances Metaverse network traffic classification by standing as the state-of-the-art solution. Yoga Suhas Kuruba Manjunath, Austin Wissborn, Mathew Szymanowski, Mushu Li, Lian Zhao, Xiao-Ping Zhang 0002 |
ICC | 6 |
| 2025 | Exo-ViHa: A Cross-Platform Exoskeleton System with Visual and Haptic Feedback for Efficient Dexterous Skill LearningabstractImitation learning has emerged as a powerful paradigm for robot skills learning. However, traditional data collection systems for dexterous manipulation face challenges, including a lack of balance between acquisition efficiency, consistency, and accuracy. To address these issues, we introduce Exo-ViHa, an innovative 3D-printed exoskeleton system that enables users to collect data from a first-person perspective while providing real-time haptic feedback. This system combines a 3D-printed modular structure with a slam camera, a motion capture glove, and a wrist-mounted camera. Various dexterous hands can be installed at the end, enabling it to simultaneously collect the posture of the end effector, hand movements, and visual data. By leveraging the first-person perspective and direct interaction, the exoskeleton enhances the task realism and haptic feedback, improving the consistency between demonstrations and actual robot deployments. In addition, it has cross-platform compatibility with various robotic arms and dexterous hands. Experiments show that the system can significantly improve the success rate and efficiency of data collection for dexterous manipulation tasks. Webpage: https://exo-viha2025.github.io/. Xintao Chao, Shilong Mu, Yushan Liu 0006, Shoujie Li, Chuqiao Lyu, Xiao-Ping Zhang 0002, Wenbo Ding 0001 |
IROS | 6 |
| 2025 | Distributional Decision Transformer: Risk-Sensitive Offline RL via Quantile-Based Critics and Stochastic ReturnabstractOffline reinforcement learning faces a critical challenge in synthesizing high-reward trajectories from suboptimal datasets while robustly handling the stochasticity inherent in real-world decision-making. While combination of return-conditioned sequence models, such as Decision Transformers (DT), and dynamics programming critics shows great potential in trajectory synthesis, their deterministic action generation and scale Q value critic often fails to distinguish intentional behavioral variability from detrimental noise, leading to suboptimal policy collapse. To address this challenge, we propose the Distributional Decision Transformer (DDT), a novel framework that unifies probabilistic return distribution modeling with autoregressive action generation. DDT introduces two key innovations: (1) a Gaussian stochastic return mechanism that reparameterizes target returns as samplable distributions, enabling diverse action candidate generation; and (2) an Implicit Quantile Network (IQN) critic embedded within the deciding loop, which evaluates actions across the full spectrum of return distributions (quantiles τ ~ U(0, 1)). In D4RL benchmarks, DDT achieves state-of-the-art performance, achieving a 91.6 average normalized score in MuJoCo locomotion and 69.3 in sparse-reward settings. The results establish DDT as a principled solution for synthesis of risk-aware trajectory in offline RL. Changxu Wei, Huaze Tang, Yixian Zhang, Chao Wang 0139, Xiao-Ping Zhang 0002, Wenbo Ding 0001 |
IROS | 5 |
| 2025 | Open3D-VQA: A Benchmark for Embodied Spatial Concept Reasoning with Multimodal Large Language Model in Open SpaceabstractSpatial reasoning is a fundamental capability of multimodal large language models (MLLMs), yet their performance in open aerial environments remains underexplored. In this work, we present Open3D-VQA, a novel benchmark for evaluating MLLMs' ability to reason about complex spatial relationships from an aerial perspective. The benchmark comprises 73k QA pairs across seven general spatial reasoning tasks, offered in multiple-choice, true/false, and short-answer formats, and supports both visual and point cloud modalities. The questions are automatically generated from spatial relations extracted from both real-world and simulated aerial scenes. Evaluation on 13 popular MLLMs reveals that: 1) Models are generally better at answering questions about relative spatial relations than absolute distances, 2) 3D LLMs fail to demonstrate significant advantages over 2D LLMs, and 3) Fine-tuning solely on the simulated dataset can significantly improve the model's spatial reasoning performance in real-world scenarios. The benchmark, generation pipeline, and evaluation toolkit are released on this page. Zile Zhou, Xuchen Liu 0001, Jianjie Fang, Chen Gao 0001, Jinqiang Cui, Yong Li 0008, Xinlei Chen, Xiao-Ping Zhang 0002 |
ACM Multimedia | 10 |
| 2025 | PID-controlled Langevin Dynamics for Faster Sampling of Generative ModelsabstractLangevin dynamics sampling suffers from extremely low generation speed, fundamentally limited by numerous fine-grained iterations to converge to the target distribution. We introduce PID-controlled Langevin Dynamics (PIDLD), a novel sampling acceleration algorithm that reinterprets the sampling process using control-theoretic principles. By treating energy gradients as feedback signals, PIDLD combines historical gradients (the integral term) and gradient trends (the derivative term) to efficiently traverse energy landscapes and adaptively stabilize, thereby significantly reducing the number of iterations required to produce high-quality samples. Our approach requires no additional training, datasets, or prior information, making it immediately integrable with any Langevin-based method. Extensive experiments across image generation and reasoning tasks demonstrate that PIDLD achieves higher quality with fewer steps, making Langevin-based generative models more practical for efficiency-critical applications. The implementation can be found at \href{https://github.com/tsinghua-fib-lab/PIDLD}{https://github.com/tsinghua-fib-lab/PIDLD}. Jianhai Shu, Jingtao Ding, Xiao-Ping Zhang 0002 |
NeurIPS | 5 |
| 2025 | SynTSBench: Rethinking Temporal Pattern Learning in Deep Learning Models for Time SeriesabstractRecent advances in deep learning have driven rapid progress in time series forecasting, yet many state-of-the-art models continue to struggle with robust performance in real-world applications, even when they achieve strong results on standard benchmark datasets. This persistent gap can be attributed to the black-box nature of deep learning architectures and the inherent limitations of current evaluation frameworks, which frequently lack the capacity to provide clear, quantitative insights into the specific strengths and weaknesses of different models, thereby complicating the selection of appropriate models for particular forecasting scenarios.To address these issues, we propose a synthetic data-driven evaluation paradigm, SynTSBench, that systematically assesses fundamental modeling capabilities of time series forecasting models through programmable feature configuration. Our framework isolates confounding factors and establishes an interpretable evaluation system with three core analytical dimensions: (1) temporal feature decomposition and capability mapping, which enables systematic evaluation of model capacities to learn specific pattern types; (2) robustness analysis under data irregularities, which quantifies noise tolerance thresholds and anomaly recovery capabilities; and (3) theoretical optimum benchmarking, which establishes performance boundaries for each pattern type—enabling direct comparison between model predictions and mathematical optima.Our experiments show that current deep learning models do not universally approach optimal baselines across all types of temporal features. Qitai Tan, Ruiwen Gu, Yilin Su, Xiao-Ping Zhang 0002 |
NeurIPS | 6 |
| 2025 | HiLTV: Hierarchical Multi-Distribution Modeling for Lifetime Value Prediction in Online GamesabstractCustomer Lifetime Value (LTV) prediction is a critical task in online games, as it helps with the formulation of refined game operation strategies, resource allocation, and personalized recommendation. Accurate LTV values enable to identify and target high-value users, enhance user retention and further long-term revenue growth. However, LTV prediction in online games faces unique challenges. Most in-app purchases (IAP) games have various fixed recharge levels, and users with different payment preferences show distinct LTV distributions. Moreover, existing methods fail to capture the multi-modal distribution of LTV values, and suffer from bias when predicting LTV values of games that users have not registered, i.e., new users. To address these challenges, we propose HiLTV, a novel hierarchical framework for LTV prediction in online games. We devise hierarchical modules to align with real-world user recharge behaviors, and a Zero-Inflated Mixture-of-Logistic (ZIMoL) loss instead of a unimodal distribution loss is adopted to better model various segments of users. We also introduce a calibration module that enables more robust predictions for new users. Both offline evaluation on real-world industrial datasets over state-of-the-art baselines and online A/B test from a leading game platform demonstrate the superior performance of our method. Aisi Zheng, Huangbin Zhang, Zhengwei Deng, Xiao-Ping Zhang 0002 |
SIGIR | 7 |
| 2025 | Segmented Learning for Metaverse Network Traffic ClassificationabstractWe propose a novel two-staged segmented learning framework to enhance network traffic classification (NTC) for 5G and beyond (B5G)-driven enhanced mobile broadband (eMBB) applications, including Metaverse traffic. The first stage improves classification speed and accuracy for eMBB traffic, and the second stage extends its capability to classify the more complex and dynamic Metaverse network traffic. We introduce Essential Vector Representation (EVR) and Frame Vector Representation (FVR) feature engineering methods. These methods reduce inference time and preserve privacy by leveraging application-level features such as transmission time, packet length, direction, and inter-arrival time. The outputs from EVR and FVR are classified using our Augmentation, Aggregation, and Retention-Online Training (A2R-OT) algorithm, which enhances adaptive online learning, improving accuracy and efficiency. Additionally, we construct a comprehensive real-world Metaverse network traffic dataset to address the lack of publicly available Metaverse traffic data. To our knowledge, this is the first framework to integrate eMBB and Metaverse traffic classification. Our approach achieves a 6% improvement over state-of-the-art solutions, advancing network traffic management (NTM) for B5G networks. Yoga Suhas Kuruba Manjunath, Lian Zhao, Xiao-Ping Zhang 0002 |
IEEE Internet Things J. | 3 |
| 2025 | CoughSlowFast: Cough Recognition With Audio and Video Signal Fusion
Mingke Feng, Guangtao Zhai, Xiao-Ping Zhang 0002, Menghan Hu |
IEEE Signal Process. Lett. | 3 |
| 2025 | Physics-Informed Diffusion Model for Complex-Valued Radar Sea Clutter GenerationabstractWe propose a physics-informed diffusion framework for radar sea clutter generation with two key innovations: a novel component consistency mechanism that preserves intrinsic statistical relationships between real and imaginary parts of complex signals, and a comprehensive physics-guided diffusion approach that systematically integrates domain-specific priors. Our framework represents complex-valued signals through a two-channel diffusion process with a KL divergence-based constraint ensuring proper distributional alignment between real and imaginary components. To enhance physical fidelity, we develop a multi-prior guidance scheme that enforces essential domain knowledge throughout the diffusion process: Doppler spectral characteristics, power spectral density, local signal smoothness, and reference pattern matching. These physics-based priors continuously guide the generation process to maintain physical correctness while ensuring numerical stability. This systematic incorporation of physical properties and constraints into the diffusion model for sea clutter generation presents a novel contribution to physics-informed signal synthesis, enabling more realistic sea clutter generation for advanced radar applications. Zhenyu Liu 0003, Xiao-Ping Zhang 0002 |
IEEE Signal Process. Lett. | 3 |
| 2025 | Frame-Level Temporal Difference Learning for Partial Deepfake Speech DetectionabstractDetecting partial deepfake speech is essential due to its potential for subtle misinformation. However, existing methods depend on costly frame-level annotations during training, limiting real-world scalability. Also, they focus on detecting transition artifacts between bonafide and deepfake segments. As deepfake generation techniques increasingly smooth these transitions, detection has become more challenging. To address this, our work introduces a new perspective by analyzing frame-level temporal differences and reveals that deepfake speech exhibits erratic directional changes and unnatural local transitions compared to bonafide speech. Based on this finding, we propose a Temporal Difference Attention Module (TDAM) that redefines partial deepfake detection as identifying unnatural temporal variations, without relying on explicit boundary annotations. A dual-level hierarchical difference representation captures temporal irregularities at both fine and coarse scales, while adaptive average pooling preserves essential patterns across variable-length inputs to minimize information loss. Our TDAM-AvgPool model achieves state-of-the-art performance, with an EER of 0.59% on the PartialSpoof dataset and 0.03% on the HAD dataset, which significantly outperforms the existing methods without requiring frame-level supervision. Menglu Li, Xiao-Ping Zhang 0002, Lian Zhao |
IEEE Signal Process. Lett. | 2 |
| 2025 | Canonical Shape Reconstruction With SE(3) Equivariance Learning for Weakly-Supervised Object Pose Estimationabstract6D object pose estimation from a single RGB-D image is a fundamental problem in computer vision and robot manipulation. Despite recent advancements, existing methods still suffer several limitations. First of all, the object shape representation extracted from the depth map is often less expressive because the object point cloud parsed from the depth map is highly incomplete due to the object self-occlusion and noisy due to the sensor artifacts. This shape representation issue further intensifies when lacking sufficient labeled data for model training, which unfortunately is another typical problem for object pose estimation considering the heavy annotation cost for real-world pose labeling. In this study, we propose to tackle the above issues in a unified way. First, we enhance the object shape representation from the partial point cloud with a novel canonical shape reconstruction module, in which an implicit canonical frame is established by incorporating the SE(3) equivariance, achieving implicit feature alignment of the partial point cloud inputs, leading to robust shape recovery. Second, based on the enhanced object representation, we further utilize the de-canonicalized and pose-dependent completed object shape as the training signal, and develop a novel weakly-supervised learning framework to leverage both labeled synthetic data and unlabeled real data to train the pose estimation model in a label-efficient way. Extensive experiments on three widely used benchmarks demonstrate the effectiveness, and superiority of our framework over state-of-the-art methods. Jun Zhou 0029, Kai Chen 0024, Mingqiang Wei, Xiao-Ping Zhang 0002, Qi Dou 0001, Harry Qin |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | A Joint Optimization Method for High-Resolution Wide-Swath SAR Imaging: Combining Signal Transmitting and Imaging Perspectives
Yu-Wei Zhuo, Jianghong Han, Xinchang Hu, Xueqian Wang 0002, Gang Li 0008, Xiao-Ping Zhang 0002, You He 0003 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | M2Restore: Mixture-of-Experts-Based Mamba-CNN Fusion Framework for All-in-One Image RestorationabstractNatural images are often degraded by complex, composite degradations such as rain, snow, and haze, which adversely impact downstream vision applications. While existing image restoration efforts have achieved notable success, they are still hindered by two critical challenges: limited generalization across dynamically varying degradation scenarios and a suboptimal balance between preserving local details and modeling global dependencies. To overcome these challenges, we propose M2Restore, a novel Mixture-of-Experts (MoE)-based Mamba-CNN fusion framework for efficient and robust all-in-one image restoration. M2Restore introduces three key contributions: First, to boost the model's generalization across diverse degradation conditions, we exploit a CLIP-guided MoE gating mechanism that fuses task-conditioned prompts with CLIP-derived semantic priors. This mechanism is further refined via cross-modal feature calibration, which enables precise expert selection for various degradation types. Second, to jointly capture global contextual dependencies and fine-grained local details, we design a dual-stream architecture that integrates the localized representational strength of CNNs with the long-range modeling efficiency of Mamba. This integration enables collaborative optimization of global semantic relationships and local structural fidelity, preserving global coherence while enhancing detail restoration. Third, we introduce an edge-aware dynamic gating mechanism that adaptively balances global modeling and local enhancement by reallocating computational attention to degradation-sensitive regions. This targeted focus leads to more efficient and precise restoration. Extensive experiments across multiple image restoration benchmarks validate the superiority of M2Restore in both visual quality and quantitative performance. Code is available at https://github.com/yz-wang/M2Restore. Yongzhen Wang 0001, Zhuoran Zheng, Xiao-Ping Zhang 0002, Mingqiang Wei |
IEEE Trans. Image Process. | 4 |
| 2025 | RSHazeDiff: A Unified Fourier-Aware Diffusion Model for Remote Sensing Image DehazingabstractHaze severely degrades the visual quality of remote sensing images and hampers the performance of road extraction, vehicle detection, and traffic flow monitoring. The emerging denoising diffusion probabilistic model (DDPM) exhibits the significant potential for dense haze removal with its strong generation ability. Since remote sensing images contain extensive small-scale texture structures, it is important to effectively restore image details from hazy images. However, current wisdom of DDPM fails to preserve image details and color fidelity well, limiting its dehazing capacity for remote sensing images. In this paper, we propose a novel unified Fourier-aware diffusion model for remote sensing image dehazing, termed RSHazeDiff. From a new perspective, RSHazeDiff explores the conditional DDPM to improve image quality in dense hazy scenarios, and it makes three key contributions. First, RSHazeDiff refines the training phase of diffusion process by performing noise estimation and reconstruction constraints in a coarse-to-fine fashion. Thus, it remedies the unpleasing results caused by the simple noise estimation constraint in DDPM. Second, by taking the frequency information as important prior knowledge during iterative sampling steps, RSHazeDiff can preserve more texture details and color fidelity in dehazed images. Third, we design a global compensated learning module to utilize the Fourier transform to capture the global dependency features of input images, which can effectively mitigate the effects of boundary artifacts when processing fixed-size patches. Experiments on both synthetic and real-world benchmarks validate the favorable performance of RSHazeDiff over state-of-the-art methods. Source code will be released athttps://github.com/jm-xiong/RSHazeDiff Jiamei Xiong, Xuefeng Yan 0001, Yongzhen Wang 0001, Wei Zhao 0039, Xiao-Ping Zhang 0002, Mingqiang Wei |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2025 | SmartSpr: A Physics-Informed Mobile Sprinkler Scheduling System for Reducing Urban Particulate Matter PollutionabstractUrban particulate pollution presents considerable public health hazards, underscoring the need for effective control measures in various cities. This paper proposes SmartSpr, a physics-informed urban mobile sprinkler scheduling system designed for enhanced efficiency in reducing particulate pollution. SmartSpr incorporates a Physics-Informed Neural Network (PINN)-based model, enriched with Bayesian optimization, to accurately simulate the impact of mobile sprinklers on particulate matter (PM) dispersion. Building on this sprinkling effect model, a selective sprinkling strategy considering the replenish process is proposed. This strategy employs a sparsity-driven decoupling simulated annealing algorithm to refine sprinkler routes, prioritizing areas with substantial environmental benefits. Extensive field experiments and simulations have validated SmartSpr, demonstrating a 64.8% reduction in prediction error of SmartSpr's sprinkling model compared to the leading baseline and an 18% enhancement in pollutant reduction efficiency of the proposed scheduling algorithm. Zijian Xiao, Zuxin Li, Xuecheng Chen, Chaopeng Hong, Xiao-Ping Zhang 0002, Xinlei Chen |
IEEE Trans. Mob. Comput. | 6 |
| 2025 | Symmetry-Informed MARL: A Decentralized and Cooperative UAV Swarm Control Approach for Communication CoverageabstractUncrewed aerial vehicle-mounted base stations (UAV-MBSs) provide flexible wireless connectivity, extending communication coverage in underserved areas. Recently, multi-agent reinforcement learning (MARL) has shown great potential for cooperative UAV swarm control to support efficient communication coverage in dynamic and complex environments. However, existing MARL-based methods often suffer from low sample efficiency due to its trial-and-error training characteristics, limiting its ability to control large UAV swarms with continuous state-action space and partial observation. We notice that UAV swarm systems in communication coverage tasks exhibit a spatial symmetry property, e.g., a rotation in the spatial observation of a UAV results in a same rotation in its optimal action. Exploiting this property, we formulate the task as a symmetric decentralized partially observable Markov decision process and introduce symmetry-informed MARL, featuring a novel network called the symmetry-informed graph neural network (SiGNN) to serve as the policy/value networks. SiGNN leverages the inherent symmetry in multi-UAV systems by embedding the symmetry into the network structure, thereby enhancing the training efficiency to handle large swarms with continuous control. Theoretical analysis shows that the SiGNN strictly preserves symmetry properties, which guarantees the effectiveness of the approach. Experiments in simulation were conducted to handle communication coverage using up to 20 UAVs with continuous control. Experimental results demonstrate that SiGNN-based MARL outperforms advanced baselines, verifying its superior sample efficiency, scalability and robustness. Rongye Shi, Xin Yu 0009, Yandong Wang 0002, Yongkai Tian, Zhenyu Liu 0003, Wenjun Wu 0001, Xiao-Ping Zhang 0002, Manuela M. Veloso |
IEEE Trans. Mob. Comput. | 7 |
| 2025 | CatUA: Catalyzing Urban Air Quality Intelligence Through Mobile Crowd-SensingabstractMobile air pollution sensing methods have emerged to collect air quality data with improved spatial and temporal resolutions. However, existing methodologies struggle to effectively process spatially mixed gas samples due to the highly dynamic fluctuations experienced by sensors, resulting in significant measurement deviations. We identify an opportunity to address this issue by exploring potential patterns within sensor measurements. To this end, we propose CatUA, a novel city-scale fine-grained air quality estimation system designed to deliver accurate mobile air quality data. First, we design AirBERT, a representation learning model specifically aimed at discerning mixed gas concentrations from sensor data. Second, we implement a Prompt-informed Training Strategy that leverages extensive unlabeled and minimal labeled city-scale data to enhance the performance of CatUA. Notably, the Auto-Prompt mechanism allows CatUA to conveniently acquire new knowledge tailored to specific downstream tasks. To ensure the practicality of CatUA, we have invested considerable effort in developing the software stack on our meticulously crafted Sensing Front-end, which has successfully gathered city-scale air quality data for over 1,200 hours. Experiments conducted on the collected data demonstrate that CatUA reduces sensing errors by 96.9% with a latency of only 44.9ms, outperforming the state-of-the-art baseline by 42.6%. Yuxuan Liu 0010, Haoyang Wang 0012, Fanhang Man, Jingao Xu, Fan Dang 0001, Chaopeng Hong, Yunhao Liu 0001, Xiao-Ping Zhang 0002, Yali Song, Qiuhua Wang, Xinlei Chen |
IEEE Trans. Mob. Comput. | 9 |
| 2025 | Human-Centered Financial Signal Analysis Based on Visual Patterns in Stock ChartsabstractThe study adopted a human-centered perspective to research the financial markets, focusing on identifying variations in eye movement patterns between professional and non-professional traders as they analyze a series of stock charts. Eye movement data was selected as the analysis target based on the hypothesis that it represents a behavioral phenotype indicative of stock analysts' cognitive processes during market analysis. Disparities were identified by conducting variance analysis and the Wilcoxon signed-rank test on statistical metrics derived from eye fixations and saccades. Psychological and behavioral economic interpretations were provided to understand the underlying reasons for these observed patterns. To showcase the practical application potential of the human-centered perspective, eye movement data and human visual characteristics were used to construct visual saliency prediction models of professional stock analysts. Leveraging this human-centered model, we developed two practical application demonstrations specifically designed to support and instruct novice traders. Based on the above demonstrations, a training program was designed that demonstrates how, with ongoing training, the non-professional traders' ability to observe stock charts improves progressively. Ji-Feng Luo, Kaixun Zhang, Xudong An, Menghan Hu, Guangtao Zhai, Xiao-Ping Zhang 0002 |
IEEE Trans. Multim. | 7 |
| 2025 | DEHand: Deformable Encoding for Photo-Realistic Free-View and Free-Pose Hand RenderingabstractInput encoding has proven crucial in the success of methods based on neural radiance field. Compared to the literature on general static scene modeling, input encoding for dynamic hand modeling has been less explored. However, this aspect is critical to the modeling of deformation and rendering, as it maps a sampled point in space to the representation containing all the information associated with dynamic hand for inferring the geometry and appearance property of this point. The design of input encoding determines how well the neural network can learn for photo-realistic hand rendering. We offer an in-depth examination of this key component and introduceDEHand, a new representation utilizingDeformableEncoding for photo-realistic free-view and free-poseHandrendering. DEHand leverages deformable encoding with a latent code map to achieve high-quality, pose-controlled rendering. Deformable encoding is achieved by adapting static input encoding techniques for the view synthesis of dynamic hands, using parametric hand mesh model as a proxy to construct encodings that map sampled points into a space capable of integrating over different poses and providing rich information for hand modeling. Our findings demonstrate that with our deformable encoding, a single Multilayer Perceptron (MLP) can achieve high-quality dynamic hand rendering, learning solely from images. Extensive experiments on InterHand2.6M validate the superior rendering quality of our method and the effectiveness of each component in our design. Yunzhi Teng, Xiaoke Huang 0001, Kejie Li, Xiao-Ping Zhang 0002, Yansong Tang |
IEEE Trans. Multim. | 4 |
| 2025 | End-to-End Deep Video Compression Based on Hierarchical Temporal Context LearningabstractEmerging learning-based video compression suffers from error propagation in long group of pictures (GOP), yielding limited coding performance. To address this problem, a novel end-to-end Deep Video Compression method based on Hierarchical Temporal Context Learning (DVCH) is proposed in this paper. DVCH aims to fully exploit temporal contexts and suppress error propagation for better coding performance. It first divides video frames into several hierarchies with different compression qualities. The frames in lower hierarchies have high compression quality, and serve as reference frames. To mine high-quality reference information, we propose a Hierarchical Temporal Context Learning (HTCL) network as the fundamental module of our DVCH. The informative temporal context features from hierarchical prediction structure can be extracted by the network. Motion vectors (MVs) between the to-be-coded frame and its reference frames are estimated by the MV Learning module and used to align the extracted contexts. The contexts are fed into Context Coding module to generate the prediction of the decoded frame. Moreover, a multi-stage training strategy is developed to solve the imbalanced training challenge. Experimental results demonstrate that the proposed DVCH exceeds x264 and other end-to-end video compression methods, regardless of objective, subjective, error propagation suppression, GOP sizes, and sequence length evaluations. As much as 49.27% bitrate savings and 2.52 dB PSNR gains can be achieved in large GOP. Kejun Wu, You Yang 0002, Qiong Liu 0001, Xiao-Ping Zhang 0002 |
IEEE Trans. Multim. | 5 |
| 2025 | Optimizing Video-Based Respiration Monitoring: Motion Artifact Reduction and Adaptive ROI SelectionabstractIn non-contact respiratory monitoring, reducing motion artifact and selecting the appropriate Region of Interest (ROI) pose significant challenges. Most motion artifact removal methods rely on signal periodicity assumptions, while respiratory signals usually are non-periodic in real-world scenarios. Existing automated ROI selection approaches are mostly primarily impacted by the texture of clothing, absence of chest landmarks, and obstruction of face. To improve the quality of respiratory signals, in this study, we propose a framework for automatic respiratory ROI selection based on video, namely, Optimizing Video-based Respiration Monitoring (OVRM), which consists of peak-trough adaptive motion artifact removal and characteristic-driven adaptive ROI selection. This motion artifact removal strategy removes motion artifacts by using a dynamic ratio-based judgment mechanism, and reconstructs signals using sinusoidal interpolation. The adaptive ROI method scores signals based on periodicity, similarity, smoothness, and energy, selecting the highest-scoring blocks as the ROIs to match respiratory signals efficiently. Experimental results, validated across four datasets, demonstrate that OVRM effectively reduces signal noise caused by subject movement and outperforms state-of-the-art non-contact respiratory monitoring algorithms. The dataset and code are publicly available at:https://github.com/zxx5058/OVRM. Xudong Tan, Mei Zhou, Menghan Hu, Zhanzhan Cheng, Nengfeng Qian, Changyin Wu, Guangtao Zhai, Xiao-Ping Zhang 0002 |
IEEE Trans. Multim. | 10 |
| 2025 | Transferability-Guided Cross-Domain Cross-Task Transfer LearningabstractWe propose two novel transferability metrics fast optimal transport-based conditional entropy (F-OTCE) and joint correspondence OTCE (JC-OTCE) to evaluate how much the source model (task) can benefit the learning of the target task and to learn more generalizable representations for cross-domain cross-task transfer learning. Unlike the original OTCE metric that requires evaluating the empirical transferability on auxiliary tasks, our metrics are auxiliary-free such that they can be computed much more efficiently. Specifically, F-OTCE estimates transferability by first solving an optimal transport (OT) problem between source and target distributions and then uses the optimal coupling to compute the negative conditional entropy (NCE) between the source and target labels. It can also serve as an objective function to enhance downstream transfer learning tasks including model finetuning and domain generalization (DG). Meanwhile, JC-OTCE improves the transferability accuracy of F-OTCE by including label distances in the OT problem, though it incurs additional computation costs. Extensive experiments demonstrate that F-OTCE and JC-OTCE outperform state-of-the-art auxiliary-free metrics by 21.1% and 25.8%, respectively, in correlation coefficient with the ground-truth transfer accuracy. By eliminating the training cost of auxiliary tasks, the two metrics reduce the total computation time of the previous method from 43 min to 9.32 and 10.78 s, respectively, for a pair of tasks. When applied in the model finetuning and DG tasks, F-OTCE shows significant improvements in the transfer accuracy in few-shot classification experiments, with up to 4.41% and 2.34% accuracy gains, respectively. Yang Tan 0004, Enming Zhang, Yang Li 0104, Shao-Lun Huang, Xiao-Ping Zhang 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | Social Physics Informed Diffusion Model for Crowd SimulationabstractCrowd simulation holds crucial applications in various domains, such as urban planning, architectural design, and traffic arrangement. In recent years, physics-informed machine learning methods have achieved state-of-the-art performance in crowd simulation but fail to model the heterogeneity and multi-modality of human movement comprehensively. In this paper, we propose a social physics-informed diffusion model named SPDiff to mitigate the above gap. SPDiff takes both the interactive and historical information of crowds in the current timeframe to reverse the diffusion process, thereby generating the distribution of pedestrian movement in the subsequent timeframe. Inspired by the well-known social physics model, i.e., Social Force, regarding crowd dynamics, we design a crowd interaction encoder to guide the denoising process and further enhance this module with the equivariant properties of crowd interactions. To mitigate error accumulation in long-term simulations, we propose a multi-frame rollout training algorithm for diffusion modeling. Experiments conducted on two real-world datasets demonstrate the superior performance of SPDiff in terms of both macroscopic and microscopic evaluation metrics. Code and appendix are available at https://github.com/tsinghua-fib-lab/SPDiff. Jingtao Ding, Yong Li 0008, Yue Wang 0007, Xiao-Ping Zhang 0002 |
AAAI | 5 |
| 2024 | THz Optical Image Recognition Method via AM-Res2Net ModelabstractTerahertz (THz) waves possess great properties such as ultra-large bandwidth, good penetration, strong reflectivity to metallic materials, and low photon energy, making them widely applicable in fields such as optical communication, optical networks, and optical imaging. Therefore, THz optical imaging, as a novel non-destructive testing technology, demonstrates significant application value in areas such as body security scanning. However, existing THz imaging systems require that targets be close to the detector in order to avoid issues like visible stripes, increased noise, reduced contrast, and blurred edges, which can significantly reduce the quality of imaging results. To solve this problem, this study proposes an AM-Res2Net model to achieve accurate THz optical image recognition. The proposed model, compared with the Res2Net model, uses multiple small kernel convolutions to replace a single large kernel convolution. Furthermore, the proposed model also replaces the activation function and introduces the attention mechanism in the residual structure. The AM-Res2Net model tested on the THz optical image set constructed in this study achieved a recognition accuracy of 94.3%, which is higher than the ResNet and Res2Net models. Jiazhen Song, Sixing Xi, Xun Guan, Xiao-Ping Zhang 0002, Zhenyu Liu 0003 |
GLOBECOM | 4 |
| 2024 | Separate or Integrated: Comprehensive Analysis of Combining Optical OTFS and OAM in RIS-assisted FSO Communication SystemsabstractPaving the way toward next-generation communications, reconfigurable intelligent surface (RIS) based free-space optical (FSO) links have emerged as a strong candidate in the unlicensed spectrum for enhanced capacity and signal quality. However, existing modulation schemes constrain the FSO links’ performance and are unsuitable for RIS-induced multipath fading. To address the issue, we firstly introduce the optical orthogonal time frequency space (O-OTFS) modulation scheme into such links, with orbital angular momentum (OAM) optical beams combined to enhance the performance. Moreover, we propose two combination systems with separate and integrated OAM-OTFS respectively, and further propose a novel quad-mode (QM)-OAM-OTFS scheme within the integrated system, demonstrating significant performance enhancement over counterparts. Numerical results indicate that the separate OAM-OTFS system enhances performance in terms of power and spectral efficiencies but reduces system robustness to turbulence with increasing OAM modes, the proposed QM-OAM-OTFS scheme effectively mitigates atmospheric turbulence effects and achieves efficient power and spectral utilization. Shuang Tang, Weijie Dai, Jianhua Pei, Xiao-Ping Zhang 0002, Jian Song 0004, Yuhan Dong |
GLOBECOM | 4 |
| 2024 | Joint Beamforming for Backscatter Integrated Sensing and CommunicationabstractIntegrated sensing and communication (ISAC) is a key technology of next generation wireless communication. Backscatter communication (BackCom) plays an important role for internet of things (IoT). Then the integration of ISAC with BackCom technology enables low-power data transmission while enhancing the system sensing ability, which is expected to provide a potentially revolutionary solution for IoT applications. In this paper, we propose a novel backscatter-ISAC (B-ISAC) system and focus on the joint beamforming design for the system. We formulate the communication and sensing model of the B-ISAC system and derive the metrics of communication and sensing performance respectively, i.e., communication rate and detection probability. We propose a joint beamforming scheme aiming to optimize the communication rate under sensing constraint and power budget. A successive convex approximation (SCA) based algorithm and an iterative algorithm are developed for solving the complicated non-convex optimization problem. Numerical results validate the effectiveness of the proposed scheme and associated algorithms. The proposed B-ISAC system has broad application prospect in IoT scenarios. Zongyao Zhao, Tiankuo Wei, Zhenyu Liu 0003, Xinke Tang, Xiao-Ping Zhang 0002, Yuhan Dong |
GLOBECOM | 5 |
| 2024 | Unified Probability Distributions of Generalized Composite Fading with Inverse-Type Distributions of Large-Scale Shadowing/FluctuationsabstractBased on novel inverse-type PDF formulae for large-scale shadowing/fluctuations, we derive novel unified probability density functions (PDFs) and moment generating function (MGF) formulae that characterize wide ranging of generalized composite fading distributions in radio frequency and free-space optical wireless communications channels. By specifying four fading parameters according to specific composite fading models, the unified PDF and MGF formulae enable us to unify and characterize wide ranging of existing and numerous novel generalized composite fading distributions. These unified PDF and MGF formulae are applicable for deriving generalized performance metrics that can in turn be further specified and evaluated according to various known and numerous generalized composite fading channel models. Chin Choy Chai, Xiao-Ping Zhang 0002 |
ICASSP | 2 |
| 2024 | Leveraging Noisy Labels of Nearest Neighbors for Label Correction and Sample SelectionabstractDealing with noisy labels (LNL) emerges as a critical challenge when applying deep learning (DL) in practical settings. Previous methodologies primarily concentrated on harnessing model predictions to mitigate the impact of noisy labels. Nevertheless, their efficacy is strongly contingent on the accuracy of model predictions, a factor that cannot be assured in the context of LNL. Our empirical analysis shows that in noisy datasets, the spatial information of latent feature representation combined with original noisy labels is more robust than the methods using model predictions. To mitigate the unreliability introduced by model predictions, we propose a novel Feature Representation method, which utilizes noisy labels of nearest neighbors for label Correction and sample Selection (FRCS). Extensive experiments on various benchmark datasets demonstrate the superiority of FRCS compared with SOTA methods. Our codes are available at https://github.com/tianfangjh/FRCS-Noisy-Labels. Yixiong Chen, Li Liu 0036, Xiaoguang Han 0001, Xiao-Ping Zhang 0002 |
ICASSP | 5 |
| 2024 | Language-Free Compositional Action Generation via Decoupling RefinementabstractComposing simple actions into complex actions is crucial yet challenging. Existing methods largely rely on language annotations to discern composable latent semantics, which is costly and labor-intensive. In this study, we introduce a novel framework to generate compositional actions without language auxiliaries. Our approach consists of three components: Action Coupling, Conditional Action Generation, and Decoupling Refinement. Action Coupling integrates two sub-actions to generate pseudo-training examples. Then, a conditional generative model, CVAE is employed to facilitate the diverse generation. Decoupling Refinement leverages a self-supervised pre-trained model MAE to ensure semantic consistency between sub-actions and compositional actions. Due to the lack of existing datasets containing both sub-actions and compositional actions, we create two new datasets, named HumanAct-C and UESTC-C. Both qualitative and quantitative assessments are conducted to show our efficacy.1 Guangyi Chen 0002, Yansong Tang, Guangrun Wang, Xiao-Ping Zhang 0002, Ser-Nam Lim |
ICASSP | 5 |
| 2024 | Optimizing Trading Strategies in Quantitative Markets Using Multi-Agent Reinforcement LearningabstractQuantitative markets are characterized by swift dynamics and abundant uncertainties, making the pursuit of profit-driven stock trading actions inherently challenging. Within this context, Reinforcement Learning (RL) — which operates on a reward-centric mechanism for optimal control — has surfaced as a potentially effective solution to the intricate financial decision-making conundrums presented. This paper delves into the fusion of two established financial trading strategies, namely the constant proportion portfolio insurance (CPPI) and the time-invariant portfolio protection (TIPP), with the multi-agent deep deterministic policy gradient (MADDPG) framework. As a result, we introduce two novel multi-agent RL (MARL) methods: CPPI-MADDPG and TIPP-MADDPG, tailored for probing strategic trading within quantitative markets. To validate these innovations, we implemented them on a diverse selection of 100 real-market shares. Our empirical findings reveal that the CPPI-MADDPG and TIPP-MADDPG strategies consistently outpace their traditional counterparts, affirming their efficacy in the realm of quantitative trading. Hengxi Zhang, Zhendong Shi, Yuanquan Hu, Wenbo Ding 0001, Ercan E. Kuruoglu, Xiao-Ping Zhang 0002 |
ICASSP | 6 |
| 2024 | Thqa: A Perceptual Quality Assessment Database for Talking HeadsabstractIn the realm of media technology, digital humans have gained prominence due to rapid advancements in computer technology. However, the manual modeling and control required for the majority of digital humans pose significant obstacles to efficient development. The speech-driven methods offer a novel avenue for manipulating the mouth shape and expressions of digital humans. Despite the proliferation of driving methods, the quality of many generated talking head (TH) videos remains a concern, impacting user visual experiences. To tackle this issue, this paper introduces the Talking Head Quality Assessment (THQA) database, featuring 800 TH videos generated through 8 diverse speechdriven methods. Extensive experiments affirm the THQA database’s richness in character and speech features. Subsequent subjective quality assessment experiments analyze correlations between scoring results and speech-driven methods, ages, and genders. In addition, experimental results show that mainstream image and video quality assessment methods have limitations for the THQA database, underscoring the imperative for further research to enhance TH video quality assessment. The THQA database is publicly accessible at https://github.com/zyj-2000/THQA. Yingjie Zhou 0003, Wei Sun 0029, Xiaohong Liu 0001, Xiongkuo Min, Zhihua Wang 0002, Xiao-Ping Zhang 0002, Guangtao Zhai |
ICIP | 7 |
| 2024 | Deviation Wing Loss for High-Performance 2D Pose EstimationabstractHeatmap regression using deep neural networks has become the dominant approach in 2D pose estimation. Nonetheless, the intrinsic variability in the flexibility of distinct keypoints engenders discernible momentum deviation among them, leading to training biases. Moreover, previous methods indiscriminately treat all pixels in a heatmap, further exacerbating biases. Regrettably, conventional loss functions, including the Mean Squared Error (MSE) loss, fail to rectify this issue adequately. Consequently, the need arises to recalibrate weights to concentrate the loss’s impact on specific regions. To this end, we introduce a novel Deviation Wing (DW) loss function for high-performance 2D pose estimation, incorporating two improvement aspects. Firstly, we use Gaussian Momentum Deviation encoding craft deviation maps, leveraging momentum deviation as a source of prior knowledge to enhance supervision over keypoints characterized by substantial momentum deviation. Subsequently, we harness the Cosine Wing function to amplify the loss concerning minor errors residing within the keypoint region and supervise errors across diverse scales. Our comprehensive empirical exploration spans multiple datasets encompassing 2D human pose and hand pose estimation. The experimental results demonstrate the efficacy of our proposed loss function in enhancing heatmap regression performances. Junliang Xing, Xinchun Yu, Xiao-Ping Zhang 0002 |
ICME | 4 |
| 2024 | PowerGest: Self-Powered Gesture Recognition for Command Input and Robotic ManipulationabstractAs human-computer interaction (HCI) advances, gesture recognition has emerged as a transformative technology for human-computer interaction. Traditional methods, often camera or glove-based, are restricted by various environmental conditions and user-specific demands, highlighting the need for more universal, non-intrusive, and sustainable solutions. Addressing this, we present PowerGest, a self-powered gesture recognition system based on a solar cell array. This innovative system leverages the dual functionalities of solar cells: energy harvesting and gesture sensing, providing an alternative to conventional methods. It integrates a designed low-powered data acquisition chip with a wireless transmission module and a user-friendly interface. PowerGest employs a series of signal processing methods and utilizes several machine learning algorithms, achieving over 97% accuracy for both numeric input and activity control gesture recognition tasks. With its broad applications in robotic control, text input, and more, PowerGest contributes to a more sustainable and intuitive HCI experience. Project demo: https://drive.google.com/drive/folders/10KEul8PAfvTUomi0JvZ8u411JyScCXQI?usp=sharing. Jiarong Li 0004, Qinghao Xu, Zhancong Xu, Changshuo Ge, Liguang Ruan, Xiaojun Liang, Wenbo Ding 0001, Weihua Gui 0001, Xiao-Ping Zhang 0002 |
ICPADS | 9 |
| 2024 | Dual-modal Tactile E-skin: Enabling Bidirectional Human-Robot Interaction via Integrated Tactile Perception and FeedbackabstractTo foster an immersive and natural human-robot interaction (HRI), the implementation of tactile perception and feedback becomes imperative, effectively bridging the conventional sensory gap. In this paper, we propose a dual-modal electronic skin (e-skin) that integrates magnetic tactile sensing and vibration feedback for enhanced HRI. The dual-modal tactile e-skin offers multi-functional tactile sensing and programmable haptic feedback, underpinned by a layered structure comprised of flexible magnetic films, soft silicone elastomer, a Hall sensor and actuator array, and a microcontroller unit. The e-skin captures the magnetic field changes caused by subtle deformations through Hall sensors, employing deep learning for accurate tactile perception. Simultaneously, the actuator array generates mechanical vibrations to facilitate haptic feedback, delivering diverse mechanical stimuli. Notably, the dual-modal e-skin is capable of transmitting tactile information bidirectionally, enabling object recognition and fine-weighing operations. This bidirectional tactile interaction framework will enhance the immersion and efficiency of interactions between humans and robots. Shilong Mu, Zenan Lin, Shoujie Li, Chenchang Li, Xiao-Ping Zhang 0002, Wenbo Ding 0001 |
ICRA | 7 |
| 2024 | SATac: A Thermoluminescence Enabled Tactile Sensor for Concurrent Perception of Temperature, Pressure, and ShearabstractMost vision-based tactile sensors use elastomer deformation to infer tactile information, which can not sense some modalities, like temperature. As an important part of human tactile perception, temperature sensing can help robots better interact with the environment. In this work, we propose a novel multi-modal vision-based tactile sensor, SATac, which can simultaneously perceive information on temperature, pressure, and shear. SATac utilizes the thermoluminescence of strontium aluminate to sense a wide range of temperatures with exceptional resolution. Additionally, the pressure and shear can also be perceived by analyzing the Voronoi diagram. A series of experiments are conducted to verify the performance of our proposed sensor. We also discuss the possible application scenarios and demonstrate how SATac could benefit robot perception capabilities. Ziwu Song, Kit Wa Sou, Shilong Mu, Dengfeng Peng, Xiao-Ping Zhang 0002, Wenbo Ding 0001 |
ICRA | 7 |
| 2024 | QUEST: Quality-informed Multi-agent Dispatching System for Optimal Mobile CrowdsensingabstractWe address the challenges in achieving optimal Quality of Information (QoI) for non-dedicated vehicular Mobile Crowdsensing (MCS) systems, by utilizing vehicles not originally designed for sensing purposes to provide real-time data while moving around the city. These challenges include the coupled sensing coverage and sensing reliability, as well as the uncertainty and time-varying vehicle status. To tackle these issues, we propose QUEST, a QUality-informed multi-agEnt diSpaTching system, that ensures high sensing coverage and sensing reliability in non-dedicated vehicular MCS. QUEST optimizes QoI by introducing a novel metric called ASQ (aggregated sensing quality), which considers both sensing coverage and sensing reliability jointly. Additionally, we design a mutual-aided truth discovery dispatching method to estimate sensing reliability and improve ASQ under uncertain vehicle statuses. Real-world data from our deployed MCS system in a metropolis is used for evaluation, demonstrating that QUEST achieves up to 26% higher ASQ improvement, leading to a reduction of reconstruction map errors by 32-65% for different reconstruction algorithms. Zuxin Li, Fanhang Man, Xuecheng Chen, Susu Xu, Fan Dang 0002, Xiao-Ping Zhang 0002, Xinlei Chen |
INFOCOM | 6 |
| 2024 | TransformLoc: Transforming MAVs into Mobile Localization Infrastructures in Heterogeneous SwarmsabstractA heterogeneous micro aerial vehicles (MAV) swarm consists of resource-intensive but expensive advanced MAVs (AMAVs) and resource-limited but cost-effective basic MAVs (BMAVs), offering opportunities in diverse fields. Accurate and real-time localization is crucial for MAV swarms, but current practices lack a low-cost, high-precision, and real-time solution, especially for lightweight BMAVs. We find an opportunity to accomplish the task by transforming AMAVs into mobile localization infrastructures for BMAVs. However, turning this insight into a practical system is non-trivial due to challenges in location estimation with BMAVs’ unknown and diverse localization errors and resource allocation of AMAVs given coupled influential factors. This study proposes TransformLoc, a new framework that transforms AMAVs into mobile localization infrastructures, specifically designed for low-cost and resource- constrained BMAVs. We first design an error-aware joint location estimation model to perform intermittent joint location estimation for BMAVs and then design a proximity-driven adaptive grouping-scheduling strategy to allocate resources of AMAVs dynamically. TransformLoc achieves a collaborative, adaptive, and cost-effective localization system suitable for large-scale heterogeneous MAV swarms. We implement TransformLoc on industrial drones and validate its performance. Results show that TransformLoc outperforms baselines including SOTA up to 68% in localization performance, motivating up to 60% navigation success rate improvement. Haoyang Wang 0012, Jingao Xu, Chenyu Zhao 0002, Zihong Lu, Yuhan Cheng, Xuecheng Chen, Xiao-Ping Zhang 0002, Yunhao Liu 0001, Xinlei Chen |
INFOCOM | 7 |
| 2024 | Interpretable Temporal Class Activation Representation for Audio Spoofing Detection
Menglu Li, Xiao-Ping Zhang 0002 |
INTERSPEECH | 2 |
| 2024 | Multidimensional Similarity Fusion for Speech Quality AssessmentabstractSince the perceptual quality of audio signals is easily to be affected by compression, transmission, noise adding, etc, it is of great significance to develop an effective audio quality assessment (AQA) method to measure end-user’s quality of experience. In this paper, we propose a full reference AQA model named Multidimensional Similarity Fusion for Audio Quality Assessment (MSF-AQA). We generalize the similarity-based image quality assessment methods for audio, then extract audio similarity features from multiple dimensions, and finally regress the multidimensional similarity features into the final quality score. The experimental results across three databases indicate that our MSF-AQA model outperforms the state-of-the-art AQA methods. Xiongkuo Min, Yuqin Cao, Xiao-Ping Zhang 0002, Guangtao Zhai |
ISCAS | 4 |
| 2024 | Translating Motion to Notation: Hand Labanotation for Intuitive and Comprehensive Hand Movement Documentation
Wenrui Yang, Xinchun Yu, Junliang Xing, Xiao-Ping Zhang 0002 |
ACM Multimedia | 5 |
| 2024 | Misaligned Over-The-Air Computation of Multi-Sensor Data with Wiener-Denoiser NetworkabstractIn data driven deep learning, distributed sensing and joint computing bring heavy load for computing and communication. To face the challenge, over-the-air computation (OAC) has been proposed for multi-sensor data aggregation, which enables the server to receive a desired function of massive sensing data during communication. However, the strict synchronization and accurate channel estimation constraints in OAC are hard to be satisfied in practice, leading to time and channel-gain misalignment. The paper formulates the misalignment problem as a non-blind image deblurring problem. At the receiver side, we first use the Wiener filter to deblur, followed by a U-Net network designed for further denoising. Our method is capable to exploit the inherent correlations in the signal data via learning, thus outperforms traditional methods in term of accuracy. Our code is available at https://github.com/auto-Dog/MOAC_deep. Mingjun Du, Sihui Zheng, Xiao-Ping Zhang 0002, Yuhan Dong |
MobiCom | 3 |
| 2024 | Foes or Friends: Embracing Ground Effect for Edge Detection on Lightweight DronesabstractDrone-based rapid and accurate environmental edge detection is highly advantageous for tasks such as disaster relief and autonomous navigation. Current methods, using radar or cameras, raise deployment costs and burden lightweight drones with high computational demands. In this paper, we propose AirTouch, a system that transforms the ground effect from a stability "foe" in traditional flight control views, into a "friend" for accurate and efficient edge detection. Our key insight is that analyzing drone sensor readings and flight commands allows us to detect ground effect changes. Such changes typically indicate the drone flying over an edge, making this information valuable for edge detection. We approach this insight through theoretical analysis, algorithm design, and implementation, fully leveraging the ground effect as a new sensing modality without compromising drone flight stability, thereby achieving accurate and efficient scene edge detection. Extensive evaluations demonstrate that our system achieves a high detection accuracy with mean detection distance errors of 0.051m, outperforming the baseline performance by 86%. Chenyu Zhao 0002, Ciyu Ruan, Jingao Xu, Haoyang Wang 0012, Jiaqi Li 0028, Jirong Zha, Zheng Yang 0002, Yunhao Liu 0001, Xiao-Ping Zhang 0002, Xinlei Chen |
MobiCom | 10 |
| 2024 | Poster: Real-time Material and Texture Recognition Using Visible Light CommunicationabstractIn response to the challenges presented by conventional material and texture recognition methods, our research introduces a system using visible light communication (VLC) technology. This approach provides a non-contact, non-destructive, dual-functional solution, overcoming the limitations of cost, safety, and environmental adaptability associated with traditional methods. Through a comprehensive design integrating hardware and software, our system utilizes VLC for precise and efficient recognition. Extensive testing confirms its effectiveness, achieving 97.7% accuracy in material identification and 93.8% in texture detection. This study highlights VLC's potential in enhancing automated recognition systems across various applications. Jiarong Li 0004, Chenxin Liang, Xiaojun Liang, Wenbo Ding 0001, Jian Song 0004, Xiao-Ping Zhang 0002 |
MobiSys | 7 |
| 2024 | Demo: SolarSense: A Self-powered Ubiquitous Gesture Recognition System for Industrial Human-Computer InteractionabstractSolarSense is a self-powered sensing system for gesture recognition using solar cell arrays, thereby offering a sustainable approach to human-computer interaction (HCI) within industrial settings. The system effectively employed the sensing and energy harvesting capabilities of solar cells, achieving over 97.0% accuracy in recognizing diverse gestures. The design incorporates a low-power wireless data acquisition chip, a signal processing framework, and a user interface to realize robotic control and text input applications. SolarSense enhances HCI with its eco-friendly and user-centric approach, which is suitable for Internet of things (IoT) scenarios. Demo: https://youtu.be/RmPolChw_c4. Jiarong Li 0004, Qinghao Xu, Qingyang Zhu, Zhancong Xu, Changshuo Ge, Liguang Ruan, H. Y. Fu 0001, Xiaojun Liang, Wenbo Ding 0001, Weihua Gui 0001, Xiao-Ping Zhang 0002 |
MobiSys | 12 |
| 2024 | MobiAir: Unleashing Sensor Mobility for City-scale and Fine-grained Air-Quality Monitoring with AirBERTabstractMobile air pollution sensing methods are developed to collect air quality data with higher spatial-temporal resolutions. However, existing methods cannot process the spatially mixed gas samples effectively due to the highly dynamic temporal and spatial fluctuations experienced by the sensor, leading to significant measurement deviations. We find an opportunity to tackle the problem by exploring the potential patterns from sensor measurements. In light of this, we propose MobiAir, a novel city-scale fine-grained air quality estimation system to deliver accurate mobile air quality data. First, we design AirBERT, a representation learning model to discern mixed gas concentrations. Second, we design a knowledge-informed training strategy leveraging massive unlabeled city-scale data to enhance the AirBERT performance. To ensure the practicality of MobiAir, we have invested significant efforts in implementing the software stack on our meticulously crafted Sensing Front-end, which has successfully gathered air quality data at a city-scale for more than 1200 hours. Experiments conducted on collected data show that MobiAir reduces sensing errors by 96.7% with only 44.9ms latency, outperforming the SOTA baseline by 39.5%. Yuxuan Liu 0010, Haoyang Wang 0012, Fanhang Man, Jingao Xu, Fan Dang 0001, Yunhao Liu 0001, Xiao-Ping Zhang 0002, Xinlei Chen |
MobiSys | 7 |
| 2024 | Multi-scale Consistency for Robust 3D Registration via Hierarchical Sinkhorn TreeabstractWe study the problem of retrieving accurate correspondence through multi-scale consistency (MSC) for robust point cloud registration. Existing works in a coarse-to-fine manner either suffer from severe noisy correspondences caused by unreliable coarse matching or struggle to form outlier-free coarse-level correspondence sets. To tackle this, we present Hierarchical Sinkhorn Tree (HST), a pruned tree structure designed to hierarchically measure the local consistency of each coarse correspondence across multiple feature scales, thereby filtering out the local dissimilar ones. In this way, we convert the modeling of MSC for each correspondence into a BFS traversal with pruning of a K-ary tree rooted at the superpoint, with its K nearest neighbors in the feature pyramid serving as child nodes. To achieve efficient pruning and accurate vicinity characterization, we further propose a novel overlap-aware Sinkhorn Distance, which retains only the most likely overlapping points for local measurement and next level exploration. The modeling process essentially involves traversing a pair of HSTs synchronously and aggregating the consistency measures of corresponding tree nodes. Extensive experiments demonstrate HST consistently outperforms the state-of-the-art methods on both indoor and outdoor benchmarks. Chengwei Ren, Yifan Feng 0001, Weixiang Zhang, Xiao-Ping Zhang 0002, Yue Gao 0002 |
NeurIPS | 4 |
| 2024 | Joint Optimization in MEC Incorporating MD Preference: A Hybrid GA and AFSA SchemeabstractMobile edge computing (MEC) is an efficient method to tackle computationally intensive tasks for mobile devices (MDs). However, current studies about MEC do not consider that different MDs have different preferences for delay and energy consumption. Thus, we propose a MD preference-based MEC computing model addressing three optimization goals: delay preference, energy preference, and their balance, tailored to diverse MD preference. This optimization problem is formulated as a mixed integer non-linear programming (MINLP) task with four optimization variables: offloading decisions, channel allocation, power allocation, and resource allocation. Additionally, we propose a novel genetic artificial fish swarm cooperative optimization algorithm (GAFSCOA) to solve this problem, which integrates genetic algorithm (GA) and artificial fish swarm algorithm (AFSA), respectively. Numerical results show our proposed model can achieve different optimization goals according to MDs’ preferences. Compared with GA and AFSA, our proposed GAFSCOA demonstrates 30.71% faster convergence speed and 46.09% better optimization results. Furthermore, in comparison with other baseline algorithms, our algorithm yields superior optimization results. Yunan Dong, Jianhua Pei, Yuhan Dong, Xiao-Ping Zhang 0002 |
VTC Fall | 5 |
| 2024 | Deep Reinforcement Learning Based Contention Window Optimization for IEEE 802.11 bnabstractThe up-to-date project authorization request (PAR) for IEEE 802.11 bn envisions achieved Ultra High Reliability capability on the basis of the Extremely High Throughput Wi-Fi (IEEE 802.11 be). Specifically, it calls for optimizing the 95th percentile of the latency distribution and MAC Protocol Data Unit (MPDU) loss while ensuring high throughput. The vision is challenging due to the competing nature of Wi-Fi channel access, especially in the case of overlapping basic service set (OBSS). This challenge gives rise to an emerging research topic of Wi-Fi, i.e., low-latency channel access. In this paper, we materialize low-latency channel access via multi-agent reinforcement learning (MARL). To meet the Wi-Fi legacy requirement, we do not drop the carrier sense multiple access with collision avoidance (CSMA/CA) protocol but resolve to tune contention window (CW) being intelligent, a critical control parameter of the protocol. The objective of intelligent adapting CW is to minimize tail latency constrained by throughput and MPDU loss, which is consistent with the PAR. The control optimization problem is then solved by MARL. The adopted method promises effectiveness in distributed learning, which is validated in many simulations covering different topologies. Extensive simulation results show that this method has an average substantive gain of more than 25%. Mingjun Du, Xiao-Ping Zhang 0002, Yuhan Dong |
VTC Spring | 3 |
| 2024 | SOScheduler: Toward Proactive and Adaptive Wildfire Suppression via Multi-UAV Collaborative SchedulingabstractMulti-UAV systems have shown immense potential in handling complex tasks in large-scale, dynamic, and cold-start (i.e., limited prior knowledge) scenarios, such as wildfire suppression. Due to the dynamic and stochastic environmental conditions, the scheduling for sensing tasks (i.e., fire monitoring) and operation tasks (i.e., fire suppression) should be executed concurrently to enable real-time information collection and timely intervention of the environment. However, the planning inclinations of sensing and operation tasks are typically inconsistent and evolve over time, complicating the task of identifying the optimal strategy for each UAV. To solve this problem, this paper proposes SOScheduler, a collaborative multi-UAV scheduling framework for integrated sensing and operation in large-scale and dynamic wildfire environments. We introduce a spatio-temporal confidence-aware assessment model to dynamically and directly pinpoint locations that can optimally enhance the understanding of environmental dynamics and operational effectiveness, as well as a priority graph-instructed scalable scheduler to coordinate multi-UAV in an efficient manner. Experiments on real multi-UAV testbeds and large-scale physical feature-based simulations show that our SOScheduler reduces the fire expansion ratio by 59% and enhances the fire coverage ratio by 190% compared to state-of-the-art (SOTA) solutions. Xuecheng Chen, Zijian Xiao, Yuhan Cheng, Chen-Chun Hsia, Haoyang Wang 0012, Jingao Xu, Susu Xu, Fan Dang 0001, Xiao-Ping Zhang 0002, Yunhao Liu 0001, Xinlei Chen |
IEEE Internet Things J. | 9 |
| 2024 | A Multimode Neuromorphic Vision Sensor With Improved Brightness Measurement Performance by Pulse Coding MethodabstractThis article proposes a multimode neuromorphic event-frame integrated vision sensor that enables event detection (ED) with simultaneous brightness measurement based on the pulse width modulation mechanism. The logarithmic voltage is directly taken as intensity information. Brightness measurement involves in-pixel voltage-to-pulse conversion and out-pixel pulse coding. The maximum event bandwidth is improved to 366 Meps by pipelining the time-prior arbiter along with the address-events grouping circuit. A wide intensity dynamic range of 105 dB can theoretically be achieved through logarithmic photoelectric conversion and pulse coding. Our sensor supports a$128\times 64$frame-like image with an improved signal-to-noise ratio of 49 dB. The experimental results indicate that the log sensitivity of the optimized logarithmic photoreceptor was measured as 164 mV/dec. The equivalent frame rate for both event and intensity reaches kilo fps, making it a promising candidate in high-speed wireless sensing applications. Zewei Ding, Qijuan Wu, Mingyu Wang 0001, Jingjing Liu 0004, Xiaoyang Zeng, Wenhong Li, Zhi Liu 0004, Xiao-Ping Zhang 0002 |
IEEE Internet Things J. | 9 |
| 2024 | Multistate Constraint Multipath-Assisted Positioning and Mismatch AlleviationabstractMultipath propagation greatly affects the accuracy of time of arrival (ToA)-based indoor positioning when line-of-sight (LOS) signals are only used. In this paper, we present a novel real-time and low computation complexity multipath-assisted ToA positioning method, namely MSC-MAP. The delays of reflected signals are taken as additional spatial observations to compensate for an insufficient number of physical transmitters to locate a moving user equipment (UE). Virtual anchors are used to model the propagation path of reflected signals, whose locations are obtained via a multi-state constraint estimator, along with the trajectory of UE. In addition, we demonstrate the mismatch problem in data association and its impact on positioning performance. To achieve real-time processing, we propose two robust multipath-assisted positioning methods with mismatch alleviation by randomly selecting subset and constraint relaxation respectively, to meet various computational complexity requirements. Simulation results show that, for the MSC-MAP method, the mean square error of the position is generally less than 0.2 m in challenging indoor environments. Among mismatch alleviation algorithms, positioning error is reduced by 69% even when the percentage of mismatched measurement data is as high as 42%. The proposed algorithms can also efficiently handle signals with non-Gaussian impairments, a common characteristic in real-world data. Moreover, these algorithms can substantially improve positioning performance while adding minimal computation time in the presence of measurement mismatches, outperforming state-of-the-art methods utilizing different data association techniques. Xueting Xu, Ao Peng, Xuemin Hong, Yixiong Zhang, Xiao-Ping Zhang 0002 |
IEEE Internet Things J. | 5 |
| 2024 | DOVE: Doodled vessel enhancement for photoacoustic angiography super resolution
Yuanzheng Ma, Wangting Zhou, Erqi Wang, Sihua Yang, Yansong Tang, Xiao-Ping Zhang 0002, Xun Guan |
Medical Image Anal. | 7 |
| 2024 | TEFISTA-Net: A learnable method for high-resolution range profile reconstruction with low-frequency ultra-wideband radar
Xueqian Wang 0002, Gang Li 0008, Xiao-Ping Zhang 0002 |
Signal Process. | 4 |
| 2024 | Blind Image Quality Assessment: A Fuzzy Neural Network for Opinion Score Distribution PredictionabstractImage quality assessment (IQA) has always been a popular research topic. There have been many methods proposed for predicting image quality, also known as the mean opinion score (MOS). However, it is worth noting that different people may assign different opinion scores to the same image. Image quality described by all subjective opinion scores can express rich subjective information about the image, such as diversity and uncertainty, which cannot be accurately described by a single MOS. Therefore, this paper proposes a fuzzy neural network to predict the opinion score distribution (OSD) of image quality. The fuzzy neural network includes three sub-networks: a feature extraction network, a feature fuzzification network, and a fuzzy learning network. First, a novel network is designed to extract image features. The extracted features are then fuzzified by fuzzy theory to model the epistemic uncertainty in the feature extraction process. Finally, the OSD of image quality is predicted using the fuzzy learning network by learning the mapping from fuzzy features to fuzzy uncertainty when rating image quality. In addition, to train the proposed fuzzy neural network, we employ a new loss function based on the quantile and the cumulative density function. We experimentally validate the feasibility and superiority of the proposed method in two aspects. On the one hand, we demonstrate the performance of the proposed method in predicting the OSD of image quality on the SJTU IQSD and KonIQ-10K databases. On the other hand, we also prove the feasibility of the proposed method in predicting the MOS of image quality on several popular IQA databases, including CSIQ, TID2013, LIVE MD, and LIVE Challenge. Xiongkuo Min, Yucheng Zhu, Xiao-Ping Zhang 0002, Guangtao Zhai |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | FlyCore: Fast Low-Frequency Coarse Registration of Large-Scale Outdoor LiDAR Point CloudsabstractFast and accurate registration of outdoor LiDAR point clouds poses a considerable challenge for their large-scale (e.g., 300 K points) and intricate (e.g., noise and outliers) distributions. In this article, we present a fast low-frequency coarse registration method for large-scale outdoor LiDAR point clouds, dubbed FlyCore. Different from existing methods, FlyCore is very fast for practical applications and bridges current refinement registration methods smoothly for their accuracy improvements. Specifically, we first construct spherical feature spaces for a pair of point clouds based on their keypoints and saliency uncertainties independently. Then, we perform harmonic decomposition on these spherical feature spaces, utilizing the low-frequency components of spherical harmonics (SHs) to implement point cloud registration. FlyCore demonstrates less sensitivity to noise and outliers compared to feature-based registration techniques. Also, FlyCore achieves exceptionally low time complexity by eliminating the need for feature matching and iterative procedures, ensuring fine alignment with only a few iterations. Experimental validations, utilizing two extensive LiDAR datasets featuring urban and natural scenarios, confirm the effectiveness and accuracy improvement of existing fine registration methods facilitated by our FlyCore. Zikuan Li, Kaijun Zhang, Zhoutao Wang, Sibo Wu, Xiao-Ping Zhang 0002, Mingqiang Wei, Jun Wang 0039 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | Exploring the Spectral Prior for Hyperspectral Image Super-ResolutionabstractIn recent years, many single hyperspectral image super-resolution methods have emerged to enhance the spatial resolution of hyperspectral images without hardware modification. However, existing methods typically face two significant challenges. First, they struggle to handle the high-dimensional nature of hyperspectral data, which often results in high computational complexity and inefficient information utilization. Second, they have not fully leveraged the abundant spectral information in hyperspectral images. To address these challenges, we propose a novel hyperspectral super-resolution network named SNLSR, which transfers the super-resolution problem into the abundance domain. Our SNLSR leverages a spatial preserve decomposition network to estimate the abundance representations of the input hyperspectral image. Notably, the network acknowledges and utilizes the commonly overlooked spatial correlations of hyperspectral images, leading to better reconstruction performance. Then, the estimated low-resolution abundance is super-resolved through a spatial spectral attention network, where the informative features from both spatial and spectral domains are fully exploited. Considering that the hyperspectral image is highly spectrally correlated, we customize a spectral-wise non-local attention module to mine similar pixels along spectral dimension for high-frequency detail recovery. Extensive experiments demonstrate the superiority of our method over other state-of-the-art methods both visually and metrically. Our code is publicly available at https://github.com/HuQ1an/SNLSR. Xinya Wang, Junjun Jiang, Xiao-Ping Zhang 0002, Jiayi Ma 0001 |
IEEE Trans. Image Process. | 4 |
| 2024 | UCL-Dehaze: Toward Real-World Image Dehazing via Unsupervised Contrastive LearningabstractWhile the wisdom of training an image dehazing model on synthetic hazy data can alleviate the difficulty of collecting real-world hazy/clean image pairs, it brings the well-known domain shift problem. From a different yet new perspective, this paper explores contrastive learning with an adversarial training effort to leverage unpaired real-world hazy and clean images, thus alleviating the domain shift problem and enhancing the network's generalization ability in real-world scenarios. We propose an effective unsupervised contrastive learning paradigm for image dehazing, dubbed UCL-Dehaze. Unpaired real-world clean and hazy images are easily captured, and will serve as the important positive and negative samples respectively when training our UCL-Dehaze network. To train the network more effectively, we formulate a new self-contrastive perceptual loss function, which encourages the restored images to approach the positive samples and keep away from the negative samples in the embedding space. Besides the overall network architecture of UCL-Dehaze, adversarial training is utilized to align the distributions between the positive samples and the dehazed images. Compared with recent image dehazing works, UCL-Dehaze does not require paired data during training and utilizes unpaired positive/negative data to better enhance the dehazing performance. We conduct comprehensive experiments to evaluate our UCL-Dehaze and demonstrate its superiority over the state-of-the-arts, even only 1,800 unpaired real-world images are used to train our network. Source code is publicly available at https://github.com/yz-wang/UCL-Dehaze. Yongzhen Wang 0001, Xuefeng Yan 0001, Fu Lee Wang, Haoran Xie 0001, Wenhan Yang, Xiao-Ping Zhang 0002, Harry Qin, Mingqiang Wei |
IEEE Trans. Image Process. | 6 |
| 2024 | SAFARI: Sparsity-Enabled Federated Learning With Limited and Unreliable CommunicationsabstractFederated learning (FL) enables edge devices to collaboratively learn a model in a distributed fashion. Many existing researches have focused on improving communication efficiency of high-dimensional models and addressing bias caused by local updates. However, most FL algorithms are either based on reliable communications or assuming fixed and known unreliability characteristics. In practice, networks could suffer from dynamic channel conditions and non-deterministic disruptions, with time-varying and unknown characteristics. To this end, in this paper we propose a sparsity-enabled FL framework with both improved communication efficiency and bias reduction, termed as SAFARI. It makes use of similarity among client models to rectify and compensate for bias that results from unreliable communications. More precisely, sparse learning is implemented on local clients to mitigate communication overhead, while to cope with unreliable communications, a similarity-based compensation method is proposed to provide surrogates for missing model updates. With respect to sparse models, we analyze SAFARI under bounded dissimilarity. It is demonstrated that SAFARI under unreliable communications is guaranteed to converge at the same rate as the standard FedAvg with perfect communications. Implementations and evaluations on the CIFAR-10 dataset validate the effectiveness of SAFARI by showing that it can achieve the same convergence speed and accuracy as FedAvg with perfect communications, with up to 60% of the model weights being pruned and a high percentage of client updates missing in each round of model updates. Yuzhu Mao, Zihao Zhao 0001, Meilin Yang, Le Liang, Yang Liu 0165, Wenbo Ding 0001, Tian Lan 0001, Xiao-Ping Zhang 0002 |
IEEE Trans. Mob. Comput. | 8 |
| 2024 | UAV-Assisted Target Tracking and Computation Offloading in USV-Based MEC NetworksabstractIn recent years, unmanned aerial vehicles (UAVs) have been widely used in ocean target tracking and image acquisition for processing. Due to the limited energy of the UAV and the high computational complexity associated with image processing tasks, a lightweight energy-saving target tracking scheme is designed for the UAV, and the unmanned surface vehicle (USV) based mobile edge computing (MEC) networks are adopted to share the computing load of the UAV. Due to the randomness of the environment, we formulate data processing, computation offloading, resource allocation, and target-tracking as a joint stochastic optimization problem. This paper investigates a two-stage optimization scheme to address the problem. Firstly, we employ a Lyapunov-based approach to convert the stochastic optimization problem into a deterministic per-time slot problem under communication and computing resources constraints. Then, we develop a real-time target tracking scheme for the UAV based on the Elman neural network. Numerical results validate that the designed tracking scheme can effectively minimize propulsion energy consumption while maintaining a high success rate in tracking. Furthermore, the proposed method balances data-related energy consumption, image detection accuracy, and stability of the data storage queue. Ziyuan Wang 0002, Jun Du 0001, Chunxiao Jiang, Yong Ren 0001, Xiao-Ping Zhang 0002 |
IEEE Trans. Mob. Comput. | 5 |
| 2024 | AQUILA: Communication Efficient Federated Learning With Adaptive Quantization in Device Selection StrategyabstractThe widespread adoption of Federated Learning (FL), a privacy-preserving distributed learning methodology, has been impeded by the challenge of high communication overheads, typically arising from the transmission of large-scale models. Existing adaptive quantization methods, designed to mitigate these overheads, operate under the impractical assumption of uniform device participation. Additionally, these methods are limited in their adaptability due to the necessity of manual quantization level selection and often overlook biases inherent in local devices' data, thereby affecting the robustness of the global model. In response, this paper introduces AQUILA (adaptivequantization in device selection strategy), a novel adaptive framework devised to effectively handle these issues, enhancing the efficiency and robustness of FL. AQUILA integrates a sophisticated device selection method that prioritizes the quality and usefulness of device updates. Utilizing the exact global model stored by devices enables a more precise device selection criterion, reduces model deviation, and limits the need for hyperparameter adjustments. Furthermore, AQUILA presents an innovative quantization criterion, optimized to improve communication efficiency while assuring model convergence. Our experiments demonstrate that AQUILA significantly decreases communication costs compared to existing methods, while maintaining comparable model performance across diverse non-homogeneous FL settings, such as Non-IID data and heterogeneous model architectures. Zihao Zhao 0001, Yuzhu Mao, Zhenpeng Shi, Yang Liu 0165, Tian Lan 0001, Wenbo Ding 0001, Xiao-Ping Zhang 0002 |
IEEE Trans. Mob. Comput. | 7 |
| 2024 | VWP:An Efficient DRL-Based Autonomous Driving ModelabstractIn this paper, a novel DRL-based model (VWP, VAE-WGAN-PPOE) is proposed to solve the problem of long training time and unsatisfactory training effect in the end-to-end autonomous driving. The model is optimized from feature extraction and algorithm decision. In feature extraction, we encode the input video by combining variational auto encoder (VAE) with wasserstein generative adversarial network (WGAN). The state dimension is reduced and the problem of mode collapse and gradient disappearance caused by generative adversarial network (GAN) training is solved. In decision algorithm, we formulate a new reward function by analyzing the factors affecting driving performance. Furthermore, we propose an enhanced algorithm PPOE based on the proximal policy optimization (PPO). In the CARLA simulator, compared with CNN and ResNet34, the convergence speed of the DRL model based on VAE-WGAN increases by 26.1% and 20.3%, the navigation task completion rate increases by 18.5% and 9.2%, and the collision rate decreases by 13.6% and 9.4%. Compared with deep deterministic policy gradient (DDPG) decision algorithm, the convergence speed of the DRL model based on PPOE increases by 23.3%, the navigation task completion rate increases by 5.0% in sunny days and 8.4% in severe weather, the collision rate decreases by 3.5% in sunny days and 6.6% in severe weather. Extensive experiments show that the proposed model enables the agent to drive safely along the navigational route in the complex environment with pedestrian and vehicle interaction, even in severe weather. Yanliang Jin, Ze-Yu Ji, Dan Zeng 0001, Xiao-Ping Zhang 0002 |
IEEE Trans. Multim. | 4 |
| 2024 | Lightweight Video-Based Respiration Rate Detection Algorithm: An Application Case on Intensive CareabstractThe video-based non-contact respiration detection technology can be used in many application scenarios to unobtrusively and ubiquitously monitor the physical state of living beings, and various researchers are currently working on this technology. The optical flow method in tandem with crossover point method is rather effective for respiration rate extraction. However, each method has one disadvantage: 1) the redundant feature points in the traditional optical flow method increase the computational effort and reduce the estimation accuracy; and 2) the traditional crossover point method suffers from crossover points unrelated to breathing movements. For these two challenges, two optimization points are proposed in this work: 1) optimize feature point space by combining spatio-temporal information; and 2) use negative feedback design to adaptively remove crossovers that are not related to respiratory movements. The performance of the proposed algorithm is validated by the Large-scale Bedside Respiration Dataset for Intensive Care (LBRD-IC), which is established using the actual surveillance videos acquired from ICU wards. The validity of the above two optimization points is verified by the ablation experiments. The influential analysis of computation time and video resolution on the performance of the proposed algorithm demonstrates that the proposed algorithm can be deployed to various application terminals to monitor the respiration rate of living organisms in real-time and with high accuracy. In addition, field measurements in the ICU ward have shown that our algorithm can measure respiratory signals of the single patient and multiple patients when only one surveillance camera is present. Xudong Tan, Menghan Hu, Guangtao Zhai, Wenfang Li, Xiao-Ping Zhang 0002 |
IEEE Trans. Multim. | 6 |
| 2024 | Hierarchical Independent Coding Scheme for Varifocal Multiview Images Based on Angular-Focal Joint PredictionabstractVarifocal multiview (VFMV) images are dense views that focus on variable focal planes. Thus, VFMV images are highly redundant in the angular, spatial and focal dimensions. In this article, the redundancies of VFMV images are analyzed and represented by full parallaxes and focal inconsistency. To exploit these distinctive redundancies, we propose a hierarchical independent coding scheme based on angular-focal joint prediction. The scheme is constructed by hierarchical independent prediction structure (HIPS) and angular-focal joint prediction (AFJP). The HIPS separates all views into several independent subdivisions and assigns different hierarchies inside each subdivision, which enhances random access capability and scalability. The AFJP conducts motion estimation and focal approximation simultaneously to predict parallaxes and focal inconsistency. Therefore, the redundancies in the angular and focal dimensions can be exploited by the proposed coding scheme. We construct a VFMV dataset with 10 test sequences for different acquisition methods. The experimental results on these test sequences demonstrate that the proposed scheme outperforms all comparison schemes in objective quality, subjective quality and random access capability. Specifically, the proposed coding scheme achieves up to 2.661 dB PSNR gains and 52.817% bitrate savings compared with the HEVC random access benchmark scheme. Kejun Wu, You Yang 0002, Qiong Liu 0001, Gangyi Jiang, Xiao-Ping Zhang 0002 |
IEEE Trans. Multim. | 5 |
| 2024 | Time-Distributed Feature Learning for Internet of Things Network Traffic ClassificationabstractDeep learning-based network traffic classification (NTC) techniques, including conventional and class-of-service (CoS) classifiers, are a popular tool that aids in the quality of service (QoS) and radio resource management for the Internet of Things (IoT) network. Holistic temporal features consist of inter-, intra-, and pseudo-temporal features within packets, between packets, and among flows, providing the maximum information on network services without depending on defined classes in a problem. Conventional spatio-temporal features in the current solutions extract only space and time information between packets and flows, ignoring the information within packets and flow for IoT traffic. Therefore, we propose a new, efficient, holistic feature extraction method for deep-learning-based NTC using time-distributed feature learning to maximize the accuracy of the NTC. We apply a time-distributed wrapper on deep-learning layers to help extract pseudo-temporal features and spatio-temporal features. Pseudo-temporal features are mathematically complex to explain since, in deep learning, a black box extracts them. However, the features are temporal because of the time-distributed wrapper; therefore, we call them pseudo-temporal features. Since our method is efficient in learning holistic-temporal features, we can extend our method to both conventional and CoS NTC. Our solution proves that pseudo-temporal and spatial-temporal features can significantly improve the robustness and performance of any NTC. We analyze the solution theoretically and experimentally on different real-world datasets. The experimental results show that the holistic-temporal time-distributed feature learning method, on average, is 13.5% more accurate than the state-of-the-art conventional and CoS classifiers. Yoga Suhas Kuruba Manjunath, Sihao Zhao, Xiao-Ping Zhang 0002, Lian Zhao |
IEEE Trans. Netw. Serv. Manag. | 3 |
| 2024 | M$^{3}$Tac: A Multispectral Multimodal Visuotactile Sensor With Beyond-Human Sensory CapabilitiesabstractTo realize the exquisite interaction and precise manipulation for the robot, in this article, we propose a multispectral multimodal visuotactile sensor named M$^{3}$Tac, which combines visible, near-infrared, and mid-infrared imaging technologies for the first time and can exceed the sensing ability of human skin in terms of resolution (719 pixels/cm$^{2}$), temperature sensing range (−20–130$^\text{o}$C), etc. The M$^{3}$Tac cannot only realize high-quality sensing of deformation, texture, force, stickiness, and temperature comparable to human skin but also can realize proximity sensing that is lacking for human skin. To achieve this, we not only design a multispectral imaging system with an elastic film whose light penetrability can be regulated by the brightness of the light, but also develop corresponding algorithms, including the pixel-level force sensing with finite element method (accuracy:$\pm$0.023N), the proximity perception (accuracy:$\pm$3.8 mm), the 3-D reconstruction (accuracy: 0.33 mm), the super-resolution temperature sensing (accuracy:$\pm 0.3^\text{o}$C), the multimodal fusion classification (accuracy: 98%), and the stickiness recognition (accuracy: 98%). Finally, we conduct experiments to verify the effectiveness and application potential of our research. Shoujie Li, Haixin Yu, Guoping Pan, Huaze Tang, Jiawei Zhang 0012, Linqi Ye, Xiao-Ping Zhang 0002, Wenbo Ding 0001 |
IEEE Trans. Robotics | 7 |
| 2023 | Unified Probability Distributions of Composite Fading for Radio Frequency and Optical Wireless ChannelsabstractWe derive novel unified probability density function (PDF) formulae of the instantaneous envelope, power and signal-to-noise ratio (SNR), for a generalized composite fading process that models the multiplicative effect of small-scale fading and large-scale shadowing/fluctuations in radio frequency and free-space optical wireless communications channels. By specifying four fading parameters according to specific composite fading models, the unified PDF formulae enable us to unify and characterize wide ranging of existing and numerous novel generalized composite fading distributions. We also present unified analysis and derive general closed form formulae for channel capacity in terms of four parameters of the unified PDF. Such generalized performance metrics can then be further specified and evaluated according to various known and numerous novel composite fading channel models. Chin Choy Chai, Xiao-Ping Zhang 0002 |
GLOBECOM | 2 |
| 2023 | Robust Deep Joint Source Channel Coding with Time-Varying NoiseabstractDeep Joint Source-Channel Coding (JSCC) has gained increased attention, asserting its significance in the communication field. However, existing Deep JSCC techniques struggle to mitigate time-varying noise due to the deep neural networks being trained beforehand and fixed. To address this issue, we propose a robust deep JSCC scheme. Firstly, a multi-network parallel structure, as well as error-correcting codes, is introduced to effectively exploit label information. Secondly, a closed-form linear encoder and decoder pair is employed at the input and output ends of the channel to deal with the varying noise, which releases the neural network from dealing with a large range of varying noise levels. Thirdly, a transfer learning algorithm is utilized for estimating real-time noise statistics, which outperforms conventional estimation methods when noise statistics are time-dependent. These three components are effectively integrated as a comprehensive transmission system. Experimental results demonstrate that our optimized scheme outperforms existing approaches in the literature. Weida Wang, Xinchun Yu, Xinyi Tong 0002, Xiao-Ping Zhang 0002, Shao-Lun Huang |
GLOBECOM | 5 |
| 2023 | Iterative Water-Filling Power and Subcarrier Allocation for Multicarrier NOMA DownlinkabstractNovel closed form formulae of iterative optimal power control and allocation, and criterion for optimal subcarrier allocation are derived for downlink of multicarrier non-orthogonal multiple access (MC-NOMA) systems. For the first time, we present closed form water-filling formulae that quantify exactly the following effects of multiple access interference for optimal power allocation of MC-NOMA downlink: (i) Each user should take into account the sum of interference from other stronger users plus the receiver Gaussian noise as an equivalent interference in each subcarrier; (ii) The sum of each user’s interference to other weaker users should be accounted as factors which determine the distinct water levels in each subcarrier of each of the other weaker users. We also propose novel iterative water-filling algorithm and projected conjugate gradient algorithm that facilitate iterative computation of optimal power allocation for MC-NOMA downlink. Chin Choy Chai, Xiao-Ping Zhang 0002 |
ICASSP | 2 |
| 2023 | TEFISTA-NET: GTD Parameter Estimation of Low-Frequency Ultra- Wideband Radar via Model-Based Deep LearningabstractThe geometrical theory of diffraction (GTD) has been widely investigated to describe the target scattering behaviors with the low-frequency ultra-wideband (LFW) radar. In this paper, we propose a new model-based deep learning method for GTD parameter estimation. The proposed method is designed by unfolding the fast iterative shrinkage thresholding algorithm (FISTA) into a deep neural network. Unlike existing methods based on compressed sensing (CS), the key parameters in our algorithm are fully learnable, avoiding nontrivial parameter tuning procedures. Our network with simple convolution operations is more computationally efficient than existing methods, which require matrix inversions or quadratic programming and have low convergence speed. A novel loss function is designed for the new network to improve the capacity of target enhancement. Experiments on simulation data show that the new method achieves higher computational efficiency while maintaining or improving the precision of GTD parameter estimation compared with existing methods. Xueqian Wang 0002, Gang Li 0008, Xiao-Ping Zhang 0002 |
ICASSP | 4 |
| 2023 | Unobtrusive Respiratory Monitoring System for Intensive CareabstractThe video-based non-contact respiration detection technology can be used in many application scenarios to unobtrusively and ubiquitously monitor the physical state of living beings, and various researchers are currently working on this technology. The optical flow method in tandem with crossover point method is rather effective for respiration rate extraction. However, each method has one disadvantage: 1) the redundant feature points in the traditional optical flow method increase the computational effort and reduce the estimation accuracy; and 2) the traditional crossover point method suffers from crossover points unrelated to breathing movements. For these two challenges, two optimization points are proposed 1) optimize feature point space by combining spatio-temporal information; and 2) use negative feedback design to adaptively remove crossovers unrelated to respiratory movements. The performance of the proposed algorithm is validated by the Large-scale Bedside Respiration Dataset for Intensive Care (LBRD-IC), which is established using the actual surveillance videos acquired from ICU wards. In addition, field measurements in the ICU ward have shown that our algorithm can measure respiratory signals of the single patient and multiple patients when only one surveillance camera is present. Xudong Tan, Menghan Hu, Guangtao Zhai, Wenfang Li, Xiao-Ping Zhang 0002 |
ICASSP | 6 |
| 2023 | Tem-adapter: Adapting Image-Text Pretraining for Video Question AnswerabstractVideo-language pre-trained models have shown remarkable success in guiding video question-answering (VideoQA) tasks. However, due to the length of video sequences, training large-scale video-based models incurs considerably higher costs than training image-based ones. This motivates us to leverage the knowledge from image-based pretraining, despite the obvious gaps between image and video domains. To bridge these gaps, in this paper, we propose Tem-adapter, which enables the learning of temporal dynamics and complex semantics by a visual Temporal Aligner and a textual Semantic Aligner. Unlike conventional pretrained knowledge adaptation methods that only concentrate on the downstream task objective, the Temporal Aligner introduces an extra language-guided autoregressive task aimed at facilitating the learning of temporal dependencies, with the objective of predicting future states based on historical clues and language guidance that describes event progression. Besides, to reduce the semantic gap and adapt the textual representation for better event description, we introduce a Semantic Aligner that first designs a template to fuse question and answer pairs as event descriptions and then learns a Transformer decoder with the whole video sequence as guidance for refinement. We evaluate Tem-adapter and different pre-train transferring methods on two VideoQA benchmarks, and the significant performance improvement demonstrates the effectiveness of our method.1 Guangyi Chen 0002, Guangrun Wang, Kun Zhang 0001, Philip Torr 0001, Xiao-Ping Zhang 0002, Yansong Tang |
ICCV | 6 |
| 2023 | Audio-Visual Quality Assessment for User Generated Content: Database and MethodabstractWith the explosive increase of User Generated Content (UGC), UGC video quality assessment (VQA) becomes more and more important for improving users’ Quality of Experience (QoE). However, most existing UGC VQA studies only focus on the visual distortions of videos, ignoring that the user’s QoE also depends on the accompanying audio signals. In this paper, we conduct the first study to address the problem of UGC audio and video quality assessment (AVQA). Specifically, we construct the first UGC AVQA database named the SJTU-UAV database, which includes 520 in-the-wild UGC audio and video (A/V) sequences, and conduct a user study to obtain the mean opinion scores of the A/V sequences. The content of the SJTU-UAV database is then analyzed from both the audio and video aspects to show the database characteristics. We also design a family of AVQA models, which fuse the popular VQA methods and audio features via support vector regressor (SVR). We validate the effectiveness of the proposed models on the three databases. The experimental results show that with the help of audio signals, the VQA models can evaluate the perceptual quality more accurately. The database will be released to facilitate further research. Yuqin Cao, Xiongkuo Min, Wei Sun 0029, Xiao-Ping Zhang 0002, Guangtao Zhai |
ICIP | 4 |
| 2023 | STEV: Stretchable Triboelectric E-skin enabled Proprioceptive Vibration Sensing for Soft RobotabstractVibration perception is essential for robotic sensing and dynamic control. Nevertheless, due to the rigorous demand for sensor conformability and stretchability, enabling soft robots with proprioceptive vibration sensing remains challenging. This paper proposes a novel liquid metal-based stretchable e-skin via a kirigami-inspired design to enable soft robot proprioceptive vibration sensing. The e-skin is fabricated into 0.1mm ultrathin thickness, ensuring its negligible influence on the overall stiffness of the soft robot. Moreover, the working mechanism of the e-skin is based on the ubiquitous triboelectrification effect, which transduces mechanical stimuli without external power supply. To demonstrate the practicability of the e-skin, we built a soft gripper consisting of three soft robotic fingers with proprioceptive vibration sensing. Our experiment shows that the gripper can accurately distinguish the grain category (six grains with the same mass, 99.9% accuracy) and the packaging quality (100% accuracy) by simply shaking the gripped bottle. In summary, a soft robotic proprioceptive vibration sensing solution is proposed; it helps soft robots to have a more comprehensive awareness of their self-state and may inspire further research on soft robotics. Kai-Chong Lei, Huaze Tang, Shoujie Li, Yuan Dai, Wenbo Ding 0001, Xiao-Ping Zhang 0002 |
ICRA | 7 |
| 2023 | Visuotactile Sensor Enabled Pneumatic Device Towards Compliant Oropharyngeal Swab SamplingabstractManual oropharyngeal (OP) swab sampling is an intensive and risky task. In this article, a novel OP swab sampling device of low cost and high compliance is designed by combining the visuotactile sensor and the pneumatic actuator-based gripper. Here, a concave visuotactile sensor called CoTac is first proposed to address the problems of high cost and poor reliability of traditional multi-axis force sensors. Besides, by imitating the doctor's fingers, a soft pneumatic actuator with a rigid skeleton structure is designed, which is demonstrated to be reliable and safe via finite element modeling and experiments. Furthermore, we propose a sampling method that adopts a compliant control algorithm based on the adaptive virtual force to enhance the safety and compliance of the swab sampling process. The effectiveness of the device has been verified through sampling experiments as well as in vivo tests, indicating great application potential. The cost of the device is around 30 US dollars and the total weight of the functional part is less than 0.1 kg, allowing the device to be rapidly deployed on various robotic arms. Shoujie Li, Mingshan He, Wenbo Ding 0001, Linqi Ye, Xueqian Wang 0001, Junbo Tan, Jinqiu Yuan, Xiao-Ping Zhang 0002 |
IROS | 8 |
| 2023 | Autonomous Swarm Robot Coordination via Mean-Field Control Embedding Multi-Agent Reinforcement LearningabstractThe learning approaches of designing a controller to guide the collective behavior of swarm robots have gained significant attention in recent years. However, the scalability of swarm robots and their inherent stochasticity complicate the control problem due to increasing complexity, unpredictability, and non-linearity. Despite considerable progress made in swarm robotics, addressing these challenges remains a significant issue. In this work, we model the stochastic dynamics of a swarm robot system and then propose a novel control framework based on a mean-field control (MFC) embedding multi-agent reinforcement learning (MARL) approach named MF-MARL to deal with these challenges. While MARL is able to deal with stochasticity statistically, we integrate MFC, allowing MF-MARL to cope with large-scale robots. Moreover, we apply statistical moments of robots' state and control action to discretize continuous input and enable MF-MARL to be applied in continuous scenarios. To demonstrate the effectiveness of MF-MARL, we evaluate the performance of the robots on a specific swarm simulation platform. The experimental results show that our algorithm outperforms the traditional algorithms both in navigation and manipulation tasks. Finally, we demonstrate the adaptability of the proposed algorithm through the component failure test. Huaze Tang, Hengxi Zhang, Zhenpeng Shi, Xinlei Chen, Wenbo Ding 0001, Xiao-Ping Zhang 0002 |
IROS | 6 |
| 2023 | Improving sparse graph attention for feature matching by informative keypoints exploration
Xingyu Jiang 0005, Xiao-Ping Zhang 0002, Jiayi Ma 0001 |
Comput. Vis. Image Underst. | 3 |
| 2023 | Increasing the Uniform Degrees of Freedom for Moving q-Dilated ArraysabstractFor$q$-dilated arrays, i.e., dilated arrays with dilation factor$q$, the uniform degrees of freedom (uDOFs) of the synthetic array after array motion can be increased by a factor of$q$compared with that of original linear arrays for$q\leq 3$. However, when$q\geq 4$, the number of uDOFs of the synthetic arrays of$q$-dilated arrays is only three using existing moving array models. To achieve a high number of uDOFs for$q$-dilated arrays with$q\geq 4$, in this article, we design a new moving array processing model in$q$-dilated arrays, which synthesizes multiple shifted arrays with displacements of half wavelength multiples. The number of shifted arrays in the proposed model is shown to be a function of the dilation factor$q$. First, we prove that the number of uDOFs of synthetic arrays of$q$-dilated arrays can be$q $times that of their original arrays for arbitrary positive integer$q$. Hence, the maximum number of detectable sources for direction-of-arrival (DOA) estimation is increased by a factor of$q $. Second, we apply the new model to two-parallel$q$-dilated arrays to estimate the 2-D DOAs, increasing the number of identifiable sources by a factor of$q$for 2-D DOA estimation. Numerical examples show the superiority of the proposed model based on$q $-dilated arrays. Shuang Li 0005, Xiao-Ping Zhang 0002, Ding-Liang Ruan |
IEEE Internet Things J. | 2 |
| 2023 | Dual Vision TransformerabstractRecent advances have presented several strategies to mitigate the computations of self-attention mechanism with high-resolution inputs. Many of these works consider decomposing the global self-attention procedure over image patches into regional and local feature extraction procedures that each incurs a smaller computational complexity. Despite good efficiency, these approaches seldom explore the holistic interactions among all patches, and are thus difficult to fully capture the global semantics. In this paper, we propose a novel Transformer architecture that elegantly exploits the global semantics for self-attention learning, namely Dual Vision Transformer (Dual-ViT). The new architecture incorporates a critical semantic pathway that can more efficiently compress token vectors into global semantics with reduced order of complexity. Such compressed global semantics then serve as useful prior information in learning finer local pixel level details, through another constructed pixel pathway. The semantic pathway and pixel pathway are integrated together and are jointly trained, spreading the enhanced self-attention information in parallel through both of the pathways. Dual-ViT is henceforth able to capitalize on global semantics to boost self-attention learning without compromising much computational complexity. We empirically demonstrate that Dual-ViT provides superior accuracy than SOTA Transformer architectures with comparable training complexity. Ting Yao 0003, Yehao Li, Yingwei Pan, Yu Wang 0060, Xiao-Ping Zhang 0002, Tao Mei 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2023 | Refine-Net: Normal Refinement Neural Network for Noisy Point CloudsabstractPoint normal, as an intrinsic geometric property of 3D objects, not only serves conventional geometric tasks such as surface consolidation and reconstruction, but also facilitates cutting-edge learning-based techniques for shape analysis and generation. In this paper, we propose a normal refinement network, called Refine-Net, to predict accurate normals for noisy point clouds. Traditional normal estimation wisdom heavily depends on priors such as surface shapes or noise distributions, while learning-based solutions settle for single types of hand-crafted features. Differently, our network is designed to refine the initial normal of each point by extracting additional information from multiple feature representations. To this end, several feature modules are developed and incorporated into Refine-Net by a novel connection module. Besides the overall network architecture of Refine-Net, we propose a new multi-scale fitting patch selection scheme for the initial normal estimation, by absorbing geometry domain knowledge. Also, Refine-Net is a generic normal estimation framework: 1) point normals obtained from other methods can be further refined, and 2) any feature module related to the surface geometric structures can be potentially integrated into the framework. Qualitative and quantitative evaluations demonstrate the clear superiority of Refine-Net over the state-of-the-arts on both synthetic and real-scanned datasets. Honghua Chen, Yingkui Zhang, Mingqiang Wei, Haoran Xie 0001, Jun Wang 0039, Tong Lu 0002, Harry Qin, Xiao-Ping Zhang 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 9 |
| 2023 | Effective Online Portfolio Selection for the Long-Short Market Using Mirror Gradient DescentabstractOnline portfolio selection has been actively studied to maximise overall returns by selecting the optimal portfolio weights using online algorithms. However, most work has focused on long-only portfolios, and developing efficient algorithms with loose portfolio constraints remains a challenge. In this letter, the classical online portfolio selection problem is reformulated to allow long/short and margin. For this problem, conventional gradient-based online algorithms face the challenges of high regret and computational complexity due to non-optimal gradients and high-dimensional projections. To tackle this, we propose a novel online algorithm that introduces mirror descent to achieve dimension-free regret in a non-Euclidean space. Specifically, a Bregman divergence is introduced to replace the$\ell _{2}$norm as a valid proximal setup for the problem to achieve uniform gradients and reduce projection computations. Furthermore, a smoothing technique is developed to reduce the variance of the gradients. The evaluation shows that our algorithm achieves low regret bound and computational complexity, which guarantees a 30% advantage over other strategies in Chinese futures market. Xiangming Li 0001, Yunzhu Chen, Neng Ye, Xiao-Ping Zhang 0002 |
IEEE Signal Process. Lett. | 5 |
| 2023 | Image Quality Score Distribution Prediction via Alpha Stable ModelabstractBased on potentially subjective and diverse image quality scores given by a group of subjects, we propose to predict the distribution of image quality scores rather than the mean opinion score (MOS) of image quality. Therefore, in this paper, we use an alpha stable model to parameterize the image quality score distribution (IQSD), and propose an objective method to predict the alpha-stable-model-based IQSD. First, the LIVE database is re-recorded. Specifically, we invite a large group of subjects (187 valid subjects) to evaluate the quality of all 808 images in the LIVE database, with their scores forming reliable IQSDs. All images in the LIVE database and their collected subjective quality scores form a new image quality assessment database, named the SJTU IQSD database. We then propose a framework and algorithm to predict the alpha-stable-model-based IQSD, in which quality features are extracted from the structural and natural statistical information of each image, and support vector regressors are trained to predict the alpha stable model parameters. Experiments carried out on the SJTU IQSD database verify the feasibility of using the alpha stable model to describe the IQSD, and the experimental results show that the alpha-stable-model-based IQSD can reflect a large amount of subjective information on image quality. We also prove that the objective alpha-stable-model-based IQSD prediction method is effective. The code and the SJTU IQSD database can be downloaded at ‘https://github.com/YixuanGao98/Image-Quality-Score-Distribution-Prediction-via-Alpha-Stable-Model.git’. Xiongkuo Min, Wenhan Zhu, Xiao-Ping Zhang 0002, Guangtao Zhai |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | Pi-ViMo: Physiology-inspired Robust Vital Sign Monitoring using mmWave RadarsabstractContinuous monitoring of human vital signs using non-contact mmWave radars is attractive due to their ability to penetrate garments and operate under different lighting conditions. Unfortunately, most prior research requires subjects to stay at a fixed distance from radar sensors and to remain still during monitoring. These restrictions limit the applications of radar vital sign monitoring in real life scenarios. In this article, we address these limitations and present Pi-ViMo, a non-contact P hysiology- i nspired Robust Vi tal Sign Mo nitoring system, using mmWave radars. We first derive a multi-scattering point model for the human body, and introduce a coherent combining of multiple scatterings to enhance the quality of estimated chest-wall movements. It enables vital sign estimations of subjects at any location in a radar’s field of view (FoV). We then propose a template matching method to extract human vital signs by adopting physical models of respiration and cardiac activities. The proposed method is capable to separate respiration and heartbeat in the presence of micro-level random body movements (RBM) when a subject is at any location within the field of view of a radar. Experiments in a radar testbed show average respiration rate errors of 6% and heart rate errors of 11.9% for the stationary subjects, and average errors of 13.5% for respiration rate and 13.6% for heart rate for subjects under different RBMs. Boyu Jiang, Rong Zheng 0001, Xiao-Ping Zhang 0002, Jun Li 0067, Qiang Xu 0005 |
ACM Trans. Internet Things | 4 |
| 2023 | LineDL: Processing Images Line-by-Line With Deep LearningabstractAlthough deep learning-based (DL-based) image processing algorithms have achieved superior performance, they are still difficult to apply on mobile devices (e.g., smartphones and cameras) due to the following reasons: 1) the high memory demand and 2) large model size. To adapt DL-based methods to mobile devices, motivated by the characteristics of image signal processors (ISPs), we propose a novel algorithm named LineDL. In LineDL, the default mode of the whole-image processing is reformulated as a line-by-line mode, eliminating the need to store large amounts of intermediate data for the whole image. An information transmission module (ITM) is designed to extract and convey the interline correlation and integrate the interline features. Furthermore, we develop a model compression method to reduce the model size while maintaining competitive performance; that is, knowledge is redefined, and compression is performed in two directions. We evaluate LineDL on general image processing tasks, including denoising and superresolution. The extensive experimental results demonstrate that LineDL achieves image quality comparable to that of state-of-the-art (SOTA) DL-based algorithms with a much smaller memory demand and competitive model size. Wenshu Chen, Liyuan Peng, Yuhao Liu 0001, Mingyu Wang 0001, Xiao-Ping Zhang 0002, Xiaoyang Zeng |
IEEE Trans. Image Process. | 6 |
| 2023 | Blind Image Quality Assessment for Pathological Microscopic Image Under Screen and Immersion ScenariosabstractThe high-quality pathological microscopic images are essential for physicians or pathologists to make a correct diagnosis. Image quality assessment (IQA) can quantify the visual distortion degree of images and guide the imaging system to improve image quality, thus raising the quality of pathological microscopic images. Current IQA methods are not ideal for pathological microscopy images due to their specificity. In this paper, we present deep learning-based blind image quality assessment model with saliency block and patch block for pathological microscopic images. The saliency block and patch block can handle the local and global distortions, respectively. To better capture the area of interest of pathologists when viewing pathological images, the saliency block is fine-tuned by eye movement data of pathologists. The patch block can capture lots of global information strongly related to image quality via the interaction between different image patches from different positions. The performance of the developed model is validated by the home-made Pathological Microscopic Image Quality Database under Screen and Immersion Scenarios (PMIQD-SIS) and cross-validated by the five public datasets. The results of ablation experiments demonstrate the contribution of the added blocks. The dataset and the corresponding code are publicly available at: https://github.com/mikugyf/PMIQD-SIS. Yifei Guo, Menghan Hu, Xiongkuo Min, Yan Wang 0036, Guangtao Zhai, Xiao-Ping Zhang 0002, Xiaokang Yang 0001 |
IEEE Trans. Medical Imaging | 7 |
| 2023 | Angel's Girl for Blind Painters: An Efficient Painting Navigation System Validated by Multimodal Evaluation ApproachabstractFor people who ardently love painting but unfortunately have visual impairments, holding a paintbrush to create a work is a very difficult task. People in this special group are eager to pick up the paintbrush, like Leonardo da Vinci, to create and make full use of their own talents. Therefore, to maximally bridge this gap, we propose a painting navigation system called “Angle’s Eyes” to assist blind people in artistic creation. The proposed system is composed of cognitive system and guidance system. The system adopts drawing board positioning based on QR code, brush navigation based on target detection and bush real-time positioning. Meanwhile, we design a simple yet efficient position information coding rule to remind the user of the current brush tip position. In addition, we design a criterion to efficiently judge whether the brush reaches the target or not. The numerous experiments are conducted to optimize and test the performance of the system. The results of real-world scenario experiments demonstrate that the developed system has great potential to help blind people with painting. This work also demonstrates that it is practicable for the blind people to feel the world through the brush in their hands. In the future, we plan to deploy “Angle’s Eyes” on the phone to make it more portable. The demo video of the proposed painting navigation system is available athttps://doi.org/10.6084/m9.figshare.9760004.v1. Menghan Hu, Qingli Li, Guangtao Zhai, Simon X. Yang, Xiao-Ping Zhang 0002, Xiaokang Yang 0001 |
IEEE Trans. Multim. | 7 |
| 2023 | Focal Stack Image Compression Based on Basis-Quadtree RepresentationabstractIn this paper, we propose an efficient compression scheme for focal stack images (FoSIs) based on a new basis-quadtree representation. In the new basis-quadtree representation, FoSIs are initially reorganized as co-located block groups in the depth dimension. In each group, selective basis blocks and adaptive quadtree partition are optimized to predict the focused or defocused co-located blocks by intra-group approximation. By solving a joint optimization problem, FoSIs can be efficiently represented by the optimal basis blocks, corresponding quadtree partition and approximation parameters, which will be compressed separately. Then, these basis blocks are stitched into several new frames (basis frames) according to their original locations and partition modes. Basis frames are compressed by our designed encoder, where the intra-group approximation is embedded into the high efficiency video coding (HEVC) encoder. Thus, the redundancies of basis blocks can be further eliminated. Finally, the approximation parameters are refined to suppress the amplified errors caused by introduced compression blur after basis frame coding. The refined parameters are compressed losslessly and multiplexed with the bitstream of the basis frames to ensure the reconstruction quality of FoSIs. Experiments on 12 test sequences demonstrate that the proposed scheme can obtain higher coding performance than the state-of-the-art comparison schemes. Specifically, the proposed scheme achieves up to 5.23 dB PSNR gains and 71.59% bitrate savings over the HEVC baseline scheme on sequences I03 and I05, respectively. Kejun Wu, You Yang 0002, Qiong Liu 0001, Xiao-Ping Zhang 0002 |
IEEE Trans. Multim. | 4 |
| 2023 | Visual-Tactile Fusion for Transparent Object Grasping in Complex BackgroundsabstractThe grasping of transparent objects is challenging but of significance to robots. In this article, a visual–tactile fusion framework for transparent object grasping in complex backgrounds is proposed, which synergizes the advantages of vision and touch, and greatly improves the grasping efficiency of transparent objects. First, we propose a multiscene synthetic grasping dataset named SimTrans12 K together with a Gaussian-mask annotation method. Next, based on the TaTa gripper, we propose a grasping network named transparent object-grasping convolutional neural network for grasping position detection, which shows good performance in both synthetic and real scenes. Inspired by human grasping, a tactile calibration method and a visual–tactile fusion classification method are designed, which improve the grasping success rate by 36.7% compared with direct grasping and the classification accuracy by 39.1%. Furthermore, the tactile height sensing module and the tactile position exploration module are added to solve the problem of grasping transparent objects in irregular and visually undetectable scenes. The experimental results demonstrate the validity of the framework. Shoujie Li, Haixin Yu, Wenbo Ding 0001, Houde Liu, Linqi Ye, Chongkun Xia, Xueqian Wang 0001, Xiao-Ping Zhang 0002 |
IEEE Trans. Robotics | 8 |
| 2022 | Segmented Learning for Class-of-Service Network Traffic ClassificationabstractClass-of-service (CoS) network traffic classification (NTC) classifies a group of similar traffic applications. The CoS classification is advantageous in resource scheduling for Internet service providers and avoids the necessity of remodelling. Our goal is to find a robust, lightweight, and fast-converging CoS classifier that uses fewer data in modelling and does not require specialized tools in feature extraction. The commonality of statistical features among the network flow segments motivates us to propose novel segmented learning that includes essential vector representation and a simple-segment method of classification. We represent the segmented traffic in the vector form using the essential vector representation (EVR). Then, the segmented traffic is modelled for classification using random forest based simple-segment method of classification (S2MC). Our solution's success relies on finding the optimal segment size and a minimum number of segments required in modelling. The solution is validated on multiple datasets for various CoS services, including virtual reality (VR). Significant findings of the research work are i) Synchronous services that require acknowledgment and request to continue communication are classified with 99 % accuracy, ii) Initial 1,000 packets in any session are good enough to model a CoS traffic for promising results, and we therefore can quickly deploy a CoS classifier, and iii) Test results remain consistent even when trained on one dataset and tested on a different dataset. In summary, our solution is the first to propose segmentation learning NTC that uses fewer features to classify most CoS traffic with an accuracy of 99 %. The implementation of our solution is available on GitHub. Yoga Suhas Kuruba Manjunath, Sihao Zhao, Hatem Abou-Zeid, Akram Bin Sediq, Ramy Atawia, Xiao-Ping Zhang 0002 |
GLOBECOM | 6 |
| 2022 | Terahertz Image Restoration Benchmarking DatasetabstractThe paper introduces a new terahertz (THz) image benchmarking dataset for THz imaging. The degradation of THz image quality is one of the main problems caused by system noise, intrinsic long-wavelength, and diffraction phenomena. In this paper, the point spread function of the THz imaging process is reconstructed firstly. The THz datasets with ground-truth and degraded images are then synthesized using the point spread function (PSF). We propose a Dense Instantiation Normalization Block (DIN Block) to reconstruct clean THz images. Based on the DIN block, a powerful multi-stage network is designed, named as DINet. DINet achieves the state-of-the-art (SOTA) restoration performance on image rain removal datasets and the proposed THz datasets. To the best of our knowledge, the THz image benchmarking dataset is the first public dataset, which is available at https://github.com/hellogry/THzDatasets Yixiong Zhang, Zhipeng Su, Jianyang Zhou 0002, Xiao-Ping Zhang 0002 |
ICASSP | 5 |
| 2022 | How Sound Affects Visual Attention in Omnidirectional VideosabstractIn this paper, we propose a new audio-visual attention dataset that records eye movement for omnidirectional videos with and without sound. We classify the videos into three types according to the number of salient objects and sound sources and analyze the impact of sound on visual attention distribution and inter-observer consistency of viewing area in different types of videos. From the quantitative and qualitative analysis, we find that visual attention will be drawn to and concentrated on the sound source with the presence of sound, especially when there are several visually salient objects and only one sound source. Also, the sound will enhance the consistency of observation areas among viewers to some extent. For more investigations on the impact of sound on visual attention and prospective audio-visual saliency model, we still need further study. Guangtao Zhai, Yucheng Zhu, Jun Zhou 0007, Xiao-Ping Zhang 0002 |
ICIP | 5 |
| 2022 | Iterative Water-Filling Power and Subcarrier Allocation for Multicarrier Non-Orthogonal Multiple Access UplinkabstractNovel closed form formulae of iterative optimal power control and allocation, and a criterion for optimal subcarrier allocation, are derived for uplink of multicarrier non-orthogonal multiple access (MC-NOMA) systems. Our new formulae provide intuitive insights on how to optimally allocate subcarriers and power to MC-NOMA (and as a special case, MC-OMA) users. In particular, the novel water-filling power allocation formulae provide a mean to iteratively compute optimal power allocation for a general NOMA cluster size. Convergence of the iterative algorithm and sum rate performance of MC-NOMA system are presented and compared with MC-OMA system. Chin Choy Chai, Xiao-Ping Zhang 0002 |
ISIT | 2 |
| 2022 | Image Quality Assessment: From Mean Opinion Score to Opinion Score DistributionabstractRecently, many methods have been proposed to predict the image quality which is generally described by the mean opinion score (MOS) of all subjective ratings given to an image. However, few efforts focus on predicting the opinion score distribution of the image quality ratings. In fact, the opinion score distribution reflecting subjective diversity, uncertainty, etc., can provide more subjective information about the image quality than a single MOS, which is worthy of in-depth study. In this paper, we propose a convolutional neural network based on fuzzy theory to predict the opinion score distribution of image quality. The proposed method consists of three main steps: feature extraction, feature fuzzification and fuzzy transfer. Specifically, we first use the pre-trained VGG16 without fully-connected layers to extract image features. Then, the extracted features are fuzzified by fuzzy theory, which is used to model epistemic uncertainty in the process of feature extraction. Finally, a fuzzy transfer network is used to predict the opinion score distribution of image quality by learning the mapping from epistemic uncertainty to the uncertainty existing in the image quality ratings. In addition, a new loss function is designed based on the subjective uncertainty of the opinion score distribution. Extensive experimental results prove the superior prediction performance of our proposed method. Xiongkuo Min, Yucheng Zhu, Jing Li 0026, Xiao-Ping Zhang 0002, Guangtao Zhai |
ACM Multimedia | 5 |
| 2022 | Heterogeneous Mean-Field Multi-Agent Reinforcement Learning for Communication Routing Selection in SAGI-NetabstractThe utilization of heterogeneous end devices such as the low earth orbit (LEO) satellite, unmanned aerial vehicles (UAVs) and ground users (GUs) deployed at different altitudes, known as the space-air-ground integrated network (SAGI-Net), can be quite promising towards a bunch of advanced applications. Whereas, the energy efficiency of the SAGI-Net communication system is a key criterion needed to be improved urgently in consideration that the inappropriate communication routing will undoubtedly cause a huge communication energy cost of the system especially with a large number of communication devices inside. In this paper, we proposed a novel communication routing selection model for the SAGI-Net system and established a heterogeneous multi-agent reinforcement learning (HMF-MARL) framework to optimize the communication energy efficiency of this system, where the mean-field theory was introduced to enhance the ability of classic MARL method while still maintaining a relatively low computational complexity. The experiment results show that the capacity of the heterogeneous multi-agent system has been improved by nearly 80% using the proposed HMF-MARL method compared with the classic MARL one, which hopefully shows the potential value on the implementation of the SAGI-Net system in the future. Hengxi Zhang, Huaze Tang, Yuanquan Hu, Xiaoli Wei, Chenye Wu, Wenbo Ding 0001, Xiao-Ping Zhang 0002 |
VTC Fall | 7 |
| 2022 | Non-local channel aggregation network for single image rain removal
Zhipeng Su, Yixiong Zhang, Xiao-Ping Zhang 0002 |
Neurocomputing | 3 |
| 2022 | Environmental Sound Classification via Time-Frequency Attention and Framewise Self-Attention-Based Deep Neural NetworksabstractEnvironmental sound classification (ESC) is crucial to understanding the surroundings in Internet of Things (IoT) applications. The state-of-the-art deep learning approaches do not have good ESC performance when there exists various clutter interference, which is common in IoT scenarios. In this article, we present a novel deep neural network framework based on time–frequency attention and framewise self-attention (TFFS-DNN). It consists of two major novel architectures: 1) gradient and 2) latent feature-based DNN to generate our time–frequency attention, which can locate the relevant time–frequency (i.e., spectral) features accurately, and self-attention normalization DNN to generate our framewise self-attentions which properly indicate the relevance of frames. By conjoining these two sorts of distinct and complementary attentions with spectrograms, we are able to identify the importance or relevance in terms of time, frequency, and frame of the sounds using TFFS-DNN, which helps in distinguishing clutter such as background as well as model interpretation to some extent. Thus, the proposed TFFS-DNN can classify environmental sounds with clutter. The evaluation using four real-world environmental sound data sets demonstrates the superior performance of the proposed framework over several state-of-the-art schemes. Notably, we achieve 79.23% classification accuracy on theUrbanSounddata set, a raw environmental sound data set that is full of clutter. The ablation study demonstrates a relative 3%–9% improvement of classification accuracy by the proposed framework over the baseline deep model. Bo Wu 0024, Xiao-Ping Zhang 0002 |
IEEE Internet Things J. | 2 |
| 2022 | Graph-Based Denoising for Respiration and Heart Rate Estimation During Sleep in Thermal VideoabstractQuality sleep is a basic human need for well-being, yet sleep deprivation has been a long-term global problem. A common type of sleep deprivation is obstrucive sleep apnea, where people repeatedly stop breathing during sleep with subsequent abnormal vital signs, namely, respiration rate and heart rate. While tremendous effort has been made for vital signs monitoring systems during sleep, existing works still lack portability for bulky and intrusive systems and reliability for consumer-level, nonintrusive systems. To bridge the gap between practicability and accuracy and facilitate Internet of Things for smart healthcare, in this article, we propose a vital signs estimation system during sleep via a thermal camera. The system first captures thermal image sequences of a sleeping subject and then processes the facial regions within the thermal images for vital signs signal extraction. Specifically, leveraging on the inherent graph structure among subregions of the facial area, we propose a graph-based, spatial–temporal signal denoising scheme. Experimental results show that the graph-based denoising scheme in our system effectively reduces the noise level introduced by cameras and subjects, and our proposed system outperforms state-of-the-art nonintrusive vital signs monitoring systems. Since the algorithm components in our system have relatively low time complexity and no model training is required, our system can be deployed efficiently at the edge devices in a smart home setting. The extracted vital signs can then be used for sleep abnormality detection and disease screening. Cheng Yang 0003, Menghan Hu, Guangtao Zhai, Xiao-Ping Zhang 0002 |
IEEE Internet Things J. | 4 |
| 2022 | Wideband Multitarget Tracking Based on Dynamic Bayesian Network Learning in an Acoustic Sensor Array NetworkabstractThe multitarget tracking (MTT) based on distributed fusion methods in an acoustic sensor array network (ASAN) is limited by the performance of measurement parameter estimators, such as the received signal strength (RSS) and the direction of arrival (DOA). For measurement parameters with low accuracy and resolution, the MTT may fail in subsequent steps, e.g., data association, because the loss of upstream information cannot be made up by downstream processing. Thus, we propose a new wideband MTT algorithm based on the dynamic Bayesian network (DBN), which treats the ASAN as an overall extended array, and directly estimates the target states from the raw acoustic data. The DBN fuses the near-field model, the acoustic propagation model. and the motion model. These submodels can be optimized by each other, improving the final estimation. Also, for each subband, target signals and the precision parameters of the sensor noise are treated as hidden random variables. Based on this, the weight of each subband can be automatically adjusted according to the accurate hidden variables. Besides, the optimization problem of the posterior probability is transformed into a graphical model learning problem. Moreover, for nonconjugate models, a novel algorithm based on Laplace approximations (LAs) with Newton’s method (NM) is developed, i.e., DBN-LA-NM, bypassing data association. In addition, the corresponding Cramer–Rao lower bound and convergence conditions are derived. The numerical simulation results show that the proposed algorithm outperforms existing MTTs based on the near-field model in terms of accuracy, convergence, and computational complexity. Field experiments further verify the feasibility of the proposed algorithm. Wenqiong Zhang, Jianfei Tong, Ming Bao, Xiao-Ping Zhang 0002, Xiaodong Li 0002 |
IEEE Internet Things J. | 4 |
| 2022 | Multitarget Tracking Based on Dynamic Bayesian Network With Reparameterized Approximate Variational InferenceabstractMultitarget tracking (MTT) is an important component of situation-awareness based on the Internet of Things (IoT). Existing algorithms mainly focus on tracking based on conventional measurements, e.g., bearings or ranges. However, measurement parameter estimations are considered in isolation, limiting the accuracy and resolution of MTT, and the related data association is an NP-hard multidimensional assignment problem. In this article, we develop a new one-step MTT algorithm based on a novel dynamic Bayesian network (DBN), i.e., DBNMTT. The new MTT algorithm directly infers target states from the raw measurement data by fusing the array signal model, the signal propagation model, and the motion model. In this new DBNMTT framework, we treat target states and conventional measurements, such as bearings and target energies as hidden random variables. The posterior joint probability optimization problem is translated into the problem of graphical model learning. In this way, we can improve the accuracy and resolution of MTT and convert the NP-hard data association problem to a hidden variable learning problem. For nonconjugate models in the DBNMTT, we develop a novel reparameterized approximation variational inference (ReAVI) approach to solve the learning problem. The ReAVI converts nonconjugate models to conjugate models with new parameters and reuses the mean-field algorithm. The performance of our proposed new MTT method, namely, DBNMTT based on ReAVI (DBNMTT-ReAVI), is analyzed on extensive simulations in challenging scenarios. The simulation results show that the DBNMTT-ReAVI algorithm is superior to conventional measurement-based MTT algorithms in several aspects, including the success probability, convergence, resolution, and accuracy. Wenqiong Zhang, Jun Zhang 0018, Ming Bao, Xiao-Ping Zhang 0002, Xiaodong Li 0002 |
IEEE Internet Things J. | 4 |
| 2022 | Sequential Doppler-Shift-Based Optimal Localization and Synchronization With TOAabstractDoppler shift is an important measurement for localization and synchronization (LAS), and is available in various practical systems. Existing studies on LAS techniques in a time-division broadcast LAS system (TDBS) only use sequential time-of-arrival (TOA) measurements from the broadcast signals. In this article, we develop a new optimal LAS method in the TDBS, namely, LAS-SDT, by taking advantage of the sequential Doppler shift and TOA measurements. It achieves higher accuracy compared with the conventional TOA-only method for user devices (UDs) with motion and clock drift. Another two variant methods, LAS-SDT-v for the case with UD velocity aiding and LAS-SDT-k for the case with UD clock drift aiding, are developed. We derive the Cramér–Rao lower bound (CRLB) for these different cases. We show analytically that the accuracies of the estimated UD position, clock offset, velocity, and clock drift are all significantly higher than those of the conventional LAS method using TOAs only. Numerical results corroborate the theoretical analysis and show the optimal estimation performance of the LAS-SDT. Sihao Zhao, Ningyan Guo, Xiao-Ping Zhang 0002, Xiaowei Cui, Mingquan Lu |
IEEE Internet Things J. | 3 |
| 2022 | Multiple-Target Localization by Millimeter-Wave Radars With Trapezoid Virtual Antenna ArraysabstractWe consider the problem of localizing multiple targets by millimeter wave (mmWave) radars with irregular antenna placement, i.e., trapezoid virtual antenna array. The goal is to estimate both the number of targets and their 3-D locations. While many well-known algorithms have been developed for either problems, they still suffer from several limitations, such as the need for a large amount of sampled radar data and high computation complexity. In this work, we develop an efficient solution by exploring the received signal structure in two steps: 1) estimating the number of targets and their ranges by extending Barone’s method to handle data from multiple antennas and 2) estimating the angle of arrival of each target by a Least-Square algorithm optimization. The proposed algorithm has been evaluated through Monte-Carlo simulations and an indoor testbed. By comparing with baseline algorithms, including 2D-FFT and multiple signal classification (MUSIC), we find that the proposed algorithm has the best performance in high signal-to-noise ratio regimes. Wei Zhao 0061, Jian-Kang Zhang 0002, Xiao-Ping Zhang 0002, Rong Zheng 0001 |
IEEE Internet Things J. | 3 |
| 2022 | Robust STAP Detection Based on Volume Cross-Correlation Function in Heterogeneous EnvironmentsabstractThe performance of moving target detection in heterogeneous environments with the traditional space-time adaptive processing (STAP) may degrade when the real clutter environments deviate from the prior assumption on the clutter distribution. In this letter, a new detector for STAP applications based on volume cross-correlation function (VCF), namely VCF-STAP, is proposed to achieve robust performance of moving target detection in heterogeneous environments. In the new VCF-STAP, the VCF is used to form a distance measure between the sample signal subspace and the target subspace without modeling the clutter distribution. Then, a new robust STAP detection statistic is constructed using this distance measure. Simulation and experimental results show that the proposed VCF-STAP achieves robust performance of moving target detection in heterogeneous environments, especially it achieves much superior detection performance compared with existing STAP methods when the real clutter environments do not satisfy their prior assumptions. Besides, it is also shown that VCF-STAP has the constant false alarm rate (CFAR) property. Zhizhuo Jiang, You He 0002, Gang Li 0008, Xiao-Ping Zhang 0002 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2022 | Dilated projection correction network based on autoencoder for hyperspectral image super-resolution
Xinya Wang, Jiayi Ma 0001, Junjun Jiang, Xiao-Ping Zhang 0002 |
Neural Networks | 4 |
| 2022 | Robust image matching via local graph structure consensus
Xingyu Jiang 0005, Xiao-Ping Zhang 0002, Jiayi Ma 0001 |
Pattern Recognit. | 3 |
| 2022 | Closed-Form Two-Way TOA Localization and Synchronization for User Devices With Motion and Clock DriftabstractA two-way time-of-arrival (TOA) system is composed of anchor nodes (ANs) and user devices (UDs). Two-way TOA measurements between AN-UD pairs are obtained via round-trip communications to achieve localization and synchronization (LAS) for a UD. Existing LAS method for a moving UD with clock drift adopts an iterative algorithm, which requires accurate initialization and has high computational complexity. In this letter, we propose a new closed-form two-way TOA LAS approach, namely CFTWLAS, which does not require initialization, has low complexity and empirically achieves optimal LAS accuracy. We first linearize the LAS problem by squaring and differencing the two-way TOA equations. We employ two auxiliary variables to simplify the problem to finding the analytical solution of quadratic equations. Due to the measurement noise, we can only obtain a raw LAS estimation from the solution of the auxiliary variables. Then, a weighted least squares step is applied to further refine the raw estimation. We analyze the theoretical error of the new CFTWLAS and show that it empirically reaches the Cramér-Rao lower bound (CRLB) with sufficient ANs under the condition of proper geometry and small noise. Numerical results in a 3D scenario verify the theoretical analysis that the estimation accuracy of the new CFTWLAS method reaches CRLB in the presented experiments when the number of ANs is large, the geometry is appropriate, and the noise is small. Unlike the iterative method whose complexity increases with the iteration count, the new CFTWLAS has constant low complexity. Sihao Zhao, Ningyan Guo, Xiao-Ping Zhang 0002, Xiaowei Cui, Mingquan Lu |
IEEE Signal Process. Lett. | 3 |
| 2022 | Gaussian-Wiener Representation and Hierarchical Coding Scheme for Focal Stack ImagesabstractFocal stack images (FoSIs) are a set of 2D images that captured one scene with serial focal depths. The redundancies of FoSIs mainly come from gradual focused depths changes rather than motion of objects. Conventional coding schemes cannot fully exploit such redundancies, leading to coding inefficiency. In this paper, we propose a new Gaussian-Wiener representation to model the gradual focused depths changes among FoSIs. In the representation, image degradation-restoration relations are utilized to describe the focus-defocus changing characteristics of FoSIs. Based on this representation, we propose a new hierarchical coding scheme for fully exploiting the inter-frame redundancies of FoSIs. In the scheme, a Gaussian-Wiener representation based inter prediction (GWR-IP) is presented by embedding Gaussian convolution and Wiener deconvolution into normal video encoder. Block-wise focus-defocus changing of FoSIs can be predicted in bi-directional manner by solving optimization problem. For higher coding efficiency, a Gaussian-Wiener representation based hierarchical prediction structure (GWR-HPS) is also designed and applied in the coding scheme. The proposed coding scheme is performed on 10 test sequences, including 5 synthetic scenes and 5 realistic scenes. Experimental results show that proposed coding scheme can obtain 2.640 dB PSNR gains and 51.830% bitrate savings on average of all test sequences in Low Delay P configuration, 2.123 dB PSNR gains and 43.975% bitrate savings in Low Delay B configuration, and 1.044 dB PSNR gains and 26.078% bitrate savings in Random Access configuration. Particularly, it achieves up to 65.544% bit rate savings and 3.901 dB PSNR increments for test sequence I09 in Low Delay P configuration. Furthermore, ablation test demonstrates that Gaussian representation contributes more on coding performance than Wiener representation and GWR-HPS. Kejun Wu, You Yang 0002, Qiong Liu 0001, Xiao-Ping Zhang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | Sequential Gesture Learning for Continuous Labanotation Generation Based on the Fusion of Graph Neural NetworksabstractLabanotation is a symbolic recording system for human movements, and also a powerful tool for protecting and spreading folk dances and other performing arts. State-of-the-art automatic Labanotation uses end-to-end methods with sequence-based skeleton representation, which cannot capture the relationship between joints and bones in the skeleton for accurate descriptions of continuous lower limb movements such as dance steps. In this paper, we propose a novel double-stream fusion method of directed graph neural networks (DGNN), combined with connectionist temporal classification (CTC), namely DFGNN-CTC, for sequential fine-grained motion recognition, such as the Labanotation generation of unsegmented dance movement. First, we extract double-stream directed graph feature, employing an orientation-normalized directed acyclic graph (ON-DAG) and an orientation-normalized temporal directed acyclic graph (ON-TDAG), to jointly model spatiotemporal properties of movement recorded in motion capture data. Then, we design a CTC-based fusion-pooling module to fuse the spatial and temporal streams encoded by two DGNNs. It concatenates and fuses the two streams to generate discriminative descriptions of each time step, and concentrates them to make per-time-step predictions of Laban gesture type, from which the CTC searches the optimal Laban symbol sequence, corresponding to elemental motions composing the movement. In this way, the new method enables much finer discrimination for similar Laban gestures with subtle differences in spatial and temporal properties through joint contextual spatiotemporal modeling so that it achieves much superior performance in continuous Labanotation generation to existing methods, which only have single-stream analysis either spatially or temporally. The experiments on two Labanotation-labelled motion capture datasets demonstrate the effectiveness of the components in the proposed method and its superiority comparing with the state-of-the-art methods, especially for lower limb movements. Ningwei Xie, Zhenjiang Miao, Xiao-Ping Zhang 0002, Wanru Xu, Min Li 0025, Jiaji Wang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | D2TNet: A ConvLSTM Network With Dual-Direction Transfer for Pan-SharpeningabstractIn this article, we propose an efficient convolutional long short-term memory (ConvLSTM) network with dual-direction transfer for pan-sharpening, termed D2TNet. We design a specially structured ConvLSTM network that allows for dual-directional communication, including multiscale information and multilevel information. On the one hand, due to the sensitivity of spatial information to scales and the sensitivity of spectral information to levels, multiscale and multilevel information is extracted to facilitate the fuller use of source images. On the other hand, ConvLSTM is employed to capture the strong dependencies between multiscale information and multilevel information. Besides, we introduce a multiscale loss to enable different scales contributing to each other to generate high-resolution multispectral images that are closer to the ground truth. Extensive experiments, including qualitative evaluation, quantitative evaluation, and efficiency comparison, are implemented to verify that our D2TNet outperforms state-of-the-art methods indeed. Meiqi Gong, Jiayi Ma 0001, Han Xu 0001, Xin Tian 0006, Xiao-Ping Zhang 0002 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | 1-Bit Radar Imaging Based on Adversarial SamplesabstractRadar imaging with 1-bit data is attractive thanks to its low storage and transmission burden. Existing 1-bit radar imaging methods cannot satisfactorily suppress the artifacts in the imaging result induced by 1-bit quantization error and noise. In this article, we propose a new 1-bit compressive sensing (CS) based algorithm, i.e., the adversarial-sample-based binary iterative hard thresholding (AS-BIHT) algorithm, to improve the 1-bit radar imaging performance. First, we formulate a parametric model for 1-bit radar imaging with a new adjustable quantization level parameter. The parametric 1-bit radar imaging model updates the imaging scene and the quantization level parameter in an iterative fashion based on adversarial samples. Then, we design a mechanism to generate adversarial samples by attacking the 1-bit radar imaging model to resist the quantization consistency condition, such that forcing quantization consistent reconstruction on adversarial samples mitigates the quantization error and noise. The quantization level parameter is then tuned based on the adversarial samples. In this way, the ability of the model to adapt to echo data contaminated by noise and quantization error is enhanced, and the artifacts are well suppressed. Simulation and experimental results on real radar data demonstrate the effectiveness of the proposed AS-BIHT algorithm in 1-bit radar imaging. Jianghong Han, Gang Li 0008, Meiya Duan, Xiao-Ping Zhang 0002 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | A Semisupervised Siamese Network for Efficient Change Detection in Heterogeneous Remote Sensing ImagesabstractChange detection in heterogeneous remote sensing images is crucial for emergencies, such as disaster assessment. Existing methods based on homogeneous transformation suffer from the high computational cost that makes the change detection tasks time-consuming. To solve this problem, this article presents a new semisupervised Siamese network (S3N) based on transfer learning. In the proposed S3N, the low- and deep-level features are separated and treated differently for transfer learning. By incorporating two identical subnetworks that are both pretrained on natural images, the proposed S3N eliminates the computational cost for learning the low-level features that are universal for both remote sensing images and natural images. As the deep-level features contain different semantics between remote sensing images and natural images, a novel transfer learning strategy is presented to train only the weights of the layers for deep-level features in the proposed S3N. The decrease in the number of network parameters to be trained reduces the demand for training samples, leading to a significant decrease in computational cost. Afterward, the thresholding method,Otsu, is applied to the difference map derived by the proposed S3N to obtain the final binary map of change detection. Three data sets including different types of heterogeneous remote sensing images are employed to evaluate the performance of the proposed S3N. The experimental results demonstrate that the proposed S3N can achieve a comparable detection performance with much lower computational cost, compared with state-of-the-art change detection algorithms. Gang Li 0008, Xiao-Ping Zhang 0002, You He 0003 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Source Localization Based on Hybrid Coarray for 1-D Mirrored Interferometric Aperture SynthesisabstractThe mirrored interferometric aperture synthesis (MIAS) is a promising technique for high-resolution observation in microwave radiometry. In this article, we present a new source localization method based on a novel hybrid coarray that can be constructed from the 1-D MIAS. The new method, namedMA-SSmethod, employs the spatial smoothing (SS) technique on the hybrid coarray of mirrored array (MA) in the MIAS. We show that it has the ability to resolve more sources than physical sensors. The theoretical analysis gives an upper bound on degrees of freedom (DOF) of$O(N^{2})$order using$N$physical sensors. Simulation and experiment results demonstrate that, compared with the discrete cosine transform (DCT) approach commonly utilized in the MIAS, the presented MA-SS method shows the superiorities on spatial resolution, sidelobe reduction, localization accuracy, and detection performance. Gang Li 0008, Xiao-Ping Zhang 0002 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | FindNet: Can You Find Me? Boundary-and-Texture Enhancement Network for Camouflaged Object DetectionabstractCamouflaged objects share very similar colors but have different semantics with the surroundings. Cognitive scientists observe that both the global contour (i.e., boundary) and the local pattern (i.e., texture) of camouflaged objects are key cues to help humans find them successfully. Inspired by the cognitive scientist's observation, we propose a novel boundary-and-texture enhancement network (FindNet) for camouflaged object detection (COD) from single images. Different from most of existing COD methods, FindNet embeds both the boundary-and-texture information into the camouflaged object features. The boundary enhancement (BE) module is leveraged to focus on the global contour of the camouflaged object, and the texture enhancement (TE) module is utilized to focus on the local pattern. The enhanced features from BE and TE, which complement each other, are combined to obtain the final prediction. FindNet performs competently on various conditions of COD, including slightly clear boundaries but very similar textures, fuzzy boundaries but slightly differentiated textures, and simultaneous fuzzy boundaries and textures. Experimental results exhibit clear improvements of FindNet over fifteen state-of-the-art methods on four benchmark datasets, in terms of detection accuracy and boundary clearness. The code will be publicly released. Peng Li 0064, Xuefeng Yan 0001, Mingqiang Wei, Xiao-Ping Zhang 0002, Harry Qin |
IEEE Trans. Image Process. | 5 |
| 2022 | A Novel Smooth Variable Structure Filter for Target Tracking Under Model UncertaintyabstractModel uncertainty is a serious challenge for robustness of tracking algorithms in radar systems. The smooth variable structure filter (SVSF) achieves error-bounded estimations for target state by scaling the magnitude of kinematic modeling error and accordingly performing a flexible switching strategy for the correction gain. However, the SVSF, without any smoothing functions, suffers from undesired chattering phenomenon since the measurement noise causes random disturbance to the identification of actual level of uncertainties, leading to obvious deterioration of tracking accuracy. In this paper, we present a new switching function for SVSF, i.e. the hyperbolic tangent function, for effective chattering suppression. Then we propose a new algorithm named as the Tanh-SVSF, which reformulates the correction gain with the new switching function, to improve the estimation accuracy for target state. A mathematical definition of SVSF chattering is proposed to quantify the chattering amplitude. It is demonstrated that the new switching function exerts a nonlinear compressing effect on the likelihood of measurement innovation and substantially reduces the disturbance of measurement noise, leading to elimination of the chattering problem. The stability of the Tanh-SVSF is analyzed, based on a proposed stability theorem and the numerical exhaustion strategy. Finally, the proposed method is tested on a simulated vehicle tracking scenario and real-world radar data from the Oxford Radar RobotCar Dataset, and shows superior performance over existing SVSF formulations and the Kalman filter, in view of tracking accuracy, track continuity and the proposed chattering indicator. Yaowen Li, Gang Li 0008, Yu Liu 0005, Xiao-Ping Zhang 0002, You He 0002 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2022 | A Hybrid SVSF Algorithm for Automotive Radar TrackingabstractThis paper concerns the robust state estimation of automotive radar targets in presence of model uncertainty. Smooth variable structure filter (SVSF) achieves error-bounded estimation for target state, even with an inaccurate description of target kinematic model. However, it suffers the undesired chattering phenomenon especially in case of a high model uncertainty level, and its performance is sensitive to a preset smoothing boundary layer parameter. In this paper, we propose a novel hybrid SVSF algorithm to handle these two problems simultaneously. First, we derive a nonlinear generalized variable smoothing boundary layer (NGVBL) parameter based on the conventional Tanh-SVSF method by minimizing the pseudo posterior estimation error covariance. Then this NGVBL is employed to realize an adaptive two-module switching strategy with respect to the uncertainty level to calculate the correction gain. If the uncertainty level is high, the undesired chattering is effectively suppressed by the standard Tanh-SVSF gain. In case of a low uncertainty level, the NGVBL is utilized to replace the preset smoothing boundary layer parameter and reformulate the correction gain. Furthermore, it is demonstrated that the NGVBL-based gain is quasi-optimal in the mean square error (MSE) sense. Accordingly, this novel NGVBL-based hybrid SVSF (NGVBL-SVSF) algorithm improves the estimation performance by avoiding parameter sensitivity in a low uncertainty level case, and maintains effective chattering suppression and robustness to increasing uncertainties. Simulation and real-world automotive radar data experiment results show that, the proposed NGVBL-SVSF outperforms existing SVSFs and the classical Kalman filter in terms of tracking accuracy and track continuity. Yaowen Li, Gang Li 0008, Yu Liu 0005, Xiao-Ping Zhang 0002, You He 0003 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2022 | Cycle-SNSPGAN: Towards Real-World Image Dehazing via Cycle Spectral Normalized Soft Likelihood Estimation Patch GANabstractImage dehazing is a common operation in autonomous driving, traffic monitoring and surveillance. Learning-based image dehazing has achieved excellent performance recently. However, it is nearly impossible to capture pairs of hazy/clean images from the real world to train an image dehazing network. Most of existing dehazing models that are learnt from synthetically generated hazy images generalize poorly on real-world hazy scenarios due to the obvious domain shift. To deal with this unpaired problem arisen by real-world hazy images, we present Cycle Spectral Normalized Soft likelihood estimation Patch Generative Adversarial Network (Cycle-SNSPGAN) for image dehazing. Cycle-SNSPGAN is an unsupervised dehazing framework to boost the generalization ability on real-world hazy images. To leverage unpaired samples of real-world hazy images without relying on their clean counterparts, we design an SN-Soft-Patch GAN and exploit a new cyclic self-perceptual loss which avoids using the ground-truth image to compute the perceptual similarity. Moreover, a significant color loss is adopted to brighten the dehazed images as human expects. Both visual and numerical results show clear improvements of the proposed Cycle-SNSPGAN over state-of-the-arts in terms of hazy-robustness and image detail recovery, with even only a small dataset training our Cycle-SNSPGAN. Code has been available athttps://github.com/yz-wang/Cycle-SNSPGAN. Yongzhen Wang 0001, Xuefeng Yan 0001, Donghai Guan, Mingqiang Wei, Yiping Chen 0002, Xiao-Ping Zhang 0002, Jonathan Li 0001 |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2022 | Rhythm-Aware Sequence-to-Sequence Learning for Labanotation Generation With Gesture-Sensitive Graph Convolutional EncodingabstractLabanotation is a professional dance notation system widely used in dance education and choreography preservation. Automatically generatingLabanotation dance scores from motion capture data can save a huge amount of manual time and effort. Recently, the sequence-to-sequence (seq2seq) model is applied to the automatic Labanotation generation. This model is based on an encoder-decoder structure, which encodes the input motion sequence to a fixed-length vector and then decodes it to generate the target sequence. However, the encoding of spatial skeleton structure of motion data is not considered in the existing work. Besides, it is challenging to align between the input motion data and the output Laban symbol sequences due to the severe imbalance of sequence lengths. Therefore, in this paper, we present a new seq2seq model for more effective Labanotation generation. In the encoder, we propose a new gesture-sensitive graph convolutional network with learned adaptive joint weights and non-physical connections to learn both spatial and temporal patterns from motion data sequences. In the decoder, we exploit motion rhythm information and propose a novel rhythm-aware attention mechanism to learn a good alignment between motion sequences and Laban symbol sequences, so that we can focus on relevant parts of the input motion sequence without searching in the whole sequence when predicting a target Laban symbol. Extensive experiments on two real-world datasets show that the proposed method achieves a better performance compared with the state-of-the-art approaches on the task of automatic Labanotation generation. Min Li 0025, Zhenjiang Miao, Xiao-Ping Zhang 0002, Wanru Xu, Cong Ma 0004, Ningwei Xie |
IEEE Trans. Multim. | 3 |
| 2021 | Exploiting Relationship for Complex-scene Image GenerationabstractThe significant progress on Generative Adversarial Networks (GANs) has facilitated realistic single-object image generation based on language input. However, complex-scene generation (with various interactions among multiple objects) still suffers from messy layouts and object distortions, due to diverse configurations in layouts and appearances. Prior methods are mostly object-driven and ignore their inter-relations that play a significant role in complex-scene images. This work explores relationship-aware complex-scene image generation, where multiple objects are inter-related as a scene graph. With the help of relationships, we propose three major updates in the generation framework. First, reasonable spatial layouts are inferred by jointly considering the semantics and relationships among objects. Compared to standard location regression, we show relative scales and distances serve a more reliable target. Second, since the relations between objects have significantly influenced an object's appearance, we design a relation-guided generator to generate objects reflecting their relationships. Third, a novel scene graph discriminator is proposed to guarantee the consistency between the generated image and the input scene graph. Our method tends to synthesize plausible layouts and objects, respecting the interplay of multiple objects in an image. Experimental results on Visual Genome and HICO-DET datasets show that our proposed method significantly outperforms prior arts in terms of IS and FID metrics. Based on our user study and visual inspection, our method is more effective in generating logical layout and appearance for complex-scenes. Tianyu Hua, Hongdong Zheng, Yalong Bai, Wei Zhang 0031, Xiao-Ping Zhang 0002, Tao Mei 0001 |
AAAI | 5 |
| 2021 | Boosting Video Representation Learning With Multi-Faceted IntegrationabstractVideo content is multifaceted, consisting of objects, scenes, interactions or actions. The existing datasets mostly label only one of the facets for model training, resulting in the video representation that biases to only one facet depending on the training dataset. There is no study yet on how to learn a video representation from multifaceted labels, and whether multifaceted information is helpful for video representation learning. In this paper, we propose a new learning framework, MUlti-Faceted Integration (MUFI), to aggregate facets from different datasets for learning a representation that could reflect the full spectrum of video content. Technically, MUFI formulates the problem as visual-semantic embedding learning, which explicitly maps video representation into a rich semantic embedding space, and jointly optimizes video representation from two perspectives. One is to capitalize on the intra-facet supervision between each video and its own label descriptions, and the second predicts the "semantic representation" of each video from the facets of other datasets as the inter-facet supervision. Extensive experiments demonstrate that learning 3D CNN via our MUFI framework on a union of four large-scale video datasets plus two image datasets leads to superior capability of video representation. The prelearnt 3D CNN with MUFI also shows clear improvements over other approaches on several downstream video applications. More remarkably, MUFI achieves 98.1%/80.9% on UCF101/HMDB51 for action recognition and 101.5% in terms of CIDEr-D score on MSVD for video captioning. Zhaofan Qiu, Ting Yao 0003, Chong-Wah Ngo, Xiao-Ping Zhang 0002, Tao Mei 0001 |
CVPR | 4 |
| 2021 | A New Image Fusion Method for Ship Target Enhancement in Spaceborne and Airborne SAR Collaboration
Xueqian Wang 0002, Gang Li 0008, Xiao-Ping Zhang 0002 |
FUSION | 4 |
| 2021 | Virtual Reality Gaming on the Cloud: A Reality CheckabstractCloud virtual reality (VR) gaming traffic characteristics such as frame size, inter-arrival time, and latency need to be carefully studied as a first step toward scalable VR cloud service provisioning. To this end, in this paper we analyze the behavior of VR gaming traffic and Quality of Service (QoS) when VR rendering is conducted remotely in the cloud. We first build a VR testbed utilizing a cloud server, a commercial VR headset, and an off-the-shelf WiFi router. Using this testbed, we collect and process cloud VR gaming traffic data from different games under a number of network conditions and fixed and adaptive video encoding schemes. To analyze the application-level characteristics such as video frame size, frame inter-arrival time, frame loss and frame latency, we develop an interval threshold based identification method for video frames. Based on the frame identification results, we present two statistical models that capture the behaviour of the VR gaming video traffic. The models can be used by researchers and practitioners to generate VR traffic models for simulations and experiments - and are paramount in designing advanced radio resource management (RRM) and network optimization for cloud VR gaming services. To the best of the authors' knowledge, this is the first measurement study and analysis conducted using a commercial cloud VR gaming platform, and under both fixed and adaptive bitrate streaming. We make our VR traffic datasets publicly available for further research by the community. Sihao Zhao, Hatem Abou-Zeid, Ramy Atawia, Yoga Suhas Kuruba Manjunath, Akram Bin Sediq, Xiao-Ping Zhang 0002 |
GLOBECOM | 6 |
| 2021 | An Attention-Seq2Seq Model Based on CRNN Encoding for Automatic Labanotation Generation from Motion Capture DataabstractLabanotation is an important notation system widely used for recording dances. Numerous methods have been proposed for automatic Labanotation generation from motion capture data. Recently, the sequence-to-sequence (seq2seq) model is proposed. However, the encoder of the model only encodes the temporal information of motion data, lacking the encoding for spatial information. And it is challenging for the decoder to align input and output sequences due to the imbalance of the sequence lengths. In this paper, we propose an attention-seq2seq model based on Convolutional Recurrent Neural Network (CRNN). The proposed model employs an encoder based on CRNN to learn the spatial-temporal information of motion data and applies an attention mechanism to align each target Laban symbol with relevant parts of the input motion data in decoding. Experiments show that the proposed method performs favorably against state-of-the-art algorithms in the automatic Labanotation generation task. Min Li 0025, Zhenjiang Miao, Xiao-Ping Zhang 0002, Wanru Xu |
ICASSP | 3 |
| 2021 | Identification of Deep Breath While Moving Forward Based on Multiple Body Regions and Graph Signal AnalysisabstractThis paper presents an unobtrusive solution that can automatically identify deep breath when a person is walking past the global depth camera. Existing non-contact breath assessments achieve satisfactory results under restricted conditions when human body stays relatively still. When someone moves forward, the breath signals detected by depth camera are hidden within signals of trunk displacement and deformation, and the signal length is short due to the short stay time, posing great challenges for us to establish models. To over-come these challenges, multiple region of interests (ROIs) based signal extraction and selection method is proposed to automatically obtain the signal informative to breath from depth video. Subsequently, graph signal analysis (GSA) is adopted as a spatial-temporal filter to wipe the components unrelated to breath. Finally, a classifier for identifying deep breath is established based on the selected breath-informative signal. In validation experiments, the proposed approach outperforms the comparative methods with the accuracy, precision, recall and F1 of 75.5%, 76.2%, 75.0% and 75.2%, respectively. This system can be extended to public places to provide timely and ubiquitous help for those who may have or are going through physical or mental trouble. Yunlu Wang, Cheng Yang 0003, Menghan Hu, Jian Zhang 0060, Qingli Li, Guangtao Zhai, Xiao-Ping Zhang 0002 |
ICASSP | 7 |
| 2021 | Optimal TOA Localization for Moving Sensor in Asymmetric NetworkabstractIn a localization system based-on asymmetric network, only one of the anchor nodes (ANs) transmits signal. A sensor node (SN) receives it and then transmits signal that is received by all ANs to form time-of-arrival (TOA) measurements. SN localization is achieved based-on these TOA measurements along with the known AN positions. Existing work all assumes the SN is stationary. This will cause extra localization error for a moving SN. We develop an optimal localization method based-on maximum likelihood (ML) estimator, namely ML-LOC, utilizing information on the SN velocity and clock drift, to determine the position of a moving SN. We analyze its localization error and derive the Cramér-Rao lower bound (CRLB). Results from numerical simulations verify its optimal performance. We implement a prototype hardware localization system based-on consumer level ultra-wide band (UWB) chips. Experiments using the real system are carried out. Results validate the performance of the proposed method and show its feasibility in real-world applications. Sihao Zhao, Xiao-Ping Zhang 0002, Xiaowei Cui, Mingquan Lu |
ICASSP | 2 |
| 2021 | Perceptual Quality Assessment for Recognizing True and Pseudo 4k ContentabstractTo meet the imperative demand for monitoring the quality of Ultra High-Definition (UHD) content in multimedia industries, we propose an efficient no-reference (NR) image quality assessment (IQA) metric to distinguish original and pseudo 4K contents and measure the quality of their quality in this paper. First, we establish a database including more than 3000 4K images composed of natural 4K images together with upscaled versions interpolated from 1080p and 720p images by fourteen algorithms. To improve computing efficiency, our model segments the input image and selects three representative patches by local variances. Then, we extract the histogram features and cut-off frequency features in the frequency domain as well as the natural scenes statistic (NSS) based features from the representative patches. Finally, we employ support vector regressor (SVR) to aggregate these extracted features as an overall quality metric to predict the quality score of the target image. Extensive experimental comparisons using seven common evaluation indicators demonstrate that the proposed model outperforms the competitive NR IQA methods and has a great ability to distinguish true and pseudo 4K images. Wenhan Zhu, Guangtao Zhai, Xiongkuo Min, Xiaokang Yang 0001, Xiao-Ping Zhang 0002 |
ICASSP | 5 |
| 2021 | Explainable Person Re-Identification with Attribute-guided Metric DistillationabstractDespite the great progress of person re-identification (ReID) with the adoption of Convolutional Neural Networks, current ReID models are opaque and only outputs a scalar distance between two persons. There are few methods providing users semantically understandable explanations for why two persons are the same one or not. In this paper, we propose a post-hoc method, named Attribute-guided Metric Distillation (AMD), to explain existing ReID models. This is the first method to explore attributes to answer: 1) what and where the attributes make two persons different, and 2) how much each attribute contributes to the difference. In AMD, we design a pluggable interpreter network for target models to generate quantitative contributions of attributes and visualize accurate attention maps of the most discriminative attributes. To achieve this goal, we propose a metric distillation loss by which the interpreter learns to decompose the distance of two persons into components of attributes with knowledge distilled from the target model. Moreover, we propose an attribute prior loss to make the interpreter generate attribute-guided attention maps and to eliminate biases caused by the imbalanced distribution of attributes. This loss can guide the interpreter to focus on the exclusive and discriminative attributes rather than the large-area but common attributes of two persons. Comprehensive experiments show that the interpreter can generate effective and intuitive explanations for varied models and generalize well under cross-domain settings. As a by-product, the accuracy of target models can be further improved with our interpreter.1 Xiaodong Chen 0011, Xinchen Liu, Wu Liu 0005, Xiao-Ping Zhang 0002, Yongdong Zhang 0001, Tao Mei 0001 |
ICCV | 4 |
| 2021 | Modeling Image Quality Score Distribution Using Alpha Stable ModelabstractIn recent years, image quality is generally described by a mean opinion score (MOS). However, we observe that an image’s quality ratings given by a group of subjects may not follow a Gaussian distribution and the image quality can not be fully described by a MOS. In this paper, we propose to describe the image quality using a parameterized distribution rather than a MOS, and an objective method is also proposed to predict the image quality score distribution (IQSD). Specifically, we selected 100 images from the LIVE database and invited a large group of subjects to evaluate the quality of these images. By analyzing the subjective quality ratings, we find that the IQSD can be well modeled by an alpha stable model and this model can reflect much more information than MOS. Therefore, we propose an algorithm to model the IQSD described by an alpha stable model. Features are extracted from images based on natural scene statistics and support vector regressors are trained to predict the IQSD described by an alpha stable model. We validate the proposed IQSD prediction model on the collected subjective quality ratings. Experimental results verify the effectiveness of the proposed algorithm in modeling the IQSD. Xiongkuo Min, Wenhan Zhu, Xiao-Ping Zhang 0002, Guangtao Zhai |
ICIP | 4 |
| 2021 | Fusion of Spaceborne and Airborne SAR Images Using Saliency and Fuzzy Logic for Vessel DetectionabstractIn the paper, we propose a new method based on multi-order superpixel-level saliency and fuzzy logic (MSSFL) to fuse spaceborne and airborne SAR images for vessel detection. First, we generate a new global regional contrast map (GRCM) by exploiting the multi-order superpixel-level saliency (MSS). In the generated GRCM, the vessel targets are well restored and the backgrounds are suppressed. Next, a new fuzzy logic approach is presented to fuse the MSS information provided by the GRCMs. This GRCM-based fuzzy fusion can further enhance the vessel target regions and filter out the inshore interference regions. Experimental results using Gaofen-3 satellite and unmanned aerial vehicle (UAV) SAR images show that the proposed MSSFL method yields higher target-to-cluster ratio (TCR) of fused images and improved detection performance compared with the commonly utilized image fusion approaches. Xueqian Wang 0002, Gang Li 0008, Xiao-Ping Zhang 0002 |
IGARSS | 4 |
| 2021 | Direction-aware Feature-level Frequency Decomposition for Single Image DerainingabstractWe present a novel direction-aware feature-level frequency decomposition network for single image deraining. Compared with existing solutions, the proposed network has three compelling characteristics. First, unlike previous algorithms, we propose to perform frequency decomposition at feature-level instead of image-level, allowing both low-frequency maps containing structures and high-frequency maps containing details to be continuously refined during the training procedure. Second, we further establish communication channels between low-frequency maps and high-frequency maps to interactively capture structures from high-frequency maps and add them back to low-frequency maps and, simultaneously, extract details from low-frequency maps and send them back to high-frequency maps, thereby removing rain streaks while preserving more delicate features in the input image. Third, different from existing algorithms using convolutional filters consistent in all directions, we propose a direction-aware filter to capture the direction of rain streaks in order to more effectively and thoroughly purge the input images of rain streaks. We extensively evaluate the proposed approach in three representative datasets and experimental results corroborate our approach consistently outperforms state-of-the-art deraining algorithms. Yidan Feng, Mingqiang Wei, Haoran Xie 0001, Yiping Chen 0002, Jonathan Li 0001, Xiao-Ping Zhang 0002, Harry Qin |
IJCAI | 7 |
| 2021 | TraND: Transferable Neighborhood Discovery for Unsupervised Cross-Domain Gait RecognitionabstractGait, i.e., the movement pattern of human limbs during locomotion, is a promising biometrie for identification of persons. Despite significant improvement in gait recognition with deep learning, existing studies still neglect a more practical but challenging scenario - unsupervised cross-domain gait recognition which aims to learn a model on a labeled dataset then adapt it to an unlabeled dataset. Due to the domain shift and class gap, directly applying a model trained on one source dataset to other target datasets usually obtains very poor results. Therefore, this paper proposes a Transferable Neighborhood Discovery (TraND) framework to bridge the domain gap for unsupervised cross-domain gait recognition. To learn effective prior knowledge for gait representation, we first adopt a backbone network pre- trained on the labeled source data in a supervised manner. Then we design an end-to-end trainable approach to automatically discover the confident neighborhoods of unlabeled samples in the latent space. During training, the class consistency indicator is adopted to select confident neighborhoods of samples based on their entropy measurements. Moreover, we explore a high- entropy-first neighbor selection strategy, which can effectively transfer prior knowledge to the target domain. Our method achieves the state-of-the-art results on two public datasets, i.e., CASIA-B and OU-LP. Jinkai Zheng, Xinchen Liu, Chenggang Yan 0001, Jiyong Zhang 0001, Wu Liu 0005, Xiao-Ping Zhang 0002, Tao Mei 0001 |
ISCAS | 6 |
| 2021 | Inter-Observer Visual Congruency in Video-ViewingabstractThere are individual differences in human visual attention between observers when viewing the same scene. Inter-observer visual congruency (IOVC) describes the dispersion between different people's visual attention areas when they observe the same stimulus. Research on the IOVC of video is interesting but lacking. In this paper, we first introduce the measurement to calculate the IOVC of video. And an eye-tracking experiment is conducted in a realistic movie-watching environment to establish a movie scene dataset. Then we propose a method to predict the IOVC of video, which employs a dual-channel network to extract and integrate content and optical flow features. The effectiveness of the proposed prediction model is validated on our dataset. And the correlation between inter-observer congruency and video emotion is analyzed. Jiaomin Yue, Dandan Zhu 0001, Xiongkuo Min, Xiao-Ping Zhang 0002, Guangtao Zhai |
VCIP | 5 |
| 2021 | Respiratory Consultant by Your Side: Affordable and Remote Intelligent Respiratory Rate and Respiratory Pattern Monitoring SystemabstractThe aim of this study is to develop an affordable and remote intelligent respiratory monitoring system. To achieve low-cost and remote measurement of respiratory signal, an RGB camera collaborated with marker tracking is used as a data acquisition sensor, and a Raspberry Pi is used as a data processing platform. To overcome challenges in actual applications, the signal processing algorithms are designed for removing sudden body movements and smoothing the raw signal. Subsequently, respiratory rate (RR) is estimated by a translational cross-point algorithm, and the respiratory pattern is identified by the recurrent neural network. For estimating RR, the translational cross-point algorithm performs better than other methods with root-mean-square error (RMSE) of 3.29 bpm. With respect to the classification of breathing patterns, the established neural network performs better than support vector machine-based classifiers with the accuracy, precision, recall, and F1 of 89.0%, 89.0%, 90.5%, and 89.0%, respectively. The obtained decision-making information and some original information are sent to the user’s smartphone via a cloud service platform. In a way, due to its low-price, noncontact, and portable merits, the established system can be seen as a “respiratory consultant” by your side. Yunlu Wang, Menghan Hu, Jian Zhang 0060, Qingli Li, Guangtao Zhai, Simon X. Yang, Xiao-Ping Zhang 0002, Xiaokang Yang 0001 |
IEEE Internet Things J. | 9 |
| 2021 | A Closed-Form Localization Method Utilizing Pseudorange Measurements From Two Nonsynchronized Positioning SystemsabstractIn a time of arrival (TOA) or pseudorange-based positioning system, user location is obtained by observing multiple anchor nodes (ANs) at known positions. Utilizing more than one positioning systems, e.g., combining global positioning system (GPS) and BeiDou navigation satellite system (BDS), brings better positioning accuracy. However, ANs from two systems are usually synchronized to two different clock sources. Different from single-system localization, an extra user-to-system clock offset needs to be handled. Existing dual-system methods either have high computational complexity or suboptimal positioning accuracy. In this article, we propose a new closed-form dual-system localization (CDL) approach that has low complexity and optimal localization accuracy. We first convert the nonlinear problem into a linear one by squaring the distance equations and employing intermediate variables. Then, a weighted least-squares (WLSs) method is used to optimize the positioning accuracy. We prove that the positioning error of the new method reaches Cramér-Rao lower bound (CRLB) in far-field conditions with small measurement noise. Simulations on 2-D and 3-D positioning scenes are conducted. Results show that, compared with the iterative approach, which has high complexity and requires a good initialization, the new CDL method does not require initialization and has lower computational complexity with comparable positioning accuracy. The numerical results verify the theoretical analysis on positioning accuracy, and show that the new CDL method has superior performance over the state-of-the-art closed-form method. Experiments using real GPS and BDS data verify the applicability of the new CDL method and the superiority of its performance in the real world. Sihao Zhao, Xiao-Ping Zhang 0002, Xiaowei Cui, Mingquan Lu |
IEEE Internet Things J. | 2 |
| 2021 | Optimal Localization With Sequential Pseudorange Measurements for Moving Users in a Time-Division Broadcast Positioning SystemabstractIn a time-division broadcast positioning system, a user device (UD) determines its position by obtaining sequential time of arrival or pseudorange measurements from signals broadcast by multiple synchronized base stations. The existing localization method using sequential pseudorange measurements and a linear clock drift model for the TDPBS, namely, LSPM-D, does not compensate the position displacement caused by the UD movement and will result in position error. In this article, depending on the knowledge of the UD velocity, we develop a set of optimal localization methods for different cases. First, for known UD velocity, we develop the optimal localization method, namely, LSPM-KVD, to compensate the movement-caused position error. We show that the LSPM-D is a special case of the LSPM-KVD when the UD is stationary with zero velocity. Second, for the case with unknown UD velocity, we develop a maximum-likelihood (ML) method to jointly estimate the UD position and velocity, namely, LSPM-UVD. Third, in the case that we have prior distribution information of the UD velocity, we present a maximum a posteriori estimator for localization, namely, LSPM-PVD. We derive the Cramér-Rao lower bound for all three estimators and analyze their localization error performance. We show that the position error of the LSPM-KVD increases as the assumed known velocity deviates from the true value. As expected, the LSPM-KVD has the smallest position error while the LSPM-PVD and the LSPM-UVD are more robust when the prior knowledge of the UD velocity is limited. Numerical results verify the theoretical analysis on the optimality and the positioning accuracy of the proposed methods. Sihao Zhao, Xiao-Ping Zhang 0002, Xiaowei Cui, Mingquan Lu |
IEEE Internet Things J. | 2 |
| 2021 | A New TOA Localization and Synchronization System With Virtually Synchronized Periodic Asymmetric Ranging NetworkabstractIn this article, we design a new time-of-arrival (TOA) system for simultaneous user device (UD) localization and synchronization with a periodic asymmetric ranging network, namely, PARN. The PARN includes one primary anchor node (PAN) transmitting and receiving signals, and many secondary ANs (SANs) only receiving signals. All the UDs can transmit and receive signals. The PAN periodically transmits sync signal and the UD transmits response signal after reception of the sync signal. Using TOA measurements from the periodic sync signal at SANs, we develop a Kalman filtering method to virtually synchronize anchor nodes (ANs) with high accuracy estimation of clock parameters. Employing the virtual synchronization, and TOA measurements from the response signal and sync signal, we then develop a maximum-likelihood (ML) approach, namely, ML-LAS, to simultaneously localize and synchronize a moving UD. We analyze the UD localization and synchronization error, and derive the Cramér-Rao lower bound (CRLB). Different from existing asymmetric ranging network-based TOA systems, the new PARN 1) uses the periodic sync signals at the SAN to exploit the temporal correlated clock information for high accuracy virtual synchronization and 2) compensates the UD movement and clock drift using various TOA measurements to achieve consistent and simultaneous localization and synchronization performance. Numerical results verify the theoretical analysis that the new system has high accuracy in AN clock offset estimation and simultaneous localization and synchronization for a moving UD. We implement a prototype hardware system and demonstrate the feasibility and superiority of the PARN in real-world applications by experiments. Sihao Zhao, Xiao-Ping Zhang 0002, Xiaowei Cui, Mingquan Lu |
IEEE Internet Things J. | 2 |
| 2021 | A Fast CFAR Algorithm Based on Density-Censoring Operation for Ship Detection in SAR ImagesabstractIn this letter, we propose a new constant false alarm rate (CFAR) detector to accelerate the existing superpixel (SP)-based CFAR detectors for ship detection in synthetic aperture radar (SAR) images. In our method, we design a new density-censoring operation to rapidly identify background clutter SPs (BCSPs) with high densities before the local CFAR detection. In this way, a large number of non-informative BCSPs are removed without time-consuming calculation of decision thresholds, and only a few candidate ship target SPs (STSPs) are retained. This reduces the computational cost of the subsequent local CFAR detection and the number of false alarms produced by it. During the local CFAR detection process for the retained candidate STSPs, we also propose an improved method to define their neighboring clutter regions (for the calculation of decision thresholds) using BCSPs identified by the density-censoring operation. Experiments on measured SAR images validate that the proposed CFAR method reduces the computational cost of commonly used SP-based CFAR methods by 75%-96% with similar or better detection performance. Xueqian Wang 0002, Gang Li 0008, Xiao-Ping Zhang 0002, You He 0003 |
IEEE Signal Process. Lett. | 3 |
| 2021 | Semidefinite Programming Two-Way TOA Localization for User Devices With Motion and Clock DriftabstractIn two-way time-of-arrival (TOA) systems, a user device (UD) obtains its position by round-trip communications to a number of anchor nodes (ANs) at known locations. The objective function of the maximum likelihood (ML) method for two-way TOA localization is nonconvex. Thus, the widely-adopted Gauss-Newton iterative method to solve the ML estimator usually suffers from the local minima problem. In this letter, we convert the original estimator into a convex problem by relaxation, and develop a new semidefinite programming (SDP) based localization method for moving UDs, namely SDP-M. Numerical result demonstrates that compared with the iterative method, which often fall into local minima, the SDP-M always converge to the global optimal solution and significantly reduces the localization error by more than 40%. It also has stable localization accuracy regardless of the UD movement, and outperforms the conventional method for stationary UDs, which has larger error with growing UD velocity. Sihao Zhao, Xiao-Ping Zhang 0002, Xiaowei Cui, Mingquan Lu |
IEEE Signal Process. Lett. | 2 |
| 2021 | Re-Synchronization Using the Hand Preceding Model for Multi-Modal Fusion in Automatic Continuous Cued Speech RecognitionabstractCued Speech (CS) is an augmented lip reading system complemented by hand coding, and it is very helpful to the deaf people. Automatic CS recognition can help communications between the deaf people and others. Due to the asynchronous nature of lips and hand movements, fusion of them in automatic CS recognition is a challenging problem. In this work, we propose a novel re-synchronization procedure for multi-modal fusion, which aligns the hand features with lips feature. It is realized by delaying hand position and hand shape with their optimal hand preceding time which is derived by investigating the temporal organizations of hand position and hand shape movements in CS. This re-synchronization procedure is incorporated into a practical continuous CS recognition system that combines convolutional neural network (CNN) with multi-stream hidden markov model (MSHMM). A significant improvement of about 4.6% has been achieved retaining 76.6% CS phoneme recognition correctness compared with the state-of-the-art architecture (72.04%), which did not take into account the asynchrony issue of multi-modal fusion in CS. To our knowledge, this is the first work to tackle the asynchronous multi-modal fusion in the automatic continuous CS recognition. Li Liu 0036, Gang Feng 0002, Denis Beautemps, Xiao-Ping Zhang 0002 |
IEEE Trans. Multim. | 4 |
| 2021 | A New Multihypothesis-Based Compressed Video Sensing Reconstruction SystemabstractMultihypothesis-based compressed video sensing scheme attracts wide attention in the research of resource-constrained video application scenarios. However, high-accuracy weight prediction of hypotheses is always challenging especially for the high-motion sequences. To solve this problem, this paper proposes a novel multihypothesis-based distributed compressed video sensing (NMH-DCVS) system. The new multihypothesis system contains two components: hypotheses acquisition, and weight prediction. First, to acquire more high-quality hypotheses, a new hypotheses acquisition scheme is proposed by constructing the search window based on the temporal, and spatial correlation of the image blocks, respectively. The optimal matching block can be quickly determined. Second, to improve the accuracy of the multihypothesis weight prediction, a new residual transforming preprocessing-based weight prediction algorithm is proposed by transforming the original hypothesis set to residual hypothesis set. The influence of the quality fluctuation of the hypotheses on prediction accuracy is effectively suppressed. Moreover, the improved hypotheses further improve the sparsity of the residual hypothesis set, leading to the additional improvement of the accuracy of the proposed residual-based weight prediction algorithm. Experiment results show that compared with the state-of-the-art methods reported in the literature, the proposed new multihypothesis system significantly improves the decoding performance both in objective, and subjective quality. Jian Chen 0002, Xiao-Ping Zhang 0002, Yonghong Kuo |
IEEE Trans. Multim. | 3 |
| 2020 | A Novel Moving Sparse Array Geometry with Increased Degrees of FreedomabstractIn this paper, we propose a novel moving sparse array geometry named dilated arrays (DAs) by extending the dilation of nested arrays to other linear array structures. The theoretical analysis of dilation to other arrays is not straightforward since the results for dilated nested arrays have been proven based on closed-form expression of their array geometry. However, other types of arrays have different, or even no closed-form expressions for their array geometries for a given number of sensors. We prove that for a DA on a moving platform, the degrees of freedom can be enhanced as three times high as that of its original array regardless of its array geometry. Numerical simulations are given to demonstrate the superior performance of different types of DAs. Shuang Li 0005, Xiao-Ping Zhang 0002 |
ICASSP | 2 |
| 2020 | A New Multihypothesis Prediction Scheme for Compressed Video Sensing ReconstructionabstractFor multihypothesis-based compressed video sensing schemes, the low accuracy of weight prediction and degradation of recovery quality for high-motion videos are open challenges. To solve this problem, this paper proposes a new multihypothesis prediction scheme. To efficiently get high-quality hypotheses, a new hypotheses acquiring method is proposed by building the search window based on the temporal and spatial correlation. To improve the accuracy of weight prediction, a residual transforming preprocessing for weight prediction is proposed. By converting the original hypotheses to residual hypotheses, the influence of quality fluctuation of hypotheses on the recovery quality is suppressed effectively. The sparsity and accuracy of the prediction model are improved efficiently. Simulation results show that a significant improvement in recovery quality is obtained in the proposed scheme compared to the state-of-the-art systems. Xiao-Ping Zhang 0002, Jian Chen 0002, Yonghong Kuo |
ICASSP | 2 |
| 2020 | Mirrored Arrays for Direction-of-Arrival EstimationabstractA mirrored array configuration, which consists of an ordinary linear array and a reflector, is proposed. The mirrored arrays can provide composable degrees of freedom (DOF) for array processing by combined measurements with the help of mechanical adjustment (e.g., movement or rotation). This composability naturally leads to DOF enhancement and facilitates its applications in direction-of-arrival (DOA) estimation. The feasibility and effectiveness of the proposed arrays are demonstrated through numerical analysis. Gang Li 0008, Xiao-Ping Zhang 0002 |
ICASSP | 3 |
| 2020 | Automatic Region Selection For Objective Sharpness Assessment Of Mobile Device PhotosabstractMobile devices are the source of a vast majority of digital photos today. Photos taken by mobile devices generally have fairly good visual quality. When evaluating high-quality mobile device photos, people have to manually zoom in to local regions to discern the subtle difference. Understandably, a global objective quality assessment method cannot perform well on such task. Therefore, local region selection is widely recognized as a prerequisite for the following quality evaluation. Clearly, subjective regions selection suffers from the drawbacks in terms of productivity, reproducibility and optimality. In this paper, we propose an automatic local region selection algorithm for sharpness measurement of mobile device photos. Specifically, local texture statistics, depth, saliency, as well as inter-pictures difference, are used as main features to select an optimal local region, in which the sharpness is then measured. For validation, we have built a largescale database for sharpness evaluation of mobile device photos, with 100 different scenes shot by several flagship mobile phones. The experimental results show that the performance of classic sharpness evaluation algorithms can be substantially improved with the region selected by the proposed algorithm. Guangtao Zhai, Wenhan Zhu, Yucheng Zhu, Xiongkuo Min, Xiao-Ping Zhang 0002, Hua Yang 0001 |
ICIP | 6 |
| 2020 | Unobtrusive and Automatic Classification of Multiple People's Abnormal Respiratory Patterns in Real Time Using Deep Neural Network and Depth CameraabstractRespiratory pattern is a representation of human breathing activity, which can reflect people's physical and psychological condition. Capturing the unexpected abnormal respiratory pattern unobtrusively of the patient or the potential patient has great significance. In the current work, we attempt to capitalize on depth camera and deep learning architecture to achieve the accurate and unobtrusive measurement of abnormal respiratory patterns, and the whole system can classify multiple people's respiratory patterns in a real-time manner. The challenges in this task are threefold: 1) the real-time online system means that the Region of Interest (ROI) needs to be located and tracked automatically; 2) the amount of real-world data is not enough for training to obtain the robust deep neural network; and 3) the intraclass variation is large and the outer class variation is small. Consequently, human joints tracking is applied to determine the location of subjects shoulder and chest. Based on the characteristics of actual respiratory signals, a novel and efficient respiratory simulation model (RSM) is proposed to generate abundant and high-quality training data. Finally, we apply a gated recurrent unit (GRU) neural network with bidirectional and attentional mechanisms (BI-AT-GRU) to classify six clinically significant respiratory patterns (Eupnea, Tachypnea, Bradypnea, Biots, Cheyne-Stokes, and Central-Apnea). The performance of the obtained BI-AT-GRU is tested by the data that is actually measured by the depth camera. The experimental results demonstrate that the proposed model can classify six different respiratory patterns with the accuracy, precision, recall, and F1 of 94.5%, 94.4%, 95.1%, and 94.8%, respectively. In comparative experiments, the obtained BI-AT-GRU specific to respiratory pattern classification outperforms the existing state-of-the-art, viz., BI-AT-LSTM, GRU, long short-term memory (LSTM), and BI-AT-GRU. Moreover, other experimental results indicate that the proposed online measuring system, deep neural network, and the modeling ideas have the potential to be extended to the large-scale applications, such as public places, sleep scenario, and office environment. The demo videos of the proposed system are available at: https://doi.org/10.6084/m9.figshare.11493666.v1. Yunlu Wang, Menghan Hu, Yuwen Zhou, Qingli Li, Nan Yao, Guangtao Zhai, Xiao-Ping Zhang 0002, Xiaokang Yang 0001 |
IEEE Internet Things J. | 7 |
| 2020 | Abnormal event detection in surveillance videos based on low-rank and compact coefficient dictionary learning
Zhenjiang Miao, Yi-Gang Cen, Xiao-Ping Zhang 0002, Linna Zhang, Shiming Chen 0001 |
Pattern Recognit. | 4 |
| 2020 | A New Approach to Construct Virtual Array With Increased Degrees of Freedom for Moving Sparse ArraysabstractIn this letter, a novel approach to construct virtual array is proposed by exploiting synthetic aperture technology and the concept of difference coarray for sparse arrays. First, we prove that the proposed method can increase the degrees of freedom (DOF) of an arbitrary moving sparse array threefold. Therefore, the maximal number of resolvable sources for direction-of-arrival (DOA) estimation is tripled as well. Furthermore, we prove that the difference coarray of synthetic arrays is a hole-free uniform linear array (ULA) if the difference coarray of original arrays is hole-free. Therefore, for ULA based DOA estimation methods, better performance can be achieved by fully employing all elements in the difference coarray of synthetic arrays. In addition, unlike the dilated nested array method, which enlarges the intersensor spacing of nested arrays threefold to triple DOF, the proposed method does not need to increase array aperture physically. Thus it is more practical to mount a sensor array on a moving platform when taking space limitation into consideration. Simulation results are given to demonstrate the superiority of the proposed method. Shuang Li 0005, Xiao-Ping Zhang 0002 |
IEEE Signal Process. Lett. | 2 |
| 2020 | SAR Image Despeckling Based on Combination of Fractional-Order Total Variation and Nonlocal Low Rank RegularizationabstractRegularization method is an effective tool for synthetic aperture radar (SAR) image despeckling. Design of the effective regularization terms describing the image priors plays a vital role in this kind of method. In this article, a new combinational regularization model for speckle reduction (CRM-SR) is proposed, in which a regularization term is elaborately designed to contain both a fractional-order total variation (FrTV) regularization and a nonlocal low rank (NLR) regularization. The new regularization model inherits both the advantages of FrTV and NLR and improves the performance of SAR despeckling and, therefore, better preserves the edges and geometrical features of the images during the despeckling process. An efficient algorithm based on alternating direction optimization is derived to solve the proposed combinational regularization model. Experimental results show that the proposed model can effectively remove SAR image speckle and preserve the geometrical features of images according to both subjective visual assessment of image quality and objective evaluation. Gang Li 0008, Yu Liu 0005, Xiao-Ping Zhang 0002 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2020 | Ship Detection in SAR Images via Local Contrast of Fisher VectorsabstractExisting superpixel-based detection algorithms for ship targets in synthetic aperture radar (SAR) images are often derived from the local contrast of intensities (i.e., the local contrast of the first-order information of superpixels) leading to deteriorating performance in low signal-to-clutter ratio (SCR) cases due to the low contrast between the intensities of targets and the clutter. In this article, we propose a new superpixel-based detector to improve the performance of ship target detection in SAR images via the local contrast of fisher vectors (LCFVs). The new LCFV-based detector exploits multiorder features of the superpixels based on the Gaussian mixture model (GMM) and accordingly improves the discrimination capability between the ship targets and the sea clutter, especially in low SCR cases. Experimental results demonstrate that the proposed LCFV-based detection algorithm provides better detection performance than the commonly used detection algorithms. Xueqian Wang 0002, Gang Li 0008, Xiao-Ping Zhang 0002, You He 0003 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2020 | DDcGAN: A Dual-Discriminator Conditional Generative Adversarial Network for Multi-Resolution Image FusionabstractIn this paper, we proposed a new end-to-end model, termed as dual-discriminator conditional generative adversarial network (DDcGAN), for fusing infrared and visible images of different resolutions. Our method establishes an adversarial game between a generator and two discriminators. The generator aims to generate a real-like fused image based on a specifically designed content loss to fool the two discriminators, while the two discriminators aim to distinguish the structure differences between the fused image and two source images, respectively, in addition to the content loss. Consequently, the fused image is forced to simultaneously keep the thermal radiation in the infrared image and the texture details in the visible image. Moreover, to fuse source images of different resolutions, e.g., a low-resolution infrared image and a high-resolution visible image, our DDcGAN constrains the downsampled fused image to have similar property with the infrared image. This can avoid causing thermal radiation information blurring or visible texture detail loss, which typically happens in traditional methods. In addition, we also apply our DDcGAN to fusing multi-modality medical images of different resolutions, e.g., a low-resolution positron emission tomography image and a high-resolution magnetic resonance image. The qualitative and quantitative experiments on publicly available datasets demonstrate the superiority of our DDcGAN over the state-of-the-art, in terms of both visual effect and quantitative metrics. Jiayi Ma 0001, Han Xu 0001, Junjun Jiang, Xiaoguang Mei, Xiao-Ping Zhang 0002 |
IEEE Trans. Image Process. | 5 |
| 2020 | A Multimodal Saliency Model for Videos With High Audio-Visual CorrespondenceabstractAudio information has been bypassed by most of current visual attention prediction studies. However, sound could have influence on visual attention and such influence has been widely investigated and proofed by many psychological studies. In this paper, we propose a novel multi-modal saliency (MMS) model for videos containing scenes with high audio-visual correspondence. In such scenes, humans tend to be attracted by the sound sources and it is also possible to localize the sound sources via cross-modal analysis. Specifically, we first detect the spatial and temporal saliency maps from the visual modality by using a novel free energy principle. Then we propose to detect the audio saliency map from both audio and visual modalities by localizing the moving-sounding objects using cross-modal kernel canonical correlation analysis, which is first of its kind in the literature. Finally we propose a new two-stage adaptive audiovisual saliency fusion method to integrate the spatial, temporal and audio saliency maps to our audio-visual saliency map. The proposed MMS model has captured the influence of audio, which is not considered in the latest deep learning based saliency models. To take advantages of both deep saliency modeling and audio-visual saliency modeling, we propose to combine deep saliency models and the MMS model via a later fusion, and we find that an average of 5% performance gain is obtained. Experimental results on audio-visual attention databases show that the introduced models incorporating audio cues have significant superiority over state-of-the-art image and video saliency models which utilize a single visual modality. Xiongkuo Min, Guangtao Zhai, Jiantao Zhou 0001, Xiao-Ping Zhang 0002, Xiaokang Yang 0001, Xin-Ping Guan |
IEEE Trans. Image Process. | 4 |
| 2020 | MEF-GAN: Multi-Exposure Image Fusion via Generative Adversarial NetworksabstractIn this paper, we present an end-to-end architecture for multi-exposure image fusion based on generative adversarial networks, termed as MEF-GAN. In our architecture, a generator network and a discriminator network are trained simultaneously to form an adversarial relationship. The generator is trained to generate a real-like fused image based on the given source images which is expected to fool the discriminator. Correspondingly, the discriminator is trained to distinguish the generated fused images from the ground truth. The adversarial relationship makes the fused image not limited to the restriction of the content loss. Therefore, the fused images are closer to the ground truth in terms of probability distribution, which can compensate for the insufficiency of single content loss. Moreover, aiming at the problem that the luminance of multi-exposure images varies greatly with spatial location, the self-attention mechanism is employed in our architecture to allow for attention-driven and long-range dependency. Thus, local distortion, confusing results, or inappropriate representation can be corrected in the fused image. Qualitative and quantitative experiments are performed on publicly available datasets, where the results demonstrate that MEF-GAN outperforms the state-of-the-art, in terms of both visual effect and objective evaluation metrics. Our code is publicly available at https://github.com/jiayi-ma/MEF-GAN. Han Xu 0001, Jiayi Ma 0001, Xiao-Ping Zhang 0002 |
IEEE Trans. Image Process. | 3 |
| 2019 | Point Cloud Denoising Based on Tensor Tucker DecompositionabstractIn this paper, we propose a new algorithm for point cloud de-noising based on the tensor Tucker decomposition. We first represent the local surface patches of a noisy point cloud to be matrices by their distances to a reference point, and stack the similar patch matrices to be a 3rd order tensor. Then we use the Tucker decomposition to compress this patch tensor to be a core tensor of smaller size. We consider this core tensor as the frequency domain and remove the noise by manipulating the hard thresholding. Finally, all the fibers of the denoised patch tensor are placed back, and the average is taken if there are more than one estimators overlapped. The experimental evaluation shows that the proposed algorithm outperforms the state-of-the-art graph Laplacian regularized (GLR) algorithm when the Gaussian noise is high (σ = 0.1), and the GLR algorithm is better in lower noise cases (σ = 0.04, 0.05, 0.08). Jianze Li, Xiao-Ping Zhang 0002 |
ICIP | 2 |
| 2019 | Homogeneous Transformation Based on Deep-Level Features in Heterogeneous Remote Sensing ImagesabstractHomogeneous transformation receives considerable attention in recent years as it is essential for change detection in heterogeneous images. However, most existing methods perform the homogeneous transformation based on low-level features. It leads to inaccurate homogeneous representations of the heterogeneous images and accordingly causes unsatisfied performance of change detection. To solve this problem, this paper presents a new model that utilizes deep- level features for homogeneous transformation instead of low-level features. Experimental results on real remote sensing data show that, the proposed method achieves an overall change detection accuracy of 95.91%, providing better performance than the existing methods based on low- level features. Gang Li 0008, Yu Liu 0005, Xiao-Ping Zhang 0002 |
IGARSS | 4 |
| 2019 | Automatic Detection of the Temporal Segmentation of Hand Movements in British English Cued Speech
Li Liu 0036, Jianze Li, Gang Feng 0002, Xiao-Ping Zhang 0002 |
INTERSPEECH | 4 |
| 2019 | A New Compressed Sensing Based Terminal-to-Cloud Video Transmission SystemabstractIt is challenging for resource-constrained wireless terminals to realize the high-efficiency video upload. This paper proposes a compressed sensing based high-efficiency video upload (CS-HEVU) system for the terminal-to-cloud network. The main contributions are as follows. First, to develop higher transmission efficiency, a new encoder sampling scheme is proposed by utilizing skip block based residual compressed sensing sampling. The redundant image blocks are effectively removed in encoding. Second, to improve the decoding quality, a local secondary reconstruction based multi-reference frames cross recovery algorithm is proposed in decoder. The user experience is improved by reducing the quality fluctuation of non-key frames. Compared with the state-of-the-art systems reported in literature, simulation results show that a significant improvement in encoding efficiency and decoding performance is obtained in the proposed scheme. Xiao-Ping Zhang 0002, Jian Chen 0002, Yonghong Kuo |
ISCAS | 2 |
| 2019 | A novel deep model with multi-loss and efficient training for person re-identification
Di Wu 0030, Si-Jia Zheng, Wenzheng Bao, Xiao-Ping Zhang 0002, Chang-an Yuan 0001, De-Shuang Huang |
Neurocomputing | 4 |
| 2019 | Deep learning-based methods for person re-identification: A comprehensive review
Di Wu 0030, Si-Jia Zheng, Xiao-Ping Zhang 0002, Chang-an Yuan 0001, Yang Zhao 0002, Yong-Jun Lin, Zhong-Qiu Zhao, Yong-Li Jiang, De-Shuang Huang |
Neurocomputing | 3 |
| 2019 | Physical Password Breaking via Thermal Sequence AnalysisabstractThe thermal camera can capture keyboard surface temperature change after a human's touch. This phenomenon may be used to steal users' passwords physically. In this paper, based on the study of thermal dynamics of keyboards, we design a password break system using an infrared thermal camera. First, we build a signal model to describe the dynamic process of temperature change on the keyboard using Newton's law of cooling. Next, we develop a maximum likelihood parameter estimation algorithm to estimate the keystroke time instants. Then, by maximizing the probability of key order arrangement, a novel password breaking algorithm is developed. Our algorithm is tested using simulated data as well as real-world data. Experiment results show that our algorithm is effective for physical password breaking using thermal characteristics. Based on our results, we discuss strategies for password protection at the end. Xiao-Ping Zhang 0002, Menghan Hu, Guangtao Zhai, Xiaokang Yang 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2019 | A High-Efficiency Compressed Sensing-Based Terminal-to-Cloud Video Transmission SystemabstractWith the rapid popularization of mobile intelligent terminals, mobile video and cloud services applications are widely used in people's lives. However, the resource-constrained characteristic of the terminals and the enormous amount of video information make the efficient terminal-to-cloud data upload a challenge. To solve the problem, this paper proposes an efficient compressed sensing-based high-efficiency video upload system for the terminal-to-cloud upload network. The system contains two main new components. First, to effectively remove the inter-frame redundant information, an encoder sampling scheme with high efficiency is developed by applying the skip block-based residual compressed sensing sampling technology. For the time-varying channel state, the encoder can adaptively allocate the sampling rate for different video frames by the proposed adaptive sampling scheme. Second, a local secondary reconstruction-based multi-reference frames cross-recovery algorithm is developed at the decoder. It further improves the reconstruction quality and reduces the quality fluctuation of the recovered video frames to improve the user experience. Compared with the state-of-the-art reference systems reported in the literature, the proposed system achieves the high-efficiency and high-quality terminal-to-cloud transmission. Xiao-Ping Zhang 0002, Jian Chen 0002, Yonghong Kuo |
IEEE Trans. Multim. | 2 |
| 2018 | Overlapping Animal Sound Classification Using Sparse RepresentationabstractIn this paper, a new method to classify the animal sound signals that are overlapped in time-frequency domain based on sparse representation is proposed. In order to obtain a discriminant sparse representation of overlapped animal sound signals, a novel dictionary atom discriminant factor is introduced. Then the proposed method generates a representation that contains crucial signal discriminant information for classification and the sparsity for sparsest representation. The experimental results show that the proposed method has a much superior performance than the conventional sparse representation based classification method for classifying the overlapped animal sound signals. Xiao-Ping Zhang 0002 |
ICASSP | 3 |
| 2018 | Modeling Thermal Sequence Signal Decreasing for Dual Modal Password BreakingabstractThe thermal camera records trace of users' touch a while after they type in the password. People's password may be stolen through a thermal camera due to this phenomenon. In this paper, we model the procedure of the thermal sequence from the view point of physical process. Based on the Newton's Law of Cooling, we set up a physical model to describe the process of the keys' temperature decreasing. Then the model is used to estimate keystroke time instants by maximizing likelihood function method. We estimate the password after getting keystroke time instants of each key. The possibility that cheaters have to steal people's password is also explored in the experiment. Based on our findings in the experiment, we give several pieces of practical advice for people to protect their password. Xiao-Ping Zhang 0002, Guangtao Zhai, Xiaokang Yang 0001, Wenhan Zhu, Xiao Gu 0001 |
ICIP | 2 |
| 2018 | Near-Infrared Fusion via Color Regularization for Haze and Color Distortion RemovalsabstractDifferent from conventional haze removal methods based on a single image, near-infrared imaging can provide two types of multimodal images: one is the near-infrared image and the other is the visible color image. These two images have different characteristics regarding color and visibility. The captured near-infrared image is haze-free, but it is grayscale, whereas the visible color image has colors, but it contains haze. There are serious discrepancies in terms of brightness and image structures between the near-infrared image and the visible color image. Due to this discrepancy, the direct use of the near-infrared image for haze removal causes a color distortion problem during near-infrared fusion. The key objective for the near-infrared fusion is therefore to remove the color distortion as well as the haze. To achieve this objective, this paper presents a new near-infrared fusion model that combines the proposed new color and depth regularizations with the conventional haze degradation model. The proposed color regularization sets the color range of the unknown haze-free image based on the combination of the two colors of the colorized near-infrared image and the captured visible color image. That is, the proposed color regularization can provide color information for the unknown haze-free color image. The new depth regularization enables the consecutively estimated depth maps not to be largely deviated, thereby transferring natural-looking colors and high visibility of the colorized near-infrared image into the preliminary dehazed version of the captured visible color image with color distortion and edge artifacts. Experimental results show that the proposed color and depth regularizations can help remove the color distortion and the haze simultaneously. The effectiveness of the proposed color regularization for the near-infrared fusion is verified by comparing it with other conventional regularizations. Chang-Hwan Son, Xiao-Ping Zhang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2018 | Quantized Kalman Filter Tracking in Directional Sensor NetworksabstractMoving target tracking in directional sensor networks has recently attracted attention by right of special directional sensing features. Unlike omnidirectional sensors, the directional sensor senses the target only in the direction of its orientation. It can provide quantized direction that indicates the presence or absence of the target in the sensing field, rather than just the analog measurement of sensing signal with respect to the detected target. A quantized Kalman filter (QKF) based on both quantized directions and analog ranging measurements is derived in the minimum mean-square error (MMSE) sense. Its performances of mean square estimation error (MSE) and complexity are also analyzed. Then, a reduced-complexity QKF of high-accuracy is pursued to facilitate its implementation. It is proved that the QKF yields a smaller MSE than the traditional extended Kalman filter (EKF) merely based on analog measurements. The posterior Cramér-Rao lower bound (PCRLB) is introduced as the performance measure. The performance advantages of the proposed QKF are demonstrated using Monte Carlo simulations in a target tracking application using ultrasonic ranging sensors. Xiaoqing Hu, Ming Bao, Xiao-Ping Zhang 0002, Sha Wen, Xiaodong Li 0002, Yu Hen Hu |
IEEE Trans. Mob. Comput. | 3 |
| 2017 | Learning a hierarchical spatio-temporal model for human activity recognitionabstractRecent works have shown that hierarchical models lead to significant improvement in human activity recognition, which can not only enhance descriptive capability, but also improve discriminative power. However, most existing methods exploit just one of the two advantages. In this paper, a new hierarchical spatio-temporal model (HSTM) is proposed to integrate feature learning into two-layer hierarchical classification model simultaneously. On the one hand, the two-layer model has sufficient descriptive capability. The bottom layer aims at capturing spatial relations in each frame and learning high-level representations, and the top layer utilizes these learned features to characterize temporal relations in the whole video sequence. On the other hand, the hierarchical model has strong discriminative power. Both spatial similarity and temporal similarity of activities are measured. Experimental results show that the HSTM can successfully recognize human activities with higher accuracies on one-person actions (KTH and UCF), human-human interactions (CASIA), and human-object interactional activities (Gupta). Wanru Xu, Zhenjiang Miao, Xiao-Ping Zhang 0002 |
ICASSP | 3 |
| 2017 | SAR image despeckling by combination of fractional-order total variation and nonlocal low rank regularizationabstractThis paper proposes a combinational regularization model for synthetic aperture radar (SAR) image despeckling. In contrast to most of the well-known regularization methods that only use one image prior property, the proposed combinational regularization model includes both fractional-order total variation (FrTV) regularization term and nonlocal low rank (NLR) regularization term. By characterizing the smoothness and nonlocal self-similarity property of the SAR image simultaneously, the proposed model, on the one hand, can better remove the noise in homogeneous regions of a noisy image, and on the other hand, can better preserve edges and geometrical features of the images during the despeckling process. Afterwards, an alternating direction method (ADM) is derived to efficiently solve the optimization problem in the proposed model. Experimental results demonstrate the good performance of the proposed model, both in removing SAR image speckles and preserving image texture and details. Gang Li 0008, Yu Liu 0005, Xiao-Ping Zhang 0002 |
ICIP | 4 |
| 2017 | Multimodal fusion via a series of transfers for noise removalabstractNear-infrared imaging has been considered as a solution to provide high quality photographs in dim lighting conditions. This imaging system captures two types of multimodal images: one is near-infrared gray image (NGI) and the other is the visible color image (VCI). NGI is noise-free but it is grayscale, whereas the VCI has colors but it contains noise. Moreover, there exist serious edge and brightness discrepancies between NGI and VCI. To deal with this problem, a new transfer-based fusion method is proposed for noise removal. Different from conventional fusion approaches, the proposed method conducts a series of transfers: contrast, detail, and color transfers. First, the proposed contrast and detail transfers aim at solving the serious discrepancy problem, thereby creating a new noise-free and detail-preserving NGI. Second, the proposed color transfer models the unknown colors from the denoised VCI via a linear transform, and then transfers natural-looking colors into the newly generated NGI. Experimental results show that the proposed transfer-based fusion method is highly successful in solving the discrepancy problem, thereby describing edges and textures clearly as well as removing noise completely on the fused images. Most of all, the proposed method is superior to conventional fusion methods and guided filtering, and even the state-of-the-art fusion methods based on scale map and layer decomposition. Chang-Hwan Son, Xiao-Ping Zhang 0002 |
ICIP | 2 |
| 2017 | Near-Infrared Coloring via a Contrast-Preserving Mapping ModelabstractNear-infrared gray images captured along with corresponding visible color images have recently proven useful for image restoration and classification. This paper introduces a new coloring method to add colors to near-infrared gray images based on a contrast-preserving mapping model. A naive coloring method directly adds the colors from the visible color image to the near-infrared gray image. However, this method results in an unrealistic image because of the discrepancies in the brightness and image structure between the captured near-infrared gray image and the visible color image. To solve the discrepancy problem, first, we present a new contrast-preserving mapping model to create a new near-infrared gray image with a similar appearance in the luminance plane to the visible color image, while preserving the contrast and details of the captured near-infrared gray image. Then, we develop a method to derive realistic colors that can be added to the newly created near-infrared gray image based on the proposed contrast-preserving mapping model. Experimental results show that the proposed new method not only preserves the local contrast and details of the captured near-infrared gray image, but also transfers the realistic colors from the visible color image to the newly created near-infrared gray image. It is also shown that the proposed near-infrared coloring can be used effectively for noise and haze removal, as well as local contrast enhancement. Chang-Hwan Son, Xiao-Ping Zhang 0002 |
IEEE Trans. Image Process. | 2 |
| 2017 | A Saliency Prior Context Model for Real-Time Object TrackingabstractReal-time object tracking has wide applications in time-critical multimedia processing areas such as motion analysis and human-computer interaction. It remains a hard problem to balance between accuracy and speed. In this paper, we present a fast real-time context-based visual tracking algorithm with a new saliency prior context (SPC) model. Based on the probability formulation, the tracking problem is solved by sequentially maximizing the computed confidence map of target location in each video frame. To handle the various cases of feature distributions generated from different targets and their contexts, we exploit low-level features as well as fast spectral analysis for saliency to build a new prior context model. Then, based on this model and a spatial context model learned online, a confidence map is computed and the target location is estimated. In addition, under this framework, the tracking procedure can be accelerated by the fast Fourier transform. Therefore, the new method generally achieves a real-time running speed. Extensive experiments show that our tracking algorithm based on the proposed SPC model achieves real-time computation efficiency with overall best performance comparing with other state-of-the-art methods. Cong Ma 0004, Zhenjiang Miao, Xiao-Ping Zhang 0002, Min Li 0025 |
IEEE Trans. Multim. | 3 |
| 2017 | A Hierarchical Spatio-Temporal Model for Human Activity RecognitionabstractThere are two key issues in human activity recognition: spatial dependencies and temporal dependencies. Most recent methods focus on only one of them, and thus do not have sufficient descriptive power to recognize complex activity. In this paper, we propose a hierarchical spatio-temporal model (HSTM) to solve the problem by modeling spatial and temporal constraints simultaneously. The new HSTM is a two-layer hidden conditional random field (HCRF), where the bottom-layer HCRF aims at describing spatial relations in each frame and learning more discriminative representations, and the top-layer HCRF utilizes these high-level features to characterize temporal relations in the whole video sequence. The new HSTM takes advantage of the bottom layer as the building blocks for the top layer and it aggregates evidence from local to global level. A novel learning algorithm is derived to train all model parameters efficiently and its effectiveness is validated theoretically. Experimental results show that the HSTM can successfully classify human activities with higher accuracies on single-person actions (UCF) than other existing methods. More importantly, the HSTM also achieves superior performance on more practical interactions, including human-human interactional activities (UT-Interaction, BIT-Interaction, and CASIA) and human-object interactional activities (Gupta video dataset). Wanru Xu, Zhenjiang Miao, Xiao-Ping Zhang 0002 |
IEEE Trans. Multim. | 3 |
| 2016 | A Weighted Variational Model for Simultaneous Reflectance and Illumination EstimationabstractWe propose a weighted variational model to estimate both the reflectance and the illumination from an observed image. We show that, though it is widely adopted for ease of modeling, the log-transformed image for this task is not ideal. Based on the previous investigation of the logarithmic transformation, a new weighted variational model is proposed for better prior representation, which is imposed in the regularization terms. Different from conventional variational models, the proposed model can preserve the estimated reflectance with more details. Moreover, the proposed model can suppress noise to some extent. An alternating minimization scheme is adopted to solve the proposed model. Experimental results demonstrate the effectiveness of the proposed model with its algorithm. Compared with other variational methods, the proposed method yields comparable or better results on both subjective and objective assessments. Xueyang Fu, Delu Zeng, Yue Huang 0001, Xiao-Ping Zhang 0002, Xinghao Ding |
CVPR | 4 |
| 2016 | Leaf Classification Utilizing a Convolutional Neural Network with a Structure of Single Connected Layer
Xiao-Ping Zhang 0002, Zhi-Kai Huang |
ICIC (2) | 3 |
| 2016 | Locally Biased Discriminative Clustering Method for Interactive Image Segmentation
Xianpeng Liang, Xiao-Ping Zhang 0002, Zhi-Kai Huang |
ICIC (2) | 2 |
| 2016 | Convolutional Neural Network Application on Leaf Classification
Yan-Hao Wu, Zhi-Kai Huang, Xiao-Ping Zhang 0002 |
ICIC (1) | 5 |
| 2016 | Saliency prior context model for visual trackingabstractIn this paper, we present a new FFT-based visual tracking algorithm based on a new model for computing prior context distribution via a spectral saliency approach. The tracking problem is formulated under a Bayesian framework where the statistical relationships between the features of the target and its spatio-temporal context are modeled. When building the context model, the prior distribution of the possible target position is an important part worth studying. To deal with various cases of distributions based on different attributes of the target and its context, we exploit low level saliency features by spectral analysis to compute prior distribution, not limited to center-surround weights. We show by extensive experiments that the performance of the new tracking algorithm based on the new saliency prior context (SPC) model achieves real-time computation efficiency with overall best location accuracy performance compared with other state-of-the-art methods. Cong Ma 0004, Zhenjiang Miao, Xiao-Ping Zhang 0002 |
ICIP | 3 |
| 2016 | A fusion-based method for single backlit image enhancementabstractIn this work, a new simple but effective fusion-based strategy for enhancing single backlit image is proposed. The fundamental idea of proposed strategy is to blend different features into a single one to improve the specific quality of image. Most of existing methods are based on the modification of histogram to enhance the contrast of low light images. However, the backlit images are different from low light images, which have wide dynamic ranges of light regions, thus the existing methods cannot achieve good enhanced results of backlit images. To improve performance of enhanced results, the proposed method considers numerous features of images and processes the dark and bright regions, respectively. Furthermore, proposed method introduces weight maps to increase the visibility. Experimental results show that proposed method is superior to existing methods, which achieves better results both in visual effects and processing time. Xueyang Fu, Xiao-Ping Zhang 0002, Xinghao Ding |
ICIP | 3 |
| 2016 | Visual data completion via local sensitive low rank tensor learningabstractThe problems of estimating missing values in visual data appear ubiquitously in computer vision applications including image inpainting, video inpainting, hyperspectral data recovery, and magnetic resonance imaging (MRI) data recovery. Recently, it is shown that tensor completion, which generalizes matrix completion to multiway data of higher order, could accurately estimate the overall data structure and achieves state-of-the-art performance for video completion. However, current tensor-based approaches implicitly assume that the partially observed video is globally low rank, which is too stringent for practical applications where the input video could include multiple heterogeneous episodes and the global correlation between frames is not high. To tackle this problem, we propose a novel local sensitive formulation of tensor learning where we assume instead that the video is inter-correlated in a local manner, leading to a representation of the observed tensor as a weighted sum of low-rank tensors. Computationally, we also design efficient scheme for solving the resulting learning problem based on the alternating direction method of multipliers (ADMM). Our experiments show improvements in prediction accuracy over classical approaches for visual data completion tasks. Qing-Yi Liu, Lin Zhu 0008, De-Shuang Huang, Xiao-Ping Zhang 0002, Zhi-Kai Huang |
IJCNN | 5 |
| 2016 | Feature extraction using maximum nonparametric margin projection
Bo Li 0002, Xiao-Ping Zhang 0002 |
Neurocomputing | 3 |
| 2016 | Constrained discriminant neighborhood embedding for high dimensional data feature extraction
Bo Li 0002, Xiao-Ping Zhang 0002 |
Neurocomputing | 3 |
| 2016 | Layer-Based Approach for Image Pair FusionabstractRecently, image pairs, such as noisy and blurred images or infrared and noisy images, have been considered as a solution to provide high-quality photographs under low lighting conditions. In this paper, a new method for decomposing the image pairs into two layers, i.e., the base layer and the detail layer, is proposed for image pair fusion. In the case of infrared and noisy images, simple naive fusion leads to unsatisfactory results due to the discrepancies in brightness and image structures between the image pair. To address this problem, a local contrast-preserving conversion method is first proposed to create a new base layer of the infrared image, which can have visual appearance similar to another base layer, such as the denoised noisy image. Then, a new way of designing three types of detail layers from the given noisy and infrared images is presented. To estimate the noise-free and unknown detail layer from the three designed detail layers, the optimization framework is modeled with residual-based sparsity and patch redundancy priors. To better suppress the noise, an iterative approach that updates the detail layer of the noisy image is adopted via a feedback loop. This proposed layer-based method can also be applied to fuse another noisy and blurred image pair. The experimental results show that the proposed method is effective for solving the image pair fusion problem. Chang-Hwan Son, Xiao-Ping Zhang 0002 |
IEEE Trans. Image Process. | 2 |
| 2016 | Blind Bleed-Through Removal for Scanned Historical Document Image With Conditional Random FieldsabstractScanned images of historical documents often suffer from bleed-through, which refers to the ink on one side seeping through the paper and appearing on the other side. In this paper, a new conditional random field (CRF)-based method is proposed to remove the bleed-through from the scanned images of historical images. The proposed method only requires the scanned image of one side, referred as a blind method. In general, the scanned historical document image is composed of three components: foreground, bleed-through, and background. By assuming Gaussian distributions of the three components, the proposed method establishes conditional probability distribution (CPD) models of the three components first. The parameters of the component CPD models are estimated based on an initial segmentation of the input image. Then, CRFs are used to capture the relations between observed pixels in the scanned image and the corresponding labels as well as the spatial relation between the adjacent labels. The belief propagation algorithm is used to calculate the probabilities of different labels for each pixel. Once the labeling is completed by choosing the most possible label for each pixel, the bleed-through component is removed from the input historical image by a random-filling inpainting algorithm. Experimental results on the real data set show that the proposed method preserves the foreground component very well and removes the bleed-through effectively. Bin Sun 0001, Shutao Li 0001, Xiao-Ping Zhang 0002, Jun Sun 0004 |
IEEE Trans. Image Process. | 3 |
| 2016 | Performance Limits and Geometric Properties of Array LocalizationabstractLocation-aware networks are of great importance and interest in both civil and military applications. This paper determines the localization accuracy of an agent, which is equipped with an antenna array and localizes itself using wireless measurements with anchor nodes, in a far-field environment. In view of the Cramér-Rao bound, we first derive the localization information for static scenarios and demonstrate that such information is a weighed sum of Fisher information matrices from each anchor-antenna measurement pair. Each matrix can be further decomposed into two parts: 1) a distance part with intensity proportional to the squared baseband effective bandwidth of the transmitted signal and 2) a direction part with intensity associated with the normalized anchor-antenna visual angle. Moreover, in dynamic scenarios, we show that the Doppler shift contributes additional direction information, with intensity determined by the agent velocity and the root mean squared time duration of the transmitted signal. In addition, two measures are proposed to evaluate the localization performance of wireless networks with different anchor-agent and array-antenna geometries, and both formulae and simulations are provided for typical anchor deployments and antenna arrays. Yanjun Han, Yuan Shen 0001, Xiao-Ping Zhang 0002, Moe Z. Win, Huadong Meng |
IEEE Trans. Inf. Theory | 3 |
| 2015 | Plant Leaf Recognition Based on Contourlet Transform and Support Vector Machine
Ze-Xue Li, Xiao-Ping Zhang 0002, Zhi-Kai Huang, Hao-Dong Zhu, Yong Gan |
ICIC (2) | 2 |
| 2015 | Hybrid Deep Learning for Plant Leaves Classification
Lin Zhu 0008, Xiao-Ping Zhang 0002, Xiaobo Zhou 0001, Zhi-Kai Huang, Yong Gan |
ICIC (2) | 3 |
| 2015 | Implementation of Plant Leaf Recognition System on ARM Tablet Based on Local Ternary Pattern
Gong-Sheng Xu, Jing-Hua Yuan, Xiao-Ping Zhang 0002, Zhi-Kai Huang, Hao-Dong Zhu, Yong Gan |
ICIC (2) | 3 |
| 2015 | Dynamic time-alignment k-means kernel clustering for time sequence clusteringabstractThis paper presents a novel method to cluster sequences by embedding a non-linear time alignment kernel function into kernel k-means. The time-alignment operation embeds the sequential pattern in the kernel function, allowing kernel k-means to be used to classify entire sequences. The method is evaluated with over 9800 videos and features from the LIRIS annotated creative commons emotional database. Our results show that the method works well in classifying sequences based on their affective content, and performs better than other unsupervised methods for clustering time series. In addition, this paper evaluates several methods abilities to map low-level features onto the valence arousal plane from the LIRIS database. The regression results also show that simple Ridge Regression had comparable performance to state-of-the-art regression methods. Joseph Santarcangelo, Xiao-Ping Zhang 0002 |
ICIP | 2 |
| 2015 | Structured feature-graph model for human activity recognitionabstractRecent works have shown that extracting and learning mid-level features lead to significant improvement in human activity recognition. Most existing methods represent activities as a collection of mid-level features and their spatio-temporal relations are completely neglected. Therefore, when scene contains interactional parts or high-level semantic actions, these mid-level features are not able to capture spatial structures as well as high order temporal relationships. In this paper, the activity is represented as a string of structured feature-graphs (SFGs) which models spatial structures and temporal structures simultaneously. A novel temporal graph kernel (TGK) is also proposed to measure similarity between two string representations. Experimental results show that our approach can successfully classify human activities with much higher accuracies for both single-person actions and human-human interactions. Wanru Xu, Zhenjiang Miao, Xiao-Ping Zhang 0002 |
ICIP | 3 |
| 2015 | Kernel-based mixture of experts models for linear regressionabstractThis paper proposes a novel kernel-based mixture of experts model for linear regression. The method is novel in that it formulates the mixture of experts model for linear regression so that kernel functions can be used. This allows the method to work directly in terms of kernels and avoids the explicit introduction of the feature vector, allowing one to use feature spaces of high, even infinite dimensionality. Other advantages of the model include the ability to take advantage of all the work related to kernels, a closed-form solution for maximization, as well as maintaining all the advantages of a linear expert. In this paper the supervised version is formulated. The model is verified and tested with simulated data. It was also found that the model had overall better performance than standard mixture of experts for regression on the well-known Boston Housing data set. Kernels used included polynomial, radial basis function and the ANOVA kernel. Joseph Santarcangelo, Xiao-Ping Zhang 0002 |
ISCAS | 2 |
| 2015 | Patch-based nonlocal dynamic MRI reconstruction with low-rank priorabstractCompressed sensing utilizes the sparsity of Magnetic resonance (MR) images to obtain accurate reconstructions from undersampled k-space data. In this paper, a novel nonlocal dynamic MRI reconstruction method with low-rank regularization is developed to exploit the spatiotemporal structural sparsity of a MRI sequence. The nonlocal prior and low rank prior are combined organically by grouping similar patches in both spatial and temporal domain. The low-rank regularization can be approximated by nuclear norm minimization solved by a singular value thresholding (SVT) method with adaptive thresholds estimation. The objective function is divided into several sub-problems that are easier to solve by alternative direction multiplier method (ADMM). Extensive experiments show that the new method outperforms commonly used classical dynamic MRI reconstruction algorithms. Liyan Sun, Jinchu Chen, Xiao-Ping Zhang 0002, Xinghao Ding |
MMSP | 3 |
| 2015 | The role of the Internet in changing industry competition
Xiao-Ping Zhang 0002 |
Inf. Manag. | 2 |
| 2015 | Nonparametric discriminant multi-manifold learning for dimensionality reduction
Bo Li 0002, Jun Li 0067, Xiao-Ping Zhang 0002 |
Neurocomputing | 3 |
| 2015 | A novel specific image scenes detection method
Yuxiang Xie, Xiao-Ping Zhang 0002, Xidao Luan, Li Liu 0002, Xin Zhang 0029 |
Multim. Tools Appl. | 2 |
| 2015 | A Probabilistic Method for Image Enhancement With Simultaneous Illumination and Reflectance EstimationabstractIn this paper, a new probabilistic method for image enhancement is presented based on a simultaneous estimation of illumination and reflectance in the linear domain. We show that the linear domain model can better represent prior information for better estimation of reflectance and illumination than the logarithmic domain. A maximum a posteriori (MAP) formulation is employed with priors of both illumination and reflectance. To estimate illumination and reflectance effectively, an alternating direction method of multipliers is adopted to solve the MAP problem. The experimental results show the satisfactory performance of the proposed method to obtain reflectance and illumination with visually pleasing enhanced results and a promising convergence rate. Compared with other testing methods, the proposed method yields comparable or better results on both subjective and objective assessments. Xueyang Fu, Yinghao Liao, Delu Zeng, Yue Huang 0001, Xiao-Ping Zhang 0002, Xinghao Ding |
IEEE Trans. Image Process. | 5 |
| 2015 | Efficient Heuristic Methods for Multimodal Fusion and Concept Fusion in Video Concept DetectionabstractSemantic models are widely used to bridge the semantic gap between low-level features and high-level features in video concept indexing. Multimodal fusion and concept fusion are two commonly used approaches in building semantic models. In the previous work, domain adaptation is neglected in multimodal fusion, and many probability maximization based and unsupervised concept fusion methods are counterintuitive since they do not incorporate subjective human intuition. In this paper, we present a new two-stage semantic model combining the multimodal fusion and the concept fusion incorporating human heuristics. In the multimodal fusion model, we employ a new generic unsupervised method, namely, domain adaptive linear combination (DALC), to update the linear combination (LC) weights by incorporating the differences of element distributions between training and testing domains. In the concept fusion model, a novel mechanical node equilibrium (NE) model is developed by using forces to model the concept correlations to update the score of concepts represented by nodes. It is intuitive and can incorporate multiple kinds of correlations simultaneously to construct more sophisticated semantic structure. Compared to other state-of-the-art supervised and unsupervised methods, the new model can use either unsupervised or supervised factors to significantly improve the mean inferred average precision (MAP) performance on all datasets. Zhenjiang Miao, Xiao-Ping Zhang 0002 |
IEEE Trans. Multim. | 3 |
| 2014 | A novel retinex based approach for image enhancement with illumination adjustmentabstractRetinex based algorithms have been widely used among in image enhancement. Since many retinex based algorithms remove illumination and regard the reflectance as enhancement, over-enhancement and unnaturalness are inevitable. In this paper, a novel retinex based image enhancement using illumination adjustment is proposed. Different from existing variational retinex models, a new model without the logarithmic transformation is established and can well preserve the edge. A fast alternating direction optimization method is used to solve this problem. After the decomposition of illumination and reflectance, a simple and effective post-processing method for illumination adjustment is adopted for the enhancement to make the result more natural. The proposed method can deal with many kinds of image, such as high dynamic range (HDR) images and non-uniform illumination images. Experimental results illustrate that the naturalness can be preserved while details are enhanced by the presented new approach. Xueyang Fu, Minghui LiWang, Yue Huang 0001, Xiao-Ping Zhang 0002, Xinghao Ding |
ICASSP | 5 |
| 2014 | Novel closed-form auxiliary variables based algorithms for sensor node localization using AOAabstractNode localization is a key issue for wireless sensor networks (WSNs). The triangulation method and the maximum likelihood (ML) estimator are usually adopted for angle of arrival (AOA) based node localization in WSNs. However, the localization accuracy of the triangulation is low, and the ML estimator requires a good initialization close to the true location to avoid the divergence problem. In this paper, we develop two efficient closed-form AOA based localization algorithms derived from effective auxiliary variables based method. First, we formulate the node localization problem as a linear least squares problem using auxiliary variables. Based on its closed-form solution, a new auxiliary variables based pseudo-linear estimator (AVPLE) is developed. Then, we further propose an auxiliary variables based total least square (AVTLS) estimator to improve the localization accuracy. In addition, we investigate the impact of the orientation of the unknown node on estimation performance of the new algorithms. Simulation results demonstrate that the new algorithms achieve much higher localization accuracy than the triangulation method and also avoid local minima and divergence problem in ML estimator. Moreover, the AVTLS estimator has higher localization accuracy than the AVPLE, and its localization accuracy remains robust when the orientation angle of the unknown node varies from 0 to 180 degrees. Huajie Shao, Xiao-Ping Zhang 0002, Zhi Wang 0003 |
ICASSP | 2 |
| 2014 | Nonparametric Discriminant Multi-manifold Learning
Bo Li 0002, Jun Li 0067, Xiao-Ping Zhang 0002 |
ICIC (1) | 3 |
| 2014 | A retinex-based enhancing approach for single underwater imageabstractSince the light is absorbed and scattered while traveling in water, color distortion, under-exposure and fuzz are three major problems of underwater imaging. In this paper, a novel retinex-based enhancing approach is proposed to enhance single underwater image. The proposed approach has mainly three steps to solve the problems mentioned above. First, a simple but effective color correction strategy is adopted to address the color distortion. Second, a variational framework for retinex is proposed to decompose the reflectance and the illumination, which represent the detail and brightness respectively, from single underwater image. An effective alternating direction optimization strategy is adopted to solve the proposed model. Third, the reflectance and the illumination are enhanced by different strategies to address the under-exposure and fuzz problem. The final enhanced image is obtained by combining use the enhanced reflectance and illumination. The enhanced result is improved by color correction, lightens dark regions, naturalness preservation, and well enhanced edges and details. Moreover, the proposed approach is a general method that can enhance other kinds of degraded image, such as sandstorm image. Xueyang Fu, Peixian Zhuang, Yue Huang 0001, Yinghao Liao, Xiao-Ping Zhang 0002, Xinghao Ding |
ICIP | 5 |
| 2014 | A new robust context-based dense CRF model for image labelingabstractFully-connected conditional random fields (CRF) models have recently been developed for image labeling task to incorporate interactions of all pairs of pixels in the image. Efficient inference in fully-connected models is very sensitive to initialization of the unary potentials. In this paper, we propose a new robust context-based fully-connected CRF model which alleviates initialization sensitivity of inference in dense CRFs. The new model integrates an extra hidden node that accounts for the overall context of the image and is connected to all other pixel nodes. By incorporating the new context node in CRF graph, we maximize probability of labeling configuration jointly with the image's context. Therefore, wrong initializations of objects that contradict the overal context could be refined. We define the context-based unary and pairwise potentials and further derive the inference algorithm for the proposed model based on the mean field approximation method. We run experiments over the benchmark MSRC image database and demonstrate that the new model improves object recognition accuracy by about 21%. We show where the conventional CRF model is impeded by wrong initialization of unary potentials, the proposed model identifies the labels correctly. Maryam Nematollahi, Xiao-Ping Zhang 0002 |
ICIP | 2 |
| 2014 | Automatic age recommendation system for children's video contentabstractThis paper presents a novel automatic method to determine the appropriate age of video content in a video database geared to children. When combined with classical features the system improves accuracy rate for more than 0.13 for the same type of classifier in determining the age category of content for children between the ages of three to six years old. The main novelty of the system is that it utilizes high level audio features related to the cognitive capacity of children. These novel features gage the cognitive ability of the intended audience by quantifying the structure of the language. These novel features include syllable rate, word rate, language complexity and noise jumps. The feature extraction methods are also novel in that we count the number of syllables and words using relatively computationally inexpensive signal processing techniques forgoing complex speech recognition. The presented new method is tested using multiple classifiers on a commercial video database. Joseph Santarcangelo, Xiao-Ping Zhang 0002 |
ISCAS | 2 |
| 2014 | Convergence analysis of multiple imputations particle filters for dealing with missing data in nonlinear problemsabstractWe apply multiple imputations particle filter (MIPF) to deal with non-linear state estimation problem in the presence of missing data. We use imputations to replace the missing data. We present the convergence analysis of MIPF and show that it is almost surely convergent.We also present examples with a nonstationary growth model and dual-sensor bearing-only tracking, which demonstrate that MIPF can effectively deal with missing data in nonlinear problems. Xiao-Ping Zhang 0002, Ahmed Shaharyar Khwaja, Ji-an Luo, Alon Shalev Housfater, Alagan Anpalagan |
ISCAS | 1 |
| 2014 | A fusion-based enhancing approach for single sandstorm imageabstractIn this paper, a novel image enhancing approach focuses on single sandstorm image is proposed. The degraded image has some problems, such as color distortion, low-visibility, fuzz and non-uniform luminance, due to the light is absorbed and scattered by particles in sandstorm. The proposed approach based on fusion principles aims to overcome the aforementioned limitations. First, the degraded image is color corrected by adopting a statistical strategy. Then two inputs, which represent different brightness, are derived only from the color corrected image by applying Gamma correction. Three weighted maps (sharpness, chromaticity and prominence), which contain important features to increase the quality of the degraded image, are computed from the derived inputs. Finally, the enhanced image is obtained by fusing the inputs with the weight maps. The proposed method is the first to adopt a fusion-based method for enhancing single sandstorm image. Experimental results show that enhanced results can be improved by color correction, well enhanced details and local contrast while promoted global brightness, increasing the visibility, naturalness preservation. Moreover, the proposed algorithm is mostly calculated by per-pixel operation, which is appropriate for real-time applications. Xueyang Fu, Yue Huang 0001, Delu Zeng, Xiao-Ping Zhang 0002, Xinghao Ding |
MMSP | 4 |
| 2014 | Classifying harmful children's content using affective analysisabstractThis paper categorizes children's videos according to an expertly assigned predefined positive or negative cognitive impact category. The method uses affective features to determine if a video belongs to an expertly assigned predefined positive or to a negative cognitive impact category. The work demonstrates that simple affective features outperform more complex systems in determining if content belongs to the positive or negative cognitive impact category. The work is tested on a set of videos that have been classified as having a short term or long term measurable negative or positive impact on cognition based on cited psychological literature. It found that affective analysis had superior performance using less features than state of the art video genre classification systems. It also found that arousal features performed better than valence features. Joseph Santarcangelo, Xiao-Ping Zhang 0002 |
MMSP | 2 |
| 2014 | Motion Parameter Estimation and Focusing From SAR Images Based on Sparse ReconstructionabstractThis letter presents a new motion parameter estimation method using synthetic aperture radar (SAR) images based on sparse reconstruction. The method uses orthogonal matching pursuit, which correlates a SAR image with elements in a reference basis. The reconstruction result provides a high-resolution focused image and corresponding motion parameters. It does not require narrow-beamwidth assumption and prior motion information, as compared with the existing method based on matching pursuit. We design the reference basis by theoretically analyzing the performance degradation due to parameter mismatch. We further calculate required position and velocity parameter resolutions in the basis. The calculated resolutions are validated by means of simulations. Imaging examples with real SAR images of a scene acquired over Ottawa, Canada, show the effectiveness of our method. Ahmed Shaharyar Khwaja, Xiao-Ping Zhang 0002 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2014 | Bayesian Nonparametric Dictionary Learning for Compressed Sensing MRIabstractWe develop a Bayesian nonparametric model for reconstructing magnetic resonance images (MRIs) from highly undersampled k -space data. We perform dictionary learning as part of the image reconstruction process. To this end, we use the beta process as a nonparametric dictionary learning prior for representing an image patch as a sparse combination of dictionary elements. The size of the dictionary and patch-specific sparsity pattern are inferred from the data, in addition to other dictionary learning variables. Dictionary learning is performed directly on the compressed image, and so is tailored to the MRI being considered. In addition, we investigate a total variation penalty term in combination with the dictionary learning model, and show how the denoising property of dictionary learning removes dependence on regularization parameters in the noisy setting. We derive a stochastic optimization algorithm based on Markov chain Monte Carlo for the Bayesian model, and use the alternating direction method of multipliers for efficiently performing total variation minimization. We present empirical results on several MRI, which show that the proposed regularization framework can improve reconstruction accuracy over other methods. Yue Huang 0001, John W. Paisley, Xinghao Ding, Xueyang Fu, Xiao-Ping Zhang 0002 |
IEEE Trans. Image Process. | 6 |
| 2014 | Illumination Robust Video Foreground Prediction Based on Color RecoveringabstractVideo foreground prediction is a technique to estimate the probability of each pixel being foreground in current frame based on a foreground segmentation result of its previous frame. Existing foreground prediction algorithms usually assume that the illumination conditions are constant for consecutive frames. Therefore, they cannot predict foreground accurately when the illumination condition changes sharply between video frames. In this paper, a new robust video foreground prediction algorithm is proposed based on color recovering, which is derived based on an observation that the illumination changes are locally smooth. By integrating color recovering with an optical flow estimation algorithm and an opacity propagation algorithm, the negative impact of the illumination changes could be removed. Experimental results show that the proposed algorithm can get more accurate results for videos with illumination changes compared with the existing foreground prediction algorithms. Yanli Wan, Zhenjiang Miao, Xiao-Ping Zhang 0002 |
IEEE Trans. Multim. | 3 |
| 2013 | A new subband information fusion method for wideband DOA estimation using sparse signal representationabstractWe present a new subband information fusion (SIF) method for wideband direction-of-arrival (DOA) estimation using single sparse signal representation of multiple frequency-based measurement vectors. The problem of wideband DOA estimation using SIF method is to jointly utilize all the frequency bin information to recover a single sparse indicative vector (SIV). The SIF method belongs to the sparse signal representation domain and therefore it will suffer from two cases of ambiguity: algebraic aliasing and spatial aliasing. We show that these two categories of ambiguity can be reduced by combining all the frequency components. The SIF algorithm is then proposed and the SIV is recovered iteratively. The numerical simulations are performed to illustrate that the SIF method has superior performances. Ji-an Luo, Xiao-Ping Zhang 0002, Zhi Wang 0003 |
ICASSP | 2 |
| 2013 | A new passive source localization method using AOA-GROA-TDOA in wireless sensor array networks and its Cramér-Rao bound analysisabstractIn this paper, a new Cramér-Rao lower bound (CRLB) is derived for passive source localization based on angles-of-arrival (AOAs), gain ratios of arrival (GROAs) and time differences of arrival (TDOAs) in a wireless sensor array network. The derived CRLB using AOA-GROA-TDOA (AGT) is reduced to the one using AOA-GROA if no coherence exists across the arrays and lower than the CRLB using AOA-only. When the coherence is considered, the CRLB using AGT measurements is consistently lower than the other known bounds using AOA-only, TDOA-only and AOA-TDOA. Ji-an Luo, Xiao-Ping Zhang 0002, Zhi Wang 0003 |
ICASSP | 2 |
| 2013 | Compressed sensing MRI with Bayesian dictionary learningabstractWe present an inversion algorithm for magnetic resonance images (MRI) that are highly undersampled in k-space. The proposed method incorporates spatial finite differences (total variation) and patch-wise sparsity through in situ dictionary learning. We use the beta-Bernoulli process as a Bayesian prior for dictionary learning, which adaptively infers the dictionary size, the sparsity of each patch and the noise parameters. In addition, we employ an efficient numerical algorithm based on the alternating direction method of multipliers (ADMM). We present empirical results on two MR images. Xinghao Ding, John W. Paisley, Yue Huang 0001, Xianbo Chen, Xiao-Ping Zhang 0002 |
ICIP | 6 |
| 2013 | Sensitivity analysis of compressed sensing ISAR imaging to rotational acceleration rate mismatchabstractIn this paper, we present a sensitivity analysis of compressed sensing inverse synthetic aperture radar imaging to a mismatch of rotational acceleration rate. This analysis shows that the mismatch leads to reconstruction performance degradation due to intra-pixel and inter-pixel shifts. Numerical examples show that increasing mismatch leads to an increase of reconstruction error due to intra-pixel shift. This error reaches a maximum constant level once the shift becomes greater than a pixel. We calculate the mismatch limits required to reduce this error. Imaging examples validate these limits by showing that the reconstructed images are well constructed if the mismatch is below these limits. Ahmed Shaharyar Khwaja, Xiao-Ping Zhang 0002 |
ICIP | 2 |
| 2013 | Pan-sharpening based on nonparametric Bayesian adaptive dictionary learningabstractPan-sharpening based on compressed sensing (CS) theory has been widely studied in recent years. In this paper, we present a novel CS-based pan-sharpening method based on nonparametric Bayesian adaptive dictionary learning. In contrast to existing optimization methods, the proposed method adaptively infers parameters such as dictionary size, patch sparsity and noise variances. In addition, high resolution multiband images, which are unavailable in practice, are not required to learn the dictionary anymore. An IKONOS satellite image is employed to validate the method. Both visual results and quality metrics demonstrate that proposed method is able to achieve higher spatial and spectral resolution simultaneously, compared with other well-known methods. Yue Huang 0001, John W. Paisley, Xinghao Ding, Xiao-Ping Zhang 0002 |
ICIP | 5 |
| 2013 | Compressed sensing SAR moving target imaging in the presence of basis mismatchabstractIn this paper, we present compressed sensing (CS) imaging of synthetic aperture radar (SAR) moving targets when there is a basis mismatch. We use Gaussian-Bernoulli prior for CS SAR reconstruction. Analysis using Cramer-Rao bounds show that using the prior results in performance improvement in the presence of mismatch compared to Laplacian prior. We further show that the more serious velocity mismatch can be dealt with using an iterative procedure. Simulation results show the effectiveness of the proposed method. Ahmed Shaharyar Khwaja, Xiao-Ping Zhang 0002 |
ISCAS | 2 |
| 2013 | Label propagation based supervised locality projection analysis for plant leaf classification
Shanwen Zhang, Ying-Ke Lei, Tianbao Dong, Xiao-Ping Zhang 0002 |
Pattern Recognit. | 4 |
| 2013 | An ICA Mixture Hidden Conditional Random Field Model for Video Event ClassificationabstractIn this paper, a hidden conditional random field (HCRF) model with independent component analysis (ICA) mixture feature functions is developed for video event classification. Video content analysis problems can be modeled using graphical models. The hidden Markov model (HMM) is a commonly used graphical model, but the HMM has several limitations such as the assumption of observation independence, the form of observation distribution and the Markov chain interaction. Unlike the HMM, the HCRF is a discriminative model without conditional independence assumption of observations, and is more suitable for video content analysis. We formulate the video content analysis problem using a new HCRF framework based on the temporal interactions between video frames. In addition, according to the non-Gaussian property of video event features, a new feature function using the likelihoods of ICA mixture components is proposed for local observation to further enhance the HCRF model. The discriminative power of the HCRF and representation power of the ICA mixture for non-Gaussian distributions are combined in the new model. The new model is applied to the challenging bowling and golf event classifications as case studies. The simulation results support the analysis that the new ICA mixture HCRF (ICAMHCRF) outperforms the existing mixture HMM models in terms of classification accuracy. Xiaofeng Wang 0008, Xiao-Ping Zhang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2012 | Compressed sensing based image formation of SAR/ISAR data in presence of basis mismatchabstractThis paper examines compressed sensing (CS) based image formation of synthetic aperture radar (SAR) and inverse synthetic aperture radar (ISAR) data for sparse scenes containing moving targets. We consider basis mismatch for the case when the basis used for reconstruction is different from the actual one in which the reconstructed data are sparse. We use orthogonal matching pursuit (OMP) algorithm for reconstruction and show using simulated data that error between original and reconstructed data increases in presence of basis mismatch. We also show that a certain level of basis mismatch in range velocity, positions and chirp rate is acceptable to achieve reasonable image formation. Ahmed Shaharyar Khwaja, Xiao-Ping Zhang 0002 |
ICIP | 2 |
| 2012 | Reconstruction of compressively sensed complex-valued terahertz dataabstractTraditionally, compressed sensing (CS) has been presented considering real-valued data. There has been some recent interest in reconstruction algorithms for CS that can handle and take advantage of complex-valued data. This kind of data is common in terahertz (THz) imaging. CS reconstruction in such a case can benefit from extra information provided by phase or real and imaginary parts of the data. In this paper, we present an algorithm based on iterative shrinkage/thesholding that takes into account this extra information for reconstruction of compressively sensed complex-valued data. We use curvelets as sparsity-promoting basis for real and imaginary parts of THz data and show using actual THz data that reconstruction performance is improved. Moreover, compared to existing methods, the proposed method is computationally efficient, flexible and suitable for large-size data. Ahmed Shaharyar Khwaja, Xiao-Ping Zhang 0002 |
ISCAS | 2 |
| 2012 | Direction-of-arrival estimation using sparse variable projection optimizationabstractWe propose a new low complexity direction-of-arrival (DOA) estimation method based on sparse variable projection (SVP) optimization. This method estimates an indicative sparse vector that indicates the locations of DOA from each visual sources corresponding to DOA sampling space and is particular useful to simplify the multiple measurement vector (MMV) problem as a single indicative sparse vector recovery problem. The indicative sparse vector can be recovered by adding additional sparsity measure information. We use ℓp(p ≤ 1) norm and smoothed approximate ℓ0norm to regularize the SVP function, so that we can formulate the SVP optimization as an unconstrained optimization problem and we solve it efficiently using quasi-Newton method. The experimental results demonstrate that our method has much lower complexity by comparing with a standard Regularized M-FOCUSS algorithm. Ji-an Luo, Xiao-Ping Zhang 0002, Zhi Wang 0003 |
ISCAS | 2 |
| 2012 | Bimodal Discriminant Projection Analysis for gait recognitionabstractAs for gait recognition, we propose a new discriminant dimensionality reduction method, named Bimodal Discriminant Projection Analysis (BDPA) algorithm. In BDPA, a weight path-based similarity measure is designed, the intra-class scatter matrix is constructed by the weight, while the inter-class scatter matrix is constructed by the heat kernel function. Compared with the classical methods, such as Multimodal Preserving Embedding (MPE) and Minimax Risk Criterion methods, the proposed method can preserve within-class neighborhood geometry and extract between-class relevant structures for recognition by minimizing the intra-class scatter and maximizing the inter-class scatter. The experimental results on real-world gait data show that BDPA is effective and feasible for gait recognition. Shanwen Zhang, Xiao-Ping Zhang 0002, Chuanlei Zhang |
MMSP | 2 |
| 2012 | Ice hockey shooting event modeling with mixture hidden Markov model
Xiaofeng Wang 0008, Xiao-Ping Zhang 0002 |
Multim. Tools Appl. | 2 |
| 2012 | Coupled Observation Decomposed Hidden Markov Model for Multiperson Activity RecognitionabstractMultiperson activity recognition in videos is a challenging task, due to the complexity of interactions among multiple persons. In this paper, a new statistical model, named coupled observation decomposed hidden Markov model (CODHMM), is presented to model multiperson activities in videos. A human activity that involves multiple persons is analyzed in two levels: the individual level that describes each individual's motion details and the interaction level that expresses the shared information among multiple persons. The two levels are modeled by two hidden Markov chains that are interdependent and interact with each other. The observation in each chain at each time slice is decomposed into subobservations according to the number of features and the number of persons. For each activity to be recognized, a CODHMM is built and model parameters are learnt by a generalized expectation maximization (EM) algorithm. Given an input video that contains an unknown activity, maximum likelihood algorithms are developed to classify it into one of the learnt activity categories. Experimental results show that the CODHMM can successfully classify human activities involving multiple persons with high accuracy and low computations. Zhenjiang Miao, Xiao-Ping Zhang 0002, Yuan Shen 0004 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2011 | A novel vector quantization-based video summarization method using independent component analysis mixture modelabstractIn this paper, we present a new independent component analysis mixture vector quantization (ICAMVQ) method to summarize the video content. In particular, independent component analysis (ICA) is applied first to explore the characteristics of video data and build a compact 2D feature space. A new ICAMVQ method is then developed to find the optimized quantization codebook in ICA subspace. The optimal codebook size is determined by Bayes information criterion (BIC). The frames that are the nearest neighbors to the quanta in the ICAMVQ codebook are sampled to summarize the video. A 2D kD-tree is employed to index the feature space and accelerate the nearest-neighbor search. Experimental results show that our method is practically effective and computationally efficient to build a video summarization system. Xiao-Ping Zhang 0002 |
ICASSP | 2 |
| 2011 | Video thumbnail extraction using video time density function and independent component analysis mixture modelabstractIn this paper, we propose a new vector quantization method to create video thumbnail. In particular, we employ video time density function (VTDF) to explore the temporal characteristics of video data first. A VTDF-based temporal quantization is then applied to segment the whole video in time domain. The optimal number of segments is obtained by a temporal mean square error (TMSE)-based criterion. We employ independent component analysis (ICA) to each temporal segment for feature extraction and build a compact 2D feature space. An ICA mixture-based vector quantization method is developed to explore the spatial characteristics of video data. The optimal number of ICA mixture components is determined by Bayes information criterion (BIC). The video frames that are the nearest neighbors to the quantization codebook are sampled to generate the video thumbnails. Experimental results show that our method is computationally efficient and practically effective to create video thumbnails. Xiao-Ping Zhang 0002 |
ICASSP | 2 |
| 2011 | A new video similarity measure model based on video time density function and dynamic programmingabstractIn this paper, we propose a novel video similarity measure model using video time density function (VTDF) and dynamic programming. First, we employ VTDF to describe the density of video activities in time domain by calculating the inter-frame mutual information. Second, a temporal partition solution is applied to divide each video sequence into equi-sized temporal segments. Third, a new VTDF based similarity measure using correlation is calculated to measure the similarity between two temporal segments. Fourth, dynamic programming is then developed to find the optimal non-linear mapping between two video sequences. A new normalized similarity measure function combing both visual characteristics and temporal information together is to evaluate the semantic similarity of two video sequences. Experimental results show that the proposed measurement model is effective to explore the semantic similarity of video sequences. Xiao-Ping Zhang 0002, Alexander C. Loui |
ICASSP | 2 |
| 2011 | A novel rate-distortion optimization method of H.264/AVC intra coderabstractIn H.264/AVC, the concept of rate-distortion optimization (RDO) mode decision has proven to be an important coding tool. But the complexity and computation load of RDO technique is extremely high. In this paper, we propose an enhanced low complexity cost function for H.264/AVC intra 4x4 mode selections. The enhanced cost function uses sum of absolute Hadamard-transformed differences (SATD) and variance of the residual block to estimate distortion part of the cost function. A threshold based large coefficients count is also used for estimating the bit-rate part. The proposed method improves the rate-distortion (RD) performance of the conventional fast cost functions while maintaining low complexity requirement. Mohammed Golam Sarwer, Q. M. Jonathan Wu, Xiao-Ping Zhang 0002 |
ICIP | 3 |
| 2011 | A content-based video fast-forward playback method using video time density function and rate distortion theoryabstractIn this paper, we propose a new video summary method using video time density function (VTDF) and rate distortion theory. The whole system has two main modules, processing and playing. In the processing part, we apply VTDF to describe the temporal dynamics of video data first. A VTDF-based temporal quantization method is then developed to find the best quanta and partition in time domain. The optimal quanta are used to extract the representative video frames. A temporal mean square error (TMSE) is introduced by using rate-distortion theory to evaluate the quantization performance. In the playing module, we develop a video player to only play all sampled frames in its intelligent fast-forward mode. The built video player can allow users to do fast-forward playback based on the semantic video content, which demonstrates the feasibility of proposed method in practice. Xiao-Ping Zhang 0002, Alexander C. Loui |
ICME | 2 |
| 2011 | A smart video player with content-based fast-forward playbackabstractIn this paper, we develop a video player to allow the users to do fast-forward playback based on the semantic video content. The whole system has two modules, processing and playing. In the processing part, we present a video time density function (VTDF) to describe the temporal dynamics of video data first. A VTDF-based temporal quantization method is then developed to find the best quanta and partition in the time domain. The optimal quanta are used to extract key frames. The optimal number of key frames is determined by a temporal mean square error (TMSE)-based criterion. In the playing module, we combine the key frame sequence and a set of parameters together and feed them into a triangle-based transition function to generate the sampled frames in a non-uniform way. A built video player will play all sampled frames in its intelligent fast-forward mode for a given fast-forward speed factor. The implementation of video player demonstrates the feasibility of proposed method in practice. Xiao-Ping Zhang 0002 |
ACM Multimedia | 2 |
| 2011 | Graphical probabilistic modeling and applications in multimedia content analysisabstractGraphical probabilistic models play an important role in modern machine learning and pattern recognition. In this half-day tutorial, we introduce the fundamentals of graphical probabilistic modeling, including Bayesian networks, Markov random fields, hidden conditional Random field, etc. We also discuss some applications of graphical probabilistic modeling in multimedia content analysis, for example, video segmentation, video event detection, video sequence matching, and image labeling. The audience will learn the basic of graphical models, and get familiar with the state of the art in multimedia content analysis systems. Xiao-Ping Zhang 0002, Zhu Liu 0001 |
ACM Multimedia | 1 |
| 2010 | A new Laplacian mixture conditional random field model for image labelingabstractIn this paper we present a novel conditional random field (CRF) model based on Laplacian mixtures for image labeling. Nature images posses many spatial regularities that can be efficiently modeled by probabilistic graphical models such as CRF. Usually hundreds of features and several types of feature functions are used together which increases computational complexity and makes the training difficult to converge. We propose a new Laplacian mixture CRF model, which simplifies the training and inference process without losing labeling accuracy. The belief propagation inference and stochastic gradient descent training are formulated accordingly for the new model. The experimental results demonstrate that the new approach achieves better classification accuracy than the baseline CRF and comparable results with the state-of-the-art complex models. Xiaofeng Wang 0008, Xiao-Ping Zhang 0002 |
ICASSP | 2 |
| 2010 | A New Hierarchical Key Frame Tree-Based Video Representation Method Using Independent Component Analysis
Xiao-Ping Zhang 0002 |
ICIC (2) | 2 |
| 2010 | Gaussian mixture vector quantization-based video summarization using independent component analysisabstractIn this paper, we propose a new Gaussian mixture vector quantization (GMVQ)-based method to summarize the video content. In particular, in order to explore the semantic characteristics of video data, we present a new feature extraction method using independent component analysis (ICA) and color histogram difference to build a compact 3D feature space first. A new GMVQ method is then developed to find the optimized quantization codebook. The optimal codebook size is determined by Bayes information criterion (BIC). The video frames that are the nearest-neighbours to the quanta in the GMVQ quantization codebook are sampled to summarize the video content. A kD-tree-based nearest-neighbour search strategy is employed to accelerate the search procedure. Experimental results show that our method is computationally efficient and practically effective to build a content-based video summarization system. Xiao-Ping Zhang 0002 |
MMSP | 2 |
| 2010 | Statistical Modeling in the Wavelet Domain for Compact Feature Extraction and Similarity Measure of ImagesabstractImage feature extraction and similarity measure in feature space are active research topics. They are basic components in a content-based image retrieval (CBIR) system. In this letter, we present a new statistical model-based image feature extraction method in the wavelet domain and a novel Kullback divergence-based similarity measure. First, a Gaussian mixture model (GMM) and a more systematic generalized Gaussian mixture model (GGMM) are employed to describe the statistical characteristics of the wavelet coefficients and the model parameters are employed to construct a compact image feature space. A nontrivial expectation-maximization (EM) algorithm for the GGMM model is derived. Subsequently, a new Kullback divergence-based similarity measure with low-computation cost is derived and analyzed. The Brodatz texture image database and some other image databases are used to evaluate the retrieval performance based on the presented new methods. Experimental results indicate that the GMM and the GGMM-based image texture features are very effective in representing multiscale image characteristics and that the new methods outperforms other conventional wavelet-based methods in retrieval performance with a comparable level of computational complexity. It is also demonstrated that for image features extracted by the new statistical models, the similarity measure based on Kullback divergence is more effective than conventional similarity measures. Xiao-Ping Zhang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2009 | A new localized superpixel Markov random field for image segmentationabstractIn this paper, we present a novel localized Markov random field (MRF) method based on superpixels for region segmentation. Early vision problems could be formulated as pixel labeling using MRF. But the local interaction in MRF is limited to pixel label comparison. We propose a new localized superpixel Markov random field (SMRF) model to incorporate local data interaction in unsupervised parameter learning. The advantages of the new model include computational efficiency by using superpixel structure and its ability to integrate local knowledge in the learning process. Quantitative evaluation and visual effects show that the new model achieves not only better segmentation accuracy but also lower computational cost than the baseline pixel based model. Xiaofeng Wang 0008, Xiao-Ping Zhang 0002 |
ICME | 2 |
| 2008 | A new implementation of trellis coded quantization based data hidingabstractThis paper discusses the construction and implementation problem of trellis coded quantization (TCQ) based data hiding. We explore the robustness and distortion of data hiding by analyzing its duality with distributed source coding. Based on our analysis a new implementation of the powerful trellis coded modulation (TCM) and TCQ data hiding scheme is presented. It simplifies the construction process with only one trellis and achieves a good tradeoff between robustness and distortion by embedding the information in the middle input of TCQ. Simulation is conducted for Gaussian, Laplacian and real image sources under additive Gaussian noise attack. The results demonstrate the effectiveness of the new implementation. Xiaofeng Wang 0008, Xiao-Ping Zhang 0002 |
ICASSP | 2 |
| 2008 | Preface to the special issue on new achievements in pervasive and interactive multimedia systems and applications
Marco Roccetti, Zhu Liu 0001, Marco Furini, Xiao-Ping Zhang 0002, Hong Heather Yu |
Multim. Tools Appl. | 4 |
| 2008 | A Novel Look-Up Table Design Method for Data Hiding With Reduced DistortionabstractLook-up table (LUT)-based data hiding is a simple and efficient technique to hide secondary information (watermark) into multimedia work for various applications such as copyright protection, transaction tracking or content annotation. This paper studies the distortion introduced by a general LUT-based data hiding. We find that designing LUT according to the distribution of host data and watermark data can greatly reduce the distortion of LUT embedding. A new practical reduced-distortion LUT design method is developed for robust data hiding. The new method is applied in a wavelet domain image data hiding system and only significant wavelet coefficients are used to embed the watermark. A Gaussian mixture model and a related expectation-maximization algorithm-based method are employed to model the statistical distribution of the host image. The statistical model is used to select significant coefficients of the host image for data hiding. The experimental results show that compared to the conventional odd-even LUT embedding method, the presented new LUT data hiding algorithm provides average 1. 5-2. 5 dB PSNR improvement and better robustness for image watermarking. Xiao-Ping Zhang 0002, Kan Li 0004, Xiaofeng Wang 0008 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2008 | An ICA Mixture Hidden Markov Model for Video Content AnalysisabstractIn this paper, a new theoretical framework based on hidden Markov model (HMM) and independent component analysis (ICA) mixture model is presented for content analysis of video, namely ICAMHMM. Unlike the Gaussian mixture observation model commonly used in conventional HMM applications, the observations in the new ICAMHMM are modeled as a mixture of non-Gaussian components. Each non-Gaussian component is formulated by an ICA mixture, reflecting the independence of different components across video frames. In addition, to construct a compact feature space to represent a video frame, ICA is applied on video frames and the ICA coefficients are used to form a compact 2-D feature subspace that makes the subsequent modeling computationally efficient. The model parameters can be identified using supervised learning by the training sequences. The new re-estimation learning formulae of iterative ICAMHMM parameter estimation are derived based on a maximum likelihood function. Employing the identified model, maximum likelihood algorithms are developed to detect and recognize video events. As a case study, golf video sequences are used to test the effectiveness of the proposed algorithm. Experimental results show that the presented method can effectively detect and recognize the recurrent event patterns in video data. The presented new ICAMHMM is generic and can be applied to sequential data analysis in other applications. Xiao-Ping Zhang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2008 | Color Image Watermarking Using Multidimensional Fourier TransformsabstractThis paper presents two vector watermarking schemes that are based on the use of complex and quaternion Fourier transforms and demonstrates, for the first time, how to embed watermarks into the frequency domain that is consistent with our human visual system. Watermark casting is performed by estimating the just-noticeable distortion of the images, to ensure watermark invisibility. The first method encodes the chromatic content of a color image into the CIE chromaticity coordinates while the achromatic content is encoded as CIE tristimulus value. Color watermarks (yellow and blue) are embedded in the frequency domain of the chromatic channels by using the spatiochromatic discrete Fourier transform. It first encodes and as complex values, followed by a single discrete Fourier transform. The most interesting characteristic of the scheme is the possibility of performing watermarking in the frequency domain of chromatic components. The second method encodes the components of color images and watermarks are embedded as vectors in the frequency domain of the channels by using the quaternion Fourier transform. Robustness is achieved by embedding a watermark in the coefficient with positive frequency, which spreads it to all color components in the spatial domain and invisibility is satisfied by modifying the coefficient with negative frequency, such that the combined effects of the two are insensitive to human eyes. Experimental results demonstrate that the two proposed algorithms perform better than two existing algorithms - ac- and discrete cosine transform-based schemes. Tsz Kin Tsui, Xiao-Ping Zhang 0002, Dimitrios Androutsos |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2007 | Joint Iterative Demodulation and Decoding of Differential Frequency Hopping SignalsabstractIn this paper, joint iterative demodulation and decoding of differential frequency hopping (DFH) signals based on the Bahl-Cocke-Jelinek-Raviv (BCJR) algorithm is considered. The DFH is a nonlinear modulation with memory in the frequency sequence of the successive transmitted symbols, which can be represented by a trellis diagram. The proposed receiving scheme is primarily composed of an a posteriori probability (APP) decoder/filter and an APP demodulator, by which the extrinsic information on the coded bits extracted from the output of the APP filter can be utilized again by the demodulator as updated a priori information. Performance over AWGN and time-varying Rayleigh flat fading channels is investigated by Monte Carlo simulation. By comparison between the symbol-by-symbol maximum a posteriori probability (SBS-MAP) detection with soft decision Viterbi decoding and the proposed method, it can be shown that considerable improvement can be achieved for both static and time-varying fading channels. Rui Zhang 0010, Xiao-Ping Zhang 0002 |
ICASSP (3) | 2 |
| 2007 | Generalized Trellis Coded Quantization for Data HidingabstractInformation theory guides us to investigate the choice of quantizers for data hiding applications. In this paper, the design of the quantizer selection rule in trellis coded quantization (TCQ) based data hiding is discussed. A novel trellis branch quantizer selection rule which changes the old state transition function and takes advantage of all trellis states is proposed so as to increase robustness against attacks. Theoretical analysis and simulation results show that the new TCQ path selection (or branch selection) method achieves better bit error rate (BER) performance in the case of Gaussian attack compared to other popular approaches. The path selection rule could also be used as a secret key to provide security for the practical data hiding. Xiaofeng Wang 0008, Xiao-Ping Zhang 0002 |
ICASSP (2) | 2 |
| 2007 | Wavelet-Based Texture Retrieval using Independent Component AnalysisabstractIn this paper, a novel approach to texture retrieval using independent component analysis (ICA) in wavelet domain is proposed. It is well recognized that the wavelet coefficients in different subbands are statistically correlated, resulting in the fact that the product of the marginal distributions of wavelet coefficients is not accurate enough to characterize the stochastic properties of texture images. To tackle this problem, we employ (ICA) in feature extraction to decorrelate the analysis coefficients in different subbands, followed by modeling the marginal distributions of the separated sources using generalized Gaussian density (GGD), and perform similarity measure based on the maximum likelihood criterion. It is demonstrated by simulation results on a database consisting of 1776 texture images that the proposed method improve the accuracy of texture image retrieval in terms of average retrieval rate, compared with the traditional method using GGD for feature extraction and Kullback-Leibler divergence for similarity measure. Rui Zhang 0010, Xiao-Ping Zhang 0002, Ling Guan |
ICIP (6) | 2 |
| 2006 | Nonlinear Fusion of Multiple Sensors with Missing DataabstractWe introduce a new algorithm, multiple imputation particle filter, to solve the problem of data fusion with missing data in nonlinear state space models. The new algorithm is then applied to the problem of fusing observations by multiple asynchronous radars. Simulated data is used demonstrate the effectiveness and performance of the fusing algorithm. Alon Shalev Housfater, Xiao-Ping Zhang 0002 |
ICASSP (4) | 2 |
| 2006 | Color Image Watermarking Using the Spatio-Chromatic Fourier TransformabstractIn this paper, a color watermarking algorithm is proposed. The method encodes the chromatic content of a color image as CIE a*b*chromaticity coordinates whereas the achromatic content is encoded as CIE L tristimulus value. Color watermarks (yellow and blue) are embedded in the frequency domain of the chromatic channels by using the Spatio Chromatic Discrete Fourier Transform (SCDFT). It first encodes a*and b*as complex values, followed by a single discrete Fourier Transform. Watermark casting is performed by estimating the Just-Noticeable distortion (JND) of the images, to ensure watermark invisibility. The most interesting characteristics of the new scheme is the possibility of performing watermarking in the frequency domain of chromatic components. Tsz Kin Tsui, Xiao-Ping Zhang 0002, Dimitrios Androutsos |
ICASSP (2) | 2 |
| 2006 | Hidden Markov Model Framework Using Independent Component Analysis Mixture ModelabstractThis paper describes a novel method for the analysis of sequential data that exhibits strong non-Gaussianities. In particular, we extend the classical continuous hidden Markov model (HMM) by modeling the observation densities as a mixture of non-Gaussian distributions. In order to obtain a parametric representation of the densities, we apply the independent component analysis (ICA) mixture model to the observations such that each non-Gaussian mixture component is associated with a standard ICA. Under this new framework, we develop the re-estimation formulas for the three fundamental HMM problems, namely, likelihood computation, state sequence estimation, and model parameter learning. The simulations also validate the theoretical results Xiao-Ping Zhang 0002 |
ICASSP (5) | 2 |
| 2006 | Video Event Detection using ICA Mixture Hidden Markov ModelsabstractIn this paper, a framework that combines feature extraction, model learning, and likelihood computation, is presented for video event detection. First, the independent component analysis (ICA) is applied to the raw feature space to extract the spatial features. Then, a framework based on ICA mixture hidden Markov models (ICAMHMM) is used to exploit the spatial and temporal characteristics of the training data. After the model is learnt, the likelihood for a given video sequence is computed and then used to classify the video into a semantic event. Golf video sequences are used for simulations. The results show that the proposed method can effectively detect semantic video events. Xiao-Ping Zhang 0002 |
ICIP | 2 |
| 2006 | Minimum Distortion Look-Up Table Based Data HidingabstractIn this paper, we present a novel data hiding scheme based on the minimum distortion look-up table (LUT) embedding that achieves good distortion-robustness performance. We first analyze the distortion introduced by LUT embedding and formulate its relationship with run constraints of LUT. Subsequently, a Viterbi algorithm is presented to find the minimum distortion LUT. Theoretical analysis and numerical results show that the new LUT design achieves not only less distortion but also more robustness than the traditional LUT based data embedding schemes Xiaofeng Wang 0008, Xiao-Ping Zhang 0002 |
ICME | 2 |
| 2006 | A Secret Key Based Multiscale Fragile Watermark in the Wavelet DomainabstractThe distribution of the wavelet coefficients in 2-D discrete wavelet transform (DWT) subspaces can be well described by a Gaussian mixture statistical model. In this paper, a secret key based fragile watermarking scheme is presented based on this statistical model. The Gaussian statistical model parameters are obtained by an expectation maximization (EM) algorithm and modified in a way to form special relationships for image authentication. The secret key is designed to securely embed a message bit stream, such as personal signatures or copyright logos, into a host image. Because of the secret embedding key, the new method is robust to most image tampering, even when the attackers are fully aware of the watermark embedding algorithms. Besides, the secret embedding key can be encrypted and embedded as a robust watermark into the same host image of the fragile watermarks for the benefit that the decoding of fragile watermarks only requires a single encryption key other than the image itself. The new method also has the advantage of changing only a few image data for watermark embedding and being able to distinguish some normal image operations such as compression from malicious to achieve a semi-fragile application Xiao-Ping Zhang 0002 |
ICME | 2 |
| 2006 | Quaternion image watermarking using the spatio-chromatic fourier coefficients analysisabstractIn this paper, a new color watermarking algorithm that uses the quaternion fourier transform (QFT) to mark the La∗b∗ components of color images is presented. First, we propose an interpretation of the QFT coefficients using the Spatio-Chromatic Fourier analysis, so the effects of any changes to the coefficients can be predicted. Next, watermark casting is performed by modifying the positive and negative coeffi-cients together. The idea is twofold: Robustness is achieved by embedding a color watermark in the coefficient with pos-itive frequency, which spreads it to all components in the spatial domain. On the other hand, invisibility is satisfied by modifying the coefficient with negative frequency, such that the combined effects of the two are insensitive to human eyes. Tsz Kin Tsui, Xiao-Ping Zhang 0002, Dimitrios Androutsos |
ACM Multimedia | 2 |
| 2006 | Multiscale Fragile Watermarking Based on the Gaussian Mixture ModelabstractIn this paper, a new multiscale fragile watermarking scheme based on the Gaussian mixture model (GMM) is presented. First, a GMM is developed to describe the statistical characteristics of images in the wavelet domain and an expectation-maximization algorithm is employed to identify GMM model parameters. With wavelet multiscale subspaces being divided into watermarking blocks, the GMM model parameters of different watermarking blocks are adjusted to form certain relationships, which are employed for the presented new fragile watermarking scheme for authentication. An optimal watermark embedding method is developed to achieve minimum watermarking distortion. A secret embedding key is designed to securely embed the fragile watermarks so that the new method is robust to counterfeiting, even when the malicious attackers are fully aware of the watermark embedding algorithm. It is shown that the presented new method can securely embed a message bit stream, such as personal signatures or copyright logos, into a host image as fragile watermarks. Compared with conventional fragile watermark techniques, this new statistical model based method modifies only a small amount of image data such that the distortion on the host image is imperceptible. Meanwhile, with the embedded message bits spreading over the entire image area through the statistical model, the new method can detect and localize image tampering. Besides, the new multiscale implementation of fragile watermarks based on the presented method can help distinguish some normal image operations such as JPEG compression from malicious image attacks and, thus, can be used for semi-fragile watermarking. Xiao-Ping Zhang 0002 |
IEEE Trans. Image Process. | 2 |
| 2005 | Video Shot Boundary Detection Using Independent Component AnalysisabstractVideo shot boundary detection is an important early stage of content-based video analysis. In this paper, a new method for shot boundary detection using independent component analysis (ICA) is presented. By projecting video frames from illumination-invariant raw feature space into low dimensional ICA subspace, each video frame is represented by a two-dimensional compact feature vector. An iterative clustering algorithm based on adaptive thresholding is developed to detect cuts and gradual transitions simultaneously in ICA subspace. Experimental results successfully validate the new method and show that it can effectively detect both abrupt transitions and gradual transitions. Xiao-Ping Zhang 0002 |
ICASSP (2) | 2 |
| 2005 | A New LUT Watermarking Scheme with Near Minimum Distortion Based on the Statistical Modeling in the Wavelet Domain
Kan Li 0004, Xiao-Ping Zhang 0002 |
ICIC (2) | 2 |
| 2005 | Automatic identification of digital video based on shot-level sequence matchingabstractTo locate a video clip in large collections is very important for retrieval applications, especially for digital rights management. In this paper, we present a novel technique for automatic identification of digital video. This new algorithm is based on dynamic programming that fully uses the temporal dimension to measure the similarity between two video sequences. A normalized chromaticity histogram is used as a feature which is illumination-invariant. Dynamic programming is applied on shot-level to find the optimal nonlinear mapping between video sequences. Two new normalized distance measures are presented for video sequence matching. One measure is based on the normalization of the optimal path found by dynamic programming. The other measure combines both the visual features and the temporal information. Experimental results show that the shot-level approach is robust to frame rate conversion, color correction, and compressions. The proposed distance measures are suitable for variable-length comparisons. Xiao-Ping Zhang 0002 |
ACM Multimedia | 2 |
| 2005 | Bidirectional labeling and registration scheme for grayscale image segmentationabstractIn this paper, we introduce a new image segmentation scheme that is based on bidirectional labeling and registration and prove that its segmentation performance is equivalent to that of the conventional watershed segmentation algorithm. The proposed bidirectional labeling and registration scheme, which we refer to as bidirectional labeling and registration scheme (BIDS), involves only linear scans of image pixels. It uses one-dimensional operations rather than the queues that are used in traditional segmentation algorithms, which are two-dimensional problems. BIDS also provides unique labels for individual homogeneous regions. In addition to achieving the same segmentation results, BIDS is four times less computationally complex than the conventional watershed by immersion technique. Lei Ma 0002, Xiao-Ping Zhang 0002, Jennie Si, Glen P. Abousleman |
IEEE Trans. Image Process. | 2 |
| 2005 | Comments on "An SVD-based watermarking scheme for protecting rightful Ownership"abstractIn a recent paper by Tan and Liu , a watermarking algorithm for digital images based on singular value decomposition (SVD) is proposed. This comment demonstrates that this watermarking algorithm is fundamentally flawed in that the extracted watermark is not the embedded watermark but determined by the reference watermark. The reference watermark generates the pair of SVD matrices employed in the watermark detector. In the watermark detection stage, the fact that the employed SVD matrices depend on the reference watermark biases the false positive detection rate such that it has a probability of one. Hence, any reference watermark that is being searched for in an arbitrary image can be found. Both theoretical analysis and experimental results are given to support our conclusion. Xiao-Ping Zhang 0002, Kan Li 0004 |
IEEE Trans. Multim. | 1 |
| 2004 | Symbol error rate evaluation for OFDM systems with MPSK modulationabstractOrthogonal frequency division multiplexing (OFDM) is the most commonly used multiplexing method in multicarrier (MC) systems. In this paper, a new closed-form formula for the symbol error rate (SER) is derived for generic OFDM systems with M-ary phase shift keying (MPSK) modulation and the optimal phase detector for each subchannel over a fading channel. This new formula can be used to evaluate both discrete Fourier transform (DFT)-based OFDM and discrete wavelet multitone (DWMT) systems. It can also be adopted as a criterion for OFDM system design. Monte-Carlo numerical simulations are performed to verify the theoretical SER analysis, and it is shown that the theoretical results from the presented new formula is consistent with simulation results. Xiao-Ping Zhang 0002 |
GLOBECOM | 2 |
| 2004 | A multiscale fragile watermark based on the Gaussian mixture model in the wavelet domainabstractThe wavelet coefficients in 2D discrete wavelet transform (DWT) subspaces have a peaky, heavy-tailed marginal distribution that can be well described by a Gaussian mixture statistical model. In this paper, a multiscale implementation of fragile watermarks based on the Gaussian mixture model is presented. The presented new method can embed a message bit stream, such as personal signatures or copyright logos, into a host image. With the embedded message bits spreading over the whole image area, the new method can detect and localize any image tampering since it will inevitably destroy a certain message bits. Compared with some other fragile watermark techniques, the statistical model based method modifies only a very small amount of image data to embed watermarks and the modification is hardly perceived by human vision because it occurs at texture edges. Besides, the multiscale implementation of fragile watermarks based on the presented method can help distinguish some normal image operations such as compression from malicious attacks, which is meaningful in terms of semi-fragile watermarking applications. Xiao-Ping Zhang 0002 |
ICASSP (3) | 2 |
| 2004 | Texture image retrieval based on a Gaussian mixture model and similarity measure using a Kullback divergenceabstractIn a content-based image retrieval (CBIR) system, indexing feature vectors and the similarity measure between feature vectors are two key factors for retrieval performance. We present a new CBIR system with statistical-model based image feature extraction in the wavelet domain and a Kullback divergence based similarity measure. A two component Gaussian mixture model (GMM) in the wavelet domain is employed and the model parameters are used to form features for image indexing. A new Kullback divergence based similarity measure is then presented for image retrieval. The experimental results demonstrate that the similarity measure based on the Kullback divergence is more effective than conventional similarity measures, such as the city-block distance and the Euclidean distance. It is shown that the new CBIR system, with the combination of the GMM and the new Kullback divergence based similarity measure, outperforms most other methods in retrieval performance for texture images, while keeping a comparable level of computational complexity. Xiao-Ping Zhang 0002 |
ICME | 2 |
| 2003 | Design of M-band complex-valued filter banks for multicarrier transmission over multipath wireless channelsabstractThe orthogonal M-band complex-valued filter banks have more flexibility in terms of frequency response. In this paper, a new general design method for orthogonal M-band complex-valued filter banks is presented for multicarrier (MC) transmission to reduce intersymbol and interchannel interference (ISI and ICI). The design includes three procedures. First, a parameterization method is used to design the baseband scaling filter. Then the M-1 band wavelet filters are deduced based on the scaling filter by a Gram-Schmidt algorithm. Thirdly, a numerical optimization method is developed to determine the free parameters and coefficients of the filter banks based on an objective to minimize the intersymbol and interchannel interference power for a multipath wireless channel. Simulation results demonstrate that the MC system based on the designed M-band complex-valued filter banks has superior performance to that of discrete Fourier transform (DFT) based orthogonal frequency division multiplexing (OFDM) and discrete wavelet multitone (DWMT) in terms of the interference power over a multipath channel. Xiao-Ping Zhang 0002 |
ICASSP (4) | 2 |
| 2003 | Bi-directional gradient labeling and registration for gray-scale image segmentationabstractWatershed is one of the commonly used methods for image segmentation. In this paper, we introduce a new segmentation scheme based on bi-directional labeling and registration and prove that its segmentation performance is equivalent to that of conventional watershed but it is much more computational efficient The bi-directional labeling and registration scheme, which will be referred to as BIDS, involves only linear scans of image pixels. It uses one dimensional operations instead of queues while traditional segmentation algorithms are two dimensional problems. BIDS also provides unique labels for each homogeneous regions. In addition to achieving the same segmentation results as conventional watershed, BIDS is four times less computationally complex than the conventional watersheds by immersion. Lei Ma 0002, Xiao-Ping Zhang 0002, Jennie Si, Glen P. Abousleman |
ICIP (1) | 2 |
| 2003 | Fragile watermark based on the Gaussian mixture model in the wavelet domain for image authenticationabstractIn this paper, a new fragile watermarking method based on statistical analysis in the wavelet domain is developed for image authentication. A two component Gaussian mixture model is developed to describe the statistical characteristics of images in the wavelet domain. Each wavelet subspace of the original image is divided into a watermarking block and a reference block. A Gaussian mixture model is then applied to both blocks to obtain their respective model parameters by an EM (expectation-maximization) algorithm. By slightly changing the wavelet coefficients (adding watermark) in the watermark block, we can adjust its model parameter to the same value as that of the reference block for authentication purposes, which constitutes the fragile watermark. Any change in the fragile-watermarked image will break the relationship of the statistical models between the watermarking block and reference block. The authentication procedure needs only a simple comparison between the model parameters of the two blocks in the watermarked image and no information about the original image. The preliminary experimental results indicate that the new watermarking scheme conforms with human perception characteristics and provides a perceptually invisible fragile watermark with fewer image data modified, compared with some other conventional fragile watermarking methods. Xiao-Ping Zhang 0002 |
ICIP (1) | 2 |
| 2003 | Content-based image retrieval using a Gaussian mixture model in the wavelet domain
Xiao-Ping Zhang 0002, Ling Guan |
VCIP | 2 |
| 2001 | Segmentation of bright targets using wavelets and adaptive thresholdingabstractA general systematic method for the detection and segmentation of bright targets is developed. We use the term "bright target" to mean a connected, cohesive object which has an average intensity distribution above that of the rest of the image. We develop an analytic model for the segmentation of targets, which uses a novel multiresolution analysis in concert with a Bayes classifier to identify the possible target areas. A method is developed which adaptively chooses thresholds to segment targets from background, by using a multiscale analysis of the image probability density function (PDF). A performance analysis based on a Gaussian distribution model is used to show that the obtained adaptive threshold is often close to the Bayes threshold. The method has proven robust even when the image distribution is unknown. Examples are presented to demonstrate the efficiency of the technique on a variety of targets. Xiao-Ping Zhang 0002, Mita D. Desai |
IEEE Trans. Image Process. | 1 |
| 1999 | Minimum component eigen-vector based classification technique with application to TM imagesabstractIn this paper, we propose a new classification technique based on the minimum component analysis (MCA) instead of the traditional principal components analysis (PCA). Most existing classification techniques based on PCA like to represent a class by its principal component. However, the principal component is not always the best choice since it has a high possibility for a class to overlap with other classes in the principal component direction. The new minimum component eigen-vector based classification technique overcomes this disadvantage by representing a class with its minimum component. In addition, a minimum likelihood decision rule is employed instead of maximum likelihood decision rule. Good performance of our technique is verified by experimental results on Kennedy Space Center (KSC) TM images. Guohui He, Mita D. Desai, Xiao-Ping Zhang 0002 |
ICASSP | 3 |
| 1998 | Nonlinear adaptive noise suppression based on wavelet transformabstractThe conventional linear adaptive filters are not effective for discriminating the transient wideband signal components from noise. A recently developed wavelet shrinkage approach is able to maintain the function local regularity while suppressing noise however, it has only been used in function estimation problems. In this paper, a new type of nonlinear filtering method for adaptive noise suppression is presented, based on shrinkage method. A new class of shrinkage functions is also presented. The filtering structure and the learning algorithm are developed. The theoretical analysis proves convergence in certain statistical sense. The numerical results of our system are presented for both the standard and the new shrinkage function and compared with the conventional linear adaptive filter based techniques. Results indicate that both the optimal solution and the learning performance are superior to the conventional methods. It is shown that our new shrinkage function performs better than the standard shrinkage function. Xiao-Ping Zhang 0002, Mita D. Desai |
ICASSP | 1 |
| 1998 | Adaptive denoising based on SURE riskabstractA new adaptive denoising method is presented based on Stein's (1981) unbiased risk estimate (SURE) and on a new class of thresholding functions. First, we present a new class of thresholding functions that has a continuous derivative while the derivative of standard soft-thresholding function is not continuous. The new thresholding functions make it possible to construct the adaptive algorithm whenever using the wavelet shrinkage method. By using the new thresholding functions, a new adaptive denoising method is presented based on SURE. Several numerical examples are given. The results indicated that for denoising applications, the proposed method is very effective in adaptively finding the optimal solution in a mean square error (MSE) sense. It is also shown that this method gives better MSE performance than those conventional wavelet shrinkage methods. Xiao-Ping Zhang 0002, Mita D. Desai |
IEEE Signal Process. Lett. | 1 |
| 1997 | Feedforward Neural Networks with Multilevel Hidden Neurons for Remotely Sensed Image ClassificationabstractArtificial neural network has been, used as a powerful tool for pattern classification. However, it is difficult to train when the data exhibit non-sparse or overlapping pattern classes which is often the case in practical applications. In this paper, we introduce the feedforward neural network with the hidden layer consisting of multilevel neurons. The convergence property of one-layer neural network with multilevel neurons is proved. The new feedforward model is inherently capable of fuzzy pattern classification of non-sparse or overlapping pattern classes. As an application, we apply the network for the classification of LANDSAT TM data. The results show that this approach produces better results compared with conventional neural networks. Zhong-yu Chen, Mita D. Desai, Xiao-Ping Zhang 0002 |
ICIP (2) | 3 |
| 1997 | Wavelet Based Automatic Thresholding for Image SegmentationabstractIn this paper, a new systematic method to segment possible target areas based on wavelet transforms is presented. We develop an analytic model for the segmentation of targets, which uses a novel multiresolution analysis in concert with a Bayesian classifier to identify the possible target areas. A method is developed which adaptively chooses thresholds to segment targets from background, by using a multiscale analysis of the image probability density function (PDF). We present examples which demonstrate the efficiency of the technique on a variety of targets. Xiao-Ping Zhang 0002, Mita D. Desai |
ICIP (1) | 1 |