Z. Jane Wang 0001

dblp:13/3672-1 · also Zhen Jane Wang 0001, Zhen Wang 0029 · DBLP profile ↗
← Back
203ranked-venue papers
10as first author
75since 2021 · last 2026
0000-0002-3791-0249ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 95 · 7 first-author · 21 since 2021Artificial intelligence and machine learning · 40 · 24 since 2021Applied, interdisciplinary, general and emerging computing · 36 · 21 since 2021Computer networks · 25 · 2 first-author · 13 since 2021Security and privacy · 8 · 1 since 2021Systems, architecture and hardware · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-authorTheory of computation · 1
YearPublicationVenuePosition
2026 Signal-aware synthesis of tissue polarization uniformity from OCT images guided by an SNR-based heuristic
abstract
Polarization-sensitive optical coherence tomography (PS-OCT) is a powerful imaging modality that captures both structural and polarization-related tissue features, offering significant diagnostic value. Among these, the degree of polarization uniformity (DOPU) is critical for characterizing tissue microstructure. However, obtaining DOPU images typically requires specialized hardware and complex system configurations. To address this limitation, we propose a knowledge-guided deep generative framework, signal attention GAN (SA-GAN), to synthesize DOPU images directly from standard OCT intensity scans. SA-GAN integrates a signal-guided attention mechanism inspired by signal-to-noise ratio (SNR) principles, enabling selective focus on regions with meaningful polarization patterns while suppressing noise-dominated areas. This design allows for the generation of accurate, high-fidelity DOPU images without additional imaging hardware. We validated SA-GAN on three independent datasets: SKIN-PSOCT, CARTILAGE-PSOCT, and the public Retinal-OCT2017 dataset. On SKIN-PSOCT, SA-GAN achieved a structural similarity index measure (SSIM) of 97.8% and a peak signal-to-noise ratio (PSNR) of 28.6 dB. On CARTILAGE-PSOCT, it reached an SSIM of 93.9% and a PSNR of 24.8 dB in cross-dataset testing. Applied to the Retinal-OCT2017 dataset, SA-GAN achieved state-of-the-art performance in a four-class retinal disease classification task. These results demonstrate the robustness and generalizability of our method. SA-GAN provides a cost-effective and practical solution to extend PS-OCT capabilities, supporting the development of intelligent imaging systems for biomedical diagnostics and digital medicine. Our code is available via this link: https://github.com/Yuhengw/SA-GAN .
Chris Zhou, Jiayue Cai, John D. W. Madden, Orlando J. Rojas, Sunil Kalia, Z. Jane Wang 0001, Daniel C. Louie, Tim K. Lee
Expert Syst. Appl.9
2026 An Extended Alamouti Code for Four-Antenna Backscatter Tags
abstract
Backscatter communications, an emerging low-cost and low-consumption green communication paradigm, has received significant attention in both industry and academia. In recent years, multiple-input multiple-output (MIMO) technology has also been integrated into backscatter communications to meet the requirements of transmission performance in IoT scenarios. However, due to the hardware limitations of backscatter tags, one of the main challenges for multi-antenna backscatter tags is to reduce the circuit complexity of space-time block codes (STBC) while maintaining transmission performance. Although low-complexity STBC for dual-antenna tags has been explored, to the best of our knowledge, research on high-performance, low-complexity STBC for four-antenna tags remains in its infancy. In this paper, we propose a 2×4 extended Alamouti code (EAC) with full-rate capability for MIMO backscatter communications. The results show that the required impedance of the proposed EAC on the backscatter tag is less than half of that of the conventional orthogonal STBC (OSTBC), which considerably reduces the complexity of the tag circuit. Additionally, we also derive the asymptotic closed form expression of the symbol error rate (SER) of the proposed EAC to obtain the system insights. The derivation results show that the achievable diversity order of the proposed EAC in 1×4×N backscatter channel is 2×min(2,N), which indicates that the proposed EAC can achieve the same diversity performance as the conventional OSTBC by setting an appropriate number of receiver antennas. Finally, numerical simulations are performed to validate the accuracy of our analysis and the superiority of the proposed EAC in SER.
Huixu Luan, Murong Lv, Jihong Wang 0001, Chen He 0002, Qianqian Zhang 0001, Z. Jane Wang 0001
IEEE Internet Things J.6
2026 GloW-VSNet: A scribble-based weakly supervised framework for global-view vitiligo lesion segmentation
abstract
Vitiligo lesion identification is essential for quantifying disease severity, monitoring disease progression and assessing treatment response, particularly for objective quantification. However, segmenting vitiligo lesions from clinical images is challenging due to indistinct borders, complex backgrounds, and image artifacts. The difficulty increases when handling small and sparse lesions in global-view photographs. Fully supervised segmentation models require extensively annotated datasets, making the labelling process time-consuming and costly. To address these challenges, we propose GloW-VSNet, a scribble-guided weakly supervised segmentation method for global-view vitiligo detection. Our approach integrates differentiable feature clustering with a spatial attention mechanism based on physician-provided scribble annotations, enabling the model to focus on relevant spatial features and improve segmentation accuracy despite background noise and artifacts. Additionally, we introduce spatial continuity optimization to preserve the natural distribution of vitiligo, enhancing segmentation consistency while reducing computational demands. Extensive experiments on two public vitiligo datasets and two private datasets demonstrate that GloW-VSNet achieves state-of-the-art performance. To our knowledge, this is the first study to explore weakly supervised global-view vitiligo segmentation, addressing a critical research gap. Our method enhances the assessment of disease severity and monitoring of treatment response through an objective assessment for real-world applications. Our code is publicly available at https://github.com/YuhanZheng0327/Weakly-Supervised-Vitiligo-Lesion-Segmentation.
Chloe Yue, Thomas Zhang, Jiayue Cai, Chunqi Chang, Harvey Lui, Sunil Kalia, Z. Jane Wang 0001, Tim K. Lee
Medical Image Anal.9
2026 DIDLM: A SLAM Dataset for Difficult Scenarios Featuring Infrared, Depth Cameras, LiDAR, 4D Radar, and Others Under Adverse Weather, Low Light Conditions, and Rough Roads
abstract
Adverse weather conditions, low-light environments, and bumpy road surfaces pose significant challenges to SLAM in robotic navigation and autonomous driving. Existing datasets in this field predominantly rely on single sensors or combinations of LiDAR, cameras, and IMUs. However, 4D millimeter-wave radar demonstrates robustness in adverse weather, infrared cameras excel in capturing details under low-light conditions, and depth images provide richer spatial information. Multi-sensor fusion methods also show potential for better adaptation to bumpy roads. Despite some SLAM studies incorporating these sensors and conditions, there remains a lack of comprehensive datasets addressing low-light environments and bumpy road conditions, or featuring a sufficiently diverse range of sensor data. In this study, we introduce a multi-sensor dataset covering challenging scenarios such as snowy weather, rainy weather, nighttime conditions, speed bumps, and rough terrains. The dataset includes rarely utilized sensors for extreme conditions, such as 4D millimeter-wave radar, infrared cameras, and depth cameras, alongside 3D LiDAR, RGB cameras, GPS, and IMU. It supports both autonomous driving and ground robot applications and provides reliable GPS/INS ground truth data, covering structured and semi-structured terrains. We evaluated various SLAM algorithms using this dataset, including RGB images, infrared images, depth images, LiDAR, and 4D millimeter-wave radar. The dataset spans a total of 18.5 km, 69 minutes, and approximately 660 GB, offering a valuable resource for advancing SLAM research under complex and extreme conditions. Our dataset is available athttps://gongweisheng.github.io/DIDLM.github.io/
Weisheng Gong, Chen He 0002, Kaijie Su, Qingyong Li, Z. Jane Wang 0001
IEEE Trans. Intell. Transp. Syst.6
2026 Modality-Agnostic Federated Learning With Adaptive Updates for Heterogeneous Medical Image Tasks
abstract
Federated learning (FL) enables collaborative model training across decentralized medical datasets while preserving data privacy. Its practical adoption remains limited due to data heterogeneity, specifically, differences in input imaging modality (e.g., CT or MRI) and client task (e.g., segmentation or classification) across participating institutions (clients). Such data heterogeneity poses significant challenges for jointly learning a unified global model that generalizes across clients with different input modality and task. To address this, we propose FedCMT, a modality-agnostic FL framework that adaptively aggregates heterogeneous client models. FedCMT supports flexible input modalities and diverse local tasks by incorporating group-wise adapters and personalized decoders that capture modality- and task-specific features. To enhance collaboration across clients, FedCMT employs a conflict-averse module that extracts modality-invariant representations and mitigates inter-client feature conflicts. FedCMT also integrates a global-to-local knowledge distillation mechanism to balance global consistency and local specialization. The proposed FedCMT maintains stability while fostering shared knowledge in diverse medical imaging modalities. We evaluate FedCMT on ten CT and MR datasets involving up to eight federated clients performing segmentation or classification tasks. Experimental results show that FedCMT consistently outperforms state-of-the-art FL baselines, yielding an average improvement of 4.76% over state-of-the-art methods and 4.01% over standalone training. These results demonstrate FedCMT as a promising adaptable FL for real-world medical image analysis.
Shaohao Rui, Z. Jane Wang 0001
IEEE Trans. Medical Imaging5
2026 Self-Supervised T2WI-Bridged Framework for Liver Segmentation and PDFF Prediction From US Images
abstract
Proton Density Fat Fraction (PDFF) is the gold standard for non-invasive fatty liver diagnosis, but its reliance on Magnetic Resonance Imaging (MRI) limits broad clinical applicability. Motivated by the accessibility of B-mode Ultrasound (US) in fatty liver assessment, we propose a novel framework for liver segmentation and PDFF prediction from US images. To enhance generalization ability despite limited paired US-PDFF data, our framework integrates a cross-task self-supervised pretext task that extracts semantic features to guide echo intensity capture, benefiting both liver segmentation and PDFF prediction. To address the noise and artifacts inherent in US images, our framework leverages T2-weighted imaging (T2WI) exclusively during training to establish a feature bridge between US and PDFF, thereby enhancing PDFF prediction. Once trained, the model relies solely on US for inference, making it a practical and cost-effective alternative to MRI-based PDFF estimation. Additionally, our framework introduces an uncertainty-augmented adversarial loss function to refine liver boundary delineation, further improving segmentation and PDFF prediction accuracy. Experimental results demonstrate that our method outperforms state-of-the-art methods in liver segmentation and PDFF prediction; and in a specific application study, our predicted PDFF achieves accuracy comparable to real PDFF for hepatic steatosis classification, highlighting its clinical potential. The full source code and detailed documentation are publicly available at https://github.com/D0ngZhang/SSTB.
Dong Zhang 0009, Qi Zeng 0004, Tim Salcudean, Z. Jane Wang 0001
IEEE Trans. Medical Imaging4
2026 CCDM: Continuous Conditional Diffusion Models for Image Generation
abstract
Continuous Conditional Generative Modeling(CCGM) estimates high-dimensional data distributions, such as images, conditioned on scalar continuous variables (aka regression labels). WhileContinuous Conditional Generative Adversarial Networks(CcGANs) were designed for this task, their instability during adversarial learning often leads to suboptimal results.Conditional Diffusion Models(CDMs) offer a promising alternative, generating more realistic images, but their diffusion processes, label conditioning, and model fitting procedures are either not optimized for or incompatible with CCGM, making it difficult to integrate CcGANs' vicinal approach. To address these issues, we introduceContinuous Conditional Diffusion Models(CCDMs), the first CDM specifically tailored for CCGM. CCDMs address existing limitations with specially designed conditional diffusion processes, a novel hard vicinal image denoising loss, a customized label embedding method, and efficient conditional sampling procedures. Through comprehensive experiments on four datasets with resolutions ranging from$64\times 64$to$192\times 192$, we demonstrate that CCDMs outperform state-of-the-art CCGM models, establishing a new benchmark. Ablation studies further validate the model design and implementation, highlighting that some widely used CDM implementations are ineffective for the CCGM task. Our code is publicly available athttps://github.com/UBCDingXin/CCDM.
Xin Ding 0004, Kao Zhang, Z. Jane Wang 0001
IEEE Trans. Multim.4
2026 Integrating Clinical Knowledge Graphs and Gradient-Based Neural Systems for Enhanced Melanoma Diagnosis via the Seven-Point Checklist
abstract
The seven-point checklist (7PCL) is a widely used diagnostic tool in dermoscopy for identifying malignant melanoma by assigning point values to seven specific attributes. However, the traditional 7PCL is limited to distinguishing between malignant melanoma and melanocytic nevi (MN) and falls short in scenarios where multiple skin diseases with appearances similar to melanoma coexist. To address this limitation, we propose a novel diagnostic framework that integrates a clinical knowledge-based topological graph (CKTG) with a gradient diagnostic strategy featuring a data-driven weighting (GD-DDW) system. The CKTG captures both the internal and external relationships among the 7PCL attributes, while the GD-DDW emulates dermatologists' diagnostic processes, prioritizing visual observation before making predictions. Additionally, we introduce a multimodal feature extraction approach leveraging a dual-attention mechanism to enhance feature extraction through cross-modal interaction and unimodal collaboration. This method incorporates meta-information to uncover interactions between clinical data and image features, ensuring more accurate and robust predictions. Our approach, evaluated on the EDRA dataset, achieved an average AUC of 88.6%, demonstrating superior performance in melanoma detection and feature prediction. This integrated system provides data-driven benchmarks for clinicians, significantly enhancing the precision of melanoma diagnosis.
Tianze Yu, Jiayue Cai, Sunil Kalia, Harvey Lui, Z. Jane Wang 0001, Tim K. Lee
IEEE Trans. Neural Networks Learn. Syst.6
2026 Multi-Frequency Radio Map Assisted Unmanned Aerial Relay for Bridging Ground D2D Networks
abstract
In the rapidly advancing realm of wireless communication, device-to-device (D2D) technology, an emerging approach for data exchange and connectivity, has been attracting increasing attention. Unmanned Aerial Vehicles (UAVs) can act as air relays or base stations, and integrate isolated D2D clusters into a cohesive network fabric in outdoor environments. However, in complex terrain, the communication signals are subject to irregular attenuation, and the signal propagation attenuation of different frequency bands in the same terrain is inconsistent. It is challenging to utilize UAVs to coverage D2D terrestrial users in complex terrain. In this paper, we propose the UAVs relaying for bridging the terrestrial D2D networks assisted by multi-frequency radio maps. From the real-world topographical data, we generate multi-frequency radio maps, which represent the distortion of different frequency band signals by rich information about land layouts. Next, we focus on the air-to-ground D2D network topology and formulate it into an optimization problem. Then, we decompose it into two subproblems. The first subproblem pertains to the design of the ground network structure. We employ the D2D frequency band radio map to assess the communication quality between user pairs, and propose a measure of D2D closeness centrality to select ‘cellular users’ that can communicate directly to a UAV. The second subproblem involves the UAVs’ deployment and the frequency selection. We present a multi-frequency radio map improved k-means method, which has lower algorithm complexity than the traversal method by reducing the utilization of the radio maps. Simulations validate the proposed scheme, demonstrating that: 1. Multi-frequency radio maps can provide efficient gains with real-world complex topography; 2. The proposed network structure and algorithm outperform other existing approaches.
Yangrui Dong, Chen He 0002, Huiyu Bai, Dusit Niyato, Z. Jane Wang 0001
IEEE Trans. Wirel. Commun.5
2025 Performance Analysis of RIS-Assisted Covert Rate-Splitting Multiple Access
abstract
This paper investigates a downlink covert communication system based on rate-splitting multiple access (RSMA), assisted by a reconfigurable intelligent surface (RIS) and a jammer. This scheme treats every user in the system as requiring covert transmission and splits each user's message stream into common and private parts to meet the covert communication demands in multi-user scenarios. We derive a closed-form approximate expression for the average minimum detection error probability (AMDEP). Through extensive simulations, the correctness of the analysis and the covertness of the communication system are validated.
Yanyu Cheng, Z. Jane Wang 0001, Dusit Niyato
GLOBECOM3
2025 PGD-Imp: Rethinking and Unleashing Potential of Classic PGD with Dual Strategies for Imperceptible Adversarial Attacks
abstract
Imperceptible adversarial attacks have recently attracted increasing research interests. Existing methods typically incorporate external modules or loss terms other than a simple lp-norm into the attack process to achieve imperceptibility, while we argue that such additional designs may not be necessary. In this paper, we rethink the essence of imperceptible attacks and propose two simple yet effective strategies to unleash the potential of PGD, the common and classical attack, for imperceptibility from an optimization perspective. Specifically, the Dynamic Step Size is introduced to find the optimal solution with minimal attack cost towards the decision boundary of the attacked model, and the Adaptive Early Stop strategy is adopted to reduce the redundant strength of adversarial perturbations to the minimum level. The proposed PGD-Imperceptible (PGD-Imp) attack achieves state-of-the-art results in imperceptible adversarial attacks for both untargeted and targeted scenarios. When performing untargeted attacks against ResNet-50, PGD-Imp attains 100% (+0.3%) ASR, 0.89 (-1.76) l2distance, and 52.93 (+9.2) PSNR with 57s (-371s) running time, significantly outperforming existing methods.
Zitong Yu, Ziqiang He, Z. Jane Wang 0001, Xiangui Kang
ICASSP4
2025 CA-UAP: Content-Agnostic Universal Adversarial Perturbation for Enhanced Generalization
abstract
Deep Neural Networks (DNNs) have been shown vulnerable to universal adversarial perturbation (UAP), which are imperceptible and capable of fooling the target model for most samples. Existing universal attack methods mainly focus on aggregating the gradient obtained from global image features to directly optimize (noise-based) or indirectly generate (generator-based) UAP. However, such methods do not yet consider improving the generalization of UAP from the perspective of making the perturbation irrelevant to image content. We note that minimizing self-similarity is helpful to make the UAP irrelevant to the image content. Therefore, we propose a novel Content-Agnostic UAP (CA-UAP), which combines global image features and local patch features to optimize UAP. Specifically, we introduce a self-similarity loss that encourages minimizing the similarity between adversarial perturbed global images and their randomly cropped local regions, making the UAP agnostic to image content and consequently enhancing UAP generalization. Extensive experiments on the ILSVRC 2012 dataset demonstrate that our proposed method outperforms existing methods in both untargeted and targeted attacks, e.g., improving the average fooling rate from 79.12% (achieved by the state-of-the-art method) to 82.25% in targeted attacks.
Ziqiang He, Jingyang Wen, Xiangui Kang, Z. Jane Wang 0001
ICASSP5
2025 INN-based Secure Steganography Using Lost Information as Adversarial Perturbations
abstract
Recently image steganography methods based on invertible neural networks (INNs) demonstrated the capability to automatically embed and extract secret messages while maintaining high visual quality in stego images. However, there remain concerns about security and invertibility of such methods. In this paper, for the first time, we introduce adversarial hiding into INN-based image steganography method to simultaneously perform steganographic embedding and adversarial perturbation generation, resulting in improved security. Our method enhances the invertibility of the INN structure: It utilizes the lost information of the INN to generate perturbations, which are then combined with the gradient of the cover image to produce an adversarial stego image. Also, a learnable noise layer is proposed to mitigate information loss caused by rounding and truncation during image storage. Therefore, the proposed method significantly improves security while enhancing extraction performance of INN-based steganography approach, as supported by our experimental results. For example, the steganalysis detection accuracy of SRNet decreases from 96.86% to 51.77% at a payload of 0.2 bits per pixel (bpp).
Fei Shang, Weixiang Zhao, Xiangui Kang, Z. Jane Wang 0001
ICASSP4
2025 Secure INN-based Steganography via Model Smoothing and Adversarial Attacks
abstract
In recent years, image steganography methods based on invertible neural networks (INNs) have received significant attention due to their invertible structure, which offers advantages in embedding and extracting secret messages. However, current INN-based image steganography methods face challenges, particularly their limited tolerance against noise interference (e.g., added Gaussian noise, adversarial perturbations, and JPEG compression) and vulnerability to detection by advanced deep steganalyzers. To address these concerns, we present a novel steganography framework that combines Median Smoothing Training (MST) with dynamic Projected Gradient Descent (d-PGD). Specifically, our method begins with employing an MST strategy during the training phase to improve the INN’s tolerance to noise, ensuring that accurate message extraction even under noise interference. Subsequently, to improve the security of INN-based steganography, we propose a d-PGD algorithm that can generate minimal adversarial perturbations capable of deceiving deep steganalyzers, thereby improving security without compromising extraction accuracy. Experimental results demonstrate that our method achieves state-of-the-art secret message extraction accuracy while significantly improving resistance against deep steganalyzers.
Weixiang Zhao, Fei Shang, Jingyang Wen, Xiangui Kang, Z. Jane Wang 0001
MMSP6
2025 Spik-NeRF: Spiking Neural Networks for Neural Radiance Fields
abstract
Spiking Neural Networks (SNNs), as a biologically inspired neural network architecture, have garnered significant attention due to their exceptional energy efficiency and increasing potential for various applications. In this work, we extend the use of SNNs to neural rendering tasks and introduce Spik-NeRF (Spiking Neural Radiance Fields). We observe that the binary spike activation map of traditional SNNs lacks sufficient information capacity, leading to information loss and a subsequent decline in the performance of spiking neural rendering models. To address this limitation, we propose the use of ternary spike neurons, which enhance the information-carrying capacity in the spiking neural rendering model. With ternary spike neurons, Spik-NeRF achieves performance that is on par with, or nearly identical to, traditional ANN-based rendering models. Additionally, we present a re-parameterization technique for inference that allows Spik-NeRF with ternary spike neurons to retain the event-driven, multiplication-free advantages typical of binary spike neurons. Furthermore, to further boost the performance of Spik-NeRF, we employ a distillation method, using an ANN-based NeRF to guide the training of our Spik-NeRF model, which is more compatible with the ternary neurons compared to the standard binary neurons. We evaluate Spik-NeRF on both realistic and synthetic scenes, and the experimental results demonstrate that Spik-NeRF achieves rendering performance comparable to ANN-based NeRF models.
Qinlong Lan, Yitian Wu, Z. Jane Wang 0001, Wanhua Li 0004, Yufei Guo 0001
NeurIPS6
2025 ESCAPE: Energy-based Selective Adaptive Correction for Out-of-distribution 3D Human Pose Estimation
Luke Bidulka, Mohsen Gholami, Jiannan Zheng, Martin J. McKeown, Z. Jane Wang 0001
Neurocomputing5
2025 Zero-Padding Space-Time Block Code for Dual-Antenna Backscatter Tag
abstract
Backscatter communications is receiving increasing attention in Internet of Things (IoT), while one major challenge for backscatter communications is that the circuit complexity for high performance tag is also high. In this article, for dual-antenna backscatter tag, we propose a zero-padding space-time block code (ZPSTBC) that can simultaneously improve the tag performance and reduce the complexity of the tag circuit. An interesting thing is that padding zero into space-time block code (STBC) is not beneficial in conventional multi-input-multi-output (MIMO) channels, but may lead to performance improvement in MIMO backscatter channels, from perspectives of error rate and energy harvesting, and also reduce the circuit complexity of the tag. This is due to characteristic of the MIMO structures and tag circuit of backscatter communications. We provide rigorous mathematical analysis and numerical simulations to illustrate the performance improvement and the tag complexity reduction caused by the proposed ZPSTBC.
Chen He 0002, Huixu Luan, Z. Jane Wang 0001
IEEE Internet Things J.3
2025 A heterogeneous federated learning framework for human activity recognition
Xinhui Yu, Arvin Tashakori, Martin J. McKeown, Z. Jane Wang 0001
Knowl. Based Syst.4
2025 VDMUFusion: A Versatile Diffusion Model-Based Unsupervised Framework for Image Fusion
abstract
Image fusion facilitates the integration of information from various source images of the same scene into a composite image, thereby benefiting perception, analysis, and understanding. Recently, diffusion models have demonstrated impressive generative capabilities in the field of computer vision, suggesting significant potential for application in image fusion. The forward process in the diffusion models requires the gradual addition of noise to the original data. However, typical unsupervised image fusion tasks (e.g., infrared-visible, medical, and multi-exposure image fusion) lack ground truth images (corresponding to the original data in diffusion models), thereby preventing the direct application of the diffusion models. To address this problem, we propose a versatile diffusion model-based unsupervised framework for image fusion, termed as VDMUFusion. In the proposed method, we integrate the fusion problem into the diffusion sampling process by formulating image fusion as a weighted average process and establishing appropriate assumptions about the noise in the diffusion model. To simplify the training process, we propose a multi-task learning framework that replaces the original noise prediction network, allowing for simultaneous prediction of noise and fusion weights. Meanwhile, our method employs joint training across various fusion tasks, which significantly improves noise prediction accuracy and yields higher quality fused images compared to training on a single task. Extensive experimental results demonstrate that the proposed method delivers very competitive performance across various image fusion tasks. The code is available at https://github.com/yuliu316316/VDMUFusion.
Yu Liu 0023, Juan Cheng 0004, Z. Jane Wang 0001, Xun Chen 0001
IEEE Trans. Image Process.4
2025 Effects of UAV Position Fluctuations on Air-to-Ground mmWave UAV Communications With Multiple Types of Blockages
abstract
Millimeter wave (mmWave)-based uncrewed aerial vehicle (UAV) communication is a promising candidate for future communications. However, hovering UAVs are susceptible to inevitable position fluctuations, while mmWave are highly sensitive to obstacles. Both factors contribute to variations in the system’s quality of service (QoS). Existing studies addressing either mmWave blockages or UAV fluctuations fail to capture their combined effects on QoS. This paper presents a tractable analytical model that establishes a theoretical relationship between UAV fluctuations and mmWave blockages (static, dynamic, and self-blockages). Closed-form expressions for reliable service probability and coverage probability are derived, providing insights into the impact of these combined factors on QoS. Monte Carlo simulations validate the theoretical analysis, showing that small fluctuations (e.g., less than 0.1 m in the studied scenario) have minimal impact on QoS, while larger fluctuations significantly degrade QoS, with various blockages further exacerbating this degradation. While the results may seem intuitive, the derived formulas reveal non-linearities and subtle dependencies, such as varying QoS sensitivity across different ranges of UAV fluctuations and mmWave blockages. Additionally, the analysis identifies feasible UAV placements to enhance the QoS.
Cunyan Ma, Xiaoya Li 0003, Yangrui Dong, Chen He 0002, Z. Jane Wang 0001
IEEE Trans. Intell. Transp. Syst.6
2024 PoseGen: Learning to Generate 3D Human Pose Dataset with NeRF
abstract
This paper proposes an end-to-end framework for generating 3D human pose datasets using Neural Radiance Fields (NeRF). Public datasets generally have limited diversity in terms of human poses and camera viewpoints, largely due to the resource-intensive nature of collecting 3D human pose data. As a result, pose estimators trained on public datasets significantly underperform when applied to unseen out-of-distribution samples. Previous works proposed augmenting public datasets by generating 2D-3D pose pairs or rendering a large amount of random data. Such approaches either overlook image rendering or result in suboptimal datasets for pre-trained models. Here we propose PoseGen, which learns to generate a dataset (human 3D poses and images) with a feedback loss from a given pre-trained pose estimator. In contrast to prior art, our generated data is optimized to improve the robustness of the pre-trained model. The objective of PoseGen is to learn a distribution of data that maximizes the prediction error of a given pre-trained model. As the learned data distribution contains OOD samples of the pre-trained model, sampling data from such a distribution for further fine-tuning a pre-trained model improves the generalizability of the model. This is the first work that proposes NeRFs for 3D human data generation. NeRFs are data-driven and do not require 3D scans of humans. Therefore, using NeRF for data generation is a new direction for convenient user-specific data generation. Our extensive experiments show that the proposed PoseGen improves two baseline models (SPIN and HybrIK) on four datasets with an average 6% relative improvement.
Mohsen Gholami, Rabab K. Ward, Z. Jane Wang 0001
AAAI3
2024 Robust Distillation via Untargeted and Targeted Intermediate Adversarial Samples
abstract
Adversarially robust knowledge distillation aims to com-press large-scale models into lightweight models while preserving adversarial robustness and natural performance on a given dataset. Existing methods typically align probability distributions of natural and adversarial samples between teacher and student models, but they overlook intermediate adversarial samples along the “adversarial path” formed by the multi-step gradient ascent of a sample towards the decision boundary. Such paths capture rich information about the decision boundary. In this paper, we propose a novel adversarially robust knowledge distillation approach by incorporating such adversarial paths into the alignment process. Recognizing the diverse impacts of intermediate adversarial samples (ranging from benign to noisy), we propose an adaptive weighting strategy to selectively em-phasize informative adversarial samples, thus ensuring efficient utilization of lightweight model capacity. Moreover, we propose a dual-branch mechanism exploiting two following insights: (i) complementary dynamics of adversar-ial paths obtained by targeted and untargeted adversarial learning, and (ii) inherent differences between the gradient ascent path from class$c_{i}$towards the nearest class bound-ary and the gradient descent path from a specific class$c_{j}$towards the decision region of$c_{i}(i\neq j)$. Comprehensive experiments demonstrate the effectiveness of our method on lightweight models under various settings.
Junhao Dong 0001, Piotr Koniusz, Junxi Chen, Z. Jane Wang 0001, Yew-Soon Ong
CVPR4
2024 UAV-Based Dynamic Object Tracking with Radio Map
abstract
Acting as dynamic base stations, unmanned aerial vehicles have a significant advantage over conventional base stations for dynamic object localization and tracking. In practice, however, the localization and tracking performance highly depends on the observation sequence of time-variant received signal strength (RSS) on UAVs from the dynamic object, which can be severely distorted by complex land layouts. In this paper, we propose dynamic object tracking by UAV with radio map, which contains abundant time-variant RSS knowledge. We generate time-varying radio maps from complex real-world topographic data. Then, we derive the grid-based method for dynamic object position estimation with the non-memoryless observations from the dynamic object. Numerical results show that the proposed algorithm can considerably improve tracking precision and efficiency for real-world topographical data.
Yangrui Dong, Fan Li 0001, Cunyan Ma, Chen He 0002, Z. Jane Wang 0001
ICASSP5
2024 Transferable and high-quality adversarial example generation leveraging diffusion model
abstract
In recent years, adversarial example methods in deep learning have proliferated. Meanwhile diffusion models have gained wide applications across various tasks due to their superior distribution reconstruction capability. By leveraging the design experience from prior adversarial examples, combining it with the modeling proficiency of diffusion models, and employing a cost function to evaluate image smoothness for regulating the regional distribution of adversarial noise, we propose a novel adversarial example design method. Experiments demonstrate that our generated adversarial examples exhibit both high attack success rate and superior image quality. The controlled distribution region of adversarial noises significantly enhances the subjective visual quality of our generated images.
Kangze Xu, Ziqiang He, Xiangui Kang, Z. Jane Wang 0001
ICME4
2024 AdvAD: Exploring Non-Parametric Diffusion for Imperceptible Adversarial Attacks
abstract
Imperceptible adversarial attacks aim to fool DNNs by adding imperceptible perturbation to the input data. Previous methods typically improve the imperceptibility of attacks by integrating common attack paradigms with specifically designed perception-based losses or the capabilities of generative models. In this paper, we propose Adversarial Attacks in Diffusion (AdvAD), a novel modeling framework distinct from existing attack paradigms. AdvAD innovatively conceptualizes attacking as a non-parametric diffusion process by theoretically exploring basic modeling approach rather than using the denoising or generation abilities of regular diffusion models requiring neural networks. At each step, much subtler yet effective adversarial guidance is crafted using only the attacked model without any additional network, which gradually leads the end of diffusion process from the original image to a desired imperceptible adversarial example. Grounded in a solid theoretical foundation of the proposed non-parametric diffusion process, AdvAD achieves high attack efficacy and imperceptibility with intrinsically lower overall perturbation strength. Additionally, an enhanced version AdvAD-X is proposed to evaluate the extreme of our novel framework under an ideal scenario. Extensive experiments demonstrate the effectiveness of the proposed AdvAD and AdvAD-X. Compared with state-of-the-art imperceptible attacks, AdvAD achieves an average of 99.9% (+17.3%) ASR with 1.34 (-0.97) $l_2$ distance, 49.74 (+4.76) PSNR and 0.9971 (+0.0043) SSIM against four prevalent DNNs with three different architectures on the ImageNet-compatible dataset. Code is available at https://github.com/XianguiKang/AdvAD.
Ziqiang He, Anwei Luo, Jianfang Hu, Z. Jane Wang 0001, Xiangui Kang
NeurIPS5
2024 TFAN: A Task-adaptive Feature Alignment Network for few-shot website fingerprinting attacks on Tor
Qiuyun Lyu, Huihui Xie, Wei Wang 0527, Yanyu Cheng, Yongqun Chen, Z. Jane Wang 0001
Comput. Secur.6
2024 APDF: An active preference-based deep forest expert system for overall survival prediction in gastric cancer
Qiucen Li, Zedong Du, Weihan Zhang, Fangming Zhong, Z. Jane Wang 0001, Zhikui Chen
Expert Syst. Appl.7
2024 Occlusion-robust FAU recognition by mining latent space of masked autoencoders
Minyang Jiang, Martin J. McKeown, Z. Jane Wang 0001
Neurocomputing4
2024 Self-supervised anatomical continuity enhancement network for 7T SWI synthesis from 3T SWI
Dong Zhang 0009, Caohui Duan, Udunna Anazodo, Z. Jane Wang 0001
Medical Image Anal.4
2024 Coverage Analysis of Single-Swarm mmWave UAV Networks Under Multiple Types of Blockages
abstract
Millimeter wave (mmWave)-based unmanned aerial vehicle (UAV) communication is susceptible to blockages, even from humans. Previous studies that primarily focused only on static blockage may not accurately characterize the system performance. This paper investigates the coverage performance of mmWave UAV networks by jointly considering multiple types of blockages under finite homogeneous Poisson point process and Binomial point process, which are commonly employed in finite area scenarios with random and fixed number of UAVs, respectively. Particularly, we derive the average line-of-sight probability and coverage probability under static, dynamic, and self blockages. Simulations verify our theoretical results, demonstrating that: the above system performance predominantly depends on self-blockage if UAVs are at high altitudes. Conversely, at relatively low altitudes, all three types of blockages impact them, with static blockage being the dominant factor. To avoid self-blockage, UAV height should satisfy$h\!\gt \!h_{R}\!+\!\frac {r_{i}}{\tan \varphi _{b}}$, where$h_{R}$is the height of the user equipment (UE),$r_{i}$is the two-dimensional distance of the UAV-UE link,$\varphi _{b}$is the elevation angle between UE and UAV. The required height is proportional to$r_{i}$and increases as distance d between the user and UE decreases, as$\varphi _{b}$is proportional to d. The findings help on designing the network parameters. To our best knowledge, this is the first work to analyze the coverage of mmWave UAV networks under multiple types of blockages.
Cunyan Ma, Xiaoya Li 0003, Chen He 0002, Jinye Peng 0001, Kun Yang 0001, Z. Jane Wang 0001
IEEE Trans. Commun.6
2024 Incentivizing Socio-Ethical Integrity in Decentralized Machine Learning Ecosystems for Collaborative Knowledge Sharing
abstract
To broaden domain knowledge and enable advanced analytics, machine learning (ML) algorithms increasingly utilize comprehensive datasets across diverse sectors. However, these disparate datasets held by various stakeholders raise concerns over data heterogeneity, privacy, and security. Decentralized ML research aims to protect data privacy and integrate knowledge bases, especially knowledge graphs, to address data heterogeneity challenges. Yet, the question of how to foster trustworthy collaborations in decentralized ML ecosystems remains underexplored. This study pioneers two innovative socio-economic mechanisms designed to ensure dependable collaborations with socio-ethical integrity within a decentralized knowledge inference framework, enabling participants to share knowledge while maintaining data privacy and ethical standards. We employ an evolutionary game theory model to analyze the dynamic interactions between requestors and workers, focusing on achieving a stable equilibrium through theoretical and numerical evaluations. Furthermore, we explore how various critical factors, such as incentive schemes and the accuracy of identifying malicious workers, influence the system's equilibrium, providing insights into optimizing collaborative efforts in decentralized ML ecosystems.
Yuanfang Chi, Jiaxiang Sun, Wei Cai 0002, Z. Jane Wang 0001, Victor C. M. Leung
IEEE Trans. Comput. Soc. Syst.5
2024 MM-Net: A MixFormer-Based Multi-Scale Network for Anatomical and Functional Image Fusion
abstract
Anatomical and functional image fusion is an important technique in a variety of medical and biological applications. Recently, deep learning (DL)-based methods have become a mainstream direction in the field of multi-modal image fusion. However, existing DL-based fusion approaches have difficulty in effectively capturing local features and global contextual information simultaneously. In addition, the scale diversity of features, which is a crucial issue in image fusion, often lacks adequate attention in most existing works. In this paper, to address the above problems, we propose a MixFormer-based multi-scale network, termed as MM-Net, for anatomical and functional image fusion. In our method, an improved MixFormer-based backbone is introduced to sufficiently extract both local features and global contextual information at multiple scales from the source images. The features from different source images are fused at multiple scales based on a multi-source spatial attention-based cross-modality feature fusion (CMFF) module. The scale diversity of the fused features is further enriched by a series of multi-scale feature interaction (MSFI) modules and feature aggregation upsample (FAU) modules. Moreover, a loss function consisting of both spatial domain and frequency domain components is devised to train the proposed fusion model. Experimental results demonstrate that our method outperforms several state-of-the-art fusion methods on both qualitative and quantitative comparisons, and the proposed fusion model exhibits good generalization capability. The source code of our fusion method will be available at https://github.com/yuliu316316.
Yu Liu 0023, Juan Cheng 0004, Z. Jane Wang 0001, Xun Chen 0001
IEEE Trans. Image Process.4
2024 Dynamic Object Tracking by Multi-UAV With Time-Variant Radio Maps
abstract
As aerial mobile receivers, unmanned aerial vehicles (UAVs) possess the capability to localize and track dynamically moving objects that broadcast signals. In practice, the performance of localization and tracking is heavily reliant on the observation sequence of time-variant received signal strength (RSS) from the dynamic object on UAVs. However, this sequence can be significantly distorted by complex land layouts. In this paper, we propose dynamic object tracking by muti-UAV with time-variant radio maps, which contain rich information on the impact of land layouts on RSS. We generate the time-variant radio maps from real-world topographic data, which can be highly complex, from the multi-UAV perspectives. Then, a grid-based method and a particle filter based method, both tailored to the observed RSS from the dynamic object and the generated time-variant radio maps, are derived for the case where both the UAVs and the object are dynamic. Simulations validate the two proposed algorithms and show that time-variant radio maps can significantly improve tracking precision and efficiency compared to other benchmarks for both the mountainous area and the urban area with highly complex layouts. In addition, the particle filter based method can reduce the computational complexity via Monte Carlo sampling of the time-variant radio maps with little sacrifice in tracking precision.
Yangrui Dong, Chen He 0002, Z. Jane Wang 0001
IEEE Trans. Wirel. Commun.3
2023 Capacity Maximization for Active RIS Assisted Outdoor-to-Indoor Communication System
abstract
In this paper, we aim to implement outdoor-to-indoor communication with the aid of an active reconfigurable intelligent surface (active-RIS), where the active-RIS allows the incoming signal from an outdoor base station (BS) to pass through the surface and be received by indoor users (UEs) after shifting phase and magnifying amplitude. Then, the problem of joint beamforming optimization of the BS and active-ITS is investigated to maximize the system capacity. A computationally efficient algorithm is developed to solve it with suboptimal solutions. Simulations indicate the superiority of the active-RIS and the effectiveness of the proposed algorithm.
Chen He 0002, Weisheng Gong, Yangrui Dong, Xie Xie, Z. Jane Wang 0001
ICASSP5
2023 Radio Map Based UAV Target Localization
abstract
Acting as the dynamic base stations, the unmanned aerial vehicles has a significant advantage over the conventional base station for target localization. When localizing a target, UAV can move toward the estimated target location after each estimation, making the received signals stronger after each iteration, and hence yield more localization precision for each iteration. In this paper, we propose target localization by UAV with radio maps as prior knowledge, which are generated from real topographic maps. We derived the minimum mean square error (MMSE) estimator for target localization by UAV with the radio map. Numerical results show that with radio map knowledge and the MMSE estimator, localization precision can be considerably improved.
Chen He 0002, Weisheng Gong, Yangrui Dong, Xie Xie, Z. Jane Wang 0001
ICASSP5
2023 Distilling and transferring knowledge via cGAN-generated samples for image classification and regression
Xin Ding 0004, Zuheng Xu, Z. Jane Wang 0001, William J. Welch
Expert Syst. Appl.4
2023 Efficient subsampling of realistic images from GANs conditional on a class or a continuous variable
Xin Ding 0004, Z. Jane Wang 0001, William J. Welch
Neurocomputing3
2023 SemiPFL: Personalized Semi-Supervised Federated Learning Framework for Edge Intelligence
abstract
Recent advances in wearable devices and Internetof-Things (IoT) have led to massive growth in sensor data generated in edge devices. Labeling such massive data for classification tasks has proven to be challenging. In addition, data generated by different users bear various personal attributes and edge heterogeneity, rendering it impractical to develop a global model that adapts well to all users. Concerns over data privacy and communication costs also prohibit centralized data accumulation and training. We propose SemiPFL that supports edge users having no label or limited labeled datasets and a sizable amount of unlabeled data that is insufficient to train a well-performing model. In this work, edge users collaborate to train a Hypernetwork in the server, generating personalized autoencoders for each user. After receiving updates from edge users, the server produces a set of base models for each user, which the users locally aggregate them using their own labeled dataset. We comprehensively evaluate our proposed framework on various public datasets from a wide range of application scenarios, from wearable health to IoT, and demonstrate that SemiPFL outperforms state-of-art federated learning frameworks under the same assumptions regarding user performance, network footprint, and computational consumption. We also show that the solution performs well for users without label or having limited labeled datasets and increasing performance for increased labeled data and number of users, signifying the effectiveness of SemiPFL for handling data heterogeneity and limited annotation. We also demonstrate the stability of SemiPFL for handling user hardware resource heterogeneity in three real-time scenarios.
Arvin Tashakori, Z. Jane Wang 0001, Peyman Servati
IEEE Internet Things J.3
2023 Automatic labeling of Parkinson's Disease gait videos with weak supervision
Mohsen Gholami, Rabab K. Ward, Ravneet Mahal, Maryam S. Mirian, Kevin Yen, Kye Won Park, Martin J. McKeown, Z. Jane Wang 0001
Medical Image Anal.8
2023 SSD-KD: A self-supervised diverse knowledge distillation method for lightweight skin lesion classification using dermoscopic images
Jiayue Cai, Tim K. Lee, Chunyan Miao, Z. Jane Wang 0001
Medical Image Anal.6
2023 Continuous Conditional Generative Adversarial Networks: Novel Empirical Losses and Label Input Mechanisms
abstract
This article focuses on conditional generative modeling (CGM) for image data with continuous, scalar conditions (termed regression labels). We propose the first model for this task which is called continuous conditional generative adversarial network (CcGAN). Existing conditional GANs (cGANs) are mainly designed for categorical conditions (e.g., class labels). Conditioning on regression labels is mathematically distinct and raises two fundamental problems: (P1) since there may be very few (even zero) real images for some regression labels, minimizing existing empirical versions of cGAN losses (a.k.a. empirical cGAN losses) often fails in practice; and (P2) since regression labels are scalar and infinitely many, conventional label input mechanisms (e.g., combining a hidden map of the generator/discriminator with a one-hot encoded label) are not applicable. We solve these problems by: (S1) reformulating existing empirical cGAN losses to be appropriate for the continuous scenario; and (S2) proposing a naive label input (NLI) mechanism and an improved label input (ILI) mechanism to incorporate regression labels into the generator and the discriminator. The reformulation in (S1) leads to two novel empirical discriminator losses, termed the hard vicinal discriminator loss (HVDL) and the soft vicinal discriminator loss (SVDL) respectively, and a novel empirical generator loss. Hence, we propose four versions of CcGAN employing different proposed losses and label input mechanisms. The error bounds of the discriminator trained with HVDL and SVDL, respectively, are derived under mild assumptions. To evaluate the performance of CcGANs, two new benchmark datasets (RC-49 and Cell-200) are created. A novel evaluation metric (Sliding Fréchet Inception Distance) is also proposed to replace Intra-FID when Intra-FID is not applicable. Our extensive experiments on several benchmark datasets (i.e., RC-49, UTKFace, Cell-200, and Steering Angle with both low and high resolutions) support the following findings: the proposed CcGAN is able to generate diverse, high-quality samples from the image distribution conditional on a given regression label; and CcGAN substantially outperforms cGAN both visually and quantitatively.
Xin Ding 0004, Zuheng Xu, William J. Welch, Z. Jane Wang 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2023 Decoupling multi-task causality for improved skin lesion segmentation and classification
Lei Song 0003, Haoqian Wang, Z. Jane Wang 0001
Pattern Recognit.3
2023 Smoothing group L1/2 regularized discriminative broad learning system for classification and regression
Dengxiu Yu, Qian Kang, Z. Jane Wang 0001, Xuelong Li 0001
Pattern Recognit.4
2023 Radio Map Assisted Multi-UAV Target Searching
abstract
When acting as dynamic receivers, unmanned aerial vehicles (UAVs) have several advantages over the conventional static base stations for received signal strength (RSS) based target searching. It is highly likely that the RSS is becoming stronger as the UAV is moving towards the estimated coordinate, and this leads to a higher estimation precision for each localization update. On the other hand, despite its low cost, low energy consumption and simple hardware requirement, RSS can be severely influenced by land layouts, therefore the assumption of a simple free space propagation model may lead to poor estimation precision and even failure, especially for an area where the topography is highly complex. In this paper, we propose multi-UAV target searching with the assistance of radio maps, which contain abundant information of the topography influence on RSS. We generate radio maps from real topographic data, and derive the minimum mean square error estimator for multi-UAV target searching with the radio map knowledge for memoryless observations. Simulations for areas with complex topography, show that the localization precision and searching efficiency can be considerably improved with the assistance of radio maps.
Chen He 0002, Yangrui Dong, Z. Jane Wang 0001
IEEE Trans. Wirel. Commun.3
2023 Joint Precoding for Active Intelligent Transmitting Surface Empowered Outdoor-to-Indoor Communication in mmWave Cellular Networks
abstract
Outdoor-to-indoor communications in millimeter-wave (mmWave) cellular networks have been one challenging research problem due to the severe attenuation and the high penetration loss caused by propagation characteristics of mmWave signals. We propose a viable solution to implement the outdoor-to-indoor mmWave communication with the aid of an active intelligent transmitting surface (active-ITS), where the active-ITS allows the incoming signal from an outdoor base station (BS) to pass through the surface and be received by indoor users (UEs) after shifting its phase and magnifying its amplitude. Then, the problem of joint precoding of the BS and active-ITS is investigated to maximize the weighted sum-rate (WSR) of the system. An efficient block coordinate descent (BCD) based algorithm is developed to solve it with the suboptimal solutions in nearly closed-forms. In addition, to reduce the size and hardware cost of active-ITSs, we provide a block-amplifying architecture to partially remove the circuit components for power-amplifying, where multiple transmissive-type elements (TEs) in each block share the same power amplifier. Simulations indicate that active-ITS has the potential of achieving a given performance with much fewer TEs compared to the passive-ITS under the same total system power consumption, which makes it suitable for application to the space-limited and aesthetic-needed scenario, and the performance degradation caused by the block-amplifying architecture is negligible.
Xie Xie, Chen He 0002, Feifei Gao 0001, Zhu Han 0001, Z. Jane Wang 0001
IEEE Trans. Wirel. Commun.6
2022 AdaptPose: Cross-Dataset Adaptation for 3D Human Pose Estimation by Learnable Motion Generation
abstract
This paper addresses the problem of cross-dataset generalization of 3D human pose estimation models. Testing a pre-trained 3D pose estimator on a new dataset results in a major performance drop. Previous methods have mainly addressed this problem by improving the diversity of the training data. We argue that diversity alone is not sufficient and that the characteristics of the training data need to be adapted to those of the new dataset such as camera view-point, position, human actions, and body size. To this end, we propose AdaptPose, an end-to-end framework that generates synthetic 3D human motions from a source dataset and uses them to fine-tune a 3D pose estimator. AdaptPose follows an adversarial training scheme. From a source 3D pose the generator generates a sequence of 3D poses and a camera orientation that is used to project the generated poses to a novel view. Without any 3D labels or camera information AdaptPose successfully learns to create synthetic 3D poses from the target dataset while only being trained on 2D poses. In experiments on the Human3.6M, MPI-INF-3DHp, 3DPW, and Ski-Pose datasets our method outperforms previous work in cross-dataset evaluations by 14% and previous semi-supervised learning methods that use partial 3D annotations by 16%.
Mohsen Gholami, Bastian Wandt, Helge Rhodin, Rabab K. Ward, Z. Jane Wang 0001
CVPR5
2022 Self-supervised 3D human pose estimation from video
Mohsen Gholami, Ahmad Rezaei, Helge Rhodin, Rabab K. Ward, Z. Jane Wang 0001
Neurocomputing5
2022 Knowledge-Based Fault Diagnosis in Industrial Internet of Things: A Survey
abstract
Industrial Internet of Things (IIoT) systems connect a plethora of smart devices, such as sensors, actuators, and controllers, to enable efficient industrial productions in manners observable and controllable by human beings. Plain model-based and data-driven diagnosis approaches can be used for fault detection and isolation of specific IIoT components. However, the physical models, signal patterns, and machine learning algorithms need to be carefully designed to describe system faults. Besides, the ever-increasing level of connectivity among devices can induce exponential complexity. Knowledge-based fault diagnosis approaches improve interoperability via ontologies so that high-level reasoning and inquiry response can be provided to nonexpert users. Therefore, knowledge-based fault diagnosis approaches are preferred over plain model-based and data-driven diagnosis approaches in recent IIoT systems. In the context of IIoT systems, this work reviews the recent progress on the construction of knowledge bases via ontologies and deductive/inductive reasoning for knowledge-based fault diagnosis. Besides, general inductive reasoning methods are discussed to shed light on their successful applications in knowledge-based fault diagnosis for IIoT systems. Following the trend of large-system decentralization, future fault diagnosis also requires decentralized implementations. Therefore, we conclude this survey by discussing several interesting open problems for decentralized knowledge-based fault diagnosis for IIoT systems.
Yuanfang Chi, Yanjie Dong 0003, Z. Jane Wang 0001, F. Richard Yu, Victor C. M. Leung
IEEE Internet Things J.3
2022 A Joint Optimization Framework for IRS-Assisted Energy Self-Sustainable IoT Networks
abstract
Energy self-sustainability is critically important for future Internet of Things (IoT) networks to support an ever-growing massive number of wireless devices with low maintenance cost and high spectrum/energy efficiency. Power-splitting (PS)-based simultaneous wireless information and power transfer (PS-SWIPT) is a promising solution to realize it. However, the performance of PS-SWIPT is severely influenced by the channel attenuation caused by the detrimental radio propagation environment. Intelligent reflecting surface (IRS) is an emerging technology that can reconfigure the incident signal with considerable array gain so as to improve the PS-SWIPT performance. Thus, in this article, we investigate the weighted sumrate (WSR) maximization problem of the IRS-assisted multi-input–multioutput (MIMO) PS-SWIPT IoT network with multiple low-power IoT PS-based devices (PSDs). The formulated problem is nonconvex and arduous to tackle due to the presence of the intricately coupled variables and the mutually exclusive constraints. To the best of our knowledge, the problem is not addressed yet and cannot be solved by employing the existing methods directly. To cope with the problem, we develop a joint optimization framework that decomposes the original problem into several subproblems that can be solved alternately. Simulation results confirm the effectiveness of IRS to improve the WSR of the PS-SWIPT energy self-sustainable IoT networks and demonstrate that the proposed algorithm outperforms benchmark methods considerably.
Xie Xie, Chen He 0002, Huixu Luan, Yangrui Dong, Kun Yang 0001, Feifei Gao 0001, Z. Jane Wang 0001
IEEE Internet Things J.7
2022 Ship Target Segmentation for SAR Images Based on Clustering Center Shift
abstract
Ship target segmentation plays an important role in synthetic aperture radar (SAR) image interpretation. However, existing segmentation methods for marine SAR images have the problem of inaccurate edge segmentation, a concern for real-world applications. In this letter, we propose a clustering center shifted adaptive target segmentation (CCSATS) method. Firstly, the proposed clustering center shift method is used to update the clustering centers of each iteration, which can quickly and accurately capture ship pixels. Then, based on regional homogeneity coefficients, we define a new similarity measurement criterion with two adaptive weight factors to ensure the homogeneity of segmentation results. Finally, neighborhood patches are used to represent pixel information, which can reduce the influence of speckle noise and enhance the target edge fitting ability. Our segmentation results of measured SAR images show that the proposed method effectively ensures segmentation accuracy. Compared with other existing methods, the proposed target segmentation method achieves better edge capture performance.
Rufei Wang, Fanyun Xu, Jifang Pei, Weibo Huo, Yulin Huang 0001, Yin Zhang 0003, Jianyu Yang 0001, Z. Jane Wang 0001
IEEE Geosci. Remote. Sens. Lett.8
2022 Cross-Domain Few-Shot Contrastive Learning for Hyperspectral Images Classification
abstract
Deep learning has achieved impressive results on Hyperspectral image (HSI) classification, which generally requires sufficient training samples and a huge number of parameters. However, it is challenging to label HSIs, and likely only a few samples are available in practice. Learning a large number of parameters by the model is also resource-intensive. This paper proposes an HSI classification model that achieves promising classification performance with fewer parameters in few-shot settings. The proposed model adopts the residual 3D-CNN as feature extraction network, and contrastive learning is introduced to learn more discriminative representations for HSIs which can conquer the obstacles from HSIs’ high inter-class similarity and large intra-class variance. The proposed few-shot contrastive learning HSI classification model is tested on five popular HSI datasets and outperforms the state-of-the-art models.
Suhua Zhang, Zhikui Chen, Dan Wang 0011, Z. Jane Wang 0001
IEEE Geosci. Remote. Sens. Lett.4
2022 Multilabel Aerial Image Classification With a Concept Attention Graph Neural Network
abstract
Compared with natural images, aerial images collected by satellite sensors/aerial cameras can provide a much larger field of view and often contain multiple objects of interest (multiple labels). There are certain limitations of existing multilabel aerial image classification methods. First, label correlations were often ignored in previous MAIC work, and thus, multilabel classifiers failed to be self-adapted. Second, existing multilabeled data sets for aerial images only cover limited images with fixed labels. Therefore, the underlying semantic correlations of labels cannot be fully included, while such correlation information is implicitly used as common knowledge by human beings. To tackle these concerns, we propose a novel multilabel classification method for aerial images. Our contributions are twofold. First, as the first attempt, label correlations are inferred from both the specific data set and ConceptNet (a popular knowledge graph for common sense). Second, based on graph neural network (GNN), we propose a novel end-to-end aerial image classification model, named the multiple label concept graph (ML-CG). ML-CG builds a concept graph to describe the semantic correlations from both the label set and the ConceptNet. We also incorporate both semantic attention and label attention in the GNN to better extract meaningful information of image labels. Compared with state-of-the-art methods, the effectiveness of the proposed method is demonstrated on both the commonly used UCM data set and a recently proposed DFC15 data set with high image resolution.
Dan Lin 0008, Jianzhe Lin, Liang Zhao 0005, Z. Jane Wang 0001, Zhikui Chen
IEEE Trans. Geosci. Remote. Sens.4
2022 Multilabel Aerial Image Classification With Unsupervised Domain Adaptation
abstract
Deep learning (DL) methods are promising for the multilabel aerial image classification (MAIC) task. However, current DL methods face a common problem: the need for large multilabeled datasets. Collecting and annotating raw aerial image datasets can be extremely time- and labor-consuming. To address this concern in MAIC, domain adaptation (DA) provides a novel solution by transferring the knowledge learned from a label-rich dataset (i.e., the source domain) to a label-scarce dataset (i.e., the target domain), while current DA models are mainly designed for single-labeled tasks. In this article, we propose a novel end-to-end MAIC model based on DA techniques, named DA-MAIC. To the best of our knowledge, this article for the first time integrates DA to tackle the label scarcity problem in the MAIC task. Specifically, the proposed DA-MAIC is composed of two main parts: the image classifier and the domain classifier. The image classifier captures task-discriminative features based on the graph convolutional network (GCN) to predict multiple image labels; and the domain classifier extracts domain-invariant representations, which mitigates the domain shift between two underlying distributions. We extensively evaluate the proposed DA-MAIC from different perspectives on three benchmark datasets, including the commonly used UCM dataset, the high-resolution AID dataset, and the recently proposed DFC15 dataset. Both quantitative and qualitative results support that the proposed DA-MAIC can generalize the source domain knowledge to new scenarios and substantially improve the classification performance on the target domain task.
Dan Lin 0008, Jianzhe Lin, Liang Zhao 0005, Z. Jane Wang 0001, Zhikui Chen
IEEE Trans. Geosci. Remote. Sens.4
2022 Rethinking Crowdsourcing Annotation: Partial Annotation With Salient Labels for Multilabel Aerial Image Classification
abstract
Annotated images are required for supervised model training and evaluation in aerial image classification. Manually annotating images is arduous and expensive, especially for aerial images, which often cover a large land area with multiple labels. A recent trend for conducting such annotation tasks is through crowdsourcing, where images are annotated by volunteers or paid workers (e.g., annotation volunteers for Open Street Map) online from scratch. However, for crowdsourcing image annotations, the quality cannot be guaranteed, and incompleteness and incorrectness are two major concerns. To address such concerns, we have a rethinking of crowdsourcing annotations: Our simple hypothesis is that if annotators only partially annotate multi-label images with salient labels they are confident in, there will be fewer annotation errors and annotators will spend less time on uncertain labels. As a pleasant surprise, with the same annotation budget, we show that a multi-label aerial image classifier supervised by images with salient annotations can outperform models supervised by fully annotated images. Our contributions are 2-fold: An active learning way is proposed to acquire salient labels for multi-label aerial images; and a novel Adaptive Temperature Associated Model (ATAM) specifically using partial annotations is proposed for multi-label aerial image classification. When tested on practical crowdsourcing aerial data, the Open Street Map (OSM) dataset, the proposed ATAM can achieve higher accuracy than state-of-the-art classification methods trained on fully annotated images. The proposed idea is promising for crowdsourcing aerial image annotation. Our code will be publicly available.
Jianzhe Lin, Tianze Yu, Z. Jane Wang 0001
IEEE Trans. Geosci. Remote. Sens.3
2022 SCIDA: Self-Correction Integrated Domain Adaptation From Single- to Multi-Label Aerial Images
abstract
Most publicly available datasets for image classification are with single labels, while images are inherently multilabeled in our daily life. Such an annotation gap makes many pretrained single-label classification models fail in practical scenarios. For aerial images, this annotation issue is more concerned: Aerial data naturally cover a relatively large land area with multiple labels, while annotated aerial datasets currently publicly available (e.g., UCM and AID) are single-labeled. As manually annotating multilabel aerial images (MAIs) would be time-/ labor-consuming, we propose a novel self-correction integrated domain adaptation (SCIDA) method for automatic multilabel learning. SCIDA is weakly supervised, i.e., automatically learning the multilabel image classification model from using massive, publicly available single-label images. To achieve this goal, we propose a novel labelwise self-correction (LWC) module to better explore underlying label correlations. This module also makes the unsupervised domain adaptation (UDA) from single-label to multilabel data possible. For model training, the proposed method uses single-label information yet requires no prior knowledge of multilabeled data and predicts labels for MAIs. Through extensive evaluations, the proposed model, which is trained with single-labeled MAI-AID-s and MAI-UCM-s datasets, achieves much better performances than comparative methods on our collected multiscene aerial image dataset. The code and data are available on GitHub (https://github.com/Ryan315/Single2multi-DA).
Tianze Yu, Jianzhe Lin, Lichao Mou, Yuansheng Hua, Xiao Xiang Zhu 0001, Z. Jane Wang 0001
IEEE Trans. Geosci. Remote. Sens.6
2022 Breast Cancer Detection Using Multimodal Time Series Features From Ultrasound Shear Wave Absolute Vibro-Elastography
abstract
In shear wave absolute vibro-elastography (S-WAVE), a steady-state multi-frequency external mechanical excitation is applied to tissue, while a time-series of ultrasound radio-frequency (RF) data are acquired. Our objective is to determine the potential of S-WAVE to classify breast tissue lesions as malignant or benign. We present a new processing pipeline for feature-based classification of breast cancer using S-WAVE data, and we evaluate it on a new data set collected from 40 patients. Novel bi-spectral and Wigner spectrum features are computed directly from the RF time series and are combined with textural and spectral features from B-mode and elasticity images. The Random Forest permutation importance ranking and the Quadratic Mutual Information methods are used to reduce the number of features from 377 to 20. Support Vector Machines and Random Forest classifiers are used with leave-one-patient-out and Monte Carlo cross-validations. Classification results obtained for different feature sets are presented. Our best results (95% confidence interval, Area Under Curve = 95%±1.45%, sensitivity = 95%, and specificity = 93%) outperform the state-of-the-art reported S-WAVE breast cancer classification performance. The effect of feature selection and the sensitivity of the above classification results to changes in breast lesion contours is also studied. We demonstrate that time-series analysis of externally vibrated tissue as an elastography technique, even if the elasticity is not explicitly computed, has promise and should be pursued with larger patient datasets. Our study proposes novel directions in the field of elasticity imaging for tissue classification.
Yanan Shao, Hoda S. Hashemi, Paula Gordon, Linda Warren, Z. Jane Wang 0001, Robert Rohling, Tim Salcudean
IEEE J. Biomed. Health Informatics5
2022 LiCaS3: A Simple LiDAR-Camera Self-Supervised Synchronization Method
abstract
Recent advances in robotics and deep learning demonstrate promising 3-D perception performances via fusing the light detection and ranging (LiDAR) sensor and camera data, where both spatial calibration and temporal synchronization are generally required. While the LiDAR–camera calibration problem has been actively studied during the past few years, LiDAR–camera synchronization has been less studied and mainly addressed by employing a conventional pipeline consisting of clock synchronization and temporal synchronization. The conventional pipeline has certain potential limitations, which have not been sufficiently addressed and could be a bottleneck for the potential wide adoption of low-cost LiDAR–camera platforms. Different from the conventional pipeline, in this article, we propose the LiCaS3, the first deep-learning-based framework, for the LiDAR–camera synchronization task via self-supervised learning. The proposed LiCaS3 does not require hardware synchronization or extra annotations and can be deployed both online and offline. Evaluated on both the KITTI and Newer College datasets, the proposed method shows promising performances. The code will be publicly available athttps://github.com/KleinYuan/LiCaS3.
Kaiwen Yuan, Mazen Abdelfattah, Z. Jane Wang 0001
IEEE Trans. Robotics4
2022 A Better Than Alamouti OSTBC for MIMO Backscatter Communications
abstract
Backscatter communications have received increasingly attention in Internet of Things (IoT) due to its advantages of requiring no internal battery and almost zero maintenance. However, the drawback of backscatter communications is that, its channel fades deeper than conventional channel. To improve the performance, multiple-input multiple-output (MIMO) techniques have been investigated for backscatter communications. By considering duty cycle, a particular behavior of backscatter communications, as a new performance criterion, we propose a$2 \times 2$orthogonal space-time block code for MIMO backscatter communications. The new code is a rotation of the classical Alamouti code, and in MIMO backscatter channels, this simple rotation will incur an improved duty cycle for both linear and nonlinear energy harvester models, as well as a same, or even better symbol error rate performance. We provide both rigorous mathematical proofs and numerical simulations to show the above results. An interesting thing is that, this new code does not outperform Alamouti code in conventional MIMO channels, which indicates that the space-time code design criteria for conventional channels is not the optimal one for backscatter channels, and existing space-time codes should be reconsidered and/or redesigned to achieve a higher performance in MIMO backscatter communications.
Huixu Luan, Xie Xie, Luyang Han, Chen He 0002, Z. Jane Wang 0001
IEEE Trans. Wirel. Commun.5
2021 Towards Universal Physical Attacks on Single Object Tracking
abstract
Recent studies show that small perturbations in video frames could misguide single object trackers. However, such attacks have been mainly designed for digital-domain videos (i.e., perturbation on full images), which makes them practically infeasible to evaluate the adversarial vulnerability of trackers in real-world scenarios. Here we made the first step towards physically feasible adversarial attacks against visual tracking in real scenes with a universal patch to camouflage single object trackers. Fundamentally different from physical object detection, the essence of single object tracking lies in the feature matching between the search image and templates, and we therefore specially design the maximum textural discrepancy (MTD), a resolution-invariant and target location-independent feature de-matching loss. The MTD distills global textural information of the template and search images at hierarchical feature scales prior to performing feature attacks. Moreover, we evaluate two shape attacks, the regression dilation and shrinking, to generate stronger and more controllable attacks. Further, we employ a set of transformations to simulate diverse visual tracking scenes in the wild. Experimental results show the effectiveness of the physically feasible attacks on SiamMask and SiamRPN++ visual trackers both in digital and physical scenes.
Kaiwen Yuan, Minyang Jiang, Ping Wang 0009, Hua Huang 0001, Z. Jane Wang 0001
AAAI7
2021 A Capsule Network Based Approach for Detection of Audio Spoofing Attacks
abstract
Audio spoofing attacks not only increasingly pose a threat to automatic speaker verification systems but also have the potential to destabilize national security (e.g., by creating fake audio of influential politicians). The main purpose of anti-spoofing is to detect fake audios synthesized by advanced methods, while current algorithms using convolutional neural networks as classifiers exposed poor generalization to the unknown attacks. In this paper, as the first attempt, we introduce a capsule network to enhance the generalization of the detection system. To make the capsule network suitable for anti-spoofing tasks, we modified the original dynamic routing algorithm to force the model to pay more attention to artifacts and thus yield better detection performance for text-to-speech/voice conversion attacks. Furthermore, replay attack detection is also investigated, and the results indicate that our proposed approach is also highly capable of detecting replay attacks.
Anwei Luo, Enlei Li, Yongliang Liu, Xiangui Kang, Z. Jane Wang 0001
ICASSP5
2021 Multi-view 3D Reconstruction with Transformers
abstract
Deep CNN-based methods have so far achieved the state of the art results in multi-view 3D object reconstruction. Despite the considerable progress, the two core modules of these methods - view feature extraction and multi-view fusion, are usually investigated separately, and the relations among multiple input views are rarely explored. Inspired by the recent great success in Transformer models, we reformulate the multi-view 3D reconstruction as a sequence-to-sequence prediction problem and propose a framework named 3D Volume Transformer. Unlike previous CNN-based methods using a separate design, we unify the feature extraction and view fusion in a single Transformer network. A natural advantage of our design lies in the exploration of view-to-view relationships using self-attention among multiple unordered inputs. On ShapeNet - a large-scale 3D reconstruction benchmark, our method achieves a new state-of-the-art accuracy in multi-view reconstruction with fewer parameters (70% less) than CNN-based methods. Experimental results also suggest the strong scaling capability of our method. Our code will be made publicly available.
Dan Wang 0011, Xinrui Cui, Xun Chen 0001, Zhengxia Zou, Tianyang Shi, Tim Salcudean, Z. Jane Wang 0001, Rabab K. Ward
ICCV7
2021 Towards Universal Physical Attacks On Cascaded Camera-Lidar 3d Object Detection Models
abstract
We propose a universal and physically realizable adversarial attack on a cascaded multi-modal deep learning network (DNN), in the context of self-driving cars. DNNs have achieved high performance in 3D object detection, but they are known to be vulnerable to adversarial attacks. These attacks have been heavily investigated in the RGB image domain and more recently in the point cloud domain, but rarely in both domains simultaneously - a gap to be filled in this paper. We use a single 3D mesh and differentiable rendering to explore how perturbing the mesh’s geometry and texture can reduce the robustness of DNNs to adversarial attacks. We attack a prominent cascaded multi-modal DNN, the Frustum-Pointnet model. Using the popular KITTI benchmark, we showed that the proposed universal multi-modal attack was successful in reducing the model’s ability to detect a car by nearly 73%. This work can aid in the understanding of what the cascaded RGB-point cloud DNN learns and its vulnerability to adversarial attacks.
Mazen Abdelfattah, Kaiwen Yuan, Z. Jane Wang 0001, Rabab K. Ward
ICIP3
2021 ShuffleCount: Task-Specific Knowledge Distillation for Crowd Counting
abstract
One promising way to improve the performance of a small deep network is knowledge distillation. Performances of smaller student models with fewer parameters and lower computational cost can be comparable to that of larger teacher models in specific computer vision tasks. Knowledge distillation is especially attractive for the high-accuracy real-time crowd counting task in our daily lives, where the computational resource can be limited and the model efficiency is extremely important. In this paper, we propose a novel task-specific knowledge distillation framework for crowd counting, named ShuffleCount. Its main contributions are two-fold: First, different from existing frameworks, our task-specific ShuffleCount effectively learns from the teacher network through hierarchic feature regulation, and better avoids negative knowledge transferred from the teacher. Second, the proposed student network, i.e., the optimized Shufflenet, shows promising performances. When tested on the benchmark dataset Shanghai Tech A, it achieves a 15% higher accuracy yet keeps low computational cost when compared with the state-of-the-art MobileCount. Our code is available online at https://github.com/JiangMinyang/CC-KD.
Minyang Jiang, Jianzhe Lin, Z. Jane Wang 0001
ICIP3
2021 CcGAN: Continuous Conditional Generative Adversarial Networks for Image Generation
Xin Ding 0004, Zuheng Xu, William J. Welch, Z. Jane Wang 0001
ICLR5
2021 Adversarial Attacks on Camera-LiDAR Models for 3D Car Detection
abstract
Most autonomous vehicles (AVs) rely on LiDAR and RGB camera sensors for perception. Using these point cloud and image data, perception models based on deep neural nets (DNNs) have achieved state-of-the-art performance in 3D detection. The vulnerability of DNNs to adversarial attacks have been heavily investigated in the RGB image domain and more recently in the point cloud domain, but rarely in both domains simultaneously. Multi-modal perception systems used in AVs can be divided into two broad types: cascaded models which use each modality independently, and fusion models which learn from different modalities simultaneously. We propose a universal and physically realizable adversarial attack for each type, and study and contrast their respective vulnerabilities to attacks. We place a single adversarial object with specific shape and texture on top of a car with the objective of making this car evade detection. Evaluating on the popular KITTI benchmark, our adversarial object made the host vehicle escape detection by each model type more than 50% of the time. The dense RGB input contributed more to the success of the adversarial attacks on both cascaded and fusion models.
Mazen Abdelfattah, Kaiwen Yuan, Z. Jane Wang 0001, Rabab K. Ward
IROS3
2021 Delving into Deep Image Prior for Adversarial Defense: A Novel Reconstruction-based Defense Framework
abstract
Deep learning based image classification models are shown vulnerable to adversarial attacks by injecting deliberately crafted noises to clean images. To defend against adversarial attacks in a training-free and attack-agnostic manner, this work proposes a novel and effective reconstruction-based defense framework by delving into deep image prior (DIP). Fundamentally different from existing reconstruction-based defenses, the proposed method analyzes and explicitly incorporates the model decision process into our defense. Given an adversarial image, firstly we map its reconstructed images during DIP optimization to the model decision space, where cross-boundary images can be detected and on-boundary images can be further localized. Then, adversarial noise is purified by perturbing on-boundary images along the reverse direction to the adversarial image. Finally, on-manifold images are stitched to construct an image that can be correctly predicted by the victim classifier. Extensive experiments demonstrate that the proposed method outperforms existing state-of-the-art reconstruction-based methods both in defending white-box attacks and defense-aware attacks. Moreover, the proposed method can maintain a high visual quality during adversarial image reconstruction.
Xin Ding 0004, Kaiwen Yuan, Ping Wang 0009, Hua Huang 0001, Z. Jane Wang 0001
ACM Multimedia7
2021 A smartly simple way for joint crowd counting and localization
Minyang Jiang, Jianzhe Lin, Z. Jane Wang 0001
Neurocomputing3
2021 OCSID: Orthogonal Accessing Control Without Spectrum Spreading for Massive RFID Network
abstract
In radio-frequency identification (RFID)-based sensors networks, each sensor is integrated with a tag and sensors may co-exist in an area of interest. Before sending data packets, the sensors send their IDs for channel reservations. However, ID collisions happen frequently at the reservation stage which leads to significant time delays, especially for massive and dense networks. In this article, by employing group theory, we show that for B = 2k, where k is a positive integer, the set (-1, +1)Band the Hadamard product 0 forms a group ((-1, +1)B, 01 that can be divided into (2B/B) disjoint subsets, each of which has B binary vectors that are mutually orthogonal. Based on this finding, we propose orthogonal coset identification (OCSID) and its generalization, query tree (QT)-OCSID, that can recover ID information from collisions, and thus considerably improve the efficiency at the reservation stage, particularly when the network is large and/or dense. In an ideal case, it can recover/decode B tags for each query. The fundamental difference between the proposed OCSID schemes and code-division multiple accessbased schemes is that, OCSID achieves orthogonal design for ID information recovery by exploiting the inherent orthogonal structure of the binary vector set, instead of spreading the spectrum. Hence, it requires a narrower frequency band, lower circuit complexity, and lower synchronization precision, which are much more preferred by hardware limited devices.
Chen He 0002, Luyang Han, Nan Chen 0005, Z. Jane Wang 0001
IEEE Internet Things J.5
2021 A deep community based approach for large scale content based X-ray image retrieval
Nandinee Fariah Haq, Mehdi Moradi, Z. Jane Wang 0001
Medical Image Anal.3
2021 Perception matters: Exploring imperceptible and transferable anti-forensics for GAN-generated fake face imagery detection
Xin Ding 0004, Yixin Yang 0001, Rabab K. Ward, Z. Jane Wang 0001
Pattern Recognit. Lett.6
2021 Attention-Aware Pseudo-3-D Convolutional Neural Network for Hyperspectral Image Classification
abstract
Convolutional neural networks (CNNs) have been applied for hyperspectral image classification recently. Among this class of deep models, 3-D CNN has been shown to be more effective by learning discriminative features from abundant spectral signatures and spatial contexts in hyperspectral imagery (HSI). However, by simply imposing 3-D CNN to HSI, a large amount of initial information might be lost in this CNN pipeline. The proposed attention-aware pseudo-3-D (AP3D) convolutional network for HSI classification is motivated by two observations. First, each dimension of the 3-D HSI is not equally important, different attention should be paid to different dimensions of the initial HSI image, especially in the first convolution operation. Second, intermediate representations of the 3-D input image at different stages in the 3-D CNN pipeline represent different levels of features and should not be neglected and abandoned. Instead, a 2-D matrix of scores for each feature map should be fed to the final softmax layer. Quantitative and qualitative results demonstrate that the proposed AP3D model outperforms the state-of-the-art HSI classification methods in agricultural and rural/urban data sets: Indian Pines, Pavia University, and Salinas Scene.
Jianzhe Lin, Lichao Mou, Xiao Xiang Zhu 0001, Xiangyang Ji, Z. Jane Wang 0001
IEEE Trans. Geosci. Remote. Sens.5
2021 Unifying Top-Down Views by Task-Specific Domain Adaptation
abstract
In this article, we aim to learn a unified representation of images from satellite/aerial/ground views by exploring their underlying correlations. Inspired by recent advances in domain adaptation (DA), we propose a novel task-specific DA method for this purpose. Different from traditional DA methods, this proposed method not only applies task-specific classifiers1but also introduces domain-specific tasks for different domains during the adaptation process. The experiments are conducted on two newly proposed ground-/satellite-to-aerial scene adaptation (GSSA) data sets. Since the semantic gap between the ground/satellite scenes and the aerial scenes is much larger than that between ground scenes, the DA task between these scenes is more challenging than traditional DA tasks. On GSSA data sets, we not only demonstrate the proposed unsupervised DA method but also explore the few-shot DA in the discussion section. The proposed method is easy to implement, and our method substantially outperforms the state-of-the-art methods on the studied data sets.We hope that the proposed method for the novel GSSA data sets can be a good baseline for future researchers. The related data sets/codes will be available online.
Jianzhe Lin, Tianze Yu, Lichao Mou, Xiao Xiang Zhu 0001, Rabab K. Ward, Z. Jane Wang 0001
IEEE Trans. Geosci. Remote. Sens.6
2021 Interpreting Bottom-Up Decision-Making of CNNs via Hierarchical Inference
abstract
With the great success of convolutional neural networks (CNNs), interpretation of their internal network mechanism has been increasingly critical, while the network decision-making logic is still an open issue. In the bottom-up hierarchical logic of neuroscience, the decision-making process can be deduced from a series of sub-decision-making processes from low to high levels. Inspired by this, we propose the Concept-harmonized HierArchical INference (CHAIN) interpretation scheme. In CHAIN, a network decision-making process from shallow to deep layers is interpreted by the hierarchical backward inference based on visual concepts from high to low semantic levels. Firstly, we learned a general hierarchical visual-concept representation in CNN layered feature space by concept harmonizing model on a large concept dataset. Secondly, for interpreting a specific network decision-making process, we conduct the concept-harmonized hierarchical inference backward from the highest to the lowest semantic level. Specifically, the network learning for a target concept at a deeper layer is disassembled into that for concepts at shallower layers. Finally, a specific network decision-making process is explained as a form of concept-harmonized hierarchical inference, which is intuitively comparable to the bottom-up hierarchical visual recognition way. Quantitative and qualitative experiments demonstrate the effectiveness of the proposed CHAIN at both instance and class levels.
Dan Wang 0011, Xinrui Cui, Xun Chen 0001, Rabab K. Ward, Z. Jane Wang 0001
IEEE Trans. Image Process.5
2021 Bridging the Gap Between 2D and 3D Contexts in CT Volume for Liver and Tumor Segmentation
abstract
Automatic liver and tumor segmentation remain a challenging topic, which subjects to the exploration of 2D and 3D contexts in CT volume. Existing methods are either only focus on the 2D context by treating the CT volume as many independent image slices (but ignore the useful temporal information between adjacent slices), or just explore the 3D context lied in many little voxels (but damage the spatial detail in each slice). These factors lead an inadequate context exploration together for automatic liver and tumor segmentation. In this paper, we propose a novel full-context convolution neural network to bridge the gap between 2D and 3D contexts. The proposed network can utilize the temporal information along the Z axis in CT volume while retaining the spatial detail in each slice. Specifically, a 2D spatial network for intra-slice features extraction and a 3D temporal network for inter-slice features extraction are proposed separately and then are guided by the squeeze-and-excitation layer that allows the flow of 2D context and 3D temporal information. To address the severe class imbalance issue in the CT volume and meanwhile improve the segmentation performance, a loss function consisting of weighted cross-entropy and jaccard distance is proposed. During the network training, the 2D and 3D contexts are learned jointly in an end-to-end way. The proposed network achieves competitive results on the Liver Tumor Segmentation Challenge (LiTS) and the 3D-IRCADB datasets. This method should be a new promising paradigm to explore the contexts for liver and tumor segmentation.
Lei Song 0003, Haoqian Wang, Z. Jane Wang 0001
IEEE J. Biomed. Health Informatics3
2021 Co-Learning Non-Negative Correlated and Uncorrelated Features for Multi-View Data
abstract
Multi-view data can represent objects from different perspectives and thus provide complementary information for data analysis. A topic of great importance in multi-view learning is to locate a low-dimensional latent subspace, where common semantic features are shared by multiple data sets. However, most existing methods ignore uncorrelated items (i.e., view-specific features) and may cause semantic bias during the process of common feature learning. In this article, we propose a non-negative correlated and uncorrelated feature co-learning (CoUFC) method to address this concern. More specifically, view-specific (uncorrelated) features are identified for each view when learning the common (correlated) feature across views in the latent semantic subspace. By eliminating the effects of uncorrelated information, useful inter-view feature correlations can be captured. We design a new objective function in CoUFC and derive an optimization approach to solve the objective with the analysis on its convergence. Experiments on real-world sensor, image, and text data sets demonstrate that the proposed method outperforms the state-of-the-art multiview learning methods.
Liang Zhao 0005, Jie Zhang 0085, Zhikui Chen, Yi Yang 0006, Z. Jane Wang 0001
IEEE Trans. Neural Networks Learn. Syst.6
2020 Auto-Generating Neural Networks with Reinforcement Learning for Multi-Purpose Image Forensics
abstract
Designing a forensic convolutional neural network (CNN) is usually based on some ad-hoc intuition and domain knowledge. Many methods to automate neural network design have been proposed for computer vision tasks, but they may not be directly applied to image forensic problems, which tend to detect weak traces signals left by image operations rather than strong image content signals. In this paper, we propose an approach to learn an optimal forensic CNN structure with reinforcement learning for detecting multiple image tampering operations. A learning agent is introduced to select CNN layers sequentially in a limited state-action space using Q-learning with an $\epsilon$-greedy strategy and experience replay. The experiments demonstrate that the auto-generated network performs better than other classic image forensic methods and shows more robustness against JPEG compression. To our knowledge, this is the first attempt to design forensic deep neural networks automatically with reinforcement learning.
Yujun Wei, Yifang Chen 0002, Xiangui Kang, Z. Jane Wang 0001, Liang Xiao 0003
ICME4
2020 Parkinson's Disease Detection from fMRI-Derived Brainstem Regional Functional Connectivity Networks
Nandinee Fariah Haq, Jiayue Cai, Tianze Yu, Martin J. McKeown, Z. Jane Wang 0001
MICCAI (7)5
2020 Dual Adversarial Network for Unsupervised Ground/Satellite-to-Aerial Scene Adaptation
abstract
Recent domain adaptation work tends to obtain a uniformed representation in an adversarial manner through joint learning of the domain discriminator and feature generator. However, this domain adversarial approach could render sub-optimal performances due to two potential reasons: First, it might fail to consider the task at hand when matching the distributions between the domains. Second, it generally treats the source and target domain data in the same way. In our opinion, the source domain data which serves the feature adaption purpose should be supplementary, whereas the target domain data mainly needs to consider the task-specific classifier. Motivated by this, we propose a dual adversarial network for domain adaptation, where two adversarial learning processes are conducted iteratively, in correspondence with the feature adaptation and the classification task respectively. The efficacy of the proposed method is first demonstrated on Visual Domain Adaptation Challenge (VisDA) 2017 challenge, and then on two newly proposed Ground/Satellite-to-Aerial Scene adaptation tasks. For the proposed tasks, the data for the same scene is collected not only by the traditional camera on the ground, but also by satellite from the out space and unmanned aerial vehicle (UAV) at the high-altitude. Since the semantic gap between the ground/satellite scene and the aerial scene is much larger than that between ground scenes, the newly proposed tasks are more challenging than traditional domain adaptation tasks. The datasets/codes can be found at https://github.com/jianzhelin/DuAN.
Jianzhe Lin, Lichao Mou, Tianze Yu, Xiao Xiang Zhu 0001, Z. Jane Wang 0001
ACM Multimedia5
2020 Xnet: Task-specific attentional domain adaptation for satellite-to-aerial scene
Jianzhe Lin, Kaiwen Yuan, Rabab K. Ward, Z. Jane Wang 0001
Neurocomputing4
2020 DT-LET: Deep transfer learning by exploring where to transfer
Jianzhe Lin, Liang Zhao 0005, Qi Wang 0009, Rabab K. Ward, Z. Jane Wang 0001
Neurocomputing5
2020 A Simple, High-Performance Space-Time Code for MIMO Backscatter Communications
abstract
In this article, we propose a 2 × 2 space-time code (STC) for backscatter communications. The interesting thing is that although the proposed code performs much worse than the classical Alamouti code in conventional channels, it achieves almost the same performance as the Alamouti code in backscatter channels. More importantly, it has a considerably lower tag circuit complexity, which is more preferred in hardware-limited backscatter devices. This article indicates that the sets of “good STCs” for conventional channels and backscatter channels are not heavily overlapped, and there may exist many simpler STCs, which were previously ignored and not reported in conventional channels, have high performance, particularly, in backscatter channels.
Chen He 0002, Huixu Luan, Xiaoya Li 0003, Cunyan Ma, Luyang Han, Z. Jane Wang 0001
IEEE Internet Things J.6
2020 Monostatic MIMO Backscatter Communications
abstract
Backscatter communications have two major antenna configurations: the bistatic configuration, in which the reader employs two different sets of antennas to transmit and receive, and the monostatic configuration, in which the reader employs one set of antennas for both transmitting and receiving. In this paper, we provide a comprehensive study on the MIMO techniques for the M × L monostatic channel. Particularly we study on the joint design of the query and the coding matrices. We show that the maximum achievable diversity order of the monostatic channel is the diversity order achieved by the block-lever unitary query and orthogonal space-time block code (BUTQ-OSTBC) design pair, which is ML 2 , exactly the half of the diversity order of the conventional MIMO channel. Then we show that uniform query, the simplest query approach, cannot achieve the maximum achievable diversity order in the monostatic channel. We generalize BUTQ-OSTBC to the general augmenting approach, and show that unitary matrix is optimal in terms of SER performance among all possible query matrices, when OSTBC is employed. The above results further indicate that in the backscatter channel, additional diversity can be obtained by varying the query signals over time slots within the channel coherent time, which is quite different from the results from the conventional MIMO channels. We verify our results by Monte Carlo simulations.
Chen He 0002, Shangdong Chen, Huixu Luan, Xiaojiang Chen, Z. Jane Wang 0001
IEEE J. Sel. Areas Commun.5
2020 An End-to-End Multi-Task Deep Learning Framework for Skin Lesion Analysis
abstract
Automatic skin lesion analysis of dermoscopy images remains a challenging topic. In this paper, we propose an end-to-end multi-task deep learning framework for automatic skin lesion analysis. The proposed framework can perform skin lesion detection, classification, and segmentation tasks simultaneously. To address the class imbalance issue in the dataset (as often observed in medical image datasets) and meanwhile to improve the segmentation performance, a loss function based on the focal loss and the jaccard distance is proposed. During the framework training, we employ a three-phase joint training strategy to ensure the efficiency of feature learning. The proposed framework outperforms state-of-the-art methods on the benchmarks ISBI 2016 challenge dataset towards melanoma classification and ISIC 2017 challenge dataset towards melanoma segmentation, especially for the segmentation task. The proposed framework should be a promising computer-aided tool for melanoma diagnosis.
Lei Song 0003, Jianzhe Lin, Z. Jane Wang 0001, Haoqian Wang
IEEE J. Biomed. Health Informatics3
2020 Novel Regional Activity Representation With Constrained Canonical Correlation Analysis for Brain Connectivity Network Estimation
abstract
Inferring brain connectivity networks from fMRI data can take place at the Region of Interest (ROI) or voxel level. With most ROI-based approaches, the signals from same-ROI voxels are simply averaged, neglecting any inhomogeneity in each ROI and assuming that the same voxels will interact with different ROIs in a similar manner. In this paper, we propose a novel method of representing ROI activity and estimating brain connectivity that takes into account the regionally-specific nature of brain activity, the spatial location of concentrated activity, and activity in other ROIs. The proposed method is able to integrate intrinsic regional structures into a network modelling framework, which we call local activity constrained canonical correlation analysis (LA-cCCA). We evaluated LA-cCCA on both simulated and real fMRI data. The simulation results demonstrated that LA-cCCA had improved accuracy of the estimated brain connectivity networks compared to the average-signal or Principal Component Analysis (PCA)-based correlation methods and the Canonical Correlation Analysis (CCA) method. We further examined the performance of LA-cCCA on real fMRI data set from the Human Connectome Project. LA-cCCA outperformed the other three approaches in terms of connectivity reproducibility. The proposed method explores the potentials of regional activity representation and is a reliable model for connectivity network estimation. It may serve as a promising tool for studying both the healthy and diseased brain.
Jiayue Cai, Aiping Liu, Martin J. McKeown, Z. Jane Wang 0001
IEEE Trans. Medical Imaging5
2020 Improving Prostate Cancer (PCa) Classification Performance by Using Three-Player Minimax Game to Reduce Data Source Heterogeneity
abstract
PCa is a disease with a wide range of tissue patterns and this adds to its classification difficulty. Moreover, the data source heterogeneity, i.e. inconsistent data collected using different machines, under different conditions, by different operators, from patients of different ethnic groups, etc., further hinders the effectiveness of training a generalized PCa classifier. In this paper, for the first time, a Generative Adversarial Network (GAN)-based three-player minimax game framework is used to tackle data source heterogeneity and to improve PCa classification performance, where a proposed modified U-Net is used as the encoder. Our dataset consists of novel high-frequency ExactVu ultrasound (US) data collected from 693 patients at five data centers. Gleason Scores (GSs) are assigned to the 12 prostatic regions of each patient. Two classification tasks: benign vs. malignant and low- vs. high-grade, are conducted and the classification results of different prostatic regions are compared. For benign vs. malignant classification, the three-player minimax game framework achieves an Area Under the Receiver Operating Characteristic (AUC) of 93.4%, a sensitivity of 95.1% and a specificity of 87.7%, respectively, representing significant improvements of 5.0%, 3.9%, and 6.0% compared to those of using heterogeneous data, which confirms its effectiveness in terms of PCa classification.
Yanan Shao, Z. Jane Wang 0001, Brian Wodlinger, Tim Salcudean
IEEE Trans. Medical Imaging2
2020 Feature-Flow Interpretation of Deep Convolutional Neural Networks
abstract
Despite the great success of deep convolutional neural networks (DCNNs) in computer vision tasks, their black-box aspect remains a critical concern. The interpretability of DCNN models has been attracting increasing attention. In this work, we propose a novel model, Feature-fLOW INterpretation (FLOWIN) model, to interpret a DCNN by its feature-flow. The FLOWIN can express deep-layer features as a sparse representation of shallow-layer features. Based on that, it distills the optimal feature-flow for the prediction of a given instance, starting from deep layers to shallow layers. Therefore, the FLOWIN can provide an instance-specific interpretation, which presents its feature-flow units and their interpretable meanings for its network decision. The FLOWIN can also give the quantitative interpretation in which the contribution of each flow unit in different layers is used to interpret the net decision. From the class-level view, we can further understand networks by studying feature-flows within and between classes. The FLOWIN not only provides the visualization of the feature-flow but also studies feature-flow quantitatively by investigating its density and similarity metrics. In our experiments, the FLOWIN is evaluated on different datasets and networks by quantitative and qualitative ways to show its interpretability.
Xinrui Cui, Dan Wang 0011, Z. Jane Wang 0001
IEEE Trans. Multim.3
2020 CHIP: Channel-Wise Disentangled Interpretation of Deep Convolutional Neural Networks
abstract
With the increasing popularity of deep convolutional neural networks (DCNNs), in addition to achieving high accuracy, it becomes increasingly important to explain how DCNNs make their decisions. In this article, we propose a CHannel-wise disentangled InterPretation (CHIP) model for visual interpretations of DCNN predictions. The proposed model distills the class-discriminative importance of channels in DCNN by utilizing sparse regularization. We first introduce network perturbation to learn the CHIP model. The proposed model is capable to not only distill the global perspective knowledge from networks but also present class-discriminative visual interpretations for the predictions of networks. It is noteworthy that the CHIP model is able to interpret different layers of networks without retraining. By combining the distilled interpretation knowledge at different layers, we further propose the Refined CHIP visual interpretation that is both high-resolution and class-discriminative. Based on qualitative and quantitative experiments on different data sets and networks, the proposed model provides promising visual interpretations for network predictions in an image classification task compared with the existing visual interpretation methods. The proposed model also outperforms the related approaches in the ILSVRC 2015 weakly supervised localization task.
Xinrui Cui, Dan Wang 0011, Z. Jane Wang 0001
IEEE Trans. Neural Networks Learn. Syst.3
2019 Fisher-Bures Adversary Graph Convolutional Networks
Ke Sun 0001, Piotr Koniusz, Z. Jane Wang 0001
UAI3
2019 Depthwise Separable Convolutional Neural Network for Image Forensics
abstract
General-purpose forensics on small image patches appears to be feasible and important, but in fact poses a challenge due to insufficient statistics. Furthermore, there is a need to develop a forensic approach that can automatically learn effective and robust features related to image forensics with high parameter efficiency. In this paper, we propose a depthwise separable convolutional neural network (CNN) for the simultaneous detection of eleven types of image manipulations in image patches. Different from the previous CNNs based on standard convolution, depthwise separable convolution is introduced in the proposed CNN to adaptively extract forensics-related features from image patches with better parameter efficiency. When compared with four state-of-the-art methods, experiments demonstrate that the proposed CNN architecture can achieve better performance, e.g., the improvement in terms of accuracy in the detection of 32 × 32 images is up to 7.33%. It also achieves significantly better overall performance for different databases and better robustness against JPEG compression.
Yifang Chen 0002, Xiangui Kang, Z. Jane Wang 0001
VCIP4
2019 Community structure detection from networks with weighted modularity
Nandinee Fariah Haq, Mehdi Moradi, Z. Jane Wang 0001
Pattern Recognit. Lett.3
2019 Medical Image Fusion via Convolutional Sparsity Based Morphological Component Analysis
abstract
In this letter, a sparse representation (SR) model named convolutional sparsity based morphological component analysis (CS-MCA) is introduced for pixel-level medical image fusion. Unlike the standard SR model, which is based on single image component and overlapping patches, the CS-MCA model can simultaneously achieve multi-component and global SRs of source images, by integrating MCA and convolutional sparse representation (CSR) into a unified optimization framework. For each source image, in the proposed fusion method, the CSRs of its cartoon and texture components are first obtained by the CS-MCA model using pre-learned dictionaries. Then, for each image component, the sparse coefficients of all the source images are merged and the fused component is accordingly reconstructed using the corresponding dictionary. Finally, the fused image is calculated as the superposition of the fused cartoon and texture components. Experimental results demonstrate that the proposed method can outperform some benchmarking and state-of-the-art SR-based fusion methods in terms of both visual perception and objective assessment.
Yu Liu 0023, Xun Chen 0001, Rabab K. Ward, Z. Jane Wang 0001
IEEE Signal Process. Lett.4
2019 Coarse-to-Fine Image DeHashing Using Deep Pyramidal Residual Learning
abstract
Image dehashing refers to the process of inferring images by inverting image hashes. Recently, image dehashing from real-valued image retrieval hashes is shown feasible using deep convolutional neural networks. However, the perceptual quality of dehashed images is challenged when real-valued hashes are quantized to less bits. Besides, the scalability to larger or color image dehashing is limited in the previous dehashing network. To this end, we propose a pyramidal long-range residual-learning network (PyLRR-Net). PyLRR-Net is a pyramidal image reconstruction network to dehash images in a progressive manner. At each image scale, we design and insert a long-range residual block to refine the coarse image reconstruction leveraging deep residual learning. Experiments on both grayscale and color image datasets show that the proposed PyLRR-Net outperforms previous work in terms of image dehashing quality, scalability, and flexibility for large and color image dehashing problems.
Rabab K. Ward, Z. Jane Wang 0001
IEEE Signal Process. Lett.3
2019 Dynamic Graph Theoretical Analysis of Functional Connectivity in Parkinson's Disease: The Importance of Fiedler Value
abstract
Graph theoretical analysis is a powerful tool for quantitatively evaluating brain connectivity networks. Conventionally, brain connectivity is assumed to be temporally stationary, whereas increasing evidence suggests that functional connectivity exhibits temporal variations during dynamic brain activity. Although a number of methods have been developed to estimate time-dependent brain connectivity, there is a paucity of studies examining the utility of brain dynamics for assessing brain disease states. Therefore, this paper aims to assess brain connectivity dynamics in Parkinson's disease (PD) and determine the utility of such dynamic graph measures as potential components to an imaging biomarker. Resting-state functional magnetic resonance imaging data were collected from 29 healthy controls and 69 PD subjects. Time-varying functional connectivity was first estimated using a sliding windowed sparse inverse covariance matrix. Then, a collection of graph measures, including the Fiedler value, were computed and the dynamics of the graph measures were investigated. The results demonstrated that PD subjects had a lower variability in the Fiedler value, modularity, and global efficiency, indicating both abnormal dynamic global integration and local segregation of brain networks in PD. Autoregressive models fitted to the dynamic graph measures suggested that Fiedler value, characteristic path length, global efficiency, and modularity were all less deterministic in PD. With canonical correlation analysis, the altered dynamics of functional connectivity networks, and particularly dynamic Fiedler value, were shown to be related with disease severity and other clinical variables including age. Similarly, Fiedler value was the most important feature for classification. Collectively, our findings demonstrate altered dynamic graph properties, and in particular the Fiedler value, provide an additional dimension upon which to non-invasively and quantitatively assess PD.
Jiayue Cai, Aiping Liu, Taomian Mi, Saurabh Garg 0004, Wade Trappe, Martin J. McKeown, Z. Jane Wang 0001
IEEE J. Biomed. Health Informatics7
2019 Multi-Scale Interpretation Model for Convolutional Neural Networks: Building Trust Based on Hierarchical Interpretation
abstract
With the rapid development of deep learning models, their performances in various tasks have improved; meanwhile, their increasingly intricate architectures make them difficult to interpret. To tackle this challenge, model interpretability is essential and has been investigated in a wide range of applications. For end users, model interpretability can be used to build trust in the deployed machine learning models. For practitioners, interpretability plays a critical role in model explanation, model validation, and model improvement to develop a faithful model. In this paper, we propose a novel Multi-scale Interpretation (MINT) model for convolutional neural networks using both the perturbation-based and the gradient-based interpretation approaches. It learns the class-discriminative interpretable knowledge from the multi-scale perturbation of feature information in different layers of deep networks. The proposed MINT model provides the coarse-scale and the fine-scale interpretations for the attention in the deep layer and specific features in the shallow layer, respectively. Experimental results show that the MINT model presents the class-discriminative interpretation of the network decision and explains the significance of the hierarchical network structure.
Xinrui Cui, Dan Wang 0011, Z. Jane Wang 0001
IEEE Trans. Multim.3
2019 ICFS Clustering With Multiple Representatives for Large Data
abstract
With the prevailing development of Cyber-physical-social systems and Internet of Things, large-scale data have been collected consistently. Mining large data effectively and efficiently becomes increasingly important to promote the development and improve the service quality of these applications. Clustering, a popular data mining technique, aims to identify underlying patterns hidden in the data. Most clustering methods assume the static data, thus they are unfavorable for analyzing large, unbalanced dynamic data. In this paper, to address this concern, we focus on incremental clustering by extending the novel [clustering by fast search (CFS) and find of density peaks] method to incrementally handle large-scale dynamic data. Specifically, we first discuss two challenges, i.e., assignment of new arriving objects and dynamic adjustment of clusters, in incremental CFS (ICFS) clustering. We then propose two ICFS clustering algorithms, ICFS with multiple representatives (ICFSMR) and the enhanced ICFSMR (E_ICFSMR) to tackle the two challenges. In ICFSMR, we explore the convex hull theory to modify the representatives identified for each cluster. E_ICFSMR improves the generality and effectiveness of ICFSMR by exploring one-time cluster adjustment strategy after integration of each data chunk. We evaluate the proposed methods with extensive experiments on four benchmark data sets, as well as the air quality and traffic monitoring time series, with comparisons to CFS and other three state-of-the-art incremental clustering methods. Experimental results demonstrate that the proposed methods outperform the compared methods in terms of both effectiveness and efficiency.
Liang Zhao 0005, Zhikui Chen, Yi Yang 0006, Liang Zou, Z. Jane Wang 0001
IEEE Trans. Neural Networks Learn. Syst.5
2019 Deep Semantic Mapping for Heterogeneous Multimedia Transfer Learning Using Co-Occurrence Data
abstract
Transfer learning, which focuses on finding a favorable representation for instances of different domains based on auxiliary data, can mitigate the divergence between domains through knowledge transfer. Recently, increasing efforts on transfer learning have employeddeepneuralnetworks (DNN) to learn more robust and higher level feature representations to better tackle cross-media disparities. However, only a few articles consider the correction and semantic matching between multi-layer heterogeneous domain networks. In this article, we propose adeep semantic mapping model forheterogeneous multimediatransferlearning (DHTL) using co-occurrence data. More specifically, we integrate the DNN withcanonicalcorrelationanalysis (CCA) to derive a deep correlation subspace as the joint semantic representation for associating data across different domains. In the proposed DHTL, a multi-layer correlation matching network across domains is constructed, in which the CCA is combined to bridge each pair of domain-specific hidden layers. To train the network, a joint objective function is defined and the optimization processes are presented. When the deep semantic representation is achieved, the shared features of the source domain are transferred for task learning in the target domain. Extensive experiments for three multimedia recognition applications demonstrate that the proposed DHTL can effectively find deep semantic representations for heterogeneous domains, and it is superior to the several existing state-of-the-art methods for deep transfer learning.
Liang Zhao 0005, Zhikui Chen, Laurence T. Yang, M. Jamal Deen, Z. Jane Wang 0001
ACM Trans. Multim. Comput. Commun. Appl.5
2018 A Rotation-Invariant Convolutional Neural Network for Image Enhancement Forensics
abstract
Many proposed complex convolutional neural network (CNN) models in image forensics are with a large number of parameters, requiring a huge number of training data and having the risk of being overfitting. Considering the desired rotation invariance in the detection of some specific image manipulations, i.e., image enhancement, we propose employing convolutional filters with an isotropic architecture in the CNN model which can significantly reduce the required number of CNN parameters. With the same weights in symmetric positions, the proposed filter can extract rotation-invariant features for image enhancement forensics. Experimental results show that the proposed rotation-invariant CNN models with much less parameters can achieve much better performance, e.g., yielding more than 13% improvement in terms of detection accuracy in Gamma correction forensics. It also achieves significantly better generalization performances on different databases and better robustness against JPEG compression when compared with the popular BayarNet in [16].
Yifang Chen 0002, Zi Xian Lyu, Xiangui Kang, Z. Jane Wang 0001
ICASSP4
2018 Robust Detection of Epileptic Seizures Using Deep Neural Networks
abstract
Robust detection of epileptic seizures in the presence of inevitable artifacts in Electroencephalogram (EEG) signals is addressed. The EEG dataset considered contains 300 signals recorded from 15 volunteers. Current seizure detection systems achieve good performance when the EEG data is entirely free of noise. However, their performance drastically decays with authentic EEG data polluted by real artifacts. We introduce a robust seizure detection method that can address clean and noisy data. The proposed method uses Long Short-Term Memory (LSTM) neural networks to extract the representative EEG features pertinent to seizures. Experimental results show that the proposed method beats existing methods by achieving 100% classification accuracy. Our method is also shown to be robust against the common EEG artifacts (e.g., muscle activities and eye-blinking) and white noise.
Ramy Hussein, Hamid Palangi, Z. Jane Wang 0001, Rabab K. Ward
ICASSP3
2018 Densely Connected Convolutional Neural Network for Multi-purpose Image Forensics under Anti-forensic Attacks
abstract
Multiple-purpose forensics has been attracting increasing attention worldwide. However, most of the existing methods based on hand-crafted features often require domain knowledge and expensive human labour and their performances can be affected by factors such as image size and JPEG compression. Furthermore, many anti-forensic techniques have been applied in practice, making image authentication more difficult. Therefore, it is of great importance to develop methods that can automatically learn general and robust features for image operation detectors with the capability of countering anti-forensics. In this paper, we propose a new convolutional neural network (CNN) approach for multi-purpose detection of image manipulations under anti-forensic attacks. The dense connectivity pattern, which has better parameter efficiency than the traditional pattern, is explored to strengthen the propagation of general features related to image manipulation detection. When compared with three state-of-the-art methods, experiments demonstrate that the proposed CNN architecture can achieve a better performance (i.e., with a 11% improvement in terms of detection accuracy under anti-forensic attacks). The proposed method can also achieve better robustness against JPEG compression with maximum improvement of 13% on accuracy under low-quality JPEG compression.
Yifang Chen 0002, Xiangui Kang, Z. Jane Wang 0001
IH&MMSec3
2018 Image Inpainting Detection Based on a Modified Formulation of Canonical Correlation Analysis
abstract
Image inpainting is a common image editing technique for filling the missing areas in images. It can be adopted to destroy the integrity of images by forgers with ulterior motives. Compared with other types of inpainting, sparsity-based inpainting assumes more general prior knowledge and is more widely used in practical applications. Although several methods for detecting exemplar-based and diffusion-based inpainting have been proposed, there is a shortage of effective scheme for detecting sparsity-based inpainting. In this paper, we proposed a novel algorithm for sparsity-based image inpainting detection. This type of inpainting has a strong effect on the coefficients of Canonical Correlation Analysis (CCA). Based on this observation, a modified objective function of CCA and a corresponding optimization algorithm are further developed to enhance the difference of inter-class in our feature set. The experiments implemented on two publicly available datasets demonstrated our method's superiority over other competitors. Particularly, unlike previous inpainting detection methods, the proposed framework has better performance in the case of JPEG compression.
Yuting Su 0001, Z. Jane Wang 0001
MMSP4
2018 Deep Transfer Learning for Hyperspectral Image Classification
abstract
Hyperspectral image (HSI) includes a vast quantities of samples, large number of bands, as well as randomly occurring redundancy. Classifying such complex data is challenging, and the classification performance generally is affected significantly by the amount of labeled training samples. Collecting such labeled training samples is labor and time consuming, motivating the idea of borrowing and reusing labeled samples from other preexisting related images. Therefore transfer learning, which can mitigate the semantic gap between existing and new HSI, has recently drawn increasing research attention. However, existing transfer learning methods for HSI which concentrated on how to overcome the divergence among images, may neglect the high level latent features during the transfer learning process. In this paper, we present two novel ideas based on this observation. We propose constructing and connecting higher level features for the source and target HSI data, to further overcome the cross-domain disparity. Different from existing methods, no priori knowledge on the target domain is needed for the proposed classification framework, and the proposed framework works for both homogeneous and heterogenous HSI data. Experimental results on real world hyperspectral images indicate the significance of the proposed method in HSI classification.
Jianzhe Lin, Rabab K. Ward, Z. Jane Wang 0001
MMSP3
2018 Robust detection of epileptic seizures based on L1-penalized robust regression of EEG signals
Ramy Hussein, Mohamed Elgendi, Z. Jane Wang 0001, Rabab K. Ward
Expert Syst. Appl.3
2018 Mid-level deep Food Part mining for food image recognition
abstract
There has been a growing interest in food image recognition for a wide range of applications. Among existing methods, mid‐level image part‐based approaches show promising performances due to their suitability for modelling deformable food parts (FPs). However, the achievable accuracy is limited by the FP representations based on low‐level features. Benefiting from the capacity to learn powerful features with labelled data, deep learning approaches achieved state‐of‐the‐art performances in several food image recognition problems. Both mid‐level‐based approaches and deep convolutional neural networks (DCNNs) approaches clearly have their respective advantages, but perhaps most importantly these two approaches can be considered complementary. As such, the authors propose a novel framework to better utilise DCNN features for food images by jointly exploring the advantages of both the mid‐level‐based approaches and the DCNN approaches. Furthermore, they tackle the challenge of training a DCNN model with the unlabelled mid‐level parts data. They accomplish this by designing a clustering‐based FP label mining scheme to generate part‐level labels from unlabelled data. They test on three benchmark food image datasets, and the numerical results demonstrate that the proposed approach achieves competitive performance when compared with existing food image recognition approaches.
Jiannan Zheng, Liang Zou, Z. Jane Wang 0001
IET Comput. Vis.3
2018 Incomplete multi-view clustering via deep semantic mapping
Liang Zhao 0005, Zhikui Chen, Yi Yang 0006, Z. Jane Wang 0001, Victor C. M. Leung
Neurocomputing4
2018 RevHashNet: Perceptually de-hashing real-valued image hashes for similarity retrieval
abstract
Image hashing has attracted increasing popularity in recent years. Some off-the-shelf image hashing methods are able to generate more compact and robust hashes for fast indexing and content-based similarity retrieval. However, the ability to infer original image contents from their real-valued image hashes has seldom been examined. Inherited from cryptographic hashing for image privacy protection, general image hashing is supposed to be a non-revertible function. Should there be a way to revert (or perceptually reconstruct) images from the corresponding real-valued image hashes? This paper explores the feasibility of perceptually image hashing reversion, and fill this gap by proposing a deep learning based framework, entitled RevHashNet. Given real-valued image hashes from certain image hashing methods, the proposed RevHashNet can automatically reconstruct perceptually similar images with respect to the original ones with high visual quality. Experiments and simulations on real image datasets support the de-hashing effectiveness of the proposed RevHashNet.
Hamid Palangi, Z. Jane Wang 0001, Haoqian Wang
Signal Process. Image Commun.3
2018 Unsupervised Multiview Nonnegative Correlated Feature Learning for Data Clustering
abstract
Multiview data, which provide complementary information for consensus grouping, are very common in real-world applications. However, synthesizing multiple heterogeneous features to learn a comprehensive description of the data samples is challenging. To tackle this problem, many methods explore the correlations among various features across different views by the assumption that all views share the common semantic information. Following this line, in this letter, we propose a new unsupervised multiview nonnegative correlated feature learning (UMCFL) method for data clustering. Different from the existing methods that only focus on projecting features from different views to a shared semantic subspace, our method learns view-specific features and captures inter-view feature correlations in the latent common subspace simultaneously. By separating the view-specific features from the shared feature representation, the effect of the individual information of each view can be removed. Thus, UMCFL can capture flexible feature correlations hidden in multiview data. A new objective function is designed and efficient optimization processes are derived to solve the proposed UMCFL. Extensive experiments on real-world multiview datasets demonstrate that the proposed UMCFL method is superior to the state-of-the-art multiview clustering methods.
Liang Zhao 0005, Zhikui Chen, Z. Jane Wang 0001
IEEE Signal Process. Lett.3
2017 Multimodal Deep Learning Approach for Joint EEG-EMG Data Compression and Classification
abstract
In this paper, we present a joint compression and classification approach of EEG and EMG signals using a deep learning approach. Specifically, we build our system based on the deep autoencoder architecture which is designed not only to extract discriminant features in the multimodal data representation but also to reconstruct the data from the latent representation using encoder-decoder layers. Since autoencoder can be seen as a compression approach, we extend it to handle multimodal data at the encoder layer, reconstructed and retrieved at the decoder layer. We show through experimental results, that exploiting both multimodal data intercorellation and intracorellation 1) Significantly reduces signal distortion particularly for high compression levels 2) Achieves better accuracy in classifying EEG and EMG signals recorded and labeled according to the sentiments of the volunteer.
Ahmed Ben Said, Amr Mohamed 0001, Tarek M. El-Fouly, Khaled A. Harras, Z. Jane Wang 0001
WCNC5
2017 Structure Preserving Transfer Learning for Unsupervised Hyperspectral Image Classification
abstract
Recent advances on remote sensing techniques allow easier access to imaging spectrometer data. Manually labeling and processing of such collected hyperspectral images (HSIs) with a vast quantities of samples and a large number of bands is labor and time consuming. To relieve these manual processes, machine learning based HSI processing methods have attracted increasing research attention. A major assumption in many machine learning problems is that the training and testing data are in the same feature space and follow the same distribution. However, this assumption doesn’t always hold true in many real world problems, especially in certain HSI processing problems with extremely insufficient or even without training samples. In this letter, we present a transfer learning framework to address this unsupervised challenge (i.e., without training samples in the target domain), by making the following three main contributions: 1) to the best of our knowledge, this is the first time for transfer learning framework to be used for the classification of totally unknown target HSI data with no training samples; 2) the characteristics of HSI are learned on dual spaces to exploit its structure knowledge to better label HSI samples; and 3) two specific new scenarios suitable for transfer learning are investigated. Experimental results on several real world HSIs support the superiority of the proposed work.
Jianzhe Lin, Chen He 0002, Z. Jane Wang 0001
IEEE Geosci. Remote. Sens. Lett.3
2017 A Framework of Camera Source Identification Bayesian Game
abstract
Image forensics with the presence of an adversary, such as the interplay between the sensor-based camera source identification (CSI) and the fingerprint-copy attack, has attracted increasing attention recently. In this paper, we propose a framework of CSI game with both complete information and incomplete information. A noise level-based counter anti-forensic method is presented to detect the potential fingerprint-copy attack, and unlike the state-of-the-art countermeasure of the triangle test, it does not need to collect the candidate image set. With the existence of countermeasure, a rational forger needs to balance the tradeoff between synthesizing source information and leaving new detectable evidence of raising the noise level of a forged image. The mixed-strategy other than the sequential-move assumption is adopted to solve the games. The Bayesian game is introduced to address the information asymmetry in practice. The Nash equilibrium of both the complete information game and Bayesian game are theoretically analyzed, and the expected Nash equilibrium payoff of a Bayesian game is obtained. Nash equilibrium receiver operating characteristic curves are adopted to evaluate the detection performance. Simulation results show that the information asymmetry can remarkably affect the final detection performance. To our knowledge, this paper is the first attempt in analyzing a Bayesian forensic game with practical information asymmetry.
Hui Zeng 0002, Jingxian Liu, Xiangui Kang, Yun Q. Shi 0001, Z. Jane Wang 0001
IEEE Trans. Cybern.6
2017 A Load-Balancing Divide-and-Conquer SVM Solver
abstract
Scaling up kernel support vector machine (SVM) training has been an important topic in recent years. Despite its theoretical elegance, training kernel SVM is impractical when facing millions of data. The divide-and-conquer (DC) strategy is a natural framework of handling gigantic problems, and the divide-and-conquer solver for kernel SVM (DC-SVM) is able to train kernel SVM with millions of data with limited time cost. However, there are some drawbacks of the DC-SVM approach. First, it used an unsupervised clustering method to partition the whole problem, which is prone to construct singular subsets, and, second, it is hard to balance the computation load between sub-problems. To address these issues, this article proposed a load-balancing partition method for kernel SVM. First, it clusters sample from one class and then assigns data samples to the cluster centers by a distance measure and construct sub-problems; in this way, it is able to control the computation load and avoid singular problems. Experimental results show that the proposed method has better load-balancing performance than DC-SVM, which implies that it is suitable for distributed and embedding systems.
Z. Jane Wang 0001, Xiangyang Ji
ACM Trans. Embed. Comput. Syst.2
2017 Illumination Variation-Resistant Video-Based Heart Rate Measurement Using Joint Blind Source Separation and Ensemble Empirical Mode Decomposition
abstract
Recent studies have demonstrated that heart rate (HR) could be estimated using video data [e.g., exploring human facial regions of interest (ROIs)] under well-controlled conditions. However, in practice, the pulse signals may be contaminated by motions and illumination variations. In this paper, tackling the illumination variation challenge, we propose an illumination-robust framework using joint blind source separation (JBSS) and ensemble empirical mode decomposition (EEMD) to effectively evaluate HR from webcam videos. The framework takes the hypotheses that both facial ROI and background ROI have similar illumination variations. The background ROI is then considered as a noise reference sensor to denoise the facial signals by using the JBSS technique to extract the underlying illumination variation sources. Further, the reconstructed illumination-resisted green channel of the facial ROI is detrended and decomposed into a number of intrinsic mode functions using EEMD to estimate the HR. Experimental results demonstrated that the proposed framework could estimate HR more accurately than the state-of-the-art methods. The Bland-Altman plots showed that it led to better agreement with HR ground truth with the mean bias 1.15 beats/min (bpm), with 95% limits from -15.43 to 17.73 bpm, and the correlation coefficient 0.53. This study provides a promising solution for realistic noncontact and robust HR measurement applications.
Juan Cheng 0004, Xun Chen 0001, Lingxi Xu, Z. Jane Wang 0001
IEEE J. Biomed. Health Informatics4
2017 Automated Detection and Segmentation of Vascular Structures of Skin Lesions Seen in Dermoscopy, With an Application to Basal Cell Carcinoma Classification
abstract
Blood vessels are important biomarkers in skin lesions both diagnostically and clinically. Detection and quantification of cutaneous blood vessels provide critical information toward lesion diagnosis and assessment. In this paper, a novel framework for detection and segmentation of cutaneous vasculature from dermoscopy images is presented and the further extracted vascular features are explored for skin cancer classification. Given a dermoscopy image, we segment vascular structures of the lesion by first decomposing the image using independent-component analysis into melanin and hemoglobin components. This eliminates the effect of pigmentation on the visibility of blood vessels. Using k-means clustering, the hemoglobin component is then clustered into normal, pigmented, and erythema regions. Shape filters are then applied to the erythema cluster at different scales. A vessel mask is generated as a result of global thresholding. The segmentation sensitivity and specificity of 90% and 86% were achieved on a set of 500 000 manually segmented pixels provided by an expert. To further demonstrate the superiority of the proposed method, based on the segmentation results, we defined and extracted vascular features toward lesion diagnosis in basal cell carcinoma (BCC). Among a dataset of 659 lesions (299 BCC and 360 non-BCC), a set of 12 vascular features are extracted from the final vessel images of the lesions and fed into a random forest classifier. When compared with a few other state-of-art methods, the proposed method achieves the best performance of 96.5% in terms of area under the curve (AUC) in differentiating BCC from benign lesions using only the extracted vascular features.
Pegah Kharazmi, Mohammed I. AlJasser, Harvey Lui, Z. Jane Wang 0001, Tim K. Lee
IEEE J. Biomed. Health Informatics4
2016 Energy Efficient EEG Monitoring System for Wireless Epileptic Seizure Detection
abstract
Wireless EEG monitoring systems have been successfully used for seizure detection outside clinical settings. The wireless EEG sensor nodes consume a considerable amount of battery energy to acquire, encode and transmit the data to the server side. In this paper, we introduce energy-efficient monitoring systems to increase the sensors' battery lifetime. Specifically, we propose a feature extraction method that is robust to artifacts and can effectively select the most discriminant features relevant to seizures. Second, we show how to use the missing at random (MAR) method to reduce the energy required at the sensor node for data transmission without compromising the seizure detection accuracy at the server side. Finally, we show how the expectation maximization (EM) method is used at the server side to accurately substitute the missing values. The performance of the proposed scheme is compared to those of the state-of-the art methods, and is shown to achieve less power consumption without compromising the seizure detection accuracy.
Ramy Hussein, Rabab K. Ward, Z. Jane Wang 0001, Amr Mohamed 0001
ICMLA3
2016 A Multimodal data fusion approach efficiently predicts disease duration in multiple sclerosis
abstract
Magnetic Resonance Imaging (MRI) biomarkers of multiple sclerosis (MS), particularly fluid-attenuated inversion-recovery (FLAIR) sequences, have long been investigated. However, advanced analytical methods, capable of fusing results from different MRI modalities, could be more informative to have enabled joint biomarkers. Here we estimate disease duration (DD) in MS subjects (n=47) based on fusion of information from Myelin Water Imaging (MWI), Diffusion Tensor Imaging (DTI), and resting state functional MRI (rsfMRI) modalities, by adapting a joint Multimodal Statistical Analysis Framework. Using this data driven, multimodal, latent variable (LV) approach, common and unique information in each dataset was acquired and their relationship with DD is analyzed through the Least Absolute Shrinkage and Selection Operator (LASSO) regression. The common components between the three modalities, but not the unique components of each modality, accurately predicted DD. To further investigate the regions importance for estimating DD, we separated the data into two groups: “early” and “late” DD depending upon their relationship to the median duration of illness (120 months). In early disease, DTI information in the Right Tapetum and the Fornix jointly with MWI information in the Left Superior Cerebellar Penducle, Right Cerebral Penducle, and Cingulum were most informative. In contrast, rsfMRI demonstrated altered connectivity throughout disease duration. Our results demonstrate the power of multimodal imaging markers in MS.
Ava Sheikhi Shoshtari, Sue-Jin Lin, Martin J. McKeown, Z. Jane Wang 0001
IJCNN4
2016 Forensics and counter anti-forensics of video inter-frame forgery
Xiangui Kang, Jingxian Liu, Hongmei Liu 0001, Z. Jane Wang 0001
Multim. Tools Appl.4
2016 Image Fusion With Convolutional Sparse Representation
abstract
As a popular signal modeling technique, sparse representation (SR) has achieved great success in image fusion over the last few years with a number of effective algorithms being proposed. However, due to the patch-based manner applied in sparse coding, most existing SR-based fusion methods suffer from two drawbacks, namely, limited ability in detail preservation and high sensitivity to misregistration, while these two issues are of great concern in image fusion. In this letter, we introduce a recently emerged signal decomposition model known as convolutional sparse representation (CSR) into image fusion to address this problem, which is motivated by the observation that the CSR model can effectively overcome the above two drawbacks. We propose a CSR-based image fusion framework, in which each source image is decomposed into a base layer and a detail layer, for multifocus image fusion and multimodal image fusion. Experimental results demonstrate that the proposed fusion methods clearly outperform the SR-based methods in terms of both objective assessment and visual quality.
Yu Liu 0023, Xun Chen 0001, Rabab K. Ward, Z. Jane Wang 0001
IEEE Signal Process. Lett.4
2016 Underdetermined Joint Blind Source Separation for Two Datasets Based on Tensor Decomposition
abstract
In this letter, we aim to jointly separate the underdetermined mixtures of latent sources from two datasets, where the number of sources exceeds the number of observations in each dataset. Currently available blind source separation (BSS) methods, including joint blind source separation (JBSS) and underdetermined blind source separation (UBSS), cannot address this underdetermined problem effectively. We exploit the second-order statistics of observations and introduce a novel BSS method, termed as underdetermined joint blind source separation (UJBSS). Considering the dependence information between two datasets, the problem of jointly estimating the mixing matrices is tackled via canonical polyadic (CP) decomposition of a specialized tensor in which a set of spatial covariance matrices are stacked. Furthermore, the estimated mixing matrices are used to recover the sources from each dataset separately. Numerical results demonstrate the competitive performance of the proposed method when compared to a commonly used JBSS method, multiset canonical correlation analysis (MCCA), and the single-set UBSS method, UBSS with free active sources (UBSS-FAS).
Liang Zou, Xun Chen 0001, Z. Jane Wang 0001
IEEE Signal Process. Lett.3
2016 An Observer/Predictor-Based Model of the User for Attaining Situation Awareness
abstract
Situation awareness (SA) is essential for the safe operation of systems involving human-automation interaction. In this paper, using the theory of functional observers, we model SA for the user interacting with a continuous-time linear time-invariant dynamical system. For systems under human control or shared control, we use the proposed model to determine the required information to be displayed in the user interface for achieving SA. The user interface provides the user with the ability to observe the continuous-time outputs of the system, as well as the ability to enter continuous-time control inputs. In some systems, due to inadequacy of the displayed information, the user may not be able to accomplish the desired task. To determine the required information to be displayed and the necessary states to be tracked, we propose a model of attaining SA for the users by modeling the user as a specific type of estimator (i.e., the extended delayed functional observer/predictor). We then evaluate what information is needed for such an estimator and how the desired functional of the states have to be expanded so that the user can precisely reconstruct and accurately predict the desired task. As an application example, we investigate the problem of controlling the depth of anesthesia during surgery and determine whether there exists a feasible combination of the expanded task and the displayed information that allows the anesthetist to precisely predict the depth of anesthesia of the patient.
Neda Eskandari, Guy Albert Dumont, Z. Jane Wang 0001
IEEE Trans. Hum. Mach. Syst.3
2016 A CNN Regression Approach for Real-Time 2D/3D Registration
abstract
In this paper, we present a Convolutional Neural Network (CNN) regression approach to address the two major limitations of existing intensity-based 2-D/3-D registration technology: 1) slow computation and 2) small capture range. Different from optimization-based methods, which iteratively optimize the transformation parameters over a scalar-valued metric function representing the quality of the registration, the proposed method exploits the information embedded in the appearances of the digitally reconstructed radiograph and X-ray images, and employs CNN regressors to directly estimate the transformation parameters. An automatic feature extraction step is introduced to calculate 3-D pose-indexed features that are sensitive to the variables to be regressed while robust to other factors. The CNN regressors are then trained for local zones and applied in a hierarchical manner to break down the complex regression task into multiple simpler sub-tasks that can be learned separately. Weight sharing is furthermore employed in the CNN regression model to reduce the memory footprint. The proposed approach has been quantitatively evaluated on 3 potential clinical applications, demonstrating its significant advantage in providing highly accurate real-time 2-D/3-D registration with a significantly enlarged capture range when compared to intensity-based methods.
Shun Miao, Z. Jane Wang 0001, Rui Liao
IEEE Trans. Medical Imaging2
2016 Block-Level Unitary Query: Enabling Orthogonal-Like Space-Time Code With Query Diversity for MIMO Backscatter RFID
abstract
Future backscatter RFID systems are required to be more reliable and data intensive for emerging fields such as Internet of Things (IoT). Motivated by this, orthogonal space-time block code (OSTBC), which is very successful in mobile communications for its low complexity and high performance, has already been investigated for MIMO backscatter RFID. Also a recently proposed scheme called unitary query was shown to be able to considerably improve the reliability of MIMO backscatter RFID by exploiting query diversity. Therefore, incorporating the classical OSTBC (at the tag end) with the unitary query (at the query end) seems promising. However, in this paper, we first show that simple direct employment of OSTBC together with unitary query incurs a linear decoding problem and eventually leads to severe performance degradation. As a redesign of the unitary query idea and the classical OSTBC specifically for MIMO backscatter RFID, we present a BUTQ-mOSTBC design pair idea by proposing the block-level unitary query (BUTQ) at the query end and the corresponding modified OSTBC (mOSTBC) at the tag end. The proposed BUTQ-mOSTBC can resolve the linear decoding problem, keep the simplicity and high performance properties of the classical OSTBC, and achieve the query diversity for the MIMO backscatter RFID.
Chen He 0002, Z. Jane Wang 0001, Chunyan Miao, Victor C. M. Leung
IEEE Trans. Wirel. Commun.2
2015 Median Filtering Forensics Based on Convolutional Neural Networks
abstract
Median filtering detection has recently drawn much attention in image editing and image anti-forensic techniques. Current image median filtering forensics algorithms mainly extract features manually. To deal with the challenge of detecting median filtering from small-size and compressed image blocks, by taking into account of the properties of median filtering, we propose a median filtering detection method based on convolutional neural networks (CNNs), which can automatically learn and obtain features directly from the image. To our best knowledge, this is the first work of applying CNNs in median filtering image forensics. Unlike conventional CNN models, the first layer of our CNN framework is a filter layer that accepts an image as the input and outputs its median filtering residual (MFR). Then, via alternating convolutional layers and pooling layers to learn hierarchical representations, we obtain multiple features for further classification. We test the proposed method on several experiments. The results show that the proposed method achieves significant performance improvements, especially in the cut-and-paste forgery detection.
Xiangui Kang, Z. Jane Wang 0001
IEEE Signal Process. Lett.4
2015 A Sparse Representation-Based Wavelet Domain Speech Steganography Method
abstract
In this paper, we present a novel speech steganography method using discrete wavelet transform and sparse decomposition to address the undetectability concern in speech steganography. The proposed speech steganography method exploits the sparse representation to embed secret messages into higher semantic levels of the cover signal, resulting in increased undetectability. The proposed method also yields improvements on both stego signal quality and embedding capacity, which are the two major requirements of a steganography algorithm. Our experimental results illustrate that the stego signals generated by the proposed method are perceptually indistinguishable from the original cover signals, quantified by both SNR and PESQ quality measures. When compared with two well-known steganography methods, the proposed method is shown to be superior on addressing major requirements of a steganography algorithm, imperceptibility, undetectability, and capacity.
Soodeh Ahani, Shahrokh Ghaemmaghami, Z. Jane Wang 0001
IEEE ACM Trans. Audio Speech Lang. Process.3
2015 Dimensionality Reduction for Hyperspectral Data Based on Class-Aware Tensor Neighborhood Graph and Patch Alignment
abstract
To take full advantage of hyperspectral information, to avoid data redundancy and to address the curse of dimensionality concern, dimensionality reduction (DR) becomes particularly important to analyze hyperspectral data. Exploring the tensor characteristic of hyperspectral data, a DR algorithm based on class-aware tensor neighborhood graph and patch alignment is proposed here. First, hyperspectral data are represented in the tensor form through a window field to keep the spatial information of each pixel. Second, using a tensor distance criterion, a class-aware tensor neighborhood graph containing discriminating information is obtained. In the third step, employing the patch alignment framework extended to the tensor space, we can obtain global optimal spectral-spatial information. Finally, the solution of the tensor subspace is calculated using an iterative method and low-dimensional projection matrixes for hyperspectral data are obtained accordingly. The proposed method effectively explores the spectral and spatial information in hyperspectral data simultaneously. Experimental results on 3 real hyperspectral datasets show that, compared with some popular vector- and tensor-based DR algorithms, the proposed method can yield better performance with less tensor training samples required.
Xuesong Wang 0001, Yuhu Cheng 0001, Z. Jane Wang 0001
IEEE Trans. Neural Networks Learn. Syst.4
2015 Unitary Query for the M×L×N MIMO Backscatter RFID Channel
abstract
A multiple-input multiple-output backscatter radio frequency identification (RFID) system consists of three operational ends: the query end (with$M$reader transmitting antennas), the tag end (with$L$tag antennas), and the receiving end (with$N$reader receiving antennas). Such an$M\times L\times N$setting in RFID can bring spatial diversity and has been studied with the use of space-time code (STC) at the tag end. Current research generally has ignored query signaling as a means to improve performance. Here we propose a novelunitary queryscheme, which creates time diversitywithin the channel coherent timeand can yield significant performance improvements. To overcome the difficulty of evaluating the performance when unitary query is employed at the query end and STC is employed at the tag end, we derive a new measure based on the ranks of certain carefully constructed matrices to show that unitary query has superior performance. Simulations show that unitary query can bring 5–10 dB gain in mid signal-to-noise ratio regimes. In addition, different from the conventional uniform query case, unitary query can also improve the performance of single-antenna tags significantly, enabling single-antenna tags with low complexity and small size to be used for high performance.
Chen He 0002, Z. Jane Wang 0001, Victor C. M. Leung
IEEE Trans. Wirel. Commun.2
2014 Physiological parameter monitoring of drivers based on video data and independent vector analysis
abstract
Although modern cars are equipped with advanced technologies to be faster, more comfortable and safer, one essential piece of the driving system, the driver, is missing in the picture. Among the physiological measures used for wellness purposes, heart rate variability has been shown to be directly associated with mental and physical status, and is easy to measure. In this paper, to maintain the driver's comfort and enhance the driving safety, we propose a non-contact, video-based approach to continuously monitor the driver's heart rate variability under real-world driving circumstances. Previously, several methods were proposed for similar goals under laboratory conditions, where simple face detectors and independent component analysis approach were used, and they may fail in both image understanding and signal processing steps under real-world circumstances in driving. Here we propose using advanced facial landmark and pose estimation, and independent vector analysis to extract heart rate variability. Our preliminary experimental results demonstrated that the proposed approach works better than the previous state-of-arts.
Z. Jane Wang 0001, Zhiqi Shen 0001
ICASSP2
2014 Time varying brain connectivity modeling using FMRI signals
abstract
Inferring brain connectivity networks has been increasingly important for understanding brain functioning. It is suggested that brain is inherently non-stationary and the dynamic patterns of brain networks may provide deeper insights into brain function. However, the majority of current models assume that brain connectivity networks have time invariant structures, neglecting the variability in brain interactions over time. To investigate time varying brain connectivity networks, a stick time varying model is presented in this paper. Simulation results demonstrate that the proposed method could improve the accuracy in estimating time-dependent connectivity patterns. It is also applied to real fMRI data set for studying time-varying resting-state brain connectivity networks.
Aiping Liu, Xun Chen 0001, Z. Jane Wang 0001, Martin J. McKeown
ICASSP3
2014 Correlation-and-bit-aware additive spread spectrum data hiding for Laplacian distributed host image signals
Xiaoqiang Zhang 0003, Z. Jane Wang 0001, Xuesong Wang 0001
Signal Process. Image Commun.2
2014 A Three-Step Multimodal Analysis Framework for Modeling Corticomuscular Activity With Application to Parkinson's Disease
abstract
Corticomuscular coupling analysis based on multiple datasets such as electroencephalography (EEG) and electromyography (EMG) signals provides a useful tool for understanding human motor control systems. A popular conventional method to assess corticomuscular coupling has been the pair-wise magnitude-squared coherence (MSC) between EEG and concomitant EMG recordings. However, there are certain limitations associated with the MSC, including the difficulty in robustly assessing group inference, only dealing with two types of datasets simultaneously and the biologically implausible assumption of pair-wise interactions. To overcome such limitations, in this paper, we propose assessing corticomuscular coupling by combining multiset canonical correlation analysis (M-CCA) and joint independent component analysis (jICA). The proposed method takes advantage of the M-CCA and jICA to ensure that the extracted components are maximally correlated across multiple datasets and meanwhile statistically independent within each dataset. Simulations were performed to illustrate the performance of the proposed method. We also applied the proposed method to concurrent EEG, EMG, and behavior data collected in a Parkinson's disease (PD) study. The results reveal highly correlated temporal patterns among the three types of signals and corresponding spatial activation patterns. In addition to the expected motor areas, the corresponding spatial activation patterns demonstrate enhanced occipital connectivity in the PD subjects, consistent with previous medical findings.
Xun Chen 0001, Z. Jane Wang 0001, Martin J. McKeown
IEEE J. Biomed. Health Informatics2
2013 Metric based Gaussian kernel learning for classification
abstract
Metric learning for KNN has attracted increasing attentions in the field of machine learning (e.g., based on the parametric form of Mahalanobis distance). A good distance metric is also the foundation for other machine learning models, for example, a Gaussian RBF kernel is constructed upon distance metric defined in the feature vector space. However, besides the KNN classifier, there is little research work on learning a good distancemetric for distance-basedmodels. In this paper, we propose a novel algorithmto learn aMahalanobis-distance type metric for Gaussian RBF kernels. We conduct experiments on 5 data sets from the UCI Machine Learning Repository database and two face recognition data sets. The classification results show that the proposed algorithm can outperform other state-of-arts on most of the data sets and achieve comparable results on the rest of data sets.
Z. Jane Wang 0001
ICASSP2
2013 A novel quantization-based watermarking approach invariant to gain attack
abstract
In this paper, a novel quantization-based information hiding approach which is invariant to gain attack is presented. For the data embedding, the host vector signal is first divided into two separate vectors. Then the ratio of the magnitude of the vectors is quantized according to the watermark data. The decoding scheme is performed blindly using the euclidean distance. The performance of the proposed method is analytically studied and assessed by simulations on artificial signals. The proposed method is applied to various test images as well. The experimental results confirm the superiority of the proposed technique against common attacks in comparison with the recently proposed methods.
Mohsen Zareian, Hamid Reza Tohidypour, Z. Jane Wang 0001
ICASSP3
2013 An Adaptive Descriptor Design for Object Recognition in the Wild
abstract
Digital images nowadays show large appearance variabilities on picture styles, in terms of color tone, contrast, vignetting, and etc. These `picture styles' are directly related to the scene radiance, image pipeline of the camera, and post processing functions (e.g., photography effect filters). Due to the complexity and nonlinearity of these factors, popular gradient-based image descriptors generally are not invariant to different picture styles, which could degrade the performance for object recognition. Given that images shared online or created by individual users are taken with a wide range of devices and may be processed by various post processing functions, to find a robust object recognition system is useful and challenging. In this paper, we investigate the influence of picture styles on object recognition by making a connection between image descriptors and a pixel mapping function g, and accordingly propose an adaptive approach based on a g-incorporated kernel descriptor and multiple kernel learning, without estimating or specifying the image styles used in training and testing. We conduct experiments on the Domain Adaptation data set, the Oxford Flower data set, and several variants of the Flower data set by introducing popular photography effects through post-processing. The results demonstrate that the proposed method consistently yields recognition improvements over standard descriptors in all studied cases.
Z. Jane Wang 0001
ICCV2
2013 Modeling the User as an Observer to Determine Display Information Requirements
abstract
For Linear Time-Invariant (LTI) systems under human control or under shared control, we develop a technique to determine the required information content of the display - that is, the device which provides the user with information about the system. Having the right information on the display can help the user to attain better understanding about the system and apply required inputs to control it. To determine the required information to be displayed, we first model the user as a specific type of observer with certain understanding regarding the derivatives of the input and output signals. We then evaluate what information has to be provided for such an observer to make the desired states of the system observable. We provide an example on modeling the pilot of a Boeing B-747 during the phase of approach to landing. In this example we also determine the required information in the display to perform a desired task.
Neda Eskandari, Z. Jane Wang 0001, Guy Albert Dumont
SMC2
2013 Compressed Binary Image Hashes Based on Semisupervised Spectral Embedding
abstract
Conventional image hashing maps invariant features of each digital image into a unique, compact, robust, and secure signature, which can be used as an index for fast content identification and copyright protection. This paper addresses an important issue of compressing the real-valued image hashes into short binary signatures, which can support fast image identification using Hamming distance metrics. The proposed binary image hashing approach presents a fundamental departure from existing methods: Prior information from virtual image distortions and attacks is explored the first time in image hash generation. More specifically, the proposed scheme takes advantages of the extended hash feature space from virtual distortions and attacks and generates the binary signature for each image based on spectral embedding. Since the objective function to learn the embedding is designed to both preserve local similarity between distorted copies of the same image and to distinguish visually distinct images, the generated binary image hash is more robust compared with the one using conventional quantization-based compression approaches. Further, the proposed method can be generalized to combine different types of image hashes to generate a fixed-length binary signature. Our experimental results demonstrate that the proposed binary image hash by combining different real-valued image hashes is more robust against various distortions and it is computationally efficient for image similarity comparison using Hamming metrics.
Z. Jane Wang 0001
IEEE Trans. Inf. Forensics Secur.2
2013 Cross-Domain Object Recognition Via Input-Output Kernel Analysis
abstract
It is of great importance to investigate the domain adaptation problem of image object recognition, because now image data is available from a variety of source domains. To understand the changes in data distributions across domains, we study both the input and output kernel spaces for cross-domain learning situations, where most labeled training images are from a source domain and testing images are from a different target domain. To address the feature distribution change issue in the reproducing kernel Hilbert space induced by vector-valued functions, we propose a domain adaptive input-output kernel learning (DA-IOKL) algorithm, which simultaneously learns both the input and output kernels with a discriminative vector-valued decision function by reducing the data mismatch and minimizing the structural error. We also extend the proposed method to the cases of having multiple source domains. We examine two cross-domain object recognition benchmark data sets, and the proposed method consistently outperforms the state-of-the-art domain adaptation and multiple kernel learning methods.
Z. Jane Wang 0001
IEEE Trans. Image Process.2
2013 Guest Editorial for Special Section on Multimodal Biomedical Imaging: Algorithms and Applications
abstract
The nine papers in this special section represent the state-of-art in the area of multimodal biomedical imaging algorithms and applications.
Tülay Adali, Z. Jane Wang 0001, Vince D. Calhoun, Tom Eichele, Martin J. McKeown, Dimitri Van De Ville
IEEE Trans. Multim.2
2013 A Joint Multimodal Group Analysis Framework for Modeling Corticomuscular Activity
abstract
Corticomuscular coupling analysis based on multiple data sets such as electroencephalography (EEG) and electromyography (EMG) signals provides a useful tool for understanding human motor control systems. Two probably most popular methods are the pair-wise magnitude-squared coherence (MSC) between EEG and simultaneously-recorded EMG signals, and partial least square (PLS). Unfortunately, MSC and PLS generally deal with only two types of data sets at the same time, while we may need to analyze more than two types of data sets. Moreover, it is not straightforward to extend MSC to the group level for combining results across subjects. Also, PLS can have the information mixing problem since only the variations in one data set are used to predict the other data set. To address these concerns, we propose a joint multimodal analysis framework for corticomuscular coupling analysis. The proposed framework models multiple data spaces simultaneously in a multidirectional fashion. Furthermore, to address the inter-subject variability concern in real-world medical applications, we extend the proposed framework from the individual subject level to the group level to obtain common corticomuscular coupling patterns across subjects. We apply the proposed framework to concurrent EEG, EMG and behavior data collected in a Parkinson's disease (PD) study. The results reveal several highly correlated temporal patterns among the three types of signals and their corresponding spatial activation patterns. In PD subjects, there are enhanced connections between occipital region and other regions, which is consistent with the previous medical finding. The proposed framework is a promising technique for performing multi-subject and multi-modal data analysis.
Xun Chen 0001, Xiang Chen 0004, Rabab K. Ward, Z. Jane Wang 0001
IEEE Trans. Multim.4
2013 An Unsupervised Hierarchical Feature Learning Framework for One-Shot Image Recognition
abstract
One-shot recognition has attracted increasing attention recently, inspired by the fact that human cognitive systems could perform recognition tasks well provided only one or a few labeled training samples, in contrast to the conventional object recognition systems that require a large number of labeled training images. One-shot recognition is a visual classification task, where only one training sample is available for each object category in the target test domain, with the help of prior-knowledge data from the source domain. In this paper, we tackle this challenging one-shot recognition problem under a more exciting setting by using only unlabeled images as prior knowledge, which requires less labeling effort than previous works which adopt fully labeled data and/or a sophisticated attribute table designed by human experts. We propose a novel unsupervised hierarchical feature learning framework to learn a feature pyramid from the prior-knowledge domain. The proposed feature learning method also could be applied across multiple feature spaces. Furthermore, we propose using pyramid-matching kernels to combine multilevel features. Examining the “Animals with Attributes” and Caltech-4 data sets in our one-shot recognition setting, we show that the proposed unsupervised feature learning approach with very limited information could achieve comparable performance to that of supervised ones.
Z. Jane Wang 0001
IEEE Trans. Multim.2
2012 Large covariance matrix estimation: Bridging shrinkage and tapering approaches
abstract
In this paper, we propose a shrinkage-to-tapering oracle (STO) estimator for estimation of large covariance matrix when the number of samples is substantially fewer than the number of variables, by combining the strength from both Steinian-type shrinkage and tapering estimators. Our contributions include: (i) Deriving the Frobenius risk and a lower bound for the spectral risk of an MMSE shrinkage estimator; (ii) Deriving a closed-form expression for the optimal coefficient of the proposed STO estimator. Simulations on auto-regression (e.g. a sparse case) and fraction Brownian motion (e.g. a non-sparse case) covariance structures are used to demonstrate the superiority of the proposed estimator.
Z. Jane Wang 0001, Martin J. McKeown
ICASSP2
2012 A hybrid RLC-Turbo codec scheme in Distributed Video Coding
abstract
Distributed Video Coding (DVC) in transform domains could yield significant coding gains over the pixel domain. While employing different coding modes in DVC, it is still difficult to fully compress the source data without increasing the complexity at the encoder side. In this paper, we first propose a particular Run-Length-Coding (RLC) coding mode to efficiently compress the continuous zero symbols to improve coding efficiency. Then a hybrid codec scheme for WZ frames by combining Turbo and RLC coding is introduced, where the mode decision is determined at the decoder side to maintain the encoder-side complexity simplicity. Simulation results show that the proposed method gains up to 1 dB when compared with the Transform domain Wyner-Ziv Coding (TDWZ).
Chun-Ling Yang, Wang-Hua Mo, Z. Jane Wang 0001
ICASSP3
2012 A tridirectional method for corticomuscular coupling analysis in Parkinson's disease
abstract
Corticomuscular coupling analysis based on multiple datasets such as electroencephalography (EEG) and electromyography (EMG) signals provides a useful tool for understanding the underlying mechanisms of human motor control systems. In this work, we propose a tridirectional statistical modeling and analysis method to identify the coupling relationships between three types of datasets. Different from conventional approaches where only two datasets are considered and the interest is to interpret one dataset by another in a unidirectional fashion, the goal in this paper is to model three data spaces simultaneously in a tridirectional fashion. To address the intersubject variability concern in real-world medical applications, we further propose a group analysis framework based on the proposed method and apply it to concurrent EEG, EMG and behavior signals collected from 8 normal subjects and 9 patients with Parkinson's disease (PD) performing a dynamic motor task. The results demonstrate highly correlated temporal patterns among the three types of signals and meaningful spatial activation patterns. The proposed approach is a promising technique for performing multi-subject and multi-modal data analysis.
Xun Chen 0001, Z. Jane Wang 0001, Martin J. McKeown
MMSP2
2012 Cross-domain object recognition by output kernel learning
abstract
It is of great importance to investigate the domain adaptation problem as vision data is now available from a variety of sources. For adapting a classifier, the first problem is how to choose ‘source’ domain. The key issue here is measuring domain similarity. In this paper, we present one of the first studies on ‘domain similarity’ measure in the context of object recognition. We introduce an output kernel divergence as a similarity measure between different data domains, and propose using it as a criterion for domain selection for better recognition accuracy. We also propose a novel domain adaptation method using a vector-valued function with learned output kernels. Fundamentally different from existing work, we focus on the shift in the output kernel space, instead of handling the distribution shift in the input feature space. In addition, those previous methods could also be applied together with ours to improve the performance further. We demonstrate the ability of the proposed model to select and adapt between different domains, and report the state-of-art results on a benchmark data set.
Z. Jane Wang 0001
MMSP2
2012 Wavelet-based gradient transform and its applications
abstract
A wavelet-based image gradient transform is proposed. The proposed transform, called multi-scale gradient transform (MSGT), obtains the first order derivative of an image in terms of the wavelet detail coefficients. While traditional methods estimate the image gradients at each wavelet scale in terms of the horizontal and vertical wavelet coefficients only, the proposed transform obtains the gradients in terms of the diagonal, as well as the horizontal and vertical wavelet coefficients. The proposed MSGT is designed to be invertible, non-redundant and computationally efficient. We demonstrate the potential applications of the proposed transform in texture feature extraction, multi-scale edge detection, image quality assessment and image watermarking.
Ehsan Nezhadarya, Rabab K. Ward, Z. Jane Wang 0001
MMSP3
2012 Perceptual Image Hashing Based on Shape Contexts and Local Feature Points
abstract
Local feature points have been widely investigated in solving problems in computer vision, such as robust matching and object detection. However, its investigation in the area of image hashing is still limited. In this paper, we propose a novel shape-contexts-based image hashing approach using robust local feature points. The contributions are twofold: 1) The robust SIFT-Harris detector is proposed to select the most stable SIFT keypoints under various content-preserving distortions. 2) Compact and robust image hashes are generated by embedding the detected local features into shape-contexts-based descriptors. Experimental results show that the proposed image hashing is robust to a wide range of distortions and attacks, due to the benefits of robust salient keypoints detection and the shape-contexts-based feature descriptors. When compared with the current state-of-the-art schemes, the proposed scheme yields better identification performances under geometric attacks such as rotation attacks and brightness changes, and provides comparable performances under classical distortions such as additive noise, blurring, and compression. Also, we demonstrate that the proposed approach could be applied for image tampering detection.
Z. Jane Wang 0001
IEEE Trans. Inf. Forensics Secur.2
2012 An Improved Multiplicative Spread Spectrum Embedding Scheme for Data Hiding
abstract
This paper presents an improved multiplicative spread spectrum (IMSS) embedding scheme for data hiding. We first analyze the error probability of the conventional multiplicative spread spectrum (MSS) scheme and derive the corresponding channel capacity and security level. It is noted that the interference effect of the host signal causes the distribution leakage and contributes to the decoding performance degradation. Since the host signal and the decoder structure information are available at the encoder side, the proposed IMSS scheme exploits both the correlation between the host signal and the watermark signal and the decoder structure in embedding the bit information to reduce the host interference effect. We can show that, compared with MSS, the proposed IMSS maintains the simple decoder structure and does not require additional information for decoding. We also analyze the decoding performances of MSS and IMSS in the presence of additional Gaussian noise. Simulation and real image results illustrate the superiority of the proposed IMSS data hiding scheme over the conventional MSS scheme.
Amir Valizadeh, Z. Jane Wang 0001
IEEE Trans. Inf. Forensics Secur.2
2011 Impact of the correlation between forward and backscatter channels on RFID system performance
abstract
The channel of radio frequency identification (RFID) can be modeled as a cascaded channel consisting of a forward link and a backscatter link. The correlation between the forward and backscatter links strongly affects the bit-error-rate (BER) performance of RFID systems. In this paper, we analytically study how the link correlation affects the BER performance of a single-input, multiple-output (SIMO) RFID system. We de rive exact forms and asymptotic forms of the BER for SIMO RFID systems with maximal ratio combining (MRC) at the reader side. The asymptotic forms provides an insight into the relationship between the link correlation and the BER perfor mance. It is shown that the BER is an increasing function of the magnitude of link correlation in the high SNR regime and is proportional to (1 + ρ2)L/(1 - ρ2)L in the extremely high SNR regime for an L branch channel.
Chen He 0002, Z. Jane Wang 0001
ICASSP2
2011 A new data hiding method using angle quantization index modulation in gradient domain
abstract
A robust data hiding scheme that embeds the watermark bits by quantizing the gradient directions of an image is proposed. By embedding the watermark in the angle, the watermark becomes robust to amplitude scaling attacks. To keep the watermark imperceptible and enhance its robustness, it is embedded in the significant gradient vectors of the image. The significant gradient vectors are obtained using discrete wavelet transform (DWT). Thus, the gradient vector at a pixel is first obtained in terms of the DWT coefficients. Then the gradient direction is quantized by modifying the DWT coefficients corresponding to the gradient vectors. Experimental results confirm that the proposed gradient direction watermarking (GDWM) method is robust to various types of attacks, specially amplitude scaling at tacks, and results in watermarked images of high fidelity.
Ehsan Nezhadarya, Z. Jane Wang 0001, Rabab K. Ward
ICASSP2
2011 Shape context based image hashing using local feature points
abstract
Local feature points have been widely utilized in solving many problems in computer vision, such as robust matching, object detection and classification, due to the fact that they can hold the intrinsic geometric structures of the image content. However, its investigation in the area of image hashing is still limited. In this paper, we propose a novel shape context based image hashing approach using local feature points, by taking advantage of the geometric invariance of local feature points such as SIFT and preserving the intrinsic structure of the image content using shape context. Experimental results clearly show that the proposed hashing is robust to various classic and malicious attacks, due to the virtue of robust salient keypoints detection as well as the shape context feature descriptors. When compared with the current state-of-art block-based image hashing schemes, such as NMF and FJLT hashing, which extract robust features using dimension reduction, experimental results show that the proposed hashing scheme yields better identification performances under geometric attacks such as rotation attacks and brightness changes, and provides comparable performances under classic distortions such as additive noise, blurring and compression.
Z. Jane Wang 0001
ICIP2
2011 Improved multiplicative spread spectrum embedding for image data hiding
abstract
In this paper, we propose an improved multiplicative spread spectrum (IMSS) embedding scheme for data hiding by exploiting the multiplicative characteristic to maintain the attractive property of having good quality in IMSS and reducing the interference effect of the host signal in the original multiplicative spread spectrum (MSS) embedding to improve the decoding performance. We first derive the optimal decoder for MSS. We then propose the IMSS approach by exploring the decoder structure of MSS to remove the interference effect of the host signal as much as possible, and derive the optimal decoder for IMSS which remains the simple structure as in MSS and does not require additional information at the receiver side. Simulation results clearly demonstrate the superiority of the proposed IMSS scheme over the MSS one.
Amir Valizadeh, Z. Jane Wang 0001
ICIP2
2011 A Bayesian Lasso via reversible-jump MCMC
Z. Jane Wang 0001, Martin J. McKeown
Signal Process.2
2011 Robust Image Watermarking Based on Multiscale Gradient Direction Quantization
abstract
We propose a robust quantization-based image watermarking scheme, called the gradient direction watermarking (GDWM), based on the uniform quantization of the direction of gradient vectors. In GDWM, the watermark bits are embedded by quantizing the angles of significant gradient vectors at multiple wavelet scales. The proposed scheme has the following advantages: 1) increased invisibility of the embedded watermark because the watermark is embedded in significant gradient vectors, 2) robustness to amplitude scaling attacks because the watermark is embedded in the angles of the gradient vectors, and 3) increased watermarking capacity as the scheme uses multiple-scale embedding. The gradient vector at a pixel is expressed in terms of the discrete wavelet transform (DWT) coefficients. To quantize the gradient direction, the DWT coefficients are modified based on the derived relationship between the changes in the coefficients and the change in the gradient direction. Experimental results show that the proposed GDWM outperforms other watermarking methods and is robust to a wide range of attacks, e.g., Gaussian filtering, amplitude scaling, median filtering, sharpening, JPEG compression, Gaussian noise, salt & pepper noise, and scaling.
Ehsan Nezhadarya, Z. Jane Wang 0001, Rabab K. Ward
IEEE Trans. Inf. Forensics Secur.2
2011 Correlation-and-Bit-Aware Spread Spectrum Embedding for Data Hiding
abstract
This paper proposes a correlation-and-bit-aware concept for data hiding by exploiting the side information at the encoder side, and we present two improved data hiding approaches based on the popular additive spread spectrum embedding idea. We first propose the correlation-aware spread spectrum (CASS) embedding scheme, which is shown to provide better watermark decoding performance than the traditional additive spread spectrum (SS) scheme. Further, we propose the correlation-aware improved spread spectrum (CAISS) embedding scheme by incorporating SS, improved spread spectrum (ISS), and the proposed correlation-and-bit-aware concept. Compared with the traditional additive SS, the proposed CASS and CAISS maintain the simplicity of the decoder. Our analysis shows that, by efficiently incorporating the side information, CASS and CAISS could significantly reduce the host effect in data hiding and improve the watermark decoding performance remarkably. To demonstrate the improved decoding performance and the robustness by employing the correlation-and-bit-aware concept, the theoretical bit-error performances of the proposed data hiding schemes in the absence and presence of additional noise are analyzed. Simulation results show the superiority of the proposed data hiding schemes over traditional SS schemes.
Amir Valizadeh, Z. Jane Wang 0001
IEEE Trans. Inf. Forensics Secur.2
2010 Asymptotic analysis of the Huberized LASSO estimator
abstract
The Huberized LASSO model, a robust version of the popular LASSO, yields robust model selection in sparse linear regression. Though its superior performance was empirically demonstrated for large variance noise, currently no theoretical asymptotic analysis has been derived for the Huberized LASSO estimator. Here we prove that the Huberized LASSO estimator is consistent and asymptotically normal distributed under a proper shrinkage rate. Our derivation shows that, unlike the LASSO estimator, its asymptotic variance is stabilized in the presence of noise with large variance. We also propose the adaptive Huberized LASSO estimator by allowing unequal penalty weights for the regression coefficients, and prove its model selection consistency. Simulations confirm our theoretical results.
Z. Jane Wang 0001, Martin J. McKeown
ICASSP2
2010 A host rejected spread spectrum embedding scheme for data hiding
abstract
In this paper, a new spread spectrum-based embedding modulation approach, which we call host rejected spread spectrum (HRSS), is presented to efficiently remove the noise-source effect of the host signal in decoding the hidden message bits. The proposed HRSS could improve the decoding performance in terms of bit error rate. To further improve the decoding performance of the HRSS under distortions (which can be modeled as additive noise), an improved host rejected spread spectrum (IHRSS) is introduced. Simulation results demonstrate the superior performances of the proposed embedding schemes for data hiding.
Amir Valizadeh, Z. Jane Wang 0001
ICASSP2
2010 FMRI group studies of brain connectivity via a group robust Lasso
abstract
Inferring effective brain connectivity from neuroimaging data such as functional Magnetic Resonance Imaging (fMRI) has been attracting increasing interest due to its critical role in understanding brain functioning. Incorporating sparsity into connectivity modeling to make models more biologically realistic and performing group analysis to deal with inter-subject variability are still challenges associated with fMRI brain connectivity modeling. To address the above two crucial challenges, the attractive computational and theoretical properties of the least absolute shrinkage and selection operator (LASSO) in sparse linear regression provide a suitable starting point. We propose a group robust LASSO (grpRLASSO) model by combining advantages of the popular group-LASSO and our recently developed robust-LASSO. Here group analysis is formulated as a grouped variable selection procedure. Superior performance of the proposed grpRLASSO in terms of group selection and robustness is demonstrated by simulations with large noise variance. The grpRLASSO is also applied to a real fMRI data set for brain connectivity study in Parkinson's disease, resulting in biologically plausible networks.
Z. Jane Wang 0001, Martin J. McKeown
ICIP2
2010 Watermark survival chance (WSC) concept for improving watermark robustness against JPEG compression
abstract
This paper presents the new concept of watermark survival chance (WSC) for improving watermark robustness. WSC provides a robustness measure for an image feature (e.g. a discrete wavelet transform (DWT) coefficient) when used for watermark embedding, and thus can provide the watermark designer with prior knowledge on robust image features. As an illustrative example, we study additive spread spectrum watermarking in the DWT domain and consider JPEG compression as the attack. WSC is obtained for each DWT coefficient/subband for different compression ratios. Based on the WSC table for JPEG compression distortion, we suggest that: Wavelet coefficients can be divided into two main categories: block boundary coefficients and block non-boundary coefficients; block boundary coefficients generally are more robust for watermark embedding than block non-boundary coefficients; larger scale wavelet coefficients are generally more robust than smaller scales; a vertical subband is slightly preferred at small and large scales, while a horizontal subband is preferred at a medium scale.
Ehsan Nezhadarya, Z. Jane Wang 0001, Rabab K. Ward
ICIP2
2010 Correlation-aware data hiding based on spread spectrum embedding
abstract
In this paper, we present a correlation-aware spread spectrum (CASS) embedding scheme for image data hiding. The basic idea is to explore the bit information and the correlation between the host signal and the watermark signature code during embedding. We show that this correlation-aware embedding approach could yield watermark decoding performance improvement at the decoder side. To reduce the interference observed in CASS and thus improving the decoding performance, we further propose the correlation-aware hybrid spread spectrum (CAHSS) data hiding scheme by incorporating the idea of improved spread spectrum (ISS). The simulation results reveal the superior watermark decoding performances of the proposed correlation-aware schemes.
Amir Valizadeh, Z. Jane Wang 0001
ICIP2
2010 Asymptotic analysis of robust LASSOs in the presence of noise with large variance
abstract
In the context of linear regression, the least absolute shrinkage and selection operator (LASSO) is probably the most popular supervised-learning technique proposed to recover sparse signals from high-dimensional measurements. Prior literature has mainly concerned itself with independent, identically distributed noise with moderate variance. In many real applications, however, the measurement errors may have heavy-tailed distributions or suffer from severe outliers, making the LASSO poorly estimate the coefficients due to its sensitivity to large error variance. To address this concern, a robust version of the LASSO is proposed, and the limiting distribution of its estimator is derived. Model selection consistency is established for the proposed robust LASSO under an adaptation procedure of the penalty weight. A parallel asymptotic analysis is derived for the Huberized LASSO, a previously proposed robust LASSO, and it is shown that the Huberized LASSO estimator preserves similar asymptotics even with a Cauchy error distribution. We show that asymptotic variances of the two robust LASSO estimators are stabilized in the presence of large variance noise, compared with the unbounded asymptotic variance of the ordinary LASSO estimator. The asymptotic analysis from the nonstochastic design is extended to the case of random design. Simulations further confirm our theoretical results.
Z. Jane Wang 0001, Martin J. McKeown
IEEE Trans. Inf. Theory2
2009 A Framework of Multiplicative Spread Spectrum Embedding for Data Hiding: Performance, Decoder and Signature Design
abstract
In this paper, we have investigated several aspects of multiplicative spread spectrum (MSS) embedding for Data Hiding. First, we analyze the probability of error of the maximum likelihood (ML) decoder of MSS and show that MSS yields a better decoding performance than that of the additive SS embedding. Further, to reduce the required information at the receiver, an alternative method for MSS is proposed and both the ML and MMSE decoders are derived. To further reduce the probability of error in decoding, we propose an approximately optimum signature code design. Simulation results show the superiority of multiplicative MSS and demonstrate enhanced performances when using the proposed signature code.
Amir Valizadeh, Z. Jane Wang 0001
GLOBECOM2
2009 Sparse multivariate autoregressive (mAR)-based partial directed coherence (PDC) for electroencephalogram (EEG) analysis
abstract
Partial directed coherence (PDC) has recently been proposed for studying brain connectivity in EEG studies. PDC provides a quantitative spectral measure of the causal relations between signals by its central use of a multivariate autoregressive (mAR) model. Yet, in real applications, the successful estimation of PDC depends on the accuracy of mAR parameter estimation, which is often sensitive to the data size and model order. In addition, it is generally believed that connections between EEG nodes (brain regions) may be sparse. To address these concerns, we propose a sparse mAR-based PDC technique where PDC estimates are computed from sparse mAR coefficient matrices derived from penalized regression. The proposed technique is applied to both simulated data and real EEG recordings, and results show enhanced stability and accuracy of the proposed technique compared to the traditional, non-sparse approach. The sparse mAR-based PDC technique is promising for analyzing brain connectivity in EEG analysis.
Joyce Chiang, Z. Jane Wang 0001, Martin J. McKeown
ICASSP2
2009 Hashing the mAR coefficients from EEG data for person authentication
abstract
Electroencephalogram (EEG) recordings of brain waves have been shown to have unique pattern for each individual and thus have potential for biometric applications. In this paper, we propose an EEG feature extraction and hashing approach for person authentication. Multi-variate autoregressive (mAR) coefficients are extracted as features from multiple EEG channels and then hashed by using our recently proposed fast Johnson-Lindenstrauss transform (FJLT)-based hashing algorithm to obtain compact hash vectors. Based on the EEG hash vectors, a naive Bayes probabilistic model is employed for person authentication. Our EEG hashing approach presents a fundamental departure from existing methods in EEG-biometry study. The promising results suggest that hashing may open new research directions and applications in the emerging EEG-based biometry area.
Chen He 0002, Z. Jane Wang 0001
ICASSP3
2009 Reduced-reference image quality assessment based on perceptual image hashing
abstract
Quality monitoring is of great importance for online media broadcasting service. Without access to the original reference image in most practical scenarios, reduced-referenced (RR) image quality assessment is a good tradeoff and generally more reliable than no-reference (NR) metrics. In this paper, we propose employing image hashing features as side information to estimate the image quality. With its monotone sensitivity to the content quality degradation (e.g. due to compression), the proposed RR quality monitoring method based on our FJLT (Fast Johnson-Lindenstrauss transform) hashing provides two advantages: the accurate image quality estimate in term of conventional objective quality measure such as PSNR, and the low data rate required for delivering the partial reference information. Experimental results demonstrate that the proposed hashing-based RR quality measure system can accurately estimate the quality degradation due to JPEG and JPEG2000, the two widely adopted compression techniques in nowadays network transmission services.
Z. Jane Wang 0001
ICIP2
2009 Image quality monitoring using spread spectrum watermarking
abstract
An improved blind image quality assessment scheme that is based on the Watson's just noticeable difference (JND) modulated spread spectrum watermarking, is proposed. For the purpose of quality monitoring, a watermark is embedded into the original image. The image quality is estimated based on the detected watermark at the receiver side. In terms of the peak signal-to-noise ratio (PSNR), the proposed method is shown to be more robust and less perceptible than simple spread spectrum watermarking. This is due to several factors: an optimum detector is used for watermark detection, the watermark is appropriately selected from a set and the detector parameter is adjusted accordingly to closely yield an empirical ideal quality curve. The proposed method was tested by finding the quality estimates of different images compressed with different quality factors. The results indicate that the method can accurately estimate the quality of the received images based on the detected watermark power.
Ehsan Nezhadarya, Z. Jane Wang 0001, Rabab K. Ward
ICIP2
2009 Minimum mean square error detector for multimessage spread spectrum embedding
abstract
In this paper, a new minimum mean square error (MMSE) based detector is presented for multi-message detection in both conventional spread spectrum (SS) watermarking and improved spread spectrum (ISS) watermarking. The proposed detector yields an improved detection performance in terms of bit error rate (BER) for both SS and ISS multiple message embedding. Simulation results demonstrate that the proposed MMSE detector outperforms the conventional detector, the matched filter detector, by orders of magnitude.
Amir Valizadeh, Z. Jane Wang 0001
ICIP2
2009 An Extended Image Hashing Concept: Content-Based Fingerprinting Using FJLT
Z. Jane Wang 0001
EURASIP J. Inf. Secur.2
2009 Controlling the False Discovery Rate of the Association/Causality Structure Learned with the PC Algorithm
Junning Li, Z. Jane Wang 0001
J. Mach. Learn. Res.2
2009 An Activity-Subspace Approach for Estimating the Integrated Input Function and Relative Distribution Volume in PET Parametric Imaging
abstract
Dynamic positron emission tomography (PET) imaging technique enables the measurement of neuroreceptor distributions corresponding to anatomic structures, and thus, allows image-wide quantification of physiological and biochemical parameters. Accurate quantification of the concentration of neuroreceptor has been the objective of many research efforts. Compartment modeling is the most widely used approach for receptor binding studies. However, current compartment-model-based methods often either require intrusive collection of accurate arterial blood measurements as the input function, or assume the existence of a reference region. To obviate the need for the input function or a reference region, in this paper, we propose to estimate the input function. We propose a novel concept of activity subspace, and estimate the input function by the analysis of the intersection of the activity subspaces. Then, the input function and the distribution volume (DV) parameter are refined and estimated iteratively. Thus, the underlying parametric image of the total DV is obtained. The proposed method is compared with a blind estimation method, iterative quadratic maximum-likelihood (IQML) via simulation, and the proposed method outperforms IQML. The proposed method is also evaluated in a brain PET dataset.
Peng Qiu, Z. Jane Wang 0001, K. J. Ray Liu, Zsolt Szabo
IEEE Trans. Inf. Technol. Biomed.2
2008 Mutual information based relevance network analysis: a Parkinson'S disease study
abstract
Monitoring the dynamics of networks in the brain is of central importance in normal and disease states. Current methods of detecting networks in the recorded EEG such as correlation and coherence only explore linear dependencies, which may be unsatisfactory. We propose applying mutual information as an alternative metric for assessing possible nonlinear statistical dependencies between EEG channels. However, EEG data are complicated by the fact that data are inherently non-stationary and also the brain may not work on the task continually. To address these concerns, we propose a novel EEG segmentation method based on the temporal dynamics of the cross-spectra of computed independent components. A real case study in Parkinson's disease and further group analysis employing ANOVA demonstrate different brain connectivity between tasks and between subject groups and also a plausible mechanism for the beneficial effects of medication used in this disease. The proposed method appears to be a promising approach for EEG analysis and warrants further study.
Pamela Wen-Hsin Lee, Z. Jane Wang 0001, Martin J. McKeown
ICASSP2
2008 Controlling the false discovery rate in modeling brain functional connectivity
abstract
Graphical models of brain functional connectivity have matured from confirming a priori hypotheses to an exploratory tool for discovering unknown connectivity. However, exploratory methods must control the error rate of "discovered" connectivity networks. Here we explore an error-rate-control method for graphical models which controls the false-discovery-rate (FDR) of the conditional-dependence relationships that a graphical model encodes. The application of this method to a group analysis of fMRI study on Parkinson's disease shows that it effectively controls the errors introduced by randomness, and yields meaningful and consistent results. The proposed approach appears promising for functional-connectivity modeling and deserves further investigation.
Junning Li, Z. Jane Wang 0001, Martin J. McKeown
ICASSP2
2008 Probabilistic Boolean Network for inferring brain connectivity using FMRI data
abstract
Recent research has suggested disrupted interactions between brain regions may contribute to some of the symptoms of Parkinson disease (PD). It is therefore important to develop models for inferring brain functional connectivity from non-invasive imaging data, such as functional magnetic resonance imaging (fMRI). In this paper, we propose applying probabilistic Boolean network (PBN) for modeling brain connectivity due to its solid stochastic properties, computational simplicity, robustness to uncertainty, and capability to deal with small-size data, typical for fMRI data sets. Applying the proposed PBN framework to real fMRI data recorded from PD subjects, we noticed that the PBN method detected statistically significant brain connectivity between region-of-interest (ROIs) in PD and normal subjects. In addition, the PBN results suggest a mechanism of the effectiveness of L-dopa, the principal treatment for PD.
Z. Jane Wang 0001, Martin J. McKeown
ICASSP2
2008 Fast Johnson-Lindenstrauss Transform for robust and secure image hashing
abstract
Dimension reduction based techniques, such as singular value decomposition (SVD) and non-negative matrix factorization (NMF), have been proved to provide excellent performance for robust and secure image hashing by retaining the essential features of the original image matrix while preventing intentional attacks. In this paper, we introduce a recently proposed low-distortion, dimension reduction technique, referred as Fast Johnson-Lindenstrauss Transform (FJLT), and propose the use of FJLT for image hashing. FJLT shares the low-distortion characteristics of a random projection but requires a much lower complexity. These two desirable properties make it suitable for image hashing. Our experiment results show that the proposed FJLT-based hash yields good robustness under a wide range of attacks. Furthermore, the influence of secret key on the proposed hashing algorithm is evaluated by receiver operating characteristics (ROC) graph, revealing the efficiency of the proposed approach.
Z. Jane Wang 0001
MMSP2
2008 A Windowed Eigenspectrum Method for Multivariate sEMG Classification During Reaching Movements
abstract
In this letter, we propose an eigenspectra-based feature extraction technique for classification of multivariate surface electromyographic (sEMG) recordings. The proposed method exploits the maximum eigenvalue vectors of the time-varying covariance patterns between sEMG channels. Together with a support vector machine (SVM) classifier, the proposed feature extraction technique is shown to be more reliable and robust, and it enhances classification between stroke and normal subjects, compared to the conventional univariate analysis methods that examine each muscle individually. In addition, analysis results show that the spatial whitening operation enhances the discriminability of eigenspectral features. This simple, easily-implemented, biologically-inspired approach is able to succinctly capture the subtle differences in muscle recruitment patterns between healthy and disease states. It appears to be a promising means to monitor motor performance in disease subjects.
Joyce Chiang, Z. Jane Wang 0001, Martin J. McKeown
IEEE Signal Process. Lett.2
2007 A Multi-Subject, Dynamic Bayesian Networks (DBNS) Framework for Brain Effective Connectivity
abstract
As dynamic connectivity is shown essential for normal brain function and is disrupted in disease, it is critical to develop models for inferring brain effective connectivity from non-invasive (e.g., fMRI) data. Increasingly, (dynamic) Bayesian network (BNs) have been suggested for this purpose due to their flexibility and suitability. However, ultimately extrapolating BN results from one subject to an entire population first requires methods meaningfully addressing inter-subject, within-group variability. Here we explore two group analysis approaches in fMRI using DBNs: one is to construct a group network based on a common structure assumption across individuals, and the other is to identify significant structure features by examining DBNs individually-trained. By investigating real fMRI data from Parkinsons disease (PD) and normal subjects performing a motor task at three progressive levels of difficulty, we noted that both methods detected statistically significant, biologically plausible connectivity between task-related region-of-interest (ROIs) that differed between the PD and normal subjects. However, the second approach was more sensitive, finding more features that were also consistent with prior neuroscience knowledge. Determining highly reproducible DBN nodes/edges across subjects seems promising for inferring altered functional connectivity within a group.
Junning Li, Z. Jane Wang 0001, Martin J. McKeown
ICASSP (1)2
2007 Local Linear Discriminant Analysis (LLDA) for Inference of Multisubject FMRI Data
abstract
Large intersubject variability is a well-described feature of fMRI studies, making inter-group inference, of critical importance for biological interpretation, difficult. Therefore, traditional approaches involve spatially transforming the data of each subject and heavily spatially smoothing the data. Here we propose an alternate method: after first defining individually-specific regions of interest (ROIs) of each subject, we utilize local linear discriminant analysis (LLDA) to jointly optimize the individually-specific and group linear combinations of ROIs that maximally discriminates between groups characterized by either disease status or task. The proposed method was applied to fMRI data recorded from eight normal subjects performing a motor task, and it was shown to successfully detect activation in multiple cortical and subcortical structures that were not present when the data were traditionally analyzed by warping the data to a common space. We suggest that the proposed method for group fMRI data analysis may be more suitable when examining co-activation in small subcortical regions susceptible to misregistration, or examining older or neurological patient populations.
Martin J. McKeown, Junning Li, Xuemei Huang 0002, Z. Jane Wang 0001
ICASSP (1)4
2007 Relevance Network Modeling for Muscle Association Pattern in Reaching Movements
abstract
Our purpose is to study how different muscles collaborate together to efficiently create a smooth, coordinated reaching movement. In the EMG literature, it has been commonplace to model the relationships between muscles using correlation and frequency-based measures such as coherence. Inspired by the observation that mutual information is a more general and reliable metric in revealing complex relationships between time series, we propose a relevance network framework for modeling temporally-aligned multi-variate sEMG recordings. Such a network can identify functional muscle associations, providing insights into the underlying motor behavior. Here we demonstrate that relevance networks can: 1) detect the effects of handedness in normal subjects, and 2) robustly detect between the healthy and stroke subjects. Specifically, the structural features of muscle associations were sensitive to handedness and disease status yet relatively robust to differences across subjects - a long-standing goal in rehabilitation research. These results warrant further study to more fully determine the extent to which the relevance networks may elucidate the complex muscle interactions in reaching movements.
Z. Jane Wang 0001, Martin J. McKeown
ICASSP (1)1
2007 Dependence network modeling for biomarker identification
abstract
MOTIVATION: Our purpose is to develop a statistical modeling approach for cancer biomarker discovery and provide new insights into early cancer detection. We propose the concept of dependence network, apply it for identifying cancer biomarkers, and study the difference between the protein or gene samples from cancer and non-cancer subjects based on mass-spectrometry (MS) and microarray data. RESULTS: Three MS and two gene microarray datasets are studied. Clear differences are observed in the dependence networks for cancer and non-cancer samples. Protein/gene features are examined three at one time through an exhaustive search. Dependence networks are constructed by binding triples identified by the eigenvalue pattern of the dependence model, and are further compared to identify cancer biomarkers. Such dependence-network-based biomarkers show much greater consistency under 10-fold cross-validation than the classification-performance-based biomarkers. Furthermore, the biological relevance of the dependence-network-based biomarkers using microarray data is discussed. The proposed scheme is shown promising for cancer diagnosis and prediction. AVAILABILITY: See supplements: http://dsplab.eng.umd.edu/~genomics/dependencenetwork/
Peng Qiu, Z. Jane Wang 0001, K. J. Ray Liu, Zhang-Zhi Hu, Cathy H. Wu
Bioinform.2
2007 Delay sensitive scheduling schemes for heterogeneous QoS over wireless networks
abstract
Future wireless networks will support the growing demands of heterogeneous and delay sensitive applications. In this paper, a users' satisfaction factor (USF) is defined to quantify quality of service (QoS) for different types of services such as voice, data, and multimedia, as well as for different delay constraints. This USF not only predicts the final delivered QoS during transmission, but also take advantages of the fact that different packets can be decoded at different time in the receivers. Based on this USF, four types of scheduling schemes considering tradeoffs between system performance and individual fairness are proposed. These schemes explore the time, channel, and multi-user diversity to guarantee quality of service and enhance the network performance. From the simulation results, the proposed scheduling schemes achieve different tradeoffs between individual fairness and high system performance for the heterogeneous and delay sensitive applications, compared with the weighted round-robin and the modified proportional fairness scheduling schemes
Zhu Han 0001, Xin Liu 0002, Z. Jane Wang 0001, K. J. Ray Liu
IEEE Trans. Wirel. Commun.3
2006 A Time-Varying Eigenspectrum/SVM Method for Semg Classification of Reaching Movements in Healthy and Stroke Subjects
abstract
A method for classification of sEMG recordings based on the time-varying covariance patterns between sEMG muscle channels is proposed. The proposed eigenspectral feature vector appears to enhance classification of sEMG patterns with an SVM classifier. The method is shown to be more reliable, robust and enhances classification between stroke and normal subjects, compared to standard analysis methods that examine each muscle individually. This simple, easily-implemented, biologically-inspired approach appears to be a promising means to monitor motor performance in healthy and disease subjects.
Joyce Chiang, Z. Jane Wang 0001, Martin J. McKeown
ICASSP (2)2
2006 Bayesian Network Modeling For Discovering "Directed Synergies" Among Muscles in Reaching Movements
abstract
Modeling the muscle activity patterns in coordinated reaching movements from surface Electromyogram (sEMG) recordings is a key challenge in motor behavior studies. Based on Bayesian Network (BN) modeling of sEMG data, this paper presents a framework for discovering and modeling muscle networks and identifying functional muscle groupings. The learned network is further explored for the purpose of classification. We demonstrate the proposed approach on reaching movements in stroke. We found that the specific muscle triples,and, are selectively recruited during reaching movements and are differentially recruited after stroke. We call these computed muscle triplets "directed synergies" to contrast with synergies that are defined by traditional covariance methods. A BN trained on a single healthy subject completely classified and detected the affected side in all stroke subjects. The proposed approach appears a promising technique for muscle network and synergy analysis in motor control.
Junning Li, Z. Jane Wang 0001, Martin J. McKeown
ICASSP (2)2
2006 An Adaptive Protocol for Cooperative Communications Achieving Asymptotic Minimum Symbol-Error-Rate
abstract
This paper investigates the protocol design issue for cooperation systems in wireless communications. A tight approximate symbol error rate (SER) for such systems is derived and analyzed. Based on such analysis, an optimum power allocation scheme is proposed by optimizing the derived approximate SER subject to fixed transmission rate and total transmit power constraints. Then, a novel adaptive protocol is proposed for cooperative communications based on minimizing the asymptotic SER (i.e. in an averaging sense under high-enough SNR regimes) of such systems. This proposed adaptive protocol is able to achieve the maximum achievable diversity gain available in such systems without sacrificing any transmission rate or the total transmit power, and optimally adapts the number of cooperation partners under the changing environments. Simulation results show that the proposed adaptive protocol provides a lower SER compared with existing protocols. In addition, the proposed adaptive protocol with optimum power allocation can remarkably enhance the SER performance in comparison with the equal power allocation scheme.
Chaiyod Pirak, Z. Jane Wang 0001, K. J. Ray Liu
ICASSP (4)2
2006 Optimum Power Allocation for Maximum-Likelihood Channel Estimation in Space-Time Coded MIMO Systems
abstract
This paper presents an optimum power allocation strategy for the maximum likelihood based channel estimation in the space-time coded multiple-input multiple-output systems employing a data-bearing approach for pilot-embedding. The corresponding channel estimation error, the Chernoff's upper bound on the detection error probability, and a lower bound on channel capacity of such systems are analyzed. Based on such analysis, the relationship between these two bounds are revealed, then a unified optimum power allocation scheme is proposed based on jointly optimizing both bounds subject to an acceptable channel estimation error. Simulation results indicate that the proposed power allocation scheme yields a better performance in terms of the error probability, whereas the equal power allocation scheme can be reasonably used as a suboptimum approach with an acceptable performance degradation. Furthermore, the unequal power allocation with more power constantly allocated to the data part yields a much better performance than the one with more power constantly allocated to the pilot part
Chaiyod Pirak, Z. Jane Wang 0001, K. J. Ray Liu, Somchai Jitapunkul
ICASSP (4)2
2006 An improved scalar quantization-based digital video watermarking scheme for H.264/AVC
abstract
Digital video watermarking has attracted a great deal of research interest in the past few years in applications such as digital fingerprinting and owner identification. H.264/AVC is the latest and most advanced video coding standard, but to this date, there are very few watermarking schemes designed for it. This is mainly due to its complexity and compression efficiency which presents a major challenge for any video watermarking approach. We developed a quantization-based video watermarking scheme, which is designed to work with H.264/AVC. Our scheme offers constant robustness at all compression rates without affecting the overall bite rate and quality of the video stream. Experimental results show that compared to existing methods, our scheme significantly outperforms existing methods under compression, transcoding, filtering, scaling, rotation and collusion attacks
Adarsh Golikeri, Panos Nasiopoulos, Z. Jane Wang 0001
ISCAS3
2006 Adaptive Pilot-Embedded Data-Bearing Approach Channel Estimation in Space-Frequency Coded MIMO-OFDM Systems
abstract
In this paper, we study the effects of model mismatch error inherently in the FFT-based channel estimation approach when considering the multipath delay profiles with non-integer delays. We propose an adaptive FFT-based channel estimator that employs the optimum number of significant taps such that the average total energy of the channels dissipating in each tap is completely captured in order to compensate the model mismatch error. Furthermore, our proposed pilot-embedded data-bearing approach is employed for joint channel estimation and data detection. Simulation results reveal that the adaptive FFT-based channel estimator is superior to the FFT-based and LS channel estimators; however, in nonquasi-static fading channels under high Doppler's shift regimes, the performances of all three channel estimators are quite close resulting from the influence of the channel mismatch error that dominates all factors causing the detection error
Chaiyod Pirak, Z. Jane Wang 0001, K. J. Ray Liu, Somchai Jitapunkul
VTC Spring2
2006 A Data-Bearing Approach for Pilot-Aiding in Space-Time Coded MIMO Systems
abstract
This paper presents a novel technique for pilot-based channel estimation and data detection by exploiting a null-space property and an orthogonality property of the data bearer and pilot matrices. The data and pilot extraction procedures can be done independently by a simple linear transformation exploiting the null-space property. The maximum-likelihood (ML) receiver employed for data detection and the unconstrained-ML estimator employed for channel estimation can be designed separately by using the orthogonality property. In addition, the linear minimum mean-squared error (LMMSE) channel estimator is also proposed to improve the performance of channel estimation. The simulation results show that, among three data bearer and pilot structures including time-multiplexing (TM)-based, ST-block-code (STBC)-based, and code-multiplexing (CM)-based structures, the CM-based structure shows superior performance over the TM-based and the STBC-based structures in term of the probability of detection error, e.g. BER, for nonquasi-static flat Rayleigh fading channels, while the performances of these three structures are quite close for quasi-static flat Rayleigh fading channels.
Chaiyod Pirak, Z. Jane Wang 0001, K. J. Ray Liu, Somchai Jitapunkul
VTC Spring2
2006 LS FFT-based channel estimators using pilot-embedded data-bearing approach in space-frequency coded MIMO-OFDM systems
abstract
Multiple-input multiple-output (MIMO) orthogonal frequency division multiplexing (OFDM) is one prominent communication system for realizing high speed data transmission services. One critical issue for such systems is channel estimation. In this paper, we first develop a pilot-embedded data-bearing (PEDB) approach for joint channel estimation and data detection. Then we propose a least square (LS) FFT-based channel estimator by employing the concept of FFT-based channel estimation to improve the performance of the PEDB-LS channel estimation. Also, the effects of model mismatch error when considering non-integer multipath delay profiles, and its performance are investigated. We further propose an adaptive LS FFT-based channel estimator that employs the optimum number of significant taps. Simulation results reveal that the adaptive LS FFT-based estimator provides superior performance under quasi-static channels or low Doppler's shift regimes
Chaiyod Pirak, Z. Jane Wang 0001, K. J. Ray Liu, Somchai Jitapunkul
WCNC2
2006 Polynomial model approach for resynchronization analysis of cell-cycle gene expression data
abstract
MOTIVATION: Identification of genes expressed in a cell-cycle-specific periodical manner is of great interest to understand cyclic systems which play a critical role in many biological processes. However, identification of cell-cycle regulated genes by raw microarray gene expression data directly is complicated by the factor of synchronization loss, thus remains a challenging problem. Decomposing the expression measurements and extracting synchronized expression will allow to better represent the single-cell behavior and improve the accuracy in identifying periodically expressed genes. RESULTS: In this paper, we propose a resynchronization-based algorithm for identifying cell-cycle-related genes. We introduce a synchronization loss model by modeling the gene expression measurements as a superposition of different cell populations growing at different rates. The underlying expression profile is then reconstructed through resynchronization and is further fitted to the measurements in order to identify periodically expressed genes. Results from both simulations and real microarray data show that the proposed scheme is promising for identifying cyclic genes and revealing underlying gene expression profiles. AVAILABILITY: Contact the authors. SUPPLEMENTARY INFORMATION: Supplementary data are available at: http://dsplab.eng.umd.edu/~genomics/syn/
Peng Qiu, Z. Jane Wang 0001, K. J. Ray Liu
Bioinform.2
2005 Performance analysis for pilot-embedded data-bearing approach in space-time coded MIMO systems
abstract
This paper evaluates the performance of the data-bearing approach for pilot-embedding for joint data detection and channel estimation in space-time (ST) coded MIMO systems. Performance measures, such as the MMSE of channel estimation, Cramer-Rao lower bound (CRLB), and the Chernoff's bound of the estimated-channel bit error rate (BER) for ST codes, are explored to examine the proposed scheme. The power allocation problem for data and pilot parts is also addressed by optimizing the probability-of-error upper-bound (PEUB) mismatched factor subject to certain constraints. Three kinds of data bearer and pilot structures are investigated via simulations, including time-multiplexing (TM)-based, ST-block-code (STBC)-based, and code-multiplexing (CM)-based data bearer and pilot matrices. Among these three structures, the CM-based scheme provides superior detection performance over the TM-based and the STBC-based schemes for nonquasistatic flat Rayleigh fading channels, while the performances of these three structures are quite close for quasi-static flat Rayleigh fading channels.
Chaiyod Pirak, Z. Jane Wang 0001, K. J. Ray Liu, Somchai Jitapunkul
ICASSP (3)2
2005 Ensemble dependence model for classification and prediction of cancer and normal gene expression data
abstract
MOTIVATION: DNA microarray technologies make it possible to simultaneously monitor thousands of genes' expression levels. A topic of great interest is to study the different expression profiles between microarray samples from cancer patients and normal subjects, by classifying them at gene expression levels. Currently, various clustering methods have been proposed in the literature to classify cancer and normal samples based on microarray data, and they are predominantly data-driven approaches. In this paper, we propose an alternative approach, a model-driven approach, which can reveal the relationship between the global gene expression profile and the subject's health status, and thus is promising in predicting the early development of cancer. RESULTS: In this work, we propose an ensemble dependence model, aimed at exploring the group dependence relationship of gene clusters. Under the framework of hypothesis-testing, we employ genes' dependence relationship as a feature to model and classify cancer and normal samples. The proposed classification scheme is applied to several real cancer datasets, including cDNA, Affymetrix microarray and proteomic data. It is noted that the proposed method yields very promising performance. We further investigate the eigenvalue pattern of the proposed method, and we discover different patterns between cancer and normal samples. Moreover, the transition between cancer and normal patterns suggests that the eigenvalue pattern of the proposed models may have potential to predict the early stage of cancer development. In addition, we examine the effects of possible model mismatch on the proposed scheme.
Peng Qiu, Z. Jane Wang 0001, K. J. Ray Liu
Bioinform.2
2005 Anti-collusion forensics of multimedia fingerprinting using orthogonal modulation
abstract
Digital fingerprinting is a method for protecting digital data in which fingerprints that are embedded in multimedia are capable of identifying unauthorized use of digital content. A powerful attack that can be employed to reduce this tracing capability is collusion, where several users combine their copies of the same content to attenuate/remove the original fingerprints. In this paper, we study the collusion resistance of a fingerprinting system employing Gaussian distributed fingerprints and orthogonal modulation. We introduce the maximum detector and the thresholding detector for colluder identification. We then analyze the collusion resistance of a system to the averaging collusion attack for the performance criteria represented by the probability of a false negative and the probability of a false positive. Lower and upper bounds for the maximum number of colluders K(max) are derived. We then show that the detectors are robust to different collusion attacks. We further study different sets of performance criteria, and our results indicate that attacks based on a few dozen independent copies can confound such a fingerprinting system. We also propose a likelihood-based approach to estimate the number of colluders. Finally, we demonstrate the performance for detecting colluders through experiments using real images.
Z. Jane Wang 0001, Min Wu 0001, H. Vicky Zhao, Wade Trappe, K. J. Ray Liu
IEEE Trans. Image Process.1
2005 Forensic analysis of nonlinear collusion attacks for multimedia fingerprinting
abstract
Digital fingerprinting is a technology for tracing the distribution of multimedia content and protecting them from unauthorized redistribution. Unique identification information is embedded into each distributed copy of multimedia signal and serves as a digital fingerprint. Collusion attack is a cost-effective attack against digital fingerprinting, where colluders combine several copies with the same content but different fingerprints to remove or attenuate the original fingerprints. In this paper, we investigate the average collusion attack and several basic nonlinear collusions on independent Gaussian fingerprints, and study their effectiveness and the impact on the perceptual quality. With unbounded Gaussian fingerprints, perceivable distortion may exist in the fingerprinted copies as well as the copies after the collusion attacks. In order to remove this perceptual distortion, we introduce bounded Gaussian-like fingerprints and study their performance under collusion attacks. We also study several commonly used detection statistics and analyze their performance under collusion attacks. We further propose a preprocessing technique of the extracted fingerprints specifically for collusion scenarios to improve the detection performance.
H. Vicky Zhao, Min Wu 0001, Z. Jane Wang 0001, K. J. Ray Liu
IEEE Trans. Image Process.3
2005 A MIMO-OFDM channel estimation approach using time of arrivals
abstract
Channel estimation is critical in designing multiple-input multiple-output (MIMO) orthogonal frequency division multiplexing (OFDM) systems for coherent detection and decoding. We propose a channel estimation scheme based on time of arrivals (TOAs) estimation. The TOAs are first determined using a variant of probabilistic data association (PDA) and employing the minimum description length principle, then the PDA is augmented by group decision feedbacks to refine the TOA estimates. With the channels of typical urban and hilly terrain delay profiles, simulation results, compared with the alternating projection (AP) and the Fourier transform based methods, show that the proposed scheme provides a promising accuracy-complexity trade-off for MIMO OFDM channel estimation.
Z. Jane Wang 0001, Zhu Han 0001, K. J. Ray Liu
IEEE Trans. Wirel. Commun.1
2004 Joint segmentation and classification of time series using class-specific features
abstract
We present an approach for the joint segmentation and classification of a time series. The segmentation is on the basis of a menu of possible statistical models: each of these must be describable in terms of a sufficient statistic, but there is no need for these sufficient statistics to be the same, and these can be as complex (for example, cepstral features or autoregressive coefficients) as fits. All that is needed is the probability density function (PDF) of each sufficient statistic under its own assumed model--presumably this comes from training data, and it is particularly appealing that there is no need at all for a joint statistical characterization of all the statistics. There is similarly no need for an a-priori specification of the number of sections, as the approach uses an appropriate penalization of an over-zealous segmentation. The scheme has two stages. In stage one, rough segmentations are implemented sequentially using a piecewise generalized likelihood ratio (GLR); in the second stage, the results from the first stage (both forward and backward) are refined. The computational burden is remarkably small, approximately linear with the length of the time series, and the method is nicely accurate in terms both of discovered number of segments and of segmentation accuracy. A hybrid of the approach with one based on Gibbs sampling is also presented; this combination is somewhat slower but considerably more accurate.
Z. Jane Wang 0001, Peter Willett 0001
IEEE Trans. Syst. Man Cybern. Part B1
2003 A resource allocation framework with credit system and user autonomy over heterogeneous wireless network
abstract
Future wireless networks will support the growing demands of heterogeneous services. Dynamic resource allocation is essential to guarantee quality of service (QoS) and enhance the network performance. We propose a novel resource allocation framework to cope with the time-varying channel conditions, co-channel interferences, and different QoS requirements in various kinds of services. We define a QoS measurement for delay sensitive applications. We introduce a credit system, where users have their autonomy to decide when and how to use their resources, and users can borrow or lend resources from the system. We also develop a simple feedback mechanism to report the system with the users' QoS satisfaction levels and channel conditions. Then the system will adapt its resource allocation strategy according to the users' feedbacks to favor the users with the bad QoS satisfaction levels or the good channels. We develop adaptive algorithms at both the user and system levels. From simulations, the proposed algorithms efficiently allocate the resources to different types of users. The users' delay constraints are satisfied and the links can survive under a long period of bad channels.
Zhu Han 0001, Z. Jane Wang 0001, K. J. Ray Liu
GLOBECOM2
2003 MIMO-OFDM channel estimation via probabilistic data association based TOAs
abstract
The multiple-input multiple-output (MIMO) orthogonal frequency division multiplexing (OFDM) is one of the most promising techniques to achieve broadband wireless communications. Channel estimation is critical in the design of MIMO OFDM systems, since channel parameters are required for coherent detection and decoding. In this work, we propose a channel estimation scheme based on the time of arrivals (TOAs) estimation. The TOAs are first determined using a variant of probabilistic data association (PDA) and employing the minimum description length principle, then the PDA is augmented by group decision feedback to refine the TOA estimates. The proposed algorithm is compared with the Fourier transform based method and the alternating projection (AP) algorithm. Simulation results show that the proposed scheme provides the performance comparable to that of AP and much better than that of the Fourier based method, meanwhile it requires much lower computational cost than that of AP.
Z. Jane Wang 0001, Zhu Han 0001, K. J. Ray Liu
GLOBECOM1
2003 Resistance of orthogonal Gaussian fingerprints to collusion attacks
abstract
Digital fingerprinting is a means to offer protection to digital data by which fingerprints embedded in the multimedia are capable of identifying unauthorized use of digital content. A powerful attack that can be employed to reduce this tracing capability is collusion. We study the collusion resistance of a fingerprinting system employing Gaussian distributed fingerprints and orthogonal modulation. We propose a likelihood-based approach to estimate the number of colluders, and introduce the thresholding detector for colluder identification. We first analyze the collusion resistance of a system to the average attack by considering the probability of a false negative and the probability of a false positive when identifying colluders. Lower and upper bounds for the maximum number of colluders are derived. We then show that the detectors are robust to different attacks. We further study different sets of performance criteria.
Z. Jane Wang 0001, Min Wu 0001, H. Vicky Zhao, K. J. Ray Liu, Wade Trappe
ICASSP (4)1
2003 The VTP test for transients of equal detectability
abstract
For detection of a permanent and precisely-modeled change in distribution of iid observations, Page's test is optimal. When employed to detect a transient change between known distributions, Page's test is a generalised likelihood ratio test (GLRT). However, the situation of interest here is of transient of unknown scale parameter: a fixed Page procedure tuned to a "short-and-loud" signal uses heavy biasing and low threshold, a combination ill-suited to a "long-but-quiet" signal. We offer an easy alternative to the standard Page: it uses a constant bias and a time-varying threshold. The idea is that the above short signals are detected quickly before post-termination data has a chance to refute them; and that evidence for a long signal is allowed to build, rather than being summarily discarded too early. Results show that the approach works quite well.
Peter Willett 0001, Z. Jane Wang 0001
ICASSP (5)2
2003 Nonlinear collusion attacks on independent fingerprints for multimedia
abstract
Digital fingerprinting is a technology for tracing the distribution of multimedia content and protecting them from unauthorized redistribution. Collusion attack is a cost effective attack against digital fingerprinting where several copies with the same content but different fingerprints are combined to remove the original fingerprints. In this paper, we investigate average and nonlinear collusion attacks of independent Gaussian fingerprints and study both their effectiveness and the perceptual quality. We also propose the bounded Gaussian fingerprints to improve the perceptual quality of the fingerprinted copies. We further discuss the tradeoff between the robustness against collusion attacks and the perceptual quality of a fingerprinting system.
H. Vicky Zhao, Min Wu 0001, Z. Jane Wang 0001, K. J. Ray Liu
ICASSP (5)3
2003 Anti-collusion of group-oriented fingerprinting
abstract
Digital fingerprinting of multimedia data involves embedding information in the content, and offers protection to the digital rights of the content by allowing illegitimate usage of the content to be identified by authorized parties. One potential threat to fingerprints is collusion, whereby a group of adversaries combine their individual copies in an attempt to remove the underlying fingerprints. Former studies indicate that collusion attacks based on a few dozen independent copies can confound a fingerprinting system that employs orthogonal modulation. However, since an adversary is more likely to collude with some users than other users, we propose a group-based fingerprinting scheme where users likely to collude with each other are assigned correlated fingerprints. We evaluate the performance of our group-based fingerprints by studying the collusion resistance of a fingerprinting system employing Gaussian distributed fingerprints. We compare the results to those of fingerprinting systems employing orthogonal modulation.
Z. Jane Wang 0001, Min Wu 0001, Wade Trappe, K. J. Ray Liu
ICME1
2003 Resistance of orthogonal Gaussian fingerprints to collusion attacks
abstract
Digital fingerprinting is a means to offer protection to digital data by which fingerprints embedded in the multimedia are capable of identifying unauthorized use of digital content. A powerful attack that can be employed to reduce this tracing capability is collusion. In this paper, we study the collusion resistance of a fingerprinting system employing Gaussian distributed fingerprints and orthogonal modulation. We propose a likelihood-based approach to estimate the number of colluders, and introduce the thresholding detector for colluder identification. We first analyze the collusion resistance of a system to the average attack by considering the probability of a false negative and the probability of a false positive when identifying colluders. Lower and upper bounds for the maximum number of colluders K/sub max/ are derived. We then show that the detectors are robust to different attacks. We further study different sets of performance criteria.
Z. Jane Wang 0001, Min Wu 0001, H. Vicky Zhao, K. J. Ray Liu, Wade Trappe
ICME1
2003 Performance of detection statistics under collusion attacks on independent multimedia fingerprints
abstract
Digital fingerprinting is a technology for tracing the distribution of multimedia content and protecting them from unauthorized redistribution. Collusion attack is a cost effective attack against digital fingerprinting where several copies with the same content but different fingerprints are combined to remove the original fingerprints. In this paper, we consider average attack and several nonlinear collusion attacks on independent Gaussian based fingerprints, and study the detection performance of several commonly used detection statistics in the literature under collusion attacks. Observing that these detection statistics are not specifically designed for collusion scenarios and do not take into account the characteristics of the newly generated fingerprints under collusion attacks, we propose pre-processing techniques to improve the detection performance of the detection statistics under collusion attacks.
H. Vicky Zhao, Min Wu 0001, Z. Jane Wang 0001, K. J. Ray Liu
ICME3
2003 Nonlinear collusion attacks on independent fingerprints for multimedia
abstract
Digital fingerprinting is a technology for tracing the distribution of multimedia content and protecting them from unauthorized redistribution. Collusion attack is a cost effective attack against digital fingerprinting where several copies with the same content but different fingerprints are combined to remove the original fingerprints. In this paper, we investigate average and nonlinear collusion attacks of independent Gaussian fingerprints and study both their effectiveness and the perceptual quality. We also propose the bounded Gaussian fingerprints to improve the perceptual quality of the fingerprinted copies. We further discuss the tradeoff between the robustness against collusion attacks and the perceptual quality of a fingerprinting system.
H. Vicky Zhao, Min Wu 0001, Z. Jane Wang 0001, K. J. Ray Liu
ICME3
2002 Fast and accurate variance-segmentation of white Gaussian data
abstract
Two new algorithms are presented for the segmentation of a white Gaussian-distributed time series having unknown but piecewise-constant variances, a problem for which only dynamic-programming (DP) approaches have. generally been suitable. The first “Sequential/MDL” includes a rough parsing via the GLR, a penalization of busy segmentations via MDL, and a refinement. The second “Gibbs Sampling” approach uses Monte Carlo ideas. From simulation it appears that both schemes are very accurate in terms of their segmentation; but that the Sequential/MDL approach is orders of magnitude lower in its computational needs both than DP or Gibbs, with Gibbs preferable to DP in this regard. The Gibbs approach can, however, be useful and efficient as a final post-processing step.
Z. Jane Wang 0001, Peter Willett 0001
ICASSP1
2001 Improved power-law detection of transients
abstract
A power-law statistic operating on DFT data has emerged as a basis for a remarkably robust detector of transient signals having unknown structure, location and strength. In this paper we offer a number of improvements to the original power-law detector. Specifically, the power-law detector requires that its data be pre-normalized and spectrally white; a CFAR and self-whitening version is developed and analyzed. Further, it is noted that transient signals tend to be contiguous both in temporal and frequency senses, and consequently new power-law detectors in the frequency and the wavelet domains are given. The resulting detectors offer exceptional performance and are extremely easy to implement. There are no parameters to tune, and they may be considered "plug-in" solutions to the transient detection problem.
Z. Jane Wang 0001, Peter Willett 0001
ICASSP1
2001 Wavelets in the frequency domain for narrowband process detection
abstract
Detecting signals that are long, weak, and narrowband is a well known and important problem in acoustic signal processing. In this paper an ad hoc scheme is developed: its stages include the DFT, a multiresolution decomposition in the frequency domain, and a GLRT. The computational load is light, and the performance is remarkably good. This is so not just in the original narrowband situation, but also, due to an inherent adaptivity to the data, in the detection of signals that are relatively broadband in nature. Generalizations are given to CFAR operation in both prewhitened and unwhitened cases, and to the detection of multi-band signals. As regards the last, it is discovered that there is little loss from over-estimating the number of bands.
Peter Willett 0001, Z. Jane Wang 0001, Roy L. Streit
ICASSP2