VLDB 2026 Research / reviewers in the wild / expert
Jiande Sun 0001
dblp:68/365-1
· DBLP profile ↗
166ranked-venue papers
9as first author
95since 2021 · last 2026
0000-0001-6157-2051ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 88 · 5 first-author · 40 since 2021Artificial intelligence and machine learning · 48 · 3 first-author · 30 since 2021Applied, interdisciplinary, general and emerging computing · 17 · 1 first-author · 14 since 2021Computer networks · 13 · 11 since 2021Systems, architecture and hardware · 6 · 5 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Security and privacy · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Progressive dynamic Taylor unfolding network for multi-modal image fusion
Xuquan Wang, Shengka Shi, Yingjie Kong, Chunan Guan, Jiande Sun 0001, Kai Zhang 0010, Jianfei Cao |
Eng. Appl. Artif. Intell. | 5 |
| 2026 | SCIA-GAN: Robust image watermarking via spatial-channel interaction attention and feature preservation
Lingchen Gu, Jun Wang 0061, Wenbo Wan, Jiande Sun 0001, Sen-Ching S. Cheung |
Expert Syst. Appl. | 6 |
| 2026 | A comprehensive survey of deep learning-based cognitive diagnosis models in education: Methods, applications, and outlook
Jing Li 0046, Yang Xu 0025, Enrique Herrera-Viedma, Hui Yu 0001, Jiande Sun 0001 |
Neurocomputing | 6 |
| 2026 | Integrated Sensing and Covert Communications With RIS Adaptive and Non-Adaptive ModesabstractIn this work, we consider an integrated sensing and covert communication (ISACC) system with a finite blocklengthLaided by a reconfigurable intelligent surface (RIS) withNelements. Specifically, with the aid of RIS, a transmitter Alice is to sense the potential existence of a target, and she is also probabilistically trying to send information to a receiver Bob covertly (i.e., trying to hide her transmissions from a warden Willie). Meanwhile, Willie is to detect whether Alice is sensing the target only or conducting ISACC. We consider two RIS operation modes, i.e., a non-adaptive mode, where RIS employs a common beamforming vector regardless of whether Alice performs sensing-only or ISACC transmissions, and an adaptive mode, where RIS dynamically switches between two beamforming vectors tailored to the transmission type. For each mode, we formulate and solve an optimization problem that jointly determines Alice’s power allocation fractionρfor covert information signals and RIS beamforming vectors to maximize the effective covert communication throughput, while satisfying sensing and covertness constraints. Our examination shows that in the non-adaptive mode, the optimalρdecreases with the blocklengthL, the number of RIS elementsN, or Alice’s transmit powerPa, reflecting a stronger trade-off between covert throughput and sensing reliability. In contrast, in the adaptive mode, the optimal solution isρ∗ = 1, indicating that Alice can fully rely on RIS adaptivity to conceal covert transmissions by switching its beamforming vectors. Numerical results further demonstrate a significant covert communication throughput gain of the adaptive mode over the non-adaptive mode, which increases with bothNandPa, highlighting the importance of RIS adaptivity in the ISACC systems. Jia Zhang 0028, Dengfeng Zhang, Jiande Sun 0001, Min Li 0008, Shihao Yan |
IEEE J. Sel. Areas Commun. | 4 |
| 2026 | Task-driven infrared and visible image fusion via detail and semantic dual injection
Kai Zhang 0010, Ludan Sun, Feng Zhang 0028, Wenbo Wan, Jiande Sun 0001 |
Neural Networks | 6 |
| 2026 | Multi-level graph attention networks for bridging the molecular structure and odor
Hongxin Xie, Jiande Sun 0001, Xiaojun Chang |
Pattern Recognit. | 2 |
| 2026 | A universal pansharpening network via spatial-spectral contrastive learning
Kai Zhang 0010, Yunlong Liu 0005, Feng Zhang 0028, Wenbo Wan, Lingchen Gu, Jiande Sun 0001 |
Pattern Recognit. | 7 |
| 2026 | Video and text semantic center alignment for text-video cross-modal retrieval
Ming Jin 0007, Huaxiang Zhang 0001, Lei Zhu 0002, Jiande Sun 0001, Li Liu 0031 |
Signal Process. Image Commun. | 4 |
| 2026 | Maximizing Computation Efficiency and Fairness of Multi-UAV Assisted MEC System Supported by RISabstractRecently, unmanned aerial vehicles (UAVs) have been widely used in mobile edge computing (MEC) systems to compute part of the computing tasks offloaded from ground equipments (GEs) due to their high mobility and flexibility. In addition, reconfigurable intelligent surfaces (RIS), as an emerging technology, can enhance the wireless propagation environment in wireless networks and improve the computation efficiency of the system. In this paper, we propose a multi-UAV-assisted MEC system supported by RIS, where GEs can partially offload tasks to UAVs for computation. A joint optimization problem is formulated to maximize the weighted computation efficiency and fairness by optimizing the user association state, offloading ratio, computing resource allocation, the trajectory of UAVs and the actual RIS phase shift design. To solve the problem, a Fuzzy C-Means-Multi-Agent Deep Deterministic Policy Gradient Alternating Iterative (FMAI) algorithm is designed. In this algorithm, we firstly design a GE-UAV Association and Variable Initialization Combine Fuzzy C-Means Clustering (GUAIFCM) algorithm to solve GE-UAV association strategy. Then we introduce the multi-agent deep deterministic policy gradient alternating iteration (MADDPGAI) algorithm to solve the computation resource allocation, the trajectory of UAVs, task allocation and RIS phase shift. The simulation results show that the proposed scheme can significantly improve the computation efficiency and fairness of RIS-assisted multi-UAV MEC system compared with the benchmark scheme. Zekun Lu, Linbo Zhai, Yujuan Jia, Meiyu Jin, Jiande Sun 0001, Zhiquan Liu 0001 |
IEEE Trans. Commun. | 6 |
| 2026 | Boosting the No-Reference Image Quality Assessment via Low-Quality Pseudo ReferencesabstractNo-reference image quality assessment (NR-IQA) aims to predict perceptual image quality without access to pristine references, which remains challenging due to diverse and complex distortions. Recent pseudo-reference-based methods attempt to mitigate this challenge but often rely on highfidelity pseudo-reference reconstruction. In contrast, this work shows that improving NR-IQA performance does not depend on reconstruction quality, but on effective representation learning, feature alignment, and deviation modeling between distorted images and pseudo references. To this end, we propose a novel NR-IQA framework that leverages low-quality pseudo references generated by a masked autoencoder with a lightweight decoder. Rather than pursuing detailed reconstruction, the pseudo reference is used to facilitate representation-level deviation modeling in a shared latent space via a cross-attention-based mechanism. Extensive experiments on multiple benchmark datasets demonstrate that the proposed method consistently outperforms state-of-the-art NR-IQA approaches while maintaining modest computational complexity. Our source code will be available at: https://github.com/jianjin008/L-IQA. Lili Meng, Yingnan Wang, Miaohui Wang, Guosheng Lin, Cheng Liang 0001, Jiande Sun 0001, Huaxiang Zhang 0001, Weisi Lin |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2026 | Robust Federated Learning Under Heterogeneity via Rank-One and Column-Sparsity ModelabstractByzantine-robust federated learning aims to maintain resilient performance in the presence of malicious attacks that can impede the convergence of learning algorithms. Although numerous robust aggregators have been developed to merge the collected gradient information in the server, they either require data homogeneity and are suboptimal for heterogeneous data, or their breakdown points—the smallest proportion of outliers that can make the aggregators fail—are not theoretically analyzed or less than 0.5. In contrast to existing aggregators, this paper formulates the aggregation process as a low-rank plus sparse decomposition model, where the low-rank component, with a rank of one, facilitates accurate gradient computation, while the sparse component, penalized by the ℓ2,0-norm, mitigates the impact of outliers. We prove that the devised rule achieves the maximum breakdown point of 0.5. Besides, we apply our aggregation rule to Byzantine-robust federated learning and employ the Polyak’s momentum to reduce gradient variance among honest workers. It is analyzed that our aggregator achieves order-optimal Byzantine-resilient federated learning for heterogeneous data. Experimental results using MNIST, Fashion-MNIST and CIFAR-10 demonstrate that the developed approach yields higher classification accuracy than the competing aggregators under different attack types and heterogeneity levels. Zhi-Yong Wang, Hao Nan Sheng, Hing-Cheung So, Jiande Sun 0001, Linqi Song, Weitao Xu |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Value-Based Proactive Caching for Sensing Data in Vehicular Networks: An Operator's PerspectiveabstractAccess to sensing data (SD) is crucial for vehicular networks to ensure safe and efficient transportation services. Given the vast volume of data involved, proactive caching required SD is a pivotal strategy for alleviating network congestion and improving data accessibility. Despite merits, existing studies predominantly address SD caching within a single slot. Therefore, these approaches lack scalability for scenarios involving multi-slots and are not well-suited for network operators who manage resources within a long-term cost budget. Moreover, the oversight of service capacity at caching nodes may result in substantial queuing delays for SD reception. To tackle these limitations, we jointly consider the problem of anchoring SD caching and allocating from an operator's perspective. A value model incorporating both temporal and spacial characteristics is given to estimate the significance of various caching decisions. Subsequently, a stochastic programming model is proposed to optimize the long-term system performance, which is converted into a series of online optimization problem by leveraging the Lyapunov method and linearized via introducing auxiliary variables. To expedite the solution, we provide a binary quantum particle swarm optimization based algorithm with quadratic time complexity. Numerical investigations demonstrate the superiority of proposed algorithms compared with other schemes in terms of energy consumption, response latency, and cache-hit ratio. Yantong Wang, Jiande Sun 0001 |
ICC | 4 |
| 2025 | Multi-Hierarchical Fine-Grained Feature Mapping Driven by Feature Contributions for Molecular Odor PredictionabstractMolecular odor prediction involves using a molecule's structure to estimate its odor. While accurate prediction remains challenging, AI models can suggest potential odors. Existing methods, however, often rely on basic descriptors or handcrafted fingerprints, which lack expressive power and hinder effective learning. Furthermore, these methods suffer from severe class imbalance, limiting the training effectiveness of AI models. To address these challenges, we propose a Feature Contribution-driven Hierarchical Multi-Feature Mapping Network (HMFNet). Specifically, we introduce a fine-grained, Local Multi-Hierarchy Feature Extraction module (LMFE) that performs deep feature extraction at the atomic level, capturing detailed features crucial for odor prediction. To enhance the extraction of discriminative atomic features, we integrate a Harmonic Modulated Feature Mapping (HMFM). This module dynamically learns feature importance and frequency modulation, improving the model's capability to capture relevant patterns. Additionally, a Global Multi-Hierarchy Feature Extraction module (GMFE) is designed to learn global features from the molecular graph topology, enabling the model to fully leverage global information and enhance its discriminative power for odor prediction. To further mitigate the issue of class imbalance, we propose a Chemically-Informed Loss (CIL). Experimental results demonstrate that our approach significantly improves performance across various deep learning models, highlighting its potential to advance molecular structure representation and accelerate the development of AI-driven technologies. Hongxin Xie, Jiande Sun 0001, Fanfu Xue, Zifei Han |
IJCAI | 2 |
| 2025 | Multimodal Inverse Attention Network with Intrinsic Discriminant Feature Exploitation for Fake News DetectionabstractMultimodal fake news detection has garnered significant attention due to its profound implications for social security. While existing approaches have contributed to understanding cross-modal consistency, they often fail to leverage modal-specific representations and explicit discrepant features. To address these limitations, we propose a Multimodal Inverse Attention Network (MIAN), a novel framework that explores intrinsic discriminative features based on news content to advance fake news detection. Specifically, MIAN introduces a hierarchical learning module that captures diverse intra-modal relationships through local-to-global and local-to-local interactions, thereby generating enhanced unimodal representations to improve the identification of fake news at the intra-modal level. Additionally, a cross-modal interaction module employs a co-attention mechanism to establish and model dependencies between the refined unimodal representations, facilitating seamless semantic integration across modalities. To explicitly extract inconsistency features, we propose an inverse attention mechanism that effectively highlights the conflicting patterns and semantic deviations introduced by fake news in both intra- and inter-modality. Extensive experiments on benchmark datasets demonstrate that MIAN significantly outperforms state-of-the-art methods, underscoring its pivotal contribution to advancing social security through enhanced multimodal fake news detection. En Yu, Jiande Sun 0001 |
IJCAI | 4 |
| 2025 | PerfSeer: An Efficient and Accurate Deep Learning Models Performance PredictorabstractPredicting the performance of deep learning (DL) models, such as execution time and resource utilization, is crucial for Neural Architecture Search (NAS), DL cluster schedulers, and other technologies that advance deep learning. The representation of a model is the foundation for its performance prediction. However, existing methods cannot comprehensively represent diverse model configurations, resulting in unsatisfactory accuracy. To address this, we represent a model as a graph that includes the topology, along with node, edge, and global features, all of which are crucial for effectively capturing the performance of the model. Based on this representation, we propose PerfSeer, a novel predictor that uses a Graph Neural Network (GNN)-based performance prediction model, SeerNet. SeerNet fully leverages the topology and various features, while incorporating optimizations such as Synergistic Max-Mean aggregation (SynMM) and Global-Node Perspective Boost (GNPB) to more effectively capture the critical performance information, enabling it to predict the performance of models accurately. Furthermore, SeerNet can be extended to SeerNet-Multi by using Project Conflicting Gradients (PCGrad), enabling efficient simultaneous prediction of multiple performance metrics without significantly affecting accuracy. We constructed a dataset containing performance metrics for 53k+ model configurations, including execution time, memory usage, and Streaming Multiprocessor (SM) utilization during both training and inference. The evaluation results show that PerfSeer outperforms nn-Meter, Brp-NAS, and DIPPM. Xinlong Zhao, Jiande Sun 0001, Jia Zhang 0028 |
IJCAI | 2 |
| 2025 | HGS_OFAT: High-fidelity Gaussian SLAM based on Optical Flow Assisted Trackingabstract3D Gaussian splatting representations have demonstrated outstanding potential for dense scene reconstruction and simultaneous localization and mapping (SLAM). However, existing multi-view filtering techniques frequently overlook critical edge information, resulting in blurred reconstruction boundaries. To address these challenges, we propose High-fidelity Gaussian SLAM based on Optical Flow Assisted Tracking (HGS_OFAT). Our approach analyzes the relationships between edge cues and pixels selected by multi-view pixel selection strategy, and employs a prudent initialization strategy to optimize Gaussian parameters—thereby reducing redundancy and improving reconstruction fidelity. Furthermore, HGS_OFAT integrates an opticalflow tracker before Gaussian-based reconstruction to obtain accurate camera poses, significantly enhancing both localization and reconstruction accuracy. Experiments on synthetic and realworld datasets demonstrate that HGS_OFAT outperforms existing methods in tracking robustness and reconstruction quality. Zhenyong Li, Shanxin Zhang, Chuanfen Feng, Jiande Sun 0001 |
MMSP | 6 |
| 2025 | MGFT: Multi-Geometric Fusion Transformer for Robust Point Cloud RegistrationabstractPoint cloud registration remains a fundamental yet challenging problem in computer vision, particularly under low-overlap conditions. We propose a Multi-Geometric Fusion Transformer (MGFT) that leverages rich geometric cues to enhance registration accuracy. MGFT integrates geometric features extracted by a Graph Convolutional Network (GCN) and Point Pair Features (PPF) into a Transformer-based framework for robust feature representation and interaction. Additionally, an overlap prediction module estimates the overlap score between the two point clouds, improving resilience in challenging scenarios. Extensive experiments on both indoor and outdoor benchmarks demonstrate that MGFT achieves state-of-the-art performance in point cloud registration tasks, particularly in low-overlap scenarios. Shanxin Zhang, Zhenyong Li, Chuanfen Feng, Jiande Sun 0001 |
MMSP | 6 |
| 2025 | RingByte: Enhancing Text-Entry Practicality via A Singular Wearable Rotating Smart Ring
Rucheng Wu, Tao Ni 0003, Zehua Sun, Jiande Sun 0001, Weitao Xu |
UIST | 4 |
| 2025 | Joint task offloading and computing resource allocation with DQN for task-dependency in Multi-access Edge Computing
Linbo Zhai, Zekun Lu, Jiande Sun 0001, Xiaole Li |
Comput. Networks | 3 |
| 2025 | Beamforming design and trajectory optimization for learning-based multi-UAV-assisted integrated sensing and communication systems
Binglin Zhao, Linbo Zhai, Jiande Sun 0001, Chuanfen Feng, Dongsheng Wu |
Comput. Networks | 3 |
| 2025 | Orientation-Aware Reversible Data Hiding With Brainstorming Optimization for UAV Aerial ImagesabstractIn recent years, with the rapid development of unmanned aerial vehicle (UAV), aerial images have extended across various industries such as intelligent building, agriculture, transportation, and Industry 4.0. Notably, the security of UAV‐assisted data acquisition during transmission has become a critical concern. The reversible data hiding (RDH) method can hide data in aerial images for transmission and ensure secure communication. In general, an aerial image may exhibit substantially different orientation regularity from a natural scene image. This casts major challenges to the RDH method, for which existing approaches lack effective mechanisms to capture such content type variations, and thus are difficult to generalize from one type to another. In this paper, the orientation‐aware selectivity mechanism is introduced to achieve an accurate orientation‐aware prediction along different directions in local regions with different structure regularity. Furthermore, we propose a progressive brainstorming optimization algorithm (BSO)‐guided optimal PSNR value strategy, which can obtain a superior perceptual performance and the corresponding thresholds by further exploring the pixel correlations within the UAV aerial images. Experimental results on the USC‐SIPI Miscellaneous dataset and two challenging aerial datasets, including the USC‐SIPI High Altitude Aerial Imagery dataset and the Kaggle dataset, demonstrate that the proposed framework enhances the imperceptibility powerfully in marked UAV aerial images and ensures sufficient embedding capacity effectively. The average PSNR of the marked image obtained by the proposed method is 63.85 dB when embedded with 30,000 bits of data, which is an improvement of 0.59 dB compared to the current state‐of‐the‐art RDH methods. Xiaodan Tai, Yannan Ren, Jing Li 0046, Jiande Sun 0001, Kai Zhang 0010, Wenbo Wan |
Int. J. Intell. Syst. | 4 |
| 2025 | MSPL: Multi-granularity Semantic Prototype Learning for occluded person re-identification
Huaxiang Zhang 0001, Lei Zhu 0002, Jiande Sun 0001, Li Liu 0031 |
Neurocomputing | 4 |
| 2025 | Computation bits maximization in multi-UAV-assisted-multi-vehicle edge computing system
Linbo Zhai, Meiyu Jin, Jiande Sun 0001, Chuanfen Feng, Zhiquan Liu 0001, Linfeng Wei, Xiaochuan Li 0001, Youlei Zhang, Jie Liu 0040 |
J. Netw. Comput. Appl. | 4 |
| 2025 | Multiscale Integration Network With Quaternion Convolution for PansharpeningabstractIn this letter, we proposed a multiscale integration network with quaternion convolution (MQ-Net) for the fusion of low spatial resolution multispectral (LRMS) and panchromatic (PAN) images. In this network, LRMS and PAN images are resampled at different scales and fed into feature fusion modules (FFMs) to merge the spatial and spectral information among them. Then, multiscale feature enhancement modules (MFEMs) are designed to sufficiently learn the spatial and spectral information at different scales. Meanwhile, we employ a quaternion convolution module (QCM) to better capture the dependencies within spectral bands of LRMS images. Then, the quaternion features are introduced into MFEMs for efficient feature enhancement. Finally, all information from different scales is integrated for the reconstruction of high LRMS images. Reduced- and full-resolution experiments are performed on GeoEye-1 and WorldView-2 satellite datasets. Compared to some state-of-the-art pansharpening methods, the proposed MQ-Net obtains better results in terms of qualitative and quantitative evaluations. The code is available athttps://github.com/RSMagneto/MQ-Net. Yingjie Kong, Xuquan Wang, Kai Zhang 0010, Hong Li 0005, Wenbo Wan, Jiande Sun 0001 |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2025 | Spatial-spectral unfolding network with mutual guidance for multispectral and hyperspectral image fusion
Kai Zhang 0010, Qinzhu Sun, Chiru Ge, Wenbo Wan, Jiande Sun 0001, Huaxiang Zhang 0001 |
Pattern Recognit. | 6 |
| 2025 | Dual Prototypes-Based Personalized Federated Adversarial Cross-Modal HashingabstractWith the rapid advances in wireless communication and IoT platforms, it is increasingly difficult to analyze relevant multi-modal data distributed across geographically diverse and heterogeneous platforms. One promising approach is to rely on federated learning to build compact cross-modal hash codes. However, existing federated learning methods easily exhibit degenerative performance in the global model due to the distributed data being derived from diverse domains. In addition, directly forcing each client to adopt the same global parameters as local parameters, without effective local training, significantly reduces the performance of each client. To overcome these challenges, we propose a novel federated adversarial cross-modal hashing, called Dual Prototypes-based personalized Federated Adversarial (DP-FeAd), which provides iterated training of shared dual prototypes. Specifically, aiming to expand local hashing models beyond their knowledge realms, DP-FeAd enables participating clients to engage in cooperative learning through two constructions: cluster prototypes and unbiased prototypes, instead of the traditional global prototypes, ensuring both generalization and stability. Specifically, the cluster prototypes are derived from local class-level prototypes and adversarially trained with local approximate hash codes to align their distributions. The unbiased prototypes are averaged from cluster prototypes and integrated into the training of local hashing models to maintain consistency across different local class-level prototypes further. The experiments conducted on two benchmark datasets demonstrate that our proposed method significantly enhances the performance of deep cross-modal hashing models in both IID (Independent and Identically Distributed) and non-IID scenarios. Lingchen Gu, Xiaojuan Shen, Jiande Sun 0001, Jing Li 0046, Zhihui Li 0001, Sen-Ching S. Cheung, Wenbo Wan |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Heterogeneous Generative Tokens and Distance-Aware Recovery Network for Occluded Person Re-IdentificationabstractIn real-world surveillance scenarios, person re-identification tasks are often seriously affected by occlusion problems, which requires the model to be able to not only extract powerful features, but also effectively recover features when they are occluded. Although existing methods disentangle visible human bodies by clustering semantic information, they often damage discriminative appearance due to the introduction of background noises. To solve this problem, we propose Heterogeneous Generative Tokens and Distance-aware Recovery (HGTDR) network, which aims to effectively extract discriminative appearance and recover the occluded body regions. HGTDR mainly contains two branches: a holistic stream and a part stream. The holistic stream utilizes ViT to capture the global context information and provide stable global features by establishing long-range relationships. In the part stream, we propose a Semantic Patch Generator (SPG), which combines the local attention mechanism to capture rich local semantics and further generate semantic patches. Further, considering the discrimination score and relevance score of semantic patches, we feed them into the proposed Adaptive Heterogeneous Semantic Token Generator (AHSTG) to gradually generate strong-response foreground and weak-response background features. In addition, to complete the features of occluded regions, the Distance-based Feature Recovery (DFR) module is designed. The module calculates the planar Euclidean distance of heterogeneous tokens and adaptively allocates the corresponding weights to dynamically recover the invisible bodies. Finally, we obtain discriminative and robust person descriptors. Extensive experiments on several challenging occluded, partial and holistic Re-ID datasets demonstrate that our proposed HGTDR network achieves superior performance and outperforms various state-of-the-art methods. Huaxiang Zhang 0001, Lei Zhu 0002, Jiande Sun 0001, Li Liu 0031 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Attention Guidance by Cross-Domain Supervision Signals for Scene Text RecognitionabstractDespite recent advances, scene text recognition remains a challenging problem due to the significant variability, irregularity and distortion in text appearance and localization. Attention-based methods have become the mainstream due to their superior vocabulary learning and observation ability. Nonetheless, they are susceptible to attention drift which can lead to word recognition errors. Most works focus on correcting attention drift in decoding but completely ignore the error accumulated during the encoding process. In this paper, we propose a novel scheme, called the Attention Guidance by Cross-Domain Supervision Signals for Scene Text Recognition (ACDS-STR), which can mitigate the attention drift at the feature encoding stage. At the heart of the proposed scheme is the cross-domain attention guidance and feature encoding fusion module (CAFM) that uses the core areas of characters to recursively guide attention to learn in the encoding process. With precise attention information sourced from CAFM, we propose a non-attention-based adaptive transformation decoder (ATD) to guarantee decoding performance and improve decoding speed. In the training stage, we fuse manual guidance and subjective learning to learn the core areas of characters, which notably augments the recognition performance of the model. Experiments are conducted on public benchmarks and show the state-of-the-art performance. The source will be available at https://github.com/xuefanfu/ACDS-STR. Fanfu Xue, Jiande Sun 0001, Yaqi Xue, Qiang Wu 0009, Lei Zhu 0002, Xiaojun Chang, Sen-Ching S. Cheung |
IEEE Trans. Image Process. | 2 |
| 2025 | Progressive Prompt-Driven Low-Light Image Enhancement With Frequency Aware LearningabstractLow-light Image Enhancement (LLIE) aims to rectify inadequate illumination conditions and achieve superior visual quality in images, which plays a pivotal role in the domain of low-level computer vision. Due to poor illumination in images, many high-frequency details are obscured, which leads to an uneven distribution of low- and high-frequency information. However, most existing LLIE methods do not pay special attention to the restoration of high-frequency detail information and some challenging-to-recover areas in images. To address this issue, we propose a novel progressive prompt-driven LLIE framework with frequency aware learning, through a two-stage coarse-to-fine learning mechanism. Specifically, the proposed method fully utilizes both the specially designed brightness-aware prompt and detail-aware prompt on the prior trained model, to achieve an excellent enhanced image that exhibits more natural brightness and richer detail information. Furthermore, the proposed frequency aware learning objective can adaptively adjust the contribution of individual pixels for image reconstruction based on the statistics of high- and low-frequency features, which enables the network to focus on learning intricate details and other challenging areas in low-light images. Extensive experimental results demonstrate the effectiveness of the proposed method, achieving superior performances to state-of-the-art methods on representative real-world and synthetic datasets. Our source code is available athttps://github.com/MSL502/PPFAL. De Cheng, Yan Li 0125, Nannan Wang 0001, Dingwen Zhang, Xinbo Gao 0001, Jiande Sun 0001 |
IEEE Trans. Multim. | 7 |
| 2025 | Texture-Content Dual Guided Network for Visible and Infrared Image FusionabstractThe preservation and enhancement of texture information is crucial for the fusion of visible and infrared images. However, most current deep neural network (DNN)-based methods ignore the differences between texture and content, leading to unsatisfactory fusion results. To further enhance the quality of fused images, we propose a texture-content dual guided (TCDG-Net) network, which produces the fused image by the guidance inferred from source images. Specifically, a texture map is first estimated jointly by combining the gradient information of visible and infrared images. Then, the features learned by the shallow feature extraction (SFE) module are enhanced with the guidance of the texture map. To effectively model the texture information in the long-range dependencies, we design the texture-guided enhancement (TGE) module, in which the texture-guided attention mechanism is utilized to capture the global similarity of the texture regions in source images. Meanwhile, we employ the content-guided enhancement (CGE) module to refine the content regions in the fused result by utilizing the complement of the texture map. Finally, the fused image is generated by adaptively integrating the enhanced texture and content information. Extensive experiments on three benchmark datasets demonstrate the effectiveness of the proposed TCDG-Net in terms of qualitative and quantitative evaluations. Besides, the fused images generated by our proposed TCDG-Net also show better performance in downstream tasks, such as objection detection and semantic segmentation. Kai Zhang 0010, Ludan Sun, Wenbo Wan, Jiande Sun 0001, Shuyuan Yang 0001, Huaxiang Zhang 0001 |
IEEE Trans. Multim. | 5 |
| 2024 | IRS-Assisted Covert Communication with a BPP Distributed Warden outside a Safety ZoneabstractIn this work, we consider an intelligent reflecting surface (IRS)-assisted covert wireless communication system, in which the location of a warden Willie follows a binomial point process (BPP) distribution within a disk centered at O with a radius of R and the warden is outside a safety zone centered at C with a radius of r. The location of the safety zone can be anywhere (inside, outside, or having intersections) relative to the disk. We first analyze the detection performance of Willie for the cases with known channel state information (CSI) and unknown CSI, respectively, based on which we determine Willie's minimum detection error rate. Then, considering the geometric randomness of Willie's position inside the disk but outside the safety zone, the average minimum detection error rate is determined and adopted in the covertness constraint in the subsequent covert system design. Our examination first shows that the covert communication rate generally increases with R. However, more interestingly, whether the covert communication rate increases or decreases with r generally depends on the relative locations of the disk, the safety zone, and Alice. For example, when both the safety zone and Alice are on the same side of the disk center O, the covert rate increases with r, and vice versa. Shihao Yan, Xiaobo Zhou 0004, Feng Shu 0002, Jiande Sun 0001, Derrick Wing Kwan Ng |
ICASSP | 5 |
| 2024 | Gradient and Brightness Guided Low-Light Enhancement with Attention-Based Self-Paced LearningabstractLow-light image enhancement aims to reconstruct images with insufficient illumination into visually appealing representations with natural brightness. While most existing methods tend to focus on enhancing illumination, they often overlook the restoration of finer details in the enhanced image. Moreover, these methods do not adequately address the varying degradation levels observed in different regions of the image. In this study, we present a gradient and brightness guided low-light image enhancement framework, which can simultaneously augment the detail and illumination during the enhancement process. Our approach involves extracting gradient information from gamma-corrected images, which offers a remarkable advantage in preserving edge details compared to direct extraction from degraded images. To further refine the enhancement process and adaptively adjust the difficulty of samples, thereby boosting learning efficiency, we introduce an attention-based self-paced learning strategy. This strategy assigns different gradient and brightness weights based on the degradation levels within different image regions. Extensive experiments demonstrate the superiority of our proposed method over state-of-the-art approaches. The code is available at https://github.com/MSL502/GBASPL. Yan Li 0125, De Cheng, Dingwen Zhang, Luofeng Zhai, Jiande Sun 0001 |
ICASSP | 7 |
| 2024 | Efficient Image Harmonization via RGB TransformationabstractImage harmonization aims to adjust the appearance of the foreground to make it harmonious with the background, thereby maintaining visual consistency in composite images. Previous deep learning-based methods have mainly focused on reconstructing harmonized images with the same size as the input composite images, often leading to complex network structures and a large number of parameters. In this paper, we propose a simple yet effective lightweight image harmonization network architecture. First, we generate a low-resolution 3-channel feature map to represent the rough variations in the RGB channels of a composite image, which is then upsampled and added to this composite image to obtain a preliminary harmonization result. Then, a refinement module is applied to refine the preliminary result and output the final harmonized image. Additionally, we design a dynamic data generation and training strategy to pre-train our model on another dataset. Experimental results on the iHarmony4 dataset show that our method indicates a significant reduction in the number of parameters compared to other methods, yet it still achieved competitive performance. Jiande Sun 0001, Wenbo Wan, Kai Zhang 0010, Jian Wang 0004 |
MMSP | 2 |
| 2024 | ID-Gait: Fine-Grained Human Gait State Recognition Using Wi-Fi Signal
Ran Lai, Mingda Han, Linlin Guo, Jia Zhang 0028, Jiande Sun 0001 |
WASA (1) | 7 |
| 2024 | Triple disentangled network with dual attention for remote sensing image fusion
Feng Zhang 0028, Guishuo Yang, Jiande Sun 0001, Wenbo Wan, Kai Zhang 0010 |
Expert Syst. Appl. | 3 |
| 2024 | EMCFN: Edge-based Multi-scale Cross Fusion Network for video frame interpolation
Zhiquan Feng, Jiande Sun 0001 |
J. Vis. Commun. Image Represent. | 4 |
| 2024 | Deep image watermarking with loss-driven modification
Wenqing Yang, Jing Li 0046, Jiande Sun 0001, Wenbo Wan |
Multim. Tools Appl. | 6 |
| 2024 | Learning spatial-spectral dual adaptive graph embedding for multispectral and hyperspectral image fusion
Xuquan Wang, Feng Zhang 0028, Kai Zhang 0010, Weijie Wang 0002, Xiong Dun, Jiande Sun 0001 |
Pattern Recognit. | 6 |
| 2024 | Building Change Detection in Earthquake: A Multiscale Interaction Network With Offset Calibration and a DatasetabstractAs one of the most destructive natural disasters, earthquakes have struck many countries around the world in recent years, causing serious economic losses. Change detection (CD) can be applied to postearthquake building CD as it can infer interested change regions from multitemporal remote sensing (RS) images. Furthermore, the CD with short imaging intervals will better satisfy the needs of the emergency rescues after earthquakes. However, the capability of current methods built on deep neural networks (DNNs) is limited because the dataset with short imaging intervals is absent. To meet postdisaster immediate relief, we create a CD dataset, the Turkey earthquake CD dataset (TUE-CD), for the detection of building collapse in the short term after an earthquake. Due to the high requirement for timeliness of postevent images, the orbit of the satellite during postevent imaging deviates from that during preevent imaging, which leads to a side-looking problem between bitemporal images. To deal with these challenges, we present a multiscale feature interaction network (MSI-Net) for efficient interaction between bitemporal features, as well as mitigating the effect of side-looking problems. Specifically, the proposed MSI-Net consists of joint cross-attention (JCA) modules, multiscale offset calibration (MOC) modules, and feature integration (FeI) modules. The JCA module unifies channel cross-attention (CCA) and spatial joint attention (SJA) for sufficient feature interaction. The MOC module further estimates the offsets to align the bitemporal image with the multiscale features. Finally, calibrated features and multiscale features are fused by FeI modules for the prediction of changed areas. The best mF1 and mIoU scores are achieved on two public datasets and the constructed TUE-CD dataset: WHU-CD (95.58%, 91.81%), CLCD (82.96%, 73.53%), and TUE-CD (78.02%, 68.48%). Experimental results demonstrate that the proposed MSI-Net provides competitive performance compared to the state-of-the-art CD methods. The TUE-CD dataset and the code of MSI-Net will be available athttps://github.com/RSMagneto/MSI-Net. Yunlong Liu 0005, Kai Zhang 0010, Chunan Guan, Shanxin Zhang, Hong Li 0005, Wenbo Wan, Jiande Sun 0001 |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2024 | Content-Guided Spatial-Spectral Integration Network for Change Detection in HR Remote Sensing ImagesabstractThe integration of spatial and spectral information is beneficial to the improvement of change detection (CD) performance. However, existing methods cannot efficiently suppress the influences of spatial and spectral differences (SDs) in unchanged areas. To address these issues, in this article, we propose a content-guided spatial–spectral integration network (CSI-Net) for the fusion of global spatial details and SD information. Specifically, the proposed CSI-Net is composed of a spatial reasoning (SR) module, an SD module, and a content-guided integration (CGI) module. In the SR module, the spatial information is learned by cascaded graph convolution (GC) blocks for global modeling. The SD module is responsible for the extraction of spectral features, by calculating the means and variances of features to reduce the impact of SDs in unchanged regions. In addition, in order to integrate the spatial–spectral features efficiently, we design a CGI module to further take advantage of their complementary information. In this module, high-level content information is introduced as a guide for proper interaction. Due to the efficient spatial–spectral fusion, the proposed CSI-Net can learn the changed features better while achieving suppression of SDs. Experimental results on LEVIR-CD, WHU-CD, and CLCD datasets demonstrate that the proposed CSI-Net produces better performance compared to state-of-the-art methods, and is applicable to different scenarios. The code of CSI-Net is available athttps://github.com/RSMagneto/CSI-Net. Yunlong Liu 0005, Feng Zhang 0028, Shanxin Zhang, Kai Zhang 0010, Jiande Sun 0001, Lorenzo Bruzzone |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | Spectral-Spatial Dual Graph Unfolding Network for Multispectral and Hyperspectral Image FusionabstractRecently, deep neural network (DNN)-based methods have achieved good results in terms of the fusion of low spatial resolution hyperspectral (LR HS) and high spatial resolution multispectral (HR MS) images. However, the spectral band correlation (SBC) and the spatial nonlocal similarity (SNS) in hyperspectral (HS) images are not sufficiently exploited by them. To model the two priors efficiently, we propose a spectral-spatial dual graph unfolding network (SDGU-Net), which is derived from the optimization of graph regularized restoration models. Specifically, we introduce spectral and spatial graphs to regularize the reconstruction of the desired high spatial resolution hyperspectral (HR HS) image. To explore the SBC and SNS priors of HS images in feature space and utilize the powerful learning ability of DNNs simultaneously, the iterative optimization of the spectral and spatial graph regularized models is unfolded as a network, which is composed of spectral and spatial graph unfolding modules. The two kinds of modules are designed according to the solutions of the spectral and spatial graph regularized models. In these modules, we employ graph convolution networks (GCNs) to capture the SBC and SNS in the fused image. Then, the learned features are integrated by the corresponding feature fusion modules and fed into the feature condense module to generate the HR HS image. We conduct extensive experiments on three benchmark datasets and the results demonstrate the effectiveness of our proposed SDGU-Net. Kai Zhang 0010, Feng Zhang 0028, Chiru Ge, Wenbo Wan, Jiande Sun 0001, Huaxiang Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2024 | DRFormer: Learning Disentangled Representation for Pan-Sharpening via Mutual Information- Based TransformerabstractIn this article, we propose a new pan-sharpening method that disentangles low spatial resolution multispectral (LRMS) and panchromatic (PAN) images in terms of sensor-specific features and common features. These features are obtained by defining mutual information (MI)-based transformers designed to achieve disentangled learning. In the proposed method, LRMS and PAN images are cross-reconstructed by cross-coupled transformers to facilitate the disentanglement of the common features and sensor-specific features. To ensure compatibility among the disentangled features, self-reconstructions of LRMS and PAN images are imposed on them, and source images are reconstructed by self-coupled transformers. In addition to the reconstruction-guided disentangled learning, we maximize the MI between the common features of LRMS and PAN images to improve the correlation of the common features from different images. We also minimize the MI between the common features and sensor-specific features from the same image to reduce the redundancy among them. Through the reconstruction and disentangled representation of source images, sensor-specific features and common features can be decomposed efficiently. Finally, all disentangled features are integrated by a fusion transformer to generate the high spatial resolution multispectral (HRMS) image. Experiments on different datasets demonstrate that the proposed method produces competitive fusion results. The code is available athttps://github.com/RSMagneto/DRFormer. Feng Zhang 0028, Kai Zhang 0010, Jiande Sun 0001, Jian Wang 0004, Lorenzo Bruzzone |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Entropy-Optimized Deep Weighted Product Quantization for Image RetrievalabstractHashing and quantization have greatly succeeded by benefiting from deep learning for large-scale image retrieval. Recently, deep product quantization methods have attracted wide attention. However, representation capability of codewords needs to be further improved. Moreover, since the number of codewords in the codebook depends on experience, representation capability of codewords is usually imbalanced, which leads to redundancy or insufficiency of codewords and reduces retrieval performance. Therefore, in this paper, we propose a novel deep product quantization method, named Entropy Optimized deep Weighted Product Quantization (EOWPQ), which not only encodes samples into the weighted codewords in a new flexible manner but also balances the codeword assignment, improving while balancing representation capability of codewords. Specifically, we encode samples using the linear weighted sum of codewords instead of a single codeword as traditionally. Meanwhile, we establish the linear relationship between the weighted codewords and semantic labels, which effectively maintains semantic information of codewords. Moreover, in order to balance the codeword assignment, that is, avoiding some codewords representing most samples or some codewords representing very few samples, we maximize the entropy of the coding probability distribution and obtain the optimal coding probability distribution of samples by utilizing optimal transport theory, which achieves the optimal assignment of codewords and balances representation capability of codewords. The experimental results on three benchmark datasets show that EOWPQ can achieve better retrieval performance and also show the improvement of representation capability of codewords and the balance of codeword assignment. Lingchen Gu, Wenbo Wan, Jiande Sun 0001 |
IEEE Trans. Image Process. | 5 |
| 2024 | Progressive Negative Enhancing Contrastive Learning for Image Dehazing and BeyondabstractImage dehazing is a pivotal preliminary step in the advancement of robust intelligent surveillance system. However, it is an extremely challenging ill-posed problem, as it faces severe information degradation when accurately restoring the clean image from its haze-polluted counterpart. This paper proposes a novel Progressive Negative Enhancing (PNE) contrastive learning mechanism to fully exploit various types of negative information, thereby facilitating the traditional positive-oriented objective function for image dehazing. The proposed method can progressively update the negative samples during model training, to steadily squeeze the restored image towards its desired clean target from various directions. Furthermore, considering the image dehazing task as a many-to-one feature mapping problem, we also make an early effort to enhance the robustness of the dehazing model under variational haze densities. Specifically, a novel density-variational dehazing network is proposed to be optimized under the consistency-regularized framework using the proposed PNE learning mechanism. The consistency regularization ensures consistent output given multi-level degraded hazy images, thereby significantly enhancing the robustness of the model in dealing with various hazy scenarios. Extensive experiments demonstrate that the proposed method exhibits superior performance over existing state-of-the-art methods. It achieves average PSNR boosts of 0.60dB, 0.28dB and 0.82dB on dehazing, deraining and desnowing tasks, respectively. The source code is available athttps://github.com/YanLi-LY/PNE-Net. De Cheng, Yan Li 0125, Dingwen Zhang, Nannan Wang 0001, Jiande Sun 0001, Xinbo Gao 0001 |
IEEE Trans. Multim. | 5 |
| 2024 | Deep Rank-N Decomposition Network for Image FusionabstractExisting deep neural network (DNN)-based image fusion methods seldom consider low-rank priors for the decomposition of source images, which cannot efficiently model base and detail components in images. To exploit the low-rank priors better, we propose a deep rank-Ndecomposition network (DRDec-Net) according to the rank-Ndecomposition of source images. Specifically, a rank-Ndecomposition model is first established by imposing low-rank priors on the base component of source images. Then, based on the decomposition model, we construct DRDec-Net, which is composed of low-rank decomposition (LRD) modules, a detail fusion (DetailF) module, and a low-rank fusion (LRF) module. In DRDec-Net, it is assumed that source images share the same base component, which is expressed as the sum of rank-1 components. We employNcascaded LRD modules to extract these rank-1 components from source images. Meanwhile, detail components are obtained by subtracting the base component from source images. Next, the extracted rank-1 components and detail components are integrated by LRF and DetailF modules to produce the base component and detail component of the fused image. Finally, the sum of the two obtained components is regarded as the fused image. Compared to some state-of-the-art methods, experimental results demonstrate that the proposed DRDec-Net can produce a better performance on three image fusion tasks, including infrared and visible images, multi-exposure images, and multi-focus images. Ludan Sun, Kai Zhang 0010, Feng Zhang 0028, Wenbo Wan, Jiande Sun 0001 |
IEEE Trans. Multim. | 5 |
| 2023 | Covert Communications Assisted by Reconfigurable Intelligent Surfaces with Discrete Phase ShiftsabstractThis work examines the covert communication performance gain achieved by deploying a reconfigurable intelligent surface (RIS) with discrete phase shifts. To this end, we first analyze the average receive signal power at a legitimate receiver Bob and a warden Willie as a function of the number of RIS reflecting elements$N$and the number of bits$d$for its discrete phase shift levels. Our analysis reveals that Bob's average power is proportional to$N^{2}$and highly depends on$d$, while Willie's average power is proportional to$N$and does not depend on$d$. This leads to the potential of enhancing the system performance via increasing$N$or$d$in the considered covert communications scenario. Specifically, the performance gain is explicitly examined by tackling the transmission outage probability from a transmitter Alice to Bob subject to a covertness constraint based on Willie's detection performance. After analyzing the transmission outage probability and Willie's total detection error rate, we determine Alice's optimal transmit power. Our explicit examination confirms the performance enhancement achieved via increasing$N$or$d$. Furthermore, our analysis shows that the covert communication performance achieved with$d=3$is already sufficiently close to that achieved with$d\rightarrow\infty$. This shows that the major benefits of deploying RIS in covert communications can be achieved by an RIS with low-resolution phase shifts. Peilin Ren, Jia Zhang 0028, Shihao Yan, Weitao Xu, Jiande Sun 0001, Naofal Al-Dhahir |
GLOBECOM | 5 |
| 2023 | S-Feature Pyramid Network and Attention Model for Drone DetectionabstractThe issue of aviation safety has always received a great of attention and focus, and birds are also an important issue in aviation safety. Nowadays, drones have emerged and share the same airspace with birds at low altitudes. The problems associated with drones should also be taken into account. For example, small drones can be misused for illegal activities and the threat from them is on the rise. Driven by this situation, we used data provided by the ICASSP Drone-vs-Bird detection Grand Challenge for drone detection and used the method of adding shallow feature pyramid network and attention model on SSD [1] (SFA-SSD) to solve the problem of drone detection in competition. Out of 30 test videos, our method was able to detect drones in 11 videos, with 8 videos scoring above 0.1 and only 3 videos scoring above 0.7. Pengcheng Dong, Chuntao Wang, Zhenyong Lu, Kai Zhang 0010, Wenbo Wan, Jiande Sun 0001 |
ICASSP | 6 |
| 2023 | Long-Short Attention Network For The Spectral Super-Resolution Of Multispectral ImagesabstractOwing to the efficiency in terms of the modeling of long-range dependencies, transformer-based spectral reconstruction methods have produced satisfactory hyperspectral (HS) images from multispectral (MS) images. Some transformer-based methods applied self-attention to all bands in the HS image to model the relationships among them, which ignore high correlations between adjacent bands and low correlations among nonadjacent ones. To learn the global relationships among all bands and the correlations between adjacent bands simultaneously, this paper proposes a long-short attention network (LSA-Net) for the spectral super-resolution of MS images. Specifically, LSA-Net is composed of cascaded long-range attention blocks and short-range attention blocks. In long-range attention blocks, the transformer is imposed on all channels by modeling each channel as a token. Then, grouped channels are fed into short-range attention blocks for correlation learning, which is inferred from the similarities among neighboring channels. With the introduction of long- and short-range attention, the relationships among spectral bands can be preserved better. Experiments on the CAVE dataset demonstrate the effectiveness of the proposed LSA-Net. The code is available at https://github.com/RSMagneto/LSA-Net. Kai Zhang 0010, Feng Zhang 0028, Jiande Sun 0001 |
ICASSP | 4 |
| 2023 | Coupling Spatial and Channel Transformer for Single Image DerainingabstractSingle image deraining is a fundamental low-level vision task, and has evolved remarkable progress with the deep learning technique. Recently, benefiting from the powerful modeling ability of long-range dependence, transformer as an alternative architecture of the dominant convolutional neural network has demonstrated large margin performance improvement in various high-level vision tasks, and has begun to be applied for low-level vision tasks. The benchmark transformer block captures long dependence via incorporating the self-attention among the spatial points of the learned feature map, and causes heavy computational workload and memory footprint quadratically increased with spatial resolutions, making it impossible to handle high-resolution images. This study proposes a novel spatial and channel coupled Transformer to jointly explore long-range dependence and correlation in both spatial and channel domains, and results in a lightweight deraining transformer model for potentially processing high-resolution images. Yuto Namba, Jiande Sun 0001, Xianhua Han |
ICIP | 2 |
| 2023 | The Multivariate Transformer Network for Mild Cognitive Impairment IdentificationabstractThe functional brain network (FBN) has emerged as a promising biomarker for classifying brain diseases. However, most existing studies ignored the extensive interactions between gray matter (GM) and white matter (WM) which would limit the performance of the process. In this paper, we presented a GM and WM FBNs based mild cognitive impairment (MCI) identification method by utilizing spatial independent component analysis (ICA) and transformer network with rs-fMRI data. Specifically, spatial ICA was performed on the GM and WM of the preprocessed fMRI data. Then the time courses of the GM and WM independent components were inputted into the transformer network with attention mechanism, which led to the optimized and extensive inter- and intra-matter functional connectivity. Finally, the fused GM and WM FBNs were constructed for MCI identification. Experimental results on ADNI dataset demonstrated the superior performance of the proposed method over the state-of-the-art methods. Jianping Qiao, Hongjia Liu, Zhishun Wang, Jiande Sun 0001 |
ICIP | 5 |
| 2023 | Effective Occlusion Suppression Network via Grouped Pose Estimation for Occluded Person Re-IdentificationabstractThe occluded target person often lead to incorrect matching results, so eliminating the interference of background clutter is a key matter. To deal with the challenging occluded person re-identification (Re-ID) tasks, this paper proposes a new grouped pose estimation occlusion generation (GPEOG) network to fully use the person grouped keypoint information and the stable non-occluded features, and make the extracted features more robust and discriminative. The local branch uses the proposed Dynamic Adaptive module (DAM) to automatically adjust the patch size during training process and provides adaptive local features. The occlusion-suppression generation branch mainly includes the proposed effective occlusion suppression (EOS) module. It judges whether the person’s keypoint groups are occluded, and then generates the masks of the corresponding parts to suppress occlusion. Extensive experiments show that the proposed method has obvious advantages compared with several state-of-the-art methods. Huaxiang Zhang 0001, Lei Zhu 0002, Jiande Sun 0001, Li Liu 0031 |
ICME | 4 |
| 2023 | Multi-scale 3D-CNN for Alzheimer's Disease ClassificationabstractThe diffusion tensor imaging (DTI) based Alzheimer's disease (AD) classification consists of the region of interest-based methods, voxel-based methods, and fiber tracer map-based methods. However, most of the studies utilize partial information of the DTI data. It is difficult to discover most discriminative biomarkers. To address this problem, this paper proposes a novel AD classification method based on 3D convolutional neural networks (3D-CNN) with multi-scale information fusion in order to uncover the differential features of AD. First, the significant voxels of anisotropy score and mean dispersion rate for each individual are extracted using Tract-Based Spatial Statistics (TBSS) method from the skeleton. Second, the texture and intensity features are extracted by utilizing radiomics method. Third, an improved 3D convolutional neural network is proposed to extract depth features from the fractional anisotropy (FA) and mean diffusivity (MD) images. Finally, the multi-scale features are linearly fused and select with LASSO algorithm. The support vector machine is utilized for five-fold cross-validation. The experimental results on the ADNI dataset with 185 DTI images containing AD and normal controls (NC) subjects show that the proposed method achieves better classification performance compared with existing methods for the AD assist diagnosis task. Hang Yan 0008, Kunlun Fang, Hao Shang, Hongjia Liu, Jiande Sun 0001, Jianping Qiao |
ISCAS | 5 |
| 2023 | Automatic Image-to-Color Point Cloud Cross-modal Registration Based on Graph Neural Networks and Iterative ReprojectionabstractImage-to-Color point cloud registration establishes a connection between two-dimensional image data and three-dimensional point cloud data, and plays a vital role in the intelligent city, autonomous driving, and robotics field. However, it is still challenging to automatically register an image and its surrounding color point clouds together, due to the special application scenario, unknown camera intrinsic parameters, lack of train data, and a large amount of noise. To overcome those issues, we propose an iterative reprojection architecture that automatically acquires the matched 2D-3D keypoints pairs between the image and the color point clouds by graph optimization method and mapping transfer first, then completes registration by Alternating Direction Method of Multipliers (ADMM). Experiments results show that the proposed method is more accurate than the manual way. Shanxin Zhang, Jiande Sun 0001, Cheng Wang 0003, Jonathan Li 0001 |
ISCAS | 3 |
| 2023 | Contextual transformer sequence-based recognition network for medical examination reports
Honglin Wan, Zongfeng Zhong, Tianping Li, Huaxiang Zhang 0001, Jiande Sun 0001 |
Appl. Intell. | 5 |
| 2023 | AFcIHNet: Attention feature-constrained network for single image information hiding
Xingwang Jia, Hua-Mei Xin 0001, Lingchen Gu, Jiande Sun 0001, Wenbo Wan |
Eng. Appl. Artif. Intell. | 5 |
| 2023 | Multi-dimensional constraints-based PPVO for high fidelity reversible data hiding
Wenxiu Liu, Lili Meng, Jiande Sun 0001, Wenbo Wan |
Expert Syst. Appl. | 4 |
| 2023 | Towards perceptual image watermarking with robust texture measurement
Yunming Zhang, Yuxin Gong, Jun Wang 0061, Jiande Sun 0001, Wenbo Wan |
Expert Syst. Appl. | 4 |
| 2023 | 3D pedestrian localization fusing via monocular camera
Jiande Sun 0001, Shanxin Zhang, Hui Yuan 0001, Huaxiang Zhang 0001, Jia Zhang 0028 |
J. Vis. Commun. Image Represent. | 2 |
| 2023 | A cover selection-based reversible data hiding method by learning cross-modal hashing
Liming Zou, Jiande Sun 0001, Wenbo Wan, Jing Li 0046, Q. M. Jonathan Wu |
Multim. Tools Appl. | 2 |
| 2023 | CanBiPT: Cancelable biometrics with physical template
Youjun Gao, Jiande Sun 0001, Huaxiang Zhang 0001, Wenbo Wan |
Pattern Recognit. Lett. | 4 |
| 2023 | Multispectral and hyperspectral image fusion based on low-rank unfolding network
Kai Zhang 0010, Feng Zhang 0028, Chiru Ge, Wenbo Wan, Jiande Sun 0001 |
Signal Process. | 6 |
| 2023 | 3D geometrical total variation regularized low-rank matrix factorization for hyperspectral image denoising
Feng Zhang 0028, Kai Zhang 0010, Wenbo Wan, Jiande Sun 0001 |
Signal Process. | 4 |
| 2023 | Video Sampled Frame Category Aggregation and Consistent Representation for Cross-Modal RetrievalabstractMany current video and text cross-modal retrieval research works focus on narrowing the semantic gap between video and text, but ignore the semantic difference between different sampled frames in the same video and the correlation of feature distribution of objects contained in different sampled frames in the same video, as a result, the features of the sampled frames in the final learned video cannot well represent the semantic features of the whole video. To overcome the shortcomings of existing studies, we first use a pre-trained video frame classification-aggregation network to make the object categories contained in different sampled frames in the same video be more close to the important object categories contained in the whole video, so as to promote the feature distribution of different sampled frames in the same video to be consistent, and increase the relevance of object features in different frames. Then we propose a video internal frame aggregation loss module to solve the problem of inconsistent feature distribution between different frame features encoded by video encoder in the same video and the aggregation feature of the sampled frame, thus enhancing the ability of video sampled frame aggregation feature representation. Experiments conducted on three common datasets MSVD, MSR-VTT and DiDeMo demonstrate the validity of the proposed approach. Ming Jin 0007, Huaxiang Zhang 0001, Lei Zhu 0002, Jiande Sun 0001, Li Liu 0031 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | A Unified Two-Stage Spatial and Spectral Network With Few-Shot Learning for PansharpeningabstractRecently, pan-sharpening methods based on deep learning (DL) have achieved state-of-the-art results. However, current existing DL-based pan-sharpening methods need to be trained repetitively for different satellite sensors to obtain satisfactory fusion performance and therefore require a large number of training images for each satellite. To deal with these issues, in this paper we propose a unified two-stage spatial and spectral network (UTSN) for pan-sharpening. A branch of networks is constructed for each different satellite, in which the spatial enhancement network (SEN) is shared to improve the spatial details in the fused images from different satellites. A spectral adjustment network (SAN) is employed to capture the spectral characteristics of the specific satellite. Through SAN, the spectral information in the intermediate image from SEN is refined to produce the final fusion results. Such a framework can integrate the datasets from different satellites together for sufficient training of SEN. The proposed method is able to achieve promising pan-sharpening results also for a new satellite with limited training images by only learning a new SAN on the few-shot datasets due to the simple but efficient structure of SAN. The experimental results show that the proposed method can produce state-of-the-art fusion results in both the standard and few-shot cases. The source code is publicly available at https://github.com/RSMagneto/UTSN. Zhi Sheng, Feng Zhang 0028, Jiande Sun 0001, Yanyan Tan, Kai Zhang 0010, Lorenzo Bruzzone |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Cross-Resolution Semi-Supervised Adversarial Learning for PansharpeningabstractExisting deep neural network (DNN)-based methods have produced good pansharpened images. However, supervised DNN-based pansharpening methods suffer from performance degradation when fusing low spatial resolution multispectral (LR MS) and panchromatic (PAN) images at full resolution. Unsupervised DNN-based methods alleviate the issue, but their training becomes difficult owing to the absence of reference images. This article establishes a novel semi-supervised framework to jointly learn the reconstructions of fused images from reduced- and full-resolution datasets. Specifically, we propose a cross-resolution semi-supervised adversarial learning network (CrossNet), which is composed of a supervised module and an unsupervised module. In these two modules, reduced- or full-resolution source images are disentangled as resolution-invariant components and resolution-aware components by reconstructing the fused images. Moreover, cross-resolution fused images are synthesized to enhance the disentanglement of the two kinds of components. Through the reconstruction of cross-resolution fused images, supervised and unsupervised modules are also coupled efficiently. Then, the semi-supervised framework can simultaneously make use of the supervised information in reduced-resolution datasets and mitigate the performance degradation via full-resolution datasets. Besides, adversarial learning is employed to improve the consistency between resolution-invariant components of source images at different resolutions. Finally, extensive experiments on QuickBird, GeoEye-1, WorldView-2, and WorldView-3 datasets demonstrate that the proposed CrossNet can produce state-of-the-art fusion results in terms of qualitative and quantitative evaluations. The source code is available athttps://github.com/RSMagneto/CrossNet. Guishuo Yang, Kai Zhang 0010, Feng Zhang 0028, Jian Wang 0004, Jiande Sun 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | Spatial-Spectral Dual Back-Projection Network for PansharpeningabstractDeep unfolding networks have obtained satisfactory performance in the pansharpening task owing to their sufficient interpretability. Inspired by the back-projection (BP) mechanism, we propose a BP-driven model, spatial-spectral dual back-project network (S2DBPN), to fuse the low spatial resolution multispectral (LR MS) and the high spatial resolution panchromatic (PAN) images by exploiting the BP in spatial and spectral domains. Specifically, the proposed S2DBPN is made up of a spatial BP network, a spectral BP network, and a reconstruction network. In the spatial BP network, spatial down- and up-projection modules are derived from BP, which is responsible for the projection of the LR MS image into the spatial domain. By analogy with the spatial BP, we reformulate the degradation between high spatial resolution multispectral (HR MS) and PAN images as spectral down- and up-projections. Then, the spectral BP network is constructed for the projection of the PAN image along the channel dimension. Finally, the features from spatial and spectral BP networks are integrated to produce the desired HR MS image through the reconstruction network. Compared to the state-of-the-art methods, extensive experiments on QuickBird, GeoEye-1, and WorldView-2 datasets demonstrate that our S2DBPN produces better HR MS images in terms of qualitative and quantitative evaluation metrics. The code of S2DBPN is released at: https://github.com/RSMagneto/S2DBPN. Kai Zhang 0010, Anfei Wang, Feng Zhang 0028, Wenbo Wan, Jiande Sun 0001, Lorenzo Bruzzone |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | Learning Deep Multiscale Local Dissimilarity Prior for PansharpeningabstractVarious deep neural networks (DNNs) have been constructed to inject the spatial information of the panchromatic (PAN) image into the low spatial resolution multispectral (LR MS) image. However, most of them ignore the local dissimilarity (LD) prior between MS and PAN images, which has a negative influence on the fused image. Considering the above-mentioned issues, we propose a deep multiscale local dissimilarity network (DMLD-Net) to learn the LD prior at different scales and enhance the spatial and spectral information in the fused image better. Specifically, we first synthesize a downsampled PAN image from the original PAN image to match the scale of the LR MS image. Then, a LD metric is designed to calculate the dissimilarity map between the two images in feature space. According to the learned dissimilarity map, we utilize a LD-guided attention block (LDGAB) to suppress the impact of LD, which filters out the dissimilar information in the features of the PAN image. To learn the LD prior between MS and PAN images sufficiently, the multiscale architecture is considered and we infer the dissimilar maps hierarchically and inject filtered features into the LR MS image progressively. Finally, the fused image is generated by a reconstruction block. Through the LD learning at different scales, reasonable spatial information is extracted from the PAN image, by which the distortions in the fused image caused by LD can be reduced efficiently. Extensive experiments are conducted on GeoEye-1 and WorldView-2 datasets and the results demonstrate the effectiveness of the proposed DMLD-Net in terms of spatial and spectral preservation. The code is available at https://github.com/RSMagneto/DMLD-Net. Kai Zhang 0010, Guishuo Yang, Feng Zhang 0028, Wenbo Wan, Man Zhou 0003, Jiande Sun 0001, Huaxiang Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2023 | Relation Changes Matter: Cross-Temporal Difference Transformer for Change Detection in Remote Sensing ImagesabstractThanks to their capability of modeling global information, transformers have been recently applied to change detection in remote sensing images. Generally, the changes in terms of shape and appearance of objects lead to relation changes among these objects in multi-temporal images. However, in this context, the attention mechanism in transformers has not been fully explored yet to learn relation changes in the observed scenes. In this paper, we analyze the relation changes in multi-temporal images and propose a cross-temporal difference (CTD) attention to capture these changes efficiently. Through the CTD attention, the changed areas are distinguished better from the unchanged areas. Based on the CTD attention, two CTD-transformer encoders are constructed to extract the features of changed areas from the embedded tokens of multi-temporal images in a cross manner. Then, the extracted features at the coarse scale are further improved to the fine-scale by the corresponding CTD-transformer decoders. In addition, consistency-perception blocks (CPBs) are designed to preserve the structures and contours of changed areas. Finally, all extracted features from multi-temporal images are concatenated to produce the desired change map. Compared to state-of-the-art methods, experimental results on LEVIR-CD, WHU-CD, and CLCD datasets demonstrate that the proposed method produces better performance. The source code is available at https://github.com/RSMagneto/CTD-Former. Kai Zhang 0010, Feng Zhang 0028, Lei Ding 0008, Jiande Sun 0001, Lorenzo Bruzzone |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | TSINIT: A Two-Stage Inpainting Network for Incomplete TextabstractAlthough there are lots of studies on scene text recognition, few of them focus on the recognition of the incomplete text. The recognition performance of existing text recognition algorithms on the incomplete text is far from the expected, and the recognition of the incomplete text is still challenging. In this paper, an end-to-end Two-Stage Inpainting Network for Incomplete Text (TSINIT) is proposed to reconstruct the incomplete text into the complete one even when the text is in various styles and with various backgrounds, and the reconstructed text can be recognized by the existing text recognition algorithms correctly. The proposed TSINIT is divided into text extraction module (TEM) and text reconstruction module (TRM) to make the inpainting only focus on the text. TEM separates the incomplete text from the background and character-like regions at the pixel level, which can reduce the ambiguity of text reconstruction caused by the background. TRM reconstructs the incomplete text towards the most possible text with the consideration of the abstract and semantic structures of the text. Furthermore, we build a synthetic incomplete text dataset (SITD), which contains contaminated and abraded text images. SITD is divided into 6 incomplete levels according to the number of pixels in the incomplete regions and the ratio of the incomplete characters to all characters. The experimental results show that the proposed method has better inpainting ability for the incomplete text compared with traditional image inpainting algorithms on the proposed SITD and real images. When using the same text recognition method, the recognition accuracy of the incomplete text on SITD can be improved much more with the help of the proposed TSINIT than with the traditional image inpainting methods. Jiande Sun 0001, Fanfu Xue, Jing Li 0046, Lei Zhu 0002, Huaxiang Zhang 0001, Jia Zhang 0028 |
IEEE Trans. Multim. | 1 |
| 2023 | Robust Coverless Image Steganography Based on Neglected Coverless Image Dataset ConstructionabstractMost of the existing image selection-based coverless image steganography methods mainly focus on improving the capacity and robustness under the assumption that the corresponding dataset is available. But they ignore how to successfully construct the coverless image dataset, which is the foundation of such methods and has a critical impact on the capacity. In this paper, a coverless image steganography is proposed that considers how to efficiently construct the coverless image dataset. In the proposed method, the CNN-based deep hash is extracted from the image and a specific mapping rule is designed to map the high-dimensional deep hash to the low-dimensional secret message. In addition, an unsupervised clustering algorithm is adopted to construct the coverless image dataset, which makes the construction of the coverless image dataset efficient and improves the robustness of the proposed steganography method. To our best knowledge, this is the first attempt to improve the construction efficiency of the coverless image dataset in the field of coverless image steganography. Experimental results show that the construction of a large coverless image dataset is feasible and reliable, and the proposed method has better robustness and higher dataset utilization rate compared with the state-of-the-art methods. Liming Zou, Jing Li 0046, Wenbo Wan, Q. M. Jonathan Wu, Jiande Sun 0001 |
IEEE Trans. Multim. | 5 |
| 2023 | ZeRGAN: Zero-Reference GAN for Fusion of Multispectral and Panchromatic ImagesabstractIn this article, we present a new pansharpening method, a zero-reference generative adversarial network (ZeRGAN), which fuses low spatial resolution multispectral (LR MS) and high spatial resolution panchromatic (PAN) images. In the proposed method, zero-reference indicates that it does not require paired reduced-scale images or unpaired full-scale images for training. To obtain accurate fusion results, we establish an adversarial game between a set of multiscale generators and their corresponding discriminators. Through multiscale generators, the fused high spatial resolution MS (HR MS) images are progressively produced from LR MS and PAN images, while the discriminators aim to distinguish the differences of spatial information between the HR MS images and the PAN images. In other words, the HR MS images are generated from LR MS and PAN images after the optimization of ZeRGAN. Furthermore, we construct a nonreference loss function, including an adversarial loss, spatial and spectral reconstruction losses, a spatial enhancement loss, and an average constancy loss. Through the minimization of the total loss, the spatial details in the HR MS images can be enhanced efficiently. Extensive experiments are implemented on datasets acquired by different satellites. The results demonstrate that the effectiveness of the proposed method compared with the state-of-the-art methods. The source code is publicly available at https://github.com/RSMagneto/ZeRGAN. Wenxiu Diao, Feng Zhang 0028, Jiande Sun 0001, Yinghui Xing, Kai Zhang 0010, Lorenzo Bruzzone |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2022 | GaitStrip: Gait Recognition via Effective Strip-Based Feature Representations and Multi-level Framework
Beibei Lin, Xianda Guo, Lincheng Li, Jiande Sun 0001, Shunli Zhang 0005, Xin Yu 0002 |
ACCV (4) | 6 |
| 2022 | Unsupervised Generative Network for Blind Hyperspectral Image Super-ResolutionabstractHyperspectral (HS) imaging sacrifices spatial resolution to ensure a high spectral resolution when capturing the detailed spectral signature at each spatial location of the scene. To compensate for this deficiency, fusing low-resolution HS (LR-HS) images with high-resolution RGB (HR-RGB) images to obtain high-resolution HS (HR-HS) images has attracted remarkable attention. Recently, deep learning-based fusion methods in a fully-supervised manner have been proven to make great progress in hyperspectral image super-resolution (HSI-SR) tasks. However, these methods require collecting a large number of training samples and constructing a non-blind prediction model to super-resolve the observations captured under controlled imaging conditions. This study proposes a novel unsupervised generative network (UGN) for learning network parameters using the observed LR-HS, HR-RGB only without the corresponding ground-truth, and designs the spatial and spectral degradation blocks to automatically learn the image degradation operations for constructing an end-to-end blind HSI SR framework. To verify the effectiveness of our proposed method, we con-duct experiments on two benchmark HS image datasets and demonstrate superior performance compared with the super-vised and unsupervised blind/non-blind SoTA methods. Zhe Liu 0039, Xianhua Han, Jiande Sun 0001, Yen-Wei Chen 0001 |
ICIP | 3 |
| 2022 | WiID: Precise WiFi-based Person Identification via Bio-electromagnetic InformationabstractFrom the perspective of privacy protection and convenience, WiFi-based person identification in wireless sensing has attracted extensive attention in recent years. In this paper, we propose a WiFi-based person IDentification (WiID) method, which can capture people’s valid physiological information from Channel State Information (CSI) of different spatial streams even when people are in motion. The key idea is to detect and extract the short-time static states from the collected CSI and achieve person identification based on these short-time signals. By designing a Motion Sensitivity Vector (MSV) conversion algorithm, WiID is able to segment CSI that carries individual physiological information automatically without the individual performing an assigned action or maintaining a specific state. As far as we know, it is the first work that enables precise person identification using people’s physiological information when people do not keep stationary. Experimental results in real-life scenarios show that WiID can achieve 92.65% of average accuracy in three different environments. Mingda Han, Linlin Guo, Jia Zhang 0028, Zihan Diao, Jiande Sun 0001 |
ICPR | 6 |
| 2022 | Robust Single Image Dehazing Based on Consistent and Contrast-Assisted ReconstructionabstractSingle image dehazing as a fundamental low-level vision task, is essential for the development of robust intelligent surveillance system. In this paper, we make an early effort to consider dehazing robustness under variational haze density, which is a realistic while under-studied problem in the research filed of singe image dehazing. To properly address this problem, we propose a novel density-variational learning framework to improve the robustness of the image dehzing model assisted by a variety of negative hazy images, to better deal with various complex hazy scenarios. Specifically, the dehazing network is optimized under the consistency-regularized framework with the proposed Contrast-Assisted Reconstruction Loss (CARL). The CARL can fully exploit the negative information to facilitate the traditional positive-orient dehazing objective function, by squeezing the dehazed image to its clean target from different directions. Meanwhile, the consistency regularization keeps consistent outputs given multi-level hazy images, thus improving the model robustness. Extensive experimental results on two synthetic and three real-world datasets demonstrate that our method significantly surpasses the state-of-the-art approaches. De Cheng, Yan Li 0125, Dingwen Zhang, Nannan Wang 0001, Xinbo Gao 0001, Jiande Sun 0001 |
IJCAI | 6 |
| 2022 | TAGAN: Texture and Attention Guided Generative Adversarial Network for Image Super ResolutionabstractSuper Resolution (SR) methods based on Generative Adversarial Networks (GANs) accomplish predominant execution in visual perception and image quality. These methods are mainly generated by traditional Peak-Signal-to-Noise-Ratio (PSNR)oriented or perceptual-driven. As the reconstruction process usually loses high frequency information, various methods aim to preserve more details. To make the details of the generated image richer, the Gradient Weight (GW) loss is introduced in the proposed method, because the gradient can reflect the texture of the image to a certain extent. The GW loss function is helpful to improve the edge and detailed texture of the generated image. Furthermore, we introduce attention mechanism to the image reconstruction block via Squeeze and Excitation Net (SENet). Attention mechanism can effectively aggregate the global features obtained by the nonlinear mapping network, and improve the channel sensitivity of the model. With the help of GW and attention mechanism, the proposed method can achieve better performance and visual quality in image texture detail restoration. The performance comparison between the state-of-the-art methods and our proposed method verifies the feasibility and reliability of the proposed method. Haitao Wang 0023, Jiande Sun 0001, Wenxiu Diao, Jing Li 0046, Kai Zhang 0010 |
ISCAS | 2 |
| 2022 | INIT: Inpainting Network for Incomplete TextabstractIn recent years, scene text recognition algorithms have achieved great progress, but they still face some challenges in practical environment, such as the incomplete scene text, which includes structurally broken characters and occluded characters, as shown in Fig. 1. Existing scene text recognition algorithms can not accurately recognize such incomplete scene text. In this paper, we design an end-to-end Inpainting Network for Incomplete Text (INIT), which can reconstruct each incomplete character into complete character. INIT can separate the text from the background and just reconstruct the text regions, which reduces the influence caused by the background. And INIT is supervised by reconstruction loss and semantic loss to reach the most likely text. Furthermore, to compensate for the absence of incomplete scene text dataset, we propose an incomplete text synthesis method and build an incomplete text dataset (SITD), which is more suitable for practical scenarios. On SITD, the recognition accuracy can achieve 82.99% by the existing text recognition method with the help of INIT, while the accuracy is only 61.99% by the same recognition method with the conventional image inpainting method. Experimental results show that the proposed method has better reconstruction ability for incomplete text compared with the existing image inpainting algorithms. Fanfu Xue, Jia Zhang 0028, Jiande Sun 0001, Jinghui Yin, Liming Zou, Jing Li 0046 |
ISCAS | 3 |
| 2022 | Energy Efficiency Optimization for RIS Assisted RSMA System over Estimated Channel
Caina Gao, Jia Zhang 0028, Linlin Guo, Lili Meng, Jiande Sun 0001 |
WASA (1) | 6 |
| 2022 | Coordinated rate splitting and power allocation in energy-spectral efficiency tradeoff-based multicell networks
Caina Gao, Jia Zhang 0028, Linlin Guo, Lili Meng, Jiande Sun 0001 |
Comput. Networks | 6 |
| 2022 | Robust watermarking based on blur-guided JND model for macrophotography imagesabstractMacrophotography Images (MPIs) have recently emerged as an active topic due to the development of mobile phone camera technology. A large number of MPIs have been rapidly increasing in many rich visual services, such as smartphones or high-definition monitors. MPIs are often composed of sharp macroimage and blur background, which exhibit different perceptual properties that often lead to different just noticeable difference (JND) estimation. Inspired by this, we formulate the blur concealment (BC) as another factor to determine the total masking effect: the interaction is relatively straightforward with a limited masking effect in the sharp regions, and is complicated with a strong masking effect in the blur parts. Furthermore, texture and orientation adaption and color information weighting are separately incorporated into the contrast masking and color masking. Finally, considering both BC and masking effects, a novel robust watermarking framework based on the proposed blur-guided JND model for MPIs, targeting at further improving the MPIs copyright protection performance. Extensive experiments on MP2020 and Blur Detection data sets show that the applicability of the proposed JND model in the scenario of perceptually MPIs watermarking, and our proposed scheme can outperform the state-of-the-art watermarking schemes by providing better robustness performance at the uniform visual quality. Wenbo Wan, Wenqian Shan, Wenxiu Liu, Zihan Diao, Jiande Sun 0001 |
Int. J. Intell. Syst. | 6 |
| 2022 | A comprehensive survey on robust image watermarking
Wenbo Wan, Jun Wang 0061, Yunming Zhang, Jing Li 0046, Hui Yu 0001, Jiande Sun 0001 |
Neurocomputing | 6 |
| 2022 | Deep Discrete Cross-Modal Hashing with Multiple Supervision
En Yu, Jiande Sun 0001, Xiaojun Chang, Huaxiang Zhang 0001, Alex Hauptmann 0001 |
Neurocomputing | 3 |
| 2022 | Coarse-to-fine dual-level attention for video-text cross modal retrieval
Ming Jin 0007, Huaxiang Zhang 0001, Lei Zhu 0002, Jiande Sun 0001, Li Liu 0031 |
Knowl. Based Syst. | 4 |
| 2022 | Single image dehazing with an independent Detail-Recovery Network
Yan Li 0125, De Cheng, Dingwen Zhang, Nannan Wang 0001, Xinbo Gao 0001, Jiande Sun 0001 |
Knowl. Based Syst. | 6 |
| 2022 | Discrete Fusion Adversarial Hashing for cross-modal retrieval
Jing Li 0046, En Yu, Xiaojun Chang, Huaxiang Zhang 0001, Jiande Sun 0001 |
Knowl. Based Syst. | 6 |
| 2022 | HLF-Net: Pansharpening Based on High- and Low-Frequency Fusion NetworksabstractMany deep neural networks have been constructed for the pansharpening task. However, the differences between the high and low frequencies in images are not considered in some DNN-based pansharpening methods. As high and low frequencies have different information of images, it is difficult for the same network to learn and reconcile the two kinds of frequencies. Considering the aforementioned differences, we propose a new pansharpening network to fuse the high and low frequencies in low spatial resolution multispectral and panchromatic images separately. Specifically, a high and low frequency fusion network is constructed, which is composed of a high-frequency fusion network and a low-frequency fusion network. In the high-frequency fusion network, skip attention is introduced into U-Net to better retain the high frequencies in feature maps. The low-frequency fusion network uses the involution to capture the dependency among the channels of feature maps. Experiments on the GeoEye-1 dataset reveal that the proposed network outperforms some state-of-the-art methods. The code can be accessed at https://github.com/RSMagneto/HLF-Net. Wenxiu Diao, Feng Zhang 0028, Haitao Wang 0023, Wenbo Wan, Jiande Sun 0001, Kai Zhang 0010 |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2022 | Unsupervised Change Detection of Multispectral Images Based on PCA and Low-Rank PriorabstractIn this letter, we propose a new unsupervised change detection method based on low-rank prior for multispectral images. It is assumed that the changed and unchanged pixels are from different subspaces due to different appearance and statistical properties. So, low-rank representation (LRR) is employed to find informative pixels from the superpixels of the difference image (DI). Besides, taking the sparsity of changed pixels in the observed scenes into consideration, the selection rule is designed to distinguish these pixels. Then, principal component analysis (PCA) is used for the training of changed and unchanged dictionaries from these pixels. Finally, the change map is estimated by comparing the reconstruction error of each pixel in DI on changed and unchanged dictionaries. By LRR, more representative pixels are found for subsequent dictionary learning, which can efficiently improve the performance of the proposed method. Experiments on multitemporal images from the Landsat satellite demonstrate the effectiveness of the proposed method. Jing Li 0046, Feng Zhang 0028, Jiande Sun 0001, Kai Zhang 0010 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2022 | Pan-Sharpening Based on Transformer With Redundancy ReductionabstractPan-sharpening methods based on deep neural network (DNN) have produced the state-of-the-art results. However, the common information in the panchromatic (PAN) image and the low spatial resolution multispectral (LRMS) image is not sufficiently explored. As PAN and LRMS images are collected from the same scene, there exists some common information among them, in addition to their respective unique information. The direct concatenation of extracted features leads to some redundancy in the feature space. To reduce the redundancy among features and exploit the global information in source images, we proposed a novel pan-sharpening method by combining the convolution neural network and transformer. Specifically, PAN and LRMS images are encoded as unique features and common features by the subnetworks consisting of convolution blocks and transformer blocks. Then, the common features are averaged and combined with unique features from source images for the reconstruction of the fused image. To extract accurate common features, the equality constraint is imposed on them. Experimental results show that the proposed method outperforms the state-of-the-art methods on both reduced-scale and full-scale datasets. The source code is available athttps://github.com/RSMagneto/TRRNet. Kai Zhang 0010, Feng Zhang 0028, Wenbo Wan, Jiande Sun 0001 |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2022 | Spatial and Spectral Extraction Network With Adaptive Feature Fusion for PansharpeningabstractPansharpening methods based on deep neural networks (DNNs) have been attracting great attention due to their powerful representation capabilities. In this article, to combine the feature maps from different subnetworks efficiently, we propose a novel pansharpening method based on a spatial and spectral extraction network (SSE-Net). Different from the other methods based on DNNs that directly concatenate the features from different subnetworks, we design adaptive feature fusion modules (AFFMs) to merge these features according to their information content. First, the spatial and spectral features are extracted by the subnetworks from low spatial resolution multispectral (LR MS) and panchromatic (PAN) images. Then, by fusing the features at different levels, the desired high spatial resolution MS (HR MS) images are generated by the fusion network consisting of AFFMs. In the fusion network, the features from different subnetworks are integrated adaptively, and the redundancy among them is reduced. Moreover, the spectral ratio loss and the gradient loss are defined to ensure the effective learning of spatial and spectral features. The spectral ratio loss captures the nonlinear relationships among the bands in the MS image to reduce the spectral distortions in the fusion result. Extensive experiments were conducted on QuickBird and GeoEye-1 satellite datasets. Visual and numerical results demonstrate that the proposed method produces better fusion results compared with literature techniques. The source code is available athttps://github.com/RSMagneto/SSE-Net. Kai Zhang 0010, Anfei Wang, Feng Zhang 0028, Wenxiu Diao, Jiande Sun 0001, Lorenzo Bruzzone |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2021 | Monocular 3D Pedestrian Localization Fusing with Bird's Eye ViewabstractIn recent years, 3D target detection and location methods in the field of autonomous driving have attracted increasing attention, but monocular 3D pedestrian localization research is still facing challenges. In this paper, a monocular pedestrian localization framework and a fine-grained location optimization method which is based on a bird's eye view are proposed. The monocular pedestrian localization framework is divided into three parts, which are coarse-grained localization, depth information reconstruction and fine-grained location optimization. In the stage of coarse-grained location, the human skeleton information is obtained from the original image by using the human skeleton point detection method, and then the pedestrian position is predicted through the method of a light-weight feed-forward neural network. In the stage of depth information reconstruction, the original image is used to reconstruct the corresponding bird's eye view with depth information through a parallel network. Finally, a fine-grained positioning optimization method makes it possible to get a more precise location with the help of the last two stages. The experimental results on the KITTI dataset show that our method has achieved better performance than the state-of-the-art methods. Shanxin Zhang, Hui Yuan 0001, Xinghai Yang, Huaxiang Zhang 0001, Jiande Sun 0001 |
ISCAS | 6 |
| 2021 | JND-aware robust image watermarking with tri-directional inter-block correlationabstractA novel block-level perceptual image watermarking framework is proposed in this study, including tri-directional correlation and a block-level just noticeable difference (JND) model. Specifically, the difference in the discrete cosine transform (DCT) coefficients of two blocks is calculated based on three directions in the neighborhood, called the tri-directional correlation (TriDC). Additionally, the representative alternating current (AC) coefficients along horizontal, vertical, and diagonal directions, which can describe structural patterns, are projected and merged for TriDC differences. Then, the difference of the DCT coefficient is modulated to a predefined zone depending on the JND-based offset. Finally, the extent of the watermarked AC coefficients is determined with perceptual JND adjustment. The experimental results demonstrate that the proposed scheme can protect most common image processing attacks; and has better robustness compared with recent zone modulation watermarking schemes and traditional watermarking methods. Yunming Zhang, Zhenhua Wang 0004, Yantong Zhan, Lili Meng, Jiande Sun 0001, Wenbo Wan |
Int. J. Intell. Syst. | 5 |
| 2021 | Variable-length image compression based on controllable learning network
Dong Zhao 0017, Jiande Sun 0001, Hongchao Zhou |
Multim. Tools Appl. | 2 |
| 2021 | Visual Security Assessment via Saliency-Weighted Structure and Orientation Similarity for Selective Encrypted ImagesabstractSelective encryption has been widely used in image privacy protection. Visual security assessment is necessary for the effectiveness and practicability of image encryption methods, and there have been a series of research studies on this aspect. However, these methods do not take into account perceptual factors. In this paper, we propose a new visual security assessment (VSA) by saliency-weighted structure and orientation similarity. Considering that the human visual perception is sensitive to the characteristics of selective encrypted images, we extract the structure and orientation feature maps, and then similarity measurements are conducted on these feature maps to generate the structure and orientation similarity maps. Next, we compute the saliency map of the original image. Then, a simple saliency-based pooling strategy is subsequently used to combine these measurements and generate the final visual security score. Extensive experiments are conducted on two public encryption databases, and the results demonstrate the superiority and robustness of our proposed VSA compared with the existing most advanced work. Zhengguo Wu, Kai Zhang 0010, Yannan Ren, Jing Li 0046, Jiande Sun 0001, Wenbo Wan |
Secur. Commun. Networks | 5 |
| 2021 | Deep Loss Driven Multi-Scale Hashing Based on Pyramid Connected NetworkabstractThanks to the great success of the deep learning, deep hashing for large-scale multimedia retrieval has made significant progress recently. However, most existing deep hashing algorithms suffer from slow convergence due to the gradient vanishing problem, caused by deep network structures and saturated activation functions. Moreover, a single convolution layer is often followed by down-sampling such as max pooling, resulting in local information loss that might affect the overall system robustness and performance. In this work, we propose a novel deep supervised hashing, Deep Loss Driven Multi-Scale Hashing (DLDMSH), which learns the high-quality approximate binary codes through an end-to-end network and improves the representative capacity of hash codes for large-scale image retrieval. Specifically, we design a Loss Driven Multi-Scale (LDMS) feature which is aggregated from convolutional feature maps. Moreover, a Pyramid Connected Convolutional Neural Network (PCNet) architecture is devised to generate LDMS feature, which inputs pairs of images during the training and outputs an image to approximate discrete values. In particular, 1 × 1 convolution kernels are applied to make a linear combination of features for realizing feature reduction, and the reduced features are fused in the fusion layer. This effectively improves the performance of deep features. A novel loss function preserving semantic information is integrated into an end-to-end learning scheme, which enhances the representative capacity of binary codes. Extensive experiments over four benchmark datasets show that DLDMSH significantly outperforms several other state-of-the-art hashing methods. Lingchen Gu, Jiande Sun 0001 |
IEEE Trans. Multim. | 4 |
| 2021 | Eye-based Recognition for User Identification on Mobile DevicesabstractUser identification is becoming more and more important for Apps on mobile devices. However, the identity recognition based on eyes, e.g., iris recognition, is rarely used on mobile devices comparing with those based on face and fingerprint due to its extra cost in hardware and complicated operations during recognition. In this article, an eye-based recognition method is designed for identity recognition on mobile devices, which can be implemented just like face recognition. In the proposed method, the eye feature is composed of the static and dynamic features, where the periocular feature extracted by deep neural network from the eye image is used as the static feature, and the motion feature of saccadic velocity is selected as the dynamic feature. The eye images can be captured by the normal camera on mobile devices just like faces, and dynamic features can provide living information to increase the difficulty of forgery. The GazeCapture dataset is used to test the proposed method, because the eye images in this dataset are captured by mobile devices during daily use. The recognition accuracy of the proposed method on the GazeCapture dataset can reach 96.87% only based on the periocular feature and can be enhanced to 97.99% when it is fused with the saccadic feature. The experiment results show that the performance of the proposed method can be comparative to that of iris recognition methods. It demonstrates that the proposed method is a practical reference for the eye-based identity recognition, and the proposed method provides one more biometric choice for mobile devices. Huiru Shao, Jing Li 0046, Jia Zhang 0028, Hui Yu 0001, Jiande Sun 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2020 | Multi-Feature Discrete Collaborative Filtering for Fast Cold-Start RecommendationabstractHashing is an effective technique to address the large-scale recommendation problem, due to its high computation and storage efficiency on calculating the user preferences on items. However, existing hashing-based recommendation methods still suffer from two important problems: 1) Their recommendation process mainly relies on the user-item interactions and single specific content feature. When the interaction history or the content feature is unavailable (the cold-start problem), their performance will be seriously deteriorated. 2) Existing methods learn the hash codes with relaxed optimization or adopt discrete coordinate descent to directly solve binary hash codes, which results in significant quantization loss or consumes considerable computation time. In this paper, we propose a fast cold-start recommendation method, called Multi-Feature Discrete Collaborative Filtering (MFDCF), to solve these problems. Specifically, a low-rank self-weighted multi-feature fusion module is designed to adaptively project the multiple content features into binary yet informative hash codes by fully exploiting their complementarity. Additionally, we develop a fast discrete optimization algorithm to directly compute the binary hash codes with simple operations. Experiments on two public recommendation datasets demonstrate that MFDCF outperforms the state-of-the-arts on various aspects. Yang Xu 0025, Lei Zhu 0002, Zhiyong Cheng 0001, Jingjing Li 0001, Jiande Sun 0001 |
AAAI | 5 |
| 2020 | G-NOMA for Energy Efficient C-RANabstractAs a green wireless network access framework, cloud radio access network (C-RAN) has the advantages of reducing energy consumption and improving spectral efficiency. In this paper, we propose to maximize the energy efficiency (EE) in the downlink of the C-RAN. We propose a design of beamforming and rate allocation based on generalized nonorthogonal multiple access (G-NOMA) by using the central cooperation and local coordination across the base transceiver stations (BTSs) in the C-RAN downlink. Our aim is to maximize the EE with the constraints of each BTS power consumption and the minimum quality of service (QoS) requirements, which we propose to solve by an efficient and low complexity successive convex approximation (SCA) algorithm. Simulation results demonstrate that the G-NOMA scheme can effectively achieve the best energy efficiency. Jia Zhang 0028, Lili Meng, Jiande Sun 0001 |
INDIN | 5 |
| 2020 | Color image watermarking based on orientation diversity and color complexity
Jun Wang 0061, Wenbo Wan, Xiao Xiao Li, Jiande Sun 0001, Huaxiang Zhang 0001 |
Expert Syst. Appl. | 4 |
| 2020 | Image super-resolution based on two-level residual learning CNN
Min Gao 0001, Xianhua Han, Jing Li 0046, Huaxiang Zhang 0001, Jiande Sun 0001 |
Multim. Tools Appl. | 6 |
| 2020 | Two-class 3D-CNN classifiers combination for video copy detection
Jing Li 0046, Huaxiang Zhang 0001, Wenbo Wan, Jiande Sun 0001 |
Multim. Tools Appl. | 4 |
| 2020 | Hash length: a neglected element
Haifeng Qi, Jing Li 0046, Qiang Wu 0009, Wenbo Wan, Jiande Sun 0001 |
Multim. Tools Appl. | 5 |
| 2020 | Saccadic trajectory-based identity authentication
Huiru Shao, Jing Li 0046, Wenbo Wan, Huaxiang Zhang 0001, Jiande Sun 0001 |
Multim. Tools Appl. | 5 |
| 2020 | Hybrid JND model-guided watermarking method for screen content images
Wenbo Wan, Jun Wang 0061, Jing Li 0046, Jiande Sun 0001, Huaxiang Zhang 0001 |
Multim. Tools Appl. | 4 |
| 2020 | Cross-modal dual subspace learning with adversarial network
Fei Shang, Huaxiang Zhang 0001, Jiande Sun 0001, Liqiang Nie, Li Liu 0031 |
Neural Networks | 3 |
| 2020 | Pattern complexity-based JND estimation for quantization watermarking
Wenbo Wan, Jun Wang 0061, Jing Li 0046, Lili Meng, Jiande Sun 0001, Huaxiang Zhang 0001 |
Pattern Recognit. Lett. | 5 |
| 2020 | Multi-class joint subspace learning for cross-modal retrieval
En Yu, Jing Li 0046, Li Wang 0148, Jia Zhang 0028, Wenbo Wan, Jiande Sun 0001 |
Pattern Recognit. Lett. | 6 |
| 2020 | Energy and Spectral Efficiency Tradeoff via Rate Splitting and Common Beamforming Coordination in Multicell NetworksabstractRate splitting (RS) has the potential to significantly enhance both energy efficiency (EE) and spectral efficiency (SE) of wireless networks. In this paper, we propose joint design of the beamforming and rate allocation to maximize both EE and SE of a downlink multicell multiple-input single-output (MISO) system with rate splitting and common beamforming coordination (RS-CBC). This design problem is formulated as a non-convex quadratically-constrained multi-objective optimization problem (MOOP). By investigating the quasi-concavity relationship between EE and SE, the formulated MOOP is transformed into a single-objective optimization problem (SOOP) to offer a tradeoff between EE and SE by maximizing EE in any achievable SE region. A series of transformations are then applied to make the SOOP tractable, after which an efficient iterative algorithm based on successive convex approximation (SCA) is proposed to solve the problem. Simulation results demonstrate the effectiveness of the proposed algorithm and unveil interesting tradeoffs between EE and SE under different parameter settings. Jia Zhang 0028, Yong Zhou 0006, Jiande Sun 0001, Naofal Al-Dhahir |
IEEE Trans. Commun. | 5 |
| 2020 | Hyperspectral Reconstruction with Redundant Camera Spectral Sensitivity FunctionsabstractHigh-resolution hyperspectral (HS) reconstruction has recently achieved significantly progress, among which the method based on the fusion of the RGB and HS images of the same scene can greatly improve the reconstruction performance compared with those based on the individually spectral or spatial enhancement. It is well known that the HS image is obtained only via the costly hypersoectral sensor, whereas the RGB images can be provided by low-price RGB cameras and the spectral sensitivity (SS) functions of RGB cameras are usually different. Thus, this study proposes a HS reconstruction, which fuses merely two RGB images with redundant spectral responses. In this work, we design a new RGB camera via shifting the SS of an existed RGB camera, which can provide similar strength of spectral response with different spectral centers of SS, and fuse the new achieved color image with an existed RGB image by a deep ResNet. Experiments validate that fusion of two existed RGB images can provide impressive HS reconstruction performance and further improvement can be achieved by integrating the color image of the simulated SS with the RGB image. Xianhua Han, Yinqiang Zheng, Jiande Sun 0001, Yen-Wei Chen 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2019 | Fusion-Supervised Deep Cross-Modal HashingabstractDeep hashing has recently received attention in cross-modal retrieval for its impressive advantages. However, existing hashing methods for cross-modal retrieval cannot fully capture the heterogeneous multi-modal correlation and exploit the semantic information. In this paper, we propose a novel Fusion-supervised Deep Cross-modal Hashing (FDCH) approach. Firstly, FDCH learns unified binary codes through a fusion hash network with paired samples as input, which effectively enhances the modeling of the correlation of heterogeneous multi-modal data. Then, these high-quality unified hash codes further supervise the training of the modality-specific hash networks for encoding out-of-sample queries. Meanwhile, both pair-wise similarity information and classification information are embedded in the hash networks under one stream framework, which simultaneously preserves cross-modal similarity and keeps semantic consistency. Experimental results on two benchmark datasets demonstrate the state-of-the-art performance of FDCH. Li Wang 0148, Lei Zhu 0002, En Yu, Jiande Sun 0001, Huaxiang Zhang 0001 |
ICME | 4 |
| 2019 | Regularized Group Sparse Discriminant Analysis for P300-Based Brain-Computer InterfaceabstractEvent-related potentials (ERPs) especially P300 are popular effective features for brain-computer interface (BCI) systems based on electroencephalography (EEG). Traditional ERP-based BCI systems may perform poorly for small training samples, i.e. the undersampling problem. In this study, the ERP classification problem was investigated, in particular, the ERP classification in the high-dimensional setting with the number of features larger than the number of samples was studied. A flexible group sparse discriminative analysis algorithm based on Moreau-Yosida regularization was proposed for alleviating the undersampling problem. An optimization problem with the group sparse criterion was presented, and the optimal solution was proposed by using the regularized optimal scoring method. During the alternating iteration procedure, the feature selection and classification were performed simultaneously. Two P300-based BCI datasets were used to evaluate our proposed new method and compare it with existing standard methods. The experimental results indicated that the features extracted via our proposed method are efficient and provide an overall better P300 classification accuracy compared with several state-of-the-art methods. Qiang Wu 0009, Yu Zhang 0009, Jiande Sun 0001, Andrzej Cichocki, Feng Gao 0008 |
Int. J. Neural Syst. | 4 |
| 2019 | Adversarial cross-modal retrieval based on dictionary learning
Fei Shang, Huaxiang Zhang 0001, Lei Zhu 0002, Jiande Sun 0001 |
Neurocomputing | 4 |
| 2019 | Weighted locality collaborative representation based on sparse subspace
Huaxiang Zhang 0001, Lei Zhu 0002, Wenbo Wan, Zhenhua Wang 0004, Qiang Wang 0015, Peilian Guo, Jiande Sun 0001 |
J. Vis. Commun. Image Represent. | 9 |
| 2019 | Semantic consistency cross-modal dictionary learning with rank constraint
Fei Shang, Huaxiang Zhang 0001, Jiande Sun 0001, Li Liu 0031 |
J. Vis. Commun. Image Represent. | 3 |
| 2019 | Supervised graph regularization based cross media retrieval with intra and inter-class correlation
Meijia Zhang, Huaxiang Zhang 0001, Junzheng Li, Li Wang 0148, Yixian Fang, Jiande Sun 0001 |
J. Vis. Commun. Image Represent. | 6 |
| 2019 | Coupled feature selection based semi-supervised modality-dependent cross-modal retrieval
En Yu, Jiande Sun 0001, Li Wang 0148, Wenbo Wan, Huaxiang Zhang 0001 |
Multim. Tools Appl. | 2 |
| 2019 | A novel coverless information hiding method based on the average pixel value of the sub-images
Liming Zou, Jiande Sun 0001, Min Gao 0001, Wenbo Wan, Brij B. Gupta |
Multim. Tools Appl. | 2 |
| 2019 | Zero-shot event detection via event-adaptive concept relevance mining
Zhihui Li 0001, Lina Yao 0001, Xiaojun Chang, Kun Zhan, Jiande Sun 0001, Huaxiang Zhang 0001 |
Pattern Recognit. | 5 |
| 2019 | Adaptive Semi-Supervised Feature Selection for Cross-Modal RetrievalabstractIn order to exploit the abundant potential information of the unlabeled data and contribute to analyzing the correlation among heterogeneous data, we propose the semi-supervised model named adaptive semi-supervised feature selection for cross-modal retrieval. First, we utilize the semantic regression to strengthen the neighboring relationship between the data with the same semantic. And the correlation between heterogeneous data can be optimized via keeping the pairwise closeness when learning the common latent space. Second, we adopt the graph-based constraint to predict accurate labels for unlabeled data, and it can also keep the geometric structure consistency between the label space and the feature space of heterogeneous data in the common latent space. Finally, an efficient joint optimization algorithm is proposed to update the mapping matrices and the label matrix for unlabeled data simultaneously and iteratively. It makes samples from different classes to be far apart, while the samples from same class lie as close as possible. Meanwhile, the l2,1-norm constraint is used for feature selection and outlier reduction when the mapping matrices are learned. In addition, we propose learning different mapping matrices corresponding to different sub-tasks to emphasize the semantic and structural information of query data. Experiment results on three datasets demonstrate that our method performs better than the state-of-the-art methods. En Yu, Jiande Sun 0001, Jing Li 0046, Xiaojun Chang, Xianhua Han, Alex Hauptmann 0001 |
IEEE Trans. Multim. | 2 |
| 2018 | Joint feature selection and graph regularization for modality-dependent cross-modal retrieval
Li Wang 0148, Lei Zhu 0002, Li Liu 0031, Jiande Sun 0001, Huaxiang Zhang 0001 |
J. Vis. Commun. Image Represent. | 5 |
| 2018 | Semi-supervised modality-dependent cross-media retrieval
Jiande Sun 0001, Peiyong Duan, Lili Meng, Yanyan Tan, Wenbo Wan, Hongchen Wu, Bin Zhang 0050, Huaxiang Zhang 0001 |
Multim. Tools Appl. | 2 |
| 2018 | View-invariant gait recognition based on kinect skeleton feature
Jiande Sun 0001, Jing Li 0046, Wenbo Wan, De Cheng, Huaxiang Zhang 0001 |
Multim. Tools Appl. | 1 |
| 2018 | Joint graph regularization based modality-dependent cross-media retrieval
Jihong Yan, Huaxiang Zhang 0001, Jiande Sun 0001, Qiang Wang 0015, Peilian Guo, Lili Meng, Wenbo Wan |
Multim. Tools Appl. | 3 |
| 2018 | Cross-media retrieval with collective deep semantic learning
Bin Zhang 0050, Lei Zhu 0002, Jiande Sun 0001, Huaxiang Zhang 0001 |
Multim. Tools Appl. | 3 |
| 2018 | Discriminative correlation hashing for supervised cross-modal retrieval
Xu Lu 0004, Huaxiang Zhang 0001, Jiande Sun 0001, Zhenhua Wang 0004, Peilian Guo, Wenbo Wan |
Signal Process. Image Commun. | 3 |
| 2017 | Appearance-based gaze block estimation via CNN classificationabstractAppearance-based gaze estimation methods have received increasing attention in the field of human-computer interaction (HCI). These methods tried to estimate the accurate gaze point via Convolutional Neural Network (CNN) model, but the estimated accuracy can't reach the requirement of gaze-based HCI when the regression model is used in the output layer of CNN. Given the popularity of button-touch-based interaction, we propose an appearance-based gaze block estimation method, which aims to estimate the gaze block, not the gaze point. In the proposed method, we relax the estimation from point to block, so that the gaze block can be estimated by CNN-based classification instead of the previous regression model. We divide the screen into square blocks to imitate the button-touch interface, and build an eye-image dataset, which contains the eye images labelled by their corresponding gaze blocks on the screen. We train the CNN model according to this dataset to estimate the gaze block by classifying the eye images. The experiments on 6- and 54-block classifications demonstrate that the proposed method has high accuracy in gaze block estimation without any calibration, and it is promising in button-touch-based interaction. Xuemei Wu, Jing Li 0046, Qiang Wu 0009, Jiande Sun 0001 |
MMSP | 4 |
| 2017 | Block-Wise Gaze Estimation Based on Binocular Images
Xuemei Wu, Jing Li 0046, Qiang Wu 0009, Jiande Sun 0001 |
PSIVT | 4 |
| 2017 | A two-stage learning approach to face recognition
Huaxiang Zhang 0001, Jiande Sun 0001, Wenbo Wan |
J. Vis. Commun. Image Represent. | 3 |
| 2017 | A Weight-Adaptive Laplacian Embedding for Graph-Based ClusteringabstractGraph-based clustering methods perform clustering on a fixed input data graph. Thus such clustering results are sensitive to the particular graph construction. If this initial construction is of low quality, the resulting clustering may also be of low quality. We address this drawback by allowing the data graph itself to be adaptively adjusted in the clustering procedure. In particular, our proposed weight adaptive Laplacian (WAL) method learns a new data similarity matrix that can adaptively adjust the initial graph according to the similarity weight in the input data graph. We develop three versions of these methods based on the L2-norm, fuzzy entropy regularizer, and another exponential-based weight strategy, that yield three new graph-based clustering objectives. We derive optimization algorithms to solve these objectives. Experimental results on synthetic data sets and real-world benchmark data sets exhibit the effectiveness of these new graph-based clustering methods. De Cheng, Feiping Nie 0001, Jiande Sun 0001, Yihong Gong |
Neural Comput. | 3 |
| 2017 | Comprehensive Feature-Based Robust Video Fingerprinting Using Tensor ModelabstractContent-based near-duplicate video detection (NDVD) is essential for effective search and retrieval, and robust video fingerprinting is a good solution for NDVD. Most existing video fingerprinting methods use a single feature or concatenate different features to generate video fingerprints, and show good performance under single-mode modifications such as noise addition and blurring. However, when they suffer combined modifications, the performance is degraded to a certain extent because such features cannot characterize the video content completely. By contrast, the assistance and consensus among different features can improve the performance of video fingerprinting. Therefore, in the present study, we mine the assistance and consensus among different features based on a tensor model, and we present a new comprehensive feature to fully use them in the proposed video fingerprinting framework. We also analyze what the comprehensive feature really is for representing the original video. In this framework, the video is initially set as a high-order tensor that consists of different features, and the video tensor is decomposed via the Tucker model with a solution that determines the number of components. Subsequently, the comprehensive feature is generated by the low-order tensor obtained from tensor decomposition. Finally, the video fingerprint is computed using this feature. A matching strategy used for narrowing the search is also proposed based on the core tensor. The robust video fingerprinting framework is resistant not only to single-mode modifications but also to their combination. Xiushan Nie, Yilong Yin, Jiande Sun 0001, Chaoran Cui |
IEEE Trans. Multim. | 3 |
| 2017 | Discrete Multimodal Hashing With Canonical Views for Robust Mobile Landmark SearchabstractMobile landmark search (MLS) recently receives increasing attention for its great practical values. However, it still remains unsolved due to two important challenges. One is high bandwidth consumption of query transmission, and the other is the huge visual variations of query images sent from mobile devices. In this paper, we propose a novel hashing scheme, named as canonical view based discrete multimodal hashing (CV-DMH), to handle these problems. First, a submodular function is designed to measure visual representativeness and redundancy of a view set. With it, canonical views, which capture key visual appearances of landmark with limited redundancy, are efficiently discovered with an iterative mining strategy. Second, multimodal sparse coding is applied to transform visual features from multiple modalities into an intermediate representation. It can robustly and adaptively characterize visual contents of varied landmark images with certain canonical views. Finally, compact binary codes are learned on intermediate representation within a tailored discrete binary embedding model which preserves visual relations of images measured with canonical views and removes the involved noises. In this part, we develop a new augmented Lagrangian multiplier (ALM) based optimization method to directly solve the discrete binary codes. We can not only explicitly deal with the discrete constraint, but also consider the bit-uncorrelated constraint and balance constraint together. The proposed solution can desirably avoid accumulated quantization errors in conventional optimization method which simply adopts two-step "relaxing+rounding'' framework. Experiments on real world landmark datasets demonstrate the superior performance of CV-DMH over several state-of-the-art methods. Lei Zhu 0002, Zi Huang, Xiaobai Liu, Xiangnan He 0001, Jiande Sun 0001, Xiaofang Zhou 0001 |
IEEE Trans. Multim. | 5 |
| 2016 | Depth propagation in 2D-to-3D conversion based on frame clusteringabstractDepth propagation plays an important role to improve the quality of converted video in 2D-to-3D conversion. During depth propagation, the propagation path is usually along the time. In this paper, a new depth propagation algorithm is proposed to improve the quality of depth propagation by changing the propagation path according to frame clustering. The frames are clustered via K-means clustering based on HSV color histogram. The frame at the center of each cluster is selected as the keyframe of the cluster and the other frames are taken as the non-keyframes attaching to the keyframe. The depth information of keyframe is propagated to the non-keyframes within the same cluster directly. The performance of the proposed depth propagation algorithm is evaluated by MSE of the propagated depth map comparing with several existing algorithms. The comparison results show that the depth propagation errors can be reduced a lot by the proposed clustering based propagation. Zhenxiao Fu, Jiande Sun 0001, Qiang Wu 0009, Jing Li 0046 |
ICASSP | 2 |
| 2016 | Gait recognition based on 3D skeleton joints captured by kinectabstract2D-video-based gait recognition techniques have been studied for decades, but there are still many challenges, one of which is the robustness against the variation of view angle. In this paper, the second generation Kinect (Kinect V2) is used as a tool to establish a 3D-skeleton-based gait database, which includes both 3D information of the skeleton joints and the corresponding 2D silhouette images captured by Kinect V2. Based on this dataset, a human walking model is built, and the static and dynamic features are extracted, which are verified to be view-invariant for gait recognition. Referring to the walking model, the gait recognition abilities for the static and dynamic features are investigated respectively and a gait recognition scheme based on the matching-level-fusion of the static and dynamic features is proposed, in which the recognition is achieved by the nearest neighbor classification method. Experiments show that the proposed scheme has robust recognition performance against the variation of view angle. Jiande Sun 0001, Jing Li 0046, Dong Zhao 0017 |
ICIP | 2 |
| 2016 | Frame rate up-conversion based on depth guided extended block matching for 3D videoabstractA frame rate up-conversion (FRUC) method for 3D video (3DV) is presented in this paper. Inspired by the fact that moving foreground objects draw more attention of the viewers, in our method depth guided extended block matching (DGE-BM) is adopted to maintain the completeness of the foreground object. We first obtain the motion vector field (MVF) of the interpolated frame via block-based bi-directional motion estimation (ME). And the blocks of the interpolated frame are classified according to the depth information. Then, the boundary blocks are divided into sub-blocks, whose motion vectors (MVs) are estimated using DGE-BM. Finally, the refined MVF is applied to do motion compensation. Experimental results show that the frame interpolation quality of the proposed method achieves significantly improvement comparing with existing algorithms. Jiande Sun 0001, Zhiquan Feng, Haokui Tang, Lingyin Wang |
VCIP | 3 |
| 2016 | Spherical torus-based video hashing for near-duplicate video detection
Xiushan Nie, Yane Chai, Jiande Sun 0001, Yilong Yin |
Sci. China Inf. Sci. | 4 |
| 2016 | Video hashing based on appearance and attention features fusion via DBN
Jiande Sun 0001, Xiaocui Liu, Wenbo Wan, Jing Li 0046, Dong Zhao 0017, Huaxiang Zhang 0001 |
Neurocomputing | 1 |
| 2016 | A novel image retrieval method based on multi-trend structure descriptor
Huaxiang Zhang 0001, Jiande Sun 0001 |
J. Vis. Commun. Image Represent. | 3 |
| 2016 | Improved logarithmic spread transform dither modulation using a robust perceptual model
Wenbo Wan, Jiande Sun 0001 |
Multim. Tools Appl. | 3 |
| 2016 | Adaptive Convolutional Neural Network and Its Application in Face Recognition
Dong Zhao 0017, Jiande Sun 0001, Guofeng Zou |
Neural Process. Lett. | 3 |
| 2015 | Improved Spread Transform Dither Modulation Using Luminance-Based JND Model
Wenhua Tang, Wenbo Wan, Jiande Sun 0001 |
ICIG (2) | 4 |
| 2015 | Investigation on the Influence of Visual Attention on Image Memorability
Wulin Wang, Jiande Sun 0001, Jing Li 0046, Qiang Wu 0009 |
ICIG (3) | 2 |
| 2015 | A calibration simplified method for gaze interaction based on using experienceabstractThe step of calibration prevents the gaze interaction from interacting naturally, which usually needs five or more calibration points to obtain the user's calibration information. In this paper, a calibration simplified method is proposed, which is based on user experience and consists of the user experience accumulation stage and the calibration simplified stage. In the user experience accumulation stage, the calibration is still needed as usual before each gaze interaction and the calibration parameters and the position of the user are stored as the calibration experience, where the calibration parameters include the actual coordinates of the calibration points and their estimation errors. In the calibration simplified stage, the user needs to fixate on only one calibration point for calibration before the calibration information can be estimated according to the user experience stored in the first stage. Even the calibration information can be estimated according to only the position of the user and the stored calibration parameters, which means the user can interact with computer via gaze directly without calibration. The simulations show that the gaze estimation accuracy of the proposed calibration simplified method can be 1.4047° with one calibration point and 1.7489° without calibration point, which are much higher than that without the user experience. Cong Niu, Jiande Sun 0001, Jing Li 0046 |
MMSP | 2 |
| 2014 | Structural similarity-based video fingerprinting for video copy detectionabstractThe authors propose a video fingerprinting method based on structural similarity and a graph model. Structural similarity‐based fingerprint generation and double‐layer matching are the two main contributions. The video is mapped to a graph with frames as its vertices, and structural similarity is proposed to compute the weights of the edges. Then, the fingerprint consisting of a match tag (a coarse fingerprint) and a fine fingerprint is generated by this graph. The match tag is generated by an independent set of this graph, and the fine fingerprint is generated by the weight matrix of the graph based on the two‐block‐dimensional discrete cosine transform. During the matching, the video can be matched at the first‐layer using the match tag to obtain a candidate set, whereas the second‐layer matching is performed in this candidate set using the fine fingerprint to find a final match. The proposed video fingerprinting method is shown to be resistant to geometric attacks on frames and impairment of transmission channels. Xiushan Nie, Wenjun Zeng 0001, Jiande Sun 0001 |
IET Image Process. | 4 |
| 2014 | Video fingerprinting based on graph model
Xiushan Nie, Jiande Sun 0001, Zhihui Xing, Xiaocui Liu |
Multim. Tools Appl. | 2 |
| 2014 | Depth-Assisted Frame Rate Up-Conversion for Stereoscopic VideoabstractIn this letter, we propose a depth-assisted frame rate up-conversion (DA-FRUC) scheme which finds applications in 3D video processing. By considering the depth cue in video plus depth representation, we categorize the blocks of the interpolated frame as depth-continuous and depth-discontinuous groups. The motion vector (MV) outliers of the depth continuous blocks are then detected and corrected by layer-constrained MV refinement method. Moreover, a depth-based adaptive interpolation and block segmentation method is proposed to deal with the disocclusion and occlusion at the boundary of the foreground area. Experimental results show that the proposed method achieves better subjective and objective performance for the interpolated color frames. Additionally, the interpolated depth frames obtained by the proposed method is more accurate, which benefit the video quality in view synthesis. Jiande Sun 0001, Yeejin Lee, Truong Q. Nguyen |
IEEE Signal Process. Lett. | 3 |
| 2014 | A Novel Distortion Model and Lagrangian Multiplier for Depth Maps CodingabstractIn three-dimensional videos (3-DV) coding systems, depth maps are not used for viewing but for rendering virtual views. Therefore, the traditional rate distortion criterion (including distortion criterion, and Lagrangian multiplier) is not suitable for depth map coding. In order to design an effective rate distortion criterion for depth maps, the relationship between the distortion of synthesized virtual view and the coding error of depth maps is analyzed in detail. Through the analysis, a polynomial model revealing the relationship between the coding error of depth maps and the distortion of synthesized virtual view is derived. Model parameters are estimated by utilizing camera parameters and features of the texture video corresponding to the depth map. Based on the model, a virtual view-based Lagrangian multiplier for depth map coding is also proposed. Experimental results demonstrated the accuracy of the model. The squared correlation coefficients between the actual distortion of virtual view and the estimated distortion are all larger than 0.98 for all tested sequences. When incorporating the proposed model and Lagrangian multiplier into the mode decision procedure of joint model version 18.5 (JM18.5) of H.264/AVC, a maximum 0.470 dB BD PSNR and an average 0.251 dB BD PSNR can be achieved. Hui Yuan 0001, Sam Kwong, Jiande Sun 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2013 | Logarithmic Spread-Transform Dither Modulation watermarking Based on Perceptual ModelabstractLogarithmic Quantization Index Modulation (LQIM) is an important extension of the original quantization-based watermarking method. However, it is well known that it is sensitive to valumetric scaling attack and easy to result in sign error after quantization and attacks. For that, in this paper, we propose a new method, namely Logarithmic Spread-Transform Dither Modulation Based on Perceptual Model (LSTDM-WM). In this regard the host signal is first projected onto a random vector and transformed using a novel Logarithmic Quantization function. Then the transformed signal is quantized regarding the watermark data and the watermarked signal is obtained by applying inverse transform to the quantized signal. The perceptual model is further exploited to adjust the quantization step adaptively for watermark embedding. Experimental results indicate that our proposed scheme overcomes two challenges cited above and has superior performance in comparison with conventional LQIM and former proposed schemes of STDM. Wenbo Wan, Jiande Sun 0001, Xiushan Nie |
ICIP | 3 |
| 2013 | Unequally Weighted Video Hashing for Copy Detection
Jiande Sun 0001, Hui Yuan 0001, Xiaocui Liu |
MMM (1) | 1 |
| 2013 | Robust video hashing based on representative-dispersive frames
Xiushan Nie, Jiande Sun 0001, Lianqi Wang |
Sci. China Inf. Sci. | 3 |
| 2013 | Visual Attention Based Temporally Weighting Method for Video HashingabstractThe video hash derived from the temporally representative frame (TRF) has attracted increasing interests recently. A temporally visual weighting (TVW) method based on visual attention is proposed for the generation of TRF in this paper. In the proposed TVW method, the visual attention regions of each frame are obtained by combining the dynamic and static attention models. The temporal weight for each frame is defined as the strength of temporal variation of visual attention regions and the TRF of a video segment can be generated by accumulating the frames by the proposed TVW method. The advantage of the TVW method is proved by the comparison experiments. The video hashes used for comparison are derived from the TRFs, which are generated based on the proposed TVW method and other existing weighting methods respectively. The experimental results show that the TVW method is helpful to enhance the robustness and discrimination of video hash. Xiaocui Liu, Jiande Sun 0001 |
IEEE Signal Process. Lett. | 2 |
| 2012 | A visual saliency based video hashing algorithmabstractA novel video hashing algorithm is proposed, which takes account of visual saliency during hash generation. In the proposed algorithm, the video hash is fused by two hashes, which are spatio-temporal hash (ST-Hash) and visual hash (V-Hash). The ST-Hash is generated based on the ordinal feature, which is formed according to the intensity difference between adjacent blocks of the temporally informative representation image (TIRI). At the same time, the representative saliency map (RSM) is constructed by the visual saliency maps in video segments. The V-Hash is formed according to the intensity difference between adjacent blocks of the RSM, and used to modulate the ST-Hash to form the final video hash. Experiments on different kinds of videos with different kinds of attacks verify that the proposed algorithm has better performance on robustness and discrimination. Jiande Sun 0001, Xiushan Nie |
ICIP | 2 |
| 2012 | Video Hashing Algorithm With Weighted Matching Based on Visual SaliencyabstractIn this letter, a novel video hashing algorithm is proposed, in which the weighted hash matching is defined in video hashing for the first time. In the proposed algorithm, the video hash is generated based on the ordinal feature derived from the temporally informative representation image (TIRI). At the same time the representative saliency map (RSM) is constructed by the visual saliency maps in video segments, and it generates the hash weights for hash matching. During hash matching, the traditional bit error rate (BER) is weighted with hash weights to form the weighted error rate (WER). WER is used to measure the similarity between different hashes. Experiments on different kinds of videos with different kinds of attacks verify the robustness and discrimination of the proposed algorithm. Jiande Sun 0001, Xiushan Nie |
IEEE Signal Process. Lett. | 1 |
| 2012 | Affine Model Based Motion Compensation Prediction for ZoomabstractZoom motion is classified into two categories, i.e., global zoom motion and local zoom motion. A simple affine motion model with four parameters is utilized to describe zoom motion efficiently based on the analyses of camera imaging principles. Based on the motion model, a basic candidate motion vector (BCMV) of a block could be derived when model parameters are confirmed. Then a set of candidate motion vectors (CMVs) could be obtained by modifying the BCMV. Thereafter, template matching is used to choose the optimal CMV (OCMV). Finally, the block is coded with the optimal CMV as an independent mode, and a rate distortion (RD) criterion is used to determine whether to use the mode or not. Experimental results demonstrate that by implementing the proposed method into Key Technology Area test platform version 2.6r1 (KTA2.6r1), a maximum -21.99% and average -8.72% bit rate savings can be achieved for videos involving zoom motion, while maintaining the same quality (evaluated by PSNR) of reconstructed videos when IPPPP coding structure is used. When Hierarchical B coding structure is employed, the maximum and average bit rate savings are - 10.48% and - 5.575% when the qualities of reconstructed videos remain unchanged. Besides, for videos involving camera rotation, translation, etc., an average -2.04% bit rate savings could also be achieved; while for videos containing common motions, an average -1.19% bit rate saving could be achieved at the same quality of reconstructed videos. Hui Yuan 0001, Jiande Sun 0001, Hechao Liu |
IEEE Trans. Multim. | 3 |
| 2011 | Step-projection-based spread transform dither modulationabstractQuantisation index modulation (QIM) is an important class of watermarking methods, which has been widely used in blind watermarking applications. It is well known that spread transform dither modulation (STDM), as an extension of QIM, has good performance in robustness against random noise and re-quantisation. However, the quantisation step-sizes used in STDM are random numbers not taking features of the image into account. The authors present a step projection-based approach to incorporate the perceptual model with STDM framework. Four implementations of the proposed algorithm are further presented according to different modified versions of the perceptual model. Experimental results indicate that the step projection-based approach can incorporate the perceptual model with STDM framework in a better way, thereby providing a significant improvement in image fidelity. Compared with the former proposed modified schemes of STDM, the author's best performed implementation provides powerful resistance against common attacks, especially in robustness against Gauss noise, salt and pepper noise and JPEG compression. Xinchao Li, Jiande Sun 0001, Wei Liu 0031 |
IET Inf. Secur. | 3 |
| 2011 | Robust Video Hashing Based on Double-Layer EmbeddingabstractA robust video hashing scheme for video content identification and authentication is proposed, which is called Double-Layer Embedding scheme. Intra-cluster Locally Linear Embedding (LLE) and inter-cluster Multi-Dimensional Scaling (MDS) are used in the scheme. Some dispersive frames of the video are first selected through graph model, and the video is partitioned into clusters based on the dispersive frames and the K-Nearest Neighbor method during the hashing. Then, the intra-cluster LLE and inter-cluster MDS are used to generate local and global hash sequences which can inherently describe the corresponding video. Experimental results show that the video hashing is resistant to geometric attacks of frames and channel impairments of transmission. Xiushan Nie, Jiande Sun 0001, Wei Liu 0031 |
IEEE Signal Process. Lett. | 3 |
| 2010 | Robust video hashing for identification based on MDSabstractVideo identification is extremely important in video browsing, database search and security. In this paper, we present a video hashing based on MDS (Multi-Dimensional Scaling) which is able to work under variable video transmission impairments and resistant to signal processing. In this method, each frame of the video is divided into blocks and compute its low and middle frequency DCT coefficients of luminance component as a disparities measurement for MDS. Then the video is mapped to two-dimensional space using MDS, and generate a robust hashing as a video signature utilizing the distances between two points mapping from frames. It found that this video hashing is resistant frame geometric attacks (rotation, shift), random noises, lossy compression and other video transmission impairments. It can be instrumental in building database search, video copy detection and watermarking applications for video. Xiushan Nie, Jiande Sun 0001 |
ICASSP | 3 |
| 2009 | Independent triangles set-based watermarking for 3D modelsabstractAn independent triangles set-based watermarking algorithm using singular value decomposition (SVD) is presented in this paper. In this algorithm, a 3D model is taken as a model that consists of triangular meshes, and we select the triangles called independent triangles set (ITS) to generate a Hankel matrix and apply SVD on it, then watermarks are embedded by modulating singular values. The watermarks that embedded by this method are resistant to similarity transformations and random noises. We prove that the proposed algorithm using ITS as embedding primitive is more robust than those using vertex coordinates discussed in other papers, and the experimental results corroborate it as well. In addition, when detect the watermarks, we only use some keys instead of the original model, so it is a blind watermarking algorithm too. Xiushan Nie, Jiande Sun 0001, Jianping Qiao |
ICME | 3 |
| 2009 | A motion location based video watermarking scheme using ICA to extract dynamic frames
Zhaowan Sun, Jiande Sun 0001, Xinghua Sun |
Neural Comput. Appl. | 3 |
| 2007 | A New Two-Stage Approach to Underdetermined Blind Source Separation using Sparse RepresentationabstractIn this paper we focus on the two-stage underdetermined blind source separation (BSS), which consists of the mixing matrix estimation stage, the first stage, and the source estimation stage, the second stage. In the first stage, both the mixing matrix and the number of sources are estimated by a new potential-function-based clustering method using a new potential function constructed by Laplacian-like window function. In the second stage, in order to overcome the disadvantage of 11-norm solution, a new sparse representation based on high-order statistics in transformed domain, which is called statistically sparse component analysis (SSCA), is proposed to recover the sources. Compared with the existing two-stage methods, the proposed approach can achieve higher reconstructed signal-to-noise ratios (SNRs). Jiande Sun 0001, Shuzhong Ba |
ICASSP (3) | 3 |
| 2007 | ICA Based Super-Resolution Face Hallucination and Recognition
Jiande Sun 0001, Xinghua Sun |
ISNN (2) | 3 |
| 2006 | A Novel Watermarking Method with Image Signature
Xiao-Li Niu, Jiande Sun 0001, Jianping Qiao |
ISNN (2) | 3 |
| 2006 | A 2DPCA-Based Video Watermarking Scheme for Resistance to Temporal Desynchronization
Jiande Sun 0001 |
ISNN (2) | 1 |
| 2006 | A blind video watermarking scheme based on ICA and shot segmentation
Jiande Sun 0001 |
Sci. China Ser. F Inf. Sci. | 1 |
| 2005 | A temporal desynchronization resilient video watermarking scheme based on independent component analysisabstractTemporal synchronization is a problem that cannot be ignored in most video watermarking applications. In this paper, a video watermarking scheme is proposed to resist temporal desynchronization through a content dependent way. In this scheme, the to-be-watermarked video is first segmented into shots. And each shot is analyzed by using independent component analysis (ICA) to get its independent feature frames (IFFs). And then the watermark is embedded into these feature frames. Simulations show that this scheme is robust not only to temporal desynchronization, e.g., frame swapping, frame dropping, fps conversion, etc, but to most of the common video processing/attacks, e.g. noising, filtering, rescaling, MPEG2-4 transcoding, and intra-video collusion. Jiande Sun 0001 |
ICIP (1) | 1 |
| 2005 | A Copy Attack Resilient Blind Watermarking Algorithm Based on Independent Feature Components
Huibo Hu, Jiande Sun 0001 |
ISNN (2) | 3 |
| 2005 | Edge Projection-Based Image Registration
Jiande Sun 0001 |
KES (2) | 3 |
| 2004 | Data Hiding in Independent Components of Video
Jiande Sun 0001, Huibo Hu |
ISNN (1) | 1 |