EDBT 2026 Demo / reviewers in the wild / expert
Cong Zou
dblp:230/8064
· DBLP profile ↗
20ranked-venue papers
9as first author
14since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 5 since 2021Computer networks · 6 · 4 first-author · 6 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-authorSecurity and privacy · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SegAuth: Semantic-Aware Multimotion Behavioral Biometric-Based Implicit AuthenticationabstractMulti-motion behavioral biometrics based implicit authentication leverages individual unique behavioral patterns for user authentication. However, traditional methods using fixed-sized time window segmentation often disrupt local temporal structures and overlook behavioral semantics. This paper investigates a semantic-aware segmentation-based implicit authentication approach, yet challenges persist in achieving semantically consistent segmentation, fixed-dimension representations for variable-length data, and robust modeling under large intra-class variance. Towards this end, we propose SegAuth, a novel semantic-aware multi-motion behavioral biometrics based implicit authentication system. Specifically, given the input raw multi-motion data, SegAuth first adopts a data-driven semantic-aware segmentation method to adaptively generate variable-sized segments, capturing fine-grained behavioral patterns for authentication. Next, SegAuth proposes a causal temporal convolutional network, which allows to learn the effective embedding of varied-sized multi-motion data segments. Finally, a multi-center deep one-class classifier–based authentication model is developed to capture behavioral representations from genuine user data characterized by high intra-class variance, allowing it to identify and authenticate behaviors that differ from normal patterns. Extensive experiments are conducted on a large-scale uncontrolled evaluation dataset. The experimental results demonstrate the state-of-the-art authentication performance of SegAuth. Zhihao Shen 0001, Chengmei Zhao, Xi Zhao 0001, Cong Zou, Jiakun Zhao |
IEEE Internet Things J. | 4 |
| 2026 | Enhancing Open-Set RFF Recognition With cGAN: Generating Multiple Unknown ClassesabstractOpen set recognition (OSR) in radio frequency fingerprint (RFF) is critical for securing Internet of Things (IoT) systems, where previously unseen or malicious devices may attempt unauthorized access. A widely used approach treats all unknown devices as a single additional class and assumes that they will produce low confidence scores during inference. However, due to the inherently subtle and highly similar RFF features across devices, this assumption often fails, leading to high false acceptance rates. To address this challenge, we propose a novel framework, called Multiple Unknown Classes Generation (MUCG), which replaces the single-class modeling of unknowns with a more expressive structure that simulates multiple distinct unknown classes. MUCG employs a conditional generative adversarial network (cGAN) guided by ideal signal priors to produce diverse and realistic unknown samples. Furthermore, we introduce a soft label perturbation (SLP) strategy that blends label semantics using Feature-wise Linear Modulation (FiLM), encouraging the generator to embed richer feature variations. Experiments on three public IoT datasets demonstrate that MUCG consistently outperforms state-of-the-art (SOTA) methods in OSR tasks, achieving superior accuracy and robustness under varying signal conditions. Haohao Sun, Cong Zou, Qiexiang Wang, Jian Wang 0030, Xudong Zhang 0001 |
IEEE Internet Things J. | 2 |
| 2026 | Multi-Motion Spatio-Temporal Graph-Based User Behavior Representation for Enhanced Smartphone SecurityabstractAs central hubs of the Internet of Everything, smart phones integrate essential functions such as payments, navigation, and IoT connectivity. However, this expanded functionality also heightens security risks. Motion dynamics biometrics, which utilizes motion patterns from user-phone interactions captured via multi-sensor data, has emerged as a promising solution for smartphone security. Offering continuous and unobtrusive protection by analyzing natural user interactions, it still faces challenges in effectively modeling the complex spatio-temporal dynamics between the user and the phone within multi-sensor data. This paper focuses on leveraging graph neural networks (GNNs) to enhance user behavior modeling for smartphone security protection by capturing the relationships within motion sensor data, but it is non-trivial due to the characteristics of complexity, asynchrony, and temporal dependencies of multi motion sensor data. Towards this end, we propose MotionGNN, a multi-motion spatial-temporal graph based behavior modeling framework for user identification and authentication. Specifically, MotionGNN first divides the input multi-motion sensor data into a sequence of segments adaptively by developing a context-aware data segmentation method. Then, MotionGNN constructs fully connected spatio-temporal graphs to model sensor dependencies and temporal dynamics. Finally, windowing graph convolutions are adopted to learn user behavior representations. To evaluate the performance of MotionGNN, we collect a large-scale dataset from real-world scenarios. Extensive experiments demonstrate the state-of-the-art performance of MotionGNN in user identi fication and authentication tasks. We also test MotionGNN for 7 days on smartphones, showing high authentication accuracy with minimal battery and memory usage, making it a reliable solution for smartphone security protection. Zhihao Shen 0001, Chengmei Zhao, Cong Zou, Xi Zhao 0001, Jiakun Zhao, Jianhua Zou |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2024 | PAD: Patch-Agnostic Defense against Adversarial Patch AttacksabstractAdversarial patch attacks present a significant threat to real-world object detectors due to their practical feasibil-ity. Existing defense methods, which rely on attack data or prior knowledge, struggle to effectively address a wide range of adversarial patches. In this paper, we show two inherent characteristics of adversarial patches, semantic in-dependence and spatial heterogeneity, independent of their appearance, shape, size, quantity, and location. Seman-tic independence indicates that adversarial patches oper-ate autonomously within their semantic context, while spatial heterogeneity manifests as distinct image quality of the patch area that differs from original clean image due to the independent generation process. Based on these observations, we propose PAD, a novel adversarial patch localization and removal method that does not require prior knowledge or additional training. PAD offers patch-agnostic de-fense against various adversarial patches, compatible with any pretrained object detectors. Our comprehensive digital and physical experiments involving diverse patch types, such as localized noise, printable, and naturalistic patches, ex-hibit notable improvements over state-of-the-art works. Our code is available at https://github.com/Lihua-Jing/PAD. Lihua Jing, Rui Wang 0032, Wenqi Ren, Xin Dong 0015, Cong Zou |
CVPR | 5 |
| 2024 | Cognitive process-driven model design: A deep learning recommendation model with textual review and contextabstractOnline reviews play a crucial role in comprehending user rating behavior and improving personalized recommendations in e-commerce. However, existing review-based recommendation systems ignore the influence of theory-driven and context information. This paper proposes the Deep Learning Recommendation Model with Textual Review and Context (DeepRM-TC), which is built upon a cognitive process-driven approach to improve the quality and interpretability of recommendations. The DeepRM-TC framework mimics the human brain's cognitive process for predicting user rating behavior. Essentially, the simulation of human cognitive processing is manifested in several aspects of the designed neural network , including treating rating prediction to an attitude question, mapping raw data in latent space as the user's belief, injecting attention mechanisms for rendering judgment and predicting rating as an answer. Furthermore, we design a three-layer coattention mechanism to adaptively match product information based on users' preferences in various contexts. This mechanism extracts finer-grained interaction information from user–product–context pairs. Extensive experiments on real datasets demonstrate that our proposed model outperforms existing state-of-the-art models. We demonstrate the importance of context information and the three-layer coattention mechanism in enhancing recommendation accuracy through ablation and hierarchic analysis, respectively. Additionally, we further validate the performance of our model through data sparseness analysis, scalability analysis, other datasets, classification analysis , and user study. Xi Zhao 0001, Ningning Liu, Zhihao Shen 0001, Cong Zou |
Decis. Support Syst. | 5 |
| 2024 | Universal domain adaptation from multiple black-box sources
Cong Zou, Xinyang Kong |
Image Vis. Comput. | 3 |
| 2024 | S2CL-Leaf Net: Recognizing Leaf Images Like Human BotanistsabstractAutomatically classifying plant leaves is a challenging fine-grained classification task because of the diversity in leaf morphology, including size, texture, shape, and venation. Although powerful deep learning-based methods have achieved great improvement in leaf classification, these methods still require a large number of well-labeled samples for supervised training, which is difficult to get. In contrast, relying on the specific coarse-to-fine classification strategy, human botanists only require a small number of samples for accurate leaf recognition. Inspired by the classification strategy of human botanists, we propose a novel S 2 CL-Leaf Net , which exploits multi-granularity clues with a hierarchical attention mechanism and boosts the learning ability with the supervised sampling contrastive learning with limited training samples to classify plant leaves as human botanists do. Specifically, to fully explore and exploit the subtle details of the leaves, a novel sampling transformation mechanism is combined with the supervised contrastive learning to enhance the network’s perception of details by amplifying the discriminative regions with a weighted sampling of different regions. Furthermore, we construct the hierarchical attention mechanism to produce attention maps of different granularity, which helps to discover details in leaves that are important for classification. Experiments are conducted on the open-access leaf datasets, including Flavia, Swedish, and LeafSnap, which prove the effectiveness of the proposed S 2 CL-Leaf Net . Cong Zou, Rui Wang 0032, Cheng Jin 0001, Sanyi Zhang, Xin Wang 0019 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2023 | The Victim and The Beneficiary: Exploiting a Poisoned Model to Train a Clean Model on Poisoned DataabstractRecently, backdoor attacks have posed a serious security threat to the training process of deep neural networks (DNNs). The attacked model behaves normally on benign samples but outputs a specific result when the trigger is present. However, compared with the rocketing progress of backdoor attacks, existing defenses are difficult to deal with these threats effectively or require benign samples to work, which may be unavailable in real scenarios. In this paper, we find that the poisoned samples and benign samples can be distinguished with prediction entropy. This inspires us to propose a novel dual-network training framework: The Victim and The Beneficiary (V&B), which exploits a poisoned model to train a clean model without extra benign samples. Firstly, we sacrifice the Victim network to be a powerful poisoned sample detector by training on suspicious samples. Secondly, we train the Beneficiary network on the credible samples selected by the Victim to inhibit backdoor injection. Thirdly, a semi-supervised suppression strategy is adopted for erasing potential backdoors and improving model performance. Furthermore, to better inhibit missed poisoned samples, we propose a strong data augmentation method, AttentionMix, which works well with our proposed V&B framework. Extensive experiments on two widely used datasets against 6 state-of-the-art attacks demonstrate that our framework is effective in preventing backdoor injection and robust to various attacks while maintaining the performance on benign samples. Our code is available at https://github.com/Zixuan-Zhu/VaB. Zixuan Zhu 0002, Rui Wang 0032, Cong Zou, Lihua Jing |
ICCV | 3 |
| 2023 | Consistency-aware Feature Learning for Hierarchical Fine-grained Visual ClassificationabstractHierarchical Fine-Grained Visual Classification (HFGVC) assigns a label sequence (e.g., ["Albatross'', "Laysan Albatross'']) with a coarse to fine hierarchy to each object. It remains challenging to achieve high accuracy and consistency due to the small inter-class difference, large intra-class variance, and difficulty in modeling relationships among classification tasks at different granularities. In this paper, we propose an effective Consistency-Aware Feature Learning (CAFL) method for HFGVC to improve prediction consistency and classification accuracy simultaneously. Our key idea is to encode the prediction consistency constraint into a weak supervision mechanism via forward deduction and backward induction over the label hierarchy. Furthermore, we develop a disentanglement and bidirectional reinforcement classification head to extract the features for the classifiers at different granularities. Together with the stop-gradient policy and attention mechanism, they enable each classifier to exploit the features from the ones at other granularities without suffering from their conflicting gradients in training. We evaluate our method on several commonly-used fine-grained public datasets, including CUB-200-2011, FGVC-Aircraft, and Stanford Cars. The results show that our method not only achieves state-of-the-art classification accuracy but also effectively reduces inconsistency errors by 50% under the hierarchical fine-grained classification setting. Rui Wang 0032, Cong Zou, Zixuan Zhu 0002, Lihua Jing |
ACM Multimedia | 2 |
| 2023 | Autoencoder for Optical Intelligent Reflecting Surface-Assisted VLC System: From Model and Data-Driven PerspectivesabstractDue to the wide and license-free bandwidth, visible light communication (VLC) functions as a potential technology to meet the exponentially expanding traffic demands in wireless communications. However, the sensitivity to obstacles and high-path loss are the key issues that practical VLC systems must carefully deal with. In this article, the utilization of optical intelligent reflecting surface (OIRS) array in VLC is considered to create additional light propagation paths, thereby achieving a remarkable performance gain. In the OIRS-assisted VLC system, though the power of the signal at receiver can be increased, the resource allocation is relatively complex. Besides, the OIRS also causes time delays among signals received via various propagation paths, which is usually overlooked in existing works. To overcome these issues, the OIRS-assisted VLC system is interpreted as an autoencoder (AE), named OIRS-AE, whose architecture is enhanced according to both the model-driven and data-driven perspectives. By this way, the processing modules at the transmitter, OIRS, and receiver, including the corresponding encoding, resource management, and decoding schemes, can be simultaneously optimized, which is expected to achieve more reliable communication. Moreover, the impact of the OIRS-induced time delay spread on system performance is explored under various situations. The simulation results show that the proposed OIRS-AE can outperform the traditional OIRS-assisted VLC systems in terms of bit error rate performance. Cong Zou, Fang Yang 0001, Shiyuan Sun 0001, Ying-Jun Angela Zhang, Jian Song 0004, Zhu Han 0001 |
IEEE Internet Things J. | 1 |
| 2023 | Underwater Wireless Optical Communication With One-Bit Quantization: A Hybrid Autoencoder and Generative Adversarial Network ApproachabstractCompared with underwater acoustic communication, underwater wireless optical communication (UWOC) has the advantages of wide communication bandwidth, high transmission speed and strong directional transmission, which make it have broad application prospects. However, the absorption and scattering effects in underwater environments cause a UWOC system to suffer from severe channel fading. To reduce the system complexity and power loss, the on- -off-keying (OOK) modulation scheme and one-bit analog-to-digital converter are considered in this work. Besides, to deal with the complex underwater optical channels and the nonlinearity of the one-bit quantization, a novel deep learning based architecture integrating the autoencoder (AE) and generative adversarial network (GAN), named hybrid AE-GAN, is developed. In the proposed hybrid AE-GAN, the generator in GAN is responsible for a generalized channel equalization, which aims to equalize both the impairments from underwater optical channels and the distortion from one-bit quantization. While the decoder in AE and the discriminator in GAN share the same network to distinguish true and fake as well as reconstruct the sent messages simultaneously. Meanwhile, the one-bit quantized signal is regarded as conditional information for GAN to assist in targeted channel equalization for various signals. Moreover, a specialized loss function and training strategy for the hybrid AE-GAN are also presented, which can optimize AE and GAN at the same time, and deal with the problem that the binarization of OOK stymies the gradient backpropagation. The simulation results indicate that the proposed hybrid AE-GAN can achieve superior performance. Cong Zou, Fang Yang 0001, Jian Song 0004, Zhu Han 0001 |
IEEE Trans. Wirel. Commun. | 1 |
| 2022 | Underwater Optical Channel Generator: A Generative Adversarial Network Based ApproachabstractAs a high-speed, large-capacity underwater communication technology, underwater wireless optical communication (UWOC) has broad development prospects. However, accurately characterizing the underwater optical channel for various underwater conditions is still a challenging task due to light absorption and scattering, which may limit further research on UWOC. To overcome these issues, this paper proposes a novel framework of the underwater optical channel impulse response (CIR) generator, named the UCIRG framework, based on the generative adversarial network. Unlike the classical Monte Carlo approach, which is a numerical simulation method taking a long time to calculate, the UCIRG framework can generate a large number of CIRs extremely close to the actual underwater optical CIRs in a short time based on limited underwater optical CIRs as training data. Besides, traditional approximation methods for underwater optical channels usually have limitations on underwater conditions, lacking generality and flexibility, hence a generalized algorithm for UCIRG is also investigated, which is valid for various underwater conditions and can further reduce training time. Moreover, the multi-carrier autoencoder based UWOC system is adopted to evaluate the performance of the trained UCIRG, and the simulation results prove that the CIRs generated by the UCIRG are extremely similar to the actual CIRs. Cong Zou, Fang Yang 0001, Jian Song 0004, Zhu Han 0001 |
IEEE Trans. Wirel. Commun. | 1 |
| 2021 | MAPS: Joint Multimodal Attention and POS Sequence Generation for Video CaptioningabstractVideo captioning is considered to be challenging due to the combination of video understanding and text generation. Recent progress in video captioning has been made mainly using methods of visual feature extraction and sequential learning. However, the syntax structure and semantic consistency of generated captions are not fully explored. Thus, in our work, we propose a novel multimodal attention based framework with Part-of-Speech (POS) sequence guidance to generate more accu-rate video captions. In general, the word sequence generation and POS sequence prediction are hierarchically jointly modeled in the framework. Specifically, different modalities including visual, motion, object and syntactic features are adaptively weighted and fused with the POS guided attention mechanism when computing the probability distributions of prediction words. Experimental results on two benchmark datasets, i.e. MSVD and MSR-VTT, demonstrate that the proposed method can not only fully exploit the information from video and text content, but also focus on the decisive feature modality when generating a word with a certain POS type. Thus, our approach boosts the video captioning performance as well as generating idiomatic captions. Cong Zou, Yaosi Hu, Zhenzhong Chen 0001, Shan Liu 0001 |
VCIP | 1 |
| 2021 | H∞ State Estimation for Round-Robin Protocol-Based Markovian Jumping Neural Networks with Mixed Time Delays
Cong Zou, Bing Li 0003, Shishi Du, Xiaofeng Chen 0009 |
Neural Process. Lett. | 1 |
| 2020 | User Conditional Hashtag Recommendation for Micro-VideosabstractWhen a user tend to publish a micro-video, hashtag recommendation aims to suggest hashtags that can reflect the theme or contents of the micro-video, and meet the user tagging preference as well. In this paper, we show how user profile and historical hashtags combined with micro-video representations can be used to perform hashtag recommendation. Specifically, a User-guided Hierarchical Multi-head Attention Network (UHMAN) is proposed to attend both image-level and video-level representations of micro-videos with user side information. We evaluate the proposed model on the dataset collected from micro-video sharing platform Musical.ly. The experimental results demonstrate the effectiveness of the proposed method. Jiayi Xie, Cong Zou |
ICME | 3 |
| 2019 | PCGAN: Partition-Controlled Human Image GenerationabstractHuman image generation is a very challenging task since it is affected by many factors. Many human image generation methods focus on generating human images conditioned on a given pose, while the generated backgrounds are often blurred. In this paper, we propose a novel Partition-Controlled GAN to generate human images according to target pose and background. Firstly, human poses in the given images are extracted, and foreground/background are partitioned for further use. Secondly, we extract and fuse appearance features, pose features and background features to generate the desired images. Experiments on Market-1501 and DeepFashion datasets show that our model not only generates realistic human images but also produce the human pose and background as we want. Extensive experiments on COCO and LIP datasets indicate the potential of our method. Rui Wang 0032, Xiaowei Tian, Cong Zou |
AAAI | 4 |
| 2019 | Weighted Focus-Attention Deep Network for Fine-grained Image ClassificationabstractFine-Grained Visual Classification (FGVC) is a challenging task, due to the small variation of visual representations from different categories. An effective solution is utilizing the bounding boxes centering the object parts to extract the discriminative representations. However, regular rectangles contains the background when the shape of the part is irregular, which may interfere with the classification. In this paper, we propose a weighted focus-attention deep network (FA-Net) to address the problem of background interference in fine-grained classification. In our FA-Net, a focus-attention module is proposed to identify the foreground region from the class activation map and remove the background. Two branches are employed to obtain the primary and secondary attention regions with focus-attention module, and a weighted layer is utilized to integrate the attention regions. Experiment results on three challenging fine-grained classification datasets (e.g., CUB-200-2011, Stanford Dogs and FGVC Aircraft) show that our FA-Net obtains state-of-the-art results and outperforms the other fine-grained algorithms. Cong Zou, Rui Wang 0032, Xiaochun Cao, Feixiao Lv |
IEEE BigData | 1 |
| 2019 | Tensor-Train Based Deep Learning Approach for Compressive Sensing in Mobile ComputingabstractSeveral recent studies have been conducted on solving the compressive sensing problem with deep learning framework , which enhance the signal recovery performance and greatly shorten the running time compared with traditional algorithms. However, as the size of signals increases, so does the neural network, which will impose large memory space and high computational complexity, making it hard to use this method on mobile devices. To address this issue, in this paper, the neural network is decomposed by Tensor-Train (TT) format, which reduces the number of parameters to a great extent. In particular, the neural network decomposed by TT format is a stacked denoising autoencoder (SDA) network, which called TT-SDA. The experiments demonstrate that, especially with low measurement rates, the proposed TT-SDA network can improve the reconstruction results, reduce the computational complexity and save the memory space. Cong Zou, Fang Yang 0001 |
IWCMC | 1 |
| 2019 | VCG: Exploiting visual contents and geographical influence for Point-of-Interest recommendation
Cong Zou, Ruifeng Ding, Zhenzhong Chen 0001 |
Neurocomputing | 2 |
| 2019 | Joint latent factors and attributes to discover interpretable preferences in recommendation
Cong Zou, Zhenzhong Chen 0001 |
Inf. Sci. | 1 |