VLDB 2026 Research / reviewers in the wild / expert
Zhenyu Weng
dblp:184/7374
· DBLP profile ↗
44ranked-venue papers
12as first author
32since 2021 · last 2026
0000-0001-7857-8687ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 26 · 7 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 19 · 8 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 4 since 2021Systems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FACT: Feature Adaptive Continual-learning Tracker for multiple object tracking
Rongzihan Song, Zhenyu Weng, Huiping Zhuang, Jinchang Ren, Yongming Chen, Zhiping Lin 0001 |
Knowl. Based Syst. | 2 |
| 2026 | Towards invariant and interpretable representations for domain generalization in time series classification
Yongming Chen, Zhenyu Weng, Bah-Hwee Gwee, Qi Cao 0002, Sirajudeen Gulam Razul, Zhiping Lin 0001 |
Pattern Recognit. | 2 |
| 2026 | Generalized Face Recognition With OcclusionabstractFace recognition under occlusion remains challenging due to masks, glasses, and other real-world obstructions that partially conceal facial information. Existing approaches typically rely on training with one or more predefined occlusion types, which limits their ability to generalize to unseen scenarios. In this letter, we propose Generalized Face Recognition with Occlusion (GFRO), an occlusion-robust framework that generalizes to diverse occlusion patterns without requiring occlusion-specific training data. GFRO is trained on partial facial views cropped from complete face images using a cross-entropy loss to learn generic representations across different partial views. A dual mixture-of-experts aggregator is then introduced to refine and integrate features from multiple branches, each handling a specific partial-view representation. Optimized with a clean–noisy contrastive loss, the aggregator aligns partial-face features with complete-face features, where each branch contains experts specializing in complementary partial-view information. Extensive experiments on multiple datasets demonstrate that GFRO generalizes effectively to both real and synthetic occlusion scenarios and achieves state-of-the-art performance compared with methods trained on occlusion-specific data under the same occlusion conditions. Dengwen Zhang, Yuxi Liu 0005, Guibo Luo, Zhenyu Weng |
IEEE Signal Process. Lett. | 4 |
| 2026 | Federated Learning for Medical Image Classification: A Comprehensive BenchmarkabstractThe federated learning (FL) paradigm is well-suited for the field of medical image analysis, as it can effectively cope with machine learning on isolated multi-center data while protecting the privacy of participating parties. However, current research on optimization algorithms in FL often focuses on limited datasets and scenarios, primarily centered around natural images, with insufficient comparative experiments in medical contexts. In this work, we conduct a comprehensive evaluation of several state-of-the-art FL algorithms in the context of medical imaging. We conduct a fair comparison of classification models trained using various FL algorithms across multiple medical imaging datasets. Additionally, we evaluate system performance metrics, such as communication cost and computational efficiency, while considering different FL architectures. Our findings show that medical imaging datasets pose substantial challenges for current FL optimization algorithms. No single algorithm consistently delivers optimal performance across all medical FL scenarios, and many optimization algorithms may under-perform when applied to these datasets. Our experiments provide a benchmark and guidance for future research and application of FL in medical imaging contexts. Furthermore, we propose an efficient and robust method that combines generative techniques using denoising diffusion probabilistic models with label smoothing to augment datasets, widely enhancing the performance of FL on classification tasks across various medical imaging datasets. Our codes are released on GitHub, offering a reliable and comprehensive benchmark for future FL studies in medical imaging. Zhekai Zhou, Guibo Luo, Zhenyu Weng, Yuesheng Zhu |
IEEE J. Biomed. Health Informatics | 4 |
| 2025 | MedDiff-FT: Data-Efficient Diffusion Model Fine-Tuning with Structural Guidance for Controllable Medical Image Synthesis
Jianhao Xie, Zhenyu Weng, Yuesheng Zhu, Guibo Luo |
MICCAI (4) | 3 |
| 2025 | Online weighted hashing for cross-modal retrieval
Zining Jiang, Zhenyu Weng, Runhao Li, Huiping Zhuang, Zhiping Lin 0001 |
Pattern Recognit. | 2 |
| 2025 | Class-Specific Prompt Learning for Vision-Language ModelsabstractThe use of learning prompts to adapt pretrained vision-language models (VLMs) for downstream tasks has gained significant attention due to its potential to reduce training costs compared to model fine-tuning through few-shot learning. Most existing methods rely on a universal prompt for all classes, as it generally delivers consistent performance across various datasets. However, a universal prompt cannot capture class-specific discriminative information. To overcome this limitation, we propose class-specific prompt learning (CPL). CPL represents the context of a prompt using two components: a base vector shared among all classes and a class-specific vector designed for individual classes. This method combines the generalization ability of the base context with the adaptability of the class-specific context. Furthermore, we introduce contrastive CPL, which enhances the ability of the prompt to capture discriminative features unique to each class. Also, we adopt the self-consistency loss to regularize the base context, enhancing its generalization ability. As a result, CPL effectively learns tailored prompts for each class. Extensive experiments demonstrate that CPL achieves superior performance over existing methods in both base-class classification and new class generalization. Runhao Li, Yongming Chen, Zhenyu Weng, Zhiping Lin 0001, Yap-Peng Tan |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Joint-Neighborhood Product Quantization for Unsupervised Cross-Modal RetrievalabstractProduct quantization (PQ) is a technique that transforms high-dimensional data into compact binary codes to reduce data storage and improve search efficiency. However, existing PQ methods separate the learning of modality-specific features from the learning of quantization codewords, resulting in suboptimal performance in cross-modal retrieval tasks. In this paper, we propose a joint-neighborhood product quantization (JNPQ) method to simultaneously learn modality-specific features and quantization codewords. To achieve this, we first introduce a cross-modal quantization contrastive learning module that preserves the inter-modal neighborhood of the original data and reduces the quantization error. Then, we design a self-neighbor contrastive learning module that enhances the intra-modal neighborhood within individual modalities. Extensive experiments demonstrate that JNPQ achieves state-of-the-art results in crossmodal retrieval when compared with other unsupervised crossmodal quantization methods. Runhao Li, Zhenyu Weng, Yongming Chen, Huiping Zhuang, Yap-Peng Tan, Zhiping Lin 0001 |
VCIP | 2 |
| 2024 | Adaptive Face Recognition for Multi-Type OcclusionsabstractDue to the prevalence of influenza outbreaks and outdoor scenarios with various obstructing decorations, recognizing faces with occlusions has become a pressing challenge to address. However, current research mainly focuses on facial recognition with one kind of occlusion and does not provide compatible solutions for different kinds of common occlusions like glasses, sunglasses, and masks. Therefore, an Adaptive Multi-Type Occluded Face Recognition Model (AMOFR) is proposed to effectively handle multiple occlusion types simultaneously in this paper. In AMOFR, a generator is developed to produce diverse occluded face images for training, achieved by simulating various occlusion types on unoccluded face images. Subsequently, an occlusion type-based adapter is formulated to address a range of occlusion scenarios, guided by prompts from a Visual-Language model. To enhance overall performance by leveraging complete facial information, a feature-level knowledge distillation loss function is implemented, facilitating joint learning of unoccluded-face and occluded-face features. Furthermore, a new sunglasses-wearing dataset (CALFW-SUNGLASSES) is generated for more comprehensive test for AMOFR and further occlusion recognition research. Experimental results on datasets containing different types of occlusions have demonstrated that AMOFR achieves significantly higher accuracy compared to other advanced face recognition models. The implementation codes of AMOFR is available athttps://github.com/LIU-YUXI/Adaptive-Multi-occlusion-Face-Recognition. Yuxi Liu 0005, Guibo Luo, Zhenyu Weng, Yuesheng Zhu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Few-Shot Contrastive Transfer Learning With Pretrained Model for Masked Face VerificationabstractFace verification has seen remarkable progress that benefits from large-scale publicly available databases. However, it remains a challenge how to generalize a pretrained face verification model to a new scenario with a limited amount of data. In many real-world applications, the training database only contains a limited number of identities with two images for each identity due to the privacy concern. In this article, we propose to transfer knowledge from a pretrained unmasked face verification model to a new model for verification between masked and unmasked faces, to meet the application requirements during the COVID-19 pandemic. To overcome the lack of intra-class diversity resulting from only a pair of masked and unmasked faces for each identity ($\text{i.e.},$two shots for each identity), a static prototype classification function is designed to learn features for masked faces by utilizing unmasked face knowledge from the pretrained model. Meanwhile, a contrastive constrained embedding function is designed to preserve unmasked face knowledge of the pretrained model during the transfer learning process. By combining these two functions, our method uses knowledge acquired from the pretrained unmasked face verification model to proceed with verification between masked and unmasked faces with a limited amount of training data. Extensive experiments demonstrate that our method can perform better than state-of-the-art methods for verification between masked and unmasked faces in the few-shot transfer learning setting. Zhenyu Weng, Huiping Zhuang, Fulin Luo, Haizhou Li 0001, Zhiping Lin 0001 |
IEEE Trans. Multim. | 1 |
| 2023 | Zero-Shot Offline Handwritten Chinese Character Recognition with Graph EmbeddingabstractHandwritten Chinese character recognition (HCCR) is a challenging topic in the field of computer vision due to its numerous categories, complex structure, shape similarity, casual writing style, and lack of training data. Some studies have attempted to develop radicals-based approaches with zero-shot recognition capacity to alleviate the data dependency problem of deep learning. However, previous studies tend to treat Chinese characters as isolated individuals and focus solely on the structural characteristics, ignoring the correlation information among Chinese characters. In this paper, we construct a graph to represent the correlation information and propose a novelty zero-shot HCCR method based on graph embedding. In addition, we present a pre-training approach for the encoder based on contrastive learning. The experimental results show that our method outperforms the other radical-based zero-shot recognition methods and also achieves a competitive performance on traditional experiment setting. Zhenyu Weng, Yuesheng Zhu |
CSCWD | 2 |
| 2023 | GKEAL: Gaussian Kernel Embedded Analytic Learning for Few-Shot Class Incremental TaskabstractFew-shot class incremental learning (FSCIL) aims to address catastrophic forgetting during class incremental learning in a few-shot learning setting. In this paper, we approach the FSCIL by adopting analytic learning, a technique that converts network training into linear problems. This is inspired by the fact that the recursive implementation (batch-by-batch learning) of analytic learning gives identical weights to that produced by training on the entire dataset at once. The recursive implementation and the weight-identical property highly resemble the FSCIL setting (phase-by-phase learning) and its goal of avoiding catastrophic forgetting. By bridging the FSCIL with the analytic learning, we propose a Gaussian kernel embedded analytic learning (GKEAL) for FSCIL. The key components of GKEAL include the kernel analytic module which allows the GKEAL to conduct FSCIL in a recursive manner, and the augmented feature concatenation module that balances the preference between old and new tasks especially effectively under the few-shot setting. Our experiments show that the GKEAL gives state-of-the-art performance on several benchmark datasets. Huiping Zhuang, Zhenyu Weng, Run He, Zhiping Lin 0001, Ziqian Zeng |
CVPR | 2 |
| 2023 | Imbalanced Conditional Conv-Transformer for Mathematical Expression Recognition
Shuaijian Ji, Zhaokun Zhou, Yuqing Wang 0006, Baishan Duan, Zhenyu Weng, Yuesheng Zhu |
ICANN (6) | 5 |
| 2023 | Style Expansion Without Forgetting for Handwritten Character Recognition
Jie Ruan, Zhenyu Weng, Jian Zhang 0018, Yuqing Wang 0006, Longhui Yu, Qiankun Gao, Yuesheng Zhu |
ICANN (3) | 2 |
| 2023 | Trapdoor Normalization with Irreversible Ownership VerificationabstractThis paper introduces a deep model watermark with an irreversible ownership verification scheme: Trapdoor Normalization (TdN), inspired by the trapdoor function in traditional cryptography. To protect intellectual property within deep models, the proposed method is able to embed ownership information into normalization layers during training. We argue and empirically validate that relevant methods are vulnerable to ambiguity attacks, where the forged watermarks can cast ambiguity over the ownership verification. The primary trait that distinguishes this work from previous ones, is its design of a bidirectional connection between watermarks and deep models. Thereby, TdN enables an irreversible ownership verification scheme that is difficult for the adversary to compromise. In this way, the proposed TdN can effectively defeat ambiguity attacks. Extensive experiments demonstrate that the proposed method is not only superior to previous state-of-the-art methods in robustness, but also has better efficiency. Zhenyu Weng, Yuesheng Zhu, Yadong Mu |
ICML | 2 |
| 2023 | TRMER: Transformer-Based End to End Printed Mathematical Expression RecognitionabstractAs a fundamental task of transcribing formula images into structural mathematical expressions, Printed Mathematical Expression Recognition (PMER) is wildly used in many fields. However, there is still a lack of an end-to-end approach toward fully exploring the spatial structure and semantic information in the formula to achieve high recognition accuracy. In this work, a Transformer-based Mathematical Expression Recognition (TRMER) model, is proposed to enhance the recognition accuracy. A Dual-Branch Encoder (DBE) is developed to extract multi-scaled feature maps from a formula image so that the spatial and semantic information can be obtained synchronously, and the different feature maps are fused with a Fusion Enhancement Module (FEM) by merging and reinforcing the spatial-semantic information. A standard transformer-based decoder is developed to decode the rich spatial-semantic information of the image and output a recognized mathematical expression in LaTex sequence. The experimental results have illustrated that the TRMER has achieved state-of-the-art recognition performance. Zhaokun Zhou, Shuaijian Ji, Yuqing Wang 0006, Zhenyu Weng, Yuesheng Zhu |
IJCNN | 4 |
| 2023 | Neighborhood Learning from Noisy Labels for Cross-Modal RetrievalabstractCross-modal retrieval methods are developed to retrieve relevant data across different modalities. Usually, super-vised cross-modal retrieval methods can achieve higher accuracy than unsupervised methods because they can utilize the semantic information provided by clean labels. However, training data with noisy labels will lead to the performance degradation of supervised cross-modal retrieval methods. In this work, we present a novel framework called Neighborhood Learning for Cross-Modal Retrieval (NLCMR) that is robust against noisy labels by exploiting the information contained in the neighbor-hood. Our NLCMR contains two main components: Clustering with Neighborhood Alignment and Neighborhood Contrastive Learning. The first component focuses on reducing the impact of noisy labels and improving clustering robustness, and the second component learns from noisy data by exploring pairwise and neighborhood information. Extensive experiments are conducted on three multi-modal datasets to demonstrate the effectiveness of NLCMR. Runhao Li, Zhenyu Weng, Huiping Zhuang, Yongming Chen, Zhiping Lin 0001 |
ISCAS | 2 |
| 2023 | POAR: Towards Open Vocabulary Pedestrian Attribute RecognitionabstractPedestrian attribute recognition (PAR) aims to predict the attributes of a target pedestrian. Recent methods often address the PAR problem by training a multi-label classifier with predefined attribute classes, but they can hardly exhaust all possible pedestrian attributes in the real world. To tackle this problem, we propose a novel Pedestrian Open-Attribute Recognition (POAR) approach by formulating the problem as a task of image-text search. Our approach employs a Transformer-based Encoder with a Masking Strategy (TEMS) to focus on the attributes of specific pedestrian parts (e.g., head, upper body, lower body, feet, etc.), and introduces a set of attribute tokens to encode the corresponding attributes into visual embeddings. Each attribute category is described as a natural language sentence and encoded by the text encoder. Then, we compute the similarity between the visual and text embeddings to find the best attribute descriptions for the input images. To handle multiple attributes of a single pedestrian, we propose a Many-To-Many Contrastive (MTMC) loss with masked tokens. In addition, we propose a Grouped Knowledge Distillation (GKD) method to minimize the disparity between visual embeddings and unseen attribute text embeddings. We evaluate our proposed method on three PAR datasets with an open-attribute setting. The results demonstrate the effectiveness of our method as a strong baseline for the POAR task. Our code is available at https://github.com/IvyYZ/POAR. Yue Zhang 0065, Suchen Wang, Shichao Kan, Zhenyu Weng, Yi-Gang Cen, Yap-Peng Tan |
ACM Multimedia | 4 |
| 2023 | Online Multi-Face Tracking With Multi-Modality Cascaded MatchingabstractTracking multiple faces online in unconstrained videos is a challenging problem as faces may appear drastically different over time and identities can be inferred only based on information available from past frames. Previous tracking methods focus on face information without reference to other modality information such as a person’s overall body appearance, leading to suboptimal performance. In this paper, we propose a new online multi-face tracking method, called online multi-face tracking with multi-modality cascaded matching (OMTMCM), to improve the tracking performance by using both face and body information. The proposed OMTMCM consists of two stages, namely detection alignment and detection association. In the first stage, a detection alignment module is designed to align face detection with body detection from the same person for the subsequent detection association. In the second stage, a cascaded matching module is designed to associate face detections across frames to locate trajectory of each target face by using both face and body information. Specifically, aligned face-body detections in the current frame are matched in a cascade manner with body and face features that are selected from past frames and stored in the designed feature memory. In this way, our method can track multiple faces online with both face and body information while eliminating the possibility of face detection and body detection from the same person being separately assigned with different identities. Experimental results demonstrate our method is on par with or better than other online tracking methods for multi-face tracking. Zhenyu Weng, Huiping Zhuang, Haizhou Li 0001, Balakrishnan Ramalingam, Mohan Rajesh Elara, Zhiping Lin 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | Attention Multihop Graph and Multiscale Convolutional Fusion Network for Hyperspectral Image ClassificationabstractConvolutional neural networks (CNNs) for hyperspectral image (HSI) classification have generated good progress. Meanwhile, graph convolutional networks (GCNs) have also attracted considerable attention by using unlabeled data, broadly and explicitly exploiting correlations between adjacent parcels. However, the CNN with a fixed square convolution kernel is not flexible enough to deal with irregular patterns, while the GCN using the superpixel to reduce the number of nodes will lose the pixel-level features, and the features from the two networks are always partial. In this paper, to make good use of the advantages of CNN and GCN, we propose a novel multiple feature fusion model termed attention multi-hop graph and multi-scale convolutional fusion network (AMGCFN), which includes two sub-networks of multi-scale fully CNN and multi-hop GCN to extract the multi-level information of HSI. Specifically, the multi-scale fully CNN aims to comprehensively capture pixel-level features with different kernel sizes, and a multi-head attention fusion module is used to fuse the multi-scale pixel-level features. The multi-hop GCN systematically aggregates the multi-hop contextual information by applying multi-hop graphs on different layers to transform the relationships between nodes, and a multi-head attention fusion module is adopted to combine the multi-hop features. Finally, we design a cross attention fusion module to adaptively fuse the features of two sub-networks. AMGCFN makes full use of multi-scale convolution and multi-hop graph features, which is conducive to the learning of multi-level contextual semantic features. Experimental results on three benchmark HSI datasets show that AMGCFN has better performance than a few state-of-the-art methods. Fulin Luo, Huiping Zhuang, Zhenyu Weng, Xiuwen Gong, Zhiping Lin 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Unsupervised Online Hashing with Multi-Bit Quantization
Zhenyu Weng, Yuesheng Zhu |
ACCV (7) | 1 |
| 2022 | Dual Branch Network Towards Accurate Printed Mathematical Expression Recognition
Yuqing Wang 0006, Zhenyu Weng, Zhaokun Zhou, Shuaijian Ji, Zhongjie Ye, Yuesheng Zhu |
ICANN (4) | 2 |
| 2022 | Multi-Teacher Knowledge Distillation for Incremental Implicitly-Refined ClassificationabstractIncremental learning methods can learn new classes continually by distilling knowledge from the last model (as a teacher model) to the current model (as a student model) in the sequentially learning process. However, these methods cannot work for Incremental Implicitly-Refined Classification (IIRC), an incremental learning extension where the incoming classes could have two granularity levels, a superclass label and a subclass label. This is because the previously learned superclass knowledge may be occupied by the sub-class knowledge learned sequentially. To solve this problem, we propose a novel Multi-Teacher Knowledge Distillation (MTKD) strategy. To preserve the subclass knowledge, we use the last model as a general teacher to distill the previous knowledge for the student model. To preserve the superclass knowledge, we use the initial model as a superclass teacher to distill the superclass knowledge as the initial model contains abundant superclass knowledge. However, distilling knowledge from two teacher models could result in the student model making some redundant predictions. We further propose a post-processing mechanism, called as Top-k prediction restriction to reduce the redundant predictions. Our experimental results on IIRC-ImageNet120 and IIRC-CIFAR100 show that the proposed method can achieve better classification accuracy compared with existing state-of-the-art methods. Longhui Yu, Zhenyu Weng, Yuqing Wang 0006, Yuesheng Zhu |
ICME | 2 |
| 2022 | ACIL: Analytic Class-Incremental Learning with Absolute Memorization and Privacy ProtectionabstractClass-incremental learning (CIL) learns a classification model with training data of different classes arising progressively. Existing CIL either suffers from serious accuracy loss due to catastrophic forgetting, or invades data privacy by revisiting used exemplars. Inspired by learning of linear problems, we propose an analytic class-incremental learning (ACIL) with absolute memorization of past knowledge while avoiding breaching of data privacy (i.e., without storing historical data). The absolute memorization is demonstrated in the sense that the CIL using ACIL given present data would give identical results to that from its joint-learning counterpart that consumes both present and historical samples. This equality is theoretically validated. The data privacy is ensured by showing that no historical data are involved during the learning process. Empirical validations demonstrate ACIL's competitive accuracy performance with near-identical results for various incremental task settings (e.g., 5-50 phases). This also allows ACIL to outperform the state-of-the-art methods for large-phase scenarios (e.g., 25 and 50 phases). Huiping Zhuang, Zhenyu Weng, Hongxin Wei, Renchunzi Xie, Kar-Ann Toh, Zhiping Lin 0001 |
NeurIPS | 2 |
| 2022 | DEMA: a distance-bounded energy-field minimization algorithm to model and layout biomolecular networks with quantitative featuresabstractSUMMARY: In biology, graph layout algorithms can reveal comprehensive biological contexts by visually positioning graph nodes in their relevant neighborhoods. A layout software algorithm/engine commonly takes a set of nodes and edges and produces layout coordinates of nodes according to edge constraints. However, current layout engines normally do not consider node, edge or node-set properties during layout and only curate these properties after the layout is created. Here, we propose a new layout algorithm, distance-bounded energy-field minimization algorithm (DEMA), to natively consider various biological factors, i.e., the strength of gene-to-gene association, the gene's relative contribution weight and the functional groups of genes, to enhance the interpretation of complex network graphs. In DEMA, we introduce a parameterized energy model where nodes are repelled by the network topology and attracted by a few biological factors, i.e., interaction coefficient, effect coefficient and fold change of gene expression. We generalize these factors as gene weights, protein-protein interaction weights, gene-to-gene correlations and the gene set annotations-four parameterized functional properties used in DEMA. Moreover, DEMA considers further attraction/repulsion/grouping coefficient to enable different preferences in generating network views. Applying DEMA, we performed two case studies using genetic data in autism spectrum disorder and Alzheimer's disease, respectively, for gene candidate discovery. Furthermore, we implement our algorithm as a plugin to Cytoscape, an open-source software platform for visualizing networks; hence, it is convenient. Our software and demo can be freely accessed at http://discovery.informatics.uab.edu/dema. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Zhenyu Weng, Zongliang Yue, Yuesheng Zhu, Jake Yue Chen |
Bioinform. | 1 |
| 2021 | Semantic-Aware Context Aggregation for Image InpaintingabstractRecent attention-based image inpainting methods have made inspiring progress by propagating distant contextual information into holes. However, they tend to generate blurry contents since the propagation process is always misled by preliminarily-recovered holes features which are not well-inferred. To handle this problem, we propose a novel semantic-aware context aggregation module (SACA) that aggregates distant contextual information from a semantic perspective by exploiting the internal semantic similarity of the input feature map. Compared with existing attention mechanisms that model the relation of all pixel-pairs, SACA can suppress the impact of misleading holes features in context aggregation and significantly reduce computation burden by learning the relation between pixels and semantics. Also, we apply SACA to both high-level and low-level feature maps in our model for generating both semantically and visually plausible results. Extensive experiments on Outdoor Scenes, CelebA and Paris StreetView datasets validate the superiority of our method compared with existing methods. Zhilin Huang, Chujun Qin, Ruixin Liu, Zhenyu Weng, Yuesheng Zhu |
ICASSP | 4 |
| 2021 | Watermarking Deep Neural Networks with Greedy ResidualsabstractDeep neural networks (DNNs) are considered as intellectual property of their corresponding owners and thus are in urgent need of ownership protection, due to the massive amount of time and resources invested in designing, tuning and training them. In this paper, we propose a novel watermark-based ownership protection method by using the residuals of important parameters. Different from other watermark-based ownership protection methods that rely on some specific neural network architectures and during verification require external data source, namely ownership indicators, our method does not explicitly use ownership indicators for verification to defeat various attacks against DNN watermarks. Specifically, we greedily select a few and important model parameters for embedding so that the impairment caused by the changed parameters can be reduced and the robustness against different attacks can be improved as the selected parameters can well preserve the model information. Also, without the external data sources for verification, the adversary can hardly cast doubts on ownership verification by forging counterfeit watermarks. The extensive experiments show that our method outperforms previous state-of-the-art methods in five tasks. Zhenyu Weng, Yuesheng Zhu |
ICML | 2 |
| 2021 | Accumulated Decoupled Learning with Gradient Staleness Mitigation for Convolutional Neural NetworksabstractGradient staleness is a major side effect in decoupled learning when training convolutional neural networks asynchronously. Existing methods that ignore this effect might result in reduced generalization and even divergence. In this paper, we propose an accumulated decoupled learning (ADL), which includes a module-wise gradient accumulation in order to mitigate the gradient staleness. Unlike prior arts ignoring the gradient staleness, we quantify the staleness in such a way that its mitigation can be quantitatively visualized. As a new learning scheme, the proposed ADL is theoretically shown to converge to critical points in spite of its asynchronism. Extensive experiments on CIFAR-10 and ImageNet datasets are conducted, demonstrating that ADL gives promising generalization results while the state-of-the-art methods experience reduced generalization and divergence. In addition, our ADL is shown to have the fastest training speed among the compared methods. Huiping Zhuang, Zhenyu Weng, Fulin Luo, Kar-Ann Toj, Haizhou Li 0001, Zhiping Lin 0001 |
ICML | 2 |
| 2021 | Sequence-Aware Graph Neural Network for Session-based RecommendationabstractSession-based recommendation (SBR) nowadays plays a vital role in many online services, aiming to predict users' next action based on anonymous sessions. Recent research of GNNs-based methods models a session as a graph via investigating complex transitions of items in a session. However, these methods do not consider sequential information of the session when aggregating item embeddings to form a session-level embedding. Most methods consider not all previous but the last one item as the interest of a user, which restricts the performance of the model. To address this problem, we propose a model named Sequence-Aware Graph Neural Network (SA-GNN) for session-based recommendation. In SA-GNN, we design a sequence-aware attention to adaptively weigh the previous items to generate a session-level embedding, which greatly improves the representation ability of the model. Also, to improve the representation ability of the item embeddings, SA-GNN harnesses the power of self-attention within the GNN layer to capture both transitions between adjacent items and long-range dependencies among all items in a session. In empirical evaluations on three public recommendation datasets, our method consistently outperforms an extensive of state-of-the-art session-based recommendation methods. Zhencheng Huang, Zhenyu Weng, Yuesheng Zhu, Zhiqiang Bai |
IJCNN | 3 |
| 2021 | Bi-encoder Network with Structure-texture Consistency for Image InpaintingabstractExisting image inpainting methods have shown their potential in filling corrupted regions with plausible contents. However, these methods tend to produce results with distorted structures or unnatural textures since they neglect the difference between structures and textures in images and jointly process these two different types of information. To solve this problem, we propose a bi-encoder network (BE-Net) that seeks to handle structure and texture information separately, and fuse them to reconstruct completed images. Specifically, BE-Net first uses two parallel encoders to infer structure and texture features of the input images respectively. Then a structure-texture consistency module (STCM) is designed to weaken artifacts and enhance visual coherency of the output images by keeping the texture features consistent with the structure features. Finally, the structure features and the texture features are fused at each level of the decoder to recover images with reasonable structures and realistic textures. Extensive experiments on Paris StreetView and CelebA datasets show the proposed approach is effective in generating realistic and visually plausible results and outperforms several state-of-the-art methods. Chujun Qin, Zhilin Huang, Ruixin Liu, Zhenyu Weng, Yuesheng Zhu |
IJCNN | 4 |
| 2021 | Learning frame-level affinity with video-level labels for weakly supervised temporal action detection
Bairong Li, Yuesheng Zhu, Ruixin Liu, Zhenyu Weng |
Neurocomputing | 4 |
| 2021 | Online Hashing With Bit Selection for Image RetrievalabstractOnline hashing methods have been intensively investigated in semantic image retrieval due to their efficiency in learning the hash functions with one pass through the streaming data. Among the online hashing methods, those based on the target codes are usually superior to others. However, the target codes in these methods are generated heuristically in advance and cannot be learned online to capture the characteristics of the data. In this paper, we propose a new online hashing method in which the target codes are constructed according to the data characteristics and are used to learn the hash functions online. By designing a metric to select the effective bits online for constructing the target codes, the learned hash functions are resistant to the bit-flipping error. At the same time, the correlation between the hash functions is also considered in the designed metric. Hence, the hash functions have low redundancy. Extensive experiments show that our method can achieve comparable or better performance than other online hashing methods on both the static database and the dynamic database. Zhenyu Weng, Yuesheng Zhu |
IEEE Trans. Multim. | 1 |
| 2020 | Efficient Querying from Weighted Binary CodesabstractBinary codes are widely used to represent the data due to their small storage and efficient computation. However, there exists an ambiguity problem that lots of binary codes share the same Hamming distance to a query. To alleviate the ambiguity problem, weighted binary codes assign different weights to each bit of binary codes and compare the binary codes by the weighted Hamming distance. Till now, performing the querying from the weighted binary codes efficiently is still an open issue. In this paper, we propose a new method to rank the weighted binary codes and return the nearest weighted binary codes of the query efficiently. In our method, based on the multi-index hash tables, two algorithms, the table bucket finding algorithm and the table merging algorithm, are proposed to select the nearest weighted binary codes of the query in a non-exhaustive and accurate way. The proposed algorithms are justified by proving their theoretic properties. The experiments on three large-scale datasets validate both the search efficiency and the search accuracy of our method. Especially for the number of weighted binary codes up to one billion, our method shows a great improvement of more than 1000 times faster than the linear scan. Zhenyu Weng, Yuesheng Zhu |
AAAI | 1 |
| 2020 | Online Hashing with Efficient Updating of Binary CodesabstractOnline hashing methods are efficient in learning the hash functions from the streaming data. However, when the hash functions change, the binary codes for the database have to be recomputed to guarantee the retrieval accuracy. Recomputing the binary codes by accumulating the whole database brings a timeliness challenge to the online retrieval process. In this paper, we propose a novel online hashing framework to update the binary codes efficiently without accumulating the whole database. In our framework, the hash functions are fixed and the projection functions are introduced to learn online from the streaming data. Therefore, inefficient updating of the binary codes by accumulating the whole database can be transformed to efficient updating of the binary codes by projecting the binary codes into another binary space. The queries and the binary code database are projected asymmetrically to further improve the retrieval accuracy. The experiments on two multi-label image databases demonstrate the effectiveness and the efficiency of our method for multi-label image retrieval. Zhenyu Weng, Yuesheng Zhu |
AAAI | 1 |
| 2020 | Real-Time Multiple Object Tracking with Discriminative FeaturesabstractTracking-by-detection methods track multiple objects by detecting the objects of interest in each frame and associating the detected objects with the tracks. By allowing object detection and appearance embedding to be learned in a shared network, recent tracking-by-detection methods can implement the tracking task in real time with the power of deep neural networks. However, they just focus on the detection stage and do not take advantage of the embedding features well in the association stage. In this paper, we exploit the discriminative embedding features in the association stage to improve the tracking performance. By combing the embedding features with the bounding boxes to associate the detected objects with the tracks, the number of identity switches during tracking can be reduced. Further, after associating the detected objects with the tracks, the embedding feature of each track is not only updated according to the associated object, but also learned to distinguish the similar detected objects. The experiments show that our method can achieve competitive tracking performance in real time compared to the state-of-the-art tracking methods. Zhenyu Weng, Yuesheng Zhu, Zhiping Lin 0001, Haizhou Li 0001 |
ICARCV | 1 |
| 2020 | Temporal Adaptive Alignment Network for Deep Video InpaintingabstractVideo inpainting aims to synthesize visually pleasant and temporally consistent content in missing regions of video. Due to a variety of motions across different frames, it is highly challenging to utilize effective temporal information to recover videos. Existing deep learning based methods usually estimate optical flow to align frames and thereby exploit useful information between frames. However, these methods tend to generate artifacts once the estimated optical flow is inaccurate. To alleviate above problem, we propose a novel end-to-end Temporal Adaptive Alignment Network(TAAN) for video inpainting. The TAAN aligns reference frames with target frame via implicit motion estimation at a feature level and then reconstruct target frame by taking the aggregated aligned reference frame features as input. In the proposed network, a Temporal Adaptive Alignment (TAA) module based on deformable convolutions is designed to perform temporal alignment in a local, dense and adaptive manner. Both quantitative and qualitative evaluation results show that our method significantly outperforms existing deep learning based methods. Ruixin Liu, Zhenyu Weng, Yuesheng Zhu, Bairong Li |
IJCAI | 2 |
| 2020 | A Disocclusion Inpainting Framework for Depth-Based View SynthesisabstractThis paper proposes a disocclusion inpainting framework for depth-based view synthesis. It consists of four modules: foreground extraction, motion compensation, improved background reconstruction, and inpainting. The foreground extraction module detects the foreground objects and removes them from both depth map and rendered video; the motion compensation module guarantees the background reconstruction model to suit for moving camera scenarios; the improved background reconstruction module constructs a stable background video by exploiting the temporal correlation information in both 2D video and its corresponding depth map; and the constructed background video and inpainting module are used to eliminate the holes in the synthesized view. The analysis and experiment indicate that the proposed framework has good generality, scalability and effectiveness, which means most of the existing background reconstruction methods and image inpainting methods can be employed or extended as the modules in our framework. Our comparison results have demonstrated that the proposed framework achieves better synthesized quality, temporal consistency, and has lower running time compared to the other methods. Guibo Luo, Yuesheng Zhu, Zhenyu Weng, Zhaotian Li |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2020 | Concatenation hashing: A relative position preserving method for learning binary codes
Zhenyu Weng, Yuesheng Zhu |
Pattern Recognit. | 1 |
| 2019 | Annular Sector Model for tracking multiple indistinguishable and deformable objects in occlusions
Guibo Luo, Zhenyu Weng, Yuesheng Zhu |
Neurocomputing | 3 |
| 2019 | A fast online spherical hashing method based on data sampling for large scale image retrieval
Zhenyu Weng, Yuesheng Zhu, Yinhe Lan, Long-Kai Huang |
Neurocomputing | 1 |
| 2017 | A new spherical hashing method in a low-dimensional isotropic spaceabstractBy using a hypersphere to group the spatially coherent data points into the same bit, spherical hashing (SPH) can achieve a good performance in approximate nearest neighbor (ANN) search. However, when the data dimensionality rises, the data becomes sparse and the hypersphere needs to increase its radius to maintain the same coverage, which makes the data points in the hypersphere less coherent. To alleviate the effect brought from the high dimensionality of the data, a new hypersphere-based hashing method is proposed. By constructing a low-dimensional isotropic space where the variance of projection along each component is equal, both the similarity and the distribution of the original data can be preserved in this space. And then, the hashing functions are learnt by SPH in this space. The experiments on SIFT1M and GIST1M datasets show that the performance of SPH can be improved by our method and is superior to other state-of-the-art hashing methods in terms of recall and mAP performance. Yinhe Lan, Zhenyu Weng, Yuesheng Zhu |
VCIP | 2 |
| 2016 | Asymmetric distance for spherical hashingabstractUsually, most of hashing methods for information retrieval have a two-step procedure, embedding the data into a low-dimensional intermediate space and then quantizing them into binary codes. In the hyperplane-based hashing methods, the distance between the data in the intermediate space can replace the Hamming distance to improve the retrieval accuracy. In this paper, a novel asymmetric distance for the hypersphere-based method is proposed to improve the accuracy of similarity search. By showing that the distance in the intermediate space can approximate the Euclidean distance between the data points, more useful information can be taken to improve the retrieval accuracy. According to the characteristics of the hypersphere-based hashing method, various asymmetric distance models are developed and described. Our experiments with two datasets have demonstrated that the proposed method can improve the retrieval accuracy of the hypersphere-based hashing methods significantly and achieve the state-of-the-art recall performance. Zhenyu Weng, Wenbin Yao, Ziqiang Sun, Yuesheng Zhu |
ICIP | 1 |
| 2016 | Diversity regularized metric learning for person re-identificationabstractMetric learning is an effective method for person re-identification. It utilizes latent factors to find a suitable space for measuring distances. In general, a small number of factors are not powerful enough to match the pedestrians while a large number of factors cause high computational cost. In this paper, to balance this trade-off, a novel diversity regularized distance metric learning method is proposed. For feature representation, the local discriminative features are extracted from the source image and an adjacency maximal constraint is developed to handle the misaligned issue. Then a diversity regularizer is used to learn a metric by making the latent factors uncorrelated so that a small amount of latent factors can preserve effectiveness in measuring distances while reducing the computational burden. Our experimental results show that the proposed method with a small amount of factors can obtain comparative or even better performance compared to the state-of-art methods. Wenbin Yao, Zhenyu Weng, Yuesheng Zhu |
ICIP | 2 |
| 2016 | Asymmetric hashing with multi-bit quantization for image retrieval
Zhenyu Weng, Ziqiang Sun, Yuesheng Zhu |
Neurocomputing | 1 |