EDBT 2026 Demo / reviewers in the wild / expert
Li Zhu 0003
dblp:74/3823-3
· DBLP profile ↗
57ranked-venue papers
1as first author
44since 2021 · last 2026
0000-0003-2136-3196ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 27 · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 25 · 17 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Systems, architecture and hardware · 2 · 2 since 2021Computer networks · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Generating Attribution Reports for Manipulated Facial Images: A Dataset and BaselineabstractJingchun Lian, Lingyu Liu, Yaxiong Wang, Yujiao Wu, Lianwei Wu, Li Zhu, Zhedong Zheng. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Jingchun Lian, Lingyu Liu, Yaxiong Wang, Yujiao Wu, Lianwei Wu, Li Zhu 0003, Zhedong Zheng |
ACL (1) | 6 |
| 2026 | Cultivating Forensic Reasoning for Generalizable Multimodal Manipulation DetectionabstractYuchen Zhang, Yaxiong Wang, Kecheng Han, Yujiao Wu, Lianwei Wu, Li Zhu, Zhedong Zheng. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yaxiong Wang, Kecheng Han, Yujiao Wu, Lianwei Wu, Li Zhu 0003, Zhedong Zheng |
ACL (1) | 6 |
| 2026 | AEGI: Anchor Event Guided Inference for TKGQA
Yuqing Fu, Yejing Wang, Li Zhu 0003, Xueming Qian, Guoshuai Zhao 0001, Xiangyu Zhao 0001 |
PAKDD (3) | 4 |
| 2026 | Minimizing the pretraining gap: Domain-aligned text-based person retrieval
Shuyu Yang, Yaxiong Wang, Li Zhu 0003, Zhedong Zheng |
Pattern Recognit. | 4 |
| 2026 | AdapSNE: Adaptive Fireworks-Optimized and Entropy-Guided Dataset Sampling for Edge DNN TrainingabstractTraining deep neural networks (DNNs) on edge devices faces challenges due to the large-scale datasets required, which are costly for edge devices, especially in large language model (LLM) tasks. To address this, a DNN-free method called Near-Memory Sampling (NMS) has been introduced. NMS reduces dimensionality and performs exemplar sampling in the reduced space, avoiding architectural bias and improving generalization. However, NMS has two limitations: 1) The mismatch between the search method and the non-monotonic property of the perplexity error function leads to the emergence of outliers; 2) Key parameter (i.e., target perplexity) is selected empirically, introducing arbitrariness and leading to uneven sampling. These two issues lead torepresentative biasof exemplars, resulting in degraded accuracy. To overcome these, we propose AdapSNE, which integrates the Fireworks Algorithm (FWA) for efficient non-monotonic search to avoid outliers and uses entropy-guided optimization for uniform sampling, ensuring representative training samples. To reduce the cost of iterative computations, we design an accelerator with custom dataflow and time-multiplexing mechanisms. Experimental results show that AdapSNE outperforms state-of-the-art methods, including both DNN-based (DQAS) and DNN-free (NMS) approaches, across small-scale image datasets, large-scale datasets, and the MMLU benchmark for LLM tasks. Boran Zhao, Hetian Liu, Zihang Yuan, Li Zhu 0003, Lina Xie, Tian Xia 0008, Wenzhe Zhao 0001, Pengju Ren |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2026 | Momentum Centroid Alignment With Temporal-Relational Disentanglement for Cross-Domain Few-Shot Action RecognitionabstractTraditional few-shot action recognition (FSAR) aims to address the problem of the scarcity of action videos, enabling the recognition of action categories with just a few labeled samples. It is generally believed that the samples in the meta-training phase and the meta-testing phase are all drawn from the same domain. However, in practical applications, they often come from different domains, which may lead to significant differences in the distribution of spatiotemporal features. Researchers have started to study the problem of cross-domain few-shot action recognition (CDFSAR). The current solution is to train the model by combining source domain video and unlabeled target domain video to improve the model’s generalization ability. In this paper, we follow this paradigm but make a more refined use of the unlabeled target domain videos to better extract transferable features. First, we decouple the source and target domain videos along the temporal dimension and extract the domain-irrelevant features in both the source and target domains. Second, in each episode, we calculate the centroid of the domain-irrelevant features of the target domain and perform a momentum update on this feature centroid. We use Cross-Attention to align the domain-irrelevant features of the source domain toward this dynamic centroid. Finally, we use these aligned source domain features for few-shot classification. Experimental results demonstrate that our approach significantly improves few-shot classification performance across diverse domain shifts, validating the effectiveness of our refined use of unlabeled target video. Our code has been published at the URL: https://github.com/cofly2014/MCA-TRD.git. Fei Guo 0010, Xuetao Zhang 0001, Qi Han 0008, Lingyu Liu, Bo Liu 0095, Li Zhu 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2026 | Look, Compare and Draw: Differential Query Transformer for Automatic Oil PaintingabstractThis work introduces a new approach to automatic oil painting that emphasizes the creation of dynamic and expressive brushstrokes. A pivotal challenge lies in mitigating the duplicate and common-place strokes, which often lead to less aesthetic outcomes. Inspired by the human painting process, i.e., observing, comparing, and drawing, we incorporate differential image analysis into a neural oil painting model, allowing the model to effectively concentrate on the incremental impact of successive brushstrokes. To operationalize this concept, we propose the Differential Query Transformer (DQ-Transformer), a new architecture that leverages differentially derived image representations enriched with positional encoding to guide the stroke prediction process. This integration enables the model to maintain heightened sensitivity to local details, resulting in more refined and nuanced stroke generation. Furthermore, we incorporate adversarial training into our framework, enhancing the accuracy of stroke prediction and thereby improving the overall realism and fidelity of the synthesized paintings. Extensive qualitative evaluations, complemented by a controlled user study, validate that our DQ-Transformer surpasses existing methods in both visual realism and artistic authenticity, typically achieving these results with fewer strokes. Lingyu Liu, Yaxiong Wang, Li Zhu 0003, Lizi Liao, Zhedong Zheng |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2025 | Beyond Walking: A Large-Scale Image-Text Benchmark for Text-Based Person Anomaly SearchabstractText-based person search aims to retrieve specific individuals across camera networks using natural language descriptions. However, current benchmarks often exhibit biases towards common actions like walking or standing, neglecting the critical need for identifying abnormal behaviors in real-world scenarios. To meet such demands, we propose a new task, text-based person anomaly search, locating pedestrians engaged in both routine or anomalous activities via text. To enable the training and evaluation of this new task, we construct a large-scale image-text Pedestrian Anomaly Behavior (PAB) benchmark, featuring a broad spectrum of actions, e.g., running, performing, playing soccer, and the corresponding anomalies, e.g., lying, being hit, and falling of the same identity. The training set of PAB comprises 1,013,605 synthesized image-text pairs of both normalities and anomalies, while the test set includes 1,978 real-world image-text pairs. To validate the potential of PAB, we introduce a cross-modal pose-aware framework, which integrates human pose patterns with identity-based hard negative pair sampling. Extensive experiments on the proposed benchmark show that synthetic training data facilitates the fine-grained behavior retrieval, and the proposed pose-aware method arrives at 84.93% recall@1 accuracy, surpassing other competitive methods. The dataset, model, and code are available at https://github.com/Shuyu-XJTU/CMP. Shuyu Yang, Yaxiong Wang, Li Zhu 0003, Zhedong Zheng |
ICCV | 3 |
| 2025 | DOGR: Towards Versatile Visual Document Grounding and Referring
Yinan Zhou, Haokun Lin, Shuyu Yang, Zhongang Qi, Chen Ma 0001, Li Zhu 0003 |
ICCV | 8 |
| 2025 | Incremental self-supervised learning based on transformer for anomaly detection and localization
Wenping Jin, Fei Guo 0010, Qi Wu 0017, Li Zhu 0003 |
Eng. Appl. Artif. Intell. | 4 |
| 2025 | DMSD-CDFSAR: Distillation from Mixed-Source Domain for Cross-Domain Few-shot Action Recognition
Fei Guo 0010, Yikang Wang, Qi Han 0008, Li Zhu 0003 |
Expert Syst. Appl. | 4 |
| 2025 | Consistency Prototype Module and Motion Compensation for few-shot action recognition (CLIP-CPM2C)
Fei Guo 0010, Yikang Wang, Qi Han 0008, Li Zhu 0003 |
Neurocomputing | 4 |
| 2025 | GSLTA-CDFSAR: Global Sequences and Local Tuples Alignment for Cross-Domain Few-Shot Action Recognition
Fei Guo 0010, Qi Han 0008, Xuetao Zhang 0001, Li Zhu 0003 |
Knowl. Based Syst. | 4 |
| 2025 | Buffer-Aided-Based Resilient State Estimation for Mobile Robot Localization With State Saturation: A Probabilistic-Encoding CaseabstractIn this article, the encoding-decoding-based resilient state estimation problem is investigated for the mobile robot localization with state saturation under the buffer-aided mechanism. To reduce the impact of the intermittent measurements on estimation performance, the buffer-aided mechanism is introduced to store the untransmitted measurements when the transmission is impermissible. The unbiased probabilistic encoding–decoding approach is employed to improve the efficiency and security of the transmission. In addition, the state saturation modeled by a signum function is considered for a more accurate depiction of real-world engineering. The main objective of this article is to establish a buffer-aided-based resilient state estimator for the state-saturated mobile robot localization subject to the probabilistic encoding–decoding, such that the minimal upper bound on the estimation error covariance is guaranteed by appropriately designing the desired estimator gain. Finally, the effectiveness of the proposed estimation scheme is evaluated on a mobile robot platform with buffer capacities of$M=2, 4, 6$. The influence of different encoding interval/gain variation coefficients on the state estimation accuracy is also discussed by conducting comparison-based analyses. Furthermore, when compared to the extended Kalman filtering, the proposed estimation scheme achieves significant improvement in estimation performance while only increasing the computational time by 62.4%. Li Zhu 0003, Weiping Ding 0001 |
IEEE Trans. Ind. Informatics | 2 |
| 2025 | Vision-Based Driving Decision Making Using Multi-Action Deep Q Network
Sheng Yuan, Yaochen Li, Li Zhu 0003, Xinnan Ma, Yuncheng Xu |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2025 | Scale Up Composed Image Retrieval Learning via Modification Text GenerationabstractComposed Image Retrieval (CIR) aims to search an image of interest using a combination of a reference image and modification text as the query. Despite recent advancements, this task remains challenging due to limited training data and laborious triplet annotation processes. To address this issue, this paper proposes to synthesize the training triplets to augment the training resource for the CIR problem. Specifically, we commence by training a modification text generator exploiting large-scale multimodal models and scale up the CIR learning throughout both the pretraining and fine-tuning stages. During pretraining, we leverage the trained generator to directly create Modification Text-oriented Synthetic Triplets (MTST) conditioned on pairs of images. For fine-tuning, we first synthesize reverse modification text to connect the target image back to the reference image. Subsequently, we devise a two-hop alignment strategy to incrementally close the semantic gap between the multimodal pair and the target image. We initially learn an implicit prototype utilizing both the original triplet and its reversed version in a cycle manner, followed by combining the implicit prototype feature with the modification text to facilitate accurate alignment with the target image. Extensive experiments validate the efficacy of the generated triplets and confirm that our proposed methodology attains competitive recall on both the CIRR and FashionIQ benchmarks. Codes and datasets will be made publicly accessible. Yinan Zhou, Yaxiong Wang, Haokun Lin, Chen Ma 0001, Li Zhu 0003, Zhedong Zheng |
IEEE Trans. Multim. | 5 |
| 2024 | Forced Exploration in Bandit ProblemsabstractThe multi-armed bandit(MAB) is a classical sequential decision problem. Most work requires assumptions about the reward distribution (e.g., bounded), while practitioners may have difficulty obtaining information about these distributions to design models for their problems, especially in non-stationary MAB problems. This paper aims to design a multi-armed bandit algorithm that can be implemented without using information about the reward distribution while still achieving substantial regret upper bounds. To this end, we propose a novel algorithm alternating between greedy rule and forced exploration. Our method can be applied to Gaussian, Bernoulli and other subgaussian distributions, and its implementation does not require additional information. We employ a unified analysis method for different forced exploration strategies and provide problem-dependent regret upper bounds for stationary and piecewise-stationary settings. Furthermore, we compare our algorithm with popular bandit algorithms on different reward distributions. Qi Han 0008, Li Zhu 0003, Fei Guo 0010 |
AAAI | 2 |
| 2024 | SFMM: Semantic-to-Frame Matching with Multi-Classifier for Few-shot Action RecognitionabstractBenefiting from the rapid development of large-scale text-image comparison models, they are also widely used in corresponding downstream tasks. In traditional action recognition tasks, the typical approach is based on neural network models for prototype comparison. However, this approach limits the models’ ability to migrate to unseen or few-shot datasets and faces problems such as insufficient data and depth information. To enhance the accuracy of the few-shot action recognition task, we propose the SFMM model. This model exploits the generalization performance of the CLIP model, which is trained on a large amount of data and complements the image features with semantic information. To circumvent the randomness in model training, we employ a multi-model stacking approach to enhance recognition. Specifically, we model the task as a text-video-action category-matching problem containing different modal information. The model’s main features include the introduction of a text-matching module for full semantic coverage and an improved feature enhancement module that fuses multimodal information through convolutional and cross-attention networks. These methods allow the model to fully utilize multimodal information and obtain highly discriminative category prototypes. Our model was tested on commonly used benchmarks, demonstrating the effectiveness of our proposed method and achieving excellent performance in action recognition tasks. Yikang Wang, Fei Guo 0010, Li Zhu 0003 |
IJCNN | 3 |
| 2024 | Joint learning of video scene detection and annotation via multi-modal adaptive context network
Litong Pan, Weiguang Sang, Hailun Luo, Pingping Wei, Li Zhu 0003 |
Expert Syst. Appl. | 7 |
| 2024 | Task-Specific Alignment and Multiple-level Transformer for few-shot action recognition
Fei Guo 0010, Li Zhu 0003, Yikang Wang |
Neurocomputing | 2 |
| 2024 | Multi-view distillation based on multi-modal fusion for few-shot action recognition (CLIP-MDMF)
Fei Guo 0010, Yikang Wang, Qi Han 0008, Wenping Jin, Li Zhu 0003 |
Knowl. Based Syst. | 5 |
| 2024 | Feature Enhancement With Reverse Distillation for Hyperspectral Anomaly DetectionabstractMost hyperspectral anomaly detection methods based on trainable parameter networks require parameter adjustments or retraining on new test scenes, resulting in high time consumption and unstable performance, limiting their practicality in large-scale anomaly detection scenarios. In this letter, we address this issue by proposing a novel feature enhancement network (FEN). FEN does not directly compute anomaly scores but enhances the background features of the original hyperspectral image (HSI), thereby improving the performance of non-training-based anomaly detection algorithms, such as those based on Mahalanobis distance. To achieve this, FEN needs to be trained on a background dataset. Once trained, it does not require retraining for new detection scenes, significantly reducing time costs. Specifically, during training on background data, a complete network is trained using a reverse distillation (RD) framework with a spectral feature alignment mechanism (SFAM) to enhance the network’s ability to express background features. For inference, a pruned version of this network, which is FEN, is applied, consisting solely of components most relevant to expressing features in the spectral dimension. This design effectively reduces redundant information, enhancing both inference efficiency and anomaly detection accuracy. Experimental results demonstrate that our method can significantly improve the performance of Mahalanobis distance-based anomaly detection methods while incurring minimal time costs. Our code is available athttps://github.com/cristianoKaKa/FERD. Wenping Jin, Feng Dang, Li Zhu 0003 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2024 | ReGO: Reference-Guided Outpainting for Scenery ImageabstractWe present ReGO (Reference-Guided Outpainting), a new method for the task of sketch-guided image outpainting. Despite the significant progress made in producing semantically coherent content, existing outpainting methods often fail to deliver visually appealing results due to blurry textures and generative artifacts. To address these issues, ReGO leverages neighboring reference images to synthesize texture-rich results by transferring pixels from them. Specifically, an Adaptive Content Selection (ACS) module is incorporated into ReGO to facilitate pixel transfer for texture compensating of the target image. Additionally, a style ranking loss is introduced to maintain consistency in terms of style while preventing the generated part from being influenced by the reference images. ReGO is a model-agnostic learning paradigm for outpainting tasks. In our experiments, we integrate ReGO with three state-of-the-art outpainting models to evaluate its effectiveness. The results obtained on three scenery benchmarks, i.e. NS6K, NS8K and SUN Attribute, demonstrate the superior performance of ReGO compared to prior art in terms of texture richness and authenticity. Our code is available at https://github.com/wangyxxjtu/ReGO-Pytorch. Yaxiong Wang, Yunchao Wei, Xueming Qian, Li Zhu 0003, Yi Yang 0001 |
IEEE Trans. Image Process. | 4 |
| 2023 | Uncertainty-Aware Image CaptioningabstractIt is well believed that the higher uncertainty in a word of the caption, the more inter-correlated context information is required to determine it. However, current image captioning methods usually consider the generation of all words in a sentence sequentially and equally. In this paper, we propose an uncertainty-aware image captioning framework, which parallelly and iteratively operates insertion of discontinuous candidate words between existing words from easy to difficult until converged. We hypothesize that high-uncertainty words in a sentence need more prior information to make a correct decision and should be produced at a later stage. The resulting non-autoregressive hierarchy makes the caption generation explainable and intuitive. Specifically, we utilize an image-conditioned bag-of-word model to measure the word uncertainty and apply a dynamic programming algorithm to construct the training pairs. During inference, we devise an uncertainty-adaptive parallel beam search technique that yields an empirically logarithmic time complexity. Extensive experiments on the MS COCO benchmark reveal that our approach outperforms the strong baseline and related methods on both captioning quality as well as decoding speed. Zhengcong Fei, Mingyuan Fan 0002, Li Zhu 0003, Junshi Huang, Xiaoming Wei, Xiaolin Wei |
AAAI | 3 |
| 2023 | Masked Auto-Encoders Meet Generative Adversarial Networks and BeyondabstractMasked Auto-Encoder (MAE) pretraining methods randomly mask image patches and then train a vision Transformer to reconstruct the original pixels based on the unmasked patches. While they demonstrates impressive performance for downstream vision tasks, it generally requires a large amount of training resource. In this paper, we introduce a novel Generative Adversarial Networks alike framework, referred to as GAN-MAE, where a generator is used to generate the masked patches according to the remaining visible patches, and a discriminator is employed to predict whether the patch is synthesized by the generator. We believe this capacity of distinguishing whether the image patch is predicted or original is benefit to representation learning. Another key point lies in that the parameters of the vision Transformer backbone in the generator and discriminator are shared. Extensive experiments demonstrate that adversarial training of GAN-MAE framework is more efficient and accordingly outperforms the standard MAE given the same model size, training data, and computation resource. The gains are substantially robust for different model sizes and datasets, in particular, a ViT-B model trained with GAN-MAE for 200 epochs outperforms the MAE with 1600 epochs on fine-tuning top-1 accuracy of ImageNet-1k with much less FLOPs. Besides, our approach also works well at transferring downstream tasks. Zhengcong Fei, Mingyuan Fan 0002, Li Zhu 0003, Junshi Huang, Xiaoming Wei, Xiaolin Wei |
CVPR | 3 |
| 2023 | Towards Unified Text-based Person Retrieval: A Large-scale Multi-Attribute and Language Search BenchmarkabstractIn this paper, we introduce a large Multi-Attribute and Language Search dataset for text-based person retrieval, called MALS, and explore the feasibility of performing pre-training on both attribute recognition and image-text matching tasks in one stone. In particular, MALS contains 1,510,330 image-text pairs, which is about 37.5 × larger than prevailing CUHK-PEDES, and all images are annotated with 27 attributes. Considering the privacy concerns and annotation costs, we leverage the off-the-shelf diffusion models to generate the dataset. To verify the feasibility of learning from the generated data, we develop a new joint Attribute Prompt Learning and Text Matching Learning (APTM) framework, considering the shared knowledge between attribute and text. As the name implies, APTM contains an attribute prompt learning stream and a text matching learning stream. (1) The attribute prompt learning leverages the attribute prompts for image-attribute alignment, which enhances the text matching learning. (2) The text matching learning facilitates the representation learning on fine-grained details, and in turn, boosts the attribute prompt learning. Extensive experiments validate the effectiveness of the pre-training on MALS, achieving state-of-the-art retrieval performance via APTM on three challenging real-world benchmarks. In particular, APTM achieves a consistent improvement of +6.96 %, +7.68%, and +16.95% Recall@1 accuracy on CUHK-PEDES, ICFG-PEDES, and RSTPReid datasets by a clear margin, respectively. The dataset, model, and code are available at https://github.com/Shuyu-XJTU/APTM. Shuyu Yang, Yinan Zhou, Zhedong Zheng, Yaxiong Wang, Li Zhu 0003, Yujiao Wu |
ACM Multimedia | 5 |
| 2023 | Diff attention: A novel attention scheme for person re-identification
Li Zhu 0003, Shuyu Yang, Yaxiong Wang |
Comput. Vis. Image Underst. | 2 |
| 2023 | Generative label fused network for image-text matching
Guoshuai Zhao 0001, Heng Shang, Yaxiong Wang, Li Zhu 0003, Xueming Qian |
Knowl. Based Syst. | 5 |
| 2023 | Self-Supervised Adversarial Video Summarizer With Context Latent Sequence LearningabstractVideo summarization attempts to create concise and complete synopsis of a video through identifying the most informative and explanatory parts while removing redundant video frames, which facilitates retrieving, managing and browsing video efficiently. Most existing video summarization approaches heavily rely on enormous high-quality human-annotated labels or fail to produce semantically meaningful video summaries with the guidance of prior information. Without any supervised labels, we propose Self-supervised Adversarial Video Summarizer (2SAVS) that exploits context latent sequence learning to generate satisfying video summary. To implement it, our model elaborates a novel pretext task of identifying latent sequences and normal frames by training self-supervised generative adversarial network (GAN) with several well-designed losses. As the core components of 2SAVS, Clip Consistency Representation (CCR) and Hybrid Feature Refinement (HFR) are developed to ensure semantic consistency and continuity of clips. Furthermore, a novel separation loss is designed to explicitly enlarge the distance between prediction frame scores to effectively enhance the model’s discriminative ability. Differently, latent sequences, additional finetune operations and generators are not required when inferring video summary. Experiments on two challenging and diverse datasets demonstrate that our approach outperforms other state-of-the-art unsupervised and weakly-supervised methods, and even produces comparable results with several excellent supervised methods. Xiangshun Li, Litong Pan, Weiguang Sang, Pingping Wei, Li Zhu 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2023 | Spatial-Spectral 1DSwin Transformer With Groupwise Feature Tokenization for Hyperspectral Image ClassificationabstractThe Hyperspectral Image (HSI) classification aims to assign each pixel to a land cover category. It is receiving increasing attention from both industry and academia. The main challenge lies in capturing reliable and informative spatial and spectral dependencies concealed in the HSI for each class. To address the challenge, we propose a Spatial-Spectral 1DSwin Transformer with Group-wise Feature Tokenization (SS1DSwin) for HSI classification. Specifically, we reveal local and hierarchical spatial-spectral relationships from two different perspectives. It mainly consists of a Group-wise Feature Tokenization Module (GFTM) and a 1DSwin Transformer with Cross-block Normalized Connection Module (TCNCM). For GFTM, we reorganize an image patch into overlapping cubes, and further generate group-wise token embeddings with Multi-head Self-Attention (MSA) to learn the local spatial-spectral relationship along the spatial dimension. For TCNCM, we adopt the shifted windowing strategy when acquiring the hierarchical spatial-spectral relationship along the spectral dimension with 1D Window based Multi-head Self-Attention (1DW-MSA) and 1D Shifted Window based Multi-head Self-Attention (1DSW-MSA), and leverage Cross-block Normalized Connection (CNC) to adaptively fuse the feature maps from different blocks. In SS1DSwin, we apply these two modules in order and predict the class label for each pixel. To test the effectiveness of the proposed method, extensive experiments are conducted on four HSI datasets, and the results indicate that SS1DSwin outperforms several current state-of-the-art methods. The source code of the proposed method is available at https://github.com/Minato252/SS1DSwin. Bicheng Li, Chuanqi Xie, Yongchuan Zhang, Aichen Wang, Li Zhu 0003 |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2023 | What You Like, What I Am: Online Dating Recommendation via Matching Individual Preferences With FeaturesabstractDating recommendation becomes a critical task since the rapid development of online dating sites and it is beneficial for users to find their ideal relationships from a large number of registered members. Different users usually have different tastes when choosing their dating partners. Therefore, it is necessary to distinguish the users personal features and preferences in dating recommendation methods. However, present approaches dont capture enough user preferences from social graph and attribute data. They also ignore user attributes, which is the complementary and consistent side information of user social graphs. In this paper, we propose a Matching Individual Preferences with Features (MIPF) model to recommend dating partners jointly using user attributes and social graphs. We aim to model user features and preferences to identify what the user has and what the user likes. We also distinguish user preferences into explicit preferences and implicit preferences. The implicit preferences are mined from social graphs, while the explicit preferences are captured from the social links. Additionally, convolutional neural networks are used to extract the latent non-linear information in user attributes. Experiments on real-world online dating datasets demonstrate our MIPF model is superior to existing methods. Xuanzhi Zheng, Guoshuai Zhao 0001, Li Zhu 0003, Jihua Zhu, Xueming Qian |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2023 | Recursive Multi-Relational Graph Convolutional Network for Automatic Photo SelectionabstractAutomatic Photo Selection (APS) is a fundamental and important task for further photo cropping and photo enhancement. As the images in a photo series normally have subtle differences, it remains challenging to surface the best photos among highly similar photos. In this work, we propose a Recursive Multi-Relational Graph Convolutional Network (RMGCN) for APS. Specifically, we explore and devise inner-relation and inter-relation graphs to learn informative representations in hierarchical manner. 1) Patch-aware Intra Graph Module (PIGM) captures visual and spatial relations between different patches to characterize the representations in an image. 2) Context-aware Inter Graph Module (CIGM) explicitly exploits mutual comparative relation between different images in a photo series. These two graphs are recursively refined each other by reasoning the graph representations. Then, our model aggregates the output of CIGM with multi-scale local features via the proposed Cross-domain Fusing Gate (CFG) to boost the discriminative ability. Besides, we formulate four companion objectives as soft constraints to improve convergence rate during training. Extensive experiments are conducted on photo-triage dataset, and superior results are reported on different metrics when comparing to the state-of-the-art methods. We also perform rigorous ablations and analysis to validate our approach. Wujiang Xu, Genan Sang, Aichen Wang, Pingping Wei, Li Zhu 0003 |
IEEE Trans. Multim. | 7 |
| 2022 | Noise Learning for Text Classification: A BenchmarkabstractNoise Learning is important in the task of text classification which depends on massive labeled data that could be error-prone. However, we find that noise learning in text classification is relatively underdeveloped: 1. many methods that have been proven effective in the image domain are not explored in text classification, 2. it is difficult to conduct a fair comparison between previous studies as they do experiments in different noise settings. In this work, we adapt four state-of-the-art methods of noise learning from the image domain to text classification. Moreover, we conduct comprehensive experiments on our benchmark of noise learning with seven commonly-used methods, four datasets, and five noise modes. Additionally, most previous works are based on an implicit hypothesis that the commonly-used datasets such as TREC, Ag-News and Chnsenticorp contain no errors. However, these datasets indeed contain 0.61% to 15.77% noise labels which we define as intrinsic noise that can cause inaccurate evaluation. Therefore, we build a new dataset Golden-Chnsenticorp( G-Chnsenticorp) without intrinsic noise to more accurately compare the effects of different noise learning methods. To the best of our knowledge, this is the first benchmark of noise learning for text classification. Bo Liu 0095, Wandi Xu, Yuejia Xiang, Lejian He, Li Zhu 0003 |
COLING | 7 |
| 2022 | PERD: Personalized Emoji Recommendation with Dynamic User PreferenceabstractEmoji recommendation is an important task to help users find appropriate emojis from thousands of candidates based on a short tweet text. Traditional emoji recommendation methods lack personalized recommendation and ignore user historical information in selecting emojis. In this paper, we propose a personalized emoji recommendation with dynamic user preference (PERD) which contains a text encoder and a personalized attention mechanism. In text encoder, a BERT model is contained to learn dense and low-dimensional representations of tweets. In personalized attention, user dynamic preferences are learned according to semantic and sentimental similarity between historical tweets and the tweet which is waiting for emoji recommendation. Informative historical tweets are selected and highlighted. Experiments are carried out on two real-world datasets from Sina Weibo and Twitter. Experimental results validate the superiority of our approach on personalized emoji recommendation. Xuanzhi Zheng, Guoshuai Zhao 0001, Li Zhu 0003, Xueming Qian |
SIGIR | 3 |
| 2022 | Smart objects recommendation based on pre-training with attention and the thing-thing relationship in social Internet of things
Li Zhu 0003, Tao Dai 0002, Kaiqi Zhang 0004, Yutian Yan |
Future Gener. Comput. Syst. | 2 |
| 2022 | Combining Non-sampling and Self-attention for Sequential Recommendation
Guangjin Chen, Guoshuai Zhao 0001, Li Zhu 0003, Zhimin Zhuo, Xueming Qian |
Inf. Process. Manag. | 3 |
| 2021 | AINet: Association Implantation for Superpixel SegmentationabstractRecently, some approaches are proposed to harness deep convolutional networks to facilitate superpixel segmentation. The common practice is to first evenly divide the image into a pre-defined number of grids and then learn to associate each pixel with its surrounding grids. However, simply applying a series of convolution operations with limited receptive fields can only implicitly perceive the relations between the pixel and its surrounding grids. Consequently, existing methods often fail to provide an effective context when inferring the association map. To remedy this issue, we propose a novel Association Implantation (AI) module to enable the network to explicitly capture the relations between the pixel and its surrounding grids. The proposed AI module directly implants the grid features to the surrounding of its corresponding central pixel, and conducts convolution on the padded window to adaptively transfer knowledge between them. With such an implantation operation, the network could explicitly harvest the pixel-grid level context, which is more in line with the target of superpixel segmentation comparing to the pixelwise relation. Furthermore, to pursue better boundary precision, we design a boundary-perceiving loss to help the network discriminate the pixels around boundaries in hidden feature level, which could benefit the subsequent inferring modules to accurately identify more boundary pixels. Extensive experiments on BSDS500 and NYUv2 datasets show that our method could achieve state-of-the-art performance. Code and pre-trained model are available at https://github.com/wangyxxjtu/AINet-ICCV2021. Yaxiong Wang, Yunchao Wei, Xueming Qian, Li Zhu 0003, Yi Yang 0001 |
ICCV | 4 |
| 2021 | Low-rank and sparse matrix factorization with prior relations for recommender systems
Jie Wang 0072, Li Zhu 0003, Tao Dai 0002, Qiannan Xu |
Appl. Intell. | 2 |
| 2021 | Saliency aware image cropping with latent region pair
Wujiang Xu, Genan Sang, Pingping Wei, Li Zhu 0003 |
Expert Syst. Appl. | 7 |
| 2021 | Social image retrieval based on topic diversityabstractAbstract Image search re-ranking is one of the most important approaches to enhance the text-based image search results. Extensive efforts have been dedicated to improve the accuracy and diversity of tag-based image retrieval. However, how to make the top-ranked results relevant and diverse is still a challenging problem. In this paper, we propose a novel method to diversify the retrieval results by latent topic analysis. We first employNMF(Non-negative Matrix Factorization) Lee and Seung (Nature 401(6755):788–791, 1999) to estimate the initial relevance score to the queryq. Then, the initial relevance score is fed into an adaptive multi-feature fusion model to learn the final relevance score. Next, the diversification process is conducted. We group all the images by semantic clustering and estimate the topic distribution of each cluster by topic analysis. The clusters are ranked based on the topic distribution vector and the final retrieval image list is obtained by a greedy selection mechanism based on the estimated relevances. Experimental results on the NUS-Wide dataset show the effectiveness of the proposed approach. Yaxiong Wang, Li Zhu 0003, Xueming Qian |
Multim. Tools Appl. | 2 |
| 2021 | Towards Better Railway Service: Passengers Counting in Railway CompartmentabstractCounting passengers in railway compartments is an essential problem for improving service quality, user experience, public security, and disaster relief in the railway system. Considering many limitations in the compartment, the infrared sensor, 3D camera, etc. are not practical in this scene. Due to the flexibility and lower cost, solutions with standard cameras attract much attention in real applications. However, since the problem with scale variation in the narrow space is different from universal detection or counting problems, the specific benchmark of dataset and methods should be provided and proposed for this task. In this paper, we provide a passenger counting dataset. Relying on this dataset, we propose a passenger counting method. The solution contains a motion supervised multi-scale representation method which provides proposals against scale variation, a spatially-temporally enhanced counting which provides precise counting numbers, and a partial proposal method which conducts methods to be utilized in reality. With the proposed solution, the passengers counting task is solved in higher accuracy and practicable in the compartment environment. In experiments, the results show that all the modules in our solution are useful and efficient, and our method outperforms in comparison with others in the compartment scene. Yuanzhi Liang, Xueming Qian, Li Zhu 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2021 | Sketch-Guided Scenery Image OutpaintingabstractThe outpainting results produced by existing approaches are often too random to meet users' requirements. In this work, we take the image outpainting one step forward by allowing users to harvest personal custom outpainting results using sketches as the guidance. To this end, we propose an encoder-decoder based network to conduct sketch-guided outpainting, where two alignment modules are adopted to impose the generated content to be realistic and consistent with the provided sketches. First, we apply a holistic alignment module to make the synthesized part be similar to the real one from the global view. Second, we reversely produce the sketches from the synthesized part and encourage them be consistent with the ground-truth ones using a sketch alignment module. In this way, the learned generator will be imposed to pay more attention to fine details and be sensitive to the guiding sketches. To our knowledge, this work is the first attempt to explore the challenging yet meaningful conditional scenery image outpainting. We conduct extensive experiments on two collected benchmarks to qualitatively and quantitatively validate the effectiveness of our approach compared with the other state-of-the-art generative models. Yaxiong Wang, Yunchao Wei, Xueming Qian, Li Zhu 0003, Yi Yang 0001 |
IEEE Trans. Image Process. | 4 |
| 2021 | Semi-Supervised Pixel-Level Scene Text Segmentation by Mutually Guided NetworkabstractIn this paper we present a new data-driven method for pixel-level scene text segmentation from a single natural image. Although scene text detection, i.e. producing a text region mask, has been well studied in the past decade, pixel-level text segmentation is still an open problem due to the lack of massive pixel-level labeled data for supervised training. To tackle this issue, we incorporate text region mask as an auxiliary data into this task, considering acquiring large-scale of labeled text region mask is commonly less expensive and time-consuming. To be specific, we propose a mutually guided network which produces a polygon-level mask in one branch and a pixel-level text mask in the other. The two branches' outputs serve as guidance for each other and the whole network is trained via a semi-supervised learning strategy. Extensive experiments are conducted to demonstrate the effectiveness of our mutually guided network, and experimental results show our network outperforms the state-of-the-art in pixel-level scene text segmentation. We also demonstrate the mask produced by our network could improve the text recognition performance besides the trivial image editing application. Chuan Wang 0001, Shan Zhao 0010, Li Zhu 0003, Kunming Luo, Yanwen Guo 0001, Jue Wang 0001, Shuaicheng Liu |
IEEE Trans. Image Process. | 3 |
| 2021 | Structural Balance of Social Internet of Things Networks with Ambiguous RelationshipsabstractIn the social Internet of Things, social networks can be built among smart objects or between smart objects and people just like human beings. One of the factors that determines the effect and efficiency of service matching in SIoT is the structure of social networks. In this paper, we exploit the theory of structural balance in signed networks to optimize SIoT network structures to provide a friendly and stable environment as well as a solid foundation for service matching of social Internet of Things. Next, besides being friends or enemies, which are traditional relationships of structural balance or structural changes in signed networks, in our research, we introduce the ambiguous relationship which is the certain state between the hostile and the friendly status and we discuss the meaning and significance of ambiguous relationship in dynamical changes of structural balance in SIoT networks. Based on previous studies, we apply an enhanced objective function and a modified approach concerning the ambiguous relationship towards the dynamical change process. Experiments show that our approach is more effective and efficient than former studies in optimizing dynamical evolution of structural balance in signed networks of SIoT. Li Zhu 0003, Haifeng Du, Kaiqi Zhang 0004, Yutian Yan, Chaobo Wang |
Wirel. Commun. Mob. Comput. | 2 |
| 2020 | Deep memory network with Bi-LSTM for personalized context-aware citation recommendation
Jie Wang 0072, Li Zhu 0003, Tao Dai 0002, Yabin Wang 0001 |
Neurocomputing | 2 |
| 2020 | Aspect-based sentiment classification with multi-attention network
Qiannan Xu, Li Zhu 0003, Tao Dai 0002, Chengbing Yan |
Neurocomputing | 2 |
| 2020 | Blind Image Deblurring Based on Local Rank
Li Zhu 0003, Jihua Zhu, Zhongyu Li 0002, Huimin Lu 0001 |
Mob. Networks Appl. | 1 |
| 2020 | Multi-view point cloud registration with adaptive convergence threshold and its application in 3D model retrieval
Yaochen Li, Li Zhu 0003 |
Multim. Tools Appl. | 5 |
| 2020 | Attentive Stacked Denoising Autoencoder With Bi-LSTM for Personalized Context-Aware Citation RecommendationabstractThe rapid growth of scientific publications brings the problem of finding appropriate citations for authors. Context-aware citation recommendation is an essential technology to overcome this obstacle when given a fragment of manuscript. In this article, we propose a novel neural network model for context-aware citation recommendation by combining stacked denoising autoencoders (SDAE) and Bi-LSTM. To obtain effective embedding for cited paper, we extend SDAE into attentive SDAE (ASDAE) by utilizing the attentive information from citation context, which essentially enhance the learning ability of original SDAE. For citation context, we devise an attentive Bi-LSTM to obtain effective embedding. Specifically, the attentive Bi-LSTM is able to extract suitable citation context and recommend citations simultaneously when given a long text, which is a issue that few papers addressed before. We also integrate personalized author information to improve the performance of recommendation. Our model is essentially a seemly integration of different types of neural network with latent variables. We derive the generative process of our model, and develop a learning algorithm based on maximum a posteriori (MAP) estimation. Experimental results on the RefSeer, ANN and DBLP datasets show that our model outperforms baseline methods. Tao Dai 0002, Li Zhu 0003, Yaxiong Wang, Kathleen M. Carley |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2020 | Skeleton Filter: A Self-Symmetric Filter for Skeletonization in Noisy Text ImagesabstractRobustly computing the skeletons of objects in natural images is difficult due to the large variations in shape boundaries and the large amount of noise in the images. Inspired by recent findings in neuroscience, we propose the Skeleton Filter, which is a novel model for skeleton extraction from natural images. The Skeleton Filter consists of a pair of oppositely oriented Gabor-like filters; by applying the Skeleton Filter in various orientations to an image at multiple resolutions and fusing the results, our system can robustly extract the skeleton even under highly noisy conditions. We evaluate the performance of our approach using challenging noisy text datasets and demonstrate that our pipeline realizes state-of-the-art performance for extracting the text skeleton. Moreover, the presence of Gabor filters in the human visual system and the simple architecture of the Skeleton Filter can help explain the strong capabilities of humans in perceiving skeletons of objects, even under dramatically noisy conditions. Xiuxiu Bai, Lele Ye, Jihua Zhu, Li Zhu 0003, Taku Komura |
IEEE Trans. Image Process. | 4 |
| 2019 | VrR-VG: Refocusing Visually-Relevant RelationshipsabstractRelationships encode the interactions among individual instances and play a critical role in deep visual scene understanding. Suffering from the high predictability with non-visual information, relationship models tend to fit the statistical bias rather than ``learning" to infer the relationships from images. To encourage further development in visual relationships, we propose a novel method to mine more valuable relationships by automatically pruning visually-irrelevant relationships. We construct a new scene graph dataset named Visually-Relevant Relationships Dataset (VrR-VG) based on Visual Genome. Compared with existing datasets, the performance gap between learnable and statistical method is more significant in VrR-VG, and frequency-based analysis does not work anymore. Moreover, we propose to learn a relationship-aware representation by jointly considering instances, attributes and relationships. By applying the representation-aware feature learned on VrR-VG, the performances of image captioning and visual question answering are systematically improved, which demonstrates the effectiveness of both our dataset and features embedding schema. Both our VrR-VG dataset and representation-aware features will be made publicly available soon. Yuanzhi Liang, Yalong Bai, Wei Zhang 0031, Xueming Qian, Li Zhu 0003, Tao Mei 0001 |
ICCV | 5 |
| 2019 | Unifying Sum and Weighted Aggregations for Efficient Yet Effective Image Representation ComputationabstractEmbedding and aggregating a set of local descriptors (e.g. SIFT) into a single vector is normally used to represent images in image search. Standard aggregation operations include sum and weighted aggregations. While showing high efficiency, sum aggregation lacks discriminative power. In contrast, weighted aggregation shows promising retrieval performance but suffers extremely high time cost. In this work, we present a general mixed aggregation method that unifies sum and weighted aggregation methods. Owing to its general formulation, our method is able to balance the trade-off between retrieval quality and image representation efficiency. Additionally, to improve query performance, we propose computing multiple weighting coefficients rather than one for each to be aggregated vector by partitioning them into several components with negligible computational cost. Extensive experimental results on standard public image retrieval benchmarks demonstrate that our aggregation method achieves state-of-the-art performance while showing over ten times speedup over baselines. Shanmin Pang, Jianru Xue, Jihua Zhu, Li Zhu 0003, Qi Tian 0001 |
IEEE Trans. Image Process. | 4 |
| 2019 | Geometry and Topology Preserving Hashing for SIFT FeatureabstractIn recent years, content-based image retrieval has been of concern because of practical needs on Internet services, especially methods that can improve retrieving speed and accuracy. The SIFT feature is a well-designed local feature. It has mature applications in feature matching and retrieval, whereas the raw SIFT feature is high dimensional, with high storage cost as well as computational cost in feature similarity measurements. Thus, we propose a hashing scheme for fast SIFT feature-based image matching and retrieval. First, a training process of the hashing function involves geometric and topological information being introduced; second, a geometry-enhanced similarity evaluation that considers both the global and details of images in evaluation is explained. Compared with state-of-the-art methods, our method achieves better performance. Chen Kang, Li Zhu 0003, Xueming Qian, Junwei Han 0001, Meng Wang 0001, Yuan Yan Tang |
IEEE Trans. Multim. | 2 |
| 2018 | Large-scale vocabularies with local graph diffusion and mode seeking
Shanmin Pang, Jianru Xue, Zhanning Gao, Lihong Zheng, Li Zhu 0003 |
Signal Process. Image Commun. | 5 |
| 2018 | Joint Hypergraph Learning for Tag-Based Image RetrievalabstractAs the image sharing websites like Flickr become more and more popular, extensive scholars concentrate on tag-based image retrieval. It is one of the important ways to find images contributed by social users. In this research field, tag information and diverse visual features have been investigated. However, most existing methods use these visual features separately or sequentially. In this paper, we propose a global and local visual features fusion approach to learn the relevance of images by hypergraph approach. A hypergraph is constructed first by utilizing global, local visual features, and tag information. Then, we propose a pseudo-relevance feedback mechanism to obtain the pseudo-positive images. Finally, with the hypergraph and pseudo relevance feedback, we adopt the hypergraph learning algorithm to calculate the relevance score of each image to the query. Experimental results demonstrate the effectiveness of the proposed approach. Yaxiong Wang, Li Zhu 0003, Xueming Qian, Junwei Han 0001 |
IEEE Trans. Image Process. | 2 |
| 2017 | Image Re-Ranking Based on Topic DiversityabstractSocial media sharing Websites allow users to annotate images with free tags, which significantly contribute to the development of the web image retrieval. Tag-based image search is an important method to find images shared by users in social networks. However, how to make the top ranked result relevant and with diversity is challenging. In this paper, we propose a topic diverse ranking approach for tag-based image retrieval with the consideration of promoting the topic coverage performance. First, we construct a tag graph based on the similarity between each tag. Then, the community detection method is conducted to mine the topic community of each tag. After that, inter-community and intra-community ranking are introduced to obtain the final retrieved results. In the inter-community ranking process, an adaptive random walk model is employed to rank the community based on the multi-information of each topic community. Besides, we build an inverted index structure for images to accelerate the searching process. Experimental results on Flickr data set and NUS-Wide data sets show the effectiveness of the proposed approach. Xueming Qian, Dan Lu 0003, Yaxiong Wang, Li Zhu 0003, Yuan Yan Tang, Meng Wang 0001 |
IEEE Trans. Image Process. | 4 |
| 2016 | Automatic multi-view registration of unordered range scans without feature extraction
Jihua Zhu, Li Zhu 0003, Zhongyu Li 0002, Chen Li 0033, Jingru Cui |
Neurocomputing | 2 |