EDBT 2026 Demo / reviewers in the wild / expert
Deyin Liu
dblp:160/6212
· DBLP profile ↗
25ranked-venue papers
9as first author
21since 2021 · last 2026
0000-0002-0371-9921ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 7 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 2 first-author · 10 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Beyond consistency: Preserving temporal structure in zero-shot video editing
Deyin Liu, Yisheng Ding, Zhe Jin 0001, Xiatian Zhu, Anjan Dutta 0001, Lin Wu 0001 |
Pattern Recognit. | 1 |
| 2026 | HAViG: Hierarchical adaptive visual grounding framework for video question answering
Lei Zhu 0005, Lingmin Pan, Siqiao Tan, Chengyuan Zhang 0001, Deyin Liu, Lin Wu 0001, Farid Boussaïd, Mohammed Bennamoun |
Pattern Recognit. | 5 |
| 2026 | Towards efficient pixel labeling for industrial anomaly detection and localization
Jingqi Wu, Lin Wu 0001, Hao Chen 0041, Deyin Liu, Haiqiang Jin |
Pattern Recognit. Lett. | 5 |
| 2026 | Accurate Industrial Anomaly Detection and Localization Using Weakly-Supervised Residual TransformersabstractRecent advancements in industrial anomaly detection (AD) have demonstrated that incorporating a small number of anomalous samples during training can significantly enhance accuracy. However, this improvement often comes at the cost of extensive annotation efforts, which are impractical for many real-world applications. In this paper, we introduce a novel framework, "Weakly-supervised RESidual $T$ ransformer" (WeakREST), designed to achieve high anomaly detection accuracy while minimizing the reliance on manual annotations. First, we reformulate the pixel-wise anomaly localization task into a block-wise classification problem. Second, we introduce a residual-based feature representation called "Positional $F$ ast $A$ nomaly $R$ esiduals" (PosFAR) which captures anomalous patterns more effectively. To leverage this feature, we adapt the Swin Transformer for enhanced anomaly detection and localization. Additionally, we propose a weak annotation approach utilizing bounding boxes and image tags to define anomalous regions. This approach establishes a semi-supervised learning context that reduces the dependency on precise pixel-level labels. To further improve the learning process, we develop a novel ResMixMatch algorithm, capable of handling the interplay between weak labels and residual-based representations. On the benchmark dataset MVTec-AD, our method achieves an Average Precision (AP) of 83.0%, surpassing the previous best result of 82.7% in the unsupervised setting. In the supervised AD setting, WeakREST attains an AP of 87.6%, outperforming the previous best of 86.0%. Notably, even when using weaker annotations such as bounding boxes, WeakREST exceeds the performance of leading methods relying on pixel-wise supervision, achieving an AP of 87.1% compared to the prior best of 86.0% on MVTec-AD. This superior performance is consistently replicated across other well-established AD datasets, including MVTec 3D, KSDD2 and Real-IAD. Code is available at: https://github.com/BeJane/Semi_REST. Jingqi Wu, Deyin Liu, Lin Wu 0001, Hao Chen 0041, Chunhua Shen |
IEEE Trans. Image Process. | 3 |
| 2026 | Multi-Modal Refined Prompting for Advancing Knowledge-Based Visual Question AnsweringabstractKnowledge-based Visual Question Answering (KB-VQA) has surfaced as a critical task in advancing AI capabilities. Despite significant progress enabled by large language models (LLMs), there are still three major challenges: (1) flawed image captions cause unreliable reasoning; (2) noisy explicit knowledge can disrupts answering; and (3) massive LLMs scale is irreplaceable to robustness. To overcome these challenges, we develop a novel approach, Multi-Modal Refined Prompting (MMRP), which generates high-quality prompts tailored for LLMs. To tackle the first challenge, a multi-faceted image captioning strategy is employed to generate detailed, contextually relevant visual descriptions. In addition, we introduce a complementary knowledge retrieval and refinement strategy to deliver concise, contextually relevant knowledge, effectively overcoming the second challenge. These enhanced image captions and explicit knowledge are then integrated into a knowledge-infused in-context prompt, effectively activating the reasoning capabilities of LLMs. Importantly, MMRP eliminates reliance on massive LLMs and avoids the need for model fine-tuning, while achieving significant improvements in answer accuracy. Extensive evaluations on the widely-used OK-VQA benchmark against 22 baselines prove the superiority of MMRP, establishing a new state-of-the-art in KB-VQA. Lei Zhu 0005, Mengxi Ying, Chengyuan Zhang 0001, Deyin Liu, Lin Wu 0001, Shichao Zhang 0001, Xuelong Li 0001 |
IEEE Trans. Multim. | 4 |
| 2025 | EntroFormer: An entropy-based sparse vision transformer for real-time semantic segmentation
Song Wang 0008, Lin Wu 0001, Deyin Liu, Lei Gao 0001, Lin Qi 0001, Guanghui Wang 0001 |
Comput. Vis. Image Underst. | 4 |
| 2025 | T-Person-GAN: Text-to-Person image generation with identity-consistency and manifold mix-up
Deyin Liu, Lin Wu 0001, Bo Li 0090, Ye Zhao 0001, ZongYuan Ge |
Expert Syst. Appl. | 1 |
| 2025 | Textual semantics enhancement adversarial hashing for cross-modal retrieval
Lei Zhu 0005, Runbing Wu, Deyin Liu, Chengyuan Zhang 0001, Lin Wu 0001, Ying Zhang 0001, Shichao Zhang 0001 |
Knowl. Based Syst. | 3 |
| 2025 | A Deep Semantic Segmentation Network With Semantic and Contextual RefinementsabstractSemantic segmentation is a fundamental task in multimedia processing, which can be used for analyzing, understanding, editing contents of images and videos, among others. To accelerate the analysis of multimedia data, existing segmentation researches tend to extract semantic information by progressively reducing the spatial resolutions of feature maps. However, this approach introduces a misalignment problem when restoring the resolution of high-level feature maps. In this paper, we design a Semantic Refinement Module (SRM) to address this issue within the segmentation network. Specifically, SRM is designed to learn a transformation offset for each pixel in the upsampled feature maps, guided by high-resolution feature maps and neighboring offsets. By applying these offsets to the upsampled feature maps, SRM enhances the semantic representation of the segmentation network, particularly for pixels around object boundaries. Furthermore, a Contextual Refinement Module (CRM) is presented to capture global context information across both spatial and channel dimensions. To balance dimensions between channel and space, we aggregate the semantic maps from all four stages of the backbone to enrich channel context information. The efficacy of these proposed modules is validated on three widely used datasets—Cityscapes, Bdd100 K, and ADE20K—demonstrating superior performance compared to state-of-the-art methods. Additionally, this paper extends these modules to a lightweight segmentation network, achieving an mIoU of 82.5% on the Cityscapes validation set with only 137.9 GFLOPs. Deyin Liu, Lin Wu 0001, Song Wang 0008, Xin Guo 0005, Lin Qi 0001 |
IEEE Trans. Multim. | 2 |
| 2025 | Deformation Field Fusion for Medical Image RegistrationabstractDeformable medical image registration is to find a series of non-linear spatial transformations to align a pair of fixed and moving voxel images. Deep learning based registration models are effective in learning differences between such image pair to obtain the deformation field which is specialized in describing non-rigid deformations in the 3D voxel context. However, existing models tend to learn either one single deformation field only or multi-stage (multi-level) deformation fields progressively arriving at a final optimal field. Actually, deformation fields resulting from different architectures or losses are capable of capturing diverse types of deformations, complementing to each other. In this article, we propose a novel framework of fusing different deformation fields to acquire an overall field to describe all-round deformations, in which multiple complementary cues regarding deformable 3D voxels can be strategically leveraged to improve the alignment of the given image pair. The key to the effect of deformation field fusion for registration lies in two aspects: the fusion network architecture and the loss function. Thus, we develop a well-designed fusion block using ingenious operations based on different types of pooling, convolution, and concatenation. Moreover, since calculating the deformation field using a conventional similarity loss cannot describe the contextual variations which are inter-dependent in each pair of fixed and moving images, we propose a novel Contrast-Structural loss to enhance the motion displacement between the image pair by calculating the similarity of pixels in density values, while being ranged in their spatial proximity. Extensive experimental results demonstrate that our proposed method achieves state-of-the-art performance on currently mainstream benchmark datasets. Haifeng Zhao 0001, Chi Zhang 0082, Deyin Liu, Lin Wu 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2024 | Towards Efficient Sparse Transformer based Medical Image RegistrationabstractDeformable medical image registration is a crucial task that involves extracting and aligning features from two images to establish precise correspondence, essentially for accurate registration. While visual transformers have propelled recent advancements in medical image analysis, training and inference with Transformers can become excessively computationally expensive, particularly due to the quadratic complexity of self-attention when handling long sequences of representations. This challenge becomes more pronounced in 3D medical image registration tasks. To tackle this issue, we propose an efficient Hierarchical Pyramid Converter for medical image registration. The proposed approach firstly capitalizes on the observation that early self-attention layers in Transformers mainly emphasize local patterns, though with limited benefits. Specifically, we employ the plain multi-layer perceptrons (MLP), i.e., Spatial shift MLP (S-MLP), in the early stages of feature extraction. This module employs a spatial offset operation to facilitate communication between patches, encoding rich local patterns and effectively reducing computational expenses. We further propose a sparse Transformer block that adaptively selects and preserves the most valuable self-attention values for feature extraction. We introduce a learnable top-k selection operator, allowing the model to selectively retain attention scores that contribute the most to each query keyword. This innovation significantly enhances feature extraction in later stages. We conducted extensive evaluations using publicly available datasets, and the experimental results confirm that our proposed method achieves state-of-the-art performance in deformable medical image registration tasks. Haifeng Zhao 0001, Quanshuang He, Deyin Liu |
CSCWD | 3 |
| 2024 | Efficient Coupling Streaming AI and Ensemble Simulations on HPC Clusters
Jiazhi Jiang, Hongbin Zhang 0006, Deyin Liu, Jiangsu Du, Xiaojiao Yao, Jinhui Wei, Pin Chen, Dan Huang 0001, Yutong Lu |
Euro-Par (1) | 3 |
| 2024 | Jacobian norm with Selective Input Gradient Regularization for interpretable adversarial defenseabstractDeep neural networks (DNNs) can be easily deceived by imperceptible alterations known as adversarial examples. These examples can lead to misclassification , posing a significant threat to the reliability of deep learning systems in real-world applications. Adversarial training (AT) is a popular technique used to enhance robustness by training models on a combination of corrupted and clean data. However, existing AT-based methods often struggle to handle transferred adversarial examples that can fool multiple defense models, thereby falling short of meeting the generalization requirements for real-world scenarios. Furthermore, AT typically fails to provide interpretable predictions, which are crucial for domain experts seeking to understand the behavior of DNNs. To overcome these challenges, we present a novel approach called Jacobian norm and Selective Input Gradient Regularization (J-SIGR). Our method leverages Jacobian normalization to improve robustness and introduces regularization of perturbation-based saliency maps, enabling interpretable predictions. By adopting J-SIGR, we achieve enhanced defense capabilities and promote high interpretability of DNNs. We evaluate the effectiveness of J-SIGR across various architectures by subjecting it to powerful adversarial attacks. Our experimental evaluations provide compelling evidence of the efficacy of J-SIGR against transferred adversarial attacks, while preserving interpretability. The project code can be found at https://github.com/Lywu-github/jJ-SIGR.git . Deyin Liu, Lin Wu 0001, Bo Li 0090, Farid Boussaïd, Mohammed Bennamoun, Xianghua Xie, Chengwu Liang |
Pattern Recognit. | 1 |
| 2023 | Image Template Matching via Dense and Consistent Contrastive LearningabstractImage template matching refers to localizing a small query image as opposed to a large reference image map. The query image a.k.a template has to be screened across every equal-sized region in the reference map to perform inner-product at pixel-level and the resulting similarity indicates the template location. Due to the domain heterogeneity between template and reference images, the matching performance degrades under dramatic appearance changes. More severely, the asymmetric matching easily leads to over-fitting by suggesting excessively false positive regions. To these ends, we propose an effective template matching method based on contrastive learning to perform a dense and consistent InfoNCEloss during matching. This can increase the matching at finer details, and thus effectively regularizes network training to prevent over-fitting. Extensive experiments on the synthetic aperture radar (SAR) and optical datasets, i.e., SEN1-2 and OS datasets demonstrate that our proposed method outperforms state-of-the-art methods by a large margin. Bo Li 0090, Lin Wu 0001, Deyin Liu, Hongyang Chen 0001, Yuanxin Ye, Xianghua Xie |
ICME | 3 |
| 2023 | LipFormer: Learning to Lipread Unseen Speakers Based on Visual-Landmark TransformersabstractLipreading refers to understanding and further translating the speech of a video speaker into textual outputs. State-of-the-art lipreading methods excel in interpreting overlap speakers, i.e., speakers appear in both training and inference. However, generalizing those methods to unseen speakers incurs catastrophic performance degradation due to the limited number of speakers in training bank as well as the dominant visual variations caused by the shape/color of lips presented by different speakers. Therefore, merely depending on the visible changes of lips tends to overfit the model. To improve to generalise, in this paper we propose to use multi-modal features, i.e., visual and landmark, to describe the lip motion while being irrespective to speaker characteristics. The proposed sentence-level framework, dubbed LipFormer, is based on visual-landmark transformer architecture wherein a lip motion stream, a facial landmark stream, and a cross-modal fusion are interconnected. More specifically, the two-stream embeddings produced by self-attention are prompted into a cross-attention module to achieve the alignment across visual and landmark variations. The resulting fused features are decoded into linguistic texts by a cascaded sequence-to-sequence translation. Extensive experiments demonstrate that our method can generalise well to unseen speakers in multiple datasets. Feng Xue 0002, Yu Li 0053, Deyin Liu, Yincen Xie, Lin Wu 0001, Richang Hong |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Asymmetric Cross-Scale Alignment for Text-Based Person SearchabstractText-based person search (TBPS) is of significant importance in intelligent surveillance, which aims to retrieve pedestrian images with high semantic relevance to a given text description. This retrieval task is characterized with both modal heterogeneity and fine-grained matching. To implement this task, one needs to extract multi-scale features from both image and text domains, and then perform the cross-modal alignment. However, most existing approaches only consider the alignment confined at their individual scales, e.g., an image-sentence or a region-phrase scale. Such a strategy adopts the presumable alignment in feature extraction, while overlooking the cross-scale alignment, e.g., image-phrase. In this paper, we present a transformer-based model to extract multi-scale representations, and perform Asymmetric Cross-Scale Alignment (ACSA) to precisely align the two modalities. Specifically, ACSA consists of a global-level alignment module and an asymmetric cross-attention module, where the former aligns an image and texts on a global scale, and the latter applies the cross-attention mechanism to dynamically align the cross-modal entities in region/image-phrase scales. Extensive experiments on two benchmark datasets CUHK-PEDES and RSTPReid demonstrate the effectiveness of our approach. Zhong Ji, Junhua Hu, Deyin Liu, Lin Wu 0001, Ye Zhao 0001 |
IEEE Trans. Multim. | 3 |
| 2023 | Verbal-Person Nets: Pose-Guided Multi-Granularity Language-to-Person GenerationabstractPerson image generation conditioned on natural language allows us to personalize image editing in a user-friendly manner. This fashion, however, involves different granularities of semantic relevance between texts and visual content. Given a sentence describing an unknown person, we propose a novel pose-guided multi-granularity attention architecture to synthesize the person image in an end-to-end manner. To determine what content to draw at a global outline, the sentence-level description and pose feature maps are incorporated into a U-Net architecture to generate a coarse person image. To further enhance the fine-grained details, we propose to draw the human body parts with highly correlated textual nouns and determine the spatial positions with respect to target pose points. Our model is premised on a conditional generative adversarial network (GAN) that translates language description into a realistic person image. The proposed model is coupled with two-stream discriminators: 1) text-relevant local discriminators to improve the fine-grained appearance by identifying the region-text correspondences at the finer manipulation and 2) a global full-body discriminator to regulate the generation via a pose-weighting feature selection. Extensive experiments conducted on benchmarks validate the superiority of our method for person image generation. Deyin Liu, Lin Wu 0001, Feng Zheng 0001, Lingqiao Liu, Meng Wang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | Generative Metric Learning for Adversarially Robust Open-world Person Re-IdentificationabstractThe vulnerability of re-identification (re-ID) models under adversarial attacks is of significant concern as criminals may use adversarial perturbations to evade surveillance systems. Unlike a closed-world re-ID setting (i.e., a fixed number of training categories), a reliable re-ID system in the open world raises the concern of training a robust yet discriminative classifier, which still shows robustness in the context of unknown examples of an identity. In this work, we improve the robustness of open-world re-ID models by proposing a generative metric learning approach to generate adversarial examples that are regularized to produce robust distance metric. The proposed approach leverages the expressive capability of generative adversarial networks to defend the re-ID models against feature disturbance attacks. By generating the target people variants and sampling the triplet units for metric learning, our learned distance metrics are regulated to produce accurate predictions in the feature metric space. Experimental results on the three re-ID datasets, i.e., Market-1501, DukeMTMC-reID, and MSMT17 demonstrate the robustness of our method. Deyin Liu, Lin Wu 0001, Richang Hong, ZongYuan Ge, Jialie Shen 0001, Farid Boussaïd, Mohammed Bennamoun |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2022 | Multi-scale Spatial Representation Learning via Recursive Hermite Polynomial NetworksabstractMulti-scale representation learning aims to leverage diverse features from different layers of Convolutional Neural Networks (CNNs) for boosting the feature robustness to scale variance. For dense prediction tasks, two key properties should be satisfied: the high spatial variance across convolutional layers, and the sub-scale granularity inside a convolutional layer for fine-grained features. To pursue the two properties, this paper proposes Recursive Hermite Polynomial Networks (RHP-Nets for short). The proposed RHP-Nets consist of two major components: 1) a dilated convolution to maintain the spatial resolution across layers, and 2) a family of Hermite polynomials over a subset of dilated grids, which recursively constructs sub-scale representations to avoid the artifacts caused by naively applying the dilation convolution. The resultant sub-scale granular features are fused via trainable Hermite coefficients to form the multi-resolution representations that can be fed into the next deeper layer, and thus allowing feature interchanging at all levels. Extensive experiments are conducted to demonstrate the efficacy of our design, and reveal its superiority over state-of-the-art alternatives on a variety of image recognition tasks. Besides, introspective studies are provided to further understand the properties of our method. Yuanbo Lin Wu, Deyin Liu, Xiaojie Guo 0001, Richang Hong |
IJCAI | 2 |
| 2022 | Pseudo-Pair Based Self-Similarity Learning for Unsupervised Person Re-IdentificationabstractPerson re-identification (re-ID) is of great importance to video surveillance systems by estimating the similarity between a pair of cross-camera person shorts. Current methods for estimating such similarity require a large number of labeled samples for supervised training. In this paper, we present a pseudo-pair based self-similarity learning approach for unsupervised person re-ID without human annotations. Unlike conventional unsupervised re-ID methods that use pseudo labels based on global clustering, we construct patch surrogate classes as initial supervision, and propose to assign pseudo labels to images through the pairwise gradient-guided similarity separation. This can cluster images in pseudo pairs, and the pseudos can be updated during training. Based on pseudo pairs, we propose to improve the generalization of similarity function via a novel self-similarity learning:it learns local discriminative features from individual images via intra-similarity, and discovers the patch correspondence across images via inter-similarity. The intra-similarity learning is based on channel attention to detect diverse local features from an image. The inter-similarity learning employs a deformable convolution with a non-local block to align patches for cross-image similarity. Experimental results on several re-ID benchmark datasets demonstrate the superiority of the proposed method over the state-of-the-arts. Lin Wu 0001, Deyin Liu, Dapeng Chen, ZongYuan Ge, Farid Boussaïd, Mohammed Bennamoun, Jialie Shen 0001 |
IEEE Trans. Image Process. | 2 |
| 2021 | Auto-encoder based structured dictionary learning for visual classification
Deyin Liu, Chengwu Liang, Shaokang Chen, Tie Yun, Lin Qi 0001 |
Neurocomputing | 1 |
| 2020 | Auto-Encoder based Structured Dictinoary LearningabstractDictionary learning and deep learning are two popular representation learning paradigms, which can be combined to boost the classification task. However, existing combination methods often learn multiple dictionaries embedded in a cascade of layers, and a specialized classifier accordingly. This may inattentively lead to overfitting and high computational cost. In this paper, we present a novel deep auto-encoding architecture to learn only a dictionary for classification. To empower the dictionary with discrimination, we construct the dictionary with class-specific sub-dictionaries, and introduce supervision by imposing category constraints. The proposed framework is inspired by a sparse optimization method, namely Iterative Shrinkage Thresholding Algorithm, which characterizes the learning process by the forward-propagation based optimization w.r.t the dictionary only, reducing the number of parameters to learn and the computational cost dramatically. Extensive experiments demonstrate the effectiveness of our method in image classification. Deyin Liu, Yuanbo Lin Wu, Qichang Hu, Lin Qi 0001 |
MMSP | 1 |
| 2020 | Medi-Care AI: Predicting medications from billing codes via robust recurrent neural networks
Deyin Liu, Lin Wu 0001, Xue Li 0001, Lin Qi 0001 |
Neural Networks | 1 |
| 2020 | Multi-task image set classification via joint representation with class-level sparsity and intra-task low-rankness
Deyin Liu, Tie Yun, Lin Qi 0001 |
Pattern Recognit. Lett. | 1 |
| 2017 | Foreign Exchange Rates Forecasting with Convolutional Neural Network
Chen Liu 0004, Weiyan Hou, Deyin Liu |
Neural Process. Lett. | 3 |