EDBT 2026 Demo / reviewers in the wild / expert
Xianglin Huang
dblp:82/6769
· DBLP profile ↗
20ranked-venue papers
1as first author
10since 2021 · last 2027
0000-0003-0324-4687ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 1 first-author · 8 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Databases, data management, data science and information retrieval · 2Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Progressive feedback-enhanced transformer for image forgery localization
Haochen Zhu, Weiheng Zhu, Gang Cao 0001, Xianglin Huang |
Expert Syst. Appl. | 4 |
| 2026 | Flow-guided cascaded transformer for consistent video colorization
Yan Zhai, Zishan Li, Zhulin Tao, Longquan Dai, Xianglin Huang |
Pattern Recognit. | 6 |
| 2026 | UDMMColor: A Unified Diffusion Model for Multi-Modal ColorizationabstractDiffusion model-based networks have been widely applied in the field of image generation and have gradually demonstrated a strong potential in image colorization tasks. However, despite the emergence of various colorization diffusion models, two major challenges remain: (1) the lack of effective control over the colorization process and (2) the prevalent issue of color bleeding. Integrating suitable conditional control can effectively alleviate these challenges. To this end, we propose a unified multi-modal diffusion model that harnesses diverse modality information to achieve flexible and high-quality colorization. Specifically, we introduce a Stroke-Adapter that extracts and integrates stroke prompt, enhancing user control over color distribution. Additionally, we design an Edge-Guided Attention mechanism to effectively inject edge information into the colorization process, significantly reducing color bleeding artifacts. Extensive comparative experiments demonstrate that our method outperforms state-of-the-art image colorization approaches in both qualitative and quantitative evaluations, achieving superior colorization results with enhanced controllability. Yan Zhai, Zerui Han, Zhulin Tao, Xianglin Huang, Jinshan Pan, Jinhui Tang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | A Lightweight and Effective Image Tampering Localization Network With Vision MambaabstractCurrent image tampering localization methods primarily rely on Convolutional Neural Networks (CNNs) and Transformers. While CNNs suffer from limited local receptive fields, Transformers offer global context modeling at the expense of quadratic computational complexity. Recently, the state space model Mamba has emerged as a competitive alternative, enabling linear-complexity global dependency modeling. Inspired by it, we propose a lightweight and effective FORensic network based on vision MAmba (ForMa) for blind image tampering localization. Firstly, ForMa captures multi-scale global features that achieves efficient global dependency modeling through linear complexity. Then the pixel-wise localization map is generated by a lightweight decoder, which employs a parameter-free pixel shuffle layer for upsampling. Additionally, a noise-assisted decoding strategy is proposed to integrate complementary manipulation traces from tampered images, boosting decoder sensitivity to forgery cues. Experimental results on 10 standard datasets demonstrate that ForMa achieves state-of-the-art generalization ability and robustness, while maintaining the lowest computational complexity. Code is available at https://github.com/multimediaFor/ForMa. Kun Guo 0010, Gang Cao 0001, Zijie Lou, Xianglin Huang, Jiaoyun Liu |
IEEE Signal Process. Lett. | 4 |
| 2025 | Multimodal Consistency Suppression Factor for Fake News DetectionabstractRecent multimodal fake news detection methods often use the consistency between textual and visual contents to determine the truth or fake of news information. Higher levels of textual-visual consistency typically lead to a greater likelihood of classifying a news item as real. However, a critical observation reveals that creators of most fake news intentionally select images that align with the textual content, thereby enhancing the credibility of the news. Consequently, high consistency between textual and visual contents alone cannot guarantee the authenticity of the information. To address this problem, we introduce a novel approach termed Multimodal Consistency-based Suppression Factor to modulate the significance of textual-visual consistency in information assessment. When the textual-visual matching is high, this suppression factor reduces the influence of consistency during the judgment process. Moreover, we use contrastive language-image pre-training (CLIP) model to extract features and measure the consistency level between modalities to guide multimodal fusion. In addition, we also use a method of compressing and fusing modal information based on variational autoencoder (VAE) to reconstruct CLIP features, learning the shared representation of different modal information of CLIP. Finally, extensive experiments were conducted on three publicly datasets, Weibo, Twitter, and Weibo21, and the results confirmed that our method outperformed the state-of-the-art methods in the field and had 0.8%, 2.6%, and 4.1% effect improvement on the accuracy rate. Zhulin Tao, Xingyu Gao 0001, Xi Wang 0014, Xianglin Huang |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2024 | Attention-enhanced joint learning network for micro-video venue classification
Bing Wang 0013, Xianglin Huang, Gang Cao 0001, Lifang Yang, Zhulin Tao |
Multim. Tools Appl. | 2 |
| 2023 | Flow-Guided Transformer for Video ColorizationabstractVideo colorization aims to add color to black-and-white films. However, propagating color information to the whole video clip accurately is a challenging task. In this paper, we propose Flow-Guided Transformer for Video Colorization (FGTVC), consisting of a Global Motion Aggregation (GMA) module, Residual modules, Flow-Guided Attention blocks (FGAB) based on encoder and decoder, to exploit the information from the neighbor patch with high similarity for each video patch colorization. Specifically, we employ Transformer to capture the long-distance dependencies between frames and learn non-local self-similarity in the frame. To overcome the shortcomings of previous optical flow-based methods, FGAB enjoys the guidance of optical flow to sample elements from spatio-temporal adjacent frames when calculating self-attention. Experiments show that the proposed FGTVC has an outstanding performance than the state-of-the-art methods. In addition, comprehensive findings demonstrate the superiority of our framework in real-world video colorization tasks. Yan Zhai, Zhulin Tao, Longquan Dai, He Wang 0054, Xianglin Huang, Lifang Yang |
ICIP | 5 |
| 2023 | Self-Supervised Learning for Multimedia RecommendationabstractLearning representations for multimedia content is critical for multimedia recommendation. Current representation learning methods roughly fall into two groups: (1) using the historical interactions to create ID embeddings of users and items, and (2) treating multi-modal data as the side information of items to enrich their ID embeddings. Each user-item interaction offers the supervisory signal to optimize the representation learning by the traditional supervised learning paradigm. Due to the overlook of the multi-modal patterns ($e.g.$, co-occurrence of visual, acoustic, textual features in micro-videos a user saw before, and her behavioral features) hidden in the data, these methods are insufficient to create powerful representations and obtain satisfactory recommendation accuracy. To capture multi-modal patterns in the data itself, we go beyond the supervised learning paradigm, and incorporate the idea of self-supervised learning (SSL) into multimedia recommendation. Specifically, SSL consists of two components: (1) data augmentation upon multi-modal contents, where we design three operators — feature dropout (FD), feature masking (FM), feature fine and coarse spaces (FAC) — to generate multiple views of individual items; and (2) contrastive learning, which differentiates the views of an item from the others’ to distill additional supervisory signals. Clearly, SSL enables us to explore and exhibit the underlying relations among modalities, thereby resulting in powerful representations. We denote the generic framework by Self-supervised Learning-guided Multimedia Recommendation (SLMRec). Extensive experiments are performed on three real-world datasets, showing that SLMRec achieves significant improvements over several state-of-the-art baselines like LightGCN [1], MMGCN [2]. Further analysis shows how SSL affects recommendation performance. Zhulin Tao, Xiaohao Liu, Yewei Xia, Xiang Wang 0010, Lifang Yang, Xianglin Huang, Tat-Seng Chua |
IEEE Trans. Multim. | 6 |
| 2022 | EliMRec: Eliminating Single-modal Bias in Multimedia RecommendationabstractThe main idea of multimedia recommendation is to introduce the profile content of multimedia documents as an auxiliary, so as to endow recommenders with generalization ability and gain better performance. However, recent studies using non-uniform datasets roughly fuse single-modal features into multi-modal features and adopt the strategy of directly maximizing the likelihood of user preference scores, leading to the single-modal bias. Owing to the defect in architecture, there is still room for improvement for recent multimedia recommendation. Xiaohao Liu, Zhulin Tao, Jiahong Shao, Lifang Yang, Xianglin Huang |
ACM Multimedia | 5 |
| 2022 | Hybrid Transformer-CNN for Real Image DenoisingabstractTransformer typically enjoys larger model capacity but higher computational loads than convolutional neural network (CNN) in vision tasks. In this letter, the advantages of such two networks are fused for achieving effective and efficient real image denoising. We propose a hybrid denoising model based on Transformer Encoder and Convolutional Decoder Network (TECDNet). The Transformer based on novel radial basis function (RBF) attention is used as encoder to improve the representation capability of overall model. In decoder, the residual CNN instead of Transformer is adopted to greatly reduce computational complexity of the whole denoising network. Extensive experimental results on real images show that TECDNet achieves the state-of-the-art denosing performance with relatively low computational cost. Mo Zhao, Gang Cao 0001, Xianglin Huang, Lifang Yang |
IEEE Signal Process. Lett. | 3 |
| 2020 | HoAFM: A High-order Attentive Factorization Machine for CTR Prediction
Zhulin Tao, Xiang Wang 0010, Xiangnan He 0001, Xianglin Huang, Tat-Seng Chua |
Inf. Process. Manag. | 4 |
| 2020 | MGAT: Multimodal Graph Attention Network for Recommendation
Zhulin Tao, Yinwei Wei, Xiang Wang 0010, Xiangnan He 0001, Xianglin Huang, Tat-Seng Chua |
Inf. Process. Manag. | 5 |
| 2020 | Multi-modal sequence model with gated fully convolutional blocks for micro-video venue classification
Wei Liu 0084, Xianglin Huang, Gang Cao 0001, Jianglong Zhang, Gege Song, Lifang Yang |
Multim. Tools Appl. | 2 |
| 2019 | Better Word Representations with Word WeightabstractAs a fundamental task of natural language processing, text classification has been widely used in various applications such as sentiment analysis and spam detection. In recent years, the continuous-valued word embedding learned by neural network attaches extensive attentions. Although word embedding achieves impressive results in capturing similarities and regularities between words, it fails to highlight important words for identifying text category. Such deficiency could be attenuated by word weight, which conveys word contribution in text categorization. Toward this end, we propose an effective text classification scheme by incorporating word weight into word embedding in this paper. Specifically, in order to enrich word representation, the bidirectional gated recurrent units (Bi-GRU) is first employed to grasp context information of words. Then the word weights yielded by term frequency (TF) are used to modulate the word representation of Bi-GRU for constructing text representation. Extensive experimental results on several large text datasets verify that the accuracy of our proposed text classification scheme outperforms the state-of-the-art ones. Gege Song, Xianglin Huang, Gang Cao 0001, Zhulin Tao, Wei Liu 0084, Lifang Yang |
MMSP | 2 |
| 2018 | Acceleration of histogram-based contrast enhancement via selective downsamplingabstractThe authors propose a general framework to accelerate the universal histogram‐based image contrast enhancement (CE) algorithms. Both spatial and grey‐level selective downsampling of digital images are adopted to decrease computational cost, while the visual quality of enhanced images is still preserved and without apparent degradation. Mapping function calibration is proposed to reconstruct the pixel mapping on the grey levels missed by downsampling. As two case studies, the accelerations of histogram equalisation (HE) and the state‐of‐the‐art global CE algorithm, i.e. spatial mutual information and PageRank (SMIRANK), are presented in detail. Both quantitative and qualitative assessment results have verified the effectiveness of their proposed CE acceleration framework. In typical tests, the computational efficiencies of HE and SMIRANK have been increased by about 3.9 and 13.5 times, respectively. Gang Cao 0001, Huawei Tian, Lifang Yu, Xianglin Huang, Yongbin Wang |
IET Image Process. | 4 |
| 2017 | Feature selection and feature learning in arousal dimension of music emotion by using shrinkage methods
Jianglong Zhang, Xianglin Huang, Lifang Yang, Shutao Sun |
Multim. Syst. | 2 |
| 2016 | A bi-directional sampling based on K-means method for imbalance text classificationabstractThis paper studies the imbalanced data classify-cation problem and proposes bi-directional sampling based on clustering (BDSK) for the imbalanced data classification. This algorithm combines SMOTE over-sampling algorithm and under-sampling algorithm based on K-Means to solve the within-class imbalance problem and the between-class imbalance problem. It not only avoid induce too much noise but also resolve the problem of shortage of sample. Experimental results on Tan corpus dataset show that the algorithm can effectively improve the classification performance on imbalanced data sets, especially in the cases when classification performance is heavily affected by class imbalance. Xianglin Huang, Sijun Qin |
ICIS | 2 |
| 2016 | Local visual similarity descriptor for describing local regionabstractMany works have devoted to exploring local region information including both the information of the local features in local region and their spatial relationships, but none of these can provide a compact representation of the information. To achieve this, we propose a new approach named Local Visual Similarity (LVS). LVS first calculates the similarities among the local features in a local region and then forms these similarities as a single vector named LVS descriptor. In our experiments, we show that LVS descriptor can preserve local region information with low dimensionality. Besides, experimental results on two public datasets also demonstrate the effectiveness of LVS descriptor. Xianglin Huang, Lifang Yang |
ICMV | 1 |
| 2016 | Shorter-is-Better: Venue Category Estimation from Micro-VideoabstractAccording to our statistics on over 2 million micro-videos, only 1.22% of them are associated with venue information, which greatly hinders the location-oriented applications and personalized services. To alleviate this problem, we aim to label the bite-sized video clips with venue categories. It is, however, nontrivial due to three reasons: 1) no available benchmark dataset; 2) insufficient information, low quality, and 3) information loss; and 3) complex relatedness among venue categories. Towards this end, we propose a scheme comprising of two components. In particular, we first crawl a representative set of micro-videos from Vine and extract a rich set of features from textual, visual and acoustic modalities. We then, in the second component, build a tree-guided multi-task multi-modal learning model to estimate the venue category for each unseen micro-video. This model is able to jointly learn a common space from multi-modalities and leverage the predefined Foursquare hierarchical structure to regularize the relatedness among venue categories. Extensive experiments have well-validated our model. As a side research contribution, we have released our data, codes and involved parameters. Jianglong Zhang, Liqiang Nie, Xiang Wang 0010, Xiangnan He 0001, Xianglin Huang, Tat-Seng Chua |
ACM Multimedia | 5 |
| 2016 | Bridge the semantic gap between pop music acoustic feature and emotion: Build an interpretable model
Jianglong Zhang, Xianglin Huang, Lifang Yang, Liqiang Nie |
Neurocomputing | 2 |