VLDB 2026 Research / reviewers in the wild / expert
Xueliang Liu
dblp:08/6294
· DBLP profile ↗
63ranked-venue papers
16as first author
31since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 41 · 11 first-author · 21 since 2021Databases, data management, data science and information retrieval · 12 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 10 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 first-author · 2 since 2021Computer networks · 5 · 1 first-author · 2 since 2021Security and privacy · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CETalk: Continuous Valence-Arousal Control for Audio-Driven 3D Talking Head Generation
Xueliang Liu |
ICIC (6) | 4 |
| 2026 | Fine-grained talking motion consistent network for realistic face generation
Jinlin Guo, Xueliang Liu |
Eng. Appl. Artif. Intell. | 3 |
| 2026 | HMR-net: Hierarchical multi-scale relation network for visible-infrared person re-identification
Ruixiang Quan, Xiaohong Li 0002, Shuo Zhuang, Xueliang Liu |
Neurocomputing | 4 |
| 2026 | Class-aware prototype augmentation and decoupled feature distillation for class-incremental learning
Chengdong Wang, Yangjun Ou, Xianfang Tang, Yuan Wu 0007, Wuxuan Shi, Xueliang Liu |
Pattern Recognit. | 6 |
| 2026 | AttriPrompt: Class Attribute-Aware Prompt Tuning for Vision-Language ModelabstractPrompt tuning has proven to be an effective alternative for fine-tuning the pre-trained vision-language models (VLMs) to downstream tasks. Among existing approaches, class-shared prompts learn a unified prompt shared across all classes, while sample-specific prompts generate distinct prompts tailored to each individual sample. However, both approaches often struggle to adequately capture the unique characteristics of underrepresented classes, particularly in imbalanced scenarios where data for tail classes is scarce. To alleviate this issue, we propose an attribute-aware prompt tuning framework that prompts a more balanced understanding for imbalance tasks by explicitly modeling critical class-level attributes. The key intuition is that, from the perspective of class, essential attributes tend to be relatively consistent across classes, regardless of sample sizes. Specifically, we build an attribute pool to learn potential semantic attributes of classes based on VLMs. For each input sample, we generate a unique attribute-aware prompt by selecting the relevant class attributes from the pool through a matching mechanism. This design enables the model to capture essential class semantics and generate informative prompts, even for classes with limited data. Additionally, we introduce a ProAdapter module to facilitate the transfer of foundational knowledge from VLMs while enhancing generalization to underrepresented classes in imbalanced settings. Extensive experiments on standard and imbalance few-shot tasks demonstrate that our model achieves superior performance especially in tail classes. Yuling Su, Xueliang Liu, Zhen Huang 0006, Yunwei Zhao, Richang Hong, Meng Wang 0001 |
IEEE Trans. Image Process. | 2 |
| 2025 | EmoGaussian High-Fidelity Emotional Talking Head Generation with 3D Gaussian Splatting
Jinlin Guo, Xueliang Liu |
ICIC (6) | 4 |
| 2025 | Deep Multi-sentence Aligned Cross-Modal Retrieval
Zhijian Lin, Sihan Gong, Xueliang Liu |
ICIG (1) | 3 |
| 2025 | CBLC-SOOD: contrastive background and label correction for semi-supervised oriented object detection
Yuxi Gong, Xueliang Liu, Liangfeng Xu |
Multim. Syst. | 4 |
| 2025 | Adapting ObjectBox for accurate hand detection
Yang Yang 0002, Xueliang Liu, Richang Hong |
Pattern Recognit. | 3 |
| 2025 | Cross-Modal Hashing via Diverse Instances MatchingabstractCross-modal hashing is a highly effective technique for searching relevant data across different modalities, owing to its low storage costs and fast similarity retrieval capability. While significant progress has been achieved in this area, prior investigations predominantly concentrate on a one-to-one feature alignment approach, where a singular feature is derived for similarity retrieval. However, the singular feature in these methods fails to adequately capture the varied multi-instance information inherent in the original data across disparate modalities. Consequently, the conventional one-to-one methodology is plagued by a semantic mismatch issue, as the rigid one-to-one alignment inhibits effective multi-instance matching. To address this issue, we propose a novel Diverse Instances Matching for Cross-modal Hashing (DIMCH), which explores the relevance between multiple instances in different modalities using a multi-instance learning algorithm. Specifically, we design a novel diverse instances learning module to extract a multi-feature set, which enables our model to capture detailed multi-instance semantics. To evaluate the similarity between two multi-feature sets, we adopt the smooth chamfer distance function, which enables our model to incorporate the conventional similarity retrieval structure. Moreover, to sufficiently exploit the supervised information from the semantic label, we adopt the weight cosine triplet loss as the objective function, which incorporates the multilevel similarity among the multi-labels into the training procedure and enables the model to mine the multi-label correlation effectively. Extensive experiments demonstrate that our diverse hashing embedding method achieves state-of-the-art performance in supervised cross-modal hashing retrieval tasks. Junfeng Tu, Xueliang Liu, Zhen Huang 0006, Yanbin Hao, Richang Hong, Meng Wang 0001 |
IEEE Trans. Image Process. | 2 |
| 2025 | Talking-DiSSM: Enhancing Temporal Consistency in Talking Face Video Generation with Bidirectional SSMsabstractGenerating temporally smooth and high-resolution videos is a crucial objective in talking face generation tasks. Diffusion-based generative models have emerged as a prime choice for these tasks due to their ability to produce high-quality outputs. To mitigate the impact of stochasticity in the diffusion process, recent research has predominantly utilized self-attention layers to extract temporal features, ensuring temporal consistency in the generated videos. However, self-attention mechanisms have computational complexity that scales quadratically with video length, leading to high computational costs. This limitation poses significant challenges when attempting to generate longer video sequences using diffusion models. To address this challenge, we propose Talking-DiSSM, an end-to-end method for generating audio-driven talking face videos using State-Space Models (SSMs). This novel framework for conditional video diffusion modeling integrates Bidirectional State-Space Models (Bi-SSM) as temporal modeling modules with linear complexity, effectively capturing complex sequential temporal information and intra-batch sequential interdependencies in videos. Additionally, we employ a simple yet effective batch-overlapped sampling strategy to process input video clips, constructing inter-batch correlations while incorporating reference face clips and landmarks as conditions to ensure stability in the generation process. Extensive experiments demonstrate that Talking-DiSSM generates temporally consistent, high-quality, and identity-preserving talking face videos synchronized with the driving audio, achieving state-of-the-art results compared to existing models. Xueliang Liu, Jinlin Guo, Richang Hong, Meng Wang 0001 |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2025 | A Unified Generative Hashing for Cross-Modal RetrievalabstractCross-modal hashing is a highly effective and efficient method for information retrieval, enabling the search for correlated data across different modality databases using compact hash codes. Conventional cross-modal hashing typically uses separate model structures for each modality and aligns approximate continuous representations of the final hash codes. These approaches not only require specialized models for each modality but also introduce a gap between the discrete hash codes and their continuous features, yielding only approximate alignment. To address these issues, we propose a unified generative cross-modal hashing method that leverages a single Uniform Mixture-of-Expert Decoder (UMoED) for both image and text modalities. UMoED streamlines cross-modal hash learning by integrating two key design elements: (1) a cross-modal representation unification module that employs unified queries to consolidate modality-specific features into a common space, and (2) an adaptive expert enhancement module that adaptively enhances feature modeling based on the input modality. Furthermore, our decoder-based hashing method outputs hash codes in a generative manner, producing precise representations of discrete codes to bridge the gap between the discrete and continuous space, thus ensuring precise alignment during similarity learning. Extensive experiments on three benchmark datasets demonstrate that the proposed method achieves the state-of-the-art performance in cross-modal hashing retrieval. Junfeng Tu, Xueliang Liu, Yanbin Hao, Richang Hong |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2024 | Exploring Robust Face-Voice Matching in Multilingual EnvironmentsabstractThis paper presents Team Xaiofei's innovative approach to exploring Face-Voice Association in Multilingual Environments (FAME) at ACM Multimedia 2024. We focus on the impact of different languages in face-voice matching by building upon Fusion and Orthogonal Projection (FOP), introducing four key components: a dual-branch structure, dynamic sample pair weighting, robust data augmentation, and score polarization strategy. Our dual-branch structure serves as an auxiliary mechanism to better integrate and provide more comprehensive information. We also introduce a dynamic weighting mechanism for various sample pairs to optimize learning. Data augmentation techniques are employed to enhance the model's generalization across diverse conditions. Additionally, score polarization strategy based on age and gender matching confidence clarifies and accentuates the final results. Our methods demonstrate significant effectiveness, achieving an equal error rate (EER) of 20.07 on the V2-EH dataset and 21.76 on the V1-EU dataset. Project page: https://github.com/cnzvan/Exploring-Robust-Face-Voice-Matching-in-Multilingual-Environments. Jiehui Tang, Xueliang Liu, Richang Hong |
ACM Multimedia | 5 |
| 2024 | A Robust Few-shot Learning Framework via Dual-branch Adversarial Noise Pretraining
Xueliang Liu, Yuling Su |
MMAsia | 2 |
| 2024 | Fine-grained Feature Assisted Cross-modal Image-text Retrieval
Chaofei Bu, Xueliang Liu, Zhen Huang 0006, Yuling Su, Junfeng Tu, Richang Hong |
PRCV (11) | 2 |
| 2024 | Complementary expert balanced learning for long-tail cross-modal retrieval
Peifang Liu, Xueliang Liu |
Multim. Syst. | 2 |
| 2024 | Reallocating and Evolving General Knowledge for Few-Shot LearningabstractLarge-scale vision-language pre-trained models like CLIP are extensively employed in few-shot tasks due to their robust generalization capabilities. Existing methods usually incorporate additional techniques to acquire knowledge for new tasks building upon the general knowledge in CLIP. However, they do not realize that the task-related knowledge might be implicitly embedded within the general knowledge well-learned. In this paper, we propose a novel framework to reallocate and evolve the general knowledge for specific few-shot tasks (REGK), mimicking the human “Attention Allocation” cognition mechanism. With a learnable mask-tuning selection, REGK focuses on selecting the task-related parameters of CLIP while learning specific few-shot knowledge without altering CLIP underlying framework. Specifically, we initially observe that inheriting the strong knowledge representation capability in CLIP is more advantageous for few-shot learning than its task-solving ability. Subsequently, a two-stage tuning framework is introduced to reallocate and control the mask-tuning on different tasks. It allows model automatically mask-tuning on different few-shot tasks with selective sparsity training. In this way, we achieve reliable transfer of task-related knowledge and effective exploration of new knowledge from limited data to enhance few-shot learning. Extensive experiments validate the superiority and potentiality of our model. Yuling Su, Xueliang Liu, Zhen Huang 0006, Richang Hong, Meng Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Modeling Hierarchical Uncertainty for Multimodal Emotion Recognition in ConversationabstractApproximating the uncertainty of an emotional AI agent is crucial for improving the reliability of such agents and facilitating human-in-the-loop solutions, especially in critical scenarios. However, none of the existing systems for emotion recognition in conversation (ERC) has attempted to estimate the uncertainty of their predictions. In this article, we present HU-Dialogue, which models hierarchical uncertainty for the ERC task. We perturb contextual attention weight values with source-adaptive noises within each modality, as a regularization scheme to model context-level uncertainty and adapt the Bayesian deep learning method to the capsule-based prediction layer to model modality-level uncertainty. Furthermore, a weight-sharing triplet structure with conditional layer normalization is introduced to detect both invariance and equivariance among modalities for ERC. We provide a detailed empirical analysis for extensive experiments, which shows that our model outperforms previous state-of-the-art methods on three popular multimodal ERC datasets. Feiyu Chen 0001, Jie Shao 0001, Anjie Zhu, Deqiang Ouyang, Xueliang Liu, Heng Tao Shen |
IEEE Trans. Cybern. | 5 |
| 2024 | Partial-Tuning Based Mixed-Modal Prototypes for Few-Shot ClassificationabstractThe significant success of machine learning models is mainly based on a large amount of data for training iterations, but this limits their generalization for few-shot data. Some existing models utilize the extensive visual and textual modal knowledge of vision-language pre-trained models (VLPs) to compensate for the data scarcity problem. However, they may suffer from a classification bias problem during the fusion of multi-modal information since that they focus on the inter-modal matching while neglecting intra-modal recognition for few-shot images. In this paper, we propose a novel few-shot model with mixed-modal prototypes by partial-tuning the VLPs for better information fusion. It aims to yield a high-quality class prototype representation by integrating the abundant multi-modal knowledge of VLPs and the specific-task information of low-shot visual data. Specifically, we introduce an image-text alignment module to ensure the consistency of the few-shot visual representation and the textual knowledge of VLPs at the feature space. A self-similar learning module is designed to excavate the local and detailed characters of specific class, which is crucial under the data scarcity. Additionally, to preserve the generalizable pre-trained knowledge in the maximum extent, we partial-tune the parameters of VLPs to adapt for the few-shot tasks. To sum up, we mix multi-modal information at the feature representation level instead of fusing multi-modal matching similarities, which effectively mitigates classification bias and ultimately enhances the model performance for few-shot data. The extensive experiments are conducted to evaluate the effectiveness of our model on 11 benchmark datasets and the results show its promising. Yuling Su, Xueliang Liu, Ye Zhao 0001, Richang Hong, Meng Wang 0001 |
IEEE Trans. Multim. | 2 |
| 2024 | Two-Step Discrete Hashing for Cross-Modal RetrievalabstractCross-modal hashing is an effective approach for information retrieval from large and heterogeneous cross-modal datasets, owing to its low storage cost and high computational speed. However, conventional cross-modal hashing techniques for generating hashing codes rely on cross-space dimensional compression, which results in two types of information loss: quantization information loss and dimension reduction loss. To address these limitations, we propose a novel method that decouples the one-step hashing (Fig.1a) strategy into two sub-steps (Fig.1b). Specifically, in the first step, we introduce a novel differentiable hash method, which utilizes a smooth hash module for binary quantization. This method allows our model to reduce the quantization information loss and make the model optimized by gradient descent. In the second step, we design a long-short Hamming space transformation approach to project the long code into a short one, which is effective in preserving the dimension information between long and short and mitigating the dimension reduction loss. We demonstrate the effectiveness of our approach through extensive experiments on several popular cross-modal datasets, achieving a significant improvement in cross-modal retrieval performance. Junfeng Tu, Xueliang Liu, Yanbin Hao, Richang Hong, Meng Wang 0001 |
IEEE Trans. Multim. | 2 |
| 2023 | Generalized Zero-Shot Learning with Noisy Labeled Data
Liqing Xu, Xueliang Liu, Yishun Jiang |
PRCV (11) | 2 |
| 2023 | Local Self-attention-based Hybrid Multiple Instance Learning for Partial Spoof Speech DetectionabstractThe development of speech synthesis technology has increased the attention toward the threat of spoofed speech. Although various high-performance spoofing countermeasures have been proposed in recent years, a particular scenario is overlooked: partially spoofed audio, where spoofed utterances may contain both spoofed and bona fide segments. Currently, the research on partially spoofed speech detection is lacking. The existing methods either train with partially spoofed speech at utterance level, resulting in gradient conflicting at the segment level, or directly train with segment level data, which requires segment labels that are difficult to obtain in practice. In this study, to better detect partially spoofed speech when only utterance labels are available, we formulate partially spoofed speech detection into a multiple instance learning (MIL) problem. The typical MIL uses a pooling layer to fuse patch scores as a whole, and we propose a hybrid MIL (H-MIL) framework based on max and log-sum-exp pooling methods, which can learn better segment representations to improve partially spoofed speech detection performance. Theoretical and experimental verification shows that H-MIL can effectively relieve the gradient conflicting and gradient vanishing problems. In addition, we analyze the local correlations between segments and introduce a local self-attention mechanism to enhance segment features, which further promotes the detection performance. In our experiments, we provide not only detection results at the segment and utterance levels but also some detailed visualization analysis, including the effect of spoof ratio and cross-dataset detection. The experimental results demonstrate the effective detection performance of our method at both the utterance and segment levels, especially when dealing with low spoof ratio attacks. The results confirm that our approach can better deal with partially spoofed speech detection than previous methods. Yupeng Zhu, Zuxing Zhao, Xueliang Liu, Jinlin Guo |
ACM Trans. Intell. Syst. Technol. | 4 |
| 2022 | Differentiable Cross-modal Hashing via Multimodal TransformersabstractCross-modal hashing aims at projecting the cross modal content into a common Hamming space for efficient search. Most existing work first encodes the samples with a deep network and then binaries the encoded feature into hashing code. However, the relative location information in the image may be lost when an image is encoded by the convolutional network, which makes it challenging to model the relationship of different modalities. Moreover, it is NP-hard to optimize the model with the discrete sign binary function popularly used in existing solutions. To address these issues, we propose a differentiable cross-modal hashing method that utilizes the multimodal transformer as the backbone to capture the location information in an image when encoding the visual content. In addition, a novel differentiable cross-modal hashing method is proposed to generate the binary code by a selecting mechanism, which could be formulated as a continuous and easily optimized problem. We perform extensive experiments on several cross modal datasets and the results show that the proposed method outperforms many existing solutions. Junfeng Tu, Xueliang Liu, Zongxiang Lin, Richang Hong, Meng Wang 0001 |
ACM Multimedia | 2 |
| 2022 | Early-Learning regularized Contrastive Learning for Cross-Modal Retrieval with Noisy LabelsabstractCross modal retrieval receives intensive attention for flexible queries between different modalities. However, in practice it is challenging to retrieve cross modal content with noisy labels. The latest research on machine learning shows that a model tends to fit cleanly labeled data at early learning stage and then memorize the data with noisy labels. Although the clustering strategy in cross modal retrieval can be utilized for alleviating outliers, the networks will rapidly overfit after clean data is fitted well and the noisy labels begin to force the cluster center drift. Motivated by these fundamental phenomena, we propose an Early Learning regularized Contrastive Learning method for Cross Modal Retrieval with Noisy Labels (ELRCMR). In the solution, we propose to project the multi-modal data to a shared feature space by contrastive learning, in which early learning regularization is employed to prevent the memorization of noisy labels when training the model, and the dynamic weight balance strategy is employed to alleviate clustering drift. We evaluated the method with extensive experiments, and the result shows the proposed method could solve the cluster drift in conventional solutions and achieve promising performance on widely used benchmark datasets. Tianyuan Xu, Xueliang Liu, Zhen Huang 0006, Dan Guo 0001, Richang Hong, Meng Wang 0001 |
ACM Multimedia | 2 |
| 2022 | Visual feature synthesis with semantic reconstructor for traditional and generalized zero-shot object classificationabstractZero-shot learning (ZSL) addresses the novel object recognition problem by leveraging semantic embedding to transfer knowledge from seen categories to unseen categories. Generative ZSL models synthesize the visual features of unseen classes and convert ZSL task into a classical supervised learning problem. These generative ZSL models are trained by using the seen classes. Although promising progress has been achieved in the ZSL and generalized zero-shot learning (GZSL) tasks. The existing approaches still suffer from a strong bias problem between unseen and seen classes, where unseen objects in the target domain tend to be recognized as seen classes in the source domain. To deal with the problem, we propose a novel named semantic consistent Wasserstein generative adversarial network (scWGAN), which uses a semantic reconstructor to reconstruct semantic embeddings from generated visual features by incorporating a novel Semantic Consistent Loss noted L rec . The Semantic Consistent Loss guides our proposed scWGAN to generate visual features that mirror the semantic relationships between seen and unseen classes. We also introduce a visual classifier to constrain visual feature generator. Extensive experiments show that the proposed approach is superior to previous state-of-the-art works under both traditional ZSL and challenging GZSL settings on six popular data sets AWA1, AWA2, CUB, APY, and SUN. Ye Zhao 0001, Xueliang Liu, Dan Guo 0001, Zhenzhen Hu 0004, Hengchang Liu, Yicong Li 0004 |
Int. J. Intell. Syst. | 3 |
| 2022 | Revisiting Local Descriptor for Improved Few-Shot ClassificationabstractFew-shot classification studies the problem of quickly adapting a deep learner to understanding novel classes based on few support images. In this context, recent research efforts have been aimed at designing more and more complex classifiers that measure similarities between query and support images but left the importance of feature embeddings seldom explored. We show that the reliance on sophisticated classifiers is not necessary, and a simple classifier applied directly to improved feature embeddings can instead outperform most of the leading methods in the literature. To this end, we present a new method, named DCAP, for few-shot classification, in which we investigate how one can improve the quality of embeddings by leveraging Dense Classification and Attentive Pooling (DCAP) . Specifically, we propose to train a learner on base classes with abundant samples to solve dense classification problem first and then meta-train the learner on plenty of randomly sampled few-shot tasks to adapt it to few-shot scenario or the test time scenario. During meta-training, we suggest to pool feature maps by applying attentive pooling instead of the widely used global average pooling to prepare embeddings for few-shot classification. Attentive pooling learns to reweight local descriptors, explaining what the learner is looking for as evidence for decision making. Experiments on two benchmark datasets show the proposed method to be superior in multiple few-shot settings while being simpler and more explainable. Code is publicly available at https://github.com/Ukeyboard/dcap/ . Richang Hong, Xueliang Liu, Mingliang Xu 0001, Qianru Sun |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2021 | NASTER: Non-local Attentional Scene Text RecognizerabstractScene text recognition has been widely investigated in computer vision. In the literature, the encoder-decoder based framework, which first encodes image into feature map and then decodes them into corresponding text sequences, have achieved great success. However, this solution fails in low-quality images, as the local visual features extracted from curved or blurred images are difficult to decode into corresponding text. To address this issue, we propose a new framework for Scene Text Recognition (STR), named Non-Local Attentional Scene Text Recognizer (NASTER). We use ResNet with Global Context Block (GC block) to extract global visual features. The global context information is then captured in parallel using the self-attention module and finally decoded by a multi-layer attention decoder with an intermediate supervision module. The proposed method achieves the state-of-the-art performances on seven benchmark datasets, demonstrating the effectiveness of our approach. Xueliang Liu, Yanbin Hao, Yunjie Ma, Richang Hong |
ICMR | 2 |
| 2021 | Learning What and When to Drop: Adaptive Multimodal and Contextual Dynamics for Emotion Recognition in ConversationabstractMulti-sensory data has exhibited a clear advantage in expressing richer and more complex feelings, on the Emotion Recognition in Conversation (ERC) task. Yet, current methods for multimodal dynamics that aggregate modalities or employ additional modality-specific and modality-shared networks are still inadequate in balancing between the sufficiency of multimodal processing and the scalability to incremental multi-sensory data type additions. This incurs a bottleneck of performance improvement of ERC. To this end, we present MetaDrop, a differentiable and end-to-end approach for the ERC task that learns module-wise decisions across modalities and conversation flows simultaneously, which supports adaptive information sharing pattern and dynamic fusion paths. Our framework mitigates the problem of modelling complex multimodal relations while ensuring it enjoys good scalability to the number of modalities. Experiments on two popular multimodal ERC datasets show that MetaDrop achieves new state-of-the-art results. Feiyu Chen 0001, Zhengxiao Sun, Deqiang Ouyang, Xueliang Liu, Jie Shao 0001 |
ACM Multimedia | 4 |
| 2021 | A lightweight multi-scale aggregated model for detecting aerial images captured by UAVs
Zhaokun Li, Xueliang Liu, Ye Zhao 0001, Bo Liu 0005, Zhen Huang 0006, Richang Hong |
J. Vis. Commun. Image Represent. | 2 |
| 2021 | A novel lossless compression framework for facial depth images in expression recognition
Chunxiao Fan 0002, Fu Li 0002, Xueliang Liu |
Multim. Tools Appl. | 4 |
| 2021 | Multi-Branch Networks for Video Super-Resolution With Dynamic Reconstruction StrategyabstractRecently, the rapid development of 2-dimensional (2D) convolutional neural networks (CNNs) has driven single image super-resolution (SISR) into a new era, owing to their powerful ability in modeling spatial relation within one single image. However, few studies focus on video super-resolution (VSR) due to the key challenge that apart from the spatial relation, the temporal dependence among consecutive low-resolution (LR) frames must be taken into consideration for better reconstruction. In this article, unlike most previous methods based on optical flow for motion compensation, 3-dimensional (3D) convolution is utilized to capture the temporal relation. Firstly, in contrast to the conventional 3D convolution which is notorious for the excessively high computational burden, we propose an efficient 3D convolutional block (E3DB) through convolution factorization principle (CFP), which significantly reduces the computing load while maximally maintaining the temporal information. Then, by taking advantage of E3DB, we propose a novel multi-resolution extraction block (MREB) which aggregates the information from multiple resolutions, leading to the stronger high-resolution representation learning and better feature extraction. Besides, based on our 3-branch architecture, instead of the simple addition or concatenation, a dynamic reconstruction strategy (DRS) is proposed to adaptively fuse the optimal information of temporal dependence from each branch. It is therefore termed as dynamic multiple branch network (DMBN). Comprehensive experiments on public benchmark datasets demonstrate the superiority of our DMBN over the current state-of-the-art methods in terms of accuracy and efficiency. Dongyang Zhang 0001, Jie Shao 0001, Zhenwen Liang, Xueliang Liu, Heng Tao Shen |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2020 | Memory-Augmented Relation Network for Few-Shot LearningabstractMetric-based few-shot learning methods concentrate on learning transferable feature embedding which generalizes well from seen categories to unseen categories under limited supervision. However, most of the methods treat each individual instance separately without considering its relationships with the others in the working context. We investigate a new metric-learning method to explicitly exploit these relationships. In particular, for an instance, we choose the samples that are visually similar from the working context, and perform weighted information propagation to attentively aggregate helpful information from the chosen samples to enhance its representation. We further formulate the distance metric as a learnable relation module which learns to compare for similarity measurement, and equip the working context with memory slots, both contributing to generality. We empirically demonstrate that the proposed method yields significant improvement over its ancestor and achieves competitive or even better performance when compared with other few-shot learning approaches on the two major benchmark datasets, i.e.mini Imagenet andtiered Imagenet. Richang Hong, Xueliang Liu, Mingliang Xu 0001, Zhengjun Zha, Meng Wang 0001 |
ACM Multimedia | 3 |
| 2020 | Weakly-Supervised Video Object Grounding by Exploring Spatio-Temporal ContextsabstractGrounding objects in visual context from natural language queries is a crucial yet challenging vision-and-language task, which has gained increasing attention in recent years. Existing work has primarily investigated this task in the context of still images. Despite their effectiveness, these methods cannot be directly migrated into the video context, mainly due to 1) the complex spatio-temporal structure of videos and 2) the scarcity of fine-grained annotations of videos. To effectively ground objects in videos is profoundly more challenging and less explored. Xun Yang 0001, Xueliang Liu, Meng Jian, Xinjian Gao, Meng Wang 0001 |
ACM Multimedia | 2 |
| 2020 | WFN-PSC: weighted-fusion network with poly-scale convolution for image dehazingabstractImage dehazing is a fundamental task for the computer vision and multimedia and usually in the face of the challenge from two aspects, i) the uneven distribution of arbitrary haze and ii) the distortion of image pixels caused by the hazed image. In this paper, we propose an end-to-end trainable framework, named Weighted-Fusion Network with Poly-Scale Convolution (WFN-PSC), to address these dehazing issues. The proposed method is designed based on the Poly-Scale Convolution (PSConv). It can extract the image feature from different scales without upsampling and downsampled, which avoids the image distortion. Beyond this, we design the spatial and channel weighted-fusion modules to make the WFN-PSC model focus on the hard dehazing parts of image from two dimensions. Specifically, we design three Part Architectures followed by the channel weighted-fusion module. Each Part Architecture consists of three PSConv residual blocks and a spatial weighted-fusion module. The experiments on the benchmark demonstrate the dehazing effectiveness of the proposed method. Furthermore, considering that image dehazing is a low-level task in the computer vision, we evaluate the dehazed image on the object detection task and the results show that the proposed method can be a good pre-processing to assist the high-level computer vision task. Lexuan Sun, Xueliang Liu, Zhenzhen Hu 0004, Richang Hong |
MMAsia | 2 |
| 2020 | Accelerating SGD using flexible variance reduction on large-scale datasets
Mingxing Tang, Linbo Qiao, Zhen Huang 0006, Xinwang Liu 0002, Yuxing Peng 0001, Xueliang Liu |
Neural Comput. Appl. | 6 |
| 2020 | Person re-identification based on multi-scale constraint network
Sishang Li, Xueliang Liu, Ye Zhao 0001, Meng Wang 0001 |
Pattern Recognit. Lett. | 2 |
| 2020 | Deep Neighborhood Component Analysis for Visual Similarity ModelingabstractLearning effective visual similarity is an essential problem in multimedia research. Despite the promising progress made in recent years, most existing approaches learn visual features and similarities in two separate stages, which inevitably limits their performance. Once useful information has been lost in the feature extraction stage, it can hardly be recovered later. This article proposes a novel end-to-end approach for visual similarity modeling, calleddeep neighborhood component analysis, which discriminatively trains deep neural networks to jointly learn visual features and similarities. Specifically, we first formulate a metric learning objective that maximizes the intra-class correlations and minimizes the inter-class correlations under the neighborhood component analysis criterion, and then train deep convolutional neural networks to learn a nonlinear mapping that projects visual instances from original feature space to a discriminative and neighborhood-structure-preserving embedding space, thus resulting in better performance. We conducted extensive evaluations on several widely used and challenging datasets, and the impressive results demonstrate the effectiveness of our proposed approach. Xueliang Liu, Xun Yang 0001, Meng Wang 0001, Richang Hong |
ACM Trans. Intell. Syst. Technol. | 1 |
| 2020 | Cross-Domain Sentiment Encoding through Stochastic Word EmbeddingabstractSentiment analysis is an important topic concerning identification of feelings, attitudes, emotions and opinions from text. To automate such analysis, a large amount of example text needs to be manually annotated for model training. This is laborious and expensive, but the cross-domain technique is a key solution to reducing the cost by reusing annotated reviews across domains. However, its success largely relies on the learning of a robust common representation space across domains. In the recent years, significant effort has been invested to improve the cross-domain representation learning by designing increasingly more complex and elaborate model inputs and architectures. We support that it is not necessary to increase design complexity as this inevitably consumes more time in model training. Instead, we propose to explore the word polarity and occurrence information through a simple mapping and encode such information more accurately whilst managing lower computational costs. The proposed approach is unique and takes advantage of the stochastic embedding technique to tackle cross-domain sentiment alignment. Its effectiveness is benchmarked with over ten data tasks constructed from two review corpora and it is compared against ten classical and state-of-the-art methods. Yanbin Hao, Tingting Mu, Richang Hong, Meng Wang 0001, Xueliang Liu, John Yannis Goulermas |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2019 | MAHCI 2019: The 2nd Workshop on Multimedia for Accessible Human Computer InterfaceabstractMultimedia technology plays a fundamental role to increase usability, and accessibility of computer interfaces in developing of advanced human-computer interaction devices. The 2nd workshop on Multimedia for Accessible Human Computer Interface (MAHCI) continues to provide a forum to both multimedia and HCI researchers to discuss the accessible human computer interface design, development, and evaluation with the state-of-the-art multimedia technology. It also enables multimedia community to expand its interaction with the HCI industry and broaden the scope of deploying multimedia technology in practical applications. The workshop features 6 papers which cover a number of novel applications and new methodologies in a half day program. Xueliang Liu, Troy McDaniel |
ACM Multimedia | 1 |
| 2019 | Social multi-modal event analysis via knowledge-based weighted topic model
Feng Xue 0002, Xueliang Liu, Tianpeng Liu, Qiang Lu 0002 |
J. Vis. Commun. Image Represent. | 3 |
| 2019 | Multi-modal max-margin supervised topic model for social event analysis
Feng Xue 0002, Shengsheng Qian, Tianzhu Zhang 0001, Xueliang Liu, Changsheng Xu |
Multim. Tools Appl. | 5 |
| 2019 | Robust guaranteed cost control for continuous-time uncertain Markov switching singular systems with mode-dependent time delays
Huanli Gao, Xueliang Liu, Fuchun Liu |
Neural Comput. Appl. | 2 |
| 2019 | Towards Accurate Georeferenced Video Search With Camera Field of View ModelingabstractPositioning data and other sensor measurements, such as camera orientation, have become important contextual features generated by mobile devices during video recording, which proved to be increasingly beneficial to video search. To enable access to videos based on their metadata (e.g., geo-properties produced by GPS and digital compass), a model representing camera field of view (FOV) is needed. Vector model of previous work, which simplifies FOV by ignoring the viewable angle for search efficiency, has become popular in georeferenced video search. However, when the viewable angle is large, many false positives and false negatives occur, which are undesirable for filtering of georeferenced video search. This paper proposes a new model, which can appropriately represent the actual FOV as a filtering step, without any false positive or false negative. Based on this model, we investigate how to process five types of overlap queries for searching videos as spatio-temporal objects. To verify the effectiveness of our model and the corresponding query processing algorithms, experiments on a real data set we collected and a large synthetic data set are conducted. The results show that the proposed model can perform much better compared with the existing vector model, and the accuracy of its search results remains almost steady, even when the angle changes. Jie Shao 0001, Gang Hu 0004, Jingkuan Song, Xueliang Liu, Heng Tao Shen |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2019 | BTDP: Toward Sparse Fusion with Block Term Decomposition Pooling for Visual Question AnsweringabstractBilinear models are very powerful in multimodal fusion tasks like Visual Question Answering. The predominant bilinear methods can all be seen as a kind of tensor-based decomposition operation that contains a key kernel called “core tensor.” Current approaches usually focus on reducing the computation complexity by applying low-rank constraint on the core tensor. In this article, we propose a novel bilinear architecture called Block Term Decomposition Pooling (BTDP), which not only maintains the advantages of previous bilinear methods but also conducts sparse bilinear interactions between modalities. Our method is based on Block Term Decompositions theory of tensor, which will result in a sparse and learnable block-diagonal core tensor for multimodal fusion. We prove that using such a block-diagonal core tensor is equivalent to conducting many “tiny” bilinear operations in different feature spaces. Thus, introducing sparsity into the bilinear operation can significantly increase the performance of feature fusion and improve VQA models. What is more, our BTDP is very flexible in design. We develop several variants of BTDP and discuss the effects of the diagonal blocks of core tensor. Extensive experiments on two challenging VQA-v1 and VQA-v2 datasets show that our BTDP method outperforms current bilinear models, achieving state-of-the-art performance. Zhiwei Fang, Jing Liu 0001, Xueliang Liu, Qu Tang, Yong Li 0034, Hanqing Lu |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2019 | A2CMHNE: Attention-Aware Collaborative Multimodal Heterogeneous Network EmbeddingabstractNetwork representation learning is playing an important role in network analysis due to its effectiveness in a variety of applications. However, most existing network embedding models focus on homogeneous networks and neglect the diverse properties such as different types of network structures and associated multimedia content information. In this article, we learn node representations for multimodal heterogeneous networks, which contain multiple types of nodes and/or links as well as multimodal content such as texts and images. We propose a novel attention-aware collaborative multimodal heterogeneous network embedding method (A 2 CMHNE), where an attention-based collaborative representation learning approach is proposed to promote the collaboration of structure-based embedding and content-based embedding, and generate the robust node representation by introducing an attention mechanism that enables informative embedding integration. In experiments, we compare our model with existing network embedding models on two real-world datasets. Our method leads to dramatic improvements in performance by 5%, and 9% compared with five state-of-the-art embedding methods on one benchmark (M10 Dataset), and on a multi-modal heterogeneous network dataset (WeChat dataset) for node classification, respectively. Experimental results demonstrate the effectiveness of our proposed method on both node classification and link prediction tasks. Jun Hu 0016, Shengsheng Qian, Quan Fang, Xueliang Liu, Changsheng Xu |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2019 | Cross-Modality Feature Learning via Convolutional AutoencoderabstractLearning robust and representative features across multiple modalities has been a fundamental problem in machine learning and multimedia fields. In this article, we propose a novel MUltimodal Convolutional AutoEncoder (MUCAE) approach to learn representative features from visual and textual modalities. For each modality, we integrate the convolutional operation into an autoencoder framework to learn a joint representation from the original image and text content. We optimize the convolutional autoencoders of different modalities jointly by exploiting the correlation between the hidden representations from the convolutional autoencoders, in particular by minimizing both the reconstructing error of each modality and the correlation divergence between the hidden feature of different modalities. Compared to the conventional solutions relying on hand-crafted features, the proposed MUCAE approach encodes features from image pixels and text characters directly and produces more representative and robust features. We evaluate MUCAE on cross-media retrieval as well as unimodal classification tasks over real-world large-scale multimedia databases. Experimental results have shown that MUCAE performs better than the state-of-the-arts methods. Xueliang Liu, Meng Wang 0001, Zhengjun Zha, Richang Hong |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2019 | On fusing the latent deep CNN feature for image classification
Xueliang Liu, Rongjie Zhang, Richang Hong, Guangcan Liu |
World Wide Web | 1 |
| 2019 | Correction to: On fusing the latent deep CNN feature for image classification
Xueliang Liu, Rongjie Zhang, Richang Hong, Guangcan Liu |
World Wide Web | 1 |
| 2018 | MAHCI 2018: The 1st Workshop on Multimedia for Accessible Human Computer InterfaceabstractIn the developing of advanced Human-Computer Interaction, multimedia technology plays a fundamental role to increase usability, and accessibility of computer interfaces. The first workshop on Multimedia for Accessible Human Computer Interface (MAHCI) provides a forum to both multimedia and HCI researchers to discuss the accessible human computer interface design, development, and evaluations with the state-of-the-art multimedia technology. It also enables multimedia community to expand its interaction with the HCI industry and broaden the scope of deploying multimedia technology in practical applications. The workshop features 5 papers which cover a number of novel applications and new methodologies in a half day program. Xueliang Liu, Benoit Huet, Jia Jia 0001 |
ACM Multimedia | 1 |
| 2017 | Adaptive Steganalysis Based on Selection Region and Combined Convolutional Neural NetworksabstractDigital image steganalysis is the art of detecting the presence of information hiding in carrier images. When detecting recently developed adaptive image steganography methods, state-of-art steganalysis methods cannot achieve satisfactory detection accuracy, because the adaptive steganography methods can adaptively embed information into regions with rich textures via the guidance of distortion function and thus make the effective steganalysis features hard to be extracted. Inspired by the promising success which convolutional neural network (CNN) has achieved in the fields of digital image analysis, increasing researchers are devoted to designing CNN based steganalysis methods. But as for detecting adaptive steganography methods, the results achieved by CNN based methods are still far from expected. In this paper, we propose a hybrid approach by designing a region selection method and a new CNN framework. In order to make the CNN focus on the regions with complex textures, we design a region selection method by finding a region with the maximal sum of the embedding probabilities. To evolve more diverse and effective steganalysis features, we design a new CNN framework consisting of three separate subnets with independent structure and configuration parameters and then merge and split the three subnets repeatedly. Experimental results indicate that our approach can lead to performance improvement in detecting adaptive steganography. Donghui Hu, Shengnan Zhou, Xueliang Liu, Yuqi Fan 0001, Lina Wang 0001 |
Secur. Commun. Networks | 4 |
| 2016 | E^2SGM E 2 S G M : Event Enrichment and Summarization by Graph Model
Xueliang Liu, Benoit Huet, Feng Wang 0036 |
MMM (2) | 1 |
| 2016 | Event analysis in social multimedia: a survey
Xueliang Liu, Meng Wang 0001, Benoit Huet |
Frontiers Comput. Sci. | 1 |
| 2016 | Linking socially contributed media with events
Xueliang Liu, Benoit Huet |
Multim. Syst. | 1 |
| 2016 | Event-based cross media question answering
Xueliang Liu, Benoit Huet |
Multim. Tools Appl. | 1 |
| 2015 | Event-Based Media Enrichment Using an Adaptive Probabilistic Hypergraph ModelabstractNowadays, with the continual development of digital capture technologies and social media services, a vast number of media documents are captured and shared online to help attendees record their experience during events. In this paper, we present a method combining semantic inference and multimodal analysis for automatically finding media content to illustrate events using an adaptive probabilistic hypergraph model. In this model, media items are taken as vertices in the weighted hypergraph and the task of enriching media to illustrate events is formulated as a ranking problem. In our method, each hyperedge is constructed using the K-nearest neighbors of a given media document. We also employ a probabilistic representation, which assigns each vertex to a hyperedge in a probabilistic way, to further exploit the correlation among media data. Furthermore, we optimize the hypergraph weights in a regularization framework, which is solved as a second-order cone problem. The approach is initiated by seed media and then used to rank the media documents using a transductive inference process. The results obtained from validating the approach on an event dataset collected from EventMedia demonstrate the effectiveness of the proposed approach. Xueliang Liu, Meng Wang 0001, Benoit Huet, Xuelong Li 0001 |
IEEE Trans. Cybern. | 1 |
| 2015 | Visual Classification by ℓ1-Hypergraph ModelingabstractVisual classification has attracted considerable research interests in the past decades. In this paper, a novel ℓ1-hypergraph model for visual classification is proposed. Hypergraph learning, as a natural extension of graph model, has been widely used in many machine learning tasks. In previous work, hypergraph is usually constructed by attribute-based or neighborhood-based methods. That is, a hyperedge is generated by connecting a set of samples sharing a same feature attribute or in a neighborhood. However, these methods are unable to explore feature space globally or sensitive to noises. To address these problems, we propose a novel hypergraph construction approach that leverages sparse representation to generate hyperedges and learns the relationship among hyperedges and their vertices. First, for each sample, a hyperedge is generated by regarding it as the centroid and linking it as well as its nearest neighbors. Then, the sparse representation method is applied to represent the centroid vertex by other vertices within the same hyperedge. The vertices with zero coefficients are removed from the hyperedge. Finally, the representation coefficients are used to define the incidence relation between the hyperedge and the vertices. In our approach, we also optimize the hyperedge weights to modulate the effects of different hyperedges. We leverage the prior knowledge on the hyperedges so that the hyperedges sharing more vertices can have closer weights, where a graph Laplacian is used to regularize the optimization of the weights. Our approach is named ℓ1-hypergraph since the ℓ1sparse representation is employed in the hypergraph construction process. The method is evaluated on various visual classification tasks, and it demonstrates promising performance. Meng Wang 0001, Xueliang Liu, Xindong Wu 0001 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2014 | On improving behavior subtractionabstractWith the popularity of monitoring devices, huge amount of surveillance data is generated in every minute. The technique for automatic analysis of monitoring videos is in urgent demand. As an extension of background subtraction, behavior subtraction succeeds in detecting the changes of scenes dynamics instead of its photometric properties. In this paper, we first propose a new algorithm in improving behavior subtraction by maximum likelihood estimate and interval estimate methods. After that we apply the improved approach to the framework of video summarization in which the goal is to condense hours of video data into a few short segments. The compressed video clips allow human to catch their interested information quickly. We finally conduct extensive experiments on real-world surveillance videos. The experimental results demonstrate its superior performance to other state-of-the-art methods. Yanbin Hao, Xueliang Liu, Richang Hong |
SMC | 4 |
| 2014 | On the automatic online collection of training data for visual event modeling
Xueliang Liu, Benoit Huet |
Multim. Tools Appl. | 1 |
| 2013 | Heterogeneous features and model selection for event-based media classificationabstractWith the rapid development of social media sites, a lot of user generated content is being shared in the Web, leading to new challenges for traditional media retrieval techniques. An event describes the happening at a specific time and place in real-world, and it is one of the most important cues for people to recall past memories. The reminder value of an event makes it extremely helpful in organizing human life. Thus, organizing media by events has recently drawn much attention within the multimedia research community. In this paper, we focus on two fundamental problems related to event based social media analysis: the study of feature importance for modeling the relation between events and media, and how to deal with missing and erroneous metadata often present in social media data. These issues are studied within an event-based media classification framework. Different learning approaches are employed to train the event models on different features. We find, through experiments on a large set of events, that the best discriminant features are tags, spatial and temporal feature. We address the missing value problem by extending the feature with an extra attribute to indicate if the values are missing. Promising results are achieved demonstrating the effectiveness of the proposed method. Xueliang Liu, Benoit Huet |
ICMR | 1 |
| 2013 | EventEnricher: a novel way to collect media illustrating eventsabstractExploiting event context to organize social media draws lots of interest from the multimedia community. In this paper, we present our system, called EventEnricher, to infer the semantics behind events and explore social media to illustrate events. We extend the set of illustrating images for a particular event by querying social media with diverse multi-modal features and subsequently pruning the results using content based visual analysis. We integrate the solution into an intelligent interface that enables the user to browse the media collection illustrating events in an easy, effective and informative way. Xueliang Liu, Benoit Huet |
ICMR | 1 |
| 2011 | Finding media illustrating eventsabstractWe present a method combining semantic inferencing and visual analysis for finding automatically media (photos and videos) illustrating events. We report on experiments validating our heuristic for mining media sharing platforms and large event directories in order to mutually enrich the descriptions of the content they host. Our overall goal is to design a web-based environment that allows users to explore and select events, to inspect associated media, and to discover meaningful, surprising or entertaining connections between events, media and people participating in events. We present a large dataset composed of semantic descriptions of events, photos and videos interlinked with the larger Linked Open Data cloud and we show the benefits of using semantic web technologies for integrating multimedia metadata. Xueliang Liu, Raphaël Troncy, Benoit Huet |
ICMR | 1 |
| 2007 | Video Collage: A Novel Presentation of Video SequenceabstractThis paper presents an automatic procedure for constructing a compact synthesized collage from a video sequence. The synthesized image called video collage, is a kind of static video summary - to select most representative images from video, to extract salient regions of interest (ROI) from these images and resize them according to their saliency, and to seamlessly arrange ROI on a given canvas with the temporal structure of video content preserved. We formulate the generation of video collage as an energy minimization problem in which each of above desirability is represented by an energy term. Unlike most existing video presentation schemes, video collage is more compact and visually appealing. We have applied video collage on several home videos and report superior performance in a user study compared with key existing approaches to video presentation. Tang Wang, Tao Mei 0001, Xian-Sheng Hua 0001, Xueliang Liu, He-Qin Zhou |
ICME | 4 |
| 2007 | Video collageabstractIn this work, we present Video Collage system, which automatically constructs a compact and visually appealing synthesized collage from a video sequence for efficient video browsing. Given a video, Video Collage is able to select the most representative images, extract salient regions of interest (ROI) from these images and resize ROI according to their saliencies, and seamlessly arrange them on a given canvas while preserving the temporal structure of video content. Furthermore, Video Collage provides a novel user interface that enables users to browse video content in a variety of more efficient ways in contrast to many existing approaches to video browsing. Xueliang Liu, Tao Mei 0001, Xian-Sheng Hua 0001, Bo Yang 0008, He-Qin Zhou |
ACM Multimedia | 1 |