EDBT 2026 Demo / reviewers in the wild / expert
Xin Liu 0011
dblp:76/1820-11
· DBLP profile ↗
93ranked-venue papers
21as first author
54since 2021 · last 2026
0000-0002-0011-6260ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 54 · 9 first-author · 34 since 2021Artificial intelligence and machine learning · 42 · 9 first-author · 28 since 2021Databases, data management, data science and information retrieval · 8 · 5 since 2021Human-computer interaction and ubiquitous computing · 4 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 1 since 2021Computer networks · 3 · 3 since 2021Security and privacy · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Enhancing Rotation-Invariant 3D Learning with Global Pose Awareness and Attention MechanismsabstractRecent advances in rotation-invariant (RI) learning for 3D point clouds typically replace raw coordinates with handcrafted RI features to ensure robustness under arbitrary rotations. However, these approaches often suffer from the loss of global pose information, making them incapable of distinguishing geometrically similar but spatially distinct structures. We identify that this limitation stems from the restricted receptive field in existing RI methods, leading to Wing–tip feature collapse, a failure to differentiate symmetric components (e.g., left and right airplane wings) due to indistinguishable local geometries. To overcome this challenge, we introduce the Shadow-informed Pose Feature (SiPF), which augments local RI descriptors with a globally consistent reference point (referred to as the “shadow”) derived from a learned shared rotation. This mechanism enables the model to preserve global pose awareness while maintaining rotation invariance. We further propose Rotation-invariant Attention Convolution (RIAttnConv), an attention-based operator that integrates SiPFs into the feature aggregation process, thereby enhancing the model’s capacity to distinguish structurally similar components. Additionally, we design a task-adaptive shadow locating module based on the Bingham distribution over unit quaternions, which dynamically learns the optimal global rotation for constructing consistent shadows. Extensive experiments on 3D classification and part segmentation benchmarks demonstrate that our approach substantially outperforms existing RI methods, particularly in tasks requiring fine-grained spatial discrimination under arbitrary rotations. Jiaxun Guo, Manar Amayri, Nizar Bouguila, Xin Liu 0011, Wentao Fan 0001 |
AAAI | 4 |
| 2026 | DRFGD: Disentangled Representation-Focused Generative Defense for Attack-Tolerant Cross-Modal HashingabstractWith the widespread deployment of cross-modal retrieval in real-world scenarios, ensuring robustness against adversarial attacks is increasingly critical. Remarkably, deep cross-modal hashing is highly vulnerable to adversarial attacks due to its discrete nature and low-dimensional hash codes, while existing defense methods often fail to suppress perturbations embedded in vulnerable features and lack the capacity to model modality-specific structural differences, resulting in suboptimal adversarial robustness. To address these challenges, we propose a novel Disentangled Representation-Focused Generative Defense (DRFGD) framework for attack-tolerant cross-modal hashing. Without altering the structure of retrieval model, DRFGD defends against adversarial attacks by disentangling input representations into adversarial-robust and adversarial-vulnerable components, by an efficient dual-branch semantic-aware encoder. Guided by such disentangled robust features, an attack-tolerant generative module is seamlessly designed to synthesize semantically aligned and perturbation-resilient examples for robust adversarial training, thereby significantly promoting collaborative defense robustness to attackers. Consequently, the semantically consistent hash codes can be well obtained to enhance adversarial robustness in complex cross-modal attacking scenarios. Extensive experiments on public benchmarks demonstrate that DRFGD substantially improves retrieval robustness under various attacking scenarios, and shows its improved defense performance in comparison with the SOTA works. Zhongqing Yu, Xin Liu 0011, Yiu-Ming Cheung, Zhikai Hu, Wentao Fan 0001, Pan Zhou 0001 |
AAAI | 2 |
| 2026 | A cooperative learning method for early fake news detection with social engagement-aware masking encoder
Pingjing Xu, Shu-Juan Peng, Xin Liu 0011, Lei Zhu 0002, Danni Yu |
Multim. Syst. | 3 |
| 2026 | Efficient image-text retrieval via bi-cross-graph learning and multi-grained alignment
Shenggang Zhou, Xin Liu 0011, Lei Zhu 0002, Shu-Juan Peng, Jixiang Du, Jianjia Cao |
Multim. Syst. | 2 |
| 2026 | MLCA: Multi-level Correlative Attacks against Deep Cross-Modal Hashing
Xiaohang Fang, Xin Liu 0011, Zhikai Hu, Yiu-Ming Cheung, Shu-Juan Peng, Xing Xu 0001 |
Pattern Recognit. | 2 |
| 2026 | FPAD: Fuzzy-Prototype-Guided Adversarial Attack and Defense for Deep Cross-Modal HashingabstractDeep cross-modal hashing models generally inherit the vulnerabilities of deep neural networks, making them susceptible to adversarial attacks and thus posing a serious security risk during real-world deployment. Current adversarial attack or defense strategies often establish a weak correlation between the hashing codes and the targeted semantic representations, and there is still a lack of related works that simultaneously consider the attack and defense for deep cross-modal hashing. To alleviate these concerns, we propose a Fuzzy-Prototype-guided Adversarial Attack and Defense (FPAD) framework to enhance the adversarial robustness of deep cross-modal hashing models. First, an adaptive fuzzy-prototype learning network (FpNet) is efficiently presented to extract a set of fuzzy-prototypes, aiming to encode the underlying semantic structure of the heterogeneous modalities in both feature and Hamming spaces. Then, these derived prototypical hash codes are heuristically employed to supervise the generation of high-quality adversarial examples, while a fuzzy-prototype rectification scheme is simultaneously designed to preserve the latent semantic consistency between the adversarial and benign examples. By mixing the adversarial samples with the original training samples as the augmented inputs, an efficient fuzzy-prototype-guided adversarial learning framework is proposed to execute the collaborative adversarial training and generate robust cross-modal hash codes with high adversarial defense capabilities, therefore resisting various attacks and benefiting various challenging cross-modal hashing tasks. Extensive experiments evaluated on benchmark datasets show that the proposed FPAD framework not only produces high-quality adversarial samples to enhance the adversarial training process, but also shows its high adversarial defense capability to benefit various cross-modal hashing tasks. The code is available at: https://github.com/yzq131/FPAD. Zhongqing Yu, Xin Liu 0011, Yiu-Ming Cheung, Lei Zhu 0002, Xing Xu 0001, Nannan Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | Multi-modal Temporal Relation Network for Video UnderstandingabstractVideo understanding endeavors to generate descriptive texts by analyzing structured semantics from dynamic visual sequences, thus facilitating context-aware reasoning and interpretation. Recent advancements primarily rely on patch-level visual–textual alignment, bridging the gap between visual and textual modalities and therefore enabling more comprehensive reasoning. While promising, they often struggle to capture object-level semantics and temporal dependencies, resulting in limited interpretability and suboptimal compositional understanding. To address these issues, we propose a novel temporal relation framework for multi-modal video understanding, dubbed VideoU-MTR, which explicitly models object-level temporal relations to facilitate fine-grained and coherent representations of cross-frame object interactions. Specifically, we introduce a query-oriented frames identification mechanism that synergistically combines textual and visual attention, allowing the model to dynamically attend to semantically relevant video content across hierarchical levels while effectively filtering out irrelevant information. Furthermore, we employ an explicit temporal relation module to capture fine-grained temporal dependencies and inter-object dynamics by modeling object-centric sequences with time-aware attention and frame-level embeddings. Additionally, we propose a cross-modal alignment adapter that aligns temporally contextualized visual features with linguistic semantics at both object and frame levels. Extensive experiments on eight benchmarks across video question answering (VideoQA), long-term video understanding (LTVU), and video captioning (VideoCap) benchmarks demonstrate that VideoU-MTR achieves superior performance compared to state-of-the-art methods. Moreover, visualization analysis further validates the effectiveness of incorporating temporal information for enhancing video comprehension. Zhixuan Wu, Quanxing Zha, Bo Cheng 0001, Xin Liu 0011, Changbao Li, Pingli Gu |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2025 | Re-Attentional Controllable Video Diffusion EditingabstractEditing videos with textual guidance has garnered popularity due to its streamlined process which mandates users to solely edit the text prompt corresponding to the source video. Recent studies have explored and exploited large-scale text-to-image diffusion models for text-guided video editing, resulting in remarkable video editing capabilities. However, they may still suffer from some limitations such as mislocated objects, incorrect number of objects. Therefore, the controllability of video editing remains a formidable challenge. In this paper, we aim to challenge the above limitations by proposing a Re-Attentional Controllable Video Diffusion Editing (ReAtCo) method. Specially, to align the spatial placement of the target objects with the edited text prompt in a training-free manner, we propose a Re-Attentional Diffusion (RAD) to refocus the cross-attention activation responses between the edited text prompt and the target video during the denoising stage, resulting in a spatially location-aligned and semantically high-fidelity manipulated video. In particular, to faithfully preserve the invariant region content with less border artifacts, we propose an Invariant Region-guided Joint Sampling (IRJS) strategy to mitigate the intrinsic sampling errors w.r.t the invariant regions at each denoising timestep and constrain the generated content to be harmonized with the invariant region content. Experimental results verify that ReAtCo consistently improves the controllability of video diffusion editing and achieves superior video editing performance. Yuanzhi Wang, Yong Li 0032, Xin Liu 0011, Zhen Cui 0001, Antoni B. Chan |
AAAI | 5 |
| 2025 | Scene Graph-Grounded Image GenerationabstractWith the beneft of explicit object-oriented reasoning capabilities of scene graphs, scene graph-to-image generation has made remarkable advancements in comprehending object coherence and interactive relations. Recent state-of-the-arts typically predict the scene layouts as an intermediate representation of a scene graph before synthesizing the image. Nevertheless, transforming a scene graph into an exact layout may restrict its representation capabilities, leading to discrepancies in interactive relationships (such as standing on, wearing, or covering) between the generated image and the input scene graph. In this paper, we propose a Scene Graph-Grounded Image Generation (SGG-IG) method to mitigate the above issues. Specifcally, to enhance the scene graph representation, we design a masked auto-encoder module and a relation embedding learning module to integrate structural knowledge and contextual information of the scene graph with a mask self-supervised manner. Subsequently, to bridge the scene graph with visual content, we introduce a spatial constraint and image-scene alignment constraint to capture the fne-grained visual correlation between the scene graph symbol representation and the corresponding image representation, thereby generating semantically consistent and high-quality images. Extensive experiments demonstrate the effectiveness of the method both quantitatively and qualitatively. Fuyun Wang, Tong Zhang 0021, Yuanzhi Wang, Xin Liu 0011, Zhen Cui 0001 |
AAAI | 5 |
| 2025 | Distribution Prototype Diffusion Learning for Open-set Supervised Anomaly DetectionabstractIn Open-set Supervised Anomaly Detection (OSAD), the existing methods typically generate pseudo anomalies to compensate for the scarcity of observed anomaly samples, while overlooking critical priors of normal samples, leading to less effective discriminative boundaries. To address this issue, we propose a Distribution Prototype Diffusion Learning (DPDL) method aimed at enclosing normal samples within a compact and discriminative distribution space. Specifically, we construct multiple learnable Gaussian prototypes to create a latent representation space for abundant and diverse normal samples and learn a Schrödinger bridge to facilitate a diffusive transition toward these prototypes for normal samples while steering anomaly samples away. Moreover, to enhance inter-sample separation, we design a dispersion feature learning way in hyper-spherical space, which benefits the identification of out-of-distribution anomalies. Experimental results demonstrate the effectiveness and superiority of our proposed DPDL, achieving state-of-the-art performance on 9 public datasets. Fuyun Wang, Tong Zhang 0021, Yuanzhi Wang, Yide Qiu, Xin Liu 0011, Zhen Cui 0001 |
CVPR | 5 |
| 2025 | STDD: Spatio-Temporal Dual Diffusion for Video GenerationabstractDiffusion probabilistic model is becoming the cornerstone of data generation, especially generating high-quality images. As an extension, video diffusion generation is in urgent need of a principled temporal-sequence diffusion way, while the spatial-domain diffusion dominates most video diffusion methods. In this work, we propose an explicit Spatio-Temporal Dual Diffusion (STDD) method by principledly extending the standard diffusion model to a spatio-temporal diffusion model for joint spatial and temporal noise propagation/reduction. Mathematically, an analysable dual diffusion process is derived to accumulate information in temporal sequence as well as spatial domain. Correspondingly, we theoretically derive a spatio-temporal probabilistic reverse diffusion process and propose an accelerated sampling way to reduce the inference cost. In principle, the spatio-temporal dual diffusion enables the information of previous frames to be transferred to the current frame, which thus could be beneficial for video consistency. Extensive experiments demonstrate that our proposed STDD is more competitive over the state-of-the-art methods in the task of video generation/prediction as well as text-to-video generation. Shuaizhen Yao, Xin Liu 0011, Zhen Cui 0001 |
CVPR | 3 |
| 2025 | ReCon: Enhancing True Correspondence Discrimination through Relation Consistency for Robust Noisy Correspondence LearningabstractCan we accurately identify the true correspondences from multimodal datasets containing mismatched data pairs? Existing methods primarily emphasize the similarity matching between the representations of objects across modalities, potentially neglecting the crucial relation consistency within modalities that are particularly important for distinguishing the true and false correspondences. Such an omission often runs the risk of misidentifying negatives as positives, thus leading to unanticipated performance degradation. To address this problem, we propose a general Relation Consistency learning framework, namely ReCon, to accurately discriminate the true correspondences among the multimodal data and thus effectively mitigate the adverse impact caused by mismatches. Specifically, ReCon leverages a novel relation consistency learning to ensure the dual-alignment, respectively of, the cross-modal relation consistency between different modalities and the intra-modal relation consistency within modalities. Thanks to such dual constrains on relations, ReCon significantly enhances its effectiveness for true correspondence discrimination and therefore reliably filters out the mismatched pairs to mitigate the risks of wrong supervisions. Extensive experiments on three widely-used benchmark datasets, including Flickr30K, MS-COCO, and Conceptual Captions, are conducted to demonstrate the effectiveness and superiority of ReCon compared with other SOTAs. The code is available at: https://github.com/qxzha/ReCon. Quanxing Zha, Xin Liu 0011, Shu-Juan Peng, Yiu-Ming Cheung, Xing Xu 0001, Nannan Wang 0001 |
CVPR | 2 |
| 2025 | Egocentric Online Action Segmentation with Behavior-Centred Feature AugmentationabstractThe Egocentric Online Action Segmentation (EOAS) task aims to sequentially segment untrimmed egocentric videos into distinct action segments in a streaming manner. Previous methods primarily focused on improving contextual information utilization, which highly relied on leveraging the prior context. However, under online constraints, the absence of post context limits the effectiveness of the prior context in learning the action semantics. Excessive reliance on prior context may lead to insufficient feature representations of current presented behavior. To tackle this problem, we propose a novel EOAS method, termed Behavior-Centred Feature Augmentation (BCFA), which consists of two key modules: (1) Behavior Prototype Learning models the common sense of each action across different surroundings, enhancing the model’s ability to capture the shared characteristics of behaviors. (2) Presented Behavior Enhancement leverages both the intrinsic characteristics of the current presented behavior itself and the common sense captured by BPL for feature enhancement, mitigating the absence of post contextualization. We evaluate our proposed BCFA method on three public EOAS benchmark datasets, GTEA, EgoProceL, and EgoPER, and demonstrate that our proposed BCFA approach outperforms recent state-of-the-art methods. Zhangye Han, Xun Jiang 0001, Zheng Wang 0044, Xin Liu 0011, Fumin Shen, Xing Xu 0001 |
ICME | 4 |
| 2025 | Noise Mitigation for Unsupervised Cross-Domain Image RetrievalabstractCross-domain image retrieval task is derived from traditional image retrieval task, wherein the model aims to find images in another domain that share the same semantic meaning. Existing methods first achieve the pseudo labels through clustering algorithm, and then perform the cross-domain alignment. However, in unsupervised scenarios, data lacking explicit semantic information can easily induce the model to produce erroneous predictions, which can significantly deteriorate existing clustering-based methods. To mitigate the influence of these noisy instances, we propose the Noise Mitigation (NM) method, utilizing information entropy to separate noisy instances from clear data. Moreover, our label adaption strategy can enhance the prediction accuracy of noisy data by leveraging clear data. We subsequently apply the explicit semantic maximization strategy, selectively construct an intermediate domain through fusing the explicit semantic in clear data, further reducing the influence of noisy instances. Our approach is evaluated on three datasets, and the experimental results demonstrate the overwhelming performance superiority of our noise mitigation strategy. Zheng Wang 0044, Xin Liu 0011, Fumin Shen, Xing Xu 0001 |
ICME | 4 |
| 2025 | Learn and Ensemble Bridge Adapters for Multi-domain Task Incremental LearningabstractMulti-domain task incremental learning (MTIL) demands models to master domain-specific expertise while preserving generalization capabilities.
Inspired by human lifelong learning, which relies on revisiting, aligning, and integrating past experiences, we propose a Learning and Ensembling Bridge Adapters (LEBA) framework.
To facilitate cohesive knowledge transfer across domains, specifically, we propose a continuous-domain bridge adaptation module, leveraging the distribution transfer capabilities of Schrödinger bridge for stable progressive learning.
To strengthen memory consolidation, we further propose a progressive knowledge ensemble strategy that revisits past task representations via a diffusion model and dynamically integrates historical adapters.
For efficiency, LEBA maintains a compact adapter pool through similarity-based selection and employs learnable weights to align replayed samples with current task semantics.
Together, these components effectively mitigate catastrophic forgetting and enhance generalization across tasks.
Extensive experiments across multiple benchmarks validate the effectiveness and superiority of LEBA over state-of-the-art methods. Ziqi Gu, Chunyan Xu, Xin Liu 0011, Yide Qiu, Zhen Cui 0001 |
NeurIPS | 4 |
| 2025 | MPDS: A Movie Posters Dataset for Image Generation with Diffusion Model
Tong Zhang 0021, Fuyun Wang, Xin Liu 0011, Zhen Cui 0001 |
PRCV (2) | 5 |
| 2025 | LPGOH: Label-prototype guided online hashing for efficient cross-modal retrieval
Shu-Juan Peng, Xueting Jiang, Xin Liu 0011, Jixiang Du, Jianjia Cao |
Knowl. Based Syst. | 3 |
| 2025 | Deciphering the Structural Code of Proteins With Deep Graph LearningabstractDeciphering the three-dimensional structure of proteins remains a grand challenge in biology and medicine, as it holds the key to understanding their biological functions and facilitating drug discovery. In this paper, we introduce DECIPHER (Deep Encoding of Cellular Interactions and Protein HiErarchical Representation), a novel deep graph learning framework for protein structure prediction. By representing proteins as graphs, where residues and atoms serve as nodes and their interactions form edges, we capture the intricate spatial relationships within these complex biomolecules. Our framework consists of two complementary modules: 1) a general protein structure prediction module that employs residue and atomic graphs to predict backbone and side-chain conformations, respectively, and utilizes SE(3) transformation for structure optimization; and 2) an antibody-specific structure prediction module that incorporates a dual-track network architecture to model sequence co-evolution and structural template information, coupled with a physics-based energy optimization process. Through extensive experiments on multiple benchmark datasets, we demonstrate that our approach significantly outperforms state-of-the-art methods, setting new standards for accuracy and efficiency in protein structure prediction. By deciphering the structural code of proteins, our work paves the way for accelerated research on protein function and opens up new avenues for rational drug design and discovery. Xiaoyi Yin, Xin Liu 0011, Zhen Cui 0001, Tong Zhang 0021 |
IEEE Trans. Comput. Biol. Bioinform. | 3 |
| 2025 | UCPM: Uncertainty-Guided Cross-Modal Retrieval With Partially Mismatched PairsabstractThe manual annotation of perfectly aligned labels for cross-modal retrieval (CMR) is incredibly labor-intensive. As an alternative, the collection of co-occurring data pairs from the Internet is a remarkably cost-effective way, but which, inevitably induces the Partially Mismatched Pairs (PMPs) and therefore significantly degrades the retrieval performance without particular treatment. Previous efforts often utilize the pair-wise similarity to filter out the mismatched pairs, and such operation is highly sensitive to mismatched or ambiguous data and thus leads to sub-optimal performance. To alleviate these concerns, we propose an efficient approach, termed UCPM, i.e., Uncertainty-guided Cross-modal retrieval with Partially Mismatched pairs, which can significantly reduce the adverse impact of mismatched data pairs. Specifically, a novel Uncertainty Guided Division (UGD) strategy is sophisticatedly designed to divide the corrupted training data into confident matched (clean), easily-identifiable mismatched (noisy) and hardly-determined hard subsets, and the derived uncertainty can simultaneously guide the informative pair learning while reducing the negative impact of potential mismatched pairs. Meanwhile, an effective Uncertainty Self-Correction (USC) mechanism is concurrently presented to accurately identify and rectify the fluctuated uncertainty during the training process, which further improves the stability and reliability of the estimated uncertainty. Besides, a Trusted Margin Loss (TML) is newly designed to enhance the discriminability between those hard pairs, by dynamically adjusting their soft margins to amplify the positive contributions of matched pairs while suppressing the negative impacts of mismatched pairs. Extensive experiments on three widely-used benchmark datasets, verify the effectiveness and reliability of UCPM compared with the existing SOTA approaches, and significantly improve the robustness in both synthetic and real-world PMPs. The code is available at: https://github.com/qxzha/UCPM. Quanxing Zha, Xin Liu 0011, Yiu-Ming Cheung, Shu-Juan Peng, Xing Xu 0001, Nannan Wang 0001 |
IEEE Trans. Image Process. | 2 |
| 2025 | MMHCL: Multi-Modal Hypergraph Contrastive Learning for RecommendationabstractThe burgeoning presence of multimodal content-sharing platforms propels the development of personalized recommender systems. Previous works usually suffer from data sparsity and cold-start problems and may fail to adequately explore semantic user–product associations from multimodal data. To address these issues, we propose a novel Multi-Modal Hypergraph Contrastive Learning (MMHCL) framework for user recommendation. For a comprehensive information exploration from user–product relations, we construct two hypergraphs, i.e., a user-to-user (u2u) hypergraph and an item-to-item (i2i) hypergraph, to mine shared preferences among users and intricate multimodal semantic resemblance among items, respectively. This process yields denser second-order semantics that are fused with first-order user–item interaction as complementary to alleviate the data sparsity issue. Then, we design a contrastive feature enhancement paradigm by applying synergistic contrastive learning. By maximizing/minimizing the mutual information between second-order (e.g., shared preference pattern for users) and first-order (information of selected items for users) embeddings of the same/different users and items, the feature distinguishability can be effectively enhanced. Compared with using sparse primary user–item interaction only, our MMHCL obtains denser second-order hypergraphs and excavates more abundant shared attributes to explore the user–product associations, which to a certain extent alleviates the problems of data sparsity and cold-start. Extensive experiments have comprehensively demonstrated the effectiveness of our method. Our code is publicly available at https://github.com/Xu107/MMHCL . Tong Zhang 0021, Fuyun Wang, Xin Liu 0011, Zhen Cui 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2024 | Unsupervised Cross-Domain Image Retrieval with Semantic-Attended Mixture-of-ExpertsabstractUnsupervised cross-domain image retrieval is designed to facilitate the retrieval between images in different domains in an unsupervised way. Without the guidance of labels, both intra-domain semantic learning and inter-domain semantic alignment pose significant challenges to the model's learning process. The resolution of these challenges relies on the accurate capture of domain-invariant semantic features by the model. Based on this consideration, we propose our Semantic-Attended Mixture of Experts (SA-MoE) model. Leveraging the proficiency of MoE network in capturing visual features, we enhance the model's focus on semantically relevant features through a series of strategies. We first utilize the self-attention mechanism of Vision Transformer to adaptively collect information with different weights on instances from different domains. In addition, we introduce contextual semantic association metrics to more accurately measure the semantic relatedness between instances. By utilizing the association metrics, secondary clustering is performed in the feature space to reinforce semantic relationships. Finally, we employ the metrics for information selection on the fused data to remove the semantic noise. We conduct extensive experiments on three widely used datasets. The consistent comparison results with existing methods indicate that our model possesses the state-of-the-art performance. Xing Xu 0001, Jingkuan Song, Xin Liu 0011, Heng Tao Shen |
SIGIR | 5 |
| 2024 | UGNCL: Uncertainty-Guided Noisy Correspondence Learning for Efficient Cross-Modal Matching
Quanxing Zha, Xin Liu 0011, Yiu-Ming Cheung, Xing Xu 0001, Nannan Wang 0001, Jianjia Cao |
SIGIR | 2 |
| 2024 | Spiking generative networks empowered by multiple dynamic experts for lifelong learning
Wentao Fan 0001, Xin Liu 0011 |
Expert Syst. Appl. | 3 |
| 2024 | Ecarnet: enhanced clue-ambiguity reasoning network for multimodal fake news detection
Shannan Zhong, Shu-Juan Peng, Xin Liu 0011, Lei Zhu 0002, Xing Xu 0001, Taihao Li |
Multim. Syst. | 3 |
| 2024 | OLCH: Online Label Consistent Hashing for streaming cross-modal retrieval
Shu-Juan Peng, Jinhan Yi, Xin Liu 0011, Yiu-Ming Cheung, Zhen Cui 0001, Taihao Li |
Pattern Recognit. | 3 |
| 2024 | Multi-Grained Attention Network With Mutual Exclusion for Composed Query-Based Image RetrievalabstractTheComposed Query-Based Image Retrieval (CQBIR)task aims to precisely obtain the preserved and modified parts, based on the multi-grained semantics learned from the composed query. Since the composed query includes a reference image and the modification text, not just a single modality, this task is more challenging than the general image retrieval tasks. Most previous methods attempt to learn preserved and modified parts via different attention modules and fuse them as a unified representation. However, these methods have two intrinsic drawbacks: 1) The different granular semantic information of the composed query is neglected, which results in the fact that learned preserved and modified parts are irrelevant to correct semantics. 2) The preserved and modified parts learned by previous methods have obvious overlaps, which may lead the model to obtain sub-optimal preserved and modified regions. To this end, we propose a novel method termedMulti-Grained Attention Network with Mutual Exclusion (MANME)to address the above problems. Our MANME method mainly consists of two components: 1) A multi-grained semantic construction for obtaining various textual and visual semantic information. 2) An attention with mutual exclusion constraint for reducing the degree of overlap between preserved and modified parts. It adequately utilizes the various granular semantic information and effectively refines the learned preserved and modified parts. Extensive experiments and further analyses on three widely used CQBIR datasets demonstrate that our proposed MANME method achieves new state-of-the-art performance on the CQBIR task. Shenshen Li, Xing Xu 0001, Xun Jiang 0001, Fumin Shen, Xin Liu 0011, Heng Tao Shen |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Learning Relationship-Enhanced Semantic Graph for Fine-Grained Image-Text MatchingabstractImage-text matching of natural scenes has been a popular research topic in both computer vision and natural language processing communities. Recently, fine-grained image-text matching has shown its significant advance in inferring the high-level semantic correspondence by aggregating pairwise region-word similarity, but it remains challenging mainly due to insufficient representation of high-order semantic concepts and their explicit connections in one modality as its matched in another modality. To tackle this issue, we propose a relationship-enhanced semantic graph (ReSG) model, which can improve the image-text representations by learning their locally discriminative semantic concepts and then organizing their relationships in a contextual order. To be specific, two tailored graph encoders, visual relationship-enhanced graph (VReG) and textual relationship-enhanced graph (TReG), are respectively exploited to encode the high-level semantic concepts of corresponding instances and their semantic relationships. Meanwhile, the representations of each graph node are optimized by aggregating semantically contextual information to enhance the node-level semantic correspondence. Further, the hard-negative triplet ranking loss, center hinge loss, and positive-negative margin loss are jointly leveraged to learn the fine-grained correspondence between the ReSG representations of image and text, whereby the discriminative cross-modal embeddings can be explicitly obtained to benefit various image-text matching tasks in a more interpretable way. Extensive experiments verify the advantages of the proposed fine-grained graph matching approach, by achieving the state-of-the-art image-text matching results on public benchmark datasets. Xin Liu 0011, Yiu-Ming Cheung, Xing Xu 0001, Nannan Wang 0001 |
IEEE Trans. Cybern. | 1 |
| 2024 | Relation-Aggregated Cross-Graph Correlation Learning for Fine-Grained Image-Text RetrievalabstractFine-grained image-text retrieval has been a hot research topic to bridge the vision and languages, and its main challenge is how to learn the semantic correspondence across different modalities. The existing methods mainly focus on learning the global semantic correspondence or intramodal relation correspondence in separate data representations, but which rarely consider the intermodal relation that interactively provide complementary hints for fine-grained semantic correlation learning. To address this issue, we propose a relation-aggregated cross-graph (RACG) model to explicitly learn the fine-grained semantic correspondence by aggregating both intramodal and intermodal relations, which can be well utilized to guide the feature correspondence learning process. More specifically, we first build semantic-embedded graph to explore both fine-grained objects and their relations of different media types, which aim not only to characterize the object appearance in each modality, but also to capture the intrinsic relation information to differentiate intramodal discrepancies. Then, a cross-graph relation encoder is newly designed to explore the intermodal relation across different modalities, which can mutually boost the cross-modal correlations to learn more precise intermodal dependencies. Besides, the feature reconstruction module and multihead similarity alignment are efficiently leveraged to optimize the node-level semantic correspondence, whereby the relation-aggregated cross-modal embeddings between image and text are discriminatively obtained to benefit various image-text retrieval tasks with high retrieval performance. Extensive experiments evaluated on benchmark datasets quantitatively and qualitatively verify the advantages of the proposed framework for fine-grained image-text retrieval and show its competitive performance with the state of the arts. Shu-Juan Peng, Xin Liu 0011, Yiu-Ming Cheung, Xing Xu 0001, Zhen Cui 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Edit Temporal-Consistent Videos with Image Diffusion ModelabstractLarge-scale text-to-image (T2I) diffusion models have been extended for text-guided video editing, yielding impressive zero-shot video editing performance. Nonetheless, the generated videos usually show spatial irregularities and temporal inconsistencies as the temporal characteristics of videos have not been faithfully modeled. In this article, we propose an elegant yet effective Temporal-Consistent Video Editing (TCVE) method to mitigate the temporal inconsistency challenge for robust text-guided video editing. In addition to the utilization of a pretrained T2I 2D Unet for spatial content manipulation, we establish a dedicated temporal Unet architecture to faithfully capture the temporal coherence of the input video sequences. Furthermore, to establish coherence and interrelation between the spatial-focused and temporal-focused components, a cohesive spatial-temporal modeling unit is formulated. This unit effectively interconnects the temporal Unet with the pretrained 2D Unet, thereby enhancing the temporal consistency of the generated videos while preserving the capacity for video content manipulation. Quantitative experimental results and visualization results demonstrate that TCVE achieves state-of-the-art performance in both video temporal consistency and video editing capability, surpassing existing benchmarks in the field. Codes are released at https://github.com/mdswyz/TCVE . Yuanzhi Wang, Yong Li 0032, Xin Liu 0011, Anbo Dai, Antoni B. Chan, Zhen Cui 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2023 | Label-Semantic-Enhanced Online Hashing for Efficient Cross-modal RetrievalabstractExisting online cross-modal hashing methods often treat the semantic label categories independently to correlate the semantically similar data instances, which intrinsically ignore the potential dependency between the label categories and thus fail to capture the discriminative information in the hash code learning process. To alleviate this concern, we explore the inter-dependency between the label categories through their co-occurrence correlation from the label set, and present an efficient Label-Semantic-Enhanced Online Hashing (LSE-OH) method for various cross-modal retrieval task. To be specific, the proposed framework integrates the instance-wise similarity and label-category affinity to incrementally learn the discriminative hash codes for the current arriving data, while updating the hash functions at a streaming manner. Further, an iterative discrete optimization algorithm is derived to mine the inter-dependency between the label categories and discriminatively learn the hash codes without relaxation. Accordingly, the hash codes are adaptively learned online with the high discriminative capability and inter-dependency, while avoiding high computation complexity to process the streaming data. Experimental results show its outstanding performance in comparison with the-state-of-arts. Xueting Jiang, Xin Liu 0011, Yiu-Ming Cheung, Xing Xu 0001, Shukai Zheng, Taihao Li |
ICME | 2 |
| 2023 | A Contrastive Method for Continual Generalized Zero-Shot Learning
Wentao Fan 0001, Xin Liu 0011, Shu-Juan Peng |
IEA/AIE (1) | 3 |
| 2023 | Unsupervised Disentanglement Learning via Dirichlet Variational Autoencoder
Kunxiong Xu, Wentao Fan 0001, Xin Liu 0011 |
IEA/AIE (1) | 3 |
| 2023 | Spiking Generative Networks in Lifelong Learning Environment
Wentao Fan 0001, Xin Liu 0011 |
IEA/AIE (1) | 3 |
| 2023 | Taking a Part for the Whole: An Archetype-agnostic Framework for Voice-Face AssociationabstractVoice-face association is generally specialized as a cross-modal cognitive matching problem, and recent attention has been paid on the feasibility of devising the computational mechanisms for recognizing such associations. Existing works are commonly resorting to the combination of contrastive learning and classification-based loss to correlate the heterogeneous datas. Nevertheless, the reliance on typical features of each category, known as archetypes, derived from the combination suffer from the weak invariance of modality-specific features within the same identity, which might induce a cross-modal joint feature space with calibration deviations. To tackle these problems, this paper presents an efficient Archetype-agnostic framework for reliable voice-face association. First, an Archetype-agnostic Subspace Merging (AaSM) method is carefully designed to perform feature calibration which can well get rid of the archetype dependence to facilitate the mutual perception of datas. Further, an efficient Bilateral Connection Re-gauging scheme is proposed to quantitatively screen and calibrate the biased datas, namely loose pairs that deviate from joint feature space. Besides, an Instance Equilibrium strategy is dynamically derived to optimize the training process on loose data pairs and significantly improve the data utilization. Through the joint exploitation of the above, the proposed framework can well associate the voice-face data to benefit various kinds of cross-modal cognitive tasks. Extensive experiments verify the superiorities of the proposed voice-face association framework and show its competitive performances with the state-of-the-arts. Guancheng Chen, Xin Liu 0011, Xing Xu 0001, Yiu-Ming Cheung, Taihao Li |
ACM Multimedia | 2 |
| 2023 | An ANN-Guided Approach to Task-Free Continual Learning with Spiking Neural Networks
Wentao Fan 0001, Xin Liu 0011 |
PRCV (8) | 3 |
| 2023 | Unsupervised meta-learning via spherical latent representations and dual VAE-GAN
Wentao Fan 0001, Hanyuan Huang, Xin Liu 0011, Shu-Juan Peng |
Appl. Intell. | 4 |
| 2023 | OMGH: Online Manifold-Guided Hashing for Flexible Cross-Modal RetrievalabstractCross-modal hashing hasrecently gained an increasing attention for its efficiency and fast retrieval speed in indexing the multimedia data across different modalities. Nevertheless, the multimedia data points often emerge in a streaming manner, and existing online methods often lack of learning capacity to handle both labeled and unlabeled data.To alleviate these concerns, this paper proposes an Online Manifold-Guided Hashing (OMGH) framework, which can incrementally learn the compact hash code of streaming data while adaptively optimizing the hash function in a streaming manner. To be specific, OMGH first exploits a matrix tri-factorization framework to learn the discriminative hash codes for streaming multi-modal data. Then, an online anchor-based manifold structure is designed to sparsely represent the old data and adaptively guide the hash code learning process, which can wellreduce the complexity in preserving the semantic correlation between the old data and streaming data. Meanwhile, such anchor-based manifold embedding is adaptive to the unsupervised and supervised learning strategies in a flexible way. Besides, an online discrete optimization method is efficiently addressed to incrementally update the hash functions and optimize the hash codes on streaming data points. As a result, the derived hash codes are more semantically meaningful for various online cross-modal retrieval tasks. Extensive experiments verify the advantages of the proposed OMGH model, by achieving and improving the state-of-the-art cross-modal retrieval performances on three benchmark datasets. Xin Liu 0011, Jinhan Yi, Yiu-Ming Cheung, Xing Xu 0001, Zhen Cui 0001 |
IEEE Trans. Multim. | 1 |
| 2022 | Inconsistency Distillation For Consistency: Enhancing Multi-View Clustering via Mutual Contrastive Teacher-Student LeaningabstractMulti-view clustering has attracted more attention recently since many real-world data are comprised of different representations or views. Recent multi-view clustering works mainly exploit the instance consistency to obtain the shared representations across different views, and apply a single-view clustering method to perform data partitions. However, these existing methods often ignore the inconsistency of instance associations within the views, which may enlarge the intra-class diversity among the views and therefore degrade the clustering performance. To address this issue, this paper proposes an efficient mutual contrastive teacher-student leaning (MC-TSL) model to enhance the multi-view clustering, which is the first attempt to study the inconsistency distillation for consistency learning. First, the proposed MC-TSL approach exploits a view-specific encoder with two heads, an instance encoding head and a semantic distillation head, respectively, for capturing the consistent and discriminative feature representations. To be specific, the former head exploits a cross-view contrastive learning method to obtain a redundancy-free consistent representation at the instance level, while the latter head designs a mutual teacher-student learning module to capture the intra-view information at semantic level. By training these two heads in an end-to-end manner, the discriminative multi-view embeddings are efficiently obtained and refined by minimizing the weighted sum of the reconstruction loss, contrastive loss and contrast distillation loss. Extensive experiments verify the superiorities of the proposed MC-TSL framework and show its competitive clustering performances. Dunqiang Liu, Shu-Juan Peng, Xin Liu 0011, Lei Zhu 0002, Zhen Cui 0001, Taihao Li |
ICDM | 3 |
| 2022 | Detach and Enhance: Learning Disentangled Cross-modal Latent Representation for Efficient Face-Voice Association and MatchingabstractMany researches in cognitive science have shown that humans often perform face-voice association for various perception tasks, and some recent data mining works have been designed in emulating such ability intelligently. Nevertheless, most methods often suffer from the degraded performance when there exist semantically irrelevant interference factors across different modalities. To alleviate this concern, this paper presents an efficient Disentangled Cross-modal Latent Representation (DCLR) method to adaptively detach the discriminative feature attributes and enhance the face-voice association. To be specific, the proposed DCLR framework consists of two-stage cross-modal disentangling process. First, the former stage employs the supervised contrastive learning to push the representations of face-voice data from the same person closer while pulling those representations of different person away. Then, the latter stage freezes all the parameters of the former stage, and further innovates a multi-layer orthogonal decoupling scheme to learn the disentangled latent representations, while filtering out the modality-dependent irrelevant factors. Besides, the cross-modal reconstruction loss is further utilized to narrow down the semantic gap between heterogeneous feature expressions. Through the joint exploitation of the above, the proposed framework can well associate the face-voice data to benefit various kinds of cross-modal perception tasks. Extensive experiments verify the superiorities of the proposed face-voice association framework and show its competitive performances. Zhenning Yu, Xin Liu 0011, Yiu-Ming Cheung, Minghang Zhu, Xing Xu 0001, Nannan Wang 0001, Taihao Li |
ICDM | 2 |
| 2022 | Prototype-based Selective Knowledge Distillation for Zero-Shot Sketch Based Image RetrievalabstractZero-Shot Sketch-Based Image Retrieval (ZS-SBIR) is an emerging research task that aims to retrieve data of new classes across sketches and images. It is challenging due to the heterogeneous distributions and the inconsistent semantics across seen and unseen classes of the cross-modal data of sketches and images. To realize knowledge transfer, the latest approaches introduce knowledge distillation, which optimizes the student network through the teacher signal distilled from the teacher network pre-trained on large-scale datasets. However, these methods often ignore the mispredictions of the teacher signal, which may make the model vulnerable when disturbed by the wrong output of the teacher network. To tackle the above issues, we propose a novel method termed Prototype-based Selective Knowledge Distillation (PSKD) for ZS-SBIR. Our PSKD method first learns a set of prototypes to represent categories and then utilizes an instance-level adaptive learning strategy to strengthen semantic relations between categories. Afterwards, a correlation matrix targeted for the downstream task is established through the prototypes. With the learned correlation matrix, the teacher signal given by transformers pre-trained on ImageNet and fine-tuned on the downstream dataset, can be reconstructed to weaken the impact of mispredictions and selectively distill knowledge on the student network. Extensive experiments conducted on three widely-used datasets demonstrate that the proposed PSKD method establishes the new state-of-the-art performance on all datasets for ZS-SBIR. Yifan Wang 0027, Xing Xu 0001, Xin Liu 0011, Weihua Ou, Huimin Lu 0004 |
ACM Multimedia | 4 |
| 2022 | Deep Adaptively-Enhanced Hashing With Discriminative Similarity Guidance for Unsupervised Cross-Modal RetrievalabstractCross-modal hashing that leverages hash functions to project high-dimensional data from different modalities into the compact common hamming space, has shown immeasurable potential in cross-modal retrieval. To ease labor costs, unsupervised cross-modal hashing methods are proposed. However, existing unsupervised methods still suffer from two factors in the optimization of hash functions: 1) similarity guidance, they barely give a clear definition of whether is similar or not between data points, leading to the residual of the redundant information; 2) optimization strategy, they ignore the fact that the similarity learning abilities of different hash functions are different, which makes the hash function of one modality weaker than the hash function of the other modality. To alleviate such limitations, this paper proposes an unsupervised cross-modal hashing method to train hash functions with discriminative similarity guidance and adaptively-enhanced optimization strategy, termed Deep Adaptively-Enhanced Hashing (DAEH). Specifically, to estimate the similarity relations with discriminability, Information Mixed Similarity Estimation (IMSE) is designed by integrating information from distance distributions and the similarity ratio. Moreover, Adaptive Teacher Guided Enhancement (ATGE) optimization strategy is also designed, which employs information theory to discover the weaker hash function and utilizes an extra teacher network to enhance it. Extensive experiments on three benchmark datasets demonstrate the superiority of the proposed DAEH against the state-of-the-arts. Yufeng Shi 0003, Xin Liu 0011, Feng Zheng 0001, Weihua Ou, Xinge You, Qinmu Peng |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | FDDH: Fast Discriminative Discrete Hashing for Large-Scale Cross-Modal RetrievalabstractCross-modal hashing, favored for its effectiveness and efficiency, has received wide attention to facilitating efficient retrieval across different modalities. Nevertheless, most existing methods do not sufficiently exploit the discriminative power of semantic information when learning the hash codes while often involving time-consuming training procedure for handling the large-scale dataset. To tackle these issues, we formulate the learning of similarity-preserving hash codes in terms of orthogonally rotating the semantic data, so as to minimize the quantization loss of mapping such data to hamming space and propose an efficient fast discriminative discrete hashing (FDDH) approach for large-scale cross-modal retrieval. More specifically, FDDH introduces an orthogonal basis to regress the targeted hash codes of training examples to their corresponding semantic labels and utilizes the ε -dragging technique to provide provable large semantic margins. Accordingly, the discriminative power of semantic information can be explicitly captured and maximized. Moreover, an orthogonal transformation scheme is further proposed to map the nonlinear embedding data into the semantic subspace, which can well guarantee the semantic consistency between the data feature and its semantic representation. Consequently, an efficient closed-form solution is derived for discriminative hash code learning, which is very computationally efficient. In addition, an effective and stable online learning strategy is presented for optimizing modality-specific projection functions, featuring adaptivity to different training sizes and streaming data. The proposed FDDH approach theoretically approximates the bi-Lipschitz continuity, runs sufficiently fast, and also significantly improves the retrieval performance over the state-of-the-art methods. The source code is released at https://github.com/starxliu/FDDH. Xin Liu 0011, Xingzhi Wang, Yiu-Ming Cheung |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2022 | Attention-Based Neural Architecture Search for Person Re-IdentificationabstractRecent years have witnessed significant progress of person reidentification (reID) driven by expert-designed deep neural network architectures. Despite the remarkable success, such architectures often suffer from high model complexity and time-consuming pretraining process, as well as the mismatches between the image classification-driven backbones and the reID task. To address these issues, we introduce neural architecture search (NAS) into automatically designing person reID backbones, i.e., reID-NAS, which is achieved via automatically searching attention-based network architectures from scratch. Different from traditional NAS approaches that originated for image classification, we design a reID-based search space as well as a search objective to fit NAS for the reID tasks. In terms of the search space, reID-NAS includes a lightweight attention module to precisely locate arbitrary pedestrian bounding boxes, which is automatically added as attention to the reID architectures. In terms of the search objective, reID-NAS introduces a new retrieval objective to search and train reID architectures from scratch. Finally, we propose a hybrid optimization strategy to improve the search stability in reID-NAS. In our experiments, we validate the effectiveness of different parts in reID-NAS, and show that the architecture searched by reID-NAS achieves a new state of the art, with one order of magnitude fewer parameters on three-person reID datasets. As a concomitant benefit, the reliance on the pretraining process is vastly reduced by reID-NAS, which facilitates one to directly search and train a lightweight reID model from scratch. Qinqin Zhou 0001, Bineng Zhong 0001, Xin Liu 0011, Rongrong Ji |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2021 | Enhancing Audio-Visual Association with Self-Supervised Curriculum LearningabstractThe recent success of audio-visual representations learning can be largely attributed to their pervasive concurrency property, which can be used as a self-supervision signal and extract correlation information. While most recent works focus on capturing the shared associations between the audio and visual modalities, they rarely consider multiple audio and video pairs at once and pay little attention to exploiting the valuable information within each modality. To tackle this problem, we propose a novel audio-visual representation learning method dubbed self-supervised curriculum learning (SSCL) under the teacher-student learning manner. Specifically, taking advantage of contrastive learning, a two-stage scheme is exploited, which transfers the cross-modal information between teacher and student model as a phased process. The proposed SSCL approach regards the pervasive property of audiovisual concurrency as latent supervision and mutually distills the structure knowledge of visual to audio data. Notably, the SSCL method can learn discriminative audio and visual representations for various downstream applications. Extensive experiments conducted on both action video recognition and audio sound recognition tasks show the remarkably improved performance of the SSCL method compared with the state-of-the-art self-supervised audio-visual representation learning methods. Jingran Zhang, Xing Xu 0001, Fumin Shen, Huimin Lu 0001, Xin Liu 0011, Heng Tao Shen |
AAAI | 5 |
| 2021 | Learning To Filter: Siamese Relation Network for Robust TrackingabstractDespite the great success of Siamese-based trackers, their performance under complicated scenarios is still not satisfying, especially when there are distractors. To this end, we propose a novel Siamese relation network, which introduces two efficient modules, i.e. Relation Detector (RD) and Refinement Module (RM). RD performs in a meta-learning way to obtain a learning ability to filter the distractors from the background while RM aims to effectively integrate the proposed RD into the Siamese framework to generate accurate tracking result. Moreover, to further improve the discriminability and robustness of the tracker, we introduce a contrastive training strategy that attempts not only to learn matching the same target but also to learn how to distinguish the different objects. Therefore, our tracker can achieve accurate tracking results when facing background clutters, fast motion, and occlusion. Experimental results on five popular benchmarks, including VOT2018, VOT2019, OTB100, LaSOT, and UAV123, show that the proposed method is effective and can achieve state-of-the-art results. The code will be available at https://github.com/hqucv/siamrn Siyuan Cheng 0003, Bineng Zhong 0001, Guorong Li, Xin Liu 0011, Zhenjun Tang, Xianxian Li, Jing Wang 0049 |
CVPR | 4 |
| 2021 | Distractor-Aware Fast Tracking via Dynamic Convolutions and MOT PhilosophyabstractA practical long-term tracker typically contains three key properties, i.e. an efficient model design, an effective global re-detection strategy and a robust distractor awareness mechanism. However, most state-of-the-art long-term trackers (e.g., Pseudo and re-detecting based ones) do not take all three key properties into account and therefore may either be time-consuming or drift to distractors. To address the issues, we propose a two-task tracking framework (named DMTrack), which utilizes two core components (i.e., one-shot detection and re-identification (re-id) association) to achieve distractor-aware fast tracking via Dynamic convolutions (d-convs) and Multiple object tracking (MOT) philosophy. To achieve precise and fast global detection, we construct a lightweight one-shot detector using a novel dynamic convolutions generation method, which provides a unified and more flexible way for fusing target information into the search field. To distinguish the target from distractors, we resort to the philosophy of MOT to reason distractors explicitly by maintaining all potential similarities’ tracklets. Benefited from the strength of high recall detection and explicit object association, our tracker achieves state-of-the-art performance on the LaSOT, Ox-UvA, TLP, VOT2018LT and VOT2019LT benchmarks and runs in real-time (3x faster than comparisons)1. Zikai Zhang 0003, Bineng Zhong 0001, Shengping Zhang, Zhenjun Tang, Xin Liu 0011, Zhaoxiang Zhang 0001 |
CVPR | 5 |
| 2021 | Multimodal Transformer Networks with Latent Interaction for Audio-Visual Event LocalizationabstractThe task of audio-visual event localization (AVEL) aims to localize a visible and audible event in a video. Previous methods first divide a video into segments and then fuse visual and acoustic features at the segment level via a co-attention mechanism. However, existing methods mostly model relations between individual visual and audio segments in a limitedly short period, which may not cover a longer video duration for better high-level event information modeling. In this paper, we proposed a novel model termed Multimodal Transformer Network with Latent Interaction (MTNLI) to tackle this problem. The proposed MTNLI model employs a multimodal Transformer structure to learn the cross-modality relationships between latent visual and audio summarizations in long segment sequences, which summarize the visual and audio segments into a small number of latent representations to avoid modeling uninformative individual visual-audio relations. The cross-modality information between the latent summarizations is propagated to fuse valuable information from both modalities, which can effectively handle large temporal inconsistent between vision and audio. Our MTNLI method achieves state-of-the-art performance on the benchmark AVE (Audio-Visual Event) dataset for the event localization task. Xing Xu 0001, Xin Liu 0011, Weihua Ou, Huimin Lu 0001 |
ICME | 3 |
| 2021 | Efficient Online Label Consistent Hashing for Large-Scale Cross-Modal RetrievalabstractExisting cross-modal hashing still faces three challenges: (1) Most batch-based methods are unsuitable for processing large-scale and streaming data. (2) Current online methods often suffer from insufficient semantic association, while lacking flexibility to learn the hash functions for varying streaming data. (3) Existing supervised methods always require much computation time or accumulate large quantization loss to learn hash codes. To address above challenges, we present an efficient Online Label Consistent Hashing (OLCH) for cross-modal retrieval, which aims to incrementally learn hash codes for the current arriving data, while updating the hash functions at a streaming manner. To be specific, an on-line semantic representation learning framework is designed to adaptively preserve the semantic similarity across different modalities, and a mini-batch online gradient descent approach associated with forward-backward splitting is developed to optimize the hash functions. Accordingly, the hash codes are adaptively learned online with the high discriminative capability, while avoiding high computation complexity to process the streaming data. Experimental results show its outstanding performance in comparison with the-state-of-arts. Jinhan Yi, Xin Liu 0011, Yiu-Ming Cheung, Xing Xu 0001, Wentao Fan 0001 |
ICME | 2 |
| 2021 | Relationship-Preserving Knowledge Distillation for Zero-Shot Sketch Based Image RetrievalabstractZero-shot sketch-based image retrieval is challenging for the modal gap between distributions of sketches and images and the inconsistency of label spaces during training and testing. Previous methods mitigate the modal gap by projecting sketches and images into a joint embedding space. Most of them also bridge seen and unseen classes by leveraging semantic embeddings, i.e., word vectors and hierarchical similarities. In this paper, we propose Relationship-Preserving Knowledge Distillation (RPKD) to study generalizable embeddings from the perspective of knowledge distillation bypassing the usage of semantic embeddings. In particular, we firstly distill the instance-level knowledge to preserve inter-class relationships without semantic similarities that require extra effort to collect. We also reconcile the contrastive relationships among instances between different embedding spaces, which is complementary to instance-level relationships. Furthermore, embedding-induced supervision, which measures the similarities of an instance to partial class embedding centers from the teacher, is developed to align the student's classification confidences. Extensive experiments conducted on three benchmark ZS-SBIR datasets, i.e., Sketchy, TU-Berlin, and QuickDraw, demonstrate the superiority of our proposed RPKD approach comparing to the state-of-the-art methods. Xing Xu 0001, Zheng Wang 0044, Fumin Shen, Xin Liu 0011 |
ACM Multimedia | 5 |
| 2021 | Online Discriminative Semantic-Preserving Hashing for Large-Scale Cross-Modal Retrieval
Jinhan Yi, Xin Liu 0011 |
PRICAI (1) | 3 |
| 2021 | Cross-Graph Attention Enhanced Multi-Modal Correlation Learning for Fine-Grained Image-Text RetrievalabstractFine-grained Image-text retrieval is challenging but vital technology in the field of multimedia analysis. Existing methods mainly focus on learning the common embedding space of images (or patches) and sentences (or words), whereby their mapping features in such embedding space can be directly measured. Nevertheless, most existing image-text retrieval works rarely consider the shared semantic concepts that potentially correlated the heterogeneous modalities, which can enhance the discriminative power of learning such embedding space. Toward this end, we propose a Cross-Graph Attention model (CGAM) to explicitly learn the shared semantic concepts, which can be well utilized to guide the feature learning process of each modality and promote the common embedding learning. More specifically, we build semantic-embedded graph for each modality, and smooth the discrepancy between two modalities via cross-graph attention model to obtain shared semantic-enhanced features. Meanwhile, we reconstruct image and text features via the shared semantic concepts and original embedding representations, and leverage multi-head mechanism for similarity calculation. Accordingly, the semantic-enhanced cross-modal embedding between image and text is discriminatively obtained to benefit the fine-grained retrieval with high retrieval performance. Extensive experiments evaluated on benchmark datasets show the performance improvements in comparison with state-of-the-arts. Xin Liu 0011, Yiu-Ming Cheung, Shu-Juan Peng, Jinhan Yi, Wentao Fan 0001 |
SIGIR | 2 |
| 2021 | Real-time video dehazing via incremental transmission learning and spatial-temporally coherent regularization
Shu-Juan Peng, Xin Liu 0011, Wentao Fan 0001, Bineng Zhong 0001, Jixiang Du |
Neurocomputing | 3 |
| 2021 | MTFH: A Matrix Tri-Factorization Hashing Framework for Efficient Cross-Modal RetrievalabstractHashing has recently sparked a great revolution in cross-modal retrieval because of its low storage cost and high query speed. Recent cross-modal hashing methods often learn unified or equal-length hash codes to represent the multi-modal data and make them intuitively comparable. However, such unified or equal-length hash representations could inherently sacrifice their representation scalability because the data from different modalities may not have one-to-one correspondence and could be encoded more efficiently by different hash codes of unequal lengths. To mitigate these problems, this paper exploits a related and relatively unexplored problem: encode the heterogeneous data with varying hash lengths and generalize the cross-modal retrieval in various challenging scenarios. To this end, a generalized and flexible cross-modal hashing framework, termed Matrix Tri-Factorization Hashing (MTFH), is proposed to work seamlessly in various settings including paired or unpaired multi-modal data, and equal or varying hash length encoding scenarios. More specifically, MTFH exploits an efficient objective function to flexibly learn the modality-specific hash codes with different length settings, while synchronously learning two semantic correlation matrices to semantically correlate the different hash representations for heterogeneous data comparable. As a result, the derived hash codes are more semantically meaningful for various challenging cross-modal retrieval tasks. Extensive experiments evaluated on public benchmark datasets highlight the superiority of MTFH under various retrieval scenarios and show its competitive performance with the state-of-the-arts. Xin Liu 0011, Zhikai Hu, Haibin Ling, Yiu-Ming Cheung |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2021 | KNN-BLOCK DBSCAN: Fast Clustering for Large-Scale DataabstractLarge-scale data clustering is an essential key for big data problem. However, no current existing approach is “optimal” for big data due to high complexity, which remains it a great challenge. In this article, a simple but fast approximate DBSCAN, namely, KNN-BLOCK DBSCAN, is proposed based on two findings: 1) the problem of identifying whether a point is a core point or not is, in fact, a kNN problem and 2) a point has a similar density distribution to its neighbors, and neighbor points are highly possible to be the same type (core point, border point, or noise). KNN-BLOCK DBSCAN uses a fast approximate kNN algorithm, namely, FLANN, to detect core-blocks (CBs), noncore-blocks, and noise-blocks within which all points have the same type, then a fast algorithm for merging CBs and assigning noncore points to proper clusters is also invented to speedup the clustering process. The experimental results show that KNN-BLOCK DBSCAN is an effective approximate DBSCAN algorithm with high accuracy, and outperforms other current variants of DBSCAN, including ρ-approximate DBSCAN and AnyDBC. Yewang Chen, Lida Zhou, Songwen Pei, Zhiwen Yu 0002, Yi Chen 0007, Xin Liu 0011, Jixiang Du, Naixue Xiong |
IEEE Trans. Syst. Man Cybern. Syst. | 6 |
| 2020 | Hearing like Seeing: Improving Voice-Face Interactions and Associations via Adversarial Deep Semantic Matching NetworkabstractMany cognitive researches have shown that human may 'see voices' or 'hear faces', and such ability can be potentially associated by machine vision and intelligence. However, this research is still under early stage. In this paper, we present a novel adversarial deep semantic matching network for efficient voice-face interactions and associations, which can well learn the correspondence between voices and faces for various cross-modal matching and retrieval tasks. Within the proposed framework, we exploit a simple and efficient adversarial learning architecture to learn the cross-modal embeddings between faces and voices, which consists of two subnetworks, respectively, for generator and discriminator. The former subnetwork is designed to adaptively discriminate the high-level semantical features between voices and faces, in which the triplet loss and multi-modal center loss are in tandem utilized to explicitly regularize the correspondences among them. The latter subnetwork is further leveraged to maximally bridge the semantic gap between the representations of voice and face data, featuring on maintaining the semantic consistency. Through the joint exploitation of the above, the proposed framework can well push representations of voice-face data from the same person closer while pulling those representations of different person away. Extensive experiments empirically show that the proposed approach involves fewer parameters and calculations, adapts various cross-modal matching tasks for voice-face data and brings substantial improvements over the state-of-the-art methods. Xin Liu 0011, Yiu-Ming Cheung, Xing Xu 0001, Bineng Zhong 0001 |
ACM Multimedia | 2 |
| 2020 | A Cooperative Tracker by Fusing Correlation Filter and Siamese Network
Bin Zhou 0004, Xin Liu 0011, Bineng Zhong 0001 |
PRCV (2) | 2 |
| 2020 | Learning Discriminative Joint Embeddings for Efficient Face and Voice AssociationabstractMany cognitive researches have shown the natural possibility of face-voice association, and such potential association has attracted much attention in biometric cross-modal retrieval domain. Nevertheless, the existing methods often fail to explicitly learn the common embeddings for challenging face-voice association tasks. In this paper, we present to learn discriminative joint embedding for face-voice association, which can seamlessly train the face subnetwork and voice subnetwork to learn their high-level semantic features, while correlating them to be compared directly and efficiently. Within the proposed approach, we introduce bi-directional ranking constraint, identity constraint and center constraint to learn the joint face-voice embedding, and adopt bi-directional training strategy to train the deep correlated face-voice model. Meanwhile, an online hard negative mining technique is utilized to discriminatively construct hard triplets in a mini-batch manner, featuring on speeding up the learning process. Accordingly, the proposed approach is adaptive to benefit various face-voice association tasks, including cross-modal verification, 1:2 matching, 1:N matching, and retrieval scenarios. Extensive experiments have shown its improved performances in comparison with the state-of-the-art ones. Xin Liu 0011, Yiu-Ming Cheung, Nannan Wang 0001, Wentao Fan 0001 |
SIGIR | 2 |
| 2020 | Fast density peak clustering for large scale data based on kNN
Yewang Chen, Xiaoliang Hu, Wentao Fan 0001, Lianlian Shen, Xin Liu 0011, Jixiang Du, Haibo Li 0005, Yi Chen 0007, Hailin Li |
Knowl. Based Syst. | 6 |
| 2020 | Semi-supervised discrete hashing for efficient cross-modal retrieval
Xingzhi Wang, Xin Liu 0011, Shu-Juan Peng, Bineng Zhong 0001, Yewang Chen, Jixiang Du |
Multim. Tools Appl. | 2 |
| 2020 | Blockchain-Enabled Contextual Online Learning Under Local Differential Privacy for Coronary Heart Disease Diagnosis in Mobile Edge ComputingabstractDue to the increasing medical data for coronary heart disease (CHD) diagnosis, how to assist doctors to make proper clinical diagnosis has attracted considerable attention. However, it faces many challenges, including personalized diagnosis, high dimensional datasets, clinical privacy concerns and insufficient computing resources. To handle these issues, we propose a novel blockchain-enabled contextual online learning model under local differential privacy for CHD diagnosis in mobile edge computing. Various edge nodes in the network can collaborate with each other to achieve information sharing, which guarantees that CHD diagnosis is suitable and reliable. To support the dynamically increasing dataset, we adopt a top-down tree structure to contain medical records which is partitioned adaptively. Furthermore, we consider patients' contexts (e.g., lifestyle, medical history records, and physical features) to provide more accurate diagnosis. Besides, to protect the privacy of patients and medical transactions without any trusted third party, we utilize the local differential privacy with randomised response mechanism and ensure blockchain-enabled information-sharing authentication under multi-party computation. Based on the theoretical analysis, we confirm that we provide real-time and precious CHD diagnosis for patients with sublinear regret, and achieve efficient privacy protection. The experimental results validate that our algorithm {outperforms} other algorithm benchmarks on running time, error rate and diagnosis accuracy. Xin Liu 0011, Pan Zhou 0001, Tie Qiu 0001, Dapeng Oliver Wu |
IEEE J. Biomed. Health Informatics | 1 |
| 2019 | Fast Semantic Preserving Hashing for Large-Scale Cross-Modal RetrievalabstractMost Cross-modal hashing methods do not sufficiently exploit the discrimination power of semantic information when learning hash codes, while often involving time-consuming training procedures for large-scale dataset. To tackle these issues, we first formulate the learning of similarity-preserving hash codes in terms of orthogonally rotating the semantic data to hamming space, and then propose a novel Fast Semantic Preserving Hashing (FSePH) approach to large-scale cross-modal retrieval. Specifically, FSePH introduces an orthonormal basis to regress the targeted hash codes of training examples to their corresponding reasonably relaxed class labels, featuring significantly reducing the quantization error. Meanwhile, an effective optimization algorithm is derived for modality-specific projection function learning and an efficient closed-form solution for hash code learning, which are computationally tractable. Extensive experiments have shown that the proposed FSePH approach runs sufficiently fast, and also significantly improves the retrieval performances over the state-of-the-arts. Xingzhi Wang, Xin Liu 0011, Shu-Juan Peng, Yiu-Ming Cheung, Zhikai Hu, Nannan Wang 0001 |
ICDM | 2 |
| 2019 | Semi-Supervised Semantic-Preserving Hashing for Efficient Cross-Modal RetrievalabstractCross-modal hashing has recently gained significant popularity to facilitate retrieval across different modalities. With limited label available, this paper presents a novel Semi-Supervised Semantic-Preserving Hashing (S3PH) for flexible cross-modal retrieval. In contrast to most semi-supervised cross-modal hashing works that need to predict the label of unlabeled data, our proposed approach groups the labeled and unlabeled data together, and integrates the relaxed latent subspace learning and semantic-preserving regularization across different modalities. Accordingly, an efficient relaxed objective function is proposed to learn the latent subspaces for both labeled and unlabeled data. Further, an orthogonal rotation matrix is efficiently learned to transform the latent subspace to hash space by minimizing the quantization error. Without sacrificing the retrieval performance, the proposed S3PH method can benefit various kinds of retrieval tasks, i.e., unsupervised, semi-supervised and supervised. Experimental results compared with several competitive algorithms show the effectiveness of the proposed method and its superiority over state-of-the-arts. Xingzhi Wang, Xin Liu 0011, Zhikai Hu, Nannan Wang 0001, Wentao Fan 0001, Jixiang Du |
ICME | 2 |
| 2019 | Triplet Fusion Network Hashing for Unpaired Cross-Modal RetrievalabstractWith the dramatic increase of multi-media data on the Internet, cross-modal retrieval has become an important and valuable task in searching systems. The key challenge of this task is how to build the correlation between multi-modal data. Most existing approaches only focus on dealing with paired data. They use pairwise relationship of multi-modal data for exploring the correlation between them. However, in practice, unpaired data are more common on the Internet but few methods pay attention to them. To utilize both paired and unpaired data, we propose a one-stream framework triplet fusion network hashing (TFNH), which mainly consists of two parts. The first part is a triplet network which is used to handle both kinds of data, with the help of zero padding operation. The second part consists of two data classifiers, which are used to bridge the gap between paired and unpaired data. In addition, we embed manifold learning into the framework for preserving both inter and intra modal similarity, exploring the relationship between unpaired and paired data and bridging the gap between them in learning process. Extensive experiments show that the proposed approach outperforms several state-of-the-art methods on two datasets in paired scenario. We further evaluate its ability of handling unpaired scenario and robustness in regard to pairwise constraint. The results show that even we discard 50% data under the setting in [19], the performance of TFNH is still better than that of other unpaired approaches and that only 70% pairwise relationships are preserved, TFNH can still outperform almost all paired approaches. Zhikai Hu, Xin Liu 0011, Xingzhi Wang, Yiu-Ming Cheung, Nannan Wang 0001, Yewang Chen |
ICMR | 2 |
| 2019 | A nonparametric Bayesian learning model using accelerated variational inference and feature selection
Wentao Fan 0001, Nizar Bouguila, Xin Liu 0011 |
Pattern Anal. Appl. | 3 |
| 2019 | Attention guided deep audio-face fusion for efficient speaker naming
Xin Liu 0011, Jiajia Geng, Haibin Ling, Yiu-Ming Cheung |
Pattern Recognit. | 1 |
| 2019 | Axially Symmetric Data Clustering Through Dirichlet Process Mixture Models of Watson DistributionsabstractThis paper proposes a Bayesian nonparametric framework for clustering axially symmetric data. Our approach is based on a Dirichlet processes mixture model with Watson distributions, which can also be considered as the infinite Watson mixture model. In this paper, first, we extend the finite Watson mixture model into its infinite counterpart based on the framework of truncated Dirichlet process mixture model with a stick-breaking representation. Second, we propose a coordinate ascent mean-field variational inference algorithm that can effectively learn the parameters of our model with closed-form solutions; Third, to cope with a massive data set, we develop a stochastic variational inference algorithm to learn the proposed model through the method of stochastic gradient ascent; Finally, the proposed nonparametric Bayesian model is evaluated through simulated axially symmetric data sets and a real-world application, namely, gene expression data clustering. Wentao Fan 0001, Nizar Bouguila, Jixiang Du, Xin Liu 0011 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2018 | Pixel-Level Character Motion Style Transfer using Conditional Adversarial NetworksabstractIn this paper, we describe a novel method for synthesizing realistic human movement in videos according to different body motion inputs, which are based on conditional GAN and Gram loss. Moreover, we present a character motion style transfer model with two-branch networks to characterize natural video sequences. The first branch is built upon convolutional LSTMs to capture spatio-temporal representations of style video, and the second branch is structured by convolutional networks to extract the spatial feature of content frame image. The entire network is constructed with encoder-decoder architecture to learn the representations for both spatial content and temporal correlations in videos, which can transform a motion style to another given style video. The main benefits of our approach lies in jointly considering the spatio-temporal correlations of motion video and establishing Gram constraint to achieve real-world character motion style transfer. The experiments demonstrate the effectiveness of our proposed motion style transfer approach on real-world video, and the generated motions with pixel-level motion style transfer are of high visual quality. Shu-Juan Peng, Xin Liu 0011 |
CGI | 3 |
| 2018 | Efficient human motion capture data annotation via multi-view spatiotemporal feature fusionabstractThe availability of large motion capture (mocap) data has sparked a great motivation for computer animation, and the task of automatically annotating complex mocap sequences plays an important role in the efficient motion analysis. To this end, this study presents an efficient human mocap data annotation approach by using multi‐view spatiotemporal feature fusion. First, the authors exploit an improved hierarchical aligned cluster analysis algorithm to divide the unknown human mocap sequence into several sub‐motion clips, and each sub‐motion clip incorporates a particular semantic meaning. Then, the two kinds of multi‐view features, namely most informative central distances and most informative geometric angles, are discriminatively extracted and temporally modelled by a Fourier temporal pyramid to complementarily characterise each motion clip. Finally, the authors utilise the discriminant correlation analysis to fuse these two types of motion features and further employ an extreme learning machine to annotate each sub‐motion clip. The extensive experiments tested on the public available database have demonstrated the effectiveness of the proposed approach in comparison with the existing counterparts. Xin Liu 0011, Shu-Juan Peng, Wentao Fan 0001, Jixiang Du |
IET Signal Process. | 1 |
| 2018 | Kernel correlation filters for visual tracking with adaptive fusion of heterogeneous cues
Bineng Zhong 0001, Gu Ouyang, Xin Liu 0011, Ziyi Chen 0001, Cheng Wang 0020 |
Neurocomputing | 5 |
| 2018 | Efficient cross-modal retrieval via flexible supervised collective matrix factorization hashing
Xin Liu 0011, Jixiang Du, Shu-Juan Peng, Wentao Fan 0001 |
Multim. Tools Appl. | 1 |
| 2017 | A hierarchical Dirichlet process mixture of GID Distributions with feature selection for spatio-temporal video modeling and segmentationabstractIn this paper, a hierarchical Dirichlet process (HDP) mixture model of generalized inverted Dirichlet (GID) distributions with an unsupervised feature selection scheme is developed. The proposed model is learned via a principled variational framework and then deployed for video modeling and segmentation. Experimental results show the merits of our developed statistical framework. Wentao Fan 0001, Nizar Bouguila, Xin Liu 0011 |
ICASSP | 3 |
| 2017 | Efficient single image dehazing and denoising: An efficient multi-scale correlated wavelet approach
Xin Liu 0011, Yiu-Ming Cheung, Xinge You, Yuan Yan Tang |
Comput. Vis. Image Underst. | 1 |
| 2017 | Automatic facial flaw detection and retouching via discriminative structure tensorabstractFacial retouching has been increasingly applied in current social media and entertainment industries. In this study, the authors propose an efficient approach to automatically detect and retouch the facial flaws by using discriminative structure tensor. First, a non‐linear structure tensor associated with saliency model is exploited to discriminatively and automatically detect the significant facial flaws. Then, a Gaussian skin model is constructed in YCbCr space and the OSTU operation is simultaneously utilised to precisely mark the facial skin regions, in which the mouth, eyebrows and nostril parts are excluded. Subsequently, diverse structure tensor is employed to discriminatively adjust the inpainting priority and propose a structure tensor‐based inpainting algorithm to retouch the detected flaws. Without manual intervention, the extensive experiments have shown its effectiveness in marking the freckles, blemishes and moles in face images, and the retouching performance is visually pleasing in comparison with state‐of‐the‐art counterparts. Xin Liu 0011, Lu Xie, Bineng Zhong 0001, Jixiang Du, Qinmu Peng |
IET Image Process. | 1 |
| 2017 | Efficient Human Motion Retrieval via Temporal Adjacent Bag of Words and Discriminative Neighborhood Preserving Dictionary LearningabstractHuman motion retrieval from motion capture data forms the fundamental basis for computer animation. In this paper, the authors propose an efficient human motion retrieval approach via temporal adjacent bag of words (TA-BoW) and discriminative neighborhood preserving dictionary learning (DNP-DL). The retrieval process includes two phases: offline training and online retrieval. In the first phase, the original skeleton model is first simplified and then pairwise joint distances are computed to characterize each motion frame. Then, a novel motion descriptor, namely TABoW, is proposed to discriminatively code the motion appearances, through which the articulated complexity and spatiotemporal dimensionality can be greatly reduced. Subsequently, by considering the neighborhood relationships of intraclass structure and the advantage of Fisher criterion, a DNP-DL method is exploited through which each human action can be discriminatively and sparsely represented by a linear combination of such dictionary atoms. In the second phase, a hierarchical retrieval mechanism is used by incorporating the sparse classification and chi-square ranking, whereby the searching range is significantly reduced. The experimental results show that the proposed human motion retrieval approach performs better than the state-of-the-art competing approaches. Xin Liu 0011, Gao-Feng He, Shu-Juan Peng, Yiu-Ming Cheung, Yuan Yan Tang |
IEEE Trans. Hum. Mach. Syst. | 1 |
| 2016 | Efficient single image dehazing via scene-adaptive segmentation and improved dark channel modelabstractIn this paper, we present an efficient single image dehazing approach via scene-adaptive segmentation and improved dark channel model. First, we detect the image depth information and segment the raw image into the close view and distant view. Then, we utilize the minimum channel image of distant view to regularize the atmospheric veil and simultaneously estimate its light value of close view within the haze-opaque area, through which the whole transmission map can be well optimized. Finally, the haze degraded image can be well restored via the atmosphere scattering model. The experimental results have shown that the proposed single image dehazing approach has significantly increased the perceptual visibility of the scene and achieved a better color fidelity visually. Xin Liu 0011, Yiu-Ming Cheung |
IJCNN | 2 |
| 2016 | Lip event detection using oriented histograms of regional optical flow and low rank affinity pursuit
Xin Liu 0011, Yiu-Ming Cheung, Yuan Yan Tang |
Comput. Vis. Image Underst. | 1 |
| 2016 | Scene-adaptive single image dehazing via opening dark channel modelabstractMany traditional dark channel prior based haze removal schemes often suffer from the colour distortion and generate halo artefacts in the remote scenes. To tackle these issues, the authors present an efficient scene‐adaptive single image dehazing approach via opening dark channel model (ODCM). First, the authors detect the image depth information and separate it into close view and distant view. Then, an ODCM is proposed to optimise the whole atmospheric veil, in which the values of close view are regularised by a minimum channel image while the distant parts are estimated by an appropriate lower constant. Accordingly, the transmission map can be further optimised by guide filter and smoothed by domain transform filter. Finally, the haze degraded image can be well restored by the atmosphere scattering model. The extensive experiments have shown that the proposed image dehazing approach has significantly increased the perceptual visibility of the scene and achieved a better colour fidelity visually. Xin Liu 0011, Yuan Yan Tang, Jixiang Du |
IET Image Process. | 1 |
| 2015 | Motion Capture Behavior Recognition via Neighborhood Preserving Dictionary LearningabstractBehavior recognition from large available motion capture data has received wide attention in the computer animation community and is growing increasingly important in recent years. In this paper, we present an efficient motion capture behavior recognition approach via neighborhood preserving dictionary learning. First, we normalize all the motion sequences in the database to make the motion to be comparable. Then, the neighborhood preserving property is exploited using Iterative Nearest Neighbors algorithm and subsequently added as a constraint condition for discriminative dictionary learning, whereby the raw motion frame can be represented as a compact set of atoms consisting of neighborhood preserving characteristics. Finally, the recognition result can be efficiently obtained by sparse coding based classification scheme. Extensive experiments tested on publicly available motion capture databases have demonstrated the accuracy and effectiveness of the proposed approach. Gao-Feng He, Shu-Juan Peng, Xin Liu 0011 |
SMC | 3 |
| 2015 | Unsupervised Facial Pose Grouping via Garbor Subspace Affinity and Self-Tuning Spectral ClusteringabstractFacial pose grouping plays an important role in the video face recognition. In this paper, we present an unsupervised facial pose grouping approach via Garbor subspace affinity and self-tuning spectral clustering. First, we utilize the local normalization method to reduce the impact of uneven illuminations, and then extract the discriminative appearance features via Gabor wavelet representation. Next, the Garbor subspace affinity method is presented to compute an affinity matrix in terms of the pair wise similarity, in which the facial frames of the same pose always share the smaller pair wise similarities. Finally, we employ the self-tuning spectral clustering algorithm to label the affinity matrix, through which the number of pose groups and the corresponding grouping results can be obtained automatically. Without any label priors, the proposed approach is able to well differentiate the distinct facial poses under uneven illuminations, and the experimental results have shown the satisfactory performances. Xin Liu 0011, Yiu-Ming Cheung |
SMC | 1 |
| 2015 | Hierarchical block-based incomplete human mocap data recovery using adaptive nonnegative matrix factorization
Shu-Juan Peng, Gao-Feng He, Xin Liu 0011, Hua-zhen Wang |
Comput. Graph. | 3 |
| 2015 | Online learning 3D context for robust visual tracking
Bineng Zhong 0001, Yingju Shen, Yan Chen 0017, Weibo Xie, Zhen Cui 0001, Hongbo Zhang 0002, Duansheng Chen, Tian Wang 0001, Xin Liu 0011, Shu-Juan Peng, Jin Gou, Jixiang Du, Jing Wang 0049, Wenming Zheng |
Neurocomputing | 9 |
| 2015 | 3D object tracking via image sets and depth-based occlusion detection
Yan Chen 0017, Yingju Shen, Xin Liu 0011, Bineng Zhong 0001 |
Signal Process. | 3 |
| 2014 | Active contours with a joint and region-scalable distribution metric for interactive natural image segmentationabstractIn this study, we present an efficient active contour with a joint and region‐scalable distribution metric for interactive natural image segmentation. First, the authors project a red–green–blue image into the CIELab colour space and employ independent component analysis to select two subspace channels. Then, by initialising the evolving curve interactively in terms of a polygonal curve or multiple polygonal curves, they compute a joint probability distribution associated with a region‐scalable mask to model the regional statistics and propose a simple but effective distribution metric to regularise the active contours. Subsequently, they convert the resultant level set function into binary pattern and find the larger 8‐connected regions as the desired objects. Finally, the selected regions are smoothed with a circular averaging filter such that the final segmentation results can be obtained. The proposed approach not only can deal with the complex appearance and intensity in homogeneity, but also has the advantages of fast convergence and easy implementation. The experiments have shown the precise and reliable segmentation results in comparison with the state‐of‐the‐art competing approaches. Xin Liu 0011, Shu-Juan Peng, Yiu-Ming Cheung, Yuan Yan Tang, Jixiang Du |
IET Image Process. | 1 |
| 2014 | Automatic mitral valve leaflet tracking in Echocardiography via constrained outlier pursuit and region-scalable active contours
Xin Liu 0011, Yiu-Ming Cheung, Shu-Juan Peng, Qinmu Peng |
Neurocomputing | 1 |
| 2014 | Automatic motion capture data denoising via filtered subspace clustering and low rank matrix approximation
Xin Liu 0011, Yiu-Ming Cheung, Shu-Juan Peng, Zhen Cui 0001, Bineng Zhong 0001, Jixiang Du |
Signal Process. | 1 |
| 2014 | Learning Multi-Boosted HMMs for Lip-Password Based Speaker VerificationabstractThis paper proposes a concept of lip motion password (simply called lip-password hereinafter), which is composed of a password embedded in the lip movement and the underlying characteristic of lip motion. It provides a double security to a visual speaker verification system, where the speaker is verified by both of the private password information and the underlying behavioral biometrics of lip motions simultaneously. Accordingly, the target speaker saying the wrong password or an impostor who knows the correct password will be detected and rejected. To this end, we shall present a multi-boosted Hidden Markov model (HMM) learning approach to such a system. Initially, we extract a group of representative visual features to characterize each lip frame. Then, an effective lip motion segmentation algorithm is addressed to segment the lip-password sequence into a small set of distinguishable subunits. Subsequently, we integrate HMMs with boosting learning framework associated with a random subspace method and data sharing scheme to formulate a precise decision boundary for these subunits verification, featuring on high discrimination power. Finally, the lip-password, whether spoken by the target speaker with the pre-registered password or not, is identified based on all the subunit verification results learned from multi-boosted HMMs. The experimental results show that the proposed approach performs favorably compared with the state-of-the-art methods. Xin Liu 0011, Yiu-Ming Cheung |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2013 | Automatic Motion Capture Data Denoising via Filtered Local Subspace Affinity and Low Rank ApproximationabstractIn this paper, we formulate the Motion capture (MoCap) data denoising problem as the concatenation of piecewise motion matrix recovery problem, in which the moving trajectories of each piecewise motion always share the similar subspace representation. To this end, we present an automatic MoCap data denoising approach based on the filtered local subspace affinity (LSA) and low rank approximation. The proposed approach does not need any physical information about the underling structure of MoCap data or require auxiliary data sets for the training priors. The experiments have shown the promising results. Shu-Juan Peng, Xin Liu 0011, Zhen Cui 0001, Zhipeng Xie, Duansheng Chen |
CAD/Graphics | 2 |
| 2012 | A multi-boosted HMM approach to lip password based speaker verificationabstractThis paper presents a multi-boosted Hidden Markov Model (HMM) approach to lip password (i.e. the password embedded in the lip motion) based speaker verification, where the speaker is verified by both of lip password and the underlying characteristics of lip motions. That is, the target speaker saying the wrong password or an impostor even knowing the correct password will be detected as well. To this end, we firstly propose an effective lip motion segmentation algorithm to segment the password sequence into a small set of discrete subunits. Then, we integrate HMMs with boosting learning framework associated with the random subspace method (RSM) and data sharing scheme (DSS) to model the segmental sequence of the input subunit discriminatively so that a precise decision boundary is formulated for these subunits verification. Finally, the speaker is verified based on all verification results of the subunits learned from multi-boosted HMMs. Experimental results show the promising results. Xin Liu 0011, Yiu-Ming Cheung |
ICASSP | 1 |
| 2012 | Subspace based active contours with a joint distribution metric for semi-supervised natural image segmentationabstractIn this paper, we present an efficient active contour with a joint distribution metric for semi-supervised natural image segmentation. Firstly, we project an RGB image into two-dimensional subspace and draw a polygon curve around the Region of Interest (ROI) as the initial evolving curve. Then, we model the regional statistics in terms of joint probability distributions and propose an effective distribution metric to regularize the active contours for evolution. Subsequently, we convert the resultant zero level set function into binary pattern and find all the 8-connected regions. Finally, the largest region is selected as the desired ROI and smoothed with a circular averaging filter so that the corresponding final segmentation result can be obtained. Meanwhile, the proposed approach also features fast convergence and easy implementation in comparison with the traditional methods, which need a laborious process of re-initializing the zero level set in terms of a sign distance function (SDF) periodically. The experiments show the promising results. Shu-Juan Peng, Xin Liu 0011, Yiu-Ming Cheung |
ICASSP | 2 |
| 2012 | A local region based approach to lip tracking
Yiu-Ming Cheung, Xin Liu 0011, Xinge You |
Pattern Recognit. | 2 |
| 2011 | A robust lip tracking algorithm using localized color active contours and deformable modelsabstractLip tracking is crucial to the success of a lipreading recognition system. This paper presents a robust lip tracking algorithm using localized color active contours and deformable models. The proposed approach utilizes a combined semiellipse as the initial evolving curve and computes the localized energies in color space for evolving such that a separation of the original lip image into lip and non-lip regions can be found. Moreover, the dynamic radius selection of the local region is presented with a 16-point deformable model (Wang et al. 2004) to achieve the lip tracking. The proposed approach is adaptive to the movement of the lips from frame to frame, and robust against the appearance of the teeth and tongue. Experiments have shown the promising results. Xin Liu 0011, Yiu-Ming Cheung |
ICASSP | 1 |
| 2011 | Active contours with a novel distribution metric for complex object segmentationabstractIn this paper, we present the efficient region-based active contours with a novel distribution metric for complex object segmentation problems. Unlike most conventional approaches, we model the regional statistics using probability distribution function and propose a simple but effective distribution metric to drive the active contours. Subsequently, the proposed approach speeds up the segmentation process without initializing the zero level set in terms of a sign distance function (SDF) and re-initializing it periodically during the evolution as used in the traditional methods. Some challenging synthetic and real-world images are utilized to evaluate the proposed segmentation algorithm. The experiments show its promising result in comparison with the existing methods. Shu-Juan Peng, Xin Liu 0011, Yiu-Ming Cheung |
ICIP | 2 |
| 2010 | A Lip Contour Extraction Method Using Localized Active Contour Model with Automatic Parameter SelectionabstractLip contour extraction is crucial to the success of a lipreading system. This paper presents a lip contour extraction algorithm using localized active contour model with the automatic selection of proper parameters. The proposed approach utilizes a minimum-bounding ellipse as the initial evolving curve to split the local neighborhoods into the local interior region and the local exterior region, respectively, and then compute the localized energy for evolving and extracting. This method is robust against the uneven illumination, rotation, deformation, and the effects of teeth and tongue. Experiments show its promising result in comparison with the existing methods. Xin Liu 0011, Yiu-Ming Cheung, Meng Li 0015, Hai-Lin Liu 0001 |
ICPR | 1 |