Shu-Juan Peng

dblp:34/10697 · DBLP profile ↗
← Back
32ranked-venue papers
8as first author
17since 2021 · last 2026
0009-0009-8205-1779ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 18 · 4 first-author · 9 since 2021Artificial intelligence and machine learning · 13 · 4 first-author · 9 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Human-computer interaction and ubiquitous computing · 2Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 A cooperative learning method for early fake news detection with social engagement-aware masking encoder
Pingjing Xu, Shu-Juan Peng, Xin Liu 0011, Lei Zhu 0002, Danni Yu
Multim. Syst.2
2026 Efficient image-text retrieval via bi-cross-graph learning and multi-grained alignment
Shenggang Zhou, Xin Liu 0011, Lei Zhu 0002, Shu-Juan Peng, Jixiang Du, Jianjia Cao
Multim. Syst.4
2026 MLCA: Multi-level Correlative Attacks against Deep Cross-Modal Hashing
Xiaohang Fang, Xin Liu 0011, Zhikai Hu, Yiu-Ming Cheung, Shu-Juan Peng, Xing Xu 0001
Pattern Recognit.5
2026 Uncertainty-guided time-frequency feature enhancement for emotion-aware speech-driven 3D facial animation
Xinfa Gong, Shu-Juan Peng, Suwen Xu
Vis. Comput.2
2025 ReCon: Enhancing True Correspondence Discrimination through Relation Consistency for Robust Noisy Correspondence Learning
abstract
Can we accurately identify the true correspondences from multimodal datasets containing mismatched data pairs? Existing methods primarily emphasize the similarity matching between the representations of objects across modalities, potentially neglecting the crucial relation consistency within modalities that are particularly important for distinguishing the true and false correspondences. Such an omission often runs the risk of misidentifying negatives as positives, thus leading to unanticipated performance degradation. To address this problem, we propose a general Relation Consistency learning framework, namely ReCon, to accurately discriminate the true correspondences among the multimodal data and thus effectively mitigate the adverse impact caused by mismatches. Specifically, ReCon leverages a novel relation consistency learning to ensure the dual-alignment, respectively of, the cross-modal relation consistency between different modalities and the intra-modal relation consistency within modalities. Thanks to such dual constrains on relations, ReCon significantly enhances its effectiveness for true correspondence discrimination and therefore reliably filters out the mismatched pairs to mitigate the risks of wrong supervisions. Extensive experiments on three widely-used benchmark datasets, including Flickr30K, MS-COCO, and Conceptual Captions, are conducted to demonstrate the effectiveness and superiority of ReCon compared with other SOTAs. The code is available at: https://github.com/qxzha/ReCon.
Quanxing Zha, Xin Liu 0011, Shu-Juan Peng, Yiu-Ming Cheung, Xing Xu 0001, Nannan Wang 0001
CVPR3
2025 MPCLA: Multi-perspective Cross-Lingual Alignment for Efficient Preference Leakage Mitigating in LLMs
Shu-Juan Peng, Zixiong Lu
PRCV (12)2
2025 LPGOH: Label-prototype guided online hashing for efficient cross-modal retrieval
Shu-Juan Peng, Xueting Jiang, Xin Liu 0011, Jixiang Du, Jianjia Cao
Knowl. Based Syst.1
2025 UCPM: Uncertainty-Guided Cross-Modal Retrieval With Partially Mismatched Pairs
abstract
The manual annotation of perfectly aligned labels for cross-modal retrieval (CMR) is incredibly labor-intensive. As an alternative, the collection of co-occurring data pairs from the Internet is a remarkably cost-effective way, but which, inevitably induces the Partially Mismatched Pairs (PMPs) and therefore significantly degrades the retrieval performance without particular treatment. Previous efforts often utilize the pair-wise similarity to filter out the mismatched pairs, and such operation is highly sensitive to mismatched or ambiguous data and thus leads to sub-optimal performance. To alleviate these concerns, we propose an efficient approach, termed UCPM, i.e., Uncertainty-guided Cross-modal retrieval with Partially Mismatched pairs, which can significantly reduce the adverse impact of mismatched data pairs. Specifically, a novel Uncertainty Guided Division (UGD) strategy is sophisticatedly designed to divide the corrupted training data into confident matched (clean), easily-identifiable mismatched (noisy) and hardly-determined hard subsets, and the derived uncertainty can simultaneously guide the informative pair learning while reducing the negative impact of potential mismatched pairs. Meanwhile, an effective Uncertainty Self-Correction (USC) mechanism is concurrently presented to accurately identify and rectify the fluctuated uncertainty during the training process, which further improves the stability and reliability of the estimated uncertainty. Besides, a Trusted Margin Loss (TML) is newly designed to enhance the discriminability between those hard pairs, by dynamically adjusting their soft margins to amplify the positive contributions of matched pairs while suppressing the negative impacts of mismatched pairs. Extensive experiments on three widely-used benchmark datasets, verify the effectiveness and reliability of UCPM compared with the existing SOTA approaches, and significantly improve the robustness in both synthetic and real-world PMPs. The code is available at: https://github.com/qxzha/UCPM.
Quanxing Zha, Xin Liu 0011, Yiu-Ming Cheung, Shu-Juan Peng, Xing Xu 0001, Nannan Wang 0001
IEEE Trans. Image Process.4
2024 Multi-view anomaly detection via hybrid instance-neighborhood aligning and cross-view reasoning
Luo Tian, Shu-Juan Peng, Xin Liu 0001, Yewang Chen, Jianjia Cao
Multim. Syst.2
2024 Ecarnet: enhanced clue-ambiguity reasoning network for multimodal fake news detection
Shannan Zhong, Shu-Juan Peng, Xin Liu 0011, Lei Zhu 0002, Xing Xu 0001, Taihao Li
Multim. Syst.2
2024 OLCH: Online Label Consistent Hashing for streaming cross-modal retrieval
Shu-Juan Peng, Jinhan Yi, Xin Liu 0011, Yiu-Ming Cheung, Zhen Cui 0001, Taihao Li
Pattern Recognit.1
2024 Relation-Aggregated Cross-Graph Correlation Learning for Fine-Grained Image-Text Retrieval
abstract
Fine-grained image-text retrieval has been a hot research topic to bridge the vision and languages, and its main challenge is how to learn the semantic correspondence across different modalities. The existing methods mainly focus on learning the global semantic correspondence or intramodal relation correspondence in separate data representations, but which rarely consider the intermodal relation that interactively provide complementary hints for fine-grained semantic correlation learning. To address this issue, we propose a relation-aggregated cross-graph (RACG) model to explicitly learn the fine-grained semantic correspondence by aggregating both intramodal and intermodal relations, which can be well utilized to guide the feature correspondence learning process. More specifically, we first build semantic-embedded graph to explore both fine-grained objects and their relations of different media types, which aim not only to characterize the object appearance in each modality, but also to capture the intrinsic relation information to differentiate intramodal discrepancies. Then, a cross-graph relation encoder is newly designed to explore the intermodal relation across different modalities, which can mutually boost the cross-modal correlations to learn more precise intermodal dependencies. Besides, the feature reconstruction module and multihead similarity alignment are efficiently leveraged to optimize the node-level semantic correspondence, whereby the relation-aggregated cross-modal embeddings between image and text are discriminatively obtained to benefit various image-text retrieval tasks with high retrieval performance. Extensive experiments evaluated on benchmark datasets quantitatively and qualitatively verify the advantages of the proposed framework for fine-grained image-text retrieval and show its competitive performance with the state of the arts.
Shu-Juan Peng, Xin Liu 0011, Yiu-Ming Cheung, Xing Xu 0001, Zhen Cui 0001
IEEE Trans. Neural Networks Learn. Syst.1
2023 A Contrastive Method for Continual Generalized Zero-Shot Learning
Wentao Fan 0001, Xin Liu 0011, Shu-Juan Peng
IEA/AIE (1)4
2023 Unsupervised meta-learning via spherical latent representations and dual VAE-GAN
Wentao Fan 0001, Hanyuan Huang, Xin Liu 0011, Shu-Juan Peng
Appl. Intell.5
2022 Inconsistency Distillation For Consistency: Enhancing Multi-View Clustering via Mutual Contrastive Teacher-Student Leaning
abstract
Multi-view clustering has attracted more attention recently since many real-world data are comprised of different representations or views. Recent multi-view clustering works mainly exploit the instance consistency to obtain the shared representations across different views, and apply a single-view clustering method to perform data partitions. However, these existing methods often ignore the inconsistency of instance associations within the views, which may enlarge the intra-class diversity among the views and therefore degrade the clustering performance. To address this issue, this paper proposes an efficient mutual contrastive teacher-student leaning (MC-TSL) model to enhance the multi-view clustering, which is the first attempt to study the inconsistency distillation for consistency learning. First, the proposed MC-TSL approach exploits a view-specific encoder with two heads, an instance encoding head and a semantic distillation head, respectively, for capturing the consistent and discriminative feature representations. To be specific, the former head exploits a cross-view contrastive learning method to obtain a redundancy-free consistent representation at the instance level, while the latter head designs a mutual teacher-student learning module to capture the intra-view information at semantic level. By training these two heads in an end-to-end manner, the discriminative multi-view embeddings are efficiently obtained and refined by minimizing the weighted sum of the reconstruction loss, contrastive loss and contrast distillation loss. Extensive experiments verify the superiorities of the proposed MC-TSL framework and show its competitive clustering performances.
Dunqiang Liu, Shu-Juan Peng, Xin Liu 0011, Lei Zhu 0002, Zhen Cui 0001, Taihao Li
ICDM2
2021 Cross-Graph Attention Enhanced Multi-Modal Correlation Learning for Fine-Grained Image-Text Retrieval
abstract
Fine-grained Image-text retrieval is challenging but vital technology in the field of multimedia analysis. Existing methods mainly focus on learning the common embedding space of images (or patches) and sentences (or words), whereby their mapping features in such embedding space can be directly measured. Nevertheless, most existing image-text retrieval works rarely consider the shared semantic concepts that potentially correlated the heterogeneous modalities, which can enhance the discriminative power of learning such embedding space. Toward this end, we propose a Cross-Graph Attention model (CGAM) to explicitly learn the shared semantic concepts, which can be well utilized to guide the feature learning process of each modality and promote the common embedding learning. More specifically, we build semantic-embedded graph for each modality, and smooth the discrepancy between two modalities via cross-graph attention model to obtain shared semantic-enhanced features. Meanwhile, we reconstruct image and text features via the shared semantic concepts and original embedding representations, and leverage multi-head mechanism for similarity calculation. Accordingly, the semantic-enhanced cross-modal embedding between image and text is discriminatively obtained to benefit the fine-grained retrieval with high retrieval performance. Extensive experiments evaluated on benchmark datasets show the performance improvements in comparison with state-of-the-arts.
Xin Liu 0011, Yiu-Ming Cheung, Shu-Juan Peng, Jinhan Yi, Wentao Fan 0001
SIGIR4
2021 Real-time video dehazing via incremental transmission learning and spatial-temporally coherent regularization
Shu-Juan Peng, Xin Liu 0011, Wentao Fan 0001, Bineng Zhong 0001, Jixiang Du
Neurocomputing1
2020 Semi-supervised discrete hashing for efficient cross-modal retrieval
Xingzhi Wang, Xin Liu 0011, Shu-Juan Peng, Bineng Zhong 0001, Yewang Chen, Jixiang Du
Multim. Tools Appl.3
2019 Fast Semantic Preserving Hashing for Large-Scale Cross-Modal Retrieval
abstract
Most Cross-modal hashing methods do not sufficiently exploit the discrimination power of semantic information when learning hash codes, while often involving time-consuming training procedures for large-scale dataset. To tackle these issues, we first formulate the learning of similarity-preserving hash codes in terms of orthogonally rotating the semantic data to hamming space, and then propose a novel Fast Semantic Preserving Hashing (FSePH) approach to large-scale cross-modal retrieval. Specifically, FSePH introduces an orthonormal basis to regress the targeted hash codes of training examples to their corresponding reasonably relaxed class labels, featuring significantly reducing the quantization error. Meanwhile, an effective optimization algorithm is derived for modality-specific projection function learning and an efficient closed-form solution for hash code learning, which are computationally tractable. Extensive experiments have shown that the proposed FSePH approach runs sufficiently fast, and also significantly improves the retrieval performances over the state-of-the-arts.
Xingzhi Wang, Xin Liu 0011, Shu-Juan Peng, Yiu-Ming Cheung, Zhikai Hu, Nannan Wang 0001
ICDM3
2018 Pixel-Level Character Motion Style Transfer using Conditional Adversarial Networks
abstract
In this paper, we describe a novel method for synthesizing realistic human movement in videos according to different body motion inputs, which are based on conditional GAN and Gram loss. Moreover, we present a character motion style transfer model with two-branch networks to characterize natural video sequences. The first branch is built upon convolutional LSTMs to capture spatio-temporal representations of style video, and the second branch is structured by convolutional networks to extract the spatial feature of content frame image. The entire network is constructed with encoder-decoder architecture to learn the representations for both spatial content and temporal correlations in videos, which can transform a motion style to another given style video. The main benefits of our approach lies in jointly considering the spatio-temporal correlations of motion video and establishing Gram constraint to achieve real-world character motion style transfer. The experiments demonstrate the effectiveness of our proposed motion style transfer approach on real-world video, and the generated motions with pixel-level motion style transfer are of high visual quality.
Shu-Juan Peng, Xin Liu 0011
CGI2
2018 Efficient human motion capture data annotation via multi-view spatiotemporal feature fusion
abstract
The availability of large motion capture (mocap) data has sparked a great motivation for computer animation, and the task of automatically annotating complex mocap sequences plays an important role in the efficient motion analysis. To this end, this study presents an efficient human mocap data annotation approach by using multi‐view spatiotemporal feature fusion. First, the authors exploit an improved hierarchical aligned cluster analysis algorithm to divide the unknown human mocap sequence into several sub‐motion clips, and each sub‐motion clip incorporates a particular semantic meaning. Then, the two kinds of multi‐view features, namely most informative central distances and most informative geometric angles, are discriminatively extracted and temporally modelled by a Fourier temporal pyramid to complementarily characterise each motion clip. Finally, the authors utilise the discriminant correlation analysis to fuse these two types of motion features and further employ an extreme learning machine to annotate each sub‐motion clip. The extensive experiments tested on the public available database have demonstrated the effectiveness of the proposed approach in comparison with the existing counterparts.
Xin Liu 0011, Shu-Juan Peng, Wentao Fan 0001, Jixiang Du
IET Signal Process.3
2018 Efficient cross-modal retrieval via flexible supervised collective matrix factorization hashing
Xin Liu 0011, Jixiang Du, Shu-Juan Peng, Wentao Fan 0001
Multim. Tools Appl.4
2017 Efficient Human Motion Retrieval via Temporal Adjacent Bag of Words and Discriminative Neighborhood Preserving Dictionary Learning
abstract
Human motion retrieval from motion capture data forms the fundamental basis for computer animation. In this paper, the authors propose an efficient human motion retrieval approach via temporal adjacent bag of words (TA-BoW) and discriminative neighborhood preserving dictionary learning (DNP-DL). The retrieval process includes two phases: offline training and online retrieval. In the first phase, the original skeleton model is first simplified and then pairwise joint distances are computed to characterize each motion frame. Then, a novel motion descriptor, namely TABoW, is proposed to discriminatively code the motion appearances, through which the articulated complexity and spatiotemporal dimensionality can be greatly reduced. Subsequently, by considering the neighborhood relationships of intraclass structure and the advantage of Fisher criterion, a DNP-DL method is exploited through which each human action can be discriminatively and sparsely represented by a linear combination of such dictionary atoms. In the second phase, a hierarchical retrieval mechanism is used by incorporating the sparse classification and chi-square ranking, whereby the searching range is significantly reduced. The experimental results show that the proposed human motion retrieval approach performs better than the state-of-the-art competing approaches.
Xin Liu 0011, Gao-Feng He, Shu-Juan Peng, Yiu-Ming Cheung, Yuan Yan Tang
IEEE Trans. Hum. Mach. Syst.3
2015 Motion Capture Behavior Recognition via Neighborhood Preserving Dictionary Learning
abstract
Behavior recognition from large available motion capture data has received wide attention in the computer animation community and is growing increasingly important in recent years. In this paper, we present an efficient motion capture behavior recognition approach via neighborhood preserving dictionary learning. First, we normalize all the motion sequences in the database to make the motion to be comparable. Then, the neighborhood preserving property is exploited using Iterative Nearest Neighbors algorithm and subsequently added as a constraint condition for discriminative dictionary learning, whereby the raw motion frame can be represented as a compact set of atoms consisting of neighborhood preserving characteristics. Finally, the recognition result can be efficiently obtained by sparse coding based classification scheme. Extensive experiments tested on publicly available motion capture databases have demonstrated the accuracy and effectiveness of the proposed approach.
Gao-Feng He, Shu-Juan Peng, Xin Liu 0011
SMC2
2015 Hierarchical block-based incomplete human mocap data recovery using adaptive nonnegative matrix factorization
Shu-Juan Peng, Gao-Feng He, Xin Liu 0011, Hua-zhen Wang
Comput. Graph.1
2015 Online learning 3D context for robust visual tracking
Bineng Zhong 0001, Yingju Shen, Yan Chen 0017, Weibo Xie, Zhen Cui 0001, Hongbo Zhang 0002, Duansheng Chen, Tian Wang 0001, Xin Liu 0011, Shu-Juan Peng, Jin Gou, Jixiang Du, Jing Wang 0049, Wenming Zheng
Neurocomputing10
2014 Active contours with a joint and region-scalable distribution metric for interactive natural image segmentation
abstract
In this study, we present an efficient active contour with a joint and region‐scalable distribution metric for interactive natural image segmentation. First, the authors project a red–green–blue image into the CIELab colour space and employ independent component analysis to select two subspace channels. Then, by initialising the evolving curve interactively in terms of a polygonal curve or multiple polygonal curves, they compute a joint probability distribution associated with a region‐scalable mask to model the regional statistics and propose a simple but effective distribution metric to regularise the active contours. Subsequently, they convert the resultant level set function into binary pattern and find the larger 8‐connected regions as the desired objects. Finally, the selected regions are smoothed with a circular averaging filter such that the final segmentation results can be obtained. The proposed approach not only can deal with the complex appearance and intensity in homogeneity, but also has the advantages of fast convergence and easy implementation. The experiments have shown the precise and reliable segmentation results in comparison with the state‐of‐the‐art competing approaches.
Xin Liu 0011, Shu-Juan Peng, Yiu-Ming Cheung, Yuan Yan Tang, Jixiang Du
IET Image Process.2
2014 Automatic mitral valve leaflet tracking in Echocardiography via constrained outlier pursuit and region-scalable active contours
Xin Liu 0011, Yiu-Ming Cheung, Shu-Juan Peng, Qinmu Peng
Neurocomputing3
2014 Automatic motion capture data denoising via filtered subspace clustering and low rank matrix approximation
Xin Liu 0011, Yiu-Ming Cheung, Shu-Juan Peng, Zhen Cui 0001, Bineng Zhong 0001, Jixiang Du
Signal Process.3
2013 Automatic Motion Capture Data Denoising via Filtered Local Subspace Affinity and Low Rank Approximation
abstract
In this paper, we formulate the Motion capture (MoCap) data denoising problem as the concatenation of piecewise motion matrix recovery problem, in which the moving trajectories of each piecewise motion always share the similar subspace representation. To this end, we present an automatic MoCap data denoising approach based on the filtered local subspace affinity (LSA) and low rank approximation. The proposed approach does not need any physical information about the underling structure of MoCap data or require auxiliary data sets for the training priors. The experiments have shown the promising results.
Shu-Juan Peng, Xin Liu 0011, Zhen Cui 0001, Zhipeng Xie, Duansheng Chen
CAD/Graphics1
2012 Subspace based active contours with a joint distribution metric for semi-supervised natural image segmentation
abstract
In this paper, we present an efficient active contour with a joint distribution metric for semi-supervised natural image segmentation. Firstly, we project an RGB image into two-dimensional subspace and draw a polygon curve around the Region of Interest (ROI) as the initial evolving curve. Then, we model the regional statistics in terms of joint probability distributions and propose an effective distribution metric to regularize the active contours for evolution. Subsequently, we convert the resultant zero level set function into binary pattern and find all the 8-connected regions. Finally, the largest region is selected as the desired ROI and smoothed with a circular averaging filter so that the corresponding final segmentation result can be obtained. Meanwhile, the proposed approach also features fast convergence and easy implementation in comparison with the traditional methods, which need a laborious process of re-initializing the zero level set in terms of a sign distance function (SDF) periodically. The experiments show the promising results.
Shu-Juan Peng, Xin Liu 0011, Yiu-Ming Cheung
ICASSP1
2011 Active contours with a novel distribution metric for complex object segmentation
abstract
In this paper, we present the efficient region-based active contours with a novel distribution metric for complex object segmentation problems. Unlike most conventional approaches, we model the regional statistics using probability distribution function and propose a simple but effective distribution metric to drive the active contours. Subsequently, the proposed approach speeds up the segmentation process without initializing the zero level set in terms of a sign distance function (SDF) and re-initializing it periodically during the evolution as used in the traditional methods. Some challenging synthetic and real-world images are utilized to evaluate the proposed segmentation algorithm. The experiments show its promising result in comparison with the existing methods.
Shu-Juan Peng, Xin Liu 0011, Yiu-Ming Cheung
ICIP1