VLDB 2026 Research / reviewers in the wild / expert
Yan Liu 0004
dblp:l/YanLiu4 · also Fiona Yan Liu
· DBLP profile ↗
121ranked-venue papers
12as first author
38since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 72 · 7 first-author · 20 since 2021Artificial intelligence and machine learning · 54 · 5 first-author · 24 since 2021Databases, data management, data science and information retrieval · 9 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Systems, architecture and hardware · 3Computer networks · 3Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Is Symbolic Music a Specific Language? Exploring Inspiration-to-Structure Machine Composition via LLMsabstractLarge Language Models (LLMs) have demonstrated remarkable proficiency in diverse tasks. This success raises a fundamental question in machine composition: Can symbolic music be considered a special form of language that can be jointly modeled with natural language for composition tasks? Recent studies validate that symbolic music can be modeled as a human language, yet composing structured music from partial symbolic inputs through natural language interaction remains underexplored. Even LLMs struggle to generate structurally coherent compositions in such hybrid input-output scenarios, highlighting a fundamental gap that calls for a domain-specific learning paradigm. To this end, we propose Inspiration-to-Structure (IoS), a cognitively inspired framework that enables LLMs to generate structured musical sections from melodic ideas. IoS employs a three-phase process—semantic, structural, and collaborative cognition—and is supported by two key components: (1) a new dataset and construction protocol called Structured Triplet Data (STD), and (2) a training method, Dual-Instance Structural Contrastive Optimization (DiSCO), designed to enhance structural awareness. Experiments show that IoS improves structural coherence by 47.8% and artistic creativity by 21.8% compared to conventional language modeling paradigm, supervised fine-tuning, and even enables smaller LLMs to surpass larger LLMs. These results suggest that symbolic music, while language-like, demands specialized modeling beyond standard language modeling paradigms. IoS enables LLMs to transform music theory knowledge into structured composition, empowering users to compose music interactively via language and advancing toward general creative AI. Zhejing Hu, Yan Liu 0004, Zhi Zhang 0004, Aiwei Zhang, Shenghua Zhong, Bruce X. B. Yu, Gong Chen 0006 |
AAAI | 2 |
| 2026 | From Pixels to Logic: A Perception-Reasoning Decomposition Framework for Open-World Referring Expression ComprehensionabstractRecent advances in Referring Expression Comprehension (REC) have been largely driven by supervised learning on curated datasets, where each expression is assumed to refer to exactly one known object. However, such assumptions rarely hold in real-world scenarios, where expressions can refer to multiple objects, fail to refer to any, or involve novel categories and complex semantics. These challenges define the task of open-world REC, which demands robust generalization and structured reasoning beyond the scope of traditional REC methods. In this work, we introduce a novel, training-free framework that decouples visual perception from linguistic reasoning to address open-world REC. Our method first transforms the visual scene into a rich textual representation using an open-vocabulary multimodal perception module. It then employs a reasoning language model to interpret the referring expression and perform explicit logical inference over the perceived scene, enabling transparent decision-making and strong generalization in open-world scenarios. Experiments on three standard REC benchmarks as well as two more challenging ones, gRefCOCO and D³, demonstrate that our framework achieves highly competitive zero-shot performance, often surpassing supervised baselines. Shenghua Zhong, Zhi Zhang 0004, Yan Liu 0004 |
AAAI | 4 |
| 2026 | MusicScaffold: Bridging Machine Efficiency and Human Growth in Adolescent Creative Education through Generative AIabstractAdolescence is marked by strong creative impulses but limited strategies for structured expression, often leading to frustration or disengagement. While generative AI lowers technical barriers and delivers efficient outputs, its role in fostering adolescents’ expressive growth has been overlooked. We propose MusicScaffold, an adolescent-centered framework that transforms classical AI roles from broad conceptualizations into stage-specific, actionable developmental scaffolds designed to make expressive strategies transparent and learnable and to support adolescents in mastering creative expression. In a four-week study with middle school students (ages 12–14), MusicScaffold enhanced cognitive specificity, behavioral regulation, and affective autonomy in music creation. By reframing generative AI as a scaffold rather than a generator, this work bridges the machine efficiency of generative systems with human growth in adolescent creativity education. Zhejing Hu, Yan Liu 0004, Zhi Zhang 0004, Gong Chen 0006, Bruce X. B. Yu, Jiannong Cao 0001 |
CHI | 2 |
| 2026 | SliceCSRef: Dual-Level Semantic Alignment for Robust Speech Referring Expression ComprehensionabstractSpeech Referring Expression Comprehension (SREC) aims to localize the object in an image referred to by a spoken natural language query. However, raw speech is continuous and noisy, and prior ASR-free methods that align full utterances with transcripts using only global supervision can overfit to spurious correlations, limiting fine-grained grounding. To address this issue, we propose SliceCSRef, a robust SREC framework that improves generalization via dual-level semantic alignment. Beyond utterance-level speech–text alignment, SliceCSRef introduces slice-wise alignment that pairs randomly sampled speech segments with weakly matched transcript spans based on their relative temporal positions, providing fine-grained supervision without additional annotations. Experiments on six benchmarks show that SliceCSRef achieves state-of-the-art performance under standard settings and consistently improves robustness under truncated speech and playback-speed variations. Shenghua Zhong, Qiao Yan, Zhijiao Xiao, Yan Liu 0004 |
ICMR | 5 |
| 2026 | SalMIM: Saliency-Guided Masked Image Modeling Network for Visual Emotion AnalysisabstractCompared to conventional image content analysis tasks, visual emotion analysis is perceived as a complex, abstract, and potentially culturally dependent endeavor. The accuracy of automatic image-based emotion recognition remains a challenge, and the most significant obstacles are the affective gap and scarcity of data, particularly labeled data. To address these challenge, this paper proposes a saliency-guided masked image modeling approach. Specifically, the proposed framework employs multi-modal large model to generate more emotional images, thereby reducing the impact of data scarcity on model performance. Subsequently, neuroimaging and behavioral studies have demonstrated that human visual attention is attracted by the emotional relevance of a stimulus. In light of this, our model employs a saliency-guided masking strategy to identify emotion-related regions for masking sampling to fit the affective gap. In contrast to the conventional approach of using the original pixel values for the reconstruction target, our model eliminates high-frequency components from the pixels, thus enhancing the generalizability of the model. The use of this unsupervised representation learning approach enables the model to exhibit outstanding recognition performance in downstream emotion recognition tasks on three standard emotion datasets. Furthermore, ablation experiments, robustness test, and visualization experiments corroborate the effectiveness of the proposed method. Weiye Peng, Shenghua Zhong, Ahmed Fares, Yan Liu 0004 |
IEEE Trans. Affect. Comput. | 4 |
| 2026 | SEmotion: Knowledge-Guided Sigmoid-Constrained Network for EEG-Based Emotion RecognitionabstractElectroencephalography (EEG)-based affective computing has achieved significant progress due to the rapidly developing learning model. However, compared to computer vision or natural language processing-related tasks, EEG data presents challenges such as a low signal-to-noise ratio, small sample size, and nonstationary properties. These factors may affect the model’s performance on cross-domain tasks, suggesting that further improvements are necessary. Existing models aim to minimize intraclass distance and maximize interclass distance to achieve an optimal solution. This can potentially have a negative impact on the model’s generalization ability because they ignore the existence of low-quality training samples that may not be suitable for strict optimization. This article presents a method for determining sample quality based on the guidance of knowledge from emotional neuroscience. The method differentiates between high- and low-quality training samples and designs a corresponding loss function to impose intraclass and interclass constraints on hyperspherical manifolds based on the quality of the samples. Therefore, it can be argued that our proposed method achieves a better balance between reducing the intraclass distance of high-quality samples and preventing the overfitting of low-quality ones. This could potentially help build a more robust EEG emotion recognition model. Numerous experiments are conducted on the SJTU Emotion EEG Dataset (SEED) and SJTU Emotion EEG Dataset (IV) (SEED-IV) datasets under cross-subject and cross-session scenarios, which show the superior performance of our proposed method and its higher recognition accuracy against adversarial attacks. Wenjie Rao, Shenghua Zhong, Zhi Zhang 0004, Yan Liu 0004 |
IEEE Trans. Comput. Soc. Syst. | 4 |
| 2026 | Mutual Generation for Cross-Domain Challenge in Stroke Patients' Motor Imagery Classification and Functional Recovery PredictionabstractThe accumulating body of research indicates that Motor Imagery (MI)-BCIs have the potential to enhance the quality of life for individuals with disabilities and to advance our understanding of brain function and rehabilitation strategies. Among these diseases, stroke is the leading cause of long-term motor disability across the globe, thereby underscoring the need for innovative rehabilitation strategies, such as MI-BCI technologies. In contrast with these expectations, the majority of existing research is built upon data obtained from healthy subjects. The construction of effective classification models for Motor Imagery tasks in patients with brain diseases, particularly stroke, remains a significant challenge. The lateralization of the left and right hemispheres is more pronounced in patients who have suffered a stroke than in healthy individuals. Moreover, the specific locations of lesions and the regions of influence result in significant variations in the electroencephalogram (EEG) data of patients with different hemiplegic sides. This paper explores the potential of generative models in addressing the issue of domain differences arising from different hemiplegic sides EEG data. Furthermore, this paper circumvents the potential adverse effects of rigorous optimization of low-quality samples on model performance through the utilization of label softening algorithm. Two MI-EEG datasets of stroke patients performing Motor Imagery tasks are used to validate our method. In comparison to both classical machine learning methods and those state-of-the-art models for MI classification, the classification model in this paper achieves a noticeable performance improvement in different data partitioning strategies, including subject-dependent and subject-independent scenarios. Each sub-module, and each designed loss function, contributes to the final performance growth. In addition, this paper also investigates the potential of the proposed framework for predicting a patient's level of functional recovery. Our findings indicate that the addition of a prediction layer to the proposed model enables the accurate prediction of functional recovery level in stroke patients. Rongrong Lu, Wenchang Deng, Tianhao Gao, Songhua Huang, Zhi Zhang 0004, Yan Liu 0004, Shenghua Zhong |
IEEE J. Biomed. Health Informatics | 6 |
| 2025 | Mixture of Knowledge Minigraph Agents for Literature Review GenerationabstractLiterature reviews play a crucial role in scientific research for understanding the current state of research, identifying gaps, and guiding future studies on specific topics. However, the process of conducting a comprehensive literature review is yet time-consuming. This paper proposes a novel framework, collaborative knowledge minigraph agents (CKMAs), to automate scholarly literature reviews. A novel prompt-based algorithm, the knowledge minigraph construction agent (KMCA), is designed to identify relations between concepts from academic literature and automatically constructs knowledge minigraphs. By leveraging the capabilities of large language models on constructed knowledge minigraphs, the multiple path summarization agent (MPSA) efficiently organizes concepts and relations from different viewpoints to generate literature review paragraphs. We evaluate CKMAs on three benchmark datasets. Experimental results show the effectiveness of the proposed method, further revealing promising applications of LLMs in scientific research. Zhi Zhang 0004, Yan Liu 0004, Shenghua Zhong, Gong Chen 0006, Yu Yang 0012, Jiannong Cao 0001 |
AAAI | 2 |
| 2025 | Inductive Cognitive Diagnosis Model Integrating the Characteristics of Learners' Answering BehaviorabstractCognitive diagnosis (CD) aims to identify the extent to which learners have mastered different knowledge concepts, which is an essential task for online education platforms. Due to the increase in educational data and advancements in deep learning technologies, numerous CD models have been proposed, achieving promising results. However, most existing models lack a comprehensive understanding of learners’ behavioral patterns, focusing primarily on global knowledge structures while overlooking fine-grained response behaviors. Therefore, we propose LAB-ICDM, an Inductive Cognitive Diagnosis Model Integrating the Characteristics of Learners’ Answering Behavior. Specifically, (i) we introduce a global relation aggregation module, which constructs a learner-centered graph to capture global representations of learners; (ii) we design a relation-aware module, which leverages a learner-exercise graph to extract fine-grained behavioral patterns, enabling a more comprehensive representation of learners’ exercise-solving behaviors and their interactions with exercises. By fusing these two modules, our model constructs a multi-view representation of learners’ knowledge states, enabling more accurate cognitive diagnosis. Extensive experiments on three real-world educational datasets demonstrate that LAB-ICDM significantly outperforms existing state-of-the-art models, showcasing superior predictive performance. Linhao Zhou, Shenghua Zhong, Yan Liu 0004 |
ECAI | 3 |
| 2025 | Improving Generated and Retrieved Knowledge Combination Through Zero-shot GenerationabstractOpen-domain Question Answering (QA) has garnered substantial interest by combining the advantages of faithfully retrieved passages and relevant passages generated through Large Language Models (LLMs). However, there is a lack of definitive labels available to pair these sources of knowledge. In order to address this issue, we propose an unsupervised and simple framework called Bi-Reranking for Merging Generated and Retrieved Knowledge (BRMGR), which utilizes re-ranking methods for both retrieved passages and LLM-generated passages. We pair the two types of passages using two separate re-ranking methods and then combine them through greedy matching. We demonstrate that BRMGR is equivalent to employing a bipartite matching loss when assigning each retrieved passage with a corresponding LLM-generated passage. The application of our model yielded experimental results from three datasets, improving their performance by +1.7 and +1.6 on NQ and WebQ datasets, respectively, and obtaining comparable result on TriviaQA dataset when compared to competitive baselines. Xinkai Du, Quanjie Han, Yan Liu 0004, Yalin Sun, Hongbo Shan, Maosong Sun 0001 |
ICASSP | 4 |
| 2025 | Attention-guided universal adversarial perturbations for EEG-based brain-computer interfaces
Shenghua Zhong, Sijia Zhao, Zhijiao Xiao, Zhi Zhang 0004, Yan Liu 0004 |
Expert Syst. Appl. | 5 |
| 2025 | Intellectual Property Protection for Deep Models: Pioneering Cross-Domain Fingerprinting SolutionsabstractThe high cost of developing high-performance deep models highlights their value as intellectual property for creators. However, it is important to consider the potential risks of theft. Although various techniques have been developed to protect the intellectual property of deep models, there is still room for improvement in terms of efficiency, comprehensiveness, and generalization. Compared with the intrusiveness of watermarking methods, fingerprinting methods do not affect the training process of the source model. Consequently, this paper proposes a fingerprinting method to address the paucity of attempts in fingerprinting methods for model protection. Our method consists of two efficient algorithms for generating fingerprinting samples, where the first one possesses the advantage of efficiency, while the second one is better in terms of robustness. The first algorithm takes a comprehensive approach to modeling the fingerprint of the deep model. The generated samples are distributed within the stable region and near the decision boundary of the model, taking into account both the duality and the conviction factors. Then, a heuristic sample perturbation algorithm is introduced, which generates a fingerprint with solid stability and generalization across multiple domains. The two algorithms proposed in this paper have been shown to be capable of withstanding attacks on intellectual property removal, detection, and evasion. They also show some advantages in terms of efficiency. In addition, the proposed method is the first to apply fingerprinting techniques in a cross-domain context. Shenghua Zhong, Zhi Zhang 0004, Yan Liu 0004 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2024 | Responding to the Call: Exploring Automatic Music Composition Using a Knowledge-Enhanced ModelabstractCall-and-response is a musical technique that enriches the creativity of music, crafting coherent musical ideas that mirror the back-and-forth nature of human dialogue with distinct musical characteristics. Although this technique is integral to numerous musical compositions, it remains largely uncharted in automatic music composition. To enhance the creativity of machine-composed music, we first introduce the Call-Response Dataset (CRD) containing 19,155 annotated musical pairs and crafted comprehensive objective evaluation metrics for musical assessment. Then, we design a knowledge-enhanced learning-based method to bridge the gap between human and machine creativity. Specifically, we train the composition module using the call-response pairs, supplementing it with musical knowledge in terms of rhythm, melody, and harmony. Our experimental results underscore that our proposed model adeptly produces a wide variety of creative responses for various musical calls. Zhejing Hu, Yan Liu 0004, Gong Chen 0006, Xiao Ma 0023, Shenghua Zhong, Qianwen Luo |
AAAI | 2 |
| 2024 | Beyond Mimicking Under-Represented Emotions: Deep Data Augmentation with Emotional Subspace Constraints for EEG-Based Emotion RecognitionabstractIn recent years, using Electroencephalography (EEG) to recognize emotions has garnered considerable attention. Despite advancements, limited EEG data restricts its potential. Thus, Generative Adversarial Networks (GANs) are proposed to mimic the observed distributions and generate EEG data. However, for imbalanced datasets, GANs struggle to produce reliable augmentations for under-represented minority emotions by merely mimicking them. Thus, we introduce Emotional Subspace Constrained Generative Adversarial Networks (ESC-GAN) as an alternative to existing frameworks. We first propose the EEG editing paradigm, editing reference EEG signals from well-represented to under-represented emotional subspaces. Then, we introduce diversity-aware and boundary-aware losses to constrain the augmented subspace. Here, the diversity-aware loss encourages a diverse emotional subspace by enlarging the sample difference, while boundary-aware loss constrains the augmented subspace near the decision boundary where recognition models can be vulnerable. Experiments show ESC-GAN boosts emotion recognition performance on benchmark datasets, DEAP, AMIGOS, and SEED, while protecting against potential adversarial attacks. Finally, the proposed method opens new avenues for editing EEG signals under emotional subspace constraints, facilitating unbiased and secure EEG data augmentation. Zhi Zhang 0004, Shenghua Zhong, Yan Liu 0004 |
AAAI | 3 |
| 2024 | Sal-Guide Diffusion: Saliency Maps Guide Emotional Image Generation through AdapterabstractThe existing text-to-image generation methods based on stable diffusion yield better results in low-semantic prompt but often neglect the generation quality of high-semantic prompt such as emotional vocabulary, resulting in poor emotional image generation. In order to address this issue, we propose a novel approach called Sal-Guide Diffusion, which leverages saliency maps to guide emotional image generation with the goal of producing superior emotionally expressive images. In order to let the saliency maps guide the diffusion process, we introduce a lightweight adapter to extract emotional information from saliency maps and incorporate it into the diffusion process. Experimental results demonstrate that our proposed method generates higher-quality images across eight emotional dimensions, excelling in both generalization, emotional congruence, and subjective preference compared to stable diffusion or similar methods. Xiangru Lin, Shenghua Zhong, Yan Liu 0004, Gong Chen 0006 |
ICME | 3 |
| 2024 | MuChin: A Chinese Colloquial Description Benchmark for Evaluating Language Models in the Field of Music
Tao Zhang 0042, Jinyang Luo, Yan Liu 0004, Ming Xi |
IJCAI | 7 |
| 2024 | Make a song curative: A spatio-temporal therapeutic music transfer model for anxiety reduction
Zhejing Hu, Gong Chen 0006, Yan Liu 0004, Xiao Ma 0023, Nianhong Guan, Xiaoying Wang 0007 |
Expert Syst. Appl. | 3 |
| 2024 | TorchEEGEMO: A deep learning toolbox towards EEG-based emotion recognition
Zhi Zhang 0004, Shenghua Zhong, Yan Liu 0004 |
Expert Syst. Appl. | 3 |
| 2024 | Lightweight improved residual network for efficient inverse tone mapping
Liqi Xue, Yongbao Song, Yan Liu 0004, Lei Zhang 0093, Xiantong Zhen, Jun Xu 0019 |
Multim. Tools Appl. | 4 |
| 2024 | EGCN++: A New Fusion Strategy for Ensemble Learning in Skeleton-Based Rehabilitation Exercise AssessmentabstractSkeleton-based exercise assessment focuses on evaluating the correctness or quality of an exercise performed by a subject. Skeleton data provide two groups of features (i.e., position and orientation), which existing methods have not fully harnessed. We previously proposed an ensemble-based graph convolutional network (EGCN) that considers both position and orientation features to construct a model-based approach. Integrating these types of features achieved better performance than available methods. However, EGCN lacked a fusion strategy across the data, feature, decision, and model levels. In this paper, we present an advanced framework, EGCN++, for rehabilitation exercise assessment. Based on EGCN, a new fusion strategy called MLE-PO is proposed for EGCN++; this technique considers fusion at the data and model levels. We conduct extensive cross-validation experiments and investigate the consistency between machine and human evaluations on three datasets: UI-PRMD, KIMORE, and EHE. Results demonstrate that MLE-PO outperforms other EGCN ensemble strategies and representative baselines. Furthermore, the MLE-PO's model evaluation scores are more quantitatively consistent with clinical evaluations than other ensemble strategies. Bruce X. B. Yu, Yan Liu 0004, Keith C. C. Chan, Chang Wen Chen |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | The Beauty of Repetition: An Algorithmic Composition Model With Motif-Level Repetition Generator and Outline-to-Music Generator in Symbolic Music GenerationabstractMost musical compositions utilize repetition as a fundamental element to create captivating aesthetic experiences. However, the potential of repetition in machine-learning-based algorithmic composition has not been thoroughly investigated. This article aims to make an initial attempt at repetition modeling by generating motif-level repetitions and integrating them into music through a combination of example-based and domain knowledge–based learning techniques. The article presents a new Motif-to-music Generation Model (MGM) that combines a motif-level repetition generator (MRG) and an outline-to-music generator (O2MG). To train this model, a new music repetition dataset (MRD) has been created, which includes 584,329 samples from various categories of motif repetition and 3,545 outline-music sequences from pop piano music. The MRG uses a Transformer encoder to learn the representation of music notes from MRD, while the repetition-aware learner in MRG takes advantage of the unique characteristics of repetitions based on music theory. The O2MG applies a novel outline-to-music learning strategy to learn the relationships among motif-level repetitions in the music and generate music based on these repetitions. The experiments show that MGM can generate a variety of beautiful repetitions with any given motif, improving the music quality and structure of machine-composed music. Zhejing Hu, Xiao Ma 0023, Yan Liu 0004, Gong Chen 0006, Yongxu Liu 0003, Roger B. Dannenberg |
IEEE Trans. Multim. | 3 |
| 2023 | Knowledge-guided Network Pruning for EEG-based Emotion RecognitionabstractWith the development of deep learning in EEG-related tasks, the complexity of learning models has gradually increased. These complex models often result in long inference times, high energy consumption, and an increased risk of overfitting. Therefore, model compression has become an important consideration. Although some EEG models have used lightweight techniques, such as separable convolution, no existing work has directly attempted to compress EEG models to reduce their complexity. In this paper, we integrate neuroscience knowledge into EEG model pruning recovery, and innovatively propose two loss functions in the learning process, the knowledge-guided region-wise loss that enforces the classification evidence consistent with the importance of the prefrontal lobe, and the knowledge-guided sample-wise loss that constrains the learning process by distinguishing the importance of different samples. Wenjie Rao, Shenghua Zhong, Zhi Zhang 0004, Yan Liu 0004 |
BIBM | 4 |
| 2023 | A Dynamic Selective Parameter Sharing Mechanism Embedded with Multi-Level Reasoning AbstractionsabstractCooperative multi-agent reinforcement learning (Co-MARL) commonly employs different parameter sharing mechanisms, such as full and partial sharing. However, imprudent application of these mechanisms can potentially constrain policy diversity and limit cooperation flexibility. Recent methods that group agents into distinct sharing categories often exhibit poor performance due to challenges in precisely differentiating agents and neglecting the issue of promoting cooperation among these categories. To address these issues, we introduce a dynamic selective parameter sharing mechanism embedded with multi-level reasoning abstractions (DSPS-MA). Our approach uses self-comparison sequences to infer agents’ abstract concepts, defining the differences between agents and allowing them to dynamically select partners to share parameters based on these abstract concepts. We also design an intrinsic reward to offer comprehensive collaboration guidance for agents, and introduce a policy cosine similarity regularization term to ensure sufficient policy diversity. Empirical evaluations demonstrate that our approach yields higher returns and faster convergence than state-of-the-art methods. Yan Liu 0004, Ying He 0006, Zhong Ming 0001, F. Richard Yu |
ECAI | 1 |
| 2023 | GLA-GCN: Global-local Adaptive Graph Convolutional Network for 3D Human Pose Estimation from Monocular Videoabstract3D human pose estimation has been researched for decades with promising fruits. 3D human pose lifting is one of the promising research directions toward the task where both estimated pose and ground truth pose data are used for training. Existing pose lifting works mainly focus on improving the performance of estimated pose, but they usually underperform when testing on the ground truth pose data. We observe that the performance of the estimated pose can be easily improved by preparing good quality 2D pose, such as fine-tuning the 2D pose or using advanced 2D pose detectors. As such, we concentrate on improving the 3D human pose lifting via ground truth data for the future improvement of more quality estimated pose data. Towards this goal, a simple yet effective model called Global-local Adaptive Graph Convolutional Network (GLA-GCN) is proposed in this work. Our GLA-GCN globally models the spatiotemporal structure via a graph representation and backtraces local joint features for 3D human pose estimation via individually connected layers. To validate our model design, we conduct extensive experiments on three benchmark datasets: Human3.6M, HumanEva-I, and MPI-INF-3DHP. Experimental results show that our GLA-GCN1implemented with ground truth 2D poses significantly outperforms state-of-the-art methods (e.g., up to 3%, 17%, and 14% error reductions on Human3.6M, HumanEva-I, and MPI-INF-3DHP, respectively). Bruce X. B. Yu, Zhi Zhang 0004, Yongxu Liu 0003, Shenghua Zhong, Yan Liu 0004, Chang Wen Chen |
ICCV | 5 |
| 2023 | A Robust Deep Learning Enhanced Monocular SLAM System for Dynamic EnvironmentsabstractSimultaneous Localization and Mapping (SLAM) has developed as a fundamental method for intelligent robot perception over the past decades. Most of the existing feature-based SLAM systems relied on traditional hand-crafted visual features and a strong static world assumption, which makes these systems vulnerable in complex dynamic environments. In this paper, we propose a robust monocular SLAM system by combining geometry-based methods with two convolutional neural networks. Specifically, a lightweight deep local feature detection network is proposed as the system front-end, which can efficiently generate keypoints and binary descriptors robust against variations in illumination and viewpoint. Besides, we propose a motion segmentation and depth estimation network for simultaneously predicting pixel-wise motion object segmentation and depth map, so that our system can easily discard dynamic features and reconstruct 3D maps without dynamic objects. The comparison against state-of-the-art methods on publicly available datasets shows the effectiveness of our system in highly dynamic environments. Yaoqing Li, Shenghua Zhong, Shuai Li 0002, Yan Liu 0004 |
ICMR | 4 |
| 2023 | MMNet: A Model-Based Multimodal Network for Human Action Recognition in RGB-D VideosabstractHuman action recognition (HAR) in RGB-D videos has been widely investigated since the release of affordable depth sensors. Currently, unimodal approaches (e.g., skeleton-based and RGB video-based) have realized substantial improvements with increasingly larger datasets. However, multimodal methods specifically with model-level fusion have seldom been investigated. In this article, we propose a model-based multimodal network (MMNet) that fuses skeleton and RGB modalities via a model-based approach. The objective of our method is to improve ensemble recognition accuracy by effectively applying mutually complementary information from different data modalities. For the model-based fusion scheme, we use a spatiotemporal graph convolution network for the skeleton modality to learn attention weights that will be transferred to the network of the RGB modality. Extensive experiments are conducted on five benchmark datasets: NTU RGB+D 60, NTU RGB+D 120, PKU-MMD, Northwestern-UCLA Multiview, and Toyota Smarthome. Upon aggregating the results of multiple modalities, our method is found to outperform state-of-the-art approaches on six evaluation protocols of the five datasets; thus, the proposed MMNet can effectively capture mutually complementary features in different RGB-D video modalities and provide more discriminative features for HAR. We also tested our MMNet on an RGB video dataset Kinetics 400 that contains more outdoor actions, which shows consistent results with those of RGB-D video datasets. Bruce X. B. Yu, Yan Liu 0004, Shenghua Zhong, Keith C. C. Chan |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Noise-robust oversampling for imbalanced data classification
Yongxu Liu 0003, Yan Liu 0004, Bruce X. B. Yu, Shenghua Zhong, Zhejing Hu |
Pattern Recognit. | 2 |
| 2023 | Bi-RRNet: Bi-level recurrent refinement network for camouflaged object detection
Yan Liu 0004, Kaihua Zhang 0001, Yaqian Zhao, Qingshan Liu 0001 |
Pattern Recognit. | 1 |
| 2023 | GANSER: A Self-Supervised Data Augmentation Framework for EEG-Based Emotion RecognitionabstractElectroencephalography (EEG)-based affective computing has a scarcity problem. As a result, it is difficult to build effective, highly accurate and stable models using machine learning algorithms, especially deep learning models. Data augmentation has recently shown performance improvements in deep learning models with increased accuracy, stability and reduced overfitting. In this paper, we propose a novel data augmentation framework, named the generative adversarial network-based self-supervised data augmentation (GANSER). As the first to combine adversarial training with self-supervised learning for EEG-based emotion recognition, the proposed framework generates high-quality and high-diversity simulated EEG samples. In particular, we utilize adversarial training to learn an EEG generator and force the generated EEG signals to approximate the distribution of real samples, ensuring the quality of the augmented samples. A transformation operation is employed to mask parts of the EEG signals and force the generator to synthesize potential EEG signals based on the unmasked parts to produce a wide variety of samples. A masking possibility during transformation is introduced as prior knowledge to generalize the classifier for the augmented sample space. Finally, numerous experiments demonstrate that our proposed method can improve emotion recognition with an increase in performance and achieve state-of-the-art results. Zhi Zhang 0004, Yan Liu 0004, Shenghua Zhong |
IEEE Trans. Affect. Comput. | 2 |
| 2023 | Can Machines Generate Personalized Music? A Hybrid Favorite-Aware Method for User Preference Music TransferabstractUser preference music transfer (UPMT) is a new problem in music style transfer that can be applied to many scenarios but remains understudied. Transferring an arbitrary song to fit a user’s preferences increases musical diversity and improves user engagement, which can greatly benefit individuals’ mental health. Most music style transfer approaches rely on data-driven methods. In general, however, constructing a large training dataset is challenging because users can rarely provide enough of their favorite songs. To address this problem, this paper proposes a novel hybrid method called User Preference Transformer (UP-Transformer) which uses prior knowledge of only one piece of a user’s favorite music. Based on the distribution of music events in the provided music, we propose a new favorite-aware loss function to fine-tune the Transformer-based model. Two steps are proposed in the transfer phase to achieve UPMT based on the extracted music pattern in a user’s favorite music. Additionally, to alleviate the problem of evaluating melodic similarity in music style transfer, we propose a new concept called pattern similarity (PS) to measure the similarity between two pieces of music. Statistical tests indicate that the results of PS are consistent with the similarity score in a qualitative experiment. Furthermore, experimental results on subjects show that the transferred music achieves better performance in musicality, similarity, and user preferences. Zhejing Hu, Yan Liu 0004, Gong Chen 0006, Yongxu Liu 0003 |
IEEE Trans. Multim. | 2 |
| 2022 | EGCN: An Ensemble-based Learning Framework for Exploring Effective Skeleton-based Rehabilitation Exercise AssessmentabstractRecently, some skeleton-based physical therapy systems have been attempted to automatically evaluate the correctness or quality of an exercise performed by rehabilitation subjects. However, in terms of algorithms and evaluation criteria, the task remains not fully explored regarding making full use of different skeleton features. To advance the prior work, we propose a learning framework called Ensemble-based Graph Convolutional Network (EGCN) for skeleton-based rehabilitation exercise assessment. As far as we know, this is the first attempt that utilizes both two skeleton feature groups and investigates different ensemble strategies for the task. We also examine the properness of existing evaluation criteria and focus on evaluating the prediction ability of our proposed method. We then conduct extensive cross-validation experiments on two latest public datasets: UI-PRMD and KIMORE. Results indicate that the model-level ensemble scheme of our EGCN achieves better performance than existing methods. Code is available: https://github.com/bruceyo/EGCN. Bruce X. B. Yu, Yan Liu 0004, Gong Chen 0006, Keith C. C. Chan |
IJCAI | 2 |
| 2022 | The Beauty of Repetition in Machine Composition ScenariosabstractRepetition, a basic form of artistic creation, appears in most musical works and delivers enthralling aesthetic experiences. However, repetition remains underexplored in terms of automatic music composition. As an initial effort in repetition modelling, this paper focuses on generating motif-level repetitions via domain knowledge-based and example-based learning techniques. A novel repetition transformer (R-Transformer) that combines a Transformer encoder and a repetition-aware learner is trained on a new repetition dataset with 584,329 samples from different categories of motif repetition. The Transformer encoder learns the representation among music notes from the repetition dataset; the novel repetition-aware learner exploits repetitions' unique characteristics based on music theory. Experiments show that, with any given motif, R-Transformer can generate a large number of variable and beautiful repetitions. With ingenious fusion of these high-quality pieces, the musicality and appeal of machine-composed music have been greatly improved. Zhejing Hu, Xiao Ma 0023, Yan Liu 0004, Gong Chen 0006, Yongxu Liu 0003 |
ACM Multimedia | 3 |
| 2022 | Detecting abnormality with separated foreground and background: Mutual Generative Adversarial Networks for video abnormal event detection
Zhi Zhang 0004, Shenghua Zhong, Ahmed Fares, Yan Liu 0004 |
Comput. Vis. Image Underst. | 4 |
| 2021 | Multimodal Fusion via Teacher-Student Network for Indoor Action RecognitionabstractIndoor action recognition plays an important role in modern society, such as intelligent healthcare in large mobile cabin hospitals. With the wide usage of depth sensors like Kinect, multimodal information including skeleton and RGB modalities brings a promising way to improve the performance. However, existing methods are either focusing on a single data modality or failed to take the advantage of multiple data modalities. In this paper, we propose a Teacher-Student Multimodal Fusion (TSMF) model that fuses the skeleton and RGB modalities at the model level for indoor action recognition. In our TSMF, we utilize a teacher network to transfer the structural knowledge of the skeleton modality to a student network for the RGB modality. With extensive experiments on two benchmarking datasets: NTU RGB+D and PKU-MMD, results show that the proposed TSMF consistently performs better than state-of-the-art single modal and multimodal methods. It also indicates that our TSMF could not only improve the accuracy of the student network but also significantly improve the ensemble accuracy. Bruce X. B. Yu, Yan Liu 0004, Keith C. C. Chan |
AAAI | 2 |
| 2021 | Video Abnormal Event Detection Via Context Cueing Generative Adversarial NetworkabstractVideo abnormal event detection is a significant research subject in intelligent surveillance, referring to identifying events that do not conform to expected behaviors. It is a challenging task due to the diversity of video content and the various definition of anomaly depending on contexts. How to grasp the nuances of anomalies from diverse video content under associated contexts? In this paper, we propose a context cueing generative adversarial network (CC-GAN) to scop attention to specific regions of interest (ROIs). Different from existing work, the designed framework takes advantage of the context cues to preserve the associations between ROIs and contexts. By intrinsic learning of static spatial layouts and apparent velocities, context cues can help to understand ROIs. Optimized in a crop-inpainting fashion, CC-GAN learns to generate ROIs matching given spatio-temporal context and discriminate the distorted or incongruous ones, corresponding to abnormal scenes. Compared with the state-of-the-art method and other abnormal event detection approaches, the proposed framework demonstrates its effectiveness, reliability, and good generalization ability across various scenes and events. Zhi Zhang 0004, Shenghua Zhong, Yan Liu 0004 |
ICME | 3 |
| 2021 | Decomposing word embedding with the capsule network
Xin Liu 0054, Qingcai Chen, Yan Liu 0004, Joanna Siebert, Baotian Hu, Xiangping Wu 0001, Buzhou Tang |
Knowl. Based Syst. | 3 |
| 2021 | Steganographer detection via a similarity accumulation graph convolutional network
Zhi Zhang 0004, Mingjie Zheng 0002, Shenghua Zhong, Yan Liu 0004 |
Neural Networks | 4 |
| 2021 | Skeleton-based human action evaluation using graph convolutional network for monitoring Alzheimer's progression
Bruce X. B. Yu, Yan Liu 0004, Keith C. C. Chan, Qintai Yang, Xiaoying Wang 0007 |
Pattern Recognit. | 2 |
| 2020 | Steganographer Detection Via Enhancement-Aware Graph Convolutional NetworkabstractSteganographer detection aims to find guilty users who hide secret information in images or other multimedia data in the social network. In existing work, the distances between users are calculated based on the distributions of all images shared by the corresponding users, then users lying an abnormal distance from others are detected as guilty users. This flattened method is difficult to grasp the nuances of the guilty and innocent users. In this paper, we are the first to propose a graph-based deep learning framework for steganographer detection. The proposed Enhancement-aware Graph Convolutional Network (EGCN) represents each user as a weighted complete graph and learns to highlight the differences between guilty users and innocent users based on the structured graph. Compared with the state-of-the-art method and other representative graph-based models, the proposed framework demonstrates its effectiveness across image domains, and even under the context of large-scale social media scenario. Zhi Zhang 0004, Mingjie Zheng 0002, Shenghua Zhong, Yan Liu 0004 |
ICME | 4 |
| 2020 | Position-Aware Safe Boundary Interpolation OversamplingabstractThe class imbalance problem is characterized by the unequal distribution of different class samples, usually resulting in a learning bias toward the majority class. In the past decades, kinds of techniques have been proposed to alleviate this problem. Among those approaches, one promising method, interpolation-based oversampling, proposes to generate synthetic minority samples based on selected reference data, which can effectively solve the skewed distribution of data samples. However, there are several unsolved issues in interpolation-based oversampling. Existing methods often suffer from noisy synthetic samples due to improper data clustering and unsatisfactory reference selection. In this paper, we propose the position-aware safe boundary interpolation oversampling algorithm (PABIO) to address such issues. We firstly introduce a combined clustering algorithm for minority samples to overcome the shortage of clustering methods which are only distance-based or density-based. Then a position-aware interpolation-based oversampling algorithm is proposed for different minority clusters. Especially, we develop a novel method to leverage the majority class information to learn a safe boundary for generating synthetic points. The proposed PABIO is evaluated on multiple imbalanced datasets classified by two base classifiers: support vector machine (SVM) and C4.5 decision tree classifier. Experimental results show that our proposed PABIO outperforms other baselines among benchmark datasets. Yongxu Liu 0003, Yan Liu 0004 |
ICPR | 2 |
| 2020 | Make Your Favorite Music Curative: Music Style Transfer for Anxiety ReductionabstractAnxiety is the most common mental problem that affects nearly 300 million individuals worldwide. The situation is even worse recently. In clinical practice, music therapy has been used for more than forty years because of its effectiveness and few side effects in emotion regulation. This paper proposes a novel style transfer model to generate the therapeutic music according to user's preference. It is widely recognized that the favorite music greatly increases the engagement of the user, hence results in much better curative effects. But in general, users can provide only one or several favorite songs, which are insufficient for the customization of therapeutic music. To address this difficulty, a new domain adaption algorithm that transfers the learning result for music genre classification to the music personalization, is designed. Targeting the joint minimization of the loss functions, three convolutional neural networks are utilized to generate the therapeutic music with only one labelled data of favorite song. The experiment on the anxiety suffers shows that the customized therapeutic music has achieved better and stable performance in anxiety reduction. Zhejing Hu, Yan Liu 0004, Gong Chen 0006, Shenghua Zhong, Aiwei Zhang |
ACM Multimedia | 2 |
| 2020 | Fusing CAMs-weighted features and temporal information for robust loop closure detectionabstractAs a key component in simultaneous localization and mapping (SLAM) system, loop closure detection (LCD) eliminates the accumulated errors by recognizing previously visited places. In recent years, deep learning methods have been proved effective in LCD. However, most of the existing methods do not make good use of the useful information provided by monocular images, which tends to limit their performance in challenging dynamic scenarios with partial occlusion by moving objects. To this end, we propose a novel workflow, which is able to combine multiple information provided by images. We first introduce semantic information into LCD by developing a local-aware Class Activation Maps (CAMs) weighting method for extracting features, which can reduce the adverse effects of moving objects. Compared with previous methods based on semantic segmentation, our method has the advantage of not requiring additional models or other complex operations. In addition, we propose two effective temporal constraint strategies, which utilize the relationship of image sequences to improve the detection performance. Moreover, we propose to use the keypoint matching strategy as the final detector to further refuse false positives. Experiments on four publicly available datasets indicate that our approach can achieve higher accuracy and better robustness than the state-of-the-art methods. Yaoqing Li, Shenghua Zhong, Tongwei Ren, Yan Liu 0004 |
MMAsia | 4 |
| 2020 | Dynamic graph convolutional network for multi-video summarization
Jiaxin Wu 0001, Shenghua Zhong, Yan Liu 0004 |
Pattern Recognit. | 3 |
| 2020 | How to Evaluate Single-Round Dialogues Like Humans: An Information-Oriented MetricabstractDeveloping a dialogue response generation system is one of important topics in natural language processing, but many obstacles are yet to be overcome before autogenerated dialogues with a human-like quality can become possible. A good evaluation method will help narrow the gap between machines and humans in dialogue generation. Unfortunately, the existing automatic evaluation methods are biased and correlate very poorly with human judgments of response quality. Such methods are incapable of assessing whether a dialogue response generation system can produce high-quality, knowledge-related and informative dialogues. In response to this challenge, we design an information-oriented framework to simulate human subjective evaluation. Using this framework, we implement a learning-based metric to evaluate the quality of a dialogue. An experimental validation demonstrates our proposed metric's effectiveness in dialogue selection and model evaluation on a Twitter dataset (in English) and a Weibo dataset (in Chinese). In addition, the metric is more relevant than the existing methods of dialogue evaluation to human subjective judgment. Shenghua Zhong, Peiqi Liu, Zhong Ming 0001, Yan Liu 0004 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2020 | A Novel Convolutional Neural Network for Image Steganalysis With Shared NormalizationabstractImage steganalysis is to discriminate innocent images (cover images) and those suspected images (stego images) with hidden messages. The task is challenging since modifications to cover images due to message hiding are extremely small. To handle this difficulty, modern approaches proposed using convolutional neural network (CNN) models to detect steganography with paired learning, i.e., cover images and their stegos are both in training set. In this paper, we explore an important technique in CNN models, the batch normalization (BN), for the task of image steganalysis in the paired learning framework. Our theoretical analysis shows that a CNN model with multiple batch normalization layers is difficult to be generalized to new data in the test set when it is well trained with paired learning. To address this problem, we propose a novel normalization technique called shared normalization (SN) in this paper. Unlike the BN layer utilizing the mini-batch mean and standard deviation to normalize each input batch, SN shares consistent statistics for training samples. Based on the proposed SN layer, we further propose a novel neural network model for image steganalysis. Extensive experiments demonstrate that the proposed network with SN layers is stable and can detect the state-of-the-art steganography with better performances than previous methods. Songtao Wu, Shenghua Zhong, Yan Liu 0004 |
IEEE Trans. Multim. | 3 |
| 2020 | Steganographer Detection via Multi-Scale Embedding Probability EstimationabstractSteganographer detection aims to identify the guilty user who utilizes steganographic methods to hide secret information in the spread of multimedia data, especially image data, from a large amount of innocent users on social networks. A true embedding probability map illustrates the probability distribution of embedding secret information in the corresponding images by specific steganographic methods and settings, which has been successfully used as the guidance for content-adaptive steganographic and steganalytic methods. Unfortunately, in real-world situation, the detailed steganographic settings adopted by the guilty user cannot be known in advance. It thus becomes necessary to propose an automatic embedding probability estimation method. In this article, we propose a novel content-adaptive steganographer detection method via embedding probability estimation. The embedding probability estimation is first formulated as a learning-based saliency detection problem and the multi-scale estimated map is then integrated into the CNN to extract steganalytic features. Finally, the guilty user is detected via an efficient Gaussian vote method with the extracted steganalytic features. The experimental results prove that the proposed method is superior to the state-of-the-art methods in both spatial and frequency domains. Shenghua Zhong, Yuantian Wang, Tongwei Ren, Mingjie Zheng 0002, Yan Liu 0004, Gangshan Wu |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2019 | MvsGCN: A Novel Graph Convolutional Network for Multi-video SummarizationabstractMulti-video summarization, which tries to generate a single summary for a collection of video, is an important task in dealing with ever-growing video data. In this paper, we are the first to propose a graph convolutional network for multi-video summarization. The novel network measures the importance and relevance of each video shot in its own video as well as in the whole video collection. The important node sampling method is proposed to emphasize the effective features which are more possible to be selected as the final video summary. Two strategies are proposed to integrate into the network to solve the inherent class imbalance problem in the task of video summarization. The loss regularization for diversity is used to encourage a diverse summary to be generated. Extensive experiments are carried out, and in comparison with traditional and recent graph models and the state-of-the-art video summarization methods, our proposed model is effective in generating a representative summary for multiple videos with good diversity. It also achieves state-of-the-art performance on two standard video summarization datasets. Jiaxin Wu 0001, Shenghua Zhong, Yan Liu 0004 |
ACM Multimedia | 3 |
| 2019 | Content-adaptive selective steganographer detection via embedding probability estimation deep networks
Mingjie Zheng 0002, Jianmin Jiang, Songtao Wu, Shenghua Zhong, Yan Liu 0004 |
Neurocomputing | 5 |
| 2019 | Insights of object proposal evaluation
Yuantian Wang, Lei Huang 0004, Tongwei Ren, Shenghua Zhong, Han Gu, Yan Liu 0004 |
Multim. Tools Appl. | 6 |
| 2018 | Information-Oriented Evaluation Metric for Dialogue Response Generation SystemsabstractDialogue response generation system is one of the hot topics in natural language processing, but it is still a long way to go before it can generate human-like dialogues. A good evaluation method will help narrow the gap between the machine and human in dialogue generation. Unfortunately, current evaluation methods cannot measure whether the dialogue response generation system is able to produce high-quality, knowledge-related, and informative dialogues. Aiming to identify and measure the existence of information in dialogues, we propose a novel automatic evaluation metric. By learning from the knowledge representation method in knowledge base, we define the heuristic rules to extract the information triples from dialogue pairs. And we design an information matching method to measure the probability of the existence of information in a dialogue. In experiments, our proposed metric demonstrates its effectiveness in dialogue selection and model evaluation on the Reddit dataset (English) and the Weibo dataset (Chinese). Peiqi Liu, Shenghua Zhong, Zhong Ming 0001, Yan Liu 0004 |
ICTAI | 4 |
| 2018 | Musicality-Novelty Generative Adversarial Nets for Algorithmic CompositionabstractAlgorithmic composition, which enables computer to generate music like human composers, has lasting charm because it intends to approximate artistic creation, most mysterious part of human intelligence. To deliver both melodious and refreshing music, this paper proposes the Musicality-Novelty Generative Adversarial Nets for algorithmic composition. With the same generator, two adversarial nets alternately optimize the musicality and novelty of the machine-composed music. A new model called novelty game is presented to maximize the minimal distance between the machine-composed music sample and any human-composed music sample in the novelty space, where all well-known human composed music products are far from each other. We implement the proposed framework using three supervised CNNs with one for generator, one for musicality critic and one for novelty critic on the time-pitch feature space. Specifically, the novelty critic is implemented by Siamese neural networks with temporal alignment using dynamic time warping. We provide empirical validations by generating the music samples under various scenarios. Gong Chen 0006, Yan Liu 0004, Shenghua Zhong |
ACM Multimedia | 2 |
| 2018 | Data Augmentation for EEG-Based Emotion Recognition with Deep Convolutional Neural Networks
Shenghua Zhong, Jianfeng Peng, Jianmin Jiang, Yan Liu 0004 |
MMM (2) | 5 |
| 2018 | On Unifying Multi-view Self-Representations for Clustering by Tensor Multi-rank Minimization
Yuan Xie 0006, Dacheng Tao, Wensheng Zhang 0002, Yan Liu 0004, Lei Zhang 0006, Yanyun Qu |
Int. J. Comput. Vis. | 4 |
| 2018 | Adaptive saliency cuts
Yuantian Wang, Tongwei Ren, Shenghua Zhong, Yan Liu 0004, Gangshan Wu |
Multim. Tools Appl. | 4 |
| 2018 | Deep residual learning for image steganalysis
Songtao Wu, Shenghua Zhong, Yan Liu 0004 |
Multim. Tools Appl. | 3 |
| 2017 | Residual convolution network based steganalysis with adaptive content suppressionabstractImage steganalysis is to discriminate innocent images and those suspected images with hidden messages. In this paper, we propose a unified Convolutional Neural Network (CNN) model for this task. In order to reliably detect modern steganographic algorithms, we design the proposed model from two aspects. For the first, different from existing CNN based steganalytic algorithms that use a predefined highpass kernel to suppress image content, we integrate the highpass filtering operation into the proposed network by building a content suppression subnetwork. For the second, we propose a novel sub-network to actively preserve the weak stego signal generated by secret messages based on residual learning, making the successive network capture the difference between cover images and stego images. Extensive experiments demonstrate that the proposed model can detect states-of-the-art steganography with much lower detection error rates than previous methods. Songtao Wu, Shenghua Zhong, Yan Liu 0004 |
ICME | 3 |
| 2017 | Implicit Visual Learning: Image Recognition via Dissipative Learning ModelabstractAccording to consciousness involvement, human’s learning can be roughly classified into explicit learning and implicit learning. Contrasting strongly to explicit learning with clear targets and rules, such as our school study of mathematics, learning is implicit when we acquire new information without intending to do so. Research from psychology indicates that implicit learning is ubiquitous in our daily life. Moreover, implicit learning plays an important role in human visual perception. But in the past 60 years, most of the well-known machine-learning models aimed to simulate explicit learning while the work of modeling implicit learning was relatively limited, especially for computer vision applications. This article proposes a novel unsupervised computational model for implicit visual learning by exploring dissipative system, which provides a unifying macroscopic theory to connect biology with physics. We test the proposed Dissipative Implicit Learning Model (DILM) on various datasets. The experiments show that DILM not only provides a good match to human behavior but also improves the explicit machine-learning performance obviously on image classification tasks. Yan Liu 0004, Yang Liu 0007, Shenghua Zhong, Songtao Wu |
ACM Trans. Intell. Syst. Technol. | 1 |
| 2016 | Programming Large-Scale Multi-Robot System with Timing ConstraintsabstractRecently years, research in multi-robot systems has attracted increasingly attentions. One important research topic is to design programming models that can facilitate the developers to programme large-scale multi-robot systems. However, existing works fail to manage the robots to perform tasks with real-time requirements. To address this issue, we propose a new programming model called RMR (Real-time Multi-Robot). RMR is a logic programming model with real-time support. On the basis of the logic programming paradigm, RMR allows the code for multi-robot system to be written from a global perspective, rather than managing a large collection of independent robots. Moreover, RMR allows developers to set timing constraints on the behaviors of an ensemble of robots, which is not implemented by state of the art. After designing RMR, we further develop a compiler and a runtime system for distributed execution of RMR programs. To evaluate the performance of RMR, we deploy it in a simulator and a test-bed, and then demonstrate RMR based on several applications. Our results indicate that RMR greatly facilitates implementing correct collaborative multi-robot applications. Shan Jiang 0005, Jiannong Cao 0001, Yan Liu 0004, Jinlin Chen, Xuefeng Liu 0001 |
ICCCN | 3 |
| 2016 | Steganalysis via Deep Residual NetworkabstractRecent studies have demonstrated that a well designed deep convolutional neural network (CNN) model achieves competitive performances on detecting the presence of secret message in digital images, compared with the classical rich model based steganalysis. In this paper, we propose to investigate a category of very deep CNN model-the deep residual network (DRN), for steganalysis. DRN is suitable for steganalysis from two aspects. For the first, the DRN model usually contains a large number of network layers, which proves to be effective to capture the complex statistics of digital images. For the second, DRN's residual learning (ResL) method actively strengthens the signal coming from secret messages, which is extremely beneficial for the discrimination between cover images and stego images. Comprehensive experiments on standard dataset show that the DRN model achieves very low detection error rates for the state of arts steganographic algorithms. It also outperforms the classical rich model method and several recently proposed CNN based methods. Songtao Wu, Shenghua Zhong, Yan Liu 0004 |
ICPADS | 3 |
| 2016 | Visual Orientation Inhomogeneity Based Convolutional Neural NetworksabstractThe details of oriented visual stimuli are better resolved when they are horizontal or vertical rather than oblique. This "oblique effect" has been researched and confirmed in numerous research studies, including behavioral studies and neurophysiological and neuroimaging findings. Although the "oblique effect" has influence in many fields, little research integrated it into computational models. In this paper, we try to explore this inhomogeneity of visual orientation based on Convolutional neural networks (CNNs) in image recognition. We validate that visual orientation inhomogeneity CNNs can achieve comparable performance with higher computational efficiency on various datasets. We can also get the conclusion that, compared with the cardinal information, oblique information is indeed less useful in natural color image recognition. Through the exploration of the proposed model on image recognition, we gain more understanding of the inhomogeneity of visual orientation. It also illuminates a wide range of opportunities for integrating the inhomogeneity of visual orientation with other computational models. Shenghua Zhong, Jiaxin Wu 0001, Yingying Zhu 0001, Peiqi Liu, Jianmin Jiang, Yan Liu 0004 |
ICTAI | 6 |
| 2016 | Learning Music Emotion Primitives via Supervised Dynamic ClusteringabstractThis paper explores a fundamental problem in music emotion analysis, i.e., how to segment the music sequence into a set of basic emotive units, which are named as emotion primitives. Current works on music emotion analysis are mainly based on the fixed-length music segments, which often leads to the difficulty of accurate emotion recognition. Short music segment, such as an individual music frame, may fail to evoke emotion response. Long music segment, such as an entire song, may convey various emotions over time. Moreover, the minimum length of music segment varies depending on the types of the emotions. To address these problems, we propose a novel method dubbed supervised dynamic clustering (SDC) to automatically decompose the music sequence into meaningful segments with various lengths. First, the music sequence is represented by a set of music frames. Then, the music frames are clustered according to the valence-arousal values in the emotion space. The clustering results are used to initialize the music segmentation. After that, a dynamic programming scheme is employed to jointly optimize the subsequent segmentation and grouping in the music feature space. Experimental results on standard dataset show both the effectiveness and the rationality of the proposed method. Yang Liu 0007, Yan Liu 0004, Gong Chen 0006 |
ACM Multimedia | 2 |
| 2016 | Perception-oriented video saliency detection via spatio-temporal attention analysis
Shenghua Zhong, Yan Liu 0004, Vincent T. Y. Ng, Yang Liu 0007 |
Neurocomputing | 2 |
| 2016 | How important is location information in saliency detection of natural images
Tongwei Ren, Yan Liu 0004, Ran Ju, Gangshan Wu |
Multim. Tools Appl. | 2 |
| 2016 | Weighted Schatten p-Norm Minimization for Image Denoising and Background SubtractionabstractLow rank matrix approximation (LRMA), which aims to recover the underlying low rank matrix from its degraded observation, has a wide range of applications in computer vision. The latest LRMA methods resort to using the nuclear norm minimization (NNM) as a convex relaxation of the nonconvex rank minimization. However, NNM tends to over-shrink the rank components and treats the different rank components equally, limiting its flexibility in practical applications. We propose a more flexible model, namely, the weighted Schatten p-norm minimization (WSNM), to generalize the NNM to the Schatten p-norm minimization with weights assigned to different singular values. The proposed WSNM not only gives better approximation to the original low-rank assumption, but also considers the importance of different rank components. We analyze the solution of WSNM and prove that, under certain weights permutation, WSNM can be equivalently transformed into independent non-convex lp-norm subproblems, whose global optimum can be efficiently solved by generalized iterated shrinkage algorithm. We apply WSNM to typical low-level vision problems, e.g., image denoising and background subtraction. Extensive experimental results show, both qualitatively and quantitatively, that the proposed WSNM can more effectively remove noise, and model the complex and dynamic scenes compared with state-of-the-art methods. Yuan Xie 0006, Shuhang Gu, Yan Liu 0004, Wangmeng Zuo, Wensheng Zhang 0002, Lei Zhang 0006 |
IEEE Trans. Image Process. | 3 |
| 2016 | Field Effect Deep Networks for Image Recognition with Incomplete DataabstractImage recognition with incomplete data is a well-known hard problem in computer vision and machine learning. This article proposes a novel deep learning technique called Field Effect Bilinear Deep Networks (FEBDN) for this problem. To address the difficulties of recognizing incomplete data, we design a novel second-order deep architecture with the Field Effect Restricted Boltzmann Machine, which models the reliability of the delivered information according to the availability of the features. Based on this new architecture, we propose a new three-stage learning procedure with field effect bilinear initialization, field effect abstraction and estimation, and global fine-tuning with missing features adjustment. By integrating the reliability of features into the new learning procedure, the proposed FEBDN can jointly determine the classification boundary and estimate the missing features. FEBDN has demonstrated impressive performance on recognition and estimation tasks in various standard datasets. Shenghua Zhong, Yan Liu 0004, Kien A. Hua |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2015 | Face recognition from a single registered image for conference socializing
Yan Liu 0004, Yang Liu 0007, Shenghua Zhong, Kien A. Hua |
Expert Syst. Appl. | 2 |
| 2015 | Visual orientation inhomogeneity based scale-invariant feature transform
Shenghua Zhong, Yan Liu 0004, Qingcai Chen |
Expert Syst. Appl. | 2 |
| 2015 | Soft-assigned bag of features for object tracking
Tongwei Ren, Zhongyan Qiu, Yan Liu 0004, Tong Yu 0001, Jia Bei |
Multim. Syst. | 3 |
| 2015 | Guest editorial: selected papers from ICIMCS 2012
Zhengjun Zha, Yan Liu 0004, Shin'ichi Satoh 0001, Xinguo Yu, Rainer Lienhart |
Multim. Syst. | 2 |
| 2015 | What Strikes the Strings of Your Heart? - Feature Mining for Music Emotion AnalysisabstractMusic can convey and evoke powerful emotions. This amazing ability has not only fascinated the general public but also attracted the researchers from different fields to discover the relationship between music and emotion. Psychologists have indicated that some specific characters of rhythm, harmony, and melody can evoke certain kinds of emotions. Those hypotheses are based on real life experience and proved by psychological paradigms on human beings. Aiming at the same target, this paper intends to design a systematic and quantitative framework, and answer three widely interested questions: 1) what are the intrinsic features embedded in music signal that essentially evoke human emotions; 2) to what extent these features influence human emotions; and 3) whether the findings from computational models are consistent with the existing research results from psychology. We formulate these tasks as a multi-label dimensionality reduction problem and propose an algorithm called multi-emotion similarity preserving embedding (ME-SPE). To adapt to the second-order music signals, we extend ME-SPE to its bilinear version. The proposed techniques show good performance in two standard music emotion datasets. Moreover, they demonstrate some interesting results for further research in this interdisciplinary topic. Yang Liu 0007, Yan Liu 0004, Kien A. Hua |
IEEE Trans. Affect. Comput. | 2 |
| 2014 | What Strikes the Strings of Your Heart?: Multi-Label Dimensionality Reduction for Music Emotion AnalysisabstractMusic can convey and evoke powerful emotions. This amazing ability has fascinated the general public and also attracted the researchers from different fields to discover the relationship between music and emotion. Psychologists have indicated that some specific characters of rhythm, harmony, melody, and also their combinations can evoke certain kinds of emotions. Their hypotheses are based on real life experience and proved by psychological paradigms on human beings. Aiming at the same target, this paper intends to design a systematic and quantitative framework, and answer three widely interested questions: 1) what are the intrinsic features embedded in music signal that essentially evoke human emotions; 2) to what extent these features influence human emotions; and 3) whether the findings from computational models are consistent with the existing research results from psychological experiments. We formulate the problem as a multi-label dimensionality reduction problem and provide the optimal solution. The proposed multi-emotion similarity preserving embedding technique not only shows better performance in two standard music emotion datasets but also demonstrates some interesting observations for further research in this interdisciplinary topic. Yang Liu 0007, Yan Liu 0004, Kien A. Hua |
ACM Multimedia | 2 |
| 2014 | Region level annotation by fuzzy based contextual cueing label propagation
Shenghua Zhong, Yan Liu 0004, Yang Liu 0007, Korris Fu-Lai Chung |
Multim. Tools Appl. | 2 |
| 2014 | Natural image denoising using evolved local adaptive filters
Ruomei Yan, Ling Shao 0001, Li Liu 0004, Yan Liu 0004 |
Signal Process. | 4 |
| 2014 | From Heuristic Optimization to Dictionary Learning: A Review and Comprehensive Comparison of Image Denoising AlgorithmsabstractImage denoising is a well explored topic in the field of image processing. In the past several decades, the progress made in image denoising has benefited from the improved modeling of natural images. In this paper, we introduce a new taxonomy based on image representations for a better understanding of state-of-the-art image denoising techniques. Within each category, several representative algorithms are selected for evaluation and comparison. The experimental results are discussed and analyzed to determine the overall advantages and disadvantages of each category. In general, the nonlocal methods within each category produce better denoising results than local ones. In addition, methods based on overcomplete representations using learned dictionaries perform better than others. The comprehensive study in this paper would serve as a good reference and stimulate new research ideas in image denoising. Ling Shao 0001, Ruomei Yan, Xuelong Li 0001, Yan Liu 0004 |
IEEE Trans. Cybern. | 4 |
| 2014 | A Unified Framework of Latent Feature Learning in Social MediaabstractThe current trend in social media analysis and application is to use the pre-defined features and devoted to the later model development modules to meet the end tasks. Representation learning has been a fundamental problem in machine learning, and widely recognized as critical to the performance of end tasks. In this paper, we provide evidence that specially learned features will addresses the diverse, heterogeneous, and collective characteristics of social media data. Therefore, we propose to transfer the focus from the model development to latent feature learning, and present a unified framework of latent feature learning on social media. To address the noisy, diverse, heterogeneous, and interconnected characteristics of social media data, the popular deep learning is employed due to its excellent abstract abilities. In particular, we instantiate the proposed framework by (1) designing a novel relational generative deep learning model to solve the social media link analysis task, and (2) developing a multimodal deep learning to lambda rank model towards the social image retrieval task. We show that the derived latent features lead to improvement in both of the social media tasks. Zhaoquan Yuan, Jitao Sang 0001, Changsheng Xu, Yan Liu 0004 |
IEEE Trans. Multim. | 4 |
| 2014 | Hybrid Manifold EmbeddingabstractIn this brief, we present a novel supervised manifold learning framework dubbed hybrid manifold embedding (HyME). Unlike most of the existing supervised manifold learning algorithms that give linear explicit mapping functions, the HyME aims to provide a more general nonlinear explicit mapping function by performing a two-layer learning procedure. In the first layer, a new clustering strategy called geodesic clustering is proposed to divide the original data set into several subsets with minimum nonlinearity. In the second layer, a supervised dimensionality reduction scheme called locally conjugate discriminant projection is performed on each subset for maximizing the discriminant information and minimizing the dimension redundancy simultaneously in the reduced low-dimensional space. By integrating these two layers in a unified mapping function, a supervised manifold embedding framework is established to describe both global and local manifold structure as well as to preserve the discriminative ability in the learned subspace. Experiments on various data sets validate the effectiveness of the proposed method. Yang Liu 0007, Yan Liu 0004, Keith C. C. Chan, Kien A. Hua |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2013 | Video Saliency Detection via Dynamic Consistent Spatio-Temporal Attention ModellingabstractHuman vision system actively seeks salient regions and movements in video sequences to reduce the search effort. Modeling computational visual saliency map provides im-portant information for semantic understanding in many real world applications. In this paper, we propose a novel video saliency detection model for detecting the attended regions that correspond to both interesting objects and dominant motions in video sequences. In spatial saliency map, we in-herit the classical bottom-up spatial saliency map. In tem-poral saliency map, a novel optical flow model is proposed based on the dynamic consistency of motion. The spatial and the temporal saliency maps are constructed and further fused together to create a novel attention model. The pro-posed attention model is evaluated on three video datasets. Empirical validations demonstrate the salient regions de-tected by our dynamic consistent saliency map highlight the interesting objects effectively and efficiency. More im-portantly, the automatically video attended regions detected by proposed attention model are consistent with the ground truth saliency maps of eye movement data. Shenghua Zhong, Yan Liu 0004, Feifei Ren, Jinghuan Zhang, Tongwei Ren |
AAAI | 2 |
| 2013 | Multi-operator Image Retargeting Based on Automatic Quality AssessmentabstractImage retargeting aims to avoid visual distortion while retaining important image content in resizing. However, no single image retargeting method can handle all cases. In this paper, we propose a novel multi-operator image retargeting approach, which utilizes an efficient and human perception based automatic quality assessment in operator selection. First, we calculate the importance map and distortion map for quality assessment. Then, we construct the resizing space and assess the performance of each operator in iterative width and/or height reduction. Finally, we select the optimal operator sequence by dynamic programming and generate the target image. Experiments demonstrate the effectiveness of the proposed approach. Zhongyan Qiu, Tongwei Ren, Yan Liu 0004, Jia Bei, Yang Yang 0222 |
ICIG | 3 |
| 2013 | Latent feature learning in social media networkabstractThe current trend in social media analysis and application is to use the pre-defined features and devoted to the later model development modules to meet the end tasks. In this work, we claim that representation is critical to the end tasks and contributes much to the model development module. We provide evidence that specially learned feature well addresses the diverse, heterogeneous and collective characteristics of social media data. Therefore, we propose to transfer the focus from the model development to latent feature learning, and present a general feature learning framework based on the popular deep architecture. In particular, following the proposed framework, we design a novel relational generative deep learning model to test the idea on link analysis tasks in the social media networks. We show that the derived latent features well embed both the media content and their observed links, leading to improvement in social media tasks of user recommendation and social image annotation. Zhaoquan Yuan, Jitao Sang 0001, Yan Liu 0004, Changsheng Xu |
ACM Multimedia | 3 |
| 2013 | Combining appearance and structural features for human action recognition
Ling Shao 0001, Xiantong Zhen, Yan Liu 0004 |
Neurocomputing | 4 |
| 2013 | Joint discriminative dimensionality reduction and dictionary learning for face recognition
Zhizhao Feng, Meng Yang 0001, Lei Zhang 0006, Yan Liu 0004, David Zhang 0001 |
Pattern Recognit. | 4 |
| 2013 | Nonlocal Hierarchical Dictionary Learning Using Wavelets for Image DenoisingabstractExploiting the sparsity within representation models for images is critical for image denoising. The best currently available denoising methods take advantage of the sparsity from image self-similarity, pre-learned, and fixed representations. Most of these methods, however, still have difficulties in tackling high noise levels or noise models other than Gaussian. In this paper, the multiresolution structure and sparsity of wavelets are employed by nonlocal dictionary learning in each decomposition level of the wavelets. Experimental results show that our proposed method outperforms two state-of-the-art image denoising algorithms on higher noise levels. Furthermore, our approach is more adaptive to the less extensively researched uniform noise. Ruomei Yan, Ling Shao 0001, Yan Liu 0004 |
IEEE Trans. Image Process. | 3 |
| 2013 | Water Reflection Recognition Based on Motion Blur Invariant Moments in Curvelet SpaceabstractWater reflection, a typical imperfect reflection symmetry problem, plays an important role in image content analysis. Existing techniques of symmetry recognition, however, cannot recognize water reflection images correctly because of the complex and various distortions caused by the water wave. Hence, we propose a novel water reflection recognition technique to solve the problem. First, we construct a novel feature space composed of motion blur invariant moments in low-frequency curvelet space and of curvelet coefficients in high-frequency curvelet space. Second, we propose an efficient algorithm including two sub-algorithms: low-frequency reflection cost minimization and high-frequency curvelet coefficients discrimination to classify water reflection images and to determine the reflection axis. Through experimenting on authentic images in a series of tasks, the proposed techniques prove effective and reliable in classifying water reflection images and detecting the reflection axis, as well as in retrieving images with water reflection. Shenghua Zhong, Yan Liu 0004, Yang Liu 0007 |
IEEE Trans. Image Process. | 2 |
| 2012 | Query-Oriented Multi-Document Summarization via Unsupervised Deep LearningabstractExtractive style query-oriented multi-document summarization generates the summary by extracting a proper set of sentences from multiple documents based on the pre-given query. This paper proposes a novel multi-document summarization framework via deep learning model. This uniform framework consists of three parts: concepts extraction, summary generation, and reconstruction validation, which work together to achieve the largest coverage of the documents content. A new query-oriented extraction technique is proposed to concentrate distributed information to hidden units layer by layer. Then, the whole deep architecture is fine-tuned by minimizing the information loss of reconstruction validation. According to the concentrated information, dynamic programming is used to seek most informative set of sentences as the summary. Experiments on three benchmark datasets demonstrate the effectiveness of the proposed framework and algorithms. Yan Liu 0004, Shenghua Zhong, Wenjie Li 0002 |
AAAI | 1 |
| 2012 | Attention Modeling for Face Recognition via Deep Learning
Shenghua Zhong, Yan Liu 0004, Korris Fu-Lai Chung |
CogSci | 2 |
| 2012 | Ordinal preserving projection: a novel dimensionality reduction method for image rankingabstractLearning to rank has been demonstrated as a powerful tool for image ranking, but the issue of the "curse of dimensionality" is a key challenge of learning a ranking model from a large image database. This paper proposes a novel dimensionality reduction algorithm named ordinal preserving projection (OPP) for learning to rank. We first define two matrices, which work in the row direction and column direction respectively. The two matrices aim at leveraging the global structure of the data set and ordinal information of the observations. By maximizing the corresponding objective functions, we can obtain two optimal projection matrices mapping original data points into low-dimensional subspace, in which both global structure and ordinal information can be preserved. The experiments are conducted on the public available MSRA-MM image data set and "Web Queries" image data set, and the experimental results demonstrate the effectiveness of the proposed method. Jing Liu 0001, Yan Liu 0004, Changsheng Xu, Qingshan Liu 0001, Hanqing Lu |
ICMR | 3 |
| 2012 | Semiconducting bilinear deep learning for incomplete image recognitionabstractImage recognition with incomplete data is a well-known hard problem in multimedia content analysis. This paper proposes a novel deep learning technique called semiconducting bilinear deep belief networks (SBDBN) by referencing human's visual cortex and intelligent perception. Inheriting from deep models, SBDBN simulates the laminar structure of human's cerebral cortex and the neural loop in human's visual areas. To address the special difficulties of image recognition with incomplete data, we design a novel second-order deep architecture with semiconducting restricted boltzmann machines. Moreover, two peaks activation of human's perception is implemented by three learning stages of semiconducting bilinear discriminant initialization, greedy layer-wise reconstruction, and global fine-tuning. Owing to exploiting the embedding information according to the reliable features rather than any completion of missing features, the proposed SBDBN has demonstrated outstanding recognition ability on two standard datasets and one constructed dataset, comparing with both incomplete image recognition techniques and existing deep learning models. Shenghua Zhong, Yan Liu 0004, Korris Fu-Lai Chung, Gangshan Wu |
ICMR | 2 |
| 2012 | S-SIFT: A Shorter SIFT without Least Discriminability Visual OrientationabstractDetection and description of local features are a classical problem in image processing and multimedia content analysis. Based on the in homogeneity of visual orientation in human visual system, we propose a novel algorithm S-SIFT to detect and describe local image features. In three stages of S-SIFT, the information from the least discriminability orientation is omitting. Compared with the standard SIFT algorithm, S-SIFT has lower dimension and provides a faster key point matching. Experiments on the standard dataset demonstrate that our algorithm yields comparable or even better results for feature detection and matching tasks. Shenghua Zhong, Yan Liu 0004, Gangshan Wu |
Web Intelligence | 2 |
| 2012 | Tensor distance based multilinear globality preserving embedding: A unified tensor based dimensionality reduction framework for image and video classification
Yang Liu 0007, Yan Liu 0004, Shenghua Zhong, Keith C. C. Chan |
Expert Syst. Appl. | 2 |
| 2012 | Relevance feedback for real-world human action retrieval
Ling Shao 0001, Jianguo Zhang 0001, Yan Liu 0004 |
Pattern Recognit. Lett. | 4 |
| 2012 | Human action segmentation and recognition via motion and shape analysis
Ling Shao 0001, Ling Ji, Yan Liu 0004, Jianguo Zhang 0001 |
Pattern Recognit. Lett. | 3 |
| 2011 | Ordinal Regression via Manifold LearningabstractOrdinal regression is an important research topic in machine learning. It aims to automatically determine the implied rating of a data item on a fixed, discrete rating scale. In this paper, we present a novel ordinal regression approach via manifold learning, which is capable of uncovering the embedded nonlinear structure of the data set according to the observations in the highdimensional feature space. By optimizing the order information of the observations and preserving the intrinsic geometry of the data set simultaneously, the proposed algorithm provides the faithful ordinal regression to the new coming data points. To offer more general solution to the data with natural tensor structure, we further introduce the multilinear extension of the proposed algorithm, which can support the ordinal regression of high order data like images. Experiments on various data sets validate the effectiveness of the proposed algorithm as well as its extension. Yang Liu 0007, Yan Liu 0004, Keith C. C. Chan |
AAAI | 2 |
| 2011 | Water reflection recognition via minimizing reflection cost based on motion blur invariant momentsabstractWater reflection, a kind of typical imperfect reflection symmetry problem, plays an important role in image content analysis. However, existing techniques of symmetry recognition cannot recognize water reflection images correctly because of the complex and various distortions caused by water wave. To address this difficulty, we construct a novel feature space which is composed of motion blur invariant moments. Moreover, we propose an efficient detection algorithm to determine the reflection axis in images with water reflection. By experimenting on real image dataset with different tasks, the proposed techniques demonstrate impressive results in the water reflection image classification, the reflection axis detection, and the retrieval of the images with water reflection. Shenghua Zhong, Yan Liu 0004, Ling Shao 0001, Korris Fu-Lai Chung |
ICMR | 2 |
| 2011 | Semi-supervised manifold ordinal regression for image rankingabstractIn this paper, we present a novel algorithm called manifold ordinal regression (MOR) for image ranking. By modeling the manifold information in the objective function, MOR is capable of uncovering the intrinsically nonlinear structure held by the image data sets. By optimizing the ranking information of the training data sets, the proposed algorithm provides faithful rating to the new coming images. To offer more general solution for the real-word tasks, we further provide the semi-supervised manifold ordinal regression (SS-MOR). Experiments on various data sets validate the effectiveness of the proposed algorithms. Yang Liu 0007, Yan Liu 0004, Shenghua Zhong, Keith C. C. Chan |
ACM Multimedia | 2 |
| 2011 | Bilinear deep learning for image classificationabstractImage classification is a well-known classical problem in multimedia content analysis. This paper proposes a novel deep learning model called bilinear deep belief network (BDBN) for image classification. Unlike previous image classification models, BDBN aims to provide human-like judgment by referencing the architecture of the human visual system and the procedure of intelligent perception. Therefore, the multi-layer structure of the cortex and the propagation of information in the visual areas of the brain are realized faithfully. Unlike most existing deep models, BDBN utilizes a bilinear discriminant strategy to simulate the "initial guess" in human object recognition, and at the same time to avoid falling into a bad local optimum. To preserve the natural tensor structure of the image data, a novel deep architecture with greedy layer-wise reconstruction and global fine-tuning is proposed. To adapt real-world image classification tasks, we develop BDBN under a semi-supervised learning framework, which makes the deep model work well when labeled images are insufficient. Comparative experiments on three standard datasets show that the proposed algorithm outperforms both representative classification models and existing deep learning techniques. More interestingly, our demonstrations show that the proposed BDBN works consistently with the visual perception of humans. Shenghua Zhong, Yan Liu 0004, Yang Liu 0007 |
ACM Multimedia | 2 |
| 2011 | Bilinear deep learning for image classificationabstractNo abstract available. Shenghua Zhong, Yan Liu 0004, Yang Liu 0007 |
ACM Multimedia | 2 |
| 2011 | A Generalized Coding Artifacts and Noise Removal Algorithm for Digitally Compressed Video Signals
Ling Shao 0001, Hui Zhang 0062, Yan Liu 0004 |
MMM (1) | 3 |
| 2011 | Tensor-based locally maximum margin classifier for image and video classification
Yang Liu 0007, Yan Liu 0004, Keith C. C. Chan |
Comput. Vis. Image Underst. | 2 |
| 2011 | Transform based spatio-temporal descriptors for human action recognition
Ling Shao 0001, Ruoyun Gao, Yan Liu 0004, Hui Zhang 0062 |
Neurocomputing | 3 |
| 2011 | Discriminative deep belief networks for visual data classification
Yan Liu 0004, Shusen Zhou, Qingcai Chen |
Pattern Recognit. | 1 |
| 2010 | Multilinear Maximum Distance Embedding Via L1-Norm OptimizationabstractDimensionality reduction plays an important role in many machine learning and pattern recognition tasks. In this paper, we present a novel dimensionality reduction algorithm called multilinear maximum distance embedding (M2DE), which includes three key components. To preserve the local geometry and discriminant information in the embedded space, M2DE utilizes a new objective function, which aims to maximize the distances between some particular pairs of data points, such as the distances between nearby points and the distances between data points from different classes. To make the mapping of new data points straightforward, and more importantly, to keep the natural tensor structure of high-order data, M2DE integrates multilinear techniques to learn the transformation matrices sequentially. To provide reasonable and stable embedding results, M2DE employs the L1-norm, which is more robust to outliers, to measure the dissimilarity between data points. Experiments on various datasets demonstrate that M2DE achieves good embedding results of high-order data for classification tasks. Yang Liu 0007, Yan Liu 0004, Keith C. C. Chan |
AAAI | 2 |
| 2010 | Rapid image retargeting based on curve-edge grid representationabstractImage retargeting technique attracts more and more attention for convenient image display on mobile devices. However, current methods can't well balance the retargeting efficiency and effectiveness, which limits their applications on the mobile devices with low computing ability. In this paper, we propose a novel image retargeting approach by combining uniform sampling and structure-aware curve-edge grid representation. We first decompose the original image into curve-edge grids by dynamic programming, and then generate the target image by uniformly sampling the pixels within the grids. The simplicity of retargeting procedure and sampling strategy enables our approach to easily achieve good computational efficiency. Furthermore, the constraint of curve-edge grid representation ensures important content emphasis and image structure preservation in the target image. Experiments on different images demonstrate the effectiveness and efficiency of our approach. Tongwei Ren, Yan Liu 0004, Gangshan Wu |
ICIP | 2 |
| 2010 | A semantic no-reference image sharpness metric based on top-down and bottom-up saliency map modelingabstractThis work presents a semantic level no-reference image sharpness/blurriness metric under the guidance of top-down & bottom-up saliency map, which is learned based on eye-tracking data by SVM. Unlike existing metrics focused on measuring the blurriness in vision level, our metric more concerns about the image content and human's intention. We integrate visual features, center priority, and semantic meaning from tag information to learn a top-down & bottom-up saliency model based on the eye-tracking data. Empirical validations on standard dataset demonstrate the effectiveness of the proposed model and metric. Shenghua Zhong, Yan Liu 0004, Yang Liu 0007, Korris Fu-Lai Chung |
ICIP | 2 |
| 2010 | Supervised manifold learning for image and video classificationabstractThis paper presents a supervised manifold learning model for dimensionality reduction in image and video classification tasks. Unlike most manifold learning models that emphasize the distance preserving, we propose a novel algorithm called maximum distance embedding (MDE), which aims to maximize the distances between some particular pairs of data points, with the intention of flattening the local nonlinearity and keeping the discriminant information simultaneously in the embedded feature space. Moreover, MDE measures the dissimilarity between data points using L1-norm distance, which is more robust to outliers than widely used Frobenius norm distance. To adapt the nature tensor structure of image and video data, we further propose the multilinear MDE (M2DE). Experiments on various datasets demonstrate that both MDE and M2DE achieve impressive embedding results of image and video data for classification tasks. Yang Liu 0007, Yan Liu 0004, Keith C. C. Chan |
ACM Multimedia | 2 |
| 2010 | Unsupervised summarization of rushes videosabstractThis paper proposes a new framework to formulate summarization of rushes video as an unsupervised learning problem. We pose the problem of video summarization as one of time-series clustering, and proposed Constrained Aligned Cluster Analysis (CACA). CACA combines kernel k-means, Dynamic Time Alignment Kernel (DTAK), and unlike previous work, CACA jointly optimizes video segmentation and shot clustering. CACA is effciently solved via dynamic programming. Experimental results on the TRECVID 2007 and 2008 BBC rushes video summarization databases validate the accuracy and effectiveness of CACA. Yang Liu 0007, Feng Zhou 0002, Wei Liu 0220, Fernando De la Torre, Yan Liu 0004 |
ACM Multimedia | 5 |
| 2010 | On the source switching problem of Peer-to-Peer streaming
Zhenhua Li 0001, Jiannong Cao 0001, Guihai Chen, Yan Liu 0004 |
J. Parallel Distributed Comput. | 4 |
| 2010 | Nonlinear dimensionality reduction with hybrid distance for trajectory representation of dynamic texture
Yang Liu 0007, Yan Liu 0004, Keith C. C. Chan |
Signal Process. | 2 |
| 2010 | Tensor Distance Based Multilinear Locality-Preserved Maximum Information EmbeddingabstractThis brief paper presents a unified framework for tensor-based dimensionality reduction (DR) with a new tensor distance (TD) metric and a novel multilinear locality-preserved maximum information embedding (MLPMIE) algorithm. Different from traditional Euclidean distance, which is constrained by the orthogonality assumption, TD measures the distance between data points by considering the relationships among different coordinates. To preserve the natural tensor structure in low-dimensional space, MLPMIE directly works on the high-order form of input data and iteratively learns the transformation matrices. In order to preserve the local geometry and to maximize the global discrimination simultaneously, MLPMIE keeps both local and global structures in a manifold model. By integrating TD into tensor embedding, TD-MLPMIE performs tensor-based DR through the whole learning procedure, and achieves stable performance improvement on various standard datasets. Yang Liu 0007, Yan Liu 0004, Keith C. C. Chan |
IEEE Trans. Neural Networks | 2 |
| 2009 | Image retargeting based on global energy optimizationabstractThis paper proposes a novel image retargeting technique based on global energy optimization. Most existing methods enhance the high energy parts of the original image by pre-defined strategies or local optimization based iterations. They can not achieve the global optimal effect in energy retainment. To solve this problem, our approach formulates image retargeting as a global optimization problem on energy. We first calculate the energy map of the original image. Then, we utilize a constrained linear programming to maximize the retained energy in retargeting. Finally, we propose a pixel fusion based method to generate the retargeted image. To make it more feasible in implementation, we further provide two strategies to reduce the time cost of our approach. We demonstrate the proposed approach by comparing with typical image retargeting methods. Tongwei Ren, Yan Liu 0004, Gangshan Wu |
ICME | 2 |
| 2009 | Tensor distance based multilinear multidimensional scaling for image and video analysisabstractThis paper presents a novel dimensionality reduction technique named Tensor Distance based Multilinear Multidimensional Scaling (TD-MMDS). First, we propose a new distance metric called Tensor Distance (TD) to build a relationship graph of data points with high-order. Then we employ an iterative strategy to sequentially learn the transformation matrices that can best keep pair-wise TDs of the high-order data in the low-dimensional embedded space. By integrating both tensor distance and tensor embedding, TD-MMDS provides a uniform framework of tensor based dimensionality reduction, which preserves the intrinsic structure of high-order data through the whole learning procedure. Experiments on standard image and video datasets validate the effectiveness of the proposed TD-MMDS. Yang Liu 0007, Yan Liu 0004 |
ACM Multimedia | 2 |
| 2009 | Image retargeting using multi-map constrained region warpingabstractImage retargeting aims to adapt images to various screens with small sizes and arbitrary aspect ratios. In this paper, we propose a novel image retargeting approach based on region warping, which emphasizes the image parts with important content while reducing the visual distortion over the whole image. First, the original image is decomposed into homogeneous regions and further represented by curve-edge trapezoid meshes. Then, two kinds of energy maps, importance map and sensitivity map, are calculated by visual attention model and weighted gradient map respectively. With mesh representation and energy map constraints, image retargeting is formulated to a constrained optimization problem of mesh vertexes relocation. Finally, the target image is generated by separately warping the regions based on the deduced optimal solution. The experiments on different images demonstrate the effective and efficiency of our algorithm. Tongwei Ren, Yan Liu 0004, Gangshan Wu |
ACM Multimedia | 2 |
| 2009 | Dimensionality reduction for heterogeneous dataset in rushes editing
Yang Liu 0007, Yan Liu 0004, Keith C. C. Chan |
Pattern Recognit. | 2 |
| 2008 | Multiple video trajectories representation using double-layer isometric feature mappingabstractThis paper proposes a novel non-linear dimensionality reduction algorithm, named double-layer isometric feature mapping (DLIso), which generates the trajectories for the video sequence containing different kinds of video clips. First, a nearest neighbor based clustering algorithm is utilized to partition the video sequence into a set of data blocks. Second, intra-cluster graphs are constructed based on the individual character of each data block to build the basic layer for DLIso. Third, the inter-cluster graph is constructed by analyzing the interrelation among these isolated data blocks to build the hyper-layer. Finally, all data points are mapped onto a unique low-dimensional feature space while preserving the corresponding relations in the double layers. Experiments on synthetic datasets as well as the real video sequences demonstrate that the low-dimensional trajectories generated by the proposed method correctly represent the semantic information of the data. Yang Liu 0007, Yan Liu 0004, Keith C. C. Chan |
ICME | 2 |
| 2008 | Full-reference quality diagnosis for video summaryabstractAs video summarization techniques have attracted more and more attention for efficient multimedia data management, objective quality assessment of video summary is desired. To address the lack of automatic evaluation techniques, this paper proposes a 3C-diagnosis algorithm to diagnose the video summary from the perspective of coverage, conciseness, and coherence. The candidate summary is first aligned against the reference summary. Then the coverage of the candidate summary is calculated according to the information bearing of the matcThing frames and the information loss of the missing frames. The conciseness is calculated based on the unwanted information contained in the candidate summary, and the coherence is calculated based on the ratio of the appearances of the frame loss for the aligned candidate summary. The proposed techniques are experimented on a standard dataset of TRECVID 2007 and show good performance. Yan Liu 0004, Yan Zhang 0031, Maosong Sun 0001, Wenjie Li 0002 |
ICME | 1 |
| 2008 | Fast Source Switching for Gossip-Based Peer-to-Peer StreamingabstractIn this paper we consider gossip-based peer-to-peer streaming applications where multiple sources exist and they work serially. More specifically, we tackle the problem of fast source switching to minimize the startup delay of the new source. We model the source switch process and formulate it into an optimization problem. Then we propose a practical greedy algorithm that can approximate the optimal solution by properly interleaving the data delivery of the old source and the new source. We perform simulations on various real-trace overlay topologies to demonstrate the effectiveness of our algorithm. The simulation results show that our proposed algorithm outperforms the normal source switch algorithm by reducing the source switch time by 20%-30% without bringing extra communication overhead, and the reduction ratio tends to increase when the network scale expands. Zhenhua Li 0001, Jiannong Cao 0001, Guihai Chen, Yan Liu 0004 |
ICPP | 4 |
| 2007 | Human-Centered Multimedia E-Learning System for Real-Time Interactive Distance EducationabstractHuman-centered computing embeds the human factors such as social and cultural awareness, human abilities, and adaptability to others' activities in the information system and methodology design. This paper proposes a human-centered multimedia E-learning system to address several classical difficulties in distance education via Internet: poor video frame quality, high requirement of network bandwidth, and inefficient interaction. We propose a novel data partition and authoring method to improve the image quality of the lecture notes, handwritings, and web pages. We use a peer-to-peer network structure to support real-time multimedia data broadcasting. The system can provide the real-time interaction between the remote students and the instructor just like in the classroom lecture. This system also records the multimedia data with an operation log file to support offline multimedia retrieval and summarization. We test our system under the practical education environment in our university for one semester. Yan Liu 0004, Yung Hoi Wah |
ICME | 1 |
| 2004 | Video feature selection using fast-converging sort-merge treeabstractHigh time complexity is a bottle-neck in video segmentation, classification, analysis, and retrieval. In This work we use a heuristic method called fast-converging sort-merge tree (FSMT) to construct automatically a hierarchy of small subsets of features that are progressively more useful for video data exploration. The method combines the virtues of a wrapper model approach for high accuracy, with those of a filter method approach for deriving the appropriate features quickly. FSMT speeds up a more fundamental method, the basic sort-merge tree (BSMT) approach, while retaining its performance. We demonstrate FSMT's high accuracy: it has a 0.001 error rate in a frame classification task on 75 minutes of instructional video, and a 0.98 precision and 0.89 recall in a segment retrieval task on 30 minutes of sports video. Additionally, FSMT is more than 80% faster than its predecessor, BSMT. Yan Liu 0004, John R. Kender |
ICME | 1 |
| 2003 | Fast scene segmentation using multi-level feature selectionabstractHigh time cost is the bottle-neck of video scene segmentation. In this paper we use a heuristic method called sort-merge feature selection to construct automatically a hierarchy of small subsets of features that are progressively more useful for segmentation. A novel combination of fastmap for dimensionality reduction and Mahalanobis distance for likelihood determination is used as induction algorithm. Because these induced feature sets from a hierarchy with increasing classification accuracy, video segments can be segmented and categorized simultaneously in a coarse-fine manner that efficiently and progressively detects and refines their temporal boundaries. We analyze the performance of these methods, and demonstrate them in the domain of long (75 minute) instructional video. Yan Liu 0004, John R. Kender |
ICME | 1 |
| 2003 | Sort-Merge Feature Selection for Video DataabstractApplying existing feature selection algorithms to video classification is impractical. A novel algorithm called Basic Sort-Merge Tree (BSMT) is proposed to choose a very small subset of features for video classification in linear time in the number of features. We reduce the cardinality of the input data by sorting the individual features by their effectiveness in categorization, and then merging pairwise these features into feature sets of cardinality two. Repeating this Sort-Merge process several times results in the learning of a small-cardinality, efficient, but highly accurate feature set. As the wrapper model, this paper exploits a novel combination of Fastmap for dimensionality reduction and Mahalanobis distance for likelihood determination. The time complexity of this induction part is linear in the number of training data. We provide theoretical proof of time cost and empirical validation of the accuracy. Yan Liu 0004, John R. Kender |
SDM | 1 |
| 2003 | Fast video segment retrieval by Sort-Merge feature selection, boundary refinement, and lazy evaluation
Yan Liu 0004, John R. Kender |
Comput. Vis. Image Underst. | 1 |
| 2002 | Semantic Extraction and Semantics-Based Annotation and Retrieval for Video Databases
Yan Liu 0004 |
Multim. Tools Appl. | 1 |