EDBT 2026 Demo / reviewers in the wild / expert
Yang Liu 0007
dblp:51/3710-7
· DBLP profile ↗
72ranked-venue papers
23as first author
17since 2021 · last 2026
0000-0002-0166-3944ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 38 · 12 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 28 · 12 first-author · 3 since 2021Databases, data management, data science and information retrieval · 10 · 5 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 6 since 2021Computer networks · 2Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PH-EMO: Decoding Emotions from the Brain Inward - EEG-Grounded Multimodal Reasoning with LLMs
Kehong Liu, Yang Liu 0007, Jiming Liu 0001 |
WWW | 2 |
| 2026 | Towards Performatively Stable Equilibria in Decision-Dependent Games for Arbitrary Data Distribution MapsabstractAbstract In decision-dependent games, multiple players optimize their decisions under data distributions that shift with their joint actions, creating complex dynamics in applications like market pricing. A practical consequence of these dynamics is the performatively stable equilibrium , where each player’s strategy is a best response under the induced distribution. Prior work relies on $$\beta $$ -smoothness, assuming Lipschitz continuity of loss function gradients with respect to data distributions, which is impractical as the data distribution maps, i.e., the relationship between joint decision and the resulting distribution shifts, are typically unknown, rendering $$\beta $$ unobtainable. To overcome this limitation, we propose a gradient-based $$\hat{\varepsilon }_i$$ -sensitivity measure. It directly quantifies the impact of decision-induced distribution shifts on decision-making and is calculable for arbitrary data distribution maps. Leveraging this measure, we derive convergence guarantees for performatively stable equilibria under a practically feasible assumption of $$\alpha $$ -strong monotonicity. Notably, we establish a linear convergence rate in finite sample scenarios when $$\alpha > 2 \sqrt{\sum _{i = 1}^n \hat{\varepsilon }_i^2}$$ , with a probability depending on sample complexity. Accordingly, we develop a sensitivity-informed repeated retraining algorithm that adjusts players’ loss functions based on the sensitivity measure to achieve the strong monotonicity. This approach ensures the game satisfies the derived convergence condition and thus guarantees convergence to performatively stable equilibria for arbitrary data distribution maps. Experiments with various data distribution maps on prediction error minimization game, Cournot competition, and revenue maximization game show that our approach outperforms state-of-the-art baselines, achieving lower losses and faster convergence, validating the theoretical convergence conditions and confirming the effectiveness of the proposed algorithm. Guangzheng Zhong, Yang Liu 0007, Jiming Liu 0001 |
Mach. Learn. | 2 |
| 2026 | Performative Prediction in the Wild: Adapting to Arbitrary Data Distribution MapsabstractAbstract Performative prediction refers to scenarios where model predictions influence the underlying data distribution they aim to predict. A desirable property in this context is performative stability , where model predictions are already optimal for the distribution they induce, indicating converged model parameters and no need for further retraining. Achieving performative stability requires characterizing the data distribution map $$\mathcal {D}(\theta )$$ , i.e., the relationship between predictions and the resulting distribution shifts. Current studies typically quantify distribution differences using metrics like $$\mathcal {W}_1$$ distance or $$\chi ^2$$ divergence, which may not provide isometric embeddings or maintain metric equivalence in practical scenarios, limiting their applicability across various data distribution maps. Moreover, the crucial smoothness parameter $$\beta $$ in existing work is often unobtainable in performative scenarios, constraining the real-world utility of current theoretical results and methods. To address these challenges, we develop an algorithm that learns a performatively stable model for arbitrary data distribution maps without requiring the joint smoothness parameter $$\beta $$ . Specifically, we introduce a new $$\hat{\varepsilon }$$ -sensitivity measure for $$\mathcal {D}(\theta )$$ , quantified by the gradient of the loss function, which naturally and directly characterizes how distribution shifts affect the optimization of the objective function. Based on this sensitivity, we formulate a $$\gamma $$ -strongly convex loss function and optimize the deployed model accordingly, where $$\gamma $$ is derived from the defined $$\hat{\varepsilon }$$ , eliminating the need for the $$\beta $$ -joint smoothness assumption. Our theoretical results guarantee the convergence of the deployed model to performative stability. Extensive experiments on synthetic and real-world datasets with diverse data distribution maps demonstrate the superiority of our method over state-of-the-art techniques in two key aspects: prediction accuracy and performative stability. Guangzheng Zhong, Yang Liu 0007, Ruichen Liu, Jiming Liu 0001 |
Mach. Learn. | 2 |
| 2026 | NIDC: General Task Backbone for Neuroimaging Analysis via Interpretable Deep ClusteringabstractClustering techniques offer strong interpretability. However, they have significant limitations in the deep learning area due to their difficulty in capturing complex data structures, such as spatial and contextual information. This issue is especially pronounced in neuroimaging research, where high-dimensional and complex data greatly restricts the feature representation capability of clustering models. Hence, we propose a Neuroimaging Deep Clustering (NIDC) backbone network. We convert 3D neuroimagings into point sets and design clustering-based paradigms for context feature aggregation, feature interaction, and feature dispatching to enable deep feature extraction from the point sets. To better capture spatial information, we propose brain spatial relative position encoding, which assists the clustering paradigm in better understanding the anatomical structure of brain tissue and the positional relationships between different regions. Additionally, we design a sample center loss function to encourage tighter clustering of labels or voxels/feature points of the same class in the feature space, aiming to suppress both inter-class and intra-class similarities. Meanwhile, NIDC preserves the interpretability of traditional clustering techniques, allowing it to uncover relationships between brain regions and trace the decision-making process of each voxel. NIDC achieves highly competitive performance across multiple datasets in various downstream tasks, emerging as a new and practical solution for neuroimaging analysis. Code is available athttps://github.com/IMCTGD/NIDC. Jiayu Ye, An Zeng, Dan Pan 0001, Jingliang Zhao, Yiqun Zhang 0006, Yang Liu 0007 |
IEEE Trans. Multim. | 7 |
| 2025 | MAP the Blockchain World: A Trustless and Scalable Blockchain Interoperability Protocol for Cross-chain ApplicationsabstractBlockchain interoperability protocols enable cross-chain asset transfers or data retrievals between isolated chains, which are considered as one of the core infrastructure for Web 3.0. However, existing protocols either face severe scalability issues due to high on-chain and off-chain cost, or suffer from trust concerns because of centralized architecture. Yinfeng Cao, Jiannong Cao 0001, Dongbin Bai, Long Wen 0002, Yang Liu 0007, Ruidong Li 0001 |
WWW | 5 |
| 2025 | W-DOE: Wasserstein Distribution-Agnostic Outlier ExposureabstractIn open-world environments, classification models should be adept at identifying out-of-distribution (OOD) data whose semantics differ from in-distribution (ID) data, leading to the emerging research in OOD detection. As a promising learning scheme, outlier exposure (OE) enables the models to learn from auxiliary OOD data, enhancing model representations in discerning between ID and OOD patterns. However, these auxiliary OOD data often do not fully represent real OOD scenarios, potentially biasing our models in practical OOD detection. Hence, we propose a novel OE-based learning method termed Wasserstein Distribution-agnostic Outlier Exposure (W-DOE), which is both theoretically sound and experimentally superior to previous works. The intuition is that by expanding the coverage of training-time OOD data, the models will encounter fewer unseen OOD cases upon deployment. In W-DOE, we achieve additional OOD data to enlarge the OOD coverage, based on a new data synthesis approach called implicit data synthesis (IDS). It is driven by our new insight that perturbing model parameters can lead to implicit data transformation, which is simple to implement yet effective to realize. Furthermore, we suggest a general learning framework to search for the synthesized OOD data that can benefit the models most, ensuring the OOD performance for the enlarged OOD coverage measured by the Wasserstein metric. Our approach comes with provable guarantees for open-world settings, demonstrating that broader OOD coverage ensures reduced estimation errors and thereby improved generalization for real OOD cases. We conduct extensive experiments across a series of representative OOD detection setups, further validating the superiority of W-DOE against state-of-the-art counterparts in the field. Bo Han 0003, Yang Liu 0007, Chen Gong 0002, Tongliang Liu, Jiming Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | Evolving meta-correlation classes for binary similarityabstractIn the field of machine learning and pattern recognition, the use of binary correlation indices is essential for accurate prediction and modelling. This work presents a novel evolutionary method to address the problem of discovering binary correlation indices in different application domains. The proposed approach introduces the concept of meta-correlation, a parametric formula representing classes of binary similarity indices, and optimizes it through an evolutionary scheme. The method has been experimented with and validated in the context of the link prediction problem based on local topological similarity (i.e. graph neighbourhood). A Differential Evolution optimization algorithm finds the evolved correlations that perform best in a given domain. Experiments conducted across different network domains have shown that the instances of the discovered meta-correlations generally outperform state-of-the-art binary correlation indices for all the experimented domains. This approach effectively explores the correlation space and can find a unique pattern that adapts to the domains under consideration. The meta-correlation classes can be applied to both topological and semantic similarity problems, taking into account local information without requiring complete knowledge of the graph. Valentina Franzoni, Giulio Biondi, Yang Liu 0007, Alfredo Milani |
Pattern Recognit. | 3 |
| 2025 | CmdVIT: A Voluntary Facial Expression Recognition Model for Complex Mental DisordersabstractFacial Expression Recognition (FER) is a critical method for evaluating the emotional states of patients with mental disorders, playing a significant role in treatment monitoring. However, due to privacy constraints, facial expression data from patients with mental disorders is severely limited. Additionally, the more complex inter-class and intra-class similarities compared to healthy individuals make accurate recognition of facial expressions challenging. Therefore, we propose a Voluntary Facial Expression Mimicry (VFEM) experiment, which collected facial expression data from schizophrenia, depression, and anxiety. This experiment establishes the first dataset designed for facial expression recognition tasks exclusively composed of patients with mental disorders. Simultaneously, based on VFEM, we propose a Vision Transformer FER model tailored for Complex mental disorder patients (CmdVIT). CmdVIT integrates crucial facial expression features through both explicit and implicit mechanisms, including explicit visual center positional encoding and implicit sparse attention center loss function. These two key components enhance positional information and minimize the facial feature space distance between conventional attention and critical attention, effectively suppressing inter-class and intra-class similarities. In various FER tasks for different mental disorders in VFEM, CmdVIT achieves more competitive performance compared to contemporary benchmark models. Our works are available at https://github.com/yjy-97/CmdVIT. Jiayu Ye, Yanhong Yu, Qingxiang Wang, Guolong Liu, An Zeng, Yiqun Zhang 0006, Yang Liu 0007, Yunshao Zheng |
IEEE Trans. Image Process. | 8 |
| 2024 | Polarized message-passing in graph neural networksabstractIn this paper, we present Polarized message-passing (PMP), a novel paradigm to revolutionize the design of message-passing graph neural networks (GNNs). In contrast to existing methods, PMP captures the power of node-node similarity and dissimilarity to acquire dual sources of messages from neighbors. The messages are then coalesced to enable GNNs to learn expressive representations from sparse but strongly correlated neighbors. Three novel GNNs based on the PMP paradigm, namely PMP graph convolutional network (PMP-GCN), PMP graph attention network (PMP-GAT), and PMP graph PageRank network (PMP-GPN) are proposed to perform various downstream tasks. Theoretical analysis is also conducted to verify the high expressiveness of the proposed PMP-based GNNs. In addition, an empirical study of five learning tasks based on 12 real-world datasets is conducted to validate the performances of PMP-GCN, PMP-GAT, and PMP-GPN. The proposed PMP-GCN, PMP-GAT, and PMP-GPN outperform numerous strong message-passing GNNs across all five learning tasks, demonstrating the effectiveness of the proposed PMP paradigm. Tiantian He 0001, Yang Liu 0007, Yew-Soon Ong, Xin Luo 0001 |
Artif. Intell. | 2 |
| 2024 | Commonality and Individuality-Based Subspace LearningabstractSubspace learning (SL) plays a key role in various learning tasks, especially those with a huge feature space. When processing multiple high-dimensional learning tasks simultaneously, it is of great importance to make use of the subspace extracted from some tasks to help learn others, so that the learning performance of all tasks can be enhanced together. To achieve this goal, it is crucial to answer the following question: How can the commonality among different learning tasks and, of equal importance, the individuality of each single learning task, be characterized and extracted from the given datasets, so as to benefit the subsequent learning, for example, classification? Existing multitask SL methods usually focused on the commonality among the given tasks, while neglecting the individuality of the learning tasks. In order to offer a more general and comprehensive framework for multitask SL, in this article, we propose a novel method dubbed commonality and individuality-based SL (CISL). First, we formally define the notions and objective functions of both commonality and individuality with respect to multiple SL tasks. Then, we design an iterative algorithm to solve the formulated objective functions, with the convergence of the algorithm being guaranteed. To show the generality of the proposed method, we theoretically analyze its connections to existing single-task and multitask SL methods. Finally, we demonstrate the necessity and effectiveness of incorporating both commonality and individuality by interpreting the learned subspaces and comparing the performance of CISL (in terms of the subsequent classification accuracy) with that of classical and state-of-the-art SL approaches on both synthetic and real-world multitask datasets. The empirical evaluation validates the effectiveness of the proposed method in characterizing the commonality and individuality for multitask SL. Jinfu Ren, Yang Liu 0007, Jiming Liu 0001 |
IEEE Trans. Cybern. | 2 |
| 2024 | A Physics-Guided Attention-Based Neural Network for Sea Surface Temperature PredictionabstractAccurate prediction of sea surface temperature (SST) is crucial in the field of oceanography, as it has a significant impact on various physical, chemical, and biological processes in the marine environment. In this study, we propose a physics-guided attention-based neural network (PANN) to address the spatiotemporal SST prediction problem. The PANN model incorporates data-driven spatiotemporal convolution operations and the underlying physical dynamics of SSTs using a cross-attention mechanism. First, we construct a spatiotemporal convolution module (SCM) using convolutional long short-term memory (ConvLSTM) to capture the spatial and temporal correlations present in the time series of the SST data. We then introduce a physical constraint module (PCM) to mimic the transport dynamics in fluids based on data assimilation techniques used to solve partial differential equations (PDEs). Consequently, we employ an attention fusion module (AFM) to effectively combine the data-driven and PDE-constrained predictions obtained from the SCM and PCM, aiming at enhancing the accuracy of the predictions. To evaluate the performance of the proposed model, we conduct short-term SST forecasts in the East China Sea (ECS) with forecast lead times ranging from one to ten days, by comparing it with several state-of-the-art models, including ConvLSTM, PredRNN, temporal convolutional transformer network (TCTN), convolutional gated recurrent unit (ConvGRU), and SwinLSTM. The experimental results demonstrate that our proposed model outperforms these models in terms of multiple evaluation metrics for short-term predictions. Benyun Shi, Liu Feng, Hailun He, Yingjian Hao, Miao Liu 0003, Yang Liu 0007, Jiming Liu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2024 | MAD-Former: A Traceable Interpretability Model for Alzheimer's Disease Recognition Based on Multi-Patch AttentionabstractThe integration of structural magnetic resonance imaging (sMRI) and deep learning techniques is one of the important research directions for the automatic diagnosis of Alzheimer's disease (AD). Despite the satisfactory performance achieved by existing voxel-based models based on convolutional neural networks (CNNs), such models only handle AD-related brain atrophy at a single spatial scale and lack spatial localization of abnormal brain regions based on model interpretability. To address the above limitations, we propose a traceable interpretability model for AD recognition based on multi-patch attention (MAD-Former). MAD-Former consists of two parts: recognition and interpretability. In the recognition part, we design a 3D brain feature extraction network to extract local features, followed by constructing a dual-branch attention structure with different patch sizes to achieve global feature extraction, forming a multi-scale spatial feature extraction framework. Meanwhile, we propose an important attention similarity position loss function to assist in model decision-making. The interpretability part proposes a traceable method that can obtain a 3D ROI space through attention-based selection and receptive field tracing. This space encompasses key brain tissues that influence model decisions. Experimental results reveal the significant role of brain tissues such as the Fusiform Gyrus (FuG) in AD recognition. MAD-Former achieves outstanding performance in different tasks on ADNI and OASIS datasets, demonstrating reliable model interpretability. Jiayu Ye, An Zeng, Dan Pan 0001, Yiqun Zhang 0006, Jingliang Zhao, Qiuping Chen, Yang Liu 0007 |
IEEE J. Biomed. Health Informatics | 7 |
| 2023 | Epidemiology-aware Deep Learning for Infectious Disease Dynamics PredictionabstractInfectious disease risk prediction plays a vital role in disease control and prevention. Recent studies in machine learning have attempted to incorporate epidemiological knowledge into the learning process to enhance the accuracy and informativeness of prediction results for decision-making. However, these methods commonly involve single-patch mechanistic models, overlooking the disease spread across multiple locations caused by human mobility. Additionally, these methods often require extra information beyond the infection data, which is typically unavailable in reality. To address these issues, this paper proposes a novel epidemiology-aware deep learning framework that integrates a fundamental epidemic component, the next-generation matrix (NGM), into the deep architecture and objective function. This integration enables the inclusion of both mechanistic models and human mobility in the learning process to characterize within- and cross-location disease transmission. From this framework, two novel methods, Epi-CNNRNN-Res and Epi-Cola-GNN, are further developed to predict epidemics, with experimental results validating their effectiveness. Mutong Liu, Yang Liu 0007, Jiming Liu 0001 |
CIKM | 2 |
| 2023 | Anatomical-Aware Point-Voxel Network for Couinaud Segmentation in Liver CT
Xukun Zhang, Yang Liu 0007, Sharib Ali, Minghao Han, Tao Liu 0050, Peng Zhai, Zhiming Cui 0001, Peixuan Zhang, Lihua Zhang 0002 |
MICCAI (3) | 2 |
| 2021 | Heterogeneous neural metric learning for spatio-temporal modeling of infectious diseases with incomplete data
Qi Tan 0002, Yang Liu 0007, Jiming Liu 0001, Benyun Shi, Shang Xia, Xiao-Nong Zhou |
Neurocomputing | 2 |
| 2021 | A comprehensive analysis of classification methods in gastrointestinal endoscopy imagingabstractGastrointestinal (GI) endoscopy has been an active field of research motivated by the large number of highly lethal GI cancers. Early GI cancer precursors are often missed during the endoscopic surveillance. The high missed rate of such abnormalities during endoscopy is thus a critical bottleneck. Lack of attentiveness due to tiring procedures, and requirement of training are few contributing factors. An automatic GI disease classification system can help reduce such risks by flagging suspicious frames and lesions. GI endoscopy consists of several multi-organ surveillance, therefore, there is need to develop methods that can generalize to various endoscopic findings. In this realm, we present a comprehensive analysis of the Medico GI challenges: Medical Multimedia Task at MediaEval 2017, Medico Multimedia Task at MediaEval 2018, and BioMedia ACM MM Grand Challenge 2019. These challenges are initiative to set-up a benchmark for different computer vision methods applied to the multi-class endoscopic images and promote to build new approaches that could reliably be used in clinics. We report the performance of 21 participating teams over a period of three consecutive years and provide a detailed analysis of the methods used by the participants, highlighting the challenges and shortcomings of the current approaches and dissect their credibility for the use in clinical settings. Our analysis revealed that the participants achieved an improvement on maximum Mathew correlation coefficient (MCC) from 82.68% in 2017 to 93.98% in 2018 and 95.20% in 2019 challenges, and a significant increase in computational speed over consecutive years. Debesh Jha, Sharib Ali, Steven Alexander Hicks, Vajira Thambawita, Hanna Borgli, Pia H. Smedsrud, Thomas de Lange, Konstantin Pogorelov, Philipp Harzig, Minh-Triet Tran, Wenhua Meng, Trung-Hieu Hoang, Danielle Dias, Tobey H. Ko, Taruna Agrawal, Olga Ostroukhova, Zeshan Khan, Muhammad Atif Tahir, Yang Liu 0007, Mathias Kirkerød, Dag Johansen, Mathias Lux, Håvard D. Johansen, Michael Riegler 0001, Pål Halvorsen |
Medical Image Anal. | 20 |
| 2021 | Demystifying Deep Learning in Predictive Spatiotemporal Analytics: An Information-Theoretic FrameworkabstractDeep learning has achieved incredible success over the past years, especially in various challenging predictive spatiotemporal analytics (PSTA) tasks, such as disease prediction, climate forecast, and traffic prediction, where intrinsic dependence relationships among data exist and generally manifest at multiple spatiotemporal scales. However, given a specific PSTA task and the corresponding data set, how to appropriately determine the desired configuration of a deep learning model, theoretically analyze the model's learning behavior, and quantitatively characterize the model's learning capacity remains a mystery. In order to demystify the power of deep learning for PSTA in a theoretically sound and explainable way, in this article, we provide a comprehensive framework for deep learning model design and information-theoretic analysis. First, we develop and demonstrate a novel interactively and integratively connected deep recurrent neural network (I2DRNN) model. I2DRNN consists of three modules: an input module that integrates data from heterogeneous sources; a hidden module that captures the information at different scales while allowing the information to flow interactively between layers; and an output module that models the integrative effects of information from various hidden layers to generate the output predictions. Second, to theoretically prove that our designed model can learn multiscale spatiotemporal dependence in PSTA tasks, we provide an information-theoretic analysis to examine the information-based learning capacity (i-CAP) of the proposed model. In so doing, we can tackle an important open question in deep learning, that is, how to determine the necessary and sufficient configurations of a designed deep learning model with respect to the given learning data sets. Third, to validate the I2DRNN model and confirm its i-CAP, we systematically conduct a series of experiments involving both synthetic data sets and real-world PSTA tasks. The experimental results show that the I2DRNN model outperforms both classical and state-of-the-art models on all data sets and PSTA tasks. More importantly, as readily validated, the proposed model captures the multiscale spatiotemporal dependence, which is meaningful in the real-world context. Furthermore, the model configuration that corresponds to the best performance on a given data set always falls into the range between the necessary and sufficient configurations, as derived from the information-theoretic analysis. Qi Tan 0002, Yang Liu 0007, Jiming Liu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2020 | Metric-Guided Multi-task Learning
Jinfu Ren, Yang Liu 0007, Jiming Liu 0001 |
ISMIS | 2 |
| 2020 | Mesoscale Anisotropically-Connected Learning
Qi Tan 0002, Yang Liu 0007, Jiming Liu 0001 |
ISMIS | 2 |
| 2020 | Contextual Correlation Preserving Multiview Featured Graph ClusteringabstractGraph clustering, which aims at discovering sets of related vertices in graph-structured data, plays a crucial role in various applications, such as social community detection and biological module discovery. With the huge increase in the volume of data in recent years, graph clustering is used in an increasing number of real-life scenarios. However, the classical and state-of-the-art methods, which consider only single-view features or a single vector concatenating features from different views and neglect the contextual correlation between pairwise features, are insufficient for the task, as features that characterize vertices in a graph are usually from multiple views and the contextual correlation between pairwise features may influence the cluster preference for vertices. To address this challenging problem, we introduce in this paper, a novel graph clustering model, dubbed contextual correlation preserving multiview featured graph clustering (CCPMVFGC) for discovering clusters in graphs with multiview vertex features. Unlike most of the aforementioned approaches, CCPMVFGC is capable of learning a shared latent space from multiview features as the cluster preference for each vertex and making use of this latent space to model the inter-relationship between pairwise vertices. CCPMVFGC uses an effective method to compute the degree of contextual correlation between pairwise vertex features and utilizes view-wise latent space representing the feature-cluster preference to model the computed correlation. Thus, the cluster preference learned by CCPMVFGC is jointly inferred by multiview features, view-wise correlations of pairwise features, and the graph topology. Accordingly, we propose a unified objective function for CCPMVFGC and develop an iterative strategy to solve the formulated optimization problem. We also provide the theoretical analysis of the proposed model, including convergence proof and computational complexity analysis. In our experiments, we extensively compare the proposed CCPMVFGC with both classical and state-of-the-art graph clustering methods on eight standard graph datasets (six multiview and two single-view datasets). The results show that CCPMVFGC achieves competitive performance on all eight datasets, which validates the effectiveness of the proposed model. Tiantian He 0001, Yang Liu 0007, Tobey H. Ko, Keith C. C. Chan, Yew-Soon Ong |
IEEE Trans. Cybern. | 2 |
| 2020 | Identifying Key Opinion Leaders in Social Media via Modality-Consistent Harmonized Discriminant EmbeddingabstractThe digital age has empowered brands with new and more effective targeted marketing tools in the form of key opinion leaders (KOLs). Because of the KOLs' unique capability to draw specific types of audience and cultivate long-term relationship with them, correctly identifying the most suitable KOLs within a social network is of great importance, and sometimes could govern the success or failure of a brand's online marketing campaigns. However, given the high dimensionality of social media data, conducting effective KOL identification by means of data mining is especially challenging. Owing to the generally multiple modalities of the user profiles and user-generated content (UGC) over the social networks, we can approach the KOL identification process as a multimodal learning task, with KOLs as a rare yet far more important class over non-KOLs in our consideration. In this regard, learning the compact and informative representation from the high-dimensional multimodal space is crucial in KOL identification. To address this challenging problem, in this paper, we propose a novel subspace learning algorithm dubbed modality-consistent harmonized discriminant embedding (MCHDE) to uncover the low-dimensional discriminative representation from the social media data for identifying KOLs. Specifically, MCHDE aims to find a common subspace for multiple modalities, in which the local geometric structure, the harmonized discriminant information, and the modality consistency of the dataset could be preserved simultaneously. The above objective is then formulated as a generalized eigendecomposition problem and the closed-form solution is obtained. Experiments on both synthetic example and a real-world KOL dataset validate the effectiveness of the proposed method. Yang Liu 0007, Zhonglei Gu, Tobey H. Ko, Jiming Liu 0001 |
IEEE Trans. Cybern. | 1 |
| 2019 | EWGAN: Entropy-Based Wasserstein GAN for Imbalanced LearningabstractIn this paper, we propose a novel oversampling strategy dubbed Entropy-based Wasserstein Generative Adversarial Network (EWGAN) to generate data samples for minority classes in imbalanced learning. First, we construct an entropyweighted label vector for each class to characterize the data imbalance in different classes. Then we concatenate this entropyweighted label vector with the original feature vector of each data sample, and feed it into the WGAN model to train the generator. After the generator is trained, we concatenate the entropy-weighted label vector with random noise feature vectors, and feed them into the generator to generate data samples for minority classes. Experimental results on two benchmark datasets show that the samples generated by the proposed oversampling strategy can help to improve the classification performance when the data are highly imbalanced. Furthermore, the proposed strategy outperforms other state-of-the-art oversampling algorithms in terms of the classification accuracy. Jinfu Ren, Yang Liu 0007, Jiming Liu 0001 |
AAAI | 2 |
| 2019 | Data Management in Supply Chain Using Blockchain: Challenges and a Case StudyabstractSupply chain management (SCM) is fundamental for gaining financial, environmental and social benefits in the supply chain industry. However, traditional SCM mechanisms usually suffer from a wide scope of issues such as lack of information sharing, long delays for data retrieval, and unreliability in product tracing. Recent advances in blockchain technology show great potential to tackle these issues due to its salient features including immutability, transparency, and decentralization. Although there are some proof-of-concept studies and surveys on blockchain-based SCM from the perspective of logistics, the underlying technical challenges are not clearly identified. In this paper, we provide a comprehensive analysis of potential opportunities, new requirements, and principles of designing blockchain-based SCM systems. We summarize and discuss four crucial technical challenges in terms of scalability, throughput, access control, data retrieval and review the promising solutions. Finally, a case study of designing blockchain-based food traceability system is reported to provide more insights on how to tackle these technical challenges in practice. Jiannong Cao 0001, Yanni Yang 0003, Cheung Leong Tung, Shan Jiang 0005, Bin Tang 0002, Yang Liu 0007, Yuming Deng |
ICCCN | 7 |
| 2018 | Bayesian Network Structure Learning: The Two-Step Clustering-Based AlgorithmabstractIn this paper we introduce a two-step clustering-based strategy, which can automatically generate prior information from data in order to further improve the accuracy and time efficiency of state-of-the-art algorithms for Bayesian network structure learning. Our clustering-based strategy is composed of two steps. In the first step, we divide the potential nodes into several groups via clustering analysis and apply Bayesian network structure learning to obtain some pre-existing arcs within each cluster. In the second step, with all the within-cluster arcs being well preserved, we learn the between-cluster structure of the given network. Experimental results on benchmark datasets show that a wide range of structure learning algorithms benefit from the proposed clustering-based strategy in terms of both accuracy and efficiency. Jiming Liu 0001, Yang Liu 0007 |
AAAI | 3 |
| 2018 | Multi-scale and Discriminative Part Detectors Based Features for Multi-label Image ClassificationabstractConvolutional neural networks (CNNs) have shown their promise for image classification task. However, global CNN features still lack geometric invariance for addressing the problem of intra-class variations and so are not optimal for multi-label image classification. This paper proposes a new and effective framework built upon CNNs to learn Multi-scale and Discriminative Part Detectors (MsDPD)-based feature representations for multi-label image classification. Specifically, at each scale level, we (i) first present an entropy-rank based scheme to generate and select a set of discriminative part detectors (DPD), and then (ii) obtain a number of DPD-based convolutional feature maps with each feature map representing the occurrence probability of a particular part detector and learn DPD-based features by using a task-driven pooling scheme. The two steps are formulated into a unified framework by developing a new objective function, which jointly trains part detectors incrementally and integrates the learning of feature representations into the classification task. Finally, the multi-scale features are fused to produce the predictions. Experimental results on PASCAL VOC 2007 and VOC 2012 datasets demonstrate that the proposed method achieves better accuracy when compared with the existing state-of-the-art multi-label classification methods. Gong Cheng 0003, Decheng Gao, Yang Liu 0007, Junwei Han 0001 |
IJCAI | 3 |
| 2018 | Joint Learning of Phenotypes and Diagnosis-Medication Correspondence via Hidden Interaction Tensor FactorizationabstractNon-negative tensor factorization has been shown effective for discovering phenotypes from the EHR data with minimal human supervision. In most cases, an interaction tensor of the elements in the EHR (e.g., diagnoses and medications) has to be first established before the factorization can be applied. Such correspondence information however is often missing. While different heuristics can be used to estimate the missing correspondence, any errors introduced will in turn cause inaccuracy for the subsequent phenotype discovery task. This is especially true for patients with multiple diseases diagnosed (e.g., under critical care). To alleviate this limitation, we propose the hidden interaction tensor factorization (HITF) where the diagnosis-medication correspondence and the underlying phenotypes are inferred simultaneously. We formulate it under a Poisson non-negative tensor factorization framework and learn the HITF model via maximum likelihood estimation. For performance evaluation, we applied HITF to the MIMIC III dataset. Our empirical results show that both the phenotypes and the correspondence inferred are clinically meaningful. In addition, the inferred HITF model outperforms a number of state-of-the-art methods for mortality prediction. Kejing Yin, William Kwok-Wai Cheung, Yang Liu 0007, Benjamin C. M. Fung, Jonathan Poon |
IJCAI | 3 |
| 2018 | Sparse Multi-label Bilinear Embedding on Stiefel Manifolds
Yang Liu 0007, Guohua Dong, Zhonglei Gu |
ISMIS | 1 |
| 2018 | Learning Perceptual Embeddings with Two Related Tasks for Joint Predictions of Media Interestingness and EmotionsabstractIntegrating media elements of various medium, multimedia is capable of expressing complex information in a neat and compact way. Early studies have linked different sensory presentation in multimedia with the perception of human-like concepts. Yet, the richness of information in multimedia makes understanding and predicting user perceptions in multimedia content a challenging task both to the machine and the human mind. This paper presents a novel multi-task feature extraction method for accurate prediction of user perceptions in multimedia content. Differentiating from the conventional feature extraction algorithms which focus on perfecting a single task, the proposed model recognizes the commonality between different perceptions (e.g., interestingness and emotional impact), and attempts to jointly optimize the performance of all the tasks through uncovered commonality features. Using both a media interestingness dataset and a media emotion dataset for user perception prediction tasks, the proposed model attempts to simultaneously characterize the individualities of each task and capture the commonalities shared by both tasks, and achieves better accuracy in predictions than other competing algorithms on real-world datasets of two related tasks: MediaEval 2017 Predicting Media Interestingness Task and MediaEval 2017 Emotional Impact of Movies Task. Yang Liu 0007, Zhonglei Gu, Tobey H. Ko, Kien A. Hua |
ICMR | 1 |
| 2018 | Motif-Aware Diffusion Network Inference
Qi Tan 0002, Yang Liu 0007, Jiming Liu 0001 |
PAKDD (3) | 2 |
| 2018 | Who is the Mr. Right for Your Brand?: - Discovering Brand Key Assets via Multi-modal Asset-aware ProjectionabstractFollowing the rising prominence of online social networks, we observe an emerging trend for brands to adopt influencer marketing, embracing key opinion leaders (KOLs) to reach potential customers (PCs) online. Owing to the growing strategic importance of these brand key assets, this paper presents a novel feature extraction method named Multi-modal Asset-aware Projection (M2A2P) to learn a discriminative subspace from the high-dimensional multi-modal social media data for effective brand key asset discovery. By formulating a new asset-aware discriminative information preserving criterion, M2A2P differentiates with the existing multi-model feature extraction algorithms in two pivotal aspects: 1) We consider brand's highly imbalanced class interest steering towards the KOLs and PCs over the irrelevant users; 2) We consider a common observation that a user is not exclusive to a single class (e.g. a KOL can also be a PC). Experiments on a real-world apparel brand key asset dataset validate the effectiveness of the proposed method. Yang Liu 0007, Tobey H. Ko, Zhonglei Gu |
SIGIR | 1 |
| 2018 | Multi-Modal Media Retrieval via Distance Metric Learning for Potential Customer DiscoveryabstractAs social media grown to become an integral part of many people's daily life, brands are quick to launch targeted social media marketing campaign to acquire new potential customers online. To facilitate the potential customer discovery process, a costly and labor intensive manual selection process is done to build a brand portfolio consisting of multimedia data relevant to the brand. To automate this process in a cost-effective way, in this paper, we propose a novel Multi-Modal Distance Metric Learning (M2DML) method, which learns a data-dependent similarity metric from multi-modal media data, aiming at assisting the brands to retrieve appropriate media data from social networks for potential customer discovery. To comprehensively model the supervised information of multi-modal data, M2DML aims to learn both the intra-modality and inter-modality distance metrics simultaneously. To further explore the unsupervised information of the dataset, M2DML aims to preserve the manifold structure of the multi-modal data. The proposed method is then formulated as a standard eigen-decomposition problem and the closed form solution is efficiently computed. Experiments on a standard multi-modal media dataset and a self-collected dataset validate the effectiveness of the proposed method. Yang Liu 0007, Zhonglei Gu, Tobey H. Ko, Jiming Liu 0001 |
WI | 1 |
| 2018 | Dimensionality Reduction in Multiple Ordinal RegressionabstractSupervised dimensionality reduction (DR) plays an important role in learning systems with high-dimensional data. It projects the data into a low-dimensional subspace and keeps the projected data distinguishable in different classes. In addition to preserving the discriminant information for binary or multiple classes, some real-world applications also require keeping the preference degrees of assigning the data to multiple aspects, e.g., to keep the different intensities for co-occurring facial expressions or the product ratings in different aspects. To address this issue, we propose a novel supervised DR method for DR in multiple ordinal regression (DRMOR), whose projected subspace preserves all the ordinal information in multiple aspects or labels. We formulate this problem as a joint optimization framework to simultaneously perform DR and ordinal regression. In contrast to most existing DR methods, which are conducted independently of the subsequent classification or ordinal regression, the proposed framework fully benefits from both of the procedures. We experimentally demonstrate that the proposed DRMOR method (DRMOR-M) well preserves the ordinal information from all the aspects or labels in the learned subspace. Moreover, DRMOR-M exhibits advantages compared with representative DR or ordinal regression algorithms on three standard data sets. Jiabei Zeng, Yang Liu 0007, Biao Leng, Zhang Xiong 0001, Yiu-Ming Cheung |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2017 | Multi-view Manifold Learning for Media Interestingness PredictionabstractMedia interestingness prediction plays an important role in many real-world applications and attracts much research attention recently. In this paper, we aim to investigate this problem from the perspective of supervised feature extraction. Specifically, we design a novel algorithm dubbed Multi-view Manifold Learning (M) to uncover the latent factors that are capable of distinguishing interesting media data from non-interesting ones. By modelling both geometry preserving criterion and discrimination maximization criterion in a unified framework, M2L learns a common subspace for data from multiple views. The analytical solution of M2L is obtained by solving a generalized eigen-decomposition problem. Experiments on the Predicting Media Interestingness Dataset validate the effectiveness of the proposed method. Yang Liu 0007, Zhonglei Gu, Yiu-Ming Cheung, Kien A. Hua |
ICMR | 1 |
| 2017 | Brand key asset discovery via cluster-wise biased discriminant projectionabstractAccurate and effective discovery of a brand's key assets, namely, Key Opinion Leaders (KOLs) and potential customers, plays an essential role in marketing campaigns. In a massive online social network, brands are challenged with identifying a small portion of key assets over an enormous volume of irrelevant users, making the problem a highly imbalanced one. Moreover, having to deal with social media data that are usually high-dimensional, the task of brand key asset discovery can be immensely expensive yet inaccurate if the information are not processed efficiently to extract representative features from the original space prior to the learning process. To address the above issues, we propose a novel method dubbed Cluster-wise Biased Discriminant Projection (CBDP) to uncover the compact and informative features from users' data for brand key asset discovery. CBDP conducts a two-layer learning procedure. In the first layer, a Discriminant Clustering (DC) scheme is developed to partition the original dataset into clusters with maximum discriminant capacity. In the second layer, a Biased Discriminant Projection (BDP) algorithm is proposed and performed on each cluster to map the high-dimensional data to the low-dimensional subspace, where the discriminant information of classes with high importance/preference is preserved. A unified mapping function of CBDP is finally established by integrating these two layers. Experiments on both synthetic examples and a real-world brand key asset dataset validate the effectiveness of the proposed method. Yang Liu 0007, Zhonglei Gu, Tobey H. Ko, Jiming Liu 0001 |
WI | 1 |
| 2017 | Joint sparse principal component analysis
Shuangyan Yi, Zhihui Lai 0001, Zhenyu He 0001, Yiu-Ming Cheung, Yang Liu 0007 |
Pattern Recognit. | 5 |
| 2017 | Implicit Visual Learning: Image Recognition via Dissipative Learning ModelabstractAccording to consciousness involvement, human’s learning can be roughly classified into explicit learning and implicit learning. Contrasting strongly to explicit learning with clear targets and rules, such as our school study of mathematics, learning is implicit when we acquire new information without intending to do so. Research from psychology indicates that implicit learning is ubiquitous in our daily life. Moreover, implicit learning plays an important role in human visual perception. But in the past 60 years, most of the well-known machine-learning models aimed to simulate explicit learning while the work of modeling implicit learning was relatively limited, especially for computer vision applications. This article proposes a novel unsupervised computational model for implicit visual learning by exploring dissipative system, which provides a unifying macroscopic theory to connect biology with physics. We test the proposed Dissipative Implicit Learning Model (DILM) on various datasets. The experiments show that DILM not only provides a good match to human behavior but also improves the explicit machine-learning performance obviously on image classification tasks. Yan Liu 0004, Yang Liu 0007, Shenghua Zhong, Songtao Wu |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2016 | Quality preserved data summarization for fast hierarchical clusteringabstractTraditional hierarchical clustering (HC) methods are not scalable with the size of databases. To address this issue, a series of summarization techniques, i.e. data bubbles (DB) and its improved versions, have been proposed to compress very large databases into representative seed points suitable for subsequent hierarchy construction. However, DB and its variants have two common drawbacks: (1) their performance is sensitive to the compression rate, and (2) their performance is sensitive to the initialization, i.e. the number and location of initialized seed points. This paper therefore proposes a new data summarization scheme, which is efficient and robust against the compression rate and initialization. In the proposed scheme, seed points are not only randomly initialized, but also trained to make them representative. After the training, a link strength network is constructed to achieve accurate hierarchy structure construction. Experiments demonstrate that the proposed method can produce high quality hierarchy structure with very high compression rate. Yiqun Zhang 0006, Yiu-Ming Cheung, Yang Liu 0007 |
IJCNN | 3 |
| 2016 | Learning Music Emotion Primitives via Supervised Dynamic ClusteringabstractThis paper explores a fundamental problem in music emotion analysis, i.e., how to segment the music sequence into a set of basic emotive units, which are named as emotion primitives. Current works on music emotion analysis are mainly based on the fixed-length music segments, which often leads to the difficulty of accurate emotion recognition. Short music segment, such as an individual music frame, may fail to evoke emotion response. Long music segment, such as an entire song, may convey various emotions over time. Moreover, the minimum length of music segment varies depending on the types of the emotions. To address these problems, we propose a novel method dubbed supervised dynamic clustering (SDC) to automatically decompose the music sequence into meaningful segments with various lengths. First, the music sequence is represented by a set of music frames. Then, the music frames are clustered according to the valence-arousal values in the emotion space. The clustering results are used to initialize the music segmentation. After that, a dynamic programming scheme is employed to jointly optimize the subsequent segmentation and grouping in the music feature space. Experimental results on standard dataset show both the effectiveness and the rationality of the proposed method. Yang Liu 0007, Yan Liu 0004, Gong Chen 0006 |
ACM Multimedia | 1 |
| 2016 | Locally learning heterogeneous manifolds for phonetic classificationabstractMost state-of-the-art phone classifiers use the same features and decision criteria for all phones, despite the fact that different broad classes are characterized by different manners and place of articulation that result in different acoustic features. This paper uses manifold learning to address structure in the acoustic space. Previous approaches to dimensionality reduction based on manifold learning assumed that the acoustic space can be characterized by a uniform manifold structure. In this paper we relax this assumption by learning different manifold structures for broad phonetic classes. Because all known classifiers make confusions between broad classes, we designed a two-level classifier in which the top level consists of a number of partially overlapping broad classes. Since the resulting classifiers are not statistically independent, we propose a new method for fusing the classifiers. Experimental results show that our two-level classifier obtained slightly better results when broad-class specific manifolds were learned, compared to a uniform manifold. However, the accuracy is still considerably lower than what could be obtained with oracle knowledge about broad class membership. From this we infer that phones do not form compact clusters in acoustic space. Heyun Huang, Yang Liu 0007, Louis ten Bosch, Bert Cranen, Lou Boves |
Comput. Speech Lang. | 2 |
| 2016 | Perception-oriented video saliency detection via spatio-temporal attention analysis
Shenghua Zhong, Yan Liu 0004, Vincent T. Y. Ng, Yang Liu 0007 |
Neurocomputing | 4 |
| 2015 | Face recognition from a single registered image for conference socializing
Yan Liu 0004, Yang Liu 0007, Shenghua Zhong, Kien A. Hua |
Expert Syst. Appl. | 3 |
| 2015 | Query-oriented unsupervised multi-document summarization via deep learning model
Shenghua Zhong, Yang Liu 0007, Bin Li 0003, Jing Long |
Expert Syst. Appl. | 2 |
| 2015 | Depth-aware salient object detection using anisotropic center-surround difference
Ran Ju, Yang Liu 0007, Tongwei Ren, Ling Ge, Gangshan Wu |
Signal Process. Image Commun. | 2 |
| 2015 | What Strikes the Strings of Your Heart? - Feature Mining for Music Emotion AnalysisabstractMusic can convey and evoke powerful emotions. This amazing ability has not only fascinated the general public but also attracted the researchers from different fields to discover the relationship between music and emotion. Psychologists have indicated that some specific characters of rhythm, harmony, and melody can evoke certain kinds of emotions. Those hypotheses are based on real life experience and proved by psychological paradigms on human beings. Aiming at the same target, this paper intends to design a systematic and quantitative framework, and answer three widely interested questions: 1) what are the intrinsic features embedded in music signal that essentially evoke human emotions; 2) to what extent these features influence human emotions; and 3) whether the findings from computational models are consistent with the existing research results from psychology. We formulate these tasks as a multi-label dimensionality reduction problem and propose an algorithm called multi-emotion similarity preserving embedding (ME-SPE). To adapt to the second-order music signals, we extend ME-SPE to its bilinear version. The proposed techniques show good performance in two standard music emotion datasets. Moreover, they demonstrate some interesting results for further research in this interdisciplinary topic. Yang Liu 0007, Yan Liu 0004, Kien A. Hua |
IEEE Trans. Affect. Comput. | 1 |
| 2014 | What Strikes the Strings of Your Heart?: Multi-Label Dimensionality Reduction for Music Emotion AnalysisabstractMusic can convey and evoke powerful emotions. This amazing ability has fascinated the general public and also attracted the researchers from different fields to discover the relationship between music and emotion. Psychologists have indicated that some specific characters of rhythm, harmony, melody, and also their combinations can evoke certain kinds of emotions. Their hypotheses are based on real life experience and proved by psychological paradigms on human beings. Aiming at the same target, this paper intends to design a systematic and quantitative framework, and answer three widely interested questions: 1) what are the intrinsic features embedded in music signal that essentially evoke human emotions; 2) to what extent these features influence human emotions; and 3) whether the findings from computational models are consistent with the existing research results from psychological experiments. We formulate the problem as a multi-label dimensionality reduction problem and provide the optimal solution. The proposed multi-emotion similarity preserving embedding technique not only shows better performance in two standard music emotion datasets but also demonstrates some interesting observations for further research in this interdisciplinary topic. Yang Liu 0007, Yan Liu 0004, Kien A. Hua |
ACM Multimedia | 1 |
| 2014 | Region level annotation by fuzzy based contextual cueing label propagation
Shenghua Zhong, Yan Liu 0004, Yang Liu 0007, Korris Fu-Lai Chung |
Multim. Tools Appl. | 3 |
| 2014 | Hybrid Manifold EmbeddingabstractIn this brief, we present a novel supervised manifold learning framework dubbed hybrid manifold embedding (HyME). Unlike most of the existing supervised manifold learning algorithms that give linear explicit mapping functions, the HyME aims to provide a more general nonlinear explicit mapping function by performing a two-layer learning procedure. In the first layer, a new clustering strategy called geodesic clustering is proposed to divide the original data set into several subsets with minimum nonlinearity. In the second layer, a supervised dimensionality reduction scheme called locally conjugate discriminant projection is performed on each subset for maximizing the discriminant information and minimizing the dimension redundancy simultaneously in the reduced low-dimensional space. By integrating these two layers in a unified mapping function, a supervised manifold embedding framework is established to describe both global and local manifold structure as well as to preserve the discriminative ability in the learned subspace. Experiments on various data sets validate the effectiveness of the proposed method. Yang Liu 0007, Yan Liu 0004, Keith C. C. Chan, Kien A. Hua |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2013 | Detecting and understanding genetic and structural features in HIV-1 B subtype V3 underlying HIV-1 co-receptor usageabstractMOTIVATION: To define V3 genetic elements and structural features underlying different HIV-1 co-receptor usage in vivo. RESULTS: By probabilistically modeling mutations in the viruses isolated from HIV-1 B subtype patients, we present a unique statistical procedure that would first identify V3 determinants associated with the usage of different co-receptors cooperatively or independently, and then delineate the complicated interactions among mutations functioning cooperatively. We built a model based on dual usage of CXCR4 and CCR5 co-receptors. The molecular basis of our statistical predictions is further confirmed by phenotypic and molecular modeling analyses. Our results provide new insights on molecular basis of different HIV-1 co-receptor usage. This is critical to optimize the use of genotypic tropism testing in clinical practice and to obtain molecular-implication for design of vaccine and new entry-inhibitors. Mengjie Chen, Valentina Svicher, Anna Artese, Giosuè Costa, Claudia Alteri, Francesco Ortuso, Lucia Parrotta, Yang Liu 0007, Carlo-Federico Perno, Stefano Alcaro, Jing Zhang 0010 |
Bioinform. | 8 |
| 2013 | Corrigendum to cross-domain video concept detection: A joint discriminative and generative active learning approach [Expert Systems with Applications 39 (15) (2012) 12220-12228]
Yang Liu 0007, Alex Hauptmann 0001, Zhang Xiong 0001 |
Expert Syst. Appl. | 3 |
| 2013 | Water Reflection Recognition Based on Motion Blur Invariant Moments in Curvelet SpaceabstractWater reflection, a typical imperfect reflection symmetry problem, plays an important role in image content analysis. Existing techniques of symmetry recognition, however, cannot recognize water reflection images correctly because of the complex and various distortions caused by the water wave. Hence, we propose a novel water reflection recognition technique to solve the problem. First, we construct a novel feature space composed of motion blur invariant moments in low-frequency curvelet space and of curvelet coefficients in high-frequency curvelet space. Second, we propose an efficient algorithm including two sub-algorithms: low-frequency reflection cost minimization and high-frequency curvelet coefficients discrimination to classify water reflection images and to determine the reflection axis. Through experimenting on authentic images in a series of tasks, the proposed techniques prove effective and reliable in classifying water reflection images and detecting the reflection axis, as well as in retrieving images with water reflection. Shenghua Zhong, Yan Liu 0004, Yang Liu 0007 |
IEEE Trans. Image Process. | 3 |
| 2012 | Smarter wheelchairs who can talk to each other: An integrated and collaborative approachabstractPervasive computing technologies can benefit the injured, disabled or elderly people in their daily lives, and smart wheelchair has been a representative of this kind of technologies. However, most existing smart wheelchairs have limitations on extensibility and flexibility of building new functionalities. One big reason is they are just stand-alone ones considering other wheelchairs and surrounding things as dummy objects. To address this issue, we proposed an integrated and collaborative approach: Smarter Wheelchairs that can “talk” to each other, and even “talk” to other things in the surrounding. Smarter Wheelchairs take advantages of smart objects deployed in living environment or hospital environment. Smarter Wheelchairs can harness functions provided by smart objects, therefore, their functionality can be flexibly extended. We have implemented and evaluated a prototype system of Smarter Wheelchairs to demonstrate the feasibility and efficiency of our approach. Junjun Kong, Jiannong Cao 0001, Yang Liu 0007, Yao Guo 0001, Weizhong Shao |
Healthcom | 3 |
| 2012 | Knowledge-based Quadratic Discriminant Analysis for phonetic classificationabstractModeling the second-order statistics of articulatory trajectories is likely to improve the performance in classifying phone segments compared to using only linear combinations of MFCCs. Nevertheless, the extremely high dimensionality of the feature space spanned by a combination of monomials of degree-1 and degree-2 makes it difficult to effectively exploit the discriminative information in the full covariance matrix. This paper proposes a novel algorithm, dubbed Knowledge-based Quadratic Discriminant Analysis (KnQDA), for reducing the number of dimensions of the space spanned by degree-1 and degree-2 monomials by using phonetic knowledge for selecting the set of degree-2 monomials that are most likely to improve classification. KnQDA seeks a trade-off between overfitting and undertraining, which further improves the learnability. Binary classifications on all pairs of phones in TIMIT show the effectiveness of the proposed method, especially on those phone pairs that overlap strongly in the linear feature space. Heyun Huang, Yang Liu 0007, Louis ten Bosch, Bert Cranen, Lou Boves |
ICASSP | 2 |
| 2012 | The Extended Co-learning Framework for Robust Object TrackingabstractRecently, object tracking has been widely studied as a binary classification problem. Semi-supervised learning is particularly suitable for improving classification accuracy when large quantities of unlabeled samples are generated (just like tracking procedure). The purpose of this paper is to fulfill robust and stable tracking by using collaborative learning, which belongs to the scope of semi-supervised learning, among three classifiers. Different from [1], random fern classifier is incorporated to deal with 2bitBP feature newly added and certain constraints are specially implemented in our framework. Besides, the way for selecting positive samples is also altered by us in order to achieve more stable tracking. Algorithm proposed in this paper is validated by tracking pedestrian and cup under occlusion. Experiments and comparison show that our algorithm can avoid drifting problem to some degree and make tracking result more robust and adaptive. Chen Gong 0002, Yang Liu 0007, Tianyu Li 0003, Jie Yang 0002, Xiangjian He |
ICME | 2 |
| 2012 | Cross-domain video concept detection: A joint discriminative and generative active learning approach
Yang Liu 0007, Alex Hauptmann 0001, Zhang Xiong 0001 |
Expert Syst. Appl. | 3 |
| 2012 | Tensor distance based multilinear globality preserving embedding: A unified tensor based dimensionality reduction framework for image and video classification
Yang Liu 0007, Yan Liu 0004, Shenghua Zhong, Keith C. C. Chan |
Expert Syst. Appl. | 1 |
| 2011 | Ordinal Regression via Manifold LearningabstractOrdinal regression is an important research topic in machine learning. It aims to automatically determine the implied rating of a data item on a fixed, discrete rating scale. In this paper, we present a novel ordinal regression approach via manifold learning, which is capable of uncovering the embedded nonlinear structure of the data set according to the observations in the highdimensional feature space. By optimizing the order information of the observations and preserving the intrinsic geometry of the data set simultaneously, the proposed algorithm provides the faithful ordinal regression to the new coming data points. To offer more general solution to the data with natural tensor structure, we further introduce the multilinear extension of the proposed algorithm, which can support the ordinal regression of high order data like images. Experiments on various data sets validate the effectiveness of the proposed algorithm as well as its extension. Yang Liu 0007, Yan Liu 0004, Keith C. C. Chan |
AAAI | 1 |
| 2011 | A ubiquitous wireless video surveillance system based on pub/subabstractIn most of the existing video surveillance systems, captured video streams by CCTV or video cameras are aggregated through cables to a central station and monitored by associated human operators. However, the high deployment cost, low scalability and reuseability still remain to be the main roadblock for their wider application. Using human to monitor events can also be unreliable. Accordingly, we designed a ubiquitous wireless video surveillance system. This system uses wireless sensor nodes to detect pre-defined events and utilizes wireless mesh network to transmit high-quality video streams. Without cables, the deployment cost is significantly decreased. Different application users can also define various events and automatically receive notification and corresponding video streams when the events occur. In addition, mobile users, even when roaming, can access this system. In this demo, we introduce the hardware and software of this system and its implementation in our intelligent transportation system testbed. Jiannong Cao 0001, Xuefeng Liu 0001, Steven Lai, Yang Zou 0002, Jun Zhang 0019, Yang Liu 0007, Chisheng Zhang |
UbiComp | 6 |
| 2011 | Globality-Locality Consistent Discriminant Analysis for Phone ClassificationabstractContains fulltext : 94442.pdf (author's version ) (Open Access) Heyun Huang, Yang Liu 0007, Jort F. Gemmeke, Louis ten Bosch, Bert Cranen, Lou Boves |
INTERSPEECH | 2 |
| 2011 | Semi-supervised manifold ordinal regression for image rankingabstractIn this paper, we present a novel algorithm called manifold ordinal regression (MOR) for image ranking. By modeling the manifold information in the objective function, MOR is capable of uncovering the intrinsically nonlinear structure held by the image data sets. By optimizing the ranking information of the training data sets, the proposed algorithm provides faithful rating to the new coming images. To offer more general solution for the real-word tasks, we further provide the semi-supervised manifold ordinal regression (SS-MOR). Experiments on various data sets validate the effectiveness of the proposed algorithms. Yang Liu 0007, Yan Liu 0004, Shenghua Zhong, Keith C. C. Chan |
ACM Multimedia | 1 |
| 2011 | Bilinear deep learning for image classificationabstractImage classification is a well-known classical problem in multimedia content analysis. This paper proposes a novel deep learning model called bilinear deep belief network (BDBN) for image classification. Unlike previous image classification models, BDBN aims to provide human-like judgment by referencing the architecture of the human visual system and the procedure of intelligent perception. Therefore, the multi-layer structure of the cortex and the propagation of information in the visual areas of the brain are realized faithfully. Unlike most existing deep models, BDBN utilizes a bilinear discriminant strategy to simulate the "initial guess" in human object recognition, and at the same time to avoid falling into a bad local optimum. To preserve the natural tensor structure of the image data, a novel deep architecture with greedy layer-wise reconstruction and global fine-tuning is proposed. To adapt real-world image classification tasks, we develop BDBN under a semi-supervised learning framework, which makes the deep model work well when labeled images are insufficient. Comparative experiments on three standard datasets show that the proposed algorithm outperforms both representative classification models and existing deep learning techniques. More interestingly, our demonstrations show that the proposed BDBN works consistently with the visual perception of humans. Shenghua Zhong, Yan Liu 0004, Yang Liu 0007 |
ACM Multimedia | 3 |
| 2011 | Bilinear deep learning for image classificationabstractNo abstract available. Shenghua Zhong, Yan Liu 0004, Yang Liu 0007 |
ACM Multimedia | 3 |
| 2011 | Dual-Mote: A Sensor Network testbed for high rate sensing-transmission and runtime evaluationabstractMost researchers encountered the following two problems when working with real Wireless Sensor Networks (WSNs): (1) Sensor nodes cannot satisfy application requirements even though the nominal sensing/transmission rates of these nodes are much higher than required. (2) In a WSN deployed in a large area, it is difficult or infeasible to get runtime performance evaluation of sensor nodes. We found out that the root reason of these two problems is resource competition, in which an operation has to wait for the resources being used by other operations. Therefore, we propose a dual-mote testbed, which is able to avoid both hardware competition on a node and wireless channel competition in a network. We have implemented the hardware, supporting protocols and tools of the testbed. The experimental results show that, compared to a general WSN, the improvement on performances such as throughput and response speed by our testbed are more than doubled. Hejun Wu, Jiannong Cao 0001, Xuefeng Liu 0001, Yang Liu 0007 |
WCNC | 4 |
| 2011 | Tensor-based locally maximum margin classifier for image and video classification
Yang Liu 0007, Yan Liu 0004, Keith C. C. Chan |
Comput. Vis. Image Underst. | 1 |
| 2010 | Multilinear Maximum Distance Embedding Via L1-Norm OptimizationabstractDimensionality reduction plays an important role in many machine learning and pattern recognition tasks. In this paper, we present a novel dimensionality reduction algorithm called multilinear maximum distance embedding (M2DE), which includes three key components. To preserve the local geometry and discriminant information in the embedded space, M2DE utilizes a new objective function, which aims to maximize the distances between some particular pairs of data points, such as the distances between nearby points and the distances between data points from different classes. To make the mapping of new data points straightforward, and more importantly, to keep the natural tensor structure of high-order data, M2DE integrates multilinear techniques to learn the transformation matrices sequentially. To provide reasonable and stable embedding results, M2DE employs the L1-norm, which is more robust to outliers, to measure the dissimilarity between data points. Experiments on various datasets demonstrate that M2DE achieves good embedding results of high-order data for classification tasks. Yang Liu 0007, Yan Liu 0004, Keith C. C. Chan |
AAAI | 1 |
| 2010 | A semantic no-reference image sharpness metric based on top-down and bottom-up saliency map modelingabstractThis work presents a semantic level no-reference image sharpness/blurriness metric under the guidance of top-down & bottom-up saliency map, which is learned based on eye-tracking data by SVM. Unlike existing metrics focused on measuring the blurriness in vision level, our metric more concerns about the image content and human's intention. We integrate visual features, center priority, and semantic meaning from tag information to learn a top-down & bottom-up saliency model based on the eye-tracking data. Empirical validations on standard dataset demonstrate the effectiveness of the proposed model and metric. Shenghua Zhong, Yan Liu 0004, Yang Liu 0007, Korris Fu-Lai Chung |
ICIP | 3 |
| 2010 | Supervised manifold learning for image and video classificationabstractThis paper presents a supervised manifold learning model for dimensionality reduction in image and video classification tasks. Unlike most manifold learning models that emphasize the distance preserving, we propose a novel algorithm called maximum distance embedding (MDE), which aims to maximize the distances between some particular pairs of data points, with the intention of flattening the local nonlinearity and keeping the discriminant information simultaneously in the embedded feature space. Moreover, MDE measures the dissimilarity between data points using L1-norm distance, which is more robust to outliers than widely used Frobenius norm distance. To adapt the nature tensor structure of image and video data, we further propose the multilinear MDE (M2DE). Experiments on various datasets demonstrate that both MDE and M2DE achieve impressive embedding results of image and video data for classification tasks. Yang Liu 0007, Yan Liu 0004, Keith C. C. Chan |
ACM Multimedia | 1 |
| 2010 | Unsupervised summarization of rushes videosabstractThis paper proposes a new framework to formulate summarization of rushes video as an unsupervised learning problem. We pose the problem of video summarization as one of time-series clustering, and proposed Constrained Aligned Cluster Analysis (CACA). CACA combines kernel k-means, Dynamic Time Alignment Kernel (DTAK), and unlike previous work, CACA jointly optimizes video segmentation and shot clustering. CACA is effciently solved via dynamic programming. Experimental results on the TRECVID 2007 and 2008 BBC rushes video summarization databases validate the accuracy and effectiveness of CACA. Yang Liu 0007, Feng Zhou 0002, Wei Liu 0220, Fernando De la Torre, Yan Liu 0004 |
ACM Multimedia | 1 |
| 2010 | Nonlinear dimensionality reduction with hybrid distance for trajectory representation of dynamic texture
Yang Liu 0007, Yan Liu 0004, Keith C. C. Chan |
Signal Process. | 1 |
| 2010 | Tensor Distance Based Multilinear Locality-Preserved Maximum Information EmbeddingabstractThis brief paper presents a unified framework for tensor-based dimensionality reduction (DR) with a new tensor distance (TD) metric and a novel multilinear locality-preserved maximum information embedding (MLPMIE) algorithm. Different from traditional Euclidean distance, which is constrained by the orthogonality assumption, TD measures the distance between data points by considering the relationships among different coordinates. To preserve the natural tensor structure in low-dimensional space, MLPMIE directly works on the high-order form of input data and iteratively learns the transformation matrices. In order to preserve the local geometry and to maximize the global discrimination simultaneously, MLPMIE keeps both local and global structures in a manifold model. By integrating TD into tensor embedding, TD-MLPMIE performs tensor-based DR through the whole learning procedure, and achieves stable performance improvement on various standard datasets. Yang Liu 0007, Yan Liu 0004, Keith C. C. Chan |
IEEE Trans. Neural Networks | 1 |
| 2009 | Tensor distance based multilinear multidimensional scaling for image and video analysisabstractThis paper presents a novel dimensionality reduction technique named Tensor Distance based Multilinear Multidimensional Scaling (TD-MMDS). First, we propose a new distance metric called Tensor Distance (TD) to build a relationship graph of data points with high-order. Then we employ an iterative strategy to sequentially learn the transformation matrices that can best keep pair-wise TDs of the high-order data in the low-dimensional embedded space. By integrating both tensor distance and tensor embedding, TD-MMDS provides a uniform framework of tensor based dimensionality reduction, which preserves the intrinsic structure of high-order data through the whole learning procedure. Experiments on standard image and video datasets validate the effectiveness of the proposed TD-MMDS. Yang Liu 0007, Yan Liu 0004 |
ACM Multimedia | 1 |
| 2009 | Dimensionality reduction for heterogeneous dataset in rushes editing
Yang Liu 0007, Yan Liu 0004, Keith C. C. Chan |
Pattern Recognit. | 1 |
| 2008 | Multiple video trajectories representation using double-layer isometric feature mappingabstractThis paper proposes a novel non-linear dimensionality reduction algorithm, named double-layer isometric feature mapping (DLIso), which generates the trajectories for the video sequence containing different kinds of video clips. First, a nearest neighbor based clustering algorithm is utilized to partition the video sequence into a set of data blocks. Second, intra-cluster graphs are constructed based on the individual character of each data block to build the basic layer for DLIso. Third, the inter-cluster graph is constructed by analyzing the interrelation among these isolated data blocks to build the hyper-layer. Finally, all data points are mapped onto a unique low-dimensional feature space while preserving the corresponding relations in the double layers. Experiments on synthetic datasets as well as the real video sequences demonstrate that the low-dimensional trajectories generated by the proposed method correctly represent the semantic information of the data. Yang Liu 0007, Yan Liu 0004, Keith C. C. Chan |
ICME | 1 |