EDBT 2026 Demo / reviewers in the wild / expert
Haoliang Sun
dblp:117/5673
· DBLP profile ↗
34ranked-venue papers
7as first author
28since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 28 · 6 first-author · 25 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 3 first-author · 10 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MTRL-CG: Multi-Task Reinforcement Learning Method with Spectral Clustering-Based Task GroupingabstractMulti-task reinforcement learning (RL) aims to enhance agent performance across multiple tasks by enabling effective knowledge transfer. However, these methods adopt a fully shared policy across all tasks without explicitly distinguishing between related and conflicting ones, making them suffer from negative interference issue, where updates beneficial to one task adversely affect others and lead to degraded overall performance. In this paper, we propose a multi-task reinforcement learning method with spectral clustering-based task grouping (MTRL-CG), which leverages spectral clustering to group related tasks and separate conflicting ones, enabling group-wise policy learning to mitigate negative interference. We first quantify inter-task affinity by measuring the influence of task-specific updates on others within a shared model, and construct an affinity matrix to capture these relationships. Spectral clustering is then applied to partition tasks via spectral embedding and k-means clustering. Each task group is trained with a dedicated policy network to promote focused learning. Built upon the Soft Actor-Critic (SAC) algorithm, MTRL-CG can be readily integrated into existing SAC-based multi-task RL methods. Extensive experiments on the Meta-World benchmark demonstrate the effectiveness of the proposed MTRL-CG method. Wenjia Meng, Haoliang Sun, Yilong Yin |
AAAI | 3 |
| 2026 | Exploiting dynamic spatio-temporal correlations for origin-destination demand predictionabstractAccurate Origin-Destination (OD) demand prediction is fundamental to intelligent transportation systems (ITS), enabling real-time traffic management, dynamic vehicle dispatch, and efficient resource allocation in urban environments. However, OD demand exhibits complex, dynamic, and highly coupled spatio-temporal patterns that remain challenging for existing models. We propose a novel Dynamic Spatio-Temporal Correlation Network (DSTCN) for OD demand forecasting. DSTCN features three key components: (1) a bidirectional demand trend modeling module (Glstm2D) that jointly learns demand evolution from both origin and destination perspectives; (2) a Transformer-based spatial similarity module (Simformer) to dynamically extract and integrate inter-regional correlations across the OD matrix; and (3) a temporal fusion and modeling module (FF-TM) that combines and processes multi-source spatio-temporal features for next-step prediction. Extensive experiments on three large-scale real-world datasets (NYC-TOD2018, NYC-TOD2019, and HZMetro) demonstrate that DSTCN consistently outperforms state-of-the-art baselines across diverse urban scenarios. Yongshun Gong, Piao Yu, Xu Zhang 0039, Xinxin Zhang 0004, Xiushan Nie, Haoliang Sun |
Expert Syst. Appl. | 6 |
| 2026 | MSFI: Multi-timescale spatio-temporal features integration in spiking neural networks
Dengfeng Xue, Chunfeng Yuan, Man Yao, Wei Liu 0153, Li Yang 0014, Bing Li 0001, Weiming Hu 0004, Haoliang Sun, Zhetao Li |
Neural Networks | 10 |
| 2026 | CausalCOMRL: Context-based offline meta-reinforcement learning with causal representationabstractContext-based offline meta-reinforcement learning (OMRL) methods have achieved appealing success by leveragingpre-collected offline datasets to develop task representations that guide policy learning. However, current context-based OMRL methods often introduce spurious correlations, where task components are incorrectly correlated due to confounders. These correlations can degrade policy performance when the confounders in the test taskdiffer from those in the training task. To address this problem, we propose CausalCOMRL, a context-based OMRL method that integrates causal representation learning. This approach uncovers causal relationships among the task components and incorporates the causal relationships into task representations, enhancing the generalizability of RL agents. We further improve the distinction of task representations from different tasks by using mutual information optimization and contrastive learning. Utilizing these causal task representations, we employSAC to optimize policies on meta-RL benchmarks. Experimental results show that CausalCOMRL achieves better performance than other methods on most benchmarks. Zhengzhe Zhang, Wenjia Meng, Haoliang Sun, Gang Pan 0001 |
Neural Networks | 3 |
| 2025 | Towards Macro-AUC Oriented Imbalanced Multi-Label Continual LearningabstractIn Continual Learning (CL), while existing work primarily focuses on the multi-class classification task, there has been limited research on Multi-Label Learning (MLL). In practice, MLL datasets are often class-imbalanced, making it inherently challenging, a problem that is even more acute in CL. Due to its sensitivity to imbalance, Macro-AUC is an appropriate and widely used measure in MLL. However, there is no research to optimize Macro-AUC in MLCL specifically. To fill this gap, in this paper, we propose a new memory replay-based method to tackle the imbalance issue for Macro-AUC-oriented MLCL. Specifically, inspired by recent theory work, we propose a new Reweighted Label-Distribution-Aware Margin (RLDAM) loss. Furthermore, to be compatible with the RLDAM loss, a new memory-updating strategy named Weight Retain Updating (WRU) is proposed to maintain the numbers of positive and negative instances of the original dataset in memory. Theoretically, we provide superior generalization analyses of the RLDAM-based algorithm in terms of Macro-AUC, separately in batch MLL and MLCL settings. This is the first work to offer theoretical generalization analyses in MLCL to our knowledge. Finally, a series of experimental results illustrate the effectiveness of our method over several baselines. Yan Zhang 0145, Guoqiang Wu, Bingzheng Wang, Teng Pang, Haoliang Sun, Yilong Yin |
AAAI | 5 |
| 2025 | SeqMvRL: A Sequential Fusion Framework for Multi-view Representation LearningabstractMulti-view representation learning integrates multiple observable views of an entity into a unified representation to facilitate downstream tasks. Current methods predominantly focus on distinguishing compatible components across views, followed by a single-step parallel fusion process. However, this parallel fusion is static in essence, overlooking potential conflicts among views and compromising representation ability. To address this issue, this paper proposes a novel Sequential fusion framework for Multi-view Representation Learning, termed SeqMvRL. Specifically, we model multi-view fusion as a sequential decision-making problem and construct a pairwise integrator (PI) and a next-view selector (NVS), which represent the environment and agent in reinforcement learning, respectively. PI merges the current fused feature with the selected view, while NVS is introduced to determine which view to fuse subsequently. By adaptively selecting the next optimal view for fusion based on the current fusion state, SeqMvRL thereby effectively reduces conflicts and enhances unified representation quality. Additionally, an elaborate novel reward function encourages the model to prioritize views that enhance the discriminability of the fused features. Experimental results demonstrate that SeqMvRL outperforms parallel fusion schemes in classification and clustering tasks. Ren Wang 0011, Haoliang Sun, Yuxiu Lin, Chuanhui Zuo, Yongshun Gong, Yilong Yin, Wenjia Meng |
CVPR | 2 |
| 2025 | Improving Generalization in Meta-Learning via Meta-Gradient AugmentationabstractMeta-learning methods typically follow a two-loop framework, where each loop potentially suffers from notorious overfitting, hindering rapid adaptation and generalization to new tasks. Existing methods address this by enhancing the mutual-exclusivity or diversity of training samples, but these data manipulation strategies are data-dependent and insufficiently flexible. This work proposes a data-independent Meta-Gradient Augmentation (MGAug) method from the perspective of gradient regularization. The key idea is first to break the rote memories by network pruning to address memorization overfitting in the inner loop, then use the gradients of pruned sub-networks to augment meta-gradients, alleviating overfitting in the outer loop. Specifically, we explore three pruning strategies, including random width pruning, random parameter pruning, and a newly proposed catfish pruning that measures a Meta-Memorization Carrying Amount (MMCA) score for each parameter and prunes high-score ones to break rote memories. The proposed MGAug is theoretically guaranteed by the generalization bound from the PAC-Bayes framework. Extensive experiments on multiple few-shot learning benchmarks validate MGAug's effectiveness and significant improvement over various meta-baselines. Ren Wang 0011, Haoliang Sun, Yuxiu Lin, Xinxin Zhang 0004, Yilong Yin |
IJCAI | 2 |
| 2025 | Spatio-temporal Prototype-based Hierarchical Learning for OD Demand PredictionabstractOrigin-Destination (OD) demand prediction is a pivotal yet highly challenging task in intelligent transportation systems, aiming to accurately forecast cross-region ridership flows within urban networks. While previous studies have focused on modeling node-to-node relationships, most of them neglect the fact that nodes (regions/stations) exhibit similar spatio-temporal (ST) patterns, which are termed as spatio-temporal prototypes. Capturing these prototypes is crucial for understanding the unified ST dependencies across the network. To bridge this gap, we propose STPro, an ST prototype-based hierarchical model with a dual-branch structure that extracts ST features from the micro and macro perspectives. At the micro level, our model learns unified ST features of individual nodes, while at the macro level, it employs dynamic clustering to identify city-wide ST prototypes, thereby uncovering latent patterns of urban mobility. Besides, we leverage different roles of nodes as origins and destinations by constructing dual O and D branches and learn the mutual information to model their intricate interactions and correlations. Extensive experiments on two public datasets demonstrate that our STPro outperforms recent state-of-the-art baselines, achieving remarkable predictive improvements in OD demand prediction. Shilu Yuan, Wenqian Mu, Ji Zhong, Meng Chen 0003, Haoliang Sun, Yongshun Gong |
IJCAI | 6 |
| 2025 | Cross-graph meta matching correction for noisy graph matching
Fangkai Li, Feiyu Pan, Wenjia Meng, Haoliang Sun, Xiushan Nie, Yilong Yin, Xiankai Lu |
Comput. Vis. Image Underst. | 4 |
| 2025 | Enhancing origin-destination flow prediction via bi-directional spatio-temporal inference and interconnected feature evolution
Piao Yu, Xu Zhang 0039, Yongshun Gong, Jian Zhang 0002, Haoliang Sun, Junjie Zhang 0002, Xinxin Zhang 0004, Yilong Yin |
Expert Syst. Appl. | 5 |
| 2025 | Variational Rectification Inference for Learning with Noisy Labels
Haoliang Sun, Qi Wei 0004, Lei Feng 0006, Yupeng Hu 0003, Fan Liu 0008, Hehe Fan, Yilong Yin |
Int. J. Comput. Vis. | 1 |
| 2025 | Correction: Variational Rectification Inference for Learning with Noisy Labels
Haoliang Sun, Qi Wei 0004, Lei Feng 0006, Yupeng Hu 0003, Fan Liu 0008, Hehe Fan, Yilong Yin |
Int. J. Comput. Vis. | 1 |
| 2025 | GeM: Gaussian embeddings with Multi-hop graph transfer for next POI recommendation
Wenqian Mu, Jiyuan Liu 0013, Yongshun Gong, Ji Zhong, Wei Liu 0007, Haoliang Sun, Xiushan Nie, Yilong Yin, Yu Zheng 0004 |
Neural Networks | 6 |
| 2025 | Diverse Teacher-Students for deep safe semi-supervised learning under class mismatch
Qikai Wang, Rundong He, Yongshun Gong, Chunxiao Ren, Haoliang Sun, Xiaoshui Huang, Yilong Yin |
Neural Networks | 5 |
| 2025 | Spatio-Temporal Multivariate Probabilistic Modeling for Traffic PredictionabstractTraffic prediction is an essential task in intelligent transportation systems dealing with complex and dynamic spatio-temporal correlations. To date, most work is focused on point estimation models, which only output a single value w.r.t an attribute of traffic data at a time, falling short of depicting diverse situations and uncertainty in future. Besides, most methods are not flexible enough to handle real complex traffic scenarios, involving missing values and non-uniformly sampled data. The interactions among different attributes of traffic data are also rarely explored explicitly. In this paper, we focus on probabilistic estimation in traffic prediction tasks, proposing a spatio-temporal multivariate probabilistic predictive model to estimate the distributions of traffic data. Specifically, we devise a multivariate spatio-temporal fusion graph block to extract spatio-temporal correlations of multiple traffic attributes at different locations. A multi-graph fusion module is designed to capture time-varying spatial relationships. We estimate the joint distributions of missing traffic data using copulas. The proposed model can simultaneously perform traffic forecasting and interpolation tasks with non-uniformly sampled data. Our experiments on two real-world traffic datasets demonstrate the advantages of our model over the state-of-the-art1. Zhibin Li 0002, Wei Liu 0007, Xinghao Yang, Haoliang Sun, Meng Chen 0003, Yu Zheng 0004, Yongshun Gong |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2024 | Spatio-temporal Graph Normalizing Flow for Probabilistic Traffic PredictionabstractWith the development of the Intelligent Transportation Systems, a great deal of work has been proposed to tackle traffic prediction tasks. Despite their good performance, most traffic prediction models are point estimation models, lacking the capability to estimate the uncertainties of future traffic data, which is crucial in practical traffic decision-making. Aiming at this problem, we combine the probabilistic estimation capabilities of conditional normalizing flows with the spatio-temporal relationship learning of spatio-temporal graphs, leading to a Spatio-Temporal Graph Normalizing Flow (STGNF) model to estimate the distribution of future traffic data. We are the first to employ the conditional normalizing flows as the backbone for probabilistic traffic prediction. Then we design a spatio-temporal graph conditional fusion network to learn the spatio-temporal relationships between future and historical traffic data, which are provided to the conditional normalizing flows as conditional information. Extensive experiments on two real-world traffic datasets demonstrate that our proposed model significantly outperforms the state-of-the-art baselines. Zhibin Li 0002, Wei Liu 0007, Haoliang Sun, Meng Chen 0003, Wenpeng Lu, Yongshun Gong |
CIKM | 4 |
| 2024 | Learning sample-aware threshold for semi-supervised learning
Qi Wei 0004, Lei Feng 0006, Haoliang Sun, Ren Wang 0011, Rundong He, Yilong Yin |
Mach. Learn. | 3 |
| 2024 | Correction: Learning sample-aware threshold for semi-supervised learning
Qi Wei 0004, Lei Feng 0006, Haoliang Sun, Ren Wang 0011, Rundong He, Yilong Yin |
Mach. Learn. | 3 |
| 2024 | MetaKernel: Learning Variational Random Features With Limited LabelsabstractFew-shot learning deals with the fundamental and challenging problem of learning from a few annotated samples, while being able to generalize well on new tasks. The crux of few-shot learning is to extract prior knowledge from related tasks to enable fast adaptation to a new task with a limited amount of data. In this paper, we propose meta-learning kernels with random Fourier features for few-shot learning, we call MetaKernel. Specifically, we propose learning variational random features in a data-driven manner to obtain task-specific kernels by leveraging the shared knowledge provided by related tasks in a meta-learning setting. We treat the random feature basis as the latent variable, which is estimated by variational inference. The shared knowledge from related tasks is incorporated into a context inference of the posterior, which we achieve via a long-short term memory module. To establish more expressive kernels, we deploy conditional normalizing flows based on coupling layers to achieve a richer posterior distribution over random Fourier bases. The resultant kernels are more informative and discriminative, which further improves the few-shot learning. To evaluate our method, we conduct extensive experiments on both few-shot image classification and regression tasks. A thorough ablation study demonstrates that the effectiveness of each introduced component in our method. The benchmark results on fourteen datasets demonstrate MetaKernel consistently delivers at least comparable and often better performance than state-of-the-art alternatives. Yingjun Du, Haoliang Sun, Xiantong Zhen, Jun Xu 0019, Yilong Yin, Ling Shao 0001, Cees Snoek |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | MetaViewer: Towards A Unified Multi-View RepresentationabstractExisting multi-view representation learning methods typically follow a specific-to-uniform pipeline, extracting latent features from each view and then fusing or aligning them to obtain the unified object representation. However, the manually pre-specified fusion functions and aligning criteria could potentially degrade the quality of the derived representation. To overcome them, we propose a novel uniform-to-specific multi-view learning framework from a meta-learning perspective, where the unified representation no longer involves manual manipulation but is automatically derived from a meta-learner named MetaViewer. Specifically, we formulated the extraction and fusion of view-specific latent features as a nested optimization problem and solved it by using a bi-level optimization scheme. In this way, MetaViewer automatically fuses view-specific features into a unified one and learns the optimal fusion scheme by observing reconstruction processes from the unified to the specific over all views. Extensive experimental results in downstream classification and clustering tasks demonstrate the efficiency and effectiveness of the proposed method. Ren Wang 0011, Haoliang Sun, Yuling Ma, Xiaoming Xi, Yilong Yin |
CVPR | 2 |
| 2023 | Fine-Grained Classification with Noisy LabelsabstractLearning with noisy labels (LNL) aims to ensure model generalization given a label-corrupted training set. In this work, we investigate a rarely studied scenario of LNL on fine-grained datasets (LNL-FG), which is more practical and challenging as large inter-class ambiguities among fine-grained classes cause more noisy labels. We empirically show that existing methods that work well for LNL fail to achieve satisfying performance for LNL-FG, arising the practical need of effective solutions for LNL-FG. To this end, we propose a novel framework called stochastic noise-tolerated supervised contrastive learning (SNSCL) that confronts label noise by encouraging distinguishable representation. Specifically, we design a noise-tolerated supervised contrastive learning loss that incorporates a weight-aware mechanism for noisy label correction and selectively updating momentum queue lists. By this mechanism, we mitigate the effects of noisy anchors and avoid inserting noisy labels into the momentum-updated queue. Besides, to avoid manually-defined augmentation strategies in contrastive learning, we propose an efficient stochastic module that samples feature embeddings from a generated distribution, which can also enhance the representation ability of deep models. SNSCL is general and compatible with prevailing robust LNL strategies to improve their performance for LNL-FG. Extensive experiments demonstrate the effectiveness of SNSCL. Qi Wei 0004, Lei Feng 0006, Haoliang Sun, Ren Wang 0011, Chenhui Guo, Yilong Yin |
CVPR | 3 |
| 2023 | Multi-View Representation Learning via View-Aware ModulationabstractMulti-view (representation) learning derives an entity's representation from its multiple observable views to facilitate various downstream tasks. The most challenging topic is how to model unobserved entities and their relationships to specific views. To this end, this work proposes a novel multi-view learning method using a View-Aware parameter Modulation mechanism, termed VAM. The key idea is to use trainable parameters as proxies for unobserved entities and views, such that modeling entity-view relationships is converted into modeling the relationship between proxy parameters. Specifically, we first build a set of trainable parameters to learn a mapping from multi-view data to the unified representation as the entity proxy. Then we learn a prototype for each view and design a Modulation Parameter Generator (MPG) that learns a set of view-aware scale and shift parameters from prototypes to modulate the entity proxy and obtain view proxies. By constraining the representativeness, uniqueness, and simplicity of the proxies and proposing an entity-view contrastive loss, parameters are alternatively updated. We end up with a set of discriminative prototypes, view proxies, and an entity proxy that are flexible enough to yield robust representations for out-of-sample entities. Extensive experiments on five datasets show that the results of our VAM outperform existing methods in both classification and clustering tasks. Ren Wang 0011, Haoliang Sun, Xiushan Nie, Yuxiu Lin, Xiaoming Xi, Yilong Yin |
ACM Multimedia | 2 |
| 2023 | Towards Accurate and Robust Domain Adaptation Under Multiple Noisy EnvironmentsabstractIn many non-stationary environments, machine learning algorithms usually confront the distribution shift scenarios. Previous domain adaptation methods have achieved great success. However, they would lose algorithm robustness in multiple noisy environments where the examples of source domain become corrupted by label noise, feature noise, or open-set noise. In this paper, we report our attempt toward achieving noise-robust domain adaptation. We first give a theoretical analysis and find that different noises have disparate impacts on the expected target risk. To eliminate the effect of source noises, we propose offline curriculum learning minimizing a newly-defined empirical source risk. We suggest a proxy distribution-based margin discrepancy to gradually decrease the noisy distribution distance to reduce the impact of source noises. We propose an energy estimator for assessing the outlier degree of open-set-noise examples to defeat the harmful influence. We also suggest robust parameter learning to mitigate the negative effect further and learn domain-invariant feature representations. Finally, we seamlessly transform these components into an adversarial network that performs efficient joint optimization for them. A series of empirical studies on the benchmark datasets and the COVID-19 screening task show that our algorithm remarkably outperforms the state-of-the-art, with over 10% accuracy improvements in some transfer tasks. Zhongyi Han, Xian-Jin Gui, Haoliang Sun, Yilong Yin, Shuo Li 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Attentional prototype inference for few-shot segmentation
Haoliang Sun, Xiankai Lu, Yilong Yin, Xiantong Zhen, Cees Snoek, Ling Shao 0001 |
Pattern Recognit. | 1 |
| 2022 | Self-Filtering: A Noise-Aware Sample Selection for Label Noise with Confidence Penalization
Qi Wei 0004, Haoliang Sun, Xiankai Lu, Yilong Yin |
ECCV (30) | 2 |
| 2022 | SNIP-FSL: Finding task-specific lottery jackpots for few-shot learning
Ren Wang 0011, Haoliang Sun, Xiushan Nie, Yilong Yin |
Knowl. Based Syst. | 2 |
| 2022 | Learning to rectify for robust learning with noisy labels
Haoliang Sun, Chenhui Guo, Qi Wei 0004, Zhongyi Han, Yilong Yin |
Pattern Recognit. | 1 |
| 2022 | Learning Transferable Parameters for Unsupervised Domain AdaptationabstractUnsupervised domain adaptation (UDA) enables a learning machine to adapt from a labeled source domain to an unlabeled target domain under the distribution shift. Thanks to the strong representation ability of deep neural networks, recent remarkable achievements in UDA resort to learning domain-invariant features. Intuitively, the goal is that a good feature representation and the hypothesis learned from the source domain can generalize well to the target domain. However, the learning processes of domain-invariant features and source hypotheses inevitably involve domain-specific information that would degrade the generalizability of UDA models on the target domain. The lottery ticket hypothesis proves that only partial parameters are essential for generalization. Motivated by it, we find in this paper that only partial parameters are essential for learning domain-invariant information. Such parameters are termed transferable parameters that can generalize well in UDA. In contrast, the rest parameters tend to fit domain-specific details and often cause the failure of generalization, which are termed untransferable parameters. Driven by this insight, we propose Transferable Parameter Learning (TransPar) to reduce the side effect of domain-specific information in the learning process and thus enhance the memorization of domain-invariant information. Specifically, according to the distribution discrepancy degree, we divide all parameters into transferable and untransferable ones in each training iteration. We then perform separate update rules for the two types of parameters. Extensive experiments on image classification and regression tasks (keypoint detection) show that TransPar outperforms prior arts by non-trivial margins. Moreover, experiments demonstrate that TransPar can be integrated into the most popular deep UDA networks and be easily extended to handle any data distribution shift scenarios. Zhongyi Han, Haoliang Sun, Yilong Yin |
IEEE Trans. Image Process. | 2 |
| 2020 | Learning to Learn Kernels with Variational Random FeaturesabstractWe introduce kernels with random Fourier features in the meta-learning framework for few-shot learning. We propose meta variational random features (MetaVRF) to learn adaptive kernels for the base-learner, which is developed in a latent variable model by treating the random feature basis as the latent variable. We formulate the optimization of MetaVRF as a variational inference problem by deriving an evidence lower bound under the meta-learning framework. To incorporate shared knowledge from related tasks, we propose a context inference of the posterior, which is established by an LSTM architecture. The LSTM-based inference network can effectively integrate the context information of previous tasks with task-specific information, generating informative and adaptive features. The learned MetaVRF can produce kernels of high representational power with a relatively low spectral sampling rate and also enables fast adaptation to new tasks. Experimental results on a variety of few-shot regression and classification tasks demonstrate that MetaVRF delivers much better, or at least competitive, performance compared to existing meta-learning alternatives. Xiantong Zhen, Haoliang Sun, Yingjun Du, Jun Xu 0019, Yilong Yin, Ling Shao 0001, Cees Snoek |
ICML | 2 |
| 2020 | Malware Classification on Imbalanced Data through Self-AttentionabstractMalware is an ever-growing threat to the Internet. New and mutated malware are appearing with increasing frequency in recent years. In the real-world scenario, new families of malware are often discovered in the cybersecurity protection system. In order to improve the protection capability of the system, it is necessary to add the newly discovered malware identification characteristics to the online system. However, the samples of newly discovered malware families are usually too small to effectively extract the characteristics of new classes, resulting in a low recognition rate of new families. The essence of this problem is a multi-classification problem based on imbalanced datasets. In this paper, we propose a self-attention based malware classification method to solve the malware classification on imbalanced datasets. An open source dataset is used to simulate the classification of malware on balanced and imbalanced datasets. Our method has reached an accuracy of 98.48%, and the F1-Score of the imbalanced Simda class has reached 89.66% on the Microsoft Kaggle dataset. Experimental results have demonstrated the effectiveness and robustness in malware classification with imbalanced datasets. Jian Xing, Xiaoyu Zhang 0002, ZiSen Oi, Ge Fu, Qian Qiang, Haoliang Sun |
TrustCom | 8 |
| 2019 | DUAL-GLOW: Conditional Flow-Based Generative Model for Modality TransferabstractPositron emission tomography (PET) imaging is an imaging modality for diagnosing a number of neurological diseases. In contrast to Magnetic Resonance Imaging (MRI), PET is costly and involves injecting a radioactive substance into the patient. Motivated by developments in modality transfer in vision, we study the generation of certain types of PET images from MRI data. We derive new flow-based generative models which we show perform well in this small sample size regime (much smaller than dataset sizes available in standard vision tasks). Our formulation, DUAL-GLOW, is based on two invertible networks and a relation network that maps the latent spaces to each other. We discuss how given the prior distribution, learning the conditional distribution of PET given the MRI image reduces to obtaining the conditional distribution between the two latent codes w.r.t. the two image types. We also extend our framework to leverage "side" information (or attributes) when available. By controlling the PET generation through "conditioning" on age, our model is also able to capture brain FDG-PET (hypometabolism) changes, as a function of age. We present experiments on the Alzheimers Disease Neuroimaging Initiative (ADNI) dataset with 826 subjects, and obtain good performance in PET image synthesis, qualitatively and quantitatively better than recent works. Haoliang Sun, Ronak Mehta, Hao Henry Zhou, Zhichun Huang, Sterling C. Johnson, Vivek Prabhakaran |
ICCV | 1 |
| 2019 | Learning the Set Graphs: Image-Set Classification Using Sparse Graph Convolutional NetworksabstractImage-set classification has recently made great progress in computer vision. Compared with traditional image classification tasks, set-based classification exhibits great challenges due to huge intra-class variability and high inter-class ambiguity. In this paper, we propose to model the image set as a graph and formulate image set classification as the graph matching task. Without relying on the strong structure assumption, we build the first end-to-end graph convolutional network, the Deep SetNet, to learn the graph structure of an image set. Specifically, the SetNet consists of one convolutional network (CNN) to sufficiently extract the discriminative vertex, one graph convolutional Network (GCN) to faithfully learn the substructure in set graphs and graph pooling layers to aggregate the vertex features from the GCN. Moreover, we propose imposing the ℓ1,2-norm based sparsity constraint to select vertex features, which largely improves the model generalization capability. Extensive experiments demonstrate that our method consistently outperforms state-of-the-art methods, showing its great effectiveness in set-based image classification. Haoliang Sun, Xiantong Zhen, Yilong Yin |
ICIP | 1 |
| 2018 | Modality-Specific Structure Preserving Hashing for Cross-Modal RetrievalabstractHashing-based methods have made great advancements in cross-modal retrieval in both computational efficiency and storage. Learning a common space from different modalities is the common strategy of hashing-based methods, however, relational and structural information between samples in each modality, namely, a modality-specific structure, is always discarded during learning. In addition, cross-modality samples sometimes suffer from inter-class ambiguity and intra-class variability because of the uncertainty of manual labeling. To address these issues, we propose a novel method named Modality-specific structure Preserving Hashing (MsPH), which learns hashes by preserving the local structure and relations between samples in each modality. Moreover, label enhancement is utilized in MsPH to address label ambiguity and variability. Extensive experiments conducted on three benchmark datasets demonstrate the superiority of MsPH under various cross-modal scenarios. Xingbo Liu, Haoliang Sun, Xiushan Nie, Chaoran Cui, Yilong Yin |
ICASSP | 2 |
| 2017 | Learning Deep Match Kernels for Image-Set ClassificationabstractImage-set classification has recently generated great popularity due to its widespread applications in computer vision. The great challenges arise from effectively and efficiently measuring the similarity between image sets with high inter-class ambiguity and huge intra-class variability. In this paper, we propose deep match kernels (DMK) to directly measure the similarity between image sets in the match kernel framework. Specifically, we build deep local match kernels between images upon arc-cosine kernels, which can faithfully characterize the similarity between images by mimicking deep neural networks, we introduce anchors to aggregate those deep local match kernels into a global match kernel between image sets, which is learned in a supervised way by kernel alignment and therefore more discriminative. The DMK provides the first match kernel framework for image-set classification, which removes specific assumptions usually required in previous approaches and is computationally more efficient. We conduct extensive experiments on four datasets for three diverse image-set classification tasks. The DMK achieves high performance and consistently surpasses state-of-the-art methods, showing its great effectiveness for image-set classification. Haoliang Sun, Xiantong Zhen, Yuanjie Zheng, Gongping Yang 0001, Yilong Yin, Shuo Li 0001 |
CVPR | 1 |