VLDB 2026 Research / reviewers in the wild / expert
Huan Liu 0012
dblp:92/309-12
· DBLP profile ↗
38ranked-venue papers
14as first author
30since 2021 · last 2026
0000-0002-7863-3751ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 24 · 10 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 7 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 1 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MIND-EEG: Multi-Granularity Integration Network With Discrete Codebook for EEG-Based Emotion RecognitionabstractEmotion recognition using electroencephalogram (EEG) signals has broad potential across various domains. EEG signals have ability to capture rich spatial information related to brain activity, yet effectively modeling and utilizing these spatial relationships remains a challenge. Existing methods struggle with simplistic spatial structure modeling, failing to capture complex node interactions, and lack generalizable spatial connection representations, failing to balance the dynamic nature of brain networks with the need for discriminative and generalizable features. To address these challenges, we propose the Multi-granularity Integration Network with Discrete Codebook for EEG-based Emotion Recognition (MIND-EEG). The framework employs a multi-granularity approach, integrating global and regional spatial information through a Global State Encoder, an Intra-Regional Functionality Encoder, and an Inter-Regional Interaction Encoder to comprehensively model brain activity. Additionally, we introduce a discrete codebook mechanism for constructing network structures via vector quantization, ensuring compact and meaningful brain network representations while mitigating over-smoothing and enhancing model generalization. The proposed framework effectively captures the dynamic and diverse nature of EEG signals, enabling robust emotion recognition. Extensive comparisons and analyses demonstrate the effectiveness of MIND-EEG, and the source code is publicly available athttps://github.com/XJTU-EEG/MIND_EEG. Yuzhe Zhang 0003, Chengxi Xie, Huan Liu 0012, Guanjian Liu, Dalin Zhang 0001 |
IEEE Trans. Affect. Comput. | 3 |
| 2025 | A Multi-Label EEG Dataset for Mental Attention State Classification in Online LearningabstractAttention is a vital cognitive process in the learning and memory environment, particularly in the context of online learning. Traditional methods for classifying attention states of online learners based on behavioral signals are prone to distortion, leading to increased interest in using electroencephalography (EEG) signals for authentic and accurate assessment. However, the field of attention state classification based on EEG signals in online learning faces challenges, including the scarcity of publicly available datasets, the lack of standardized data collection paradigms, and the requirement to consider the interplay between attention and other psychological states. In light of this, we present the Multi-label EEG dataset for classifying Mental Attention states (MEMA) in online learning. We meticulously designed a reliable and standard experimental paradigm with three attention states: neutral, relaxing, and concentrating, considering human physiological and psychological characteristics. This paradigm collected EEG signals from 20 subjects, each participating in 12 trials, resulting in 1,060 minutes of data. Emotional state labels, basic personal information, and personality traits were also collected to investigate the relation-ship between attention and other psychological states. Extensive quantitative and qualitative analysis, including a multi-label correlation study, validated the quality of the EEG attention data. The MEMA dataset and analysis provide valuable insights for advancing research on attention in online learning. The dataset is publicly available at https://github.com/XJTU-EEG/MEMA. Huan Liu 0012, Yuzhe Zhang 0003, Guanjian Liu, Xinxin Du, Haochong Wang, Dalin Zhang 0001 |
ICASSP | 1 |
| 2025 | FCAT-Diff: Flexible and Consistent Appearance Transfer Based on Training-free Diffusion Model
Zhengyi Gong, Ming-Wen Shao, Chang Liu 0115, Huan Liu 0012 |
Comput. Graph. | 5 |
| 2025 | DiffRA: universal restorative adversarial attack based on diffusion model
Ming-Wen Shao, Lingzhuang Meng, Huan Liu 0012, Xiaodong Tan 0002 |
Multim. Syst. | 4 |
| 2025 | SeBIR: Semantic-guided burst image restoration
Huan Liu 0012, Ming-Wen Shao, Yecong Wan, Yuexian Liu, Kai Shang 0001 |
Neural Networks | 1 |
| 2025 | Masked contrastive graph representation learning for age estimation
Yuntao Shou, Xiangyong Cao, Huan Liu 0012, Deyu Meng |
Pattern Recognit. | 3 |
| 2025 | LibEER: A Comprehensive Benchmark and Algorithm Library for EEG-Based Emotion RecognitionabstractEEG-based emotion recognition (EER) has gained significant attention due to its potential for understanding and analyzing human emotions. While recent advancements in deep learning techniques have substantially improved EER, the field lacks a convincing benchmark and comprehensive open-source libraries. This absence complicates fair comparisons between models and creates reproducibility challenges for practitioners, which collectively hinder progress. To address these issues, we introduce LibEER, a comprehensive benchmark and algorithm library designed to facilitate fair comparisons in EER. LibEER carefully selects popular and powerful baselines, harmonizes key implementation details across methods, and provides a standardized codebase in PyTorch. By offering a consistent evaluation framework with standardized experimental settings, LibEER enables unbiased assessments of seventeen representative deep learning models for EER across the six most widely used datasets. Additionally, we conduct a thorough, reproducible comparison of model performance and efficiency, providing valuable insights to guide researchers in the selection and design of EER models. Moreover, we make observations and in-depth analysis on the experiment results and identify current challenges in this community. We hope that our work will not only lower entry barriers for newcomers to EEG-based emotion recognition but also contribute to the standardization of research in this domain, fostering steady development. The library and source code are publicly available athttps://github.com/XJTU-EEG/LibEER. Huan Liu 0012, Shusen Yang, Yuzhe Zhang 0003, Mengze Wang, Fanyu Gong, Chengxi Xie, Guanjian Liu, Zejun Liu, Yong-Jin Liu 0001, Bao-Liang Lu, Dalin Zhang 0001 |
IEEE Trans. Affect. Comput. | 1 |
| 2025 | A Low-Rank Matching Attention Based Cross-Modal Feature Fusion Method for Conversational Emotion RecognitionabstractConversational emotion recognition (CER) is an important research topic in human-computer interactions. Although recent advancements in transformer-based cross-modal fusion methods have shown promise in CER tasks, they tend to overlook the crucial intra-modal and inter-modal emotional interaction or suffer from high computational complexity. To address this, we introduce a novel and lightweight cross-modal feature fusion method called Low-Rank Matching Attention Method (LMAM). LMAM effectively captures contextual emotional semantic information in conversations while mitigating the quadratic complexity issue caused by the self-attention mechanism. Specifically, by setting a matching weight and calculating inter-modal features attention scores row by row, LMAM requires only one-third of the parameters of self-attention methods. We also employ the low-rank decomposition method on the weights to further reduce the number of parameters in LMAM. As a result, LMAM offers a lightweight model while avoiding overfitting problems caused by a large number of parameters. Moreover, LMAM is able to fully exploit the intra-modal emotional contextual information within each modality and integrates complementary emotional semantic information across modalities by computing and fusing similarities of intra-modal and inter-modal features simultaneously. Experimental results verify the superiority of LMAM compared with other popular cross-modal fusion methods on the premise of being more lightweight. Also, LMAM can be embedded into any existing state-of-the-art CER methods in a plug-and-play manner, and can be applied to other multi-modal recognition tasks, e.g., session recommendation and humour detection, demonstrating its remarkable generalization ability. Yuntao Shou, Huan Liu 0012, Xiangyong Cao, Deyu Meng, Bo Dong 0001 |
IEEE Trans. Affect. Comput. | 2 |
| 2025 | LLM-Enhanced Multi-Teacher Knowledge Distillation for Modality-Incomplete Emotion Recognition in Daily HealthcareabstractThe critical importance of monitoring and recognizing human emotional states in healthcare has led to a surge in proposals for EEG-based multimodal emotion recognition in recent years. However, practical challenges arise in acquiring EEG signals in daily healthcare settings due to stringent data acquisition conditions, resulting in the issue of incomplete modalities. Existing studies have turned to knowledge distillation as a means to mitigate this problem by transferring knowledge from multimodal networks to unimodal ones. However, these methods are constrained by the use of a single teacher model to transfer integrated feature extraction knowledge, particularly concerning spatial and temporal features in EEG data. To address this limitation, we propose a multi-teacher knowledge distillation framework enhanced with a Large Language Model (LLM), aimed at facilitating effective feature learning in the student network by transferring knowledge of extracting integrated features. Specifically, we employ an LLM as the teacher for extracting temporal features and a graph convolutional neural network for extracting spatial features. To further enhance knowledge distillation, we introduce causal masking and a confidence indicator into the LLM to facilitate the transfer of the most discriminative features. Extensive testing on the DEAP and MAHNOB-HCI datasets demonstrates that our model outperforms existing methods in the modality-incomplete scenario. This study underscores the potential application of large models in this field. Yuzhe Zhang 0003, Huan Liu 0012, Yang Xiao 0014, Mohammed Amoon, Dalin Zhang 0001, Di Wang 0004, Shusen Yang, Hiok Chai Quek |
IEEE J. Biomed. Health Informatics | 2 |
| 2025 | A Unified Optimal Transport Framework for Cross-Modal Retrieval With Noisy LabelsabstractCross-modal retrieval (CMR) aims to establish interaction between different modalities, among which supervised CMR is emerging due to its flexibility in learning semantic category discrimination. Despite the remarkable performance of previous supervised CMR methods, much of their success can be attributed to the well-annotated data. However, even for unimodal data, precise annotation is expensive and time-consuming, and it becomes more challenging with the multimodal scenario. In practice, massive multimodal data are collected from the Internet with coarse annotation, which inevitably introduces noisy labels. Training with such misleading labels would bring two key challenges-enforcing the multimodal samples to align incorrect semantics and widen the heterogeneous gap, resulting in poor retrieval performance. To tackle these challenges, this work proposes UOT-RCL, a unified framework based on optimal transport (OT) for robust CMR. First, we propose a semantic alignment based on partial OT to progressively correct the noisy labels, where a novel cross-modal consistent cost function is designed to blend different modalities and provide precise transport cost. Second, to narrow the discrepancy in multimodal data, an OT-based relation alignment is proposed to infer the semantic-level cross-modal matching. Both of these components leverage the inherent correlation among multimodal data to facilitate effective cost function. The experiments on three widely used CMR datasets demonstrate that our UOT-RCL surpasses the state-of-the-art approaches and significantly improves the robustness against noisy labels. Haochen Han, Minnan Luo, Huan Liu 0012, Jun Liu 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Learning Multimodal Attention Mixed with Frequency Domain Information as Detector for Fake News DetectionabstractDetecting fake news on social media has become a crucial task in combating online misinformation and countering malicious propaganda. Existing methods rely on semantic consistency across modalities to fuse features and determine news authenticity. However, cunning fake news publisher manipulate image to ensure a high level of semantic consistency between news post and image, making it more difficult to distinguish fake news. To this end, we propose MHFFD (Mixed High-Frequency Feature Detector), a novel fake news detection framework that utilizes token-level semantic consistency evaluation to identify key elements in news content and provide guidance for discovering image manipulation and learning better news representations. Extensive experiments demonstrate that MHFFD outperforms state-of-the-art methods on two widely used fake news detection datasets. Further research also validates the effectiveness of token-level semantic alignment and manipulation detection. Zihan Ma 0001, Huan Liu 0012, Zhi Zeng 0001, Xiang Zhao 0002, Minnan Luo |
ICME | 2 |
| 2024 | Mitigating World Biases: A Multimodal Multi-View Debiasing Framework for Fake News Video DetectionabstractShort videos turn into an important channel for public sharing, as well as they've become a fertile ground for fake news. Fake news video detection is to judge the veracity of news based on its different modal information, such as video, audio, text, image and social context information. Current detection models tend to learn the multimodal dataset biases within spurious correlations between news modalities and veracity labels as shortcuts, rather than learning how to integrate the multimodal information behind them to reason, resulting in seriously degrading their detection and generalization capabilities. To address this issues, we propose a Multimodal Multi-View Debiasing (MMVD) framework, which makes the first attempt to mitigate various multimodal biases for fake news video detection. Inspired by people's misleading situations by multimodal short videos, we summarize three cognitive biases: static, dynamic and social biases. MMVD put forward a multi-view causal reasoning strategy to learn unbiased dependencies within the cognitive biases, thus enhancing the unbiased prediction of multimodal videos. The extensive experimental results show that the MMVD could improve the detection performance of multimodal fake news video. Studies also confirm that our MMVD can mitigate multiple biases on complex real-world scenarios and improve generalization ability of fake news video detection. Zhi Zeng 0001, Minnan Luo, Xiangzheng Kong, Huan Liu 0012, Hao Yang 0042, Zihan Ma 0001, Xiang Zhao 0002 |
ACM Multimedia | 4 |
| 2024 | Frequency-aware network for low-light image enhancement
Kai Shang 0001, Ming-Wen Shao, Yuanjian Qiao 0001, Huan Liu 0012 |
Comput. Graph. | 4 |
| 2024 | When guided diffusion model meets zero-shot image super-resolution
Huan Liu 0012, Ming-Wen Shao, Kai Shang 0001, Yuanjian Qiao 0001, Shuigen Wang |
Eng. Appl. Artif. Intell. | 1 |
| 2024 | Semantics-Guided Contrastive Network for Zero-Shot Object DetectionabstractZero-shot object detection (ZSD), the task that extends conventional detection models to detecting objects from unseen categories, has emerged as a new challenge in computer vision. Most existing approaches tackle the ZSD task with a strict mapping-transfer strategy that may lead to suboptimal ZSD results: 1) the learning process of these models neglects the available semantic information on unseen classes, which can easily bias towards the seen categories; 2) the original visual feature space is not well-structured for the ZSD task due to the lack of discriminative information. To address these issues, we develop a novel Semantics-Guided Contrastive Network for ZSD, named ContrastZSD, a detection framework that first brings contrastive learning mechanism into the realm of zero-shot detection. Particularly, ContrastZSD incorporates two semantics-guided contrastive learning subnets that contrast between region-category and region-region pairs respectively. The pairwise contrastive tasks take advantage of supervision signals derived from both the ground truth label and class similarity information. By performing supervised contrastive learning over those explicit semantic supervision, the model can learn more knowledge about unseen categories to avoid the bias problem to seen concepts, while optimizing the visual data structure to be more discriminative for better visual-semantic alignment. Extensive experiments are conducted on two popular benchmarks for ZSD, i.e., PASCAL VOC and MS COCO. Results show that our method outperforms the previous state-of-the-art on both ZSD and generalized ZSD tasks. Caixia Yan, Xiaojun Chang, Minnan Luo, Huan Liu 0012, Xiaoqin Zhang 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | Self-Supervised EEG Representation Learning for Robust Emotion RecognitionabstractEmotion recognition based on electroencephalography (EEG) is becoming a growing concern of researchers due to its various applications and portable devices. Existing methods are mainly dedicated to EEG feature representation and have made impressive progress. However, the problem of scarce labels restricts their further promotion. In light of this, we propose a self-supervised framework with contrastive learning for robust EEG-based emotion recognition, which can effectively leverage both readily available unlabeled EEG signals and labeled ones to learn highly discriminative EEG features. Firstly, we construct a specific pretext task according to the sequential non-stationarity of emotional EEG signals for contrastive learning, which aims at extracting pseudo-label information from all EEG data. Meanwhile, we propose a novel negative segment selection algorithm to reduce the noise of unlabeled data during the contrastive learning process. Secondly, to mitigate the overfitting issue induced by a small number of labeled samples during learning, we originate a loss function with label smoothing regularization that can guide the model to learn generalizable features. Extensive experiments over three benchmark datasets demonstrate the effectiveness and superiority of our model on EEG-based emotion recognition task. Besides, the generalization and robustness of the model have also been proved through sufficient experiments. Huan Liu 0012, Yuzhe Zhang 0003, Xuxu Chen, Dalin Zhang 0001, Rui Li 0073, Tao Qin 0002 |
ACM Trans. Sens. Networks | 1 |
| 2023 | Spatially-Aware Human-Object Interaction Detection with Cross-Modal Enhancement
Gaowen Liu, Huan Liu 0012, Caixia Yan, Rui Li 0073, Sizhe Dang |
ICONIP (5) | 2 |
| 2023 | An Efficient Frequency Domain Separation Network for Paired and Unpaired Image Super-ResolutionabstractAlthough existing super-resolution (SR) techniques have made great progress, they are often tailored for either paired or unpaired scenery, thus may result in poor migration ability. In this work, we propose a generalized Frequency Domain Separation Network (FDSNet) for both paired and unpaired SR settings. Firstly, through statistical analysis, we found that real-world low-resolution (LR) images and high-resolution (HR) images differ greatly in high frequencies but less in low frequencies. Inspired by this, we perform high and low-frequency separation of LR images and guide our model to reconstruct the HR contents in the different frequency domains. Then, according to the varying attention on frequencies of traditional CNN and Transformer models, we design a parallel pipeline: LFNet based on Transformer for low-frequency feature extraction, and HFNet based on CNN for high frequencies. In LFNet, to further alleviate the high complexity and data dependency of Transformer, Simplified Multi-head Self Attention (SMSA) is proposed at a low computational cost. And original MLP is replaced by our Spatial Enhancement MLP (SEMLP) to take full advantage of local spatial contexts. Finally, to further facilitate frequency separation and learning, a Frequency attention block is designed to impose guidance on high frequencies. Experiments indicate that our FDSNet achieves promising performance in terms of quantitative and qualitative evaluations while enjoying a faster speed and much fewer parameters. Huan Liu 0012, Ming-Wen Shao, Yuanjian Qiao 0001, Fukang Liu |
IJCNN | 1 |
| 2023 | Mutual channel prior guided dual-domain interaction network for single image raindrop removal
Yuanjian Qiao 0001, Ming-Wen Shao, Huan Liu 0012, Kai Shang 0001 |
Comput. Graph. | 3 |
| 2023 | Graph neural networks induced by concept lattices for classification
Ming-Wen Shao, Weizhi Wu 0001, Huan Liu 0012 |
Int. J. Approx. Reason. | 4 |
| 2023 | Adaptive one-stage generative adversarial network for unpaired image super-resolution
Ming-Wen Shao, Huan Liu 0012, Jianxin Yang, Feilong Cao |
Neural Comput. Appl. | 2 |
| 2023 | Image Super-Resolution Using a Simple Transformer Without Pretraining
Huan Liu 0012, Ming-Wen Shao, Chao Wang 0102, Feilong Cao |
Neural Process. Lett. | 1 |
| 2023 | Unpaired image super-resolution using a lightweight invertible neural network
Huan Liu 0012, Ming-Wen Shao, Yuanjian Qiao 0001, Yecong Wan, Deyu Meng |
Pattern Recognit. | 1 |
| 2023 | Dual-channel graph contrastive learning for self-supervised graph-level representation learning
Zhenfei Luo, Yixiang Dong, Huan Liu 0012, Minnan Luo |
Pattern Recognit. | 4 |
| 2023 | Social Image-Text Sentiment Classification With Cross-Modal Consistency and Knowledge DistillationabstractSocial media sentiment analysis, which aims to evaluate the attitudes of online users based on their posts, has attracted significant research attention due to its successful application in the field of social media monitoring. It is a beneficial way to utilize multimodal information uploaded by users in order to improve sentiment classification ability. However, existing multimodal fusion-based approaches continue to face difficulties due to the issues of between-modality semantic inconsistency and missing modality. To address these issues, we propose a cross-modal consistency modeling-based knowledge distillation framework for image–text sentiment classification of social media data. Specifically, we design a hybrid curriculum learning strategy to measure the semantic consistency of multimodal data, then gradually train all image–text pairs from easy to hard, which can effectively handle the massive amounts of noise caused by inconsistencies between image and text data on social media. Moreover, in order to alleviate the problem of missing images in unimodal posts, we propose a privileged feature distillation method, in which the teacher model additionally considers images as privileged features, to transfer the visual knowledge to the student model, thereby enhancing the accuracy for text sentiment classification. Extensive experiments conducted over three real-world social media datasets demonstrate the effectiveness and superiority of the proposed multimodal sentiment analysis model. Huan Liu 0012, Jianping Fan 0007, Caixia Yan, Tao Qin 0002 |
IEEE Trans. Affect. Comput. | 1 |
| 2023 | EEG-Based Emotion Recognition With Emotion Localization via Hierarchical Self-AttentionabstractEmotion recognition based on electroencephalography (EEG) has attracted significant attention due to its wide range of applications, especially in Human-Computer Interaction(HCI). Previous research treats different segments of EEG signals uniformly, ignoring the fact that emotions are unstable and discrete during an extended period. In this paper, we propose a novel two-step spatial-temporal emotion recognition framework. First, considering that the human emotion has not only ”short-term continuity” but also ”long-term similarity”, we propose a hierarchical self-attention network to jointly model local and global temporal information, so as to localize most related segments and reduce the influence of noise at the temporal level. Second, in order to extract discriminative features at the spatial level to enhance the emotion recognition performance, we further employ the squeeze-and-excitation module (SE module) along with the channel correlation loss (CC-Loss) to select the most task-related channels. We also define a new task calledemotion localization, which aims to localize fragments with stronger emotions. We evaluate the proposed method on the proposed emotion localization task and typical emotion recognition task with three publicly available datasets, i.e., SEED, DEAP, and MAHNOB-HCI. The experimental results demonstrate that the proposed approach outperforms state-of-the-art methods. Yuzhe Zhang 0003, Huan Liu 0012, Dalin Zhang 0001, Xuxu Chen, Tao Qin 0002 |
IEEE Trans. Affect. Comput. | 2 |
| 2023 | Deep Fuzzy Clustering Transformer: Learning the General Property of Corruptions for Degradation-Agnostic Multitask Image RestorationabstractFor the sake of eliminating multiple degradations, most existing multitask image restoration methods prefer to learn the properties of each degradation type, which is often accompanied by a bloated model size and a heavy learning burden. To tackle the aforementioned issues, in this article, we propose to treat multiple degradations uniformly to achieve degradation type-agnostic multitask image restoration. We observe that the degradations in different spatial locations are always morphologically similar while the background sceneries vary greatly. In accordance with the aforementioned observation, we decouple the degradation features and the background features by an efficient fuzzy clustering method. The degradation features contain all the diverse degradation information, while the images are recovered from the decoupled background features. In practice, we discover a uniformity between the fuzzy C-means algorithm and cross attention and propose a deep fuzzy clustering transformer to achieve degradation type-agnostic background extraction via feature map clustering based on spatial distribution characteristics. Furthermore, to capture the spatial distribution properties of an image, an efficient global attention tree (GAT) is devised to provide a global spatial receptive field for the clustering process. By virtue of the quadtree structure, the proposed GATs enable more efficient global modeling than existing methods. Our experimental analysis showed that the proposed method outperformed the state-of-the-art models in terms of both efficiency and performance. Yuanshuo Cheng, Ming-Wen Shao, Yecong Wan, Yue-Xian Liu, Huan Liu 0012, Deyu Meng |
IEEE Trans. Fuzzy Syst. | 5 |
| 2021 | Reliable Recommendation with Review-level ExplanationsabstractThe quality of user-generated reviews is significant for users to understand recommendation results and make online purchasing decisions correctly. However, the reliability of a review, which captures the likelihood that a review is benign, is ignored by many studies. The low reliability reviews cause a recommendation system's unsatisfying performance. Especially the fake reviews written by fraudulent users mislead the system into generating error recommendation results and explanations, which confuse customers and deprive customers of confidence in the system. In this paper, we propose a model, Reliable Recommendation with Review-level Explanations (RRRE), which detects reliable reviews and improves the performance of the explainable recommendation system as well. Recognizing the textual content of reviews, user-item interactions are valuable features for both rating prediction and reliability prediction. RRRE builds a uniform framework to predict rating scores and reliability scores simultaneously. Firstly, RRRE embeds user preferences and item profiles, which are extracted from textual and interactive features, into the representation of the review. Secondly, the supervised information of two subtasks is jointly combined. It makes the optimization of RRRE faster and better. Finally, the reviews with both high reliability scores and rating scores are given to customers as reliable explanations. To the best of our knowledge, we are the first to consider the reliability of reviews for improving explainable recommender system. And the experimental results confirm this idea and show that our model outperforms other baseline methods on Yelp and Amazon datasets. Yanzhang Lyu, Hongzhi Yin, Jun Liu 0002, Mengyue Liu, Huan Liu 0012, Shizhuo Deng |
ICDE | 5 |
| 2021 | MutualRec: Joint friend and item recommendations with mutualistic attentional graph neural networks
Yang Xiao 0014, Qingqi Pei, Tingting Xiao, Lina Yao 0001, Huan Liu 0012 |
J. Netw. Comput. Appl. | 5 |
| 2021 | DMDIT: Diverse multi-domain image-to-image translation
Ming-Wen Shao, Youcai Zhang, Huan Liu 0012, Chao Wang 0102, Xun Shao |
Knowl. Based Syst. | 3 |
| 2020 | Dual-stream generative adversarial networks for distributionally robust zero-shot learning
Huan Liu 0012, Lina Yao 0001, Minnan Luo, Hongke Zhao, Yanzhang Lyu |
Inf. Sci. | 1 |
| 2020 | Memory transformation networks for weakly supervised visual classification
Huan Liu 0012, Minnan Luo, Xiaojun Chang, Caixia Yan, Lina Yao 0001 |
Knowl. Based Syst. | 1 |
| 2020 | Progressive generative adversarial networks with reliable sample identification
Minnan Luo, Huan Liu 0012 |
Pattern Recognit. Lett. | 3 |
| 2018 | Detecting global and local topics via mining twitter data
Huan Liu 0012, Yong Ge 0001, Rongcheng Lin |
Neurocomputing | 1 |
| 2018 | Top-k multi-class SVM using multiple features
Caixia Yan, Minnan Luo, Huan Liu 0012, Zhihui Li 0001 |
Inf. Sci. | 3 |
| 2018 | An efficient multi-feature SVM solver for complex event detection
Huan Liu 0012, Zhihui Li 0001, Tao Qin 0002, Lei Zhu 0002 |
Multim. Tools Appl. | 1 |
| 2018 | Exploring open information via event networkabstractAbstract It is a challenging task to discover information from a large amount of data in an open domain.1In this paper, an event network framework is proposed to address this challenge. It is in fact an empirical construct for exploring open information, composed of three steps: document event detection, event network construction and event network analysis. First, documents are clustered into document events for reducing the impact of noisy and heterogeneous resources. Secondly, linguistic units (e.g., named entities or entity relations) are extracted from each document event and combined into an event network, which enables content-oriented retrieval. Then, in the final step, techniques such as social network or complex network can be applied to analyze the event network for exploring open information. In the implementation section, we provide examples of exploring open information via event network. Yanping Chen 0010, Feng Tian 0002, Huan Liu 0012 |
Nat. Lang. Eng. | 4 |
| 2017 | How Unlabeled Web Videos Help Complex Event Detection?abstractThe lack of labeled exemplars is an important factor that makes the task of multimedia event detection (MED) complicated and challenging. Utilizing artificially picked and labeled external sources is an effective way to enhance the performance of MED. However, building these data usually requires professional human annotators, and the procedure is too time-consuming and costly to scale. In this paper, we propose a new robust dictionary learning framework for complex event detection, which is able to handle both labeled and easy-to-get unlabeled web videos by sharing the same dictionary. By employing the lq-norm based loss jointly with the structured sparsity based regularization, our model shows strong robustness against the substantial noisy and outlier videos from open source. We exploit an effective optimization algorithm to solve the proposed highly non-smooth and non-convex problem. Extensive experiment results over standard datasets of TRECVID MEDTest 2013 and TRECVID MEDTest 2014 demonstrate the effectiveness and superiority of the proposed framework on complex event detection. Huan Liu 0012, Minnan Luo, Dingwen Zhang, Xiaojun Chang, Cheng Deng 0002 |
IJCAI | 1 |