Guang-Neng Hu

dblp:165/3128 · also Guangneng Hu · DBLP profile ↗
← Back
16ranked-venue papers
6as first author
11since 2021 · last 2026
0000-0001-5239-8151ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 3 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 5 · 4 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Dual-Seed Evolutionary Algorithm for Noise Optimization in Diffusion Models
abstract
Diffusion models have emerged as state-of-the-art generative methods, particularly excelling in conditional tasks such as prompt-driven image synthesis. While recent research emphasizes the pivotal role of noise seeds in enhancing text-image alignment and generating human-preferred outputs,these works predominantly rely on random Gaussian noise or heuristic local adjustments, , overlooking the potential of global optimization trategies to systematically improve generation quality. To bridge this gap, we propose Seed Optimization based on Evolution (SOE), a hybrid framework that integrates global evolutionary search with local semantic refinement. The global evolutionary stage conducts seed selection by jointly optimizing text-image alignment (via CLIP-Score) and human preference estimation (via ImageReward), while the local stage employs diffusion inversion to inject conditional semantics into the noise seed. Together, these components constitute a model-agnostic, training-free optimization framework for conditional diffusion models. Extensive experiments across various diffusion models demonstrate that SOE consistently improves semantic fidelity and visual quality, highlighting its generalizability and potential as a plug-and-play enhancement for generative diffusion pipelines.
Yuzheng Tan, Yuan He 0011, Yao Zhu 0003, Tianlin Huo, Huanqian Yan, Hang Su 0006, Guang-Neng Hu
AAAI8
2025 ICM-Assistant: Instruction-tuning Multimodal Large Language Models for Rule-based Explainable Image Content Moderation
abstract
Controversial contents largely inundate the Internet, infringing various cultural norms and child protection standards. Traditional Image Content Moderation (ICM) models fall short in producing precise moderation decisions for diverse standards, while recent multimodal large language models (MLLMs), when adopted to general rule-based ICM, often produce classification and explanation results that are inconsistent with human moderators. Aiming at flexible, explainable, and accurate ICM, we design a novel rule-based dataset generation pipeline, decomposing concise human-defined rules and leveraging well-designed multi-stage prompts to enrich short explicit image annotations. Our ICM-Instruct dataset includes detailed moderation explanation and moderation Q-A pairs. Built upon it, we create our ICM-Assistant model in the framework of rule-based ICM, making it readily applicable in real practice. Our ICM-Assistant model demonstrates exceptional performance and flexibility. Specifically, it significantly outperforms existing approaches on various sources, improving both the moderation classification (36.8% on average) and moderation explanation quality (26.6% on average) consistently over existing MLLMs. Caution: Content includes offensive language or images.
Mengyang Wu, Yuzhi Zhao, Jialun Cao, Mingjie Xu, Zhongming Jiang, Qinbin Li, Guang-Neng Hu, Shengchao Qin, Chi-Wing Fu
AAAI8
2025 SEFE: Superficial and Essential Forgetting Eliminator for Multimodal Continual Instruction Tuning
abstract
Multimodal Continual Instruction Tuning (MCIT) aims to enable Multimodal Large Language Models (MLLMs) to incrementally learn new tasks without catastrophic forgetting, thus adapting to evolving requirements. In this paper, we explore the forgetting caused by such incremental training, categorizing it into superficial forgetting and essential forgetting. Superficial forgetting refers to cases where the model’s knowledge may not be genuinely lost, but its responses to previous tasks deviate from expected formats due to the influence of subsequent tasks’ answer styles, making the results unusable. On the other hand, essential forgetting refers to situations where the model provides correctly formatted but factually inaccurate answers, indicating a true loss of knowledge. Assessing essential forgetting necessitates addressing superficial forgetting first, as severe superficial forgetting can conceal the model’s knowledge state. Hence, we first introduce the Answer Style Diversification (ASD) paradigm, which defines a standardized process for data style transformations across different tasks, unifying their training sets into similarly diversified styles to prevent superficial forgetting caused by style shifts. Building on this, we propose RegLoRA to mitigate essential forgetting. RegLoRA stabilizes key parameters where prior knowledge is primarily stored by applying regularization to LoRA’s weight update matrices, enabling the model to retain existing competencies while remaining adaptable to new tasks. Experimental results demonstrate that our overall method, SEFE, achieves state-of-the-art performance.
Jinpeng Chen 0003, Runmin Cong, Yuzhi Zhao, Hongzheng Yang, Guang-Neng Hu, Horace Ho-Shing Ip, Sam Kwong
ICML5
2025 Advanced dialog state tracking with noetic graphs for complex human-machine interactions
Sitong Yan, Guang-Neng Hu, Chengen Lai
Pattern Recognit.4
2024 Towards More Faithful Natural Language Explanation Using Multi-Level Contrastive Learning in VQA
abstract
Natural language explanation in visual question answer (VQA-NLE) aims to explain the decision-making process of models by generating natural language sentences to increase users' trust in the black-box systems. Existing post-hoc methods have achieved significant progress in obtaining a plausible explanation. However, such post-hoc explanations are not always aligned with human logical inference, suffering from the issues on: 1) Deductive unsatisfiability, the generated explanations do not logically lead to the answer; 2) Factual inconsistency, the model falsifies its counterfactual explanation for answers without considering the facts in images; and 3) Semantic perturbation insensitivity, the model can not recognize the semantic changes caused by small perturbations. These problems reduce the faithfulness of explanations generated by models. To address the above issues, we propose a novel self-supervised Multi-level Contrastive Learning based natural language Explanation model (MCLE) for VQA with semantic-level, image-level, and instance-level factual and counterfactual samples. MCLE extracts discriminative features and aligns the feature spaces from explanations with visual question and answer to generate more consistent explanations. We conduct extensive experiments, ablation analysis, and case study to demonstrate the effectiveness of our method on two VQA-NLE benchmarks.
Chengen Lai, Shiqi Meng, Sitong Yan, Guang-Neng Hu
AAAI6
2024 MAMO: Multi-Task Architecture Learning via Multi-Objective and Gradients Mediative Kernel
abstract
Multi-task learning (MTL) is effective in solving multiple related tasks simultaneously by sharing knowledge. However, a key challenge hindering its applications is the task interference problem where different tasks compete with each other, leading to the gradients conflicts during optimization and suffering from negative transfer. One thread is to manipulate task gradients by adjusting conflicting directions, ignoring architecture learning. Another thread is to learn architectures by generating task-exclusive modules, ignoring all-task balances. We address the problem by proposing a novel Multi-task Architecture learning model via Multi-Objective (MAMO) optimization. It achieves the goal in two steps. First, for the competing tasks detected during architecture learning, MAMO automatically generates a new module of gradient mediative kernel (GMK). Second, MAMO finds a Pareto optimal solution that balances all tasks during model parameter learning. MAMO outperforms various MTL baselines on benchmarks with an effective model size. It is model-agnostic and can be integrated into other SOTA methods to promote their performance. Extensive ablation study is conducted to understand the working of MAMO.
Yuzheng Tan, Guang-Neng Hu
ECAI2
2024 Improving Vision and Language Concepts Understanding with Multimodal Counterfactual Samples
Chengen Lai, Sitong Yan, Guang-Neng Hu
ECCV (69)4
2024 DANTE: Dialog graph enhanced prompt learning for conversational question answering over KGs
Sitong Yan, Guang-Neng Hu, Chengen Lai
Knowl. Based Syst.4
2023 TITAN : Task-oriented Dialogues with Mixed-Initiative Interactions
abstract
In multi-domain task-oriented dialogue systems, users proactively propose a series of domain-specific requests that can often be under-or over-specified, sometimes with ambiguous and cross-domain demands. System-sided initiative would be necessary to identify certain situations and appropriately interact with users to resolve them. However, most existing task-oriented dialogue systems fail to consider such mixed-initiative interaction strategies, performing low efficiency and poor collaboration ability in human-computer conversation. In this paper, we construct a multi-domain task-oriented dialogue dataset with mixed-initiative strategies named TITAN from the large-scale dialogue corpus MultiWOZ 2.1. It contains a total of 1,800 human-human conversations where the system can either ask clarification questions actively or provides relevant information to address failure situations and implicit user requests. We report the results of several baseline models on system response generation and dialogue act prediction to assess the performance of SOTA methods on TITAN. These models can capture mixed-initiative dialogue acts, while remaining the deficiency to actively generate implicit requests and accurately provide alternative information, suggesting ample room for improvement in future studies.
Sitong Yan, Shiqi Meng, Guang-Neng Hu
IJCAI5
2021 TrNews: Heterogeneous User-Interest Transfer Learning for News Recommendation
abstract
We investigate how to solve the cross-corpus news recommendation for unseen users in the future.This is a problem where traditional content-based recommendation techniques often fail.Luckily, in real-world recommendation services, some publisher (e.g., Daily news) may have accumulated a large corpus with lots of consumers which can be used for a newly deployed publisher (e.g., Political news).To take advantage of the existing corpus, we propose a transfer learning model (dubbed as TrNews) for news recommendation to transfer the knowledge from a source corpus to a target corpus.To tackle the heterogeneity of different user interests and of different word distributions across corpora, we design a translator-based transfer-learning strategy to learn a representation mapping between source and target corpora.The learned translator can be used to generate representations for unseen users in the future.We show through experiments on real-world datasets that TrNews is better than various baselines in terms of four metrics.We also show that our translator is effective among existing transfer strategies.
Guang-Neng Hu, Qiang Yang 0001
EACL1
2021 Dual Side Deep Context-aware Modulation for Social Recommendation
abstract
Social recommendation is effective in improving the recommendation performance by leveraging social relations from online social networking platforms. Social relations among users provide friends’ information for modeling users’ interest in candidate items and help items expose to potential consumers (i.e., item attraction). However, there are two issues haven’t been well-studied: Firstly, for the user interests, existing methods typically aggregate friends’ information contextualized on the candidate item only, and this shallow context-aware aggregation makes them suffer from the limited friends’ information. Secondly, for the item attraction, if the item’s past consumers are the friends of or have a similar consumption habit to the targeted user, the item may be more attractive to the targeted user, but most existing methods neglect the relation enhanced context-aware item attraction.
Bairan Fu, Wenming Zhang, Guang-Neng Hu, Xinyu Dai, Shujian Huang, Jiajun Chen 0001
WWW3
2019 Transfer Meets Hybrid: A Synthetic Approach for Cross-Domain Collaborative Filtering with Text
abstract
Collaborative Filtering (CF) is the key technique for recommender systems. CF exploits user-item behavior interactions (e.g., clicks) only and hence suffers from the data sparsity issue. One research thread is to integrate auxiliary information such as product reviews and news titles, leading to hybrid filtering methods. Another thread is to transfer knowledge from source domains such as improving the movie recommendation with the knowledge from the book domain, leading to transfer learning methods. In real-world applications, a user registers for multiple services across websites. Thus it motivates us to exploit both auxiliary and source information for recommendation in this paper. To achieve this, we propose a Transfer Meeting Hybrid (TMH) model for cross-domain recommendation with unstructured text. The proposed TMH model attentively extracts useful content from unstructured text via a memory network and selectively transfers knowledge from a source domain via a transfer network. On two real-world datasets, TMH shows better performance in terms of three ranking metrics by comparing with various baselines. We conduct thorough analyses to understand how the text content and transferred knowledge help the proposed model.
Guang-Neng Hu, Yu Zhang 0006, Qiang Yang 0001
WWW1
2018 CoNet: Collaborative Cross Networks for Cross-Domain Recommendation
abstract
The cross-domain recommendation technique is an effective way of alleviating the data sparse issue in recommender systems by leveraging the knowledge from relevant domains. Transfer learning is a class of algorithms underlying these techniques. In this paper, we propose a novel transfer learning approach for cross-domain recommendation by using neural networks as the base model. In contrast to the matrix factorization based cross-domain techniques, our method is deep transfer learning, which can learn complex user-item interaction relationships. We assume that hidden layers in two base networks are connected by cross mappings, leading to the collaborative cross networks (CoNet). CoNet enables dual knowledge transfer across domains by introducing cross connections from one base network to another and vice versa. CoNet is achieved in multi-layer feedforward networks by adding dual connections and joint loss functions, which can be trained efficiently by back-propagation. The proposed model is thoroughly evaluated on two large real-world datasets. It outperforms baselines by relative improvements of 7.84% in NDCG. We demonstrate the necessity of adaptively selecting representations to transfer. Our model can reduce tens of thousands training examples comparing with non-transfer methods and still has the competitive performance with them.
Guang-Neng Hu, Yu Zhang 0006, Qiang Yang 0001
CIKM1
2018 Collaborative Filtering with Topic and Social Latent Factors Incorporating Implicit Feedback
abstract
Recommender systems (RSs) provide an effective way of alleviating the information overload problem by selecting personalized items for different users. Latent factors-based collaborative filtering (CF) has become the popular approaches for RSs due to its accuracy and scalability. Recently, online social networks and user-generated content provide diverse sources for recommendation beyond ratings. Although social matrix factorization (Social MF) and topic matrix factorization (Topic MF) successfully exploit social relations and item reviews, respectively; both of them ignore some useful information. In this article, we investigate the effective data fusion by combining the aforementioned approaches. First, we propose a novel model MR3 to jointly model three sources of information (i.e., ratings, item reviews, and social relations) effectively for rating prediction by aligning the latent factors and hidden topics. Second, we incorporate the implicit feedback from ratings into the proposed model to enhance its capability and to demonstrate its flexibility. We achieve more accurate rating prediction on real-life datasets over various state-of-the-art methods. Furthermore, we measure the contribution from each of the three data sources and the impact of implicit feedback from ratings, followed by the sensitivity analysis of hyperparameters. Empirical studies demonstrate the effectiveness and efficacy of our proposed model and its extension.
Guang-Neng Hu, Xinyu Dai, Feng-Yu Qiu, Tao Li 0001, Shujian Huang, Jiajun Chen 0001
ACM Trans. Knowl. Discov. Data1
2017 Integrating Reviews into Personalized Ranking for Cold Start Recommendation
Guang-Neng Hu, Xinyu Dai
PAKDD (2)1
2015 A Synthetic Approach for Recommendation: Combining Ratings, Social Relations, and Reviews
Guang-Neng Hu, Xinyu Dai, Yunya Song, Shujian Huang, Jiajun Chen 0001
IJCAI1