Huan Rong

dblp:175/6356 · DBLP profile ↗
← Back
41ranked-venue papers
10as first author
38since 2021 · last 2026
0000-0003-0542-1827ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 24 · 3 first-author · 23 since 2021Databases, data management, data science and information retrieval · 7 · 7 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Systems, architecture and hardware · 3 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Personalized dialogue generation through knowledge expansion and in-context learning
Zhewen Wang, Tinghuai Ma, Huan Rong
Appl. Intell.3
2026 ReST-Pre: Event prediction by spatial-temporal structural replay on generative implicit event pattern induction
Tinghuai Ma, Huan Rong
Expert Syst. Appl.3
2026 Mixed-order relation learning spatio-temporal graph neural network for weather forecasting
Yuming Su, Tinghuai Ma, Huan Rong, Baobao Pan, Xuejian Huang, Mohamed Magdy Abdel Wahab
Expert Syst. Appl.3
2026 Brain-TVR: A multi-granularity structure-aware text-video retrieval framework inspired by human brain episodic memory neural mechanisms
Huan Rong, Canyang Liu
Neurocomputing2
2026 BKUF: A Novel Real-time Rumor Detection Method Integrating Background Knowledge and User Features
abstract
Real-time rumor detection methods that do not rely on propagation features have emerged as an effective strategy to curb the spread of misinformation. To address the pressing challenge of enhancing semantic understanding of short texts and extracting latent user features in real-time rumor detection, this article proposes a novel approach that integrates B ackground K nowledge and U ser F eatures (BKUF). First, relevant background knowledge is extracted from an external knowledge graph through knowledge distillation. To accommodate different granularities of knowledge, we design two fusion strategies: one based on graph attention networks and the other on co-attention mechanisms, effectively enriching the semantic representation of the text. In addition to traditional user features, we further introduce two novel latent user attributes—rationality and professionalism—which are inferred from users’ historical posts. Finally, the enhanced semantic and user features are adaptively integrated and passed into a multi-layer perceptron for classification. Experiments conducted on four widely used public rumor datasets—Weibo, PHEME, Twitter15, and Twitter16—show that our method achieves accuracies of 92.8%, 84.9%, 81.5%, and 82.7%, respectively, outperforming state-of-the-art baselines.
Xuejian Huang, Tinghuai Ma, Huan Rong, Gan Zhou, Najla Al-Nabhan
ACM Trans. Asian Low Resour. Lang. Inf. Process.3
2026 BiCaution: Bridging Forward-Backward Interventional and Counterfactual Causal Establishment for Event Graph-Based Abductive Reasoning
abstract
Many complex systems like our society can be considered as the dependency among a series of events, where the event graph can properly depict the uncertainty by branches aggregating into or separating from event nodes. Consequently, targeting the decisive node in event graph as the plausible hypothesis to explain causality on occurrence between indirectly connected events (i.e.,event abductive reasoning) can facilitate evolutionary pattern mining. Previous works focused on extrapolating the best hypothesis based on observed associations between events. However, causal observations at higher causality levels that can provide additional causal information are still underutilized. To obtain more complete causal establishment on the whole event graph, we propose a more challenging task called the event graph-based abductive reasoning (EGAR), which may suffer fromnonmonotonicdefeasibility (i.e., seem to be correct may not necessarily be right), along with the chain-based transition problem due to the varying contexts between intermediary nodes. Therefore, we proposeBiCautionto resolve EGAR task and thoroughly mine thelatent causality in graph. Specifically, we model the abstract probabilistic graph as prototype, on which the pearl causal hierarchy (PCH) has been imposed across the causation ladder ofassociational,interventional, andcounterfactual. Thecore innovationofBiCautionlies in its “three-anchor, multijump” mechanism, which navigatesforward/backwardpaths by projecting event triples to different causality levels, and aggregating new observations on causal establishment into the representation of original triples to be scored. Benefiting from the above principle of in-graph multilayer observation on causal establishment, our proposedBiCautioncan not only outperform existing counterparts on graph-based event abductive reasoning in different graph sizes but can also resist different causality “error” characteristics, with the capacity to process long-chain event abductive reasoning.
Huan Rong, Tinghuai Ma, Yongyi Jiang
IEEE Trans. Comput. Soc. Syst.2
2026 Knowledge-Enhanced Dynamic Scene Graph Attention Network for Fake News Video Detection
abstract
With the rapid rise of short video social platforms, the spread of fake news videos has become a global challenge. Short videos, which integrate multiple modalities such as text, images, and audio, have a powerful visual and auditory impact, making fake news more prone to widespread dissemination and causing serious societal consequences. However, the complex fusion of multimodal information in fake news videos, coupled with editing artifacts that often blur the distinction between real and fake content, presents considerable challenges to traditional detection methods. To address these challenges, this paper proposes a fake news video detection method based on the Knowledge-Enhanced Dynamic Scene Graph Attention Network (KDSGAT). This method captures temporal correlations and local semantic differences in visual scenes by leveraging dynamic scene graph networks, while enhancing semantic understanding through knowledge distillation from external knowledge graphs. Specifically, we first use pre-trained models such as BERT, HuBERT, and Swin Transformer to extract text semantic features, audio emotion features, and visual features, respectively. Next, we apply an unbiased scene graph generation approach to convert keyframes from the video into scene graphs, which are then processed by the dynamic scene graph attention network to capture temporal correlations and local semantic variations within the scene graph sequences. Finally, co-attention is used to interactively fuse multimodal features, enabling precise detection of fake news in videos. We conduct extensive experiments on two real-world datasets from short video social platforms, FakeSV and FakeTT. The results show that our method outperforms state-of-the-art baselines, improving accuracy by 1.86% and 2.68% on the two datasets, respectively. The source code and data are available athttps://github.com/xuejianhuang/KDSGAT-FNVD.
Xuejian Huang, Tinghuai Ma, Hao Tang 0005, Huan Rong
IEEE Trans. Multim.4
2025 TADST: reconstruction with spatio-temporal feature fusion for deviation-based time series anomaly detection
Tinghuai Ma, Huan Rong, Xuejian Huang, Chaoming Wang
Appl. Intell.3
2025 Dual evidence enhancement and text-image similarity awareness for multimodal rumor detection
Xuejian Huang, Tinghuai Ma, Huan Rong, Yuming Su
Eng. Appl. Artif. Intell.3
2025 Multi-axis fusion with optimal transport learning for multimodal aspect-based sentiment analysis
Tinghuai Ma, Huan Rong, Liyuan Gao, Yu-Feng Zhang, Victor S. Sheng
Expert Syst. Appl.3
2025 RTA: A reinforcement learning-based temporal knowledge graph question answering model
Tinghuai Ma, Huan Rong, Yexin Bian
Neurocomputing4
2025 EDCM-EA: event prediction based on event development context mining considering event arguments
Zhiheng Gong, Huan Rong, Zhongfeng Chen, Yixiang Tang, Victor S. Sheng
Multim. Syst.2
2025 Adaptive Graph Structure Learning Neural Rough Differential Equations for Multivariate Time Series Forecasting
abstract
Multivariate time series forecasting has extensive applications in urban computing, such as financial analysis, weather prediction, and traffic forecasting. Using graph structures to model the complex correlations among variables in time series, and leveraging graph neural networks and recurrent neural networks for temporal aggregation and spatial propagation stage, has shown promise. However, traditional methods’ graph structure node learning and discrete neural architecture are not sensitive to issues such as sudden changes, time variance, and irregular sampling often found in real-world data. To address these challenges, we propose a method calledAdaptiveGraph structureLearning neuralRoughDifferentialEquations (AGLRDE). Specifically, we combine dynamic and static graph structure learning to adaptively generate a more robust graph representation. Then we employ a spatio-temporal encoder-decoder based on Neural Rough Differential Equations (Neural RDE) to model spatio-temporal dependencies. Additionally, we introduce a path reconstruction loss to constrain the path generation stage. We conduct experiments on six benchmark datasets, demonstrating that our proposed method outperforms existing state-of-the-art methods. The results show that AGLRDE effectively handles aforementioned challenges, significantly improving the accuracy of multivariate time series forecasting.
Yuming Su, Tinghuai Ma, Huan Rong, Mohamed Magdy Abdel Wahab
IEEE Trans. Big Data3
2025 Multiview Spatio-Temporal Learning With Dual Dynamic Graph Convolutional Networks for Rumor Detection
abstract
Detecting rumors on social networks is increasingly important due to their rapid dissemination and negative societal impact. The structural characteristics of propagation play a crucial role in rumor detection. However, most current graph neural network-based methods focus on spatial structural features, overlooking the temporal structural features or exploring spatio-temporal features from a single perspective, failing to comprehensively and finely learn representations of dynamic events. Therefore, this article proposes a multiview spatio-temporal feature learning method based on dual dynamic graph convolutional networks. First, dynamic graphs of information propagation and user interactions are constructed based on retweet and reply relationships. Second, BERT is utilized to extract semantic features of content, serving as initial node representations for the information propagation graph, while social features of users serve as initial node representations for the user interaction graph. Subsequently, dual graph convolutional networks are employed to learn representations of graph structures at different time steps. Finally, a time fusion unit based on cross-attention is devised to facilitate the learning and fusion of the spatio-temporal features from the two dynamic graphs. Experimental results on two real-world social network rumor datasets, PHEME and Weibo, demonstrate that our method outperforms all compared baseline methods and enables early detection of rumors.
Xuejian Huang, Tinghuai Ma, Wenwen Jin, Huan Rong, Xintong Xie
IEEE Trans. Comput. Soc. Syst.4
2025 Hypergraph-based multimodal adaptive fusion for emotion recognition in conversation
Xintong Xie, Tinghuai Ma, Huan Rong
J. Supercomput.4
2025 CogLign: Interpretable Text Sentiment Determination by Aligning Cognition Between EEG-Derived Brain Graph and Text-Derived Knowledge Graph
abstract
Nowadays, detecting sentiment or emotion from user generated texts has been intensively studied in natural language understanding, especially via neural-based models based on text representation. However, the interpretability on how could the final text sentiment be determined by neural-based text representation has not been thoroughly unfolded yet. Consequently, in this paper, we proposeCogLignwhich injects theneural-cognitionderived from Electroencephalogram (EEG)-signal into theneural-basedtext sentiment analysis model, aimed at learning the activation of brain regions stimulated by different sentiments, so as to guide our proposedCogLignto make proper determination on text sentiment in brain-like way. Specifically, on the one hand, the given videos in different sentiments have been watched bysubjects, during which the EEG-signals are monitored to construct brain connectivity pattern asbrain graph(BG), attaining more obvious sentiment response on brain region activation forneural-cognition. On the other hand, we interpret the video-plots (or video-semantics) along timeline into text, where the entire video-interpreted-text will bestrictly boundwith the wholeEEG-signal-sequencebysegmentvia the fixed size oftime-window. Then, entities and relations are extracted from the video-interpreted-text to constructknowledge graph(KG), depicting text semantics. Next, mapping fromentities(or nodes) inKGtoEEG-Electrodes(or nodes) inBG, further dated back to different brain regions, has been learned viacognition alignmentbetween the EEG-derivedBGand text-derivedKG. In this way, by aligningneural cognitionfrombrain graphwith thesemantic cognitionfromknowledge graph, our proposed frameworkCogLigncan not only achieve the overall best sentiment analysis performance on thevideo-interpreted-text, but can also detect brain connectivity patterns in different sentiments more consistent with the prior conclusion of brain region sentiment preference, revealing competitiveinterpretabilityon text sentiment determination.
Huan Rong, Wenxuan Ji, Tinghuai Ma, Weiyi Ding, Victor S. Sheng
IEEE Trans. Knowl. Data Eng.1
2024 CLFFRD: Curriculum Learning and Fine-grained Fusion for Multimodal Rumor Detection
abstract
In an era where rumors can propagate rapidly across social media platforms such as Twitter and Weibo, automatic rumor detection has garnered considerable attention from both academia and industry. Existing multimodal rumor detection models often overlook the intricacies of sample difficulty, e.g., text-level difficulty, image-level difficulty, and multimodal-level difficulty, as well as their order when training. Inspired by the concept of curriculum learning, we propose the Curriculum Learning and Fine-grained Fusion-driven multimodal Rumor Detection (CLFFRD) framework, which employs curriculum learning to automatically select and train samples according to their difficulty at different training stages. Furthermore, we introduce a fine-grained fusion strategy that unifies entities from text and objects from images, enhancing their semantic cohesion. We also propose a novel data augmentation method that utilizes linear interpolation between textual and visual modalities to generate diverse data. Additionally, our approach incorporates deep fusion for both intra-modality (e.g., text entities and image objects) and inter-modality (e.g., CLIP and social graph) features. Extensive experimental results demonstrate that CLFFRD outperforms state-of-the-art models on both English and Chinese benchmark datasets for rumor detection in social media.
Fan Xu 0002, Bowei Zou, AiTi Aw, Huan Rong
LREC/COLING5
2024 Multi-modal anchor adaptation learning for multi-modal summarization
Zhongfeng Chen, Zhenyu Lu 0002, Huan Rong, Chuanjun Zhao, Fan Xu 0002
Neurocomputing3
2024 Automatic medical report generation combining contrastive learning and feature difference
Chongwen Lyu, Cheng-Jian Qiu, Kai Han 0006, Saisai Li, Victor S. Sheng, Huan Rong, Yuqing Song 0001, Yi Liu 0114, Zhe Liu 0004
Knowl. Based Syst.6
2024 KGCDP-T: Interpreting knowledge graphs into text by content ordering and dynamic planning with three-level reconstruction
Huan Rong, Tinghuai Ma, Di Jin 0001, Victor S. Sheng
Knowl. Based Syst.1
2024 Multization: Multi-Modal Summarization Enhanced by Multi-Contextually Relevant and Irrelevant Attention Alignment
abstract
This article focuses on the task of Multi-Modal Summarization with Multi-Modal Output for China JD.COM e-commerce product description containing both source text and source images. In the context learning of multi-modal (text and image) input, there exists a semantic gap between text and image, especially in the cross-modal semantics of text and image. As a result, capturing shared cross-modal semantics earlier becomes crucial for multi-modal summarization. However, when generating the multi-modal summarization, based on the different contributions of input text and images, the relevance and irrelevance of multi-modal contexts to the target summary should be considered, so as to optimize the process of learning cross-modal context to guide the summary generation process and to emphasize the significant semantics within each modality. To address the aforementioned challenges, Multization has been proposed to enhance multi-modal semantic information by multi-contextually relevant and irrelevant attention alignment. Specifically, a Semantic Alignment Enhancement mechanism is employed to capture shared semantics between different modalities (text and image), so as to enhance the importance of crucial multi-modal information in the encoding stage. Additionally, the IR-Relevant Multi-Context Learning mechanism is utilized to observe the summary generation process from both relevant and irrelevant perspectives, so as to form a multi-modal context that incorporates both text and image semantic information. The experimental results in the China JD.COM e-commerce dataset demonstrate that the proposed Multization method effectively captures the shared semantics between the input source text and source images, and highlights essential semantics. It also successfully generates the multi-modal summary (including image and text) that comprehensively considers the semantics information of both text and image.
Huan Rong, Zhongfeng Chen, Zhenyu Lu 0002, Fan Xu 0002, Victor S. Sheng
ACM Trans. Asian Low Resour. Lang. Inf. Process.1
2024 GCMA: An Adaptive Multiagent Reinforcement Learning Framework With Group Communication for Complex and Similar Tasks Coordination
abstract
Coordinating multiple agents with diverse tasks and changing goals without interference is a challenge. Multi-Agent Reinforcement Learning (MARL) aims to develop effective communication and joint policies using group learning. Some of the previous approaches required each agent to maintain a set of networks independently, resulting in no consideration of interactions. Joint communication work causes agents receiving information unrelated to their own tasks. Currently, agents with different task divisions are often grouped by action tendency, but this can lead to poor dynamic grouping. This paper presents a two-phase solution for multiple agents, addressing these issues. The first phase develops heterogeneous agent communication joint policies using a Group Communication MARL framework (GCMA). The framework employs a periodic grouping strategy, reducing exploration and communication redundancy by dynamically assigning agent group hidden features through hyper-network and graph communication. The scheme efficiently utilizes resources for adapting to multiple similar tasks. In the second phase, each agent's policy network is distilled into a generalized simple network, adapting to similar tasks with varying quantities and sizes. GCMA is tested in complex environments like StarCraft II and UAV take-off, showing its well-performing for large-scale, coordinated tasks. It shows GCMA's effectiveness for solid generalization in multi-task tests with simulated pedestrians.
Kexing Peng, Tinghuai Ma, Huan Rong, Yurong Qian, Najla Al-Nabhan
IEEE Trans. Games4
2024 Enhancing Collaboration in Heterogeneous Multiagent Systems Through Communication Complementary Graph
abstract
Heterogeneous multiagent systems are characterized by diverse task distributions, which are prevalent in practical scenarios, such as distributed decision making and robotic collaboration. A significant challenge in these systems is the constraint of limited observations, where each agent has access only to partial information. Many studies facilitate information exchange by employing shared parameters among agents. However, this approach is generally more effective for homogeneous systems where agents have similar observation or action spaces. In heterogeneous systems, indiscriminate parameter sharing can significantly increase the exploration cost required for effective adaptation. To address this challenge, we propose a novel communication complementary graph model (CCGM) for enhancing collaboration in heterogeneous multiagent systems. Our approach builds upon the training framework of heterogeneous agent reinforcement learning (HARL) with trust region learning and nonparameter sharing. This model utilizes advantage function decomposition and sequential updates to promote policy convergence. Within this framework, we introduce a novel communication method inspired by signaling games, where agents acting as receivers, process messages from other agents alongside their own observations. CCGM aligns the messages with observations in a graph-based communication module, which establishes communication relationships and supplements observational information. Subsequently, agents generate self-interested information, which they then share with others as senders. We evaluate our algorithm across various environments, including multiagent particle environments (MPE) and multiagent MuJoCo (MAMuJoCo) robot experiments. The results demonstrate the effectiveness of CCGM in enhancing HARL-based algorithms.
Kexing Peng, Tinghuai Ma, Huan Rong
IEEE Trans. Cybern.4
2024 Temporal patterns decomposition and Legendre projection for long-term time series forecasting
Tinghuai Ma, Yuming Su, Huan Rong, Alaa Abd El-Raouf Mohamed Khalil, Mohamed Magdy Abdel Wahab, Benjamin Kwapong Osibo
J. Supercomput.4
2024 CoBjeason: Reasoning Covered Object in Image by Multi-Agent Collaboration Based on Informed Knowledge Graph
abstract
Object detection is a widely studied problem in existing works. However, in this paper, we turn to a more challenging problem of “ Covered Object Reasoning ”, aimed at reasoning the category label of target object in the given image particularly when it has been totally covered (or invisible ). To resolve this problem, we propose CoBjeason to seize the opportunity when visual reasoning meets the knowledge graph, where “ empirical cognition ” on common visual contexts have been incorporated as knowledge graph to conduct reinforced multi-hop reasoning via two collaborative agents. Such two agents, for one thing, stand at the covered object (or unknown entity ) to observe the surrounding visual cues in the given image and gradually select entities and relations from the global gallery-level knowledge graph which contains entity-pairs frequently occurring across the entire image-collection, so as to infer the main structure of image-level knowledge graph forward expanded from the unknown entity . In turn, for another, based on the reasoned image-level knowledge graph, the semantic context among entities will be aggregated backward into unknown entity to select an appropriate entity from the global gallery-level knowledge graph as the reasoning result. Moreover, such two agents will collaborate with each other, securing that the above Forward & Backward Reasoning will step towards the same destination of the higher performance on covered object reasoning. To our best knowledge, this is the first work on Covered Object Reasoning with Knowledge Graphs and reinforced Multi-Agent collaboration. Particularly, our study on Covered Object Reasoning and the proposed model CoBjeason could offer novel insights into more basic Computer Vision (CV) tasks, such as Semantic Segmentation with better understanding on the current scene when some objects are blurred or covered, Visual Question Answering with enhancement on the inference in more complicated visual context when some objects are covered or invisible, and Image Caption Generation with the augmentation on the richness of visual context for images containing partially visible objects. The improvement on the above basic CV tasks can further refine more complicated ones involved with nuanced visual interpretation like Autonomous Driving, where the recognition and reasoning on partially visible or covered object are critical. According to the experimental results, our proposed CoBjeason can achieve the best overall ranking performance on covered object reasoning compared with other models, meanwhile enjoying the advantage of lower “ exploration cost ”, with the insensitivity against the long-tail covered objects and the acceptable time complexity.
Huan Rong, Minfeng Qian, Tinghuai Ma, Di Jin 0001, Victor S. Sheng
ACM Trans. Knowl. Discov. Data1
2024 Three-stage Transferable and Generative Crowdsourced Comment Integration Framework Based on Zero- and Few-shot Learning with Domain Distribution Alignment
abstract
Online shopping has become a crucial way to encourage daily consumption, where the User-generated, or crowdsourced product comments, can offer a broad range of feedback on e-commerce products. As a result, integrating critical opinions or major attitudes from the crowdsourced comments can provide valuable feedback for marketing strategy adjustment or product-quality monitoring. Unfortunately, the scarcity of annotated ground truth on the integrated comment, or the limited gold integration reference, has incurred the infeasibility of the regular supervised-learning-based comment integration. To resolve this problem, in this article, inspired by the principle of Transfer Learning, we propose a three-stage transferable and generative crowdsourced comment integration framework ( TTGCIF ) based on zero-and-few-shot learning with the support of domain distribution alignment. The proposed framework aims at generating abstractive integrated comment in target domain via the enhanced neural text generation model, by referring the available integration resource in related source domains, to avoid the exhausted effort on resource annotation devoted to the target domain. Specifically, at the first stage, to enhance the domain transferability, representations on the crowdsourced comments have been aligned up between the source and target domain, by minimizing the domain distribution discrepancy in the kernel space. At the second stage, Zero-shot comment integration mechanism has been adopted to deal with the dilemma that none of the gold integration reference may be available in target domain. In other words, taking the sample-level semantic prototype as input, the enhanced neural text generation model in TTGCIF is trained to learn data semantic association among different domains via semantic prototype transduction, so that the “ unlabeled ” crowdsourced comments in target domain can be associated with existing integration references in related source domains. At the third stage, based on the parameters trained at the second stage, fast domain adaptation mechanism in a Few-shot manner has also been adopted by seeking most potential parameters along the gradient direction constrained by instances across multiple source domains. In this way, parameters in TTGCIF can be sensitive to any alteration on training data, ensuring that even if only few annotated resource in target domain are available for “Fine-tune,” TTGCIF can still react promptly to achieve effective target domain adaptation. According to the experimental results, TTGCIF can achieve the best transferable product comment integration performance in target domain, with fast and stable domain adaption effect depending on no more than 10% annotated resource in target domain. More importantly, even if TTGCIF has not been fine-tuned on the target domain, yet by referring to the available integration resource in related source domains, the integrated comments generated by TTGCIF on the target domain are still superior to those generated by models already fine-tuned on the target domain.
Huan Rong, Tinghuai Ma, Victor S. Sheng, Yang Zhou 0001, Mznah Al-Rodhaan
ACM Trans. Knowl. Discov. Data1
2024 FuFaction: Fuzzy Factual Inconsistency Correction on Crowdsourced Documents With Hybrid-Mask at the Hidden-State Level
abstract
Nowadays, crowdsourced documents like Wikipedia pages and comments on products are all over the Internet. However, documents generated by crowdsourcing participants may contain inconsistent facts, implicit semantics and fabricated contents, thus threatening the trustworthiness of information content security available in Internet. To address this problem, we propose FuFaction, enabled by an enhanced observation mechanism based on the notion of hybrid-mask consisting of a hard-mask and a soft-mask, to eliminate factual inconsistencies on crowdsourced documents at the hidden-state level (or in a fuzzy way), according to the given evidence retrieved from an external open domain. Specifically, instead of focusing on a specific category of factual inconsistency, FuFaction captures anomalous hidden-states between a crowdsourced document and evidence obtained via a reverse-attention mechanism, where a hard-mask controls the attending direction as bidirectional and unidirectional for better understanding on semantics. Then, a soft-mask is generated with the help of the hard-masked reverse-attention to revise or mask anomalous hidden-states on the crowdsourced document. Afterwards, the masked hidden-states are further refined by a cross reverse-attention and factual consistency reinforcement strategy, based on which a new crowdsourced document with higher factual consistency is generated via neural text generation. According to our experimental results, FuFaction can effectively deal with the fuzzy factual inconsistencies on crowdsourced documents, achieving the overall best performance in terms of factual consistency metrics with a little higher (yet still competitive) editing cost on literal vocabulary, so as to reflect factually consistent semantics supported by the given evidence.
Huan Rong, Gongchi Chen, Tinghuai Ma, Victor S. Sheng, Elisa Bertino
IEEE Trans. Knowl. Data Eng.1
2023 A Self-play and Sentiment-Emphasized Comment Integration Framework Based on Deep Q-Learning in a Crowdsourcing Scenario : Extended Abstract
abstract
Crowdsourcing is a sourcing model where individuals or organizations obtain goods and services from a large, relatively open and often rapidly evolving group of internet users. The most common way that crowdsourcing can facilitate machine learning is to annotate instances with labels [1] . However, the same instance may have inconsistent class labels, in the eyes of various annotators. Therefore, current efforts in crowdsourcing mainly focus on the truth inference or label integration, to remove inconsistent labels or to alleviate biased labeling. In turn, instances with the integrated labels could facilitate the training on machine learning models. The future direction of crowdsourcing is to apply more fine-grained truth inference methods to different application domains [2] . Consequently, we evolve toward another challenging problem of comment integration. That is, how can we integrate or summarize the core opinions of multiple product comments obtained from users, rather than the discrete labels.
Huan Rong, Victor S. Sheng, Tinghuai Ma, Yang Zhou 0001, Mznah Al-Rodhaan
ICDE1
2023 Cascaded multi-point regression Network for high-quality generic lesion detection
Huan Rong, Victor S. Sheng, Yuqing Song 0001, Cheng-Jian Qiu, Kai Han 0006, Zhe Liu 0004
Expert Syst. Appl.2
2023 An effective multimodal representation and fusion method for multimodal intent recognition
Xuejian Huang, Tinghuai Ma, Huan Rong, Najla Al-Nabhan
Neurocomputing5
2023 A privacy-preserving trajectory data synthesis framework based on differential privacy
Tinghuai Ma, Huan Rong, Najla Al-Nabhan
J. Inf. Secur. Appl.3
2023 AGRCNet: communicate by attentional graph relations in multi-agent reinforcement learning for traffic signal control
Tinghuai Ma, Kexing Peng, Huan Rong, Yurong Qian
Neural Comput. Appl.3
2023 SPK-CG: Siamese Network based Posterior Knowledge Selection Model for Knowledge Driven Conversation Generation
abstract
Building a human-computer conversational system that can communicate with humans is a research hotspot in the field of artificial intelligence. Traditional dialogue systems tend to produce irrelevant and non-information responses, which reduce people’s interest in engaging in a conversation. This often leads to boring conversations. To alleviate this problem, many researchers use external knowledge to assist conversation generation. The accuracy of knowledge selection is the prerequisite to ensure the quality of knowledge conversation. This approach has worked positively to a certain extent, but generally only searches knowledge information based on entity words themselves, without considering the specific conversation context. Therefore, if irrelevant knowledge is retrieved, the quality of conversation generation will be reduced. Motivated by this, we propose a novel neural knowledge-based conversation generation model, namedSiamese Network based Posterior Knowledge Selection Model for Knowledge Driven Conversation Generation (SPK-CG). We have designed a novel knowledge selection mechanism to obtain knowledge information that is highly relevant to the context of the conversation. Specifically, the posterior knowledge distribution is used as a soft label to make the prior distribution consistent with the posterior distribution in the training process. At the same time, in order to narrow the gap between prior and posterior distributions and improve the accuracy of knowledge selection, we leverage siamese network and design multi-granularity matching module for knowledge selection. Compared with previous knowledge-based models, our method can select more appropriate knowledge and use the selected knowledge to generate responses that are more relevant to the conversation context. Extensive automatic and human evaluations demonstrate that our model has advantages over previous baselines.
Tinghuai Ma, Huan Rong, Najla Al-Nabhan
ACM Trans. Asian Low Resour. Lang. Inf. Process.3
2022 MDMN: Multi-task and Domain Adaptation based Multi-modal Network for early rumor detection
Honghao Zhou, Tinghuai Ma, Huan Rong, Yurong Qian, Yuan Tian 0003, Najla Al-Nabhan
Expert Syst. Appl.3
2022 A Novel Sentiment Polarity Detection Framework for Chinese
abstract
Nowadays, mining opinions or sentiment from online user-generated text has become a research hot spot. Although a large amount of lexicon-based Chinese polarity detection works have been done, the existing methods have one common flaw: that even the same word can have opposite polarities among different seed lexicons. This is known as polarity fuzziness. To enhance the performance of Chinese sentiment polarity detection, we start from a two-aspect lexicon expansion so that the polarity fuzziness can be avoided. Specifically, we detect sentiment polarity for new words and revise sentiment polarity for words already defined in seed lexicons. Then, we formulate a novel sentiment polarity detection framework for Chinese (SPDFC) with more attention to fine-grained sentiment processing, which is involved in symmetrical mapping, sentiment feature pruning and text representation. In this way, words’ polarity can be directly taken as features, penetrating further in the polarity detection phase. According to our experimental results, the proposed SPDFC framework can achieve the best overall performance from the perspective of Chinese polarity detection, sentiment feature pruning, and text representation compared to other classical and state-of-the-art methods.
Tinghuai Ma, Huan Rong, Yongsheng Hao, Jie Cao 0011, Yuan Tian 0003, Mznah Al-Rodhaan
IEEE Trans. Affect. Comput.2
2022 T-BERTSum: Topic-Aware Text Summarization Based on BERT
abstract
In the era of social networks, the rapid growth of data mining in information retrieval and natural language processing makes automatic text summarization necessary. Currently, pretrained word embedding and sequence to sequence models can be effectively adapted in social network summarization to extract significant information with strong encoding capability. However, how to tackle the long text dependence and utilize the latent topic mapping has become an increasingly crucial challenge for these models. In this article, we propose a topic-aware extractive and abstractive summarization model named T-BERTSum, based on Bidirectional Encoder Representations from Transformers (BERTs). This is an improvement over previous models, in which the proposed approach can simultaneously infer topics and generate summarization from social texts. First, the encoded latent topic representation, through the neural topic model (NTM), is matched with the embedded representation of BERT, to guide the generation with the topic. Second, the long-term dependencies are learned through the transformer network to jointly explore topic inference and text summarization in an end-to-end manner. Third, the long short-term memory (LSTM) network layers are stacked on the extractive model to capture sequence timing information, and the effective information is further filtered on the abstractive model through a gated network. In addition, a two-stage extractive–abstractive model is constructed to share the information. Compared with the previous work, the proposed model T-BERTSum focuses on pretrained external knowledge and topic mining to capture more accurate contextual representations. Experimental results on the CNN/Daily mail and XSum datasets demonstrate that our proposed model achieves new state-of-the-art results while generating consistent topics compared with the most advanced method.
Tinghuai Ma, Huan Rong, Yurong Qian, Yuan Tian 0003, Najla Al-Nabhan
IEEE Trans. Comput. Soc. Syst.3
2022 A Self-Play and Sentiment-Emphasized Comment Integration Framework Based on Deep Q-Learning in a Crowdsourcing Scenario
abstract
Crowdsourcing is a hotspot research field which can facilitate machine learning by collecting labels to train models. Consequently, the state-of-the-art research efforts in crowdsourcing focus on truth inference or label integration, to remove inconsistent labels or to alleviate biased labeling. In turn, the integrated labels will be used to fine-tune machine learning models. Particularly, in this paper, we change the target of truth inference in crowdsourcing from discrete labels to multiple comments given by online participants, that is, the integration of the crowdsourced comments. For such a goal, we propose aSelf-play andSentiment-EmphasizedCommentIntegrationFramework (SSECIF), based on deepQ-learning, with three unique features. First, our framework SSECIF can generate the comment integration in a totally self-play way, without relying on the ground truth generated by human effort. Second, the integrated comment generated by SSECIF can include salient content with low redundancy. Third, the proposed framework SSECIF has emphasized, with a higher intensity, the sentiment in the integrated comment, in order to reflect the attitude or opinion more obviously. Extensive evaluation on real-world datasets demonstrates that SSECIF has achieved the best overall performance in terms of both effectiveness and efficiency, compared with the state-of-the-art methods.
Huan Rong, Victor S. Sheng, Tinghuai Ma, Yang Zhou 0001, Mznah Al-Rodhaan
IEEE Trans. Knowl. Data Eng.1
2021 Dual-path CNN with Max Gated block for text-based person re-identification
Tinghuai Ma, Huan Rong, Yurong Qian, Yuan Tian 0003, Najla Al-Nabhan
Image Vis. Comput.3
2019 Deep rolling: A novel emotion prediction model for a multi-participant communication context
Huan Rong, Tinghuai Ma, Jie Cao 0011, Yuan Tian 0003, Abdullah Al-Dhelaan, Mznah Al-Rodhaan
Inf. Sci.1
2018 A novel subgraph K+ -isomorphism method in social network based on graph similarity detection
Huan Rong, Tinghuai Ma, Meili Tang, Jie Cao 0011
Soft Comput.1
2016 Detect structural-connected communities based on BSCHEF in C-DBLP
abstract
Summary Chinese Digital Bibliography & Library Project (C‐DBLP) is a huge and real‐life co‐author social network in China, rarely cited by published paper. It contains a large amount of ground‐truth community structure with distinguished research topics. Despite the fact that rich studies on community detection have been conducted with gains of practically fruitful algorithms, unfortunately, with the coming of ‘Big Data’ era and speedy development of mobile devices, social networks like C‐DBLP have incredibly expanded on nodes and edges, as a result, because of massive data cardinality, a large portion of community detection methods consume memory resource excessively. Therefore, in this work, we select Based on Structural Connection Hierarchical Exploration (BSCHE) algorithm to partition nodes in C‐DBLP because of its O(n) time cost, fast enough to process massive data, and its novel physical meaning of similarity between nodes defined by structural connection and availability. In addition, in order to avoid huge memory resource consumption caused by ‘Big Data’ of C‐DBLP, we strengthen BSCHE as a framework (BSCHEF) by our proposed ‘count‐pointer‐strategy’ imitated from incremental batch process to detect co‐author communities on C‐DBLP. The experiment results show that BSCHEF can find sets of communities onC‐DBLPmore effectively with the highest modularity value and the least execution time compared to other clustering algorithm. Copyright © 2015 John Wiley & Sons, Ltd.
Tinghuai Ma, Huan Rong, Changhong Ying, Yuan Tian 0003, Abdullah Al-Dhelaan, Mznah Al-Rodhaan
Concurr. Comput. Pract. Exp.2