Hua Xu 0003

dblp:31/4114-3 · DBLP profile ↗
← Back
86ranked-venue papers
3as first author
35since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 70 · 3 first-author · 25 since 2021Graphics, computer vision, multimedia, augmented reality and games · 20 · 14 since 2021Databases, data management, data science and information retrieval · 6 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Ellipsoid-Based Decision Boundaries for Open Intent Classification
abstract
Textual open intent classification is crucial for real-world dialogue systems, enabling robust detection of unknown user intents without prior knowledge and contributing to the robustness of the system. While adaptive decision boundary methods have shown great potential by eliminating manual threshold tuning, existing approaches assume isotropic distributions of known classes, restricting boundaries to balls and overlooking distributional variance along different directions. To address this limitation, we propose EliDecide, a novel method that learns ellipsoid decision boundaries with varying scales along different feature directions. First, we employ supervised contrastive learning to obtain a discriminative feature space for known samples. Second, we apply learnable matrices to parameterize ellipsoids as the boundaries of each known class, offering greater flexibility than spherical boundaries defined solely by centers and radii. Third, we optimize the boundaries via a novelly designed dual loss function that balances empirical and open-space risks: expanding boundaries to cover known samples while contracting them against synthesized pseudo-open samples. Our method achieves state-of-the-art performance on multiple text intent benchmarks and further on a question classification dataset. The flexibility of the ellipsoids demonstrates superior open intent detection capability and strong potential for generalization to more text classification tasks in diverse complex open-world scenarios.
Yuetian Zou, Hanlei Zhang, Hua Xu 0003, Long Xiao
AAAI3
2026 HEQP: A Hypergraph Neural Network-Based Evolutionary Method for Large-Scale QCQPs
abstract
Machine learning-based optimization frameworks have attracted increasing attention for accelerating the solution of large-scale quadratically constrained quadratic programs (QCQPs) by exploiting shared problem structure across instances. However, existing machine learning (ML) frameworks often rely on the assumption of parametric models and large-scale solvers. This article introduces HEQP, a hypergraph neural network-based evolutionary optimization framework for large-scale QCQPs. This framework features two main components: 1) hypergraph-based neural prediction, which predicts optimal solutions for QCQPs without assumptions of models; and 2) evolutionary large neighborhood search (Evo-LNS), which employs a McCormick relaxation-based repair strategy to search and apply crossover on neighborhood solutions using a small-scale solver. We further show that our framework is equivalent to the interior-point method (IPM), a polynomial-time algorithm, for quadratic programming. Experiments on two types of benchmark problems and 13 large-scale real-world instances from the QPLIB illustrate that our framework outperforms state-of-the-art solvers (including Gurobi, SCIP, and SHOT) in both solution quality and time efficiency, highlighting the efficiency of ML-based optimization frameworks for QCQPs.
Zhixiao Xiong, Huigen Ye, Hua Xu 0003, Carlos A. Coello Coello
IEEE Trans. Cybern.3
2026 Light-EvoOPT: A Lightweight Evolutionary Optimization Framework for Ultralarge-Scale Mixed Integer Linear Programs
abstract
Machine Learning (ML)-based optimization frameworks emerge as a promising technique for solving large-scale Mixed Integer Linear Programs (MILPs), as they can capture the mapping between problem structures and optimal solutions to expedite their solution process. However, existing solution frameworks often suffer from high model computation costs, incomplete problem reduction, and reliance on large-scale solvers, leading to performance bottlenecks in ultra-large-scale problems with complex constraints. To address these issues, this paper proposes Light-EvoOPT, a Lightweight Evolutionary Optimization Framework for Ultra-Large-Scale Mixed Integer Linear Programs, which can be divided into four stages: (1) Problem Formulation for problem division to reduce model computational costs, (2) Model-based Initial Solution Prediction for predicting and constructing the initial solution using a small-scale training dataset, (3) Problem Reduction for both variable and constraint reduction, and (4) Evolutionary Optimization for current solution improvement employing a lightweight optimizer. Experiments on four benchmark datasets with tens of millions of variables and constraints and a real-world problem show that the proposed framework based on the sole use of a lightweight optimizer, trained on only one-thousandth of the scale of ultra-large-scale problems, is able to outperform state-of-the-art ML-based frameworks and advanced solvers (e.g. Gurobi) within a specified computational time, validating the feasibility and effectiveness of our proposed ML-based evolutionary optimization framework for ultra-large-scale MILPs.
Huigen Ye, Hua Xu 0003, Carlos A. Coello Coello
IEEE Trans. Evol. Comput.2
2025 LLM-Guided Semantic Relational Reasoning for Multimodal Intent Recognition
abstract
Understanding human intents from multimodal signals is critical for analyzing human behaviors and enhancing human-machine interactions in real-world scenarios.However, existing methods exhibit limitations in their modality-level reliance, constraining relational reasoning over fine-grained semantics for complex intent understanding.This paper proposes a novel LLM-Guided Semantic Relational Reasoning (LGSRR) method, which harnesses the expansive knowledge of large language models (LLMs) to establish semantic foundations that boost smaller models' relational reasoning performance.Specifically, an LLM-based strategy is proposed to extract fine-grained semantics as guidance for subsequent reasoning, driven by a shallow-to-deep Chain-of-Thought (CoT) that autonomously uncovers, describes, and ranks semantic cues by their importance without relying on manually defined priors.Besides, we formally model three fundamental types of semantic relations grounded in logical principles and analyze their nuanced interplay to enable more effective relational reasoning.Extensive experiments on multimodal intent and dialogue act recognition tasks demonstrate LGSRR's superiority over state-of-theart methods, with consistent performance gains across diverse semantic understanding scenarios.The complete data and code are available at https://github.com/thuiar/LGSRR.
Qianrui Zhou, Hua Xu 0003, Xinzhi Dong, Hanlei Zhang
EMNLP2
2025 MILPBench: A Large-scale Benchmark Test Suite for Mixed Integer Linear Programming Problems
abstract
Mixed-integer linear programming (MILP) is a cornerstone of optimization with applications across numerous domains. However, the development and evaluation of MILP-solving algorithms are hindered by existing benchmark datasets, which are often limited in scale, lack diversity, and are poorly structured, making them inadequate for systematic testing across different solving approaches, especially for machine learning (ML)-based methods. To address these issues, we introduce MILPBench, a large-scale benchmark suite comprising 100,000 MILP instances organized into 60 well-categorized classes. Using structural properties and embedding similarity metrics, we developed a novel classification framework to ensure both intra-class homogeneity and inter-class diversity. In addition to the dataset, MILPBench includes a comprehensive baseline library featuring 15 mainstream solving methods, spanning traditional solvers, heuristic algorithms, and ML-based approaches. This design enables rigorous and standardized evaluation of MILP-solving algorithms under diverse conditions. Extensive benchmarking demonstrates the utility of MILPBench as a scalable and versatile testbed for advancing MILP research, fostering innovation in solver development, and bridging the gap between optimization and machine learning.
Huigen Ye, Yaoyang Cheng, Hua Xu 0003, Zhiguang Cao, Hanzhang Qin
GECCO3
2025 Large Language Model-driven Large Neighborhood Search for Large-Scale MILP Problems
abstract
Large Neighborhood Search (LNS) is a widely used method for solving large-scale Mixed Integer Linear Programming (MILP) problems. The effectiveness of LNS crucially depends on the choice of the search neighborhood. However, existing strategies either rely on expert knowledge or computationally expensive Machine Learning (ML) approaches, both of which struggle to scale effectively for large problems. To address this, we propose LLM-LNS, a novel Large Language Model (LLM)-driven LNS framework for large-scale MILP problems. Our approach introduces a dual-layer self-evolutionary LLM agent to automate neighborhood selection, discovering effective strategies with scant small-scale training data that generalize well to large-scale MILPs. The inner layer evolves heuristic strategies to ensure convergence, while the outer layer evolves evolutionary prompt strategies to maintain diversity. Experimental results demonstrate that the proposed dual-layer agent outperforms state-of-the-art agents such as FunSearch and EOH. Furthermore, the full LLM-LNS framework surpasses manually designed LNS algorithms like ACP, ML-based LNS methods like CL-LNS, and large-scale solvers such as Gurobi and SCIP. It also achieves superior performance compared to advanced ML-based MILP optimization frameworks like GNN&GBDT and Light-MILPopt, further validating the effectiveness of our approach.
Huigen Ye, Hua Xu 0003, Yaoyang Cheng
ICML2
2025 Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark
abstract
Multimodal language analysis is a rapidly evolving field that leverages multiple modalities to enhance the understanding of high-level semantics underlying human conversational utterances. Despite its significance, little research has investigated the capability of multimodal large language models (MLLMs) to comprehend cognitive-level semantics. In this paper, we introduce MMLA, a comprehensive benchmark specifically designed to address this gap. MMLA comprises over 61K multimodal utterances drawn from both staged and real-world scenarios, covering six core dimensions of multimodal semantics: intent, emotion, dialogue act, sentiment, speaking style, and communication behavior. We evaluate eight mainstream branches of LLMs and MLLMs using three methods: zero-shot inference, supervised fine-tuning, and instruction tuning. Extensive experiments reveal that even fine-tuned models achieve only about 60~70% accuracy, underscoring the limitations of current MLLMs in understanding complex human language. We believe that MMLA will serve as a solid foundation for exploring the potential of large language models in multimodal language analysis and provide valuable resources to advance this field. The datasets and code are open-sourced at https://github.com/thuiar/MMLA.
Hanlei Zhang, Hua Xu 0003, Yeshuang Zhu, Peiwu Wang, Haige Zhu, Jie Zhou 0016, Jinchao Zhang 0001
NeurIPS3
2025 Multimodal Classification and Out-of-Distribution Detection for Multimodal Intent Understanding
Hanlei Zhang, Qianrui Zhou, Hua Xu 0003, Jianhua Su, Roberto Evans, Kai Gao 0006
IEEE Trans. Multim.3
2024 Token-Level Contrastive Learning with Modality-Aware Prompting for Multimodal Intent Recognition
abstract
Multimodal intent recognition aims to leverage diverse modalities such as expressions, body movements and tone of speech to comprehend user's intent, constituting a critical task for understanding human language and behavior in real-world multimodal scenarios. Nevertheless, the majority of existing methods ignore potential correlations among different modalities and own limitations in effectively learning semantic features from nonverbal modalities. In this paper, we introduce a token-level contrastive learning method with modality-aware prompting (TCL-MAP) to address the above challenges. To establish an optimal multimodal semantic environment for text modality, we develop a modality-aware prompting module (MAP), which effectively aligns and fuses features from text, video and audio modalities with similarity-based modality alignment and cross-modality attention mechanism. Based on the modality-aware prompt and ground truth labels, the proposed token-level contrastive learning framework (TCL) constructs augmented samples and employs NT-Xent loss on the label token. Specifically, TCL capitalizes on the optimal textual semantic insights derived from intent labels to guide the learning processes of other modalities in return. Extensive experiments show that our method achieves remarkable improvements compared to state-of-the-art methods. Additionally, ablation analyses demonstrate the superiority of the modality-aware prompt over the handcrafted prompt, which holds substantial significance for multimodal prompt learning. The codes are released at https://github.com/thuiar/TCL-MAP.
Qianrui Zhou, Hua Xu 0003, Hanlei Zhang, Kai Gao 0006
AAAI2
2024 Unsupervised Multimodal Clustering for Semantics Discovery in Multimodal Utterances
abstract
Discovering the semantics of multimodal utterances is essential for understanding human language and enhancing human-machine interactions.Existing methods manifest limitations in leveraging nonverbal information for discerning complex semantics in unsupervised scenarios.This paper introduces a novel unsupervised multimodal clustering method (UMC), making a pioneering contribution to this field.UMC introduces a unique approach to constructing augmentation views for multimodal data, which are then used to perform pre-training to establish well-initialized representations for subsequent clustering.An innovative strategy is proposed to dynamically select high-quality samples as guidance for representation learning, gauged by the density of each sample's nearest neighbors.Besides, it is equipped to automatically determine the optimal value for the top-K parameter in each cluster to refine sample selection.Finally, both high-and low-quality samples are used to learn representations conducive to effective clustering.We build baselines on benchmark multimodal intent and dialogue act datasets.UMC shows remarkable improvements of 2-6% scores in clustering metrics over state-of-the-art methods, marking the first successful endeavor in this domain.The complete code and data are available at https://github.com/thuiar/UMC.
Hanlei Zhang, Hua Xu 0003, Xin Wang 0220, Kai Gao 0006
ACL (1)2
2024 Light-MILPopt: Solving Large-scale Mixed Integer Linear Programs with Lightweight Optimizer and Small-scale Training Dataset
abstract
Machine Learning (ML)-based optimization approaches emerge as a promising technique for solving large-scale Mixed Integer Linear Programs (MILPs). However, existing ML-based frameworks suffer from high model computation complexity, weak problem reduction, and reliance on large-scale optimizers and large training datasets, resulting in performance bottlenecks for large-scale MILPs. This paper proposes Light-MILPopt, a lightweight large-scale optimization framework that only uses a lightweight optimizer and small training dataset to solve large-scale MILPs. Specifically, Light-MILPopt can be divided into four stages: Problem Formulation for problem division to reduce model computational costs, Model-based Initial Solution Prediction for predicting and constructing the initial solution using a small-scale training dataset, Problem Reduction for both variable and constraint reduction, and Data-driven Optimization for current solution improvement employing a lightweight optimizer. Experimental evaluations on four large-scale benchmark MILPs and a real-world case study demonstrate that Light-MILPopt, leveraging a lightweight optimizer and small training dataset, outperforms the state-of-the-art ML-based optimization framework and advanced large-scale solvers (e.g. Gurobi, SCIP). The results and further analyses substantiate the ML-based framework's feasibility and effectiveness in solving large-scale MILPs.
Huigen Ye, Hua Xu 0003
ICLR2
2024 MIntRec2.0: A Large-scale Benchmark Dataset for Multimodal Intent Recognition and Out-of-scope Detection in Conversations
abstract
Multimodal intent recognition poses significant challenges, requiring the incorporation of non-verbal modalities from real-world contexts to enhance the comprehension of human intentions. However, most existing multimodal intent benchmark datasets are limited in scale and suffer from difficulties in handling out-of-scope samples that arise in multi-turn conversational interactions. In this paper, we introduce MIntRec2.0, a large-scale benchmark dataset for multimodal intent recognition in multi-party conversations. It contains 1,245 high-quality dialogues with 15,040 samples, each annotated within a new intent taxonomy of 30 fine-grained classes, across text, video, and audio modalities. In addition to more than 9,300 in-scope samples, it also includes over 5,700 out-of-scope samples appearing in multi-turn contexts, which naturally occur in real-world open scenarios, enhancing its practical applicability. Furthermore, we provide comprehensive information on the speakers in each utterance, enriching its utility for multi-party conversational research. We establish a general framework supporting the organization of single-turn and multi-turn dialogue data, modality feature extraction, multimodal fusion, as well as in-scope classification and out-of-scope detection. Evaluation benchmarks are built using classic multimodal fusion methods, ChatGPT, and human evaluators. While existing methods incorporating nonverbal information yield improvements, effectively leveraging context information and detecting out-of-scope samples remains a substantial challenge. Notably, powerful large language models exhibit a significant performance gap compared to humans, highlighting the limitations of machine learning methods in the advanced cognitive intent understanding task. We believe that MIntRec2.0 will serve as a valuable resource, providing a pioneering foundation for research in human-machine conversational interactions, and significantly facilitating related applications. The full dataset and codes are available for use at https://github.com/thuiar/MIntRec2.0.
Hanlei Zhang, Xin Wang 0220, Hua Xu 0003, Qianrui Zhou, Kai Gao 0006, Jianhua Su, Jinyue Zhao
ICLR3
2024 Multimodal Consistency-Based Teacher for Semi-Supervised Multimodal Sentiment Analysis
abstract
Multimodal sentiment analysis holds significant importance within the realm of human-computer interaction. Due to the ease of collecting unlabeled online resources compared to the high costs associated with annotation, it becomes imperative for researchers to develop semi-supervised methods that leverage unlabeled data to enhance model performance. Existing semi-supervised approaches, particularly those applied to trivial image classification tasks, are not suitable for multimodal regression tasks due to their reliance on task-specific augmentation and thresholds designed for classification tasks. To address this limitation, we propose the Multimodal Consistency-based Teacher (MC-Teacher), which incorporates consistency-based pseudo-label technique into semi-supervised multimodal sentiment analysis. In our approach, we first propose synergistic consistency assumption which focus on the consistency among bimodal representation. Building upon this assumption, we develop a learnable filter network that autonomously learns how to identify misleading instances instead of threshold-based methods. This is achieved by leveraging both the implicit discriminant consistency on unlabeled instances and the explicit guidance on constructed training data with labeled instances. Additionally, we design the self-adaptive exponential moving average strategy to decouple the student and teacher networks, utilizing a heuristic momentum coefficient. Through both quantitative and qualitative experiments on two benchmark datasets, we demonstrate the outstanding performances of the proposed MC-Teacher approach. Furthermore, detailed analysis experiments and case studies are provided for each crucial component to intuitively elucidate the inner mechanism and further validate their effectiveness.
Jingliang Fang, Hua Xu 0003, Kai Gao 0006
IEEE ACM Trans. Audio Speech Lang. Process.3
2024 High-Dimensional Multi-Objective Bayesian Optimization With Block Coordinate Updates: Case Studies in Intelligent Transportation System
abstract
Many transportation system problems can be formulated as high-dimensional expensive multi-objective problems. They are challenging for Gaussian process-based Bayesian optimization methods to find the Pareto fronts due to the curse of dimensionality and the boundary issue in the acquisition function optimization. This paper presents a multi-objective Bayesian optimization method with block coordinate updates, Block-MOBO, to solve high-dimensional expensive multi-objective problems. Block-MOBO first partitions the decision variable space into different blocks, each of which includes a low-dimensional multi-objective problem. At each iteration, one block is considered and the decision variables not in this block are approximated by context-vector generation embedded with the Pareto prior knowledge thus promoting convergence. To tackle the boundary issue, we present$\epsilon $-greedy acquisition function in a Bayesian and multi-objective fashion, which recommends candidates either from the exploitation-exploration trade-off perspective or with probability$\epsilon $from the Pareto dominance relationship perspective. We verify the effectiveness of Block-MOBO by comparing it with other multi-objective Bayesian methods on two real-world optimization problems in transportation system and three multi-objective synthetic test suites. The experimental results show that Block-MOBO can find more evenly distributed and non-dominated solutions in the whole search space with lower complexity compared with other state-of-the-art approaches. Our analyses illustrate that block coordinate updates and$\epsilon $-greedy acquisition function contribute to computational complexity reduction and convergence-diversity trade-offs, respectively.
Hua Xu 0003, Zeqiu Zhang
IEEE Trans. Intell. Transp. Syst.2
2024 A Clustering Framework for Unsupervised and Semi-Supervised New Intent Discovery
abstract
New intent discovery is of great value to natural language processing, allowing for a better understanding of user needs and providing friendly services. However, most existing methods struggle to capture the complicated semantics of discrete text representations when limited or no prior knowledge of labeled data is available. To tackle this problem, we propose a novel clustering framework, USNID, forunsupervised andsemi-supervisednewintentdiscovery, which has three key technologies. First, it fully utilizes of unsupervised or semi-supervised data to mine shallow semantic similarity relations and provide well-initialized representations for clustering. Second, it designs a centroid-guided clustering mechanism to address the issue of cluster allocation inconsistency and provide high-quality self-supervised targets for representation learning. Third, it captures high-level semantics in unsupervised or semi-supervised data to discover fine-grained intent-wise clusters by optimizing both cluster-level and instance-level objectives. We also propose an effective method for estimating the cluster number in open-world scenarios without knowing the number of new intents beforehand. USNID performs exceptionally well on several benchmark intent datasets, achieving new state-of-the-art results in unsupervised and semi-supervised new intent discovery and demonstrating robust performance with different cluster numbers.
Hanlei Zhang, Hua Xu 0003, Xin Wang 0220, Kai Gao 0006
IEEE Trans. Knowl. Data Eng.2
2024 Noise Imitation Based Adversarial Training for Robust Multimodal Sentiment Analysis
abstract
As an inevitable phenomenon in real-world applications, data imperfection has emerged as one of the most critical challenges for multimodal sentiment analysis. However, existing approaches tend to overly focus on a specific type of imperfection, leading to performance degradation in real-world scenarios where multiple types of noise exist simultaneously. In this work, we formulate the imperfection with the modality feature missing at the training period and propose the noise intimation based adversarial training framework to improve the robustness against various potential imperfections at the inference period. Specifically, the proposed method first uses temporal feature erasing as the augmentation for noisy instances construction and exploits the modality interactions through the self-attention mechanism to learn multimodal representation for original-noisy instance pairs. Then, based on paired intermediate representation, a novel adversarial training strategy with semantic reconstruction supervision is proposed to learn unified joint representation between noisy and perfect data. For experiments, the proposed method is first verified with the modality feature missing, the same type of imperfection as the training period, and shows impressive performance. Moreover, we show that our approach is capable of achieving outstanding results for other types of imperfection, including modality missing, automation speech recognition error and attacks on text, highlighting the generalizability of our model. Finally, we conduct case studies on general additive distribution, which introduce background noise and blur into raw video clips, further revealing the capability of our proposed method for real-world applications.
Hua Xu 0003, Kai Gao 0006
IEEE Trans. Multim.3
2024 Meta Noise Adaption Framework for Multimodal Sentiment Analysis With Feature Noise
abstract
Improving the robustness of models against feature noise has emerged as one of the most crucial research topics in the field of multimodal sentiment analysis. Recent studies assume that the training instances are free of noise and develop either translation or reconstruction based method under the guidance of perfect training data for robust testing time performance. However, such an ideal assumption neglects the potential presence of the feature noise in training instances and inevitably results in degradation for the scenario where high-quality training instances are unavailable. In order to achieve robust training with noisy instances, we propose the Meta Noise Adaption (Meta-NA) learning strategy, a meta learning method accumulating the experience of dealing with various types of feature noise. Specifically, we first formulate the tasks distribution where each task is corresponding to one specific pattern of noise, and propose the feature adaption module adding on the unimodal encoder in late fusion based architecture. Through an nested online optimization between the auxiliary feature adaption module and the late fusion backbone modules, the proposed method can leverage shared knowledge across different noisy source tasks and learn how to learn from the noisy instances for robust testing performances. Extensive experiments are conducted on two benchmark multimodal sentiment analysis datasets, namely MOSI and CH-SIMS v2. The results demonstrate that our proposed method can rapidly adapt to various unseen types of feature noise and outperforms all baseline methods, particularly when the training instances are limited.
Baozheng Zhang, Hua Xu 0003, Kai Gao 0006
IEEE Trans. Multim.3
2024 Crossmodal Translation Based Meta Weight Adaption for Robust Image-Text Sentiment Analysis
abstract
Image-Text Sentiment Analysis task has garnered increased attention in recent years due to the surge in user-generated content on social media platforms. Previous research efforts have made noteworthy progress by leveraging the affective concepts shared between vision and text modalities. However, emotional cues may reside exclusively within one of the prevailing modalities, owing to modality independent nature and the potential absence of certain modalities. In this study, we aim to emphasize the significance of modality-independent emotional behaviors, in addition to the modality-invariant behaviors. To achieve this, we propose a novel approach called Crossmodal Translation-Based Meta Weight Adaption (CTMWA). Specifically, our approach involves the construction of the crossmodal translation network, which serves as the encoder. This architecture captures the shared concepts between vision content and text, empowering the model to effectively handle scenarios where either the vision or textual modality is missing. Building upon the translation-based framework, we introduce the strategy of unimodal weight adaption. Leveraging the meta-learning paradigm, our proposed strategy gradually learns to acquire unimodal weights for individual instances from a few hand-crafted meta instances with unimodal annotations. This enables us to modulate the gradients of each modality encoder based on the discrepancy between modalities during model training. Extensive experiments are conducted on three benchmark image-text sentiment analysis datasets, namely MVSA-Single, MVSA-Multiple, and TumEmo. The empirical results demonstrate that our proposed approach achieves the highest performance across all conventional image-text databases. Furthermore, experiments under modality missing settings and case study for reliable sentiment prediction are also conducted further exhibiting superior robustness as well as reliability of the propose approach.
Baozheng Zhang, Hua Xu 0003, Kai Gao 0006
IEEE Trans. Multim.3
2023 Robust-MSA: Understanding the Impact of Modality Noise on Multimodal Sentiment Analysis
abstract
Improving model robustness against potential modality noise, as an essential step for adapting multimodal models to real-world applications, has received increasing attention among researchers. For Multimodal Sentiment Analysis (MSA), there is also a debate on whether multimodal models are more effective against noisy features than unimodal ones. Stressing on intuitive illustration and in-depth analysis of these concerns, we present Robust-MSA, an interactive platform that visualizes the impact of modality noise as well as simple defence methods to help researchers know better about how their models perform with imperfect real-world data.
Huisheng Mao, Baozheng Zhang, Hua Xu 0003
AAAI3
2023 Adaptive Constraint Partition Based Optimization Framework for Large-Scale Integer Linear Programming (Student Abstract)
abstract
Integer programming problems (IPs) are challenging to be solved efficiently due to the NP-hardness, especially for large-scale IPs. To solve this type of IPs, Large neighborhood search (LNS) uses an initial feasible solution and iteratively improves it by searching a large neighborhood around the current solution. However, LNS easily steps into local optima and ignores the correlation between variables to be optimized, leading to compromised performance. This paper presents a general adaptive constraint partition-based optimization framework (ACP) for large-scale IPs that can efficiently use any existing optimization solver as a subroutine. Specifically, ACP first randomly partitions the constraints into blocks, where the number of blocks is adaptively adjusted to avoid local optima. Then, ACP uses a subroutine solver to optimize the decision variables in a randomly selected block of constraints to enhance the variable correlation. ACP is compared with LNS framework with different subroutine solvers on four IPs and a real-world IP. The experimental results demonstrate that in specified wall-clock time ACP shows better performance than SCIP and Gurobi.
Huigen Ye, Hua Xu 0003
AAAI3
2023 GNN&GBDT-Guided Fast Optimizing Framework for Large-scale Integer Programming
abstract
The latest two-stage optimization framework based on graph neural network (GNN) and large neighborhood search (LNS) is the most popular framework in solving large-scale integer programs (IPs). However, the framework can not effectively use the embedding spatial information in GNN and still highly relies on large-scale solvers in LNS, resulting in the scale of IP being limited by the ability of the current solver and performance bottlenecks. To handle these issues, this paper presents a GNN&GBDT-guided fast optimizing framework for large-scale IPs that only uses a small-scale optimizer to solve large-scale IPs efficiently. Specifically, the proposed framework can be divided into three stages: Multi-task GNN Embedding to generate the embedding space, GBDT Prediction to effectively use the embedding spatial information, and Neighborhood Optimization to solve large-scale problems fast using the small-scale optimizer. Extensive experiments show that the proposed framework can solve IPs with millions of scales and surpass SCIP and Gurobi in the specified wall-clock time using only a small-scale optimizer with 30% of the problem size. It also shows that the proposed framework can save 99% of running time in achieving the same solution quality as SCIP, which verifies the effectiveness and efficiency of the proposed framework in solving large-scale IPs.
Huigen Ye, Hua Xu 0003
ICML2
2023 Trustworthy machine reading comprehension with conditional adversarial calibration
Zhijing Wu 0002, Hua Xu 0003
Appl. Intell.2
2023 Learning Discriminative Representations and Decision Boundaries for Open Intent Detection
abstract
Open intent detection is a significant problem in natural language understanding, which aims to identify the unseen open intent while ensuring known intent identification performance. However, current methods face two major challenges. Firstly, they struggle to learn friendly representations to detect the open intent with prior knowledge of only known intents. Secondly, there is a lack of an effective approach to obtaining specific and compact decision boundaries for known intents. To address these issues, this paper presents an original framework called DA-ADB, which successively learns distance-aware intent representations and adaptive decision boundaries for open intent detection. Specifically, we first leverage distance information to enhance the distinguishing capability of the intent representations. Then, we design a novel loss function to obtain appropriate decision boundaries by balancing both empirical and open space risks. Extensive experiments demonstrate the effectiveness of the proposed distance-aware and boundary learning strategies. Compared to state-of-the-art methods, our framework achieves substantial improvements on three benchmark datasets. Furthermore, it yields robust performance with varying proportions of labeled data and known categories. The full data and codes are available for use athttps://github.com/thuiar/TEXTOIR.
Hanlei Zhang, Hua Xu 0003, Shaojie Zhao, Qianrui Zhou
IEEE ACM Trans. Audio Speech Lang. Process.2
2022 An End-to-End Traditional Chinese Medicine Constitution Assessment System Based on Multimodal Clinical Feature Representation and Fusion
abstract
Traditional Chinese Medicine (TCM) constitution is a fundamental concept in TCM theory. It is determined by multimodal TCM clinical features which, in turn, are obtained from TCM clinical information of image (face, tongue, etc.), audio (pulse and voice), and text (inquiry) modality. The auto assessment of TCM constitution is faced with two major challenges: (1) learning discriminative TCM clinical feature representations; (2) jointly processing the features using multimodal fusion techniques. The TCM Constitution Assessment System (TCM-CAS) is proposed to provide an end-to-end solution to this task, along with auxiliary functions to aid TCM researchers. To improve the results of TCM constitution prediction, the system combines multiple machine learning algorithms such as facial landmark detection, image segmentation, graph neural networks and multimodal fusion. Extensive experiments are conducted on a four-category multimodal TCM constitution dataset, and the proposed method achieves state-of-the-art accuracy. Provided with datasets containing annotations of diseases, the system can also perform automatic disease diagnosis from a TCM perspective.
Huisheng Mao, Baozheng Zhang, Hua Xu 0003, Kai Gao 0006
AAAI3
2022 An In-depth Interactive and Visualized Platform for Evaluating and Analyzing MRC Models
abstract
Machine Reading Comprehension (MRC) has made leaps and bounds when focusing on answering questions. However, since the existing accuracy-based evaluation metrics are agnostic to the nuances of neural networks, the true understanding and inferencing abilities of MRC models remain largely unknown. To address the above limitations, InDepth-Eva-MRC, an interactive and visualized platform, is proposed to provide analysis from cognitive fine-grained for MRC models. Concretely, the platform makes post-hoc systems to explain the behavior of MRC models. On the one hand, it analyzes the linguistic bias via performances with different linguistic properties. On the other hand, it performs skill-based analysis methods based on the modified test samples and semi-automatically generated test samples. Furthermore, through its detailed and interactive visualizations, the platform offers in-depth results analysis and model comparison from cognitive fine-grained. A screencast video and additional external material are available on https://github.com/thuiar/InDepth-Eva-MRC.
Zhijing Wu 0002, Jingliang Fang, Hua Xu 0003, Kai Gao 0006
CIKM3
2022 Make Acoustic and Visual Cues Matter: CH-SIMS v2.0 Dataset and AV-Mixup Consistent Module
abstract
Multimodal sentiment analysis (MSA), which supposes to improve text-based sentiment analysis with associated acoustic and visual modalities, is an emerging research area due to its potential applications in Human-Computer Interaction (HCI). However, existing researches observe that the acoustic and visual modalities contribute much less than the textual modality, termed as text-predominant. Under such circumstances, in this work, we emphasize making non-verbal cues matter for the MSA task. Firstly, from the resource perspective, we present the CH-SIMS v2.0 dataset, an extension and enhancement of the CH-SIMS. Compared with the original dataset, the CH-SIMS v2.0 doubles its size with another 2121 refined video segments containing both unimodal and multimodal annotations and collects 10161 unlabelled raw video segments with rich acoustic and visual emotion-bearing context to highlight non-verbal cues for sentiment prediction. Secondly, from the model perspective, benefiting from the unimodal annotations and the unsupervised data in the CH-SIMS v2.0, the Acoustic Visual Mixup Consistent (AV-MC) framework is proposed. The designed modality mixup module can be regarded as an augmentation, which mixes the acoustic and visual modalities from different videos. Through drawing unobserved multimodal context along with the text, the model can learn to be aware of different non-verbal contexts for sentiment prediction. Our evaluations demonstrate that both CH-SIMS v2.0 and AV-MC framework enable further research for discovering emotion-bearing acoustic and visual cues and pave the path to interpretable end-to-end HCI applications for real-world scenarios. The full dataset and code are available for use at https://github.com/thuiar/ch-sims-v2.
Huisheng Mao, Zhiyun Liang, Wanqiuyue Yang, Yuanzhe Qiu, Tie Cheng, Xiaoteng Li, Hua Xu 0003, Kai Gao 0006
ICMI9
2022 MIntRec: A New Dataset for Multimodal Intent Recognition
abstract
Multimodal intent recognition is a significant task for understanding human language in real-world multimodal scenes. Most existing intent recognition methods have limitations in leveraging the multimodal information due to the restrictions of the benchmark datasets with only text information. This paper introduces a novel dataset for multimodal intent recognition (MIntRec) to address this issue. It formulates coarse-grained and fine-grained intent taxonomies based on the data collected from the TV series Superstore. The dataset consists of 2,224 high-quality samples with text, video, and audio modalities and has multimodal annotations among twenty intent categories. Furthermore, we provide annotated bounding boxes of speakers in each video segment and achieve an automatic process for speaker annotation. MIntRec is helpful for researchers to mine relationships between different modalities to enhance the capability of intent recognition. We extract features from each modality and model cross-modal interactions by adapting three powerful multimodal fusion methods to build baselines. Extensive experiments show that employing the non-verbal modalities achieves substantial improvements compared with the text-only modality, demonstrating the effectiveness of using multimodal information for intent recognition. The gap between the best-performing methods and humans indicates the challenge and importance of this task for the community. The full dataset and codes are available for use at https://github.com/thuiar/MIntRec.
Hanlei Zhang, Hua Xu 0003, Xin Wang 0220, Qianrui Zhou, Shaojie Zhao, Jiayan Teng
ACM Multimedia2
2022 An adaptive batch Bayesian optimization approach for expensive multi-objective problems
Hua Xu 0003, Yuan Yuan 0004, Zeqiu Zhang
Inf. Sci.2
2022 GAR-Net: A Graph Attention Reasoning Network for conversation understanding
Hua Xu 0003, Yunfeng Xu, Jiyun Zou, Kai Gao 0006
Knowl. Based Syst.1
2022 Co-attentive multi-task convolutional neural network for facial expression recognition
Wenmeng Yu, Hua Xu 0003
Pattern Recognit.2
2021 Learning Modality-Specific Representations with Self-Supervised Multi-Task Learning for Multimodal Sentiment Analysis
abstract
Representation Learning is a significant and challenging task in multimodal learning. Effective modality representations should contain two parts of characteristics: the consistency and the difference. Due to the unified multimodal annota- tion, existing methods are restricted in capturing differenti- ated information. However, additional unimodal annotations are high time- and labor-cost. In this paper, we design a la- bel generation module based on the self-supervised learning strategy to acquire independent unimodal supervisions. Then, joint training the multimodal and uni-modal tasks to learn the consistency and difference, respectively. Moreover, dur- ing the training stage, we design a weight-adjustment strat- egy to balance the learning progress among different sub- tasks. That is to guide the subtasks to focus on samples with the larger difference between modality supervisions. Last, we conduct extensive experiments on three public multimodal baseline datasets. The experimental results validate the re- liability and stability of auto-generated unimodal supervi- sions. On MOSI and MOSEI datasets, our method surpasses the current state-of-the-art methods. On the SIMS dataset, our method achieves comparable performance than human- annotated unimodal labels. The full codes are available at https://github.com/thuiar/Self-MM.
Wenmeng Yu, Hua Xu 0003, Jiele Wu
AAAI2
2021 Deep Open Intent Classification with Adaptive Decision Boundary
abstract
Open intent classification is a challenging task in dialogue systems. On the one hand, it should ensure the quality of known intent identification. On the other hand, it needs to detect the open (unknown) intent without prior knowledge. Current models are limited in finding the appropriate decision boundary to balance the performances of both known intents and the open intent. In this paper, we propose a post-processing method to learn the adaptive decision boundary (ADB) for open intent classification. We first utilize the labeled known intent samples to pre-train the model. Then, we automatically learn the adaptive spherical decision boundary for each known class with the aid of well-trained features. Specifically, we propose a new loss function to balance both the empirical risk and the open space risk. Our method does not need open intent samples and is free from modifying the model architecture. Moreover, our approach is surprisingly insensitive with less labeled data and fewer known intents. Extensive experiments on three benchmark datasets show that our method yields significant improvements compared with the state-of-the-art methods.
Hanlei Zhang, Hua Xu 0003, Ting-En Lin
AAAI2
2021 Discovering New Intents with Deep Aligned Clustering
abstract
Discovering new intents is a crucial task in dialogue systems. Most existing methods are limited in transferring the prior knowledge from known intents to new intents. These methods also have difficulties in providing high-quality supervised signals to learn clustering-friendly features for grouping unlabeled intents. In this work, we propose an effective method (Deep Aligned Clustering) to discover new intents with the aid of limited known intent data. Firstly, we leverage a few labeled known intent samples as prior knowledge to pre-train the model. Then, we perform k-means to produce cluster assignments as pseudo-labels. Moreover, we propose an alignment strategy to tackle the label inconsistency problem during clustering assignments. Finally, we learn the intent representations under the supervision of the aligned pseudo-labels. With an unknown number of new intents, we predict the number of intent categories by eliminating low-confidence intent-wise clusters. Extensive experiments on two benchmark datasets show that our method is more robust and achieves substantial improvements over the state-of-the-art methods.
Hanlei Zhang, Hua Xu 0003, Ting-En Lin, Rui Lyu
AAAI2
2021 Transformer-based Feature Reconstruction Network for Robust Multimodal Sentiment Analysis
abstract
Improving robustness against data missing has become one of the core challenges in Multimodal Sentiment Analysis (MSA), which aims to judge speaker sentiments from the language, visual, and acoustic signals. In the current research, translation-based methods and tensor regularization methods are proposed for MSA with incomplete modality features. However, both of them fail to cope with random modality feature missing in non-aligned sequences. In this paper, a transformer-based feature reconstruction network (TFR-Net) is proposed to improve the robustness of models for the random missing in non-aligned modality sequences. First, intra-modal and inter-modal attention-based extractors are adopted to learn robust representations for each element in modality sequences. Then, a reconstruction module is proposed to generate the missing modality features. With the supervision of SmoothL1Loss between generated and complete sequences, TFR-Net is expected to learn semantic-level features corresponding to missing features. Extensive experiments on two public benchmark datasets show that our model achieves good results against data missing across various missing modality combinations and various missing degrees.
Wei Li 0232, Hua Xu 0003, Wenmeng Yu
ACM Multimedia3
2021 Representation iterative fusion based on heterogeneous graph neural network for joint entity and relation extraction
Hua Xu 0003, Xiaoteng Li, Kai Gao 0006
Knowl. Based Syst.2
2020 Constrained Self-Supervised Clustering for Discovering New Intents (Student Abstract)
abstract
Discovering new user intents is an emerging task in the dialogue system. In this paper, we propose a self-supervised clustering method that can naturally incorporate pairwise constraints as prior knowledge to guide the clustering process and does not require intensive feature engineering. Extensive experiments on three benchmark datasets show that our method can yield significant improvements over strong baselines.
Ting-En Lin, Hua Xu 0003, Hanlei Zhang
AAAI2
2020 Discovering New Intents via Constrained Deep Adaptive Clustering with Cluster Refinement
abstract
Identifying new user intents is an essential task in the dialogue system. However, it is hard to get satisfying clustering results since the definition of intents is strongly guided by prior knowledge. Existing methods incorporate prior knowledge by intensive feature engineering, which not only leads to overfitting but also makes it sensitive to the number of clusters. In this paper, we propose constrained deep adaptive clustering with cluster refinement (CDAC+), an end-to-end clustering method that can naturally incorporate pairwise constraints as prior knowledge to guide the clustering process. Moreover, we refine the clusters by forcing the model to learn from the high confidence assignments. After eliminating low confidence assignments, our approach is surprisingly insensitive to the number of clusters. Experimental results on the three benchmark datasets show that our method can yield significant improvements over strong baselines. 1
Ting-En Lin, Hua Xu 0003, Hanlei Zhang
AAAI2
2020 Combining Fine-Tuning with a Feature-Based Approach for Aspect Extraction on Reviews (Student Abstract)
abstract
One key task of fine-grained sentiment analysis on reviews is to extract aspects or features that users have expressed opinions on. Generally, fine-tuning BERT with sophisticated task-specific layers can achieve better performance than only extend one extra task-specific layer (e.g., a fully-connected + softmax layer) since not all tasks can easily be represented by Transformer encoder architecture and special task-specific layer can capture task-specific features. However, BERT fine-tuning may be unstable on a small-scale dataset. Besides, in our experiments, directly fine-tuning BERT on extending sophisticated task-specific layers did not take advantage of the features of task-specific layers and even restrict the performance of BERT module. To address the above consideration, this paper combines Fine-tuning with a feature-based approach to extract aspect. To the best of our knowledge, this is the first paper to combine fine-tuning with a feature-based approach for aspect extraction.
Hua Xu 0003, Xiaomin Sun 0001, Guangcan Tao
AAAI2
2020 A Multi-Task Learning Machine Reading Comprehension Model for Noisy Document (Student Abstract)
abstract
Current neural models for Machine Reading Comprehension (MRC) have achieved successful performance in recent years. However, the model is too fragile and lack robustness to tackle the imperceptible adversarial perturbations to the input. In this work, we propose a multi-task learning MRC model with a hierarchical knowledge enrichment to further improve the robustness for noisy document. Our model follows a typical encode-align-decode framework. Additionally, we apply a hierarchical method of adding background knowledge into the model from coarse-to-fine to enhance the language representations. Besides, we optimize our model by jointly training the answer span and unanswerability prediction, aiming to improve the robustness to noise. Experiment results on benchmark datasets confirm the superiority of our method, and our method can achieve competitive performance compared with other strong baselines.
Zhijing Wu 0002, Hua Xu 0003
AAAI2
2020 Multi-Channel Convolutional Neural Networks with Adversarial Training for Few-Shot Relation Classification (Student Abstract)
abstract
The distant supervised (DS) method has improved the performance of relation classification (RC) by means of extending the dataset. However, DS also brings the problem of wrong labeling. Contrary to DS, the few-shot method relies on few supervised data to predict the unseen classes. In this paper, we use word embedding and position embedding to construct multi-channel vector representation and use the multi-channel convolutional method to extract features of sentences. Moreover, in order to alleviate few-shot learning to be sensitive to overfitting, we introduce adversarial learning for training a robust model. Experiments on the FewRel dataset show that our model achieves significant and consistent improvements on few-shot RC as compared with baselines.
Yuxiang Xie, Hua Xu 0003, Congcong Yang, Kai Gao 0006
AAAI2
2020 CM-BERT: Cross-Modal BERT for Text-Audio Sentiment Analysis
abstract
Multimodal sentiment analysis is an emerging research field that aims to enable machines to recognize, interpret, and express emotion. Through the cross-modal interaction, we can get more comprehensive emotional characteristics of the speaker. Bidirectional Encoder Representations from Transformers (BERT) is an efficient pre-trained language representation model. Fine-tuning it has obtained new state-of-the-art results on eleven natural language processing tasks like question answering and natural language inference. However, most previous works fine-tune BERT only base on text data, how to learn a better representation by introducing the multimodal information is still worth exploring. In this paper, we propose the Cross-Modal BERT (CM-BERT), which relies on the interaction of text and audio modality to fine-tune the pre-trained BERT model. As the core unit of the CM-BERT, masked multimodal attention is designed to dynamically adjust the weight of words by combining the information of text and audio modality. We evaluate our method on the public multimodal sentiment analysis datasets CMU-MOSI and CMU-MOSEI. The experiment results show that it has significantly improved the performance on all the metrics over previous baselines and text-only finetuning of BERT. Besides, we visualize the masked multimodal attention and proves that it can reasonably adjust the weight of words by introducing audio modality information.
Kaicheng Yang 0005, Hua Xu 0003, Kai Gao 0006
ACM Multimedia2
2020 Deep reinforcement learning for robust emotional classification in facial expression recognition
Hua Xu 0003
Knowl. Based Syst.2
2020 Improving the robustness of machine reading comprehension model with hierarchical knowledge and auxiliary unanswerability prediction
Zhijing Wu 0002, Hua Xu 0003
Knowl. Based Syst.2
2020 Heterogeneous graph neural networks for noisy few-shot relation classification
Yuxiang Xie, Hua Xu 0003, Jiaoe Li, Congcong Yang, Kai Gao 0006
Knowl. Based Syst.2
2020 Finding structural hole spanners based on community forest model and diminishing marginal utility in large scale social networks
Hua Xu 0003, Yunfeng Xu, Junhui Deng, Juan Gu, Jie Lai, Jiangtao Hu, Xiaoshuai Yu, Lidong Gu, Yanling Wei 0003, Yichao Xiao, Junhao Lu
Knowl. Based Syst.2
2019 Deep Unknown Intent Detection with Margin Loss
abstract
Identifying the unknown (novel) user intents that have never appeared in the training set is a challenging task in the dialogue system.In this paper, we present a two-stage method for detecting unknown intents.We use bidirectional long short-term memory (BiLSTM) network with the margin loss as the feature extractor.With margin loss, we can learn discriminative deep features by forcing the network to maximize inter-class variance and to minimize intra-class variance.Then, we feed the feature vectors to the density-based novelty detection algorithm, local outlier factor (LOF), to detect unknown intents.Experiments on two benchmark datasets show that our method can yield consistent improvements compared with the baseline methods.
Ting-En Lin, Hua Xu 0003
ACL (1)2
2019 A joint model of extended LDA and IBTM over streaming Chinese short texts
abstract
With the prevalent of short texts, discovering the topics within them has become an important task. Biterm Topic Model (BTM) is more suitable to discover topics on short texts than traditional topic models. However, there are still some challenges that dealing short texts with BTM will always ignor e the document-topic semantic information and lack the true intentions of users. In addition, it is a static method and can not manage streaming short texts when a new one arrives immediately. In order to keep document-topic information and get the topic distribution of a new short text at once, we propose a joint model based on online algorithms of Latent Dirichlet Allocation (LDA) and BTM, which combines the merits of both models. Not only does it alleviate the sparsity when addressing short texts with the online algorithm of BTM, namely Incremental Biterm Topic Model (IBTM), but also keeps document-topic information with extended LDA. And considering the differences between English and Chinese text in writing, we use combined words in short texts as key words to extend the length of short texts and keep the true intensions of users. As shown in the experiment results on two real world datasets, our method is better than other baseline methods. In the end, we explain an application of our method in the task of discovering user interest tags.
Longxia Zhu, Hua Xu 0003, Yunfeng Xu, Jia Li 0025, Junhui Deng, Xiaomin Sun 0001, Xiaoli Bai
Intell. Data Anal.2
2019 A post-processing method for detecting unknown intent of dialogue system via pre-trained deep neural network classifier
Ting-En Lin, Hua Xu 0003
Knowl. Based Syst.2
2018 Two-Stage Attention Network for Aspect-Level Sentiment Classification
Kai Gao 0006, Hua Xu 0003, Chengliang Gao, Xiaomin Sun 0001, Junhui Deng
ICONIP (4)2
2018 Attention-Based BiLSTM Network with Lexical Feature for Emotion Classification
abstract
Emotion classification is an important task for identifying users' emotional expressions in text. Though a variety of neural models have been proposed nowadays, these models mainly focus on modeling the content of words or characters without fully employing the emotional features in lexical features, especially the features of part-of -speech (POS). In this paper, we reveal that the information of POS as well as that of words is important for identifying the type of emotion in a given text. We propose two simple models to fully learn the emotional features of the POS of words. Every model consists of the long short-term memory (LSTM) network as input encoders and the component of attention mechanism. One model is to concatenate the POS tags of vectors into the hidden states of representations generated by LSTM as raw feature representations and put them into the component of attention mechanism to generate the text representation toward a special emotion. The other is to use both LSTM and attention mechanism to model the context representation of words and those of POS tags respectively and concatenate these context representations as the text representation toward a special emotion. We conduct some experiments on datasets for evaluation and demonstrate the effectiveness of our model, where the datasets consist of the open-source dataset from NLPCC& 2014 and the dataset of manual annotation. Experimental results show that our models can achieve outstanding performance for emotion classification in Chinese Weibo texts and outperform classical baselines.
Kai Gao 0006, Hua Xu 0003, Chengliang Gao, Hanyong Hao, Junhui Deng, Xiaomin Sun 0001
IJCNN2
2018 Topic Discovery for Streaming Short Texts with CTM
abstract
Short texts are prevalent on today’s Web, especially with the emergence of social media. However, how to discover the topics of streaming short texts has become an important task for many content analysis applications. Conventional topic models such as Probabilistic Latent Semantic Analysis (PLSA) and Latent Dirichlet Allocation (LDA) will suffer from sparsity problem when we infer the latent topics from short texts with them. The reason is that they derive topics from document-level word co-occurrence by modeling each document as a mixture of topics. Different from the above idea, Biterm Topic Model (BTM) discovers topics in short texts by directly modeling the generation of word co-occurrence patterns in the whole corpus. But semantic information is lacking for short texts. In this paper, in order to alleviate the sparsity problem, keep the semantic information of documents and get the latent topic information of streaming short texts immediately, we propose a joint topic model for Chinese streaming short texts (CTM) based on the online algorithms of LDA and BTM. Experiments on short texts from Sina Weibo show that our joint topic model can discover more precise topics and carry out more applications. In addition, considering the preprocessing in Chinese text is different from English and errors in extracting key phrases, we use a combined word method to extend the length of short texts and reduce errors in extracting key phrases.
Yunfeng Xu, Hua Xu 0003, Longxia Zhu, Hanyong Hao, Junhui Deng, Xiaomin Sun 0001, Xiaoli Bai
IJCNN2
2018 Objective Reduction in Many-Objective Optimization: Evolutionary Multiobjective Approaches and Comprehensive Analysis
abstract
Many-objective optimization problems bring great difficulties to the existing multiobjective evolutionary algorithms, in terms of selection operators, computational cost, visualization of the high-dimensional tradeoff front, and so on. Objective reduction can alleviate such difficulties by removing the redundant objectives in the original objective set, which has become one of the most important techniques in many-objective optimization. In this paper, we suggest to view objective reduction as a multiobjective search problem and introduce three multiobjective formulations of the problem, where the first two formulations are both based on preservation of the dominance structure and the third one utilizes the correlation between objectives. For each multiobjective formulation, a multiobjective objective reduction algorithm is proposed by employing the nondominated sorting genetic algorithm II to generate a Pareto front of nondominated objective subsets that can offer decision support to the user. Moreover, we conduct a comprehensive analysis of two major categories of objective reduction approaches based on several theorems, with the aim of revealing their strengths and limitations. Lastly, the performance of the proposed multiobjective algorithms is studied extensively on various benchmark problems and two real-world problems. Numerical results and comparisons are then shown to highlight the effectiveness and superiority of the proposed multiobjective algorithms over existing state-of-the-art approaches in the related field.
Yuan Yuan 0004, Yew-Soon Ong, Abhishek Gupta 0001, Hua Xu 0003
IEEE Trans. Evol. Comput.4
2017 Optimize collapsed Gibbs sampling for biterm topic model by alias method
abstract
With the popularity of social networks, such as mi-croblogs and Twitter, topic inference for short text is increasingly significant and essential for many content analysis tasks. Biterm topic model (BTM) is superior to conventional topic models in uncovering latent semantic relevance for short text. However, Gibbs sampling employed by BTM is very time consuming when inferring topics, especially for large-scale datasets. It requires O{K) operations per sample for K topics, where K denotes the number of topics in the corpus. In this paper, we propose an acceleration algorithm of BTM, FastBTM, using an efficient sampling method for BTM which only requires O(1) amortized time while the traditional ones scale linearly with the number of topics. FastBTM is based on Metropolis-Hastings and alias method, both of which have been widely adopted in latent Dirichlet allocation (LDA) model and achieved outstanding speedup. We carry out a number of experiments on Tweets2011 Collection dataset and Enron dataset, indicating that our method is robust enough for both short texts and normal documents. Our work can be approximately 9 times faster than traditional Gibbs sampling method per iteration, when setting K = 1000. The source code of FastBTM can be obtained from https://github.com/paperstudy/FastBTM.
Xingwei He 0002, Hua Xu 0003, Xiaomin Sun 0001, Junhui Deng, Xiaoli Bai, Jia Li 0025
IJCNN2
2017 ABiRCNN with neural tensor network for answer selection
abstract
Answer selection is a very important task in domain question answering. However, because of the word variety between questions and answers, there exists the lexical gap between questions and answers, which is the major challenge in question answer matching. In this work, in order to overcome the lexical gap, we propose an attention based bidirectional gated convolution with neural tensor network (ABiRCNN+NTN), which can improve the representations for both questions and answers and model their interactions with a neural tensor network. We carry out large-scale experiments on answer selection dataset, InsuranceQA and achieve new state-of-the-art results on InsuranceQA dataset. The experimental results demonstrate that our model can effectively capture the complex semantic relations between questions and answers and encode them in a more effective way. The source code of our work can be obtained from https://github.com/paperstudy/AnswerSelection.
Xingwei He 0002, Hua Xu 0003, Xiaomin Sun 0001, Junhui Deng, Jia Li 0025
IJCNN2
2017 On the need of hierarchical emotion classification: Detecting the implicit feature using constrained topic model
abstract
Nowadays in China, Sina Weibo has become the most popular microblog platform and researches about it are proposed increasingly. In this paper, the problem of emotion classification of Weibo’s posts is addressed in a hierarchical way using a constrained topic model and Support Vector Regression (SVR ). Based on this topic model which is variation of Latent Dirichlet Allocation (LDA), an implicit emotion detection algorithm is proposed to identify the underlying emotions. Meanwhile, the constraints are generated based on prior knowledge extraction approaches to compact LDA in order to generate domain-specified topics. Furthermore, a hierarchical emotion structure is employed to classify emotions more precisely into 19 classes. This hierarchy can meet different research granularities. The whole architecture is proposed aimed at alleviating the pain of misclassification caused by feature imbalance and decreasing the labor cost. The experiment results validate that our model outperforms traditional methods with precision, recall and F-scores.
Hua Xu 0003, Xiaoli Bai
Intell. Data Anal.2
2017 FastBTM: Reducing the sampling time for biterm topic model
Xingwei He 0002, Hua Xu 0003, Jia Li 0025, Linlin Yu
Knowl. Based Syst.2
2016 Hyperbolic linear units for deep convolutional neural networks
abstract
Recently, rectified linear units (ReLUs) have been used to solve the vanishing gradient problem. Their use has led to state-of-the-art results in various problems such as image classification. In this paper, we propose the hyperbolic linear units (HLUs) which not only speed up learning process in deep convolutional neural networks but also obtain better performance in image classification tasks. Unlike ReLUs, HLUs have inheriently negative values which could make mean unit outputs closer to zero. Mean unit outputs close to zero means we can speed up the learning process because they bring the normal gradient close to the natural gradient. Indeed, the difference called bias shift between natural gradient and the normal gradient is related to the mean activation of input units. Experiments with three popular CNN architectures, LeNet, Inception network and ResNet on various benchmarks including MNIST, CIFAR-10 and CIFAR-100 demonstrate that our proposed HLUs achieve significant improvement compared to other commonly used activation functions1.
Jia Li 0025, Hua Xu 0003, Junhui Deng, Xiaomin Sun 0001
IJCNN2
2016 Tweet modeling with LSTM recurrent neural networks for hashtag recommendation
abstract
The hash symbol, called a hashtag, is used to mark the keyword or topic in a tweet. It was created organically by users as a way to categorize messages. Hashtags also provide valuable information for many research applications such as sentiment classification and topic analysis. However, only a small number of tweets are manually annotated. Therefore, an automatic hashtag recommendation method is needed to help users tag their new tweets. Previous methods mostly use conventional machine learning classifiers such as SVM or utilize collaborative filtering technique. A bottleneck of these approaches is that they all use the TF-IDF scheme to represent tweets and ignore the semantic information in tweets. In this paper, we also regard hashtag recommendation as a classification task but propose a novel recurrent neural network model to learn vector-based tweet representations to recommend hashtags. More precisely, we use a skip-gram model to generate distributed word representations and then apply a convolutional neural network to learn semantic sentence vectors. Afterwards, we make use of the sentence vectors to train a long short-term memory recurrent neural network (LSTM-RNN). We directly use the produced tweet vectors as features to classify hashtags without any feature engineering. Experiments on real world data from Twitter to recommend hashtags show that our proposed LSTM-RNN model outperforms state-of-the-art methods and LSTM unit also obtains the best performance compared to standard RNN and gated recurrent unit (GRU).
Jia Li 0025, Hua Xu 0003, Xingwei He 0002, Junhui Deng, Xiaomin Sun 0001
IJCNN2
2016 Grasp the implicit features: Hierarchical emotion classification based on topic model and SVM
abstract
Microblog post has been a hot research source for emotion classification in recent years. However, due to bloggers' free narrative style and topics' timeliness, the data from microblog post is usually implicit and imbalanced. In this paper, the problems of emotion classification in Chinese microblog posts are solved in a hierarchical way using a knowledge-based topic model and Support Vector Machine(SVM). Based on topic model, an implicit feature detection algorithm is proposed to identify the latent emotions underlying the microblog posts. Meanwhile, a hierarchical emotion structure is employed to classify emotions into 19 classes of four levels by SVM. This structure can meet different research requirements at three granularities. The experiment results validate that our model can achieve better performance in terms of precision, recall and F-scores.
Hua Xu 0003, Jiushuo Wang, Xiaomin Sun 0001, Junhui Deng
IJCNN2
2016 User-IBTM: An Online Framework for Hashtag Suggestion in Twitter
Jia Li 0025, Hua Xu 0003
WAIM (2)2
2016 Suggest what to tag: Recommending more precise hashtags based on users' dynamic interests and streaming tweet content
Jia Li 0025, Hua Xu 0003
Knowl. Based Syst.2
2016 Finding overlapping community from social networks based on community forest model
Yunfeng Xu, Hua Xu 0003, Dongwen Zhang
Knowl. Based Syst.2
2016 BitHash: An efficient bitwise Locality Sensitive Hashing method with applications
Wenhao Zhang 0003, Jianqiu Ji, Jun Zhu 0001, Jianmin Li 0001, Hua Xu 0003, Bo Zhang 0010
Knowl. Based Syst.5
2016 A New Dominance Relation-Based Evolutionary Algorithm for Many-Objective Optimization
abstract
Many-objective optimization has posed a great challenge to the classical Pareto dominance-based multiobjective evolutionary algorithms (MOEAs). In this paper, an evolutionary algorithm based on a new dominance relation is proposed for many-objective optimization. The proposed evolutionary algorithm aims to enhance the convergence of the recently suggested nondominated sorting genetic algorithm III by exploiting the fitness evaluation scheme in the MOEA based on decomposition, but still inherit the strength of the former in diversity maintenance. In the proposed algorithm, the nondominated sorting scheme based on the introduced new dominance relation is employed to rank solutions in the environmental selection phase, ensuring both convergence and diversity. The proposed algorithm is evaluated on a number of well-known benchmark problems having 3-15 objectives and compared against eight state-of-the-art algorithms. The extensive experimental results show that the proposed algorithm can work well on almost all the test functions considered in this paper, and it is compared favorably with the other many-objective optimizers. Additionally, a parametric study is provided to investigate the influence of a key parameter in the proposed algorithm.
Yuan Yuan 0004, Hua Xu 0003, Bo Wang 0051, Xin Yao 0001
IEEE Trans. Evol. Comput.2
2016 Balancing Convergence and Diversity in Decomposition-Based Many-Objective Optimizers
abstract
The decomposition-based multiobjective evolutionary algorithms (MOEAs) generally make use of aggregation functions to decompose a multiobjective optimization problem into multiple single-objective optimization problems. However, due to the nature of contour lines for the adopted aggregation functions, they usually fail to preserve the diversity in high-dimensional objective space even by using diverse weight vectors. To address this problem, we propose to maintain the desired diversity of solutions in their evolutionary process explicitly by exploiting the perpendicular distance from the solution to the weight vector in the objective space, which achieves better balance between convergence and diversity in many-objective optimization. The idea is implemented to enhance two well-performing decomposition-based algorithms, i.e., MOEA, based on decomposition and ensemble fitness ranking. The two enhanced algorithms are compared to several state-of-the-art algorithms and a series of comparative experiments are conducted on a number of test problems from two well-known test suites. The experimental results show that the two proposed algorithms are generally more effective than their predecessors in balancing convergence and diversity, and they are also very competitive against other existing algorithms for solving many-objective optimization problems.
Yuan Yuan 0004, Hua Xu 0003, Bo Wang 0051, Bo Zhang 0010, Xin Yao 0001
IEEE Trans. Evol. Comput.2
2015 Scale adaptive reproduction operator for decomposition based estimation of distribution algorithm
abstract
Multi-objective evolutionary algorithm based on decomposition (MOEA/D) uses crossover operator which often either breaks the building blocks or mix them ineffectively. Multi-objective estimation of distribution algorithm based on decomposition (MEDA/D) evolves a probability vector for each sub-problem to guide the search instead of using crossover operator.However, since the number of the weight vectors in the neighborhood of each weight vector is relatively small and MEDA/D does not provide a way to maintain diversity, the performance of MEDA/D is limited. To overcome the drawbacks of MEDA/D, we proposed a new reproduction operator. This operator could promote diversity. We introduced it into MOEA/D framework and the new algorithm is called s-MEDA/D. We also prove that the parameter newly introduced has physical significance and the reproduction operator is not susceptible to the scale of the problem. The s-MEDA/D was tested on nine instances of the 0/1 multi-objective knapsack problem. Empirical evaluation suggests that the proposed algorithm is effective and efficient.
Bo Wang 0051, Hua Xu 0003, Yuan Yuan 0004
CEC2
2015 An Experimental Investigation of Variation Operators in Reference-Point Based Many-Objective Optimization
abstract
Reference-point based multi-objective evolutionary algorithms (MOEAs) have shown promising performance in many-objective optimization. However, most of existing research within this area focused on improving the environmental selection procedure, and little work has been done on the effect of variation operators. In this paper, we conduct an experimental investigation of variation operators in a typical reference-point based MOEA, i.e., NSGA-III. First, we provide a new NSGA-III variant, i.e., NSGA-III-DE, which introduces differential evolution (DE) operator into NSGA-III, and we further examine the effect of two main control parameters in NSGA-III-DE. Second, we have an experimental analysis of the search behavior of NSGA-III-DE and NSGA-III. We observe that NSGA-III-DE is generally better at exploration whereas NSGA-III normally has advantages in exploitation. Third, based on this observation, we present two other NSGA-III variants, where DE operator and genetic operators are simply combined to reproduce solutions. Experimental results on several benchmark problems show that very encouraging performance can be achieved by three suggested new NSGA-III variants. Our work also indicates that the performance of NSGA-III is significantly bottlenecked by its variation operators, providing opportunities for the study of the other alternative ones.
Yuan Yuan 0004, Hua Xu 0003, Bo Wang 0051
GECCO2
2015 Emotion Cause Detection for Chinese Micro-Blogs Based on ECOCC Model
Kai Gao 0006, Hua Xu 0003, Jiushuo Wang
PAKDD (2)2
2015 A rule-based approach to emotion cause detection for Chinese micro-blogs
Kai Gao 0006, Hua Xu 0003, Jiushuo Wang
Expert Syst. Appl.2
2015 A novel disjoint community detection algorithm for social networks based on backbone degree and expansion
Yunfeng Xu, Hua Xu 0003, Dongwen Zhang
Expert Syst. Appl.2
2015 Hierarchical emotion classification and emotion component analysis on chinese micro-blog posts
Hua Xu 0003, Jiushuo Wang
Expert Syst. Appl.1
2015 Implicit feature identification in Chinese reviews using explicit topic mining model
Hua Xu 0003
Knowl. Based Syst.1
2015 Multiobjective Flexible Job Shop Scheduling Using Memetic Algorithms
abstract
In this paper, we propose new memetic algorithms (MAs) for the multiobjective flexible job shop scheduling problem (MO-FJSP) with the objectives to minimize the makespan, total workload, and critical workload. The problem is addressed in a Pareto manner, which aims to search for a set of Pareto optimal solutions. First, by using well-designed chromosome encoding/decoding scheme and genetic operators, the nondominated sorting genetic algorithm II (NSGA-II) is adapted for the MO-FJSP. Then, our MAs are developed by incorporating a novel local search algorithm into the adapted NSGA-II, where some good individuals are chosen from the offspring population for local search using a selection mechanism. Furthermore, in the proposed local search, a hierarchical strategy is adopted to handle the three objectives, which mainly considers the minimization of makespan, while the concern of the other two objectives is reflected in the order of trying all the possible actions that could generate the acceptable neighbor. In the experimental studies, the influence of two alternative acceptance rules on the performance of the proposed MAs is first examined. Afterwards, the effectiveness of key components in our MAs is verified, including genetic search, local search, and the hierarchical strategy in local search. Finally, extensive comparisons are carried out with the state-of-the-art methods specially presented for the MO-FJSP on well-known benchmark instances. The results show that the proposed MAs perform much better than all the other algorithms.
Yuan Yuan 0004, Hua Xu 0003
IEEE Trans Autom. Sci. Eng.2
2014 Quantum-inspired evolutionary algorithm with linkage learning
abstract
The quantum-inspired evolutionary algorithm (QEA) uses several quantum computing principles to optimize problems on a classical computer. QEA possesses a number of quantum individuals, which are all probability vectors. They work well for linear problems but fail on problems with strong interactions among variables. Moreover, many optimization problems have multiple global optima. And because of the genetic drift, these problems are difficult for evolutionary algorithms to find all global optima. Local and global migration that QEA uses to synchronize different individuals prevent QEA from finding multiple optima. To overcome these difficulties, we proposed a quantum-inspired evolutionary algorithm with linkage learning (QEALL). QEALL uses a modified concept-guide operator based on low order statistics to learn linkage. We also replaced the migration procedure by a niching technology to prevent genetic drift, accordingly to find all global optima and to expedite convergence speed. The performance of QEALL was tested on a number of benchmarks including both unimodal and multimodal problems. Empirical evaluation suggests that the proposed algorithm is effective and efficient.
Bo Wang 0051, Hua Xu 0003, Yuan Yuan 0004
IEEE Congress on Evolutionary Computation2
2014 An improved NSGA-III procedure for evolutionary many-objective optimization
abstract
Many-objective (four or more objectives) optimization problems pose a great challenge to the classical Pareto-dominance based multi-objective evolutionary algorithms (MOEAs), such as NSGA-II and SPEA2. This is mainly due to the fact that the selection pressure based on Pareto-dominance degrades severely with the number of objectives increasing. Very recently, a reference-point based NSGA-II, referred as NSGA-III, is suggested to deal with many-objective problems, where the maintenance of diversity among population members is aided by supplying and adaptively updating a number of well-spread reference points. However, NSGA-III still relies on Pareto-dominance to push the population towards Pareto front (PF), leaving room for the improvement of its convergence ability. In this paper, an improved NSGA-III procedure, called θ-NSGA-III, is proposed, aiming to better tradeoff the convergence and diversity in many-objective optimization. In θ-NSGA-III, the non-dominated sorting scheme based on the proposed θ-dominance is employed to rank solutions in the environmental selection phase, which ensures both convergence and diversity. Computational experiments have shown that θ-NSGA-III is significantly better than the original NSGA-III and MOEA/D on most instances no matter in convergence and overall performance.
Yuan Yuan 0004, Hua Xu 0003, Bo Wang 0051
GECCO2
2014 Evolutionary many-objective optimization using ensemble fitness ranking
abstract
In this paper, a new framework, called ensemble fitness ranking (EFR), is proposed for evolutionary many-objective optimization that allows to work with different types of fitness functions and ensemble ranking schemes. The framework aims to rank the solutions in the population more appropriately by combing the ranking results from many simple individual rankers. As to the form of EFR, it can be regarded as an extension of average and maximum ranking methods which have been shown promising for many-objective problems. The significant change is that EFR adopts more general fitness functions instead of objective functions, which would make it easier for EFR to balance the convergence and diversity in many-objective optimization. In the experimental studies, the influence of several fitness functions and ensemble ranking schemes on the performance of EFR is fist investigated. Afterwards, EFR is compared with two state-of-the-art methods (MOEA/D and NSGA-III) on well-known test problems. The computational results show that EFR significantly outperforms MOEA/D and NSGA-III on most instances, especially for those having a high number of objectives.
Yuan Yuan 0004, Hua Xu 0003, Bo Wang 0051
GECCO2
2014 Constrained-hLDA for Topic Discovery in Chinese Microblogs
Hua Xu 0003, Xiaoqiu Huang 0002
PAKDD (2)2
2014 Box office prediction based on microblog
Jingfei Du, Hua Xu 0003, Xiaoqiu Huang 0002
Expert Syst. Appl.2
2014 Text-based emotion classification using emotion cause extraction
Weiyuan Li, Hua Xu 0003
Expert Syst. Appl.2
2013 Implicit Feature Detection via a Constrained Topic Model and SVM
abstract
Implicit feature detection, also known as implicit feature identification, is an essential aspect of feature-specific opinion mining but previous works have often ignored it.We think, based on the explicit sentences, several Support Vector Machine (SVM) classifiers can be established to do this task.Nevertheless, we believe it is possible to do better by using a constrained topic model instead of traditional attribute selection methods.Experiments show that this method outperforms the traditional attribute selection methods by a large margin and the detection task can be completed better.
Hua Xu 0003, Xiaoqiu Huang 0002
EMNLP2
2013 A memetic algorithm for the multi-objective flexible job shop scheduling problem
abstract
In this paper, a new memetic algorithm (MA) is proposed for the muti-objective flexible job shop scheduling problem (MO-FJSP) with the objectives to minimize the makespan, total workload and critical workload. By using well-designed chromosome encoding/decoding scheme and genetic operators, the non-dominated sorting genetic algorithm II (NSGA-II) is first adapted for the MO-FJSP. Then the MA is developed by incorporating a novel local search algorithm into the adapted NSGA-II, where several mechanisms to balance the genetic search and local search are employed. In the proposed local search, a hierarchical strategy is adopted to handle the three objectives, which mainly considers the minimization of makespan, while the concern of the other two objectives is reflected in the order of trying all the possible actions that could generate the acceptable neighbor. Experimental results on well-known benchmark instances show that the proposed MA outperforms significantly two off-the-shelf multi-objective evolutionary algorithms and four state-of-the-art algorithms specially proposed for the MO-FJSP.
Yuan Yuan 0004, Hua Xu 0003
GECCO2
2012 HHS/LNS: An integrated search method for flexible job shop scheduling
abstract
The flexible job shop scheduling problem (FJSP) is a generalization of the classical job shop scheduling problem (JSP), where each operation is allowed to be processed by any machine from a given set, rather than one specified machine. In this paper, two algorithm modules, namely, hybrid harmony search (HHS) and large neighborhood search (LNS) are developed for the FJSP with makespan criterion. The HHS is an evolutionary-based algorithm with the memetic paradigm, while the LNS is typical of constraint-based approaches. To form a stronger search mechanism, an integrated search method is proposed for the FJSP based on the two algorithms, which starts with the HHS, and then the solution is further improved by the LNS. Computational simulations and comparisons demonstrate that, the proposed HHS alone can effectively solve some medium to large FJSP instances, when integrated with the LNS, it shows competitive performance with state-of-the-art algorithms on very hard and large-scale problems, some new upper bounds among the unsolved benchmark instances have even been found.
Yuan Yuan 0004, Hua Xu 0003
IEEE Congress on Evolutionary Computation2
2012 Weakness Finder: Find product weakness from Chinese reviews by using aspects based sentiment analysis
Wenhao Zhang 0003, Hua Xu 0003
Expert Syst. Appl.2
2011 Questionnaires-based skin attribute prediction using Elman neural network
Hua Xu 0003, Wenhao Zhang 0003, Xincheng Hu
Neurocomputing2
2009 Duple-EDA and sample density balancing
Yunpeng Cai, Hua Xu 0003, Xiaomin Sun 0001, Peifa Jia, ZeHua Liu
Sci. China Ser. F Inf. Sci.2
2007 Cross entropy and adaptive variance scaling in continuous EDA
abstract
This paper deals with the adaptive variance scaling issue incontinuous Estimation of Distribution Algorithms. A phenomenon is discovered that current adaptive variance scaling method in EDA suffers from imprecise structure learning. A new type of adaptation method is proposed to overcome this defect. The method tries to measure the difference between the obtained population and the prediction of the probabilistic model, then calculate the scaling factor by minimizing the cross entropy between these two distributions. This approach calculates the scaling factor immediately rather than adapts it incrementally. Experiments show that this approach extended the class of problems that can be solved, and improve the search efficiency in some cases. Moreover, the proposed approach features in that each decomposed subspace can be assigned an individual scaling factor, which helps to solve problems with special dimension property.
Yunpeng Cai, Xiaomin Sun 0001, Hua Xu 0003, Peifa Jia
GECCO3