Xindong You

dblp:97/3955 · DBLP profile ↗
← Back
22ranked-venue papers
5as first author
16since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 3 first-author · 8 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-authorComputer networks · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Semantic-driven seasonal data classification: An artificial intelligence-enabled cost-effective storage system
Xueqiang Lv, Yunchao Gong, Xiao Qin 0001, Xindong You
Eng. Appl. Artif. Intell.6
2026 Innovative weighted clustering for categorical matrix-object data: new distance and cluster center considering data distribution
Liqin Yu, Fuyuan Cao, Likun Lu, Xindong You
J. Supercomput.4
2025 CE-DCVSI: Multimodal relational extraction based on collaborative enhancement of dual-channel visual semantic information
Yunchao Gong, Xueqiang Lv, Zangtai Cai, Yuzhong Chen 0003, Zhaojun Wang, Xindong You
Expert Syst. Appl.8
2025 Stronger interaction brings better performance: fine-grained alignment between different modalities for MNER
Xindong You, Jiang Jian, Xueqiang Lv
J. Supercomput.1
2024 TL-RelD: Tight-Loose Pairwise Loss for Object Re-Identification
Changwang Mei, Xindong You, Shangzhi Teng, Xueqiang Lyu
PRCV (12)2
2024 GNN-Based Multimodal Named Entity Recognition
abstract
Abstract The Multimodal Named Entity Recognition (MNER) task enhances the text representations and improves the accuracy and robustness of named entity recognition by leveraging visual information from images. However, previous methods have two limitations: (i) the semantic mismatch between text and image modalities makes it challenging to establish accurate internal connections between words and visual representations. Besides, the limited number of characters in social media posts leads to semantic and contextual ambiguity, further exacerbating the semantic mismatch between modalities. (ii) Existing methods employ cross-modal attention mechanisms to facilitate interaction and fusion between different modalities, overlooking fine-grained correspondences between semantic units of text and images. To alleviate these issues, we propose a graph neural network approach for MNER (GNN-MNER), which promotes fine-grained alignment and interaction between semantic units of different modalities. Specifically, to mitigate the issue of semantic mismatch between modalities, we construct corresponding graph structures for text and images, and leverage graph convolutional networks to augment text and visual representations. For the second issue, we propose a multimodal interaction graph to explicitly represent the fine-grained semantic correspondences between text and visual objects. Based on this graph, we implement deep-level feature fusion between modalities utilizing graph attention networks. Compared with existing methods, our approach is the first to extend graph deep learning throughout the MNER task. Extensive experiments on the Twitter multimodal datasets validate the effectiveness of our GNN-MNER.
Yunchao Gong, Xueqiang Lv, Xindong You, Yuzhong Chen 0003
Comput. J.4
2024 CSEA: A Fine-Grained Framework of Climate-Season-Based Energy-Aware in Cloud Storage Systems
abstract
Abstract Continuous data scale growth increases energy consumption and operating cost that cannot be ignored in cloud storage systems. Previous studies have shown that analyzing the characteristics of I/O access and mining data features is effective for reasonable data distribution in storage systems. The granularity and criterion of classification are the key factors in determining the data distribution. To decrease energy consumption and operating cost, this paper puts forward a fine-grained framework of the climatic-season-based energy-aware in cloud storage system called CSEA. The framework concludes the following three aspects: (i) data feature mining. CSEA discovers potential data features by analyzing data access to provide help with data classification. (ii) K-means clustering algorithm. CSEA uses an unsupervised data classification algorithm in machine learning to divide data into categories based on seasonal characteristics by gathering real I/O access. (iii) data distribution of fine-grained. On the basis of seasonal features, CSEA fuses regional features to further refine the data distribution granularity to save on energy consumption and operating cost. Simulation experiments using extended CloudSimDisk and the constructed mathematical models indicate that CSEA reduces the energy consumption and operating cost compared with the single data classification standard and coarse-grained data distribution.
Xueqiang Lv, Haojie Ge, Xindong You
Comput. J.5
2024 Cost-effective data classification storage through text seasonal features
abstract
Data classification storage has emerged as an effective strategy, harnessing the diverse performance attributes of storage devices and orchestrating a harmonious equilibrium between energy consumption, cost considerations, and user accessibility. As research on emerging storage media (e.g. Non-Volatile Memory (NVM)) delves deeper, in scenarios characterized by dynamically evolving storage demands, conventional paradigms of rigid classification strategy and Solid-State Drive (SSD) and Hard Disk Drive (HDD) storage architectures fall short of addressing such complex situations. In this paper, we propose an effective data classification storage method using text seasonal features based on the traditional access frequency analysis strategy. First, to procure richer semantic information, we employ external knowledge actualizing the short-text feature expansion. Then, we leverage the ensemble learning stacking method optimized models to improve the accuracy of data classification based on seasonal features. Additionally, following the seasonal feature, we further classify the data into hot, warm, and cold and place them separately on the NVM, SSD, and HDD, which conserves storage energy consumption and operational costs while ensuring the quality of user access. The experimental results demonstrate that data classification accuracy can reach more than 95.10%, and the energy consumption and operating cost can be reduced by more than 30.22% and 8.73%, respectively.
Xueqiang Lv, Yunchao Gong, Taifu Yuan, Xindong You
Future Gener. Comput. Syst.6
2024 Multimodal heterogeneous graph entity-level fusion for named entity recognition with multi-granularity visual guidance
Yunchao Gong, Xueqiang Lv, Zhaojun Wang, Xindong You
J. Supercomput.6
2024 A relation enhanced model for temporal knowledge graph alignment
Zhaojun Wang, Xindong You, Xueqiang Lv
J. Supercomput.2
2023 HBert: A Long Text Processing Method Based on BERT and Hierarchical Attention Mechanisms
abstract
With the emergence of a large-scale pre-training model based on the transformer model, the effect of all-natural language processing tasks has been pushed to a new level. However, due to the high complexity of the transformer's self-attention mechanism, these models have poor processing ability for long text. Aiming at solving this problem, a long text processing method named HBert based on Bert and hierarchical attention neural network is proposed. Firstly, the long text is divided into multiple sentences whose vectors are obtained through the word encoder composed of Bert and the word attention layer. And the article vector is obtained through the sentence encoder that is composed of transformer and sentence attention. Then the article vector is used to complete the subsequent tasks. The experimental results show that the proposed HBert method achieves good results in text classification and QA tasks. The F1 value is 95.7% in longer text classification tasks and 75.2% in QA tasks, which are better than the state-of-the-art model longformer.
Xueqiang Lv, Zhaonan Liu, Xindong You
Int. J. Semantic Web Inf. Syst.5
2023 Research on the Generation of Patented Technology Points in New Energy Based on Deep Learning
abstract
Effective extraction of patent technology points in new energy fields is profitable, which motivates technological innovation and facilitates patent transformation and application. However, since patent data exists the ununiform distribution of technology points information, long length of term, and long sentences, technology point extraction faces the dilemmas of poor readability and logic confusion. To mitigate these problems, the article proposes a method to generate patent technology points called IGPTP—a two-stage strategy, which fuses the advantage of extractive and generative ways. IGPTP utilizes the RoBERTa+CNN model to obtain the key sentences of text and takes the output as input of UNILM (unified pre-trained language model). Simultaneously, it takes a multi-strategies integration technique to enhance the quality of patent technology points by combining the copy mechanism and external knowledge guidance model. Substantial experimental results manifest that IGPTP outperforms the current mainstream models, which can generate more coherent and richer text.
Haixiang Yang, Xindong You, Xueqiang Lv
Int. J. Semantic Web Inf. Syst.2
2022 MQDS: An energy saving scheduling strategy with diverse QoS constraints towards reconfigurable cloud storage systems
Xindong You, Dawei Sun 0001, Xueqiang Lv, Shang Gao 0003, Rajkumar Buyya
Future Gener. Comput. Syst.1
2021 SS6: Online Short-Code RAID-6 Scaling by Optimizing New Disk Location and Data Migration
abstract
Abstract Thanks to excellent reliability, availability, flexibility and scalability, redundant arrays of independent (or inexpensive) disks (RAID) are widely deployed in large-scale data centers. RAID scaling effectively relieves the storage pressure of the data center and increases both the capacity and I/O parallelism of storage systems. To regain load balancing among all disks including old and new, some data usually are migrated from old disks to new disks. Owing to unique parity layouts of erasure codes, traditional scaling approaches may incur high migration overhead on RAID-6 scaling. This paper proposes an efficient approach based Short-Code for RAID-6 scaling. The approach exhibits three salient features: first, SS6 introduces $\tau $ to determine where new disks should be inserted. Second, SS6 minimizes migration overhead by delineating migration areas. Third, SS6 reduces the XOR calculation cost by optimizing parity update. The numerical results and experiment results demonstrate that (i) SS6 reduces the amount of data migration and improves the scaling performance compared with Round-Robin and Semi-RR under offline, (ii) SS6 decreases the total scaling time against Round-Robin and Semi-RR under two real-world I/O workloads (iii) the user average response time of SS6 is better than the other two approaches during scaling and after scaling.
Xindong You, Xueqiang Lv
Comput. J.2
2021 K-ear: Extracting data access periodic characteristics for energy-aware data clustering and storing in cloud storage systems
abstract
Abstract Rapid increase in energy consumption is a serious problem in cloud storage systems. Data accessed in large‐scale storage systems usually exhibit temporal and spatial characteristics, which make it possible to reduce energy consumption by clustering data with similar access characteristics for storage in the same zone of cloud storage systems. Existing works usually only focus on the frequency of data access. However, widely existing phenomena show data access with seasonal and tidal characteristics in cloud storage systems. The seasonal and tidal characteristics of data access are extracted thoroughly in this paper. According to the extracted data access characteristics, energy‐aware data clustering through a machine learning algorithm (K‐ear) is proposed. K‐ear classifies data into five seasonal categories according to their seasonal access characteristics and then classifies every seasonal category into three tidal categories according to its tidal access characteristics. The 15 classified categories are stored in different storage zones with different energy and performance modes. Simulation experiments using CloudSimDisk with the constructed mathematic models demonstrate that the proposed K‐ear algorithm is more energy‐efficient than the default data clustering algorithms in Hadoop and the classical data clustering storage strategy according to the data access frequency (Striping‐Based Energy‐Aware Strategy).
Xindong You, Dawei Sun 0001, Xunyun Liu, Xueqiang Lv, Rajkumar Buyya
Concurr. Comput. Pract. Exp.1
2021 HS6: An Efficient H-Code RAID-6 Scaling by Optimizing Data Migrating and Parity Updating
Xindong You, Xueqiang Lv, Muyuan Li
J. Supercomput.2
2020 Sentiment Analysis of Film Reviews Based on Deep Learning Model Collaborated with Content Credibility Filtering
Xindong You, Xueqiang Lv, Shangqian Zhang, Dawei Sun 0001, Shang Gao 0003
CollaborateCom (1)1
2020 Dynamic redirection of real-time data streams for elastic stream computing
Dawei Sun 0001, Shang Gao 0003, Xunyun Liu, Xindong You, Rajkumar Buyya
Future Gener. Comput. Syst.4
2020 Distant Supervised Relation Extraction via DiSAN-2CNN on a Feature Level
abstract
At present, the mainstream distant supervised relation extraction methods existed problems: the coarse granularity for coding the context feature information; the difficulty in capturing the long-term dependency in the sentence, and the difficulty in coding prior knowledge of structures are major issues. To address these problems, we propose a distant supervised relation extraction model via DiSAN-2CNN on feature level, in which multi-dimension self-attention mechanism is utilized to encode the features of the words and DiSAN-2CNN is used to encode the sentence to obtain the long-term dependency, the prior knowledge of the structure, the time sequence, and the entity dependence in the sentence. Experiments conducted on the NYT-Freebase benchmark dataset demonstrate that the proposed DiSAN-2CNN on a feature level model achieves better performance than the current two state-of-art distant supervised relation extraction models PCNN+ATT and ResCNN-9, and it has d generalization ability with the least artificial feature engineering.
Xueqiang Lv, Huixin Hou, Xindong You, Junmei Han
Int. J. Semantic Web Inf. Syst.3
2019 Relation Extraction Toward Patent Domain Based on Keyword Strategy and Attention+BiLSTM Model (Short Paper)
Xueqiang Lv, Xiangru Lv, Xindong You, Zhian Dong, Junmei Han
CollaborateCom3
2018 QGLG Automatic Energy Gear-Shifting Mechanism with Flexible QoS Constraint in Cyber-Physical Systems: Designing, Analysis, and Evaluation
abstract
This article describes how with the continuous expansion on the volume of data produced by sensors in Cyber Physical Systems, the scale of the cloud storage system has become larger. This will lead to the problems of a high energy consumption rate and a low utilization becoming a serious issue. In order to enhance the effective energy consumption, reduce the invalid energy consumption, and supply more flexible QoS for users in CPS, this article proposes an automatic energy gear-shifting mechanism with flexible QoS constraints (QGLG). The QGLG predicts system load of the follow-up period through a support vector machine model. According to the current system load, the predicted load, and the flexible QoS, QGLG automatically up-shifts and down-shifts among nodes. Substantive results from the simulation experiments done on GridSim show that the QGLG can achieve energy consumption reduction while satisfying the user's flexible QoS requirements. Compared with a similar energy-reducing mechanism, QGLG has its obvious advantage when considering the requirements of user with energy saved notwithstanding.
Xindong You, Yeli Li, Zhenyang Zhu, Lifeng Yu, Dawei Sun 0001
J. Database Manag.1
2010 Power aware job scheduling with QoS guarantees based on feedback control
abstract
With the scale of computing system increases, power consumption has become the major challenge to system performance, reliability and IT management costs. Specifically, system performance and reliability, described by various Quality of Service(QoS) metrics, cannot be guaranteed if the objective is to minimize the total power consumption solely, despite of the violations of QoS. Various methods have been developed to control power consumption to avoid system failures and thermal emergencies through coarse-grained designs. However, the existing methods can be improved and more power can be saved if fine-grained job level adaptation is integrated into them. In this paper a feedback control based power aware job scheduling algorithm is proposed to minimize power consumption in computing system and to provide QoS guarantees. In the proposed algorithm, jobs are scheduled according to the realtime and historical power consumption as well as the QoS requirements. Simulations and experiments on real multi core computing system show that the power potential of the system can be deeply explored while still providing QoS guarantees and the performance degradation is acceptable. The experiment results also show that fine-grained job-level power aware scheduling can achieve better power/performance balancing between multiple processors or cores than coarse-grained methods.
Congfeng Jiang, Xianghua Xu, Jian Wan 0001, Xindong You, Ritai Yu
IWQoS5