Xin Li 0137

dblp:09/1365-137 · DBLP profile ↗
← Back
13ranked-venue papers
1as first author
8since 2021 · last 2026
0000-0003-2411-5374ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Human-computer interaction and ubiquitous computing · 2Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 EAP-LSTM: A Bi-LSTM-Based Deep Learning Framework for Quantitatively Predicting Enhancer Activity in Drosophila and Human Cell Lines
abstract
Enhancer activity plays a critical role in gene regulation, influencing various biological processes such as development and disease progression. Accurate prediction of enhancer activity is essential for understanding the mechanisms underlying gene regulation and enhancer function. This study introduces a novel deep learning framework, EAP-LSTM (Enhancer Activity Prediction based on Bi-LSTM), to quantitatively predict enhancer activity across different species and cell lines. The model integrates multiple feature modules, including Word2Vec-based representations of DNA sequences, reverse complement k-mer, mismatch k-mer features, and epigenomic data. Evaluated on six cell lines, including five human cell lines (A549, HCT116, HepG2, K562, and MCF-7) and one Drosophila cell line (S2), EAP-LSTM consistently outperforms state-of-the-art models, such as DeepSTARR and HEAP, in all datasets. For example, on the K562 dataset, EAP-LSTM achieves a Pearson correlation coefficient (PCC) of 0.7944, outperforming DeepSTARR and HEAP by 13.65% and 2.73%, respectively. In addition, EAP-LSTM demonstrates strong performance in small-sample learning scenarios, showing clear improvements compared with baseline models. Furthermore, the study investigates the role of transcription factor binding sites (TFBSs) within enhancer regions, identifying critical motifs associated with enhancer activity. These findings not only improve enhancer prediction accuracy but also provide valuable insights into the molecular mechanisms underlying enhancer function.
Lichang Dai, Yu Dou, Xin Li 0137, Chang Lu 0015, Hao Wu 0062
IEEE J. Biomed. Health Informatics4
2025 HAP: Hybrid Adaptive Parallelism for Efficient Mixture-of-Experts Inference
abstract
Current inference systems for Mixture-of-Experts (MoE) models primarily employ static parallelization strategies. However, these static approaches cannot consistently achieve optimal performance across different inference scenarios, as they lack the flexibility to adapt to varying computational requirements. In this work, we propose HAP (Hybrid Adaptive Parallelism), a novel method that dynamically selects hybrid parallel strategies to enhance MoE inference efficiency. The fundamental innovation of HAP lies in hierarchically decomposing MoE architectures into two distinct computational modules: the Attention module and the Expert module, each augmented with a specialized inference latency simulation model. This decomposition promotes the construction of a comprehensive search space for seeking model parallel strategies. By leveraging Integer Linear Programming (ILP), HAP could solve the optimal hybrid parallel configurations to maximize inference efficiency under varying computational constraints. Our experiments demonstrate that HAP consistently determines parallel configurations that achieve comparable or superior performance to the TP strategy prevalent in mainstream inference systems. Compared to the TP-based inference, HAP-based inference achieves speedups of$1.68 \times, 1.77 \times$, and$1.57 \times$on A100, A6000, and V100 GPU platforms, respectively. Furthermore, HAP showcases remarkable generalization capability, maintaining performance effectiveness across diverse MoE model configurations, including Mixtral and Qwen series models.
Xianzhi Yu, Zongyuan Zhan, Wulong Liu, Zekun Yin, Xin Li 0137
ICPADS9
2025 A novel deep learning framework with dynamic tokenization for identifying chromatin interactions along with motif importance investigation
abstract
A comprehensive understanding of chromatin interaction networks is crucial for unraveling the regulatory mechanisms of gene expression. While various computational methods have been developed to predict chromatin interactions and address the limitations and high costs of high-throughput experimental techniques, their performance is often overestimated due to the specificity of chromatin interaction data. In this study, we proposed Inter-Chrom, a novel deep learning model integrating dynamic tokenization, DNABERT's word embedding, and the efficient channel attention mechanism to identify chromatin interactions using sequence and genomic features, leveraging a newly curated dataset. Experimental results demonstrate that Inter-Chrom outperforms existing methods on three cell line datasets. Additionally, we proposed a novel method for calculating motif importance and analyzed the motifs with high importance scores identified through this method, including those that have been extensively studied and others that have received limited attention to date. Inter-Chrom's robustness for input variations and superior ability to leverage sequence features position it as a powerful tool for advancing chromatin interaction research. The source code of Inter-Chrom is freely available at https://github.com/HaoWuLab-Bioinformatics/Inter-Chrom.
Liangcan Li, Xin Li 0137, Hao Wu 0062
Briefings Bioinform.2
2025 RabbitTrim: An Efficient and Versatile Trimmer on Multi-Core Platforms
abstract
Trimming is an essential step in sequencing data processing. However, many existing trimming tools, such as Trimmomatic and Ktrim, are limited by suboptimal implementations and fail to fully leverage the computational power of modern multi-core platforms. To address this, we introduce RabbitTrim, a highly optimized and versatile trimming tool that fully supports the functionalities of Trimmomatic and Ktrim. RabbitTrim's performance is enhanced through efficient I/O strategies, parallel (de)compression engines, block-based memory pools, bitwise operations, and vectorization techniques. Compared to Trimmomatic, RabbitTrim (in trimmomatic mode) achieves speedups ranging from 1.8x to 6.0x for plain FASTQ files and 3.7x to 14.0x for gzip-compressed FASTQ files on a 48-core Intel server. Similarly, compared to Ktrim, RabbitTrim (in ktrim mode) achieves speedups ranging from 1.5x to 2.5x for plain FASTQ files and 2.7x to 5.6x for gzip-compressed FASTQ files on the same server. Moreover, RabbitTrim is able to process 101 GB gzip-compressed sequencing data in only 5 minutes while Trimmomatic requires at least 21 minutes.
Zekun Yin, Lifeng Yan, Fangjin Zhu, Xin Li 0137, Xiaohui Duan, Bertil Schmidt
IEEE Trans. Comput. Biol. Bioinform.7
2024 Iteratively Calibrating Prompts for Unsupervised Diverse Opinion Summarization
abstract
Diverse opinion summarization aims to generate a summary that captures multiple opinions in texts. Although large language models (LLMs) have become the main choice for this task, the performance is highly depend on prompts. In this paper, we propose a self-evaluation based prompt calibration framework to stimulate LLM for generating high quality summary. It adopts the reinforcement learning mechanism to calibrate prompts for maximizing the reward of summary. The framework contains three parts. In the prompt construction part, we design the prompt that contains topic, task instruction and key opinion reference. The topic indicates the main focus of documents, the instruction describes the task with natural language and the key opinion reference is the explicit constraint on the expected opinions. In the reward part, for each summary, its coverage score and diversity score are used to represent the semantic coverage to the source documents and the inter opinion differences, respectively. The prompt calibration part selects the sentences in generated summaries to calibrate the prompts for the next iteration. With this framework, we use a LLM with 7B parameters to generate summaries, which outperforms large GPT-4 and multiple strong baselines. The ablation studies indicate the effectiveness of the iterative calibration process. We analyze the opinion difference in terms of the tendencies of sentences in summaries and use the Natural Language Inference (NLI)-based method to evaluate the faithfulness of summaries. Experiment results show that our method generates summaries with high opinion difference and faithfulness.
Jian Wang 0118, Yuqing Sun 0001, Xin Li 0137
ECAI4
2024 RabbitTrim: Highly Optimized Trimming of Illumina Sequencing Data on Multi-core Platforms
Zekun Yin, Lifeng Yan, Fangjin Zhu, Xiaohui Duan, Xin Li 0137, Bertil Schmidt
ISBRA (2)7
2023 SGRU: A High-Performance Structured Gated Recurrent Unit for Traffic Flow Prediction
abstract
Traffic flow prediction is an essential task in constructing smart cities and is a typical Multivariate Time Series (MTS) Problem. Recent research has abandoned Gated Recurrent Units (GRU) and utilized dilated convolutions or temporal slicing for feature extraction, and they have the following drawbacks: (1) Dilated convolutions fail to capture the features of adjacent time steps, resulting in the loss of crucial transitional data. (2) The connections within the same temporal slice are strong, while the connections between different temporal slices are too loose. In light of these limitations, we emphasize the importance of analyzing a complete time series repeatedly and the crucial role of GRU in MTS. Therefore, we propose SGRU: Structured Gated Recurrent Units, which involve structured GRU layers and non-linear units, along with multiple layers of time embedding to enhance the model’s fitting performance. We evaluate our approach on four publicly available California traffic datasets: PeMS03, PeMS04, PeMS07, and PeMS08 for regression prediction. Experimental results demonstrate that our model outperforms baseline models with average improvements of 11.7%, 18.6%, 18.5%, and 12.0% respectively.
Wenfeng Zhang, Xin Li 0137, Ti Wang, Honglei Gao
ICPADS2
2023 A Comprehensive Memory Management Framework for CPU-FPGA Heterogenous SoCs
abstract
Efficient utilization of restrained memory resources is of paramount importance in CPU-FPGA heterogeneous multiprocessor system-on-chip (HMPSoC)-based system design for memory-intensive applications. State-of-the-art high level synthesis (HLS) tools rely on the system programmers to manually determine the data placement within the complex memory hierarchy. Different data placement policies may lead to different system performance, and finding an optimal data placement policy is a nontrivial problem. For instance, we show counter-intuitive results that traditional frequency and locality-based data placement strategy designed for CPU architecture leads to nonoptimal system performance in CPU-FPGA HMPSoCs. In this work, we first propose an automatic data placement framework for field programmable gate array (FPGA) kernels to determine whether each array object should be accessed via the on-chip BRAM, shared CPU L2-cache, or DDR memory to achieve the optimal performance. Moreover, we find that when the CPU kernel and the FPGA kernel are executed in parallel, memory contentions may degrade the performance and the optimal data placement policy designed for the FPGA kernel alone will not achieve the optimal overall system performance. In this article, we proposed to use cache partitioning to alleviate the impact brought by memory contentions. We extend the framework designed for FPGA by adding the cross-layer memory contentions analysis to automatically generate an optimal data placement policy and cache partitioning mechanism for the parallel executing kernels. The proposed data placement framework can be seamlessly integrated with the commercial Vivado HLS. The experimental results on the Zedboard platform show an average$1.5\times $performance speedup for FPGA kernels compared with a greedy-based allocation strategy. When FPGA kernels and CPU kernels are executed in parallel, the FPGA kernel and the CPU kernel have a performance speedup of$1.62\times $and$1.10\times $on average, respectively.
Zelin Du, Qianling Zhang, Mao Lin, Shiqing Li, Xin Li 0137, Lei Ju 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2020 SWMapper: Scalable Read Mapper on SunWay TaihuLight
abstract
With the rapid development of next-generation sequencing (NGS) technologies, high throughput sequencing platforms continuously produce large amounts of short read DNA data at low cost. Read mapping is a performance-critical task, being one of the first stages required for many different types of NGS analysis pipelines. We present SWMapper — a scalable and efficient read mapper for the Sunway TaihuLight supercomputer. A number of optimization techniques are proposed to achieve high performance on its heterogeneous architecture which are centered around a memory-efficient succinct hash index data structure including seed filtration, duplicate removal, dynamic scheduling, asynchronous data transfer, and overlapping I/O and computation. Furthermore, a vectorized version of the banded Myers algorithm for pairwise alignment with 256-bit vector registers is presented to fully exploit the computational power of the SW26010 processor. Our performance evaluation shows that SWMapper using all 4 compute groups of a single Sunway TaihuLight node outperforms S-Aligner on the same hardware by a factor of 6.2. In addition, compared the state-of-the-art CPU-based mappers RazerS3, BitMapper2, and Hobbes3 running on a 4-core Xeon W-2123v3 CPU, SWMapper achieves speedups of 26.5, 7.8, and 2.6, respectively. Our optimizations achieve an aggregated speedup of 11 compared to the naïve implementation on one compute group of an SW26010 processor as well as a strong scaling efficiency of 74% on 128 compute groups.
Xiaohui Duan, Xiangxu Meng, Xin Li 0137, Bertil Schmidt
ICPP4
2018 Collaboratively Learning Latent Factors and Correlations for New Paper Influence Predication
abstract
There are an increasing number of papers published every year. It is desired for researchers to find the new high-quality papers, which is also a challenging task due to the lack of citation information. In this paper, we propose a novel method to predicate a new paper influence by collaboratively learning the latent vectors of paper features and correlations. We propose the concept topic related authority to integrate the dynamic topic model with paper citations so as to learn how content and authors influence a paper quality. We adopt the Factorization Machine method to collaboratively learn the latent vectors of correlations between different paper features. Comparing with traditional methods, it does not require the citation information to evaluate a paper quality, which is appropriate for new published papers. We conduct extensive evaluation against a real dataset crawled from ACM Digital Library. The results show that our method outperforms the other methods.
Yuqing Sun 0001, Xin Li 0137
CSCWD3
2018 Modeling User Intrinsic Characteristic on Social Media for Identity Linkage
abstract
Most users on social media have intrinsic characteristics, such as interests and political views, that can be exploited to identify and track them. It raises privacy and identity issues in online communities. In this paper we investigate the problem of user identity linkage on two behavior datasets collected from different experiments. Specifically, we focus on user linkage based on users' interaction behaviors with respect to content topics. We propose an embedding method to model a topic as a vector in a latent space so as to interpret its deep semantics. Then a user is modeled as a vector based on his or her interactions with topics. The embedding representations of topics are learned by optimizing the joint-objective: the compatibility between topics with similar semantics, the discriminative abilities of topics to distinguish identities, and the consistency of the same user's characteristics fromtwo datasets. The effectiveness of our method is verified on real-life datasets and the results show that it outperforms related methods.
Xianqi Yu, Yuqing Sun 0001, Elisa Bertino, Xin Li 0137
GROUP4
2018 Position prediction system based on spatio-temporal regularity of object mobility
Xin Li 0137, Chongsheng Yu, Lei Ju 0001, Lei Dou, Yuqing Sun 0001
Inf. Syst.1
2016 Implement and Optimization of Indoor Positioning System Based on Wi-Fi Signal
Chongsheng Yu, Xin Li 0137, Lei Dou, Yuqing Sun 0001, Zhiyue Cao
ICA3PP2