Cheng Ling

dblp:85/7641 · DBLP profile ↗
← Back
16ranked-venue papers
5as first author
4since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 6 · 3 first-authorSystems, architecture and hardware · 4 · 2 first-authorDatabases, data management, data science and information retrieval · 4 · 4 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Computer networks · 1Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2025 Enhancing Sequential Recommender with Large Language Models for Joint Video and Comment Recommendation
Bowen Zheng 0005, Enze Liu 0005, Chen Yang 0032, Enyang Bai, Cheng Ling, Han Li 0005, Wayne Xin Zhao, Ji-Rong Wen
RecSys6
2025 Generative Recommender with End-to-End Learnable Item Tokenization
abstract
Generative recommender systems have gained increasing attention as an innovative approach that directly generates item identifiers for recommendation tasks. Despite their potential, a major challenge is the effective construction of item identifiers that align well with recommender systems. Current approaches often treat item tokenization and generative recommendation training as separate processes, which can lead to suboptimal performance. To overcome this issue, we introduce ETEGRec, a novel End-To-End Generative Recommender that unifies item tokenization and generative recommendation into a cohesive framework. Built on a dual encoder-decoder architecture, ETEGRec consists of an item tokenizer and a generative recommender. To enable synergistic interaction between these components, we propose a recommendation-oriented alignment strategy, which includes two key optimization objectives: sequence-item alignment and preference-semantic alignment. These objectives tightly couple the learning processes of the item tokenizer and the generative recommender, fostering mutual enhancement. Additionally, we develop an alternating optimization technique to ensure stable and efficient end-to-end training of the entire framework. Extensive experiments demonstrate the superior performance of our approach compared to traditional sequential recommendation models and existing generative recommendation baselines. Our code is available at https://github.com/RUCAIBox/ETEGRec.
Enze Liu 0005, Bowen Zheng 0005, Cheng Ling, Lantao Hu, Han Li 0005, Wayne Xin Zhao
SIGIR3
2024 Modeling User Fatigue for Sequential Recommendation
abstract
Recommender systems filter out information that meets user interests. However, users may be tired of the recommendations that are too similar to the content they have been exposed to in a short historical period, which is the so-called user fatigue. Despite the significance for a better user experience, user fatigue is seldom explored by existing recommenders. In fact, there are three main challenges to be addressed for modeling user fatigue, including what features support it, how it influences user interests, and how its explicit signals are obtained. In this paper, we propose to model user Fatigue in interest learning for sequential Recommendations (FRec). To address the first challenge, based on a multi-interest framework, we connect the target item with historical items and construct an interest-aware similarity matrix as features to support fatigue modeling. Regarding the second challenge, built upon feature cross, we propose a fatigue-enhanced multi-interest fusion to capture long-term interest. In addition, we develop a fatigue-gated recurrent unit for short-term interest learning, with temporal fatigue representations as important inputs for constructing update and reset gates. For the last challenge, we propose a novel sequence augmentation to obtain explicit fatigue signals for contrastive learning. We conduct extensive experiments on real-world datasets, including two public datasets and one large-scale industrial dataset. Experimental results show that FRec can improve AUC and GAUC up to 0.026 and 0.019 compared with state-of-the-art models, respectively. Moreover, large-scale online experiments demonstrate the effectiveness of FRec for fatigue reduction. Our codes are released at https://github.com/tsinghua-fib-lab/SIGIR24-FRec.
Nian Li 0001, Xin Ban, Cheng Ling, Chen Gao 0001, Lantao Hu, Peng Jiang 0002, Kun Gai, Yong Li 0008, Qingmin Liao
SIGIR3
2022 Multi-granularity Fatigue in Recommendation
abstract
Personalized recommendation aims to provide appropriate items according to user preferences mainly from their behaviors. Excessive homogeneous user behaviors on similar items will lead to fatigue, which may decrease user activeness and degrade user experience. However, existing models seldom consider user fatigue in recommender systems. In this work, we propose a novel multi-granularity fatigue, modeling user fatigue from coarse to fine. Specifically, we focus on the recommendation feed scenario, where the underexplored global session fatigue and coarse-grained taxonomy fatigue have large impacts. We conduct extensive analyses to demonstrate the characteristics and influence of different types of fatigues in real-world recommender systems. In experiments, we verify the effectiveness of multi-granularity fatigue in both offline and online evaluations. Currently, the fatigue-enhanced model has also been deployed on a widely-used recommendation system of WeChat.
Ruobing Xie, Cheng Ling, Feng Xia 0006, Leyu Lin
CIKM2
2020 The Likelihood Prediction of Phylogenetic Trees based on Artificial Neural Network: a new perspective and preliminary attempt
abstract
Bayesian Evolutionary Analysis Sampling Trees (BEAST) is a widely spread phylogenetic inference tool using empirical evolution models and Bayesian statistics. However, the cost of calculating the likelihood function for massive sampled trees is very expensive, resulting in long execution time. For accelerating the process, this paper proposes a likelihood prediction model based on Artificial Neural Network (ANN) using the deep neighbor information between nodes from the topology representations of historical evolution trees. The experimental results indicate that the proposed method achieves 1.2-5.9x speedup factors on obtaining the likelihood probabilities in BEAST.
Cheng Ling
BIBM2
2020 A novel synonymous processing method based on amino acid substitution matrics for the classification of G-protein-coupled receptors
abstract
Extracting valuable features and filtering out redundancy are the key challenges to determine the overall classification performance for G-protein-coupled receptors (GPCRs). In this study, we consider improving the feature synonym problem, and put forward a novel feature knowledge mining strategy based on functional word clustering and integration. The essence behind the method is the novel feature knowledge mining strategy. Through evaluating the independence of each candidate feature using the evolutionary hypothesis based on residue substitution matrices, clustering candidate features, and fusing them by retaining the main functional words, the proposed strategy adds a layer between the feature extraction layer and the prediction layer. Based on the proposed method, four classic machine learning algorithms in conjunction with the feature extraction method were applied to classify GPCRs at all family levels. Surprisingly, these classifiers achieve considerable performance in almost all evaluation criteria which indicated the validity and superiority of the proposed molecular evolution based feature extraction method.
Cheng Ling, Yitian Shen, Jingyang Gao
BIBM1
2020 Deep Feedback Network for Recommendation
abstract
Both explicit and implicit feedbacks can reflect user opinions on items, which are essential for learning user preferences in recommendation. However, most current recommendation algorithms merely focus on implicit positive feedbacks (e.g., click), ignoring other informative user behaviors. In this paper, we aim to jointly consider explicit/implicit and positive/negative feedbacks to learn user unbiased preferences for recommendation. Specifically, we propose a novel Deep feedback network (DFN) modeling click, unclick and dislike behaviors. DFN has an internal feedback interaction component that captures fine-grained interactions between individual behaviors, and an external feedback interaction component that uses precise but relatively rare feedbacks (click/dislike) to extract useful information from rich but noisy feedbacks (unclick). In experiments, we conduct both offline and online evaluations on a real-world recommendation system WeChat Top Stories used by millions of users. The significant improvements verify the effectiveness and robustness of DFN. The source code is in https://github.com/qqxiaochongqq/DFN.
Ruobing Xie, Cheng Ling, Yalong Wang, Rui Wang 0068, Feng Xia 0006, Leyu Lin
IJCAI2
2018 cnnCNV: A Sensitive and Efficient Method for Detecting Copy Number Variation based on Convolutional Neural Networks
Miaosen Ding, Jingyang Gao, Cheng Ling, Liwei Gao
BIBM3
2018 Classification of G-protein Coupled Receptors Based on Semi-navïe Bayesian Inference
Cheng Ling, Lin Yue, Jingyang Gao
BIBM1
2017 An efficient CNN-based classification on G-protein Coupled Receptors using TF-IDF and N-gram
abstract
Protein sequence classification is increasingly crucial in the current “biological information sciences” epoch, where researchers hammer at functional genomics and proteomics technologies for predicting the function of large-scale new proteins. This has sparked interest in the methods which do not rely on traditional sequence alignment, but prefer machine learning approaches. In this paper, we present a Convolutional Neural Network (CNN) based method to perform the classification on the different levels of G-protein Coupled Receptors (GPCRs). The method is implemented in conjunction with an improved feature extraction method and TF-IDF feature weighting strategy. Experimental results indicate that the proposed method makes significant improvements over previous methods, which attains an accuracy of up to 98.34%, 98.13% and 96.47% in the classification of family level, subfamily level I and II, respectively. In comparison to the other well-known classification methods for GPCRs, the classification error rate of the proposed method is reduced by of at least 55.14% (family level), 72.86% (level I) and 52.63% (Level II).
Cheng Ling, Jingyang Gao
ISCC2
2016 A high-precision shallow Convolutional Neural Network based strategy for the detection of Genomic Deletions
abstract
Genomic Deletion holds the largest proportion of the structural variation (SV). There are many methods for detection of SVs using next-generation data, such as Pindel, SVseq2, BreakDancer, DELLY and so on. However, each method has advantages on only some kind of SVs. For deletions, existing tools usually produce variations with accurate results under 0.5. In this paper, we present CNNdel, a tool based on shallow convolutional neural network to detect genomic deletions with real data from the 1000 Genomes Project. The experimental results show that the accuracy and sensitivity is both improved compared with other existing methods.
Jing Wang 0043, Cheng Ling, Jingyang Gao
BIBM2
2016 MrBayes 3.2.6 on Tianhe-1A: A High Performance and Distributed Implementation of Phylogenetic Analysis
abstract
Phylogenetic analysis has achieved extraordinary results in domains like species delimitation and evolutionary biology. An essential element behind this success has been the introduction of high performance computing techniques in the step of estimating the phylogenetic likelihoods. This paper describes the design and implementation of a distributed and CPU-GPU based heterogeneous computing system on parallelizing the analysis. The parallelization has been implemented in the state-of-the-art version of MrBayes, a widespread phylogeny reconstruction program. We benchmarked the method and another two GPU-based methods by using 8 distributed computing nodes on Tianhe-1A. The experimental results indicate that the proposed method outstrips BEAGLE and the nMC3 method by speedup factors of up to 1.98× and 1.68×, respectively. In comparison to the serially implemented MrBayes, a peak speedup of 188× is finally achieved by using 8 Tesla M 2050 GPUs. The proposed method is publicly available to facilitate further research on phylogenetic analysis.
Cheng Ling, Arong Luo, Jingyang Gao
ICPADS1
2016 MrBayes tgMC3++: A High Performance and Resource-Efficient GPU-Oriented Phylogenetic Analysis Method
abstract
MrBayes is a widespread phylogenetic inference tool harnessing empirical evolutionary models and Bayesian statistics. However, the computational cost on the likelihood estimation is very expensive, resulting in undesirably long execution time. Although a number of multi-threaded optimizations have been proposed to speed up MrBayes, there are bottlenecks that severely limit the GPU thread-level parallelism of likelihood estimations. This study proposes a high performance and resource-efficient method for GPU-oriented parallelization of likelihood estimations. Instead of having to rely on empirical programming, the proposed novel decomposition storage model implements high performance data transfers implicitly. In terms of performance improvement, a speedup factor of up to 178 can be achieved on the analysis of simulated datasets by four Tesla K40 cards. In comparison to the other publicly available GPU-oriented MrBayes, the tgMC3++ method (proposed herein) outperforms the tgMC3(v1.0), nMC3(v2.1.1) and oMC3(v1.00) methods by speedup factors of up to 1.6, 1.9 and 2.9, respectively. Moreover, tgMC3++ supports more evolutionary models and gamma categories, which previous GPU-oriented methods fail to take into analysis.
Cheng Ling, Tsuyoshi Hamada, Jingyang Gao, Guoguang Zhao, Donghong Sun, Weifeng Shi
IEEE ACM Trans. Comput. Biol. Bioinform.1
2015 Optimizing the Bayesian Inference of Phylogeny on Graphic Processors
abstract
Searching for the evolutionary relationships between groups of organism has become a routine procedure in molecular biology. MrBayes is a popular model based phylogenetic inference tool using Bayesian statistics. Unfortunately, the computational cost is very high, resulting in undesirably long execution time. In this paper, we present what we believe the fastest solution of the MrBayes MC3 algorithm running on off-the-shelf graphic processors. The performance benefits are offered by the multi-granularity parallelism model, coarse-grained GPU kernel system, efficient thread arrangement strategy and GPU code level optimizations. MrBayes goMC3 (proposed herein) provides a significant performance improvement over the sequential MrBayes MC3 by a speedup of up to 48× when using single Tesla C2075 GPU card, whereas a speedup factor of 77× can be achieved when using dual GPUs. In comparison to the state-of-the-art version of other publicly available GPU implementations of MrBayes MC3, the cumulative optimizations adopted in goMC3 resulted in a speedup of up 2.5× over oMC3 (v1.0), 1.75× over tgMC3 (v1.0) and 1.46× over nMC3(v2.1.1) for realistic empirical biological datasets. Besides, experimental results indicated that goMC3 outstrips these GPU implementations on the analysis of simulated datasets composed of ultra-large-scale sequences. As a consequence, the reported performance improvement of goMC3 is significant and appears to scale well with increasing dataset sizes.
Cheng Ling, Chunbao Zhou, Arong Luo, Guoguang Zhao, Tsuyoshi Hamada, Xiaoyan Zhu 0001
CCGRID1
2015 SparkSW: Scalable Distributed Computing System for Large-Scale Biological Sequence Alignment
abstract
The Smith-Waterman (SW) algorithm is universally used for a database search owing to its high sensitively. The widespread impact of the algorithm is reflected in over 8000 citations that the algorithm has received in the past decades. However, the algorithm is prohibitively high in terms of time and space complexity, and so poses significant computational challenges. Apache Spark is an increasingly popular fast big data analytics engine, which has been highly successful in implementing large-scale data-intensive applications on commercial hardware. This paper presents the first ever reported system that implements the SW algorithm on Apache Spark based distributed computing framework, with a couple of off-the-shelf workstations, which is named as SparkSW. The scalability and load-balancing efficiency of the system are investigated by realistic ultra-large database from the state-of-the-art UniRef100. The experimental results indicate that 1) SparkSW is load-balancing for parallel adaptive on workloads and scales extremely well with the increases of computing resource, 2) SparkSW provides a fast and universal option high sensitively biological sequence alignments. The success of SparkSW also reveals that Apache Spark framework provides an efficient solution to facilitate coping with ever increasing sizes of biological sequence databases, especially generated by second-generation sequencing technologies.
Guoguang Zhao, Cheng Ling, Donghong Sun
CCGRID2
2010 Highly efficient mapping of the Smith-Waterman algorithm on CUDA-compatible GPUs
abstract
This paper describes a multi-threaded parallel design and implementation of the Smith-Waterman (SW) algorithm on graphic processing units (GPUs) with NVIDIA corporation's Compute Unified Device Architecture (CUDA). Central to this is a divide and conquer approach which divides the computation of a whole pairwise sequence alignment matrix into multiple sub-matrices (or parallelograms) each running efficiently on the available hardware resources of the GPU in hand, with temporary intermediate data stored in global memory. Moreover, we use thread warps and padding techniques in order to decrease the cost of thread synchronization, as well as loop unrolling in order to reduce the cost of conditional branches. While intermediate data is stored in global memory for large queries, the most inner loop in our implementation will only access shared memory and registers. As a result of these optimizations, our implementation of the SW algorithm achieves a throughput ranging between 9.09 GCUPS (Giga Cell Update per Second) and 12.71 GCUPS on a single-GPU version, and a throughput between 29.46 GCUPS and 43.05 GCUPS on a quad-GPU platform. Compared with the best GPU implementation of the SW algorithm reported to date, our implementation achieves up to 46 % improvement in speed. The source code of our implementation is available in the public domain for Bioinformaticians to benefit from its performance.
Keisuke Dohi, Khaled Benkrid, Cheng Ling, Tsuyoshi Hamada, Yuichiro Shibata
ASAP3