EDBT 2026 Demo / reviewers in the wild / expert
Li Yang 0012
dblp:09/3925-12
· DBLP profile ↗
18ranked-venue papers
3as first author
12since 2021 · last 2026
0000-0002-8929-7554ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Security and privacy · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Interactive Hybrid Rice Breeding with Parametric Dual ProjectionabstractHybrid rice breeding crossbreeds different rice lines and cultivates the resulting hybrids in fields to select those with desirable agronomic traits, such as higher yields. Recently, genomic selection has emerged as an efficient way for hybrid rice breeding. It predicts the traits of hybrids based on their genes, which helps exclude many undesired hybrids, largely reducing the workload of field cultivation. However, due to the limited accuracy of genomic prediction models, breeders still need to combine their experience with the models to identify regulatory genes that control traits and select hybrids, which remains a time-consuming process. To ease this process, in this paper, we proposed a visual analysis method to facilitate interactive hybrid rice breeding. Regulatory gene identification and hybrid selection naturally ensemble a dual-analysis task. Therefore, we developed a parametric dual projection method with theoretical guarantees to facilitate interactive dual analysis. Based on this dual projection method, we further developed a gene visualization and a hybrid visualization to verify the identified regulatory genes and hybrids. The effectiveness of our method is demonstrated through the quantitative evaluation of the parametric dual projection method, identified regulatory genes and desired hybrids in the case study, and positive feedback from breeders. Changjian Chen, Fei Lyu 0007, Zhuo Tang, Li Yang 0012, Kenli Li 0001 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2025 | Enhancing Small-Scale Dataset Expansion with Triplet-Connection-based Sample Re-WeightingabstractThe performance of computer vision models in certain real-world applications, such as medical diagnosis, is often limited by the scarcity of available images. Expanding datasets using pre-trained generative models is an effective solution. However, due to the uncontrollable generation process and the ambiguity of natural language, noisy images may be generated. Re-weighting is an effective way to address this issue by assigning low weights to such noisy images. We first theoretically analyze three types of supervision for the generated images. Based on the theoretical analysis, we develop TriReWeight, a triplet-connection-based sample re-weighting method to enhance generative data augmentation. Theoretically, TriReWeight can be integrated with any generative data augmentation methods and never downgrade their performance. Moreover, its generalization approaches the optimal in the order O(√d ln (n)/n). Our experiments validate the correctness of the theoretical analysis and demonstrate that our method outperforms the existing SOTA methods by 7.9% on average over six natural image datasets and by 3.4% on average over three medical datasets. We also experimentally validate that our method can enhance the performance of different generative data augmentation methods. Ting Xiang, Changjian Chen, Zhuo Tang, Fei Lyu 0007, Li Yang 0012, Jiapeng Zhang 0001, Kenli Li 0001 |
ACM Multimedia | 6 |
| 2025 | Boosting multi-document summarization with hierarchical graph convolutional networks
Li Yang 0012, Wenming Luo, Zhuo Tang |
Neurocomputing | 2 |
| 2023 | A Real-Time Partition Generation Mechanism for Data Skew Mitigation in Spark Computing Environment
Li Yang 0012, Zhechang Hu, Zhuo Tang |
J. Grid Comput. | 1 |
| 2023 | Parallel incremental association rule mining framework for public opinion analysis
Li Yang 0012, Sheng You, Zhuo Tang |
Inf. Sci. | 2 |
| 2023 | A Network Load Perception Based Task Scheduler for Parallel Distributed Data Processing SystemsabstractIn parallel distributed data processing frameworks like Spark and Flink, task scheduling has a great impact on cluster performance. Though task Scheduling has proven to be an NP-complete problem, a large number of researchers have proposed many heuristic rules to obtain approximate optimal solutions. But most of them ignore the fact that the resource requirements of tasks are dynamically changing during its runtime. Considering the overall task entire lives, the CPU utilization is often lower during the data transfer. Especially for most distributed data processing platforms, data transmission is time-consuming, which usually resulting in low overall CPU utilization. Similarly, network throughput during task calculations is also low in some cases. In this article, we propose a network load variation perception based heuristic task scheduling algorithm, and based on this implement a dual-phase pipeline task scheduler (D2PTS) from the perspective of dynamic resource requirements that aims at maximizing cluster resource utilization, as a supplement to existing data-parallel frameworks. D2PTS divides the states of task into two phases: network-intensive and network-free. To improve the overall resource utilities, this article proposes different algorithms to evaluate the execution time of network sensitive and network free phases respectively. When an executing task is in the network-free phase, D2PTS can additionally schedule a new network-intensive task at the right time. Under this scheduling policy, the two tasks sharing the same CPU core can be executed as a coarse-grained pipeline. This execution method can start tasks earlier and improve resource utilization. Finally, we have implemented our model prototype on Spark 2.4.3 and conducted a number of experiments to evaluate the performance of our model. Experimental results show that D2PTS can not only minimize application makespan, but also improve resource utilization. Zhuo Tang, Zhanfei Xiao, Li Yang 0012, Kailin He, Kenli Li 0001 |
IEEE Trans. Cloud Comput. | 3 |
| 2022 | AEML: An Acceleration Engine for Multi-GPU Load-Balancing in Distributed Heterogeneous EnvironmentabstractFor the rapid growth computation requirements in big data and artificial intelligence area, CPU-GPU heterogeneous clusters can provide more powerful computing capacity compared to CPU clusters. The number of GPUs on single computing node is scalable, which greatly improves the computing capacity of the cluster under the condition of limited cluster size. However, there is a lack of the effective load-balancing scheduling model in multi-GPU hardware environment. This paper proposes AEML, an acceleration engine for multi-GPU load-balancing in distributed heterogeneous environment. AEML can effectively integrate GPUs into distributed processing framework and achieve great load-balance among multiple heterogeneous GPUs. We propose a heterogeneous task execution model based on multiple GPUs and multiple streams (MGMS), which can effectively balance the workload of multiple GPUs. MGMS model utilizes four core techniques: a fine-grained task mapping mechanism, a device resource unified management scheme, a novel resource-aware GPU task scheduling strategy and a feedback-based streams adjustment scheme. The implementation of AEML system is based on Spark 2.4.1 and NVIDIA CUDA 10.0. We comprehensively evaluate the performance of AEML with multiple typical benchmarks. Experimental results show that AEML can fully exploit the computing power of GPUs and achieve great load-balance among multiple heterogeneous GPUs. Zhuo Tang, Lifan Du, Li Yang 0012, Kenli Li 0001 |
IEEE Trans. Computers | 4 |
| 2022 | IncGraph: An Improved Distributed Incremental Graph Computing Model and Framework Based on Spark GraphXabstractThe excavated information will become obsolete when the data changes in dynamic graphs. To compute the up-to-date results, the graph algorithm has to re-compute the entire data from scratch, which will consume huge computation time and resources. To reduce the cost of such calculations, this paper proposes a model called IncGraph to support incremental iterative computation over dynamic graphs. Different from the way of traditional iteration, IncGraph executes the graph algorithm through reusing the results of the previous graph and performs computation on the part of the graph that has changed. IncGraph has two critical components: (1) an incremental iterative computation model that consists of two steps: an incremental step to calculate the results on the changed vertices of the graph, and a merge step to calculate the results on the entire graph by using the results of the previous graph and the incremental step; and (2) an incremental update method to accelerate the iterative process within the iterative graph algorithm. We implement IncGraph model on GraphX and evaluate its performance by using several representative iterative graph algorithms: PageRank, Connected components, and Single Source Shortest Path. The results show that compared with the traditional iteration, when adding the 100k of vertices in different size data sets, the performance optimization ratio of IncGraph is 31.79 percent averagely, and 50.2 percent maximum; and when the percentage of added vertices varied from 0.01 to 10 percent in different data sets, the performance optimization ratio of IncGraph varied from 19.9 to 66.1 percent. Moreover, the result errors of IncGraph is small and can be neglected. Zhuo Tang, Mengsi He, Zhongming Fu, Li Yang 0012 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2021 | FSLM: An Intelligent Few-Shot Learning Model Based on Siamese Networks for IoT TechnologyabstractAs an important application of the Internet of Things (IoT) devices, sentiment analysis has been paid more attention with the rapid development of artificial intelligence. As a widely used method in artificial intelligence applications, traditional deep learning methods need massive data for training. However, due to the limitations of hardware, IoT devices have deficiencies in processing big data. In the case of insufficient sample size, how to carry out a machine learning method for IoT devices has become a common concern of the industry. In order to perform sentiment analysis on text with few data samples from the IoT devices, we propose FSLM, which is an intelligent few-shot learning model based on Siamese networks. The FSLM model consists of two self-attention models with the same parameters, which are divided into two parts. First, for two input texts, a self-attention model is used to extract sentiment features, and then the Mahalanobis distance is adopted to measure the similarity between two feature vectors to determine whether they belong to the same category. The FSLM is tested on the Amazon Review Sentiment Classification (ARSC) data set. The extensive experimental results on this data set demonstrate that the FSLM model has better accuracy and robustness for text sentiment analysis than other main existing models with a small number of samples. Li Yang 0012, Ying Li 0033, Jin Wang 0001, Naixue Xiong |
IEEE Internet Things J. | 1 |
| 2021 | A Polishing Robot Force Control System Based on Time Series Data in Industrial Internet of ThingsabstractInstalling a six-dimensional force/torque sensor on an industrial arm for force feedback is a common robotic force control strategy. However, because of the high price of force/torque sensors and the closedness of an industrial robot control system, this method is not convenient for industrial mass production applications. Various types of data generated by industrial robots during the polishing process can be saved, transmitted, and applied, benefiting from the growth of the industrial internet of things (IIoT). Therefore, we propose a constant force control system that combines an industrial robot control system and industrial robot offline programming software for a polishing robot based on IIoT time series data. The system mainly consists of four parts, which can achieve constant force polishing of industrial robots in mass production. (1) Data collection module. Install a six-dimensional force/torque sensor at a manipulator and collect the robot data (current series data, etc.) and sensor data (force/torque series data). (2) Data analysis module. Establish a relationship model based on variant long short-term memory which we propose between current time series data of the polishing manipulator and data of the force sensor. (3) Data prediction module. A large number of sensorless polishing robots of the same type can utilize that model to predict force time series. (4) Trajectory optimization module. The polishing trajectories can be adjusted according to the prediction sequences. The experiments verified that the relational model we proposed has an accurate prediction, small error, and a manipulator taking advantage of this method has a better polishing effect. Chen Zhang 0027, Zhuo Tang, Kenli Li 0001, Jianzhong Yang, Li Yang 0012 |
ACM Trans. Internet Techn. | 5 |
| 2021 | An Incremental Iterative Acceleration Architecture in Distributed Heterogeneous Environments With GPUs for Deep LearningabstractThe parallel computing capabilities of GPUs have a significant impact on computationally intensive iterative tasks. Offloading part or all of the deep learning tasks from the CPU to the GPU for execution is mainstream. However, a large number of redundant iterative calculations exist in the iterative process of computing tasks. Therefore, we propose a GPU-based distributed incremental iterative computing architecture that can make full use of distributed parallel computing and GPU memory structure. The architecture supports deep learning and other computationally intensive iterative applications by optimizing data placement and reducing redundant iterative calculations. To support block-based data partitioning and coalesced memory access on GPUs, we propose GDataSet, an abstract data set. The GPU incremental iteration manager called GTracker is designed to be responsible for GDataSet cache management on the GPU. In order to solve the limitation of on-chip memory size, we propose a variable sliding window mechanism. It improves the hit rate of cache access and the speed of data access by realizing the best block arrangement between on-chip memory and off-chip memory. Besides, a communication channel based on an incremental iterative model is designed to support data transmission and task communication in cluster computing. Finally, we implement the proposed architecture based on Spark 2.4.1 and CUDA 10.0. Comparative experiments with widely used computationally intensive iterative applications (K-means, LSTM, etc.) show that the incremental iterative acceleration architecture can significantly improve the efficiency of iterative computing. Zhuo Tang, Lifan Du, Li Yang 0012 |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2021 | Using Conditional Random Fields to Optimize a Self-Adaptive Bell-LaPadula Model in Control SystemsabstractOnce defined, the access control policies and regulations would never be changed in a running and state transition process. However, it will give attackers the possibility of discovering vulnerabilities in the system, and the control systems lack the ability of dynamic perception of security state and risk, causing the systems to be exposed to risks. In this article, a dynamic Bell-LaPadula (BLP) model is proposed. The conditional random field (CRF) is introduced into the BLP model to optimize the rules. First, the model formalizes the security attributes, states of system, transition rules, and constraint models on the basis of the state transition of CRFs. After the historical system access logs are processed as the original dataset, a feature selection method is proposed to extract the requests and current states as feature vectors. Second, this article presents a rules training algorithm based on L-BFGS to implement the study and training of datasets, and then marks the logs in the test set through Viterbi algorithm automatically. On the base of these, a rule generation algorithm is proposed to dynamically adjust the access control rules based on the current security status and events of the system. Third, the security of CRFs-BLP is proved by theoretical analysis. Finally, the validity and accuracy of the model are verified by estimating the value of the precision, recall, and F1-score. As the system threats are shown to be decreased obviously from these experiments, this dynamic model can decrease the vulnerabilities and risk effectively. Li Yang 0012, Jin Wang 0001, Zhuo Tang, Naixue Xiong |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2020 | Dynamic memory-aware scheduling in spark computing environment
Zhuo Tang, Ailing Zeng, Li Yang 0012, Kenli Li 0001 |
J. Parallel Distributed Comput. | 4 |
| 2020 | ImRP: A Predictive Partition Method for Data Skew Alleviation in Spark Streaming Environment
Zhongming Fu, Zhuo Tang, Li Yang 0012, Kenli Li 0001, Keqin Li 0001 |
Parallel Comput. | 3 |
| 2020 | Word-Character Graph Convolution Network for Chinese Named Entity RecognitionabstractRecent researches try to integrate word information into the character-based Chinese NER by modifying the structure of the standard BiLSTM-CRF model. They follow the paradigm of explicitly modeling forward and backward sequences, adopting an LSTM variant that takes both characters and words as input for each direction. Though enriching the representations, these models cannot fully exploit the interaction between future and past contexts. In this paper, we propose a novel word-character graph convolution network (WC-GCN) which uses a cross GCN block to simultaneously process the word-character directed acyclic graphs (DAGs) of two directions. To improve the capture of long-distance dependency, a global attention GCN block is introduced to learn node representations conditioned on a global context. In both blocks, unlike previous works where each word is attached to its associated character or taken as a shortcut between LSTM cells, words and characters are treated equally as nodes in the graph and have their instance-specific representations. Experiments on four widely used datasets show that our proposed model can work standalone or with the standard BiLSTM. Both forms can outperform previous LSTM-based models without training on extra corpora while only an external lexicon and its corresponding pretrained character and word embeddings are needed. Zhuo Tang, Boyan Wan, Li Yang 0012 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2020 | An Optimal Locality-Aware Task Scheduling Algorithm Based on Bipartite Graph Modelling for Spark ApplicationsabstractIn the distributed computing framework of Spark, cross-node/rack data transfer produced by map tasks and reduce tasks are common problems resulting in performance degradation, such as prolonging of entire execution time and network congestion. To address these problems, this article utilizes the bipartite graph modelling to propose an optimal locality-aware task scheduling algorithm. By considering global optimality, the algorithm can generate the optimal scheduling solution for both the map tasks and the reduce tasks for data locality. Because of the different communication modes, this article uses a unified graph to model the map task scheduling and the reduce task scheduling respectively. Then, by calculating the communication cost matrix of tasks, we formulate an optimal task scheduling scheme to minimize overall communication cost and transform the problem as the well-known graph problem: minimum weighted bipartite matching (MWBM), which can be resolved by Kuhn-Munkres algorithm. In addition, this article proposes a locality-aware executor allocation strategy to improve the data locality further. We implement our algorithm and strategy in Spark-2.4.1 and evaluate its performance using several representative micro-benchmarks, macro-benchmarks, and HiBench benchmark suite. The experimental results verify that by reducing the network traffic and access latency, the proposed algorithm can improve the job performance substantially compared to some other task scheduling algorithms. Zhongming Fu, Zhuo Tang, Li Yang 0012, Chubo Liu |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2018 | A Self-Adaptive Bell-LaPadula Model Based on Model Training With Historical Access LogsabstractIn currently popular access control models, the security policies and regulations never change in the running system process once they are identified, which makes it possible for attackers to find the vulnerabilities in a system, resulting in the lack of ability to perceive the system security status and risks in a dynamic manner and exposing the system to such risks. By introducing the maximum entropy (MaxENT) models into the rule optimization for the Bell-LaPadula (BLP) model, this paper proposes an improved BLP model with the self-learning function: MaxENT-BLP. This model first formalizes the security properties, system states, transformational rules, and a constraint model based on the states transition of the MaxENT. After handling the historical system access logs as the original data sets, this model extracts the user requests, current states, and decisions to act as the feature vectors. Second, we use k -fold cross validation to divide all vectors into a training set and a testing set. In this paper, the model training process is based on the Broyden-Fletcher-Goldfarb-Shanno algorithm. And this model contains a strategy update algorithm to adjust the access control rules dynamically according to the access and decision records in a system. Third, we prove that MaxENT-BLP is secure through theoretical analysis. By estimating the precision, recall, and F1-score, the experiments show the availability and accuracy of this model. Finally, this paper provides the process of model training based on deep learning and discussions regarding adversarial samples from the malware classifiers. We demonstrate that MaxENT-BLP is an appropriate choice and has the ability to help running information systems to avoid more risks and losses. Zhuo Tang, Xiaofei Ding, Ying Zhong 0008, Li Yang 0012, Keqin Li 0001 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2017 | A parallel k-means clustering algorithm based on redundance elimination and extreme points optimization employing MapReduceabstractSummary When facing massive statistical data, the k‐means algorithm is very difficult to satisfy the need of data processing as it lacks an effective parallel mechanism. This paper proposes an improved k‐means algorithm (IMR‐KCA) to conduct clustering analysis based on medical data employing MapReduce computing framework. Through analyzing the defects of vast redundancy in the traditional k‐means algorithms, a selection model is firstly proposed to simplify the computations with multiple clustering centers. Based on several proposed theorems, we prove the correctness of this selection model. Second, this paper provides a method to calculate the distances from extreme points to central points, and the original Euclidean distance is replaced with Manhattan distance. For this simplification, a group of theorems are proposed to prove the correctness. Next, we provide a group of implementation algorithms to complete the parallelism of the clustering computation employing the MapReduce framework. Finally, the experimental results illustrate that IMR‐KCA is more reliable and efficient than the direct parallelization of the traditional clustering algorithms based on MapReduce. Copyright © 2017 John Wiley & Sons, Ltd. Zhuo Tang, Kunkun Liu, Jinbo Xiao, Li Yang 0012 |
Concurr. Comput. Pract. Exp. | 4 |