Nana Du

dblp:197/1064 · DBLP profile ↗
← Back
7ranked-venue papers
4as first author
3since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2Computer networks · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2025 Scheduling DAG-structured workloads based on whale optimization algorithm
abstract
Abstract Many computing workloads in big data and machine learning applications are structured as directed acyclic graphs (DAG) and deployed on PC clusters for parallel execution using multiple physical or virtual machines. The scheduling of such workloads is critical to the application performance such as execution time and a plethora of techniques have been developed, taking into account various aspects such as data locality, network bandwidth, and server capability. We formulate DAG-structured workload scheduling as a nonlinear integer programming (NIP) problem and prove it to be NP-complete. Our empirical study reveals a positive correlation between scheduling plan distance (SPD) and finish time gap (FTG), and based on this finding, we propose a running time gap strategy (RTGS) to tackle this scheduling problem in multiprocessor environments. RTGS follows the main optimization strategy in the family of whale optimization algorithm (WOA). We derive a new function and use a greedy algorithm to generate an effective scheduling plan in RTGS. Extensive experiments with real production traces from Alibaba on simulation environments and realistic Hadoop environments show that our approach significantly improves the stability of WOA when applied to the scheduling problem of DAG-structured workloads, and also reduces the workload completion time by up to $$93\%$$ 93 % in comparison with seven state-of-the-art baseline algorithms.
Nana Du, Yudong Ji, Chase Qishi Wu, Aiqin Hou, Weike Nie
J. Supercomput.1
2024 UJPS: Urgent Job Priority Scheduling in Hadoop YARN
abstract
The rapidly increasing demand for big data processing has necessitated the development of advanced scheduling policies that can effectively accommodate urgent job requirements. This paper presents the Urgent Job Priority Scheduler (UJPS) for Hadoop YARN, aimed at handling urgent jobs efficiently in big data processing. UJPS uses an Aging model to cut waiting times and prevent job starvation, a Dynamic Priority model for urgency-based prioritization, and a Container Load model to boost data locality and efficiency. Tested on Hadoop with benchmark tasks, UJPS outperforms five advanced schedulers, lowering waiting times by up to 81.42% and reducing job runtime by 32.90%. It prioritizes urgent tasks while ensuring overall efficiency, offering benefits to organizations using Hadoop YARN for timely job execution.
Nana Du, Aiqin Hou, Chase Qishi Wu, Weike Nie
HPCC1
2023 Dynamic Priority Job Scheduling on a Hadoop YARN Platform
abstract
In Hadoop’s big data processing systems, YARN is responsible for resource management and job scheduling. The built-in job scheduling algorithms in YARN are simple to execute, but have some limitations such as job starvation, excessive server load, and load imbalance. In this paper, we propose a new Hybrid Dynamic Priority job Scheduling algorithm (HDPS) to address these limitations. HDPS dynamically adjusts the priority of a job as its waiting time increases to prevent job starvation. It also features a task assignment strategy designed specifically to address data locality by considering the available resources of servers and the distribution of data blocks stored on servers to reduce data transfer time and improve job execution efficiency. We implement and integrate HDPS into YARN and conduct experiments in a real Hadoop system using built-in benchmark test cases of Hadoop. Experimental results show that HDPS exhibits comprehensive superior performance over existing algorithms in terms of execution efficiency and load balance.
Nana Du, Yudong Ji, Aiqin Hou, Chase Qishi Wu, Weike Nie
ICPADS1
2020 Recommendation of Academic Papers based on Heterogeneous Information Networks
abstract
The rapid advance in science and technology is made possible by research conduct and breakthroughs in a wide range of fields, which have resulted in a large number of academic papers. Searching through the enormous literature to find relevant information of one's research interest has become an increasingly important yet challenging problem for many researchers. Most existing methods for academic paper recommendation are based on the analysis of paper contents and only meet with limited success. We propose a novel method based on heterogeneous information networks for academic paper recommendation, referred to as HNPR. This method considers the citation relationship between papers, the collaboration relationship between authors, and the research area information of papers to construct two types of heterogeneous information networks. In such networks, a random walk-based strategy is used to simulate natural sentences for the discovery of relevance between two papers according to a mature natural language processing model. Extensive experimental results using real data in public digital libraries show that HNPR significantly improves the accuracy of academic paper recommendation in comparison with traditional content-based recommendation methods.
Nana Du, Jun Guo 0020, Chase Qishi Wu, Aiqin Hou, Zimin Zhao, Daguang Gan
AICCSA1
2019 Towards robust controller placement in software-defined networks against links failure
Nana Du, Ruifang Zhang, Chaobo Yan
IM2
2018 Classification for Social Media Short Text Based on Word Distributed Representation
Dexin Zhao, Zhi Chang, Nana Du, Shutao Guo
WISA3
2017 Keyword Extraction for Social Media Short Text
abstract
With the booming development of social media in recent years, researchers have begun to pay more attention to extracting personal profiles from information. Keyword extraction plays an important role in extracting personal profiles. However, most of the previous studies are only valid for ordinary text, but not ideal for social media short text. In this paper, we propose an improved method for keyword extraction based on Word2vec and Textrank to solve the unique problem of social media short text. Our approach uses the Word2vec to capture the semantic features between words in selected text, and meanwhile naturally fuses the word frequency, semantic relation and directional relation into Textrank to extract keywords. We conduct the experiments on the three datasets. The experimental results show the superior performance of our method in keyword extraction.
Dexin Zhao, Nana Du, Zhi Chang
WISA2