Zijie Zhong

dblp:343/8920 · DBLP profile ↗
← Back
8ranked-venue papers
4as first author
8since 2021 · last 2026
0009-0003-8681-6077ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Historical reliability-based dual contrastive hashing for robust cross-modal retrieval with noisy labels
Haixiao Huang, Likang Peng, Chao Su 0003, Zijie Zhong, Da Rao, Dezhong Peng, Xu Wang 0028
Neurocomputing5
2026 HRADP: A Heat-Recirculation-Aware Data Placement Approach for Energy-Efficient Cloud Data Centers
abstract
With the proliferation of cloud computing, the exponential growth of data amount requires continuous expansion of storage capacity to meet the storage demand, consequently resulting in higher energy consumption in data centers. However, many of the conventional data placement strategies strive to save energy by optimizing the distribution of data requests, while overlooking the impact of heat recirculation among data nodes. To bridge this gap, we propose a heat-recirculation-aware data placement approach, termed HRADP, designed to optimize data placement in data centers while reducing energy consumption. First, based on the extent of heat recirculation, HRADP places data to the upper limit of the disk allowance for each data node. Second, during energy allocation, HRADP further eliminates disks with high current workloads to prevent localized hotspots. Finally, during request scheduling, HRADP selects dormant data nodes according to the request size, minimizing the startup energy consumption caused by data nodes waking up with small requests. We implement the HRADP approach on a data center simulation platform, CloudSim, and its performance with state-of-the-art, including the Energy-efficient and Thermal-aware Data Placement (ETDP) algorithm, Thermal-aware file assignment technique (TIGER), Storage and rack-sensitive replica placement algorithm (SRS), and Hadoop Distributed File System (HDFS). The experimental results reveal that HRADP revamps the cooling supply temperature, total energy, and data throughput by averages of 0.08%-0.6%, 8.22%-53.03%, and 9.58%-64.59%, respectively.
Jie Li 0067, Yuhui Deng 0001, Zijie Zhong, Geyong Min
IEEE Trans. Computers3
2025 Mix-of-Granularity: Optimize the Chunking Granularity for Retrieval-Augmented Generation
abstract
Integrating information from various reference databases is a major challenge for Retrieval-Augmented Generation (RAG) systems because each knowledge source adopts a unique data structure and follows different conventions. Retrieving from multiple knowledge sources with one fixed strategy usually leads to under-exploitation of information. To mitigate this drawback, inspired by Mix-of-Expert, we introduce Mix-of-Granularity (MoG), a method that dynamically determines the optimal granularity of a knowledge source based on input queries using a router. The router is efficiently trained with a newly proposed loss function employing soft labels. We further extend MoG to MoG-Graph (MoGG), where reference documents are pre-processed as graphs, enabling the retrieval of distantly situated snippets. Experiments demonstrate that MoG and MoGG effectively predict optimal granularity levels, significantly enhancing the performance of the RAG system in downstream tasks. The code of both MoG and MoGG will be made public.
Zijie Zhong, Xiaoya Cui, Xiaofan Zhang 0012, Zengchang Qin
COLING1
2025 SyntheT2C: Generating Synthetic Data for Fine-Tuning Large Language Models on the Text2Cypher Task
abstract
Integrating Large Language Models (LLMs) with existing Knowledge Graph (KG) databases presents a promising avenue for enhancing LLMs’ efficacy and mitigating their “hallucinations”. Given that most KGs reside in graph databases accessible solely through specialized query languages (e.g., Cypher), it is critical to connect LLMs with KG databases by automating the translation of natural language into Cypher queries (termed as “Text2Cypher” task). Prior efforts tried to bolster LLMs’ proficiency in Cypher generation through Supervised Fine-Tuning (SFT). However, these explorations are hindered by the lack of annotated datasets of Query-Cypher pairs, resulting from the labor-intensive and domain-specific nature of such annotation. In this study, we propose SyntheT2C, a methodology for constructing a synthetic Query-Cypher pair dataset, comprising two distinct pipelines: (1) LLM-based prompting and (2) template-filling. SyntheT2C is applied to two medical KG databases, culminating in the creation of a synthetic dataset, MedT2C. Comprehensive experiments demonstrate that the MedT2C dataset effectively enhances the performance of backbone LLMs on Text2Cypher task via SFT. Both the SyntheT2C codebase and the MedT2C dataset will be released.
Zijie Zhong, Linqing Zhong, Zhaoze Sun, Qingyun Jin, Zengchang Qin, Xiaofan Zhang 0012
COLING1
2025 CADReN: Contextual Anchor-Driven Relational Network for Controllable Cross-Graphs Node Importance Estimation
Zijie Zhong, Yunhui Zhang, Ziyi Chang, Zengchang Qin
PAKDD (1)1
2025 An Energy-Aware Virtual Machine Scheduling Approach for Cloud Data Centers
abstract
The reduction of energy consumption will be even more urgent in cloud data centers due to the explosive increase of application data. Virtual machine (VM) integration is a relatively standard technology currently applied for computing facilities of data centers. However, excessive VM consolidation can easily lead to local hot spots that lower the energy efficiency and reliability of data centers. In addition, on account of the impact of heat recirculation in data centers, the traditional VM scheduling strategy cannot comprehensively ponder optimizing the holistic data center energy, which encompasses both server energy and cooling energy. To handle these issues, we proposedEAVMS- an Energy-Aware VM Scheduling approach for minimizing the holistic energy consumption of data centers. EAVMS adopts a two-phase approach to gain energy efficiency while guaranteeing QoS. First, EAVMS leverages a Blended Genetic algorithm and Simulated Annealing algorithm (BGSA) to optimize the initial placement of VMs. Second, EAVMS utilizes a dynamic migration algorithm to achieve effective migration by setting a maximum server temperature threshold without violating the service level agreement (SLA) that cuts down energy consumption by moderating the hot spots of servers. We conducted extensive experiments using two real-world traces (i.e., PlanetLab and Google Cluster datasets) to evaluate the effectiveness of EAVMS. The experimental results unveil that our approach is capable of saving 3.23$ \%$–43.07$ \%$in the holistic energy consumption of cloud data centers with only a tiny service performance degradation compared to other state-of-the-art alternatives (e.g., MJPM, GRANITE, TAS, XINT-GA, and Random).
Jie Li 0067, Yuhui Deng 0001, Zijie Zhong, Zhaorui Wu, Shujie Pang, Lin Cui 0001, Geyong Min
IEEE Trans. Sustain. Comput.3
2024 Towards Energy-Efficient and Thermal-Aware Data Placement for Storage Clusters
abstract
The explosion of large-scale data has increased the scale and capacity of storage clusters in data centers, leading to huge power consumption issues. Cloud providers can effectively promote the energy efficiency of data centers by employing energy-aware data placement techniques, which primarily encompass storage cluster's power and cooling power. Traditional data placement approaches do not diminish the overall power consumption of the data center due to the heat recirculation effect between storage nodes. To fill this gap, we build an elaborate thermal-aware data center model. Then we propose two energy-efficient thermal-aware data placement strategies, ETDP-I and ETDP-II, to reduce the overall power consumption of the data center. The principle of our proposed algorithm is to utilize a greedy algorithm to calculate the optimal disk sequence at the minimum total power of the data center and then place the data into the optimal disk sequence. We implement these two strategies in a cloud computing simulation platform based on CloudSim. Experimental results unveil that ETDA-I and ETDP-II outperform MinTin-G and MinTout-G in terms of the supplied temperature of CRAC, storage nodes power, cooling cost, and total power consumption of the data center. In particular, ETDP-I and ETDP-II algorithms can save about 9.46%-38.93% of the overall power consumption compared to MinTout-G and MinTin-G algorithms.
Jie Li 0067, Yuhui Deng 0001, Zhifeng Fan, Zijie Zhong, Geyong Min
IEEE Trans. Sustain. Comput.4
2022 A Heat-Recirculation-Aware Data Placement Strategy towards Data Centers
abstract
The development of cloud computing leads to an exponential growth of data, which requires expanding the storage capacity to meet the storage needs, in exchange the energy consumption of the data center will also increase. Many traditional data placement schemes attempt to achieve energy consumption minimization by optimizing the distribution of data requests, however, ignoring the impact of heat recirculation in data placement. To fill this gap, we propose a heat-recirculation-aware data placement strategy called HRADP to achieve optimized data placement to data centers. Furthermore, our strategy can minimize the overall energy consumption of data centers and improve throughput. Specifically, we consider heat recirculation between data nodes by regulating the energy consumption limit of each data node. We implement this data placement strategy on a data center simulation platform, CloudSim, and compare it with two data placement strategies, TIGER and HDFS. The experimental results unveil that HRADP achieves 2.6x - 25.7x performance improvement in the same energy consumption.
Zijie Zhong, Yuhui Deng 0001, Jie Li 0067
ICPADS1