Yang Che

dblp:269/5589 · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
5since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
YearPublicationVenuePosition
2025 Inter-Well Active Magnetic Ranging with Temporal and Interaction Network
abstract
Inter-well distance measurement is crucial for ensuring safety in oil and gas drilling operations, particularly during blowout emergencies where rescue wells are drilled to mitigate risks. Traditional methods, including GPS and ultrasonic ranging, are ineffective in downhole environments. While magnetic ranging methods such as Passive Magnetic Ranging and Active Magnetic Ranging (AMR) offer solutions, they suffer from limitations like magnetic interference and the inability to handle complex well structures. In this paper, we propose a novel inter-well distance prediction method that integrates AMR with Temporal and Interaction Network (TINet). Firstly, to construct TINet, we conduct data collection experiments using active magnetic ranging tools, resulting in a calibrated dataset that includes both magnetic and nonmagnetic data. Subsequently, we apply the collected dataset to train TINet, which involves enhancing feature extraction through an improved sample convolution and interaction network, followed by the integration of the extracted features into a long short-term memory unit. This allows TINet to effectively utilize historical information while capturing the temporal and spatial relationships within the data. Moreover, to further improve prediction accuracy, we introduce a custom loss function, adaptive error scaling loss, which balances relative and absolute errors. The experimental results indicate that our method demonstrates significant improvements compared to traditional methods across various evaluation metrics and distance scales.
Zelong Hao, Yang Che
SDM3
2023 High-Level Data Abstraction and Elastic Data Caching for Data-Intensive AI Applications on Cloud-Native Platforms
abstract
Nowdays, it is prevalent to train deep learning models in cloud-native platforms that actively leverage containerization and orchestration technologies for high elasticity, low and flexible operation cost, and many other benefits. However, it also faces new challenges and our work is focusing on those related to I/O throughput for training, including complex data access, lack of matching dynamic I/O requirement, and inefficient I/O resource scheduling across different jobs. We proposeFluid, a cloud-native platform that provides DL training jobs with high-level data abstraction calledFluid Datasetto access training data from heterogeneous sources with elastic data acceleration. In addition, it comes with an on-the-fly cache system autoscaler that can match the online training speed and increase the number of cache replicas adaptively to alleviate I/O bottlenecks. To improve the overall performance of multiple DL jobs, Fluid co-orchestrate the data cache and DL jobs by arranging job scheduling in an appropriate order and can also schedule data cache and DL jobs on the same node to realize cache affinity. Experimental results show significant performance improvement of each individual DL job which uses dynamic computing resources with Fluid. For scheduling multiple DL jobs with same datasets, Fluid achieves around 2x performance speedup when integrated with existing widely-used and cutting-edge scheduling solutions through the appropriate job scheduling order. Besides, the cache affinity scheduling policy also improves job execution performance significantly. Fluid is now an open source project hosted by Cloud Native Computing Foundation (CNCF) with many production adopters.
Rong Gu 0001, Yang Che, Haipeng Dai 0001, Haojun Hou, Li Yi 0003, Yihua Huang 0001, Guihai Chen
IEEE Trans. Parallel Distributed Syst.3
2022 Fluid: Dataset Abstraction and Elastic Acceleration for Cloud-native Deep Learning Training Jobs
abstract
Nowdays, it is prevalent to train deep learning (DL) models in cloud-native platforms that actively leverage containerization and orchestration technologies for high elasticity, low and flexible operation cost, and many other benefits. However, it also faces new challenges and our work is focusing on those related to I/O throughput for training, including complex data access with complicated performance tuning, lack of cache capacity with specialized hardware to match its high and dynamic I/O requirement, and inefficient I/O resource sharing across different training jobs. We propose Fluid, a cloud-native platform that provides DL training jobs with a data abstraction called Fluid Dataset to access training data from heterogeneous sources in a unified manner with transparent and elastic data acceleration powered by auto-tuned cache runtimes. In addition, it comes with an on-the-fly cache system autoscaler that can intelligently scale up and down the cache capacity to match the online training speed of each individual DL job. To improve the overall performance of multiple DL jobs, Fluid can co-orchestrate the data cache and DL jobs by arranging job scheduling in an appropriate order. Our experimental results show significant performance improvement of each individual DL job which uses dynamic computing resources with Fluid. In addition, for scheduling multiple DL jobs with same datasets, Fluid gives around 2x performance speedup when integrated with existing widely-used and cutting-edge scheduling solutions. Fluid is now an open source project hosted by Cloud Native Computing Foundation (CNCF) with adopters in production including Alibaba Cloud, Tencent Cloud, Weibo.com, China Telecom, etc.
Rong Gu 0001, Yang Che, Haojun Hou, Haipeng Dai 0001, Li Yi 0003, Guihai Chen, Yihua Huang 0001
ICDE4
2022 Octopus-DF: Unified DataFrame-based cross-platform data analytic system
Rong Gu 0001, Zhaokang Wang, Yang Che, Yihua Huang 0001
Parallel Comput.5
2022 Liquid: Intelligent Resource Estimation and Network-Efficient Scheduling for Deep Learning Jobs on Distributed GPU Clusters
abstract
Deep learning (DL) is becoming increasingly popular in many domains, including computer vision, speech recognition, self-driving automobiles, etc. GPU can train DL models efficiently but is expensive, which motivates users to share GPU resource to reduce money costs in practice. To ensure efficient sharing among multiple users, it is necessary to develop efficient GPU resource management and scheduling solutions. However, existing ones have several shortcomings. First, they require the users to specify the job resource requirement which is usually quite inaccurate and leads to cluster resource underutilization. Second, when scheduling DL jobs, they rarely take the cluster network characteristics into consideration, resulting in low job execution performance. To overcome the above issues, we propose Liquid, an efficient GPU resource management platform for DL jobs with intelligent resource requirement estimation and scheduling. First, we propose a regression model based method for job resource requirement estimation to avoid users over-allocating computing resources. Second, we propose intelligent cluster network-efficient scheduling methods in both immediate and batch modes based on the above resource requirement estimation techniques. Third, we further propose three system-level optimizations, including pre-scheduling data transmission, fine-grained GPU sharing, and event-driven communication. Experimental results show that our Liquid can accelerate the job execution speed by 18% on average and shorten the average job completion time (JCT) by 21% compared with cutting-edge solutions. Moreover, the proposed optimization methods are effective in various scenarios.
Rong Gu 0001, Yuquan Chen, Haipeng Dai 0001, Guihai Chen, Yang Che, Yihua Huang 0001
IEEE Trans. Parallel Distributed Syst.7