EDBT 2026 Demo / reviewers in the wild / expert
Ruini Xue
dblp:09/2138
· DBLP profile ↗
21ranked-venue papers
7as first author
11since 2021 · last 2026
0000-0003-1802-5188ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 5 first-author · 3 since 2021Artificial intelligence and machine learning · 6 · 6 since 2021Software engineering, systems software and programming languages · 2Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Computer networks · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GRASP: Fine-grained and Adaptive Sampled Simulation for GPU Performance ModelingabstractGPUs have emerged as foundational platforms for high-performance computing and machine learning. However, rapid architectural evolution challenges existing performance evaluations. Cycle-accurate simulation offers high precision but is prohibitively slow for large-scale design exploration. Sampling-based acceleration techniques struggle with behavior characterization, stable-state detection, and execution time prediction, thus constraining their applicability in complex GPU workloads. Ruini Xue, Lingwei Chao, Zhenxing Huang, Peilin Cai, Jiangying Xue, Tianyu Xiong |
ICS | 1 |
| 2026 | Adaptive probabilistic transformer for medical image segmentation
Tahir Kamal, Wenhong Tian, Jinyu Guo, Muhammad Shafiq 0006, Shuaihong Jiang, Ruini Xue |
Expert Syst. Appl. | 6 |
| 2026 | Accelerating long-context inference of large language models via dynamic attention load balancing
Jie Ou, Jinyu Guo, Shuaihong Jiang, Ruini Xue, Wenhong Tian, Rajkumar Buyya |
Knowl. Based Syst. | 5 |
| 2025 | SimPoint+: More Stable, Accurate and Efficient Program Analysis
Jiangying Xue, Tianyu Xiong, Lingwei Chao, Ruini Xue |
Euro-Par (2) | 4 |
| 2025 | A Neural Network-Based Pipeline Parallel Strategy Solver for Heterogeneous EnvironmentsabstractThe widespread application of large language models(LLMs) has made distributed training increasingly important, especially pipeline parallelism, which is a fundamental technique for ultra-large-scale LLMs. Current research in this field mainly employs combinatorial optimization algorithms such as dynamic programming. However, as the problem size increases, these methods become difficult to solve quickly in large-scale scenarios due to their high search time. Online optimization algorithms that combine neural networks with reinforcement learning require real-time interaction with the cluster environment to obtain feedback, resulting in high resource overhead and low search efficiency. Moreover, current research lacks studies on heterogeneous computing environments, which are frequently used by small research teams. To address these issues, we designed a novel Neural Network-based Pipeline Parallel strategy solver (NN-Piper) for heterogeneous environments. NN-Piper can perceive computational and communication costs, the number of stages to be divided, and the number of micro-batches. In addition, it can directly provide the strategy for allocating specific devices to each pipeline stage. To avoid an online training process that requires interaction with the cluster environment, we propose the Virtual Contrastive Training Algorithm (VCTA) to enable efficient training of NN-Piper without collecting large amounts of real data. After training, NN-Piper can be transferred to many different scenarios without further training or fine-tuning, and it can search for strategies within a few dozen milliseconds. Compared with the state-of-the-art method, NN-Piper can improve the training speed on average by 16-25% in different environments for the transformer-based models. Jie Ou, Jinyu Guo, Yueming Chen, Shuaihong Jiang, Ruini Xue, Wenhong Tian |
IJCNN | 5 |
| 2025 | Low-Rank Decomposition Assisted Quantization and Inference Compensation for Quality Large Language Model InferenceabstractLarge Language Models (LLMs) have demonstrated exceptional performance on natural language processing tasks. However, these models are computationally intensive and require substantial hardware resources for deployment. Quantization has emerged as a popular technique for LLM deployment, reducing memory requirements, but it results in accuracy degradation, particularly when using low-bit quantization. To mitigate this accuracy loss, we introduce Low-Rank Compensation (LoRC), a novel compensation mechanism that aims to recover the performance drop caused by quantization. Additionally, we propose Low-Rank Quantization (LoRQ), which further reduces the quantization-induced loss by adaptively adjusting weights at the element-wise level to help LLMs accommodate quantized computations. LoRC focuses on compensating for accuracy loss during inference, LoRQ integrates low-rank compensation directly into the quantization process, and they do not need end-to-end fine-tuning with LLM. Furthermore, we propose the Rank-α Addition Strategy (RαAS) to combine LoRC into the inference framework, which improves inference accuracy without increasing inference latency. Experimental results show that our method outperforms the state-of-the-art OmniQuant by 1.89% on several common zero-shot datasets under the W4A4 setting of the widely-used LLaMA. Through the joint design of algorithms and systems, our techniques can be easily integrated into the FlexGen inference framework without introducing additional inference latency, thereby maintaining high throughput while improving accuracy. Jie Ou, Jinyu Guo, Shuaihong Jiang, Zhaokun Wang, Yueming Chen, Ruini Xue, Wenhong Tian |
IJCNN | 6 |
| 2024 | LDTR: Transformer-based lane detection with anchor-chain representationabstractDespite recent advances in lane detection methods, scenarios with limited- or no-visual-clue of lanes due to factors such as lighting conditions and occlusion remain challenging and crucial for automated driving. Moreover, current lane representations require complex post-processing and struggle with specific instances. Inspired by the DETR architecture, we propose LDTR, a transformer-based model to address these issues. Lanes are modeled with a novel anchor-chain, regarding a lane as a whole from the beginning, which enables LDTR to handle special lanes inherently. To enhance lane instance perception, LDTR incorporates a novel multi-referenced deformable attention module to distribute attention around the object. Additionally, LDTR incorporates two line IoU algorithms to improve convergence efficiency and employs a Gaussian heatmap auxiliary branch to enhance model representation capability during training. To evaluate lane detection models, we rely on Fréchet distance, parameterized Fl-score, and additional synthetic metrics. Experimental results demonstrate that LDTR achieves state-of-the-art performance on well-known datasets. Zhongyu Yang, Tengfei Xing, Runbo Hu, Pengfei Xu 0013, Ruini Xue |
Comput. Vis. Media | 8 |
| 2024 | A joint learning method with consistency-aware for low-resolution facial expression recognition
Yuanlun Xie, Wenhong Tian, Ruini Xue, Zhiyuan Zha, Bihan Wen |
Expert Syst. Appl. | 4 |
| 2023 | CANet: Curved Guide Line Network with Adaptive Decoder for Lane DetectionabstractLane detection is challenging due to the complicated onroad scenarios and line deformation from different camera perspectives. Lots of solutions were proposed, but can not deal with "corner lanes" well. To address this problem, this paper proposes a new top-down deep learning lane detection approach, CANet. A lane instance is first responded by the heatmap on the U-shaped "curved guide line" at global semantic level, thus the corresponding features of each lane are aggregated at the response point. Then CANet obtains the heatmap response of the entire lane through conditional convolution, and finally decodes the point set to describe lanes via adaptive decoder. The prototype is implemented with Pytorch, and evaluated against 3 well-known datasets extensively. The experimental results show that CANet reaches SOTA in different metrics. Zhongyu Yang, Tengfei Xing, Runbo Hu, Pengfei Xu 0013, Ruini Xue |
ICASSP | 8 |
| 2022 | SPAC: Scalable Pattern Approximate Counting in Graph Mining
Ruini Xue, Shengbo Liu, Wenhong Tian |
ICA3PP | 1 |
| 2021 | Angular Triplet Loss-based Camera Network for ReIDabstractPerson re-identification (ReID) is a challenging cross-camera retrieval task to identify pedestrians. Many complex network structures are proposed recently and many of them concentrate on multi-branch features to achieve high performance. However, they are too heavy-weight to deploy in realworld applications. Additionally, pedestrian images are often captured by different surveillance cameras, so the varied lights, perspectives and resolutions result in inevitable multi-camera domain gaps for ReID. To address these issues, this paper proposes ATCN, a simple but effective angular triplet loss-based camera network, which is able to achieve compelling performance with only global features. In ATCN, a novel angular distance is introduced to learn a more discriminative feature representation in the embedding space. Meanwhile, a lightweight camera network is designed to transfer global features to more discriminative features. ATCN is designed to be simple and flexible so it can be easily deployed in practice. The experiment results on various benchmark datasets show that ATCN outperforms many SOTA approaches. Ruini Xue, Zenglin Xu |
IJCNN | 2 |
| 2018 | LMCC: Lazy Message and Centralized Cache for Asynchronous Graph Computing
Ruini Xue, Zhibin Dong, Wei Su 0005 |
ICA3PP (2) | 1 |
| 2017 | HybridFS - A High Performance and Balanced File System Framework with Multiple Distributed File SystemsabstractIn the big data era, the distributed file system is getting more and more significant due to the characteristics of its scale-out capability, high availability, and high performance. Different distributed file systems may have different design goals. For example, some of them are designed to have good performance for small file operations, such as GlusterFS, while some of them are designed for large file operations, such as Hadoop distributed file system. With the divergence of big data applications, a distributed file system may provide good performance for some applications but fails for some other applications, that is, there has no universal distributed file system that can produce good performance for all applications. In this paper, we propose a hybrid file system framework, HybridFS, which can deliver satisfactory performance for all applications. HybridFS is composed of multiple distributed file systems with the integration of advantages of these distributed file systems. In HybridFS, on top of multiple distributed file systems, we have designed a metadata management server to perform three functions: file placement, partial metadata store, and dynamic file migration. The file placement is performed based on a decision tree. The partial metadata store is performed for files whose size is less than a few hundred Bytes to increase throughput. The dynamic file migration is performed to balance the storage usage of distributed file systems without throttling performance. We have implemented HybridFS in java on eight nodes and choose Ceph, HDFS, and GlusterFS as designated distributed file systems. The experimental results show that, in the best case, HybridFS can have up to 30% performance improvement of read/write operations over a single distributed file system. In addition, if the difference of storage usage among multiple distributed file systems is less than 40%, the performance of HybridFS is guaranteed, that is, no performance degradation. Yongwei Wu 0001, Ruini Xue, Tse-Chuan Hsu, Yeh-Ching Chung |
COMPSAC (1) | 3 |
| 2016 | Replichard: Towards Tradeoff between Consistency and Performance for MetadataabstractMetadata scalability is critical for distributed systems as the storage scale is growing rapidly. Because of the strict consistency requirement of metadata, many existing metadata services utilize a fundamentally unscalable design for the sake of easy management, while others provide improved scalability but lead to unacceptable latency and management complexity. Without delivering scalable performance, metadata will be the bottleneck of the entire system. Based on the observation that real file dependencies are few, and there are usually more idempotent than non-idempotent operations, we propose a practical strategy, Replichard, allowing a tradeoff between metadata consistency and scalable performance. Replichard provides metadata services through a cluster of metadata servers, in which a flexible consistency scheme is adopted: strict consistency for non-idempotent operations with dynamic write-lock sharding, and relaxed consistency with accuracy estimations of return values where consistency for idempotent requests is relaxed to achieve high throughput. Write-locks are dynamically created at subtree-level and designated to independent metadata servers in an application-oriented manner. A subtree metadata update that occurs on a particular server is replicated to all metadata servers conforming to the application "start-end" semantics, resulting in an eventually consistent namespace. An asynchronous notification mechanism is also devised to enable users to deal with potential stale reads from operations of relaxed consistency. A prototype was implemented based on HDFS, and the experimental results show promising scalability and performance for both micro benchmarks and various real-world applications written in Pig, Hive and MapReduce. Ruini Xue, Lixiang Ao |
ICS | 2 |
| 2015 | BOLAS: Bipartite-Graph Oriented Locality-Aware Scheduling for MapReduce TasksabstractTask scheduling is critical to reduce the make span of MapReduce jobs. It is an effective approach for scheduling optimization by improving the data locality, which involves attempting to locate a task and its related data block on the same node. However, recent studies have been insufficient in addressing the locality issue. This paper proposes BOLAS, a MapReducetask scheduling algorithm, which models the scheduling processes a bipartite-graph matching problem trying best to assign data block to the nearest task. By considering the divergence of node performance of distribution of data blocks in MapReduce cluster, BOLAS can achieve a high degree of data locality, guarantee minimal data transfer during execution, and reduces a job's makespan subsequently. As a dynamic algorithm, BOLAS solves the model using Kuhn-Munkres optimal matching algorithm, and can be deployed in either homogeneous or heterogeneous environments. In this study, BOLAS was implemented as a plug in for Hadoop, and the experimental results indicate that BOLAScan localize nearly 100% of the map tasks and reduce the execution time by up to 67.1%. Ruini Xue, Shengli Gao, Lixiang Ao, Zhongyang Guan |
ISPDC | 1 |
| 2015 | COMET: Client-Oriented METadata Service for Highly Available Distributed File SystemsabstractHighly available metadata services of distributed file systems are essential to cloud applications. However, existing highly available metadata designs lack client-oriented features that treat metadata discriminately, leading to a single metadata fault domain and low availability. After investigating the workload characteristics of Hadoop, we propose Client-Oriented METadata (COMET), a novel highly available metadata service design that divides and distributes metadata into independent regions in terms of clients. These regions are isolated fault domains inherently, and failures in one region will not break file operations in other regions. A prototype of COMET was implemented based on HDFS, and the experimental results show that COMET can significantly improve metadata availability of HDFS without obvious performance degradation. It can also deliver scalable performance and faster metadata recovery due to its decentralized architecture. Ruini Xue, Lixiang Ao, Zhongyang Guan |
SBAC-PAD | 1 |
| 2015 | A Service Framework for Scientific Workflow Management in the CloudabstractCloud computing is an emerging computing paradigm that can offer unprecedented scalability and resources on demand, and is getting more and more adoption in the science community, while scientific workflow management systems provide essential support such as management of data and task dependencies, job scheduling and execution, provenance tracking, etc., to scientific computing. As we are entering into a “big data” era, it is imperative to migrate scientific workflow management systems into the cloud to manage the ever increasing data scale and analysis complexity. We propose a reference service framework for integrating scientific workflow management systems into various cloud platforms, which consists of eight major components, including Cloud Workflow Management Service, Cloud Resource Manager, etc., and six interfaces between them. We also present a reference framework for the implementation of Cloud Resource Manager, which is responsible for the provisioning and management of virtual resources in the cloud. We discuss our implementation of the framework by integrating the Swift scientific workflow management system with the OpenNebula and Eucalyptus cloud platforms, and demonstrate the capability of the solution using a NASA MODIS image processing workflow and a production deployment on the Science@Guoshi network with support for the Montage image mosaic workflow. Yong Zhao 0009, Youfu Li 0002, Ioan Raicu, Shiyong Lu, Cui Lin, Wenhong Tian, Ruini Xue |
IEEE Trans. Serv. Comput. | 8 |
| 2013 | An Energy-Efficient Online Parallel Scheduling Algorithm for Cloud Data CentersabstractThis paper considers online energy-efficient scheduling of real-time virtual machines (VMs) for Cloud data centers. Each request is associated with a starttime, a end-time, a processing time and demand for a Physical Machine (PM) capacity. The goal is to schedule all of the requests non-preemptively in their start-timeend- time windows, subjecting to PM capacity constraints, such that total busy time of all used PMs is minimized (called MinTBT-ON for abbreviation). This problem is a fundamental scheduling problem for parallel jobs allocation on mutliple machines, it has important applications in power-aware scheduling in cloud computing, optical network design and customer service systems and other related areas. Offline scheduling to minimize busy time is NP-hard already in the special case where all jobs have the same processing time and can be scheduled in a fixed time interval. One best-known result for MinTBT-ON problem is a g-competitive algorithm for general instances using First-Fit algorithm for unit-size jobs, where g is the total capacity of a PM. In this paper, a B-competitive algorithm, GRID is proposed and proved for general case, where B is a natural number and 1 < B < g. More results are obtained and applied to Cloud computing to improve energy-efficiency. Wenhong Tian, Ruini Xue, Qin Xiong, Yunjun Hu |
SERVICES | 2 |
| 2010 | A Rate and Resource Detection Based Receive Buffer Adaptation Approach for High-Speed Data TransportationabstractWith the development of computing devices and networks, several efficient and high performance UDP-based protocols have been proposed and employed in recently emerging computing paradigms, e.g., pervasive or cloud computing, to transport large data. However, since the server in such protocols uses a fixed-size memory buffer to hold received packets before handling them, the buffer will be exhausted if these packets cannot be handled as fast as they arrive, impairing the performance dramatically even if there is plenty of free memory. To solve this problem, we propose a Rate and Resource Detection Based Buffer Adaptation Approach (RRDA). RRDA collects the difference between the server receiving and processing rate and the amount of free memory periodically. Based on these information, RRDA decides whether the receive buffer should be resized and if so, to what extent to adjust. RRDA can not only avoid the exhaustion of receive buffer when the server load is heavy, but also can free unnecessary memory when the load is low. Experimental results show that RRDA can reduce the occurrence of buffer exhaustion by a factor of 10 and improve the throughput remarkably, compared with the fixed-size buffer scheme. Hao Liu 0006, Yaoxue Zhang, Yue-Zhi Zhou, Ruini Xue |
ICCCN | 4 |
| 2009 | MPIWiz: subgroup reproducible replay of mpi applicationsabstractMessage Passing Interface (MPI) is a widely used standard for managing coarse-grained concurrency on distributed computers. Debugging parallel MPI applications, however, has always been a particularly challenging task due to their high degree of concurrent execution and non-deterministic behavior. Deterministic replay is a potentially powerful technique for addressing these challenges, with existing MPI replay tools adopting either data-replay or order-replay approaches. Unfortunately, each approach has its tradeoffs. Data-replay generates substantial log sizes by recording every communication message. Order-replay generates small logs, but requires all processes to be replayed together. We believe that these drawbacks are the primary reasons that inhibit the wide adoption of deterministic replay as the critical enabler of cyclic debugging of MPI applications. Ruini Xue, Xuezheng Liu, Ming Wu 0007, Zheng Zhang 0001, Geoffrey M. Voelker |
PPoPP | 1 |
| 2008 | CprFS: a user-level file system to support consistent file states for checkpoint and restartabstractCheckpoint and Restart (CPR) is becoming critical to large scale parallel computers, whose Mean Time Between Failures (MTBF) may be much shorter than the execution times of the applications. The CPR mechanism should be able to store and recover the states of virtual memory, communication and files for the applications in a consistent way. Ruini Xue |
ICS | 1 |