Xinyuan Lin

dblp:250/8399 · DBLP profile ↗
← Back
6ranked-venue papers
0as first author
6since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Click, Share, Learn: Teaching Data Science Using Apache Texera
Sarah Asad, Kun Woo Park, Xinyuan Lin, Jiadong Bai, Shengquan Ni, Yicong Huang 0002, Chen Li 0001
EDBT4
2026 HybridSpec: Exploiting Hybrid-Bonding Memory to Accelerate LLM Serving Through Heterogeneous Architecture and Speculative Decoding
Zongle Huang, Wenbin Jia, Yaolei Li, Xinyuan Lin, Shupei Fan, Shuwen Deng, Yongpan Liu
ISCA4
2025 MPICC: Multiple-Precision Inter-Combined MAC Unit with Stochastic Rounding for Ultra-Low-Precision Training
abstract
Recent studies have proved the feasibility of ultra-low-precision (≤ 8-bit) training. However, most of the existing operational circuits support only a few higher precisions (FP16, FP32, etc.) for multiplication, and the bit width for accumulation cannot be reduced, which results in low area and power efficiencies. In this paper, we propose MPICC, a multiple-precision Multiply-Accumulate (MAC) unit designed for ultra-low-precision training. It supports inter-combined computations across 18 different precision combinations, including LOG4 (radix-4 FP4), FP4, FP6, FP8, and INT4. It also reduces the accumulation precision from FP16 to FP12 through an optimized Stochastic Rounding (SR) strategy to further save logic resources. Moreover, a low-cost emulating controller, which time-division multiplexed the low-precision MAC unit, is also designed to accomplish high-precision computations for critical DNN layers. Compared with the existing multiple-precision computing units, the area/energy efficiencies of this design are improved by 1.17×/1.19× at FP8, and 4.69×/3.64× at FP4, respectively. The SR strategy further reduces the area/power consumption by 15.6%/14.9% of the floating-point accumulator.
Leran Huang, Yongpan Liu, Xinyuan Lin, Chenhan Wei, Wenyu Sun, Zengwei Wang, Boran Cao, Xiaoxia Fu
ASP-DAC3
2024 DCU-CHK: checkpointing for large-scale CPU-DCU heterogeneous computing systems
Xinyuan Lin, Yi Liu 0013
CCF Trans. High Perform. Comput.2
2024 Pasta: A Cost-Based Optimizer for Generating Pipelining Schedules for Dataflow DAGs
abstract
Data analytics tasks are often formulated as data workflows represented as directed acyclic graphs (DAGs) of operators. The recent trend of adopting machine learning (ML) techniques in workflows results in increasingly complicated DAGs with many operators and edges. Compared to the operator-at-a-time execution paradigm, pipelined execution has benefits of reducing the materialization cost of intermediate results and allowing operators to produce results early, which are critical in iterative analysis on large data volumes. Correctly scheduling a workflow DAG for pipelined execution is non-trivial due to the richer semantics of operators and the increasing complexity of DAGs. Several existing data systems adopt simple heuristics to solve the problem without considering costs such as materialization sizes. In this paper, we systematically study the problem of scheduling a workflow DAG for pipelined execution, and develop a novel cost-based optimizer called Pasta for generating a high-quality schedule. The Pasta optimizer is not only general and applicable to a wide variety of cost functions, but also capable of utilizing properties inherent in a broad class of cost functions to improve its performance significantly. We conducted a thorough evaluation of developed techniques on real-world workflows and show the efficiency and efficacy of these solutions.
Yicong Huang 0002, Xinyuan Lin, Avinash Kumar 0004, Sadeem Alsudais, Chen Li 0001
Proc. ACM Manag. Data3
2024 Texera: A System for Collaborative and Interactive Data Analytics Using Workflows
abstract
Domain experts play an important role in data science, as their knowledge can unlock valuable insights from data. As they often lack technical skills required to analyze data, they need collaborations with technical experts. In these joint efforts, productive collaborations are critical not only in the phase of constructing a data science task, but more importantly, during the execution of a task. This need stems from the inherent complexity of data science, which often involves user-defined functions or machine-learning operations. Consequently, collaborators want various interactions during runtime, such as pausing/resuming the execution, inspecting an operator's state, and modifying an operator's logic. To achieve the goal, in the past few years we have been developing an open-source system called Texera to support collaborative data analytics using GUI-based workflows as cloud services. In this paper, we present a holistic view of several important design principles we followed in the design and implementation of the system. We focus on different methods of sending messages to running workers, how these methods are adopted to support various runtime interactions from users, and their trade-offs on both performance and consistency. These principles enable Texera to provide powerful user interactions during a workflow execution to facilitate efficient collaborations in data analytics.
Zuozhi Wang, Yicong Huang 0002, Shengquan Ni, Avinash Kumar 0004, Sadeem Alsudais, Xinyuan Lin, Yunyan Ding, Chen Li 0001
Proc. VLDB Endow.7