Xingyang Li

dblp:276/8145 · DBLP profile ↗
← Back
19ranked-venue papers
2as first author
18since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 8 · 7 since 2021Systems, architecture and hardware · 6 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Turning Failures into Value: Negative Experience Replay for RLVR via Confidence Gating and Boundary Failure Sampling
abstract
Jialiang Guo, Fucheng Xiong, Xu He, Haodong Zhao, Xingyang li, Ke Zeng, Xunliang Cai. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Jialiang Guo, Fucheng Xiong, Haodong Zhao, Xingyang Li
ACL (1)5
2026 BalanceGS: Algorithm-System Co-design for Efficient 3D Gaussian Splatting Training on GPU
abstract
D Gaussian Splatting (3DGS) has emerged as a promising 3D reconstruction technique. The traditional 3DGS training pipeline follows three sequential steps: Gaussian densification, Gaussian projection, and color splatting. Despite its promising reconstruction quality, this conventional approach suffers from three critical inefficiencies: (1) Skewed density allocation during Gaussian densification. The adaptive densification strategy in 3DGS makes skewed Gaussian allocation across dense and sparse regions. The number of Gaussians of dense regions can be $100 \times$ that of sparse regions, leading to Gaussian redundancy. (2) Imbalanced computation workload during Gaussian projection. The traditional one-to-one allocation mechanism between threads and pixels results in execution time discrepancies between threads, leading to $\sim 20 \%$ latency overhead. (3) Fragmented memory access during color splatting. Discrete storage of colors in memory fails to take advantage of data locality with fragmented memory access, resulting in $\sim 2.0 \times$ color memory access time. To tackle the above challenges, we introduce BalanceGS, the algorithm-system co-design for efficient training in 3DGS. (1) At the algorithm level, we propose heuristic workload-sensitive Gaussian density control to automatically balance point distributions - removing 80% redundant Gaussians in dense regions while filling gaps in sparse areas. (2) At the system level, we propose Similarity-based Gaussian sampling and merging, which replaces the static one-to-one thread-pixel mapping with adaptive workload distribution - threads now dynamically process variable numbers of Gaussians based on local cluster density. (3) At the mapping level, we propose reordering-based memory access mapping strategy that restructures RGB storage and enables batch loading in shared memory. Extensive experiments demonstrate that compared with 3DGS, our approach achieves a $1.44 \times$ training speedup on a NVIDIA A100 GPU with negligible quality degradation.
Jinhao Li 0006, Xingyang Li, Guohao Dai 0001
ASP-DAC6
2026 Nebula: Infinite-Scale 3D Gaussian Splatting in VR via Collaborative Rendering and Accelerated Stereo Rasterization
abstract
3D Gaussian splatting (3DGS) has drawn significant attention in the architectural community recently. However, current architectural designs often overlook the 3DGS scalability, making them fragile for extremely large-scale 3DGS. Meanwhile, the VR bandwidth requirement makes it impossible to deliver high-fidelity and smooth VR content from the cloud.
Zheng Liu 0022, Xingyang Li, Anbang Wu, Jieru Zhao, Fangxin Liu, Yiming Gan, Jingwen Leng, Yu Feng 0007
ASPLOS (2)3
2026 OpenACM: An Open-Source SRAM-Based Approximate CiM Compiler
abstract
The rise of data-intensive AI workloads has exacerbated the "memory wall" bottleneck. Digital Compute-in-Memory (DCiM) using SRAM offers a scalable solution, but its vast design space makes manual design impractical, creating a need for automated compilers. A key opportunity lies in approximate computing, which leverages the error tolerance of AI applications for significant energy savings. However, existing DCiM compilers focus on exact arithmetic, failing to exploit this optimization. This paper introduces OpenACM, the first open-source, accuracy-aware compiler for SRAM-based approximate DCiM architectures. OpenACM bridges the gap between application error tolerance and hardware automation. Its key contribution is an integrated library of accuracy-configurable multipliers (exact, tunable approximate, and logarithmic), enabling designers to make fine-grained accuracy-energy trade-offs. The compiler automates the generation of the DCiM architecture, integrating a transistor-level customizable SRAM macro with variation-aware characterization into a complete, open-source physical design flow based on OpenROAD and the FreePDK45 library. This ensures full reproducibility and accessibility, removing dependencies on proprietary tools. Experimental results on representative convolutional neural networks (CNNs) demonstrate that OpenACM achieves energy savings of up to 64% with negligible loss in application accuracy. The framework is available on OpenACM:URL.
JunHao Ma, Xingyang Li, Yule Sheng, Bochang Wang, Yiheng Wu, Shan Shen, Daying Sun
DATE3
2026 DCI-SiteDTA: drug-target affinity prediction based on binding sites detection and site-aware dual cross-interaction block
abstract
BACKGROUND: Predicting the binding affinity between drugs and proteins is crucial for accelerating drug discovery. However, traditional research methods typically treat binding site detection and affinity prediction as two separate tasks, lacking solutions that integrate both into a unified deep learning framework. Furthermore, existing approaches exhibit limitations in two aspects: (1) for binding site detection, fine-grained features within drug target binding pockets need to be encoded; and (2) for drug-target affinity (DTA) prediction, the fusion of drug and targets needs to consider a multidimensional fusion guided by binding sites, rather than via concatenation or simple attention mechanisms. RESULTS: To address the aforementioned challenges, we propose a drug-target affinity prediction based on binding sites detection and site-aware dual cross-interaction block(DCI-SiteDTA). Specifically, a multi-scale feature fusion is employed to extract both local and global contextual information from protein residues to get the binding site vector. Then, with a binding site-guided strategy, we introduce a dual cross-interaction fusion block. Designed to model drug-target interactions, this module constructs multilevel representations based on drug, target, and binding site information to capture their fundamental biological patterns of interaction and integration. Finally, our proposed DCI-SiteDTA model is evaluated on Davis and KIBA benchmark datasets. The experimental results reveal that our proposed model yields better accuracy on both binding site detection and DTA prediction. CONCLUSIONS: Our proposed model with the binding sites-guided strategy and dual cross-interaction block improves the binding site detection and affinity prediction of drug-target pairs.
Xingyang Li, Yuni Zeng
BMC Bioinform.2
2025 Harnessing Conventional Video Processing Insights for Emerging 3D Video Generation Models: A Comprehensive Attention-aware Way
abstract
Video Generation Models based on 3D full attention (3D-VGMs) have significantly enhanced video quality. However, their inference overhead remains substantial, primarily due to the high computational cost of the attention mechanism, which accounts for over 75% of computations. Inspired by the success of conventional video processing, where video compression exploits similarities among patches, we point out that the attention mechanism can also harness the benefits from similarities among tokens. Nonetheless, two critical problems arise: (1) How can similarities be efficiently acquired in real-time? (2) How can workload balance be maintained when similar tokens are randomly distributed? To address these problems and leverage similarities for 3DVGMs, we propose Simpicker, a comprehensive attentionaware algorithm-hardware co-design for 3D-VGMs. Our core methodology is to fully utilize similarities in attention through both coarse-grained and fine-grained approaches while adopting dynamic adaptive strategies to leverage them. From the algorithm perspective, we propose a speculation-based similarity exploitation algorithm, allowing real-time importance speculation on the frame level, which is coarse-grained, and the token level, which is fine-grained. From the micro-architecture perspective, we propose a buffered lookup table-based (LUT-based) multiplication architecture for FP-INT multiplication and further eliminate potential bank conflicts to accelerate unimportant attention computation. From the mapping perspective, SimPicker proposes an adaptive grouping strategy in speculation to tame workload imbalance caused by randomly distributed similar tokens and allow seamless integration of our algorithms. Extensive experiments show that Simpicker achieves an average of $5.21 \times 1.45 \times$ speedup and $17.92 \times 1.63 \times$ energy efficiency compared to the NVIDIA A100 GPU and the state-of-the-art accelerators.
Tianlang Zhao, Jun Liu 0117, Xingyang Li, Li Ding 0012, Jinhao Li 0006, Shuaiheng Li, Jinbo Hu, Guohao Dai 0001
DAC3
2025 SLTarch: Towards Scalable Point-Based Neural Rendering by Taming Workload Imbalance and Memory Irregularity
abstract
Rendering is critical in fields like 3D modeling, AR/VR, and autonomous driving, where high-quality, real-time output is essential. Point-based neural rendering (PBNR) offers a photorealistic and efficient alternative to conventional methods, yet it is still challenging to achieve real-time rendering on mobile platforms. We pinpoint two major bottlenecks in PBNR pipelines: LoD search and splatting. LoD search suffers from workload imbalance and irregular memory access, making it inefficient on off-the-shelf GPUs. Meanwhile, splatting introduces severe warp divergence across GPU threads due to its inherent sparsity.To tackle these challenges, we propose SLTarch, an algorithm-architecture co-designed framework. At its core, SLTarch introduces SLTree, a dedicated subtree-based data structure, and LTcore, a specialized hardware architecture tailored for efficient LoD search. Additionally, we co-design a divergence-free splatting algorithm with our simple yet principled hardware augmentation, SPcore, to existing PBNR accelerators. Compared to a mobile GPU, SLTarch achieves 3.9× speedup and 98% energy savings with negligible architecture overhead. Compared to existing accelerator designs, SLTarch achieves 1.8× speedup with 54% energy savings.
Xingyang Li, Yu Feng 0007, Yiming Gan, Jieru Zhao, Zihan Liu 0002, Jingwen Leng, Minyi Guo
ICCAD1
2025 OpenYield: An Open-Source SRAM Yield Analysis and Optimization Benchmark Suite
abstract
Static Random-Access Memory (SRAM) yield analysis is essential for semiconductor innovation, yet research progress faces a critical challenge: the large gap between simplified academic models and the complexities observed in practice. The lack of open, higher-fidelity benchmarks has hindered reproducibility and transferability, as promising academic techniques often fail to carry over to more realistic settings. We present OpenYield, an open-source ecosystem that aims to narrow this gap through three contributions: (i) An SRAM circuit generator that explicitly incorporates second-order effects (interconnect/line parasitics, inter-cell leakage coupling, and peripheralcircuit variations) that are commonly omitted in academic studies. (ii) A standardized evaluation platform with a simple interface and baseline yield-analysis implementations to enable fair comparisons and reproducible research on these higherfidelity circuits. (iii) An optimization platform for transistor-level sizing under these models, supporting reproducible studies of robustness/efficiency trade-offs. OpenYield aims to foster more reproducible and transferable progress in SRAM-yield research. The framework is publicly available at OpenYield:URL.
Shan Shen, Xingyang Li, Zhuohua Liu, Junhao Ma, Yiheng Wu, Yuquan Sun, Wei W. Xing
ICCD2
2025 Radial Attention: 𝒪(n log n) Sparse Attention with Energy Decay for Long Video Generation
Xingyang Li, Tianle Cai, Haocheng Xi, Shuo Yang 0011, Yujun Lin 0001, Lvmin Zhang, Jinbo Hu, Kelly Peng, Maneesh Agrawala, Ion Stoica, Kurt Keutzer, Song Han 0003
NeurIPS1
2025 Dual-stream collaborative network (DSCN): Multimodal sentiment analysis via modality-invariant and modality-specific representation learning
Shengbing Chen, Dongmei Zhou, Chao Wei 0007, Xingyang Li
Neurocomputing4
2025 Lexicographic Dual-Objective Path Finding in Multi-Agent Systems
abstract
Path finding in multi-agent systems aims to identify collision-free and cost-optimized paths for all agents with distinct start and goal positions. It poses challenging optimization problems. Existing research typically treats all agents equally, overlooking their differences in practical scenarios where they undertake the tasks of varying importance. In many application scenarios, agents must be differentiated into critical (c-agents) and acritical ones (a-agents) due to premium/general service, no-loaded/full-loaded states, and urgent/non-urgent tasks. Facing this practical need, this work focuses on multi-agent systems in which the different importance of agents must be considered; and tackles a lexicographic dual-objective variant of path-finding problem. The consideration of c-agents makes the concerned problem more useful yet more challenging than basic multi-agent path-finding problems. Different from existing multi-agent path planning methods that minimize the sum-of-costs of all agents, we optimize two objectives with preferences to emphasize the influence of c-agents on the system. The primary one is to minimize the sum-of-costs of c-agents and the secondary one is to minimize that of a-agents. As existing methods are inadequate for this unique challenge, we adapt a conflict-based search framework and design new two-level lexicographic dual-objective optimization methods to deal with it. A high level is responsible for iteratively expanding a search tree and adding constraints to resolve conflicts among agents. A low level is responsible for finding the path of each agent for the node newly expanded in the high level. By conducting numerous computational experiments, we verify the great performance of the presented methods in solving the concerned problem. We further develop a prototype system incorporating our methods and make it public to promote their practical application. This research contributes valuable insights and solutions to pathfinding challenges in multi-agent systems with critical and acritical agents. Note to Practitioners—This work addresses a multi-agent path finding problem in multi-agent systems involving both critical and acritical agents. The former are more important than the latter since they are assigned to perform more important tasks. This is a common scenario in manufacturing and service environments. We propose a two-level lexicographic dual-objective optimization framework to deal with the problem and underscore the importance of c-agents in the path finding problem. Three solution approaches are designed for practitioners to select based on their practical application needs. The computational experimental results and statistical analyses highlight the exceptional performance of our proposed approaches. In order to facilitate the practical application, we further provide a prototype system with our proposed approaches embedded, which is openly accessible for practitioners to realize their specific applications.
Shixin Liu, MengChu Zhou, Xingyang Li, Xiaochun Yang 0001
IEEE Trans Autom. Sci. Eng.5
2025 Multi-Mobile-Robot Transport and Production Integrated System Optimization
abstract
A production workshop with mobile robots can be considered as a hybrid system consisting of a production system and a transportation one. Mobile robots are responsible for transferring production tasks among the machines of a production system and constitute a multi-robot transport system. It is highly coupled with a production system because of the interdependency that exists between production scheduling and mobile robot assignment. In this work, we study their integrated optimization problem for a mobile robot-based job shop with blocking properties. Its aim is to minimize total completion time as an objective function to improve overall operational efficiency. We consider the speed of a robot that varies according to whether it is loaded or not. We formulate this new problem into a mixed integer linear program to provide an algebraic description. Then, we propose a constraint programming method to solve it with high efficiency. The superiority of constraint programming over mixed integer linear programming in terms of the number of variables and constraints is analyzed. Numerous experiments on benchmark examples show that constraint programming can well handle the concerned problem. Under a one-hour time limit, it can exactly solve its instances while mixed integer linear programming cannot. Under a one-minute time limit, it obtains much better solutions than mixed integer linear programming and heuristic strategies, thus implying its high potential to be put into industrial applications. Note to Practitioners—The integration of mobile robots into production workshops has emerged as a pivotal strategy to enhance the operational efficiency of an advanced manufacturing system. This integration transforms the traditional job shop into a hybrid system, where mobile robots play a crucial role in transferring production tasks among machines. This brings a unique challenge to practitioners due to the intricate interdependencies between production scheduling and mobile robot assignment. The focus of our study is on the optimization of a mobile robot-based job shop to improve its overall operational efficiency. Our approach involves formulating this complex problem as a mixed-integer linear program, thereby providing a concise mathematical representation. We propose a constraint programming method to solve the problem efficiently. Through numerous experiments on benchmark examples, our findings indicate that the proposed constraint programming method can well solve the concerned problem given long or short solution time. This underscores its high potential for practical implementation in industrial scenarios.
Xingyang Li, Shixin Liu, MengChu Zhou, Xiaochun Yang 0001
IEEE Trans Autom. Sci. Eng.2
2024 A Dynamic Operational Optimization Method for Robotic Mobile Fulfillment Systems with Inventory Discrepancy Events
abstract
A Robotic Mobile Fulfillment System (RMFS) is an emerging “cargos-to-person” picking system that relies on the broom of Internet of Things (IoT) technology. It aims to offer significant enhancements in order picking efficiency. However, dynamic disturbances, such as inventory discrepancies arising from errors in receiving, shipping, and handling of goods, often disrupt its operations, leading to degraded service and increased operational costs, thereby affecting overall system performance. Traditional optimization solutions may necessitate adaptations or overhauls in response to such disturbances. This paper introduces a proactive multi-pathway response algorithm tailored to mitigating dynamic disturbances in RMFS, particularly concerning inventory discrepancies. We extend an open-source simulation framework to evaluate the performance of the proposed algorithm and conduct a comparative analysis of dynamic systems. Experimental results indicate that our proposed algorithm can effectively improve the processing efficiency of abnormal orders with the minimal system-wide impact, highlighting its potential to well address dynamic and abnormal events in smart warehouses.
Huai Ma, Xingyang Li, Shixin Liu
SMC4
2024 Lexicographic Multi-objective Order Picking Optimization for Robotic Mobile Fulfillment Systems
abstract
In light of advancements in artificial intelligence, the Internet of Things, and mechatronics, robots are increasingly integrated into e-commerce warehouses to enable smart order picking solutions and foster intelligent automation. A robot-assisted order picking process revolutionizes the traditional labor-intensive person-to-goods order picking technology, leading to a goods-to-person (G2P) smart warehouse. Within it, robots transport pods to predefined picking stations, where human pickers retrieve the requested goods from these pods to fulfill customer orders. The allocation of pods to robots and the scheduling of picking operations are key optimization issues in G2P order picking systems. Although they play a key role in improving operational efficiency, existing research has paid limited attention to their joint optimization. This study considers a lexicographic multi-objective optimization problem to shorten the order picking cycles under the premise of optimizing the picking efficiency evaluated by makespan. We build a mixed integer program for the newly proposed problem and develop a matheuristic algorithm by integrating a commodity-order model into a metaheuristic algorithm to solve it. Experimental results show that the proposed method can significantly shorten the total order picking cycles while keeping the minimum makespan. It outperforms a recent state-of-the-art algorithm. This work emphasizes the importance of joint optimization within G2P smart warehouses and reveals the high potential of the proposed method to be used in practice.
Hanying Wang, Xingyang Li, Shixin Liu
SMC4
2024 HRGUNet: A novel high-resolution generative adversarial network combined with an improved UNet method for brain tumor segmentation
Dongmei Zhou, Hao Luo 0014, Xingyang Li, Shengbing Chen
J. Vis. Commun. Image Represent.3
2023 Pixel-Wise Reconstruction of Private Data in Split Federated Learning
Xingyang Li, Wenjian He
ICICS2
2022 Conflict-Based Search and Improvement Strategies for Solving a New Lexicographic Bi-Objective Multi-Agent Path Finding Problem
abstract
Multi-Agent Path Finding (MAPF) is an important problem with a variety of applications. Its aim is to find collision-free paths for agents having separate start and goal positions. This work proposes a new lexicographic bi-objective MAPF considering different task types, where agents are divided into two kinds to perform critical and acritical tasks. This is common in practical intelligent warehousing scenarios where a critical/acritical-task-performing agent (called c-agent and a-agent, respectively) may represent a full-load/no-load one or the one conducting urgent/non-urgent tasks. The primary objective is to minimize the sum-of-costs of c-agents, while the secondary objective is to minimize the sum-of-costs of a-agents. Two MAPF algorithms are modified to fit and solve the concerned problem for the first time. Moreover, four improvement strategies are embedded to the proposed algorithms and proved to be effective in solving MAPF problems with different task types.
Xingyang Li, MengChu Zhou, Shixin Liu
SMC2
2021 M6A2Target: a comprehensive database for targets of m6A writers, erasers and readers
abstract
N6-methyladenosine (m6A) is the most abundant posttranscriptional modification in mammalian mRNA molecules and has a crucial function in the regulation of many fundamental biological processes. The m6A modification is a dynamic and reversible process regulated by a series of writers, erasers and readers (WERs). Different WERs might have different functions, and even the same WER might function differently in different conditions, which are mostly due to different downstream genes being targeted by the WERs. Therefore, identification of the targets of WERs is particularly important for elucidating this dynamic modification. However, there is still no public repository to host the known targets of WERs. Therefore, we developed the m6A WER target gene database (m6A2Target) to provide a comprehensive resource of the targets of m6A WERs. M6A2Target provides a user-friendly interface to present WER targets in two different modules: 'Validated Targets', referred to as WER targets identified from low-throughput studies, and 'Potential Targets', including WER targets analyzed from high-throughput studies. Compared to other existing m6A-associated databases, m6A2Target is the first specific resource for m6A WER target genes. M6A2Target is freely accessible at http://m6a2target.canceromics.org.
Shuang Deng, Hongwan Zhang, Kaiyu Zhu, Xingyang Li, Xuefei Liu, Dongxin Lin, Zhixiang Zuo
Briefings Bioinform.4
2020 CrossICC: iterative consensus clustering of cross-platform gene expression data without adjusting batch effect
abstract
Unsupervised clustering of high-throughput gene expression data is widely adopted for cancer subtyping. However, cancer subtypes derived from a single dataset are usually not applicable across multiple datasets from different platforms. Merging different datasets is necessary to determine accurate and applicable cancer subtypes but is still embarrassing due to the batch effect. CrossICC is an R package designed for the unsupervised clustering of gene expression data from multiple datasets/platforms without the requirement of batch effect adjustment. CrossICC utilizes an iterative strategy to derive the optimal gene signature and cluster numbers from a consensus similarity matrix generated by consensus clustering. This package also provides abundant functions to visualize the identified subtypes and evaluate subtyping performance. We expected that CrossICC could be used to discover the robust cancer subtypes with significant translational implications in personalized care for cancer patients. AVAILABILITY AND IMPLEMENTATION: The package is implemented in R and available at GitHub (https://github.com/bioinformatist/CrossICC) and Bioconductor (http://bioconductor.org/packages/release/bioc/html/CrossICC.html) under the GPL v3 License.
Qi Zhao 0009, Yu Sun 0050, Hongwan Zhang, Xingyang Li, Kaiyu Zhu, Zexian Liu, Jian Ren 0002, Zhixiang Zuo
Briefings Bioinform.5