EDBT 2026 Demo / reviewers in the wild / expert
Ping Yin
dblp:26/6924
· DBLP profile ↗
16ranked-venue papers
4as first author
11since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 3 since 2021Computer networks · 4 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MFS: An Efficient Model Family Serving System for LLMsabstractLLM serving providers typically offer a suite of structurally similar models, known as model families, such as the open-source Llama2 series featuring 7B, 13B, and 70B models. While numerous optimizations for LLM serving have been proposed, the potential for leveraging synergies between models within the same family has not been thoroughly explored. This paper introduces MFS, an innovative multi-tiered LLM model family serving system to exploit the structural similarities and parameter redundancies across different scales of models within a family. By utilizing a novel fine-tuning technique called Knowledge Precipitation, MFS restructures the largest model in a family to encapsulate smaller models within its architecture, enabling a unified multi-tiered serving pipeline. Based on the multi-tiered model, MFS realizes a highly parallelized tiered-level batching approach, significantly enhancing system efficiency. It also enables the sharing of intermediate features and KV-cache between models and facilitates multi-level sampling techniques during the inference phase. Experimental results demonstrate that MFS achieves substantial improvements over existing methods, including a 56.1% reduction in end-to-end token generation latency and a 47.8% decrease in GPU memory footprint without compromising the quality of generated content. Yunxuan Zhang, Hao Wang 0116, Han Tian, Liu Yang 0008, Xudong Liao, Wenxue Li 0004, Ping Yin, Bowen Liu 0002, Kai Chen 0005 |
EuroSys | 7 |
| 2026 | A deep learning-based 3D reconstruction framework for Hongcun village using cultural-spatial attention networks
Yuye Gong, Dongliang Huang, Nengshun Peng, Lerong Li, Ping Yin |
Multim. Syst. | 6 |
| 2026 | Learning Heterogeneous Mixture of Scene Experts for Large-Scale Neural Radiance FieldsabstractRecent Neural Radiance Field (NeRF) methods on large-scale scenes have demonstrated promising results and underlined the importance of scene decomposition for scalable NeRFs. Although these methods achieved reasonable scalability, there are several critical problems remaining unexplored in the existing large-scale NeRF modeling methods, i.e., learnable decomposition, modeling scene heterogeneity, and modeling efficiency. In this paper, we introduce Switch-NeRF++, a Heterogeneous Mixture of Hash Experts (HMoHE) network that addresses these challenges within a unified framework. Our framework is a highly scalable NeRF that learns heterogeneous decomposition and heterogeneous Neural Radiance Fields efficiently for large-scale scenes in an end-to-end manner. In our framework, a gating network learns to decompose scenes into partitions and allocates 3D points to specialized NeRF experts. This gating network is co-optimized with the experts by our proposed Sparsely Gated Mixture of Experts (MoE) NeRF framework. Our network architecture incorporates a hash-based gating network and distinct heterogeneous hash experts. The hash-based gating efficiently learns the decomposition of the large-scale scene. The distinct heterogeneous hash experts consist of hash grids of different resolution ranges. This enables effective learning of the heterogeneous representation of different decomposed scene parts within large-scale complex scenes. These design choices make our framework an end-to-end and highly scalable NeRF solution for real-world large-scale scene modeling to achieve both quality and efficiency. We evaluate our accuracy and scalability on existing large-scale NeRF datasets. Additionally, we also introduce a new dataset with very large-scale scenes ($ {>} 6.5\,\text{km}^{2}$>6.5km2) from UrbanBIS. Extensive experiments demonstrate that our approach can be easily scaled to various large-scale scenes and achieve state-of-the-art scene rendering accuracy. Furthermore, our method exhibits significant efficiency gains, with an 8x acceleration in training and a 16x acceleration in rendering compared to the best-performing competitor Switch-NeRF. Zhenxing Mi, Ping Yin, Dan Xu 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2026 | Toward Fine-Grained Load Balancing With Congested-Flow Isolation in Lossless DatacentersabstractRemote Direct Memory Access (RDMA) over Converged Ethernet (RoCE) cooperating with Priority Flow Control (PFC) has been widely deployed in production datacenters to enable low latency, lossless transmission. At the same time, modern datacenters typically offer parallel transmission paths between any pair of end-hosts, underscoring the importance of load balancing. However, the well-studied load balancing mechanisms designed for lossy datacenter networks (DCNs) are ill-suited for such lossless environments. Through extensive experiments, we are among the first to comprehensively inspect the interactions between PFC and load balancing, and uncover that existing fine-grained rerouting schemes can be counterproductive to spread the congested flows among more paths, further aggravating PFC’s head-of-line (HoL) blocking. Motivated by this, we present FLB, a Fine-grained Load Balancing scheme for lossless DCNs. At its core, FLB employs threshold-free rerouting to effectively balance traffic load and improve link utilization during normal conditions and leverages timely congested flow isolation to eliminate HoL blocking on non-congested flows when congestion occurs. To handle complex multi-bottleneck scenarios, we further introduce FLB*, which incorporates an enhanced congestion-point-aware isolation mechanism using Congestion Point Identifiers (CPI) to eliminate HoL blocking among different congested flows.We have fully implemented a FLB prototype, and our evaluation results show that FLB reduces PFC PAUSE rate by up to 96% and avoids HoL blocking, translating to up to 45% improvement in goodput over CONGA+DCQCN and 40%, 36%, 29% and 18% reduction in average flow completion time (FCT) over LetFlow+Swift, MP-RDMA, Proteus+DCQCN and LetFlow+PCN, respectively. Jinbin Hu 0001, Siyao Li, Wenxue Li 0004, Xiangzhou Liu, Bowen Liu 0002, Ping Yin, Mengyu Ma, Jin Wang 0001, Jianxin Wang 0001, Jiawei Huang 0001, Kai Chen 0005 |
IEEE Trans. Netw. | 6 |
| 2025 | CEIO: A Cache-Efficient Network I/O Architecture for NIC-CPU Data PathsabstractEfficient Input/Output (I/O) data path between NICs and CPUs/DRAMs is critical for supporting datacenter applications with high-performance network transmission, especially as link speed scales to 100Gbps and beyond. Traditional I/O acceleration strategies, such as Data Direct I/O (DDIO) and Remote Direct Memory Access (RDMA), perform suboptimally due to the inefficient utilization of the Last-Level Cache (LLC). This paper presents CEIO, a novel cache-efficient network I/O architecture that employs proactive rate control and elastic buffering to achieve zero LLC misses in the I/O data path while ensuring the effectiveness of DDIO and RDMA under various network conditions. We have implemented CEIO on commodity SmartNICs and incorporated it into widely-used DPDK and RDMA libraries. Experiments with well-optimized RPC framework and distributed file system under realistic workloads demonstrate that CEIO achieves up to 2.9× higher throughput and 1.9× lower P99.9 latency over prior work. Bowen Liu 0002, Qijing Li, Zhuobin Huang, Yijun Sun, Wenxue Li 0004, Junxue Zhang 0001, Ping Yin, Kai Chen 0005 |
SIGCOMM | 8 |
| 2025 | FLB: Fine-grained Load Balancing for Lossless Datacenter Networks
Jinbin Hu 0001, Wenxue Li 0004, Xiangzhou Liu, Bowen Liu 0002, Ping Yin, Jianxin Wang 0001, Jiawei Huang 0001, Kai Chen 0005 |
USENIX ATC | 6 |
| 2025 | Balanced ID-OOD tradeoff transfer makes query based detectors good few shot learnersabstractFine-tuning is a popular approach to solve the few-shot object detection problem. In this paper, we attempt to introduce a new perspective on it. We formulate the few-shot novel tasks as a type of distribution shifted from its ground-truth distribution. We introduce the concept of imaginary placeholder masks to show that this distribution shift is essentially a composite of in-distribution(ID) and out-of-distribution(OOD) shifts. Our empirical investigation results show that it is significant to balance the trade-off between adapting to the available few-shot distribution and keeping the distribution-shift robustness of the pre-trained model. We explore improvements in the few-shot fine-tuning transfer in the few-shot object detection(FSOD) settings from three aspects. First, we explore the LinearProbe-Finetuning(LP-FT) technique to balance this trade-off to mitigate the feature distortion problem. Second, we explore the effectiveness of utilizing the protection freezing strategy for query-based object detectors to keep their OOD robustness. Third, we try to utilize ensembling methods to circumvent the feature distortion. All these techniques are integrated into a whole method called BIOT(Balanced ID-OOD Transfer). Evaluation results show that our method is simple yet effective and general to tap the FSOD potential of query-based object detectors. It outperforms the current SOTA method in many FSOD settings and has a promising scaling capability. Yuantao Yin, Ping Yin, Siqing Sun, Xiaobo An |
High Confid. Comput. | 2 |
| 2024 | Privacy Protection for Image Sharing Using Reversible Adversarial ExamplesabstractOnline image sharing on social media platforms faces information leakage due to deep learning-aided privacy attacks. To avoid these attacks, this paper proposes a privacy protection mechanism for image sharing without changing the visual effect, which is based on reversible adversarial examples. Specifically, social media platform users can change the class activation feature to convert the original image into an adversarial image before sharing. When users want to restore the adversarial image to the original image, they can use an improved generative adversarial network model to restore it. The experimental results prove that the conversion model in this paper can effectively prevent privacy attacks from analyzing and stealing users' private information while having no visual impact. At the same time, the proposed restoration model can restore the adversarial examples with high accuracy. Ping Yin, Wei Chen 0006, Jiaxi Zheng, Lifa Wu |
ICC | 1 |
| 2023 | Runtime Row/Column Activation Pruning for ReRAM-based Processing-in-Memory DNN AcceleratorsabstractResistive random access memory (ReRAM)-based processing-in-memory (PIM) DNN accelerators have shown great potential in improving model efficiency and saving energy. To further improve memory and computation efficiency, model weight sparsity has been widely explored in ReRAM-based accelerator designs. However, these optimized accelerators rarely touched the model activation sparsity. In this paper, we observe that there exist plenty sparse rows/columns in the DNN model activation matrix which have negligible effect on accuracy, termed as insensitive rows/columns. Pruning them has little impact on model accuracy but would have significant potential to improve the performance and energy efficiency of DNN accelerators. Therefore, we propose a new ReRAM-based PIM accelerator, named as RapPIM, to take advantage of the model activation sparsity. In RapPIM, we first propose an insensitive activation rows/columns pruning method to search and prune the insensitive rows/columns. Then, we present an activation low-bits skipping strategy and a forward propagation delay hiding strategy to further improve model performance and minimize the latency of activation pruning on forward propagation. Our evaluations with several well-known DNN models show that the RapPIM achieves up to 2.40× speedup and 44.82% power reduction compared with the state-of-the-art ReRAM-based accelerator. Xikun Jiang, Zhaoyan Shen, Siqing Sun, Ping Yin, Zhiping Jia, Lei Ju 0001, Zhiyong Zhang 0006, Dongxiao Yu |
ICCAD | 4 |
| 2022 | Two-dimensional dynamic time warping algorithm for matrices similarityabstractDynamic Time Warping (DTW algorithm) provides an effective method to obtain the similarity between unequal-sized signals. However, it cannot directly deal with high-dimensional samples such as matrices. Expanding a matrix to one dimensional vector as the input data of DTW will decrease the measure accuracy because of the losing of position information in the matrix. Aiming at this problem, a two-dimensional dynamic time warping algorithm (2D-DTW) is proposed in this paper to directly measure the similarity between matrices. In 2D-DTW algorithm, a three dimensional distance-cuboid is constructed, and its mapped distance matrix is defined by cutting and compressing the distance-cuboid. By introducing the dynamic programming theory to search the shortest warping path in the mapped matrix, the corresponding shortest distance can be obtained as the expected similarity measure. The experimental results suggest that the performance of 2D-DTW distance is superior to the traditional Euclidean distance and can improve the similarity accuracy between matrices by introducing the warping alignment mechanisms. 2D-DTW algorithm extends the application ranges of traditional DTW and is especially suitable for high-dimensional data. Cuifang Gao, Wanqiang Shen, Ping Yin |
Intell. Data Anal. | 4 |
| 2022 | Non-Interlaced Dynamic Time Warping for Distance Between Matrixes
Cuifang Gao, Ping Yin |
Neural Process. Lett. | 3 |
| 2019 | Strengthening the Positive Effect of Viral MarketingabstractIn traditional viral marketing, the goal is to reach out to the maximum number of people. However, some studies have demonstrated that spreading a product indiscriminately in a network can cause some counter effect because it may reach people who evaluate it negatively. In this paper, we study how to make use of social networks to avoid negative people so that the 'positive effect' of viral marketing can be maximized, and this optimization problem is called Strengthening the Positive Effect (SPE). SPE has a non-monotone and non-submodular objective function, and it is NP-hard to be approximately solved with any positive factor. Although SPE is almost impossible to solve approximately, we make the pioneer contribution by discovering that: 1) The almost optimal solution is obtainable in some network; 2) For the general network, a polynomial algorithm that yields a multiplicative guarantee is also possible under a reasonable assumption. We test our solution on various realworld social networks with a comprehensive set of experiments. The result affirms that besides its performance analyzability, our solution is more scalable than the current heuristic. Yuqing Zhu 0002, Ping Yin, Deying Li 0001, Bill Lin 0001 |
ICDCS | 2 |
| 2019 | A Study of the Contribution of Information Technology on the Growth of Tourism Economy Using Cross-Sectional DataabstractInformation technology (IT) has dramatically changed tourism industry, particularly in facilitating and improving information discovery and dissemination in tourism industry. Prior research has identified the key role that IT plays in the development of tourism industry. However, IT is examined solely, while other factors that drive the development of tourism industry are neglected in these studies. Guided by the new economic growth theory, this paper integrates fixed capital, labor, and IT and further examine how they together affect the development of tourism industry. Based on the cross-sectional tourism data of 2014 in 30 provinces in China, this article runs gray correlation analysis first and then applies a Cobb-Douglas production function model. The results indicate that fixed capital, labor, and IT are all correlated with the total tourism revenue at a certain degree, and that the development of tourism industry replies more on labor and IT than on fixed capital. Labor is the factor that makes the largest contribution to the development of tourism industry. Ping Yin, Xianrong Zheng |
J. Glob. Inf. Manag. | 1 |
| 2018 | MeDEStrand: an improved method to infer genome-wide absolute methylation levels from DNA enrichment dataabstractBACKGROUND: DNA methylation of CpG dinucleotides is an essential epigenetic modification that plays a key role in transcription. Widely used DNA enrichment-based methods offer high coverage for measuring methylated CpG dinucleotides, with the lowest cost per CpG covered genome-wide. However, these methods measure the DNA enrichment of methyl-CpG binding, and thus do not provide information on absolute methylation levels. Further, the enrichment is influenced by various confounding factors in addition to methylation status, for example, CpG density. Computational models that can accurately derive absolute methylation levels from DNA enrichment data are needed. RESULTS: We developed "MeDEStrand," a method that uses a sigmoid function to estimate and correct the CpG bias from enrichment results to infer absolute DNA methylation levels. Unlike previous methods, which estimate CpG bias based on reads mapped at the same genomic loci, MeDEStrand processes the reads for the positive and negative DNA strands separately. We compared the performance of MeDEStrand to that of three other state-of-the-art methods "MEDIPS," "BayMeth," and "QSEA" on four independent datasets generated using immortalized cell lines (GM12878 and K562) and human primary cells (foreskin fibroblasts and mammary epithelial cells). Based on the comparison of the inferred absolute methylation levels from MeDIP-seq data and the corresponding reduced-representation bisulfite sequencing data from each method, MeDEStrand showed the best performance at high resolution of 25, 50, and 100 base pairs. CONCLUSIONS: The MeDEStrand tool can be used to infer whole-genome absolute DNA methylation levels at the same cost of enrichment-based methods with adequate accuracy and resolution. R package MeDEStrand and its tutorial is freely available for download at https://github.com/jxu1234/MeDEStrand.git . Jingting Xu, Shimeng Liu, Ping Yin, Serdar Bulun |
BMC Bioinform. | 3 |
| 2017 | Improving Backpressure-based Adaptive Routing via Incremental Expansion of Routing ChoicesabstractBackpressure-based adaptive routing algorithms have been studied extensively in the literature. Although backpressure-based adaptive routing algorithms have been shown to be network-wide throughput optimal, they typically have poor delay performance under light or moderate loads because packets may be sent over unnecessarily long routes. Further, backpressure-based algorithms have required every node to compute differential backlogs for every destination queue with the corresponding destination queue at every adjacent node. This computation is expensive given the large number of possible pairwise differential backlogs and requires many exchanges of backlog information between adjacent nodes. In this paper, we propose new backpressure-based adaptive routing algorithms that only use shortest-path routes to destinations when they are sufficient to accommodate the given traffic load, but the proposed algorithms will incrementally expand routing choices as needed to accommodate increasing traffic loads. We show analytically by means of fluid analysis that the proposed algorithms retain network-wide throughput optimality, and we show empirically by means of simulations that our proposed algorithms provide substantial improvements in delay performance. Our evaluations further show that in practice, our approach dramatically reduces the number of pairwise differential backlogs that have to be computed and the amount of corresponding backlog information that has to be exchanged because routing choices are only incrementally expanded as needed. Ping Yin, Sen Yang 0001, Jun (Jim) Xu, Jim Dai, Bill Lin 0001 |
ANCS | 1 |
| 2005 | Cluster based real-time rendering system for large terrain datasetabstractIn this paper we propose an efficient system for rendering large terrain data in a cluster environment. Both geometry and texture data are stored in the server, constructed and optimized as pre-computed static quadtree. Hierarchical view frustum culling and view-dependent texture and geometry refinement is performed through a top-down traversal algorithm in the client. We guarantee overall geometric continuity and exploit programmable graphics hardware to achieve low CPU overload and high triangles throughput. Speculative terrain data prefetching is exploited for hiding network delay. Large datasets can be rendered in real-time with our system and the performance is almost independent of the dataset size. Ping Yin, Jiaoying Shi |
CAD/Graphics | 1 |