VLDB 2026 Research / reviewers in the wild / expert
Yiping Yao
dblp:36/5447 · also Yi-Ping Yao
· DBLP profile ↗
33ranked-venue papers
1as first author
17since 2021 · last 2026
0009-0008-8158-5501ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 7 since 2021Systems, architecture and hardware · 10 · 3 since 2021Human-computer interaction and ubiquitous computing · 8 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Sociologically-informed opinion prediction: Fusing bounded confidence theory with TabTransformer
Jiao Luo, Yiping Yao |
Neurocomputing | 2 |
| 2026 | Veritas: Structuring and verifying LLM knowledge for logically consistent behavior tree generation in LLM-based agentsabstractBehavior trees (BTs) have been widely adopted in autonomous task planning because of their modularity and reactivity. Recently, automatic BT generation based on Large Language Models (LLMs) has attracted growing research attention. However, synthesizing BTs for long-horizon tasks without relying on predefined expert rules remains an open problem, and it poses two key challenges: ensuring the logical consistency of the generated BTs, and maintaining their factual alignment with the ground truth of the environment. To address these challenges, this paper presents Veritas, a novel verification-driven framework for automatically generating logically consistent BTs. Veritas integrates STRIPS-like symbolic operators with a multi-layered verification mechanism, thereby transforming the planning process into a coherent chain of logical derivations. We further introduce Veritas+, which augments this framework with a memory module that accumulates both successful and failed execution experiences, enabling dynamic self-correction and improving factual consistency. We evaluate the framework on 67 long-horizon tasks in Minecraft and 116 real-world tasks in AndroidWorld. Experimental results show that Veritas and Veritas+ are highly effective and significantly outperform state-of-the-art baselines. Kejia Wan, Songyi Lu, Yiping Yao, Xinhai Xu |
Inf. Process. Manag. | 6 |
| 2025 | Temporal Adaptive Neural Message Passing for Opinion DynamicsabstractAccurate prediction of opinion evolution is crucial for understanding public opinion dynamics. The neural message passing mechanisms, which exhibit conceptual similarities with opinion dynamics, are capable of autonomously updating parameters through gradient-based optimization techniques, thereby establishing themselves as a surrogate model for investigating the evolution of opinions. However, existing methods face three key limitations: (1) Social weights between individuals rely solely on node attribute similarity, ignoring temporal dynamics; (2) Temporal features are modeled only through time derivatives, and lacking robustness to external noise; (3) There is a lack of research on real-world opinion data with network structures, especially open-source datasets. To tackle these challenges, we propose Temporal Adaptive Neural Message Passing, a deep learning framework that addresses opinion dynamics through adaptive message passing, noise-aware data fusion to enhance long-term forecasting. Furthermore, we release three large-scale real-world datasets to serve as benchmarks for future research. For evaluation, we conducted experiments on three real-world and three synthetic datasets, comparing our method with existing mechanistic models (including hybrid models) and non-mechanistic models. The results demonstrate that our method achieves the best performance across all datasets. Lizhen Ou, Yiping Yao, Kai Chen 0020 |
ECAI | 2 |
| 2025 | Hippocampal-like Sequential Editing for Continual Knowledge Updates in Large Language ModelsabstractLarge language models (LLMs) are now pivotal in real-world applications. Model editing has emerged as a promising paradigm for efficiently modifying LLMs without full retraining. However, current editing approaches face significant limitations due to parameter drift, which stems from inconsistencies between newly edited knowledge and the model's existing knowledge. In sequential editing scenarios, cumulative drifts progressively lead to model collapse characterized by general capability degradation and balance between acquiring new knowledge and catastrophic forgetting of existing knowledge. Drawing inspiration from the hippocampal trisynaptic circuit for continual memorizing and forgetting, we propose a Hippocampal-like Sequential Editing (HSE) framework that designs the unlearning of obsolete knowledge, domain-specific knowledge update separation and replay for edited knowledge. Specifically, the HSE framework designs three core mechanisms: (1) Machine unlearning selectively erases outdated knowledge to facilitate integration of new information, (2) Fisher Information Matrix-guided parameter updates prevents cross-domain knowledge interference, and (3) Parameter replay consolidates long-term editing memory through lightweight and global replay of editing data in a parametric form. Theoretical analysis demonstrates that HSE achieves smaller generalization error bounds, more stable convergence and higher computational efficiency. Experimental results validate its effective balance between acquiring new knowledge and mitigating catastrophic forgetting, maintaining or even slightly enhancing general capabilities. In practical applications, experiments confirm its effectiveness in multi-domain hallucination mitigation, healthcare knowledge injecting, and societal bias reduction. Quntian Fang, Zhen Huang 0006, Zhiliang Tian, Minghao Hu 0001, Dongsheng Li 0001, Yiping Yao, Xinyue Fang, Menglong Lu, Guotong Geng |
NeurIPS | 6 |
| 2025 | A Transformer-Enhanced Stochastic Bounded Confidence Model for Opinion Dynamics
Jiao Luo, Yiping Yao, Yu Xuan Peng |
SIGSIM-PADS | 2 |
| 2024 | An Adaptive Approach For Parallel Discrete Event Simulation Thread Pool PredictionabstractThe number of threads in the thread pool is a critical factor that significantly influences the efficiency of parallel execution in Parallel Discrete Event Simulation (PDES). However, current methodologies, including such as static configuration, iterative search, solution based on system state information, and machine learning primarily cater to the domain of parallel computing. These approaches fail to consider PDES-specific characteristics like logical clock synchronization and the interaction among various simulation parameters, thereby posing challenges in accurately predicting the optimal number of threads required for achieving peak performance in PDES. In response to this, this paper proposes an adaptive multivariate power coefficient probability prediction method that considers the interrelated parameters in PDES which collectively influence runtime behavior. The method models various factors affecting PDES efficiency as multivariate power coefficients, incorporating a probability model and bias term to capture the variability of the simulation system. It constructs a nonlinear correlation model between the number of threads and simulation speedup, and derives the number of threads by calculating the optimal speedup of PDES. Experimental results demonstrate that this method achieves an average relative error of 6.4% in predicting the optimal number of threads, with performance improvement achieved by utilizing these predicted threads reaching 93.28% compared to ideal performance improvement relative to serial execution. JiMing Su, Yiping Yao, Feng Zhu 0009 |
ECMS | 2 |
| 2024 | An Agent-based Model of Opinion Dynamics with Hierarchical ThinkingabstractOpinion dynamics studies the principles governing the evolution of collective opinion, offering valuable insights into the comprehension of social phenomena and forecasting group behavior. However, existing opinion dynamics models often overlook the impact of both opinion climate and cognitive capacities on interactive behaviors, thus causing simulation outcomes diverge from real-world observations. Addressing this gap, we propose a novel opinion dynamics model based on hierarchical thinking to describe the opinion evolution on social networks. Individuals are classified into different levels according to their cognitive abilities. They act with bounded rationality at their respective levels to optimize both the promotion of personal opinions and the avoidance of cyberbullying. Through simulation analysis, we found the crucial role of users with high levels of hierarchical thinking. They can discern the opinion climate and articulate their opinion, acting as bridges in the evolution of public opinion. Their opinion can reach the bounded confidence range of more people, thereby enabling polarization to shift to consensus under the same conditions. Furthermore, this effect is independent of individual inherent attributes, which is more in line with real-life scenarios. Lizhen Ou, Yiping Yao, Jiao Luo |
SMC | 2 |
| 2024 | Hierarchical sort-based parallel algorithm for dynamic interest matching
Yiping Yao, Lizhen Ou, Kai Chen 0020 |
J. Parallel Distributed Comput. | 2 |
| 2023 | Opinion Formation Forecasts in Social Networks: A Graph Convolutional Neural Network ApproachabstractOpinion formation in social networks has a significant impact on the society. Consequently, accurate opinion formation forecasts that can be effectively controlled and handled are particularly desirable. Current opinion prediction is primarily based on two methods: a. modeling and simulation and b. machine learning. Although modeling and simulation approaches have the advantage of incorporating a priori knowledge, such as topological information about social networks, existing studies suggest that the predictive accuracy of such models may be lower than expected. Neural network-based methods demonstrate improved performance; however, incorporating a priori knowledge in these models is difficult. To address these limitations, in this study, we develop a graph convolutional neural network-based opinion evolution prediction method, called GCNN. The method follows the opinion dynamics in the model training process to directly involve the topology of the network in opinion prediction. This mechanism shrinks the search space of the neural network parameters and improves the accuracy rate by exploiting the network topology as a priori knowledge. To validate our method, we perform experimental validation on nine synthetic datasets and one real dataset. The results reveal that our method outperforms benchmark models. In addition, our method exhibits superior accuracy rates compared to other candidate models, even when trained on only one-sixth of the training data typically used by other models. Lizhen Ou, Yiping Yao, Haozhe Yuan, Li-li Chen |
DS-RT | 2 |
| 2023 | ContextAD: Context-Aware Acronym Disambiguation with Siamese BERT NetworkabstractAcronym disambiguation is the process of determining the correct expansion of an acronym in given context, which can assist many downstream natural language processing tasks. Typically, existing methods on this task will directly perform semantic comparisons between the candidate expansions and the original sentence, ignoring the relevance of contextual information to expansions. To solve this issue, this paper proposes a context‐aware acronym disambiguation method with Siamese BERT network (ContextAD). First, we combine each candidate expansion with corresponding acronym’s context to form a new sentence set. Then, the new and original sentences are input into a Siamese BERT network that can obtain the semantic similarity. The new sentences and the separate candidate expansions are input into the Siamese BERT network, respectively, along with the original sentences, which can obtain another semantic similarity. Finally, the two different semantic similarities are combined to determine the most suitable expansion. We quantify the improvement of our proposed ContextAD model against a state‐of‐the‐art baseline using the public dataset of the shared tasks of acronym disambiguation (AD) held under AAAI‐2021 workshop on SDU and show that it achieves a better performance based on the same BERT model. Lizhen Ou, Yiping Yao, Xueshan Luo, Xinmeng Li, Kai Chen 0020 |
Int. J. Intell. Syst. | 2 |
| 2023 | Opinion-aware information diffusion model based on multivariate marked Hawkes process
Haoming Zhang 0001, Yiping Yao, Jiefan Zhu |
Knowl. Based Syst. | 2 |
| 2022 | Fixed-wing UAV Kinematics Model using Direction Restriction for Formation Cooperative Flight
Yuxuan Fang, Yiping Yao, Feng Zhu 0009, Kai Chen 0020 |
SIMULTECH | 2 |
| 2022 | An Opinion Dynamics Model with Cross-Link InteractionabstractAt present, online social networks have become the main platform for people to express their opinions and interact with them, which has a great impact on the evolution of opinions. Therefore, the research on the evolution of opinions in online social networks has become a current hotspot. Opinion dynamics is an important tool to study the evolution law of opinions in the network. In the opinion dynamics model for online social networks, agents often can only interact with others who have links on the network. However, in reality, the interaction of agents’ opinions is not limited to agents with links on the network. Therefore, due to the lack of sufficient interaction, the traditional model cannot reflect the true final state of public opinion evolution. In view of the above situation, we propose a novel cross-link interaction mechanism which enables agents interact opinion with others without the limit of network inks and use machine learning methodology to get the cross-link interaction distance. After that, we introduce the mechanism to the bounded confidence opinion dynamics model. With this mechanism, the interaction of the agent will be more reasonable and more like social behavior. The simulation results show that the proposed model fits the real data better than traditional models and even under a very small bounded confidence value, agents opinions will still be around high-influence agents’ opinions. Haoming Zhang 0001, Yiping Yao, Jiefan Zhu |
SMC | 3 |
| 2021 | A Universal Construction to implement Concurrent Data Structure for NUMA-muticoreabstractUniversal constructions are attractive as they can turn a sequential implementation of any data structure into a concurrent implementation. However, existing universal constructions have limitations, such as imposing high copying overhead, or poor scalability on NUMA systems mainly due to their lack of NUMA-aware design principles. To overcome these limitations, this paper introduces CR, a universal construction that provides highly scalable updates on NUMA systems while offering fast read-side performance. CR achieves NUMA-awareness by utilizing delegation within a NUMA node and a global shared log to maintain the consistency of replicas of data structures across nodes. Using CR does not require expertise in concurrent data structure design. Our evaluation shows that CR has up to 11.2 times better performance compared to a state-of-the-art universal construction CX on our tested sequential data structures. To demonstrate the effectiveness and applicability of CR, we have applied CR to an in-memory database system. The database shows up to 18.1 times better performance compared to the original version. Zhengming Yi, Yiping Yao, Kai Chen 0020 |
ICPP | 2 |
| 2021 | A Parallel Hierarchical Sort-based Interest Matching AlgorithmabstractInterest management is a filtering technique to reduce communication in simulation. It involves a process called "interest matching" to identify intersections between two sets of d-dimensional axis-parallel rectangles. Because of frequent demands in simulation execution, interest matching becomes a bottleneck as the problem size grows. However, classical interest matching algorithms, mainly designed for serial processing, do not take advantage of modern multicore processors' computing power. Recent parallel interest matching algorithms can fill the gap, but there is scope for improvement. In this paper, we propose a parallel hierarchical sort-based interest matching algorithm. It embeds subscription regions into an interest management tree and allows update regions compare with nodes of the tree to find results in parallel. The association between adjacent nodes and the hierarchical relation between parent-child nodes can serve to eliminate unnecessary operations. Moreover, we also provide proof to confirm the correctness and a detailed analysis of time-complexity. The experimental results demonstrate that the proposed algorithm can achieve better performance than state-of-art algorithms. Yiping Yao, Feng Zhu 0009, Bin Chen 0003, Wentong Cai 0001 |
SIGSIM-PADS | 2 |
| 2021 | Simulation Runtime Prediction Approach based on Stacking Ensemble Learning
Yiping Yao, Feng Zhu 0009, Kai Chen 0020 |
SIMULTECH | 2 |
| 2021 | A stealing mechanism for delegation methods
Zhengming Yi, Yiping Yao |
J. Supercomput. | 2 |
| 2020 | A barrier optimization framework for NUMA multi-core systemabstractSummary Parallel program performance often critically depends on barrier performance. In modern NUMA multi‐core machines, barrier synchronization performance is significantly affected by cache‐coherence communication between cores, especially when the scale of NUMA systems is large, complex interconnected networks, memory hierarchies, and cache‐coherence protocols make optimization of barrier algorithm hard. We propose a general barrier optimization framework on NUMA multi‐core machines. The framework splits the barrier into three stages: the barrier arrival within a NUMA node, the barrier arrival across the NUMA nodes, and the wakeup, providing an opportunity to optimize the communication pattern and the cache‐line placement in each stage. To reduce remote communication traffic, we introduce a coordinator per NUMA node. In addition, we implement two barrier algorithms based on the framework. Finally, we show the superiority of the barrier algorithms within our framework over other barrier algorithms and show how to translate a barrier algorithm into a performance model to help make an optimal tradeoff design. Experiments were conducted on three NUMA multi‐core platforms and the results show that the barrier algorithm optimized within our framework is sufficient to deliver as good or better performance than state‐of‐art approaches on NUMA multi‐core machines. Zhengming Yi, Yiping Yao |
Concurr. Comput. Pract. Exp. | 3 |
| 2020 | A scalable lock on NUMA multicoreabstractSummary Modern NUMA multicore architectures exhibit complicated memory behavior, such as cache coherence invalidation and nonuniform memory access where the access from a core to its local memory is significantly faster than crossnode access to memory on a different NUMA node. The complicated memory behavior has a large impact on the efficiency of locking synchronization, which affects the performance of parallel applications. Prior works offer several efficient designs to improve locking performance such as delegation schemes. However, the existing delegation schemes either occupy computing cores or provide nonscalable performance, or offer less portability. In this work, we present a NUMA‐aware delegation lock that occupies no cores while offering scalable performance under high contention for NUMA multicore machines. The new lock is a variant of an efficient FFWD lock, and inherits its performance features, such as buffering responses within a NUMA node to minimize cache coherence traffic. Unlike FFWD, the new lock employs hierarchical NUMA‐aware memory allocation and NUMA‐aware dynamic server thread technique, to reduce crossnode communication between client and server threads. Our evaluation shows that the new lock outperforms FFWD under high contention, achieving the significant performance gains when compared with other state‐of‐the‐art locks. Zhengming Yi, Yiping Yao |
Concurr. Comput. Pract. Exp. | 2 |
| 2019 | HMalloc: A Hybrid, Scalable, and Lock-Free Memory Allocator for Multi-Threaded ApplicationsabstractAn efficient multi-threaded memory allocator has great impact on performance of applications with frequent memory allocation and deallocation operations. Currently, most popular memory allocators ignore the difference between thread local and shared memory, and manage them in a unified manner, which cannot make good use of the spatiotemporal locality of memory access. To solve this problem, this paper proposes a hybrid and lock-free multi-threaded memory allocator, named HMalloc. The allocator separates local memory from shared memory. There is no false sharing and lock contentions in local memory allocation and deallocation process. Moreover, coalescence-free is used to optimize this process. Further, a flag-based shared memory allocation and deallocation method is proposed to achieve lock-free shared memory management. Experimental results show that HMalloc can achieve significant performance improvement when compared with existing well-known memory allocators. Tianlin Li, Yiping Yao, Zhongwei Lin |
ICPADS | 2 |
| 2019 | An efficient virtual machine allocation algorithm for parallel and distributed simulation applicationsabstractSummary Allocating appropriate resource for parallel and distributed simulation (PADS) applications in clouds is an intuitive way to improve their execution efficiency. However, the heterogeneity of virtual machine (VMs) in clouds with respect to both their computing power and network latency influences the execution efficiency of PADS applications on different combinations of VMs. Besides, frequent synchronization is one of the characteristics during the execution of PADS applications, which seriously challenges the prediction of the influence of VMs' computing power and network latency on their execution efficiency, and makes allocating appropriate VMs difficult as a result. This paper first proposes a revivification‐based prediction model (ERP), which revives the execution based on statistical data from actual execution of PADS applications to predict the running time of PADS applications on different combinations of VMs. Then, an ERP‐based Allocation algorithm, namely, ERPA, is raised to optimize VMs allocation to minimize the running time of PADS applications in clouds. A series of experiments are conducted to compare the proposed ERPA with three resource allocation algorithms, ie, Gang‐scheduling‐based, Makespan‐based, and Max‐Min‐based algorithms, and the experimental results demonstrate the advantage of ERPA in improving execution efficiency of PADS applications in clouds. In particular, for communication‐sensitive PADS applications, the advantage of ERPA is more significant. Yiping Yao, Huangke Chen, Tianlin Li, Menglong Lin |
Concurr. Comput. Pract. Exp. | 2 |
| 2018 | A Binary Search Enhanced Sort-based Interest Matching AlgorithmabstractIn distributed simulation, communication based on publish/subscribe will generate large amount of irrelevant data transmissions, and thereby degrading the performance. To solve the problem, HLA standard defines data distribution management to filter unnecessary communication. Among several famous interest matching algorithms, the sort-based algorithm has been proven to be the most efficient method in most scenarios. However, the potential of existing sort-based algorithm has not been fully exploited, due to the overhead of sorting the bounds can be further reduced and a portion of unnecessary bit operations can be eliminated. In this paper, we propose a binary search enhanced sort-based interest matching algorithm (BSSIM). Based on a different sufficient and necessary condition to judge interval overlapping, the size of list to be sorted can be remarkably reduced. Moreover, unnecessary bit operations can be eliminated by binary searches. Experimental results show that BSSIM algorithm outperforms the sort-based algorithm, and approximately 64%-159% performance improvement can be achieved at different scenarios. Tianlin Li, Yiping Yao, Feng Zhu 0009 |
SIGSIM-PADS | 3 |
| 2017 | ARM-K: A Methodology for Mining Associations of Traffic Congested LinksabstractTraffic congestion has become a worldwide problem, seriously restricting the performance of transport network. Association of congested links is an essential factor in the formation of traffic congestion, while it lacks enough research. In this paper, a methodology ARM-K combining K-means clustering and association rules mining is proposed to discover the relations of congested links. Specifically, the method digs out the association of two congested links and further extends the 2-tuple relations into a relation graph. As a result, the relation graph and its topological-order profile the associations of congested links. The experiments including congestion prediction and congestion dispersion are devised for the proposed method based on mobile data derived from the traffic simulator VISSIM. The results show relatively high prediction accuracy (above 0.75) and obvious improvement of network performance after dispersion (more than 12.5 percentage-point drop of travel time index), indicating that ARM-K can effectively explores the associations of congested links and provide valuable information for congestion management. Wenyu Xu, Yiping Yao |
MDM | 2 |
| 2016 | Development and Experimentation of PDES-based Analytic SimulationabstractParallel-discrete-event-simulation-based analytic simulation (PAS) is an effective approach to studycomplex issues and analyze complex systems.But the complexity and high demand for credibility of PAS make its development and experimentation quite different from traditional information systems. Firstly, this article briefly introduces analytic simulation concept and the difference with training simulation. And then five computational characteristics which cause the huge computation demand are summarized: multi-sample, multi-entity, as fast as possible, synchronization for constraint of causality and complex model calculation. According to these characteristics, a "Sample, Entity, Model" three-level-Parallelization solution(SEMP) is introduced for PAS.The solution can be used to fully exploit the parallelization of PAS and utilize the computing resources in different levels, which is able to meet the growing computation demand of PAS. Finally, in order to improve the development efficiency and credibility of PAS application, based on the accumulation of several years' R&D, we conclude a summary of development and experimentation flow of PAS, and propose four additional VV&A principles to improve credibility, which can be used to guide the development and experimentation of PDES-based analytic simulation. Yiping Yao, Dong Meng, Qingjun Qu, Zhiwen Jiang |
SIGSIM-PADS | 1 |
| 2015 | Can MIC Find Its Place in the Field of PDES? An Early Performance Evaluation of PDES Simulator on Intel Many Integrated Cores CoprocessorabstractThe widespread utilization of many-core processors offers a good opportunity for Parallel Discrete Events Simulation (PDES) to obtain a better execution performance. As one of the newly introduced many-core processors, the Intel Xeon Phi coprocessor based on Many Integrated Core (MIC) architecture integrates about 60 optimized x86 cores within a PCB board, reaching a peak performance of 1.0 TFLOPS. Furthermore, benefiting from using x86 architecture cores, the MIC coprocessor is fully compatible with almost all programs designed for general purpose CPUs, which makes it easy to run simulation progress on MIC. There have been many works on performance evaluation and optimization of PDES simulator using Graphic Processing Unit (GPU) or Tilera or other many-core processors, yet almost no related works on Phi are published until now. In this article, an early performance evaluation of the well-known PDES simulator ROSS and its POSIX thread version ROSS-MT was conducted based on a computing node composed of two Intel Xeon multi-core CPUs and one Phi coprocessor, using the classical PDES benchmark PHOLD and its extended version by adding different event granularities. Experiment results show that the pure MPI based ROSS performs poorly on MIC coprocessor, indicating that it would not be feasible for common PDES applications. Though ROSS-MT has a much better performance on MIC, the computation potential of MIC is still hardly fully explored. Furthermore, with the event granularity becomes larger, performance of this benchmark exhibits a "fall of cliff", which turns it into a computation dominant application. However, the entire performance on MIC coprocessor is still worse than that on host. After reasoning the problems, we vectorized the code of event handler to better use the Vector Processing Unit (VPU) of MIC coprocessor, which brings us a peak speedup of 9.7X, showing that MIC coprocessor is able to find its place in the PDES field. At last, according to our evaluation work, we provide some advices to further exploit the power of MIC coprocessor for PDES applications. Huilong Chen, Yiping Yao, Dong Meng, Feng Zhu 0009, Yuewen Fu |
DS-RT | 2 |
| 2015 | A WordNet-Based Parameter Configuration Assistance Technology in Simulation ApplicationabstractThe parameter configuration is frequent and necessary in the simulation application development. During the parameter configuration, the candidate parameters are numerous and of various types, selecting the correct candidate parameter by hand is of heavy workload and error-prone, thus hampering the progress of the simulation application development. Aiming at this problem, this paper puts forward a Word Net-based parameter configuration assistance technology, which calculates the similarity of the target parameter and candidate parameters based on Word Net one by one, from three aspects: parameter name, parameter type and parameter description. And then rank the candidate parameters in a descending order according to the similarity value, providing the users for easier selection. The test result shows that this method can effectively avoid frequent selection, thus improving the efficiency of parameters configuration. Yiping Yao |
DS-RT | 2 |
| 2015 | A high performance framework for modeling and simulation of large-scale complex systems
Feng Zhu 0009, Yiping Yao, Dan Chen 0001 |
Future Gener. Comput. Syst. | 2 |
| 2013 | An expansion-aided synchronous conservative time management algorithm on GPUabstractThe graphic processing unit (GPU) brings an opportunity to implement large scale simulations in an economical way. GPU's performance relies on high parallelism, but using synchronous conservative time management algorithm for discrete event simulation will meet the scenarios with limited parallelism. This conflict leads to bad performance even though the application itself has high parallelism. To solve this problem, we propose an expansion-aided synchronous conservative time management algorithm. It uses runtime information to enlarge the time bound of "safe" events, and uses an expansion method to import "safe" events. By interleaving a series of expansions with event computation, more events can be assembled to be processed in parallel. Moreover, a simulated annealing algorithm is adopted to control the number of expansions. It helps achieve stable performance under different conditions by finding a balance between low parallelism and unnecessary expansions. Experiments demonstrate that the proposed algorithm can achieve up to a 30% performance improvement. Yiping Yao, Feng Zhu 0009 |
SIGSIM-PADS | 2 |
| 2011 | HSK: A Hierarchical Parallel Simulation Kernel for Multicore PlatformabstractThe development of CPU has stepped into the era of multi-core. Due to lack of support on thread level, most of the simulation platform can not take full advantage of multicore. To fulfill this gap, we proposed a hierarchical parallel simulation kernel(HSK) model. The model has two layers. The first layer, named process kernel, was responsible for managing all thread kernels on second layer. The second layer is a group of thread kernels, which were responsible for scheduling and advancing logical processes. Each thread kernel was mapped onto an executing thread to advance simulation parallel. In addition, two algorithms were proposed to support high performance: (1) To improve the communication efficiency between threads, we proposed a pointer-based communication mechanism. By using buffers, synchronization between threads can be annihilated. (2) To eliminate redundant Lower Bound on Time Stamp(LBTS) computation and not to interrupt thread execution, we employ an approximate method to compute LBTS asynchronously. A proof of validity was presented. The execution performance of HSK was demonstrated by a series of simulation experiments with a modified phold model. The HSK can achieve good speedup for applications, especially with coarse-grained event. Yiping Yao |
ISPA | 2 |
| 2011 | Accelerating the Requirement Space Exploration through Coarse-Grained Parallel Execution
Zhongwei Lin, Yiping Yao |
NPC | 2 |
| 2010 | CommPar: A Community-Based Model Partitioning Approach for Large-Scale Networked Social Dynamics SimulationabstractEfficient large-scale simulation on multiple processors is essential for social dynamics study but still has been proved to be a challenge. Community structure is a ubiquitous property of social networks. It has significant influence on its dynamics and leads the selection of model partition algorithms a critical performance issue. However, the underlying community structure is not well exploited by existing approaches of load-balancing optimizations, which discounted their effectiveness. This paper proposes COMMPAR, a community-based model partitioning approach, which utilizes the community information of social networks for performance tuning. It contains a two-phased network model partitioning as follows: first, community detection algorithm is employed to discover community structure residing in large-scale social networks, second, those communities are further equally partitioned to achieve an appropriate configuration of simulation execution, and facilitates mapping of the communities onto multiple computer processors. Eventually, the experimental results of a random-walk dynamics simulation show that COMMPAR significantly outperforms several existing partitioning approaches, and can efficiently reduce the overhead of interprocessor communications. Bonan Hou, Yiping Yao |
DS-RT | 2 |
| 2009 | EDEVS : A Scalable DEVS Formalism for Event-Scheduling Based Parallel and Distributed SimulationsabstractScalability is very important for parallel and distributed simulations. Several techniques have been proposed to develop scalable synchronization strategies, communication services or fundamental algorithms, while little has been seen to deal with the modeling stage of the application. Learning from the HPC (High Performance Computing) lesson, it is clear that the time spent in developing a simulation application must be considered in evaluating the scalability of the application. There are many discrete event simulation platforms built for large parallel and distributed simulations, such as SPEEDES (Synchronous Parallel Environment for Emulation and Discrete Event Simulation), GTW (Georgia tech Time Warp), and YHSUPE, etc. They take Event-Scheduling as their modeling paradigm and have achieved great runtime performance, but lack in providing efficient modeling methods. To deal with this issue, a component-based specification, which can support hierarchical decomposition of large models and facilitate model reuse, is presented. This paper extends the DEVS (Discrete Event simulation specification) and proposes a component-based formalism, called EDEVS (Event-Scheduling Discrete Event simulation Specification) for the existing Event-Scheduling parallel and distributed simulation platforms. Yiping Yao, Shaoliang Peng |
DS-RT | 2 |
| 2009 | Human Flesh Search Model Incorporating Network Expansion and GOSSIP with FeedbackabstractWith the development of on-line forum technology and the pervasive participation of the public, the Human Flesh Search is becoming an arising phenomenon which makes a great impact on our daily life. There arose big research interests in social, legal issues resulted from HFS, however, very little work has been conducted to understand how it comes into being and how it dynamically evolves. This paper proposes a modeling and simulation approach incorporating network expansion and GOSSIP propagation with feedback for a better understanding of the human flesh search phenomenon. Based on the acquisition and analysis of the netizens' surfing behavior data, the evolution of the HFS is modeled as a network growth process with proper dynamic input, which is characterized by heavy-tail and burst-oriented distribution, modeling as a Weibulloid process. Then, an improved GOSSIP model with feedback is proposed to represent the information propagation, processing and aggregation during the HFS. New insights for HFS are gained through a set of simulation experiments. Bonan Hou, Yiping Yao, Laibin Yan |
DS-RT | 3 |