VLDB 2026 Research / reviewers in the wild / expert
Ziang Hu
dblp:12/6862
· DBLP profile ↗
19ranked-venue papers
2as first author
5since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 13 · 1 first-authorArtificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Software engineering, systems software and programming languages · 4Security and privacy · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Distributed systems · 42% Memory systems · 32% Parallel and multicore computing · 11% | |
| Databases, data mining, and information retrieval
1 paper |
Graph data management · 100% |
Topics — the 14 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Graph data management › graph processing
asynchronous graph processing |
0.3 | 1 | 2017 | CoRAL: Confined Recovery in Distributed Asynchronous Graph Processing · ASPLOS 2017 |
Graph data management
distributed graph processing |
0.3 | 1 | 2017 | CoRAL: Confined Recovery in Distributed Asynchronous Graph Processing · ASPLOS 2017 |
Distributed systems
fault tolerance |
0.3 | 1 | 2017 | CoRAL: Confined Recovery in Distributed Asynchronous Graph Processing · ASPLOS 2017 |
Distributed systems › fault tolerance
rollback recovery |
0.3 | 1 | 2017 | CoRAL: Confined Recovery in Distributed Asynchronous Graph Processing · ASPLOS 2017 |
Memory systems
cache coherence |
0.2 | 1 | 2014 | PREDATOR: predictive false sharing detection · PPoPP 2014 |
Memory systems › cache coherence
false sharing |
0.2 | 1 | 2014 | PREDATOR: predictive false sharing detection · PPoPP 2014 |
Memory systems › cache coherence
false sharing detection |
0.2 | 1 | 2014 | PREDATOR: predictive false sharing detection · PPoPP 2014 |
Performance modeling and evaluation
performance analysis tools |
0.2 | 1 | 2014 | PREDATOR: predictive false sharing detection · PPoPP 2014 |
Distributed systems › distributed algorithms › distributed snapshot
consistent snapshots |
0.1 | 1 | 2017 | CoRAL: Confined Recovery in Distributed Asynchronous Graph Processing · ASPLOS 2017 |
Distributed systems
distributed coordination |
0.1 | 1 | 2017 | CoRAL: Confined Recovery in Distributed Asynchronous Graph Processing · ASPLOS 2017 |
Parallel and multicore computing › synchronization
fine-grain synchronization |
0.1 | 1 | 2007 | Synchronization state buffer: supporting efficient fine-grain synchronization on many-core architectures · ISCA 2007 |
Processor architecture and microarchitecture
many-core architecture |
0.1 | 1 | 2007 | Synchronization state buffer: supporting efficient fine-grain synchronization on many-core architectures · ISCA 2007 |
Parallel and multicore computing
synchronization |
0.1 | 1 | 2007 | Synchronization state buffer: supporting efficient fine-grain synchronization on many-core architectures · ISCA 2007 |
Parallel and multicore computing › thread-level parallelism
multithreaded applications |
0.1 | 1 | 2014 | PREDATOR: predictive false sharing detection · PPoPP 2014 |
Methods — techniques the papers use, named apart from their topics
lightweight checkpointing · 0.6confined recovery · 0.6predictive modeling · 0.2detailed simulation · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ComSFI: A Community-Aware Approach to Early Rumor Propagation PredictionabstractInformation propagation prediction in social networks remains challenging due to the complex interactions between community structures. Existing models typically overlook multi-peaked cascade patterns that emerge from community-specific dynamics, especially during critical early propagation stages. To address these limitations, we propose ComSFI (Community-aware Susceptible-Forwarding-Immune), a dynamic prediction framework that captures both intra- and inter-community information diffusion. ComSFI makes three key contributions: (1) community-specific diffusion modeling through individualized SFI modules, (2) cross-community propagation modeling with a learnable threshold mechanism that identifies cascade initiation timing, and (3) integration of user interest profiles to estimate propagation willingness across community boundaries. Experiments on Weibo, Douban, and Memetracker datasets demonstrate that ComSFI outperforms state-of-the-art baselines by 3%-12% in cascade size prediction and temporal pattern accuracy, with particular effectiveness in early-stage rumor propagation prediction—a critical application for timely intervention. Our results establish ComSFI as a versatile framework for analyzing and predicting complex information diffusion patterns in real-world networks. Wei Zhou 0019, Ziang Hu, Jizhong Han, Tao Guo 0006 |
ICTAI | 3 |
| 2025 | Multi-Directional Dynamical Framework for Information Propagation PredictionabstractAccurately predicting information propagation in social networks is crucial for applications such as social media analysis and rumor monitoring. Dynamical approaches, valued for their interpretability and ability to capture the underlying mechanisms of propagation, have become a leading research direction. However, most existing methods focus exclusively on forward prediction, overlooking valuable information from other temporal directions. Consequently, these approaches face two critical limitations: (1) inadequacy in handling missing values prevalent in real-world scenarios with sparse and irregular observations; and (2) susceptibility to error accumulation that significantly degrades long-term prediction accuracy. To address these limitations, we propose the Multi-Directional Dynamical Framework (MDDF), which integrates forward, backward, and interpolation dynamics to fully leverage observations from all temporal directions. MDDF formulates the propagation process as a multi-boundary value problem, enabling observations at any time point to constrain predictions across the entire timeline. We further introduce rigorous mathematical formulations for backward and interpolation dynamics in stochastic propagation processes, a unified framework to reconcile multi-directional predictions, and an adaptive ensemble mechanism that dynamically weights each component based on data characteristics. Extensive experiments on real-world datasets demonstrate that MDDF not only surpasses existing methods in prediction accuracy, but also exhibits exceptional robustness to missing data and long-term forecasting. Notably, MDDF achieves up to 27.4% lower MAE than the best baseline under high data sparsity, and reduces long-term prediction errors by up to 25% compared to forward-only approaches. Ziang Hu, Guang Wang 0001, Jizhong Han, Tao Guo 0006 |
TrustCom | 2 |
| 2024 | CSFI for Social Media: Understanding and Predicting Cross-Community Information PropagationabstractSocial platform users are intricately interconnected, forming a complex social system. Personalized recommendations help users access information from diverse sources, fostering community communication. Micro-level dynamic analysis technology offers a scientific approach to measuring and predicting information transmission. While fundamental dissemination principles are studied, cross-community communication scenarios are under-researched. This study presents a community-centric communication dynamics model (CSFI) to explain information transmission within and across communities on social media. Tested on Sina Weibo data, the model improves retweet prediction accuracy by 11.3 % over baselines. Accurately predicting information propagation on social media is crucial for public opinion analysis. This research aids in predicting dissemination trajectories, reflecting public sentiment, and guiding targeted information strategies or interventions. Wei Zhou 0019, Ziang Hu, Jizhong Han, Tao Guo 0006 |
ICTAI | 3 |
| 2023 | EPT: A data-driven transformer model for earthquake prediction
Ziang Hu, Pin Wu, Haiwang Huang, Jiansheng Xiang |
Eng. Appl. Artif. Intell. | 2 |
| 2022 | PEKIN: Prompt-Based External Knowledge Integration Network for Rumor Detection on Social Media
Ziang Hu, Zongzhen Liu |
PRICAI (2) | 1 |
| 2018 | Analysis of classic algorithms on highly-threaded many-core architectures
Lin Ma 0007, Roger D. Chamberlain, Kunal Agrawal 0001, Chen Tian 0002, Ziang Hu |
Future Gener. Comput. Syst. | 5 |
| 2017 | CoRAL: Confined Recovery in Distributed Asynchronous Graph ProcessingabstractExisting distributed asynchronous graph processing systems employ checkpointing to capture globally consistent snapshots and rollback all machines to most recent checkpoint to recover from machine failures. In this paper we argue that recovery in distributed asynchronous graph processing does not require the entire execution state to be rolled back to a globally consistent state due to the relaxed asynchronous execution semantics. We define the properties required in the recovered state for it to be usable for correct asynchronous processing and develop CoRAL, a lightweight checkpointing and recovery algorithm. First, this algorithm carries out confined recovery that only rolls back graph execution states of the failed machines to affect recovery. Second, it relies upon lightweight checkpoints that capture locally consistent snapshots with a reduced peak network bandwidth requirement. Our experiments using real-world graphs show that our technique recovers from failures and finishes processing 1.5x to 3.2x faster compared to the traditional asynchronous checkpointing and recovery mechanism when failures impact 1 to 6 machines of a 16 machine cluster. Moreover, capturing locally consistent snapshots significantly reduces intermittent high peak bandwidth usage required to save the snapshots -- the average reduction in 99th percentile bandwidth ranges from 22% to 51% while 1 to 6 snapshot replicas are being maintained. Keval Vora, Chen Tian 0002, Rajiv Gupta 0001, Ziang Hu |
ASPLOS | 4 |
| 2017 | SceneMan: Bridging mobile apps with system energy manager via scenario notificationabstractPower management on current mobile devices relies on OS modules known as DVFS governors. However, existing governors determine system configuration only based on low-level information such as CPU load without any input about application-level behaviors. In particular, there exists no communication from mobile apps to energy managers. We find that information about app usage scenarios (e.g., gaming, video chatting) can usually help energy manager perform a better job and achieve more energy savings. Although app-level energy optimizations have been proposed, they generally focus on single usage scenarios and do not address optimization across multiple scenarios. In this paper, we propose SceneMan, an energy optimization framework for mobile apps based on usage scenario notification. SceneMan has three components: an API, a scenario notifier, and an energy manager. The key idea is to make energy managers aware of app-level scenarios. At runtime, apps notify the energy manager about their usage scenarios with provided APIs used by developers. The energy manager then takes appropriate actions to minimize energy consumption of the running scenario while meeting performance requirements. Energy optimization across scenarios can thus be easily achieved. The framework requires little extra programming effort and can help apps achieve better energy efficiency in a transparent way. We implement our system on a Nexus 6 smartphone and test it with 13 real-world apps under 2 usage scenarios, namely, gaming and video chatting. We achieve up to 33.2% energy savings with a worst-case performance loss of 5.1%. Li Li 0064, Jun Wang 0077, Handong Ye, Ziang Hu |
ISLPED | 5 |
| 2014 | Code Layout Optimization for Defensiveness and Politeness in Shared CacheabstractCode layout optimization seeks to reorganize the instructions of a program to better utilize the cache. On multicore, parallel executions improve the throughput but may significantly increase the cache contention, because the co-run programs share the cache and in the case of hyper-threading, the instruction cache. In this paper, we extend the reference affinity model for use in whole-program code layout optimization. We also implement the temporal relation graph (TRG) model used in prior work for comparison. For code reorganization, we have developed both function reordering and inter-procedural basic-block reordering. We implement the two models and the two transformations in the LLVM compiler. Experimental results on a set of benchmarks show frequently 20% to 50% reduction in instruction cache misses. By better utilizing the shared cache, the new techniques magnify the throughput improvement of hyper-threading by 8%. Pengcheng Li 0001, Hao Luo 0007, Chen Ding 0001, Ziang Hu, Handong Ye |
ICPP | 4 |
| 2014 | PREDATOR: predictive false sharing detectionabstractFalse sharing is a notorious problem for multithreaded applications that can drastically degrade both performance and scalability. Existing approaches can precisely identify the sources of false sharing, but only report false sharing actually observed during execution; they do not generalize across executions. Because false sharing is extremely sensitive to object layout, these detectors can easily miss false sharing problems that can arise due to slight differences in memory allocation order or object placement decisions by the compiler. In addition, they cannot predict the impact of false sharing on hardware with different cache line sizes. Tongping Liu, Chen Tian 0002, Ziang Hu, Emery D. Berger |
PPoPP | 3 |
| 2010 | Efficient compilation of fine-grained SPMD-threaded programs for multicore CPUsabstractIn this paper we describe techniques for compiling fine-grained SPMD-threaded programs, expressed in programming models such as OpenCL or CUDA, to multicore execution platforms. Programs developed for manycore processors typically express finer thread-level parallelism than is appropriate for multicore platforms. We describe options for implementing fine-grained threading in software, and find that reasonable restrictions on the synchronization model enable significant optimizations and performance improvements over a baseline approach. We evaluate these techniques in a production-level compiler and runtime for the CUDA programming model targeting modern CPUs. Applications tested with our tool often showed performance parity with the compiled C version of the application for single-thread performance. With modest coarse-grained multithreading typical of today's CPU architectures, an average of 3.4x speedup on 4 processors was observed across the test applications. John A. Stratton, Vinod Grover, Jaydeep Marathe, Bastiaan Aarts, Mike Murphy, Ziang Hu, Wen-Mei W. Hwu |
CGO | 6 |
| 2007 | Optimizing the Fast Fourier Transform on a Multi-core ArchitectureabstractThe rapid revolution in microprocessor chip architecture due to multicore technology is presenting unprecedented challenges to the application developers as well as system software designers: how to best exploit the parallelism potential due to such multi-core architectures? In this paper, we report an in-depth study on such challenges based on our experience of optimizing the fast Fourier transform (FFT) on the IBM Cyclops-64 chip architecture - a large-scale multi-core chip architecture consisting 160 thread units, associated memory banks and an interconnection network that connect them together in a shared memory organization. We demonstrate how multi-core architectures like the C64 could be used to achieve a high performance implementation of FFT both in 1D and 2D cases. We analyze the optimization challenges and opportunities including problem decomposition, load balancing, work distribution, and data-reuse, together with the exploiting of the C64 architecture features such as the multi-level of memory hierarchy and large register files. Furthermore, the experience learned during the hand-tuned optimization process have provided valuable guidance in our compiler optimization design and implementation. The main contributions of this paper include: 1) our study demonstrates that successful optimization for C64-like large-scale multi-core architectures requires a careful analysis that can identify certain domain-specific features of a target application (e.g. FFT) and match them well with some key multi-core architecture features; 2) our optimization, assisted with hand-tuned process, provided quantitative evidence on the importance of each optimization identified in 1); 3) automatic optimization by our compiler, the design and implementation of which is guided by the feedbacks from 1) and 2), shows excellent results that are often comparable to the results derived from our time-consuming hand-tuned code. Long Chen 0020, Ziang Hu, Junmin Lin, Guang R. Gao |
IPDPS | 2 |
| 2007 | Exploring a Multithreaded Methodology to Implement a Network Communication Protocol on the Cyclops-64 Multithreaded ArchitectureabstractThe IBM Cyclops-64 (C64) chip employs a multithreaded architecture that integrates a large number of hardware thread units on a single chip. A cellular supercomputer is being developed based on a 3D-mesh connection of the C64 chips. This paper introduces the Cyclops datagram protocol (CDP) developed for the C64 supercomputer system. CDP is inspired by the TCP/IP protocol, yet simpler and more compact. The implementation of CDP leverages the abundant hardware thread-level parallelism provided by the C64 multithreaded architecture. The main contributions of this paper are: (1) We have completed a design and implementation of CDP that is used as the fundamental communication infrastructure for the C64 supercomputer system. (2) CDP successfully exploits the massive thread-level parallelism provided on the C64 hardware, achieving good performance scalability; (3) CDP is quite efficient. Its peak throughput reaches 884Mbps on the gigabit Ethernet, even it is running at the user-level on a single-processor Linux machine; (4) Extensive application test cases are passed and no reliability problems have been reported. Ge Gan, Ziang Hu, Juan del Cuvillo, Guang R. Gao |
IPDPS | 2 |
| 2007 | On the Role of Deterministic Fine-Grain Data Synchronization for Scientific Applications: A Revisit in the Emerging Many-Core EraabstractThe design of microprocessor chip for high-end computing systems is moving towards many-core architectures with 10s or 100+ processing units. An important class of the target applications for such architectures are scientific numerical computations, many of which are intrinsically deterministic - that is for a given input a fixed output (result) should be produced no matter how the program is parallelized. It is critical that the read-after-write data dependencies in such programs should be implemented correctly and efficiently via fine-grain data synchronization. In this paper, we investigate the parallelization of three representative scientific computation kernels using fine-grain data synchronization supported by an recently proposed architectural mechanism for many-core chips, called synchronization state buffer (SSB). Using detailed simulation on a simulator for the IBM 160-core Cyclops-64 chip architecture with the SSB extension, our experiments demonstrate significant performance advantage of using fine-grain data synchronization based parallelization schemes for scientific workloads. Weirong Zhu, Ziang Hu, Guang R. Gao |
IPDPS | 2 |
| 2007 | On the Role of Deterministic Fine-Grain Data Synchronization for Scientific Applications: A Revisit in the Emerging Many-Core EraabstractThe design of microprocessor chip for high-end computing systems is moving towards many-core architectures with 10s or 100+ processing units. An important class of the target applications for such architectures are scientific numerical computations, many of which are intrinsically deterministic - that is for a given input a fixed output (result) should be produced no matter how the program is parallelized. It is critical that the read-after-write data dependencies in such programs should be implemented correctly and efficiently via fine-grain data synchronization. In this paper, we investigate the parallelization of three representative scientific computation kernels using fine-grain data synchronization supported by an recently proposed architectural mechanism for many-core chips, called synchronization state buffer (SSB). Using detailed simulation on a simulator for the IBM 160-core Cyclops-64 chip architecture with the SSB extension, our experiments demonstrate significant performance advantage of using fine-grain data synchronization based parallelization schemes for scientific workloads. Weirong Zhu, Ziang Hu, Guang R. Gao |
IPDPS | 2 |
| 2007 | Synchronization state buffer: supporting efficient fine-grain synchronization on many-core architecturesabstractEfficient fine-grain synchronization is extremely important to effectively harness the computational power of many-core architectures. However, designing and implementing finegrain synchronization in such architectures presents several challenges, including issues of synchronization induced overhead, storage cost, scalability, and the level of granularity to which synchronization is applicable. This paper proposes the Synchronization State Buffer (SSB), a scalable architectural design for fine-grain synchronization that efficiently performs synchronizations between concurrent threads. The design of SSB is motivated by the following observation: at any instance during the parallel execution only a small fraction of memory locations are actively participating in synchronization. Based on this observation we present a fine-grain synchronization design that records and manages the states of frequently synchronized data using modest hardware support. We have implemented the SSB design in the context of the 160-core IBM Cyclops-64 architecture. Using detailed simulation, we present our experience for a set of benchmarks with different workload characteristics. Weirong Zhu, Vugranam C. Sreedhar, Ziang Hu, Guang R. Gao |
ISCA | 3 |
| 2006 | Optimization of Dense Matrix Multiplication on IBM Cyclops-64: Challenges and Experiences
Ziang Hu, Juan del Cuvillo, Weirong Zhu, Guang R. Gao |
Euro-Par | 1 |
| 2005 | Performance Modelling and Optimization of Memory Access on Cellular Computer Architecture Cyclops64
Yanwei Niu, Ziang Hu, Kenneth E. Barner, Guang R. Gao |
NPC | 2 |
| 2005 | Madd Operation Aware Redundancy EliminationabstractOn general purpose computer architectures, the optimization of redundancy elimination almost always improves the cycle count. We argue that a specific consideration should be taken when applying this optimization to embedded architectures that feature multiply-add(MADD) instruction. This paper presents a redundancy elimination algorithm with MADD operation aware consideration. It produces optimized results for both code size and cycle count. The algorithm is integrated into KylinC compiler, a compiler for embedded systems developed at the University of Delaware. Experimental results demonstrate that the cycle counts of the benchmark programs are reduced on average 8% and the code sizes are reduced on average 5.27%. Haiping Wu, Ziang Hu, Joseph B. Manzano, Guang R. Gao |
Int. J. Softw. Eng. Knowl. Eng. | 2 |