Haichuan Wang

dblp:49/1599 · DBLP profile ↗
← Back
14ranked-venue papers
5as first author
6since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 2 first-author · 5 since 2021Systems, architecture and hardware · 3 · 2 first-authorSoftware engineering, systems software and programming languages · 3 · 2 first-authorComputer networks · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Generative modeling · 52% Reinforcement learning · 48%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Environmental and earth informatics · 44% Smart cities and intelligent transportation · 44% Computational social science and digital humanities · 13%
Theoretical computer science
2 papers
Algorithmic game theory and mechanism design · 66% Mathematical optimization · 34%
Computer architecture, parallel and distributed computing, and storage systems
3 papers
Parallel and multicore computing · 55% Cloud and datacenter computing · 20% Memory systems · 11%
Computer networks
1 paper
Edge and fog computing · 100%
Software engineering, system software, and programming languages
1 paper
Compilers and program optimization · 44% Runtime systems and virtual machines · 44% Programming languages and type systems · 13%

Topics — the 21 heaviest of 22, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
flow matching
1.922026
Generative AI Against Poaching: Latent Composite Flow Matching for Poaching Prediction · AAAI 2026
Composite Flow Matching for Reinforcement Learning with Shifted-Dynamics Data · NeurIPS 2025
Environmental and earth informatics › conservation
poaching prediction
1.012026
Generative AI Against Poaching: Latent Composite Flow Matching for Poaching Prediction · AAAI 2026
Smart cities and intelligent transportation
spatio-temporal prediction
1.012026
Generative AI Against Poaching: Latent Composite Flow Matching for Poaching Prediction · AAAI 2026
Algorithmic game theory and mechanism design
equilibrium analysis
1.012026
The Publication Choice Problem · AAAI 2026
Machine learning › Reinforcement learning › non-stationary reinforcement learning
dynamics shift
0.912025
Composite Flow Matching for Reinforcement Learning with Shifted-Dynamics Data · NeurIPS 2025
Machine learning › Reinforcement learning
offline reinforcement learning
0.912025
Composite Flow Matching for Reinforcement Learning with Shifted-Dynamics Data · NeurIPS 2025
Edge and fog computing
distributed learning
0.612022
Joint Data Collection and Resource Allocation for Distributed Machine Learning at the Edge · IEEE Trans. Mob. Comput. 2022
Computational social science and digital humanities
science of science
0.312026
The Publication Choice Problem · AAAI 2026
Mathematical optimization
optimal transport
0.312025
Composite Flow Matching for Reinforcement Learning with Shifted-Dynamics Data · NeurIPS 2025
Mathematical optimization › optimal transport
wasserstein distance
0.312025
Composite Flow Matching for Reinforcement Learning with Shifted-Dynamics Data · NeurIPS 2025
Runtime systems and virtual machines › interpreter
interpreter optimization
0.212015
Vectorization of apply to reduce interpretation overhead of R · OOPSLA 2015
Compilers and program optimization
vectorization
0.212015
Vectorization of apply to reduce interpretation overhead of R · OOPSLA 2015
Cloud and datacenter computing › resource management
resource management and scheduling
0.212022
Joint Data Collection and Resource Allocation for Distributed Machine Learning at the Edge · IEEE Trans. Mob. Comput. 2022
Parallel and multicore computing › parallel programming runtimes
runtime systems and scheduling
0.112012
A work-stealing scheduler for X10's task parallelism with suspension · PPoPP 2012
Parallel and multicore computing › parallel programming models
task parallelism
0.112012
A work-stealing scheduler for X10's task parallelism with suspension · PPoPP 2012
Parallel and multicore computing › task scheduling › dynamic scheduling
work-stealing scheduler
0.112012
A work-stealing scheduler for X10's task parallelism with suspension · PPoPP 2012
Memory systems
memory wall
0.112009
Allocation wall: a limiting factor of Java applications on emerging multi-core platforms · OOPSLA 2009
Performance modeling and evaluation
workload characterization
0.112009
Allocation wall: a limiting factor of Java applications on emerging multi-core platforms · OOPSLA 2009
Programming languages and type systems
dynamic languages
0.112015
Vectorization of apply to reduce interpretation overhead of R · OOPSLA 2015
Parallel and multicore computing › parallel computing
parallel programming languages
0.012012
A work-stealing scheduler for X10's task parallelism with suspension · PPoPP 2012
Processor architecture and microarchitecture
chip multiprocessor
0.012009
Allocation wall: a limiting factor of Java applications on emerging multi-core platforms · OOPSLA 2009

Methods — techniques the papers use, named apart from their topics

occupancy-based detection model · 2.0game theory · 2.0equilibrium analysis · 2.0composite flow · 2.0wasserstein distance · 1.7optimal transport · 1.7flow matching · 1.7active data collection · 1.7mixed-integer nonlinear programming · 1.1approximation algorithm · 1.1function vectorization · 0.2data transformation · 0.2work stealing · 0.1conditional atomic blocks · 0.1performance analysis · 0.1
YearPublicationVenuePosition
2026 Generative AI Against Poaching: Latent Composite Flow Matching for Poaching Prediction
abstract
Poaching poses significant threats to biodiversity. A valuable step in reducing poaching is to forecast poacher behavior, which can inform patrol deployment and other conservation interventions. Existing poaching prediction methods based on linear models or decision trees lack the expressivity to capture complex, nonlinear spatiotemporal patterns. Recent advances in generative modeling, particularly flow matching, offer a more flexible alternative. However, training such models on real-world poaching data faces two central obstacles: imperfect detection of poaching events and limited data. To address imperfect detection, we integrate flow matching with an occupancy-based detection model and train the flow in latent space to infer the underlying occupancy state. To mitigate data scarcity, we adopt a composite flow initialized from a linear-model prediction rather than random noise which is the standard in diffusion models, injecting prior knowledge and improving generalization. Evaluations on datasets from two national parks in Uganda show consistent gains in predictive accuracy.
Haichuan Wang, Charles A. Emogor, Vincent Börsch-Supan, Lily Xu, Milind Tambe
AAAI2
2026 The Publication Choice Problem
abstract
Researchers strategically choose where to submit their work in order to maximize its impact, and these publication decisions in turn determine venues' impact factors. To analyze how individual publication choices both respond to and shape venue impact, we introduce a game-theoretic framework - coined the Publication Choice Problem - that captures this two‐way interplay. We show the existence of a pure-strategy equilibrium in the Publication Choice Problem and its uniqueness under binary researcher types. Our characterizations of the equilibrium properties offer insights about what publication behaviors better indicate a researcher's impact level. Through equilibrium analysis, we further investigate how labeling papers with ``spotlight'' affects the impact factor of venues in the research community. Our analysis shows that competitive venue labeling top papers with ``spotlight'' may decrease the overall impact of other venues in the community, while less competitive venues with ``spotlight'' labeling have an opposite impact.
Haichuan Wang
AAAI1
2025 Finite-Horizon Single-Pull Restless Bandits: An Efficient Index Policy For Scarce Resource Allocation
Guojun Xiong, Haichuan Wang, Yuqi Pan, Saptarshi Mandal, Sanket Shah, Niclas Boehmer, Milind Tambe
AAMAS2
2025 Composite Flow Matching for Reinforcement Learning with Shifted-Dynamics Data
abstract
Incorporating pre-collected offline data can substantially improve the sample efficiency of reinforcement learning (RL), but its benefits can break down when the transition dynamics in the offline dataset differ from those encountered online. Existing approaches typically mitigate this issue by penalizing or filtering offline transitions in regions with large dynamics gap. However, their dynamics-gap estimators often rely on KL divergence or mutual information, which can be ill-defined when offline and online dynamics have mismatched support. To address this challenge, we propose CompFlow, a principled framework built on the theoretical connection between flow matching and optimal transport. Specifically, we model the online dynamics as a conditional flow built upon the output distribution of a pretrained offline flow, rather than learning it directly from a Gaussian prior. This composite structure provides two advantages: (1) improved generalization when learning online dynamics under limited interaction data, and (2) a well-defined and stable estimate of the dynamics gap via the Wasserstein distance between offline and online transitions. Building on this dynamics-gap estimator, we further develop an optimistic active data collection strategy that prioritizes exploration in high-gap regions, and show theoretically that it reduces the performance gap to the optimal policy. Empirically, CompFlow consistently outperforms strong baselines across a range of RL benchmarks with shifted-dynamics data.
Haichuan Wang, Tonghan Wang 0001, Guojun Xiong, Milind Tambe
NeurIPS2
2025 Robust Optimization with Diffusion Models for Green Security
abstract
In green security, defenders must forecast adversarial behavior-such as poaching, illegal logging, and illegal fishing-to plan effective patrols. These behavior are often highly uncertain and complex. Prior work has leveraged game theory to design robust patrol strategies to handle uncertainty, but existing adversarial behavior models primarily rely on Gaussian processes or linear models, which lack the expressiveness needed to capture intricate behavioral patterns. To address this limitation, we propose a conditional diffusion model for adversary behavior modeling, leveraging its strong distribution-fitting capabilities. To the best of our knowledge, this is the first application of diffusion models in the green security domain. Integrating diffusion models into game-theoretic optimization, however, presents new challenges, including a constrained mixed strategy space and the need to sample from an unnormalized distribution to estimate utilities. To tackle these challenges, we introduce a mixed strategy of mixed strategies and employ a twisted Sequential Monte Carlo (SMC) sampler for accurate sampling. Theoretically, our algorithm is guaranteed to converge to an \(\epsilon\)-equilibrium with high probability using a finite number of iterations and samples. Empirically, we evaluate our approach on both synthetic and real-world poaching datasets, demonstrating its effectiveness.
Haichuan Wang, Yuqi Pan, Cheol Woo Kim, Mingxiao Song, Alayna Nguyen, Tonghan Wang 0001, Milind Tambe
UAI2
2022 Joint Data Collection and Resource Allocation for Distributed Machine Learning at the Edge
abstract
Under the paradigm of edge computing, the enormous data generated at the network edge can be processed locally. To make full utilization of these widely distributed data, we focus on an edge computing system that conducts distributed machine learning using gradient-descent based approaches. To ensure the system’s performance, there are two major challenges: how to collect data from multiple data source nodes for training jobs and how to allocate the limited resources on each edge server among these jobs. In this paper, we jointly consider the two challenges for distributed training (without service requirement), aiming to maximize the system throughput while ensuring the system’s quality of service (QoS). Specifically, we formulate the joint problem as a mixed-integer non-linear program, which is NP-hard, and propose an efficient approximation algorithm. Furthermore, we take service placement into consideration for diverse training jobs and propose an approximation algorithm. We also analyze that our proposed algorithm can achieve the constant bipartite approximation under many practical situations. We build a test-bed to evaluate the effectiveness of our proposed algorithm in a practical scenario. Extensive simulation results and testing results show that the proposed algorithms can improve the system throughput 56-69 percent compared with the conventional algorithms.
Min Chen 0033, Haichuan Wang, Zeyu Meng, Hongli Xu 0001, Yang Xu 0020, Jianchun Liu, He Huang 0001
IEEE Trans. Mob. Comput.2
2020 GDS: General Distributed Strategy for Functional Dependency Discovery Algorithms
Peizhong Wu, Wei Yang 0011, Haichuan Wang, Liusheng Huang
DASFAA (1)3
2015 Vectorization of apply to reduce interpretation overhead of R
abstract
R is a popular dynamic language designed for statistical computing. Despite R's huge user base, the inefficiency in R's language implementation becomes a major pain-point in everyday use as well as an obstacle to apply R to solve large scale analytics problems. The two most common approaches to improve the performance of dynamic languages are: implementing more efficient interpretation strategies and extending the interpreter with Just-In-Time (JIT) compiler. However, both approaches require significant changes to the interpreter, and complicate the adoption by development teams as a result. This paper presents a new approach to improve execution efficiency of R programs by vectorizing the widely used Apply class of operations. Apply accepts two parameters: a function and a collection of input data elements. The standard implementation of Apply iteratively invokes the input function with each element in the data collection. Our approach combines data transformation and function vectorization to convert the looping-over-data execution of the standard Apply into a single invocation of a vectorized function that contains a sequence of vector operations over the input data. This conversion can significantly speed-up the execution of Apply operations in R by reducing the number of interpretation steps. We implemented the vectorization transformation as an R package. To enable the optimization, all that is needed is to invoke the package, and the user can use a normal R interpreter without any changes. The evaluation shows that the proposed method delivers significant performance improvements for a collection of data analysis algorithm benchmarks. This is achieved without any native code generation and using only a single-thread of execution.
Haichuan Wang, David A. Padua, Peng Wu 0001
OOPSLA1
2014 Optimizing R VM: Allocation Removal and Path Length Reduction via Interpreter-level Specialization
Haichuan Wang, Peng Wu 0001, David A. Padua
CGO1
2012 A work-stealing scheduler for X10's task parallelism with suspension
abstract
The X10 programming language is intended to ease the programming of scalable concurrent and distributed applications. X10 augments a familiar imperative object-oriented programming model with constructs to support light-weight asynchronous tasks as well as execution across multiple address spaces. A crucial aspect of X10's runtime system is the scheduling of concurrent tasks. Work-stealing schedulers have been shown to efficiently load balance fine-grain divide-and-conquer task-parallel program on SMPs and multicores. But X10 is not limited to shared-memory fork-join parallelism. X10 permits tasks to suspend and synchronize by means of conditional atomic blocks and remote task invocations.
Olivier Tardieu, Haichuan Wang
PPoPP2
2010 Using the middle tier to understand cross-tier delay in a multi-tier application
abstract
Understanding the cause of poor performance in a multi-tier enterprise application is challenging, because a performance bottleneck on any tier may cause the whole system to be under utilized, and to fail its throughput or quality of service goals. This paper presents an approach that focuses on the application server to identify bottlenecks in a multi-tier application that are caused by tiers other then the application server. The approach uses a performance tool, named SLICE, that selectively tracks method invocations that cross tier boundaries, and extracts contextual information associated with these invocations. SLICE also collects information from the operating system's scheduler to determine when a thread is blocked. Using the contextual information from method invocations and the information of when a thread is blocked from the operating system, SLICE computes cross tier delay. Experiments on DayTrader, a multi-tier application, show that performance bottlenecks caused by clients or database servers can be identified using cross tier delay.
Haichuan Wang, Qiming Teng, Xiao Zhong, Peter F. Sweeney
IPDPS1
2009 Tale in the Multi-Core Era: Is Java Still Competitive to Host SIP Applications?
abstract
Multi-core platforms are becoming prevailing in telecom infrastructures, and many SIP(Session Initiation Protocol) enabled applications are using Java as the development language and runtime environment. It is important to understand the workload characteristics and performance issues of Java based SIP stack over multi-core platforms. In this paper, two RFC 3261 compliant Java based SIP stacks, the JAIN-SIP and a proprietary SIP stack are studied in depth. Especially we focused on the typical performance issues caused by either Java language features or the workload of SIP protocol, including scalability of SIP stack, resource contention issue and GC impact with SIP memory usage pattern. Some optimization techniques and their drawbacks are also discussed along with the performance evaluation. It turns out that due to the complex combination of SIP protocol semantics and Java language features, the performance of Java based SIP stack tends to be far from competitive on multicore platforms. Extensive optimization and certain improvements of program structure and object management policy may help, however, may, on the other hand, sacrifice some key values of the java language, e.g. easy development.
Kai Zheng 0003, Haichuan Wang, Zhiguo Gao
ICC3
2009 Allocation wall: a limiting factor of Java applications on emerging multi-core platforms
abstract
Multi-core processors are widely used in computer systems. As the performance of microprocessors greatly exceeds that of memory, the memory wall becomes a limiting factor. It is important to understand how the large disparity of speed between processor and memory influences the performance and scalability of Java applications on emerging multi-core platforms.
Kai Zheng 0003, Haichuan Wang, Ling Shao 0002
OOPSLA4
2004 Linear generalization probe samples for face recognition
Haichuan Wang, Liming Zhang 0001
Pattern Recognit. Lett.1