VLDB 2026 Research / reviewers in the wild / expert
Jiachuan Wang
dblp:00/3645
· DBLP profile ↗
14ranked-venue papers
6as first author
10since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 3 first-author · 7 since 2021Databases, data management, data science and information retrieval · 9 · 4 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Efficient Zero-Shot and Label-free Log Anomaly Detection for Resource-Constrained Systems
Zuohan Wu, Jiachuan Wang, Libin Zheng 0001, Shuangyin Li |
ICDE | 2 |
| 2026 | Proficient Graph Neural Network Design by Accumulating Knowledge on Large Language ModelsabstractHigh-level automation is increasingly critical in AI, driven by rapid advances in large language models (LLMs) and AI agents. However, LLMs, despite their general reasoning power, struggle significantly in specialized, data-sensitive tasks such as designing Graph Neural Networks (GNNs). This difficulty arises from (1) the inherent knowledge gaps in modeling the intricate, varying relationships between graph properties and suitable architectures and (2) the external noise from misleading descriptive inputs, often resulting in generic or even misleading model suggestions. Achieving proficiency in designing data-aware models—defined as the meta-level capability to systematically accumulate, interpret, and apply data-specific design knowledge—remains challenging for existing automated approaches, due to their inefficient construction and application of meta-knowledge. To achieve meta-level proficiency, we propose DesiGNN, a knowledge-centered framework that systematically converts past model design experience into structured, fine-grained knowledge priors well-suited for meta-learning with LLMs. To account for the inherent variability and external noise, DesiGNN aligns empirical property filtering from extensive benchmarks with adaptive elicitation of literature insights via LLMs. By constructing a solid meta-knowledge between unseen graph understanding and known effective architecture patterns, DesiGNN can deliver top-5.77% initial model proposals for unseen datasets within seconds and achieve consistently superior performance with minimal search cost compared to baselines. Hanmo Liu, Shimin Di, Jiachuan Wang, Lei Chen 0002, Xiaofang Zhou 0001 |
WSDM | 5 |
| 2025 | Structuring Benchmark into Knowledge Graphs to Assist Large Language Models in Retrieving and Designing ModelsabstractIn recent years, the design and transfer of neural network models have been widely studied due to their exceptional performance and capabilities. However, the complex nature of datasets and the vast architecture space pose significant challenges for both manual and automated algorithms in creating high-performance models. Inspired by researchers who design, train, and document the performance of various models across different datasets, this paper introduces a novel schema that transforms the benchmark data into a Knowledge Benchmark Graph (KBG), which primarily stores the facts in the form of performance(data, model). Constructing the KBG facilitates the structured storage of design knowledge, aiding subsequent model design and transfer. However, it is a non-trivial task to retrieve or design suitable neural networks based on the KBG, as real-world data are often off the records. To tackle this challenge, we propose transferring existing models stored in KBG by establishing correlations between unseen and previously seen datasets. Given that measuring dataset similarity is a complex and open-ended issue, we explore the potential for evaluating the correctness of the similarity function. Then, we further integrate the KBG with Large Language Models (LLMs), assisting LLMs to think and retrieve existing model knowledge in a manner akin to humans when designing or transferring models. We demonstrate our method specifically in the context of Graph Neural Network (GNN) architecture design, constructing a KBG (with 26,206 models, 211,669 performance records, and 2,540,064 facts) and validating the effectiveness of leveraging the KBG to promote GNN architecture design. Hanmo Liu, Shimin Di, Jiachuan Wang, Xiaofang Zhou 0001, Lei Chen 0002 |
ICLR | 5 |
| 2025 | Adapting Pretrained Language Models for Citation Classification via Self-Supervised Contrastive LearningabstractCitation classification, which identifies the intention behind academic citations, is pivotal for scholarly analysis. Previous works suggest fine-tuning pretrained language models (PLMs) on citation classification datasets, reaping the reward of the linguistic knowledge they gained during pretraining. However, directly fine-tuning for citation classification is challenging due to labeled data scarcity, contextual noise, and spurious keyphrase correlations. In this paper, we present a novel framework, Citss, that adapts the PLMs to overcome these challenges. Citss introduces self-supervised contrastive learning to alleviate data scarcity, and is equipped with two specialized strategies to obtain the contrastive pairs: sentence-level cropping, which enhances focus on target citations within long contexts, and keyphrase perturbation, which mitigates reliance on specific keyphrases. Compared with previous works that are only designed for encoder-based PLMs, Citss is carefully developed to be compatible with both encoder-based PLMs and decoder-based LLMs, to embrace the benefits of enlarged pretraining. Experiments with three benchmark datasets with both encoder-based PLMs and decoder-based LLMs demonstrate our superiority compared to the previous state of the art. Our code is available at: github.com/LITONG99/Citss Tong Li 0017, Jiachuan Wang, Shuangyin Li, Lei Chen 0002 |
KDD (2) | 2 |
| 2024 | Cross-Domain-Aware Worker Selection with Training for Crowdsourced AnnotationabstractAnnotation through crowdsourcing draws incremental attention, which relies on an effective selection scheme given a pool of workers. Existing methods propose to select workers based on their performance on tasks with ground truth, while two important points are missed. 1) The historical performances of workers in other tasks. In real-world scenarios, workers need to solve a new task whose correlation with previous tasks is not well-known before the training, which is called cross-domain. 2) The dynamic worker performance as workers will learn from the ground truth. In this paper, we consider both factors in designing an allocation scheme named cross-domain-aware worker selection with training approach. Our approach proposes two estimation modules to both statistically analyze the cross-domain correlation and simulate the learning gain of workers dynamically. A framework with a theoretical analysis of the worker elimination process is given. To validate the effectiveness of our methods, we collect two novel real-world datasets and generate synthetic datasets. The experiment results show that our method outperforms the baselines on both real-world and synthetic datasets. Yushi Sun, Jiachuan Wang, Peng Cheng 0003, Libin Zheng 0001, Lei Chen 0002, Jian Yin 0001 |
ICDE | 2 |
| 2024 | Learning from Emergence: A Study on Proactively Inhibiting the Monosemantic Neurons of Artificial Neural NetworksabstractRecently, emergence has received widespread attention from the research community along with the success of large-scale models.Different from the literature, we hypothesize a key factor that promotes the performance during the increase of scale: the reduction of monosemantic neurons that can only form one-to-one correlations with specific features.Monosemantic neurons tend to be sparser and have negative impacts on the performance in large models.Inspired by this insight, we propose an intuitive idea to identify monosemantic neurons and inhibit them.However, achieving this goal is a non-trivial task as there is no unified quantitative evaluation metric and simply banning monosemantic neurons does not promote polysemanticity in neural networks.Therefore, we first propose a new metric to measure the monosemanticity of neurons with the guarantee of efficiency for online computation, then introduce a theoretically supported method to suppress monosemantic neurons and proactively promote the ratios of polysemantic neurons in training neural networks.We validate our conjecture that monosemanticity brings about performance change at different model scales on a variety of neural networks and benchmark datasets in different areas, including language, image, and physics simulation tasks.Further experiments validate our analysis and theory regarding the inhibition of monosemanticity. Jiachuan Wang, Shimin Di, Lei Chen 0002, Charles Wang Wai Ng |
KDD | 1 |
| 2023 | Noise2Info: Noisy Image to Information of Noise for Self-Supervised Image DenoisingabstractUnsupervised image denoising has been proposed to alleviate the widespread noise problem without requiring clean images. Existing works mainly follow the self-supervised way, which tries to reconstruct each pixel x of noisy images without the knowledge of x. More recently, some pioneer works further emphasize the importance of x and propose to weigh the information extracted from x and other pixels when recovering x. However, such a method is highly sensitive to the standard deviation σnof noise injected to clean images, where σnis inaccessible without knowing clean images. Thus, it is unrealistic to assume that σnis known for pursuing high model performance.To alleviate this issue, we propose Noise2Info to extract the critical information, the standard deviation σnof injected noise, only based on the noisy images. Specifically, we first theoretically provide an upper bound on σn, while the bound requires clean images. Then, we propose a novel method to estimate the bound of σnby only using noisy images. Besides, we prove that the difference between our estimation with the true deviation goes smaller as the model training. Empirical studies show that Noise2Info is effective and robust on benchmark data sets and closely estimates the standard deviation of noise during model training. Jiachuan Wang, Shimin Di, Lei Chen 0002, Charles Wang Wai Ng |
ICCV | 1 |
| 2022 | A New Class of Polynomial Activation Functions of Deep Learning for Precipitation ForecastingabstractPrecipitation forecasting, modeled as an important chaotic system in earth system science, is not explicitly solved with theory-driven models. In recent years, deep learning models have achieved great success in various applications including rainfall prediction. However, these models work in an image processing manner regardless of the nature of a physical system. We found that the non-linearity relationships learned by deep learning models, which mostly rely on the activation functions, are commonly weighted piecewise continuous functions with bounded first-order derivatives. In contrast, the polynomial is one of the most widely used classes of functions for theory-driven models, applied to numerical approximation, dynamic system modeling, etc.. Researchers started to use the polynomial activation functions (Pacs in short) for neural networks from the 1990s. In recent years, with bloomed researches that apply deep learning to scientific problems, it is weird that such a powerful class of basis functions is rarely used. In this paper, we investigate it and argue that, even though polynomials are good at information extraction, it is too fragile to train stably. We finally solve its serious data flow explosion problem with Chebyshev polynomials and prepended normalization, which enables networks to go deep with Pacs. To enhance the robustness of training, a normalization called Range Norm is further proposed. Performance on synthetic dataset and summer precipitation prediction task validates the necessity of such a class of activation functions to simulate complex physical mechanisms. The new tool for deep learning enlightens a new way of automatic theoretical physics analysis. Jiachuan Wang, Lei Chen 0002, Charles Wang Wai Ng |
WSDM | 1 |
| 2022 | Online Ridesharing with Meeting PointsabstractNowadays, ridesharing becomes a popular commuting mode. Dynamically arriving riders post their origins and destinations, then the platform assigns drivers to serve them. In ridesharing, different groups of riders can be served by one driver if their trips can share common routes. Recently, many ridesharing companies (e.g., Didi and Uber) further propose a new mode, namely "ridesharing with meeting points". Specifically, with a short walking distance but less payment, riders can be picked up and dropped off around their origins and destinations, respectively. In addition, meeting points enables more flexible routing for drivers, which can potentially improve the global profit of the system. In this paper, we first formally define the Meeting-Point-based Online Ridesharing Problem (MORP). We prove that MORP is NP-hard and there is no polynomial-time deterministic algorithm with a constant competitive ratio for it. We notice that a structure of vertex set, k -skip cover, fits well to the MORP. k -skip cover tends to find the vertices (meeting points) that are convenient for riders and drivers to come and go. With meeting points, MORP tends to serve more riders with these convenient vertices. Based on the idea, we introduce a convenience-based meeting point candidates selection algorithm. We further propose a hierarchical meeting-point oriented graph (HMPO graph), which ranks vertices for assignment effectiveness and constructs k -skip cover to accelerate the whole assignment process. Finally, we utilize the merits of k -skip cover points for ridesharing and propose a novel algorithm, namely SMDB, to solve MORP. Extensive experiments on real and synthetic datasets validate the effectiveness and efficiency of our algorithms. Jiachuan Wang, Peng Cheng 0003, Libin Zheng 0001, Lei Chen 0002, Wenjie Zhang 0001 |
Proc. VLDB Endow. | 1 |
| 2021 | Privacy-Preserving Batch-based Task Assignment in Spatial Crowdsourcing with Untrusted ServerabstractIn this paper, we study the privacy-preserving task assignment problem in spatial crowdsourcing, where the locations of both workers and tasks, prior to their release to the server, are perturbed with Geo-Indistinguishability (a differential privacy notion for location-based systems). Different from the previously studied online setting, where each task is assigned immediately upon arrival, we target the batch-based setting, where the server maximizes the number of successfully assigned tasks after a batch of tasks arrive. To achieve this goal, we propose the k-Switch solution, which first divides the workers into small groups based on the perturbed distance between workers/tasks, and then utilizes Homomorphic Encryption (HE) based secure computation to enhance the task assignment. Furthermore, we expedite HE-based computation by limiting the size of the small groups under k. Extensive experiments demonstrate that, in terms of the number of successfully assigned tasks, the k-Switch solution improves batch-based baselines by 5.9X and the existing online solution by 1.74X, with no privacy leak. Maocheng Li, Jiachuan Wang, Libin Zheng 0001, Peng Cheng 0003, Lei Chen 0002, Xuemin Lin 0001 |
CIKM | 2 |
| 2020 | Demand-Aware Route Planning for Shared Mobility ServicesabstractThe dramatic development of shared mobility in food delivery, ridesharing, and crowdsourced parcel delivery has drawn great concerns. Specifically, shared mobility refers to transferring or delivering more than one passenger/package together when their traveling routes have common sub-routes or can be shared. A core problem for shared mobility is to plan a route for each driver to fulfill the requests arriving dynamically with given objectives. Previous studies greedily and incrementally insert each newly coming request to the most suitable worker with a minimum travel cost increase, which only considers the current situation and thus not optimal. In this paper, we propose a demand-aware route planning (DARP) for shared mobility services. Based on prediction, DARP tends to make optimal route planning with more information about requests in the future. We prove that the DARP problem is NP-hard, and further show that there is no polynomial-time deterministic algorithm with a constant competitive ratio for the DARP problem unless P=NP. Hence, we devise an approximation algorithm to realize the insertion operation for our goal. With the insertion algorithm, we devise a prediction based solution for the DARP problem. Extensive experiment results on real datasets validate the effectiveness and efficiency of our technique. Jiachuan Wang, Peng Cheng 0003, Libin Zheng 0001, Lei Chen 0002, Xuemin Lin 0001, Zheng Wang 0010 |
Proc. VLDB Endow. | 1 |
| 2007 | Genetically generated double-level fuzzy controller with a fuzzy adjustment strategyabstractThis paper describes the use of a genetic algorithm (GA) in tuning a double-level modular fuzzy logic controller (DLMFLC), which can expand its control working zone to a larger spectrum than a single-level FLC. The first-level FLCs are tuned by a GA so that the input parameters of their membership functions and fuzzy rules are optimized according to their individual working zones. The second-level FLC is then used to adjust contributions of the first-level FLCs to the final output signal of the whole controller, i.e., DLMFLC, so that it can function in a wider spectrum covering all individual working zones of the first-level FLCs. The second-level FLC is again optimized by a GA. An inverted pendulum system (IPS) is used to demonstrate the feasibility of the approach. Sofiane Achiche, Zhun Fan, Ali Gürcan Özkil, Torben Sørensen, Jiachuan Wang, Erik D. Goodman |
GECCO | 6 |
| 2005 | Knowledge interaction with genetic programming in mechatronic systems design using bond graphsabstractThis paper describes a unified network synthesis approach for the conceptual stage of mechatronic systems design using bond graphs. It facilitates knowledge interaction with evolutionary computation significantly by encoding the structure of a bond graph in a genetic programming tree representation. On the one hand, since bond graphs provide a succinct set of basic design primitives for mechatronic systems modeling, it is possible to extract useful modular design knowledge discovered during the evolutionary process for design creativity and reusability. On the other hand, design knowledge gained from experience can be incorporated into the evolutionary process to improve the topologically open-ended search capability of genetic programming for enhanced search efficiency and design feasibility. This integrated knowledge-based design approach is demonstrated in a quarter-car suspension control system synthesis and a MEMS bandpass filter design application. Jiachuan Wang, Zhun Fan, Janis P. Terpenny, Erik D. Goodman |
IEEE Trans. Syst. Man Cybern. Part C | 1 |
| 2004 | Hierarchical evolutionary synthesis of MEMSabstractWe discuss the hierarchy that is involved in a typical MEMS design and how evolutionary approaches can be used to automate the hierarchical design and synthesis process for MEMS. At the system level, the approach combining bond graphs and genetic programming can lead to satisfactory design candidates of system level models that meet the predefined behavioral specifications for designers to tradeoff. At the physical layout synthesis level, the selection of geometric parameters for component devices is formulated as a constrained optimization problem and addressed using a constrained GA approach. A multiple-resonator microsystem design is used to illustrate the integrated design automation idea using evolutionary approaches. Zhun Fan, Erik D. Goodman, Jiachuan Wang, Ronald C. Rosenberg, Kisung Seo, Jianjun Hu |
IEEE Congress on Evolutionary Computation | 3 |