VLDB 2026 Research / reviewers in the wild / expert
Wenlu Zhang
dblp:117/4246
· DBLP profile ↗
21ranked-venue papers
5as first author
9since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 7 · 3 first-author · 1 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-authorSystems, architecture and hardware · 4 · 4 since 2021Computer networks · 4 · 4 since 2021Artificial intelligence and machine learning · 2 · 1 first-authorSecurity and privacy · 1Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Borrow Counter: A Generic Sketch Framework for Non-uniform Flow Estimation
Jiawei Huang 0001, Sitan Li, Zirong Wei, Shengwen Zhou, Wenlu Zhang, Jin Ye 0003 |
ICNP | 6 |
| 2024 | D2T: Dynamic Dual Threshold Policy of Shared-Memory in Data Center SwitchesabstractNowadays the data center switches employ the on-chip shared buffer to absorb bursts and avoid packet loss during transient congestion. However, as the buffer-per-port-per-Gbps in production data centers decreases, it becomes more challenging to provide efficient buffer management to meet the requirements of heterogeneous traffic. We observe that typical shared buffer management policies have two steps: first, they identify short flows arriving at ports and then allocate more buffer room for these ports. Unfortunately, the lack of isolation between long and short flows leads to increased queue buildup and even packet loss of short flows. To address this limitation, we propose D2T, which uses different queue length thresholds for long and short flows. Specifically, we first design a compact data structure to distinguish between long and short flows. Then when two kinds of flows coexist at the same port, the threshold of long flows will decrease to absorb the bursty short flows. We implement D2T at a P4- programmable switch and large-scale simulations. The results demonstrate that D2T reduces both average and tail flow completion times (FCT) of short flows by up to 29% and 62% compared with the state-of-the-art policies, respectively. Jiawei Huang 0001, Hui Li 0120, Jingling Liu, Wenlu Zhang, Yijun Li 0002, Sitan Li, Shengwen Zhou, Ping Zhong 0002, Jianxin Wang 0001, Wanchun Jiang, Yong Cui 0001 |
ICDCS | 7 |
| 2024 | Influence of Cyclic Pneumatic Brake on the Longitudinal Dynamics of Heavy-Haul Combined TrainsabstractHeavy-haul transportation plays a significant role in railway freight. Limited by the complex topography and resource distribution, China’s heavy-haul railways contain a great number of long downward slopes, which is the main challenge for heavy-haul combined trains. Constrained by the pneumatic brake system of the train, cyclic pneumatic brake needs to be applied to provide adequate braking capability, when the train is running on a long and steep downward slope. The subject of this work is 20,000-tonne heavy-haul combined trains, consisting of 1 locomotive +105 wagons + 1 locomotive + 105 wagons with end of train device (abbreviated as 1+1+ EOT). The test data show that the switching process from service lap to brake release is quite dangerous when applying cyclic pneumatic brake. During the brake release process, the asymmetric forces on the front and rear vehicles can result in the exacerbation of longitudinal forces. This paper establishes a simulation model of the heavy-haul combined trains’ cyclic braking process. Meanwhile, we focus on the release process of the cyclic pneumatic brake, and investigate the effect of the initial speed of dynamic braking adjustment, value and time of dynamic braking application, and differentiated collaborative control of leader and follower locomotives on the longitudinal train dynamics (LTD). The comparison of these simulation results reveals how differentiated collaborative control strategy works on the longitudinal dynamics of heavy-haul combined trains, laying the foundation for the optimization of longitudinal dynamic performance during cyclic pneumatic brake. Wenlu Zhang |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2023 | CAST: An Intricate-Scene Aware Adaptive Bitrate Approach for Video Streaming via Parallel Training
Weihe Li, Jiawei Huang 0001, Jingling Liu, Wenlu Zhang, Wenjun Lyu, Jianxin Wang 0001 |
ICA3PP (4) | 5 |
| 2023 | Achieving High Accuracy and Fast Speed for Sketch CompressionabstractTo reduce the communication overhead in distributed sketch system, it is desirable to compress sketches before uploading. However, current sketch compression approaches hardly achieve high speed of compression procedure and low error of compressed sketches at the same time. In this paper, we take a clean slate approach to design a sketch compression scheme called as Fast-Mapping that achieves both fast compression speed and high accuracy. Based on the prior statistics knowledge of bucket data distribution, Fast-Mapping compresses the similar buckets to obtain high accuracy. We also theoretically derive the compression error bound of Fast-Mapping. The experimental results show that, Fast-Mapping achieves higher speed and lower error than the-state-of-art works. Jin Ye 0003, Yuanchao Shan, Wenlu Zhang, Sitan Li, Jiawei Huang 0001 |
ICC | 3 |
| 2023 | PA-Sketch: A Fast and Accurate Sketch for Differentiated Flow EstimationabstractDue to the ability to maintain good accuracy and high throughput with limited memory resources, sketch has gained wide deployment and application for approximate flow estimation. However, most existing sketch approaches ignore the distinctions between flow priorities, though the high-priority flows are relatively scarce but hold significant information. Therefore, a class of priority-aware sketches has appeared recently to provide differentiated measurement accuracy for flows with different priorities. Unfortunately, it is challenging for these priority-aware sketches to strike a good balance between accuracy and throughput. To address this issue, we propose a priority-adaptive architecture PA-Sketch, which utilizes priority-aware hash to dynamically allocate appropriate numbers of hash functions for different flows according to their priorities. For the scenarios we experimented, we observed that PA-Sketch significantly improves accuracy while minimizing the hash overhead. Compared to the state-of-the-art priority-aware sketches, PA-Sketch achieves around 4.83x higher accuracy and 1.83x higher F1 score for high-priority flows on average, meanwhile maintaining slight accuracy loss for low-priority flows. Sitan Li, Jiawei Huang 0001, Wenlu Zhang |
ICNP | 3 |
| 2023 | ChainSketch: An Efficient and Accurate Sketch for Heavy Flow DetectionabstractIdentifying heavy flows is essential for network management. However, it is challenging to detect heavy flow quickly and accurately under the highly dynamic traffic and rapid growth of network capacity. Existing heavy flow detection schemes can make a trade-off in efficiency, accuracy and speed. However, these schemes still require memory large enough to obtain acceptable performance. To address this issue, we propose ChainSketch, which has the advantages of good memory efficiency, high accuracy and fast detection. Specifically, ChainSketch uses the selective replacement strategy to mitigate the over-estimation issue. Meanwhile, ChainSketch utilizes the hash chain and compact structure to improve memory efficiency. We implement the ChainSketch on OVS platform, P4-based testbed and large-scale simulations to process heavy hitter and heavy changer detection. The results of trace-driven tests show that, ChainSketch greatly improves the F1-score by up to$3.43\times $compared with the state-of-the-art solutions especially for small memory. Jiawei Huang 0001, Wenlu Zhang, Yijun Li 0002, Jin Ye 0003, Jianxin Wang 0001 |
IEEE/ACM Trans. Netw. | 2 |
| 2022 | UA-Sketch: An Accurate Approach to Detect Heavy Flow based on Uninterrupted ArrivalabstractHeavy flow detection in enormous network traffic is a critical task for network measurement. Due to the limited memory size and high link capacity, accurate detection of heavy flows becomes challenging in large-scale networks. Almost all existing approaches of detecting heavy flows use single-dimension statistics of flow size to make flow-replacement decisions. However, under the mass number of small flows, the heavy flows are prone to be frequently and mistakenly replaced, resulting in unsatisfactory accuracy. To solve this problem, we reveal that the number of uninterrupted arrival packets is a useful metric in identifying flow types. We further propose UA-Sketch that expels small flows and protects heavy ones according to the multiple-dimension statistics including both estimated flow size and number of uninterrupted arrival packets. The test results of trace-driven simulations and OVS experiments show that, even under small memory, UA-Sketch achieves higher accuracy than the existing works, with the F1 Score by up to 2.1 ×. Jin Ye 0003, Wenlu Zhang, Guihao Chen, Yuanchao Shan, Yijun Li 0002, Weihe Li, Jiawei Huang 0001 |
ICPP | 3 |
| 2022 | Prediction Modeling for Application-Specific Communication Architecture Design of Optical NoCabstractMulti-core systems-on-chip are becoming state-of-the-art. Therefore, there is a need for a fast and energy-efficient interconnect to take full advantage of the computational capabilities. Integration of silicon photonics with a traditional electrical interconnect in a Network-on-Chip (NoC) proposes a promising solution for overcoming the scalability issues of electrical interconnect. In this article, we derive and evaluate prediction modeling techniques for the design space exploration (DSE) of application-specific communication architectures for an Optical Network-on-Chip (ONoC). Our proposed model accurately predicts network packet latency, contention delay, and the static and dynamic energy consumption of the network. This work specifically addresses the challenge of accurately estimating performance metrics of the entire design space without having to perform time-consuming and computationally intensive exhaustive simulations. The proposed technique, based on machine learning (ML), can build accurate prediction models using only 10% to 50% (best case and worst case) of the entire design space. The accuracy, expressed as R 2 (Coefficient of Determination) is 0.99901, 0.99967, 0.99996, and 0.99999 for network packet latency, contention delay, static energy consumption, and dynamic energy consumption, respectively, in six different benchmarks from the Splash-2 benchmark suite, chosen among 6 different machine learning prediction models. Jelena Trajkovic, Sara Karimi, Samantha Hangsan, Wenlu Zhang |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2020 | Improving Fitness Levels of Individuals with Autism Spectrum Disorder: A Preliminary Evaluation of Real-Time Interactive Heart Rate Visualization to Motivate Engagement in Physical Activity
Bo Fu 0005, Jimmy Chao, Melissa Bittner, Wenlu Zhang, Mehrdad Aliasgari |
ICCHP (2) | 4 |
| 2020 | An Analytics Framework for Heuristic Inference Attacks against Industrial Control SystemsabstractIndustrial control systems (ICS) of critical infrastructure are increasingly connected to the Internet for remote site management at scale. However, cyber attacks against ICS - especially at the communication channels between human-machine interface (HMIs) and programmable logic controllers (PLCs) - are increasing at a rate which outstrips the rate of mitigation. In this paper, we introduce a vendor-agnostic analytics framework which allows security researchers to analyse attacks against ICS systems, even if the researchers have zero control automation domain knowledge or are faced with a myriad of heterogenous ICS systems. Unlike existing works that require expertise in domain knowledge and specialised tool usage, our analytics framework does not require prior knowledge about ICS communication protocols, PLCs, and expertise of any network penetration testing tool. Using `digital twin' scenarios comprising industry-representative HMIs, PLCs and firewalls in our test lab, our framework's steps were demonstrated to successfully implement a stealthy deception attack based on false data injection attacks (FDIA). Furthermore, our framework also demonstrated the relative ease of attack dataset collection, and the ability to leverage well-known penetration testing tools. We also introduce the concept of `heuristic inference attacks', a new family of attack types on ICS which is agnostic to PLC and HMI brands/models commonly deployed in ICS. Our experiments were also validated on a separate ICS dataset collected from a cyber-physical scenario of water utilities. Finally, we utilized time complexity theory to estimate the difficulty for the attacker to conduct the proposed packet analyses, and recommended countermeasures based on our findings. Taejun Choi, Guangdong Bai, Ryan Kok Leong Ko, Naipeng Dong, Wenlu Zhang, Shunyao Wang |
TrustCom | 5 |
| 2020 | Deep Model Based Transfer and Multi-Task Learning for Biological Image AnalysisabstractA central theme in learning from image data is to develop appropriate representations for the specific task at hand. Thus, a practical challenge is to determine what features are appropriate for specific tasks. For example, in the study of gene expression patterns inDrosophila, texture features were particularly effective for determining the developmental stages from in situ hybridization images. Such image representation is however not suitable for controlled vocabulary term annotation. Here, we developed feature extraction methods to generate hierarchical representations for ISH images. Our approach is based on the deep convolutional neural networks that can act on image pixels directly. To make the extracted features generic, the models were trained using a natural image set with millions of labeled examples. These models were transferred to the ISH image domain. To account for the differences between the source and target domains, we proposed a partial transfer learning scheme in which only part of the source model is transferred. We employed multi-task learning method to fine-tune the pre-trained models with labeled ISH images. Results showed that feature representations computed by deep models based on transfer and multi-task learning significantly outperformed other methods for annotating gene expression patterns at different stage ranges. Wenlu Zhang, Rongjian Li, Qian Sun 0002, Sudhir Kumar 0001, Jieping Ye, Shuiwang Ji |
IEEE Trans. Big Data | 1 |
| 2015 | Deep Model Based Transfer and Multi-Task Learning for Biological Image AnalysisabstractA central theme in learning from image data is to develop appropriate image representations for the specific task at hand. Traditional methods used handcrafted local features combined with high-level image representations to generate image-level representations. Thus, a practical challenge is to determine what features are appropriate for specific tasks. For example, in the study of gene expression patterns in Drosophila melanogaster, texture features based on wavelets were particularly effective for determining the developmental stages from in situ hybridization (ISH) images. Such image representation is however not suitable for controlled vocabulary (CV) term annotation because each CV term is often associated with only a part of an image. Here, we developed problem-independent feature extraction methods to generate hierarchical representations for ISH images. Our approach is based on the deep convolutional neural networks (CNNs) that can act on image pixels directly. To make the extracted features generic, the models were trained using a natural image set with millions of labeled examples. These models were transferred to the ISH image domain and used directly as feature extractors to compute image representations. Furthermore, we employed multi-task learning method to fine-tune the pre-trained models with labeled ISH images, and also extracted features from the fine-tuned models. Experimental results showed that feature representations computed by deep models based on transfer and multi-task learning significantly outperformed other methods for annotating gene expression patterns at different stage ranges. We also demonstrated that the intermediate layers of deep models produced the best gene expression pattern representations. Wenlu Zhang, Rongjian Li, Qian Sun 0002, Sudhir Kumar 0001, Jieping Ye, Shuiwang Ji |
KDD | 1 |
| 2015 | Evolutionary soft co-clustering: formulations, algorithms, and applications
Wenlu Zhang, Rongjian Li, Daming Feng, Andrey N. Chernikov, Nikos Chrisochoides, Christopher Osgood, Shuiwang Ji |
Data Min. Knowl. Discov. | 1 |
| 2015 | Sparsity Learning Formulations for Mining Time-Varying DataabstractTraditional clustering and feature selection methods consider the data matrix as static. However, the data matrices evolve smoothly over time in many applications. A simple approach to learn from these time-evolving data matrices is to analyze them separately. Such strategy ignores the time-dependent nature of the underlying data. In this paper, we propose two formulations for evolutionary co-clustering and feature selection based on the fused Lasso regularization. The evolutionary co-clustering formulation is able to identify smoothly varying hidden block structures embedded into the matrices along the temporal dimension. Our formulation is very flexible and allows for imposing smoothness constraints over only one dimension of the data matrices. The evolutionary feature selection formulation can uncover shared features in clustering from time-evolving data matrices. We show that the optimization problems involved are non-convex, non-smooth and non-separable. To compute the solutions efficiently, we develop a two-step procedure that optimizes the objective function iteratively. We evaluate the proposed formulations using the Allen Developing Mouse Brain Atlas data. Results show that our formulations consistently outperform prior methods. Rongjian Li, Wenlu Zhang, Yao Zhao 0001, Zhenfeng Zhu, Shuiwang Ji |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2014 | Deep Learning Based Imaging Data Completion for Improved Brain Disease Diagnosis
Rongjian Li, Wenlu Zhang, Heung-Il Suk, Li Wang 0026, Jiang Li 0001, Dinggang Shen, Shuiwang Ji |
MICCAI (3) | 2 |
| 2014 | Automated identification of cell-type-specific genes in the mouse brain by image computing of expression patternsabstractBACKGROUND: Differential gene expression patterns in cells of the mammalian brain result in the morphological, connectional, and functional diversity of cells. A wide variety of studies have shown that certain genes are expressed only in specific cell-types. Analysis of cell-type-specific gene expression patterns can provide insights into the relationship between genes, connectivity, brain regions, and cell-types. However, automated methods for identifying cell-type-specific genes are lacking to date. RESULTS: Here, we describe a set of computational methods for identifying cell-type-specific genes in the mouse brain by automated image computing of in situ hybridization (ISH) expression patterns. We applied invariant image feature descriptors to capture local gene expression information from cellular-resolution ISH images. We then built image-level representations by applying vector quantization on the image descriptors. We employed regularized learning methods for classifying genes specifically expressed in different brain cell-types. These methods can also rank image features based on their discriminative power. We used a data set of 2,872 genes from the Allen Brain Atlas in the experiments. Results showed that our methods are predictive of cell-type-specificity of genes. Our classifiers achieved AUC values of approximately 87% when the enrichment level is set to 20. In addition, we showed that the highly-ranked image features captured the relationship between cell-types. CONCLUSIONS: Overall, our results showed that automated image computing methods could potentially be used to identify cell-type-specific genes in the mouse brain. Rongjian Li, Wenlu Zhang, Shuiwang Ji |
BMC Bioinform. | 2 |
| 2013 | Evolutionary Soft Co-ClusteringabstractWe consider the mining of hidden block structures from time-varying data using evolutionary co-clustering. Existing methods are based on the spectral learning framework, thus lacking a probabilistic interpretation. To overcome this limitation, we develop a probabilistic model for evolutionary co-clustering in this paper. The proposed model assumes that the observed data are generated via a two-step process that depends on the historic co-clusters, thereby capturing the temporal smoothness in a probabilistically principled manner. We develop an EM algorithm to perform maximum likelihood parameter estimation. An appealing feature of the proposed probabilistic model is that it leads to soft co-clustering assignments naturally. To the best of our knowledge, our work represents the first attempt to perform evolutionary soft co-clustering. We evaluate the proposed method on both synthetic and real data sets. Experimental results show that our method consistently outperforms prior approaches based on spectral method. Shuiwang Ji, Wenlu Zhang |
SDM | 2 |
| 2013 | A mesh generation and machine learning framework for Drosophila gene expression pattern image analysisabstractBACKGROUND: Multicellular organisms consist of cells of many different types that are established during development. Each type of cell is characterized by the unique combination of expressed gene products as a result of spatiotemporal gene regulation. Currently, a fundamental challenge in regulatory biology is to elucidate the gene expression controls that generate the complex body plans during development. Recent advances in high-throughput biotechnologies have generated spatiotemporal expression patterns for thousands of genes in the model organism fruit fly Drosophila melanogaster. Existing qualitative methods enhanced by a quantitative analysis based on computational tools we present in this paper would provide promising ways for addressing key scientific questions. RESULTS: We develop a set of computational methods and open source tools for identifying co-expressed embryonic domains and the associated genes simultaneously. To map the expression patterns of many genes into the same coordinate space and account for the embryonic shape variations, we develop a mesh generation method to deform a meshed generic ellipse to each individual embryo. We then develop a co-clustering formulation to cluster the genes and the mesh elements, thereby identifying co-expressed embryonic domains and the associated genes simultaneously. Experimental results indicate that the gene and mesh co-clusters can be correlated to key developmental events during the stages of embryogenesis we study. The open source software tool has been made available at http://compbio.cs.odu.edu/fly/. CONCLUSIONS: Our mesh generation and machine learning methods and tools improve upon the flexibility, ease-of-use and accuracy of existing methods. Wenlu Zhang, Daming Feng, Rongjian Li, Andrey N. Chernikov, Nikos Chrisochoides, Christopher Osgood, Charlotte Konikoff, Stuart J. Newfeld, Sudhir Kumar 0001, Shuiwang Ji |
BMC Bioinform. | 1 |
| 2013 | A Probabilistic Latent Semantic Analysis Model for Coclustering the Mouse Brain AtlasabstractThe mammalian brain contains cells of a large variety of types. The phenotypic properties of cells of different types are largely the results of distinct gene expression patterns. Therefore, it is of critical importance to characterize the gene expression patterns in the mammalian brain. The Allen Developing Mouse Brain Atlas provides spatiotemporal in situ hybridization gene expression data across multiple stages of mouse brain development. It provides a framework to explore spatiotemporal regulation of gene expression during development. We employ a graph approximation formulation to cocluster the genes and the brain voxels simultaneously for each time point. We show that this formulation can be expressed as a probabilistic latent semantic analysis (PLSA) model, thereby allowing us to use the expectation-maximization algorithm for PLSA to estimate the coclustering parameters. To provide a quantitative comparison with prior methods, we evaluate the coclustering method on a set of standard synthetic data sets. Results indicate that our method consistently outperforms prior methods. We apply our method to cocluster the Allen Developing Mouse Brain Atlas data. Results indicate that our clustering of voxels is more consistent with classical neuroanatomy than those of prior methods. Our analysis also yields sets of genes that are co-expressed in a subset of the brain voxels. Shuiwang Ji, Wenlu Zhang, Rongjian Li |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2012 | A sparsity-inducing formulation for evolutionary co-clusteringabstractTraditional co-clustering methods identify block structures from static data matrices. However, the data matrices in many applications are dynamic; that is, they evolve smoothly over time. Consequently, the hidden block structures embedded into the matrices are also expected to vary smoothly along the temporal dimension. It is therefore desirable to encourage smoothness between the block structures identified from temporally adjacent data matrices. In this paper, we propose an evolutionary co-clustering formulation for identifying co-cluster structures from time-varying data. The proposed formulation encourages smoothness between temporally adjacent blocks by employing the fused Lasso type of regularization. Our formulation is very flexible and allows for imposing smoothness constraints over only one dimension of the data matrices, thereby enabling its applicability to a large variety of settings. The optimization problem for the proposed formulation is non-convex, non-smooth, and non-separable. We develop an iterative procedure to compute the solution. Each step of the iterative procedure involves a convex, but non-smooth and non-separable problem. We propose to solve this problem in its dual form, which is convex and smooth. This leads to a simple gradient descent algorithm for computing the dual optimal solution. We evaluate the proposed formulation using the Allen Developing Mouse Brain Atlas data. Results show that our formulation consistently outperforms methods without the temporal smoothness constraints. Shuiwang Ji, Wenlu Zhang, Jun Liu 0003 |
KDD | 2 |