VLDB 2026 Research / reviewers in the wild / expert
Guojun Wu
dblp:200/2455
· DBLP profile ↗
12ranked-venue papers
7as first author
9since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 4 first-author · 4 since 2021Databases, data management, data science and information retrieval · 6 · 4 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Theory of computation · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | iDOMO: identification of drug combinations via multi-set operations for treating diseasesabstractCombination therapy has become increasingly important for treating complex diseases which often involve multiple pathways and targets. However, experimental screening of drug combinations is costly and time-consuming. The availability of large-scale transcriptomic datasets (e.g. CMap and LINCS) from in vitro drug treatment experiments makes it possible to computationally predict drug combinations with synergistic effects. Towards this end, we developed a computational approach, termed Identification of Drug Combinations via Multi-Set Operations (iDOMO), to predict drug synergy based on multi-set operations of drug and disease gene signatures. iDOMO quantifies the synergistic effect of a pair of drugs by taking into account the combination's beneficial and detrimental effects on treating a disease. We evaluated iDOMO, in a DREAM Challenge dataset with the matched, pre- and post-treatment gene expression data and cell viability information. We further evaluated the performance of iDOMO by concordance index and Spearman correlation on predicting the Highest Single Agency (HSA) synergy scores for four most common cancer types in two large-scale drug combination databases, showing that iDOMO significantly outperformed two existing popular drug combination approaches including the Therapeutic Score and the SynergySeq Orthogonality Score. Application of iDOMO to triple-negative breast cancer (TNBC) identified drug pairs with potential synergistic effects, with the combination of trifluridine and monobenzone being the most synergistic. Our in vitro experiments confirmed that the top predicted drug combination exerted a significant synergistic effect in inhibiting TNBC cell growth. In summary, iDOMO is an effective method for the in silico screening of synergistic drug combinations and will be a valuable tool for the development of novel therapeutics for complex diseases. Xianxiao Zhou, Guojun Wu |
Briefings Bioinform. | 4 |
| 2025 | A categorical equivalence between Q-domains and interpolative generalized Q-closure spaces
Guojun Wu, Wei Yao 0004, Qingguo Li |
Fuzzy Sets Syst. | 1 |
| 2025 | sL-approximation spaces capture sL-domainsabstractAbstract In this paper, by means of upper approximation operators in rough set theory, we study representations for sL-domains and its special subclasses. We introduce the concepts of sL-approximation spaces, L-approximation spaces, and bc-approximation spaces, which are special types of CF-approximation spaces. We prove that the collection of CF-closed sets in an sL-approximation space (resp., an L-approximation space, a bc-approximation space) ordered by set-theoretic inclusion is an sL-domain (resp., an L-domain, a bc-domain); conversely, every sL-domain (resp., L-domain, bc-domain) is order-isomorphic to the collection of CF-closed sets of an sL-approximation space (resp., an L-approximation space, a bc-approximation space). Consequently, we establish an equivalence between the category of sL-domains (resp., L-domains) with Scott continuous mappings and that of sL-approximation spaces (resp., L-approximation spaces) with CF-approximable relations. Guojun Wu, Luoshan Xu, Wei Yao 0004 |
Math. Struct. Comput. Sci. | 1 |
| 2025 | StreakNet-Arch: An Anti-Scattering Network-Based Architecture for Underwater Carrier LiDAR-Radar ImagingabstractIn this paper, we introduce StreakNet-Arch, a real-time, end-to-end binary-classification framework based on our self-developed Underwater Carrier LiDAR-Radar (UCLR) that embeds Self-Attention and our novel Double Branch Cross Attention (DBC-Attention) to enhance scatter suppression. Under controlled water tank validation conditions, StreakNet-Arch with Self-Attention or DBC-Attention outperforms traditional bandpass filtering and achieves higher $F_{1}$ scores than learning-based MP networks and CNNs at comparable model size and complexity. Real-time benchmarks on an NVIDIA RTX 3060 show a constant Average Imaging Time (54 to 84 ms) regardless of frame count, versus a linear increase (58 to 1,257 ms) for conventional methods. To facilitate further research, we contribute a publicly available streak-tube camera image dataset contains 2,695,168 real-world underwater 3D point cloud data. More importantly, we validate our UCLR system in a South China Sea trial, reaching an error of 46mm for 3D target at 1,000 m depth and 20 m range. Source code and data are available at https://github.com/BestAnHongjun/StreakNet. Xuelong Li 0001, Hongjun An, Haofei Zhao, Guangying Li, Bo Liu 0111, Guanghua Cheng, Guojun Wu, Zhe Sun 0007 |
IEEE Trans. Image Process. | 8 |
| 2023 | DNet: Distributional Network for Distributional Individualized Treatment EffectsabstractThere is a growing interest in developing methods to estimate individualized treatment effects (ITEs) for various real-world applications, such as e-commerce and public health. This paper presents a novel architecture, called DNet, to infer distributional ITEs. DNet can learn the entire outcome distribution for each treatment, whereas most existing methods primarily focus on the conditional average treatment effect and ignore the conditional variance around its expectation. Additionally, our method excels in settings with heavy-tailed outcomes and outperforms state-of-the-art methods in extensive experiments on benchmark and real-world datasets. DNet has also been successfully deployed in a widely used mobile app with millions of daily active users. Guojun Wu, Xiaoxiang Lv, Shikai Luo, Chengchun Shi, Hongtu Zhu |
KDD | 1 |
| 2023 | Learning Lightweight Neural Networks via Channel-Split Recurrent ConvolutionabstractLightweight neural networks refer to deep networks with small numbers of parameters, which can be deployed in resource-limited hardware such as embedded systems. To learn such lightweight networks effectively and efficiently, in this paper we propose a novel convolutional layer, namely Channel-Split Recurrent Convolution (CSR-Conv), where we split the output channels to generate data sequences with length T as the input to the recurrent layers with shared weights. As a consequence, we can construct lightweight convolutional networks by simply replacing (some) linear convolutional layers with CSR-Conv layers. We prove that under mild conditions the model size decreases with the rate of $O\left( {\frac{1}{{{T^2}}}} \right)$. Empirically we demonstrate the state-of-the-art performance using VGG-16, ResNet-50, ResNet-56, ResNet-110, DenseNet-40, MobileNet, and EfficientNet as backbone networks on CIFAR-10 and ImageNet. Codes can be found on https://github.com/tuaxon/CSR_Conv. Guojun Wu, Xin Zhang 0098, Xun Zhou 0001, Christopher G. Brinton, Zhenming Liu |
WACV | 1 |
| 2021 | Deep Incremental RNN for Learning Sequential Data: A Lyapunov Stable Dynamical SystemabstractWith the recent advances in mobile sensing technologies, large amounts of sequential data are collected, such as vehicle GPS records, stock prices, sensor data from air quality detectors. Recurrent neural networks (RNNs) have been studied extensively to learn complex patterns for sequential data, with applicatons in natural language processing for sentence prediction/completion, human activity recognition for predicting or classifying human activities. However, there are many practical issues when training RNNs, e.g., vanishing and exploding gradients often occur due to the repeatability of network weights, etc. In this paper, we study the training stability in deep recurrent neural networks (RNNs), and propose a novel network, namely, deep incremental RNN (DIRNN). In contrast to the literature, we prove that DIRNN is essentially a Lyapunov stable dynamical system where there is no vanishing or exploding gradient in training. To demonstrate the applicability in practice, we also propose a novel implementation, namely TinyRNN, that sparsifies the transition matrices in DIRNN using weighted random permutations to reduce the model sizes. We evaluate our approach on seven benchmark datasets, and achieve state-of-the-art results. Demo code is provided in the supplementary file. Guojun Wu, Yun Yue, Xun Zhou 0001 |
ICDM | 2 |
| 2021 | SBO-RNN: Reformulating Recurrent Neural Networks via Stochastic Bilevel OptimizationabstractIn this paper we consider the training stability of recurrent neural networks (RNNs) and propose a family of RNNs, namely SBO-RNN, that can be formulated using stochastic bilevel optimization (SBO). With the help of stochastic gradient descent (SGD), we manage to convert the SBO problem into an RNN where the feedforward and backpropagation solve the lower and upper-level optimization for learning hidden states and their hyperparameters, respectively. We prove that under mild conditions there is no vanishing or exploding gradient in training SBO-RNN. Empirically we demonstrate our approach with superior performance on several benchmark datasets, with fewer parameters, less training data, and much faster convergence. Code is available at https://zhang-vislab.github.io. Yun Yue, Guojun Wu, Haichong K. Zhang |
NeurIPS | 3 |
| 2021 | Mining Spatio-Temporal Reachable Regions With Multiple Sources over Massive Trajectory DataabstractGiven a set of user-specified locations and a massive trajectory dataset, the task of mining spatio-temporal reachable regions aims at finding which road segments are reachable from these locations within a given temporal period based on the historical trajectories. Determining such spatio-temporal reachable regions with high accuracy is vital for many urban applications, such as location-based recommendations and advertising. Traditional approaches to answering such queries essentially perform a distance-based range query over the given road network, which does not consider dynamic travel time at different time of day. By contrast, we propose a data-driven approach to formulate the problem as mining actual reachable regions based on a real historical trajectory dataset. Efficient algorithms for the Single-location spatio-temporal reachability Query (S-Query) and the Union-of-multi-location spatio-temporal reachability Query (U-Query) were presented in our recent work. In this paper, we extend the previous ideas by introducing a new type of reachability query with multiple sources, namely, the Intersection-of-multi-location spatio-temporal reachability Query (I-Query). As we demonstrate, answering I-Queries efficiently is generally more computationally challenging than answering either S-Queries or U-Queries because I-Queries involve complicated intersect conditions. We propose two new algorithms called the Intersection-of-Multi-location Query Maximum Bounding region search (I-MQMB) algorithm and the I-Query Trace Back Search (I-TBS) algorithm to efficiently answer I-Queries, which utilize an indexing schema composed of a spatio-temporal index and a connection index. We evaluate our system extensively by using a large-scale real taxi trajectory dataset that records taxi rides in Shenzhen, China. Our results demonstrate that the proposed approach reduces the running time of I-Queries by 50 percent on average compared to the baseline method. Yichen Ding, Xun Zhou 0001, Guojun Wu, Jie Bao 0003, Yu Zheng 0004, Jun Luo 0007 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2020 | A Joint Inverse Reinforcement Learning and Deep Learning Model for Drivers' Behavioral PredictionabstractUsers' behavioral predictions are crucially important for many domains including major e-commerce companies, ride-hailing platforms, social networking, and education. The success of such prediction strongly depends on the development of representation learning that can effectively model the dynamic evolution of user's behavior. This paper aims to develop a joint framework of combining inverse reinforcement learning (IRL) with deep learning (DL) regression model, called IRL-DL, to predict drivers' future behavior in ride-hailing platforms. Specifically, we formulate the dynamic evolution of each driver as a sequential decision-making problem and then employ IRL as representation learning to learn the preference vector of each driver. Then, we integrate drivers' preference vector with their static features (e.g., age, gender) and other attributes to build a regression model (e.g., LTSM-neural network) to predict drivers' future behavior. We use an extensive driver data set obtained from a ride-sharing platform to verify the effectiveness and efficiency of our IRL-DL framework, and results show that our IRL-DL framework can achieve consistent and remarkable improvements over models without drivers' preference vectors. Guojun Wu, Shikai Luo, Jieping Ye, Xiaohu Qie, Hongtu Zhu |
CIKM | 1 |
| 2018 | Human-Centric Urban Transit Evaluation and PlanningabstractPublic transits, such as buses and subway lines, offer affordable ride-sharing services and reduce the road network traffic, thus have significant impacts in mitigating the urban traffic congestion problem. However, it is non-trivial to evaluate a new transit plan, such as a new bus route or a new subway line, of its future ridership prior to actual deployment, since the travel preferences of passengers along the planned routes may vary. In this paper, we make the first attempt to model passengers' preferences of making various transit choices using a Markov Decision Process (MDP). Moreover, we develop a novel inverse preference learning algorithm to infer the passengers' preferences and predict the future human behavior changes, e.g., ridership, of a new urban transit plan before its deployment. We validate our proposed framework using a unique real-world dataset (from Shenzhen, China) with three subway lines opened during the data time span. With the data collected from both before and after the transit plan deployments, Our evaluation results demonstrated that the proposed framework can predict the ridership with only 19.8% relative error, which is 23%-51% lower than other baseline approaches. Guojun Wu, Jie Bao 0003, Yu Zheng 0004, Jieping Ye, Jun Luo 0007 |
ICDM | 1 |
| 2017 | Mining Spatio-Temporal Reachable Regions over Massive Trajectory DataabstractMining spatio-temporal reachable regions aims to find a set of road segments from massive trajectory data, that are reachable from a user-specified location and within a given temporal period. Accurately extracting such spatiotemporal reachable area is vital in many urban applications, e.g., (i) location-based recommendation, (ii) location-based advertising, and (iii) business coverage analysis. The traditional approach of answering such queries essentially performs a distance-based range query over the given road network, which have two main drawbacks: (i) it only works with the physical travel distances, where the users usually care more about dynamic traveling time, and (ii) it gives the same result regardless of the querying time, where the reachable area could vary significantly with different traffic conditions. Motivated by these observations, we propose a data-driven approach to formulate the problem as mining actual reachable region based on real historical trajectory dataset. The main challenge in our approach is the system efficiency, as verifying the reachability over the massive trajectories involves huge amount of disk I/Os. In this paper, we develop two indexing structures: 1) spatio-temporal index (ST-Index) and 2) connection index (Con-Index) to reduce redundant trajectory data access operations. We also propose a novel query processing algorithm with: 1) maximum bounding region search, which directly extracts a small searching region from the index structure and 2) trace back search, which refines the search results from the previous step to find the final query result. Moreover, our system can also efficiently answer the spatio-temporal reachability query with multiple query locations by skipping the overlapped area search. We evaluate our system extensively using a large-scale real taxi trajectory data in Shenzhen, China, where results demonstrate that the proposed algorithms can reduce 50%-90% running time over baseline algorithms. Guojun Wu, Yichen Ding, Jie Bao 0003, Yu Zheng 0004, Jun Luo 0007 |
ICDE | 1 |