VLDB 2026 Research / reviewers in the wild / expert
Weijia Zhang 0001
dblp:158/5387-1
· DBLP profile ↗
20ranked-venue papers
8as first author
14since 2021 · last 2025
0000-0001-8103-5325ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 6 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 5 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Partial Label Causal Representation Learning for Instance-Dependent Supervision and Domain GeneralizationabstractPartial label learning (PLL) addresses situations where each training example is associated with a set of candidate labels, among which only one corresponds to the true class label. As the candidate labels often come from crowdsourced workers, their generation is inherently dependent on the features of the instance. Existing PLL methods primarily aim to resolve these ambiguous labels to enhance classification accuracy, overlooking the opportunity to use this feature dependency for causal representation learning. This focus on accuracy can make PLL systems vulnerable to stylistic variations and shifts in domain. In this paper, we explore the learning of causal representations within an instance-dependent PLL framework, introducing a new approach that uncovers identifiable latent representations. By separating content from style in the identified causal representation, we introduce CausalPLL+, an algorithm for instance-dependent PLL based on causal representation. Our algorithm performs exceptionally well in terms of both classification accuracy and generalization robustness. Qualitative and quantitative experiments on instance-dependent PLL benchmarks and domain generalization tasks verify the effectiveness of our approach. Weijia Zhang 0001, Min-Ling Zhang |
AAAI | 2 |
| 2025 | Instance-dependent label noise learning via separating style from content
Han-Wen Deng, Weijia Zhang 0001, Min-Ling Zhang |
Pattern Recognit. Lett. | 2 |
| 2025 | Disentangled Representation Learning for Causal Inference With InstrumentsabstractLatent confounders are a fundamental challenge for inferring causal effects from observational data. The instrumental variable (IV) approach is a practical way to address this challenge. Existing IV-based estimators need a known IV or other strong assumptions, such as the existence of two or more IVs in the system, which limits the application of the IV approach. In this article, we consider a relaxed requirement, which assumes there is an IV proxy in the system without knowing which variable is the proxy. We propose a variational autoencoder (VAE)-based disentangled representation learning method to learn an IV representation from a dataset with latent confounders and then utilize the IV representation to obtain an unbiased estimation of the causal effect from the data. Extensive experiments on synthetic and real-world data have demonstrated that the proposed algorithm outperforms the existing IV-based estimators and VAE-based estimators. Debo Cheng, Jiuyong Li, Lin Liu 0003, Ziqi Xu 0001, Weijia Zhang 0001, Jixue Liu, Thuc Duy Le |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | Deep Copula-Based Survival Analysis for Dependent Censoring with Identifiability GuaranteesabstractCensoring is the central problem in survival analysis where either the time-to-event (for instance, death), or the time-to censoring (such as loss of follow-up) is observed for each sample. The majority of existing machine learning-based survival analysis methods assume that survival is conditionally independent of censoring given a set of covariates; an assumption that cannot be verified since only marginal distributions is available from the data. The existence of dependent censoring, along with the inherent bias in current estimators has been demonstrated in a variety of applications, accentuating the need for a more nuanced approach. However, existing methods that adjust for dependent censoring require practitioners to specify the ground truth copula. This requirement poses a significant challenge for practical applications, as model misspecification can lead to substantial bias. In this work, we propose a flexible deep learning-based survival analysis method that simultaneously accommodate for dependent censoring and eliminates the requirement for specifying the ground truth copula. We theoretically prove the identifiability of our model under a broad family of copulas and survival distributions. Experiments results from a wide range of datasets demonstrate that our approach successfully discerns the underlying dependency structure and significantly reduces survival estimation bias when compared to existing methods. Weijia Zhang 0001, Chun Kai Ling, Xuanhui Zhang |
AAAI | 1 |
| 2024 | Proposal Feature Learning Using Proposal Relations for Weakly Supervised Object DetectionabstractWeakly Supervised Object Detection (WSOD) trains detectors using only image-level annotations. Most existing WSOD models are based on pre-computed proposals and do not fully explore the relations of proposals. In this work, we address this limitation by proposing two approaches of Proposal Feature Learning for WSOD (PFL-WSOD), which effectively capture intra-proposal relations and inter-proposal relations respectively, thus improving proposal representation. To extract intra-proposal relations, we propose to utilize Self-Attention on Single Proposal for capturing relations inside each proposal. For inter-proposal relations, we propose Salient Region Banks by capturing a unique type of inter-proposal relation called deep inclusion, which significantly improves proposal representation when used in synergy with contrastive learning. Experimental results on benchmarks demonstrate the effectiveness of our methods. Zhaofei Wang, Weijia Zhang 0001, Min-Ling Zhang |
ICME | 2 |
| 2024 | Exploiting Conjugate Label Information for Multi-Instance Partial-Label Learning
Wei Tang 0017, Weijia Zhang 0001, Min-Ling Zhang |
IJCAI | 2 |
| 2024 | Multi-Instance Partial-Label Learning with Margin AdjustmentabstractMulti-instance partial-label learning (MIPL) is an emerging learning framework where each training sample is represented as a multi-instance bag associated with a candidate label set. Existing MIPL algorithms often overlook the margins for attention scores and predicted probabilities, leading to suboptimal generalization performance. A critical issue with these algorithms is that the highest prediction probability of the classifier may appear on a non-candidate label. In this paper, we propose an algorithm named MIPLMA, i.e., Multi-Instance Partial-Label learning with Margin Adjustment, which adjusts the margins for attention scores and predicted probabilities. We introduce a margin-aware attention mechanism to dynamically adjust the margins for attention scores and propose a margin distribution
loss to constrain the margins between the predicted probabilities on candidate and non-candidate label sets. Experimental results demonstrate the superior performance of MIPLMA over existing MIPL algorithms, as well as other well-established multi-instance learning algorithms and partial-label learning algorithms. Wei Tang 0017, Yin-Fang Yang, Zhaofei Wang, Weijia Zhang 0001, Min-Ling Zhang |
NeurIPS | 4 |
| 2024 | Multi-instance partial-label learning: towards exploiting dual inexact supervision
Wei Tang 0017, Weijia Zhang 0001, Min-Ling Zhang |
Sci. China Inf. Sci. | 2 |
| 2024 | Special Issue Editorial on "The Innovative Use of Data Science to Transform How We Work and Live"
Yee Ling Boo, Manik Gupta, Weijia Zhang 0001, Philippe Fournier-Viger |
Data Sci. Eng. | 3 |
| 2023 | Disambiguated Attention Embedding for Multi-Instance Partial-Label LearningabstractIn many real-world tasks, the concerned objects can be represented as a multi-instance bag associated with a candidate label set, which consists of one ground-truth label and several false positive labels. Multi-instance partial-label learning (MIPL) is a learning paradigm to deal with such tasks and has achieved favorable performances. Existing MIPL approach follows the instance-space paradigm by assigning augmented candidate label sets of bags to each instance and aggregating bag-level labels from instance-level labels. However, this scheme may be suboptimal as global bag-level information is ignored and the predicted labels of bags are sensitive to predictions of negative instances. In this paper, we study an alternative scheme where a multi-instance bag is embedded into a single vector representation. Accordingly, an intuitive algorithm named DEMIPL, i.e., Disambiguated attention Embedding for Multi-Instance Partial-Label learning, is proposed. DEMIPL employs a disambiguation attention mechanism to aggregate a multi-instance bag into a single vector representation, followed by a momentum-based disambiguation strategy to identify the ground-truth label from the candidate label set. Furthermore, we introduce a real-world MIPL dataset for colorectal cancer classification. Experimental results on benchmark and real-world datasets validate the superiority of DEMIPL against the compared MIPL and partial-label learning approaches. Wei Tang 0017, Weijia Zhang 0001, Min-Ling Zhang |
NeurIPS | 2 |
| 2022 | Multi-Instance Causal Representation Learning for Instance Label Prediction and Out-of-Distribution GeneralizationabstractMulti-instance learning (MIL) deals with objects represented as bags of instances and can predict instance labels from bag-level supervision. However, significant performance gaps exist between instance-level MIL algorithms and supervised learners since the instance labels are unavailable in MIL. Most existing MIL algorithms tackle the problem by treating multi-instance bags as harmful ambiguities and predicting instance labels by reducing the supervision inexactness. This work studies MIL from a new perspective by considering bags as auxiliary information, and utilize it to identify instance-level causal representations from bag-level weak supervision. We propose the CausalMIL algorithm, which not only excels at instance label prediction but also provides robustness to distribution change by synergistically integrating MIL with identifiable variational autoencoder. Our approach is based on a practical and general assumption: the prior distribution over the instance latent representations belongs to the non-factorized exponential family conditioning on the multi-instance bags. Experiments on synthetic and real-world datasets demonstrate that our approach significantly outperforms various baselines on instance label prediction and out-of-distribution generalization tasks. Weijia Zhang 0001, Xuanhui Zhang, Hanwen Deng, Min-Ling Zhang |
NeurIPS | 1 |
| 2022 | Imbalanced volunteer engagement in cultural heritage crowdsourcing: a task-related exploration based on causal inference
Xuanhui Zhang, Weijia Zhang 0001, Yuxiang Zhao 0001, Qinghua Zhu 0002 |
Inf. Process. Manag. | 2 |
| 2021 | Treatment Effect Estimation with Disentangled Latent FactorsabstractMuch research has been devoted to the problem of estimating treatment effects from observational data; however, most methods assume that the observed variables only contain confounders, i.e., variables that affect both the treatment and the outcome. Unfortunately, this assumption is frequently violated in real-world applications, since some variables only affect the treatment but not the outcome, and vice versa. Moreover, in many cases only the proxy variables of the underlying confounding factors can be observed. In this work, we first show the importance of differentiating confounding factors from instrumental and risk factors for both average and conditional average treatment effect estimation, and then we propose a variational inference approach to simultaneously infer latent factors from the observed variables, disentangle the factors into three disjoint sets corresponding to the instrumental, confounding, and risk factors, and use the disentangled factors for treatment effect estimation. Experimental results demonstrate the effectiveness of the proposed method on a wide range of synthetic, benchmark, and real-world datasets. Weijia Zhang 0001, Lin Liu 0003, Jiuyong Li |
AAAI | 1 |
| 2021 | Non-I.I.D. Multi-Instance Learning for Predicting Instance and Bag Labels with Variational Auto-EncoderabstractMulti-instance learning is a type of weakly supervised learning. It deals with tasks where the data is a set of bags and each bag is a set of instances. Only the bag labels are observed whereas the labels for the instances are unknown. An important advantage of multi-instance learning is that by representing objects as a bag of instances, it is able to preserve the inherent dependencies among parts of the objects. Unfortunately, most existing algorithms assume all instances to be identically and independently distributed, which violates real-world scenarios since the instances within a bag are rarely independent. In this work, we propose the Multi-Instance Variational Autoencoder (MIVAE) algorithm which explicitly models the dependencies among the instances for predicting both bag labels and instance labels. Experimental results on several multi-instance benchmarks and end-to-end medical imaging datasets demonstrate that MIVAE performs better than state-of-the-art algorithms for both instance label and bag label prediction tasks. Weijia Zhang 0001 |
IJCAI | 1 |
| 2020 | Robust Multi-Instance Learning with Stable InstancesabstractMulti-instance learning (MIL) deals with tasks where data is represented by a set of bags and each bag is described by a set of instances. Unlike standard supervised learning, only the bag labels are observed whereas the label for each instance is not available to the learner. Previous MIL studies typically follow the i.i.d. assumption, that the training and test samples are independently drawn from the same distribution. However, such assumption is often violated in real-world applications. Efforts have been made towards addressing distribution changes by importance weighting the training data with the density ratio between the training and test samples. Unfortunately, models often need to be trained without seeing the test distributions. In this paper we propose possibly the first framework for addressing distribution change in MIL without requiring access to the unlabeled test data. Our framework builds upon identifying a novel connection between MIL and the potential outcome framework in causal effect estimation. Experimental results on synthetic distribution change datasets, real-world datasets with synthetic distribution biases and real distributional biased image classification datasets validate the effectiveness of our approach. Weijia Zhang 0001, Lin Liu 0003, Jiuyong Li |
ECAI | 1 |
| 2020 | MONET: a toolbox integrating top-performing methods for network modularizationabstractSUMMARY: We define a disease module as a partition of a molecular network whose components are jointly associated with one or several diseases or risk factors thereof. Identification of such modules, across different types of networks, has great potential for elucidating disease mechanisms and establishing new powerful biomarkers. To this end, we launched the 'Disease Module Identification (DMI) DREAM Challenge', a community effort to build and evaluate unsupervised molecular network modularization algorithms. Here, we present MONET, a toolbox providing easy and unified access to the three top-performing methods from the DMI DREAM Challenge for the bioinformatics community. AVAILABILITY AND IMPLEMENTATION: MONET is a command line tool for Linux, based on Docker and Singularity containers; the core algorithms were written in R, Python, Ada and C++. It is freely available for download at https://github.com/BergmannLab/MONET.git. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Mattia Tomasoni, Jake Crawford, Weijia Zhang 0001, Sarvenaz Choobdar, Daniel Marbach, Sven Bergmann |
Bioinform. | 4 |
| 2018 | miRBaseConverter: an R/Bioconductor package for converting and retrieving miRNA name, accession, sequence and family information in different versions of miRBaseabstractBACKGROUND: miRBase is the primary repository for published miRNA sequence and annotation data, and serves as the "go-to" place for miRNA research. However, the definition and annotation of miRNAs have been changed significantly across different versions of miRBase. The changes cause inconsistency in miRNA related data between different databases and articles published at different times. Several tools have been developed for different purposes of querying and converting the information of miRNAs between different miRBase versions, but none of them individually can provide the comprehensive information about miRNAs in miRBase and users will need to use a number of different tools in their analyses. RESULTS: We introduce miRBaseConverter, an R package integrating the latest miRBase version 22 available in Bioconductor to provide a suite of functions for converting and retrieving miRNA name (ID), accession, sequence, species, version and family information in different versions of miRBase. The package is implemented in R and available under the GPL-2 license from the Bioconductor website ( http://bioconductor.org/packages/miRBaseConverter/ ). A Shiny-based GUI suitable for non-R users is also available as a standalone application from the package and also as a web application at http://nugget.unisa.edu.au:3838/miRBaseConverter . miRBaseConverter has a built-in database for querying miRNA information in all species and for both pre-mature and mature miRNAs defined by miRBase. In addition, it is the first tool for batch querying the miRNA family information. The package aims to provide a comprehensive and easy-to-use tool for miRNA research community where researchers often utilize published miRNA data from different sources. CONCLUSIONS: The Bioconductor package miRBaseConverter and the Shiny-based web application are presented to provide a suite of functions for converting and retrieving miRNA name, accession, sequence, species, version and family information in different versions of miRBase. The package will serve a wide range of applications in miRNA research and could provide a full view of the miRNAs of interest. Taosheng Xu, Lin Liu 0003, Junpeng Zhang 0001, Weijia Zhang 0001, Jie Gui, Kui Yu, Jiuyong Li, Thuc Duy Le |
BMC Bioinform. | 6 |
| 2018 | Estimating heterogeneous treatment effect by balancing heterogeneity and fitnessabstractBACKGROUND: Estimating heterogeneous treatment effect is a fundamental problem in biological and medical applications. Recently, several recursive partitioning methods have been proposed to identify the subgroups that respond differently towards a treatment, and they rely on a fitness criterion to minimize the error between the estimated treatment effects and the unobservable ground truths. RESULTS: In this paper, we propose that a heterogeneity criterion, which maximizes the differences of treatment effects among the subgroups, also needs to be considered. Moreover, we show that better performances can be achieved when the fitness and the heterogeneous criteria are considered simultaneously. Selecting the optimal splitting points then becomes a multi-objective problem; however, a solution that achieves optimal in both aspects are often not available. To solve this problem, we propose a multi-objective splitting procedure to balance both criteria. The proposed procedure is computationally efficient and fits naturally into the existing recursive partitioning framework. Experimental results show that the proposed multi-objective approach performs consistently better than existing ones. CONCLUSION: Heterogeneity should be considered with fitness in heterogeneous treatment effect estimation, and the proposed multi-objective splitting procedure achieves the best performance by balancing both criteria. Weijia Zhang 0001, Thuc Duy Le, Lin Liu 0003, Jiuyong Li |
BMC Bioinform. | 1 |
| 2017 | Mining heterogeneous causal effects for personalized cancer treatmentabstractMOTIVATION: Cancer is not a single disease and involves different subtypes characterized by different sets of molecules. Patients with different subtypes of cancer often react heterogeneously towards the same treatment. Currently, clinical diagnoses rather than molecular profiles are used to determine the most suitable treatment. A molecular level approach will allow a more precise and informed way for making treatment decisions, leading to a better survival chance and less suffering of patients. Although many computational methods have been proposed to identify cancer subtypes at molecular level, to the best of our knowledge none of them are designed to discover subtypes with heterogeneous treatment responses. RESULTS: In this article we propose the Survival Causal Tree (SCT) method. SCT is designed to discover patient subgroups with heterogeneous treatment effects from censored observational data. Results on TCGA breast invasive carcinoma and glioma datasets have shown that for each subtype identified by SCT, the patients treated with radiotherapy exhibit significantly different relapse free survival pattern when compared to patients without the treatment. With the capability to identify cancer subtypes with heterogeneous treatment responses, SCT is useful in helping to choose the most suitable treatment for individual patients. AVAILABILITY AND IMPLEMENTATION: Data and code are available at https://github.com/WeijiaZhang24/SurvivalCausalTree . CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Weijia Zhang 0001, Thuc Duy Le, Lin Liu 0003, Zhi-Hua Zhou, Jiuyong Li |
Bioinform. | 1 |
| 2014 | Multi-Instance Learning with Distribution ChangeabstractMulti-instance learning deals with tasks where each example is a bag of instances, and the bag labels of training data are known whereas instance labels are unknown. Most previous studies on multi-instance learning assumed that the training and testing data are from the same distribution; however, this assumption is often violated in real tasks. In this paper, we present possibly the first study on multi-instance learning with distribution change. We propose the MICS approach by considering both bag-level and instance-level distribution change. Experiments show that MICS is almost always significantly better than many state-of-the-art multi-instance learning algorithms when distribution change occurs; and even when there is no distribution change, their performances are still comparable. Weijia Zhang 0001, Zhi-Hua Zhou |
AAAI | 1 |