Cong Su

dblp:164/3904 · DBLP profile ↗
← Back
8ranked-venue papers
5as first author
7since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021Computer networks · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Counterfactual Fairness with Imperfect Causal Graphs
abstract
Fairness-aware machine learning aims to build predictive models that comply with fairness requirements, particularly concerning sensitive attributes such as race, gender, and age. Among causality-based fairness notions, counterfactual fairness is widely adopted for its individual-level guarantees, requiring that an individual’s predicted outcome remains unchanged in a counterfactual world where its sensitive attribute is altered. However, existing methods critically assume that the true causal graph is fully known, which is rarely the case in practice. Moreover, counterfactual fairness suffers from inherent identifiability limitations, as counterfactual quantities cannot always be uniquely estimated from observational data, especially under incomplete causal knowledge. To address these challenges, we propose a principled framework (CF-ICG) for counterfactual fairness under imperfectly known causal graphs, e.g., Completed Partially Directed Acyclic Graphs (CPDAGs). We first introduce a criterion to determine the identifiability, and bound the counterfactual quantities under CPDAGs. Building upon this, we develop an efficient local algorithm that avoids the exhaustive enumeration of all DAGs, ensuring robustness against worst-case fairness violations. Experimental results on synthetic and real-world datasets demonstrate the practical effectiveness and theoretical soundness of CF-ICG.
Cong Su, Qiaoyu Tan, Carlotta Domeniconi, Jun Wang 0035, Guoxian Yu
AAAI1
2026 Autonomous markerless augmented reality for perforator visualization in DIEP surgery
Zitan Shi, Cong Su, Jian Yin 0010, Bo Guan 0004, Jianchang Zhao
Appl. Intell.3
2025 AN-IHP: Incompatible Herb Pairs Prediction by Attention Networks
abstract
The adverse drug-drug interaction (DDI) is a crucial safety concern in drug development. The intricate combinations of traditional Chinese medicine (TCM), while powerful in their therapeutic potentials, also harbor potential risks when incompatibly paired. Some methods have been proposed to infer incompatible herb pairs (IHPs), but most of them focus on revealing and analyzing the adverse reactions of known IHPs, despite that there are still a number of undiscovered IHPs at intervals. This paper introduces a deep attention network (AN-IHP) that effectively exploits diverse types of data for IHPs prediction. AN-IHP designs an attention-aggregation block to learn the ingredient-level features towards herbs and use similarity profiles to represent the efficacy and property. Then it defines commonality and specificity constraints to enhance the representations from different types of features. After that, it makes dynamic representation fusion across herb pairs using a gated attention unit (GAU) and leverages a deep neural network (DNN) to predict IHPs. The experimental results on the collected IHPTCM dataset demonstrate that AN-IHP outperforms competitive methods. AN-IHP provides interpretability for analyzing IHPs at the ingredient level, proves beneficial for wet-lab experiments. It is also capable of predicting DDIs.
Cong Su, Jun Wang 0035, Xin Li 0251, Guoxian Yu
IEEE Trans. Comput. Biol. Bioinform.2
2025 Multi-Dimensional Causality Fairness Learning
Cong Su, Guoxian Yu, Jun Wang 0035, Wei Guo 0017, Yongqing Zheng, Carlotta Domeniconi
IEEE Trans. Knowl. Data Eng.1
2024 Multi-Dimensional Fair Federated Learning
abstract
Federated learning (FL) has emerged as a promising collaborative and secure paradigm for training a model from decentralized data without compromising privacy. Group fairness and client fairness are two dimensions of fairness that are important for FL. Standard FL can result in disproportionate disadvantages for certain clients, and it still faces the challenge of treating different groups equitably in a population. The problem of privately training fair FL models without compromising the generalization capability of disadvantaged clients remains open. In this paper, we propose a method, called mFairFL, to address this problem and achieve group fairness and client fairness simultaneously. mFairFL leverages differential multipliers to construct an optimization objective for empirical risk minimization with fairness constraints. Before aggregating locally trained models, it first detects conflicts among their gradients, and then iteratively curates the direction and magnitude of gradients to mitigate these conflicts. Theoretical analysis proves mFairFL facilitates the fairness in model development. The experimental evaluations based on three benchmark datasets show significant advantages of mFairFL compared to seven state-of-the-art baselines.
Cong Su, Guoxian Yu, Jun Wang 0035, Hui Li 0048, Qingzhong Li, Han Yu 0001
AAAI1
2024 Causality-Based Fair Multiple Decision by Response Functions
abstract
A recent trend of fair machine learning is to build a decision model subjected to causality-based fairness requirements, which concern with the causality between sensitive attributes and decisions. Almost all (if not all) solutions focus on a single fair decision model and assume no hidden confounder to model causal effects in a too simplified way. However, multiple interdependent decision models are actually used and discrimination may transmit among them. The hidden confounder is another inescapable fact and causal effects cannot be computed from observational data in the unidentifiable situation. To address these problems, we propose a method called CMFL (Causality-based Multiple Fairness Learning). CMFL parameterizes the causal model by response-function variables, whose distributions capture the randomness of causal models. CMFL treats each classifier as a soft intervention to infer the post-intervention distribution, and combines the fairness constraints with the classification loss to train multiple decision classifiers. In this way, all classifiers can make approximately fair decisions. Experiments on synthetic and benchmark datasets confirm its effectiveness, the response-function variables can deal with the unidentifiable issue and hidden confounders.
Cong Su, Guoxian Yu, Yongqing Zheng, Jun Wang 0035, Zhengtian Wu, Xiangliang Zhang 0001, Carlotta Domeniconi
ACM Trans. Knowl. Discov. Data1
2021 Cost-effective multi-instance multilabel active learning
abstract
Multi-instance multi-label (MIML) Active Learning (M2AL) aims to improve the learner while reducing the cost as much as possible by querying informative labels of complex bags composed of diverse instances. Existing M2AL solutions suffer high query costs for scrutinizing all relevant labels of MIML samples, querying excessive bag–label or instance–label pairs. To address these issues, a Cost-effective M2AL solution (CM2AL) is presented. CM2AL first selects the most informative bag–label pairs by leveraging uncertainty, label correlations, label space sparsity, and informativeness from queried instances of the bag, and thus avoids scrutinizing all labels. Next, it queries the most probably positive instance–label pairs of the selected bag–label pair. Particularly, if the feedback is positive, the bag is positively annotated with the label. For negative feedback, it further leverages the label of the neighborhood bags and the label of the nearby instances of queried instances of this bag, if the suggested labels from bag- and instance-levels disagree, CM2AL temporally gives up querying this bag–label pair and moves to another most informativeness one; otherwise, it takes the agreed label to annotate the bag, which further saves the cost by avoiding the excessive query. Extensive experiments on MIML data sets from diverse domains show that CM2AL can more reduce the cost while managing a better performance than state-of-the-art methods, the collaboration between bags and instances contributes to the saved cost.
Cong Su, Zhongmin Yan, Guoxian Yu
Int. J. Intell. Syst.1
2019 Self-Adaptive Probabilistic Sampling for Elephant Flows Detection
abstract
Sampling traffic traces to collect data from Internet nodes offers a new approach to reduce the traffic amount of sampling and the memory depletion. However, the loss of accuracy in traffic detection may be large due to the impertinent sampling probability. Many elephant flows detection methods suffer the high time consumption and the poor detection results from the empirical sampling. In this paper, the self-adaptive sampling method for elephant flow detection is investigated based on the heavy-tailed distribution of Internet. Specifically, the SaPS (Self-adaptive Probabilistic Sampling) algorithm is proposed to capture the characteristics of heavy-tailed flows through periodically calculating the kurtosis of tailedness for the flows. We employ this algorithm to provide simplified but representative samples to four well- known detection algorithms. The results of extensive simulations show that our sampling algorithm achieves the high performance in terms of time and memory consumption while maintaining a high accuracy through dynamically adjusting sampling probability for elephant flows detection.
Jun Tao 0003, Yizheng Li, Zhaoyue Wang, Pengkun Xu, Cong Su
GLOBECOM5