VLDB 2026 Research / reviewers in the wild / expert
Kang Gu
dblp:187/3055
· DBLP profile ↗
8ranked-venue papers
3as first author
5since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 1 first-author · 1 since 2021Security and privacy · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Gradient-Free Privacy Leakage in Federated Language Models through Selective Weight TamperingabstractFederated learning (FL) has become a key component in various language modeling applications such as machine translation, next-word prediction, and medical record analysis. These applications are trained on datasets from many FL participants that often include privacy-sensitive data, such as healthcare records, phone/credit card numbers, login credentials, etc. Although FL enables computation without necessitating clients to share their raw data, existing works show that privacy leakage is still probable in federated language models. In this paper, we present two novel findings on the leakage of privacy-sensitive user data from federated large language models without requiring access to gradients. Firstly, we make a key observation that model snapshots from the intermediate rounds in FL can cause greater privacy leakage than the final trained model. Secondly, we identify that a malicious FL participant can aggravate the leakage by tampering with the model's selective weights that are responsible for memorizing the sensitive training data of some other clients, even without any cooperation from the server. Our best-performing method increases the membership inference recall by 29% and achieves up to 71% private data reconstruction, evidently outperforming existing attacks that consider much stronger adversary capabilities. Lastly, we recommend a balanced suite of techniques for an FL client to defend against such privacy risk. Md. Rafi Ur Rashid, Vishnu Asutosh Dasu, Kang Gu, Najrin Sultana, Shagufta Mehnaz |
Proc. Priv. Enhancing Technol. | 3 |
| 2025 | Robust Unlearning for Large Language Models
Kang Gu, Md. Rafi Ur Rashid, Najrin Sultana, Shagufta Mehnaz |
PAKDD (5) | 1 |
| 2023 | DocGraphLM: Documental Graph Language Model for Information ExtractionabstractAdvances in Visually Rich Document Understanding (VrDU) have enabled information extraction and question answering over documents with complex layouts. Two tropes of architectures have emerged-transformer-based models inspired by LLMs, and Graph Neural Networks. In this paper, we introduce DocGraphLM, a novel framework that combines pre-trained language models with graph semantics. To achieve this, we propose 1) a joint encoder architecture to represent documents, and 2) a novel link prediction approach to reconstruct document graphs. DocGraphLM predicts both directions and distances between nodes using a convergent joint loss function that prioritizes neighborhood restoration and downweighs distant node detection. Our experiments on three SotA datasets show consistent improvement on IE and QA tasks with the adoption of graph features. Moreover, we report that adopting the graph features accelerates convergence in the learning process druing training, despite being solely constructed through link prediction. Dongsheng Wang 0005, Armineh Nourbakhsh, Kang Gu, Sameena Shah |
SIGIR | 4 |
| 2023 | Towards Sentence Level Inference Attack Against Pre-trained Language ModelsabstractIn recent years, pre-trained language models (e.g., BERT and GPT) have shown the superior capability of textual representation learning, benefiting from their large architectures and massive training corpora. The industry has also quickly embraced language models to develop various downstream NLP applications. For example, Google has already used BERT to improve its search system. The utility of the language embeddings also brings about potential privacy risks. Prior works have revealed that an adversary can either identify whether a keyword exists or gather a set of possible candidates for each word in a sentence embedding. However, these attacks cannot recover coherent sentences which leak high-level semantic information from the original text. To demonstrate that the adversary can go beyond the word-level attack, we present a novel decoder-based attack, which can reconstruct meaningful text from private embeddings after being pre-trained on a public dataset of the same domain. This attack is more challenging than a word-level attack due to the complexity of sentence structures. We comprehensively evaluate our attack in two domains and with different settings to show its superiority over the baseline attacks. Quantitative experimental results show that our attack can identify up to 3.5X of the number of keywords identified by the baseline attacks. Although our method reconstructs high-quality sentences in many cases, it often produces lower-quality sentences as well. We discuss these cases and the limitations of our method in detail Kang Gu, Ehsanul Kabir, Neha Ramsurrun, Soroush Vosoughi, Shagufta Mehnaz |
Proc. Priv. Enhancing Technol. | 1 |
| 2022 | Going Beyond Accuracy: Interpretability Metrics for CNN Representations of Physiological SignalsabstractThough deep neural networks such as convolutional neural networks (CNNs) have achieved high predictive accuracy on many tasks in health domain, the lack of explanation and reasoning behind the predictions can cause adverse consequences in many “high stakes” applications such as predicting mortality in ICU. Existing popular explanation methods, including GradCam, LIME, SHAP and so on, are designed for image and text data, but not for physiological data in the form of time series like electrocardiograms (ECG). Moreover, these methods can only generate interpretable visualizations for qualitative evaluations while cannot satisfy the need for quantitative evaluations. To fill this gap, we introduce a novel generalized framework for quantitatively explaining CNN representations for physiological signals called ENPhyS. Specifically, inspired by visual concepts (e.g. parts, objects and scenes), we first define physiological concepts which are drawn from existing clinical knowledge (e.g. P wave in ECG). ENPhys further measures the interpretability of CNN by the alignment between a hidden CNN unit and each physiological concept. Finally, we propose novel interpretability metrics —Consistency, Confidence, Completeness and C-Score —to quantify the interpretability of CNNs from different perspectives. We evaluate our metrics on three state-of-the-art CNN models and on three types of physiological signals. Results from user study suggest that our metrics are in line with human intuition of interpretability. Kang Gu, Temiloluwa Prioleau, Soroush Vosoughi |
ICPR | 1 |
| 2020 | Gesture Recognition Through sEMG with Wearable Device Based on Deep Learning
Shu Shen, Kang Gu, Xinrong Chen, Caixia Lv, Ruchuan Wang 0001 |
Mob. Networks Appl. | 2 |
| 2018 | Spatial Pyramid Pooling Mechanism in 3D Convolutional Network for Sentence-Level ClassificationabstractIn this paper, we investigate the usage of the convolutional neural network (CNN) to propose a novel end-to-end language processing structure to model textual data for this task. In particular, we propose a 3D CNN structure for the task, which is featured by spatial pyramid pooling (SPP). To our knowledge, it is the first time that 3D convolution and SPP structure are applied together in language processing issues. Compared with methods of 2D CNNs, the proposed method can effectively and efficiently capture the complicated internal relations in sentences. Furthermore, in previous work, the issue of sentence length variety is usually addressed by padding zero to make all sentences vectors to a fixed length, which causes too much redundant and useless noise. Inspired by the SPP structure for object detection in image processing, this issue can be well handled with the SPP, which divides the sentences into several length sections for respective pooling processing. Experiments are conducted for the task of sentence classification as well as relation classification. Experiments on Stanford Treebank, TREC, subj, and Yelp datasets demonstrate that our proposed method can outperform other state-of-the-art models, with respect to classification accuracy. Auxiliary attempts to leverage our method to SemEval-2010 Task 8 dataset further substantiate the model's capability of extracting features efficiently. Xi Ouyang, Kang Gu, Pan Zhou 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2017 | Jointly Attentive Spatial-Temporal Pooling Networks for Video-Based Person Re-identificationabstractPerson Re-Identification (person re-id) is a crucial task as its applications in visual surveillance and human-computer interaction. In this work, we present a novel joint Spatial and Temporal Attention Pooling Network (ASTPN) for video-based person re-identification, which enables the feature extractor to be aware of the current input video sequences, in a way that interdependency from the matching items can directly influence the computation of each other's representation. Specifically, the spatial pooling layer is able to select regions from each frame, while the attention temporal pooling performed can select informative frames over the sequence, both pooling guided by the information from distance matching. Experiments are conduced on the iLIDS-VID, PRID-2011 and MARS datasets and the results demonstrate that this approach outperforms existing state-of-art methods. We also analyze how the joint pooling in both dimensions can boost the person re-id performance more effectively than using either of them separately 1. Shuangjie Xu, Yu Cheng 0001, Kang Gu, Yang Yang 0002, Shiyu Chang, Pan Zhou 0001 |
ICCV | 3 |