Tianjie Ju

dblp:319/4246 · DBLP profile ↗
← Back
14ranked-venue papers
4as first author
14since 2021 · last 2026
0009-0006-6978-1935ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 4 first-author · 11 since 2021Computer networks · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Dr.V : A Hierarchical Perception-Temporal-Cognition Framework to Diagnose Video Hallucination by Fine-Grained Spatial-Temporal Grounding
Meng Luo 0010, Shengqiong Wu, Liqiang Jing, Tianjie Ju, Jinxiang Lai, Tianlong Wu, Xinya Du, Siyuan Yan, Jiebo Luo 0001, William Yang Wang, Hao Fei 0001, Mong-Li Lee, Wynne Hsu
Int. J. Comput. Vis.4
2025 On Path to Multimodal Generalist: General-Level and General-Bench
abstract
The Multimodal Large Language Model (MLLM) is currently experiencing rapid growth, driven by the advanced capabilities of language-based LLMs. Unlike their specialist predecessors, existing MLLMs are evolving towards a Multimodal Generalist paradigm. Initially limited to understanding multiple modalities, these models have advanced to not only comprehend but also generate across modalities. Their capabilities have expanded from coarse-grained to fine-grained multimodal understanding and from supporting singular modalities to accommodating a wide array of or even arbitrary modalities. To assess the capabilities of various MLLMs, a diverse array of benchmark test sets has been proposed. This leads to a critical question: Can we simply assume that higher performance across tasks indicates a stronger MLLM capability, bringing us closer to human-level AI? We argue that the answer is not as straightforward as it seems. In this project, we introduce an evaluation framework to delineate the capabilities and behaviors of current multimodal generalists. This framework, named General-Level, establishes 5-scale levels of MLLM performance and generality, offering a methodology to compare MLLMs and gauge the progress of existing systems towards more robust multimodal generalists and, ultimately, towards AGI (Artificial General Intelligence). Central to our framework is the use of Synergy as the evaluative criterion, categorizing capabilities based on whether MLLMs preserve synergy across comprehension and generation, as well as across multimodal interactions. To evaluate the comprehensive abilities of various generalists, we present a massive multimodal benchmark, General-Bench, which encompasses a broader spectrum of skills, modalities, formats, and capabilities, including over 700 tasks and 325,800 instances. The evaluation results that involve over 100 existing state-of-the-art MLLMs uncover the capability rankings of generalists, highlighting the challenges in reaching genuine AI. We expect this project to pave the way for future research on next-generation multimodal foundation models, providing a robust infrastructure to accelerate the realization of AGI. Project Page: https://generalist.top/, Leaderboard: https://generalist.top/leaderboard/, Benchmark: https://huggingface.co/General-Level/.
Hao Fei 0001, Yuan Zhou 0016, Juncheng Li 0006, Xiangtai Li, Qingshan Xu 0001, Bobo Li 0001, Shengqiong Wu, Yaoting Wang, Junbao Zhou, Jiahao Meng, Liangtao Shi, Minghe Gao, Daoan Zhang, Zhiqi Ge, Siliang Tang, Kaihang Pan, Yaobo Ye, Haobo Yuan, Tao Zhang 0042, Weiming Wu, Tianjie Ju, Zixiang Meng, Shilin Xu 0001, Liyu Jia, Meng Luo 0010, Jiebo Luo 0001, Tat-Seng Chua, Shuicheng Yan, Hanwang Zhang
ICML23
2025 Watch Out Your Album! On the Inadvertent Privacy Memorization in Multi-Modal Large Language Models
abstract
Multi-Modal Large Language Models (MLLMs) have exhibited remarkable performance on various vision-language tasks such as Visual Question Answering (VQA). Despite accumulating evidence of privacy concerns associated with task-relevant content, it remains unclear whether MLLMs inadvertently memorize private content that is entirely irrelevant to the training tasks. In this paper, we investigate how randomly generated task-irrelevant private content can become spuriously correlated with downstream objectives due to partial mini-batch training dynamics, thus causing inadvertent memorization. Concretely, we randomly generate task-irrelevant watermarks into VQA fine-tuning images at varying probabilities and propose a novel probing framework to determine whether MLLMs have inadvertently encoded such content. Our experiments reveal that MLLMs exhibit notably different training behaviors in partial mini-batch settings with task-irrelevant watermarks embedded. Furthermore, through layer-wise probing, we demonstrate that MLLMs trigger distinct representational patterns when encountering previously seen task-irrelevant knowledge, even if this knowledge does not influence their output during prompting. Our code is available at https://github.com/illusionhi/ProbingPrivacy.
Tianjie Ju, Hao Fei 0001, Zhenyu Shao, Yubin Zheng, Haodong Zhao, Mong-Li Lee, Wynne Hsu, Zhuosheng Zhang 0001, Gongshen Liu
ICML1
2025 FedDEAP: Adaptive Dual-Prompt Tuning for Multi-Domain Federated Learning
abstract
Federated learning (FL) enables multiple clients to collaboratively train machine learning models without exposing local data, balancing performance and privacy. However, domain shift and label heterogeneity across clients often hinder the generalization of the aggregated global model. Recently, large-scale vision-language models like CLIP have shown strong zero-shot classification capabilities, raising the question of how to effectively fine-tune CLIP across domains in a federated setting. In this work, we propose an adaptive federated prompt tuning framework, FedDEAP, to enhance CLIP's generalization in multi-domain scenarios. Our method includes the following three key components: (1) To mitigate the loss of domain-specific information caused by label-supervised tuning, we disentangle semantic and domain-specific features in images by using semantic and domain transformation networks with unbiased mappings; (2) To preserve domain-specific knowledge during global prompt aggregation, we introduce a dual-prompt design with a global semantic prompt and a local domain prompt to balance shared and personalized information; (3) To maximize the inclusion of semantic and domain information from images in the generated text features, we align textual and visual representations under the two learned transformations to preserve semantic and domain consistency. Theoretical analysis and extensive experiments on four datasets demonstrate the effectiveness of our method in enhancing the generalization of CLIP for federated image recognition across multiple domains.
Yubin Zheng, Pak-Hei Yeung, Tianjie Ju, Peng Tang 0002, Weidong Qiu, Jagath C. Rajapakse
ACM Multimedia4
2024 Investigating Multi-Hop Factual Shortcuts in Knowledge Editing of Large Language Models
abstract
Tianjie Ju, Yijin Chen, Xinwei Yuan, Zhuosheng Zhang, Wei Du, Yubin Zheng, Gongshen Liu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Tianjie Ju, Xinwei Yuan, Zhuosheng Zhang 0001, Yubin Zheng, Gongshen Liu
ACL (1)1
2024 Federated Semi-supervised Learning for Medical Image Segmentation with Intra-client and Inter-client Consistency
abstract
Medical image segmentation plays a vital role in medical image analysis. However, it is impractical to build a large-scale centralized segmentation dataset due to the privacy of medical images. Federated learning (FL) aims to train a shared model of isolated clients without local data exchange which aligns well with the scarcity and privacy characteristics of medical images. Moreover, there is a large amount of unlabeled data in clients due to the difficulty in annotating medical images. Federated semi-supervised learning (FSSL) can leverage the unlabeled data of clients to improve the performance of the global model. Many existing FSSL methods apply the complicated semi-supervised learning protocols and some of them neglect the problem of data heterogeneity in FL. In this paper, we propose a novel federated semi-supervised learning framework for medical image segmentation incorporating intra-client and inter-client consistency learning. The intra-client consistency learning can introduce global data noise in data augmentation which can improve the generalization ability of the model and reduce the impact of data heterogeneity. The inter-client consistency learning is proposed to expand the feature search space and learn the ensemble knowledge of different clients. The two consistency learning mechanisms are achieved with the assistance of a Variational Autoencoder (VAE) trained collaboratively by clients. The experimental results illustrate that our method outperforms the state-of-the-art methods under different FSSL settings. The code is available at https://github.com/zyb98/FV2IC.
Yubin Zheng, Peng Tang 0002, Tianjie Ju, Weidong Qiu, Jagath C. Rajapakse
BIBM3
2024 Backdoor NLP Models via AI-Generated Text
abstract
Backdoor attacks pose a critical security threat to natural language processing (NLP) models by establishing covert associations between trigger patterns and target labels without affecting normal accuracy. Existing attacks usually disregard fluency and semantic fidelity of poisoned text, rendering the malicious data easily detectable. However, text generation models can produce coherent and content-relevant text given prompts. Moreover, potential differences between human-written and AI-generated text may be captured by NLP models while being imperceptible to humans. More insidious threats could arise if attackers leverage latent features of AI-generated text as trigger patterns. We comprehensively investigate backdoor attacks on NLP models using AI-generated poisoned text obtained via continued writing or paraphrasing, exploring three attack scenarios: data, model and pre-training. For data poisoning, we fine-tune generators with attribute control to enhance the attack performance. For model poisoning, we leverage downstream tasks to derive specialized generators. For pre-training poisoning, we train multiple attribute-based generators and align their generated text with pre-defined vectors, enabling task-agnostic migration attacks. Experiments demonstrate that our method achieves effective attacks while maintaining fluency and semantic similarity across all scenarios. We hope this work can raise awareness of the security risks hidden in AI-generated text.
Tianjie Ju, Gaolei Li, Gongshen Liu
LREC/COLING2
2024 How Large Language Models Encode Context Knowledge? A Layer-Wise Probing Study
abstract
Previous work has showcased the intriguing capability of large language models (LLMs) in retrieving facts and processing context knowledge. However, only limited research exists on the layer-wise capability of LLMs to encode knowledge, which challenges our understanding of their internal mechanisms. In this paper, we devote the first attempt to investigate the layer-wise capability of LLMs through probing tasks. We leverage the powerful generative capability of ChatGPT to construct probing datasets, providing diverse and coherent evidence corresponding to various facts. We employ \mathcal V-usable information as the validation metric to better reflect the capability in encoding context knowledge across different layers. Our experiments on conflicting and newly acquired knowledge show that LLMs: (1) prefer to encode more context knowledge in the upper layers; (2) primarily encode context knowledge within knowledge-related entity tokens at lower layers while progressively expanding more knowledge within other tokens at upper layers; and (3) gradually forget the earlier context knowledge retained within the intermediate layers when provided with irrelevant evidence. Code is publicly available at https://github.com/Jometeorie/probing_llama.
Tianjie Ju, Weiwei Sun 0001, Xinwei Yuan, Zhaochun Ren, Gongshen Liu
LREC/COLING1
2024 On the Robustness of Editing Large Language Models
abstract
Large language models (LLMs) have played a pivotal role in building communicative AI, yet they encounter the challenge of efficient updates.Model editing enables the manipulation of specific knowledge memories and the behavior of language generation without retraining.However, the robustness of model editing remains an open question.This work seeks to understand the strengths and limitations of editing methods, facilitating practical applications of communicative AI.We focus on three key research questions.RQ1: Can edited LLMs behave consistently resembling communicative AI in realistic situations?RQ2: To what extent does the rephrasing of prompts lead LLMs to deviate from the edited knowledge memory?RQ3: Which knowledge features are correlated with the performance and robustness of editing?Our empirical studies uncover a substantial disparity between existing editing methods and the practical application of LLMs.On rephrased prompts that are flexible but common in realistic applications, the performance of editing experiences a significant decline.Further analysis shows that more popular knowledge is memorized better, easier to recall, and more challenging to edit effectively.
Xinbei Ma, Tianjie Ju, Jiyang Qiu, Zhuosheng Zhang 0001, Hai Zhao 0001, Lifeng Liu, Yulong Wang 0004
EMNLP2
2024 Fine-Grained Contrastive Learning for Pulmonary Nodule Classification
abstract
Lung cancer is one of the most threatening human diseases which develops from malignant pulmonary nodules. The accurate classification of benign and malignant pulmonary nodules is important for formulating treatment plans and improving the survival rate of lung cancer patients. However, detecting and classifying pulmonary nodules pose a challenge due to their small region of interest (ROI) and diverse patterns. Many deep learning methods struggle to extract the valid feature representations of small pulmonary nodules which results in poor classification performance. In this work, we propose a novel fine-grained contrastive learning method for pulmonary nodule classification. Traditional contrastive learning utilizes positive and negative sample pairs to learn good representations of images. However, different pulmonary nodule patterns contain similar features that are vital for nodule classification. We propose using attributes which are categorial features labeled by experts to adjust the importance of sample pairs in contrastive learning. The method can guide the model to understand the degree of similarity and difference of pulmonary nodules and obtain their meaningful representations. Due to the small ROI of pulmonary nodules, we discard CNN backbone and use Vision Transformer (ViT) to learn the correlation of adjacent small-size slices. Compared with other deep learning methods, our method achieves the state-of-the-art performance in five metrics. In addition, we transfer the representations of pulmonary nodules learned by our model to a new dataset. The model obtains competitive performance without additional domain expertise which proves the prospect of our model in transfer learning.
Yubin Zheng, Peng Tang 0002, Tianjie Ju, Weidong Qiu
IJCNN3
2023 Neural Linguistic Steganography with Controllable Security
abstract
Information hiding is an art and science with a long history and is widely used in covert communication. There are many ways to hide secret data in image, audio, and video. However, relatively few systems can hide information in text. Generative text steganography is a promising topic in natural language text infor-mation hiding. Previous generative text steganography methods use a fixed candidate pool generation rule, and they cannot effec-tively control the security of the generated text. The perceptual-imperceptibility and statistical-imperceptibility conflict effect also causes the poor quality of the steganographic text generated by previous generative text steganography methods. Moreover, pre-vious generative text steganography approaches barely discuss the robustness of steganographic text. This paper proposes a security controllable text steganography method that can generate natural-looking steganographic text with a statistical distribution that matches the natural language distribution. The proposed method combines the metrics of per-ceptual-imperceptibility and statistical-imperceptibility to calcu-late the combined distortion. It selects the tokens with the smallest combined distortion to construct a candidate pool at each time step. Moreover, the maximum combined distortion threshold is set when embedding secret messages to ensure controllable security. We conducted several experiments to evaluate the proposed model from the perspectives of embedding rate, perceptual-impercepti-bility, statistical-imperceptibility, and anti-attack ability. The ex-perimental results show that the proposed method can generate smooth and readable steganographic sentences with good re-sistance to steganalysis and high robustness.
Tianhe Lu, Gongshen Liu, Ru Zhang 0002, Tianjie Ju
IJCNN4
2023 Robust Secret Data Hiding for Transformer-based Neural Machine Translation
abstract
Hiding secret information in text is a research area of significant importance and a great challenge. In recent years, there have been huge developments and exciting advances in generation-based text information hiding techniques. Current generative text information hiding methods mainly establish correspondence between token and secret bits based on probability distributions given by language models. However, the semantic control of such methods is weak, and their robustness is not discussed. In this paper, we investigate an end-to-end generation-based text information hiding scheme. The proposed method uses a sequence-to-sequence model with adversarial training as a machine translation model. It converts the secret information into an embedding vector to be added to each position of the hidden state representation of the source language text, which in turn allows the model to automatically learn to produce translation results with the embedded secret information without using fixed rules. The semantics of the text with embedded secret messages obtained by translation can be controlled by the meaning of the source language text. Our experiments show that the proposed method can embed the secret message into the translation results with little loss of the translation quality and is robust to active attacks such as word deletion or synonym substitution.
Tianhe Lu, Gongshen Liu, Ru Zhang 0002, Tianjie Ju
IJCNN5
2022 TIMS: A Novel Approach for Incrementally Few-Shot Text Instance Selection via Model Similarity
abstract
Large-scale pre-trained models' demand for high-quality instances forces people to consider how to select instances for annotation with limited resources. Nonetheless, little attention has been paid to the scenario where the number of instances that ultimately need to be annotated is agnostic. Meanwhile, the anisotropy of the sentence vector output by pre-trained models makes it hard to represent the instance itself well. Faced with the two challenges, we propose an incrementally few-shot instance selection approach (TIMS) based on model similarity and outlier detection, which suits the starting step of active learning well and serves as a better benchmark for few-shot learning. Specifically, TIMS determines the representative candidate set by calculating the similarity between changes in model parameters caused by each instance and by the full dataset. Meanwhile, Isolation Forest is adopted to select instances from the candidate set for annotation, which prevents selected instances from being too similar. Comprehensive experiments on WikiLingua & SQuAD show that TIMS outperforms other algorithms across almost every circumstance. It inspires us that the proper implementation of model similarity detection and outlier detection is of great help to select representative instances incrementally.
Tianjie Ju, Han Liao, Gongshen Liu
IJCNN1
2021 FLAG: Flow Representation Generator based on Self-supervised Learning for Encrypted Traffic Classification
abstract
Due to its excellent ability in learning features from large scale raw data, deep learning (DL) has attracted much attention for encrypted traffic classification. However, most DL-based traffic classifiers usually rely on enormous labeled samples. Motivated by this, we investigate a self-supervised traffic classifier (FLAG) without sacrifice of identification accuracy, only depending on small labeled traffic samples and highly available unlabeled traffic samples. Specifically, focusing on local short-term characteristics of traffic, we design a preprocessing algorithm, termed as N-phrase Extration, to convert unlabeled raw traffic dataset into sequences of high-frequency phrases as input of Bidirectional Encoder. On account of their significance, potential timing characteristics from input sequences are mined by Bidirectional Encoder and embedded into robust representations with distributed vectors to enhance classifier’s performance significantly. Our comprehensive experiments indicate FLAG can achieve 98.65% in 100% of dataset and 98.07% in 10% of dataset in terms of true positive rate in UNB ISCX VPN-nonVPN dataset, which are better than p-FP, FS-Net and Deep Packet.
Wenting Wei, Tianjie Ju, Han Liao, Weike Zhao, Huaxi Gu
APNet2