Jindong Li 0002

dblp:38/10174-2 · DBLP profile ↗
← Back
9ranked-venue papers
3as first author
9since 2021 · last 2026
0009-0007-2228-3696ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 3 first-author · 9 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Transfer learning and domain adaptation · 32% Vision and language · 24% Language models and text generation · 18%
Human-computer interaction and pervasive computing
1 paper
Human-AI interaction · 100%

Topics — the 15 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language
vision-language model
1.122026
ScreenAgent: A Vision Language Model-driven Computer Control Agent · IJCAI 2024
Beyond Retraining: Training-Free Unknown Class Filtering for Source-Free Open Set Domain Adaptation of Vision-Language Models · AAAI 2026
Machine learning › Representation and self-supervised learning
associative memory
1.012026
HeLa-Mem: Hebbian Learning and Associative Memory for LLM Agents · ACL (1) 2026
Machine learning › Representation and self-supervised learning › representation learning › discrete representation learning
discrete tokenization
1.012026
Discrete Tokenization for Multimodal LLMs: A Comprehensive Survey · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Machine learning › Transfer learning and domain adaptation
domain adaptation
1.012026
CLIP-Powered Domain Generalization and Domain Adaptation: A Comprehensive Survey · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Machine learning › Transfer learning and domain adaptation
domain generalization
1.012026
CLIP-Powered Domain Generalization and Domain Adaptation: A Comprehensive Survey · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Natural language and speech › Language models and text generation
LLM agents
1.012026
HeLa-Mem: Hebbian Learning and Associative Memory for LLM Agents · ACL (1) 2026
Natural language and speech › Language models and text generation › LLM agents
long-term memory
1.012026
HeLa-Mem: Hebbian Learning and Associative Memory for LLM Agents · ACL (1) 2026
Computer vision › Vision and language › vision-language model
multimodal large language model
1.012026
Discrete Tokenization for Multimodal LLMs: A Comprehensive Survey · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Machine learning › Transfer learning and domain adaptation › domain adaptation
open-set domain adaptation
1.012026
Beyond Retraining: Training-Free Unknown Class Filtering for Source-Free Open Set Domain Adaptation of Vision-Language Models · AAAI 2026
Machine learning › Transfer learning and domain adaptation › domain adaptation › open-set domain adaptation
source-free open-set domain adaptation
1.012026
Beyond Retraining: Training-Free Unknown Class Filtering for Source-Free Open Set Domain Adaptation of Vision-Language Models · AAAI 2026
Machine learning › Trustworthy machine learning › open-world recognition › open-set recognition
unknown class detection
1.012026
Beyond Retraining: Training-Free Unknown Class Filtering for Source-Free Open Set Domain Adaptation of Vision-Language Models · AAAI 2026
Computer vision › Vision and language › vision-language model
CLIP
0.622026
CLIP-Powered Domain Generalization and Domain Adaptation: A Comprehensive Survey · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Beyond Retraining: Training-Free Unknown Class Filtering for Source-Free Open Set Domain Adaptation of Vision-Language Models · AAAI 2026
Natural language and speech › Language models and text generation
large language model
0.312026
Discrete Tokenization for Multimodal LLMs: A Comprehensive Survey · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Machine learning › Reinforcement learning
memory architectures
0.312026
HeLa-Mem: Hebbian Learning and Associative Memory for LLM Agents · ACL (1) 2026
Computer vision › Vision and language
vision-language pretraining
0.312026
CLIP-Powered Domain Generalization and Domain Adaptation: A Comprehensive Survey · IEEE Trans. Pattern Anal. Mach. Intell. 2026

Methods — techniques the papers use, named apart from their topics

vision-language model · 1.5vector quantization · 1.0survey · 1.0knowledge distillation · 1.0hebbian learning · 1.0graph-based memory · 1.0gaussian mixture model · 1.0codebook learning · 1.0box-cox transform · 1.0SVD · 1.0
YearPublicationVenuePosition
2026 Beyond Retraining: Training-Free Unknown Class Filtering for Source-Free Open Set Domain Adaptation of Vision-Language Models
abstract
Vision-language models (VLMs) have gained widespread attention for their strong zero-shot capabilities across numerous downstream tasks. However, these models assume that each test image’s class label is drawn from a predefined label set and lack a reliable mechanism to reject samples from emerging unknown classes when only unlabeled data are available. To address this gap, open-set domain adaptation methods retrain models to push potential unknowns away from known clusters. Yet, some unknown samples remain stably anchored to specific known classes in the VLM feature space due to semantic relevance, which is termed as Semantic Affinity Anchoring (SAA). Forcibly repelling these samples unavoidably distorts the native geometry of VLMs and degrades performance. Meanwhile, existing score‑based unknown detectors use simplistic thresholds and suffer from threshold sensitivity, resulting in sub‑optimal performance. To address aforementioned issues, we propose VLM-OpenXpert, which comprises two training‑free, plug‑and‑play inference modules. SUFF performs SVD on high-confidence unknowns to extract a low-rank "unknown subspace". Each sample’s projection onto this subspace is weighted and softly removed from its feature, suppressing unknown components while preserving semantics. BGAT corrects score skewness via a Box–Cox transform, then fits a bimodal Gaussian mixture to adaptively estimate the optimal threshold balancing known-class recognition and unknown-class rejection. Experiments on 9 benchmarks and three backbones (CLIP, SigLIP, ALIGN) under Source-Free OSDA settings show that our training-free pipeline matches or outperforms retraining-heavy state-of-the-art methods, establishing a powerful lightweight inference calibration paradigm for open-set VLM deployment.
Yongguang Li, Jindong Li 0002, Qi Wang 0078, Qianli Xing 0002, Runliang Niu, Sheng-Sheng Wang 0001, Menglin Yang 0001
AAAI2
2026 HeLa-Mem: Hebbian Learning and Associative Memory for LLM Agents
abstract
Long-term memory is a critical challenge for Large Language Model agents, as fixed context windows cannot preserve coherence across extended interactions.Existing memory systems encode conversation history as embedding vectors and retrieve information through semantic similarity.This paradigm fails to capture the associative structure of human memory, wherein related experiences progressively strengthen interconnections through repeated co-activation.Inspired by cognitive neuroscience, we identify three mechanisms central to biological memory: association, consolidation, and spreading activation, which remain largely absent in current research.To bridge this gap, we propose HeLa-Mem, a bio-inspired memory architecture that models memory as a dynamic graph with Hebbian learning dynamics.HeLa-Mem employs a dual-level organization: (1) an episodic memory graph that evolves through co-activation patterns, and (2) a semantic memory store populated via Hebbian Distillation, wherein a Reflective Agent identifies densely connected memory hubs and distills them into structured, reusable semantic knowledge.This dual-path design leverages both semantic similarity and learned associations, mirroring the episodic-semantic distinction in human cognition.Experiments on LoCoMo demonstrate superior performance across four question categories while using significantly fewer context tokens.Code is available on GitHub.
Jinchang Zhu, Jindong Li 0002, Jiahong Liu 0001, Menglin Yang 0001
ACL (1)2
2026 Data-efficient CLIP-powered dual-branch networks for source-free unsupervised domain adaptation
Yongguang Li, Yueqi Cao, Jindong Li 0002, Qi Wang 0078, Sheng-Sheng Wang 0001
Expert Syst. Appl.3
2026 HC-GLAD: Dual hyperbolic contrastive learning for unsupervised graph-level anomaly detection
Yali Fu, Jindong Li 0002, Jiahong Liu 0001, Qianli Xing 0002, Qi Wang 0078, Irwin King
Neural Networks2
2026 Discrete Tokenization for Multimodal LLMs: A Comprehensive Survey
abstract
The rapid advancement of large language models (LLMs) has intensified the need for effective mechanisms to transform continuous multimodal data into discrete representations suitable for language-based processing. Discrete tokenization, with vector quantization (VQ) as a central approach, offers both computational efficiency and compatibility with LLM architectures. Despite its growing importance, there is a lack of a comprehensive survey that systematically examines VQ techniques in the context of LLM-based systems. This work fills this gap by presenting the first structured taxonomy and analysis of discrete tokenization methods designed for LLMs. We categorize 8 representative VQ variants that span classical and modern paradigms and analyze their algorithmic principles, training dynamics, and integration challenges with LLM pipelines. Beyond algorithm-level investigation, we discuss existing research in terms of classical applications without LLMs, LLM-based single-modality systems, and LLM-based multimodal systems, highlighting how quantization strategies influence alignment, reasoning, and generation performance. In addition, we identify key challenges including codebook collapse, unstable gradient estimation, and modality-specific encoding constraints. Finally, we discuss emerging research directions such as dynamic and task-adaptive quantization, unified tokenization frameworks, and biologically inspired codebook learning. This survey bridges the gap between traditional vector quantization and modern LLM applications, serving as a foundational reference for the development of efficient and generalizable multimodal systems.
Jindong Li 0002, Yali Fu, Jiahong Liu 0001, Linxiao Cao, Wei Ji 0008, Menglin Yang 0001, Irwin King, Ming-Hsuan Yang 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2026 CLIP-Powered Domain Generalization and Domain Adaptation: A Comprehensive Survey
abstract
As machine learning evolves, domain generalization (DG) and domain adaptation (DA) have become crucial for improving model robustness across diverse environments. Contrastive Language-Image Pretraining (CLIP) plays a central role in these tasks, offering strong zero-shot capabilities that allow models to operate effectively in unseen domains. Yet, despite CLIP's growing influence, no comprehensive survey has systematically examined its applications in DG and DA, underscoring the need for this review. This survey provides a unified and in-depth overview of CLIP-driven DG and DA. Before reviewing methods, we establish precise and complete scenario definitions covering source accessibility (SA vs. SF), source number (SS vs. MS), and label relations (CS, PS, OS, OPS), forming a coherent taxonomy that structures all subsequent analyses. For DG, we categorize methods into prompt optimization techniques that enhance task alignment and architectures that leverage CLIP as a backbone for transferable feature extraction. For DA, we examine both source-available approaches that rely on labeled source data and source-free approaches operating primarily on target-domain samples, emphasizing the knowledge transfer mechanisms that enable adaptation across heterogeneous settings. We further provide consolidated trend analyses for both DG and DA, revealing overarching patterns, methodological principles, and scenario-dependent behaviors. We then discuss key challenges such as realistic deployment scenarios, LLM knowledge integration, multimodal fusion, interpretability, and catastrophic forgetting, and outline future directions for developing scalable and trustworthy CLIP-based DG and DA systems. By synthesizing existing studies and highlighting critical gaps, this survey offers actionable insights for researchers and practitioners, motivating new strategies for leveraging CLIP to advance domain robustness in real-world scenarios.
Jindong Li 0002, Yongguang Li, Yali Fu, Jiahong Liu 0001, Yixin Liu 0001, Menglin Yang 0001, Irwin King
IEEE Trans. Pattern Anal. Mach. Intell.1
2025 GLADMamba: Unsupervised Graph-Level Anomaly Detection Powered by Selective State Space Model
Yali Fu, Jindong Li 0002, Qi Wang 0078, Qianli Xing 0002
ECML/PKDD (1)2
2024 ScreenAgent: A Vision Language Model-driven Computer Control Agent
Runliang Niu, Jindong Li 0002, Shiqi Wang 0006, Yali Fu, Xiyu Hu, Xueyuan Leng, He Kong 0004, Yi Chang 0001, Qi Wang 0078
IJCAI2
2023 CVTGAD: Simplified Transformer with Cross-View Attention for Unsupervised Graph-Level Anomaly Detection
Jindong Li 0002, Qianli Xing 0002, Qi Wang 0078, Yi Chang 0001
ECML/PKDD (1)1