Jinhe Bi

dblp:361/7122 · DBLP profile ↗
← Back
9ranked-venue papers
1as first author
9since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 1 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Efficient and distributed learning · 27% Language models and text generation · 24% Trustworthy machine learning · 14%
Human-computer interaction and pervasive computing
1 paper
Health and well-being technologies · 100%

Topics — the 27 heaviest of 30, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
federated learning
2.032026
HYPERION: Fine-Grained Hypersphere Alignment for Robust Federated Graph Learning · NeurIPS 2025
FedBiP: Heterogeneous One-Shot Federated Learning with Personalized Latent Diffusion Models · CVPR 2025
DAWN: Distributed LLM Multi-Agent Workflow Synthesis · AAAI 2026
Computer vision › Vision and language › vision-language model
multimodal large language model
2.022026
ASCD: Attention-Steerable Contrastive Decoding for Reducing Hallucination in MLLM · AAAI 2026
AUVIC: Adversarial Unlearning of Visual Concepts for Multi-modal Large Language Models · AAAI 2026
Machine learning › Efficient and distributed learning › federated learning
federated graph learning
1.222026
HYPERION: Fine-Grained Hypersphere Alignment for Robust Federated Graph Learning · NeurIPS 2025
DAWN: Distributed LLM Multi-Agent Workflow Synthesis · AAAI 2026
Natural language and speech › Language models and text generation › model steering › language model steering
attention steering
1.012026
ASCD: Attention-Steerable Contrastive Decoding for Reducing Hallucination in MLLM · AAAI 2026
Natural language and speech › Language models and text generation › decoding
contrastive decoding
1.012026
ASCD: Attention-Steerable Contrastive Decoding for Reducing Hallucination in MLLM · AAAI 2026
Natural language and speech › Language models and text generation › decoding
decoding strategy
1.012026
ASCD: Attention-Steerable Contrastive Decoding for Reducing Hallucination in MLLM · AAAI 2026
Natural language and speech › Language models and text generation
hallucination mitigation
1.012026
ASCD: Attention-Steerable Contrastive Decoding for Reducing Hallucination in MLLM · AAAI 2026
Machine learning › Trustworthy machine learning
interpretability
1.012026
ASCD: Attention-Steerable Contrastive Decoding for Reducing Hallucination in MLLM · AAAI 2026
Natural language and speech › Language models and text generation
LLM agents
1.012026
Memory-R1: Enhancing Large Language Model Agents to Manage and Utilize Memories via Reinforcement Learning · ACL (1) 2026
Machine learning › Trustworthy machine learning
machine unlearning
1.012026
AUVIC: Adversarial Unlearning of Visual Concepts for Multi-modal Large Language Models · AAAI 2026
Machine learning › Efficient and distributed learning
memory management
1.012026
Memory-R1: Enhancing Large Language Model Agents to Manage and Utilize Memories via Reinforcement Learning · ACL (1) 2026
Natural language and speech › Language models and text generation
retrieval-augmented generation
1.012026
PsyPARSE: Retrieval-Augmented Slow Thinking for Personalized Empathetic Counseling · AAAI 2026
Machine learning › Generative modeling
diffusion model
0.912025
FedBiP: Heterogeneous One-Shot Federated Learning with Personalized Latent Diffusion Models · CVPR 2025
Machine learning › Graph learning
graph neural network
0.912025
HYPERION: Fine-Grained Hypersphere Alignment for Robust Federated Graph Learning · NeurIPS 2025
Machine learning › Efficient and distributed learning › federated learning
heterogeneous federated learning
0.912025
FedBiP: Heterogeneous One-Shot Federated Learning with Personalized Latent Diffusion Models · CVPR 2025
Machine learning › Representation and self-supervised learning › representation learning › embedding learning › geometric embedding
hyperspherical embedding
0.912025
HYPERION: Fine-Grained Hypersphere Alignment for Robust Federated Graph Learning · NeurIPS 2025
Machine learning › Generative modeling › diffusion model
latent diffusion model
0.912025
FedBiP: Heterogeneous One-Shot Federated Learning with Personalized Latent Diffusion Models · CVPR 2025
Machine learning › Trustworthy machine learning › robustness
learning with noisy labels
0.912025
HYPERION: Fine-Grained Hypersphere Alignment for Robust Federated Graph Learning · NeurIPS 2025
Machine learning › Efficient and distributed learning › federated learning › communication-efficient federated learning
one-shot federated learning
0.912025
FedBiP: Heterogeneous One-Shot Federated Learning with Personalized Latent Diffusion Models · CVPR 2025
Machine learning › Efficient and distributed learning
parameter-efficient fine-tuning
0.912025
LLaVA Steering: Visual Instruction Tuning with 500x Fewer Parameters through Modality Linear Representation-Steering · ACL (1) 2025
Machine learning › Generative modeling › diffusion model › diffusion model adaptation
personalized diffusion model
0.912025
FedBiP: Heterogeneous One-Shot Federated Learning with Personalized Latent Diffusion Models · CVPR 2025
Machine learning › Trustworthy machine learning
robustness
0.912025
HYPERION: Fine-Grained Hypersphere Alignment for Robust Federated Graph Learning · NeurIPS 2025
Computer vision › Vision and language › vision-language model › multimodal large language model
visual instruction tuning
0.912025
LLaVA Steering: Visual Instruction Tuning with 500x Fewer Parameters through Modality Linear Representation-Steering · ACL (1) 2025
Natural language and speech › Language models and text generation › LLM agents
tool use
0.312026
Memory-R1: Enhancing Large Language Model Agents to Manage and Utilize Memories via Reinforcement Learning · ACL (1) 2026
Health and well-being technologies › mental health
mental health support
0.312026
PsyPARSE: Retrieval-Augmented Slow Thinking for Personalized Empathetic Counseling · AAAI 2026
Privacy and data protection › privacy regulation
right to be forgotten
0.312026
AUVIC: Adversarial Unlearning of Visual Concepts for Multi-modal Large Language Models · AAAI 2026
Machine learning › Efficient and distributed learning › federated learning
data heterogeneity
0.312025
FedBiP: Heterogeneous One-Shot Federated Learning with Personalized Latent Diffusion Models · CVPR 2025

Methods — techniques the papers use, named apart from their topics

multi-turn rollouts · 2.0machine unlearning · 2.0large language model · 2.0adversarial perturbation · 2.0structural gravity · 1.0patient-therapist agent simulation · 1.0parametric resonance · 1.0gromov-wasserstein distance · 1.0contrastive decoding · 1.0attention steering · 1.0SVD-based denoising · 1.0
YearPublicationVenuePosition
2026 AUVIC: Adversarial Unlearning of Visual Concepts for Multi-modal Large Language Models
abstract
Multimodal Large Language Models (MLLMs) achieve impressive performance once optimized on massive datasets. Such datasets often contain sensitive or copyrighted content, raising significant data privacy concerns. Regulatory frameworks mandating the 'right to be forgotten' drive the need for machine unlearning. This technique allows for the removal of target data without resource-consuming retraining. However, while well-studied for text, visual concept unlearning in MLLMs remains underexplored. A primary challenge is precisely removing a target visual concept without disrupting model performance on related entities. To address this, we introduce AUVIC, a novel visual concept unlearning framework for MLLMs. AUVIC applies adversarial perturbations to enable precise forgetting. This approach effectively isolates the target concept while avoiding unintended effects on similar entities. To evaluate our method, we construct VCUBench. It is the first benchmark designed to assess visual concept unlearning in group contexts. Experimental results demonstrate that AUVIC achieves state-of-the-art target forgetting rates while incurs minimal performance degradation on non-target concepts.
Jinhe Bi, Yan Xia 0003, Jindong Gu, Volker Tresp
AAAI4
2026 DAWN: Distributed LLM Multi-Agent Workflow Synthesis
abstract
Large language models (LLMs) have recently empowered multi-agent systems (MAS) to achieve remarkable advances in collaborative reasoning and complex task automation. The effectiveness of these systems fundamentally depends on the design of adaptive communication graphs—the underlying workflows that coordinate agent interactions. However, in real-world scenarios, strict privacy constraints often silo data across organizations, and client distributions are highly non-IID, posing major challenges for synthesizing such workflows. In this work, we are the first to systematically study distributed multi-agent workflow synthesis under these privacy and heterogeneity constraints, and we introduce the Difficulty-Based Skew (DBS) benchmark to emulate such challenging environments. Drawing inspiration from federated graph learning (FGL)—which has primarily focused on classification over static graphs—we identify a critical gap: existing FGL methods do not address the generative design of communication topologies. We reveal two fundamental obstacles to generative workflow synthesis in this setting: (i) workflow specialization conflict, where agents optimized for different task distributions generate incompatible communication patterns that resist meaningful aggregation, and (ii) structural communication shift, where locally optimal agent interaction graphs fail to compose into globally coherent multi-agent workflows. To address these challenges, we propose DAWN, a federated framework that integrates two key innovations: Parametric Resonance, which robustly aggregates heterogeneous local updates via layer-wise SVD-based denoising and alignment, and Structural Gravity, which regularizes local workflow generation by penalizing the Fusion Gromov-Wasserstein distance to a set of prototype communication graphs, ensuring global structural coherence without stifling local adaptation. Experiments on the DBS benchmark show that DAWN surpasses baselines in global task success and reduces inter-client graph divergence, laying a solid foundation for privacy-preserving, adaptive MAS workflow design in heterogeneous settings.
Guancheng Wan, Xiaoran Shang, Eric Hanchen Jiang, Guibin Zhang, Jinhe Bi, Yunpu Ma, Zaixi Zhang, Ke Liang 0006, Wenke Huang 0003
AAAI7
2026 ASCD: Attention-Steerable Contrastive Decoding for Reducing Hallucination in MLLM
abstract
Multimodal large language models (MLLMs) frequently hallucinate by over-committing to spurious visual cues. Prior remedies–Visual and Instruction Contrastive Decoding (VCD, ICD)–mitigate this issue, yet the mechanism remains opaque. We first empirically show that their improvements systematically coincide with redistributions of cross-modal attention. Building on this insight, we propose Attention-Steerable Contrastive Decoding (ASCD), which directly steers the attention scores during decoding. ASCD combines (i) positive steering, which amplifies automatically mined text-centric heads–stable within a model and robust across domains–with (ii) negative steering, which dampens on-the-fly identified critical visual tokens. The method incurs negligible runtime/memory overhead and requires no additional training. Across five MLLM backbones and three decoding schemes, ASCD reduces hallucination on POPE, CHAIR, and MMHal-Bench by up to 38.2% while improving accuracy on standard VQA benchmarks, including MMMU, MM-VET, ScienceQA, TextVQA, and GQA. These results position attention steering as a simple, model-agnostic, and principled route to safer, more faithful multimodal generation.
Aniri, Jinhe Bi, Sören Pirk, Yunpu Ma
AAAI3
2026 PsyPARSE: Retrieval-Augmented Slow Thinking for Personalized Empathetic Counseling
abstract
The escalating global demand for mental health services highlights the potential of Large Language Models (LLMs) in psychological counseling. However, current LLM-based approaches, particularly fine-tuned models, are constrained by data distribution biases, leading to limited therapeutic diversity and personalization. Crucially, they often lack anticipatory empathetic reasoning, struggle to foresee patient emotional responses beyond immediate dialogue history, and incur substantial computational costs. To address these limitations, we propose PsyPARSE, a novel training-free framework for psychological counseling that emulates the deliberate and empathetic reasoning of human counselors. PsyPARSE integrates Multi-Therapy Retrieval-Augmented Generation (RAG) to overcome data biases and provide highly personalized therapeutic approaches tailored to individual patient attributes. Pioneering the first multi-stage slow-thinking engine in mental health LLMs, PsyPARSE employs Multi-Turn Rollouts to identify optimal therapeutic paths and through anticipating patient reactions, optimizes empathetic responses, thereby ensuring genuinely empathetic and impactful responses in complex, long-dialogue interactions. Operating as a plug-and-play solution, PsyPARSE avoids the computational burden of fine-tuning. We establish a comprehensive LLM-based patient-therapist agent simulation framework for evaluation. Extensive experiments demonstrate that PsyPARSE significantly enhances the capabilities of various LLM baselines, achieving superior personalization and deeper empathy compared to both fine-tuned and other training-free methods. This work offers an efficient, adaptable, and scalable solution to advance mental health support.
Pukun Zhao, Jinhe Bi, Huacan Wang, Ronghao Chen
AAAI4
2026 Memory-R1: Enhancing Large Language Model Agents to Manage and Utilize Memories via Reinforcement Learning
abstract
Sikuan Yan, Xiufeng Yang, Zuchao Huang, Ercong Nie, Zifeng Ding, Zonggen Li, Xiaowen Ma, Jinhe Bi, Kristian Kersting, Jeff Z. Pan, Hinrich Schuetze, Volker Tresp, Yunpu Ma. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Sikuan Yan, Xiufeng Yang, Zuchao Huang, Ercong Nie, Zifeng Ding, Zonggen Li, Xiaowen Ma 0001, Jinhe Bi, Kristian Kersting, Jeff Z. Pan, Hinrich Schütze, Volker Tresp, Yunpu Ma
ACL (1)8
2025 LLaVA Steering: Visual Instruction Tuning with 500x Fewer Parameters through Modality Linear Representation-Steering
abstract
Jinhe Bi, Yujun Wang, Haokun Chen, Xun Xiao, Artur Hecker, Volker Tresp, Yunpu Ma. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Jinhe Bi, Xun Xiao, Artur Hecker, Volker Tresp, Yunpu Ma
ACL (1)1
2025 FedBiP: Heterogeneous One-Shot Federated Learning with Personalized Latent Diffusion Models
abstract
One-Shot Federated Learning (OSFL), a special decentralized machine learning paradigm, has recently gained significant attention. OSFL requires only a single round of client data or model upload, which reduces communication costs and mitigates privacy threats compared to traditional FL. Despite these promising prospects, existing methods face challenges due to client data heterogeneity and limited data quantity when applied to real-world OSFL systems. Recently, Latent Diffusion Models (LDM) have shown remarkable advancements in synthesizing high-quality images through pretraining on large-scale datasets, thereby presenting a potential solution to overcome these issues. However, directly applying pretrained LDM to heterogeneous OSFL results in significant distribution shifts in synthetic data, leading to performance degradation in classification models trained on such data. This issue is particularly pronounced in rare domains, such as medical imaging, which are underrepresented in LDM’s pretraining data. To address this challenge, we propose Federated Bi-Level Personalization (FedBiP), which personalizes the pretrained LDM at both instance-level and concept-level. Hereby, FedBiP synthesizes images following the client’s local data distribution without compromising the privacy regulations. FedBiP is also the first approach to simultaneously address feature space heterogeneity and client data scarcity in OSFL. Our method is validated through extensive experiments on three OSFL benchmarks with feature space heterogeneity, as well as on challenging medical and satellite image datasets with label heterogeneity. The results demonstrate the effectiveness of FedBiP, which substantially outperforms other OSFL methods. Our code is available at https://github.com/HaokunChen245/FedBiP.
Hang Li 0010, Jinhe Bi, Gengyuan Zhang, Philip Torr 0001, Jindong Gu, Denis Krompass, Volker Tresp
CVPR4
2025 Backdoor Cleaning without External Guidance in MLLM Fine-tuning
abstract
Multimodal Large Language Models (MLLMs) are increasingly deployed in fine-tuning-as-a-service (FTaaS) settings, where user-submitted datasets adapt general-purpose models to downstream tasks. This flexibility, however, introduces serious security risks, as malicious fine-tuning can implant backdoors into MLLMs with minimal effort. In this paper, we observe that backdoor triggers systematically disrupt cross-modal processing by causing abnormal attention concentration on non-semantic regions—a phenomenon we term **attention collapse**. Based on this insight, we propose **Believe Your Eyes (BYE)**, a data filtering framework that leverages attention entropy patterns as self-supervised signals to identify and filter backdoor samples. BYE operates via a three-stage pipeline: (1) extracting attention maps using the fine-tuned model, (2) computing entropy scores and profiling sensitive layers via bimodal separation, and (3) performing unsupervised clustering to remove suspicious samples. Unlike prior defenses, BYE equires no clean supervision, auxiliary labels, or model modifications. Extensive experiments across various datasets, models, and diverse trigger types validate BYE's effectiveness: it achieves near-zero attack success rates while maintaining clean-task performance, offering a robust and generalizable solution against backdoor threats in MLLMs.
Xuankun Rong, Wenke Huang 0003, Jian Liang 0003, Jinhe Bi, Xun Xiao, Yiming Li 0004, Bo Du 0001, Mang Ye
NeurIPS4
2025 HYPERION: Fine-Grained Hypersphere Alignment for Robust Federated Graph Learning
abstract
Robust Federated Graph Learning (FGL) provides an effective decentralized framework for training Graph Neural Networks (GNNs) in noisy-label environments. However, the subtlety of noise during training presents formidable obstacles for developing robust FGL systems. Previous robust FL approaches neither adequately constrain edge-mediated error propagation nor account for intra-class topological differences. At the client level, we innovatively demonstrate that hyperspherical embedding can effectively capture graph structures in a fine-grained manner. Correspondingly, our method effectively addresses the aforementioned issues through fine-grained hypersphere alignment. Moreover, we uncover undetected noise arising from localized perspective constraints and propose the geometric-aware hyperspherical purification module at the server level. Combining both level strategies, we present our robust FGL framework,**HYPERION**, which operates all components within a unified hyperspherical space. **HYPERION** demonstrates remarkable robustness across multiple datasets, for instance, achieving a 29.7\% $\uparrow$ F1-macro score with 50\%-pair noise on Cora. The code is available for anonymous access at \url{https://anonymous.4open.science/r/Hyperion-NeurIPS/}.
Frank Wan, Xiaoran Shang, Guibin Zhang, Jinhe Bi, Liangtao Zheng, Yanbiao Ma, Wenke Huang 0003, Bo Du 0001
NeurIPS5