Xinting Hu

dblp:222/7753 · DBLP profile ↗
← Back
20ranked-venue papers
7as first author
18since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 6 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 first-author · 9 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Benchmarking TCR-pMHC structure prediction: a unified evaluation and CDR3-based functional insights
abstract
Interactions between T cell receptors (TCRs) and peptide-major histocompatibility complexes (pMHCs) are central to adaptive immunity. Recent advances in structure prediction tools have enabled atomic-level modeling of TCR-pMHC interactions. However, the lack of systematic evaluation forces practitioners to invest substantial resources in selecting appropriate tools. Here, we present a comprehensive benchmark of TCR-pMHC structure prediction with 70 previously unseen complexes and 13 models spanning MSA-based, PLM-based, and docking-based approaches, revealing the superior modeling accuracy and docking quality of MSA-based methods, especially AlphaFold3. To further enhance the utility of AlphaFold3 predictions, we identify the pLDDT score of the TCR CDR3 region as an informative indicator of both structural correctness and functional relevance. Specifically, it enables up to 4.3% Top-1 success gain through reranking and captures mutation-induced affinity changes in 75.3% of cases. Overall, our analysis would facilitate the practical usage of immune structure prediction models and guide the advancement of these models.
Jiadong Lu, Xinyuan Zhu, Xinting Hu, Fuli Feng
Briefings Bioinform.3
2025 Unlearning Concepts in Diffusion Model via Concept Domain Correction and Concept Preserving Gradient
abstract
Text-to-image diffusion models have achieved remarkable success in generating photorealistic images. However, the inclusion of sensitive information during pre-training poses significant risks. Machine Unlearning (MU) offers a promising solution to eliminate sensitive concepts from these models. Despite its potential, existing MU methods face two main challenges: 1) limited generalization, where concept erasure is effective only within the unlearned set, failing to prevent sensitive concept generation from out-of-set prompts; and 2) utility degradation, where removing target concepts significantly impacts the model's overall performance. To address these issues, we propose a novel concept domain correction framework named \textbf{DoCo} (\textbf{Do}main \textbf{Co}rrection). By aligning the output domains of sensitive and anchor concepts through adversarial training, our approach ensures comprehensive unlearning of target concepts. Additionally, we introduce a concept-preserving gradient surgery technique that mitigates conflicting gradient components, thereby preserving the model's utility while unlearning specific concepts. Extensive experiments across various instances, styles, and offensive concepts demonstrate the effectiveness of our method in unlearning targeted concepts with minimal impact on related concepts, outperforming previous approaches even for out-of-distribution prompts.
Yongliang Wu, Shiji Zhou, Mingzhuo Yang, Lianzhe Wang, Heng Chang, Xinting Hu, Xu Yang 0021
AAAI7
2025 Personalized Generation In Large Model Era: A Survey
abstract
Yiyan Xu, Jinghao Zhang, Alireza Salemi, Xinting Hu, Wenjie Wang, Fuli Feng, Hamed Zamani, Xiangnan He, Tat-Seng Chua. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Yiyan Xu, Alireza Salemi, Xinting Hu, Wenjie Wang 0007, Fuli Feng, Hamed Zamani, Xiangnan He 0001, Tat-Seng Chua
ACL (1)4
2025 PersonaHOI: Effortlessly Improving Face Personalization in Human-Object Interaction Generation
abstract
We introduce PersonaHOI, a training-and tuning-free framework that fuses a general StableDiffusion model with a personalized face diffusion (PFD) model to generate identity-consistent human-object interaction (HOI) images. While existing PFD models have advanced significantly, they often overemphasize facial features at the expense of full-body coherence, PersonaHOI introduces an additional StableDiffusion (SD) branch guided by HOI-oriented text inputs. By incorporating cross-attention constraints in the PFD branch and spatial merging at both latent and residual levels, PersonaHOI preserves personalized facial details while ensuring interactive non-facial regions. Experiments, validated by a novel interaction alignment metric, demonstrate the superior realism and scalability of PersonaHOI, establishing a new standard for practical personalized face with HOI generation. Code is available at https://github.com/JoyHuYY1412/PersonaHOI.
Xinting Hu, Haoran Wang 0008, Jan Eric Lenssen, Bernt Schiele
CVPR1
2025 Mimic In-Context Learning for Multimodal Tasks
abstract
Recently, In-context Learning (ICL) has become a significant inference paradigm in Large Multimodal Models (LMMs), utilizing a few in-context demonstrations (ICDs) to prompt LMMs for new tasks. However, the synergistic effects in multimodal data increase the sensitivity of ICL performance to the configurations of ICDs, stimulating the need for a more stable and general mapping function. Mathematically, in Transformer-based models, ICDs act as "shift vectors" added to the hidden states of query tokens. Inspired by this, we introduce Mimic In-Context Learning (MimIC) to learn stable and generalizable shift effects from ICDs. Specifically, compared with some previous shift vector-based methods, MimIC more strictly approximates the shift effects by integrating lightweight learnable modules into LMMs with four key enhancements: 1) inserting shift vectors after attention layers, 2) assigning a shift vector to each attention head, 3) making shift magnitude query-dependent, and 4) employing a layer-wise alignment loss. Extensive experiments on two LMMs (Idefics-9b and Idefics2-8b-base) across three multimodal tasks (VQAv2, OK-VQA, Captioning) demonstrate that MimIC outperforms existing shift vector-based methods. The code is available at https://github.com/Kamichanw/MimIC.
Yuchu Jiang, Jiale Fu, Chenduo Hao, Xinting Hu, Yingzhe Peng, Xin Geng 0001, Xu Yang 0004
CVPR4
2025 Number it: Temporal Grounding Videos like Flipping Manga
abstract
Video Large Language Models (Vid-LLMs) have made remarkable advancements in comprehending video content for QA dialogue. However, they struggle to extend this visual understanding to tasks requiring precise temporal localization, known as Video Temporal Grounding (VTG). To address this, we introduce Number-Prompt (NumPro), a novel method that empowers Vid-LLMs to bridge visual comprehension with temporal grounding by adding unique numerical identifiers to each video frame. Treating a video as a sequence of numbered frame images, NumPro transforms VTG into an intuitive process: flipping through manga panels in sequence. This allows Vid-LLMs to “read” event timelines, accurately linking visual content with cor responding temporal information. Our experiments demonstrate that NumPro significantly boosts VTG performance of top-tier Vid-LLMs without additional computational cost. Furthermore, fine-tuning on a NumPro-enhanced dataset defines a new state-of-the-art for VTG, surpassing previous top-performing methods by up to 6.9% in mIoU for moment retrieval and 8.5% in mAP for highlight detection. The code is available at https://github.com/yongliang-wu/NumPro.
Yongliang Wu, Xinting Hu, Yizhou Zhou, Fengyun Rao, Bernt Schiele, Xu Yang 0004
CVPR2
2025 Edit360: 2D Image Edits to 3D Assets From Any Angle
abstract
Recent advances in diffusion models have significantly improved image generation and editing, but extending these capabilities to 3D assets remains challenging, especially for fine-grained edits that require multi-view consistency. Existing methods typically restrict editing to predetermined viewing angles, severely limiting their flexibility and practical applications. We introduce Edit360, a tuning-free framework that extends 2D modifications to multi-view consistent 3D editing. Built upon video diffusion models, Edit360 enables user-specific editing from arbitrary viewpoints while ensuring structural coherence across all views. The framework selects anchor views for 2D modifications and propagates edits across the entire 360-degree range. To achieve this, Edit360 introduces a novel Anchor-View Editing Propagation mechanism, which effectively aligns and merges multi-view information within the latent and attention spaces of diffusion models. The resulting edited multi-view sequences facilitate the reconstruction of high-quality 3D assets, enabling customizable 3D content creation.
Junchao Huang, Xinting Hu, Shaoshuai Shi, Zhuotao Tian, Li Jiang 0009
ICCV2
2025 Automated Multi-Aircraft Rerouting Under Convective Weather Using Policy-Shared Deep Reinforcement Learning
abstract
Automated decision support tools are increasingly important for assisting pilots and air traffic controllers in managing aircraft operations under complex scenarios such as convective weather, especially given the increasing traffic density and the limits of human cognitive capacity. Existing automation methods based on geometric heuristics or optimization are efficient and interpretable, but fail to coordinate multiple aircraft and adapt to rapidly evolving airspace conditions. This research presents a decentralized deep reinforcement learning (DRL) framework for multi-aircraft rerouting in thunderstorm-affected environments. Each aircraft acts as an autonomous agent and learns a shared policy via Independent Deep Deterministic Policy Gradient (IDDPG). During training, agents optimize a shared multi-objective reward that encodes safety and efficiency. By learning from diverse multi-agent scenarios, the shared policy captures transferable coordination patterns, enabling agents to generalize across traffic densities and storm configurations. The evaluation results for both simulated and real world airspace scenarios show that the proposed method reduces the aircraft conflict rates to below 1 %, maintains a success of over 95 % in reaching exit waypoints, and produces smoother trajectories with more organized traffic flows compared to baseline methods. These findings demonstrate the potential of the method as a reliable and scalable AI-based decision support tool for real-time multi-aircraft rerouting in convective weather conditions.
Xinting Hu, Bizhao Pang, Mingcheng Zhang, Sameer Alam, Guglielmo Lulli 0001
ICTAI1
2025 Decentralized Deep Reinforcement Learning for Cooperative Multi-Agent Flight Trajectory Planning in Adverse Weather
Bizhao Pang, Xinting Hu, Mingcheng Zhang, Sameer Alam, Guglielmo Lulli 0001
AAMAS2
2025 DRC: Enhancing Personalized Image Generation via Disentangled Representation Composition
abstract
Personalized image generation has emerged as a promising direction in multimodal content creation. It aims to synthesize images tailored to individual style preferences (e.g. color schemes, character appearances, layout) and semantic intentions (e.g. emotion, action, scene contexts) by leveraging user-interacted history images and multimodal instructions. Despite notable progress, existing methods -- whether based on diffusion models, large language models, or Large Multimodal Models (LMMs) -- struggle to accurately capture and composite user style preferences and semantic intentions. In particular, the state-of-the-art LMM-based method suffers from the entanglement of visual features, leading to Guidance Collapse, where the generated images fail to preserve user-preferred styles or reflect the specified semantics.
Yiyan Xu, Wuqiang Zheng, Wenjie Wang 0007, Fengbin Zhu, Xinting Hu, Yang Zhang 0072, Fuli Feng, Tat-Seng Chua
ACM Multimedia5
2025 KRIS-Bench: Benchmarking Next-Level Intelligent Image Editing Models
abstract
Recent advances in multi-modal generative models have enabled significant progress in instruction-based image editing. However, while these models produce visually plausible outputs, their capacity for knowledge-based reasoning editing tasks remains under-explored. In this paper, We introduce KRIS-Bench (Knowledge-based Reasoning in Image-editing Systems Benchmark), a diagnostic benchmark designed to assess models through a cognitively informed lens. Drawing from educational theory, KRIS-Bench categorizes editing tasks across three foundational knowledge types: Factual, Conceptual, and Procedural. Based on this taxonomy, we design 22 representative tasks spanning 7 reasoning dimensions and release 1,267 high-quality annotated editing instances. To support fine-grained evaluation, we propose a comprehensive protocol that incorporates a novel Knowledge Plausibility metric, enhanced by knowledge hints and calibrated through human studies. Empirical results on nine state-of-the-art models reveal significant gaps in reasoning performance, highlighting the need for knowledge-centric benchmarks to advance the development of intelligent image editing systems.
Yongliang Wu, Zonghui Li, Xinting Hu, Xinyu Ye, Xianfang Zeng, Bernt Schiele, Ming-Hsuan Yang 0001, Xu Yang 0004
NeurIPS3
2025 A multi-aircraft co-operative trajectory planning model under dynamic thunderstorm cells using decentralized deep reinforcement learning
Bizhao Pang, Xinting Hu, Mingcheng Zhang, Sameer Alam, Guglielmo Lulli 0001
Adv. Eng. Informatics2
2024 Training Vision Transformers for Semi-Supervised Semantic Segmentation
abstract
We present S4Former, a novel approach to training Vision Transformers for Semi-Supervised Semantic Segmentation (S4). At its core, S4Former employs a Vision Transformer within a classic teacher-student framework, and then leverages three novel technical ingredients: PatchShuffle as a parameter-free perturbation technique, Patch-Adaptive Self-Attention (PASA) as a fine-grainedfeature modulation method, and the innovative Negative Class Ranking (NCR) regularization loss. Based on these regu-larization modules aligned with Transformer-specific char-acteristics across the image input, feature, and output di-mensions, S4Former exploits the Transformer's ability to capture and differentiate consistent global contextual information in unlabeled images. Overall, S4 Former not only defines a new state of the art in S4 but also maintains a streamlined and scalable architecture. Being readily compatible with existing frameworks, S4 Former achieves strong improvements (up to 4.9%) on benchmarks like Pascal VOC 2012, COCO, and Cityscapes, with varying numbers of labeled data. The code is at https://github.com/JoyHuYY1412/S4Former.
Xinting Hu, Li Jiang 0009, Bernt Schiele
CVPR1
2024 MTA-CLIP: Language-Guided Semantic Segmentation with Mask-Text Alignment
Anurag Das, Xinting Hu, Li Jiang 0009, Bernt Schiele
ECCV (54)2
2024 LIVE: Learnable In-Context Vector for Visual Question Answering
abstract
As language models continue to scale, Large Language Models (LLMs) have exhibited emerging capabilities in In-Context Learning (ICL), enabling them to solve language tasks by prefixing a few in-context demonstrations (ICDs) as context. Inspired by these advancements, researchers have extended these techniques to develop Large Multimodal Models (LMMs) with ICL capabilities. However, applying ICL usually faces two major challenges: 1) using more ICDs will largely increase the inference time and 2) the performance is sensitive to the selection of ICDs. These challenges are further exacerbated in LMMs due to the integration of multiple data types and the combinational complexity of multimodal ICDs. Recently, to address these challenges, some NLP studies introduce non-learnable In-Context Vectors (ICVs) which extract useful task information from ICDs into a single vector and then insert it into the LLM to help solve the corresponding task. However, although useful in simple NLP tasks, these non-learnable methods fail to handle complex multimodal tasks like Visual Question Answering (VQA). In this study, we propose \underline{\textbf{L}}earnable \underline{\textbf{I}}n-Context \underline{\textbf{Ve}}ctor (LIVE) to distill essential task information from demonstrations, improving ICL performance in LMMs. Experiments show that LIVE can significantly reduce computational costs while enhancing accuracy in VQA tasks compared to traditional ICL and other non-learnable ICV methods.
Yingzhe Peng, Chenduo Hao, Xinting Hu, Jiawei Peng 0001, Xin Geng 0001, Xu Yang 0021
NeurIPS3
2022 On Non-Random Missing Labels in Semi-Supervised Learning
Xinting Hu, Yulei Niu, Chunyan Miao, Xian-Sheng Hua 0001, Hanwang Zhang
ICLR1
2021 Distilling Causal Effect of Data in Class-Incremental Learning
abstract
We propose a causal framework to explain the catastrophic forgetting in Class-Incremental Learning (CIL) and then derive a novel distillation method that is orthogonal to the existing anti-forgetting techniques, such as data replay and feature/label distillation. We first 1) place CIL into the framework, 2) answer why the forgetting happens: the causal effect of the old data is lost in new training, and then 3) explain how the existing techniques mitigate it: they bring the causal effect back. Based on the causal framework, we propose to distill the Colliding Effect between the old and the new data, which is fundamentally equivalent to the causal effect of data replay, but without any cost of replay storage. Thanks to the causal effect analysis, we can further capture the Incremental Momentum Effect of the data stream, removing which can help to retain the old effect overwhelmed by the new data effect, and thus alleviate the forgetting of the old class in testing. Extensive experiments on three CIL benchmarks: CIFAR-100, ImageNet-Sub&Full, show that the proposed causal effect distillation can improve various state-of-the-art CIL methods by a large margin (0.72%–9.06%).1
Xinting Hu, Kaihua Tang, Chunyan Miao, Xian-Sheng Hua 0001, Hanwang Zhang
CVPR1
2021 Cooperative Path Planning of UAVs & UGVs for a Persistent Surveillance Task in Urban Environments
abstract
There have been many applications of drones in urban environments, such as delivery, rescue, and surveillance. In a persistent surveillance task, the drones sometimes cannot complete it independently when some regions are required to be covered on the ground. For this purpose, unmanned aerial vehicles and unmanned ground vehicles (UAVs & UGVs) system is introduced to perform such a task in this article, and the goal is to generate the circular paths for the drones and the UGVs, respectively, to minimize their travel time of realizing a complete coverage. First, the cooperative path planning problem of UAVs & UGVs is formulated into a large-scale 0-1 optimization problem, in which the on-off states of the discrete points are to be optimized. Second, a hybrid algorithm integrating the estimation of distribution algorithm (EDA) and the genetic algorithm (GA) algorithm is proposed to solve the problem. The advantages of EDA and GA in the global and local search are fully taken considering the demands in different phases of the iterative process. A simple sweep-based approach is employed to determine the optimal sequence of passing the open points. Then, an online local adjustment strategy is also applied to address the changes of the requirements on covering the ground area. Simulation results demonstrate that the UAVs & UGVs system can enhance the efficiency of the task. The hybrid EDA-GA algorithm can greatly improve the performance of EDA and GA in terms of the quality and the stability of solutions. The online adjustment strategy is effective to maintain a complete coverage while minimizing the impact on the circular paths.
Yu Wu 0009, Shaobo Wu, Xinting Hu
IEEE Internet Things J.3
2020 Learning to Segment the Tail
abstract
Real-world visual recognition requires handling the extreme sample imbalance in large-scale long-tailed data. We propose a “divide&conquer” strategy for the challenging LVIS task: divide the whole data into balanced parts and then apply incremental learning to conquer each one. This derives a novel learning paradigm: class-incremental few-shot learning, which is especially effective for the challenge evolving over time: 1) the class imbalance among the old class knowledge review and 2) the few-shot data in new-class learning. We call our approach Learning to Segment the Tail (LST). In particular, we design an instance-level balanced replay scheme, which is a memory-efficient approximation to balance the instance-level samples from the old-class images. We also propose to use a meta-module for new-class learning, where the module parameters are shared across incremental phases, gaining the learning-to-learn knowledge incrementally, from the data-rich head to the data-poor tail. We empirically show that: at the expense of a little sacrifice of head-class forgetting, we can gain a significant 8.3% AP improvement for the tail classes with less than 10 instances, achieving an overall 2.0% AP boost for the whole 1,230 classes.
Xinting Hu, Kaihua Tang, Jingyuan Chen 0003, Chunyan Miao, Hanwang Zhang
CVPR1
2018 Risk modeling and optimization approach for system protection communication networks
abstract
System Protection Communication Network (SPCN) is a new type of high-speed, real-time, secure and reliable communication network proposed in China supporting services such as AC/DC control, pumped storage control etc. In order to reduce the impact of SPCN failure on electric power system, this paper proposes a risk modeling and optimization approach. Firstly, we build a risk model to analyze the dynamic link and service risk from aspects of failure probability and its impact value. Then, we construct a risk optimization problem aiming at minimizing the link risk balance degree with service quality and risk constraints, and propose improved genetic algorithm to solve it. Based on part of network topology from a Chinese province, simulation results show that the proposed approach can make SPCN more reliable comparing to other methods when link failure occurs.
Xinting Hu, Wenjing Li 0001, Peng Yu 0001, Fangzheng Chen
NOMS1