Tianyang Han

dblp:254/6445 · DBLP profile ↗
← Back
11ranked-venue papers
1as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 1 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Vision and language · 24% Deep learning architectures and training · 21% Trustworthy machine learning · 21%
Network and information security
1 paper
Blockchain and cryptocurrency security · 100%
Software engineering, system software, and programming languages
1 paper
Software maintenance and evolution · 100%

Topics — the 14 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language › vision-language model
multimodal large language model
2.642025
Personalized Visual Instruction Tuning · ICLR 2025
The Instinctive Bias: Spurious Images lead to Illusion in MLLMs · EMNLP 2024
Strengthening Multimodal Large Language Model with Bootstrapped Preference Optimization · ECCV (33) 2024
Machine learning › Deep learning architectures and training › biologically plausible learning › feedback alignment
backpropagation alternative
1.012026
MAN++: Scaling Momentum Auxiliary Network for Supervised Local Learning in Vision Tasks · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Machine learning › Deep learning architectures and training › neural network training › local learning
supervised local learning
1.012026
MAN++: Scaling Momentum Auxiliary Network for Supervised Local Learning in Vision Tasks · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Blockchain and cryptocurrency security › smart contract security
vulnerability detection
1.012026
Fine-Grained Detection of Java Cross-Library Vulnerability Propagation by Extracting Semantic Constraints From Security Patches · IEEE Trans. Dependable Secur. Comput. 2026
Natural language and speech › Question answering and dialogue systems
personalized dialogue
0.912025
Personalized Visual Instruction Tuning · ICLR 2025
Computer vision › Face, body and person analysis
person identification
0.912025
Personalized Visual Instruction Tuning · ICLR 2025
Natural language and speech › Language models and text generation
preference optimization
0.812024
Strengthening Multimodal Large Language Model with Bootstrapped Preference Optimization · ECCV (33) 2024
Machine learning › Trustworthy machine learning
robustness
0.812024
The Instinctive Bias: Spurious Images lead to Illusion in MLLMs · EMNLP 2024
Machine learning › Trustworthy machine learning › robustness
spurious correlation
0.812024
The Instinctive Bias: Spurious Images lead to Illusion in MLLMs · EMNLP 2024
Machine learning › Generative modeling
visual illusion
0.812024
The Instinctive Bias: Spurious Images lead to Illusion in MLLMs · EMNLP 2024
Machine learning › Deep learning architectures and training › neural network training
local learning
0.312026
MAN++: Scaling Momentum Auxiliary Network for Supervised Local Learning in Vision Tasks · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Machine learning › Efficient and distributed learning
memory-efficient training
0.312026
MAN++: Scaling Momentum Auxiliary Network for Supervised Local Learning in Vision Tasks · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Software maintenance and evolution › software ecosystems
package dependencies
0.312026
Fine-Grained Detection of Java Cross-Library Vulnerability Propagation by Extracting Semantic Constraints From Security Patches · IEEE Trans. Dependable Secur. Comput. 2026
Software maintenance and evolution
software ecosystems
0.312026
Fine-Grained Detection of Java Cross-Library Vulnerability Propagation by Extracting Semantic Constraints From Security Patches · IEEE Trans. Dependable Secur. Comput. 2026

Methods — techniques the papers use, named apart from their topics

semantic constraint extraction · 2.0security patch analysis · 2.0momentum auxiliary network · 1.0learnable scaling bias · 1.0exponential moving average · 1.0instruction tuning · 0.9fine-tuning · 0.9safety alignment · 0.8preference optimization · 0.8bootstrapping · 0.8benchmark construction · 0.8adversarial training · 0.8
YearPublicationVenuePosition
2026 MAN++: Scaling Momentum Auxiliary Network for Supervised Local Learning in Vision Tasks
abstract
End-to-end backpropagation remains the dominant training paradigm in deep learning, yet it suffers from inherent drawbacks, including update locking, high GPU memory consumption, and limited biological plausibility. Supervised local learning alleviates these issues by dividing the network into multiple blocks and training each block independently with an auxiliary network. However, gradient isolation also weakens the influence of downstream representations on earlier blocks, often resulting in a clear accuracy gap to end-to-end training. We propose Momentum Auxiliary Network++ (MAN++), a scalable framework that improves supervised local learning via a lightweight parameter-space transfer between adjacent blocks. MAN++ employs the exponential moving average (EMA) of parameters from adjacent blocks to propagate contextual information across the network. To address feature mismatches arising from direct EMA parameter transfer, we introduce a learnable scaling bias, which compensates feature statistics mismatch and stabilizes the transfer. Extensive experiments on image classification, object detection, and semantic segmentation across multiple architectures illustrate that MAN++ achieves accuracy on par with end-to-end training while substantially reducing GPU memory usage. These results position MAN++ as a practical and effective alternative to conventional backpropagation, offering new insights into scalable supervised local learning for vision tasks.
Junhao Su, Hengyu Shi, Tianyang Han, Yurui Qiu, Junfeng Luo, Xiaoming Wei, Jialin Gao
IEEE Trans. Pattern Anal. Mach. Intell.4
2026 Fine-Grained Detection of Java Cross-Library Vulnerability Propagation by Extracting Semantic Constraints From Security Patches
Fute Sun, Lei Zhang 0096, Zhiyu Wu, Tianyang Han, Min Yang 0002
IEEE Trans. Dependable Secur. Comput.6
2025 Personalized Visual Instruction Tuning
abstract
Recent advancements in multimodal large language models (MLLMs) have demonstrated significant progress; however, these models exhibit a notable limitation, which we refer to as "face blindness." Specifically, they can engage in general conversations but fail to conduct personalized dialogues targeting at specific individuals. This deficiency hinders the application of MLLMs in personalized settings, such as tailored visual assistants on mobile devices, or domestic robots that need to recognize members of the family. In this paper, we introduce Personalized Visual Instruction Tuning (PVIT), a novel data curation and training framework designed to enable MLLMs to identify target individuals within an image and engage in personalized and coherent dialogues. Our approach involves the development of a sophisticated pipeline that autonomously generates training data containing personalized conversations. This pipeline leverages the capabilities of various visual experts, image generation models, and (multi-modal) large language models. To evaluate the personalized potential of MLLMs, we present a benchmark called P-Bench, which encompasses various question types with different levels of difficulty. The experiments demonstrate a substantial personalized performance enhancement after fine-tuning with our curated dataset.
Renjie Pi, Jianshu Zhang 0003, Tianyang Han, Rui Pan 0002, Tong Zhang 0001
ICLR3
2025 An adaptive traffic signal control scheme with Proximal Policy Optimization based on deep reinforcement learning for a single intersection
Guoshan Zhang, Qiaoli Yang, Tianyang Han
Eng. Appl. Artif. Intell.4
2024 Strengthening Multimodal Large Language Model with Bootstrapped Preference Optimization
Renjie Pi, Tianyang Han, Wei Xiong 0015, Runtao Liu, Rui Pan 0002, Tong Zhang 0001
ECCV (33)2
2024 The Instinctive Bias: Spurious Images lead to Illusion in MLLMs
abstract
Large language models (LLMs) have recently experienced remarkable progress, where the advent of multi-modal large language models (MLLMs) has endowed LLMs with visual capabilities, leading to impressive performances in various multi-modal tasks.However, those powerful MLLMs such as GPT-4V still fail spectacularly when presented with certain image and text inputs.In this paper, we identify a typical class of inputs that baffles MLLMs, which consist of images that are highly relevant but inconsistent with answers, causing MLLMs to suffer from visual illusion.To quantify the effect, we propose CorrelationQA, the first benchmark that assesses the visual illusion level given spurious images.This benchmark contains 7,308 text-image pairs across 13 categories.Based on the proposed CorrelationQA, we conduct a thorough analysis on 9 mainstream MLLMs, illustrating that they universally suffer from this instinctive bias to varying degrees.We hope that our curated benchmark and evaluation results aid in better assessments of the MLLMs' robustness in the presence of misleading images.The code and datasets are available at https://github.com/MasaiahHan/CorrelationQA.Known for its distinctive black and white stripes, this African equine is closely related to horses and donkeys, .... what is it?
Tianyang Han, Qing Lian, Rui Pan 0002, Renjie Pi, Shizhe Diao, Tong Zhang 0001
EMNLP1
2024 MLLM-Protector: Ensuring MLLM's Safety without Hurting Performance
abstract
Renjie Pi, Tianyang Han, Jianshu Zhang, Yueqi Xie, Rui Pan, Qing Lian, Hanze Dong, Jipeng Zhang, Tong Zhang. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Renjie Pi, Tianyang Han, Jianshu Zhang 0003, Yueqi Xie, Rui Pan 0002, Qing Lian, Hanze Dong, Tong Zhang 0001
EMNLP2
2024 A Coordination Graph Based Framework for Network Traffic Signal Control
abstract
The efficiency of road networks affects the daily activities of each stakeholder. Multi-agent reinforcement learning (MARL) has emerged as a method for managing network traffic signal control (TSC). It treats each intersection as an agent and coordinates their actions to enhance overall performance. A critical issue is enabling agents to appropriately and systematically respond to network demand changes. In response, this study proposes a coordination graph-based framework. It considers two adjacent intersections as a pair and updates coordination graphs periodically based on observed demand patterns, determining which intersection pairs should be coordinated. Within this framework, an adaptive TSC method based on reinforcement learning is designed for isolated intersections. Furthermore, paired intersections are jointly controlled using a modified max-plus algorithm. The coordination graph is solved considering factors such as traffic demand and intersection spacing, employing a decomposition method named “snake game solver”. Experimental results show that the individual learning scheme resulted in robust control and quick adaptability to traffic fluctuations. However, the coordination learning scheme only led to improvements when the inter-demand between intersections was sufficiently high and the spacing was short. The numerical study suggests that this control framework could enhance network efficiency compared to other MARL-TSC methods.
Hong Zhu 0013, Fengmei Sun, Keshuang Tang, Tianyang Han, Junping Xiang
IEEE Trans. Intell. Transp. Syst.4
2023 A Dual-Flow Attentive Network With Feature Crossing for Chained Trip Purpose Inference
abstract
Trip purpose is essential information supporting tasks in intelligent transportation systems, such as travel behaviour comprehension, location-based service, and urban planning. The observation of trip purpose is a necessary aspect of travel surveys. However, owing to the sampling volume, survey budget, and survey frequency, relying solely on travel surveys in the era of big data is a difficult task. There has long been a demand for an accurate, generalizable, and robust inference method for trip purposes. Although existing studies contributed significant efforts to improve the trip purpose inference, the potential of leveraging the trip chain is insufficient. The spatial correlations and chaining patterns hidden in travelled zones are worthy of further exploration. The unequal importance within trip chains has not been clearly represented. Additionally, complex activity-zone mutual interdependence has not been considered in previous models. Herein, we propose a framework-Dual-FlowAttentive Network with FeatureCrossing (DACross), specifically for inferring the chained trip purpose. We form trip chains innovatively that treat trip activities and travelled geographic zones as two chains with mutual interactions. We propose DACross, which consists of two parallel attentive branches and a co-attentive feature crossing module, for fully learning the intra- and inter-chain dependencies. We conducted extensive experiments on four large-scale real-world datasets to evaluate not only the performance of DACross but also the generalizability of the proposed framework among different cities and scenarios. Notably, the Experimental results prove the overall superiority of the proposed DACross.
Suxing Lyu, Tianyang Han, Xingyu Luo, Takahiko Kusakabe
IEEE Trans. Intell. Transp. Syst.2
2022 A plug-in memory network for trip purpose classification
abstract
Trip purpose plays a critical role in reflecting human mobility behavior. However, it is relatively difficult to determine. With the rapid growth of urban mobility and big mobile data, utilizing these data for trip purpose classification has been a long-term objective to enhance travel demand and behavior models used in urban planning. Although studies on this topic have been extensively conducted, most past research preferred relying on traveler attributes or long-term travel histories to achieve accurate results. These data could be privacy sensitive and often do not satisfy real-world scenarios. This study addresses the problem of classifying trip purpose by only space activity information to avoid privacy conflict. 1) External memories are collected from factorized components based on the non-negative Tucker decomposition scheme. 2) These memories are extended by the cross-attention mechanism to achieve feature augmentation. 3) Subsequently, a novel concept called "latent mode alignment" is proposed. By leveraging the linear characteristics of external memories, geographic contextual latent modes are represented and matched with travel activities; this procedure is called "alignment." 4) The gate mechanism controls the eventual outputs for update. The proposed plug-in memory network (PMN), combined with baseline models, effectively outperforms the original settings. Moreover, combination models are validated with strong tolerance through missing data tests, which are common and problematic in real-world scenarios. The proposed PMN is a plug-and-play design that is easy to combine with newly developed classification models, and other memory collection methods can be expected.
Suxing Lyu, Tianyang Han, Yuuki Nishiyama, Kaoru Sezaki, Takahiko Kusakabe
SIGSPATIAL/GIS2
2022 DEMOC: a deep embedded multi-omics learning approach for clustering single-cell CITE-seq data
abstract
Advances in single-cell RNA sequencing (scRNA-seq) technologies has provided an unprecedent opportunity for cell-type identification. As clustering is an effective strategy towards cell-type identification, various computational approaches have been proposed for clustering scRNA-seq data. Recently, with the emergence of cellular indexing of transcriptomes and epitopes by sequencing (CITE-seq), the cell surface expression of specific proteins and the RNA expression on the same cell can be captured, which provides more comprehensive information for cell analysis. However, existing single cell clustering algorithms are mainly designed for single-omic data, and have difficulties in handling multi-omics data with diverse characteristics efficiently. In this study, we propose a novel deep embedded multi-omics clustering with collaborative training (DEMOC) model to perform joint clustering on CITE-seq data. Our model can take into account the characteristics of transcriptomic and proteomic data, and make use of the consistent and complementary information provided by different data sources effectively. Experiment results on two real CITE-seq datasets demonstrate that our DEMOC model not only outperforms state-of-the-art single-omic clustering methods, but also achieves better and more stable performance than existing multi-omics clustering methods. We also apply our model on three scRNA-seq datasets to assess the performance of our model in rare cell-type identification, novel cell-subtype detection and cellular heterogeneity analysis. Experiment results illustrate the effectiveness of our model in discovering the underlying patterns of data.
Guanhua Zou, Yilong Lin, Tianyang Han, Le Ou-Yang
Briefings Bioinform.3