Yiran Zhu

dblp:170/3813 · DBLP profile ↗
← Back
13ranked-venue papers
6as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-author · 5 since 2021Security and privacy · 4 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Computer networks · 1Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-authorTheory of computation · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Video Echoed in Music: Semantic, Temporal, and Rhythmic Alignment for Video-to-Music Generation
abstract
Video-to-Music generation seeks to generate musically appropriate background music that enhances audiovisual immersion for videos. However, current approaches suffer from two critical limitations: 1) incomplete representation of video details, leading to weak alignment, and 2) inadequate temporal and rhythmic correspondence, particularly in achieving precise beat synchronization. To address the challenges, we propose Video Echoed in Music (VeM), a latent music diffusion that generates high-quality soundtracks with semantic, temporal, and rhythmic alignment for input videos. To capture video details comprehensively, VeM employs a hierarchical video parsing that acts as a music conductor, orchestrating multi-level information across modalities. Modality-specific encoders, coupled with a storyboard-guided cross-attention mechanism (SG-CAtt), integrate semantic cues while maintaining temporal coherence through position and duration encoding. For rhythmic precision, the frame-level transition-beat aligner and adapter (TB-As) dynamically synchronize visual scene transitions with music beats. We further contribute a novel video-music paired dataset sourced from e-commerce advertisements and video-sharing platforms, which imposes stricter transition-beat synchronization requirements. Meanwhile, we introduce novel metrics tailored to the task. Experimental results demonstrate superiority, particularly in semantic relevance and rhythmic precision.
Xinyi Tong 0001, Yiran Zhu, Jishang Chen, Chunru Zhan, Tianle Wang 0007, Sirui Zhang, Nian Liu 0003, Tiezheng Ge, Duo Xu 0004, Xin Jin 0015, Feng Yu 0032, Song-Chun Zhu
AAAI2
2026 BINALIGNER: Aligning Binary Code for Cross-Compilation Environment Diffing
Yiran Zhu, Zhenguang Liu, Lorenzo Cavallaro
NDSS1
2025 HER2 Expression Prediction with Flexible Multi-Modal Inputs via Dynamic Bidirectional Reconstruction
abstract
In the field of HER2 expression level assessment for breast cancer, clinical evaluations often rely on the synergistic analysis of both H&E and IHC stained images. However, acquiring dual-modality images for the same patient is frequently hindered by complex clinical workflows and high costs, resulting in missing modalities. To address this challenge, we propose an adaptive bimodal input prediction framework that flexibly supports both single-modality and dual-modality inputs. This framework employs a dynamic branch selection mechanism to overcome the rigid dependency of existing models on complete inputs, enabling accurate predictions using either H&E or IHC images alone, while retaining the ability for joint inference when both modalities are available. The core technical innovations include: a missing modality branch selector that dynamically activates either a modality completion process or an end-to-end dual-modality inference pipeline based on the available input; and a cross-modal generative adversarial network (CM-GAN) that facilitates context-aware reconstruction of the missing modality in the feature space. This design improves the prediction accuracy from 71.44% to 94.25% when using single-modality H&E images, significantly mitigating performance degradation caused by incomplete information. Experimental results demonstrate that the proposed framework achieves a prediction accuracy of 95.09% with full dual-modality input and maintains a high reliability of 90.28% under single-modality conditions. By adopting this ''dual-modality preferred, single-modality compatible'' flexible architecture, healthcare institutions can achieve near dual-modality accuracy without mandating synchronized acquisition of both image types. This is particularly valuable for regions with limited IHC staining infrastructure, offering a cost-effective clinical solution and substantially enhancing the accessibility of HER2 expression level assessment.
Wei Yang 0034, Yiran Zhu, Weizhen Li, Yunyue Pan, Chengchang Pan, Honggang Qi
ACM Multimedia4
2025 Defending Data Inference Attacks Against Machine Learning Models by Mitigating Prediction Distinguishability
abstract
Neural networks are vulnerable to data inference attacks, including the membership inference attack, the model inversion attack, and the attribute inference attack. In this paper, we proposePurifierto defend against membership inference attacks by quantifying the differences between dataset members and non-members in three dimensions: individual shape, statistical distribution, and prediction label.Purifierinvolves transforming the confidence scores produced by the target classifier, resulting in purified confidence scores that are indistinguishable across the dimensions above. We conduct experiments on widely-used datasets and models. The results show thatPurifieroffers robust defense against membership inference attacks with superior efficacy compared to prior defense techniques while maintaining minimal utility degradation (e.g., less than 0.7% classification accuracy drop of most datasets). Additionally, our extended experiments explore the effectiveness ofPurifierin defending against the model inversion attack and the attribute inference attack.
Yiran Zhu, ChuXiao Xiang, Ruite Xu, Lijin Wang, Fan Zhang 0010, Jiarong Xu, Zhan Qin
IEEE Trans. Dependable Secur. Comput.2
2024 APIBeh: Learning Behavior Inclination of APIs for Malware Classification
abstract
Malware classification involves categorizing mal-ware samples based on their characteristics. While deep learning techniques applied to malware execution traces, mainly API calls, have shown potential in this field, they still perform poorly. This is primarily because they treat all APIs equally and train classifiers directly on native APIs, which inadequately capture the under-lying family-related semantics. In this paper, we first investigate the behaviors of multiple malware families and observe that different families exhibit divergent behaviors, with each family consistently favoring certain behaviors over time. Motivated by this, we propose APIBeh, a new embedding method designed to enhance malware classification. APIBeh first utilizes Benignity Degree Algorithm to identify and exclude insignificant, likely benign APIs from sequences. Then, it introduces the concept of Behavior Inclination, which quantifies the association between an API and malicious behaviors, facilitating high-level behavior encoding for each API. This Behavior Inclination embedding is then concatenated with raw embedding to represent an API, and fed into a DL model for classifier training. Experimental results show that APIBeh outperforms existing embedding methods in classification performance, e.g., 3.18% boost in weighted f1-score over a recent study using word2vec. In addition, it offers robustness to concept drift and adversarial attacks.
Lei Cui 0003, Yiran Zhu, Junnan Yin, Zhiyu Hao, Wei Wang 0428, Peng Liu 0044, Xiao-chun Yun
ISSRE2
2022 Selective Hypergraph Convolutional Networks for Skeleton-based Action Recognition
abstract
In skeleton-based action recognition, Graph Convolutional Networks (GCNs) have achieved remarkable performance since the skeleton representation of human action can be naturally modeled by the graph structure. Most of the existing GCN-based methods extract skeleton features by exploiting single-scale joint information, while neglecting the valuable multi-scale contextual information. Besides, the commonly used strided convolution in temporal dimension could evenly filters out the keyframes we expect to preserve and leads to the loss of keyframe information. To address these issues, we propose a novel Selective Hypergraph Convolution Network, dubbed Selective-HCN, which stacks two key modules: Selective-scale Hypergraph Convolution (SHC) and Selective-frame Temporal Convolution (STC). The SHC module represents the human skeleton as the graph and hypergraph to fully extract multi-scale information, and selectively fuse features at various scales. Instead of traditional strided temporal convolution, the STC module can adaptively select keyframes and filter redundant frames according to the importance of the frames. Extensive experiments on two challenging skeleton action benchmarks, i.e., NTU-RGB+D and Skeleton-Kinetics, demonstrate the superiority and effectiveness of our proposed method.
Yiran Zhu, Guangji Huang, Xing Xu 0001, Yanli Ji, Fumin Shen
ICMR1
2022 Black box attack and network intrusion detection using machine learning for malicious traffic
Yiran Zhu, Lei Cui 0003, Zhenquan Ding, Yongji Liu, Zhiyu Hao
Comput. Secur.1
2022 CruParamer: Learning on Parameter-Augmented API Sequences for Malware Detection
abstract
Learning on execution behaviour, i.e., sequences of API calls, is proven to be effective in malware detection. In this paper, we present CruParamer, a deep neural network based malware detection approach for Windows platform that performs learning on sequences of parameter-augmented APIs. It first employs rule-based and clustering-based classification to assess the sensitivity of a parameter to malicious behaviour, and further labels the API following the run-time parameters with varying degrees of sensitivities. Then, it encodes the APIs by concatenating the native embedding and the sensitive embedding of labelled APIs, for characterizing the relationship between successive labelled APIs and their correspondence in terms of security semantics. Finally, it feeds the sequences of API embedding into the deep neural network for training a binary classifier to detect malware. In addition to presenting the design, we have implemented CruParamer and evaluated it on two datasets. The results demonstrate that CruParamer outperforms naïve models when taking raw APIs as input, proving the effectiveness of CruParamer. Moreover, we have evaluated the impact ofmimicryand adversarial attacks on our model, and the results verify the robustness of CruParamer.
Zhiyu Hao, Lei Cui 0003, Yiran Zhu, Zhenquan Ding, Yongji Liu
IEEE Trans. Inf. Forensics Secur.5
2021 Solving the Home Service Assignment, Routing, and Appointment Scheduling (H-SARA) Problem with Uncertainties
abstract
The Home Service Assignment, Routing, and Appointment scheduling (H-SARA) problem integrates the strategic fleet-sizing, tactical assignment, operational vehicle routing and scheduling problems at different decision levels, with a single period planning horizon and uncertainty (stochasticity) from the service duration, travel time, and customer cancellation rate. We propose a stochastic mixed-integer linear programming model for the H-SARA problem. Additionally, a reduced deterministic version is introduced which allows to solve small-scale instances to optimality with two acceleration approaches. For larger instances, we develop a tailored two-stage decision support system that provides high-quality and in-time solutions based on information revealed at different stages. Our solution method aims to reduce various costs under stochasticity, to create reasonable routes with balanced workload and team-based customer service zones, and to increase customer satisfaction by introducing a two-stage appointment notification system updated at different time stages before the actual service. Our two-stage heuristic is competitive to CPLEX’s exact solution methods in providing time and cost-effective decisions and can update previously-made decisions based on an increased level of information. Results show that our two-stage heuristic is able to tackle reasonable-size instances and provides good-quality solutions using less time compared to the deterministic and stochastic models on the same set of simulated instances.
Syu-Ning Johnn, Yiran Zhu, Andrés Miniguano-Trujillo, Akshay Gupte
ATMOS2
2021 Graph Convolutional Hourglass Networks for Skeleton-Based Action Recognition
abstract
Graph convolution networks (GCNs) have become the mainstream framework for the skeleton-based action recognition task, since the skeleton representation of human action can be naturally modeled by the graph structure. Generally, most of the existing GCN based models extract and aggregate skeleton features by exploiting single-scale joint information, while neglecting the valuable multi-scale information such as part and body features in the skeleton. To address this issue, we propose a novel Graph Convolutional Hourglass Network (GCHN) model, which is scalable by stacking several basic modules of Graph Convolutional Hourglass Block (GCHB). Each GCHB module consists of the sequential operations of graph convolution, graph pooling and graph unpooling, which can promote the interaction of multi-scale information in the skeleton and effectively improve the recognition performance. Extensive experiments on the challenging NTU-RGB+D and Kinetics-Skeleton datasets demonstrate that the proposed GCHN model achieves state-of-the-art performance.
Yiran Zhu, Xing Xu 0001, Yanli Ji, Fumin Shen, Heng Tao Shen, Huimin Lu 0001
ICME1
2021 PoseGTAC: Graph Transformer Encoder-Decoder with Atrous Convolution for 3D Human Pose Estimation
abstract
Graph neural networks (GNNs) have been widely used in the 3D human pose estimation task, since the pose representation of a human body can be naturally modeled by the graph structure. Generally, most of the existing GNN-based models utilize the restricted receptive fields of filters and single-scale information, while neglecting the valuable multi-scale contextual information. To tackle this issue, we propose a novel Graph Transformer Encoder-Decoder with Atrous Convolution, named PoseGTAC, to effectively extract multi-scale context and long-range information. In our proposed PoseGTAC model, Graph Atrous Convolution (GAC) and Graph Transformer Layer (GTL), respectively for the extraction of local multi-scale and global long-range information, are combined and stacked in an encoder-decoder structure, where graph pooling and unpooling are adopted for the interaction of multi-scale information from local to global (e.g., part-scale and body-scale). Extensive experiments on the Human3.6M and MPI-INF-3DHP datasets demonstrate that the proposed PoseGTAC model exceeds all previous methods and achieves state-of-the-art performance.
Yiran Zhu, Xing Xu 0001, Fumin Shen, Yanli Ji, Lianli Gao, Heng Tao Shen
IJCAI1
2020 An Adversarial Learning Model for Intrusion Detection in Real Complex Network Environments
Ying Zhong 0008, Yiran Zhu, Xia Yin 0001, Xingang Shi, Keqin Li 0001
WASA (1)2
2019 Skeleton Tree based Non-rigid 3D Shape Retrieval
abstract
We propose a skeleton tree based method for classifying and retrieving non-rigid shapes. Firstly, based on the extracted skeletons of a non-rigid shape and geodesic distance computation, the center point in skeleton is defined and detected. Then, a skeleton tree is constructed based on the connection between the center point and other discrete points in skeleton. After that, a correspondence between the skeleton tree and the area distribution of the non-rigid shape is established. The skeleton tree features are achieved. The advantages of our method can be summarized as follows: (1) Scale-Invariant; (2) Low computational complexity; (3) Automatic topology repair. The experimental results show that our method is more accurate than existing methods.
Yiran Zhu, Jiaqi Kang, Chenlei Lv, Yanping Xue, Xingce Wang, Zhongke Wu
VINCI1