Junyu Shi

dblp:195/2109 · DBLP profile ↗
← Back
13ranked-venue papers
4as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021Computer networks · 2 · 1 since 2021Security and privacy · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 EmbryoDiff: A Conditional Diffusion Framework with Multi-Focal Feature Fusion for Fine-Grained Embryo Developmental Stage Recognition
abstract
Identification of fine-grained embryo developmental stages during In Vitro Fertilization (IVF) is crucial for assessing embryo viability. Although recent deep learning methods have achieved promising accuracy, existing discriminative models fail to utilize the distributional prior of embryonic development to improve accuracy. Moreover, their reliance on single-focal information leads to incomplete embryonic representations, making them susceptible to feature ambiguity under cell occlusions. To address these limitations, we propose EmbryoDiff, a two-stage diffusion-based framework that formulates the task as a conditional sequence denoising process. Specifically, we first train and freeze a frame-level encoder to extract robust multi-focal features. In the second stage, we introduce a Multi-Focal Feature Fusion Strategy that aggregates information across focal planes to construct a 3D-aware morphological representation, effectively alleviating ambiguities arising from cell occlusions. Building on this fused representation, we derive complementary semantic and boundary cues and design a Hybrid Semantic-Boundary Condition Block to inject them into the diffusion-based denoising process, enabling accurate embryonic stage classification. Extensive experiments on two benchmark datasets show that our method achieves state-of-the-art results. Notably, with only a single denoising step, our model obtains the best average test performance, reaching 82.8% and 81.3% accuracy on the two datasets, respectively.
Zhengjie Zhang, Junyu Shi, Lijiang Liu, Qiang Nie
AAAI3
2026 Morphology Prior Enhanced Teeth Segmentation for High-Resolution Oral Scans
abstract
Deep learning methods have been proposed for tooth segmentation on high-resolution intra-oral scans (IOS) that plays a crucial role in clinical dental practice. However, they generally segment teeth in a low-resolution data with a fixed receptive field and generate final segmentation by up-sampling interpolation, and neglect teeth's morphology priors: their similar dental arch structures and significantly different curvatures in different parts of each tooth. They thus lack adaptability to different parts of each tooth, and show less accurate segmentation of boundary points between teeth and gums due to the up-sampling computation. Further, cluttered poses of IOS limit their generalization and usability of teeth location and geometric information. To address these limitations, a morphology prior enhanced teeth segmentation framework is proposed in this paper. Firstly, a robust preprocessing is introduced to align poses of different IOS by computing their dental arch orientations, thereby improving segmentation generalization and usability of IOS geometric information. Secondly, a decomposition-merging strategy is designed to avoid the up-sampling limitation, which decomposes an IOS into multiple low-resolution data and merges their segmentation outcomes into a high-resolution result. Thirdly, an innovative module integrating semantic and geometric features is proposed to adaptively select deformable receptive fields. It geometrically samples within a variable probability space to construct receptive fields with varied graph relationships for different points, facilitating adaptive segmentation of different parts of each tooth. Experimental results on 6238 IOS from four centers demonstrate that our method significantly outperforms 11 state-of-the-art methods, achieving a 6.93% enhancement for cross-center testing.
Yuxian Jiang, Xiuying Wang 0001, Tao Yang 0037, Changkai Ji, Lanshan He, Yusheng Liu 0001, Junyu Shi, Huayan Guo, Lisheng Wang
IEEE J. Biomed. Health Informatics9
2025 Improving Generalization of Universal Adversarial Perturbation via Dynamic Maximin Optimization
abstract
Deep neural networks (DNNs) are susceptible to universal adversarial perturbations (UAPs). These perturbations are meticulously designed to fool the target model universally across all sample classes. Unlike instance-specific adversarial examples (AEs), generating UAPs is more complex because they must be generalized across a wide range of data samples and models. Our research reveals that existing universal attack methods, which optimize UAPs using DNNs with static model parameter snapshots, do not fully leverage the potential of DNNs to generate more effective UAPs. Rather than optimizing UAPs against static DNN models with a fixed training set, we suggest using dynamic model-data pairs to generate UAPs. In particular, we introduce a dynamic maximin optimization strategy, aiming to optimize the UAP across a variety of optimal model-data pairs. We term this approach DM-UAP. DM-UAP utilizes an iterative max-min-min optimization framework that refines the model-data pairs, coupled with a curriculum UAP learning algorithm to examine the combined space of model parameters and data thoroughly. Comprehensive experiments on the ImageNet dataset demonstrate that the proposed DM-UAP markedly enhances both cross-sample universality and cross-model transferability of UAPs. Using only 500 samples for UAP generation, DM-UAP outperforms the state-of-the-art approach with an average increase in fooling ratio of 12.108%.
Yechao Zhang, Yingzhe Xu, Junyu Shi, Leo Yu Zhang, Shengshan Hu, Yanjun Zhang 0002
AAAI3
2025 GenM3: Generative Pretrained Multi-Path Motion Model for Text Conditional Human Motion Generation
Junyu Shi, Lijiang Liu, Jinni Zhou, Qiang Nie
ICCV1
2025 Time-Lapse Video-Based Embryo Grading via Complementary Spatial-Temporal Pattern Mining
Junyu Shi, Yanmei Xiao, Manxi Jiang, Qiang Nie
MICCAI (13)3
2024 Why Does Little Robustness Help? A Further Step Towards Understanding Adversarial Transferability
abstract
Adversarial examples for deep neural networks (DNNs) are transferable: examples that successfully fool one white-box surrogate model can also deceive other black-box models with different architectures. Although a bunch of empirical studies have provided guidance on generating highly transferable adversarial examples, many of these findings fail to be well explained and even lead to confusing or inconsistent advice for practical use.In this paper, we take a further step towards understanding adversarial transferability, with a particular focus on surrogate aspects. Starting from the intriguing "little robustness" phenomenon, where models adversarially trained with mildly perturbed adversarial samples can serve as better surrogates for transfer attacks, we attribute it to a trade-off between two dominant factors: model smoothness and gradient similarity. Our research focuses on their joint effects on transferability, rather than demonstrating the separate relationships alone. Through a combination of theoretical and empirical analyses, we hypothesize that the data distribution shift induced by off-manifold samples in adversarial training is the reason that impairs gradient similarity.Building on these insights, we further explore the impacts of prevalent data augmentation and gradient regularization on transferability and analyze how the trade-off manifests in various training methods, thus building a comprehensive blueprint for the regulation mechanisms behind transferability. Finally, we provide a general route for constructing superior surrogates to boost transferability, which optimizes both model smoothness and gradient similarity simultaneously, e.g., the combination of input gradient regularization and sharpness-aware minimization (SAM), validated by extensive experiments. In summary, we call for attention to the united impacts of these two factors for launching effective transfer attacks, rather than optimizing one while ignoring the other, and emphasize the crucial role of manipulating surrogate models.
Yechao Zhang, Shengshan Hu, Leo Yu Zhang, Junyu Shi, Xiaogeng Liu, Hai Jin 0001
SP4
2024 Gradient multi-foci networks for 3D skeleton-based human motion prediction
Junyu Shi, Jianqi Zhong, Zhiquan He, Wenming Cao 0001
Neural Comput. Appl.1
2024 Multi-Semantics Aggregation Network Based on the Dynamic-Attention Mechanism for 3D Human Motion Prediction
abstract
Graph convolutional network-based methods have recently shown promising performance in skeleton-based data processing. However, these methods have two critical issues in skeleton-based motion prediction tasks: First, graph modeling of motion poses is based on the fixed graph according to the physical connection of human joints and ignores the exploration of deep implicit information based on human dynamic kinetics. Second, existing methods usually use motion information in a single semantic space to model the whole motion sequences, underestimating diverse semantic patterns for improving the modeling ability. To address the first issue, we propose the Attention-based Dynamic Graph Convolution method, which tries to capture implicit semantic information dynamically. To address the second issue, we propose the Kinematic-based Semantics Aggregation Block (KSAB), which combines various semantic features from four semantic perspectives to rich motion representation. Integrating the above two designs, we propose a novel Multi-Semantics Aggregation Network (MANet), resulting in more comprehensive feature extraction in dynamic implicit semantics learning to enhance motion prediction. Extensive experiments are conducted to validate the effectiveness of MANet, which outperforms state-of-the-art methods by 10.9%, 6.6%, and 19.6% in terms of MPJPE for motion prediction on Human3.6M, CMU Mocap, and 3DPW datasets, respectively.
Junyu Shi, Jianqi Zhong, Wenming Cao 0001
IEEE Trans. Multim.1
2023 Joint Detection and Association for End-to-End Multi-object Tracking
Ye Li 0024, Junyu Shi, Xinzhong Wang, Guangqiang Yin, Zhiguo Wang 0004
Neural Process. Lett.3
2022 Shielding Federated Learning: Mitigating Byzantine Attacks with Less Constraints
abstract
Federated learning is a newly emerging distributed learning framework that facilitates the collaborative training of a shared global model among distributed participants with their privacy preserved. However, federated learning systems are vulnerable to Byzantine attacks from malicious participants, who can upload carefully crafted local model updates to degrade the quality of the global model and even leave a backdoor. While this problem has received significant attention recently, current defensive schemes heavily rely on various assumptions, such as a fixed Byzantine model, availability of participants' local data, minority attackers, IID data distribution, etc. To relax those constraints, this paper presents Robust-FL, the first prediction-based Byzantine-robust federated learning scheme where none of the assumptions is leveraged. The core idea of the Robust-FL is exploiting historical global model to construct an estimator based on which the local models will be filtered through similarity detection. We then cluster local models to adaptively adjust the acceptable differences between the local models and the estimator such that Byzantine users can be identified. Extensive experiments over different datasets show that our approach achieves the following advantages simultaneously: (i) independence of participants' local data, (ii) tolerance of majority attackers, (iii) generalization to variable Byzantine model.
Jianrong Lu, Shengshan Hu, Junyu Shi, Leo Yu Zhang, Man Zhou 0004, Yifeng Zheng 0001
MSN5
2022 Challenges and Approaches for Mitigating Byzantine Attacks in Federated Learning
abstract
Recently emerged federated learning (FL) is an attractive distributed learning framework in which numerous wireless end-user devices can train a global model with the data remained autochthonous. Compared with the traditional machine learning framework that collects user data for centralized storage, which brings huge communication burden and concerns about data privacy, this approach can not only save the network bandwidth but also protect the data privacy. Despite the promising prospect, Byzantine attack, an intractable threat in conventional distributed network, is discovered to be rather efficacious against FL as well. In this paper, we conduct a comprehensive investigation of the state-of-the-art strategies for defending against Byzantine attacks in FL. We first provide a taxonomy for the existing defense solutions according to the techniques they used, followed by an across-the-board comparison and discussion. Then we propose a new Byzantine attack method called weight attack to defeat those defense schemes, and conduct experiments to demonstrate its threat. The results show that existing defense solutions, although abundant, are still far from fully protecting FL. Finally, we indicate possible countermeasures for weight attack, and highlight several challenges and future research directions for mitigating Byzantine attacks in FL.
Junyu Shi, Shengshan Hu, Jianrong Lu, Leo Yu Zhang
TrustCom1
2018 Improved synchronisation algorithm based on reconstructed correlation function for BOC modulation in satellite navigation and positioning system
abstract
With the development of the new generation satellite navigation and positioning systems utilise binary offset carrier (BOC) modulation to improve inter‐operability. The main shortcoming of BOC modulation is its ambiguity of searching for a multi‐peaked auto‐correlation function (ACF). In this study, an improved unambiguous synchronisation algorithm based on compensated correlation reconstructed technique (CCRT) is proposed for the arbitrary‐order BOC, which is a new modulation technique for navigation modernisation. The key technique used in this study is to generate local reference signals with different shape vectors. By recombining the piecewise correlation functions of the step‐shape coded symbol, a reconstructed correlation function which has only one peak is obtained, and the energy loss is compensated by the ACF. In the proposed algorithm, different shape vectors are employed in different BOC modulation signals. The simulation results demonstrate that the ambiguity problem is completely eliminated. Compared with traditional synchronisation methods, the proposed algorithm shows outstanding detection and multipath mitigation performance. Moreover, the CCRT can be utilised for arbitrary stage BOC modulation signals.
Hailiang Xiong, Songhua Wang, Shu Gong, Meixuan Peng, Junyu Shi
IET Commun.5
2018 Transferring deep knowledge for object recognition in Low-quality underwater videos
Xin Sun 0003, Junyu Shi, Lipeng Liu, Junyu Dong, Claudia Plant, Huiyu Zhou 0001
Neurocomputing2