Yiyang Su

dblp:271/8192 · DBLP profile ↗
← Back
12ranked-venue papers
4as first author
12since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 SuperPhys-Net: A Physics-Informed Super-Resolution Electromagnetic Simulator for Nanophotonic Devices
abstract
The rapid advancement of photonic integrated circuits is driving innovations in interconnect, computing, and sensing applications. This progress has led to the development of nanophotonic waveguide devices with complex geometries, offering greater design flexibility and a wider range of functional applications. However, electromagnetic (EM) simulation imposes a heavy computational burden during the design and validation phases. This significantly hampers design iteration speed and scalability. Although existing data-driven methods and physics-informed neural networks have shown promise for simpler structures, they fall short for highly complex geometries, limiting the automation of photonic device design. To address these issues, we present the SuperPhys-Net framework. This innovative approach enhances coarse-grid simulation results through super-resolution and integrates physical constraints to generate fine-grid solutions that adhere to physical laws. Our model demonstrates outstanding performance across complex nanophotonic waveguide devices with varying dimensions, achieving a 72.61% improvement in accuracy over current state-of-the-art models. Additionally, it reduces computational time by 76.09% compared to standard finite-difference frequency-domain solvers, all while maintaining exceptional accuracy across all scales.
Yiyang Su, Guohao Dai 0003, Yuzhe Ma, Yeyu Tong
DATE1
2026 Dynamic Surgery Video Summarization With Balancing Informativeness and Diversity
abstract
Surgery video summarization can help medical professionals quickly gain the insight into the surgical process for the surgical education and skill evaluation. However, existing methods are unable to efficiently summarize information to satisfy medical professionals. Since it is challenging to summarize the video while balancing the information richness and diversity. In this paper, we propose a dynamic surgery video summarization framework (DSVS). We first used a multitask learning network to perceive and comprehend surgical action triplet components and phases. An information contribution module then measures the frame-level importance using the predicted triplets. A two-stage strategy which involves phase recognition and change-point detection further applied to divide each phase of the surgical videos into shots. Finally, A multi-objective zero-one programming model is formulated to select the optimal subset of shots by simultaneously maximizing intra-shot information contribution and minimizing inter-shot information similarity. Experimental results on two surgical video datasets show the framework can generate summaries that encompass crucial and diverse content. Clinical validations indicate the framework is capable of summarizing the information expected by surgeons. The source code can be found at https://github.com/syypretend/DSVS.
Hao Wang 0081, Yiyang Su, Xuefei Song, Xianqun Fan, Shuai Ding 0001
IEEE Trans. Medical Imaging2
2025 SapiensID: Foundation for Human Recognition
abstract
Existing human recognition systems often rely on separate, specialized models for face and body analysis, limiting their effectiveness in real-world scenarios where pose, visibility, and context vary widely. This paper introduces SapiensID, a unified model that bridges this gap, achieving robust performance across diverse settings. SapiensID introduces (i) Retina Patch (RP), a dynamic patch generation scheme that adapts to subject scale and ensures consistent tokenization of regions of interest, (ii) a masked recognition model (MRM) that learns from variable token length, and (iii) Semantic Attention Head (SAH), an module that learns pose-invariant representations by pooling features around key body parts. To facilitate training, we introduce WebBody4M, a large-scale dataset capturing diverse poses and scale variations. Extensive experiments demonstrate that SapiensID achieves state-of-the-art results on various body ReID benchmarks, outperforming specialized models in both short-term and long-term scenarios while remaining competitive with dedicated face recognition systems. Furthermore, SapiensID establishes a strong baseline for the newly introduced challenge of Cross Pose-Scale ReID, demonstrating its ability to generalize to complex, real-world conditions. Project Link
Dingqiang Ye, Yiyang Su, Feng Liu 0037, Xiaoming Liu 0002
CVPR3
2025 HAMoBE: Hierarchical and Adaptive Mixture of Biometric Experts for Video-Based Person ReID
Yiyang Su, Yunping Shi
ICCV1
2025 A Quality-Guided Mixture of Score-Fusion Experts Framework for Human Recognition
abstract
Whole-body biometric recognition is a challenging multimodal task that integrates various biometric modalities, including face, gait, and body. This integration is essential for overcoming the limitations of unimodal systems. Traditionally, whole-body recognition involves deploying different models to process multiple modalities, achieving the final outcome by score-fusion (e.g., weighted averaging of similarity matrices from each model). However, these conventional methods may overlook the variations in score distributions of individual modalities, making it challenging to improve final performance. In this work, we present \textbf{Q}uality-guided \textbf{M}ixture of score-fusion \textbf{E}xperts (QME), a novel framework designed for improving whole-body biometric recognition performance through a learnable score-fusion strategy using a Mixture of Experts (MoE). We introduce a novel pseudo-quality loss for quality estimation with a modality-specific Quality Estimator (QE), and a score triplet loss to improve the metric performance. Extensive experiments on multiple whole-body biometric datasets demonstrate the effectiveness of our proposed approach, achieving state-of-the-art results across various metrics compared to baseline methods. Our method is effective for multimodal and multi-model, addressing key challenges such as model misalignment in the similarity score domain and variability in data quality.
Yiyang Su, Anil K. Jain 0001, Xiaoming Liu 0002
ICCV2
2025 A Sparse Spatial Spectrum Reconstruction Algorithm in the Second-Order Statistical Domain
abstract
Traditional sparse recovery-based DOA estimation algorithms can achieve high-precision DOA estimation. However, these algorithms process the signal in receive domain, which has a large complexity. Moreover, traditional algorithms are proposed based on scalar sensor arrays, which cannot exploit the unique polarization information of electromagnetic waves. Therefore, in this paper, a sparse reconstruction algorithm in the second-order statistical domain is proposed for a more complex polarizationsensitive arrays. We firstly reconstruct the received domain signal to achieve parameter decoupling, then obtain the DOA estimation by reconstructing the sparse power spectrum, and finally obtain the polarization parameter by searching the spectral peaks. Simulation experiments demonstrate that the proposed method has higher estimation accuracy and lower complexity than the traditional algorithms.
Yiyang Su
VTC2025-Spring3
2024 KeyPoint Relative Position Encoding for Face Recognition
abstract
In this paper, we address the challenge of making ViT models more robust to unseen affine transformations. Such robustness becomes useful in various recognition tasks such as face recognition when image alignment failures occur. We propose a novel method called KP-RPE, which leverages key points (e.g. facial landmarks) to make ViT more resilient to scale, translation, and pose variations. We begin with the observation that Relative Position Encoding (RPE) is a good way to bring affine transform generalization to ViTs. RPE, however, can only inject the model with prior knowledge that nearby pixels are more important than far pixels. Keypoint RPE (KP-RPE) is an extension of this principle, where the significance of pixels is not solely dictated by their proximity but also by their relative positions to specific keypoints within the image. By anchoring the significance of pixels around keypoints, the model can more effectively retain spatial relationships, even when those relationships are disrupted by affine transformations. We show the merit of KP-RPE inface and gait recognition. The experimental results demonstrate the effectiveness in improving face recognition performance from low-quality images, particularly where alignment is prone to failure. Code and pre-trained models are available.
Yiyang Su, Feng Liu 0037, Xiaoming Liu 0002
CVPR2
2024 Open-Set Biometrics: Beyond Good Closed-Set Models
Yiyang Su, Feng Liu 0037, Anil K. Jain 0001, Xiaoming Liu 0002
ECCV (62)1
2024 FarSight: A Physics-Driven Whole-Body Biometric System at Large Distance and Altitude
abstract
Whole-body biometric recognition is an important area of research due to its vast applications in law enforcement, border security, and surveillance. This paper presents the end-to-end design, development and evaluation of FarSight, an innovative software system designed for whole-body (fusion of face, gait and body shape) biometric recognition. FarSight accepts videos from elevated platforms and drones as input and outputs a candidate list of identities from a gallery. The system is designed to address several challenges, including (i) low-quality imagery, (ii) large yaw and pitch angles, (iii) robust feature extraction to accommodate large intra-person variabilities and large inter-person similarities, and (iv) the large domain gap between training and test sets. FarSight combines the physics of imaging and deep learning models to enhance image restoration and biometric feature encoding. We test FarSight’s effectiveness using the newly acquired IARPA Biometric Recognition and Identification at Altitude and Range (BRIAR) dataset. Notably, FarSight demonstrated a substantial performance increase on the BRIAR dataset, with gains of +11.82% Rank-20 identification and +11.30% TAR@1% FAR.
Feng Liu 0037, Ryan Ashbaugh, Nicholas Chimitt, Najmul Hassan, Ali Hassani 0001, Ajay Jaiswal, Zhiyuan Mao, Christopher Perry, Yiyang Su, Pegah Varghaei, Kai Wang 0058, Stanley H. Chan, Arun Ross, Humphrey Shi, Zhangyang Wang, Xiaoming Liu 0002
WACV11
2024 Low-Rank Tensor Completion Pansharpening Based on Haze Correction
abstract
Pansharpening refers to the fusion between a multispectral (MS) image with abundant spectral information and a panchromatic (PAN) image with high spatial resolution to obtain a high spatial resolution multispectral (HRMS) image. The traditional pansharpening methods often ignore the effect of path-radiation caused by scattering from different atmospheric components, and the few methods that introduce haze correction only calibrate each band of the MS image individually, without exploring the intrinsic correlation among different bands. To address this problem, low rank tensor completion pansharpening based on haze correction (LRTCP) is proposed. The haze-line prior is first introduced into the joint haze correction of MS and PAN images, and obtain the pre-modulated images with the help of the improved high-pass modulation (HPM) injection scheme. We then use tensor completion to simulate the degradation problem by applying low-tubal-rank tensor complementation to the process of reconstructing HRMS images, thus constructing a low rank tensor completion pansharpening model based on haze correction. Finally, the alternating direction multiplier (ADMM) is employed to find the solution of the proposed approach, producing the final fusion result. Comprehensive qualitative and quantitative assessment of reduced- and full-resolution datasets from different satellites shows that the proposed method outperforms the state-of-the-art methods.
Peng Wang 0030, Yiyang Su, Bo Huang 0001, Daiyin Zhu, Alexandr Nedzved, Viktor V. Krasnoproshin, Henry Leung 0001
IEEE Trans. Geosci. Remote. Sens.2
2023 Multispectral Pansharpening Based on High-Pass Modulation Injection Model with Difference Factor
abstract
This paper proposes a multispectral pansharpening based on high-pass modulation injection model with difference factor (HPM-DF), which is dedicated to solving the problem of difference between multispectral image and panchromatic image acquired at different moments. In the high-pass modulation injection model, we introduce a difference factor and use the alternating direction method of multipliers (ADMM) to fully analyze the difference variability to derive the final fusion product. Experiments assessed at both reduced and full resolution show that the proposed method can acquire better performance than the traditional pansharpening methods.
Yiyang Su, Peng Wang 0030, Xiwang Zhang
IGARSS1
2023 ChatGPT-Powered Hierarchical Comparisons for Image Classification
abstract
The zero-shot open-vocabulary setting poses challenges for image classification. Fortunately, utilizing a vision-language model like CLIP, pre-trained on image-text pairs, allows for classifying images by comparing embeddings. Leveraging large language models (LLMs) such as ChatGPT can further enhance CLIP’s accuracy by incorporating class-specific knowledge in descriptions. However, CLIP still exhibits a bias towards certain classes and generates similar descriptions for similar classes, disregarding their differences. To address this problem, we present a novel image classification framework via hierarchical comparisons. By recursively comparing and grouping classes with LLMs, we construct a class hierarchy. With such a hierarchy, we can classify an image by descending from the top to the bottom of the hierarchy, comparing image and text embeddings at each level. Through extensive experiments and analyses, we demonstrate that our proposed approach is intuitive, effective, and explainable. Code will be released upon publication.
Yiyang Su, Xiaoming Liu 0002
NeurIPS2