Yushan Zhang

dblp:50/10702 · DBLP profile ↗
← Back
18ranked-venue papers
8as first author
12since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 5 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorSoftware engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2026 Reduced Complexity Blind Recognition Method of LDPC Codes Over a Candidate Set
abstract
Adaptive modulation and coding (AMC) systems require the transmission of control signals, thereby reducing overall system transmission efficiency. The channel coding blind recognition technique is key to solving this problem. This paper proposes a reduced-complexity method for blind recognition of low-density parity-check (LDPC) coding parameters within a given candidate set, thereby enhancing the work efficiency of AMC systems. This paper applies a method based on code rate for classification evaluation, circumventing superfluous calculations for several candidate parity-check matrices. Furthermore, computational complexity is reduced by applying the offset min-sum algorithm (OMSA) to the parity-check stage. Subsequently, the Z-score is used to measure the difference between the actual data and the theoretical distribution. Compared with the best existing recognition methods, the proposed algorithm offers clear advantages in computational complexity and is virtually identical in recognition performance.
Zhuolun Wu, Yushan Zhang, Wei Zhang 0055, Yanyan Liu 0001
IEEE Signal Process. Lett.2
2025 Zero-Shot 4D Lidar Panoptic Segmentation
abstract
Zero-shot 4D segmentation and recognition of arbitrary objects in Lidar is crucial for embodied navigation, with applications ranging from streaming perception to semantic mapping and localization. However, the primary challenge in advancing research and developing generalized, versatile methods for spatio-temporal scene understanding in Lidar lies in the scarcity of datasets that provide the necessary diversity and scale of annotations. To overcome these challenges, we propose SAL-4D (Segment Anything in Lidar- 4D), a method that utilizes multi-modal robotic sensor setups as a bridge to distill recent developments in Video Object Segmentation (VOS) in conjunction with off-the-shelf Vision-Language foundation models to Lidar. We utilize VOS models to pseudo-label tracklets in short video sequences, annotate these tracklets with sequence-level CLIP tokens, and lift them to the 4D Lidar space using calibrated multi-modal sensory setups to distill them to our SAL-4D model. Due to temporal consistent predictions, we outperform prior art in 3D Zero-Shot Lidar Panoptic Segmentation (LPS) over 5 PQ, and unlock Zero-Shot 4D-LPS.
Yushan Zhang, Aljosa Osep, Laura Leal-Taixé, Tim Meinhardt
CVPR1
2025 Towards Improved Deep Metric Learning via Unsupervised Object Location
abstract
Deep Metric Learning (DML) aims at learning the representation of fine-grained data, which plays a vital role in fine-grained retrieval and classification applications. Previous studies have shown that cropping an object based on its bounding box (BBox) can significantly enhance performance. However, manually annotating the BBox is costly. In this paper, we propose to predict the BBox unsupervisedly, termed Towards Improved Deep Metric Learning via Unsupervised Object Location (DML-OL). DML-OL first proposes a BBoxNN to predict the BBox of the object. Building on this, DML-OL further introduces a CRNN, which crops and resizes the image based on the predicted BBox. Unlike the non-differentiable naive crop operator, the CRNN is designed to be fully differentiable. This differentiability property enables the model to be trained end-to-end using the pretext task without the BBox labels. We evaluate our proposed DML-OL on two strong baselines, and the results show that DML-OL outperforms the compared methods. Additionally, we demonstrate that DML-OL can predict the BBox accurately without the BBox labels.
Changxin Ye, Yushan Zhang, Wei Huangfu, Cheng Deng 0002
ICME2
2025 DeltaFlow: An Efficient Multi-frame Scene Flow Estimation Method
abstract
Previous dominant methods for scene flow estimation focus mainly on input from two consecutive frames, neglecting valuable information in the temporal domain. While recent trends shift towards multi-frame reasoning, they suffer from rapidly escalating computational costs as the number of frames grows. To leverage temporal information more efficiently, we propose DeltaFlow ($\Delta$Flow), a lightweight 3D framework that captures motion cues via a $\Delta$ scheme, extracting temporal features with minimal computational cost, regardless of the number of frames. Additionally, scene flow estimation faces challenges such as imbalanced object class distributions and motion inconsistency. To tackle these issues, we introduce a Category-Balanced Loss to enhance learning across underrepresented classes and an Instance Consistency Loss to enforce coherent object motion, improving flow accuracy. Extensive evaluations on the Argoverse 2, Waymo and nuScenes datasets show that $\Delta$Flow achieves state-of-the-art performance with up to 22\% lower error and $2\times$ faster inference compared to the next-best multi-frame supervised method, while also demonstrating a strong cross-domain generalization ability. The code is open-sourced at https://github.com/Kin-Zhang/DeltaFlow along with trained model weights.
Qingwen Zhang, Yushan Zhang, Yixi Cai, Olov Andersson, Patric Jensfelt
NeurIPS3
2025 Blind Recognition Algorithm of RS-SPC Concatenated Codes Based on Single-Error Correction
abstract
This letter presents an RS-SPC concatenated code blindrecognition algorithm based on the single-error correction. The algorithm corrects the least reliable bit of the single parity check (SPC) codewords based on the parity check characteristics, thereby increasing correct Reed-Solomon (RS) codewords and laying the foundation for improving the recognition probability. In addition, this algorithm combines threshold judgement with the matrix recording method, thereby eliminating unnecessary iterative operations under the condition that accurate recognition is possible. At the same time, it employs probability theory as a theoretical basis to quantify the degree of dispersion of the data through sample variance. The experimental results demonstrate that the recognition probability of this algorithm is superior to that of all other algorithms. When the codeword error rate (CER) is 0.5, RS(15,9)-SPC(4,3) still has a recognition probability of 20%. For RS(255,239)-SPC(8,7), the gain of the proposed algorithm exceeds 1.3dB compared to the upper bound of the recognition probability.
Zhuolun Wu, Wei Zhang 0055, Yihan Wang 0002, Yushan Zhang, Yanyan Liu 0001
IEEE Signal Process. Lett.4
2024 DP-Discriminator: A Differential Privacy Evaluation Tool Based on GAN
abstract
Differential privacy has become increasingly popular in private machine learning applications due to its provable ability to limit information leakage. However, there are often vulnerabilities in the practical implementation of differentially private algorithms, making it necessary to have effective tools for evaluating them before deployment. Unfortunately, the current state of the art classifier-based evaluation tools for differential privacy are still weakly distinguishable and need to be improved. In this paper, we propose a DP-Discriminator to automatically detect the ξ-differential distinguishability (ξ-DD) for specific algorithms, which is able to efficiently discover violations of differential privacy. Specially, we give a new attack definition of ξ-DD, based on a mathematical observation, which conduce to find a larger ξ. In addition, the proposed DP-Discriminator learns the overall distribution of features across samples depending on the ability of the generating adversarial network to capture the latent features. Benefiting from powerful classifiers, DP-Discriminator is able to automatically and accurately evaluate differential privacy with minimal time consumption. The experimental results demonstrate the effectiveness of the proposed method in estimating the ξ-DD for various practical randomized algorithms. For example, the latest work that detects the ξ-DD of the algorithm RAPPOR(0.4-DP) is 0.301, whereas our tool detects ξ-DD=0.369, with an error that is one order of magnitude lower.
Yushan Zhang, Xiaoyan Liang, Ruizhong Du
CF1
2024 DiffSF: Diffusion Models for Scene Flow Estimation
abstract
Scene flow estimation is an essential ingredient for a variety of real-world applications, especially for autonomous agents, such as self-driving cars and robots. While recent scene flow estimation approaches achieve reasonable accuracy, their applicability to real-world systems additionally benefits from a reliability measure. Aiming at improving accuracy while additionally providing an estimate for uncertainty, we propose DiffSF that combines transformer-based scene flow estimation with denoising diffusion models. In the diffusion process, the ground truth scene flow vector field is gradually perturbed by adding Gaussian noise. In the reverse process, starting from randomly sampled Gaussian noise, the scene flow vector field prediction is recovered by conditioning on a source and a target point cloud. We show that the diffusion process greatly increases the robustness of predictions compared to prior approaches resulting in state-of-the-art performance on standard scene flow estimation benchmarks. Moreover, by sampling multiple times with different initial states, the denoising process predicts multiple hypotheses, which enables measuring the output uncertainty, allowing our approach to detect a majority of the inaccurate predictions. The code is available at https://github.com/ZhangYushan3/DiffSF.
Yushan Zhang, Bastian Wandt, Maria Magnusson, Michael Felsberg
NeurIPS1
2024 High-fidelity Pseudo-labels for Boosting Weakly-Supervised Segmentation
abstract
Image-level weakly-supervised semantic segmentation (WSSS) reduces the usually vast data annotation cost by surrogate segmentation masks during training. The typical approach involves training an image classification network using global average pooling (GAP) on convolutional feature maps. This enables the estimation of object locations based on class activation maps (CAMs), which identify the importance of image regions. The CAMs are then used to generate pseudo-labels, in the form of segmentation masks, to supervise a segmentation model in the absence of pixel-level ground truth. Our work is based on two techniques for improving CAMs; importance sampling, which is a substitute for GAP, and the feature similarity loss, which utilizes a heuristic that object contours almost always align with color edges in images. However, both are based on the multinomial posterior with softmax, and implicitly assume that classes are mutually exclusive, which turns out suboptimal in our experiments. Thus, we reformulate both techniques based on binomial posteriors of multiple independent binary problems. This has two benefits; their performance is improved and they become more general, resulting in an add-on method that can boost virtually any WSSS method. This is demonstrated on a wide variety of baselines on the PASCAL VOC dataset, improving the region similarity and contour quality of all implemented stateof-the-art methods. Experiments on the MS COCO dataset further show that our proposed add-on is well-suited for large-scale settings. Our code implementation is available at https://github.com/arvijj/hfpl.
Arvi Jonnarth, Yushan Zhang, Michael Felsberg
WACV2
2023 Leveraging Optical Flow Features for Higher Generalization Power in Video Object Segmentation
abstract
We propose to leverage optical flow features for higher generalization power in semi-supervised video object segmentation. Optical flow is usually exploited as additional guidance information in many computer vision tasks. However, its relevance in video object segmentation was mainly in unsupervised settings or using the optical flow to warp or refine the previously predicted masks. Different from the latter, we propose to directly leverage the optical flow features in the target representation. We show that this enriched representation improves the encoder-decoder approach to the segmentation task. A model to extract the combined information from the optical flow and the image is proposed, which is then used as input to the target model and the decoder network. Unlike previous methods, e.g. in tracking where concatenation is used to integrate information from image data and optical flow, a simple yet effective attention mechanism is exploited in our work. Experiments on DAVIS 2017 and YouTube-VOS 2019 show that integrating the information extracted from optical flow into the original image branch results in a strong performance gain, especially in unseen classes which demonstrates its higher generalization power.
Yushan Zhang, Andreas Robinson, Maria Magnusson, Michael Felsberg
ICIP1
2023 GMSF: Global Matching Scene Flow
abstract
We tackle the task of scene flow estimation from point clouds. Given a source and a target point cloud, the objective is to estimate a translation from each point in the source point cloud to the target, resulting in a 3D motion vector field. Previous dominant scene flow estimation methods require complicated coarse-to-fine or recurrent architectures as a multi-stage refinement. In contrast, we propose a significantly simpler single-scale one-shot global matching to address the problem. Our key finding is that reliable feature similarity between point pairs is essential and sufficient to estimate accurate scene flow. We thus propose to decompose the feature extraction step via a hybrid local-global-cross transformer architecture which is crucial to accurate and robust feature representations. Extensive experiments show that the proposed Global Matching Scene Flow (GMSF) sets a new state-of-the-art on multiple scene flow estimation benchmarks. On FlyingThings3D, with the presence of occlusion points, GMSF reduces the outlier percentage from the previous best performance of 27.4% to 5.6%. On KITTI Scene Flow, without any fine-tuning, our proposed method shows state-of-the-art performance. On the Waymo-Open dataset, the proposed method outperforms previous methods by a large margin. The code is available at https://github.com/ZhangYushan3/GMSF.
Yushan Zhang, Johan Edstedt, Bastian Wandt, Per-Erik Forssén, Maria Magnusson, Michael Felsberg
NeurIPS1
2023 MFA: Multi-layer Feature-aware Attack for Object Detection
abstract
Physical adversarial attacks can mislead detectors in real-world scenarios and have attracted increasing attention. However, most existing works manipulate the detector’s final outputs as attack targets while ignoring the inherent characteristics of objects. This can result in attacks being trapped in model-specific local optima and reduced transferability. To address this issue, we propose a Multi-layer Feature-aware Attack (MFA) that considers the importance of multi-layer features and disrupts critical object-aware features that dominate decision-making across different models. Specifically, we leverage the location and category information of detector outputs to assign attribution scores to different feature layers. Then, we weight each feature according to their attribution results and design a pixel-level loss function in the opposite optimized direction of object detection to generate adversarial camouflages. We conduct extensive experiments in both digital and physical worlds on ten outstanding detection models and demonstrate the superior performance of MFA in terms of attacking capability and transferability. Our code is available at: \url{https://github.com/ChenWen1997/MFA}.
Yushan Zhang, Yuehuan Wang
UAI2
2023 Pinolo: Detecting Logical Bugs in Database Management Systems with Approximate Query Synthesis
Zongyin Hao, Quanfeng Huang, Chengpeng Wang 0001, Yushan Zhang, Rongxin Wu, Charles Zhang 0001
USENIX ATC5
2020 An Experimental Method to Estimate Running Time of Evolutionary Algorithms for Continuous Optimization
abstract
Running time analysis is a fundamental problem of critical importance in evolutionary computation. However, the analysis results have rarely been applied to advanced evolutionary algorithms (EAs) in practice, let alone their variants for continuous optimization. In this paper, an experimental method is proposed for analyzing the running time of EAs that are widely used for solving continuous optimization problems. Based on Glivenko-Cantelli theorem, the proposed method simulates the distribution of gain, which is introduced by average gain model to characterize progress during the optimization process. Data fitting techniques are subsequently adopted to obtain a desired function for further analyses. To verify the validity of the proposed method, experiments were conducted to estimate the upper bounds on expected first hitting time of various evolutionary strategies, such as (1, $\lambda $ ) evolution strategy, standard evolution strategy, covariance matrix adaptation evolution strategy, and its improved variants. The results suggest that all estimated upper bounds are correct. Backed up by the proposed method, state-of-the-art EAs for continuous optimization will have identical results about the running time as simplified schemes, which will bridge the gap between theoretical foundation and applications of evolutionary computation.
Han Huang 0002, Junpeng Su, Yushan Zhang
IEEE Trans. Evol. Comput.3
2019 Runtime Analysis of Pigeon-Inspired Optimizer Based on Average Gain Model
abstract
The pigeon-inspired optimization (PIO) algorithm is a novel swarm intelligence optimizer inspired by the homing behaviors of pigeons. Although PIO has demonstrated effectiveness and superiority in numerous fields, there are few results about the theoretical foundation of PIO. This paper employs the average gain model to estimate the upper bound for the expected first hitting time of PIO in continuous optimization. The case study and experiment result indicate that our theoretical analysis is applicable to the general case where the population size and problem size are both larger than 1, which is close to the practical situation.
Yushan Zhang, Han Huang 0002, Zhou Hong
CEC1
2019 Theoretical analysis of the convergence property of a basic pigeon-inspired optimizer in a continuous search space
Yushan Zhang, Han Huang 0002, Hongyue Wu
Sci. China Inf. Sci.1
2019 EdSketch: execution-driven sketching for Java
Jinru Hua, Yushan Zhang, Yuqun Zhang, Sarfraz Khurshid
Int. J. Softw. Tools Technol. Transf.2
2018 Running-time analysis of evolutionary programming based on Lebesgue measure of searching space
Yushan Zhang, Han Huang 0002, Guiwu Hu
Neural Comput. Appl.1
2011 A Method to Improve Performance of Heteroassociative Morphological Memories
Naiqin Feng, Yushan Zhang, Lianhui Ao, Shuangxi Wang
ICIC (2)2