Yujing Sun 0001

dblp:64/8656-1 · DBLP profile ↗
← Back
21ranked-venue papers
8as first author
18since 2021 · last 2026
0000-0003-0819-296XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 13 · 6 first-author · 10 since 2021Artificial intelligence and machine learning · 10 · 10 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 3 since 2021
YearPublicationVenuePosition
2026 LIMA: Towards building a non-invasive and stealthy real-world adversarial attack model for traffic sign recognition systems
Yujing Sun 0001, Canjian Jiang, You Jiang, Hezhong Pan, Siu-Ming Yiu, Zoe Lin Jiang
Neural Networks3
2026 OptimalCap: Efficient and Robust LiDAR-Based Motion Capture in Free Environments
abstract
LiDAR-based human motion capture holds great promise for large-scale, unconstrained environments. However, existing approaches often rely on clean, pre-segmented point clouds and struggle with noisy or dynamic scenes, limiting their practical applicability. We propose OptimalCap, a robust and efficient LiDAR-based framework that integrates hierarchical skeletal modeling and kinematic-aware temporal optimization to enable accurate, coherent, and real-time multi-human motion capture. To support training and evaluation under realistic disturbances, we also introduce NoiseMotion, a large-scale synthetic dataset simulating human-object interactions in noisy environments. Extensive experiments on public and synthetic benchmarks demonstrate that OptimalCap achieves state-of-the-art accuracy, robustness, and temporal consistency, while supporting over 20 individuals, at 60 FPS and up to 100 meters, setting a new standard for scalable, real-world LiDAR-based motion capture.
Yiming Ren 0001, Yujing Sun 0001, Yichen Yao 0001, Xiaoxiao Long, Xinge Zhu, Siu-Ming Yiu, Yuexin Ma
IEEE Trans. Pattern Anal. Mach. Intell.2
2025 UniDemoiré: Towards Universal Image Demoiréing with Data Generation and Synthesis
abstract
Image demoiréing poses one of the most formidable challenges in image restoration, primarily due to the unpredictable and anisotropic nature of moiré patterns. Limited by the quantity and diversity of training data, current methods tend to overfit to a single moiré domain, resulting in performance degradation for new domains, and restricting their robustness in real-world applications. In this paper, we propose a universal image demoiréing solution, UniDemoiré, which has superior generalization capability. Notably, we propose innovative and effective data generation and synthesis methods that can automatically provide vast high-quality moiré images to train a universal demoiréing model. Our extensive experiments demonstrate the cutting-edge performance and broad potential of our approach for generalized image demoiréing.
Zemin Yang, Yujing Sun 0001, Xidong Peng, Siu-Ming Yiu, Yuexin Ma
AAAI2
2025 Extreme Two-View Geometry From Object Poses with Diffusion Models
Yujing Sun 0001, Caiyi Sun, Yuan Liu 0025, Yuexin Ma, Siu-Ming Yiu
CVM (2)1
2025 FreqPolicy: Frequency Autoregressive Visuomotor Policy with Continuous Tokens
abstract
Learning effective visuomotor policies for robotic manipulation is challenging, as it requires generating precise actions while maintaining computational efficiency. Existing methods remain unsatisfactory due to inherent limitations in the essential action representation and the basic network architectures. We observe that representing actions in the frequency domain captures the structured nature of motion more effectively: low-frequency components reflect global movement patterns, while high-frequency components encode fine local details. Additionally, robotic manipulation tasks of varying complexity demand different levels of modeling precision across these frequency bands. Motivated by this, we propose a novel paradigm for visuomotor policy learning that progressively models hierarchical frequency components. To further enhance precision, we introduce continuous latent representations that maintain smoothness and continuity in the action space. Extensive experiments across diverse 2D and 3D robotic manipulation benchmarks demonstrate that our approach outperforms existing methods in both accuracy and efficiency, showcasing the potential of a frequency-domain autoregressive framework with continuous tokens for generalized robotic manipulation.Code is available at https://github.com/4DVLab/Freqpolicy
Yiming Zhong 0001, Chuyang Xiao, Zemin Yang, Youzhuo Wang, Ye Shi 0001, Yujing Sun 0001, Xinge Zhu, Yuexin Ma
NeurIPS8
2025 Generalizable Single-View Object Pose Estimation by Two-Side Generating and Matching
abstract
In this paper, we present a novel generalizable object pose estimation method to determine the object pose using only one RGB image. Unlike traditional approaches that rely on instance-level object pose estimation and necessitate extensive training data, our method offers generalization to unseen objects without extensive training, operates with a single reference image of the object, and eliminates the need for 3D object models or multiple views of the object. These characteristics are achieved by utilizing a diffusion model to generate novel-view images and conducting a two-sided matching on these generated images. Quantitative experiments demonstrate the superiority of our method over existing pose estimation techniques across both synthetic and real-world datasets. Remarkably, our approach maintains strong performance even in scenarios with significant viewpoint changes, highlighting its robustness and versatility in challenging conditions. The code will be released at https://github.com/scy639/Gen2SM.
Yujing Sun 0001, Caiyi Sun, Yuan Liu 0025, Yuexin Ma, Siu-Ming Yiu
WACV1
2025 A real world attack model combining LED modulation and attention-superpixel guidance
You Jiang, Yuqiao Luo, Canjian Jiang, Yinglong Liao, Yujing Sun 0001, Siu-Ming Yiu, Chuanyi Liu, Zoe Lin Jiang
Expert Syst. Appl.7
2025 Imperceptible Physical Attack Against Face Recognition Systems via LED Illumination Modulation
abstract
Although face recognition starts to play an important role in our daily life, we need to pay attention that data-driven face recognition vision systems are vulnerable to adversarial attacks. However, current digital adversarial attacks and physical adversarial attacks both have drawbacks, with the former ones impractical and the latter one conspicuous, high-computational and low-executable. To address the issues, we propose a practical, executable, stealthy and low computational adversarial attack based on LED illumination modulation. To fool the systems, the proposed attack generates physically imperceptible luminance changes to human eyes through fast intensity modulation of scene LED illumination and uses the rolling shutter effect of CMOS image sensors in face recognition systems to implant luminance information perturbation to the captured face images. In summary, we present a denial-of-service (DoS) attack for face detection and an evasion attack for face verification. We also evaluate their effectiveness against wellknown face detection models, Dlib, MTCNN and RetinaFace, and face verification models, Dlib, FaceNet, and ArcFace. The extensive physical experiments show that the success rates of DoS attacks against face detection models reach 97.67%, 100%, and 100%, respectively, and the success rates of evasion attacks against all face verification models reach 100%.
Canjian Jiang, You Jiang, Puxi Lin, Zhaojie Chen, Yujing Sun 0001, Siu-Ming Yiu, Zoe Lin Jiang
IEEE Trans. Big Data6
2025 Restoration of Recaptured Screen Images With a Divide and Conquer Strategy
abstract
Moiré patterns in recaptured screen images are image defects that can affect image quality to an extreme extent. Different from other image defects, moiré artefacts can vary greatly in scales, colours and shapes. Such moiré patterns mix with image content and disturb image features of different scales in different ways, making moiré pattern removal a challenging task. In this paper, we present a novel divide-and-conquer strategy to solve the problem. In the divide stage, we innovatively decompose images into different layers, as well as into structure components and detail components. Then in the conquer stage, guided by the layers retrieved from the divide stage, we can restore coarse and fine image components independently, which greatly improve the demoiréing performance. Our strategy outperforms state-of-arts in both quantitative and qualitative evaluations.
Yujing Sun 0001, Hao Xiong 0002, Siu-Ming Yiu
IEEE Trans. Big Data1
2024 HUNTER: Unsupervised Human-Centric 3D Detection via Transferring Knowledge from Synthetic Instances to Real Scenes
abstract
Human-centric 3D scene understanding has recently drawn increasing attention, driven by its critical impact on robotics. However, human-centric real-life scenarios are extremely diverse and complicated, and humans have intri-cate motions and interactions. With limited labeled data, supervised methods are difficult to generalize to general scenarios, hindering real-life applications. Mimicking human intelligence, we propose an unsupervised 3D detection method for human-centric scenarios by transferring the knowledge from synthetic human instances to real scenes. To bridge the gap between the distinct data representations and feature distributions of synthetic models and real point clouds, we introduce novel modules for effective instance-to-scene representation transfer and synthetic-to-real feature alignment. Remarkably, our method exhibits superior performance compared to current state-of-the-art techniques, achieving 87.8% improvement in mAP and closely approaching the performance of fully supervised methods (62.15 mAP vs. 69.02 mAP) on HuCenLife Dataset.
Yichen Yao 0001, Zimo Jiang, Yujing Sun 0001, Zhencai Zhu, Xinge Zhu, Runnan Chen, Yuexin Ma
CVPR3
2024 WildRefer: 3D Object Localization in Large-Scale Dynamic Scenes with Multi-modal Visual Data and Natural Language
Zhenxiang Lin, Xidong Peng, Peishan Cong, Ge Zheng, Yujing Sun 0001, Yuenan Hou, Xinge Zhu, Sibei Yang, Yuexin Ma
ECCV (46)5
2024 Learning to Adapt SAM for Segmenting Cross-Domain Point Clouds
Xidong Peng, Runnan Chen, Feng Qiao 0001, Lingdong Kong, Youquan Liu, Yujing Sun 0001, Xinge Zhu, Yuexin Ma
ECCV (43)6
2024 LiveHPS++: Robust and Coherent Motion Capture in Dynamic Free Environment
Yiming Ren 0001, Yichen Yao 0001, Xiaoxiao Long, Yujing Sun 0001, Yuexin Ma
ECCV (29)5
2024 FreeReg: Image-to-Point Cloud Registration Leveraging Pretrained Diffusion Models and Monocular Depth Estimators
abstract
Matching cross-modality features between images and point clouds is a fundamental problem for image-to-point cloud registration. However, due to the modality difference between images and points, it is difficult to learn robust and discriminative cross-modality features by existing metric learning methods for feature matching. Instead of applying metric learning on cross-modality data, we propose to unify the modality between images and point clouds by pretrained large-scale models first, and then establish robust correspondence within the same modality. We show that the intermediate features, called diffusion features, extracted by depth-to-image diffusion models are semantically consistent between images and point clouds, which enables the building of coarse but robust cross-modality correspondences. We further extract geometric features on depth maps produced by the monocular depth estimator. By matching such geometric features, we significantly improve the accuracy of the coarse correspondences produced by diffusion features. Extensive experiments demonstrate that without any task-specific training, direct utilization of both features produces accurate image-to-point cloud registration. On three public indoor and outdoor benchmarks, the proposed method averagely achieves a 20.6 percent improvement in Inlier Ratio, a $3.0\times$ higher Inlier Number, and a 48.6 percent improvement in Registration Recall than existing state-of-the-arts. The code and additional results are available at \url{https://whu-usi3dv.github.io/FreeReg/}.
Haiping Wang 0004, Yuan Liu 0025, Bing Wang 0013, Yujing Sun 0001, Zhen Dong 0005, Wenping Wang 0001, Bisheng Yang
ICLR4
2024 Gait Recognition in Large-scale Free Environment via Single LiDAR
abstract
Human gait recognition is crucial in multimedia, enabling identification through walking patterns without direct interaction, enhancing the integration across various media forms in real-world applications like smart homes, healthcare and non-intrusive security. LiDAR's ability to capture depth makes it pivotal for robotic perception and holds promise for real-world gait recognition. In this paper, based on a single LiDAR, we present the Hierarchical Multi-representation Feature Interaction Network (HMRNet) for robust gait recognition. Prevailing LiDAR-based gait datasets primarily derive from controlled settings with predefined trajectory, remaining a gap with real-world scenarios. To facilitate LiDAR-based gait recognition research, we introduce FreeGait, a comprehensive gait dataset from large-scale, unconstrained settings, enriched with multi-modal and varied 2D/3D data. Notably, our approach achieves state-of-the-art performance on prior dataset (SUSTech1K) and on FreeGait. https://4dvlab.github.io/project_page/FreeGait.html
Yiming Ren 0001, Peishan Cong, Yujing Sun 0001, Jingya Wang 0001, Lan Xu 0003, Yuexin Ma
ACM Multimedia4
2024 Towards Practical Human Motion Prediction with LiDAR Point Clouds
abstract
Human motion prediction is crucial for human-centric multimedia understanding and interacting. Current methods typically rely on ground truth human poses as observed input, which is not practical for real-world scenarios where only raw visual sensor data is available. To implement these methods in practice, a pre-phrase of pose estimation is essential. However, such two-stage approaches often lead to performance degradation due to the accumulation of errors. Moreover, reducing raw visual data to sparse keypoint representations significantly diminishes the density of information, resulting in the loss of fine-grained features. In this paper, we propose LiDAR-HMP, the first single-LiDAR-based 3D human motion prediction approach, which receives the raw LiDAR point cloud as input and forecasts future 3D human poses directly. Building upon our novel structure-aware body feature descriptor, LiDAR-HMP adaptively maps the observed motion manifold to future poses and effectively models the spatial-temporal correlations of human motions for further refinement of prediction results. Extensive experiments show that our method achieves state-of-the-art performance on two public benchmarks and demonstrates remarkable robustness and efficacy in real-world deployments. https://4dvlab.github.io/project_page/LiDARHMP.html
Yiming Ren 0001, Yichen Yao 0001, Yujing Sun 0001, Yuexin Ma
ACM Multimedia4
2023 BitAnalysis: A Visualization System for Bitcoin Wallet Investigation
abstract
Bitcoin is gaining ever increasing popularity. However, professional skills are required if people want to check bitcoin transaction information from the blockchain. As pointed out in a recent study, there is a lack of tools to support effective interactive investigation of bitcoin transactions. Therefore, we present a novel visualization system,BitAnalysis, for interactive bitcoin wallet investigation. The analytical and visualization functions ofBitAnalysisare defined and developed by following the advice and requirements of a group of entrepreneurs and regulators of bitcoin-related business.BitAnalysisprovides a rich set of functions and intuitive visual interfaces for the users, such as law-enforcement officers and regulators, to effectively visualize and analyze the transactions of a bitcoin wallet (i.e., a cluster of bitcoin addresses) and its related wallets, to track the flow of bitcoins, and to identify wallet correlation using our novel clustering functions. To achieve these functions, we have designed new visualization techniques for presenting bitcoin transactions information and introduced theconnection diagramandbitcoin flow mapas new ways of analyzing, tracking and monitoring the trading activities of a cluster of closely related wallets. We also present an extensive user study that validated the effectiveness and usability ofBitAnalysis.
Yujing Sun 0001, Hao Xiong 0002, Siu-Ming Yiu, Kwok-Yan Lam
IEEE Trans. Big Data1
2021 Understanding deep face anti-spoofing: from the perspective of data
Yujing Sun 0001, Hao Xiong 0002, Siu-Ming Yiu
Vis. Comput.1
2018 Moiré Photo Restoration Using Multiresolution Convolutional Neural Networks
abstract
Digital cameras and mobile phones enable us to conveniently record precious moments. While digital image quality is constantly being improved, taking high-quality photos of digital screens still remains challenging because the photos are often contaminated with moiré patterns, a result of the interference between the pixel grids of the camera sensor and the device screen. Moiré patterns can severely damage the visual quality of photos. However, few studies have aimed to solve this problem. In this paper, we introduce a novel multiresolution fully convolutional network for automatically removing moiré patterns from photos. Since a moiré pattern spans over a wide range of frequencies, our proposed network performs a nonlinear multiresolution analysis of the input image before computing how to cancel moiré artefacts within every frequency band. We also create a large-scale benchmark dataset with 100,000+ image pairs for investigating and evaluating moiré pattern removal algorithms. Our network achieves state-of-the-art performance on this dataset in comparison to existing learning architectures for image restoration problems.
Yujing Sun 0001, Yizhou Yu, Wenping Wang 0001
IEEE Trans. Image Process.1
2018 Image Structure Retrieval via L0 Minimization
abstract
Retrieving salient structure from textured images is an important but difficult problem in computer vision because texture, which can be irregular, anisotropic, non-uniform and complex, shares many of the same properties as structure. Observing that salient structure in a textured image should be piece-wise smooth, we present a method to retrieve such structures using an minimization of a modified form of the relative total variation metric. Thanks to the characteristics shared by texture and small structures, our method is effective at retrieving structure based on scale as well. Our method outperforms state-of-art methods in texture removal as well as scale-space filtering. We also demonstrate our method's ability in other applications such as edge detection, clip art compression artifact removal, and inverse half-toning.
Yujing Sun 0001, Scott Schaefer, Wenping Wang 0001
IEEE Trans. Vis. Comput. Graph.1
2015 Denoising point sets via L0 minimization
Yujing Sun 0001, Scott Schaefer, Wenping Wang 0001
Comput. Aided Geom. Des.1