EDBT 2026 Demo / reviewers in the wild / expert
Yujing Sun 0001
dblp:64/8656-1
· DBLP profile ↗
21ranked-venue papers
8as first author
18since 2021 · last 2026
0000-0003-0819-296XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 13 · 6 first-author · 10 since 2021Artificial intelligence and machine learning · 10 · 10 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LIMA: Towards building a non-invasive and stealthy real-world adversarial attack model for traffic sign recognition systems
Yujing Sun 0001, Canjian Jiang, You Jiang, Hezhong Pan, Siu-Ming Yiu, Zoe Lin Jiang |
Neural Networks | 3 |
| 2026 | OptimalCap: Efficient and Robust LiDAR-Based Motion Capture in Free EnvironmentsabstractLiDAR-based human motion capture holds great promise for large-scale, unconstrained environments. However, existing approaches often rely on clean, pre-segmented point clouds and struggle with noisy or dynamic scenes, limiting their practical applicability. We propose OptimalCap, a robust and efficient LiDAR-based framework that integrates hierarchical skeletal modeling and kinematic-aware temporal optimization to enable accurate, coherent, and real-time multi-human motion capture. To support training and evaluation under realistic disturbances, we also introduce NoiseMotion, a large-scale synthetic dataset simulating human-object interactions in noisy environments. Extensive experiments on public and synthetic benchmarks demonstrate that OptimalCap achieves state-of-the-art accuracy, robustness, and temporal consistency, while supporting over 20 individuals, at 60 FPS and up to 100 meters, setting a new standard for scalable, real-world LiDAR-based motion capture. Yiming Ren 0001, Yujing Sun 0001, Yichen Yao 0001, Xiaoxiao Long, Xinge Zhu, Siu-Ming Yiu, Yuexin Ma |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | UniDemoiré: Towards Universal Image Demoiréing with Data Generation and SynthesisabstractImage demoiréing poses one of the most formidable challenges in image restoration, primarily due to the unpredictable and anisotropic nature of moiré patterns. Limited by the quantity and diversity of training data, current methods tend to overfit to a single moiré domain, resulting in performance degradation for new domains, and restricting their robustness in real-world applications. In this paper, we propose a universal image demoiréing solution, UniDemoiré, which has superior generalization capability. Notably, we propose innovative and effective data generation and synthesis methods that can automatically provide vast high-quality moiré images to train a universal demoiréing model. Our extensive experiments demonstrate the cutting-edge performance and broad potential of our approach for generalized image demoiréing. Zemin Yang, Yujing Sun 0001, Xidong Peng, Siu-Ming Yiu, Yuexin Ma |
AAAI | 2 |
| 2025 | Extreme Two-View Geometry From Object Poses with Diffusion Models
Yujing Sun 0001, Caiyi Sun, Yuan Liu 0025, Yuexin Ma, Siu-Ming Yiu |
CVM (2) | 1 |
| 2025 | FreqPolicy: Frequency Autoregressive Visuomotor Policy with Continuous TokensabstractLearning effective visuomotor policies for robotic manipulation is challenging, as it requires generating precise actions while maintaining computational efficiency. Existing methods remain unsatisfactory due to inherent limitations in the essential action representation and the basic network architectures. We observe that representing actions in the frequency domain captures the structured nature of motion more effectively: low-frequency components reflect global movement patterns, while high-frequency components encode fine local details. Additionally, robotic manipulation tasks of varying complexity demand different levels of modeling precision across these frequency bands. Motivated by this, we propose a novel paradigm for visuomotor policy learning that progressively models hierarchical frequency components. To further enhance precision, we introduce continuous latent representations that maintain smoothness and continuity in the action space. Extensive experiments across diverse 2D and 3D robotic manipulation benchmarks demonstrate that our approach outperforms existing methods in both accuracy and efficiency, showcasing the potential of a frequency-domain autoregressive framework with continuous tokens for generalized robotic manipulation.Code is available at https://github.com/4DVLab/Freqpolicy Yiming Zhong 0001, Chuyang Xiao, Zemin Yang, Youzhuo Wang, Ye Shi 0001, Yujing Sun 0001, Xinge Zhu, Yuexin Ma |
NeurIPS | 8 |
| 2025 | Generalizable Single-View Object Pose Estimation by Two-Side Generating and MatchingabstractIn this paper, we present a novel generalizable object pose estimation method to determine the object pose using only one RGB image. Unlike traditional approaches that rely on instance-level object pose estimation and necessitate extensive training data, our method offers generalization to unseen objects without extensive training, operates with a single reference image of the object, and eliminates the need for 3D object models or multiple views of the object. These characteristics are achieved by utilizing a diffusion model to generate novel-view images and conducting a two-sided matching on these generated images. Quantitative experiments demonstrate the superiority of our method over existing pose estimation techniques across both synthetic and real-world datasets. Remarkably, our approach maintains strong performance even in scenarios with significant viewpoint changes, highlighting its robustness and versatility in challenging conditions. The code will be released at https://github.com/scy639/Gen2SM. Yujing Sun 0001, Caiyi Sun, Yuan Liu 0025, Yuexin Ma, Siu-Ming Yiu |
WACV | 1 |
| 2025 | A real world attack model combining LED modulation and attention-superpixel guidance
You Jiang, Yuqiao Luo, Canjian Jiang, Yinglong Liao, Yujing Sun 0001, Siu-Ming Yiu, Chuanyi Liu, Zoe Lin Jiang |
Expert Syst. Appl. | 7 |
| 2025 | Imperceptible Physical Attack Against Face Recognition Systems via LED Illumination ModulationabstractAlthough face recognition starts to play an important role in our daily life, we need to pay attention that data-driven face recognition vision systems are vulnerable to adversarial attacks. However, current digital adversarial attacks and physical adversarial attacks both have drawbacks, with the former ones impractical and the latter one conspicuous, high-computational and low-executable. To address the issues, we propose a practical, executable, stealthy and low computational adversarial attack based on LED illumination modulation. To fool the systems, the proposed attack generates physically imperceptible luminance changes to human eyes through fast intensity modulation of scene LED illumination and uses the rolling shutter effect of CMOS image sensors in face recognition systems to implant luminance information perturbation to the captured face images. In summary, we present a denial-of-service (DoS) attack for face detection and an evasion attack for face verification. We also evaluate their effectiveness against wellknown face detection models, Dlib, MTCNN and RetinaFace, and face verification models, Dlib, FaceNet, and ArcFace. The extensive physical experiments show that the success rates of DoS attacks against face detection models reach 97.67%, 100%, and 100%, respectively, and the success rates of evasion attacks against all face verification models reach 100%. Canjian Jiang, You Jiang, Puxi Lin, Zhaojie Chen, Yujing Sun 0001, Siu-Ming Yiu, Zoe Lin Jiang |
IEEE Trans. Big Data | 6 |
| 2025 | Restoration of Recaptured Screen Images With a Divide and Conquer StrategyabstractMoiré patterns in recaptured screen images are image defects that can affect image quality to an extreme extent. Different from other image defects, moiré artefacts can vary greatly in scales, colours and shapes. Such moiré patterns mix with image content and disturb image features of different scales in different ways, making moiré pattern removal a challenging task. In this paper, we present a novel divide-and-conquer strategy to solve the problem. In the divide stage, we innovatively decompose images into different layers, as well as into structure components and detail components. Then in the conquer stage, guided by the layers retrieved from the divide stage, we can restore coarse and fine image components independently, which greatly improve the demoiréing performance. Our strategy outperforms state-of-arts in both quantitative and qualitative evaluations. Yujing Sun 0001, Hao Xiong 0002, Siu-Ming Yiu |
IEEE Trans. Big Data | 1 |
| 2024 | HUNTER: Unsupervised Human-Centric 3D Detection via Transferring Knowledge from Synthetic Instances to Real ScenesabstractHuman-centric 3D scene understanding has recently drawn increasing attention, driven by its critical impact on robotics. However, human-centric real-life scenarios are extremely diverse and complicated, and humans have intri-cate motions and interactions. With limited labeled data, supervised methods are difficult to generalize to general scenarios, hindering real-life applications. Mimicking human intelligence, we propose an unsupervised 3D detection method for human-centric scenarios by transferring the knowledge from synthetic human instances to real scenes. To bridge the gap between the distinct data representations and feature distributions of synthetic models and real point clouds, we introduce novel modules for effective instance-to-scene representation transfer and synthetic-to-real feature alignment. Remarkably, our method exhibits superior performance compared to current state-of-the-art techniques, achieving 87.8% improvement in mAP and closely approaching the performance of fully supervised methods (62.15 mAP vs. 69.02 mAP) on HuCenLife Dataset. Yichen Yao 0001, Zimo Jiang, Yujing Sun 0001, Zhencai Zhu, Xinge Zhu, Runnan Chen, Yuexin Ma |
CVPR | 3 |
| 2024 | WildRefer: 3D Object Localization in Large-Scale Dynamic Scenes with Multi-modal Visual Data and Natural Language
Zhenxiang Lin, Xidong Peng, Peishan Cong, Ge Zheng, Yujing Sun 0001, Yuenan Hou, Xinge Zhu, Sibei Yang, Yuexin Ma |
ECCV (46) | 5 |
| 2024 | Learning to Adapt SAM for Segmenting Cross-Domain Point Clouds
Xidong Peng, Runnan Chen, Feng Qiao 0001, Lingdong Kong, Youquan Liu, Yujing Sun 0001, Xinge Zhu, Yuexin Ma |
ECCV (43) | 6 |
| 2024 | LiveHPS++: Robust and Coherent Motion Capture in Dynamic Free Environment
Yiming Ren 0001, Yichen Yao 0001, Xiaoxiao Long, Yujing Sun 0001, Yuexin Ma |
ECCV (29) | 5 |
| 2024 | FreeReg: Image-to-Point Cloud Registration Leveraging Pretrained Diffusion Models and Monocular Depth EstimatorsabstractMatching cross-modality features between images and point clouds is a fundamental problem for image-to-point cloud registration. However, due to the modality difference between images and points, it is difficult to learn robust and discriminative cross-modality features by existing metric learning methods for feature matching. Instead of applying metric learning on cross-modality data, we propose to unify the modality between images and point clouds by pretrained large-scale models first, and then establish robust correspondence within the same modality. We show that the intermediate features, called diffusion features, extracted by depth-to-image diffusion models are semantically consistent between images and point clouds, which enables the building of coarse but robust cross-modality correspondences. We further extract geometric features on depth maps produced by the monocular depth estimator. By matching such geometric features, we significantly improve the accuracy of the coarse correspondences produced by diffusion features. Extensive experiments demonstrate that without any task-specific training, direct utilization of both features produces accurate image-to-point cloud registration. On three public indoor and outdoor benchmarks, the proposed method averagely achieves a 20.6 percent improvement in Inlier Ratio, a $3.0\times$ higher Inlier Number, and a 48.6 percent improvement in Registration Recall than existing state-of-the-arts. The code and additional results are available at \url{https://whu-usi3dv.github.io/FreeReg/}. Haiping Wang 0004, Yuan Liu 0025, Bing Wang 0013, Yujing Sun 0001, Zhen Dong 0005, Wenping Wang 0001, Bisheng Yang |
ICLR | 4 |
| 2024 | Gait Recognition in Large-scale Free Environment via Single LiDARabstractHuman gait recognition is crucial in multimedia, enabling identification through walking patterns without direct interaction, enhancing the integration across various media forms in real-world applications like smart homes, healthcare and non-intrusive security. LiDAR's ability to capture depth makes it pivotal for robotic perception and holds promise for real-world gait recognition. In this paper, based on a single LiDAR, we present the Hierarchical Multi-representation Feature Interaction Network (HMRNet) for robust gait recognition. Prevailing LiDAR-based gait datasets primarily derive from controlled settings with predefined trajectory, remaining a gap with real-world scenarios. To facilitate LiDAR-based gait recognition research, we introduce FreeGait, a comprehensive gait dataset from large-scale, unconstrained settings, enriched with multi-modal and varied 2D/3D data. Notably, our approach achieves state-of-the-art performance on prior dataset (SUSTech1K) and on FreeGait. https://4dvlab.github.io/project_page/FreeGait.html Yiming Ren 0001, Peishan Cong, Yujing Sun 0001, Jingya Wang 0001, Lan Xu 0003, Yuexin Ma |
ACM Multimedia | 4 |
| 2024 | Towards Practical Human Motion Prediction with LiDAR Point CloudsabstractHuman motion prediction is crucial for human-centric multimedia understanding and interacting. Current methods typically rely on ground truth human poses as observed input, which is not practical for real-world scenarios where only raw visual sensor data is available. To implement these methods in practice, a pre-phrase of pose estimation is essential. However, such two-stage approaches often lead to performance degradation due to the accumulation of errors. Moreover, reducing raw visual data to sparse keypoint representations significantly diminishes the density of information, resulting in the loss of fine-grained features. In this paper, we propose LiDAR-HMP, the first single-LiDAR-based 3D human motion prediction approach, which receives the raw LiDAR point cloud as input and forecasts future 3D human poses directly. Building upon our novel structure-aware body feature descriptor, LiDAR-HMP adaptively maps the observed motion manifold to future poses and effectively models the spatial-temporal correlations of human motions for further refinement of prediction results. Extensive experiments show that our method achieves state-of-the-art performance on two public benchmarks and demonstrates remarkable robustness and efficacy in real-world deployments. https://4dvlab.github.io/project_page/LiDARHMP.html Yiming Ren 0001, Yichen Yao 0001, Yujing Sun 0001, Yuexin Ma |
ACM Multimedia | 4 |
| 2023 | BitAnalysis: A Visualization System for Bitcoin Wallet InvestigationabstractBitcoin is gaining ever increasing popularity. However, professional skills are required if people want to check bitcoin transaction information from the blockchain. As pointed out in a recent study, there is a lack of tools to support effective interactive investigation of bitcoin transactions. Therefore, we present a novel visualization system,BitAnalysis, for interactive bitcoin wallet investigation. The analytical and visualization functions ofBitAnalysisare defined and developed by following the advice and requirements of a group of entrepreneurs and regulators of bitcoin-related business.BitAnalysisprovides a rich set of functions and intuitive visual interfaces for the users, such as law-enforcement officers and regulators, to effectively visualize and analyze the transactions of a bitcoin wallet (i.e., a cluster of bitcoin addresses) and its related wallets, to track the flow of bitcoins, and to identify wallet correlation using our novel clustering functions. To achieve these functions, we have designed new visualization techniques for presenting bitcoin transactions information and introduced theconnection diagramandbitcoin flow mapas new ways of analyzing, tracking and monitoring the trading activities of a cluster of closely related wallets. We also present an extensive user study that validated the effectiveness and usability ofBitAnalysis. Yujing Sun 0001, Hao Xiong 0002, Siu-Ming Yiu, Kwok-Yan Lam |
IEEE Trans. Big Data | 1 |
| 2021 | Understanding deep face anti-spoofing: from the perspective of data
Yujing Sun 0001, Hao Xiong 0002, Siu-Ming Yiu |
Vis. Comput. | 1 |
| 2018 | Moiré Photo Restoration Using Multiresolution Convolutional Neural NetworksabstractDigital cameras and mobile phones enable us to conveniently record precious moments. While digital image quality is constantly being improved, taking high-quality photos of digital screens still remains challenging because the photos are often contaminated with moiré patterns, a result of the interference between the pixel grids of the camera sensor and the device screen. Moiré patterns can severely damage the visual quality of photos. However, few studies have aimed to solve this problem. In this paper, we introduce a novel multiresolution fully convolutional network for automatically removing moiré patterns from photos. Since a moiré pattern spans over a wide range of frequencies, our proposed network performs a nonlinear multiresolution analysis of the input image before computing how to cancel moiré artefacts within every frequency band. We also create a large-scale benchmark dataset with 100,000+ image pairs for investigating and evaluating moiré pattern removal algorithms. Our network achieves state-of-the-art performance on this dataset in comparison to existing learning architectures for image restoration problems. Yujing Sun 0001, Yizhou Yu, Wenping Wang 0001 |
IEEE Trans. Image Process. | 1 |
| 2018 | Image Structure Retrieval via L0 MinimizationabstractRetrieving salient structure from textured images is an important but difficult problem in computer vision because texture, which can be irregular, anisotropic, non-uniform and complex, shares many of the same properties as structure. Observing that salient structure in a textured image should be piece-wise smooth, we present a method to retrieve such structures using an minimization of a modified form of the relative total variation metric. Thanks to the characteristics shared by texture and small structures, our method is effective at retrieving structure based on scale as well. Our method outperforms state-of-art methods in texture removal as well as scale-space filtering. We also demonstrate our method's ability in other applications such as edge detection, clip art compression artifact removal, and inverse half-toning. Yujing Sun 0001, Scott Schaefer, Wenping Wang 0001 |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2015 | Denoising point sets via L0 minimization
Yujing Sun 0001, Scott Schaefer, Wenping Wang 0001 |
Comput. Aided Geom. Des. | 1 |