VLDB 2026 Research / reviewers in the wild / expert
Xiaokun Sun
dblp:59/10595
· DBLP profile ↗
11ranked-venue papers
5as first author
10since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 3 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MVP: Enhancing Video Large Language Models via Self-supervised Masked Video PredictionabstractReinforcement learning based post-training paradigms for Video Large Language Models (VideoLLMs) have achieved significant success by optimizing for visual-semantic tasks such as captioning or VideoQA.However, while these approaches effectively enhance perception abilities, they primarily target holistic content understanding, often lacking explicit supervision for intrinsic temporal coherence and inter-frame correlations.This tendency limits the models' ability to capture intricate dynamics and fine-grained visual causality.To explicitly bridge this gap, we propose a novel posttraining objective: Masked Video Prediction (MVP).By requiring the model to reconstruct a masked continuous segment from a set of challenging distractors, MVP forces the model to attend to the sequential logic and temporal context of events.To support scalable training, we introduce a scalable data synthesis pipeline capable of transforming arbitrary video corpora into MVP training samples, and further employ Group Relative Policy Optimization (GRPO) with a fine-grained reward function to enhance the model's understanding of video context and temporal properties.Comprehensive evaluations demonstrate that MVP enhances video reasoning capabilities by directly reinforcing temporal reasoning and causal understanding. Xiaokun Sun, Zezhong Wu, Zewen Ding |
ACL (1) | 1 |
| 2026 | DreamBarbie: Text to Barbie-Style 3D AvatarsabstractTo integrate digital humans into everyday life, there is a strong demand for generating high-quality, fine-grained disentangled 3D avatars that support expressive animation and simulation capabilities, ideally from low-cost textual inputs. Although text-driven 3D avatar generation has made significant progress by leveraging 2D generative priors, existing methods still struggle to fulfill all these requirements simultaneously. To address this challenge, we propose DreamBarbie, a novel text-driven framework for generating animatable 3D avatars with separable shoes, accessories, and simulation-ready garments, truly capturing the iconic "Barbie doll" aesthetic. The core of our framework lies in an expressive 3D representation combined with appropriate modeling constraints. Unlike prior methods, we use G-Shell to uniformly model watertight components (e.g., bodies, shoes) and non-watertight garments. By reformulating boundaries as euclidean field intersections instead of manifold geodesics, we propose an SDF-based initialization and a hole regularization loss that together achieve a $100\times$100× speedup and stable open topology without image input. These disentangled 3D representations are then optimized by specialized expert diffusion models tailored to each domain, ensuring high-fidelity outputs. To mitigate geometric artifacts and texture conflicts when combining different expert models, we further propose several effective geometric losses and strategies. Extensive experiments demonstrate that DreamBarbie outperforms existing methods in both dressed human and outfit generation. Our framework further enables diverse applications, including apparel combination, editing, expressive animation, and physical simulation. Xiaokun Sun, Zhenyu Zhang 0005, Ying Tai, Hao Tang 0005, Zili Yi, Jian Yang 0003 |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2025 | StrandHead: Text to Hair-Disentangled 3D Head Avatars Using Human-Centric Priors
Xiaokun Sun, Ying Tai, Jian Yang 0003, Zhenyu Zhang 0005 |
ICCV | 1 |
| 2024 | Statistical Characterization of Polarimetric Time-Frequency Coherent IndicatorabstractThe polarimetric time-frequency coherent indicator has been proposed and effectively applied, especially for target detection. In this paper, based on the complex Wishart distribution, an associated asymptotic probability for the statistical test of the polarimetric time-frequency coherence is analytically derived. The goodness of fit for the probability density function is tested with Monte Carlo simulated data. It can be observed that the estimated probability density function curve fit well with the theoretical value. Canbin Hu, Hongyun Chen, Xiaokun Sun, Zekai Yun, Laurent Ferro-Famil |
IGARSS | 3 |
| 2024 | SPLM-Net:Large Scene SAR Image Registration Based on Point and Line Matching NetworkabstractDue to the complex distribution of point features in large-scene synthetic aperture radar (SAR) images, it is challenging to achieve subsequent accurate and robust registration of the images. This paper proposes a SAR point-and-line matching network (SPLM-Net) based on joint matching of sparse key points and line segments. SPLM-Net uses Graph Neural Networks (GNNs) processing to unify points, lines, and their descriptors into a wireframe structure to provide a comprehensive and local discriminant description of the image, and uses dual-softmax mechanism to determine the final feature assignment. SPLM-Net makes up for the shortcoming of high error in point feature matching of large-scale scene images. Experimental results show that the results obtained using this algorithm have significantly improved visual effects and evaluation metrics. Xiaokun Sun, Zekai Yun, Canbin Hu, Hongyun Chen |
IGARSS | 1 |
| 2024 | PolSAR Image Registration Combining Siamese Multiscale Attention Network and Joint FilterabstractPolarimetric synthetic aperture radar (PolSAR) is an active microwave imaging system. Due to the coherence characteristic of PolSAR imaging, inherent coherent speckle noise exists in PolSAR images. The registration of PolSAR images is severely affected by speckle noise. Therefore, we first propose a joint filter that combines Refined-Lee filtering and polarimetric whitening filtering (PWF). The filter first applies Refined-Lee filtering to PolSAR images, which greatly reduces the speckle noise while maintaining high-resolution detailed information of the texture, and then uses PWF to normalize and whiten the polarimetric matrix to further limit the interference of speckle noise. After that, the binary robust invariant scalable keypoints (BRISK) algorithm is used to extract high-quality keypoints from the denoised PolSAR image. Then a novel Siamese Multiscale Attention Network (SMAN) is designed, which uses attention modules to construct feature descriptors with different scales. To fully utilize polarimetric information, we adopt the polarimetric covariance matrix and three polarimetric features as inputs to the network and the Second Order Similarity (SOS) as the loss function to train the network. In the keypoint matching stage, we present to use the symmetric displacement distance to further constrain the keypoint pairs obtained by the initial matching, which improves the accuracy of matching keypoint pairs. Experimental results show that our proposed method can effectively reduce the interference of speckle noise and overcome non-linear differences, geometric distortions, and differences in polarimetric scattering information to achieve accurate PolSAR image registration. Deliang Xiang, Huaiyue Ding, Xiaokun Sun, Jianda Cheng, Canbin Hu, Yi Su 0003 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Sidelobe Suppression for High-Resolution SAR Imagery Based on Spectral Reshaping and Feature Statistical DifferenceabstractSidelobe suppression is a crucial preprocessing technique for Synthetic Aperture Radar (SAR) image analysis. The presence of strong scattering targets generates sidelobes that can interfere the targets with relatively weak scattering. The overlapping of multiple strong cross-shaped sidelobes may even generate fake targets, significantly influencing the accuracy of SAR target detection and recognition. Among existing sidelobe suppression methods, the Spectral Reshaping Sidelobe Reduction (SRSR) method has shown promising results. It separates the mainlobe and sidelobes through altering the sidelobe direction while preserving the SAR image resolution. However, this method exhibits limitations in effectively suppressing strong cross-shaped sidelobes. It also introduces additional sidelobes, blurring the surroundings of the scattering points. This paper proposes an improved SRSR method to resolve this disadvantage. It constructs a feature image representing the superposition of sidelobes. This is achieved by analyzing the statistical differences of complex data between orthogonal and non-orthogonal sidelobe regions before and after spectral reshaping. Further modulus selection ensures that the feature image only contains the sidelobe information that needs to be eliminated. The proposed method successfully resolves the drawback of introducing new sidelobes in the original SRSR while achieving better suppression of strong cross-shaped sidelobes. Experimental results on airborne and spaceborne SAR images demonstrate that the proposed method outperforms other state-of-the-art techniques. Improved peak sidelobe ratio (PSLR) and integrated sidelobe ratio (ISLR) in both range and azimuth directions and smaller image entropy can be achieved by our method. Due to its superior sidelobe suppression capability, the SAR images processed by our method exhibit significantly improved accuracy in target detection. Deliang Xiang, Wenhang Li, Xiaokun Sun, Huaijun Wang, Yi Su 0003 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Two-Stage Registration of SAR Images With Large Distortion Based on Superpixel SegmentationabstractWhen the geometric distortion of the SAR images to be registered is large, the spatial correspondence between the feature points of the two images will change significantly. Hence, the registration of SAR images with large geometric distortion is challenging. To solve this problem, a two-stage registration method of SAR images with large distortion based on superpixel segmentation is proposed in this paper. Firstly, the two SAR images are coarsely registered by geographic coordinate referencing. After coarse registration, superpixel segmentation is performed on the two SAR images respectively. Next, in the superpixel neighborhood of the reference image, we slide the corresponding superpixel template of the sensed image, finding its position with the highest similarity in the reference image. Compared with the traditional fixed-size template, the superpixel template can segment the distorted region more effectively. Meanwhile, with the help of the adaptive threshold detector proposed in this paper, the regions with varying degrees of distortion can be distinguished based on the similarity. Further, the geometry mapping relationship is calculated for the regions with different distortion degrees in the images respectively, and the corresponding feature points in different images are accurately matched to complete the fine registration. Finally, the registration results of regions with different degrees of distortion are fused to obtain the final SAR image registration results. Experimental results based on Sentinel-1 data show that the registration accuracy of the proposed method can reach within 1 pixel. Deliang Xiang, Huaiyue Ding, Jianda Cheng, Xiaokun Sun |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | Learning Semantic-Aware Disentangled Representation for Flexible 3D Human Body Editingabstract3D human body representation learning has received increasing attention in recent years. However, existing works cannot flexibly, controllably and accurately represent human bodies, limited by coarse semantics and unsatisfactory representation capability, particularly in the absence of supervised data. In this paper, we propose a human body representation with fine-grained semantics and high reconstruction-accuracy in an unsupervised setting. Specifically, we establish a correspondence between latent vectors and geometric measures of body parts by designing a part-aware skeleton-separated decoupling strategy, which facilitates controllable editing of human bodies by modifying the corresponding latent codes. With the help of a bone-guided auto-encoder and an orientation-adaptive weighting strategy, our representation can be trained in an unsupervised manner. With the geometrically meaningful latent space, it can be applied to a wide range of applications, from human body editing to latent code interpolation and shape style transfer. Experimental results on public datasets demonstrate the accurate reconstruction and flexible editing abilities of the proposed method. The code will be available at http://cic.tju.edu.cn/faculty/likun/projects/SemanticHuman. Xiaokun Sun, Qiao Feng 0001, Xiongzheng Li, Yukun Lai, Jing-Yu Yang 0002, Kun Li 0001 |
CVPR | 1 |
| 2023 | Learning to Infer Inner-Body Under Clothing From Monocular VideoabstractAccurately estimating the human inner-body under clothing is very important for body measurement, virtual try-on and VR/AR applications. In this article, we propose the first method to allow everyone to easily reconstruct their own 3D inner-body under daily clothing from a self-captured video with the mean reconstruction error of 0.73cm within 15s. This avoids privacy concerns arising from nudity or minimal clothing. Specifically, we propose a novel two-stage framework with a Semantic-guided Undressing Network (SUNet) and an Intra-Inter Transformer Network (IITNet). SUNet learns semantically related body features to alleviate the complexity and uncertainty of directly estimating 3D inner-bodies under clothing. IITNet reconstructs the 3D inner-body model by making full use of intra-frame and inter-frame information, which addresses the misalignment of inconsistent poses in different frames. Experimental results on both public datasets and our collected dataset demonstrate the effectiveness of the proposed method. The code and dataset is available for research purposes at http://cic.tju.edu.cn/faculty/likun/projects/Inner-Body. Xiongzheng Li, Xiaokun Sun, Haibiao Xuan, Yukun Lai, Yingdi Xie, Jing-Yu Yang 0002, Kun Li 0001 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2019 | Moving Target Velocity Estimation Using Multi-Azimuth Angle ModeabstractBased on the multiple azimuth squint angles mode, a novel moving target velocity estimation method is proposed, including both along range and along azimuth directions. The acquisition geometry of multi-azimuth angle mode is given first with variation of antenna azimuth squint angles. Then, range curvature is corrected in range-frequency domain after range compression and the residual range curvature can be neglected which is smaller than a range bin. Furthermore, azimuth velocity is estimated by the slope variation of the range walk trajectory from two certain observations and then is the range velocity estimation based on the slope from one observation. Velocity estimation accuracy is also analyzed, considering the errors caused by azimuth squint angle and trajectory slope. The effectiveness of the proposed strategy is demonstrated by experimental results. Jie Chen 0009, Wei Yang 0004, Zhirong Men, Rui Zhang 0100, Xiaokun Sun |
IGARSS | 6 |