Shikun Sun

dblp:293/2733 · DBLP profile ↗
← Back
16ranked-venue papers
5as first author
16since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 3 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 TEENet: An Effective Clinical Detection Network for Identifying Spontaneous Echo Contrast Automatically
abstract
Spontaneous Echo Contrast (SEC) is a swirling smoke-like echo phenomenon in Transesophageal Echocardiography (TEE) videos caused by slow blood flow and hypercoagulable states. It is a significant indicator for assessing thromboembolic risk. However, current SEC identification requires extensive manual intervention, leading to low accuracy, high costs, and subjectivity. To address these issues, we propose TEENet, an effective clinical detection network for identifying SEC in TEE videos. Specifically, TEENet first generates attention maps for the input clips to highlight important regions and integrates Convolutional Neural Network with the Multi-Head Self-Attention to capture spatiotemporal representations. Furthermore, to enhance the classification performance across different SEC severity grades, we introduce an auxiliary classification module, which simultaneously utilizes the main classification head and auxiliary classification heads. Notably, we constructed a comprehensive dataset of 1106 TEE videos collected during clinical examinations performed at the First Affiliated Hospital of Soochow University from 2018 to 2023, providing a solid foundation for the development and validation of TEENet. Extensive experimental results demonstrate that our proposed network achieves the highest SEC identification accuracy of 92.4$\pm$1.3% compared to other spatiotemporal representation networks such as SlowFastR50 (89.6$\pm$0.7%) and TimeSformer (74.9$\pm$1.8%), which shows strong potential for effective auxiliary diagnosis in clinical practice.
Zhiwen Wu, Fei Gu 0001, Shikun Sun, Changsheng Ma
IEEE J. Biomed. Health Informatics4
2025 Entropy-Adaptive Diffusion Policy Optimization with Dynamic Step Alignment
Renye Yan, Jikang Cheng, Yaozhong Gan, Shikun Sun, Yunfan Yang, Ling Liang 0003, Jinlong Lin, Yeshuang Zhu, Jie Zhou 0001, Junliang Xing, Yimao Cai, Ru Huang 0001
ICCV4
2025 Minimal Impact ControlNet: Advancing Multi-ControlNet Integration
abstract
With the advancement of diffusion models, there is a growing demand for high-quality, controllable image generation, particularly through methods that utilize one or multiple control signals based on ControlNet. However, in current ControlNet training, each control is designed to influence all areas of an image, which can lead to conflicts when different control signals are expected to manage different parts of the image in practical applications. This issue is especially pronounced with edge-type control conditions, where regions lacking boundary information often represent low-frequency signals, referred to as silent control signals. When combining multiple ControlNets, these silent control signals can suppress the generation of textures in related areas, resulting in suboptimal outcomes. To address this problem, we propose Minimal Impact ControlNet. Our approach mitigates conflicts through three key strategies: constructing a balanced dataset, combining and injecting feature signals in a balanced manner, and addressing the asymmetry in the score function’s Jacobian matrix induced by ControlNet. These improvements enhance the compatibility of control signals, allowing for freer and more harmonious generation in areas with silent control signals.
Shikun Sun, Zixuan Wang 0026, Xubin Li, Tiezheng Ge, Zijie Ye, Xiaoyu Qin 0001, Junliang Xing, Bo Zheng 0007, Jia Jia 0001
ICLR1
2025 V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos
abstract
Automatic video commentary systems are widely used on multimedia social media platforms to extract factual information about video content. However, current systems may overlook essential paralinguistic cues, including emotion and attitude, which are critical for fully conveying the meaning of visual content. The absence of these cues can limit user understanding or, in some cases, distort the video’s original intent. Expressive speech effectively conveys these cues and enhances the user’s comprehension of videos. Building on these insights, this paper explores the usage of vision-context-aware expressive speech in enhancing users’ understanding of videos in video commentary systems1. Firstly, our formatting study indicates that semantic-only speech can lead to ambiguity, and misaligned emotions between speech and visuals may distort content interpretation. To address this, we propose a method called vision-context-aware speech synthesis (V-CASS). It analyzes para-linguistic cues from visuals using a vision-language model and leverages a knowledge-infused language model to guide the expressive speech model in generating context-aligned speech. User studies show that V-CASS enhances emotional and attitudinal resonance, as well as user audio-visual understanding and engagement, with 74.68% of participants preferring the system. Finally, we explore the potential of our method in helping blind and low-vision users navigate web videos, improving universal accessibility.
Qixin Wang 0002, Songtao Zhou, Zeyu Jin, Chenglin Guo, Shikun Sun, Xiaoyu Qin 0001
IJCNN5
2025 DEPO: Enhancing E-commerce Image Background Generation with Short Trajectory Direct Expected Preference Optimization
Shikun Sun, Zixuan Wang 0026, Xiaoyu Qin 0001, Tiezheng Ge, Bo Zheng 0007, Jia Jia 0001
ACM Multimedia1
2025 Dual-Flow: Transferable Multi-Target, Instance-Agnostic Attacks via In-the-wild Cascading Flow Optimization
Shikun Sun, Jianshu Li, Junliang Xing
NeurIPS2
2025 PVCsNet : A Specialized Artificial Intelligence-Based Model to Classify Premature Ventricular Contractions From ECG Images
abstract
Premature ventricular complexes (PVCs) are irregularities in heart rhythm where the ventricles contract earlier than expected, disrupting the normal cardiac cycle. Identifying the origin of PVCs before surgery is crucial as it can reduce operation duration, lower radiation exposure, and potentially enhance ablation success rates. Current detection methods face limitations in accuracy and data processing, often requiring large datasets and complex interpretations. This study presents PVCsNet, a deep-learning network specifically designed for classifying premature ventricular complexes (PVCs) in ECG images. It incorporates residual structures and attention mechanisms to enhance classification performance. PVCsNet consists of four 3 × 3 convolutional layers as feature extractors, followed by residual connections and attention blocks. This design enables the network to map image features to class probability distributions, enhancing performance even with limited data. Our experimental results demonstrate that using the SE Block with MaxPool and a ratio of 4, PVCsNet achieves an overall accuracy of 94.49%, with high precision in critical categories and a moderate parameter size. We successfully categorize the data into six distinct classes based on their origin locations in the heart: right ventricular outflow tract (RVOT), left ventricular outflow tract (LVOT), papillary muscle (PM), valvular annulus (VA), summit, and His-Purkinje system (HPS). Among these, RVOT is the most common and crucial origin of PVCs. PM and HPS are also significant origins due to their clinical implications. This study demonstrates the potential of PVCsNet in clinical diagnostics, providing promising results in classifying ECG images and contributing to future medical research and diagnosis.
Biren Guo, Fei Gu 0001, Zeyang Zhang 0003, Shikun Sun
IEEE J. Biomed. Health Informatics5
2024 DanceCamera3D: 3D Camera Movement Synthesis with Music and Dance
abstract
Choreographers determine what the dances look like, while cameramen determine the final presentation of dances. Recently, various methods and datasets have show-cased the feasibility of dance synthesis. However, camera movement synthesis with music and dance remains an un-solved challenging problem due to the scarcity of paired data. Thus, we present DCM, a new multi-modal 3D dataset, which for the first time combines camera movement with dance motion and music audio. This dataset encom-passes 108 dance sequences (3.2 hours) of paired dance-camera-music data from the anime community, covering 4 music genres. With this dataset, we uncover that dance camera movement is multifaceted and human-centric, and possesses multiple influencing factors, making dance camera synthesis a more challenging task compared to camera or dance synthesis alone. To overcome these difficulties, we propose DanceCamera3D, a transformer-based diffusion model that incorporates a novel body attention loss and a condition separation strategy. For evaluation, we devise new metrics measuring camera movement quality, diversity, and dancer fidelity. Utilizing these metrics, we conduct extensive experiments on our DCM dataset, providing both quantitative and qualitative evidence showcasing the effectiveness of our DanceCamera3D model. Code and video demos are available at https://github.com/Carmenw1203/DanceCamera3D-Official.
Zixuan Wang 0026, Jia Jia 0001, Shikun Sun, Haozhe Wu, Jiaqing Zhou, Jiebo Luo 0001
CVPR3
2024 Inner Classifier-Free Guidance and Its Taylor Expansion for Diffusion Models
abstract
Classifier-free guidance (CFG) is a pivotal technique for balancing the diversity and fidelity of samples in conditional diffusion models. This approach involves utilizing a single model to jointly optimize the conditional score predictor and unconditional score predictor, eliminating the need for additional classifiers. It delivers impressive results and can be employed for continuous and discrete condition representations. However, when the condition is continuous, it prompts the question of whether the trade-off can be further enhanced. Our proposed inner classifier-free guidance (ICFG) provides an alternative perspective on the CFG method when the condition has a specific structure, demonstrating that CFG represents a first-order case of ICFG. Additionally, we offer a second-order implementation, highlighting that even without altering the training policy, our second-order approach can introduce new valuable information and achieve an improved balance between fidelity and diversity for Stable Diffusion.
Shikun Sun, Longhui Wei, Zhicai Wang, Zixuan Wang 0026, Junliang Xing, Jia Jia 0001, Qi Tian 0001
ICLR1
2024 PlacidDreamer: Advancing Harmony in Text-to-3D Generation
abstract
Recently, text-to-3D generation has attracted significant attention, resulting in notable performance enhancements. Previous methods utilize end-to-end 3D generation models to initialize 3D Gaussians, multi-view diffusion models to enforce multi-view consistency, and text-to-image diffusion models to refine details with score distillation algorithms. However, these methods exhibit two limitations. Firstly, they encounter conflicts in generation directions since different models aim to produce diverse 3D assets. Secondly, the issue of over-saturation in score distillation has not been thoroughly investigated and solved. To address these limitations, we propose PlacidDreamer, a text-to-3D framework that harmonizes initialization, multi-view generation, and text-conditioned generation with a single multi-view diffusion model, while simultaneously employing a novel score distillation algorithm to achieve balanced saturation. To unify the generation direction, we introduce the Latent-Plane module, a training-friendly plug-in extension that enables multi-view diffusion models to provide fast geometry reconstruction for initialization and enhanced multi-view images to personalize the text-to-image diffusion model. To address the over-saturation problem, we propose to view score distillation as a multi-objective optimization problem and introduce the Balanced Score Distillation algorithm, which offers a Pareto Optimal solution that achieves both rich details and balanced saturation. Extensive experiments validate the outstanding capabilities of our PlacidDreamer. The code is available at https://github.com/HansenHuang0823/PlacidDreamer.
Shuo Huang 0005, Shikun Sun, Zixuan Wang 0026, Xiaoyu Qin 0001, Yanmin Xiong, Yuan Zhang 0020, Pengfei Wan 0001, Di Zhang 0026, Jia Jia 0001
ACM Multimedia2
2024 DanceCamAnimator: Keyframe-Based Controllable 3D Dance Camera Synthesis
abstract
Synthesizing camera movements from music and dance is highly challenging due to the contradicting requirements and complexities of dance cinematography. Unlike human movements, which are always continuous, dance camera movements involve both continuous sequences of variable lengths and sudden drastic changes to simulate the switching of multiple cameras. However, in previous works, every camera frame is equally treated and this causes jittering and unavoidable smoothing in post-processing. To solve these problems, we propose to integrate animator dance cinematography knowledge by formulating this task as a three-stage process: keyframe detection, keyframe synthesis, and tween function prediction. Following this formulation, we design a novel end-to-end dance camera synthesis framework DanceCamAnimator, which imitates human animation procedures and shows powerful keyframe-based controllability with variable lengths. Extensive experiments on the DCM dataset demonstrate that our method surpasses previous baselines quantitatively and qualitatively. Code will be available at https://github.com/Carmenw1203/DanceCamAnimator-Official.
Zixuan Wang 0026, Xiaoyu Qin 0001, Shikun Sun, Songtao Zhou, Jia Jia 0001, Jiebo Luo 0001
ACM Multimedia4
2024 Skinned Motion Retargeting with Dense Geometric Interaction Perception
abstract
Capturing and maintaining geometric interactions among different body parts is crucial for successful motion retargeting in skinned characters. Existing approaches often overlook body geometries or add a geometry correction stage after skeletal motion retargeting. This results in conflicts between skeleton interaction and geometry correction, leading to issues such as jittery, interpenetration, and contact mismatches. To address these challenges, we introduce a new retargeting framework, MeshRet, which directly models the dense geometric interactions in motion retargeting. Initially, we establish dense mesh correspondences between characters using semantically consistent sensors (SCS), effective across diverse mesh topologies. Subsequently, we develop a novel spatio-temporal representation called the dense mesh interaction (DMI) field. This field, a collection of interacting SCS feature vectors, skillfully captures both contact and non-contact interactions between body geometries. By aligning the DMI field during retargeting, MeshRet not only preserves motion semantics but also prevents self-interpenetration and ensures contact preservation. Extensive experiments on the public Mixamo dataset and our newly-collected ScanRet dataset demonstrate that MeshRet achieves state-of-the-art performance. Code available at https://github.com/abcyzj/MeshRet.
Zijie Ye, Jia-Wei Liu, Shikun Sun, Zheng Shou 0001
NeurIPS4
2023 MSNet: A Deep Architecture Using Multi-Sentiment Semantics for Sentiment-Aware Image Style Transfer
abstract
Sentiment plays an essential role in people’s perception of images. To incorporate the sentiment information into the image style transfer task for better sentiment-aware performance, we introduce a new task named sentiment-aware image style transfer. To solve this problem, we first introduce a novel Multi-Sentiment Semantics Space (MSS-Space) to capture the non-deterministic and complicated nature of sentiment semantics. With the MSS-Space, we establish tight associations between the visual attributes of images and the multi-sentiment semantics by minimizing their distance in MSS-Space and then propose the Multi-Sentiment Style Transfer Net (MSNet). Experiments demonstrate that, compared with three competing models, our proposed MSNet generates more explicit images and better preserves the integrity of salient objects, local details, and multi-sentiment. In particular, our model outperforms the state-of-the-art by +28.72% in terms of the top-3 accuracy on average.
Shikun Sun, Jia Jia 0001, Haozhe Wu, Zijie Ye, Junliang Xing
ICASSP1
2023 Salient Co-Speech Gesture Synthesizing with Discrete Motion Representation
abstract
Synthesizing co-speech gestures is challenging because the mapping from speech to gesticulation is inherently non-deterministic. When giving talks, people conduct not only gentle and rhythmic motions but also abrupt and salient gesticulations. Most previous research efforts, however, ignore this nature of co-speech gestures and synthesize deterministic results, producing over-smoothed movements with limited expressiveness. To address this issue, we propose a new co-speech gesture generation approach that produces high-quality salient gesticulations. Specifically, we build a discrete motion representation (DMR) space to bridge the speech-gesture mapping and the gesture generation stages. The incorporation of DMR enables random sampling in motion space and avoids the over-smooth problem in speech-gesture mapping. Based on DMR, we devise a novel multi-modal co-speech gesture synthesis model with temporal attention (MCGT). MCGT explicitly models DMR’s categorical distribution conditioned on the speech context, which captures complex context patterns and produces more salient gesticulations in sync with the context. In addition, we construct a new benchmark for evaluating salient motion quality in co-speech gestures, containing a large-scale co-speech gesture dataset with salient gesticulations. We also introduce a new metric, referred to as salient motion similarity, to evaluate the salient motion quality. Experiments demonstrate superior results from our approach over several competing baselines.
Zijie Ye, Jia Jia 0001, Haozhe Wu, Shuo Huang 0005, Shikun Sun, Junliang Xing
ICASSP5
2023 SDDM: Score-Decomposed Diffusion Models on Manifolds for Unpaired Image-to-Image Translation
abstract
Recent score-based diffusion models (SBDMs) show promising results in unpaired image-to-image translation (I2I). However, existing methods, either energy-based or statistically-based, provide no explicit form of the interfered intermediate generative distributions. This work presents a new score-decomposed diffusion model (SDDM) on manifolds to explicitly optimize the tangled distributions during image generation. SDDM derives manifolds to make the distributions of adjacent time steps separable and decompose the score function or energy guidance into an image "denoising" part and a content "refinement" part. To refine the image in the same noise level, we equalize the refinement parts of the score function and energy guidance, which permits multi-objective optimization on the manifold. We also leverage the block adaptive instance normalization module to construct manifolds with lower dimensions but still concentrated with the perturbed reference image. SDDM outperforms existing SBDM-based methods with much fewer diffusion steps on several I2I benchmarks.
Shikun Sun, Longhui Wei, Junliang Xing, Jia Jia 0001, Qi Tian 0001
ICML1
2022 AI Carpet: Automatic Generation of Aesthetic Carpet Pattern
abstract
Stylized pattern generation is challenging and has received increasing attention in recent studies. However, it requires further exploration in pattern generation that matches the given scenes. This paper proposes a demonstration that automatically generates carpet patterns with input home scenes and other user preferences. Besides meeting the individual needs of users and providing highly editable output, the critical challenge of the system is to make the output pattern coordinate with the input home scene, which distinguishes our approach from others. The carpets generated by the system also permit easy modification or extension to various interior design styles.
Xingqi Wang 0003, Zeyu Jin, Shikun Sun, Jia Jia 0001
ACM Multimedia5