Feilin Han

dblp:138/5236 · DBLP profile ↗
← Back
12ranked-venue papers
2as first author
8since 2021 · last 2026
0000-0001-7463-2252ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Software engineering, systems software and programming languages · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 U2UData+: A Scalable Swarm UAVs Autonomous Flight Dataset for Embodied Long-horizon Tasks
abstract
Swarm UAV autonomous flight for Embodied Long-Horizon (ELH) tasks is crucial for advancing the low-altitude economy. However, existing methods focus only on specific basic tasks due to dataset limitations, failing in real-world deployment for ELH tasks. ELH tasks are not mere concatenations of basic tasks, requiring handling long-term dependencies, maintaining embodied persistent states, and adapting to dynamic goal shifts. This paper presents U2UData+, the first large-scale swarm UAV autonomous flight dataset for ELH tasks and the first scalable swarm UAV data online collection and algorithm closed-loop verification platform. The dataset is captured by 15 UAVs in autonomous collaborative flights for ELH tasks, comprising 12 scenes, 720 traces, 120 hours, 600 seconds per trajectory, 4.32M LiDAR frames, and 12.96M RGB frames. This dataset also includes brightness, temperature, humidity, smoke, and airflow values covering all flight routes. The platform supports the customization of simulators, UAVs, sensors, flight algorithms, formation modes, and ELH tasks. Through a visual control window, this platform allows users to collect customized datasets through one-click deployment online and to verify algorithms by closed-loop simulation. U2UData+ also introduces an ELH task for wildlife conservation and provides comprehensive benchmarks with 9 SOTA models.
Tongtong Feng, Xin Wang 0019, Feilin Han, Wenwu Zhu 0001
AAAI3
2025 "You'll Be Alice Adventuring in Wonderland!" Processes, Challenges, and Opportunities of Creating Animated Virtual Reality Stories
abstract
Animated virtual reality (VR) stories, combining the presence of VR and the artistry of computer animation, offer a compelling way to deliver messages and evoke emotions. Motivated by the growing demand for immersive narrative experiences, more creators are creating animated VR stories. However, a holistic understanding of their creation processes and challenges involved in crafting these stories is still limited. Based on semi-structured interviews with 21 animated VR story creators, we identify ten common stages in their end-to-end creation processes, ranging from idea generation to evaluation, which form diverse workflows that are story-driven or visual-driven. Additionally, we highlight nine unique issues that arise during the creation process, such as a lack of reference material for multi-element plots, the absence of specific functionalities for story integration, and inadequate support for audience evaluation. We compare the creation of animated VR stories to general XR applications and distill several future research opportunities.
Linping Yuan, Feilin Han, Liwenhan Xie, Jian Zhao 0010, Huamin Qu
CHI2
2025 VideoDreamer: Customized Multi-Subject Text-to-Video Generation With Disen-Mix Finetuning on Language-Video Foundation Models
abstract
Customized text-to-video generation aims to generate text-guided videos with user-given subjects, which has gained increasing attention. However, existing works are primarily limited to single-subject oriented text-to-video generation, leaving the more challenging problem of customized multi-subject generation unexplored. In this paper, we fill this gap and propose a novel VideoDreamer framework, which can generate temporally consistent text-guided videos that faithfully preserve the visual features of the given multiple subjects. Specifically, VideoDreamer adopts the pretrained Stable Diffusion with temporal modules as its base video generator, taking the power of the text-to-image model to generate diversified content. The video generator is further customized for multi-subjects, which leverages the proposed Disen-Mix Finetuning and Human-in-the-Loop Re-finetuning strategy, to tackle the attribute binding problem of multi-subject generation. Additionally, we present a disentangled motion customization strategy to finetune the temporal modules so that we can generate videos with both customized subjects and motions. To evaluate the performance of customized multi-subject text-to-video generation, we introduce the MultiStudioBench benchmark. Extensive experiments demonstrate the remarkable ability of VideoDreamer to generate videos with new content such as new events and backgrounds, tailored to the customized multiple subjects.
Hong Chen 0011, Xin Wang 0019, Guanning Zeng, Yipeng Zhang 0003, Yuwei Zhou, Feilin Han, Yaofei Wu, Wenwu Zhu 0001
IEEE Trans. Multim.6
2024 The Correlation Analysis Between Cybersickness and Postural Behavior in Immersive VR Experience
abstract
Cybersickness detection is one of the primary tasks in Virtual Reality (VR) content production. The existing subjective and objective studies on cybersickness give few guiding implications to VR content creators. To do experimental verification on previous hypotheses and propose design guidelines, this paper investigates the relationship between cybersickness and postural behavior, by analyzing the surface electromyography (sEMG) signals and hand movement videos. We conducted a user study to build the sEMG-video Cybersickness Benchmark Dataset (sEMG-CBD) and employed statistical analysis to summarize the regular pattern of participants’ dizziness status under VR experiences. The results indicate that the fluctuations of cybersickness correlate positively with the extent of forearm sEMG signals and hand movements. The preliminary analysis implies the potentiality of sEMG-based cybersickness detection being used as one of the significant representations of VR viewing experience, which could contribute to VR content production.
Ying Zhong 0007, Ke-Ao Zhao, Fangming Zhao, Feilin Han
ICME6
2024 U2UData: A Large-scale Cooperative Perception Dataset for Swarm UAVs Autonomous Flight
abstract
Modern perception systems for autonomous flight are sensitive to occlusion and have limited long-range capability, which is a key bottleneck in improving low-altitude economic task performance. Recent research has shown that the UAV-to-UAV (U2U) cooperative perception system has great potential to revolutionize the autonomous flight industry. However, the lack of a large-scale dataset is hindering progress in this area. This paper presents U2UData, the first large-scale cooperative perception dataset for swarm UAVs autonomous flight. The dataset was collected by three UAVs flying autonomously in the U2USim, covering a 9 km$^2$ flight area. It comprises 315K LiDAR frames, 945K RGB and depth frames, and 2.41M annotated 3D bounding boxes for 3 classes. It also includes brightness, temperature, humidity, smoke, and airflow values covering all flight routes. U2USim is the first real-world mapping swarm UAVs simulation environment. It takes Yunnan Province as the prototype and includes 4 terrains, 7 weather conditions, and 8 sensor types. U2UData introduces two perception tasks: cooperative 3D object detection and cooperative 3D object tracking. This paper provides comprehensive benchmarks of recent cooperative perception algorithms on these tasks.
Tongtong Feng, Xin Wang 0019, Feilin Han, Wenwu Zhu 0001
ACM Multimedia3
2024 U2USim - A UAV Telepresence Simulation Platform with Multi-agent Sensing and Dynamic Environment
Feilin Han, Xin Wang 0019, Ke-Ao Zhao, Ying Zhong 0007, Ziyi Su, Tongtong Feng, Wenwu Zhu 0001
ACM Multimedia1
2024 Dance2MIDI: Dance-driven multi-instrument music generation
abstract
Dance-driven music generation aims to generate musical pieces conditioned on dance videos. Previous works focus on monophonic or raw audio generation, while the multi-instrument scenario is under-explored. The challenges associated with dance-driven multi-instrument music (MIDI) generation are twofold: (i) lack of a publicly available multi-instrument MIDI and video paired dataset and (ii) the weak correlation between music and video. To tackle these challenges, we have built the first multi-instrument MIDI and dance paired dataset (D2MIDI). Based on this dataset, we introduce a multi-instrument MIDI generation framework (Dance2MIDI) conditioned on dance video. Firstly, to capture the relationship between dance and music, we employ a graph convolutional network to encode the dance motion. This allows us to extract features related to dance movement and dance style. Secondly, to generate a harmonious rhythm, we utilize a transformer model to decode the drum track sequence, leveraging a cross-attention mechanism. Thirdly, we model the task of generating the remaining tracks based on the drum track as a sequence understanding and completion task. A BERT-like model is employed to comprehend the context of the entire music piece through self-supervised learning. We evaluate the music generated by our framework trained on the D2MIDI dataset and demonstrate that our method achieves state-of-the-art performance.
Feilin Han
Comput. Vis. Media5
2022 Evaluating the Effect of Cinematography on the Viewing Experience in Immersive Environment
abstract
Cinematic Virtual Reality (CVR) is an increasingly popular digital art production technology that could enhance the sense of presence when a viewer explores immersive environments. There are three important viewing-experience-related aspects, attention, sustainability, and guidance, which can be affected by the cinematography principles. Attention indicates whether the viewer is focusing on the storytelling-related region or not. Sustainability refers to viewers' ability to continuously watch the CVR content, and guidance affects the understanding of the narrative. In this paper, we conducted within-subject repeated-measures experiments on 22 participants in an HMD-based immersive environment, to explore the correlation between viewing experience and comprehensive factors. According to experimental results, we suggest an attention-comfort-understanding analysis paradigm for directing the CVR shot, which could help creators effectively attract viewers' attention, minimize the cybersickness, and deepen their understanding of narratives.
Feilin Han, Ying Zhong 0007, Minxi Zhou
ICME1
2018 Fine-Grained Grocery Product Recognition by One-Shot Learning
abstract
Fine-grained grocery product recognition via camera is a challenging task to identify the visually similar products with subtle differences by using single-shot training examples. To address this issue? we present a novel hybrid classification approach that combines feature-based matching and one-shot deep learning with a coarse-to-fine strategy. The candidate regions of product instances are first detected and coarsely labeled by recurring features in product images without any training. Then, attention maps are generated to guide the classifier to focus on fine discriminative details by magnifying the influences of the features in the candidate regions of interest (ROI) and suppressing the interferences of the features outside, improving the accuracy of fine-grained grocery products recognition effectively. Our framework also performs a good adaptability which allows existing classifier to be refined without retraining for new coming product classes. As an additional contribution, we collect a new grocery product database with 102 classes from 2 stores. Extensive experiments demonstrate that our approach outperforms the state-of-the-art methods.
Weidong Geng, Feilin Han, Jiangke Lin, Liuyi Zhu, Jieming Bai, Suzhen Wang 0001, Zhangjiong Lai
ACM Multimedia2
2017 Image mosaicking for oversized documents with a multi-camera rig
abstract
We provide a method to obtain high resolution images of oversized documents by mosaicking their images captured by a rig equipped with multiple consumer-level cameras. We only need to calibrate the system once, and then it gives accurate full-size document mosaicking results repeatedly and robustly afterwards. During calibration, we first determine global image registration information based on homography from camera calibration results; then we refine the alignment information locally with smoothed piecewise affine transformations based on feature matches. The resulting image transformation and warping information are stored for composition of new document images. We evaluate our method against the state-of-the-art image stitching algorithms and software to demonstrate its advantages on suppressing seams and artifacts.
Zhen Wang 0003, Feilin Han, Weidong Geng
SERA2
2017 An improved saliency detection method based on non-uniform quantification and channel-weighted color distance
Aili Han, Feilin Han, Yahui Yuan
Multim. Tools Appl.2
2016 Marker-Less 3D Human Motion Capture with Monocular Image Sequence and Height-Maps
Yu Du 0016, Yongkang Wong, Feilin Han, Yilin Gui, Zhen Wang 0003, Mohan Kankanhalli, Weidong Geng
ECCV (4)4