Yu He 0001

dblp:16/3418-1 · DBLP profile ↗
← Back
10ranked-venue papers
3as first author
10since 2021 · last 2026
0000-0002-0357-681XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 BTKD++: Beyond Teachers by Critically Distilling Knowledge from Teacher's Bias
abstract
Abstract Existing knowledge distillation methods indiscriminately transfer knowledge from teacher networks, including output-level decisional biases, i.e., incorrect final predictions that can mislead student learning and limit student performance. We challenge this paradigm by proposing BTKD++, a framework that systematically filters and rectifies teacher’s output-level biased knowledge into corrective signals. Our approach partitions training data into Easy Tasks (correct teacher predictions) and Hard Tasks (incorrect predictions), then applies bias elimination and rectification modules orchestrated by dynamic learning curriculum. We provide an interpretive information-theoretic abstraction to explain the observed competence-threshold phenomenon, under which bias rectification becomes more effective when teacher errors contain sufficiently structured corrective information. BTKD++ demonstrates broad applicability across classification, detection, and segmentation tasks when task outputs are equipped with suitable probabilistic interfaces, and shows consistent effectiveness across CNNs, Transformers, and State-Space Models. Extensive experiments show consistent student-teacher transcendence, establishing new state-of-the-art results. This work redefines knowledge distillation from blind mimicry to critical learning, proving that students can surpass teachers through principled bias correction. The source code is available at https://github.com/smartyige/BTKD .
Jianhua Zhang 0002, Yu He 0001, Xu Cheng 0003, Xiufeng Liu 0001, Shengyong Chen, Houxiang Zhang, Ruyu Liu
Int. J. Comput. Vis.3
2025 Strategies for reducing motion sickness in virtual reality through improved handheld controller movements
abstract
As technology advances, user demand for immersive and authentic information presentation rises. Traditional 2D displays and interactions fail to meet modern standards, while virtual reality (VR) is gaining attention for its immersive experience. However, using a controller for VR movement can cause dizziness due to mismatched visual and vestibular cues, impacting the VR experience. This paper analyzes the main causes of VR-induced vertigo and develops improved handheld controller movement strategies. These strategies adjust the user’s pitch angle and field of view in real time or map the user’s real-world head acceleration to the virtual character. By intelligently adjusting the controller-to-VR display mapping, these methods reduce vertigo. In addition, this paper also verified the actual effects of these designs through a series of experiments, and conducted detailed data analysis on the degree of user vertigo. The experimental results showed that using a specific improved handheld controller movement design can significantly improve the user’s comfort in the VR environment, effectively reducing the occurrence of vertigo and discomfort.
Khang Yeu Tang, Juhong Wang, Yu He 0001, Sen-Zhe Xu 0001, Song-Hai Zhang
Graph. Model.4
2024 AdaPIP: Adaptive picture-in-picture guidance for 360° film watching
abstract
360° videos enable viewers to watch freely from different directions but inevitably prevent them from perceiving all the helpful information. To mitigate this problem, picture-in-picture (PIP) guidance was proposed using preview windows to show regions of interest (ROIs) outside the current view range. We identify several drawbacks of this representation and propose a new method for 360° film watching called AdaPIP. AdaPIP enhances traditional PIP by adaptively arranging preview windows with changeable view ranges and sizes. In addition, AdaPIP incorporates the advantage of arrow-based guidance by presenting circular windows with arrows attached to them to help users locate the corresponding ROIs more efficiently. We also adapted AdaPIP and Outside-In to HMD-based immersive virtual reality environments to demonstrate the usability of PIP-guided approaches beyond 2D screens. Comprehensive user experiments on 2D screens, as well as in VR environments, indicate that AdaPIP is superior to alternative methods in terms of visual experiences while maintaining a comparable degree of immersion.
Yi-Xiao Li, Guan Luo, Yi-Ke Xu, Yu He 0001, Song-Hai Zhang
Comput. Vis. Media4
2023 Color-Correlated Texture Synthesis for Hybrid Indoor Scenes
Yu He 0001, Yi-Han Jin, Ying-Tian Liu, Baoli Lu
CAD/Graphics1
2023 Enhancing Ocean Scene Video Captioning with Multimodal Pre-Training and Video-Swin-Transformer
abstract
With the success of multimodal pre-training models in the video-language field and various downstream tasks, previous multimodal models used 3DCNN networks as video feature extractors, which have limitations in interacting and fusing with text features. This paper proposes a multimodal pre-training model that utilizes a Video-Swin-Transformer-based network to encode both video and text data, to achieve better performance in video understanding. The model consists of four modules: video encoder, text encoder, interact encoder, and caption decoder to accomplish the task of ocean scene video captioning. A dataset of ocean scene videos, including various content types such as sea surfaces and shores, is also constructed. The training process is divided into two stages: pre-training and fine-tuning. Pre-training is performed on the Howto100m dataset to allow the model to learn video captions in natural scenes and complete video-language matching tasks. The fine-tuning stage is then performed on the ocean1000 dataset to better understand the events and content in ocean scene videos and generate captions that conform to ocean scene video descriptions. The model achieves satisfying results on both the public dataset YouCook2 and the proprietary dataset Ocean1000, demonstrating its ability in video-text information fusion and interaction.
Meng Zhao 0001, Fan Shi 0001, Meng'en Zhang, Yu He 0001, Shengyong Chen
IECON5
2023 VMesh: Hybrid Volume-Mesh Representation for Efficient View Synthesis
abstract
With the emergence of neural radiance fields (NeRFs), view synthesis quality has reached an unprecedented level. Compared to traditional mesh-based assets, this volumetric representation is more powerful in expressing scene geometry but inevitably suffers from high rendering costs and can hardly be involved in further processes like editing, posing significant difficulties in combination with the existing graphics pipeline. In this paper, we present a hybrid volume-mesh representation, VMesh, which depicts an object with a textured mesh along with an auxiliary sparse volume. VMesh retains the advantages of mesh-based assets, such as efficient rendering and compact storage, while also incorporating the ability to represent subtle geometric structures provided by the volumetric counterpart. VMesh can be obtained from multi-view images of an object and renders at 2K 60FPS on common consumer devices with high fidelity, unleashing new opportunities for real-time immersive applications.
Yan-Pei Cao 0001, Chen Wang 0049, Yu He 0001, Ying Shan, Song-Hai Zhang
SIGGRAPH Asia4
2022 NeRFReN: Neural Radiance Fields with Reflections
abstract
Neural Radiance Fields (NeRF) has achieved unprece-dented view synthesis quality using coordinate-based neu-ral scene representations. However, NeRF's view depen-dency can only handle simple reflections like highlights but cannot deal with complex reflections such as those from glass and mirrors. In these scenarios, NeRF models the virtual image as real geometries which leads to inaccurate depth estimation, and produces blurry renderings when the multi-view consistency is violated as the reflected objects may only be seen under some of the viewpoints. To over-come these issues, we introduce NeRFReN, which is built upon NeRF to model scenes with reflections. Specifically, we propose to split a scene into transmitted and reflected components, and model the two components with separate neural radiance fields. Considering that this decomposition is highly under-constrained, we exploit geometric priors and apply carefully-designed training strategies to achieve reasonable decomposition results. Experiments on various self-captured scenes show that our method achieves high-quality novel view synthesis and physically sound depth es-timation results while enabling scene editing applications.
Linchao Bao, Yu He 0001, Song-Hai Zhang
CVPR4
2022 Smoothness preserving layout for dynamic labels by hybrid optimization
abstract
Stable label movement and smooth label trajectory are critical for effective information understanding. Sudden label changes cannot be avoided by whatever forced directed methods due to the unreliability of resultant force or global optimization methods due to the complex trade-off on the different aspects. To solve this problem, we proposed a hybrid optimization method by taking advantages of the merits of both approaches. We first detect the spatial-temporal intersection regions from whole trajectories of the features, and initialize the layout by optimization in decreasing order by the number of the involved features. The label movements between the spatial-temporal intersection regions are determined by force directed methods. To cope with some features with high speed relative to neighbors, we introduced a force from future, called temporal force, so that the labels of related features can elude ahead of time and retain smooth movements. We also proposed a strategy by optimizing the label layout to predict the trajectories of features so that such global optimization method can be applied to streaming data. ELECTRONIC SUPPLEMENTARY MATERIAL: Supplementary material is available in the online version of this article at 10.1007/s41095-021-0231-y.
Yu He 0001, Song-Hai Zhang
Comput. Vis. Media1
2022 Context-Consistent Generation of Indoor Virtual Environments Based on Geometry Constraints
abstract
In this article, we propose a system that can automatically generate immersive and interactive virtual reality (VR) scenes by taking real-world geometric constraints into account. Our system can not only help users avoid real-world obstacles in virtual reality experiences, but also provide context-consistent contents to preserve their sense of presence. To do so, our system first identifies the positions and bounding boxes of scene objects as well as a set of interactive planes from 3D scans. Then context-consistent virtual objects that have similar geometric properties to the real ones can be automatically selected and placed into the virtual scene, based on learned object association relations and layout patterns from large amounts of indoor scene configurations. We regard virtual object replacement as a combinatorial optimization problem, considering both geometric and contextual consistency constraints. Quantitative and qualitative results show that our system can generate plausible interactive virtual scenes that highly resemble real environments, and have the ability to keep the sense of presence for users in their VR experiences.
Yu He 0001, Ying-Tian Liu, Yi-Han Jin, Song-Hai Zhang, Yukun Lai, Shi-Min Hu 0001
IEEE Trans. Vis. Comput. Graph.1
2021 MageAdd: Real-Time Interaction Simulation for Scene Synthesis
abstract
While recent researches on computational 3D scene synthesis have achieved impressive results, automatically synthesized scenes do not guarantee satisfaction of end users. On the other hand, manual scene modelling can always ensure high quality, but requires a cumbersome trial-and-error process. In this paper, we bridge the above gap by presenting a data-driven 3D scene synthesis framework that can intelligently infer objects to the scene by incorporating and simulating user preferences with minimum input. While the cursor is moved and clicked in the scene, our framework automatically selects and transforms suitable objects into scenes in real time. This is based on priors learnt from the dataset for placing different types of objects, and updated according to the current scene context. Through extensive experiments we demonstrate that our framework outperforms the state-of-the-art on result aesthetics, and enables effective and efficient user interactions.
Shao-Kui Zhang, Yi-Xiao Li, Yu He 0001, Yongliang Yang 0002, Song-Hai Zhang
ACM Multimedia3