Yihua Cheng

dblp:226/4171 · DBLP profile ↗
← Back
34ranked-venue papers
12as first author
27since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 20 · 8 first-author · 16 since 2021Artificial intelligence and machine learning · 16 · 8 first-author · 13 since 2021Computer networks · 6 · 1 first-author · 4 since 2021Systems, architecture and hardware · 5 · 2 first-author · 4 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Force-Aware 3D Contact Modeling for Stable Grasp Generation
abstract
Contact-based grasp generation plays a crucial role in various applications. Recent methods typically focus on the geometric structure of objects, producing grasps with diverse hand poses and plausible contact points. However, these approaches often overlook the physical attributes of the grasp, specifically the contact force, leading to reduced stability of the grasp. In this paper, we focus on stable grasp generation using explicit contact force predictions. First, we define a force-aware contact representation by transforming the normal force value into discrete levels and encoding it using a one-hot vector. Next, we introduce force-aware stability constraints. We define the stability problem as an acceleration minimization task and explicitly relate stability with contact geometry by formulating the underlying physical constraints. Finally, we present a pose optimizer that systematically integrates our contact representation and stability constraints to enable stable grasp generation. We show that these constraints can help identify key contact points for stability which provide effective initialization and guidance for optimization towards a stable grasp. Experiments are carried out on two public benchmarks, showing that our method brings about 20% improvement in stability metrics and adapts well to novel objects.
Zhuo Chen 0028, Zhongqun Zhang, Yihua Cheng, Ales Leonardis, Hyung Jin Chang
AAAI3
2026 RTGaze: Real-Time 3D-Aware Gaze Redirection from a Single Image
abstract
Gaze redirection methods aim to generate realistic human face images with controllable eye movement. However, recent methods often struggle with 3D consistency, efficiency, or quality, limiting their practical applications. In this work, we propose RTGaze, a real-time and high-quality gaze redirection method. Our approach learns a gaze-controllable facial representation from face images and gaze prompts, then decodes this representation via neural rendering for gaze redirection. Additionally, we distill face geometric priors from a pretrained 3D portrait generator to enhance generation quality. We evaluate RTGaze both qualitatively and quantitatively, demonstrating state-of-the-art performance in efficiency, redirection accuracy, and image quality across multiple datasets. Our system achieves real-time, 3D-aware gaze redirection with a feedforward network (~0.06 sec/image), making it 800× faster than the previous state-of-the-art 3D-aware methods.
Hengfei Wang, Zhongqun Zhang, Yihua Cheng, Hyung Jin Chang
AAAI3
2026 DroidSpeak: KV Cache Sharing Across Fine-tuned Model Variants
Yuhan Liu 0004, Shaoting Feng, Zhuohan Gu, Kuntai Du, Hanchen Li, Yihua Cheng, Junchen Jiang, Shan Lu 0001, Madan Musuvathi, Esha Choukse
NSDI8
2025 Single-view Image to Novel-view Generation for Hand-Object Interactions
abstract
Hand-object interaction modeling from a single RGB image is a significantly challenging task. Previous works typically reconstruct hand-object interactions as texture-less meshes, ignoring photo-realistic image generation. In this work, we introduce the HO123, a novel method to synthesize novel-view hand-object interaction images from a single image. To this end, we first train a 2D diffusion prior. Given the camera pose in novel views, our approach transfers the camera information into explicit hand representations, including hand depth and skeleton images. We propose a global hand embedding to control the diffusion model based on these hand representations. We then learn a 3D Gaussian splatting for novel-view rendering using the diffusion prior. However, occluded objects present a persistent challenge. To address this issue, we further introduce local hand embedding, where a contact field is defined in the 3D Gaussian Splatting. We leverage contact information to guide the rendering in the contact field. Extensive experiments on the HO3D and DexYCB datasets demonstrate that our method significantly outperforms state-of-the-art novel-view synthesis for hand-object interactions.
Zhongqun Zhang, Yihua Cheng, Eduardo Pérez-Pellitero, Yiren Zhou, Jiankang Deng, Hyung Jin Chang, Jifei Song
AAAI2
2025 Earth+: On-Board Satellite Imagery Compression Leveraging Historical Earth Observations
abstract
Due to limited downlink (satellite-to-ground) capacity, over 90% of the images captured by the earth-observation satellites are not downloaded to the ground. To overcome the downlink limitation, we present Earth+, a new on-board satellite imagery compression system that identifies and downloads only changed areas in each image compared to latest on-board reference images of the same location. The key of Earth+ is that it obtains latest on-board reference images by letting the ground stations upload images recently captured by all satellites in the constellation. To our best knowledge, Earth+ is the first system that leverages images across an entire satellite constellation to enable more images to be downloaded to the ground (by better satellite imagery compression). Our evaluation shows that to download images of the same area, Earth+ can reduce the downlink usage by 3.3× compared to state-of-the-art on-board image compression techniques without sacrificing imagery quality or using more resources (downlink, computation or storage).
Kuntai Du, Yihua Cheng, Peder A. Olsen, Shadi A. Noghabi, Junchen Jiang
ASPLOS (1)2
2025 3D Prior Is All You Need: Cross-Task Few-shot 2D Gaze Estimation
abstract
3D and 2D gaze estimation share the fundamental objective of capturing eye movements but are traditionally treated as two distinct research domains. In this paper, we introduce a novel cross-task few-shot 2D gaze estimation approach, aiming to adapt a pre-trained 3D gaze estimation network for 2D gaze prediction on unseen devices using only a few training images. This task is highly challenging due to the domain gap between 3D and 2D gaze, unknown screen poses, and limited training data. To address these challenges, we propose a novel framework that bridges the gap between 3D and 2D gaze. Our framework contains a physics-based differentiable projection module with learnable parameters to model screen poses and project 3D gaze into 2D gaze. The framework is fully differentiable and can integrate into existing 3D gaze networks without modifying their original architecture. Additionally, we introduce a dynamic pseudo-labelling strategy for flipped images, which is particularly challenging for 2D labels due to unknown screen poses. To overcome this, we reverse the projection process by converting 2D labels to 3D space, where flipping is performed. Notably, this 3D space is not aligned with the camera coordinate system, so we learn a dynamic transformation matrix to compensate for this misalignment. We evaluate our method on MPIIGaze, EVE, and GazeCapture datasets, collected respectively on laptops, desktop computers, and mobile devices. The superior performance highlights the effectiveness of our approach, and demonstrates its strong potential for real-world applications.
Yihua Cheng, Hengfei Wang, Zhongqun Zhang, Boeun Kim, Feng Lu 0005, Hyung Jin Chang
CVPR1
2025 Trajectory Mamba: Efficient Attention-Mamba Forecasting Model Based on Selective SSM
abstract
Motion prediction is crucial for autonomous driving, as it enables accurate forecasting of future vehicle trajectories based on historical inputs. This paper introduces Trajectory Mamba, a novel efficient trajectory prediction framework based on the selective state-space model (SSM). Conventional attention-based models face the challenge of computational costs that grow quadratically with the number of targets, hindering their application in highly dynamic environments. In response, we leverage the SSM to redesign the self-attention mechanism in the encoder-decoder architecture, thereby achieving linear time complexity. To address the potential reduction in prediction accuracy resulting from modifications to the attention mechanism, we propose a joint polyline encoding strategy to better capture the associations between static and dynamic contexts, ultimately enhancing prediction accuracy. Additionally, to balance prediction accuracy and inference speed, we adopted the decoder that differs entirely from the encoder. Through cross-state space attention, all target agents share the scene context, allowing the SSM to interact with the shared scene representation during decoding, thus inferring different trajectories over the next prediction steps. Our model achieves state-of-the-art results in terms of inference speed and parameter efficiency on both the Argoverse 1 and Argoverse 2 datasets. It demonstrates a four-fold reduction in FLOPs compared to existing methods and reduces parameter count by over 40% while surpassing the performance of the vast majority of previous methods. These findings validate the effectiveness of Trajectory Mamba in trajectory prediction tasks.
Yihua Cheng, Kezhi Wang
CVPR2
2025 PersonaBooth: Personalized Text-to-Motion Generation
abstract
This paper introduces Motion Personalization, a new task that generates personalized motions aligned with text descriptions using several basic motions containing Persona. To support this novel task, we introduce a new large-scale motion dataset called PerMo (PersonaMotion), which captures the unique personas of multiple actors. We also propose a multi-modal finetuning method of a pretrained motion diffusion model called PersonaBooth. PersonaBooth addresses two main challenges: i) A significant distribution gap between the persona-focused PerMo dataset and the pretraining datasets, which lack persona-specific data, and ii) the difficulty of capturing a consistent persona from the motions vary in content (action type). To tackle the dataset distribution gap, we introduce a persona token to accept new persona features and perform multi-modal adaptation for both text and visuals during finetuning. To capture a consistent persona, we incorporate a contrastive learning technique to enhance intra-cohesion among samples with the same persona. Furthermore, we introduce a context-aware fusion mechanism to maximize the integration of persona cues from multiple input motions. PersonaBooth outperforms state-of-the-art motion style transfer methods, establishing a new benchmark for motion personalization.
Boeun Kim, Hea In Jeong, JungHoon Sung, Yihua Cheng, Jeongmin Lee 0007, Ju Yong Chang, Sang-Il Choi, Younggeun Choi 0001, Saim Shin, Hyung Jin Chang
CVPR4
2025 CacheBlend: Fast Large Language Model Serving for RAG with Cached Knowledge Fusion
abstract
Large language models (LLMs) often incorporate multiple text chunks in their inputs to provide the necessary contexts. To speed up the prefill of the long LLM inputs, one can pre-compute the KV cache of a text and re-use the KV cache when the context is reused as the prefix of another LLM input. However, the reused text chunks are not always the input prefix, which makes precomputed KV caches not directly usable since they ignore the text's cross-attention with the preceding texts. Thus, the benefits of reusing KV caches remain largely unrealized.
Hanchen Li, Yuhan Liu 0004, Siddhant Ray, Yihua Cheng, Qizheng Zhang, Kuntai Du, Shan Lu 0001, Junchen Jiang
EuroSys5
2025 Multi-Hypothesis 3D Hand Mesh Recovering from a Single Blurry Image
abstract
Recovery of 3D hand mesh from blurry hand images is challenging due to the ambiguity. Most existing works attempt to solve this issue by exploiting physical and temporal constraints. However, those works ignore the fact that multiple feasible solutions exist. In this paper, we propose a two-stage Multi-Hypothesis Hand Mesh Recovery network, consisting of a generation and selection model. In the first stage, the generation model explicitly extracts the temporal information with an unfolder. Then, a multi-hypothesis Transformer generates multiple diverse hypotheses with a lightweight hypothesis embedding set. In the second stage, the selection model selects a subset of good-quality hypotheses. We additionally combine the classifying and ranking loss to better align with the target of the selection model. Extensive experiments show that the proposed method produces much more accurate results on blurry images. Source code is available at https://github.com/RandSF/Multi_Hypothesis_BlurHandNet.
Rongyu Chen, Zhongqun Zhang, Yihua Cheng, Hyung Jin Chang
ICME4
2025 PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications
abstract
Besides typical generative applications, like ChatGPT, GitHub Copilot, and Cursor, we observe an emerging trend that LLMs are increasingly used in traditional discriminative tasks, such as recommendation, credit verification, and data labeling. The key characteristic of these emerging use cases is that the LLM generates only a single output token, rather than an arbitrarily long sequence of tokens. We refer to this as a prefill-only workload. However, since existing LLM engines assume arbitrary output lengths, they fail to leverage the unique properties of prefill-only workloads. In this paper, we present PrefillOnly, the first LLM inference engine that improves the inference throughput and latency by fully embracing the properties of prefill-only workloads. First, since it generates only one token, PrefillOnly only needs to store the KV cache of only the last computed layer, rather than of all layers. This drastically reduces the GPU memory footprint of LLM inference and allows handling long inputs without using solutions that reduce throughput, such as cross-GPU KV cache parallelization. Second, because the output length is fixed, rather than arbitrary, PrefillOnly can precisely determine the job completion time (JCT) of each prefill-only request before it starts. This enables efficient JCT-aware scheduling policies such as shortest prefill first. PrefillOnly can process up to 4× larger queries per second without inflating the average and P99 latency.
Kuntai Du, Bowen Wang 0016, Chen Zhang 0001, Qing Lan, Hejian Sang, Yihua Cheng, Yifan Qiao 0002, Ion Stoica, Junchen Jiang
SOSP7
2025 Behavior-Aware Knowledge-Embedded Model for Driver Attention Prediction
abstract
Accurately predicting driver attention is crucial for enhancing advanced driving assistance systems and autonomous vehicles, attracting increasing research interest. Most existing approaches, rooted in general, task-free saliency detection, adopt data-driven paradigms to correlate bottom-up environmental situations with attention distributions. However, they often overlook the complex top-down task-driven aspects of driver attention that are fundamental for the safe navigation of driving tasks, leading to limitations in handling real-world scenarios. In this paper, we take an initial step to explore and introduce BKnet, a Behavior-aware Knowledge-embedded model that innovatively integrates driving behaviors and empirical knowledge. Specifically, inspired by the human long-term cognitive process, we introduce a novel knowledge memory mechanism. It dynamically associates varied traffic scenarios with consistent driving behaviors, fostering the generation of robust behavior-aware empirical knowledge representations. To this end, BKnet facilitates a nuanced and comprehensive simulation of drivers’ attention mechanisms, driven synergistically by both top-down and bottom-up processes. Additionally, we further contribute to the field by collecting a novel Behavior-Aware Driver Attention (BADA) dataset. To the best of our knowledge, BADA is the first attention dataset explicitly incorporated into real-world driving behavior tasks from multiple drivers. Lastly, comprehensive experiments underscore BKnet’s superiority over existing state-of-the-art approaches and validate the effectiveness and necessity of integrating behavior-aware knowledge into driver attention prediction.
Yuchen Zhou 0002, Chao Gou, Zipeng Guo, Yihua Cheng, Hyung Jin Chang
IEEE Trans. Circuits Syst. Video Technol.4
2024 What Do You See in Vehicle? Comprehensive Vision Solution for In-Vehicle Gaze Estimation
abstract
Driver's eye gaze holds a wealth of cognitive and intentional cues crucial for intelligent vehicles. Despite its sig-nificance, research on in-vehicle gaze estimation remains limited due to the scarcity of comprehensive and well-annotated datasets in real driving scenarios. In this pa-per, we present three novel elements to advance in-vehicle gaze research. Firstly, we introduce IVGaze, a pioneering dataset capturing in-vehicle gaze, collected from 125 sub-jects and covering a large range of gaze and head poses within vehicles. In this dataset, we propose a new vision-based solution for in-vehicle gaze collection, introducing a refined gaze target calibration method to tackle annotation challenges. Second, our research focuses on in-vehicle gaze estimation leveraging the IvGaze. In-vehicle face images often suffer from low resolution, prompting our in-troduction of a gaze pyramid transformer that leverages transformer-based multilevel features integration. Expanding upon this, we introduce the dual-stream gaze pyramid transformer (GazeDPTR). Employing perspective transfor-mation, we rotate virtual cameras to normalize images, uti-lizing camera pose to merge normalized and original images for accurate gaze estimation. GazeDPTR shows state-of-the-art performance on the IVGaze dataset. Thirdly, we explore a novel strategy for gaze zone classification by extending the GazeDPTR. A foundational tri-plane and project gaze onto these planes are newly defined. Leveraging both positional features from the projection points and visual attributes from images, we achieve superior performance compared to relying solely on visual features, sub-stantiating the advantage of gaze estimation. The project is available at https://yihua.zone/work/ivgaze.
Yihua Cheng, Yaning Zhu, Zongji Wang, Hongquan Hao, Yongwei Liu, Shiqing Cheng, Xi Wang 0021, Hyung Jin Chang
CVPR1
2024 NL2Contact: Natural Language Guided 3D Hand-Object Contact Modeling with Diffusion Model
Zhongqun Zhang, Hengfei Wang, Ziwei Yu, Yihua Cheng, Angela Yao, Hyung Jin Chang
ECCV (28)4
2024 TextGaze: Gaze-Controllable Face Generation with Natural Language
abstract
Generating face image with specific gaze information has attracted considerable attention in recent years. Existing approaches typically input gaze values directly for face generation, which is unnatural and requires annotated gaze datasets for training, thereby limiting its application. In this paper, we present a novel gaze-controllable face generation task that overcomes these limitations. Our approach inputs textual descriptions that describe human gaze and head behavior and generates corresponding face images. Our work first introduces a text-of-gaze dataset containing over 90k text descriptions spanning a dense distribution of gaze and head poses. We further propose a gaze-controllable text-to-face method. Our method contains a sketch-conditioned face diffusion module and a model-based sketch diffusion module. We define a face sketch based on facial landmarks and eye segmentation map. It provides a structured and detailed foundation for generating facial images. The face diffusion module generates face images from the face sketch, and the sketch diffusion module employs a 3D face model to generate face sketch from text description. Experiments on the FFHQ dataset show the effectiveness of our method. Our dataset is available at https://github.com/hengfei-wang/TextGaze.
Hengfei Wang, Zhongqun Zhang, Yihua Cheng, Hyung Jin Chang
ACM Multimedia3
2024 GRACE: Loss-Resilient Real-Time Video through Neural Codecs
Yihua Cheng, Hanchen Li, Anton Arapin, Qizheng Zhang, Yuhan Liu 0004, Kuntai Du, Francis Y. Yan, Amrita Mazumdar, Nick Feamster, Junchen Jiang
NSDI1
2024 CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving
abstract
As large language models (LLMs) take on complex tasks, their inputs are supplemented with longer contexts that incorporate domain knowledge. Yet using long contexts is challenging as nothing can be generated until the whole context is processed by the LLM. While the context-processing delay can be reduced by reusing the KV cache of a context across different inputs, fetching the KV cache, which contains large tensors, over the network can cause high extra network delays.
Yuhan Liu 0004, Hanchen Li, Yihua Cheng, Siddhant Ray, Qizheng Zhang, Kuntai Du, Shan Lu 0001, Ganesh Ananthanarayanan, Michael Maire, Henry Hoffmann, Ari Holtzman, Junchen Jiang
SIGCOMM3
2024 Multi-Modal Gaze Following in Conversational Scenarios
abstract
Gaze following estimates gaze targets of in-scene person by understanding human behavior and scene information. Existing methods usually analyze scene images for gaze following. However, compared with visual images, audio also provides crucial cues for determining human behavior. This suggests that we can further improve gaze following considering audio cues. In this paper, we explore gaze following tasks in conversational scenarios. We propose a novel multimodal gaze following framework based on our observation "audiences tend to focus on the speaker". We first leverage the correlation between audio and lips, and classify speakers and listeners in a scene. We then use the identity information to enhance scene images and propose a gaze candidate estimation network. The network estimates gaze candidates from enhanced scene images and we use MLP to match subjects with candidates as classification tasks. Existing gaze following datasets focus on visual images while ignore audios. To evaluate our method, we collect a conversational dataset, VideoGazeSpeech (VGS), which is the first gaze following dataset including images and audio. Our method significantly outperforms existing methods in VGS datasets. The visualization result also prove the advantage of audio cues in gaze following tasks. Our work will inspire more researches in multi-modal gaze following estimation.
Zhongqun Zhang, Nora Horanyi, Jaewon Moon, Yihua Cheng, Hyung Jin Chang
WACV5
2024 Appearance-Based Gaze Estimation With Deep Learning: A Review and Benchmark
abstract
Human gaze provides valuable information on human focus and intentions, making it a crucial area of research. Recently, deep learning has revolutionized appearance-based gaze estimation. However, due to the unique features of gaze estimation research, such as the unfair comparison between 2D gaze positions and 3D gaze vectors and the different pre-processing and post-processing methods, there is a lack of a definitive guideline for developing deep learning-based gaze estimation algorithms. In this paper, we present a systematic review of the appearance-based gaze estimation methods using deep learning. First, we survey the existing gaze estimation algorithms along the typical gaze estimation pipeline: deep feature extraction, deep learning model design, personal calibration and platforms. Second, to fairly compare the performance of different approaches, we summarize the data pre-processing and post-processing methods, including face/eye detection, data rectification, 2D/3D gaze conversion and gaze origin conversion. Finally, we set up a comprehensive benchmark for deep learning-based gaze estimation. We characterize all the public datasets and provide the source code of typical gaze estimation algorithms. This paper serves not only as a reference to develop deep learning-based gaze estimation methods, but also a guideline for future gaze estimation research.
Yihua Cheng, Haofei Wang 0001, Yiwei Bao, Feng Lu 0005
IEEE Trans. Pattern Anal. Mach. Intell.1
2023 High-Fidelity Eye Animatable Neural Radiance Fields for Human Face
Hengfei Wang, Zhongqun Zhang, Yihua Cheng, Hyung Jin Chang
BMVC3
2023 Raising the Level of Abstraction for Time-State Analytics With the Timeline Framework
Henry Milner, Yihua Cheng, Jibin Zhan, Hui Zhang 0001, Vyas Sekar, Junchen Jiang, Ion Stoica
CIDR2
2023 Online Profiling and Adaptation of Quality Sensitivity for Internet Video
abstract
A key to video streaming systems is knowing how sensitive quality of experience (QoE) is to quality metrics (e.g., buffering ratio and average bitrate). In the conventional wisdom, such quality sensitivity should be profiled by offline user studies because QoE is equally sensitive to quality metrics everywhere for an entire genre of videos. However, recent studies show that quality sensitivity varies substantially both across videos and within a video, giving rise to a new potential for improving QoE and serving more users without using more bandwidth. Unfortunately, offline profiling cannot capture the variability of quality sensitivity within a new video (e.g., a new TV show episode or live sports event), if users join to watch it within a short time window.
Yihua Cheng, Hui Zhang 0001, Junchen Jiang
SoCC1
2023 DVGaze: Dual-View Gaze Estimation
abstract
Gaze estimation methods estimate gaze from facial appearance with a single camera. However, due to the limited view of a single camera, the captured facial appearance cannot provide complete facial information and thus complicate the gaze estimation problem. Recently, camera devices are rapidly updated. Dual cameras are affordable for users and have been integrated in many devices. This development suggests that we can further improve gaze estimation performance with dual-view gaze estimation. In this paper, we propose a dual-view gaze estimation network (DV-Gaze). DV-Gaze estimates dual-view gaze directions from a pair of images. We first propose a dual-view interactive convolution (DIC) block in DV-Gaze. DIC blocks exchange dual-view information during convolution in multiple feature scales. It fuses dual-view features along epipolar lines and compensates for the original feature with the fused feature. We further propose a dual-view transformer to estimate gaze from dual-view features. Camera poses are encoded to indicate the position information in the transformer. We also consider the geometric relation between dual-view gaze directions and propose a dual-view gaze consistency loss for DV-Gaze. DV-Gaze achieves state-of-the-art performance on ETH-XGaze and EVE datasets. Our experiments also prove the potential of dual-view gaze estimation. We release codes in https://github.com/yihuacheng/DVGaze.
Yihua Cheng, Feng Lu 0005
ICCV1
2023 POLYCORN: Data-driven Cross-layer Multipath Networking for High-speed Railway through Composable Schedulerlets
Yunzhe Ni, Feng Qian 0001, Taide Liu, Yihua Cheng, Zhiyao Ma, Jing Wang 0077, Gang Huang 0001, Xuanzhe Liu, Chenren Xu
NSDI4
2022 PureGaze: Purifying Gaze Feature for Generalizable Gaze Estimation
abstract
Gaze estimation methods learn eye gaze from facial features. However, among rich information in the facial image, real gaze-relevant features only correspond to subtle changes in eye region, while other gaze-irrelevant features like illumination, personal appearance and even facial expression may affect the learning in an unexpected way. This is a major reason why existing methods show significant performance degradation in cross-domain/dataset evaluation. In this paper, we tackle the cross-domain problem in gaze estimation. Different from common domain adaption methods, we propose a domain generalization method to improve the cross-domain performance without touching target samples. The domain generalization is realized by gaze feature purification. We eliminate gaze-irrelevant factors such as illumination and identity to improve the cross-domain performance. We design a plug-and-play self-adversarial framework for the gaze feature purification. The framework enhances not only our baseline but also existing gaze estimation methods directly and significantly. To the best of our knowledge, we are the first to propose domain generalization methods in gaze estimation. Our method achieves not only state-of-the-art performance among typical gaze estimation methods but also competitive results among domain adaption methods. The code is released in https://github.com/yihuacheng/PureGaze.
Yihua Cheng, Yiwei Bao, Feng Lu 0005
AAAI1
2022 Gaze Estimation using Transformer
abstract
Recent work has proven the effectiveness of transformers in many computer vision tasks. However, the performance of transformers in gaze estimation is still unexplored. In this paper, we employ transformers and assess their effectiveness for gaze estimation. We consider two forms of vision transformer which are pure transformers and hybrid transformers. We first follow the popular ViT and employ a pure transformer to estimate gaze from images. On the other hand, we preserve the convolutional layers and integrate CNNs as well as transformers. The transformer serves as a component to complement CNNs. We compare the performance of the two transformers in gaze estimation. The Hybrid transformer significantly outperforms the pure transformer in all evaluation datasets with fewer parameters. We further conduct experiments to assess the effectiveness of the hybrid transformer and explore the advantage of the self-attention mechanism. Experiments show the hybrid transformer can achieve state-of-the-art performance in all benchmarks with pre-training. To facilitate further research, we release codes and models in https://github.com/yihuacheng/GazeTR.
Yihua Cheng, Feng Lu 0005
ICPR1
2021 Critique of "Planetary Normal Mode Computation: Parallel Algorithms, Performance, and Reproducibility" by SCC Team From Peking University
abstract
Shi et al. (2018) proposed a highly parallel polynomial filtering eigensolver for the computation of planetary normal modes. As a challenge at the Student Cluster Competition in The International Conference for High Performance Computing, Networking, Storage and Analysis (SC19), we reproduce the computational efficiency of the polynomial filtering eigensolver on our Intel Xeon machine. We present the weak scalability, scaling of runtime with model size (in a fixed interval) and the strong scalability results in this report.
Yihua Cheng, Zejia Fan, Jing Mai, Yifan Wu 0005, Pengcheng Xu 0005, Yuxuan Yan, Zhenxin Fu, Yun Liang 0001
IEEE Trans. Parallel Distributed Syst.1
2020 A Coarse-to-Fine Adaptive Network for Appearance-Based Gaze Estimation
abstract
Human gaze is essential for various appealing applications. Aiming at more accurate gaze estimation, a series of recent works propose to utilize face and eye images simultaneously. Nevertheless, face and eye images only serve as independent or parallel feature sources in those works, the intrinsic correlation between their features is overlooked. In this paper we make the following contributions: 1) We propose a coarse-to-fine strategy which estimates a basic gaze direction from face image and refines it with corresponding residual predicted from eye images. 2) Guided by the proposed strategy, we design a framework which introduces a bi-gram model to bridge gaze residual and basic gaze direction, and an attention component to adaptively acquire suitable fine-grained feature. 3) Integrating the above innovations, we construct a coarse-to-fine adaptive network named CA-Net and achieve state-of-the-art performances on MPIIGaze and EyeDiap.
Yihua Cheng, Shiyao Huang, Fei Wang 0032, Chen Qian 0006, Feng Lu 0005
AAAI1
2020 Adaptive Feature Fusion Network for Gaze Tracking in Mobile Tablets
abstract
Recently, many multi-stream gaze estimation methods have been proposed. They estimate gaze from eye and face appearances and achieve reasonable accuracy. However, most of the methods simply concatenate the features extracted from eye and face appearance. The feature fusion process has been ignored. In this paper, we propose a novel Adaptive Feature Fusion Network (AFF-Net), which performs gaze tracking task in mobile tablets. We stack two-eye feature maps and utilize Squeeze-and-Excitation layers to adaptively fuse two-eye features according to their similarity on appearance. Meanwhile, we also propose Adaptive Group Normalization to recalibrate eye features with the guidance of facial feature. Extensive experiments on both GazeCapture and MPIIFaceGaze datasets demonstrate consistently superior performance of the proposed method.
Yiwei Bao, Yihua Cheng, Yunfei Liu 0001, Feng Lu 0005
ICPR2
2020 A First Look at Disconnection-Centric TCP Performance on High-Speed Railways
abstract
High-speed rail (HSR) systems potentially provide a more efficient way of door-to-door transportation than airplane. However, they also pose unprecedented challenges in delivering seamless Internet service for on-board passengers. In this paper, we conduct the first large-scale disconnection-centric measurement study of TCP performance over LTE on HSR. Our measurement targets the main HSR route in China operating at 300/350 km/h. We performed extensive data collection obtaining 378.3 GB data collected over 56639 km of trips. Leveraging such a unique dataset, we measure important performance metrics such as TCP goodput, latency and loss rate across different congestion control algorithm, mobile carrier, and different train speed. We further develop the LTE disconnection taxonomy, and conduct a in-depth correlation study between TCP stall and LTE disconnection. Our findings reveal the networking performance on today's HSR environment “in the wild”, as well as identify several root causes of performance inefficiencies, which together highlight the need to develop dedicated protocol mechanisms that are friendly to extreme mobility.
Chenren Xu, Jing Wang 0077, Zhiyao Ma, Yihua Cheng, Yunzhe Ni, Wangyang Li, Feng Qian 0001, Yuanjie Li
IEEE J. Sel. Areas Commun.4
2020 Gaze Estimation by Exploring Two-Eye Asymmetry
abstract
Eye gaze estimation is increasingly demanded by recent intelligent systems to facilitate a range of interactive applications. Unfortunately, learning the highly complicated regression from a single eye image to the gaze direction is not trivial. Thus, the problem is yet to be solved efficiently. Inspired by the two-eye asymmetry as two eyes of the same person may appear uneven, we propose the face-based asymmetric regression-evaluation network (FARE-Net) to optimize the gaze estimation results by considering the difference between left and right eyes. The proposed method includes one face-based asymmetric regression network (FAR-Net) and one evaluation network (E-Net). The FAR-Net predicts 3D gaze directions for both eyes and is trained with the asymmetric mechanism, which asymmetrically weights and sums the loss generated by two-eye gaze directions. With the asymmetric mechanism, the FAR-Net utilizes the eyes that can achieve high performance to optimize network. The E-Net learns the reliabilities of two eyes to balance the learning of the asymmetric mechanism and symmetric mechanism. Our FARENet achieves leading performances on MPIIGaze, EyeDiap and RT-Gene datasets. Additionally, we investigate the effectiveness of FARE-Net by analyzing the distribution of errors and ablation study.
Yihua Cheng, Xucong Zhang, Feng Lu 0005, Yoichi Sato 0001
IEEE Trans. Image Process.1
2019 An Active-Passive Measurement Study of TCP Performance over LTE on High-speed Rails
abstract
High-speed rail (HSR) systems potentially provide a more efficient way of door-to-door transportation than airplane. However, they also pose unprecedented challenges in delivering seamless Internet service for on-board passengers. In this paper, we conduct a large-scale active-passive measurement study of TCP performance over LTE on HSR. Our measurement targets the HSR routes in China operating at above 300 km/h. We performed extensive data collection through both controlled setting and passive monitoring, obtaining 1732.9 GB data collected over 135719 km of trips. Leveraging such a unique dataset, we measure important performance metrics such as TCP goodput, latency, loss rate, as well as key characteristics of TCP flows, application breakdown, and users' behaviors. We further quantitatively study the impact of frequent cellular handover on HSR networking performance, and conduct in-depth examination of the performance of two widely deployed transport-layer protocols: TCP CUBIC and TCP BBR. Our findings reveal the performance of today's commercial HSR networks "in the wild'', as well as identify several performance inefficiencies, which motivate us to design a simple yet effective congestion control algorithm based on BBR to further boost the throughput by up to 36.5%. They together highlight the need to develop dedicated protocol mechanisms that are friendly to extreme mobility.
Jing Wang 0077, Yufan Zheng, Yunzhe Ni, Chenren Xu, Feng Qian 0001, Wangyang Li, Wantong Jiang, Yihua Cheng, Yuanjie Li, Xiufeng Xie
MobiCom8
2018 Appearance-Based Gaze Estimation via Evaluation-Guided Asymmetric Regression
Yihua Cheng, Feng Lu 0005, Xucong Zhang
ECCV (14)1
2018 Student Cluster Competition 2017, Team Peking University: Reproducing vectorization of the Tersoff multi-body potential on the Intel Broadwell architecture
Zhenxin Fu, Lei Yang 0031, Wenbin Hou, Yifan Wu 0005, Yihua Cheng, Yun Liang 0001
Parallel Comput.6