VLDB 2026 Research / reviewers in the wild / expert
Ye Yan 0001
dblp:85/6777-1
· DBLP profile ↗
33ranked-venue papers
0as first author
32since 2021 · last 2026
0000-0002-7514-0505ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 11 since 2021Computer networks · 9 · 8 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DBMIF: a deep balanced multimodal iterative fusion framework for air- and bone-conduction speech enhancement
Yilei Wu, Changyan Zheng, Yakun Zhang 0002, Chengshi Zheng, Ye Yan 0001, Erwei Yin |
Appl. Intell. | 7 |
| 2026 | MPFNet: A Multi-Prior Fusion Network With a Progressive Training Strategy for Micro-Expression RecognitionabstractMicro-expression recognition (MER), a critical subfield of affective computing, presents greater challenges than macro-expression recognition due to its brief duration and low intensity. While incorporating prior knowledge has been shown to enhance MER performance, existing methods predominantly rely on simplistic, singular sources of prior knowledge, failing to fully exploit multi-source information. This paper introduces the Multi-Prior Fusion Network (MPFNet), leveraging a progressive training strategy to optimize MER tasks. We propose two complementary encoders: the Generic Feature Encoder (GFE) and the Advanced Feature Encoder (AFE), both based on Inflated 3D ConvNets (I3D) with Coordinate Attention (CA) mechanisms, to improve the model's ability to capture spatiotemporal and channel-specific features. Inspired by developmental psychology, we present two variants of MPFNet—MPFNet-P and MPFNet-C—corresponding to two fundamental modes of infant cognitive development: parallel and hierarchical processing. These variants enable the evaluation of different strategies for integrating prior knowledge. Extensive experiments demonstrate that MPFNet significantly improves MER accuracy while maintaining balanced performance across categories, achieving accuracies of 0.811, 0.924, and 0.857 on the SMIC, CASME II, and SAMM datasets, respectively. To the best of our knowledge, our approach achieves state-of-the-art performance on the SMIC and SAMM datasets. The source code is available at:https://github.com/Mac0504/MPFNet. Shaokai Zhao, Dongdong Zhou, Zhiguo Luo, Liang Xie 0012, Ye Yan 0001, Erwei Yin |
IEEE Trans. Affect. Comput. | 7 |
| 2025 | A Multi-Prior Fusion Network for Video-based Micro-Expression RecognitionabstractThe analysis of facial micro-expressions (MEs) has emerged as a significant application and topic within the field of image and video processing. However, challenges persist due to the brief duration and subtle intensity of these spontaneous expressions. This paper presents a novel multi-prior fusion network (MPFNet) for ME recognition based on a progressive training strategy. During the prior learning phase, our model is trained using a dual-stream architecture to capture both generic and advanced ME features. In the classification phase, we merge two pre-trained models with complementary prior knowledge and employ weighted fusion for classification within a meta-learning framework. Additionally, this study employs inflated 3D ConvNets (I3D) as a feature encoder and integrates Coordinate Attention (CA) blocks to enhance the automatic learning of spatiotemporal and channel features of ME video sequences. Extensive experiments conducted on three benchmark datasets validate the effectiveness of our model. Shaokai Zhao, Liang Xie 0012, Erwei Yin, Ye Yan 0001 |
ICASSP | 6 |
| 2025 | Self-supervised Contrastive Pre-training for Dry Electrode EEG Emotion Recognition via Cross Device Representation ConsistencyabstractThe use of dry electrode electroencephalography (EEG) systems holds significant importance in advancing the everyday application of emotion recognition. However, adapting it to real-world applications faces unique challenges due to low signal-to-noise ratios and unreliable emotion labels. To address these challenges, we propose a Cross-Device Representation Consistency (CDRC) pre-training paradigm for dry EEG emotion recognition, where the self-supervised signal is provided by the distance between representations embedded in wet and dry EEG components and trained via contrastive estimation. Specifically, we employ a dual-branch embedding prediction task coupled with contrastive feature alignment module to extract robust and distinctive features from dry electrode EEG signals. We evaluate our model on an available emotional dataset PaDWEED, extensive experiments demonstrate that CDRC performs comparably to fully supervised training and achieves state-of-the-art results compared to several self-supervised approaches. Moreover, the remarkable performance on subject-independent tasks highlights its effectiveness in addressing and mitigating subject variability. Meihong Zhang, Shaokai Zhao, Zhiguo Luo, Liang Xie 0012, Tiejun Liu, Dezhong Yao 0001, Ye Yan 0001, Erwei Yin |
ICASSP | 7 |
| 2025 | FI-HGR: A Robust Hand Gesture Recognition System Based on Wearable Data Glove and Multimodal Fusion AlgorithmabstractHand gesture recognition (HGR) plays a crucial role in human-computer interaction systems within the Internet of Things (IoT). Recent HGR methods often rely on vision-based images or videos, which are limited in terms of occluded fingers and high computational cost due to complex neural networks. In contrast, wearable sensors like inertial measurement units (IMUs) and flexible sensors can handle hand self-obscuration. However, there are two unresolved issues. First, using a single modality is hard to balance high precision and low latency. Second, existing multimodal-based approaches lack deep inter-modal coupling to effectively address IMU drift and mechanical coupling of flexible sensors. To address these problems, we propose FI-HGR (HGR based on flexible and inertial data). FI-HGR comprises a sensor-integrated data glove and a novel Cascaded Complementary-Stochastic Fusion Algorithm (CS-Algorithm). CS-Algorithm employs six Mahony filters to estimate the state quaternion of each IMU, along with an Extended Kalman Filter that continuously corrects IMU drift based on the index finger’s bending angle sensed by a flexible sensor. This design allows a single flexible sensor to calibrate multiple IMUs and introduces a hard constraint, resolving the sensor drift problems that previous methods cannot. Based on the CS-Algorithm outputs, precise finger bending angles are estimated in real time. Subjective and objective experimental results show that our approach effectively addresses IMU drift and mechanical coupling in flexible sensors, reduces gesture tracking error to approximately 3.4∘, and significantly improves both recognition accuracy and operational efficiency. Tao Zhen, Buyuan Zhang, Dezhong Yao 0001, Liang Xie 0012, Ye Yan 0001, Erwei Yin |
IEEE Internet Things J. | 8 |
| 2025 | PanoGen++: Domain-adapted text-guided panoramic environment generation for vision-and-language navigation
Dongliang Zhou, Liang Xie 0012, Ye Yan 0001, Erwei Yin |
Neural Networks | 5 |
| 2025 | Neural Chinese silent speech recognition with facial electromyography
Liang Xie 0012, Yakun Zhang 0002, Meishan Zhang, Changyan Zheng, Ye Yan 0001, Erwei Yin |
Speech Commun. | 7 |
| 2025 | Identifying Stable EEG Patterns in Manipulation Task for Negative Emotion RecognitionabstractNegative emotion recognition during manipulation task plays crucial role in human-machine interaction, where diverse cognitive variables coexist and influence each other. However, traditional emotion experiments often overemphasize emotion induction while overlooking other practical cognitive tasks, which leads participants to suffer from simplistic emotional experiences and ultimately compromises the real-world applicability of the emotional data collected. To incorporate critical cognitive variables into emotion elicitation, we utilize joystick-based real-time emotion annotation to encourage subjects to continuously feel emotional intensity, to advisedly decide when to manipulate the joystick, and to physically operate it. Consequently, at least two essential cognitive variables—decision-making and action—are integrated into emotion perception. Following this, we develop a novel negative emotion dataset called CRED, which includes a variety of physiological data, particularly Electroencephalograph (EEG). To assess the stability of emotional EEG patterns, we employ strict statistical analysis and a dual-branch transformer (DBT) with the gradient-based attribution method on the proposed CRED. Additionally, two well-known public datasets (SEED and SEED-V) are used to verify the DBT. Compared to traditional methods, DBT improves classification accuracy by approximately 5% on CRED and by around 2% on the public datasets. The experimental results indicate that the occipital lobe plays a crucial role in the discrimination of negative emotions; the critical frequency bands vary between the five emotions in the CRED. Specifically, the low-delta rhythm is associated with anger, while fear is influenced by both theta and alpha rhythms; disgust is found to be significant in the theta rhythm; and for neutral emotions, both low-delta and alpha rhythms are identified as crucial. In summary, our findings demonstrate the existence of stable emotional EEG patterns when additional cognitive variables are involved. Shaokai Zhao, Liang Xie 0012, Zhiguo Luo, Dongdong Zhou, Ye Yan 0001, Erwei Yin |
IEEE Trans. Affect. Comput. | 7 |
| 2025 | PVEye: A Large Posture-Variant Eye Tracking Dataset for Head-Mounted AR DevicesabstractEye tracking technology, essential for enhancing user experience in virtual reality (VR) and augmented reality (AR) devices, has been widely incorporated into advanced head-mounted devices like the Apple Vision Pro and PICO 4 Pro, becoming a standard feature. However, dedicated eye tracking datasets for such devices are severely lacking, with existing datasets commonly facing issues like camera skew and low resolution, particularly failing to adequately consider the diversity in wearing postures. To address this gap, we have developed the Posture-Variant Eye Tracking Dataset (PVEye), which includes 11,044,800 high-resolution near-eye images from 104 participants, showcasing a rich variety of wearing postures. This dataset aims to advance the development and application of appearance-based eye tracking methods. Utilizing this dataset, our evaluations demonstrate that the appearance-based method, particularly the NVGaze model, provides improved accuracy and robustness compared to the traditional feature-based method. Crucially, our experiments indicate that variations in wearing posture can significantly impact eye tracking performance, with posture-related errors contributing approximately 45% to the overall error variance. Moreover, the study delves into the specific impact of calibration and other critical factors on eye tracking performance, offering insights for further optimization of tracking effectiveness. Xiaowei Bai, Liang Xie 0012, Yingxi Li, Qining Wang, Ye Yan 0001, Erwei Yin |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2024 | Two-Stream Vision Swin Transformer for Video-based Eye Movement Detection
Xiaowei Bai, Zhenyu Fang, Liang Xie 0012, Ye Yan 0001, Erwei Yin |
CogSci | 6 |
| 2024 | Landmark-Guided Cross-Speaker Lip Reading with Mutual Information RegularizationabstractLip reading, the process of interpreting silent speech from visual lip movements, has gained rising attention for its wide range of realistic applications. Deep learning approaches greatly improve current lip reading systems. However, lip reading in cross-speaker scenarios where the speaker identity changes, poses a challenging problem due to inter-speaker variability. A well-trained lip reading system may perform poorly when handling a brand new speaker. To learn a speaker-robust lip reading model, a key insight is to reduce visual variations across speakers, avoiding the model overfitting to specific speakers. In this work, in view of both input visual clues and latent representations based on a hybrid CTC/attention architecture, we propose to exploit the lip landmark-guided fine-grained visual clues instead of frequently-used mouth-cropped images as input features, diminishing speaker-specific appearance characteristics. Furthermore, a max-min mutual information regularization approach is proposed to capture speaker-insensitive latent representations. Experimental evaluations on public lip reading datasets demonstrate the effectiveness of the proposed approach under the intra-speaker and inter-speaker conditions. Linzhi Wu, Yakun Zhang 0002, Changyan Zheng, Tiejun Liu, Liang Xie 0012, Ye Yan 0001, Erwei Yin |
LREC/COLING | 7 |
| 2024 | Complex-Valued Gabor-Attention Residual Fusion Network for Iris RecognitionabstractIris recognition has gained significant attention in identity verification due to the unique, stable texture patterns in iris. Successfully extracting these patterns is essential for quick and precise identification. Although deep learning methods have automated the iris recognition, they predominantly rely on real-valued networks that overlook the complex-valued representation of iris texture. This means they cannot effectively process phase and amplitude information, and fail to integrate domain-specific knowledge of iris, thereby not fully capturing the intricate details of the iris texture. Inspired by classical manual methods that efficiently harness the complex-valued representation of the iris to extract both amplitude and phase information. We integrate Gabor filters with complex-valued neural networks, propose a Complex-Valued Gabor-Attention Residual Fusion Network (GRFN) tailored for iris recognition, aiming to comprehensively capture the iris texture’s multi-scale and multi-orientation phase and amplitude features. The GRFN incorporates adaptive Gabor Complex-Valued Convolution Kernels (GCVK) to introduce a Gabor attention mechanism focused on iris biometric characteristics. Furthermore, we propose a novel residual feature fusion approach that selects and merges local and global features across multiple directions and scales, mitigating model degradation and enhancing the network’s ability to extract iris texture features effectively. Extensive experiments show that the proposed network outperforms the state-of-the-art performance on two benchmark datasets. Zhuoru Li, Xiaowei Bai, Yingxi Li, Zhenyu Fang, Liang Xie 0012, Ye Yan 0001, Erwei Yin |
ECAI | 8 |
| 2024 | GloveTyping: A Hand Gesture Recognition System for Text Input Using a Hierarchical Framework with Attention Mechanism
Tao Zhen, Pengfei Ren 0001, Liang Xie 0012, Ye Yan 0001, Erwei Yin |
ICONIP (5) | 7 |
| 2024 | Trajectory-based Calibration for Optical See-Through Head-Mounted Displays Without Alignment
Shaohua Zhao, Wei Chen 0092, Zhongchen Shi, Liang Xie 0012, Ye Yan 0001, Erwei Yin |
PRCV (6) | 6 |
| 2024 | Challenges and solutions for vision-based hand gesture interpretation: A review
Kun Gao 0002, Liang Xie 0012, Ye Yan 0001, Erwei Yin |
Comput. Vis. Image Underst. | 7 |
| 2024 | MVINS: Tightly Coupled Mocap-Visual-Inertial Fusion for Global and Drift-Free Pose EstimationabstractAugmented reality (AR), a prominent application within the Internet of Things (IoT) domain, demands high-performance pose estimation. Presently, the visual-inertial navigation system (VINS) is acknowledged as an essential method for providing 6-DoF poses. However, VINS builds the local frame at random during the system initialization stage, making it difficult to establish a connection with the global frame. In addition, VINS is prone to drifting. In this paper, we propose an innovative method that tightly couples markerless motion capture (Mocap) with vision and an IMU to achieve global and drift-free pose estimation for AR glasses. To address the issue of pose initialization and establish a connection between the IMU and Mocap, we introduce a coarse-to-fine initialization strategy, enabling data fusion for Mocap, vision, and the IMU under a unified global frame. Furthermore, we formulate the Mocap factor alongside the visual and inertial factors and integrate them into a factor graph framework to constrain the system states. With a spatiotemporal calibration method, the IMU-Mocap extrinsic parameter and time offset are calibrated online to improve the pose estimation accuracy. Experimental evaluations in real-world experiments demonstrate the capability of our method to accurately estimate drift-free poses in the global frame. Compared to the state-of-the-art VINS-Fusion, ORB-SLAM3, and GVIS, we achieve improvements of 81%, 42%, and 33% in translation accuracy and improvements of 58%, 33%, and 72% in rotation accuracy, respectively. Moreover, we also evaluate our system for the EuRoC dataset, further indicating the effectiveness of the proposed work. Liang Xie 0012, Wei Wang 0076, Zhongchen Shi, Wei Chen 0092, Ye Yan 0001, Erwei Yin |
IEEE Internet Things J. | 6 |
| 2024 | Progressively global-local fusion with explicit guidance for accurate and robust 3d hand pose reconstruction
Kun Gao 0002, Pengfei Ren 0001, Tao Zhen, Liang Xie 0012, Zhongkui Li, Ye Yan 0001, Erwei Yin |
Knowl. Based Syst. | 8 |
| 2024 | Trajectory-based alignment for optical see-through HMD calibrationabstractAbstract In order to align the virtual and real content precisely through augmented reality devices, especially in optical see-through head-mounted displays (OST-HMD), it is necessary to calibrate the device before using it. However, most existing methods estimated the parameters via 3D-2D correspondences based on the 2D alignment, which is cumbersome, time-consuming, theoretically complex, and results in insufficient robustness. To alleviate this issue, in this paper, we propose an efficient and simple calibration method based on the principle of directly calculating the projection transformation between virtual space and the real world via 3D-3D alignment. The proposed method merely needs to record the motion trajectory of the cube-marker in the real and virtual world, and then calculate the transformation matrix between the virtual space and the real world by aligning the two trajectories in the observed view. There are two advantages associated with the proposed method. First, the operation is simple. Theoretically, the user only needs to perform four alignment operations for calibration without changing the rotation variation. Second, the trajectory can be easily distributed throughout the entire observation view, resulting in more robust calibration results. To validate the effectiveness of the proposed method, we conducted extensive experiments on our self-built optical see-through head-mounted display (OST-HMD) device. The experimental results show that the proposed method can achieve better calibration results than other calibration methods. Lingling Chen, Shaohua Zhao, Wei Chen 0092, Zhongchen Shi, Liang Xie 0012, Ye Yan 0001, Erwei Yin |
Multim. Tools Appl. | 6 |
| 2024 | Real-Time Gaze Tracking via Head-Eye Cues on Head Mounted DevicesabstractGaze is a crucial element in human-computer interaction and plays an increasingly vital role in promoting the adoption of head-mounted devices (HMDs). Existing gaze tracking methods for HMDs either demand user calibration or face challenges in balancing accuracy and speed, compromising the overall user experience. In this paper, we introduce a novel strategy for real-time, calibration-free gaze tracking using joint head-eye cues on HMDs. Initially, we create a multimodal gaze tracking dataset named HE-Gaze, encompassing synchronized eye images and 6DoF head movement data, addressing a gap in the current data landscape. Statistical analyses unveil the correlation between head movements and gaze positions. Building on these insights, we introduce the hierarchical head-eye coordinated gaze tracking model (HHE-Tracker), which incorporates two lightweight branches to encode input eye images and head sequences efficiently. It combines encoded head velocity and posture features with eye features across various scales to infer gaze position. HHE-Tracker was implemented on a commercial HMD, and its performance was assessed in unconstrained scenarios. The results demonstrate the HHE-Tracker's capability to accurately estimate gaze positions in real-time. In comparison to the state-of-the-art gaze tracking algorithm, HHE-Tracker exhibits commendable accuracy (3.47$^{\circ }$) and a 40-fold speedup (81FPSon a Snapdragon 845 SoC). Yingxi Li, Xiaowei Bai, Liang Xie 0012, Feng Lu 0005, Feitian Zhang, Ye Yan 0001, Erwei Yin |
IEEE Trans. Mob. Comput. | 7 |
| 2024 | Vision-Language Navigation With Beam-Constrained Global NormalizationabstractVision-language navigation (VLN) is a challenging task, which guides an agent to navigate in a realistic environment by natural language instructions. Sequence-to-sequence modeling is one of the most prospective architectures for the task, which achieves the agent navigation goal by a sequence of moving actions. The line of work has led to the state-of-the-art performance. Recently, several studies showed that the beam-search decoding during the inference can result in promising performance, as it ranks multiple candidate trajectories by scoring each trajectory as a whole. However, the trajectory-level score might be seriously biased during ranking. The score is a simple averaging of individual unit scores of the target-sequence actions, and these unit scores could be incomparable among different trajectories since they are calculated by a local discriminant classifier. To address this problem, we propose a global normalization strategy to rescale the scores at the trajectory level. Concretely, we present two global score functions to rerank all candidates in the output beam, resulting in more comparable trajectory scores. In this way, the bias problem can be greatly alleviated. We conduct experiments on the benchmark room-to-room (R2R) dataset of VLN to verify our method, and the results show that the proposed global method is effective, providing significant performance than the corresponding baselines. Our final model can achieve competitive performance on the VLN leaderboard. Liang Xie 0012, Meishan Zhang, Ye Yan 0001, Erwei Yin |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2023 | Data Volume-Aware Computation Task Scheduling for Smart Grid Data Analytic ApplicationsabstractEmerging smart grid applications analyze large amounts of data collected from millions of meters and systems to facilitate distributed monitoring and real-time control tasks. However, current parallel data processing systems are designed for common applications, unaware of the massive volume of the collected data, causing long data transfer delay during the computation and slow response time of smart grid systems. A promising direction to reduce delay is to jointly schedule computation tasks and data transfers. We identify that the smart grid data analytic jobs require the intermediate data among different computation stages to be transmitted orderly to avoid network congestion. This new feature prevents current scheduling algorithms from being efficient. In this work, an integrated computing and communication task scheduling scheme is proposed. The mathematical formulation of smart grid data analytic jobs scheduling problem is given, which is unsolvable by existing optimization methods due to the strongly coupled constraints. Several techniques are combined to linearize it for adapting the Branch and Cut method. Based on the topological information in the job graph, the Topology Aware Branch and Cut method is further proposed to speed up searching for optimal solutions. Numerical results demonstrate the effectiveness of the proposed method. Binquan Guo, Hongyan Li 0001, Ye Yan 0001, Zhou Zhang 0004, Peng Wang 0044 |
ICC | 3 |
| 2023 | Online Network Slicing for Real Time Applications in Large-scale Satellite NetworksabstractIn this work, we investigate resource allocation strategy for real time communication (RTC) over satellite networks with virtual network functions. Enhanced by inter-satellite links (ISLs), in-orbit computing and network virtualization technologies, large-scale satellite networks promise global coverage at low-latency and high-bandwidth for RTC applications with diversified functions. However, realizing RTC with specific function requirements using intermittent ISLs, requires efficient routing methods with fast response times. We identify that such a routing problem over time-varying graph can be formulated as an integer linear programming problem. The branch and bound method incurs$\mathcal{O}(\vert \mathcal{L}^{\tau}\vert \cdot(3\vert \mathcal{V}^{\tau}\vert+\vert \mathcal{L}^{\tau}\vert )^{\vert \mathcal{L}^{\tau}\vert })$time complexity, where$\vert \mathcal{V}^{\tau}\vert$is the number of nodes, and$\vert \mathcal{L}^{\tau}\vert$is the number of links during time interval$\tau$. By adopting a k-shortest path-based algorithm, the theoretical worst case complexity becomes$O(\vert \mathcal{V}^{\tau}\vert !\vert \mathcal{V}^{\tau}\vert ^{3})$. Although it runs fast in most cases, its solution can be sub-optimal and may not be found, resulting in compromised acceptance ratio in practice. To overcome this, we further design a graph-based algorithm by exploiting the special structure of the solution space, which can obtain the optimal solution in polynomial time with a computational complexity of$\mathrm{O}(3\vert \mathcal{L}^{\tau}\vert +(2\log\vert \mathcal{V}^{\tau}\vert +1)\vert \mathcal{V}^{T}\vert )$. Simulations conducted on starlink constellation with thousands of satellites corroborate the effectiveness of the proposed algorithm. Binquan Guo, Hongyan Li 0001, Zhou Zhang 0004, Ye Yan 0001 |
ICC | 4 |
| 2023 | Grounded Entity-Landmark Adaptive Pre-training for Vision-and-Language NavigationabstractCross-modal alignment is one key challenge for Vision-and-Language Navigation (VLN). Most existing studies concentrate on mapping the global instruction or single sub-instruction to the corresponding trajectory. However, another critical problem of achieving fine-grained alignment at the entity level is seldom considered. To address this problem, we propose a novel Grounded Entity-Landmark Adaptive (GELA) pre-training paradigm for VLN tasks. To achieve the adaptive pre-training paradigm, we first introduce grounded entity-landmark human annotations into the Room-to-Room (R2R) dataset, named GEL-R2R. Additionally, we adopt three grounded entity-landmark adaptive pre-training objectives: 1) entity phrase prediction, 2) landmark bounding box prediction, and 3) entity-landmark semantic alignment, which explicitly supervise the learning of fine-grained cross-modal alignment between entity phrases and environment landmarks. Finally, we validate our model on two downstream benchmarks: VLN with descriptive instructions (R2R) and dialogue instructions (CVDN). The comprehensive experiments show that our GELA model achieves state-of-the-art results on both tasks, demonstrating its effectiveness and generalizability. Liang Xie 0012, Yakun Zhang 0002, Meishan Zhang, Ye Yan 0001, Erwei Yin |
ICCV | 5 |
| 2023 | Auxiliary Fine-grained Alignment Constraints for Vision-and-Language NavigationabstractVision-and-Language Navigation (VLN) requires a visual agent to navigate in photo-realistic environments following instructions. Fine-grained cross-modal alignment is one critical challenge in VLN because the agent needs to focus on a particular sub-part within the complete instruction for the next movement. However, previous work failed to implement explicit supervision for matching the sub-trajectory to the corresponding sub-instruction. In this paper, we propose Auxiliary Fine-grained Alignment Constraints (AFAC) to facilitate decision-making learning during navigation. AFAC consists of two constraints, i.e., Attention Alignment Constraint (AAC) and Representation Alignment Constraint (RAC), which produce additional supervising signals from the perspective of attention and representation respectively. We test our method on the Landmark-RxR benchmark and achieve state-of-the-art results both in seen and unseen environments. Ruqiang Huang, Yakun Zhang 0002, Yingjie Cen, Liang Xie 0012, Ye Yan 0001, Erwei Yin |
ICME | 6 |
| 2023 | One-Stage Wireframe Parsing in Fish-Eye Images
Ruqiang Huang, Zhongchen Shi, Wei Chen 0092, Liang Xie 0012, Ye Yan 0001, Erwei Yin |
PRCV (11) | 6 |
| 2023 | Resource-Constraint Network Selection for IoT Under the Unknown and Dynamic Heterogeneous Wireless EnvironmentabstractWe investigate the problem of network selection for the Internet of Things (IoT) to maximize the Quality of Experience (QoE) in a heterogeneous wireless environment. Different from the traditional network access approaches with the assumption that the network state information (NSI) is static and known a priori, a scenario where the NSI of networks is unknown and dynamic to IoT devices is considered. Due to hardware limitations in IoT, the device has a limited resource budget and consumes resources, e.g., the energy, during the process of network access. To maximize the cumulative QoE before resource exhausts, the device should learn and estimate the NSI of networks, and make appropriate decisions to balance the network selection and resource consumption. To address this issue, we formulate the problem as a combination of multiarmed bandit and optimization problems and propose two algorithms NCSA and network selection algorithm (NSA). Moreover, we consider the fact that the device has various traffic types and propose an algorithm MT-NSA. Theoretical analysis shows that the regret of all algorithms has a sublinear relationship with the resource budget. The effectiveness is validated by simulations. Zuohong Xu, Zhou Zhang 0004, Shilian Wang, Ye Yan 0001, Qian Cheng 0001 |
IEEE Internet Things J. | 4 |
| 2023 | DOS by Dynamic Groups: a Coalition Formation Game Perspective
Jia Xie, Zhou Zhang 0004, Ye Yan 0001, Hailin Zhang 0001 |
Mob. Networks Appl. | 3 |
| 2022 | Optimal Job Scheduling and Bandwidth Augmentation in Hybrid Data Center NetworksabstractOptimizing data transfers is critical for improving job performance in data-parallel frameworks. In the hybrid data center with both wired and wireless links, reconfigurable wireless links can provide additional bandwidth to speed up job execution. However, it requires the scheduler and transceivers to make joint decisions under coupled constraints. In this work, we identify that the joint job scheduling and bandwidth augmentation problem is a complex mixed integer nonlinear problem, which is not solvable by existing optimization methods. To address this bottleneck, we transform it into an equivalent problem based on the coupling of its heuristic bounds, the revised data transfer representation and non-linear constraints decoupling and reformulation, such that the optimal solution can be efficiently acquired by the Branch and Bound method. Based on the proposed method, the performance of job scheduling with and without bandwidth augmentation is studied. Experiments show that the performance gain depends on multiple factors, especially the data size. Compared with existing solutions, our method can averagely reduce the job completion time by up to 10% under the setting of production scenario. Binquan Guo, Zhou Zhang 0004, Ye Yan 0001, Hongyan Li 0001 |
GLOBECOM | 3 |
| 2022 | Improved Word-level Lipreading with Temporal Shrinkage Network and NetVLADabstractIn most word-level lipreading architectures of recent years, temporal feature extraction module tend to employ Multi-scale Temporal Convolution Network (MS-TCN). In our experiments, we have noticed it is hard for MS-TCN to deal with noise information that may contain in image sequences. In order to solve the problems, we propose a lipreading architecture based on temporal shrinkage network and NetVLAD. We first propose Temporal Shrinkage Unit according to Residual Shrinkage Network and then replace temporal convolution unit with it. The improved network which named Multi-scale Temporal Shrinkage Network (MS-TSN) could focus more on relevant information. Following with MS-TSN that deals with noise frames, NetVLAD is proposed to integrate local information into global feature. Compared with Global Average Pooling, NetVLAD could extract key features by clustering. Our experiments on Lipreading in the Wild (LRW) show that the architecture we propose achieves an accuracy of 89.41%, attaining new state-of-the-art in word-level lipreading. In addition, we build a new Mandarin Chinese lipreading dataset named MCLR-100 and verify our proposed architecture on it. Tao Luo 0010, Yakun Zhang 0002, Mingwu Song, Liang Xie 0012, Ye Yan 0001, Erwei Yin |
ICMI | 6 |
| 2022 | Real-time Gaze Tracking with Head-eye Coordination for Head-mounted DisplaysabstractHigh-accuracy, low-latency gaze tracking is becoming one of the indispensable features in augmented reality (AR) head-mounted devices (HMDs). Researchers have proposed different approaches to predict gaze positions from eye images. However, since only the eye modality is focused, these appearance-based algorithms are still struggle to trade off the accuracy and running speed in HMDs. In this paper, we propose a lightweight multi-modal network (HE-Tracker) to regress gaze positions. By fusing head-movement features with eye features, HE-Tracker achieves comparable accuracy (3.655° in all subjects) and $27 \times$ speedup (48 fps in the specialized AR HMD) compared to the state-of-the-art gaze tracking algorithm. We further demonstrate that when applying our head-eye coordination strategy to other baseline models, all these models achieve at least 6.36% performance improvement without a pronounced effect on running speed. Moreover, we construct HE-Gaze, the first multi-modal dataset with eye images and head-movement data for near-eye gaze tracking. This dataset is currently made of 757,360 frames and 15 persons, providing an opportunity to foster research in multi-modal gaze tracking approaches. Our dataset is available at DOWNLOAD LINK1. Lingling Chen, Yingxi Li, Xiaowei Bai, Yongqiang Hu, Mingwu Song, Liang Xie 0012, Ye Yan 0001, Erwei Yin |
ISMAR | 8 |
| 2022 | Auto calibration of multi-camera system for human pose estimationabstractAbstract The multi‐camera calibration is an essential step for many spatially aware applications, such as robotic navigation, augmented reality, and 3D human pose estimation. Traditional calibration methods use off‐the‐shelf checkerboards or triangles as the known world coordinate system and their corresponding corners are set as control points, which heavily depends on specific calibration patterns and is not suitable for calibration pattern‐denied environments. In this paper, an automatic calibration method is proposed to calibrate the multi‐camera system without the aid of a known calibration pattern. The key idea of the proposed method is that the authors consider the human body, which is always available, as the counterpart of the calibration pattern. The authors’ approach starts with binocular camera calibration, in which the extrinsic and intrinsic parameters are calculated in order and followed by a joint optimisation. With the results of each pair of binocular camera calibration, the multi‐camera system calibration is carried out in three steps: (i) parameters initialisation, (ii) extrinsic parameters optimisation, and (iii) jointly optimising intrinsic and extrinsic parameters. Since the authors’ approach does not require additional calibration patterns except for one visible person, it is flexible and easy to be implemented. Real experiments are conducted in different scenes, camera angles, and camera settings. Human pose estimation with the multi‐camera system is additionally performed for exhaustive experiments. The experimental results demonstrate that the authors’ method shows superior performance than the traditional method with the aid of a specific calibration pattern. Lingling Chen, Liang Xie 0012, Jian Yin 0025, Shuwei Gan, Ye Yan 0001, Erwei Yin |
IET Comput. Vis. | 6 |
| 2021 | A Fusion Framework to Enhance sEMG-Based Gesture Recognition Using TD and FD Features
Yao Luo, Tao Luo 0010, Qianchen Xia, Huijiong Yan, Liang Xie 0012, Ye Yan 0001, Erwei Yin |
ICONIP (6) | 6 |
| 2020 | Distributed Scheduling in Wireless Multiple Decode-and-forward Relay Networks
Zhou Zhang 0004, Ye Yan 0001, Zuohong Xu |
Mob. Networks Appl. | 2 |