VLDB 2026 Research / reviewers in the wild / expert
Jiyang Yu
dblp:52/7702
· DBLP profile ↗
22ranked-venue papers
11as first author
9since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 6 first-author · 6 since 2021Artificial intelligence and machine learning · 9 · 7 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 1 since 2021Systems, architecture and hardware · 3 · 2 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SensorFlow: Sensor and Image Fused Video StabilizationabstractWe present SensorFlow, a novel image and sensor fusion framework for robust, high-quality video stabilization. We start with sensor-based pre-stabilization that smooths out large-scale camera motion. A new angular velocity domain optimization has been introduced to achieve frame rate in-variance. We then feed the stabilized optical flows into an occlusion-aware 3D CNN that infers dense warp fields to remove residual translation and jitter. To further avoid dis-tortion, we propose a novel masking scheme to determine the disoccluded and dynamic regions in optical flow and in-paint them with spatially smooth flow vectors. Our method is appealing as it shares both the dense warping field's flex-ibility to correct complex motions and the robustness of sen-sor data for arbitrarily challenging scenes. We have vali-dated its effectiveness and demonstrated our solution out-performs state-of-the-art alternatives via extensive ablation studies and quantitative comparisons. Jiyang Yu, Fuhao Shi, Chia-Kai Liang |
WACV | 1 |
| 2025 | Monocular Visual SLAM With Adjusting Neural Radiance Fields for 3-D Reconstruction in Planetary EnvironmentsabstractIn planetary environments, conducting autonomous exploration tasks requires rovers to autonomously navigate the scene and achieve a detailed understanding of the terrain. Vision-based simultaneous localization and mapping (SLAM), which utilizes compact and low-power visual sensors for autonomous exploration, offers significant advantages in hardware deployment. While several methods have been proposed to apply visual navigation in planetary scenarios, they often rely on aerial imagery from orbiters and high-resolution DEMs for assistance. Additionally, accurate camera poses are typically required for dense matching during scene reconstruction, and the inability to perform loop closure significantly limits the performance of visual SLAM. Here, we propose a monocular visual SLAM approach combined with an adjusted neural radiance field for autonomous navigation and 3D reconstruction in planetary environments. Our approach solely relies on visual images as input and leverages the powerful learning capabilities of neural radiance fields to adapt to unseen scenes while simultaneously regressing both camera poses and scene representations. The estimated depth maps and poses can be further used for 3D reconstruction, assisting planetary exploration missions. The proposed method was tested on the Devon and MADMAX datasets that simulate planetary environments and achieved remarkable results. Even under the fixed rover navigation perspective, our pose estimation accuracy outperforms classical visual SLAM and other deep learning-based SLAM methods. Additionally, our novel view synthesis results exhibit quality comparable to those in terrestrial scenes. Comparisons with MVS techniques in terms of 3D reconstruction demonstrate that our approach recovers finer surface details. We also applied our method to the Perseverance rover dataset and achieved satisfactory positioning and reconstruction results in a real Martian environment, proving the practical feasibility of our method. Rong Huang 0001, Chen Liu 0040, Huan Xie 0001, Jiyang Yu, Yusheng Xu, Zhen Ye 0009, Xiaohua Tong |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Parallel accelerated computing architecture for dim target tracking on-boardabstractAbstract The real‐time tracking process of dim targets in space is mainly achieved through the correlation and prediction of dots after the detection and calculation process. The on‐board calculation of the tracking needs to be completed in milliseconds, and it needs to reach the microsecond level at high frame rates. For real‐time tracking of dim targets in space, it is necessary to achieve universal tracking calculation acceleration in response to different space regions and complex backgrounds, which poses high requirements for engineering implementation architecture. This paper designs a Kalman filter calculation based on digital logic parallel acceleration architecture for real‐time solution of dim target tracking on‐board. A unified architecture of Vector Processing Element (VPE) was established for the calculation of Kalman filtering matrix, and an array computing structure based on VPE was designed to decompose the entire filtering process and form a parallel pipelined data stream. The prediction errors under different fixed point bit widths were analyzed and deduced, and the guidance methods for selecting the optimal bit width based on the statistical results were provided. The entire design was engineered based on Xilinx's XC7K325T, resulting in an energy efficiency improvement compared to previous designs. The single iteration calculation time does not exceed 0.7 microseconds, which can meet the current high frame rate target tracking requirements. The effectiveness of this design has been verified through simulation of random trajectory data, which is consistent with the theoretical calculation error. Jiyang Yu, Xianjie Wang |
Comput. Intell. | 1 |
| 2024 | EI-MVSNet: Epipolar-Guided Multi-View Stereo Network With Interval-Aware LabelabstractRecent learning-based methods demonstrate their strong ability to estimate depth for multi-view stereo reconstruction. However, most of these methods directly extract features via regular or deformable convolutions, and few works consider the alignment of the receptive fields between views while constructing the cost volume. Through analyzing the constraint and inference of previous MVS networks, we find that there are still some shortcomings that hinder the performance. To deal with the above issues, we propose an Epipolar-Guided Multi-View Stereo Network with Interval-Aware Label (EI-MVSNet), which includes an epipolar-guided volume construction module and an interval-aware depth estimation module in a unified architecture for MVS. The proposed EI-MVSNet enjoys several merits. First, in the epipolar-guided volume construction module, we construct cost volume with features from aligned receptive fields between different pairs of reference and source images via epipolar-guided convolutions, which take rotation and scale changes into account. Second, in the interval-aware depth estimation module, we attempt to supervise the cost volume directly and make depth estimation independent of extraneous values by perceiving the upper and lower boundaries, which can achieve fine-grained predictions and enhance the reasoning ability of the network. Extensive experimental results on two standard benchmarks demonstrate that our EI-MVSNet performs favorably against state-of-the-art MVS methods. Specifically, our EI-MVSNet ranks$1_{st}$on both intermediate and advanced subsets of the Tanks and Temples benchmark, which verifies the high precision and strong robustness of our model. Tianzhu Zhang 0001, Jiyang Yu, Feng Wu 0001 |
IEEE Trans. Image Process. | 4 |
| 2023 | SE-ORNet: Self-Ensembling Orientation-Aware Network for Unsupervised Point Cloud Shape CorrespondenceabstractUnsupervised point cloud shape correspondence aims to obtain dense point-to-point correspondences between point clouds without manually annotated pairs. However, humans and some animals have bilateral symmetry and various orientations, which lead to severe mispredictions of symmetrical parts. Besides, point cloud noise disrupts consistent representations for point cloud and thus degrades the shape correspondence accuracy. To address the above issues, we propose a Self-Ensembling ORientation-aware Network termed SE-ORNet. The key of our approach is to exploit an orientation estimation module with a domain adaptive discriminator to align the orientations of point cloud pairs, which significantly alleviates the mispredictions of symmetrical parts. Additionally, we design a self-ensembling framework for unsupervised point cloud shape correspondence. In this framework, the disturbances of point cloud noise are overcome by perturbing the inputs of the student and teacher networks with different data augmentations and constraining the consistency of predictions. Extensive experiments on both human and animal datasets show that our SE-ORNet can surpass state-of-the-art unsupervised point cloud shape correspondence methods. Jiacheng Deng 0002, Chuxin Wang, Jiahao Lu 0001, Tianzhu Zhang 0001, Jiyang Yu |
CVPR | 6 |
| 2023 | Adaptive Spot-Guided Transformer for Consistent Local Feature MatchingabstractLocal feature matching aims at finding correspondences between a pair of images. Although current detector-free methods leverage Transformer architecture to obtain an impressive performance, few works consider maintaining local consistency. Meanwhile, most methods struggle with large scale variations. To deal with the above issues, we propose Adaptive Spot-Guided Transformer (ASTR) for local feature matching, which jointly models the local consistency and scale variations in a unified coarse-to-fine architecture. The proposed ASTR enjoys several merits. First, we design a spot-guided aggregation module to avoid interfering with irrelevant areas during feature aggregation. Second, we design an adaptive scaling module to adjust the size of grids according to the calculated depth information at fine stage. Extensive experimental results on five standard benchmarks demonstrate that our ASTR performs favorably against state-of-the-art methods. Our code will be released on https://astr2023.github.io. Jiahuan Yu, Tianzhu Zhang 0001, Jiyang Yu, Feng Wu 0001 |
CVPR | 5 |
| 2022 | Memory-Augmented Non-Local Attention for Video Super-ResolutionabstractIn this paper, we propose a simple yet effective video super-resolution method that aims at generating highfidelity high-resolution (HR) videos from low-resolution (LR) ones. Previous methods predominantly leverage temporal neighbor frames to assist the super-resolution of the current frame. Those methods achieve limited performance as they suffer from the challenges in spatial frame alignment and the lack of useful information from similar LR neighbor frames. In contrast, we devise a cross-frame non-local attention mechanism that allows video superresolution without frame alignment, leading to being more robust to large motions in the video. In addition, to acquire general video prior information beyond neighbor frames, and to compensate for the information loss caused by large motions, we design a novel memory-augmented attention module to memorize general video details during the superresolution training. We have thoroughly evaluated our work on various challenging datasets. Compared to other recent video super-resolution approaches, our method not only achieves significant performance gains on large motion videos but also shows better generalization. Our source code and the new Parkour benchmark dataset is available at https://github.com/jiy173/MANA. Jiyang Yu, Jingen Liu, Liefeng Bo, Tao Mei 0001 |
CVPR | 1 |
| 2021 | Real-Time Selfie Video StabilizationabstractWe propose a novel real-time selfie video stabilization method. Our method is completely automatic and runs at 26 fps. We use a 1D linear convolutional network to directly infer the rigid moving least squares warping which implicitly balances between the global rigidity and local flexibility. Our network structure is specifically designed to stabilize the background and foreground at the same time, while providing optional control of stabilization focus (relative importance of foreground vs. background) to the users. To train our network, we collect a selfie video dataset with 1005 videos, which is significantly larger than previous selfie video datasets. We also propose a grid approximation to the rigid moving least squares that enables the real-time frame warping. Our method is fully automatic and produces visually and quantitatively better results than previous real-time general video stabilization methods. Compared to previous offline selfie video methods, our approach produces comparable quality with a speed improvement of orders of magnitude. Our code and selfie video dataset is available at https://github.com/jiy173/selfievideostabilization. Jiyang Yu, Ravi Ramamoorthi, Ke-Li Cheng, Michel Sarkis, Ning Bi |
CVPR | 1 |
| 2021 | Selfie Video StabilizationabstractWe propose a novel algorithm for stabilizing selfie videos. Our goal is to automatically generate stabilized video that has optimal smooth motion in the sense of both foreground and background. The key insight is that non-rigid foreground motion in selfie videos can be analyzed using a 3D face model, and background motion can be analyzed using optical flow. We use second derivative of temporal trajectory of selected pixels as the measure of smoothness. Our algorithm stabilizes selfie videos by minimizing the smoothness measure of the background, regularized by the motion of the foreground. Experiments show that our method outperforms state-of-the-art general video stabilization techniques in selfie videos. Jiyang Yu, Ravi Ramamoorthi |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2020 | Learning Video Stabilization Using Optical FlowabstractWe propose a novel neural network that infers the per-pixel warp fields for video stabilization from the optical flow fields of the input video. While previous learning based video stabilization methods attempt to implicitly learn frame motions from color videos, our method resorts to optical flow for motion analysis and directly learns the stabilization using the optical flow. We also propose a pipeline that uses optical flow principal components for motion inpainting and warp field smoothing, making our method robust to moving objects, occlusion and optical flow inaccuracy, which is challenging for other video stabilization methods. Our method achieves quantitatively and visually better results than the state-of-the-art optimization based and deep learning based video stabilization methods. Our method also gives a ~3x speed improvement compared to the optimization based methods. Jiyang Yu, Ravi Ramamoorthi |
CVPR | 1 |
| 2020 | A Target Detection Algorithm of Neural Network Based on Histogram StatisticsabstractAiming at the problems of poor adaptability of traditional target detection algorithms and high computational resources of deep learning algorithms, a BP neural network target detection algorithm based on histogram statistics is proposed. It is based on the principle that similar areas have similar histograms. In this algorithm, the two-dimensional image information converts to the one-dimensional histogram information. We establish a three-layer neural network model, and the histogram is used as the input of the BP neural network. Compared to the traditional target detection algorithms, its complexity is low, and its efficiency and accuracy is high. The experimental results show that the fewer classification categories, the higher target detection probability. The computational complexity of the BP neural network is low, so the computational efficiency is quite high. The accuracy of target recognition is higher than 97% with SAR and optical images. Yalong Pang, Luyuan Wang, Jiyang Yu, Bowen Cheng, Zongling Li |
IGARSS | 4 |
| 2019 | Robust Video Stabilization by Optimization in CNN Weight SpaceabstractWe propose a novel robust video stabilization method. Unlike traditional video stabilization techniques that involve complex motion models, we directly model the appearance change of the frames as the dense optical flow field of consecutive frames. We introduce a new formulation of the video stabilization task based on first principles, which leads to a large scale non-convex problem. This problem is hard to solve, so previous optical flow based approaches have resorted to heuristics. In this paper, we propose a novel optimization routine that transfers this problem into the convolutional neural network parameter domain. While we exploit the general benefits of CNNs, including standard gradient-based optimization techniques, our method is a new approach to using CNNs purely as an optimizer rather than learning from data.Our method trains the CNN from scratch on each specific input example, and intentionally overfits the CNN parameters to produce the best result on the input example. By solving the problem in the CNN weight space rather than directly for image pixels, we make it a viable formulation for video stabilization. Our method produces both visually and quantitatively better results than previous work, and is robust in situations acknowledged as limitations in current state-of-the-art methods. Jiyang Yu, Ravi Ramamoorthi |
CVPR | 1 |
| 2019 | SJARACNe: a scalable software tool for gene network reverse engineering from big dataabstractSUMMARY: Over the last two decades, we have observed an exponential increase in the number of generated array or sequencing-based transcriptomic profiles. Reverse engineering of biological networks from high-throughput gene expression profiles has been one of the grand challenges in systems biology. The Algorithm for the Reconstruction of Accurate Cellular Networks (ARACNe) represents one of the most effective and widely-used tools to address this challenge. However, existing ARACNe implementations do not efficiently process big input data with thousands of samples. Here we present an improved implementation of the algorithm, SJARACNe, to solve this big data problem, based on sophisticated software engineering. The new scalable SJARACNe package achieves a dramatic improvement in computational performance in both time and memory usage and implements new features while preserving the network inference accuracy of the original algorithm. Given that large-sampled transcriptomic data is increasingly available and ARACNe is extremely demanding for network reconstruction, the scalable SJARACNe will allow even researchers with modest computational resources to efficiently construct complex regulatory and signaling networks from thousands of gene expression profiles. AVAILABILITY AND IMPLEMENTATION: SJARACNe is implemented in C++ (computational core) and Python (pipelining scripting wrapper, ≥3.6.1). It is freely available at https://github.com/jyyulab/SJARACNe. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Alireza Khatamian, Evan O. Paull, Andrea Califano, Jiyang Yu |
Bioinform. | 4 |
| 2018 | Selfie Video Stabilization
Jiyang Yu, Ravi Ramamoorthi |
ECCV (5) | 1 |
| 2016 | ScreenBEAM: a novel meta-analysis algorithm for functional genomics screens via Bayesian hierarchical modelingabstractMOTIVATION: Functional genomics (FG) screens, using RNAi or CRISPR technology, have become a standard tool for systematic, genome-wide loss-of-function studies for therapeutic target discovery. As in many large-scale assays, however, off-target effects, variable reagents' potency and experimental noise must be accounted for appropriately control for false positives. Indeed, rigorous statistical analysis of high-throughput FG screening data remains challenging, particularly when integrative analyses are used to combine multiple sh/sgRNAs targeting the same gene in the library. METHOD: We use large RNAi and CRISPR repositories that are publicly available to evaluate a novel meta-analysis approach for FG screens via Bayesian hierarchical modeling, Screening Bayesian Evaluation and Analysis Method (ScreenBEAM). RESULTS: Results from our analysis show that the proposed strategy, which seamlessly combines all available data, robustly outperforms classical algorithms developed for microarray data sets as well as recent approaches designed for next generation sequencing technologies. Remarkably, the ScreenBEAM algorithm works well even when the quality of FG screens is relatively low, which accounts for about 80-95% of the public datasets. AVAILABILITY AND IMPLEMENTATION: R package and source code are available at: https://github.com/jyyu/ScreenBEAM. CONTACT: [email protected], [email protected], [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jiyang Yu, Andrea Califano |
Bioinform. | 1 |
| 2016 | Thread-Aware Adaptive Prefetcher on Multicore Systems: Improving the Performance for Multithreaded WorkloadsabstractMost processors employ hardware data prefetching techniques to hide memory access latencies. However, the prefetching requests from different threads on a multicore processor can cause severe interference with prefetching and/or demand requests of others. The data prefetching can lead to significant performance degradation due to shared resource contention on shared memory multicore systems. This article proposes a thread-aware data prefetching mechanism based on low-overhead runtime information to tune prefetching modes and aggressiveness, mitigating the resource contention in the memory system. Our solution has three new components: (1) a self-tuning prefetcher that uses runtime feedback to dynamically adjust data prefetching modes and arguments of each thread, (2) a filtering mechanism that informs the hardware about which prefetching request can cause shared data invalidation and should be discarded, and (3) a limiter thread acceleration mechanism to estimate and accelerate the critical thread which has the longest completion time in the parallel region of execution. On a set of multithreaded parallel benchmarks, our thread-aware data prefetching mechanism improves the overall performance of 64-core system by 13% over a multimode prefetch baseline system with two-level cache organization and conventional modified, exclusive, shared, and invalid-based directory coherence protocol. We compare our approach with the feedback directed prefetching technique and find that it provides 9% performance improvement on multicore systems, while saving the memory bandwidth consumption. Peng Liu 0016, Jiyang Yu, Michael C. Huang 0001 |
ACM Trans. Archit. Code Optim. | 2 |
| 2016 | Minimal BRDF sampling for two-shot near-field reflectance acquisitionabstractWe develop a method to acquire the BRDF of a homogeneous flat sample from only two images, taken by a near-field perspective camera, and lit by a directional light source. Our method uses the MERL BRDF database to determine the optimal set of lightview pairs for data-driven reflectance acquisition. We develop a mathematical framework to estimate error from a given set of measurements, including the use of multiple measurements in an image simultaneously, as needed for acquisition from near-field setups. The novel error metric is essential in the near-field case, where we show that using the condition-number alone performs poorly. We demonstrate practical near-field acquisition of BRDFs from only one or two input images. Our framework generalizes to configurations like a fixed camera setup, where we also develop a simple extension to spatially-varying BRDFs by clustering the materials. Zexiang Xu, Jannik Boll Nielsen, Jiyang Yu, Henrik Wann Jensen, Ravi Ramamoorthi |
ACM Trans. Graph. | 3 |
| 2014 | A Thread-Aware Adaptive Data PrefetcherabstractMost processors employ hardware data prefetching to hide memory access latencies. However the prefetching requests from different threads on a multi-core processor can cause severe interference with prefetching and/or demand requests of others. The data prefetching can lead to significant performance degradation due to shared resource contention on shared memory multi-core systems. This paper proposes a thread-aware data prefetching mechanism based on low-overhead run-time information to tune prefetching modes and aggressiveness, mitigating the resource contention in the memory system. Our solution has two new components: 1) a filtering mechanism that informs the hardware about which prefetching requests can cause shared data invalidation and should be discarded, and 2) a self-tuning prefetcher that uses run-time feedback to adjust each thread's data prefetching mode and arguments. On a set of parallel benchmarks, our thread-aware data prefetching mechanisms improve the overall performance of 64-core system by 11% and reduce the energy-delay product by 13% over a multi-mode prefetch baseline system with a two level cache organization and a conventional MESI-based directory coherence protocol. We compare our approach to the feedback directed prefetching (FDP) technique and find that it provides better performance on multi-core systems, while reducing the energy delay product. Jiyang Yu, Peng Liu 0016 |
ICCD | 1 |
| 2013 | An improved constant coefficient multiplication algorithm based on cascaded adder graph
He Chen 0004, Xiujie Qu, Long Pang, Jiyang Yu |
Sci. China Inf. Sci. | 4 |
| 2013 | A novel conflict-free parallel memory access scheme for FFT constant geometry architectures
Cuimei Ma, He Chen 0004, Jiyang Yu |
Sci. China Inf. Sci. | 3 |
| 2013 | Improved Goldschmidt division method using mapping of divisors
Wen Yan 0007, Xiujie Qu, He Chen 0004, Jiyang Yu |
Sci. China Inf. Sci. | 4 |
| 2013 | An efficient protocol with synchronization accelerator for multi-processor embedded systems
Jiyang Yu, Peng Liu 0016, Chunming Huang, Yingtao Jiang, Qingdong Yao |
Parallel Comput. | 1 |