EDBT 2026 Demo / reviewers in the wild / expert
Yijin Li
dblp:178/6879
· DBLP profile ↗
25ranked-venue papers
6as first author
22since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 5 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 3 first-author · 13 since 2021Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | From Threads to Tiles: T2T, a Compiler for CUDA-to-NPU Translation via 2D VectorizationabstractCUDA’s programming model, exposing massive parallelism via fine-grained scalar threads, has become the de facto standard for GPU computing. Concurrently, NPUs are emerging as highly efficient accelerators, but their architecture is fundamentally different, relying on coarse-grained, explicit 2-D tile-based instructions. This creates a critical challenge: bridging the semantic gap "From Threads to Tiles". A direct translation is infeasible, as it requires lifting the implicit parallelism of CUDA’s scalar model into the explicit, multi-dimensional vector space of NPUs, a problem we formalize as a lifting challenge.This paper introduces T2T, a compiler framework that automates this "Threads to Tiles" translation via the 2-D Vectorization technique. T2T first transforms a CUDA kernel’s implicit SIMT parallelism into a structured, explicit loop nest via our Unified Parallelism Abstraction (UPA), making the parallelism analyzable. From this representation, T2T’s core vectorization engine systematically selects optimal pairs of loops and maps them onto the NPU’s 2-D tile instructions to maximize hardware utilization. To ensure correctness and handle performance-critical CUDA features, a final set of semantics-preserving optimizations is applied, including efficient control-flow management and vectorization of warp-level intrinsics.We implement T2T based on Polygeist and evaluate representative NPU architectures. On a diverse set of benchmarks, kernels translated by T2T achieve up to 73% of native CUDA performance on an A100 GPU and outperform baseline translation approaches by up to 6.9×. Our work demonstrates that a systematic, compiler-driven approach to 2-D vectorization is a principled and high-performance path for porting the rich CUDA ecosystem to the evolving landscape of NPU accelerators. Shuaijiang Li, Ying Liu 0055, Shuoming Zhang, Yijin Li, Yangyu Zhang, Runyu Zhou, Xiyu Shi, Chunwei Xia, Yuan Wen, Xiaobing Feng 0002, Huimin Cui |
CGO | 6 |
| 2026 | SYCL-MLU: unifying SIMT and SIMD in heterogeneous programming
Runyu Zhou, Yijin Li, En Shao, Ziyan Xie, Huimin Cui |
CCF Trans. High Perform. Comput. | 2 |
| 2026 | Dual-stream attention-guided learning for weakly supervised whole slide image classification
Daoxi Cao, Hangbei Cheng, Yijin Li, Ruolin Zhou, Xuehan Zhang, Binwei Li, Xuancheng Gu, Xueyu Liu, Yongfei Wu |
Eng. Appl. Artif. Intell. | 3 |
| 2025 | OpticFusion: Multi-Modal Neural Implicit 3D Reconstruction of Microstructures by Fusing White Light Interferometry and Optical MicroscopyabstractWhite Light Interferometry (WLI) is a precise optical tool for measuring the 3D topography of microstructures. However, conventional WLI cannot capture the natural color of a sample's surface, which is essential for many microscale research applications that require both 3D geometry and color information. Previous methods have attempted to overcome this limitation by modifying WLI hardware and analysis software, but these solutions are often costly. In this work, we address this challenge from a computer vision multi-modal reconstruction perspective for the first time. We introduce OpticFusion, a novel approach that uses an additional digital optical microscope (OM) to achieve 3D reconstruction with natural color textures using multi-view WLI and OM images. Our method employs a two-step data association process to obtain the poses of WLI and OM data. By leveraging the neural implicit representation, we fuse multi-modal data and apply color decomposition technology to extract the sample's natural color. Tested on our multi-modal dataset of various microscale samples, Optic-Fusion achieves detailed 3D reconstructions with color textures. Our method provides an effective tool for practical applications across numerous microscale research fields. The source code and our real-world dataset are available at https://github.conz/zju3dv/OpticFusion. Yijin Li, Guofeng Zhang 0001 |
3DV | 2 |
| 2025 | GS-DiT: Advancing Video Generation with Dynamic 3D Gaussian Fields through Efficient Dense 3D Point Trackingabstract4D video control is essential in video generation as it enables the use of sophisticated lens techniques, such as multicamera shooting and dolly zoom, which are currently unsupported by existing methods. Training a video Diffusion Transformer (DiT) directly to control 4D content requires expensive multi-view videos. Inspired by Monocular Dynamic novel View Synthesis (MDVS) that optimizes a 4D representation and renders videos according to different 4D elements, such as camera pose and object motion editing, we bring dynamic 3D Gaussian fields to video generation. Specifically, we propose a novel framework that constructs dynamic 3D Gaussian fields with dense 3D point tracking and renders the Gaussian field for all video frames. Then we finetune a pretrained DiT to generate videos following the guidance of the rendered video, dubbed as GS-DiT. To boost the training of the GS-DiT, we also propose an efficient Dense 3D Point Tracking (D3D-PT) method for the dynamic 3D Gaussian field construction. Our D3D-PT outperforms SpatialTracker, the state-of-the-art sparse 3D point tracking method, in accuracy and accelerates the inference speed by two orders of magnitude. During the inference stage, GS-DiT can generate videos with the same dynamic content while adhering to different camera parameters, addressing a significant limitation of current video generation models. GS-DiT demonstrates strong generalization capabilities and extends the 4D controllability of Gaussian splatting to video generation beyond just camera poses. It supports advanced cinematic effects through the manipulation of the Gaussian field and camera intrinsics, making it a powerful tool for creative video production. Demos are available at https://wkbian.github.io/Projects/GS-DiT/. Weikang Bian, Xiaoyu Shi 0002, Yijin Li, Fu-Yun Wang, Hongsheng Li 0001 |
CVPR | 4 |
| 2025 | BlinkTrack: Feature Tracking Over 80 FPS via Events and ImagesabstractEvent cameras, known for their high temporal resolution and ability to capture asynchronous changes, have gained significant attention for their potential in feature tracking, especially in challenging conditions. However, event cameras lack the fine-grained texture information that conventional cameras provide, leading to error accumulation in tracking. To address this, we propose a novel framework, BlinkTrack, which integrates event data with grayscale images for high-frequency feature tracking. Our method extends the traditional Kalman filter into a learning-based framework, utilizing differentiable Kalman filters in both event and image branches. This approach improves single-modality tracking and effectively solves the data association and fusion from asynchronous event and image data. We also introduce new synthetic and augmented datasets to better evaluate our model. Experimental results indicate that BlinkTrack significantly outperforms existing methods, exceeding 80 FPS with multi-modality data and 100 FPS with preprocessed event data. Codes and dataset are available at https://github.com/ColieShen/BlinkTrack. Yichen Shen 0004, Yijin Li, Guanglin Li 0005, Hujun Bao, Zhaopeng Cui, Guofeng Zhang 0001 |
ICCV | 2 |
| 2025 | ETO+: Revisit the Refinement Stage in Efficient Feature MatchingabstractRecent feature matching approaches like ETO have focused on developing lightweight matching algorithms for real-time applications. However, their lack of cross-image feature interaction and sufficient refinement often lead to a decline in matching accuracy. To address these challenges, we propose ETO+, a novel and accurate feature matching algorithm that incorporates a lightweight yet efficient bidirectional interaction module and multi-stage refinement. Specifically, we introduce Trans-CNN, a bidirectional feature interaction module that integrates CNN- and transformer-based techniques to enhance both intra-image feature refinement and inter-image feature fusion, all while maintaining a comparable computational cost. Furthermore, by leveraging the inherent sparsity of local feature matching, we propose an efficient strategy to adaptively reallocate computational resources within the network. Additionally, we design an adaptive loss function that mitigates the impact of large matching errors, thereby improving overall robustness. Extensive experiments on widely used datasets demonstrate that our approach achieves a strong balance between accuracy and computational efficiency. It outperforms ETO by 7.9 in AUC@5 on MegaDepth, respectively, while being about 40% faster than E-LoFTR. Junjie Ni, Yichen Shen 0004, Yijin Li, Hongjia Zhai, Hujun Bao, Guofeng Zhang 0001 |
IROS | 3 |
| 2025 | CoProSketch: Controllable and Progressive Sketch Generation with Diffusion Model
Ruohao Zhan, Yijin Li, Yisheng He, Yichen Shen 0004, Zilong Dong, Guofeng Zhang 0001 |
ACM Multimedia | 2 |
| 2025 | Guided Attention Mechanism in Multi-turn Dialogue Summarization
Yijin Li, Guohua Zhu, Zicai Xia, Yupei Zheng |
PRICAI (4) | 1 |
| 2024 | BlinkVision: A Benchmark for Optical Flow, Scene Flow and Point Tracking Estimation Using RGB Frames and Events
Yijin Li, Yichen Shen 0004, Weikang Bian, Xiaoyu Shi 0002, Fu-Yun Wang, Keqiang Sun, Hujun Bao, Zhaopeng Cui, Guofeng Zhang 0001, Hongsheng Li 0001 |
ECCV (67) | 1 |
| 2024 | ZoLA: Zero-Shot Creative Long Animation Generation with Short Video Model
Fu-Yun Wang, Guanglu Song, Weikang Bian, Yijin Li, Yu Liu 0015, Hongsheng Li 0001 |
ECCV (45) | 7 |
| 2024 | A Global Depth-Range-Free Multi-View Stereo Transformer Network with Pose EmbeddingabstractIn this paper, we propose a novel multi-view stereo (MVS) framework that gets rid of the depth range prior. Unlike recent prior-free MVS methods that work in a pair-wise manner, our method simultaneously considers all the source images. Specifically, we introduce a Multi-view Disparity Attention (MDA) module to aggregate long-range context information within and across multi-view images. Considering the asymmetry of the epipolar disparity flow, the key to our method lies in accurately modeling multi-view geometric constraints. We integrate pose embedding to encapsulate information such as multi-view camera poses, providing implicit geometric constraints for multi-view disparity feature fusion dominated by attention. Additionally, we construct corresponding hidden states for each source image due to significant differences in the observation quality of the same pixel in the reference frame across multiple source frames. We explicitly estimate the quality of the current pixel corresponding to sampled points on the epipolar line of the source image and dynamically update hidden states through the uncertainty estimation module. Extensive results on the DTU dataset and Tanks\&Temple benchmark demonstrate the effectiveness of our method. Yitong Dong, Yijin Li, Weikang Bian, Hujun Bao, Zhaopeng Cui, Hongsheng Li 0001, Guofeng Zhang 0001 |
NeurIPS | 2 |
| 2024 | ETO: Efficient Transformer-based Local Feature Matching by Organizing Multiple Homography HypothesesabstractWe tackle the efficiency problem of learning local feature matching.Recent advancements have given rise to purely CNN-based and transformer-based approaches, each augmented with deep learning techniques. While CNN-based methods often excel in matching speed, transformer-based methods tend to provide more accurate matches. We propose an efficient transformer-based network architecture for local feature matching.This technique is built on constructing multiple homography hypotheses to approximate the continuous correspondence in the real world and uni-directional cross-attention to accelerate the refinement. On the YFCC100M dataset, our matching accuracy is competitive with LoFTR, a state-of-the-art transformer-based architecture, while the inference speed is boosted to 4 times, even outperforming the CNN-based methods.Comprehensive evaluations on other open datasets such as Megadepth, ScanNet, and HPatches demonstrate our method's efficacy, highlighting its potential to significantly enhance a wide array of downstream applications. Junjie Ni, Guofeng Zhang 0001, Guanglin Li 0005, Yijin Li, Hujun Bao |
NeurIPS | 4 |
| 2023 | PATS: Patch Area Transportation with Subdivision for Local Feature MatchingabstractLocal feature matching aims at establishing sparse correspondences between a pair of images. Recently, detector-free methods present generally better performance but are not satisfactory in image pairs with large scale differences. In this paper, we propose Patch Area Transportation with Subdivision (PATS) to tackle this issue. Instead of building an expensive image pyramid, we start by splitting the original image pair into equal-sized patches and gradually resizing and subdividing them into smaller patches with the same scale. However, estimating scale differences between these patches is non-trivial since the scale differences are determined by both relative camera poses and scene structures, and thus spatially varying over image pairs. Moreover, it is hard to obtain the ground truth for real scenes. To this end, we propose patch area transportation, which enables learning scale differences in a self-supervised manner. In contrast to bipartite graph matching, which only handles one-to-one matching, our patch area transportation can deal with many-to-many relationships. PATS improves both matching accuracy and coverage, and shows superior performance in downstream tasks, such as relative pose estimation, visual localization, and optical flow estimation. The source code is available at https://zju3dv.github.io/pats/. Junjie Ni, Yijin Li, Hongsheng Li 0001, Hujun Bao, Zhaopeng Cui, Guofeng Zhang 0001 |
CVPR | 2 |
| 2023 | Multi-Modal Neural Radiance Field for Monocular Dense SLAM with a Light-Weight ToF SensorabstractLight-weight time-of-flight (ToF) depth sensors are compact and cost-efficient, and thus widely used on mobile devices for tasks such as autofocus and obstacle detection. However, due to the sparse and noisy depth measurements, these sensors have rarely been considered for dense geometry reconstruction. In this work, we present the first dense SLAM system with a monocular camera and a light-weight ToF sensor. Specifically, we propose a multi-modal implicit scene representation that supports rendering both the signals from the RGB camera and light-weight ToF sensor which drives the optimization by comparing with the raw sensor inputs. Moreover, in order to guarantee successful pose tracking and reconstruction, we exploit a predicted depth as an intermediate supervision and develop a coarse-to-fine optimization strategy for efficient learning of the implicit representation. At last, the temporal information is explicitly exploited to deal with the noisy signals from light-weight ToF sensors to improve the accuracy and robustness of the system. Experiments demonstrate that our system well exploits the signals of light-weight ToF sensors and achieves competitive results both on camera tracking and dense scene reconstruction. Project page: https://zju3dv.github.io/tof_slam/. Yijin Li, Yanbin Teng, Hujun Bao, Guofeng Zhang 0001, Yinda Zhang 0001, Zhaopeng Cui |
ICCV | 2 |
| 2023 | BlinkFlow: A Dataset to Push the Limits of Event-Based Optical Flow EstimationabstractEvent cameras provide high temporal precision, low data rates, and high dynamic range visual perception, which are well-suited for optical flow estimation. While data-driven optical flow estimation has obtained great success in RGB cameras, its generalization performance is seriously hindered in event cameras mainly due to the limited and biased training data. In this paper, we present a novel simulator, BlinkSim, for the fast generation of large-scale data for event-based optical flow. BlinkSim incorporates a configurable rendering engine alongside an event simulation suite. By leveraging the wealth of current 3D assets, the rendering engine enables us to automatically build up thousands of scenes with different objects, textures, and motion patterns and render very high-frequency images for realistic event data simulation. Based on BlinkSim, we construct a large training dataset and evaluation benchmark BlinkFlow that contains sufficient, diversiform, and challenging event data with optical flow ground truth. Experiments show that BlinkFlow improves the generalization performance of state-of-the-art methods by more than 40% on average and up to 90%. Moreover, we further propose an Event-based optical Flow transFormer (E-FlowFormer) architecture. Powered by our BlinkFlow, E-FlowFormer outperforms the SOTA methods by up to 91% on the MVSEC dataset and 14% on the DSEC dataset and presents the best generalization performance. The source code and data are available at https://zju3dv.github.io/blinkflow/. Yijin Li, Xiaoyu Shi 0002, Hongsheng Li 0001, Hujun Bao, Zhaopeng Cui, Guofeng Zhang 0001 |
IROS | 1 |
| 2023 | Context-PIPs: Persistent Independent Particles Demands Context Features
Weikang Bian, Xiaoyu Shi 0002, Yitong Dong, Yijin Li, Hongsheng Li 0001 |
NeurIPS | 5 |
| 2022 | DELTAR: Depth Estimation from a Light-Weight ToF Sensor and RGB Image
Yijin Li, Wenqi Dong, Hujun Bao, Guofeng Zhang 0001, Yinda Zhang 0001, Zhaopeng Cui |
ECCV (1) | 1 |
| 2022 | Neural rendering in a room: amodal 3D understanding and free-viewpoint rendering for the closed scene composed of pre-captured objectsabstractWe, as human beings, can understand and picture a familiar scene from arbitrary viewpoints given a single image, whereas this is still a grand challenge for computers. We hereby present a novel solution to mimic such human perception capability based on a new paradigm of amodal 3D scene understanding with neural rendering for a closed scene. Specifically, we first learn the prior knowledge of the objects in a closed scene via an offline stage, which facilitates an online stage to understand the room with unseen furniture arrangement. During the online stage, given a panoramic image of the scene in different layouts, we utilize a holistic neural-rendering-based optimization framework to efficiently estimate the correct 3D scene layout and deliver realistic free-viewpoint rendering. In order to handle the domain gap between the offline and online stage, our method exploits compositional neural rendering techniques for data augmentation in the offline training. The experiments on both synthetic and real datasets demonstrate that our two-stage design achieves robust 3D scene understanding and outperforms competing methods by a large margin, and we also show that our realistic free-viewpoint rendering enables various applications, including scene touring and editing. Code and data are available on the project webpage: https://zju3dv.github.io/nr_in_a_room/. Bangbang Yang, Yinda Zhang 0001, Yijin Li, Zhaopeng Cui, Sean Ryan Fanello, Hujun Bao, Guofeng Zhang 0001 |
ACM Trans. Graph. | 3 |
| 2021 | VS-Net: Voting With Segmentation for Visual LocalizationabstractVisual localization is of great importance in robotics and computer vision. Recently, scene coordinate regression based methods have shown good performance in visual localization in small static scenes. However, it still estimates camera poses from many inferior scene coordinates. To address this problem, we propose a novel visual localization framework that establishes 2D-to-3D correspondences between the query image and the 3D map with a series of learnable scene-specific landmarks. In the landmark generation stage, the 3D surfaces of the target scene are over-segmented into mosaic patches whose centers are regarded as the scene-specific landmarks. To robustly and accurately recover the scene-specific landmarks, we propose the Voting with Segmentation Network (VS-Net) to segment the pixels into different landmark patches with a segmentation branch and estimate the landmark locations within each patch with a landmark location voting branch. Since the number of landmarks in a scene may reach up to 5000, training a segmentation network with such a large number of classes is both computation and memory costly for the commonly used cross-entropy loss. We propose a novel prototype-based triplet loss with hard negative mining, which is able to train semantic segmentation networks with a large number of labels efficiently. Our proposed VS-Net is extensively tested on multiple public benchmarks and can outperform stateof-the-art visual localization methods. Code and models are available at https://github.com/zju3dv/VS-Net. Yijin Li, Bangbang Yang, Xiaowei Zhou 0001, Hujun Bao, Guofeng Zhang 0001, Hongsheng Li 0001 |
CVPR | 3 |
| 2021 | Graph-based Asynchronous Event Processing for Rapid Object RecognitionabstractDifferent from traditional video cameras, event cam- eras capture asynchronous events stream in which each event encodes pixel location, trigger time, and the polarity of the brightness changes. In this paper, we introduce a novel graph-based framework for event cameras, namely SlideGCN. Unlike some recent graph-based methods that use groups of events as input, our approach can efficiently process data event-by-event, unlock the low latency nature of events data while still maintaining the graph’s structure internally. For fast graph construction, we develop a radius search algorithm, which better exploits the partial regular structure of event cloud against k-d tree based generic methods. Experiments show that our method reduces the computational complexity up to 100 times with respect to current graph-based methods while keeping state-of-the-art performance on object recognition. Moreover, we verify the superiority of event-wise processing with our method. When the state becomes stable, we can give a prediction with high confidence, thus making an early recognition. Yijin Li, Bangbang Yang, Zhaopeng Cui, Hujun Bao, Guofeng Zhang 0001 |
ICCV | 1 |
| 2021 | Learning Object-Compositional Neural Radiance Field for Editable Scene RenderingabstractImplicit neural rendering techniques have shown promising results for novel view synthesis. However, existing methods usually encode the entire scene as a whole, which is generally not aware of the object identity and limits the ability to the high-level editing tasks such as moving or adding furniture. In this paper, we present a novel neural scene rendering system, which learns an object-compositional neural radiance field and produces realistic rendering with editing capability for a clustered and real-world scene. Specifically, we design a novel two-pathway architecture, in which the scene branch encodes the scene geometry and appearance, and the object branch encodes each standalone object conditioned on learnable object activation codes. To survive the training in heavily cluttered scenes, we propose a scene-guided training strategy to solve the 3D space ambiguity in the occluded regions and learn sharp boundaries for each object. Extensive experiments demonstrate that our system not only achieves competitive performance for static scene novel-view synthesis, but also produces realistic rendering for object-level editing. Bangbang Yang, Yinda Zhang 0001, Yinghao Xu 0001, Yijin Li, Hujun Bao, Guofeng Zhang 0001, Zhaopeng Cui |
ICCV | 4 |
| 2020 | NIID-Net: Adapting Surface Normal Knowledge for Intrinsic Image Decomposition in Indoor ScenesabstractIntrinsic image decomposition, i.e., decomposing a natural image into a reflectance image and a shading image, is used in many augmented reality applications for achieving better visual coherence between virtual contents and real scenes. The main challenge is that the decomposition is ill-posed, especially in indoor scenes where lighting conditions are complicated, while real training data is inadequate. To solve this challenge, we propose NIID-Net, a novel learning-based framework that adapts surface normal knowledge for improving the decomposition. The knowledge learned from relatively more abundant data for surface normal estimation is integrated into intrinsic image decomposition in two novel ways. First, normal feature adapters are proposed to incorporate scene geometry features when decomposing the image. Secondly, a map of integrated lighting is proposed for propagating object contour and planarity information during shading rendering. Furthermore, this map is capable of representing spatially-varying lighting conditions indoors. Experiments show that NIID-Net achieves competitive performance in reflectance estimation and outperforms all previous methods in shading estimation quantitatively and qualitatively. The source code of our implementation is released at https://github.com/zju3dv/NIID-Net. Jundan Luo, Yijin Li, Xiaowei Zhou 0001, Guofeng Zhang 0001, Hujun Bao |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2017 | Smooth path planning for autonomous parking systemabstractIn this paper, we present a path planning algorithm for autonomous parking system. We focus on the kinematics of the car-like vehicle and improve the conventional geometric parking algorithm by proposing a new curve element named linearly steering spiral. A path planning algorithm based on smooth path searching and optimizing are presented. This method can generate smooth paths incrementally, and the reference control signals can be deduced directly once the path is determined. A simple closed-loop controller is designed in order to deal with the uncertainty from various aspects. Simulations are implemented in different scenarios including obstacle-free and cluttered environments, moreover, a real-world online experiment is executed. The results indicate that the proposed method achieves good performance on both accuracy and computational cost. Yi Yang 0009, Lu Zhang 0047, Xin Qu, Jinzhou Lei, Yijin Li, Jianhang Wang |
Intelligent Vehicles Symposium | 5 |
| 2016 | Similarity Recoverable, Format-Preserving String Encryption
Yijin Li, Wendy Hui Wang |
APWeb (2) | 1 |