VLDB 2026 Research / reviewers in the wild / expert
Weikang Bian
dblp:252/4248
· DBLP profile ↗
10ranked-venue papers
4as first author
8since 2021 · last 2025
0000-0001-9986-3348ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 5 since 2021Security and privacy · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
8 papers |
3D vision · 52% Generative modeling · 29% Deep learning architectures and training · 10% | |
| Computer graphics and multimedia
3 papers |
Visual content generation and editing · 54% Computer animation and physical simulation · 23% Geometric modeling and processing · 18% | |
| Network and information security
2 papers |
Web and mobile security · 50% Malware analysis · 50% |
Topics — the 27 heaviest of 29, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › 3D vision › motion estimation
optical flow |
1.4 | 2 | 2024 | BlinkVision: A Benchmark for Optical Flow, Scene Flow and Point Tracking Estimation Using RGB Frames and Events · ECCV (67) 2024 VideoFlow: Exploiting Temporal Cues for Multi-frame Optical Flow Estimation · ICCV 2023 |
Machine learning › Generative modeling
video generation |
1.0 | 2 | 2024 | ZoLA: Zero-Shot Creative Long Animation Generation with Short Video Model · ECCV (45) 2024 Phased Consistency Models · NeurIPS 2024 |
Computer vision › 3D vision
novel view synthesis |
0.9 | 1 | 2025 | GS-DiT: Advancing Video Generation with Dynamic 3D Gaussian Fields through Efficient Dense 3D Point Tracking · CVPR 2025 |
Visual content generation and editing › video generation
4d video generation |
0.9 | 1 | 2025 | GS-DiT: Advancing Video Generation with Dynamic 3D Gaussian Fields through Efficient Dense 3D Point Tracking · CVPR 2025 |
Visual content generation and editing
video generation |
0.9 | 1 | 2025 | GS-DiT: Advancing Video Generation with Dynamic 3D Gaussian Fields through Efficient Dense 3D Point Tracking · CVPR 2025 |
Computer vision › 3D vision
motion estimation |
0.9 | 2 | 2023 | VideoFlow: Exploiting Temporal Cues for Multi-frame Optical Flow Estimation · ICCV 2023 Context-PIPs: Persistent Independent Particles Demands Context Features · NeurIPS 2023 |
Machine learning › Generative modeling › diffusion model
consistency model |
0.8 | 1 | 2024 | Phased Consistency Models · NeurIPS 2024 |
Computer vision › 3D vision
depth estimation |
0.8 | 1 | 2024 | A Global Depth-Range-Free Multi-View Stereo Transformer Network with Pose Embedding · NeurIPS 2024 |
Machine learning › Generative modeling
diffusion model |
0.8 | 1 | 2024 | Phased Consistency Models · NeurIPS 2024 |
Computer vision › Video understanding and tracking
feature tracking |
0.8 | 1 | 2024 | BlinkVision: A Benchmark for Optical Flow, Scene Flow and Point Tracking Estimation Using RGB Frames and Events · ECCV (67) 2024 |
Computer vision › 3D vision › 3d reconstruction
multi-view stereo |
0.8 | 1 | 2024 | A Global Depth-Range-Free Multi-View Stereo Transformer Network with Pose Embedding · NeurIPS 2024 |
Machine learning › Deep learning architectures and training › transformer
multi-view transformer |
0.8 | 1 | 2024 | A Global Depth-Range-Free Multi-View Stereo Transformer Network with Pose Embedding · NeurIPS 2024 |
Computer vision › 3D vision
pose estimation |
0.8 | 1 | 2024 | A Global Depth-Range-Free Multi-View Stereo Transformer Network with Pose Embedding · NeurIPS 2024 |
Computer vision › 3D vision
scene flow estimation |
0.8 | 1 | 2024 | BlinkVision: A Benchmark for Optical Flow, Scene Flow and Point Tracking Estimation Using RGB Frames and Events · ECCV (67) 2024 |
Machine learning › Generative modeling › diffusion model
text-to-image generation |
0.8 | 1 | 2024 | Phased Consistency Models · NeurIPS 2024 |
Machine learning › Deep learning architectures and training
transformer |
0.8 | 1 | 2024 | A Global Depth-Range-Free Multi-View Stereo Transformer Network with Pose Embedding · NeurIPS 2024 |
Machine learning › Generative modeling › video generation
zero-shot video generation |
0.8 | 1 | 2024 | ZoLA: Zero-Shot Creative Long Animation Generation with Short Video Model · ECCV (45) 2024 |
Computer animation and physical simulation
animation generation |
0.8 | 1 | 2024 | ZoLA: Zero-Shot Creative Long Animation Generation with Short Video Model · ECCV (45) 2024 |
Computer vision › 3D vision › motion estimation › optical flow
multi-frame optical flow |
0.7 | 1 | 2023 | VideoFlow: Exploiting Temporal Cues for Multi-frame Optical Flow Estimation · ICCV 2023 |
Computer vision › Video understanding and tracking › object tracking › probabilistic tracking
particle tracking |
0.7 | 1 | 2023 | Context-PIPs: Persistent Independent Particles Demands Context Features · NeurIPS 2023 |
Computer vision › 3D vision › feature matching
correspondence learning |
0.6 | 1 | 2022 | NeuralMarker: A Framework for Learning General Marker Correspondence · ACM Trans. Graph. 2022 |
Geometric modeling and processing
shape correspondence |
0.6 | 1 | 2022 | NeuralMarker: A Framework for Learning General Marker Correspondence · ACM Trans. Graph. 2022 |
Web and mobile security
web security |
0.4 | 1 | 2020 | MineThrottle: Defending against Wasm In-Browser Cryptojacking · WWW 2020 |
Malware analysis › malware detection
cryptojacking detection |
0.4 | 1 | 2019 | Poster: Detecting WebAssembly-based Cryptocurrency Mining · CCS 2019 |
Computer vision › 3D vision
event-based vision |
0.2 | 1 | 2024 | BlinkVision: A Benchmark for Optical Flow, Scene Flow and Point Tracking Estimation Using RGB Frames and Events · ECCV (67) 2024 |
Machine learning › Generative modeling › video generation
text-to-video generation |
0.2 | 1 | 2024 | Phased Consistency Models · NeurIPS 2024 |
Virtual and augmented reality › augmented reality
augmented reality applications |
0.2 | 1 | 2022 | NeuralMarker: A Framework for Learning General Marker Correspondence · ACM Trans. Graph. 2022 |
Methods — techniques the papers use, named apart from their topics
gaussian splatting · 1.7diffusion transformer · 1.7dense 3d point tracking · 1.7short video model · 1.5uncertainty estimation · 0.8pose embedding · 0.8phased consistency training · 0.8multi-view disparity attention · 0.8multi-step refinement · 0.8iterative flow refinement · 0.7symmetric epipolar distance loss · 0.6structure from motion · 0.6neural network · 0.6throttling · 0.4dynamic analysis · 0.4dynamic instruction execution trace analysis · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | GS-DiT: Advancing Video Generation with Dynamic 3D Gaussian Fields through Efficient Dense 3D Point Trackingabstract4D video control is essential in video generation as it enables the use of sophisticated lens techniques, such as multicamera shooting and dolly zoom, which are currently unsupported by existing methods. Training a video Diffusion Transformer (DiT) directly to control 4D content requires expensive multi-view videos. Inspired by Monocular Dynamic novel View Synthesis (MDVS) that optimizes a 4D representation and renders videos according to different 4D elements, such as camera pose and object motion editing, we bring dynamic 3D Gaussian fields to video generation. Specifically, we propose a novel framework that constructs dynamic 3D Gaussian fields with dense 3D point tracking and renders the Gaussian field for all video frames. Then we finetune a pretrained DiT to generate videos following the guidance of the rendered video, dubbed as GS-DiT. To boost the training of the GS-DiT, we also propose an efficient Dense 3D Point Tracking (D3D-PT) method for the dynamic 3D Gaussian field construction. Our D3D-PT outperforms SpatialTracker, the state-of-the-art sparse 3D point tracking method, in accuracy and accelerates the inference speed by two orders of magnitude. During the inference stage, GS-DiT can generate videos with the same dynamic content while adhering to different camera parameters, addressing a significant limitation of current video generation models. GS-DiT demonstrates strong generalization capabilities and extends the 4D controllability of Gaussian splatting to video generation beyond just camera poses. It supports advanced cinematic effects through the manipulation of the Gaussian field and camera intrinsics, making it a powerful tool for creative video production. Demos are available at https://wkbian.github.io/Projects/GS-DiT/. Weikang Bian, Xiaoyu Shi 0002, Yijin Li, Fu-Yun Wang, Hongsheng Li 0001 |
CVPR | 1 |
| 2024 | BlinkVision: A Benchmark for Optical Flow, Scene Flow and Point Tracking Estimation Using RGB Frames and Events
Yijin Li, Yichen Shen 0004, Weikang Bian, Xiaoyu Shi 0002, Fu-Yun Wang, Keqiang Sun, Hujun Bao, Zhaopeng Cui, Guofeng Zhang 0001, Hongsheng Li 0001 |
ECCV (67) | 5 |
| 2024 | ZoLA: Zero-Shot Creative Long Animation Generation with Short Video Model
Fu-Yun Wang, Guanglu Song, Weikang Bian, Yijin Li, Yu Liu 0015, Hongsheng Li 0001 |
ECCV (45) | 6 |
| 2024 | A Global Depth-Range-Free Multi-View Stereo Transformer Network with Pose EmbeddingabstractIn this paper, we propose a novel multi-view stereo (MVS) framework that gets rid of the depth range prior. Unlike recent prior-free MVS methods that work in a pair-wise manner, our method simultaneously considers all the source images. Specifically, we introduce a Multi-view Disparity Attention (MDA) module to aggregate long-range context information within and across multi-view images. Considering the asymmetry of the epipolar disparity flow, the key to our method lies in accurately modeling multi-view geometric constraints. We integrate pose embedding to encapsulate information such as multi-view camera poses, providing implicit geometric constraints for multi-view disparity feature fusion dominated by attention. Additionally, we construct corresponding hidden states for each source image due to significant differences in the observation quality of the same pixel in the reference frame across multiple source frames. We explicitly estimate the quality of the current pixel corresponding to sampled points on the epipolar line of the source image and dynamically update hidden states through the uncertainty estimation module. Extensive results on the DTU dataset and Tanks\&Temple benchmark demonstrate the effectiveness of our method. Yitong Dong, Yijin Li, Weikang Bian, Hujun Bao, Zhaopeng Cui, Hongsheng Li 0001, Guofeng Zhang 0001 |
NeurIPS | 4 |
| 2024 | Phased Consistency ModelsabstractConsistency Models (CMs) have made significant progress in accelerating the generation of diffusion models. However, their application to high-resolution, text-conditioned image generation in the latent space remains unsatisfactory. In this paper, we identify three key flaws in the current design of Latent Consistency Models~(LCMs). We investigate the reasons behind these limitations and propose Phased Consistency Models (PCMs), which generalize the design space and address the identified limitations. Our evaluations demonstrate that PCMs outperform LCMs across 1--16 step generation settings. While PCMs are specifically designed for multi-step refinement, they achieve comparable 1-step generation results to previously state-of-the-art specifically designed 1-step methods. Furthermore, we show the methodology of PCMs is versatile and applicable to video generation, enabling us to train the state-of-the-art few-step text-to-video generator. Our code is available at https://github.com/G-U-N/Phased-Consistency-Model. Fu-Yun Wang, Alexander William Bergman, Dazhong Shen, Peng Gao 0007, Michael Lingelbach, Keqiang Sun, Weikang Bian, Guanglu Song, Yu Liu 0015, Xiaogang Wang 0001, Hongsheng Li 0001 |
NeurIPS | 8 |
| 2023 | VideoFlow: Exploiting Temporal Cues for Multi-frame Optical Flow EstimationabstractWe introduce VideoFlow, a novel optical flow estimation framework for videos. In contrast to previous methods that learn to estimate optical flow from two frames, VideoFlow concurrently estimates bi-directional optical flows for multiple frames that are available in videos by sufficiently exploiting temporal cues.We first propose a TRi-frame Optical Flow (TROF) module that estimates bi-directional optical flows for the center frame in a three-frame manner. The information of the frame triplet is iteratively fused onto the center frame. To extend TROF for handling more frames, we further propose a MOtion Propagation (MOP) module that bridges multiple TROFs and propagates motion features between adjacent TROFs. With the iterative flow estimation refinement, the information fused in individual TROFs can be propagated into the whole sequence via MOP. By effectively exploiting video information, VideoFlow presents extraordinary performance, ranking 1st on all public benchmarks. On the Sintel benchmark, VideoFlow achieves 1.649 and 0.991 average end-point-error (AEPE) on the final and clean passes, a 15.1% and 7.6% error reduction from the best published results (1.943 and 1.073 from FlowFormer++). On the KITTI-2015 benchmark, VideoFlow achieves an F1-all error of 3.65%, a 19.2% error reduction from the best published result (4.52% from FlowFormer++). Code is released at https://github.com/XiaoyuShi97/VideoFlow. Xiaoyu Shi 0002, Weikang Bian, Dasong Li, Manyuan Zhang, Ka Chun Cheung, Simon See, Hongwei Qin, Jifeng Dai, Hongsheng Li 0001 |
ICCV | 3 |
| 2023 | Context-PIPs: Persistent Independent Particles Demands Context Features
Weikang Bian, Xiaoyu Shi 0002, Yitong Dong, Yijin Li, Hongsheng Li 0001 |
NeurIPS | 1 |
| 2022 | NeuralMarker: A Framework for Learning General Marker CorrespondenceabstractWe tackle the problem of estimating correspondences from a general marker, such as a movie poster, to an image that captures such a marker. Conventionally, this problem is addressed by fitting a homography model based on sparse feature matching. However, they are only able to handle plane-like markers and the sparse features do not sufficiently utilize appearance information. In this paper, we propose a novel framework NeuralMarker, training a neural network estimating dense marker correspondences under various challenging conditions, such as marker deformation, harsh lighting, etc. Deep learning has presented an excellent performance in correspondence learning once provided with sufficient training data. However, annotating pixel-wise dense correspondence for training marker correspondence is too expensive. We observe that the challenges of marker correspondence estimation come from two individual aspects: geometry variation and appearance variation. We, therefore, design two components addressing these two challenges in NeuralMarker. First, we create a synthetic dataset FlyingMarkers containing marker-image pairs with ground truth dense correspondences. By training with FlyingMarkers, the neural network is encouraged to capture various marker motions. Second, we propose the novel Symmetric Epipolar Distance (SED) loss, which enables learning dense correspondence from posed images. Learning with the SED loss and the cross-lighting posed images collected by Structure-from-Motion (SfM), NeuralMarker is remarkably robust in harsh lighting environments and avoids synthetic image bias. Besides, we also propose a novel marker correspondence evaluation method circumstancing annotations on real marker-image pairs and create a new benchmark. We show that NeuralMarker significantly outperforms previous methods and enables new interesting applications, including Augmented Reality (AR) and video editing. Xiaokun Pan, Weihong Pan, Weikang Bian, Ka Chun Cheung, Guofeng Zhang 0001, Hongsheng Li 0001 |
ACM Trans. Graph. | 4 |
| 2020 | MineThrottle: Defending against Wasm In-Browser CryptojackingabstractIn-browser cryptojacking is an urgent threat to web users, where an attacker abuses the users’ computing resources without obtaining their consent. In-browser mining programs are usually developed in WebAssembly (Wasm) for its great performance. Several prior works have measured cryptojacking in the wild and proposed detection methods using static features and dynamic features. However, there exists no good defense mechanism within the user’s browser to stop the malicious drive-by mining behavior. Weikang Bian, Wei Meng 0001, Mingxue Zhang 0001 |
WWW | 1 |
| 2019 | Poster: Detecting WebAssembly-based Cryptocurrency MiningabstractIn-browser cryptojacking is an emerging threat to web users. The attackers can abuse the users' computation resources to perform cryptocurrency mining without obtaining their consent. Moreover, the new web feature -WebAssembly (Wasm)- enables efficient in-browser cryptocurrency mining and has been commonly used in mining applications. In this work, we use the dynamic Wasm instruction execution trace to model the behavior of different Wasm applications. We observe that the cryptocurrency mining Wasm programs exhibit very different execution traces from other Wasm programs (e.g., games). Based on our findings, we propose a novel browser-based methodology to detect in-browser Wasm-based cryptojacking. Weikang Bian, Wei Meng 0001, Yi Wang 0004 |
CCS | 1 |