Weikang Bian

dblp:252/4248 · DBLP profile ↗
← Back
10ranked-venue papers
4as first author
8since 2021 · last 2025
0000-0001-9986-3348ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 5 since 2021Security and privacy · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
3D vision · 52% Generative modeling · 29% Deep learning architectures and training · 10%
Computer graphics and multimedia
3 papers
Visual content generation and editing · 54% Computer animation and physical simulation · 23% Geometric modeling and processing · 18%
Network and information security
2 papers
Web and mobile security · 50% Malware analysis · 50%

Topics — the 27 heaviest of 29, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision › motion estimation
optical flow
1.422024
BlinkVision: A Benchmark for Optical Flow, Scene Flow and Point Tracking Estimation Using RGB Frames and Events · ECCV (67) 2024
VideoFlow: Exploiting Temporal Cues for Multi-frame Optical Flow Estimation · ICCV 2023
Machine learning › Generative modeling
video generation
1.022024
ZoLA: Zero-Shot Creative Long Animation Generation with Short Video Model · ECCV (45) 2024
Phased Consistency Models · NeurIPS 2024
Computer vision › 3D vision
novel view synthesis
0.912025
GS-DiT: Advancing Video Generation with Dynamic 3D Gaussian Fields through Efficient Dense 3D Point Tracking · CVPR 2025
Visual content generation and editing › video generation
4d video generation
0.912025
GS-DiT: Advancing Video Generation with Dynamic 3D Gaussian Fields through Efficient Dense 3D Point Tracking · CVPR 2025
Visual content generation and editing
video generation
0.912025
GS-DiT: Advancing Video Generation with Dynamic 3D Gaussian Fields through Efficient Dense 3D Point Tracking · CVPR 2025
Computer vision › 3D vision
motion estimation
0.922023
VideoFlow: Exploiting Temporal Cues for Multi-frame Optical Flow Estimation · ICCV 2023
Context-PIPs: Persistent Independent Particles Demands Context Features · NeurIPS 2023
Machine learning › Generative modeling › diffusion model
consistency model
0.812024
Phased Consistency Models · NeurIPS 2024
Computer vision › 3D vision
depth estimation
0.812024
A Global Depth-Range-Free Multi-View Stereo Transformer Network with Pose Embedding · NeurIPS 2024
Machine learning › Generative modeling
diffusion model
0.812024
Phased Consistency Models · NeurIPS 2024
Computer vision › Video understanding and tracking
feature tracking
0.812024
BlinkVision: A Benchmark for Optical Flow, Scene Flow and Point Tracking Estimation Using RGB Frames and Events · ECCV (67) 2024
Computer vision › 3D vision › 3d reconstruction
multi-view stereo
0.812024
A Global Depth-Range-Free Multi-View Stereo Transformer Network with Pose Embedding · NeurIPS 2024
Machine learning › Deep learning architectures and training › transformer
multi-view transformer
0.812024
A Global Depth-Range-Free Multi-View Stereo Transformer Network with Pose Embedding · NeurIPS 2024
Computer vision › 3D vision
pose estimation
0.812024
A Global Depth-Range-Free Multi-View Stereo Transformer Network with Pose Embedding · NeurIPS 2024
Computer vision › 3D vision
scene flow estimation
0.812024
BlinkVision: A Benchmark for Optical Flow, Scene Flow and Point Tracking Estimation Using RGB Frames and Events · ECCV (67) 2024
Machine learning › Generative modeling › diffusion model
text-to-image generation
0.812024
Phased Consistency Models · NeurIPS 2024
Machine learning › Deep learning architectures and training
transformer
0.812024
A Global Depth-Range-Free Multi-View Stereo Transformer Network with Pose Embedding · NeurIPS 2024
Machine learning › Generative modeling › video generation
zero-shot video generation
0.812024
ZoLA: Zero-Shot Creative Long Animation Generation with Short Video Model · ECCV (45) 2024
Computer animation and physical simulation
animation generation
0.812024
ZoLA: Zero-Shot Creative Long Animation Generation with Short Video Model · ECCV (45) 2024
Computer vision › 3D vision › motion estimation › optical flow
multi-frame optical flow
0.712023
VideoFlow: Exploiting Temporal Cues for Multi-frame Optical Flow Estimation · ICCV 2023
Computer vision › Video understanding and tracking › object tracking › probabilistic tracking
particle tracking
0.712023
Context-PIPs: Persistent Independent Particles Demands Context Features · NeurIPS 2023
Computer vision › 3D vision › feature matching
correspondence learning
0.612022
NeuralMarker: A Framework for Learning General Marker Correspondence · ACM Trans. Graph. 2022
Geometric modeling and processing
shape correspondence
0.612022
NeuralMarker: A Framework for Learning General Marker Correspondence · ACM Trans. Graph. 2022
Web and mobile security
web security
0.412020
MineThrottle: Defending against Wasm In-Browser Cryptojacking · WWW 2020
Malware analysis › malware detection
cryptojacking detection
0.412019
Poster: Detecting WebAssembly-based Cryptocurrency Mining · CCS 2019
Computer vision › 3D vision
event-based vision
0.212024
BlinkVision: A Benchmark for Optical Flow, Scene Flow and Point Tracking Estimation Using RGB Frames and Events · ECCV (67) 2024
Machine learning › Generative modeling › video generation
text-to-video generation
0.212024
Phased Consistency Models · NeurIPS 2024
Virtual and augmented reality › augmented reality
augmented reality applications
0.212022
NeuralMarker: A Framework for Learning General Marker Correspondence · ACM Trans. Graph. 2022

Methods — techniques the papers use, named apart from their topics

gaussian splatting · 1.7diffusion transformer · 1.7dense 3d point tracking · 1.7short video model · 1.5uncertainty estimation · 0.8pose embedding · 0.8phased consistency training · 0.8multi-view disparity attention · 0.8multi-step refinement · 0.8iterative flow refinement · 0.7symmetric epipolar distance loss · 0.6structure from motion · 0.6neural network · 0.6throttling · 0.4dynamic analysis · 0.4dynamic instruction execution trace analysis · 0.4
YearPublicationVenuePosition
2025 GS-DiT: Advancing Video Generation with Dynamic 3D Gaussian Fields through Efficient Dense 3D Point Tracking
abstract
4D video control is essential in video generation as it enables the use of sophisticated lens techniques, such as multicamera shooting and dolly zoom, which are currently unsupported by existing methods. Training a video Diffusion Transformer (DiT) directly to control 4D content requires expensive multi-view videos. Inspired by Monocular Dynamic novel View Synthesis (MDVS) that optimizes a 4D representation and renders videos according to different 4D elements, such as camera pose and object motion editing, we bring dynamic 3D Gaussian fields to video generation. Specifically, we propose a novel framework that constructs dynamic 3D Gaussian fields with dense 3D point tracking and renders the Gaussian field for all video frames. Then we finetune a pretrained DiT to generate videos following the guidance of the rendered video, dubbed as GS-DiT. To boost the training of the GS-DiT, we also propose an efficient Dense 3D Point Tracking (D3D-PT) method for the dynamic 3D Gaussian field construction. Our D3D-PT outperforms SpatialTracker, the state-of-the-art sparse 3D point tracking method, in accuracy and accelerates the inference speed by two orders of magnitude. During the inference stage, GS-DiT can generate videos with the same dynamic content while adhering to different camera parameters, addressing a significant limitation of current video generation models. GS-DiT demonstrates strong generalization capabilities and extends the 4D controllability of Gaussian splatting to video generation beyond just camera poses. It supports advanced cinematic effects through the manipulation of the Gaussian field and camera intrinsics, making it a powerful tool for creative video production. Demos are available at https://wkbian.github.io/Projects/GS-DiT/.
Weikang Bian, Xiaoyu Shi 0002, Yijin Li, Fu-Yun Wang, Hongsheng Li 0001
CVPR1
2024 BlinkVision: A Benchmark for Optical Flow, Scene Flow and Point Tracking Estimation Using RGB Frames and Events
Yijin Li, Yichen Shen 0004, Weikang Bian, Xiaoyu Shi 0002, Fu-Yun Wang, Keqiang Sun, Hujun Bao, Zhaopeng Cui, Guofeng Zhang 0001, Hongsheng Li 0001
ECCV (67)5
2024 ZoLA: Zero-Shot Creative Long Animation Generation with Short Video Model
Fu-Yun Wang, Guanglu Song, Weikang Bian, Yijin Li, Yu Liu 0015, Hongsheng Li 0001
ECCV (45)6
2024 A Global Depth-Range-Free Multi-View Stereo Transformer Network with Pose Embedding
abstract
In this paper, we propose a novel multi-view stereo (MVS) framework that gets rid of the depth range prior. Unlike recent prior-free MVS methods that work in a pair-wise manner, our method simultaneously considers all the source images. Specifically, we introduce a Multi-view Disparity Attention (MDA) module to aggregate long-range context information within and across multi-view images. Considering the asymmetry of the epipolar disparity flow, the key to our method lies in accurately modeling multi-view geometric constraints. We integrate pose embedding to encapsulate information such as multi-view camera poses, providing implicit geometric constraints for multi-view disparity feature fusion dominated by attention. Additionally, we construct corresponding hidden states for each source image due to significant differences in the observation quality of the same pixel in the reference frame across multiple source frames. We explicitly estimate the quality of the current pixel corresponding to sampled points on the epipolar line of the source image and dynamically update hidden states through the uncertainty estimation module. Extensive results on the DTU dataset and Tanks\&Temple benchmark demonstrate the effectiveness of our method.
Yitong Dong, Yijin Li, Weikang Bian, Hujun Bao, Zhaopeng Cui, Hongsheng Li 0001, Guofeng Zhang 0001
NeurIPS4
2024 Phased Consistency Models
abstract
Consistency Models (CMs) have made significant progress in accelerating the generation of diffusion models. However, their application to high-resolution, text-conditioned image generation in the latent space remains unsatisfactory. In this paper, we identify three key flaws in the current design of Latent Consistency Models~(LCMs). We investigate the reasons behind these limitations and propose Phased Consistency Models (PCMs), which generalize the design space and address the identified limitations. Our evaluations demonstrate that PCMs outperform LCMs across 1--16 step generation settings. While PCMs are specifically designed for multi-step refinement, they achieve comparable 1-step generation results to previously state-of-the-art specifically designed 1-step methods. Furthermore, we show the methodology of PCMs is versatile and applicable to video generation, enabling us to train the state-of-the-art few-step text-to-video generator. Our code is available at https://github.com/G-U-N/Phased-Consistency-Model.
Fu-Yun Wang, Alexander William Bergman, Dazhong Shen, Peng Gao 0007, Michael Lingelbach, Keqiang Sun, Weikang Bian, Guanglu Song, Yu Liu 0015, Xiaogang Wang 0001, Hongsheng Li 0001
NeurIPS8
2023 VideoFlow: Exploiting Temporal Cues for Multi-frame Optical Flow Estimation
abstract
We introduce VideoFlow, a novel optical flow estimation framework for videos. In contrast to previous methods that learn to estimate optical flow from two frames, VideoFlow concurrently estimates bi-directional optical flows for multiple frames that are available in videos by sufficiently exploiting temporal cues.We first propose a TRi-frame Optical Flow (TROF) module that estimates bi-directional optical flows for the center frame in a three-frame manner. The information of the frame triplet is iteratively fused onto the center frame. To extend TROF for handling more frames, we further propose a MOtion Propagation (MOP) module that bridges multiple TROFs and propagates motion features between adjacent TROFs. With the iterative flow estimation refinement, the information fused in individual TROFs can be propagated into the whole sequence via MOP. By effectively exploiting video information, VideoFlow presents extraordinary performance, ranking 1st on all public benchmarks. On the Sintel benchmark, VideoFlow achieves 1.649 and 0.991 average end-point-error (AEPE) on the final and clean passes, a 15.1% and 7.6% error reduction from the best published results (1.943 and 1.073 from FlowFormer++). On the KITTI-2015 benchmark, VideoFlow achieves an F1-all error of 3.65%, a 19.2% error reduction from the best published result (4.52% from FlowFormer++). Code is released at https://github.com/XiaoyuShi97/VideoFlow.
Xiaoyu Shi 0002, Weikang Bian, Dasong Li, Manyuan Zhang, Ka Chun Cheung, Simon See, Hongwei Qin, Jifeng Dai, Hongsheng Li 0001
ICCV3
2023 Context-PIPs: Persistent Independent Particles Demands Context Features
Weikang Bian, Xiaoyu Shi 0002, Yitong Dong, Yijin Li, Hongsheng Li 0001
NeurIPS1
2022 NeuralMarker: A Framework for Learning General Marker Correspondence
abstract
We tackle the problem of estimating correspondences from a general marker, such as a movie poster, to an image that captures such a marker. Conventionally, this problem is addressed by fitting a homography model based on sparse feature matching. However, they are only able to handle plane-like markers and the sparse features do not sufficiently utilize appearance information. In this paper, we propose a novel framework NeuralMarker, training a neural network estimating dense marker correspondences under various challenging conditions, such as marker deformation, harsh lighting, etc. Deep learning has presented an excellent performance in correspondence learning once provided with sufficient training data. However, annotating pixel-wise dense correspondence for training marker correspondence is too expensive. We observe that the challenges of marker correspondence estimation come from two individual aspects: geometry variation and appearance variation. We, therefore, design two components addressing these two challenges in NeuralMarker. First, we create a synthetic dataset FlyingMarkers containing marker-image pairs with ground truth dense correspondences. By training with FlyingMarkers, the neural network is encouraged to capture various marker motions. Second, we propose the novel Symmetric Epipolar Distance (SED) loss, which enables learning dense correspondence from posed images. Learning with the SED loss and the cross-lighting posed images collected by Structure-from-Motion (SfM), NeuralMarker is remarkably robust in harsh lighting environments and avoids synthetic image bias. Besides, we also propose a novel marker correspondence evaluation method circumstancing annotations on real marker-image pairs and create a new benchmark. We show that NeuralMarker significantly outperforms previous methods and enables new interesting applications, including Augmented Reality (AR) and video editing.
Xiaokun Pan, Weihong Pan, Weikang Bian, Ka Chun Cheung, Guofeng Zhang 0001, Hongsheng Li 0001
ACM Trans. Graph.4
2020 MineThrottle: Defending against Wasm In-Browser Cryptojacking
abstract
In-browser cryptojacking is an urgent threat to web users, where an attacker abuses the users’ computing resources without obtaining their consent. In-browser mining programs are usually developed in WebAssembly (Wasm) for its great performance. Several prior works have measured cryptojacking in the wild and proposed detection methods using static features and dynamic features. However, there exists no good defense mechanism within the user’s browser to stop the malicious drive-by mining behavior.
Weikang Bian, Wei Meng 0001, Mingxue Zhang 0001
WWW1
2019 Poster: Detecting WebAssembly-based Cryptocurrency Mining
abstract
In-browser cryptojacking is an emerging threat to web users. The attackers can abuse the users' computation resources to perform cryptocurrency mining without obtaining their consent. Moreover, the new web feature -WebAssembly (Wasm)- enables efficient in-browser cryptocurrency mining and has been commonly used in mining applications. In this work, we use the dynamic Wasm instruction execution trace to model the behavior of different Wasm applications. We observe that the cryptocurrency mining Wasm programs exhibit very different execution traces from other Wasm programs (e.g., games). Based on our findings, we propose a novel browser-based methodology to detect in-browser Wasm-based cryptojacking.
Weikang Bian, Wei Meng 0001, Yi Wang 0004
CCS1