EDBT 2026 Demo / reviewers in the wild / expert
Sijie Zhao
dblp:300/5422
· DBLP profile ↗
10ranked-venue papers
5as first author
10since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Generative modeling · 41% 3D vision · 13% Efficient and distributed learning · 12% | |
| Computer graphics and multimedia
3 papers |
Computational photography and imaging · 53% Audio and music processing · 35% Visual content generation and editing · 12% |
Topics — the 25 heaviest of 27, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Vision and language › vision-language model
multimodal large language model |
1.4 | 2 | 2024 | Making LLaMA SEE and Draw with SEED Tokenizer · ICLR 2024 GPT4Tools: Teaching Large Language Model to Use Tools via Self-instruction · NeurIPS 2023 |
Computer vision › 3D vision
depth estimation |
0.9 | 1 | 2025 | DepthCrafter: Generating Consistent Long Depth Sequences for Open-world Videos · CVPR 2025 |
Machine learning › Generative modeling
diffusion model |
0.9 | 1 | 2025 | DepthCrafter: Generating Consistent Long Depth Sequences for Open-world Videos · CVPR 2025 |
Computer vision › 3D vision › depth estimation
video depth estimation |
0.9 | 1 | 2025 | DepthCrafter: Generating Consistent Long Depth Sequences for Open-world Videos · CVPR 2025 |
Machine learning › Generative modeling › diffusion model
video diffusion model |
0.9 | 1 | 2025 | DepthCrafter: Generating Consistent Long Depth Sequences for Open-world Videos · CVPR 2025 |
Machine learning › Generative modeling › autoregressive model
autoregressive image generation |
0.8 | 1 | 2024 | Making LLaMA SEE and Draw with SEED Tokenizer · ICLR 2024 |
Machine learning › Deep learning architectures and training
convolutional neural network |
0.8 | 1 | 2024 | UniRepLKNet: A Universal Perception Large-Kernel ConvNet for Audio, Video, Point Cloud, Time-Series and Image Recognition · CVPR 2024 |
Computer vision › Image recognition and object detection
image classification |
0.8 | 1 | 2024 | UniRepLKNet: A Universal Perception Large-Kernel ConvNet for Audio, Video, Point Cloud, Time-Series and Image Recognition · CVPR 2024 |
Machine learning › Deep learning architectures and training › convolutional neural network › convolutional neural network architecture
large-kernel convnet |
0.8 | 1 | 2024 | UniRepLKNet: A Universal Perception Large-Kernel ConvNet for Audio, Video, Point Cloud, Time-Series and Image Recognition · CVPR 2024 |
Machine learning › Generative modeling
latent space regularization |
0.8 | 1 | 2024 | CV-VAE: A Compatible Video VAE for Latent Generative Video Models · NeurIPS 2024 |
Machine learning › Generative modeling › diffusion model › video diffusion model
latent video diffusion model |
0.8 | 1 | 2024 | CV-VAE: A Compatible Video VAE for Latent Generative Video Models · NeurIPS 2024 |
Machine learning › Generative modeling › autoregressive model
multimodal autoregressive model |
0.8 | 1 | 2024 | Making LLaMA SEE and Draw with SEED Tokenizer · ICLR 2024 |
Machine learning › Generative modeling
variational autoencoder |
0.8 | 1 | 2024 | CV-VAE: A Compatible Video VAE for Latent Generative Video Models · NeurIPS 2024 |
Machine learning › Generative modeling
video generation |
0.8 | 1 | 2024 | CV-VAE: A Compatible Video VAE for Latent Generative Video Models · NeurIPS 2024 |
Audio and music processing › audio analysis › audio content analysis
audio recognition |
0.8 | 1 | 2024 | UniRepLKNet: A Universal Perception Large-Kernel ConvNet for Audio, Video, Point Cloud, Time-Series and Image Recognition · CVPR 2024 |
Natural language and speech › Language models and text generation
instruction tuning |
0.7 | 1 | 2023 | GPT4Tools: Teaching Large Language Model to Use Tools via Self-instruction · NeurIPS 2023 |
Natural language and speech › Language models and text generation › LLM agents
tool use |
0.7 | 1 | 2023 | GPT4Tools: Teaching Large Language Model to Use Tools via Self-instruction · NeurIPS 2023 |
Computational photography and imaging › depth estimation
depth reconstruction |
0.6 | 1 | 2022 | Fisher Information Guidance for Learned Time-of-Flight Imaging · CVPR 2022 |
Computational photography and imaging
time-of-flight imaging |
0.6 | 1 | 2022 | Fisher Information Guidance for Learned Time-of-Flight Imaging · CVPR 2022 |
Machine learning › Efficient and distributed learning › model quantization
bit-width allocation |
0.5 | 1 | 2021 | Distribution-Aware Adaptive Multi-Bit Quantization · CVPR 2021 |
Machine learning › Efficient and distributed learning
model compression |
0.5 | 1 | 2021 | Distribution-Aware Adaptive Multi-Bit Quantization · CVPR 2021 |
Machine learning › Efficient and distributed learning › model compression › quantization
multi-bit quantization |
0.5 | 1 | 2021 | Distribution-Aware Adaptive Multi-Bit Quantization · CVPR 2021 |
Computer vision › 3D vision › 3d object recognition
point cloud recognition |
0.2 | 1 | 2024 | UniRepLKNet: A Universal Perception Large-Kernel ConvNet for Audio, Video, Point Cloud, Time-Series and Image Recognition · CVPR 2024 |
Machine learning › Efficient and distributed learning › parameter-efficient fine-tuning
low-rank adaptation |
0.2 | 1 | 2023 | GPT4Tools: Teaching Large Language Model to Use Tools via Self-instruction · NeurIPS 2023 |
Machine learning › Efficient and distributed learning › model compression
pruning |
0.1 | 1 | 2021 | Distribution-Aware Adaptive Multi-Bit Quantization · CVPR 2021 |
Methods — techniques the papers use, named apart from their topics
three-stage training · 1.7segment-wise stitching · 1.7modality-related preprocessing · 1.5spatio-temporal compression · 0.8latent space regularization · 0.8large-kernel convolution · 0.8large kernel convolution · 0.8instruction tuning · 0.8image tokenization · 0.8autoregressive transformer · 0.8low-rank adaptation · 0.7fisher information · 0.6dual-branch reconstruction network · 0.6differentiable physical imaging modeling · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | DepthCrafter: Generating Consistent Long Depth Sequences for Open-world VideosabstractEstimating video depth in open-world scenarios is challenging due to the diversity of videos in appearance, content motion, camera movement, and length. We present DepthCrafter, an innovative method for generating temporally consistent long depth sequences with intricate details for open-world videos, without requiring any supplementary information such as camera poses or optical flow. The generalization ability to open-world videos is achieved by training the video-to-depth model from a pretrained image-to-video diffusion model, through our meticulously designed three-stage training strategy. Our training approach enables the model to generate depth sequences with variable lengths at one time, up to 110 frames, and harvest both precise depth details and rich content diversity from realistic and synthetic datasets. We also propose an inference strategy that can process extremely long videos through segment-wise estimation and seamless stitching. Comprehensive evaluations on multiple datasets reveal that DepthCrafter achieves state-of-the-art performance in open-world video depth estimation under zero-shot settings. Furthermore, DepthCrafter facilitates various downstream applications, including depth-based visual effects and conditional video generation. Wenbo Hu 0002, Xiangjun Gao, Xiaoyu Li 0002, Sijie Zhao, Xiaodong Cun, Yong Zhang 0034, Long Quan, Ying Shan |
CVPR | 4 |
| 2025 | VegeDiff: Latent Diffusion Model for Geospatial Vegetation ForecastingabstractIn the context of global climate change and frequent extreme weather events, forecasting future geospatial vegetation states under these conditions is of significant importance. The vegetation change process is influenced by the complex interplay between dynamic meteorological variables and static environmental variables, leading to high levels of uncertainty. Existing deterministic methods are inadequate in addressing this uncertainty and fail to accurately model the impact of these variables on vegetation, resulting in blurry and inaccurate forecasting results. To address these issues, VegeDiff is proposed for the geospatial vegetation forecasting task. To our best knowledge, VegeDiff is the first to employ a diffusion model to probabilistically capture the uncertainties in vegetation change processes, enabling the generation of clear and accurate future vegetation states. VegeDiff also separately models the global impact of dynamic meteorological variables and the local effects of static environmental variables, thus accurately modeling the impact of these variables. Extensive experiments on geospatial vegetation forecasting tasks demonstrate the effectiveness of VegeDiff. By capturing the uncertainties in vegetation changes and modeling the complex influence of relevant variables, VegeDiff outperforms existing deterministic methods, providing clear and accurate forecasting results of future vegetation states. Interestingly, this study demonstrate the potential of VegeDiff in applications of forecasting future vegetation states from multiple aspects and exploring the impact of meteorological variables on vegetation dynamics. The code of this work will be available at https://github.com/walking-shadow/Official VegeDiff. Sijie Zhao, Hao Chen 0045, Xueliang Zhang 0002, Pengfeng Xiao, Lei Bai 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | UniRepLKNet: A Universal Perception Large-Kernel ConvNet for Audio, Video, Point Cloud, Time-Series and Image RecognitionabstractLarge-kernel convolutional neural networks (ConvNets) have recently received extensive research attention, but two unresolved and critical issues demand further investigation. 1) The architectures of existing large-kernel ConvNets largely follow the design principles of conventional ConvNets or transformers, while the architectural design for large-kernel ConvNets remains under-addressed. 2) As transformers have dominated multiple modalities, it re-mains to be investigated whether ConvNets also have a strong universal perception ability in domains beyond vision. In this paper, we contribute from two aspects. 1) We propose four architectural guidelines for designing large- kernel ConvNets, the core of which is to exploit the essential characteristics of large kernels that distinguish them from small kernels - they can see wide without going deep. Fol-lowing such guidelines, our proposed large-kernel ConvNet shows leading performance in image recognition (ImageNet accuracy of 88.0%, ADE20K mIoU of 55.6%, and COCO box AP of 56.4%), demonstrating better performance and higher speed than the recent powerful competitors. 2) We discover large kernels are the key to unlocking the exceptional performance of ConvNets in domains where they were originally not proficient. With certain modality-related pre-processing approaches, the proposed model achieves state- of-the-art performance on time-series forecasting and audio recognition tasks even without modality-specific customization to the architecture. All the code and models are publicly available on GitHub and Huggingface. Xiaohan Ding, Yixiao Ge, Sijie Zhao, Lin Song 0002, Xiangyu Yue 0001, Ying Shan |
CVPR | 4 |
| 2024 | Making LLaMA SEE and Draw with SEED TokenizerabstractThe great success of Large Language Models (LLMs) has expanded the potential of multimodality, contributing to the gradual evolution of General Artificial Intelligence (AGI). A true AGI agent should not only possess the capability to perform predefined multi-tasks but also exhibit emergent abilities in an open-world context. However, despite the considerable advancements made by recent multimodal LLMs, they still fall short in effectively unifying comprehension and generation tasks, let alone open-world emergent abilities. We contend that the key to overcoming the present impasse lies in enabling text and images to be represented and processed interchangeably within a unified autoregressive Transformer. To this end, we introduce $\textbf{SEED}$, an elaborate image tokenizer that empowers LLMs with the ability to $\textbf{SEE}$ and $\textbf{D}$raw at the same time. We identify two crucial design principles: (1) Image tokens should be independent of 2D physical patch positions and instead be produced with a $\textit{1D causal dependency}$, exhibiting intrinsic interdependence that aligns with the left-to-right autoregressive prediction mechanism in LLMs. (2) Image tokens should capture $\textit{high-level semantics}$ consistent with the degree of semantic abstraction in words, and be optimized for both discriminativeness and reconstruction during the tokenizer training phase. With SEED tokens, LLM is able to perform scalable multimodal autoregression under its original training recipe, i.e., next-word prediction. SEED-LLaMA is therefore produced by large-scale pretraining and instruction tuning on the interleaved textual and visual data, demonstrating impressive performance on a broad range of multimodal comprehension and generation tasks. More importantly, SEED-LLaMA has exhibited compositional emergent abilities such as multi-turn in-context multimodal generation, acting like your AI assistant. The code (training and inference) and models are released in https://github.com/AILab-CVC/SEED. Yuying Ge, Sijie Zhao, Ziyun Zeng, Yixiao Ge, Chen Li 0046, Xintao Wang 0002, Ying Shan |
ICLR | 2 |
| 2024 | CV-VAE: A Compatible Video VAE for Latent Generative Video ModelsabstractSpatio-temporal compression of videos, utilizing networks such as Variational Autoencoders (VAE), plays a crucial role in OpenAI's SORA and numerous other video generative models. For instance, many LLM-like video models learn the distribution of discrete tokens derived from 3D VAEs within the VQVAE framework, while most diffusion-based video models capture the distribution of continuous latent extracted by 2D VAEs without quantization. The temporal compression is simply realized by uniform frame sampling which results in unsmooth motion between consecutive frames. Currently, there lacks of a commonly used continuous video (3D) VAE for latent diffusion-based video models in the research community. Moreover, since current diffusion-based approaches are often implemented using pre-trained text-to-image (T2I) models, directly training a video VAE without considering the compatibility with existing T2I models will result in a latent space gap between them, which will take huge computational resources for training to bridge the gap even with the T2I models as initialization. To address this issue, we propose a method for training a video VAE of latent video models, namely CV-VAE, whose latent space is compatible with that of a given image VAE, e.g., image VAE of Stable Diffusion (SD). The compatibility is achieved by the proposed novel latent space regularization, which involves formulating a regularization loss using the image VAE. Benefiting from the latent space compatibility, video models can be trained seamlessly from pre-trained T2I or video models in a truly spatio-temporally compressed latent space, rather than simply sampling video frames at equal intervals. To improve the training efficiency, we also design a novel architecture for the video VAE. With our CV-VAE, existing video models can generate four times more frames with minimal finetuning. Extensive experiments are conducted to demonstrate the effectiveness of the proposed video VAE. Sijie Zhao, Yong Zhang 0034, Xiaodong Cun, Shaoshu Yang, Muyao Niu, Xiaoyu Li 0002, Wenbo Hu 0002, Ying Shan |
NeurIPS | 1 |
| 2024 | RS-Mamba for Large Remote Sensing Image Dense PredictionabstractContext modeling is critical for remote sensing image dense prediction tasks. Nowadays, the growing size of very-high-resolution (VHR) remote sensing images poses challenges in effectively modeling context. While transformer-based models possess global modeling capabilities, they encounter computational challenges when applied to large VHR images due to their quadratic complexity. The conventional practice of cropping large images into smaller patches results in a notable loss of contextual information. To address these issues, we propose the remote sensing Mamba (RSM) for dense prediction tasks in large VHR remote sensing images. RSM is specifically designed to capture the global context of remote sensing images with linear complexity, facilitating the effective processing of large VHR images. Considering that the land covers in remote sensing images are distributed in arbitrary spatial directions due to characteristics of remote sensing over-head imaging, the RSM incorporates an omnidirectional selective scan module (OSSM) to globally model the context of images in multiple directions, capturing large spatial features from various directions. We designed simple yet effective models based on RSM, achieving state-of-the-art performance on dense prediction tasks in VHR remote sensing images without fancy training strategies. Extensive experiments on semantic segmentation (SS) and change detection (CD) tasks across various land covers demonstrate the effectiveness of the proposed RSM. Leveraging the linear complexity and global modeling capabilities, RSM achieves better efficiency and accuracy than transformer-based models on large remote sensing images. Interestingly, we also demonstrated that our model generally performs better with a larger image size on dense prediction tasks. Sijie Zhao, Hao Chen 0045, Xueliang Zhang 0002, Pengfeng Xiao, Lei Bai 0001, Wanli Ouyang |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | GPT4Tools: Teaching Large Language Model to Use Tools via Self-instructionabstractThis paper aims to efficiently enable Large Language Models (LLMs) to use multi-modal tools.
The advanced proprietary LLMs, such as ChatGPT and GPT-4, have shown great potential for tool usage through sophisticated prompt engineering.
Nevertheless, these models typically rely on prohibitive computational costs and publicly inaccessible data.
To address these challenges, we propose the GPT4Tools based on self-instruct to enable open-source LLMs, such as LLaMA and OPT, to use tools.
It generates an instruction-following dataset by prompting an advanced teacher with various multi-modal contexts.
By using the Low-Rank Adaptation (LoRA) optimization, our approach facilitates the open-source LLMs to solve a range of visual problems, including visual comprehension and image generation.
Moreover, we provide a benchmark to evaluate the ability of LLMs to use tools, which is performed in both zero-shot and fine-tuning ways.
Extensive experiments demonstrate the effectiveness of our method on various language models, which not only significantly improves the accuracy of invoking seen tools, but also enables the zero-shot capacity for unseen tools. Rui Yang 0041, Lin Song 0002, Sijie Zhao, Yixiao Ge, Xiu Li 0001, Ying Shan |
NeurIPS | 4 |
| 2023 | Exchanging Dual-Encoder-Decoder: A New Strategy for Change Detection With Semantic Guidance and Spatial LocalizationabstractChange detection is a critical task in earth observation applications. Recently, deep-learning-based methods have shown promising performance and are quickly adopted in change detection. However, the widely used multiple encoders and single decoder (MESD) as well as dual-encoder–decoder (DED) architectures still struggle to effectively handle change detection well. The former has problems of bitemporal feature interference in the feature-level fusion, while the latter is inapplicable to intraclass change detection (ICCD) and multiview building change detection (MVBCD). To solve these problems, we propose a new strategy with an exchanging DED (EDED) structure for binary change detection with semantic guidance and spatial localization. The proposed strategy solves the problems of bitemporal feature inference in MESD by fusing bitemporal features in the decision level and the inapplicability in DED by determining changed areas using bitemporal semantic features. We build a binary change detection model based on this strategy and then validate and compare it with 18 state-of-the-art change detection methods on six datasets in three scenarios, including ICCD datasets (CDD and SYSU), single-view building change detection (SVBCD) datasets (WHU, LEVIR-CD, and LEVIR-CD+), and an MVBCD dataset (NJDS). The experimental results demonstrate that our model achieves superior performance with high efficiency and outperforms all benchmark methods with F1-scores of 97.77%, 83.07%, 94.86%, 92.33%, 91.39%, and 74.35% on CDD, SYSU, WHU, LEVIR-CD, LEVIR-CD+, and NJDS datasets, respectively. The code of this work will be available athttps://github.com/NJU-LHRS/official-SGSLN. Sijie Zhao, Xueliang Zhang 0002, Pengfeng Xiao, Guangjun He |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Fisher Information Guidance for Learned Time-of-Flight ImagingabstractIndirect Time-of-Flight (ToF) imaging is widely applied in practice for its superiorities on cost and spatial resolution. However, lower signal-to-noise ratio (SNR) of measurement leads to larger error in ToF imaging, especially for imaging scenes with strong ambient light or long distance. In this paper, we propose a Fisher-information guided framework to jointly optimize the coding functions (light modulation and sensor demodulation functions) and the reconstruction network of iToF imaging, with the super-vision of the proposed discriminative fisher loss. By introducing the differentiable modeling of physical imaging process considering various real factors and constraints, e.g., light-falloff with distance, physical implementability of coding functions, etc., followed by a dual-branch depth reconstruction neural network, the proposed method could learn the optimal iToF imaging system in an end-to-end manner. The effectiveness of the proposed method is extensively verified with both simulations and prototype experiments. Jiaqu Li, Tao Yue 0003, Sijie Zhao |
CVPR | 3 |
| 2021 | Distribution-Aware Adaptive Multi-Bit QuantizationabstractIn this paper, we explore the compression of deep neural networks by quantizing the weights and activations into multi-bit binary networks (MBNs). A distribution-aware multi-bit quantization (DMBQ) method that incorporates the distribution prior into the optimization of quantization is proposed. Instead of solving the optimization in each iteration, DMBQ search the optimal quantization scheme over the distribution space beforehand, and select the quantization scheme during training using a fast lookup table based strategy. Based upon DMBQ, we further propose loss-guided bit-width allocation (LBA) to adaptively quantize and even prune the neural network. The first-order Taylor expansion is applied to build a metric for evaluating the loss sensitivity of the quantization of each channel, and automatically adjust the bit-width of weights and activations channel-wisely. We extend our method to image classification tasks and experimental results show that our method not only outperforms state-of-the-art quantized networks in terms of accuracy but also is more efficient in terms of training time compared with state-of-the-art MBNs, even for the extremely low bit width (below 1-bit) quantization cases. Sijie Zhao, Tao Yue 0003 |
CVPR | 1 |