VLDB 2026 Research / reviewers in the wild / expert
Siqi Kou
dblp:356/3521
· DBLP profile ↗
7ranked-venue papers
4as first author
7since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 4 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Generative modeling · 38% Efficient and distributed learning · 17% Language models and text generation · 17% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Computational science and engineering · 100% |
Topics — the 18 heaviest of 18, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling
diffusion model |
2.3 | 3 | 2025 | Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads · ICML 2025 BayesDiff: Estimating Pixel-wise Uncertainty in Diffusion via Bayesian Inference · ICLR 2024 Phasic Content Fusing Diffusion Model with Directional Distribution Consistency for Few-Shot Model Adaption · ICCV 2023 |
Machine learning › Generative modeling
autoregressive model |
0.9 | 1 | 2025 | Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads · ICML 2025 |
Computer vision › Vision and language › vision-language generation
interleaved image-text generation |
0.9 | 1 | 2025 | Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads · ICML 2025 |
Machine learning › Efficient and distributed learning › KV cache management
KV cache compression |
0.9 | 1 | 2025 | MatryoshkaKV: Adaptive KV Compression via Trainable Orthogonal Projection · ICLR 2025 |
Natural language and speech › Language models and text generation
large language model inference |
0.9 | 1 | 2025 | MatryoshkaKV: Adaptive KV Compression via Trainable Orthogonal Projection · ICLR 2025 |
Machine learning › Efficient and distributed learning
model compression |
0.9 | 1 | 2025 | MatryoshkaKV: Adaptive KV Compression via Trainable Orthogonal Projection · ICLR 2025 |
Machine learning › Generative modeling
multimodal generation |
0.9 | 1 | 2025 | Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads · ICML 2025 |
Natural language and speech › Language models and text generation
multimodal language model |
0.9 | 1 | 2025 | Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads · ICML 2025 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
bayesian inference |
0.8 | 1 | 2024 | BayesDiff: Estimating Pixel-wise Uncertainty in Diffusion via Bayesian Inference · ICLR 2024 |
Machine learning › Generative modeling › diffusion model
consistency model |
0.8 | 1 | 2024 | CLLMs: Consistency Large Language Models · ICML 2024 |
Natural language and speech › Language models and text generation › large language model inference
decoding acceleration |
0.8 | 1 | 2024 | CLLMs: Consistency Large Language Models · ICML 2024 |
Machine learning › Deep learning architectures and training › neural operator
fourier neural operator |
0.8 | 1 | 2024 | Amortized Fourier Neural Operators · NeurIPS 2024 |
Machine learning › Efficient and distributed learning
inference efficiency |
0.8 | 1 | 2024 | CLLMs: Consistency Large Language Models · ICML 2024 |
Machine learning › Deep learning architectures and training
neural operator |
0.8 | 1 | 2024 | Amortized Fourier Neural Operators · NeurIPS 2024 |
Machine learning › Trustworthy machine learning
uncertainty estimation |
0.8 | 1 | 2024 | BayesDiff: Estimating Pixel-wise Uncertainty in Diffusion via Bayesian Inference · ICLR 2024 |
Computational science and engineering › scientific machine learning › physics-informed machine learning › physics-informed neural networks
partial differential equation solving |
0.8 | 1 | 2024 | Amortized Fourier Neural Operators · NeurIPS 2024 |
Computational science and engineering
scientific machine learning |
0.8 | 1 | 2024 | Amortized Fourier Neural Operators · NeurIPS 2024 |
Machine learning › Generative modeling › generative model evaluation
distribution consistency |
0.7 | 1 | 2023 | Phasic Content Fusing Diffusion Model with Directional Distribution Consistency for Few-Shot Model Adaption · ICCV 2023 |
Methods — techniques the papers use, named apart from their topics
vector quantization · 0.9principal component analysis · 0.9orthogonal projection · 0.9matryoshka learning · 0.9low-rank projection · 0.9diffusion · 0.9autoregressive modeling · 0.9orthogonal embedding · 0.8multilayer perceptron · 0.8laplace approximation · 0.8kolmogorov-arnold network · 0.8consistency training · 0.8bayesian inference · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MatryoshkaKV: Adaptive KV Compression via Trainable Orthogonal ProjectionabstractKV cache has become a *de facto* technique for the inference of large language models (LLMs), where tensors of shape (layer number, head number, sequence length, feature dimension) are introduced to cache historical information for self-attention.
As the size of the model and data grows, the KV cache can, yet, quickly become a bottleneck within the system in both storage and memory transfer.
To address this, prior studies usually focus on the first three axes of the cache tensors for compression.
This paper supplements them, focusing on the feature dimension axis,
by utilizing low-rank projection matrices to transform the cache features into spaces with reduced dimensions.
We begin by investigating the canonical orthogonal projection method for data compression through principal component analysis (PCA).
We identify the drawback of PCA projection that model performance degrades rapidly under relatively low compression rates (less than 60%).
This phenomenon is elucidated by insights derived from the principles of attention mechanisms.
To bridge the gap, we propose to directly tune the orthogonal projection matrix on the continual pre-training or supervised fine-tuning datasets with an elaborate Matryoshka learning strategy.
Thanks to such a strategy, we can adaptively search for the optimal compression rates for various layers and heads given varying compression budgets.
Compared to Multi-head Latent Attention (MLA), our method can easily embrace pre-trained LLMs and hold a smooth tradeoff between performance and compression rate.
We witness the high data efficiency of our training procedure and find that our method can sustain over 90\% performance with an average KV cache compression rate of 60% (and up to 75% in certain extreme scenarios) for popular LLMs like LLaMA2 and Mistral. Bokai Lin, Zihao Zeng, Zipeng Xiao, Siqi Kou, Xiaofeng Gao 0001, Zhijie Deng |
ICLR | 4 |
| 2025 | Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific HeadsabstractWe introduce Orthus, a unified multimodal model that excels in generating interleaved images and text from mixed-modality inputs by simultaneously handling discrete text tokens and continuous image features under the AR modeling principle. The continuous treatment of visual signals minimizes the information loss while the fully AR formulation renders the characterization of the correlation between modalities straightforward. Orthus leverages these advantages through its modality-specific heads—one regular language modeling (LM) head predicts discrete text tokens and one diffusion head generates continuous image features. We devise an efficient strategy for building Orthus—by substituting the Vector Quantization (VQ) operation in the existing unified AR model with a soft alternative, introducing a diffusion head, and tuning the added modules to reconstruct images, we can create an Orthus-base model effortlessly (e.g., within 72 A100 GPU hours). Orthus-base can further embrace post-training to craft lengthy interleaved image-text, reflecting the potential for handling intricate real-world tasks. For visual understanding and generation, Orthus achieves a GenEval score of 0.58 and an MME-P score of 1265.8 using 7B parameters, outperforming competing baselines including Show-o and Chameleon. Siqi Kou, Jiachun Jin, Jian Jia, Quan Chen 0006, Peng Jiang 0002, Zhijie Deng |
ICML | 1 |
| 2025 | Which Data Attributes Stimulate Math and Code Reasoning? An Investigation via Influence FunctionsabstractLarge language models (LLMs) have demonstrated remarkable reasoning capabilities in math and coding, often bolstered by post-training on the chain-of-thoughts (CoTs) generated by stronger models.
However, existing strategies for curating such training data predominantly rely on heuristics, limiting generalizability and failing to capture subtleties underlying in data.
To address these limitations, we leverage influence functions to systematically attribute LLMs' reasoning ability on math and coding to individual training examples, sequences, and tokens, enabling deeper insights into effective data characteristics.
Our Influence-based Reasoning Attribution (Infra) uncovers nontrivial cross-domain effects across math and coding tasks: high-difficulty math examples improve both math and code reasoning, while low-difficulty code tasks most effectively benefit code reasoning.
Based on these findings, we introduce a simple yet effective dataset reweighting strategy by flipping task difficulty, which doubles AIME24 accuracy from 10\% to 20\% and boosts LiveCodeBench accuracy from 33.8\% to 35.3\% for Qwen2.5-7B-Instruct.
Moreover, our fine-grained attribution reveals that the sequence-level exploratory behaviors enhance reasoning performance in both math and code, and the token-level influence patterns are distinct for math and code reasoning: the former prefers natural language logic connectors and the latter emphasizes structural syntax. Siqi Kou, Qingyuan Tian, Zihao Zeng, Zhijie Deng |
NeurIPS | 1 |
| 2024 | BayesDiff: Estimating Pixel-wise Uncertainty in Diffusion via Bayesian InferenceabstractDiffusion models have impressive image generation capability, but low-quality generations still exist, and their identification remains challenging due to the lack of a proper sample-wise metric. To address this, we propose BayesDiff, a pixel-wise uncertainty estimator for generations from diffusion models based on Bayesian inference. In particular, we derive a novel uncertainty iteration principle to characterize the uncertainty dynamics in diffusion, and leverage the last-layer Laplace approximation for efficient Bayesian inference. The estimated pixel-wise uncertainty can not only be aggregated into a sample-wise metric to filter out low-fidelity images but also aids in augmenting successful generations and rectifying artifacts in failed generations in text-to-image tasks. Extensive experiments demonstrate the efficacy of BayesDiff and its promise for practical applications. Siqi Kou, Dequan Wang, Chongxuan Li, Zhijie Deng |
ICLR | 1 |
| 2024 | CLLMs: Consistency Large Language ModelsabstractJacobi decoding shows promise for more efficient LLM inference as it breaks the sequential nature of the LLM decoding process and transforms it into more parallelizable computation. However, in practice, it achieves little speedup compared to traditional autoregressive (AR) decoding, primarily because Jacobi decoding seldom accurately predicts more than one token in a single fixed-point iteration step. To address this, we develop a new approach aimed at realizing fast convergence from any state to the fixed point in a Jacobi trajectory. This is accomplished by refining the target LLM to consistently predict the fixed point given any state as input. Extensive experiments demonstrate the effectiveness of our method, showing 2.4$\times$ to 3.4$\times$ improvements in generation speed while preserving generation quality across both domain-specific and open-domain benchmarks. Siqi Kou, Lanxiang Hu, Zhezhi He, Zhijie Deng, Hao Zhang 0025 |
ICML | 1 |
| 2024 | Amortized Fourier Neural OperatorsabstractFourier Neural Operators (FNOs) have shown promise for solving partial differential equations (PDEs).
Typically, FNOs employ separate parameters for different frequency modes to specify tunable kernel integrals in Fourier space, which, yet, results in an undesirably large number of parameters when solving high-dimensional PDEs.
A workaround is to abandon the frequency modes exceeding a predefined threshold, but this limits the FNOs' ability to represent high-frequency details and poses non-trivial challenges for hyper-parameter specification.
To address these, we propose AMortized Fourier Neural Operator (AM-FNO), where an amortized neural parameterization of the kernel function is deployed to accommodate arbitrarily many frequency modes using a fixed number of parameters.
We introduce two implementations of AM-FNO, based on the recently developed, appealing Kolmogorov–Arnold Network (KAN) and Multi-Layer Perceptrons (MLPs) equipped with orthogonal embedding functions respectively.
We extensively evaluate our method on diverse datasets from various domains and observe up to 31\% average improvement compared to competing neural operator baselines. Zipeng Xiao, Siqi Kou, Zhongkai Hao, Bokai Lin, Zhijie Deng |
NeurIPS | 2 |
| 2023 | Phasic Content Fusing Diffusion Model with Directional Distribution Consistency for Few-Shot Model AdaptionabstractTraining a generative model with limited number of samples is a challenging task. Current methods primarily rely on few-shot model adaption to train the network. However, in scenarios where data is extremely limited (less than 10), the generative network tends to overfit and suffers from content degradation. To address these problems, we propose a novel phasic content fusing few-shot diffusion model with directional distribution consistency loss, which targets different learning objectives at distinct training stages of the diffusion model. Specifically, we design a phasic training strategy with phasic content fusion to help our model learn content and style information when t is large, and learn local details of target domain when t is small, leading to an improvement in the capture of content, style and local details. Furthermore, we introduce a novel directional distribution consistency loss that ensures the consistency between the generated and source distributions more efficiently and stably than the prior methods, preventing our model from overfitting. Finally, we propose a cross-domain structure guidance strategy that enhances structure consistency during domain adaptation. Theoretical analysis, qualitative and quantitative experiments demonstrate the superiority of our approach in few-shot generative model adaption tasks compared to state-of-the-art methods. The source code is available at: https://github.com/sjtuplayer/few-shot-diffusion. Jiangning Zhang, Liang Liu 0007, Ran Yi 0002, Siqi Kou, Haokun Zhu, Xu Chen 0024, Yabiao Wang, Chengjie Wang 0001, Lizhuang Ma |
ICCV | 5 |