VLDB 2026 Research / reviewers in the wild / expert
Shiqian Su
dblp:155/0896
· DBLP profile ↗
3ranked-venue papers
1as first author
2since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Vision and language · 45% Representation and self-supervised learning · 37% Deep learning architectures and training · 17% |
Topics — the 6 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Representation and self-supervised learning
multimodal representation learning |
0.9 | 1 | 2025 | HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding · CVPR 2025 |
Computer vision › Vision and language
vision-language model |
0.9 | 1 | 2025 | HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding · CVPR 2025 |
Computer vision › Vision and language › cross-modal alignment › visual-semantic alignment
vision-language representation alignment |
0.9 | 1 | 2025 | HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding · CVPR 2025 |
Machine learning › Deep learning architectures and training
attention mechanism |
0.8 | 1 | 2024 | Learning 1D Causal Visual Representation with De-focus Attention Networks · NeurIPS 2024 |
Machine learning › Representation and self-supervised learning
visual representation |
0.8 | 1 | 2024 | Learning 1D Causal Visual Representation with De-focus Attention Networks · NeurIPS 2024 |
Computer vision › Vision and language
multimodal understanding |
0.2 | 1 | 2024 | Learning 1D Causal Visual Representation with De-focus Attention Networks · NeurIPS 2024 |
Methods — techniques the papers use, named apart from their topics
knowledge distillation · 0.9instruction tuning · 0.9learnable bandpass filters · 0.8drop path regularization · 0.8auxiliary loss · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language EmbeddingabstractThe rapid advance of Large Language Models (LLMs) has catalyzed the development of Vision-Language Models (VLMs). Monolithic VLMs, which avoid modality-specific encoders, offer a promising alternative to the compositional ones but face the challenge of inferior performance. Most existing monolithic VLMs require tuning pre-trained LLMs to acquire vision abilities, which may degrade their language capabilities. To address this dilemma, this paper presents a novel high-performance monolithic VLM named HoVLE. We note that LLMs have been shown to be capable of interpreting images when image embeddings are aligned with text embeddings. The challenge for current monolithic VLMs actually lies in the lack of a holistic embedding module for both vision and language inputs. Therefore, HoVLE introduces a holistic embedding module that converts visual and textual inputs into a shared space, allowing LLMs to process images in the same way as texts. Furthermore, a multi-stage training strategy is carefully designed to empower the holistic embedding module. It is first trained to distill visual features from a pre-trained vision encoder and text embeddings from the LLM, enabling large-scale training with unpaired random images and text tokens. The whole model further undergoes next-token prediction on multi-modal data to align the embeddings. Finally, an instruction-tuning stage is incorporated. Our experiments show that HoVLE achieves performance close to leading compositional models on various benchmarks, outperforming previous monolithic models by a large margin. Chenxin Tao, Shiqian Su, Xizhou Zhu, Zhe Chen 0017, Wenhai Wang, Lewei Lu, Gao Huang 0001, Yu Qiao 0001, Jifeng Dai |
CVPR | 2 |
| 2024 | Learning 1D Causal Visual Representation with De-focus Attention NetworksabstractModality differences have led to the development of heterogeneous architectures for vision and language models. While images typically require 2D non-causal modeling, texts utilize 1D causal modeling. This distinction poses significant challenges in constructing unified multi-modal models. This paper explores the feasibility of representing images using 1D causal modeling. We identify an "over-focus" issue in existing 1D causal vision models, where attention overly concentrates on a small proportion of visual tokens. The issue of "over-focus" hinders the model's ability to extract diverse visual features and to receive effective gradients for optimization. To address this, we propose De-focus Attention Networks, which employ learnable bandpass filters to create varied attention patterns. During training, large and scheduled drop path rates, and an auxiliary loss on globally pooled features for global understanding tasks are introduced. These two strategies encourage the model to attend to a broader range of tokens and enhance network optimization. Extensive experiments validate the efficacy of our approach, demonstrating that 1D causal visual representation can perform comparably to 2D non-causal representation in tasks such as global perception, dense prediction, and multi-modal understanding. Code shall be released. Chenxin Tao, Xizhou Zhu, Shiqian Su, Lewei Lu, Changyao Tian, Gao Huang 0001, Hongsheng Li 0001, Yu Qiao 0001, Jie Zhou 0001, Jifeng Dai |
NeurIPS | 3 |
| 2004 | A robust face detection methodabstractA new face detection method based on learning is proposed in this paper, it has three properties: first, it uses not only the local facial feature but also the global facial feature to design weak classifiers, a new kind of global facial feature called as the unified average face feature (UAFF) is proposed; second, it uses two kinds of rectangle feature as the local feature, different from other methods, these local features are selected and calculated only in the partial regions of face; third, these weak classifiers corresponding to the global facial features and the local facial features are combined and trained by our novel cascade classifier training algorithm to construct a cascade face detector. Because of these properties, our face detector is robust and generalizes well. Experimental results show that, with a small number of features, it can reach higher detection rate while maintain lower false alarm rate. Moreover, it can detect faces with partial occlusion. Shiqian Su |
ICIG | 1 |