EDBT 2026 Demo / reviewers in the wild / expert
Sijia Cai
dblp:140/7615
· DBLP profile ↗
18ranked-venue papers
4as first author
13since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 4 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 4 first-author · 7 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | PerLDiff: Controllable Street View Synthesis Using Perspective-Layout Diffusion Model
Hualian Sheng, Sijia Cai, Bing Deng, Qiao Liang 0002, Wen Li 0001, Jieping Ye, Shuhang Gu |
ICCV | 3 |
| 2025 | TAU-106K: A New Dataset for Comprehensive Understanding of Traffic AccidentabstractMultimodal Large Language Models (MLLMs) have demonstrated impressive performance in general visual understanding tasks. However, their potential for high-level, fine-grained comprehension, such as anomaly understanding, remains unexplored. Focusing on traffic accidents, a critical and practical scenario within anomaly understanding, we investigate the advanced capabilities of MLLMs and propose TABot, a multimodal MLLM specialized for accident-related tasks. To facilitate this, we first construct TAU-106K, a large-scale multimodal dataset containing 106K traffic accident videos and images collected from academic benchmarks and public platforms. The dataset is meticulously annotated through a video-to-image annotation pipeline to ensure comprehensive and high-quality labels. Building upon TAU-106K, we train TABot using a two-step approach designed to integrate multi-granularity tasks, including accident recognition, spatial-temporal grounding, and an auxiliary description task to enhance the model's understanding of accident elements. Extensive experiments demonstrate TABot's superior performance in traffic accident understanding, highlighting not only its capabilities in high-level anomaly comprehension but also the robustness of the TAU-106K benchmark. Our code and data will be available at https://github.com/cool-xuan/TABot. Yixuan Zhou 0001, Long Bai 0012, Sijia Cai, Bing Deng, Xing Xu 0001, Heng Tao Shen |
ICLR | 3 |
| 2025 | EchoShot: Multi-Shot Portrait Video GenerationabstractVideo diffusion models substantially boost the productivity of artistic workflows with high-quality portrait video generative capacity. However, prevailing pipelines are primarily constrained to single-shot creation, while real-world applications urge multiple shots with identity consistency and flexible content controllability. In this work, we propose EchoShot, a native and scalable multi-shot framework for portrait customization built upon a foundation video diffusion model. To start with, we propose shot-aware position embedding mechanisms within the video diffusion transformer architecture to model inter-shot variations and establish intricate correspondence between multi-shot visual content and their textual descriptions. This simple yet effective design enables direct training on multi-shot video data without introducing additional computational overhead. To facilitate model training within multi-shot scenarios, we construct PortraitGala, a large-scale and high-fidelity human-centric video dataset featuring cross-shot identity consistency and fine-grained captions such as facial attributes, outfits, and dynamic motions. To further enhance applicability, we extend EchoShot to perform reference image-based personalized multi-shot generation and long video synthesis with infinite shot counts. Extensive evaluations demonstrate that EchoShot achieves superior identity consistency as well as attribute-level controllability in multi-shot portrait video generation. Notably, the proposed framework demonstrates potential as a foundational paradigm for general multi-shot video modeling. Project page: https://johnneywang.github.io/EchoShot-webpage. Jiahao Wang 0004, Hualian Sheng, Sijia Cai, Weizhan Zhang, Caixia Yan, Yachuang Feng, Bing Deng, Jieping Ye |
NeurIPS | 3 |
| 2025 | Multi-sensor system deployment planning method for underwater surveillance based on formation characteristics
Zheping Yan, Sijia Cai, Shuping Hou, Mingyao Zhang |
Ad Hoc Networks | 2 |
| 2025 | CT3D++: Improving 3D Object Detection with Keypoint-Induced Channel-wise Transformer
Hualian Sheng, Sijia Cai, Na Zhao 0004, Bing Deng, Qiao Liang 0002, Minjian Zhao, Jieping Ye |
Int. J. Comput. Vis. | 2 |
| 2024 | RoScenes: A Large-Scale Multi-view 3D Dataset for Roadside Perception
Xiaosu Zhu, Hualian Sheng, Sijia Cai, Bing Deng, Shaopeng Yang, Qiao Liang 0002, Ken Chen 0005, Lianli Gao, Jingkuan Song, Jieping Ye |
ECCV (41) | 3 |
| 2024 | Versatile correlation learning for size-robust generalized counting: A new perspective
Hanqing Yang 0002, Sijia Cai, Bing Deng, Mohan Wei, Yu Zhang 0018 |
Knowl. Based Syst. | 2 |
| 2024 | Context-Aware and Semantic-Consistent Spatial Interactions for One-Shot Object Detection Without Fine-TuningabstractOne-shot object detection (OSOD) without fine-tuning has recently garnered considerable attention and research focus. It aims to directly detect novel-class objects in the target image by providing merely one support image patch without undergoing the fine-tuning stage. However, most existing methods adopt image pair matching regardless of the scale inconsistency and spatial semantic mismatch of image pairs, which limits their ability to acquire high-quality target-support related features. This paper addresses these limitations by incorporating cross-scale contexts and semantic-consistent cues that are robust against the challenges of scarce and ambiguous matching. Specifically, we first introduce a simple yet effective Aggregation-Transformer-based Pyramid (ATP) module to explore the long-range cross-scale spatial interactions by employing the customized size-aware aggregation approach and the vanilla transformer encoder, thus the coarse-to-fine local image patterns are optimally utilized. Furthermore, we formulate the 4D contrastive cross-correlation tensor for instance-level features matching and suggest a Geometric Consistent Correlation (GCC) module that utilizes the bidirectional spatial-aware convolutions to extract the long-range semantic correspondences for target-support pairs. Additionally, a Channel Contrastive Learning (CCL) branch is adopted to complement the inter-channel interactions between target-support pairs for the GCC module. Extensive experiments demonstrate that our approach significantly outperforms the previous state-of-the-art methods by 6.5% and 2.1% on PASCAL VOC and COCO datasets for unseen classes, respectively. Hanqing Yang 0002, Sijia Cai, Bing Deng, Jieping Ye, Guosheng Lin, Yu Zhang 0018 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Visual Topic Semantic Enhanced Machine Translation for Multi-Modal Data Efficiency
Sijia Cai, Bei-Xiang Shi, Zhihong Chong |
J. Comput. Sci. Technol. | 2 |
| 2023 | PDR: Progressive Depth Regularization for Monocular 3D Object DetectionabstractAccurately predicting object depth is a key challenge in monocular 3D detection task. The perspective projection principle used by most state-of-the-art approaches demands a complex balance between the ratio-form depth estimation and 2D-3D geometric regularizations, and thus can lead to sub-optimal solutions. In this paper, we propose a novel synergistic scheme that can achieve better trade-off among these competing objectives. Our main proposal is a progressive depth regularization (PDR) architecture that splits the overall training process into three sequential depth estimation steps to gradually remove the unwanted deviations induced by the over-regularization. Specifically, our model first learns the coarse depth with the conventional perspective projection and combines the coarse-to-fine generation to reduce the search space of 2D projection height prediction. We then deactivate individual supervision on 2D projection height prediction and introduces a new auxiliary 3D physical height prediction to relax the 2D and 3D regularizations, respectively. Consequently, our PDR leads to more precise depth estimation by mitigating the inherent ambiguities in the geometric priors of perspective projection through progressive regularization relaxation. Extensive experiments on both KITTI and Rope3D benchmark show that our PDR delivers strong performance gains as compared to the previous methods. Hualian Sheng, Sijia Cai, Na Zhao 0004, Bing Deng, Minjian Zhao, Gim Hee Lee |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Balanced and Hierarchical Relation Learning for One-shot Object DetectionabstractInstance-level feature matching is significantly important to the success of modern one-shot object detectors. Re-cently, the methods based on the metric-learning paradigm have achieved an impressive process. Most of these works only measure the relations between query and target objects on a single level, resulting in suboptimal performance overall. In this paper, we introduce the balanced and hierarchical learning for our detector. The contributions are two-fold: firstly, a novel Instance-level Hierarchical Relation (IHR) module is proposed to encode the contrastive-level, salient-level, and attention-level relations simultane-ously to enhance the query-relevant similarity representation. Secondly, we notice that the batch training of the IHR module is substantially hindered by the positive-negative sample imbalance in the one-shot scenario. We then in-troduce a simple but effective Ratio-Preserving Loss (RPL) to protect the learning of rare positive samples and sup-press the effects of negative samples. Our loss can adjust the weight for each sample adaptively, ensuring the desired positive-negative ratio consistency and boosting query-related IHR learning. Extensive experiments show that our method outperforms the state-of-the-art method by 1.6% and 1.3% on PASCAL VOC and MS COCO datasets for unseen classes, respectively. The code will be available at https://github.com/hero-y/BHRL. Hanqing Yang 0002, Sijia Cai, Hualian Sheng, Bing Deng, Jianqiang Huang 0001, Xian-Sheng Hua 0001, Yu Zhang 0018 |
CVPR | 2 |
| 2022 | Rethinking IoU-based Optimization for Single-stage 3D Object Detection
Hualian Sheng, Sijia Cai, Na Zhao 0004, Bing Deng, Jianqiang Huang 0001, Xian-Sheng Hua 0001, Minjian Zhao, Gim Hee Lee |
ECCV (9) | 2 |
| 2021 | Improving 3D Object Detection with Channel-wise TransformerabstractThough 3D object detection from point clouds has achieved rapid progress in recent years, the lack of flexible and high-performance proposal refinement remains a great hurdle for existing state-of-the-art two-stage detectors. Previous works on refining 3D proposals have relied on human-designed components such as keypoints sampling, set abstraction and multi-scale feature fusion to produce powerful 3D object representations. Such methods, however, have limited ability to capture rich contextual dependencies among points. In this paper, we leverage the high-quality region proposal network and a Channel-wise Transformer architecture to constitute our two-stage 3D object detection framework (CT3D) with minimal hand-crafted design. The proposed CT3D simultaneously performs proposal-aware embedding and channel-wise context aggregation for the point features within each proposal. Specifically, CT3D uses proposal’s keypoints for spatial contextual modelling and learns attention propagation in the encoding module, mapping the proposal to point embeddings. Next, a new channel-wise decoding module enriches the query-key interaction via channel-wise re-weighting to effectively merge multi-level contexts, which contributes to more accurate object predictions. Extensive experiments demonstrate that our CT3D method has superior performance and excellent scalability. Remarkably, CT3D achieves the AP of 81.77% in the moderate car category on the KITTI test 3D detection benchmark, outperforms state-of-the-art 3D detectors. Hualian Sheng, Sijia Cai, Yuan Liu 0017, Bing Deng, Jianqiang Huang 0001, Xian-Sheng Hua 0001, Minjian Zhao |
ICCV | 2 |
| 2018 | Weakly-Supervised Video Summarization Using Variational Encoder-Decoder and Web Prior
Sijia Cai, Wangmeng Zuo, Larry Davis 0001, Lei Zhang 0006 |
ECCV (14) | 1 |
| 2017 | Higher-Order Integration of Hierarchical Convolutional Activations for Fine-Grained Visual CategorizationabstractThe success of fine-grained visual categorization (FGVC) extremely relies on the modeling of appearance and interactions of various semantic parts. This makes FGVC very challenging because: (i) part annotation and detection require expert guidance and are very expensive; (ii) parts are of different sizes; and (iii) the part interactions are complex and of higher-order. To address these issues, we propose an end-to-end framework based on higherorder integration of hierarchical convolutional activations for FGVC. By treating the convolutional activations as local descriptors, hierarchical convolutional activations can serve as a representation of local parts from different scales. A polynomial kernel based predictor is proposed to capture higher-order statistics of convolutional activations for modeling part interaction. To model inter-layer part interactions, we extend polynomial predictor to integrate hierarchical activations via kernel fusion. Our work also provides a new perspective for combining convolutional activations from multiple layers. While hypercolumns simply concatenate maps from different layers, and holistically-nested network uses weighted fusion to combine side-outputs, our approach exploits higher-order intra-layer and inter-layer relations for better integration of hierarchical convolutional features. The proposed framework yields more discriminative representation and achieves competitive results on the widely used FGVC datasets. Sijia Cai, Wangmeng Zuo, Lei Zhang 0006 |
ICCV | 1 |
| 2016 | A Probabilistic Collaborative Representation Based Approach for Pattern ClassificationabstractConventional representation based classifiers, ranging from the classical nearest neighbor classifier and nearest subspace classifier to the recently developed sparse representation based classifier (SRC) and collaborative representation based classifier (CRC), are essentially distance based classifiers. Though SRC and CRC have shown interesting classification results, their intrinsic classification mechanism remains unclear. In this paper we propose a probabilistic collaborative representation framework, where the probability that a test sample belongs to the collaborative subspace of all classes can be well defined and computed. Consequently, we present a probabilistic collaborative representation based classifier (ProCRC), which jointly maximizes the likelihood that a test sample belongs to each of the multiple classes. The final classification is performed by checking which class has the maximum likelihood. The proposed ProCRC has a clear probabilistic interpretation, and it shows superior performance to many popular classifiers, including SRC, CRC and SVM. Coupled with the CNN features, it also leads to state-of-the-art classification results on a variety of challenging visual datasets. Sijia Cai, Lei Zhang 0006, Wangmeng Zuo, Xiangchu Feng |
CVPR | 1 |
| 2016 | Efficient Background Modeling Based on Sparse Representation and Outlier Iterative RemovalabstractBackground modeling is a critical component for various vision-based applications. Most traditional methods tend to be inefficient when solving large-scale problems. In this paper, we introduce sparse representation into the task of large-scale stable-background modeling, and reduce the video size by exploring its discriminative frames. A cyclic iteration process is then proposed to extract the background from the discriminative frame set. The two parts combine to form our sparse outlier iterative removal (SOIR) algorithm. The algorithm operates in tensor space to obey the natural data structure of videos. Experimental results show that a few discriminative frames determine the performance of the background extraction. Furthermore, SOIR can achieve high accuracy and high speed simultaneously when dealing with real video sequences. Thus, SOIR has an advantage in solving large-scale tasks. Linhao Li, Qinghua Hu, Sijia Cai |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2014 | Support Vector Guided Dictionary Learning
Sijia Cai, Wangmeng Zuo, Lei Zhang 0006, Xiangchu Feng, Ping Wang 0072 |
ECCV (4) | 1 |