EDBT 2026 Demo / reviewers in the wild / expert
Jianlou Si
dblp:159/3878
· DBLP profile ↗
16ranked-venue papers
4as first author
11since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 12 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Style4D-Bench: A Benchmark Suite for 4D StylizationabstractWe introduce Style4D-Bench, the first benchmark suite specifically designed for 4D stylization, with the goal of standardizing evaluation and facilitating progress in this emerging area. Style4D-Bench comprises: 1) a strong baseline that make an initial attempt for 4D stylization, 2) a comprehensive evaluation protocol measuring spatial fidelity, temporal coherence, and multi-view consistency through both perceptual and quantitative metrics, and 3) a curated collection of high-resolution dynamic 4D scenes with diverse motions and complex backgrounds. To establish a strong baseline, we present Style4D, a novel framework built upon 4D Gaussian Splatting. It consists of three key components: a basic 4DGS scene representation to capture reliable geometry, a Style Gaussian Representation that leverages lightweight per-Gaussian MLPs for temporally and spatially aware appearance control, and a Holistic Geometry-Preserved Style Transfer module designed to enhance spatio-temporal consistency via contrastive coherence learning and structural content preservation. Extensive experiments on Style4D-Bench demonstrate that Style4D achieves state-of-the-art performance in 4D stylization, producing fine-grained stylistic details with stable temporal dynamics and consistent multi-view rendering. We expect Style4D-Bench to become a valuable resource for benchmarking and advancing research in stylized rendering of dynamic 3D scenes. Beiqi Chen, Haitang Feng, Jian-Huang Lai, Jianlou Si, Guangcong Wang |
AAAI | 5 |
| 2025 | TokensGen: Harnessing Condensed Tokens for Long Video GenerationabstractGenerating consistent long videos is a complex challenge: while diffusion-based generative models generate visually impressive short clips, extending them to longer durations often leads to memory bottlenecks and long-term inconsistency. In this paper, we propose TokensGen, a novel two-stage framework that leverages condensed tokens to address these issues. Our method decomposes long video generation into three core tasks: (1) inner-clip semantic control, (2) long-term consistency control, and (3) inter-clip smooth transition. First, we train To2V (Token-to-Video), a short video diffusion model guided by text and video tokens, with a Video Tokenizer that condenses short clips into semantically rich tokens. Second, we introduce T2To (Text-to-Token), a video token diffusion transformer that generates all tokens at once, ensuring global consistency across clips. Finally, during inference, an adaptive FIFO-Diffusion strategy seamlessly connects adjacent clips, reducing boundary artifacts and enhancing smooth transitions. Experimental results demonstrate that our approach significantly enhances long-term temporal and content coherence without incurring prohibitive computational overhead. By leveraging condensed tokens and pre-trained short video models, our method provides a scalable, modular solution for long video generation, opening new possibilities for storytelling, cinematic production, and immersive simulations. Please see our project page at https://vicky0522.github.io/tokensgen-webpage/ . Wenqi Ouyang, Zeqi Xiao, Danni Yang, Yifan Zhou 0001, Shuai Yang 0001, Lei Yang 0045, Jianlou Si, Xingang Pan |
ICCV | 7 |
| 2025 | Trajectory attention for fine-grained video motion controlabstractRecent advancements in video generation have been greatly driven by video diffusion models, with camera motion control emerging as a crucial challenge in creating view-customized visual content. This paper introduces trajectory attention, a novel approach that performs attention along available pixel trajectories for fine-grained camera motion control. Unlike existing methods that often yield imprecise outputs or neglect temporal correlations, our approach possesses a stronger inductive bias that seamlessly injects trajectory information into the video generation process. Importantly, our approach models trajectory attention as an auxiliary branch alongside traditional temporal attention. This design enables the original temporal attention and the trajectory attention to work in synergy, ensuring both
precise motion control and new content generation capability, which is critical when the trajectory is only partially available. Experiments on camera motion control for images and videos demonstrate significant improvements in precision and long-range consistency while maintaining high-quality generation. Furthermore, we show that our approach can be extended to other video motion control tasks, such as first-frame-guided video editing, where it excels in maintaining content consistency over large spatial and temporal ranges. Zeqi Xiao, Wenqi Ouyang, Yifan Zhou 0001, Shuai Yang 0001, Lei Yang 0045, Jianlou Si, Xingang Pan |
ICLR | 6 |
| 2024 | I2VEdit: First-Frame-Guided Video Editing via Image-to-Video Diffusion ModelsabstractThe remarkable generative capabilities of diffusion models have motivated extensive research in both image and video editing. Compared to video editing which faces additional challenges in the time dimension, image editing has witnessed the development of more diverse, high-quality approaches and more capable software like Photoshop. In light of this gap, we introduce a novel and generic solution that extends the applicability of image editing tools to videos by propagating edits from a single frame to the entire video using a pre-trained image-to-video model. Our method, dubbed I2VEdit, adaptively preserves the visual and motion integrity of the source video depending on the extent of the edits, effectively handling global edits, local edits, and moderate shape changes, which existing methods cannot fully achieve. At the core of our method are two main processes: Coarse Motion Extraction to align basic motion patterns with the original video, and Appearance Refinement for precise adjustments using fine-grained attention matching. We also incorporate a skip-interval strategy to mitigate quality degradation from auto-regressive generation across multiple video clips. Experimental results demonstrate our framework’s superior performance in fine-grained video editing, proving its capability to produce high-quality, temporally consistent outputs. Our website is at https://i2vedit.github.io/. Wenqi Ouyang, Lei Yang 0045, Jianlou Si, Xingang Pan |
SIGGRAPH Asia | 4 |
| 2023 | Amodal Instance Segmentation via Prior-Guided ExpansionabstractAmodal instance segmentation aims to infer the amodal mask, including both the visible part and occluded part of each object instance. Predicting the occluded parts is challenging. Existing methods often produce incomplete amodal boxes and amodal masks, probably due to lacking visual evidences to expand the boxes and masks. To this end, we propose a prior-guided expansion framework, which builds on a two-stage segmentation model (i.e., Mask R-CNN) and performs box-level (resp., pixel-level) expansion for amodal box (resp., mask) prediction, by retrieving regression (resp., flow) transformations from a memory bank of expansion prior. We conduct extensive experiments on KINS, D2SA, and COCOA cls datasets, which show the effectiveness of our method. Junjie Chen 0008, Li Niu 0002, Jianfu Zhang 0003, Jianlou Si, Chen Qian 0006, Liqing Zhang 0001 |
AAAI | 4 |
| 2023 | Object Part Parsing with Hierarchical Dual TransformerabstractObject part parsing involves segmenting objects into semantic parts, which has drawn great attention recently. The current methods ignore the specific hierarchical structure of the object, which can be used as strong prior knowledge. To address this, we propose the Hierarchical Dual Transformer (HDTR) to explore the contribution of the typical structural priors of the object parts. HDTR first generates the pyramid multi-granularity pixel representations under the supervision of the object part parsing maps at different semantic levels and then assigns each region an initial part embedding. Moreover, HDTR generates an edge pixel representation to extend the capability of the network to capture detailed information. Afterward, we design a Hierarchical Part Transformer to upgrade the part embeddings to their hierarchical counterparts with the assistance of the multi-granularity pixel representations. Next, we propose a Hierarchical Pixel Transformer to infer the hierarchical information from the part embeddings to enrich the pixel representations. Note that both transformer decoders rely on the structural relations between object parts, i.e., dependency, composition, and decomposition relations. The experiments on five large-scale datasets, i.e., LaPa, CelebAMask-HQ, CIHP, LIP and Pascal Animal, demonstrate that our method sets a new state-of-the-art performance for object part parsing. Jianlou Si, Naihao Liu, Li Niu 0002, Chen Qian 0006 |
ACM Multimedia | 2 |
| 2023 | Taming the Power of Diffusion Models for High-Quality Virtual Try-On with Appearance FlowabstractVirtual try-on is a critical image synthesis task that aims to transfer clothes from one image to another while preserving the details of both humans and clothes. While many existing methods rely on Generative Adversarial Networks (GANs) to achieve this, flaws can still occur, particularly at high resolutions. Recently, the diffusion model has emerged as a promising alternative for generating high-quality images in various applications. However, simply using clothes as a condition for guiding the diffusion model to inpaint is insufficient to maintain the details of the clothes. To overcome this challenge, we propose an exemplar-based inpainting approach that leverages a warping module to guide the diffusion model's generation effectively. The warping module performs initial processing on the clothes, which helps to preserve the local details of the clothes. We then combine the warped clothes with clothes-agnostic person image and add noise as the input of diffusion model. Additionally, the warped clothes is used as local conditions for each denoising process to ensure that the resulting output retains as much detail as possible. Our approach, namely Diffusion-based Conditional Inpainting for Virtual Try-ON(DCI-VTON), effectively utilizes the power of the diffusion model, and the incorporation of the warping module helps to produce high-quality and realistic virtual try-on results. Experimental results on VITON-HD demonstrate the effectiveness and superiority of our method. Source code and trained models will be publicly released at: https://github.com/bcmi/DCI-VTON-Virtual-Try-On. Junhong Gou, Jianfu Zhang 0003, Jianlou Si, Chen Qian 0006, Liqing Zhang 0001 |
ACM Multimedia | 4 |
| 2023 | ASHFormer: Axial and Sliding Window-Based Attention With High-Resolution Transformer for Automatic Stratigraphic CorrelationabstractThe stratigraphic correlation of well logs is crucial for characterizing subsurface reservoirs. However, due to the complexity of well logs and the huge amount of well data, manual correlation is time- and resource-intensive. Hence, various computerized stratigraphic correlation methods have been developed, especially regarding convolutional neural networks (CNNs). Recently, Transformer, a self-attention system that evolved from Natural Language Processing (NLP), has attained state-of-the-art performance over CNNs in a variety of domains because of its ability to perceive global features. We propose the Axial and Sliding window based attention with High-resolution Transformer (ASHFormer), combining the High-Resolution Network (HRNet) with an Axial and Sliding window self-attention Block (ASBlock) intended for stratigraphic correlation of well logs. ASBlock includes three different forms of Multi-Head Self-Attentions (MHSA), including sliding-window attention, horizontal-axis attention, and vertical-axis attention, therefore, it is possible to retrieve well logs’ long-range and local information. The experiments show that ASHFormer predicts more accurate stratigraphic correlation results than HRNet and CMT (a Transformer combining CNN and self-attention). The usefulness of the Transformer for well log feature extraction and automatic stratigraphic correlation is demonstrated by ASHFormer’s 9.74% improvement in correlation accuracy over HRNet with the same architecture. Naihao Liu, Rongchang Liu, Jinghuai Gao, Jianlou Si, Hao Wu 0047 |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2022 | Weak-shot Semantic Segmentation by Transferring Semantic Affinity and Boundary
Li Niu 0002, Jianlou Si, Chen Qian 0006, Liqing Zhang 0001 |
BMVC | 3 |
| 2022 | Weak-shot Semantic Segmentation via Dual Similarity TransferabstractSemantic segmentation is a practical and active task, but severely suffers from the expensive cost of pixel-level labels when extending to more classes in wider applications. To this end, we focus on the problem named weak-shot semantic segmentation, where the novel classes are learnt from cheaper image-level labels with the support of base classes having off-the-shelf pixel-level labels. To tackle this problem, we propose a dual similarity transfer framework, which is built upon MaskFormer to disentangle the semantic segmentation task into single-label classification and binary segmentation for each proposal. Specifically, the binary segmentation sub-task allows proposal-pixel similarity transfer from base classes to novel classes, which enables the mask learning of novel classes. We also learn pixel-pixel similarity from base classes and distill such class-agnostic semantic similarity to the semantic masks of novel classes, which regularizes the segmentation model with pixel-level semantic relationship across images. In addition, we propose a complementary loss to facilitate the learning of novel classes. Comprehensive experiments on the challenging COCO-Stuff-10K and ADE20K datasets demonstrate the effectiveness of our method. Junjie Chen 0008, Li Niu 0002, Jianlou Si, Chen Qian 0006, Liqing Zhang 0001 |
NeurIPS | 4 |
| 2021 | Video Semantic Segmentation via Sparse Temporal TransformerabstractCurrently, video semantic segmentation mainly faces two challenges: 1) the demand of temporal consistency; 2) the balance between segmentation accuracy and inference efficiency. For the first challenge, existing methods usually use optical flow to capture the temporal relation in consecutive frames and maintain the temporal consistency, but the low inference speed by means of optical flow limits the real-time applications. For the second challenge, flow based key frame warping is one mainstream solution. However, the unbalanced inference latency of flow-based key frame warping makes it unsatisfactory for real-time applications. Considering the segmentation accuracy and inference efficiency, we propose a novel Sparse Temporal Transformer (STT) to bridge temporal relation among video frames adaptively, which is also equipped with query selection and key selection. The key selection and query selection strategies are separately applied to filter out temporal and spatial redundancy in our temporal transformer. Specifically, our STT can reduce the time complexity of temporal transformer by a large margin without harming the segmentation accuracy and temporal consistency. Experiments on two benchmark datasets, Cityscapes and Camvid, demonstrate that our method achieves the state-of-the-art segmentation accuracy and temporal consistency with comparable inference speed. Jiangtong Li, Wentao Wang 0009, Junjie Chen 0008, Li Niu 0002, Jianlou Si, Chen Qian 0006, Liqing Zhang 0001 |
ACM Multimedia | 5 |
| 2020 | DMVOS: Discriminative Matching for Real-time Video Object SegmentationabstractThough recent methods on semi-supervised video object segmentation (VOS) have achieved an appreciable improvement of segmentation accuracy, it is still hard to get an adequate speed-accuracy balance when facing real-world application scenarios. In this work, we propose Discriminative Matching for real-time Video Object Segmentation (DMVOS), a real-time VOS framework with high-accuracy to fill this gap. Based on the matching mechanism, our framework introduces discriminative information through the Isometric Correlation module and the Instance Center Offset module. Specifically, the isometric correlation module learns a pixel-level similarity map with semantic discriminability, and the instance center offset module is applied to exploit the instance-level spatial discriminability. Experiments on two benchmark datasets show that our model achieves state-of-the-art performance with extremely fast speed, for example, J&F of 87.8% on DAVIS-2016 validation set with 35 milliseconds per frame. Peisong Wen, Ruolin Yang 0001, Qianqian Xu 0001, Chen Qian 0006, Qingming Huang, Runmin Cong, Jianlou Si |
ACM Multimedia | 7 |
| 2018 | Dual Attention Matching Network for Context-Aware Feature Sequence Based Person Re-IdentificationabstractTypical person re-identification (ReID) methods usually describe each pedestrian with a single feature vector and match them in a task-specific metric space. However, the methods based on a single feature vector are not sufficient enough to overcome visual ambiguity, which frequently occurs in real scenario. In this paper, we propose a novel end-to-end trainable framework, called Dual ATtention Matching network (DuATM), to learn context-aware feature sequences and perform attentive sequence comparison simultaneously. The core component of our DuATM framework is a dual attention mechanism, in which both intrasequence and inter-sequence attention strategies are used for feature refinement and feature-pair alignment, respectively. Thus, detailed visual cues contained in the intermediate feature sequences can be automatically exploited and properly compared. We train the proposed DuATM network as a siamese network via a triplet loss assisted with a decorrelation loss and a cross-entropy loss. We conduct extensive experiments on both image and video based ReID benchmark datasets. Experimental results demonstrate the significant advantages of our approach compared to the state-of-the-art methods. Jianlou Si, Honggang Zhang 0002, Chun-Guang Li, Jason Kuen, Xiangfei Kong, Alex Chichung Kot, Gang Wang 0012 |
CVPR | 1 |
| 2018 | Spatial Pyramid-Based Statistical Features for Person Re-Identification: A Comprehensive EvaluationabstractPerson re-identification (Re-Id) across nonoverlapping camera views is one of challenging problems in surveillance video analysis. The difficulties in person Re-Id mainly come from the large appearance variations caused by camera view angle, human pose, illumination, and occlusion. Recently, extensive efforts have been cast into addressing this problem by developing invariant features or discriminative distance metrics. However, there is still a lack of systematic evaluations on the pipeline for feature extraction and combination. In this paper, we propose a spatial pyramid-based statistical feature extraction framework as a unified pipeline of feature extraction and combination for person Re-Id, and systematically evaluate the configuration details in feature extraction and the fusion strategies in feature combination. Extensive experiments on benchmark datasets demonstrate the critical components in feature extraction. Moreover, by combining multiple features, our proposed approach can yield state-of-the-art performance. It should be mentioned that our approach achieves rank 1 matching rate of 45.8% on dataset VIPeR and 61.5% on dataset CUHK01, respectively. Jianlou Si, Honggang Zhang 0002, Chun-Guang Li, Jun Guo 0002 |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2015 | Regularization in metric learning for person re-identificationabstractMetric learning plays a critical role in person re-identification problem. Unfortunately, due to the small size of training data, the metric learning used in this scenario suffers from over-fitting which leads to degenerated performance. In this paper, we investigate the effect of regularization in metric learning for person re-identification. Concretely we formulate the distance function from three perspectives and hence present four different regularized metric learning methods. Experiments on two popular benchmark data sets VIPeR and CUHK01 validate the effectiveness of our proposed regularization approaches. Jianlou Si, Honggang Zhang 0002, Chun-Guang Li |
ICIP | 1 |
| 2014 | Person re-identification via region-of-interest based featuresabstractPerson re-identification is still a challenging task due to large visual appearance variations caused by illumination, background, viewpoints and poses in multi-camera surveillance. To address these challenges, many methods have been proposed. In this paper, we present an efficient method, called Region-of-Interest based Features (ROIF), via combining textural and chromatic features. It consists of two main phases - region-of-interest exploration from image and features extraction from ROI. Experimental results on the database VIPeR show that our method can yield promising accuracy with a quite cheap time cost. Jianlou Si, Honggang Zhang 0002, Chun-Guang Li |
VCIP | 1 |