EDBT 2026 Demo / reviewers in the wild / expert
Xin Jin 0023
dblp:68/3340-23
· DBLP profile ↗
16ranked-venue papers
6as first author
13since 2021 · last 2025
0000-0001-5508-7957ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 9 · 4 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Unified Arbitrary-Time Video Frame Interpolation and PredictionabstractVideo frame interpolation and prediction aim to synthesize frames in-between and subsequent to existing frames, respectively. Despite being closely-related, these two tasks are traditionally studied with different model architectures, or same architecture but individually trained weights. Furthermore, while arbitrary-time interpolation has been extensively studied, the value of arbitrary-time prediction has been largely overlooked. In this work, we present uniVIP - unified arbitrary-time Video Interpolation and Prediction. Technically, we firstly extend an interpolation-only network for arbitrary-time interpolation and prediction, with a special input channel for task (interpolation or prediction) encoding. Then, we show how to train a unified model on common triplet frames. Our uniVIP provides competitive results for video interpolation, and outperforms existing state-of-the-arts for video prediction. Codes will be available at: https://github.com/srcn-ivl/uniVIP Xin Jin 0023, Longhai Wu, Ilhyun Cho, Cheul-Hee Hahm |
ICASSP | 1 |
| 2025 | Exploring Simple Siamese Network for High-Resolution Video Quality AssessmentabstractIn the research of video quality assessment (VQA), two-branch network [1] has emerged as a promising solution. It decouples VQA with separate technical and aesthetic branches to measure the perception of low-level distortions and high-level semantics respectively. However, we argue that while technical and aesthetic perspectives are complementary, the technical perspective itself should be measured in semantic-aware manner. We hypothesize that existing technical branch struggles to perceive the semantics of high-resolution videos, as it is trained on local mini-patches sampled from videos. This issue can be hidden by apparently good results on low-resolution videos, but indeed becomes critical for high-resolution VQA. This work introduces SiamVQA, a simple but effective Siamese network for high-resolution VQA. SiamVQA shares weights between technical and aesthetic branches, enhancing the semantic perception ability of technical branch to facilitate technical-quality representation learning. Furthermore, it integrates a dual cross-attention layer for fusing technical and aesthetic features. SiamVQA achieves state-of-the-art accuracy on high-resolution benchmarks, and competitive results on lower-resolution benchmarks. Codes will be available at: https://github.com/srcn-ivl/SiamVQA Guotao Shen, Ziheng Yan, Xin Jin 0023, Longhai Wu, Ilhyun Cho, Cheul-Hee Hahm |
ICASSP | 3 |
| 2025 | UPR-Net: A Unified Pyramid Recurrent Network for Video Frame Interpolation
Xin Jin 0023, Longhai Wu, Youxin Chen, Jayoon Koo, Cheul-Hee Hahm |
Int. J. Comput. Vis. | 1 |
| 2024 | Alignment-aware Patch-level Routing for Dynamic Video Frame Interpolation
Ban Chen, Xin Jin 0023, Longhai Wu, Ilhyun Cho, Cheul-Hee Hahm |
BMVC | 2 |
| 2024 | Dynamic Video Frame Interpolation with Integrated Difficulty Pre-AssessmentabstractVideo frame interpolation (VFI) has witnessed great progress in recent years. However, existing VFI models still struggle to achieve a good trade-off between accuracy and efficiency. Accurate VFI models typically rely on heavy compute to process all samples, ignoring the fact that easy samples with small motion or clear texture can be well addressed by a fast VFI model and do not require such heavy compute. In this paper, we present a dynamic VFI pipeline with integrated pre-assessment of interpolation difficulty. Specifically, it leverages a difficulty pre-assessment model to measure the difficulty level of interpolating input frames, and then dynamically selects an accurate or a fast VFI model for frame interpolation. Furthermore, we contribute a large-scale annotated dataset to train our VFI difficulty pre-assessment model. Extensive experiments show that our dynamic VFI pipeline can achieve an excellent trade-off between accuracy and efficiency, by feeding hard samples to accurate model, and passing easy samples through fast model. Ban Chen, Xin Jin 0023, Youxin Chen, Longhai Wu, Jayoon Koo, Cheul-Hee Hahm |
ICASSP | 2 |
| 2024 | In Pursuit of Causal Label Correlations for Multi-label Image RecognitionabstractMulti-label image recognition aims to predict all objects present in an input image. A common belief is that modeling the correlations between objects is beneficial for multi-label recognition. However, this belief has been recently challenged as label correlations may mislead the classifier in testing, due to the possible contextual bias in training. Accordingly, a few of recent works not only discarded label correlation modeling, but also advocated to remove contextual information for multi-label image recognition. This work explicitly explores label correlations for multi-label image recognition based on a principled causal intervention approach. With causal intervention, we pursue causal label correlations and suppress spurious label correlations, as the former tend to convey useful contextual cues while the later may mislead the classifier. Specifically, we decouple label-specific features with a Transformer decoder attached to the backbone network, and model the confounders which may give rise to spurious correlations by clustering spatial features of all training images. Based on label-specific features and confounders, we employ a cross-attention module to implement causal intervention, quantifying the causal correlations from all object categories to each predicted object category. Finally, we obtain image labels by combining the predictions from decoupled features and causal label correlations. Extensive experiments clearly validate the effectiveness of our approach for multi-label image recognition in both common and cross-dataset settings. Xin Jin 0023, Yisu Ge |
NeurIPS | 2 |
| 2024 | SiSe: Simultaneous and Sequential Transformers for multi-label activity recognition
Xin Jin 0023 |
Pattern Recognit. | 2 |
| 2023 | A Unified Pyramid Recurrent Network for Video Frame InterpolationabstractFlow-guided synthesis provides a common framework for frame interpolation, where optical flow is estimated to guide the synthesis of intermediate frames between consecutive inputs. In this paper, we present UPR-Net, a novel Unified Pyramid Recurrent Network for frame interpolation. Cast in a flexible pyramid framework, UPR-Net exploits lightweight recurrent modules for both bi-directional flow estimation and intermediate frame synthesis. At each pyramid level, it leverages estimated bi-directional flow to generate forward-warped representations for frame synthesis; across pyramid levels, it enables iterative refinement for both optical flow and intermediate frame. In particular, we show that our iterative synthesis strategy can significantly improve the robustness of frame interpolation on large motion cases. Despite being extremely lightweight (1.7M parameters), our base version of UPR-Net achieves excellent performance on a large range of benchmarks. Code and trained models of our UPR-Net series are available at: https://github.com/srcn-iv1/UPR-Net. Xin Jin 0023, Longhai Wu, Youxin Chen, Jayoon Koo, Cheul-Hee Hahm |
CVPR | 1 |
| 2023 | Enhanced Bi-directional Motion Estimation for Video Frame InterpolationabstractWe propose a simple yet effective algorithm for motion-based video frame interpolation. Existing motion-based interpolation methods typically rely on an off-the-shelf optical flow model or a U-Net based pyramid network for motion estimation, which either suffer from large model size or limited capacity in handling various challenging motion cases. In this work, we present a novel compact model to simultaneously estimate the bi-directional motions between input frames. It is designed by carefully adapting the ingredients (e.g., warping, correlation) in optical flow research for simultaneous bi-directional motion estimation within a flexible pyramid recurrent framework. Our motion estimator is extremely lightweight (15x smaller than PWC-Net), yet enables reliable handling of large and complex motion cases. Based on estimated bi-directional motions, we employ a synthesis network to fuse forward-warped representations and predict the intermediate frame. Our method achieves excellent performance on a broad range of frame interpolation benchmarks. Code and trained models are available at https://github.com/srcn-ivl/EBME. Xin Jin 0023, Longhai Wu, Guotao Shen, Youxin Chen, Jayoon Koo, Cheul-Hee Hahm |
WACV | 1 |
| 2022 | Delving deep into spatial pooling for squeeze-and-excitation networks
Xin Jin 0023, Yanping Xie, Xiu-Shen Wei, Borui Zhao, Xiaoyang Tan |
Pattern Recognit. | 1 |
| 2022 | A Lightweight Encoder-Decoder Path for Deep Residual NetworksabstractIn this article, we present a novel lightweight path for deep residual neural networks. The proposed method integrates a simple plug-and-play module, i.e., a convolutional encoder-decoder (ED), as an augmented path to the original residual building block. Due to the abstract design and ability of the encoding stage, the decoder part tends to generate feature maps where highly semantically relevant responses are activated, while irrelevant responses are restrained. By a simple elementwise addition operation, the learned representations derived from the identity shortcut and original transformation branch are enhanced by our ED path. Furthermore, we exploit lightweight counterparts by removing a portion of channels in the original transformation branch. Fortunately, our lightweight processing does not cause an obvious performance drop but brings a computational economy. By conducting comprehensive experiments on ImageNet, MS-COCO, CUB200-2011, and CIFAR, we demonstrate the consistent accuracy gain obtained by our ED path for various residual architectures, with comparable or even lower model complexity. Concretely, it decreases the top-1 error of ResNet-50 and ResNet-101 by 1.22% and 0.91% on the task of ImageNet classification and increases the mmAP of Faster R-CNN with ResNet-101 by 2.5% on the MS-COCO object detection task. The code is available at https://github.com/Megvii-Nanjing/ED-Net. Xin Jin 0023, Yanping Xie, Xiu-Shen Wei, Borui Zhao, Xiaoyang Tan, Yang Yu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2021 | HCE: Hierarchical Context Embedding for Region-Based Object DetectionabstractState-of-the-art two-stage object detectors apply a classifier to a sparse set of object proposals, relying on region-wise features extracted by RoIPool or RoIAlign as inputs. The region-wise features, in spite of aligning well with the proposal locations, may still lack the crucial context information which is necessary for filtering out noisy background detections, as well as recognizing objects possessing no distinctive appearances. To address this issue, we present a simple but effective Hierarchical Context Embedding (HCE) framework, which can be applied as a plug-and-play component, to facilitate the classification ability of a series of region-based detectors by mining contextual cues. Specifically, to advance the recognition of context-dependent object categories, we propose an image-level categorical embedding module which leverages the holistic image-level context to learn object-level concepts. Then, novel RoI features are generated by exploiting hierarchically embedded context information beneath both whole images and interested regions, which are also complementary to conventional RoI features. Moreover, to make full use of our hierarchical contextual RoI features, we propose the early-and-late fusion strategies (i.e., feature fusion and confidence fusion), which can be combined to boost the classification accuracy of region-based detectors. Comprehensive experiments demonstrate that our HCE framework is flexible and generalizable, leading to significant and consistent improvements upon various region-based detectors, including FPN, Cascade R-CNN, Mask R-CNN and PA-FPN. With simple modification, our HCE framework can be conveniently adapted to fit the structure of one-stage detectors, and achieve improved performance for SSD, RetinaNet and EfficientDet. Xin Jin 0023, Borui Zhao, Xiaoqin Zhang 0002, Yanwen Guo 0001 |
IEEE Trans. Image Process. | 2 |
| 2021 | Disentangling, Embedding and Ranking Label Cues for Multi-Label Image RecognitionabstractMulti-label image recognition is a fundamental but challenging computer vision and multimedia task. Great progress has been achieved by exploiting label correlations among these multiple labels associated with a single image, which is the most crucial issue for multi-label image recognition. In this paper, to explicitly model label correlations, we propose a unified deep learning framework to Disentangle, Embed and Rank (DER) the corresponding label cues. Specifically, we first obtain class-aware disentangled maps (CADMs) by reforming deep activations in accordance with the class-specific recognition weights. Then, after transforming CADMs into the corresponding label vectors, we propose an embedding operation from a metric learning perspective to pull the relevant label vectors together and push irrelevant label vectors away. Furthermore, a ranking operation is employed, which aims to accurately and robustly measure the similarity/dissimilarity of these label vectors. Our model can be trained in an end-to-end manner with only image-level supervision, during which the proposed embedding and ranking operations can contribute to the CADMs learning through back-propagation. In addition, the obtained CADMs are aggregated and further used as an essential feature stream for the final multi-label classification. We conduct extensive experiments on three commonly used multi-label benchmark datasets. Quantitative results show that our model can significantly and consistently outperform previous competitive methods. Moreover, qualitative analysis of our DER proposal also reveals the effectiveness of our proposed model. Quan Cui, Xiu-Shen Wei, Xin Jin 0023, Yanwen Guo 0001 |
IEEE Trans. Multim. | 4 |
| 2020 | Exploring Categorical Regularization for Domain Adaptive Object DetectionabstractIn this paper, we tackle the domain adaptive object detection problem, where the main challenge lies in significant domain gaps between source and target domains. Previous work seeks to plainly align image-level and instance-level shifts to eventually minimize the domain discrepancy. However, they still overlook to match crucial image regions and important instances across domains, which will strongly affect domain shift mitigation. In this work, we propose a simple but effective categorical regularization framework for alleviating this issue. It can be applied as a plug-and-play component on a series of Domain Adaptive Faster R-CNN methods which are prominent for dealing with domain adaptive detection. Specifically, by integrating an image-level multi-label classifier upon the detection backbone, we can obtain the sparse but crucial image regions corresponding to categorical information, thanks to the weakly localization ability of the classification manner. Meanwhile, at the instance level, we leverage the categorical consistency between image-level predictions (by the classifier) and instance-level predictions (by the detection head) as a regularization factor to automatically hunt for the hard aligned instances of target domains. Extensive experiments of various domain shift scenarios show that our method obtains a significant performance gain over original Domain Adaptive Faster R-CNN detectors. Furthermore, qualitative visualization and analyses can demonstrate the ability of our method for attending on the key regions/instances targeting on domain adaptation. Our code is open-source and available at https://github.com/Megvii-Nanjing/CR-DA-DET. Chang-Dong Xu, Xing-Ran Zhao, Xin Jin 0023, Xiu-Shen Wei |
CVPR | 3 |
| 2020 | Hierarchical Context Embedding for Region-Based Object Detection
Xin Jin 0023, Borui Zhao, Xiu-Shen Wei, Yanwen Guo 0001 |
ECCV (21) | 2 |
| 2019 | Multi-Label Image Recognition with Joint Class-Aware Map Disentangling and Label Correlation EmbeddingabstractMulti-label image recognition is a fundamental but challenging computer vision task. Great progress has been achieved by exploring the label correlation among these multiple labels which is the most crucial issue for multi-label recognition. In this paper, we propose a unified deep learning framework to jointly disentangle class-specific maps corresponding to discriminative category-wise information and then evaluate the label co-occurrence of these maps. Specifically, after obtaining the general deep image features and conducting multi-label classification, we employ the classification weights to reform the feature maps into class-aware disentangled maps (CADMs). Then, based on CADMs, we first transfer them into label vectors and then formulate the label correlation dependency from an embedding perspective. The whole model is driven by both the classification loss and the label correlation embedding loss, which is end-to-end trainable with only image-level supervisions. Extensive quantitative results of two benchmark multi-label image datasets show our model consistently outperforms other competing methods by a large margin. Meanwhile, qualitative analyses also demonstrate our model can effectively capture relatively pure class-aware maps and model label correlation dependency as well. Xiu-Shen Wei, Xin Jin 0023, Yanwen Guo 0001 |
ICME | 3 |