EDBT 2026 Demo / reviewers in the wild / expert
Mingyuan Jiu
dblp:122/0982
· DBLP profile ↗
24ranked-venue papers
15as first author
15since 2021 · last 2026
0000-0002-4868-0709ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 16 · 11 first-author · 10 since 2021Artificial intelligence and machine learning · 12 · 7 first-author · 8 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Less Is Better: Sparse Instance Learning for Cross-Domain Few-Shot Object DetectionabstractCross-Domain Few-Shot Object Detection (CD-FSOD) is an extremely challenging task due to the inherent data scarcity and substantial domain shift between the source and target domains. Existing methods often suffer from overfitting and noisy feature representations, which hinder the construction of discriminative class prototypes in the target domain. In this paper, we propose a novel framework with sparse instance learning (SI-ViTO) for CD-FSOD, which leverages instance sparsity to achieve a better detection with less representation. SI-ViTO adopts a dual-stage sparsity module, consisting of instance feature sparsity not only on the few-shot support images but also on the query images. This dual sparsity enables the model to effectively preserve salient foreground semantics and simultaneously to filter out redundant or noisy information. Furthermore, a new prototype calibration strategy is also used to dynamically refine the class prototypes with query instances to accelerate prototype adaptation. Extensive experimental results on CD-FSOD benchmarks show that SI-ViTO outperforms the state-of-the-art methods, demonstrating that less discriminative representations yield better cross-domain few-shot object detection performance than more abundant ones. Yali Huang, Hongru Zhao, Mingyuan Jiu, Hichem Sahbi |
AAAI | 6 |
| 2026 | TrajKD: Distilling Knowledge via Adaptive Trajectory Curriculum and Dynamic Weighting
Mingyuan Jiu, Mi Guo, Hongru Zhao, Mingliang Xu 0001 |
ICPR (10) | 1 |
| 2026 | Sparse 3D Object Detection via Local Geometric Refinement and Dynamic Context Perception
Bingxi Chen, Xuemeng Li, Mi Guo, Mingyuan Jiu, Shupan Li |
ICPR (9) | 6 |
| 2026 | Calibrate and Aggregate: Cross-Modal Retrieval with Distribution Alignment and Token ReductionabstractCross-modal image-text retrieval remains a fundamental challenge at the intersection of vision and language. While existing methods leveraging large-scale pre-trained models such as CLIP and BERT have achieved notable progress, they often struggle with the inherent modality gap, background clutter in images, and computational inefficiency. In this paper, we propose an end-to-end framework for efficient cross-modal retrieval using Distribution Alignment and Token Reduction (DATR). The model employs a feature extraction network to extract multi-level representations, which are preliminarily aligned via a distribution calibration module. This module maps features into a Gaussian latent space and maximizes inter-modal mutual information through InfoNCE loss, effectively bridging the semantic gap. Furthermore, we incorporate a differentiable token reduction module inspired by GroupViT to dynamically cluster redundant visual tokens into semantic groups, significantly reducing computational overhead. Similarity computation integrates global and local alignments for robust matching. Extensive experiments on Flickr30K and MSCOCO datasets demonstrate state-of-the-art performance, with image-to-text retrieval R@1 reaching 89.6% and text-to-image retrieval R@1 achieving 77.2% on Flickr30K. The source codes are available at https://github.com/Wuziyi123/DATR. Mingyuan Jiu, Hongru Zhao, Hichem Sahbi, Mingliang Xu 0001 |
ICMR | 3 |
| 2026 | Deep Convolutional Primal-Dual Network for Image DeblurringabstractImage deblurring is a challenging image task, which is regarded as a classical inverse problem. Deep primal-dual proximal network (DeepPDNet) is recently proposed which unrolls the Condat-Vũ primal-dual splitting algorithm as a feed-forward network and it has demonstrated excellent restoration performance. However, the feature patterns in the DeepPDNet are well manually designed and thus the network is not implemented in an efficient convolutional fashion. In this work, we revisit the DeepPDNet and extend it in three respects: i) the convolution and pooling operators as well as their associating adjoint operations are studied in the primal-dual algorithm, and then a deep convolutional primal-dual network (DeepConvPDNet) and its full variant with skips are proposed to preserve the optimization consistence of primal-dual Condat-Vũ algorithm; ii) two (cascade vs parallel) variants of the networks are designed according to the structure of convolutional kernels; iii) rather than that the blur kernels are given as prior knowledge, they can be encoded by a set of convolutional layers and deconvolutional layers for their conjugate, resulting to a full learnable deep convolutional primal-dual neural network.We investigate the proposed networks on the MNIST dataset, the grayscale and color version of BSD dataset and GoPro dataset for image deblurring. Extensive experiments are conducted to validate the performance of the proposed networks, and promising results in term of PSNR and SSIM are obtained in comparison with twelve methods including state-of-the-art methods (e.g. Restormer, DRUNet, and DeblurGAN), which validated its effectiveness. Mingyuan Jiu, Mingjing Peng, Fanfan Zhang, Shupan Li, Hongru Zhao, Rongrong Ji, Mingliang Xu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | BiDAFuse: Bimodal Differences-Aware Attentive Network for Infrared and Visible Image Fusion
Keyu Sun, Mingyuan Jiu, Shupan Li, Hongru Zhao |
ICIG (3) | 2 |
| 2025 | SE-D3FNet: A LiDAR-Camera Fusion Network with SE Attention and Dynamic 3D Focal Loss for 3D Object Detectionabstract3D object detection is a critical problem in the field of computer vision, widely applied in autonomous driving, robotic navigation, and other domains. Although modern detectors have achieved success in singlesensor object detection, they remain vulnerable to complex environments due to the limitations of single-sensor modalities. We propose SE-D3FNet, a multi-modal fusion framework for 3D object detection that integrates Squeeze-and-Excitation (SE) channel attention and a dynamic 3D focal loss to significantly improve detection accuracy. We present an enhanced feature extraction network termed SE-ResBlock, which demonstrates superior capability in capturing global contextual information. The loss function is also optimized to more accurately capture targets with poor recognition rates. Experimental results on the KITTI benchmark demonstrate that our proposed 3D object detection algorithm achieves superior performance for the car category compared to existing methods. Mingyuan Jiu, Shupan Li, Hongru Zhao, Mingliang Xu 0001 |
ICPADS | 3 |
| 2025 | PD-YOLOv11s: An End-to-End Paper Surface Detection for Specific Visible Angle DefectabstractSurface defect detection plays a critical role in the paper manufacturing process. However, some defects are only visible from a specific angle, which challenges accurate defect recognition. We propose an innovative video defect dataset and an end-to-end detection method named PD-YOLOv11s to address this issue. We use frame differencing and Gaussian background subtraction in the defect dataset to extract inter-frame information from the video. For PD-YOLOv11s, we improve YOLOv11s by the following: (1) PBottleneck replaces the C3k2 structure to reduce the number of parameters, (2) the DSK attention mechanism is added to the end of each backbone output module to extract features better. PD-YOLOv11s achieves the following performance metrics: 8.7M parameters, 95.2% recall, 95.0% precision, 95.1% F1 score, 98.5% mAP50, and 62.8% mAP50:95. Compared to other methods (SSD, FCOS, Faster-RCNN, etc.), this approach significantly improves both accuracy and parameter efficiency, demonstrating its effectiveness in surface defect detection. Shupan Li, Hanlin Zhu, Xiangrong Zhong, Mingyuan Jiu, Mingliang Xu 0001 |
IJCNN | 4 |
| 2025 | Towards Robust Multimodal Domain Generalization via Modality-Domain Joint Adversarial TrainingabstractMultimodal Domain Generalization (MMDG) aims to enhance the robustness of multimodal models against distribution shifts in unseen target domains. Unlike unimodal domain generalization methods, which primarily focus on mitigating domain bias within individual modalities, MMDG faces unique challenges, notably modality heterogeneity (divergent feature spaces) and stability discrepancy (varying sensitivity to domain shifts). To tackle these challenges, we propose Modality-Domain Joint Adversarial Training, a unified framework that addresses these challenges through two key innovations: (1) a tri-discriminator adversarial module that mitigates domain biases in both modality-specific and multimodal representations, while suppressing modality-heterogeneous patterns in the representation space; and (2) a stability-aware dynamic weighting mechanism that adaptively balances modality contributions based on cross-domain stability, reducing reliance on unstable modalities. Additionally, we provide the first theoretical error bound for MMDG, offering a theoretical foundation that supports the effectiveness of our approach. Our approach achieves state-of-the-art performance on the EPIC-Kitchens and HAC datasets while using 75.2% fewer parameters than previous MMDG methods. The source code is available at https://github.com/lihongzhao99/MMDG-Joint-Adversarial-Training. Hongzhao Li, Hualei Wan, Liangzhi Zhang, Mingyuan Jiu, Shupan Li, Mingliang Xu 0001, Muhammad Haris Khan |
ACM Multimedia | 4 |
| 2025 | RRGMambaFormer: A hybrid Transformer-Mamba architecture for radiology report generation
Hongzhao Li, Siwei Liu 0001, Xiaoheng Jiang, Mingyuan Jiu, Yang Lu 0016, Shupan Li, Mingliang Xu 0001 |
Expert Syst. Appl. | 5 |
| 2024 | Deep Multi-order Context-Aware Kernel Network for Multi-label Classification
Mingyuan Jiu, Hailong Zhu, Hichem Sahbi |
ICPR (3) | 1 |
| 2024 | Sparse Context Transformer for Few-Shot Object Detection
Mingyuan Jiu, Hichem Sahbi, Xiaoheng Jiang, Mingliang Xu 0001 |
PRICAI (4) | 1 |
| 2022 | Context-aware deep kernel networks for image annotation
Mingyuan Jiu, Hichem Sahbi |
Neurocomputing | 1 |
| 2022 | Alternative Design of DeepPDNet in the Context of Image RestorationabstractThis work designs an image restoration deep network relying on unfolded Chambolle-Pock primal-dual iterations. Each layer of our network is built from Chambolle-Pock iterations when specified for minimizing a sum of a$\ell _2$-norm data-term and an analysis sparse prior. The parameters of our network are the step-sizes of the Chambolle-Pock scheme and the linear operator involved in sparsity-based penalization, including implicitly the regularization parameter. A backpropagation procedure is fully described. Preliminary experiments illustrate the good behavior of such a deep primal-dual network in the context of image restoration on BSD68 database. Mingyuan Jiu, Nelly Pustelnik |
IEEE Signal Process. Lett. | 1 |
| 2021 | DHCN: Deep Hierarchical Context Networks For Image AnnotationabstractContext modeling is one of the most fertile sub-fields of visual recognition which aims at designing discriminant image representations while incorporating their intrinsic and extrinsic relationships. However, the potential of context modeling is currently under-explored and most of the existing solutions are either context-free or restricted to simple handcrafted geometric relationships.We introduce in this paper DHCN: a novel Deep Hierarchical Context Network that leverages different sources of contexts including geometric and semantic relationships. The proposed method is based on the minimization of an objective function mixing a fidelity term, a context criterion and a regularizer. The solution of this objective function defines the architecture of a bi-level hierarchical context network; the first level of this network captures scene geometry while the second one corresponds to semantic relationships. We solve this representation learning problem by training its underlying deep network whose parameters correspond to the most influencing bi-level contextual relationships and we evaluate its performances on image annotation using the challenging ImageCLEF benchmark. Mingyuan Jiu, Hichem Sahbi |
ICASSP | 1 |
| 2020 | End-to-End Deep Kernel Map Design for Image AnnotationabstractDeep kernel map networks have shown excellent performances in various classification problems including image annotation. Their general recipe consists in aggregating several layers of singular value decompositions (SVDs) - that map data from input spaces into high dimensional spaces - while preserving the similarity of the underlying kernels. However, the potential of these deep map networks has not been fully explored as the original setting of these networks focuses mainly on the approximation quality of their kernels and ignores their discrimination power.In this paper, we introduce a novel “end-to-end” design for deep kernel map learning that balances the approximation quality of kernels and their discrimination power. Our method proceeds in two steps; first, layerwise SVD is applied in order to build initial deep kernel map approximations and then an “end-to-end” supervised learning is employed to further enhance their discrimination power while maintaining their efficiency. Extensive experiments, conducted on the challenging ImageCLEF annotation benchmark, show the high efficiency and the out-performance of this two-step process with respect to different related methods. Mingyuan Jiu, Hichem Sahbi |
ICIP | 1 |
| 2019 | Deep representation design from deep kernel networks
Mingyuan Jiu, Hichem Sahbi |
Pattern Recognit. | 1 |
| 2018 | Deep Context Networks for Image AnnotationabstractContext plays an important role in visual pattern recognition as it provides complementary clues for different learning tasks including image classification and annotation. In the particular scenario of kernel learning, the general recipe of context-based kernel design consists in learning positive semi-definite similarity functions that return high values not only when data share similar content but also similar context. However, in spite of having a positive impact on performance, the use of context in these kernel design methods has not been fully explored; indeed, context has been handcrafted instead of being learned. In this paper, we introduce a novel context-aware kernel design framework based on deep learning. Our method discriminatively learns spatial geometric context as the weights of a deep network (DN). The architecture of this network is fully determined by the solution of an objective function that mixes content, context and regularization, while the parameters of this network determine the most relevant (discriminant) parts of the learned context. We apply this context and kernel learning framework to image classification using the challenging ImageCLEF Photo Annotation benchmark; the latter shows that our deep context learning provides highly effective kernels for image classification as corroborated through extensive experiments. Mingyuan Jiu, Hichem Sahbi |
ICPR | 1 |
| 2017 | Nonlinear Deep Kernel Learning for Image AnnotationabstractMultiple kernel learning (MKL) is a widely used technique for kernel design. Its principle consists in learning, for a given support vector classifier, the most suitable convex (or sparse) linear combination of standard elementary kernels. However, these combinations are shallow and often powerless to capture the actual similarity between highly semantic data, especially for challenging classification tasks such as image annotation. In this paper, we redefine multiple kernels using deep multi-layer networks. In this new contribution, a deep multiple kernel is recursively defined as a multi-layered combination of nonlinear activation functions, each one involves a combination of several elementary or intermediate kernels, and results into a positive semi-definite deep kernel. We propose four different frameworks in order to learn the weights of these networks: supervised, unsupervised, kernel-based semisupervised and Laplacian-based semi-supervised. When plugged into support vector machines (SVMs), the resulting deep kernel networks show clear gain, compared to several shallow kernels for the task of image annotation. Extensive experiments and analysis on the challenging ImageCLEF photo annotation benchmark, the COREL5k database and the Banana dataset validate the effectiveness of the proposed method. Mingyuan Jiu, Hichem Sahbi |
IEEE Trans. Image Process. | 1 |
| 2016 | Laplacian deep kernel learning for image annotationabstractSemi-supervised learning seeks to build accurate classification machines by taking advantage of both labeled and unlabeled data. This learning scheme is useful especially when labeled data are scarce while unlabeled ones are abundant. Among the existing semi-supervised learning algorithms, Laplacian support vector machines (SVMs) are known to be particularly powerful but their success is highly dependent on the choice of kernels., In this paper, we propose an algorithm that designs kernels as a part of Laplacian SVM learning. The proposed kernels correspond to deep multi-layered combinations of elementary kernels which capture simple - linear - as well as intricate - nonlinear - relationships between data. Our optimization process finds both the parameters of the deep kernels and the Laplacian SVMs in a unified framework resulting into highly discriminative and accurate classifiers. When applied to the challenging ImageCLEF2013 Photo Annotation benchmark, the proposed deep kernels show significant and consistent gain compared to existing elementary kernels as well as standard multiple kernels. Mingyuan Jiu, Hichem Sahbi |
ICASSP | 1 |
| 2016 | Deep kernel map networks for image annotationabstractDeep multiple kernel learning is a powerful technique that selects and deeply combines multiple elementary kernels in order to provide the best performance on a given classification task. This technique, particularly effective, becomes intractable when handling large scale datasets; indeed, multiple nonlinear kernel combinations are time and memory demanding., In this paper, we propose a new framework that significantly reduces the complexity of deep multiple kernels. Given a deep kernel network (DKN), our method designs its equivalent deep map network (DMN), using multi-layer explicit maps that approximate the initial DKN with a high precision. When combined with support vector machines, the design of DMN preserves high classification accuracy compared to its underlying DKN while being (at least) an order of magnitude faster. Experiments conducted on the challenging Im-ageCLEF2013 annotation benchmark, show that the proposed DMN is indeed effective and highly efficient. Mingyuan Jiu, Hichem Sahbi |
ICASSP | 1 |
| 2015 | Semi supervised deep kernel design for image annotationabstractIt is commonly agreed that the success of support vector machines (SVMs), is highly dependent on the choice of particular similarity functions referred to as kernels. The latter are usually handcrafted or designed using appropriate optimization schemes. Multiple kernel learning (MKL) is one possible scheme that designs kernels as sparse or convex linear combinations of existing elementary functions. However, this results into shallow kernels, which are powerless to capture the right similarity between data, especially when content of these data is highly semantic. In this paper, we redefine multiple kernels using a deep architecture. In this new formulation, a global kernel is learned as a multi-layered linear combination of activation functions, each one involves a combination of several elementary or intermediate functions on multiple features. We propose three different settings to learn the weights of these kernel combinations; supervised, unsupervised and semi-supervised. When plugged into SVMs, the resulting deep multiple kernels show a gain, compared to shallow kernels, for the challenging task of image annotation using the ImageCLEF benchmark. Mingyuan Jiu, Hichem Sahbi |
ICASSP | 1 |
| 2014 | Evaluation of video activity localizations integrating quality and quantity measurements
Christian Wolf 0001, Eric Lombardi, Julien Mille, Oya Çeliktutan, Mingyuan Jiu, Emre Dogan, Gonen Eren, Moez Baccouche, Emmanuel Dellandréa, Charles-Edmond Bichot, Christophe Garcia, Bülent Sankur |
Comput. Vis. Image Underst. | 5 |
| 2014 | Human body part estimation from depth images via spatially-constrained deep learning
Mingyuan Jiu, Christian Wolf 0001, Graham W. Taylor, Atilla Baskurt |
Pattern Recognit. Lett. | 1 |