EDBT 2026 Demo / reviewers in the wild / expert
Zenglin Shi
dblp:187/5383
· DBLP profile ↗
33ranked-venue papers
10as first author
26since 2021 · last 2026
0000-0002-1889-1409ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 21 · 6 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 5 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 1 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LoRA in LoRA: Towards Parameter-Efficient Architecture Expansion for Continual Visual Instruction TuningabstractContinual Visual Instruction Tuning (CVIT) enables Multimodal Large Language Models (MLLMs) to incrementally learn new tasks over time. However, this process is challenged by catastrophic forgetting, where performance on previously learned tasks deteriorates as the model adapts to new ones. A common approach to mitigate forgetting is architecture expansion, which introduces task-specific modules to prevent interference. Yet, existing methods often expand entire layers for each task, leading to significant parameter overhead and poor scalability. To overcome these issues, we introduce LoRA in LoRA (LiLoRA), a highly efficient architecture expansion method tailored for CVIT in MLLMs. LiLoRA shares the LoRA matrix A across tasks to reduce redundancy, applies an additional low-rank decomposition to matrix B to minimize task-specific parameters, and incorporates a cosine-regularized stability loss to preserve consistency in shared representations over time. Extensive experiments on a diverse CVIT benchmark show that LiLoRA consistently achieves superior performance in sequential task learning while significantly improving parameter efficiency compared to existing approaches. Chang Che, Pengwan Yang, Cheems Wang, Hui Ma 0011, Zenglin Shi |
AAAI | 6 |
| 2026 | Tail Task Risk Minimization in Meta-Learning From Theoretical Advances to Practical StrategiesabstractMeta learning is a promising paradigm in the era of large models, and task distributional robustness has become an indispensable consideration in real-world scenarios. Recent advances have examined the effectiveness of tail task risk minimization in fast adaptation robustness improvement. This work contributes to more theoretical investigations and practical enhancements in the field. Specifically, we reduce the distributionally robust strategy to a max-min optimization problem, constitute the Stackelberg equilibrium as the solution concept, and estimate the convergence rate. Under certain scenarios, we incorporate the diversity regularizer into the acquisition criteria design during active subset selection and further improve meta learners' comprehensive generalization under tail risk minimization. In the presence of tail risk, we further derive the generalization bound, establish connections with estimated quantiles, systematically analyze the diversity regularizer's impacts, and practically improve the studied strategy. Accordingly, extensive evaluations on tasks such as few-shot sinusoid regression, system identification, image classification, and meta reinforcement learning, along with experiments on multimodal large models, demonstrate the significance, robustness and scalability of our proposal. Yiqin Lv, Wumei Du, Zenglin Shi, Cheems Wang, Meng Wang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2026 | Boosting Semi-Supervised Learning With Entropy-Guided Adaptive Reward MaximizationabstractExisting semi-supervised learning (SSL) methods rely predominantly on pseudo-labeling and consistency regularization to leverage unlabeled data, demonstrating significant performance improvements. However, we pinpoint that these methods suffer from a confidence-for-weighting issue, overvaluing high-confidence pseudo-labels while undervaluing low-confidence yet informative samples that are critical for robust generalization. In this paper, we introduce EntropyMatch, an entropy-driven SSL framework that redefines sample importance through prediction entropy rather than confidence alone. EntropyMatch employs a bidirectional weighting strategy: upward exploitation exploits reliable hard samples to refine decision boundaries while downward exploration cautiously explores uncertain ones to reduce noise. Additionally, EntropyMatch features an adaptive training mechanism that aligns with model maturity, shifting focus from safe exploration to strategic exploitation as training progresses. Experiments on eight benchmarks across various SSL tasks-spanning image classification, facial expression recognition, and human action recognition-validate EntropyMatch's robustness and effectiveness. It consistently achieves state-of-the-art results, notably matching state-of-the-art LION's performance on RAF-DB with just half the labeled data, demonstrating superior data efficiency and generalization. Anyang Tong, Zenglin Shi, Zhun Zhong, Meng Wang 0001 |
IEEE Trans. Image Process. | 2 |
| 2026 | IAMAgent: Toward an Interactive and Adaptive Multi-Agent System for Image RestorationabstractExisting image restoration and enhancement (IRE) methods suffer from three fundamental limitations: 1) they present a high technical barrier, requiring expert knowledge and lacking intuitive natural language control; 2) they are inflexible and poorly adaptable, as models are typically designed for single, specific degradations and fail on complex or mixed real-world scenarios; and 3) they lack interactivity and ignore subjectivity, operating as "closed-box" tools that cannot incorporate human feedback or understand nuanced user intentions. To overcome these challenges, we pioneer a novel paradigm: a Multi-Agent System (MAS) for interactive and adaptive image restoration. We design and implement a prototype system, Interactive and Adaptive Multi-Agent System (IAMAgent), which orchestrates a team of specialized agents to collaboratively solve complex IRE tasks. At its core, a Manager Agent, driven by a Large Language Model, interprets user commands, devises strategies, and allocates sub-tasks. It directs a Perception Agent for degradation diagnosis, a suite of specialized Execution Agents that encapsulate various low-level vision models, and a Critique Agent for automated quality assessment. This collaborative framework enables an innovative, language-driven, and human-in-the-loop optimization process. Our work is the first to introduce the MAS paradigm to the IRE domain, transforming it from a collection of static tools into a dynamic, user-centric, and intelligent system. We demonstrate that IAMAgent not only significantly enhances restoration performance and adaptability but also bridges the critical gap between high-level human intention and low-level vision tasks. Yanyan Wei, Yilin Zhang 0012, Jiahuan Ren, Xiaogang Xu 0002, Zenglin Shi, Zhao Zhang 0001, Meng Wang 0001 |
IEEE Trans. Image Process. | 6 |
| 2026 | Random Dense Knowledge Distillation for Continual LearningabstractContinual Learning (CL), involving sequential training on diverse tasks, often faces catastrophic forgetting. While knowledge distillation–based approaches exhibit notable success in preventing forgetting, we pinpoint a limitation in their ability to distill the cumulative knowledge of all the previous tasks. To remedy this, we propose Random Dense Knowledge Distillation (RDKD). RDKD uses a task pool to track the model’s capabilities. It partitions the output logits of the model into dense groups, each corresponding to a task in the task pool. It then distills all tasks’ knowledge using all groups. However, using all the groups can be computationally expensive, so we also suggest random group selection in each optimization step. Moreover, we propose an adaptive weighting scheme, which balances the learning of new classes and the retention of old classes, based on the count and similarity of the classes. Our RDKD outperforms recent state-of-the-art baselines across diverse benchmarks and scenarios. Empirical analysis underscores RDKD’s ability to enhance model stability, promotes flatter minima for improved generalization, and remains robust across various memory budgets and task orders. Moreover, it seamlessly integrates with other CL methods to boost performance and proves versatile in offline scenarios like model compression. Jie Chu, Yunpeng Wu, Zenglin Shi |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2025 | CCMPlus: Leveraging Latent Causal Relationships Among Web Services for Traffic PredictionabstractPredicting web service traffic is crucial for system operation tasks including dynamic resource scaling, anomaly detection, and fraud detection. Web service traffic is characterized by frequent and drastic fluctuations over time and are influenced by heterogeneous user behaviors, making accurate prediction a challenging task. Previous research has extensively explored statistical approaches, and neural networks to mine features from preceding service traffic time series for prediction. However, these methods have largely overlooked the latent causal relationships between services. Drawing inspiration from causality in ecological systems, we empirically recognize the causal relationships between web services. To leverage these relationships for improved traffic prediction, we propose an effective neural network module, CCMPlus, designed to extract causal relationship features across services. This module can be seamlessly integrated with existing time series models to consistently enhance the performance of traffic predictions. We theoretically justify that the causal correlation matrix generated by the CCMPlus module captures causal relationships among services. Empirical results on real-world datasets from Microsoft Azure, Alibaba Group, and Ant Group confirm that our method surpasses state-of-the-art approaches in Mean Squared Error and Mean Absolute Error for predicting service traffic time series. These findings highlight the efficacy of feature representations from the CCMPlus module. Mingzhe Xing, Zenglin Shi, Matthew B. Blaschko, Yinliang Yue, Marie-Francine Moens |
ECAI | 3 |
| 2025 | SMoLoRa: Exploring and Defying Dual Catastrophic Forgetting in Continual Visual Instruction Tuning
Chang Che, Zenglin Shi |
ICCV | 5 |
| 2025 | MobileIE: An Extremely Lightweight and Effective ConvNet for Real-Time Image Enhancement on Mobile DevicesabstractRecent advancements in deep neural networks have driven significant progress in image enhancement (IE). However, deploying deep learning models on resource-constrained platforms, such as mobile devices, remains challenging due to high computation and memory demands. To address these challenges and facilitate real-time IE on mobile, we introduce an extremely lightweight Convolutional Neural Network (CNN) framework with around 4K parameters. Our approach integrates reparameterization with an Incremental Weight Optimization strategy to ensure efficiency. Additionally, we enhance performance with a Feature Self-Transform module and a Hierarchical Dual-Path Attention mechanism, optimized with a Local Variance-Weighted loss. With this efficient framework, we are the first to achieve real-time IE inference at up to 1,100 frames per second (FPS) while delivering competitive image quality, achieving the best trade-off between speed and performance across multiple IE tasks. The code will be available at https://github.com/AVC2-UESTC/MobileIE.git. Hailong Yan, Ao Li 0007, Xiangtao Zhang, Zhe Liu 0019, Zenglin Shi, Ce Zhu, Le Zhang 0001 |
ICCV | 5 |
| 2025 | Visual-Oriented Fine-Grained Knowledge Editing for MultiModal Large Language ModelsabstractKnowledge editing aims to efficiently and cost-effectively correct inaccuracies and update outdated information. Recently, there has been growing interest in extending knowledge editing from Large Language Models (LLMs) to Multimodal Large Language Models (MLLMs), which integrate both textual and visual information, introducing additional editing complexities. Existing multimodal knowledge editing works primarily focus on text-oriented, coarse-grained scenarios, failing to address the unique challenges posed by multimodal contexts. In this paper, we propose a visual-oriented, fine-grained multimodal knowledge editing task that targets precise editing in images with multiple interacting entities. We introduce the Fine-Grained Visual Knowledge Editing (FGVEdit) benchmark to evaluate this task. Moreover, we propose a Multimodal Scope Classifier-based Knowledge Editor (MSCKE) framework. MSCKE leverages a multimodal scope classifier that integrates both visual and textual information to accurately identify and update knowledge related to specific entities within images. This approach ensures precise editing while preserving irrelevant information, overcoming the limitations of traditional text-only editing methods. Extensive experiments on the FGVEdit benchmark demonstrate that MSCKE outperforms existing methods, showcasing its effectiveness in solving the complex challenges of multimodal knowledge editing. Leijiang Gu, Xun Yang 0001, Zhangling Duan, Zenglin Shi, Meng Wang 0001 |
ICCV | 5 |
| 2025 | Precise Localization of Memories: A Fine-grained Neuron-level Knowledge Editing Technique for LLMsabstractKnowledge editing aims to update outdated information in Large Language Models (LLMs). A representative line of study is locate-then-edit methods, which typically employ causal tracing to identify the modules responsible for recalling factual knowledge about entities. However, we find these methods are often sensitive only to changes in the subject entity, leaving them less effective at adapting to changes in relations. This limitation results in poor editing locality, which can lead to the persistence of irrelevant or inaccurate facts, ultimately compromising the reliability of LLMs. We believe this issue arises from the insufficient precision of knowledge localization. To address this, we propose a Fine-grained Neuron-level Knowledge Editing (FiNE) method that enhances editing locality without affecting overall success rates. By precisely identifying and modifying specific neurons within feed-forward networks, FiNE significantly improves knowledge localization and editing. Quantitative experiments demonstrate that FiNE efficiently achieves better overall performance compared to existing techniques, providing new insights into the localization and modification of knowledge within LLMs. Haowen Pan, Xiaozhi Wang, Yixin Cao 0002, Zenglin Shi, Xun Yang 0001, Juan-Zi Li, Meng Wang 0001 |
ICLR | 4 |
| 2025 | Prompt to Restore, Restore to Prompt: Cyclic Prompting for Universal Adverse Weather RemovalabstractUniversal adverse weather removal (UAWR) seeks to address various weather degradations within a unified framework. Recent methods are inspired by prompt learning using pre-trained vision-language models (e.g., CLIP), leveraging degradation-aware prompts to facilitate weather-free image restoration, yielding significant improvements. In this work, we propose CyclicPrompt, an innovative cyclic prompt approach designed to enhance the effectiveness, adaptability, and generalizability of UAWR. CyclicPrompt comprises two key components: 1) a composite context prompt that integrates weather-related information and context-aware representations into the network to guide restoration. This prompt differs from previous methods by marrying learnable input-conditional vectors with weather-specific knowledge, thereby improving adaptability across various degradations and 2) the erase-and-paste mechanism, after the initial guided restoration, substitutes weather-specific knowledge with constrained restoration priors, inducing high-quality weather-free concepts into the composite prompt to further fine-tune the restoration process. Therefore, we can form a cyclic "Prompt-Restore-Prompt" pipeline that adeptly harnesses weather-specific knowledge, textual contexts, and reliable textures. Extensive experiments on synthetic and real-world datasets validate the superior performance of CyclicPrompt. The code is available at: https://github.com/RongxinL/CyclicPrompt. Rongxin Liao, Feng Li 0037, Yanyan Wei, Zenglin Shi, Le Zhang 0001, Huihui Bai 0001, Meng Wang 0001 |
IEEE Trans. Image Process. | 4 |
| 2025 | Unveiling the Tapestry: The Interplay of Generalization and Forgetting in Continual LearningabstractIn artificial intelligence (AI), generalization refers to a model's ability to perform well on out-of-distribution data related to the given task, beyond the data it was trained on. For an AI agent to excel, it must also possess the continual learning capability, whereby an agent incrementally learns to perform a sequence of tasks without forgetting the previously acquired knowledge to solve the old tasks. Intuitively, generalization within a task allows the model to learn underlying features that can readily be applied to novel tasks, facilitating quicker learning and enhanced performance in subsequent tasks within a continual learning framework. Conversely, continual learning methods often include mechanisms to mitigate catastrophic forgetting, ensuring that knowledge from earlier tasks is retained. This preservation of knowledge over tasks plays a role in enhancing generalization for the ongoing task at hand. Despite the intuitive appeal of the interplay of both abilities, existing literature on continual learning and generalization has proceeded separately. In the preliminary effort to promote studies that bridge both fields, we first present empirical evidence showing that each of these fields has a mutually positive effect on the other. Next, building upon this finding, we introduce a simple and effective technique known as shape-texture consistency regularization (STCR), which caters to continual learning. STCR learns both shape and texture representations for each task, consequently enhancing generalization and thereby mitigating forgetting. Remarkably, extensive experiments validate that our STCR, can be seamlessly integrated with existing continual learning methods, including replay-free approaches. Its performance surpasses these continual learning methods in isolation or when combined with established generalization techniques by a large margin. Our data and source code are available at https://github.com/ZhangLab-DeepNeuroCogLab/distillation-style-cnn. Zenglin Shi, Jie Jing 0001, Ying Sun 0001, Joo-Hwee Lim, Mengmi Zhang |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | Training-free Object Counting with PromptsabstractThis paper tackles the problem of object counting in images. Existing approaches rely on extensive training data with point annotations for each object, making data collection labor-intensive and time-consuming. To overcome this, we propose a training-free object counter that treats the counting task as a segmentation problem. Our approach leverages the Segment Anything Model (SAM), known for its high-quality masks and zero-shot segmentation capability. However, the vanilla mask generation method of SAM lacks class-specific information in the masks, resulting in inferior counting accuracy. To overcome this limitation, we introduce a prior-guided mask generation method that incorporates three types of priors into the segmentation process, enhancing efficiency and accuracy. Additionally, we tackle the issue of counting objects specified through text by proposing a two-stage approach that combines reference object selection and prior-guided mask generation. Extensive experiments on standard datasets demonstrate the competitive performance of our training-free counter compared to learning-based approaches. This paper presents a promising solution for counting objects in various scenarios without the need for extensive data collection and counting-specific training. Code is available at https://github.com/shizenglin/training-free-object-counter. Zenglin Shi, Ying Sun 0001, Mengmi Zhang |
WACV | 1 |
| 2024 | Focus for Free in Density-Based Counting
Zenglin Shi, Pascal Mettes, Cees Snoek |
Int. J. Comput. Vis. | 1 |
| 2024 | Target-Guided Diffusion Models for Unpaired Cross-Modality Medical Image TranslationabstractIn a clinical setting, the acquisition of certain medical image modality is often unavailable due to various considerations such as cost, radiation, etc. Therefore, unpaired cross-modality translation techniques, which involve training on the unpaired data and synthesizing the target modality with the guidance of the acquired source modality, are of great interest. Previous methods for synthesizing target medical images are to establish one-shot mapping through generative adversarial networks (GANs). As promising alternatives to GANs, diffusion models have recently received wide interests in generative tasks. In this paper, we propose a target-guided diffusion model (TGDM) for unpaired cross-modality medical image translation. For training, to encourage our diffusion model to learn more visual concepts, we adopted a perception prioritized weight scheme (P2W) to the training objectives. For sampling, a pre-trained classifier is adopted in the reverse process to relieve modality-specific remnants from source data. Experiments on both brain MRI-CT and prostate MRI-US datasets demonstrate that the proposed method achieves a visually realistic result that mimics a vivid anatomical section of the target organ. In addition, we have also conducted a subjective assessment based on the synthesized samples to further validate the clinical value of TGDM. Yimin Luo, Qinyu Yang, Ziyi Liu 0010, Zenglin Shi, Weimin Huang 0002, Guoyan Zheng, Jun Cheng 0003 |
IEEE J. Biomed. Health Informatics | 4 |
| 2024 | Multiview Clustering With Propagating Information BottleneckabstractIn many practical applications, massive data are observed from multiple sources, each of which contains multiple cohesive views, called hierarchical multiview (HMV) data, such as image-text objects with different types of visual and textual features. Naturally, the inclusion of source and view relationships offers a comprehensive view of the input HMV data and achieves an informative and correct clustering result. However, most existing multiview clustering (MVC) methods can only process single-source data with multiple views or multisource data with single type of feature, failing to consider all the views across multiple sources. Observing the rich closely related multivariate (i.e., source and view) information and the potential dynamic information flow interacting among them, in this article, a general hierarchical information propagation model is first built to address the above challenging problem. It describes the process from optimal feature subspace learning (OFSL) of each source to final clustering structure learning (CSL). Then, a novel self-guided method named propagating information bottleneck (PIB) is proposed to realize the model. It works in a circulating propagation fashion, so that the resulting clustering structure obtained from the last iteration can "self-guide" the OFSL of each source, and the learned subspaces are in turn used to conduct the subsequent CSL. We theoretically analyze the relationship between the cluster structures learned in the CSL phase and the preservation of relevant information propagated from the OFSL phase. Finally, a two-step alternating optimization method is carefully designed for optimization. Experimental results on various datasets show the superiority of the proposed PIB method over several state-of-the-art methods. Shizhe Hu, Zenglin Shi, Zhengzheng Lou, Yangdong Ye |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | DualMatch: Robust Semi-supervised Learning with Dual-Level Interaction
Lanzhe Guo, Zenglin Shi |
ECML/PKDD (5) | 4 |
| 2023 | Approximately Learning Quantum Automata
Wenjing Chu, Shuo Chen 0010, Marcello M. Bonsangue, Zenglin Shi |
TASE | 4 |
| 2022 | Non-local Attention Improves Description Generation for Retinal ImagesabstractAutomatically generating medical reports from retinal images is a difficult task in which an algorithm must generate semantically coherent descriptions for a given retinal image. Existing methods mainly rely on the input image to generate descriptions. However, many abstract medical concepts or descriptions cannot be generated based on image information only. In this work, we integrate additional information to help solve this task; we observe that early in the diagnosis process, ophthalmologists have usually written down a small set of keywords denoting important information. These keywords are then subsequently used to aid the later creation of medical reports for a patient. Since these keywords commonly exist and are useful for generating medical reports, we incorporate them into automatic report generation. Since we have two types of inputs expert-defined unordered keywords and images - effectively fusing features from these different modalities is challenging. To that end, we propose a new keyword-driven medical report generation method based on a non-local attention-based multi-modal feature fusion approach, TransFuser, which is capable of fusing features from different types of inputs based on such attention. Our experiments show the proposed method successfully captures the mutual information of keywords and image content. We further show our proposed keyword-driven generation model reinforced by the TransFuser is superior to baselines under the popular text evaluation metrics BLEU, CIDEr, and ROUGE. Trans-Fuser Github: https://github.com/Jhhuangkay/Non-local-Attention-ImprovesDescription-Generation-for-Retinal-Images. Jia-Hong Huang, Ting-Wei Wu, Chao-Han Huck Yang, Zenglin Shi, I-Hung Lin, Jesper Tegnér, Marcel Worring |
WACV | 4 |
| 2022 | On Measuring and Controlling the Spectral Bias of the Deep Image PriorabstractAbstract The deep image prior showed that a randomly initialized network with a suitable architecture can be trained to solve inverse imaging problems by simply optimizing it’s parameters to reconstruct a single degraded image. However, it suffers from two practical limitations. First, it remains unclear how to control the prior beyond the choice of the network architecture. Second, training requires an oracle stopping criterion as during the optimization the performance degrades after reaching an optimum value. To address these challenges we introduce a frequency-band correspondence measure to characterize the spectral bias of the deep image prior, where low-frequency image signals are learned faster and better than high-frequency counterparts. Based on our observations, we propose techniques to prevent the eventual performance degradation and accelerate convergence. We introduce a Lipschitz-controlled convolution layer and a Gaussian-controlled upsampling layer as plug-in replacements for layers used in the deep architectures. The experiments show that with these changes the performance does not degrade during optimization, relieving us from the need for an oracle stopping criterion. We further outline a stopping criterion to avoid superfluous computation. Finally, we show that our approach obtains favorable results compared to current approaches across various denoising, deblocking, inpainting, super-resolution and detail enhancement tasks. Code is available at https://github.com/shizenglin/Measure-and-Control-Spectral-Bias . Zenglin Shi, Pascal Mettes, Subhransu Maji, Cees Snoek |
Int. J. Comput. Vis. | 1 |
| 2022 | DMIB: Dual-Correlated Multivariate Information Bottleneck for Multiview ClusteringabstractMultiview clustering (MVC) has recently been the focus of much attention due to its ability to partition data from multiple views via view correlations. However, most MVC methods only learn either interfeature correlations or intercluster correlations, which may lead to unsatisfactory clustering performance. To address this issue, we propose a novel dual-correlated multivariate information bottleneck (DMIB) method for MVC. DMIB is able to explore both interfeature correlations (the relationship among multiple distinct feature representations from different views) and intercluster correlations (the close agreement among clustering results obtained from individual views). For the former, we integrate both view-shared feature correlations discovered by learning a shared discriminative feature subspace and view-specific feature information to fully explore the interfeature correlation. This allows us to attain multiple reliable local clustering results of different views. Following this, we explore the intercluster correlations by learning the shared mutual information over different local clusterings for an improved global partition. By integrating both correlations, we formulate the problem as a unified information maximization function and further design a two-step method for optimization. Moreover, we theoretically prove the convergence of the proposed algorithm, and discuss the relationships between our method and several existing clustering paradigms. The experimental results on multiple datasets demonstrate the superiority of DMIB compared to several state-of-the-art clustering methods. Shizhe Hu, Zenglin Shi, Yangdong Ye |
IEEE Trans. Cybern. | 2 |
| 2021 | Social Fabric: Tubelet Compositions for Video Relation DetectionabstractThis paper strives to classify and detect the relationship between object tubelets appearing within a video as a 〈subject-predicate-object〉 triplet. Where existing works treat object proposals or tubelets as single entities and model their relations a posteriori, we propose to classify and detect predicates for pairs of object tubelets a priori. We also propose Social Fabric: an encoding that represents a pair of object tubelets as a composition of interaction primitives. These primitives are learned over all relations, resulting in a compact representation able to localize and classify relations from the pool of co-occurring object tubelets across all timespans in a video. The encoding enables our two-stage network. In the first stage, we train Social Fabric to suggest proposals that are likely interacting. We use the Social Fabric in the second stage to simultaneously finetune and predict predicate labels for the tubelets. Experiments demonstrate the benefit of early video relation modeling, our encoding and the two-stage architecture, leading to a new state-of-the-art on two benchmarks. We also show how the encoding enables query-by-primitive-example to search for spatio-temporal video relations. Code: https://github.com/shanshuo/Social-Fabric. Shuo Chen 0010, Zenglin Shi, Pascal Mettes, Cees Snoek |
ICCV | 2 |
| 2021 | Ordered or Orderless: A Revisit for Video Based Person Re-IdentificationabstractIs recurrent network really necessary for learning a good visual representation for video based person re-identification (VPRe-id)? In this paper, we first show that the common practice of employing recurrent neural networks (RNNs) to aggregate temporal-spatial features may not be optimal. Specifically, with a diagnostic analysis, we show that the recurrent structure may not be effective learn temporal dependencies than what we expected and implicitly yields an orderless representation. Based on this observation, we then present a simple yet surprisingly powerful approach for VPRe-id, where we treat VPRe-id as an efficient orderless ensemble of image based person re-identification problem. More specifically, we divide videos into individual images and re-identify person with ensemble of image based rankers. Under the i.i.d. assumption, we provide an error bound that sheds light upon how could we improve VPRe-id. Our work also presents a promising way to bridge the gap between video and image based person re-identification. Comprehensive experimental evaluations demonstrate that the proposed solution achieves state-of-the-art performances on multiple widely used datasets (iLIDS-VID, PRID 2011, and MARS). Le Zhang 0001, Zenglin Shi, Joey Tianyi Zhou, Ming-Ming Cheng, Yun Liu 0011, Jiawang Bian, Zeng Zeng, Chunhua Shen |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2021 | Nonlinear Regression via Deep Negative Correlation LearningabstractNonlinear regression has been extensively employed in many computer vision problems (e.g., crowd counting, age estimation, affective computing). Under the umbrella of deep learning, two common solutions exist i) transforming nonlinear regression to a robust loss function which is jointly optimizable with the deep convolutional network, and ii) utilizing ensemble of deep networks. Although some improved performance is achieved, the former may be lacking due to the intrinsic limitation of choosing a single hypothesis and the latter may suffer from much larger computational complexity. To cope with those issues, we propose to regress via an efficient "divide and conquer" manner. The core of our approach is the generalization of negative correlation learning that has been shown, both theoretically and empirically, to work well for non-deep regression problems. Without extra parameters, the proposed method controls the bias-variance-covariance trade-off systematically and usually yields a deep regression ensemble where each base model is both "accurate" and "diversified." Moreover, we show that each sub-problem in the proposed method has less Rademacher Complexity and thus is easier to optimize. Extensive experiments on several diverse and challenging tasks including crowd counting, personality analysis, age estimation, and image super-resolution demonstrate the superiority over challenging baselines as well as the versatility of the proposed method. The source code and trained models are available on our project page: https://mmcheng.net/dncl/. Le Zhang 0001, Zenglin Shi, Ming-Ming Cheng, Yun Liu 0011, Jiawang Bian, Joey Tianyi Zhou, Guoyan Zheng, Zeng Zeng |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2021 | Correction to "Nonlinear Regression via Deep Negative Correlation Learning"abstractReports on changes to the author information presented in the above named paper. Le Zhang 0001, Zenglin Shi, Ming-Ming Cheng, Yun Liu 0011, Jiawang Bian, Joey Tianyi Zhou, Guoyan Zheng, Zeng Zeng |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2021 | Unsharp Mask Guided FilteringabstractThe goal of this paper is guided image filtering, which emphasizes the importance of structure transfer during filtering by means of an additional guidance image. Where classical guided filters transfer structures using hand-designed functions, recent guided filters have been considerably advanced through parametric learning of deep networks. The state-of-the-art leverages deep networks to estimate the two core coefficients of the guided filter. In this work, we posit that simultaneously estimating both coefficients is suboptimal, resulting in halo artifacts and structure inconsistencies. Inspired by unsharp masking, a classical technique for edge enhancement that requires only a single coefficient, we propose a new and simplified formulation of the guided filter. Our formulation enjoys a filtering prior from a low-pass filter and enables explicit structure transfer by estimating a single coefficient. Based on our proposed formulation, we introduce a successive guided filtering network, which provides multiple filtering results from a single network, allowing for a trade-off between accuracy and efficiency. Extensive ablations, comparisons and analysis show the effectiveness and efficiency of our formulation and network, resulting in state-of-the-art results across filtering tasks like upsampling, denoising, and cross-modality filtering. Code is available at https://github.com/shizenglin/Unsharp-Mask-Guided-Filtering. Zenglin Shi, Yunlu Chen, Efstratios Gavves, Pascal Mettes, Cees Snoek |
IEEE Trans. Image Process. | 1 |
| 2019 | Counting With Focus for FreeabstractThis paper aims to count arbitrary objects in images. The leading counting approaches start from point annotations per object from which they construct density maps. Then, their training objective transforms input images to density maps through deep convolutional networks. We posit that the point annotations serve more supervision purposes than just constructing density maps. We introduce ways to repurpose the points for free. First, we propose supervised focus from segmentation, where points are converted into binary maps. The binary maps are combined with a network branch and accompanying loss function to focus on areas of interest. Second, we propose supervised focus from global density, where the ratio of point annotations to image pixels is used in another branch to regularize the overall density estimation. To assist both the density estimation and the focus from segmentation, we also introduce an improved kernel size estimator for the point annotations. Experiments on six datasets show that all our contributions reduce the counting error, regardless of the base network, resulting in state-of-the-art accuracy using only a single network. Finally, we are the first to count on WIDER FACE, allowing us to show the benefits of our approach in handling varying object scales and crowding levels. Code is available at https://github.com/shizenglin/Counting-with-Focus-for-Free. Zenglin Shi, Pascal Mettes, Cees Snoek |
ICCV | 1 |
| 2019 | Multidimensional Balance-Based Cluster Boundary Detection for High-Dimensional DataabstractThe balance of neighborhood space around a central point is an important concept in cluster analysis. It can be used to effectively detect cluster boundary objects. The existing neighborhood analysis methods focus on the distribution of data, i.e., analyzing the characteristic of the neighborhood space from a single perspective, and could not obtain rich data characteristics. In this paper, we analyze the high-dimensional neighborhood space from multiple perspectives. By simulating each dimension of a data point's k nearest neighbors space ( k NNs) as a lever, we apply the lever principle to compute the balance fulcrum of each dimension after proving its inevitability and uniqueness. Then, we model the distance between the projected coordinate of the data point and the balance fulcrum on each dimension and construct the DHBlan coefficient to measure the balance of the neighborhood space. Based on this theoretical model, we propose a simple yet effective cluster boundary detection algorithm called Lever. Experiments on both low- and high-dimensional data sets validate the effectiveness and efficiency of our proposed algorithm. Xiaofeng Cao 0002, Baozhi Qiu, Zenglin Shi, Guandong Xu, Jianliang Xu |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2018 | Crowd Counting With Deep Negative Correlation LearningabstractDeep convolutional networks (ConvNets) have achieved unprecedented performances on many computer vision tasks. However, their adaptations to crowd counting on single images are still in their infancy and suffer from severe over-fitting. Here we propose a new learning strategy to produce generalizable features by way of deep negative correlation learning (NCL). More specifically, we deeply learn a pool of decorrelated regressors with sound generalization capabilities through managing their intrinsic diversities. Our proposed method, named decorrelated ConvNet (D-ConvNet), is end-to-end-trainable and independent of the backbone fully-convolutional network architectures. Extensive experiments on very deep VGGNet as well as our customized network structure indicate the superiority of D-ConvNet when compared with several state-of-the-art methods. Our implementation will be released at https://github.com/shizenglin/Deep-NCL. Zenglin Shi, Le Zhang 0001, Yun Liu 0011, Yangdong Ye, Ming-Ming Cheng, Guoyan Zheng |
CVPR | 1 |
| 2018 | Bayesian VoxDRN: A Probabilistic Deep Voxelwise Dilated Residual Network for Whole Heart Segmentation from 3D MR Images
Zenglin Shi, Guodong Zeng, Le Zhang 0001, Xiahai Zhuang, Lei Li 0020, Guang Yang 0006, Guoyan Zheng |
MICCAI (4) | 1 |
| 2018 | Multiscale Multitask Deep NetVLAD for Crowd CountingabstractDeep convolutional networks (CNNs) reign undisputed as the new de-facto method for computer vision tasks owning to their success in visual recognition task on still images. However, their adaptations to crowd counting have not clearly established their superiority over shallow models. Existing CNNs turn out to be self-limiting in challenging scenarios such as camera illumination changing, partial occlusions, diverse crowd distributions, and perspective distortions for crowd counting because of their shallow structure. In this paper, we introduce a dynamic augmentation technique to train a much deeper CNN for crowd counting. In order to decrease overfitting caused by limited number of training samples, multitask learning is further employed to learn generalizable representations across similar domains. We also propose to aggregate multiscale convolutional features extracted from the entire image into a compact single vector representation amenable to efficient and accurate counting by way of “Vector of Locally Aggregated Descriptors” (VLAD). The “deeply supervised” strategy is employed to provide additional supervision signal for bottom layers for further performance improvement. Experimental results on three benchmark crowd datasets show that our method achieves better performance than the existing methods. Our implementation will be released at https://github.com/shizenglin/Multitask-Multiscale-Deep-NetVLAD. Zenglin Shi, Le Zhang 0001, Yangdong Ye |
IEEE Trans. Ind. Informatics | 1 |
| 2018 | Collective Density Clustering for Coherent Motion DetectionabstractCoherent motion detection remains a challenging problem due to the inherent complexity and vast diversity found in crowded scenes. Inspired by divide-and-conquer strategy, we desire to detect coherent motion from both local and global level. In this study, a novel collective density clustering (CDC) method is proposed to detect local and global coherent motion. We creatively define a collective density to discover underlying ordered density estimation, and subsequently a novel collective clustering algorithm is introduced, which is able to identify collective subgroups rapidly. Considering the complex interaction among subgroups, we present a hierarchical Union-Find-based collective merging algorithm to recognize coherent motion by merging collective subgroups. Our method is very efficient and effective. Experimental results on several challenging video datasets demonstrate that the proposed CDC achieves better results than state-of-the-art works, and multiple times or even tens of times faster. The proposed framework shows potential to be further applied to other problems (e.g., affine motion segmentation), related to local and global clustering. Yunpeng Wu, Yangdong Ye, Zenglin Shi |
IEEE Trans. Multim. | 4 |
| 2016 | Rank-based pooling for deep convolutional neural networks
Zenglin Shi, Yangdong Ye, Yunpeng Wu |
Neural Networks | 1 |