EDBT 2026 Demo / reviewers in the wild / expert
Yu Liu 0012
dblp:97/2274-12
· DBLP profile ↗
66ranked-venue papers
19as first author
38since 2021 · last 2026
0000-0002-2067-9175ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 51 · 15 first-author · 31 since 2021Artificial intelligence and machine learning · 28 · 10 first-author · 14 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Learning 3D Occupancy from Beam Overlap in 2D Rotating mmWave RadarabstractRobust 3D perception under adverse weather is critical for autonomous systems. While mmWave Radars are inherently weather-resistant, conventional 2D rotating Radar sensors lack direct elevation resolution, limiting their 3D perception ability. Although 4D imaging radars can provide elevation information, they typically suffer from limited coverage and range. In this work, we exploit a key observation about mechanically rotating 2D mmWave Radars: in each sweep, an overlap exists between adjacent azimuth beam coverage due to the width of the main lobe, which makes the reflected intensity difference imply object materials and geometric shapes, including elevation. With this observation, we propose a method that learns 3D occupancy by disentangling bird’s-eye view (BEV) layout and elevation estimation from one frame Radar scan. Specifically, we partition one sweep into two interleaved subsets, corresponding to overlapping beam directions, and utilize them to infer coarse geometric structure through spatial differences and intensity patterns. Extensive quantitative and qualitative evaluations on two real-world datasets demonstrate that our proposed method outperforms existing baselines. The codes will be publicly available. Ruifeng Nie, Long Ma 0002, Chengpei Xu, Yu Liu 0012, Weimin Wang 0007 |
AAAI | 5 |
| 2026 | RSOD: Reliability-Guided Sonar Image Object Detection with Extremely Limited LabelsabstractObject detection in sonar images is a key technology in underwater detection systems. Compared to natural images, sonar images contain fewer texture details and are more susceptible to noise, making it difficult for non-experts to distinguish subtle differences between classes. This leads to their inability to provide precise annotation data for sonar images. Therefore, designing effective object detection methods for sonar images with extremely limited labels is particularly important. To address this, we propose a teacher-student framework called RSOD, which aims to fully learn the characteristics of sonar images and develop a pseudo-label strategy suitable for these images to mitigate the impact of limited labels. First, RSOD calculates a reliability score by assessing the consistency of the teacher's predictions across different views. To leverage this score, we introduce an object mixed pseudo-label method to tackle the shortage of labeled data in sonar images. Finally, we optimize the performance of the student by implementing a reliability-guided adaptive constraint. By taking full advantage of unlabeled data, the student can perform well even in situations with extremely limited labels. Notably, on the UATD dataset, our method, using only 5% of labeled data, achieves results that can compete against those of our baseline algorithm trained on 100% labeled data. We also collected a new dataset to provide more valuable data for research in the field of sonar. Chengzhou Li, Guanchen Meng, Qi Jia 0001, Jinyuan Liu 0001, Zhu Liu 0004, Yu Liu 0012, Zhongxuan Luo, Xin Fan 0001 |
AAAI | 8 |
| 2026 | SWG-Fusion: Soft weather-guided multimodal fusion with VLM-assistance for BEV object detection under harsh weather
Weimin Wang 0007, Ruifeng Nie, Yingchi Liu, Long Ma 0002, Chengpei Xu, Qi Jia 0001, Yu Liu 0012, Na Lei |
Pattern Recognit. | 7 |
| 2026 | Model-aware ellipse detection via parametric correlation learningabstractEllipse detection presents a significant challenge in computer vision and pattern recognition, often hindered by traditional parameter regression methods that fail to account for the unique geometric characteristics and complex parameter interactions of ellipses. These limitations frequently result in imprecise detections, notably with small or partially occluded ellipses. To overcome these challenges, we propose EDNet, a novel ellipse detection network that exploits the geometric properties of ellipses, thus moving beyond the reliance on internal textures. EDNet improves ellipse detection by refining the loss function to better capture the relationship between the error and each parameters during training. It features a LoG-like Edge Detection Module (LEDM) and an Edge Guided Module (EGM) for precise boundary extraction and multi-scale feature enhancement. Additionally, an auxiliary component estimates ellipse vertices, boosting accuracy for occluded ellipses. Experimental results on two wildly-used benchmark datasets demonstrate that EDNet achieves significant improvements, with an average detection accuracy increase of 6% and 10% over leading state-of-the-art models. • We concentrate on the geometric characteristics of ellipse detection via Edge Detection Module and Edge Guided Module. • We design an auxiliary head for the estimation of four ellipse vertices, invoking additional feature attention on these pivotal points. • We establish the relations between the error and geometric characteristics of the ellipse by a model-aware loss function. Qi Jia 0001, Zezheng Liu, Yu Liu 0012, Yi Wang 0037, Xinwei Xue, Weimin Wang 0007 |
Signal Process. | 3 |
| 2025 | As Pseudo-Label Free as Possible: Leveraging Adaptive Feature Generation for Sparsely Annotated Object DetectionabstractCompared to fully supervised object detection, training with sparse annotations typically leads to a decline in performance due to insufficient feature diversity. Existing sparsely annotated object detection (SAOD) methods often rely on pseudo-labeling strategies, but these pseudo-labels tend to introduce noise under extreme sparsity. To simultaneously avoid the impact of pseudo-label noise and enhance feature diversity, we propose a novel Adaptive Feature Generation (AdaptFG) model that generates features based on class names. This model integrates a pre-trained CLIP into a VAE-based feature generator, with its core innovation being an Adaptor that adaptively maps CLIP’s semantic embeddings to the object detector domain. Additionally, we introduce inter-class relationship reasoning in detector, which effectively mitigates misclassifications stemming from similar features. Extensive experimental results demonstrate that AdaptFG consistently outperforms state-of-the-art SAOD methods on the PASCAL VOC and MS COCO benchmarks. Shuilian Yao, Yu Liu 0012, Qi Jia 0001 |
AAAI | 2 |
| 2025 | Physics-Guided Sonar Image Fine-grained Recognition under Scarce AnnotationsabstractSonar image recognition is a key technology in underwater exploration systems. Compared with natural images, sonar images have fewer texture details and are easily affected by heavy noise, making it more challenging for specialists to distinguish the subtle differences among classes. In view of this, studying fine-grained classification methods for sonar images with scarce annotations is of significant importance. To address this issue, we propose a Physics-Guided Teacher-Student (PGTS) framework to explore the unique physical information of sonar images while simultaneously mitigating the effects of limited annotations. First, PGTS reconstructs sonar signals through physical simulation and a specially designed physics-guided feature generation module, which allows it to bypass the time-consuming physical simulation during inference. Then, we design a multi-modal teacher model combines the reconstructed sonar signals and sonar images to extract discriminative features to generate robust pseudo labels for fine-grained target categories. Finally, the knowledge is transferred to a single-modal student model through consistency loss. Under the joint constraints of the teacher model and the reconstructed sonar physical signals, the student model continuously improves its performance in annotation-scarce scenarios. Notably, when merely 1% of the data is labeled, our method outperforms other state-of-the-art approaches by 12.46% in terms of accuracy. Chengzhou Li, Qi Jia 0001, Jinyuan Liu 0001, Zhiying Jiang, Longhan Feng, Yu Liu 0012, Zhongxuan Luo, Xin Fan 0001 |
ACM Multimedia | 7 |
| 2025 | 3D-NLM: Voxel-based non-local means for 3D point cloud noise detection and smoothing
Weimin Wang 0007, Yu Liu 0012, Qiong Chang |
Comput. Graph. | 4 |
| 2025 | Bilevel progressive homography estimation via correlative region-focused transformer
Qi Jia 0001, Xiaomei Feng, Wei Zhang 0339, Yu Liu 0012, Nan Pu, Nicu Sebe |
Comput. Vis. Image Underst. | 4 |
| 2025 | Refining Pseudo Labeling via Multi-Granularity Confidence Alignment for Unsupervised Cross Domain Object DetectionabstractMost state-of-the-art object detection methods suffer from poor generalization due to the domain shift between training and testing datasets. To resolve this challenge, unsupervised cross domain object detection is proposed to learn an object detector for an unlabeled target domain by transferring knowledge from an annotated source domain. Promising results have been achieved via Mean Teacher, however, pseudo labeling which is the bottleneck of mutual learning remains to be further explored. In this study, we find that confidence misalignment of the predictions, including category-level overconfidence, instance-level task confidence inconsistency, and image-level confidence misfocusing, leading to the injection of noisy pseudo labels in the training process, will bring suboptimal performance. Considering the above issue, we present a novel general framework termed Multi-Granularity Confidence Alignment Mean Teacher (MGCAMT) for unsupervised cross domain object detection, which alleviates confidence misalignment across category-, instance-, and image-levels simultaneously to refine pseudo labeling for better teacher-student learning. Specifically, to align confidence with accuracy at category level, we propose Classification Confidence Alignment (CCA) to model category uncertainty based on Evidential Deep Learning (EDL) and filter out the category incorrect labels via an uncertainty-aware selection strategy. Furthermore, we design Task Confidence Alignment (TCA) to mitigate the instance-level misalignment between classification and localization by enabling each classification feature to adaptively identify the optimal feature for regression. Finally, we develop imagery Focusing Confidence Alignment (FCA) adopting another way of pseudo label learning, i.e., we use the original outputs from the Mean Teacher network for supervised learning without label assignment to achieve a balanced perception of the image's spatial layout. When these three procedures are integrated into a single framework, they mutually benefit to improve the final performance from a cooperative learning perspective. Extensive experiments across multiple scenarios demonstrate that our method outperforms large foundational models, and surpasses other state-of-the-art approaches by a large margin. Jiangming Chen, Li Liu 0002, Wanxia Deng, Zhen Liu 0004, Yu Liu 0012, Yingmei Wei, Yongxiang Liu |
IEEE Trans. Image Process. | 5 |
| 2025 | Rectangling for Stitched Image via Pixel-Wise Deformation LearningabstractImage rectangling involves filling in the blanks created during image stitching through deformation techniques. However, existing methods still struggle with incomplete filling and distortion of content, ultimately affecting the overall visual impression and potentially hindering subsequent tasks such as recognition. In this work, we design a pixel-wise deformation framework that utilizes explicit edge guidance to maintain consistency of texture and structure, yielding rectangular images with natural structure. Specifically, we decouple motion into region-level and pixel-level components through uniform mesh warping and pixel-wise deformation to precisely rearrange the spatial distribution of all pixels. Uniform deformation preserves local structure within divided patches, while pixel-wise motion coordinates the consistency between patches. Their combination provides robust and accurate pixel-wise offsets for structure-preserved rectangling. To further bolster the consistency of structure and texture, we leverage edge information to establish structural constraints and design an edge-guided enhancement module to aid in restoring fine texture details. Additionally, stitched images encompass both meaningful content and blank spaces, we innovatively incorporate a mask predictor, which acts as a guiding beacon, directing the network's attention solely towards content-rich regions to facilitate precise pixel-wise motion estimation. Experimental results demonstrate that our approach achieves state-of-the-art performance in rectifying irregular boundaries while contributing to downstream visual perception tasks. Xiaomei Feng, Qi Jia 0001, Yu Liu 0012, Weimin Wang 0007, Yuqing Liu 0001, Xinwei Xue |
IEEE Trans. Multim. | 3 |
| 2024 | Novel Class Discovery for Ultra-Fine-Grained Visual CategorizationabstractUltra-fine-grained visual categorization (Ultra-FGVC) aims at distinguishing highly similar sub-categories within fine-grained objects, such as different soybean cultivars. Compared to traditional fine-grained visual categorization, Ultra-FGVC encounters more hurdles due to the small inter-class and large intra-class variation. Given these challenges, relying on human annotation for Ultra-FGVC is impractical. To this end, our work introduces a novel task termed Ultra-Fine-Grained Novel Class Discovery (UFG-NCD), which leverages partially annotated data to identify new categories of unlabeled images for Ultra-FGVC. To tackle this problem, we devise a Region-Aligned Proxy Learning (RAPL) framework, which comprises a Channel-wise Region Alignment (CRA) module and a Semi-Supervised Proxy Learning (SemiPL) strategy. The CRA module is designed to extract and utilize discriminative features from local regions, facilitating knowledge transfer from labeled to unlabeled classes. Furthermore, SemiPL strengthens representation learning and knowledge transfer with proxy-guided supervised learning and proxy-guided contrastive learning. Such techniques leverage class distribution information in the embedding space, improving the mining of subtle differences between labeled and unlabeled ultra-fine-grained classes. Extensive experiments demonstrate that RAPL significantly outperforms baselines across various datasets, indicating its effectiveness in handling the challenges of UFG-NCD. Code is available at https://github.com/SSDUT-Caiyq/UFG-NCD. Yu Liu 0012, Yaqi Cai, Qi Jia 0001, Binglin Qiu, Weimin Wang 0007, Nan Pu |
CVPR | 1 |
| 2024 | Depth-Guided Dominant Plane Perception for Unsupervised Homography EstimationabstractHomography describes the mapping relations of the same plane across views. In scenarios with multiple planes, single homography estimation aims to obtain the optimal solution generated by the largest consistent plane to obey the coplanar constraints. However, existing methods typically consider all planes equally, neglecting the negative impact of regions that differ significantly from the largest approximate planar areas (dominant plane). In this work, we propose a depth-guided dominant plane perception network to achieve unsupervised homography estimation with additional attention on the dominant plane. Specifically, we leverage the depth-wise prior to adaptively detecting the approximate dominant plane, invoking essential scene structures for unsupervised homography estimation. Then, we enhance the corresponding features of the dominant plane and explore their correlations through a specially designed perceptual module. Finally, we employ dominant plane perception on multi-scale features progressively to estimate the homography in a coarse-to-fine manner. Extensive experiments on a large parallax dataset demonstrate that our method improves the alignment performance by 10.29%, yielding more accurate alignment than previous competitive methods. Xiaomei Feng, Qi Jia 0001, Yu Liu 0012, Xin Fan 0001, Longin Jan Latecki |
ICASSP | 3 |
| 2024 | Attribution-Based Scanline Perturbation Attack on 3d Detectors of Lidar Point CloudsabstractLiDAR point cloud data is widely utilized in autonomous driving systems and has significantly improved the 3D detection performance with well-designed deep neural network models. However, due to the complexity of real-world environments and model vulnerability, false detections or malicious attacks may cause severe accidents in unseen situations. In this paper, we propose a novel attack approach, Attribution-based Scanline Perturbation (ASP), an efficient and physically possible adversarial attack method for 3D detectors. ASP first utilizes attribution methods to identify critical points for the detection model and perturbs them along the laser beams by simulating the situation in which particles exist between the LiDAR sensor and objects, which can actually occur in snow or sandstorm weather. Extensive experiments on practical 3D detectors validate the effectiveness of our approach in misleading the model and causing both false and missed detections. Ziyang Yu 0004, Qiong Chang, Yu Liu 0012, Weimin Wang 0007 |
ICASSP | 4 |
| 2024 | CSCNet: Class-Specified Cascaded Network for Compositional Zero-Shot LearningabstractAttribute and object (A-O) disentanglement is a fundamental and critical problem for Compositional Zero-shot Learning (CZSL), whose aim is to recognize novel A-O compositions based on foregone knowledge. Existing methods based on disentangled representation learning lose sight of the contextual dependency between the A-O primitive pairs. Inspired by this, we propose a novel A-O disentangled framework for CZSL, namely Class-specified Cascaded Network (CSC-Net). The key insight is to firstly classify one primitive and then specifies the predicted class as a priori for guiding another primitive recognition in a cascaded fashion. To this end, CSCNet constructs Attribute-to-Object and Object-to- Attribute cascaded branches, in addition to a composition branch modeling the two primitives as a whole. Notably, we devise a parametric classifier (ParamCls) to improve the matching between visual and semantic embeddings. By improving the A-O disentanglement, our framework achieves superior results than previous competitive methods. Yanyi Zhang, Qi Jia 0001, Xin Fan 0001, Yu Liu 0012 |
ICASSP | 4 |
| 2024 | Sketch-Based 3D Shape Retrieval With Multi-View Fusion TransformerabstractSketch-based 3D shape retrieval aims to retrieve similar 3D shapes given a 2D sketch query. Although this task has been studied for years, the inherent cross-modal gap and data imbalance between 2D sketches and 3D shapes remain challenging. To address the problems, we propose a simple and effective framework based on Multi-view Fusion Transformer. To be specific, we project 3D shapes into twelve distinct views, and their CNN features are combined with position embeddings, passing together into a transformer encoder to learn view weights. Then we process them through average pooling and MLPs to obtain the final 3D shape representation. Furthermore, to narrow the data imbalance between 2D sketches and 3D shapes, affine transformation and elastic deformation are fully utilized for sketch augmentation, so as to extract more comprehensive sketch features for feature matching with the multi-view 3D shape representation. Extensive experiments on SHREC13, SHREC14 and PART-SHREC14 datasets demonstrate our method achieves superior performance than previous competitive methods. Cunjuan Zhu, Dongdong Cui, Qi Jia 0001, Weimin Wang 0007, Yu Liu 0012, Michael S. Lew |
ICASSP | 5 |
| 2024 | Fuzzy Boundary-Guided Network for Camouflaged Object DetectionabstractCamouflaged object detection (COD) is a challenging task that identifies camouflaged objects from highly similar backgrounds. Existing methods typically treat the whole object equally while neglecting the indistinguishable regions that require more attention than other regions. In this paper, we propose a Fuzzy Boundary-Guided Network (FBG-Net) for camouflaged object detection, which mimics the human behavior that pays more attention to these low-confidence regions when observing objects. Specifically, we devise two main building blocks: (1) Mixed Semantics Aggregation Module (MSAM) to integrate boundary and texture features cumulatively in the high-to-low scales, and (2) Fuzzy Boundary-Guided Module (FBGM) to locate and enhance the low-confidence regions under the guidance of fuzzy boundary. Extensive experiments demonstrate the effectiveness of FBG-Net with superior performance to existing state-of-the-art methods. Code is available at https://github.com/YAOSL98/FBG-Net. Qi Jia 0001, Shuilian Yao, Youcan Xu, Yu Liu 0012, Dehao Kong, Longin Jan Latecki |
ICME | 4 |
| 2024 | Joint edge detection learning for recurrent homography estimationabstractHomography estimation plays a pivotal role in aligning image pairs across multiple viewpoints. Existing methods focus mainly on texture alignment, whereas overlooking the influence of geometric structures, thereby resulting in inaccurate homography estimation. In this paper, we propose a novel recurrent homography estimation framework with joint edge detection learning. We find that edge detection explores extra anchors for homography estimation, and meanwhile homography provides complementary information of cross views for edge detection refinement. Unlike traditional edge detection applied to individual images, our approach establishes structural consistency constraints to reinforce mutual edges while suppressing unreliable structures. Specifically, the detected edges guide and enhance the texture features through a specifically designed edge-aware fusion module. Ultimately, we recurrently compute the correlation of fusion features from small to large scales for homography regression. Our experimental results demonstrate that the proposed method reduces the matching error by 41.7% than state-of-the-art methods. Furthermore, our network excels in detecting edges with extensive details even under dramatic perspective changes. Code is available at https://github.com/edmandzhao/edge-detection-for-RHE. Qi Jia 0001, Zikun Zhao, Xiaomei Feng, Jinyuan Liu 0001, Yu Liu 0012, Xinwei Xue |
ICME | 5 |
| 2024 | MergeNet: Explicit Mesh Reconstruction from Sparse Point Clouds via Edge PredictionabstractThis paper introduces a novel method for reconstructing meshes from sparse point clouds by predicting edge connection. Existing implicit methods usually produce superior smooth and watertight meshes due to the isosurface extraction algorithms (e.g., Marching Cubes). However, these methods become memory and computationally intensive with increasing resolution. Explicit methods are more efficient by directly forming the face from points. Nevertheless, the challenge of selecting appropriate faces from enormous candidates often leads to undesirable faces and holes. Moreover, the reconstruction performance of both approaches tends to degrade when the point cloud gets sparse. To this end, we propose MEsh Reconstruction via edGE (MergeNet), which converts mesh reconstruction into local connectivity prediction problems. Specifically, MergeNet learns to extract the features of candidate edges and regress their distances to the underlying surface. Consequently, the predicted distance is utilized to filter out edges that lay on surfaces. Finally, the meshes are reconstructed by refining the triangulations formed by these edges. Extensive experiments on synthetic and real-scanned datasets demonstrate the superiority of MergeNet to SoTA explicit methods. Weimin Wang 0007, Yingxu Deng, Zezeng Li, Yu Liu 0012, Na Lei |
ICME | 4 |
| 2024 | Unseen No More: Unlocking the Potential of CLIP for Generative Zero-shot HOI DetectionabstractZero-shot human-object interaction (HOI) detector is capable of generalizing to HOI categories even not encountered during training. Inspired by the impressive zero-shot capabilities offered by CLIP, latest methods strive to leverage CLIP embeddings for improving zero-shot HOI detection. However, these embedding-based methods train the classifier on seen classes only, inevitably resulting in seen-unseen confusion for the model during inference. Besides, we find that using prompt-tuning and adapters further increases the gap between seen and unseen accuracy. To tackle this challenge, we present the first generation-based model using CLIP for zero-shot HOI detection, coined HOIGen. It allows to unlock the potential of CLIP for feature generation instead of feature extraction only. To achieve it, we develop a CLIP-injected feature generator in accordance with the generation of human, object and union features. Then, we extract realistic features of seen samples and mix them with synthetic features together, allowing the model to train seen and unseen classes jointly. To enrich the HOI scores, we construct a generative prototype bank in a pairwise HOI recognition branch, and a multi-knowledge prototype bank in an image-wise HOI recognition branch, respectively. Extensive experiments on HICO-DET benchmark demonstrate our HOIGen achieves superior performance for both seen and unseen classes under various zero-shot settings, compared with other top-performing methods. Code is available at: https://github.com/soberguo/HOIGen Yu Liu 0012, Weimin Wang 0007, Qi Jia 0001 |
ACM Multimedia | 2 |
| 2024 | Two Teachers Are Better Than One: Semi-supervised Elliptical Object Detection by Dual-Teacher Collaborative Guidance
Yu Liu 0012, Longhan Feng, Qi Jia 0001, Zezheng Liu, Zi-Huang Cao |
ACM Multimedia | 1 |
| 2024 | PMGNet: Disentanglement and entanglement benefit mutually for compositional zero-shot learning
Yu Liu 0012, Yanyi Zhang, Qi Jia 0001, Weimin Wang 0007, Nan Pu, Nicu Sebe |
Comput. Vis. Image Underst. | 1 |
| 2024 | Edge-Aware Correlation Learning for Unsupervised Progressive Homography EstimationabstractHomography estimation aligns image pairs in cross-views, which is a crucial and fundamental computer vision problem. Existing methods only consider correspondences of texture features for homography estimation, leading to unpleasant artifacts and misalignments introduced by mismatches, especially for low-texture image pairs. In contrast to others, we introduce intuitive structural information as an additional clue that is more sensitive to human vision and low-texture scenarios. In this paper, we propose an edge-aware unsupervised progressive network that couples texture and edge correlation to comprehensively explore potential matching features for homography estimation. To explore robust edge and texture features, we employ a multiscale network to capture feature pyramids with different receptive fields. Then, we design an edge-aware correlation module tailored for homography regression, which plugs in multiscale features to capture accurate correlation maps. Specifically, the edge-aware correlation module leverages the feature-selecting strategy for edge features to capture discriminative matching edges and further guides the texture correlation unit to focus on correctly matched textures. Finally, we leverage multiscale edge-aware correlation maps to predict homography progressively from coarse to fine. Experimental results demonstrate that our proposed method improves PSNR by 11.09% on the real large parallax dataset and reduces matching error by 32.04% on the synthetic COCO dataset, yielding more accurate alignment results than previous state-of-the-art methods. Xiaomei Feng, Qi Jia 0001, Zikun Zhao, Yu Liu 0012, Xinwei Xue, Xin Fan 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | COCA: COllaborative CAusal Regularization for Audio-Visual Question AnsweringabstractAudio-Visual Question Answering (AVQA) is a sophisticated QA task, which aims at answering textual questions over given video-audio pairs with comprehensive multimodal reasoning. Through detailed causal-graph analyses and careful inspections of their learning processes, we reveal that AVQA models are not only prone to over-exploit prevalent language bias, but also suffer from additional joint-modal biases caused by the shortcut relations between textual-auditory/visual co-occurrences and dominated answers. In this paper, we propose a COllabrative CAusal (COCA) Regularization to remedy this more challenging issue of data biases. Specifically, a novel Bias-centered Causal Regularization (BCR) is proposed to alleviate specific shortcut biases by intervening bias-irrelevant causal effects, and further introspect the predictions of AVQA models in counterfactual and factual scenarios. Based on the fact that the dominated bias impairing model robustness for different samples tends to be different, we introduce a Multi-shortcut Collaborative Debiasing (MCD) to measure how each sample suffers from different biases, and dynamically adjust their debiasing concentration to different shortcut correlations. Extensive experiments demonstrate the effectiveness as well as backbone-agnostic ability of our COCA strategy, and it achieves state-of-the-art performance on the large-scale MUSIC-AVQA dataset. Mingrui Lao, Nan Pu, Yu Liu 0012, Erwin M. Bakker, Michael S. Lew |
AAAI | 3 |
| 2023 | Learning Pixel-wise Alignment for Unsupervised Image StitchingabstractImage stitching aims to align a pair of images in the same view. Generating precise alignment with natural structures is challenging for image stitching, as there is no wider field-of-view image as a reference, especially in non-coplanar practical scenarios. In this paper, we propose an unsupervised image stitching framework, breaking through the coplanar constraints in homography estimation, yielding accurate pixel-wise alignment under limited overlapping regions. First, we generate a global transformation by an iterative dense feature matching combined with an error control strategy to alleviate the difference introduced by large parallax. Second, we propose a pixel-wise warping network embedded within a large-scale feature extractor and a correlative feature enhancement module to explicitly learn correspondences between the inputs, and generate accurate pixel-level offsets upon novel constraints on both overlapping and non-overlapping regions. Notably, we leverage the pixel-level offsets in the overlapping area to guide the adjustment in the non-overlapping area upon content and structure consistency constraints, rendering a natural transition between two regions and distortions suppression over the entire stitched image. The proposed method achieves state-of-the-art performance that surpasses both traditional and deep learning approaches by a large margin. It also achieves the shortest execution time and has the best generalization ability on the traditional dataset. Qi Jia 0001, Xiaomei Feng, Yu Liu 0012, Xin Fan 0001, Longin Jan Latecki |
ACM Multimedia | 3 |
| 2023 | Multi-Domain Lifelong Visual Question Answering via Self-Critical DistillationabstractVisual Question Answering (VQA) has achieved significant success over the last few years, while most studies focus on training a VQA model on a stationary domain (e.g., a given dataset). In real-world application scenarios, however, these methods are often inefficient because VQA systems are always supposed to extend their knowledge and meet the ever-changing demands of users. In this paper, we introduce a new and challenging multi-domain lifelong VQA task, dubbed MDL-VQA, which encourages the VQA model to continuously learn across multiple domains while mitigating the forgetting on previously-learned domains. Furthermore, we propose a novel replay-free Self-Critical Distillation (SCD) framework tailor-made for MDL-VQA, which alleviates forgetting issue via transferring previous-domain knowledge from teacher to student models. First, we propose to introspect the teacher's understanding over original and counterfactual samples, thereby creating informative instance-relevant and domain-relevant knowledge for logits-based distillation. Second, on the side of feature-based distillation, we propose to introspect the reasoning behavior of student model to establish the harmful domain-specific knowledge acquired in current domain, and further leverage the metric learning strategy to encourage student to learn useful knowledge in new domain. Extensive experiments demonstrate that SCD framework outperforms state-of-the-art competitors with different training orders. Mingrui Lao, Nan Pu, Yu Liu 0012, Zhun Zhong, Erwin M. Bakker, Nicu Sebe, Michael S. Lew |
ACM Multimedia | 3 |
| 2023 | Broaden Your Positives: A General Rectification Approach for Novel Class Discovery
Yaqi Cai, Nan Pu, Qi Jia 0001, Weimin Wang 0007, Yu Liu 0012 |
PRCV (4) | 5 |
| 2023 | Deep Learning for Instance Retrieval: A SurveyabstractIn recent years a vast amount of visual content has been generated and shared from many fields, such as social media platforms, medical imaging, and robotics. This abundance of content creation and sharing has introduced new challenges, particularly that of searching databases for similar content - Content Based Image Retrieval (CBIR) - a long-established research area in which improved efficiency and accuracy are needed for real-time retrieval. Artificial intelligence has made progress in CBIR and has significantly facilitated the process of instance search. In this survey we review recent instance retrieval works that are developed based on deep learning algorithms and techniques, with the survey organized by deep feature extraction, feature embedding and aggregation methods, and network fine-tuning strategies. Our survey considers a wide variety of recent methods, whereby we identify milestone work, reveal connections among various methods and present the commonly used benchmarks, evaluation results, common challenges, and propose promising future directions. Wei Chen 0072, Yu Liu 0012, Weiping Wang 0002, Erwin M. Bakker, Theodoros Georgiou 0001, Paul W. Fieguth, Li Liu 0002, Michael S. Lew |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Lifelong Fine-Grained Image RetrievalabstractFine-grained image retrieval has been extensively explored in a zero-shot manner. A deep model is trained on the seen part and then evaluated the generalization performance on the unseen part. However, this setting is infeasible for many real-world applications since (1) the retrieval dataset can be non-fixed so that new data are added constantly, and (2) data samples of the seen categories are also common in practice and are important for evaluation. In this paper, we explore lifelong fine-grained image retrieval (LFGIR), which learns continuously on a sequence of new tasks with data from different datasets. We first use knowledge distillation to minimize catastrophic forgetting on old tasks. Training continuously on different datasets causes large domain shifts between the old and new tasks while image retrieval is sensitive to even small shifts in the features. This tends to weaken the effectiveness of knowledge distillation by the frozen teacher. To mitigate the impact of domain shifts, we use the network inversion method to generate images of the old tasks. In addition, we design an on-the-fly teacher which transfers knowledge captured on a new task to the student to improve better generalization performance, thereby achieving a better balance between old and new tasks in the end. We name the whole framework as Dual Knowledge Distillation (DKD), whose efficacy is demonstrated by extensive experimental results on sequential tasks including seven datasets. Wei Chen 0072, Haoyang Xu, Nan Pu, Yu Liu 0012, Mingrui Lao, Weiping Wang 0002, Li Liu 0002, Michael S. Lew |
IEEE Trans. Multim. | 4 |
| 2023 | Residual Tuning: Toward Novel Category Discovery Without LabelsabstractDiscovering novel visual categories from a set of unlabeled images is a crucial and essential capability for intelligent vision systems since it enables them to automatically learn new concepts with no need for human-annotated supervision anymore. To tackle this problem, existing approaches first pretrain a neural network with a set of labeled images and then fine-tune the network to cluster unlabeled images into a few categorical groups. However, their unified feature representation hits a tradeoff bottleneck between feature preservation on labeled data and feature adaptation on unlabeled data. To circumvent this bottleneck, we propose a residual-tuning approach, which estimates a new residual feature from the pretrained network and adds it with a previous basic feature to compute the clustering objective together. Our disentangled representation approach facilitates adjusting visual representations for unlabeled images and overcoming forgetting old knowledge acquired from labeled images, with no need of replaying the labeled images again. In addition, residual-tuning is an efficient solution, adding few parameters and consuming modest training time. Our results on three common benchmarks show consistent and considerable gains over other state-of-the-art methods, and further reduce the performance gap to the fully supervised learning setup. Moreover, we explore two extended scenarios, including using fewer labeled classes and continually discovering more unlabeled sets, where the results further signify the advantages and effectiveness of our residual-tuning approach against previous approaches. Our code is available at https://github.com/liuyudut/ResTune. Yu Liu 0012, Tinne Tuytelaars |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2022 | Segment, Magnify and Reiterate: Detecting Camouflaged Objects the Hard WayabstractIt is challenging to accurately detect camouflaged objects from their highly similar surroundings. Existing methods mainly leverage a single-stage detection fashion, while neglecting small objects with low-resolution fine edges requires more operations than the larger ones. To tackle camouflaged object detection (COD), we are inspired by humans attention coupled with the coarse-to-fine detection strategy, and thereby propose an iterative refinement framework, coined SegMaR, which integrates Segment, Magnify and Reiterate in a multi-stage detection fashion. Specifically, we design a new discriminative mask which makes the model attend on the fixation and edge regions. In addition, we leverage an attention-based sampler to magnify the object region progressively with no need of enlarging the image size. Extensive experiments show our SegMaR achieves remarkable and consistent improvements over other state-of-the-art methods. Especially, we surpass two competitive methods 7.4% and 20.0% respectively in average over standard evaluation metrics on small camouflaged objects. Additional studies provide more promising insights into Seg-MaR, including its effectiveness on the discriminative mask and its generalization to other network architectures. Code is available at https://github.com/dlut-dimt/SegMaR. Qi Jia 0001, Shuilian Yao, Yu Liu 0012, Xin Fan 0001, Risheng Liu, Zhongxuan Luo |
CVPR | 3 |
| 2022 | Meta Reconciliation Normalization for Lifelong Person Re-IdentificationabstractLifelong person re-identification (LReID) is a challenging and emerging task, which concerns the ReID capability on both seen and unseen domains after learning across different domains continually. Existing works on LReID are devoted to introducing commonly-used lifelong learning approaches, while neglecting a serious side effect caused by using normalization layers in the context of domain-incremental learning. In this work, we aim to raise awareness of the importance of training proper batch normalization layers by proposing a new meta reconciliation normalization (MRN) method specifically designed for tackling LReID. Our MRN consists of grouped mixture standardization and additive rectified rescaling components, which are able to automatically maintain an optimal balance between domain-dependent and domain-independent statistics, and even adapt MRN for different testing instances. Furthermore, inspired by synaptic plasticity in human brain, we present a MRN-based meta-learning framework for mining the meta-knowledge shared across different domains, even without replaying any previous data, and further improve the model's LReID ability with theoretical analyses. Our method achieves new state-of-the-art performances on both balanced and imbalanced LReID benchmarks. Nan Pu, Yu Liu 0012, Wei Chen 0072, Erwin M. Bakker, Michael S. Lew |
ACM Multimedia | 2 |
| 2022 | Feature Estimations Based Correlation Distillation for Incremental Image RetrievalabstractDeep learning for fine-grained image retrieval in an incremental context is less investigated. In this paper, we explore this task to realize the model’s continuous retrieval ability. That means, the model enables to perform well on new incoming data and reduce forgetting of the knowledge learned on preceding old tasks. For this purpose, we distill semantic correlations knowledge among the representations extracted from the new data only so as to regularize the parameters updates using the teacher-student framework. In particular, for the case of learning multiple tasks sequentially, aside from the correlations distilled from the penultimate model, we estimate the representations for all prior models and further their semantic correlations by using the representations extracted from the new data. To this end, the estimated correlations are used as an additional regularization and further prevent catastrophic forgetting over all previous tasks, and it is unnecessary to save the stream of models trained on these tasks. Extensive experiments demonstrate that the proposed method performs favorably for retaining performance on the already-trained old tasks and achieving good accuracy on the current task when new data are added at once or sequentially. Wei Chen 0072, Yu Liu 0012, Nan Pu, Weiping Wang 0002, Li Liu 0002, Michael S. Lew |
IEEE Trans. Multim. | 2 |
| 2021 | Lifelong Person Re-Identification via Adaptive Knowledge AccumulationabstractPerson re-identification (ReID) methods always learn through a stationary domain that is fixed by the choice of a given dataset. In many contexts (e.g., lifelong learning), those methods are ineffective because the domain is continually changing in which case incremental learning over multiple domains is required potentially. In this work we explore a new and challenging ReID task, namely lifelong person re-identification (LReID), which enables to learn continuously across multiple domains and even generalise on new and unseen domains. Following the cognitive processes in the human brain, we design an Adaptive Knowledge Accumulation (AKA) framework that is endowed with two crucial abilities: knowledge representation and knowledge operation. Our method alleviates catastrophic forgetting on seen domains and demonstrates the ability to generalize to unseen domains. Correspondingly, we also provide a new and large-scale benchmark for LReID. Extensive experiments demonstrate our method outperforms other competitors by a margin of 5.8% mAP in generalising evaluation. The codes will be available at https://github.com/TPCD/LifelongReID. Nan Pu, Wei Chen 0072, Yu Liu 0012, Erwin M. Bakker, Michael S. Lew |
CVPR | 3 |
| 2021 | A Language Prior Based Focal Loss for Visual Question AnsweringabstractAccording to current research, one of the major challenges in Visual Question Answering (VQA) models is the overdependence on language priors (and neglect of the visual modality). VQA models tend to predict answers only based on superficial correlations between the first few words in question and frequency of related answer candidates. To address this issue, we propose a novel Language Prior based Focal Loss (LP-Focal Loss) by rescaling the standard cross entropy loss. Specifically, we employ a question-only branch to capture the language biases for each answer candidate based on the corresponding question input. Then, the LP-Focal Loss dynamically assigns lower weights to biased answers when computing the training loss, thereby reducing the contribution of more-biased instances in the train split. Extensive experiments show that the LP-Focal Loss can be generally applied to common baseline VQA models, and achieves significantly better performance on the VQA-CP v2 dataset, with an overall 18% accuracy boost over benchmark models. Mingrui Lao, Yanming Guo, Yu Liu 0012, Michael S. Lew |
ICME | 3 |
| 2021 | Underwater Species Detection using Channel Sharpening AttentionabstractWith the continuous exploration of marine resources, underwater artificial intelligent robots play an increasingly important role in the fish industry. However, the detection of underwater objects is a very challenging problem due to the irregular movement of underwater objects, the occlusion of sand and rocks, the diversity of water illumination, and the poor visibility and low color contrast in the underwater environment. In this article, we first propose a real-world underwater object detection dataset (UODD), which covers more than 3K images of the most common aquatic products. Then we propose Channel Sharpening Attention Module (CSAM) as a plug-and-play module to further fuse high-level image information, providing the network with the privilege of selecting feature maps. Fusion of original images through CSAM can improve the accuracy of detecting small and medium objects, thereby improving the overall detection accuracy. We also use Water-Net as a preprocessing method to remove the haze and color cast in complex underwater scenes, which shows a satisfactory detection result on small-sized objects. In addition, we use the class weighted loss as the training loss, which can accurately describe the relationship between classification and precision of bounding boxes of targets, and the loss function converges faster during the training process. Experimental results show that the proposed method reaches a maximum AP of 50.1%, outperforming other traditional and state-of-the-art detectors. In addition, our model only needs an average inference time of 25.4 ms per image, which is quite fast and might suit the real-time scenario. Lihao Jiang, Yi Wang 0037, Qi Jia 0001, Shengwei Xu, Yu Liu 0012, Xin Fan 0001, Risheng Liu, Xinwei Xue, Ruili Wang 0001 |
ACM Multimedia | 5 |
| 2021 | From Superficial to Deep: Language Bias driven Curriculum Learning for Visual Question AnsweringabstractMost Visual Question Answering (VQA) models are faced with language bias when learning to answer a given question, thereby failing to understand multimodal knowledge simultaneously. Based on the fact that VQA samples with different levels of language bias contribute differently for answer prediction, in this paper, we overcome the language prior problem by proposing a novel Language Bias driven Curriculum Learning (LBCL) approach, which employs an easy-to-hard learning strategy with a novel difficulty metric Visual Sensitive Coefficient (VSC). Specifically, in the initial training stage, the VQA model mainly learns the superficial textual correlations between questions and answers (easy concept) from more-biased examples, and then progressively focuses on learning the multimodal reasoning (hard concept) from less-biased examples in the following stages. The curriculum selection of examples on different stages is according to our proposed difficulty metric VSC, which is to evaluate the difficulty driven by the language bias of each VQA sample. Furthermore, to avoid the catastrophic forgetting of the learned concept during the multi-stage learning procedure, we propose to integrate knowledge distillation into the curriculum learning framework. Extensive experiments show that our LBCL can be generally applied to common VQA baseline models, and achieves remarkably better performance on the VQA-CP v1 and v2 datasets, with an overall 20% accuracy boost over baseline models. Mingrui Lao, Yanming Guo, Yu Liu 0012, Wei Chen 0072, Nan Pu, Michael S. Lew |
ACM Multimedia | 3 |
| 2021 | Multi-stage hybrid embedding fusion network for visual question answering
Mingrui Lao, Yanming Guo, Nan Pu, Wei Chen 0072, Yu Liu 0012, Michael S. Lew |
Neurocomputing | 5 |
| 2021 | Integrating information theory and adversarial learning for cross-modal retrievalabstractAccurately matching visual and textual data in cross-modal retrieval has been widely studied in the multimedia community. To address these challenges posited by the heterogeneity gap and the semantic gap, we propose integrating Shannon information theory and adversarial learning. In terms of the heterogeneity gap, we integrate modality classification and information entropy maximization adversarially. For this purpose, a modality classifier (as a discriminator) is built to distinguish the text and image modalities according to their different statistical properties. This discriminator uses its output probabilities to compute Shannon information entropy, which measures the uncertainty of the modality classification it performs. Moreover, feature encoders (as a generator) project uni-modal features into a commonly shared space and attempt to fool the discriminator by maximizing its output information entropy. Thus, maximizing information entropy gradually reduces the distribution discrepancy of cross-modal features, thereby achieving a domain confusion state where the discriminator cannot classify two modalities confidently. To reduce the semantic gap, Kullback-Leibler (KL) divergence and bi-directional triplet loss are used to associate the intra- and inter-modality similarity between features in the shared space. Furthermore, a regularization term based on KL-divergence with temperature scaling is used to calibrate the biased label classifier caused by the data imbalance issue. Extensive experiments with four deep models on four benchmarks are conducted to demonstrate the effectiveness of the proposed approach. Wei Chen 0072, Yu Liu 0012, Erwin M. Bakker, Michael S. Lew |
Pattern Recognit. | 2 |
| 2020 | On the Exploration of Incremental Learning for Fine-grained Image Retrieval
Wei Chen 0072, Yu Liu 0012, Weiping Wang 0002, Tinne Tuytelaars, Erwin M. Bakker, Michael S. Lew |
BMVC | 2 |
| 2020 | A Novel Baseline for Zero-shot Learning via Adversarial Visual-Semantic Embedding
Yu Liu 0012, Tinne Tuytelaars |
BMVC | 1 |
| 2020 | More Classifiers, Less Forgetting: A Generic Multi-classifier Paradigm for Incremental Learning
Yu Liu 0012, Sarah Parisot, Gregory Slabaugh, Xu Jia 0012, Ales Leonardis, Tinne Tuytelaars |
ECCV (26) | 1 |
| 2020 | Dual Gaussian-based Variational Subspace Disentanglement for Visible-Infrared Person Re-IdentificationabstractVisible-infrared person re-identification (VI-ReID) is a challenging and essential task in night-time intelligent surveillance systems. Except for the intra-modality variance that RGB-RGB person re-identification mainly overcomes, VI-ReID suffers from additional inter-modality variance caused by the inherent heterogeneous gap. To solve the problem, we present a carefully designed dual Gaussian-based variational auto-encoder (DG-VAE), which disentangles an identity-discriminable and an identity-ambiguous cross-modality feature subspace, following a mixture-of-Gaussians (MoG) prior and a standard Gaussian distribution prior, respectively. Disentangling cross-modality identity-discriminable features leads to more robust retrieval for VI-ReID. To achieve efficient optimization like conventional VAE, we theoretically derive two variational inference terms for the MoG prior under the supervised setting, which not only restricts the identity-discriminable subspace so that the model explicitly handles the cross-modality intra-identity variance, but also enables the MoG distribution to avoid posterior collapse. Furthermore, we propose a triplet swap reconstruction (TSR) strategy to promote the above disentangling process. Extensive experiments demonstrate that our method outperforms state-of-the-art methods on two VI-ReID datasets. Codes will be available at https://github.com/TPCD/DG-VAE. Nan Pu, Wei Chen 0072, Yu Liu 0012, Erwin M. Bakker, Michael S. Lew |
ACM Multimedia | 3 |
| 2020 | A Deep Multi-Modal Explanation Model for Zero-Shot LearningabstractZero-shot learning (ZSL) has attracted significant attention due to its capabilities of classifying new images from unseen classes. To perform the classification task for ZSL, learning visual and semantic embeddings has been the main research approach in existing literature. At the same time, generating complementary explanations to justify the classification decision has remained largely unexplored. In this paper, we propose to address a new and challenging task, namely explainable zero-shot learning (XZSL), which aims to generate visual and textual explanations to support the classification decision. To accomplish this task, we build a novel Deep Multi-modal Explanation (DME) model that incorporates a joint visual-attribute embedding module and a multi-channel explanation module in an end-to-end fashion. In contrast to existing ZSL approaches, our visual-attribute embedding is associated not only with the decision, but also with new visual and textual explanations. For visual explanations, we first capture several attribute activation maps (AAM) and then merge them into a class activation map (CAM) that visually infers which region of an image is relevant to the class. Textual explanations are generated from the multi-channel explanation module, jointly integrating three long short-term memory models (LSTMs) each of which is conditioned on a different feature representation. Additionally, we suggest that the DME model can retain explanatory consistency for similar instances and explanatory diversity for diverse instances. We conduct qualitative and quantitative experiments to assess the model for ZSL classification and explanation. Specifically, the ablation studies verify the effectiveness of the components in our model. Our results on three well-known datasets are competitive with prior approaches. More importantly, the joint training of our embedding and explanation modules demonstrates mutual performance improvements between ZSL classification and explanation. We shed more light on DME to analyze and diagnose its advantages and limitations. Yu Liu 0012, Tinne Tuytelaars |
IEEE Trans. Image Process. | 1 |
| 2019 | An Empirical Study on Sensor-aware Design of Convolutional Neural Networks for P300 Speller in Brain Computer InterfaceabstractA Brain Computer Interface (BCI) character speller allows human-beings to directly spell characters using eye-gazes, thereby building communication between the human brain and a computer. Convolutional Neural Networks (CNNs) have achieved state-of-the-art results on the BCI character spelling accuracy. Unfortunately, to the best of our knowledge, it has not been studied whether the CNN should be designed differently to increase the spelling accuracy when the number of sensors used to acquire EEG signals is different. This paper performs an empirical study to investigate this issue. First, we show a motivational example which motivates us for this investigation. Then, we propose a method to design CNNs according to the number of sensors used in the BCI character speller. This method automatically configures a parametric CNN we have devised according to the given number of sensors. Experimental results on six datasets show that we need to design different CNNs when different number of sensors are used for the acquisition of EEG signals. Experimental results also show that our designed sensor-aware CNNs outperform other CNNs in terms of spelling accuracy in most cases. Our CNNs can increase the spelling accuracy achieved by other CNNs with up to 34%. Hongchang Shan, Yu Liu 0012, Todor P. Stefanov |
HSI | 2 |
| 2019 | Ensemble of Convolutional Neural Networks for P300 Speller in Brain Computer Interface
Hongchang Shan, Yu Liu 0012, Todor P. Stefanov |
ICANN (4) | 2 |
| 2019 | Domain Uncertainty Based On Information Theory for Cross-Modal Hash RetrievalabstractCross-modal hash retrieval has received considerable interest in the area of deep learning. Here hash codes of data of different modalities are learned where pair-wise loss functions control feature similarity in a shared embedding space. In this paper we improve on feature similarity by using Shannon's information entropy with respect to the modality information that is left in learning superior hash codes. We introduce a novel network for predicting the domain from the learned features while the protagonist network uses a loss function based on Shannon's information entropy to learn to maximize the domain uncertainty and therefore the information content. Additionally, according to the number of common labels between each similar image-text pair, we define a multi-level similarity matrix as supervisory information, which constrains all similar pairs with different weights. We show with extensive experiments that our novel approach to domain uncertainty leads to a cross-modal hash retrieval that outperforms the state-of-the-art. Wei Chen 0072, Nan Pu, Yu Liu 0012, Erwin M. Bakker, Michael S. Lew |
ICME | 3 |
| 2019 | CycleMatch: A cycle-consistent embedding network for image-text matching
Yu Liu 0012, Yanming Guo, Li Liu 0002, Erwin M. Bakker, Michael S. Lew |
Pattern Recognit. | 1 |
| 2019 | SwapGAN: A Multistage Generative Approach for Person-to-Person Fashion Style TransferabstractFashion style transfer has attracted significant attention because it both has interesting scientific challenges and it is also important to the fashion industry. This paper focuses on addressing a practical problem in fashion style transfer, person-to-person clothing swapping, which aims to visualize what the person would look like with the target clothes worn on another person instead of dressing them physically. This problem remains challenging due to varying pose deformations between different person images. In contrast to traditional nonparametric methods that blend or warp the target clothes for the reference person, in this paper we propose a multistage deep generative approach named SwapGAN that exploits three generators and one discriminator in a unified framework to fulfill the task end-to-end. The first and second generators are conditioned on a human pose map and a segmentation map, respectively, so that we can simultaneously transfer the pose style and the clothes style. In addition, the third generator is used to preserve the human body shape during the image synthesis process. The discriminator needs to distinguish two fake image pairs from the real image pair. The entire SwapGAN is trained by integrating the adversarial loss and the mask-consistency loss. The experimental results on the DeepFashion dataset demonstrate the improvements of SwapGAN over other existing approaches through both quantitative and qualitative evaluations. Moreover, we conduct ablation studies on SwapGAN and provide a detailed analysis about its effectiveness. Yu Liu 0012, Wei Chen 0072, Li Liu 0002, Michael S. Lew |
IEEE Trans. Multim. | 1 |
| 2018 | A Dual Prediction Network for Image CaptioningabstractGeneral captioning practice involves a single forward prediction, with the aim of predicting the word in the next timestep given the word in the current timestep. In this paper, we present a novel captioning framework, namely Dual Prediction Network (DPN), which is end-to-end trainable and addresses the captioning problem with dual predictions. Specifically, the dual predictions consist of a forward prediction to generate the next word from the current input word, as well as a backward prediction to reconstruct the input word using the predicted word. DPN has two appealing properties: 1) By introducing an extra supervision signal on the prediction, DPN can better capture the interplay between the input and the target; 2) Utilizing the reconstructed input, DPN can make another new prediction. During the test phase, we average both predictions to formulate the final target sentence. Experimental results on the MS COCO dataset demonstrate that, benefiting from the reconstruction step, both generated predictions in DPN outperform the predictions of methods based on the general captioning practice (single forward prediction), and averaging them can bring a further accuracy boost. Overall, DPN achieves competitive results with state-of-the-art approaches, across multiple evaluation metrics. Yanming Guo, Yu Liu 0012, Maaike de Boer, Li Liu 0002, Michael S. Lew |
ICME | 2 |
| 2018 | An Extensive Study of Cycle-Consistent Generative Networks for Image-to-Image TranslationabstractImage-to-image translation between different domains has been an important research direction, with the aim of arbitrarily manipulating the source image content to become similar to a target image. Recently, cycle-consistent generative network (CycleGAN) has become a fundamental approach for general-purpose image-to-image translation, while almost no work has examined what factors may influence its performance. To provide more insights, we propose two new models roughly based on CycleGAN, namely Long CycleGAN and Nest CycleGAN. First, Long CycleGAN cascades several generators to perform the domain translation in a long cycle. It shows the benefit of stacking more generators on the generation quality. In addition to the long cycle, Nest CycleGAN develops new inner cycles to bridge intermediate generators directly, which can help constrain the unsupervised mappings. In the experiments, we conduct qualitative and quantitative comparisons for tasks including photo↔label, photo↔sketch, and photo colorization. The quantitative and qualitative results demonstrate the effectiveness of our two proposed models. Yu Liu 0012, Yanming Guo, Wei Chen 0072, Michael S. Lew |
ICPR | 1 |
| 2018 | A Simple Convolutional Neural Network for Accurate P300 Detection and Character Spelling in Brain Computer InterfaceabstractA Brain Computer Interface (BCI) character speller allows human-beings to directly spell characters using eye-gazes, thereby building communication between the human brain and a computer. Convolutional Neural Networks (CNNs) have shown better performance than traditional machine learning methods for BCI signal recognition and its application to the character speller. However, current CNN architectures limit further accuracy improvements of signal detection and character spelling and also need high complexity to achieve competitive accuracy, thereby preventing the use of CNNs in portable BCIs. To address these issues, we propose a novel and simple CNN which effectively learns feature representations from both raw temporal information and raw spatial information. The complexity of the proposed CNN is significantly reduced compared with state-of-the-art CNNs for BCI signal detection. We perform experiments on three benchmark datasets and compare our results with those in previous research works which report the best results. The comparison shows that our proposed CNN can increase the signal detection accuracy by up to 15.61% and the character spelling accuracy by up to 19.35%. Hongchang Shan, Yu Liu 0012, Todor P. Stefanov |
IJCAI | 2 |
| 2018 | Learning Fluid FlowsabstractComputational Fluid Dynamics (CFD) simulations are able to produce complex and large outputs that accurately describe the physical properties of fluids and gases in various domains, such as air flow around a car, or the multi-phase flow inside an internal combustion engine. The simulation results, i.e. the flow fields, are often too complex to be analyzed directly. With the increasing number of simulations as well as their complexity, there is a need of automated processes that can analyze these complex outputs. In this paper, inspired by the success of convolutional neural networks (CNNs) in Computer Vision, we apply for the first time CNNs on CFD output. We show their capabilities in capturing and processing flow patterns. Furthermore, we design a novel CNN architecture tailored to the data produced by CFD simulations, as well as two conventional architectures and compare them. We propose and construct a new dataset of turbulent flow, within the application domain of steady flow around passenger cars. We use that dataset to evaluate and compare the proposed methods, on different tasks that depend on flow patterns. Finally, we compare our methods with a baseline k-nearest neighbor approach, tuned to be comparable to the state-of-the-art. Theodoros Georgiou 0001, Markus Olhofer, Yu Liu 0012, Thomas Bäck, Michael S. Lew |
IJCNN | 4 |
| 2018 | CNN-RNN: a large-scale hierarchical image classification frameworkabstractObjects are often organized in a semantic hierarchy of categories, where fine-level categories are grouped into coarse-level categories according to their semantic relations. While previous works usually only classify objects into the leaf categories, we argue that generating hierarchical labels can actually describe how the leaf categories evolved from higher level coarse-grained categories, thus can provide a better understanding of the objects. In this paper, we propose to utilize the CNN-RNN framework to address the hierarchical image classification task. CNN allows us to obtain discriminative features for the input images, and RNN enables us to jointly optimize the classification of coarse and fine labels. This framework can not only generate hierarchical labels for images, but also improve the traditional leaf-level classification performance due to incorporating the hierarchical information. Moreover, this framework can be built on top of any CNN architecture which is primarily designed for leaf-level classification. Accordingly, we build a high performance network based on the CNN-RNN paradigm which outperforms the original CNN (wider-ResNet) and also the current state-of-the-art. In addition, we investigate how to utilize the CNN-RNN framework to improve the fine category classification when a fraction of the training data is only annotated with coarse labels. Experimental results demonstrate that CNN-RNN can use the coarse-labeled training data to improve the classification of fine categories, and in some cases it even surpasses the performance achieved by fully annotated training data. This reveals that, CNN-RNN can alleviate the challenge of specialized and expensive annotation of fine labels. Yanming Guo, Yu Liu 0012, Erwin M. Bakker, Yuanhao Guo, Michael S. Lew |
Multim. Tools Appl. | 2 |
| 2018 | Fusion that matters: convolutional fusion networks for visual recognitionabstractIn recent years, deep learning has been successfully applied to diverse multimedia research areas, with the aim of learning powerful and informative representations for a variety of visual recognition tasks. In this work, we propose convolutional fusion networks (CFN) to integrate multi-level deep features and fuse a richer visual representation. Despite recent advances in deep fusion networks, they still have limitations due to expensive parameters and weak fusion modules. Instead, CFN uses 1 × 1 convolutional layers and global average pooling to generate side branches with few parameters, and employs a locally-connected fusion module, which can learn adaptive weights for different side branches and form a better fused feature. Specifically, we introduce three key components of the proposed CFN, and discuss its differences from other deep models. Moreover, we propose fully convolutional fusion networks (FCFN) that are an extension of CFN for pixel-level classification applied to several tasks, such as semantic segmentation and edge detection. Our experiments demonstrate that CFN (and FCFN) can achieve promising performance by consistent improvements for both image-level and pixel-level classification tasks, compared to a plain CNN. We release our codes on https://github.com/yuLiu24/CFN . Also, we make a live demo ( goliath.liacs.nl ) using a CFN model trained on the ImageNet dataset. Yu Liu 0012, Yanming Guo, Theodoros Georgiou 0001, Michael S. Lew |
Multim. Tools Appl. | 1 |
| 2018 | Learning visual and textual representations for multimodal matching and classification
Yu Liu 0012, Li Liu 0002, Yanming Guo, Michael S. Lew |
Pattern Recognit. | 1 |
| 2018 | Bag of Surrogate Parts Feature for Visual RecognitionabstractConvolutional neural networks (CNNs) have attracted significant attention in visual recognition. Several recent studies have shown that, in addition to the fully connected layers, the features derived from the convolutional layers of CNNs can also achieve promising performance in image classification tasks. In this paper, we propose a new feature from the convolutional layers, called Bag of Surrogate Parts (BoSP), and its spatial variant, Spatial-BoSP (S-BoSP). The main idea is, we assume the feature maps in the convolutional layers as surrogate parts, and densely sample and assign image regions to these surrogate parts by observing the activation values. Together with BoSP/S-BoSP, we further propose another two schemes to enhance the performance: scale pooling and global-part prediction. Scale pooling aims to handle the objects with different scales and deformations, and global-part prediction combines the predictions of global and part features. By conducting extensive experiments on generic object, fine-grained object and scene datasets, we find the proposed scheme can not only achieve superior performance to the fully connected feature, but also produces competitive or, in some cases, remarkably better performance than the state of the art. Yanming Guo, Yu Liu 0012, Songyang Lao, Erwin M. Bakker, Michael S. Lew |
IEEE Trans. Multim. | 2 |
| 2017 | Learning a Recurrent Residual Fusion Network for Multimodal Matching
Yu Liu 0012, Yanming Guo, Erwin M. Bakker, Michael S. Lew |
ICCV | 1 |
| 2017 | Improving the discrimination between foreground and background for semantic segmentationabstractOne challenging problem in semantic segmentation is due to the erroneous predictions between categorical foreground and cluttered background. To address it, we propose to utilize a fused loss function to train a fully convolutional network, which aims to enhance the discrimination between foreground and background in images. In addition, we propose a pixel objectness (POS) to measure the importance of pixels. POS is able to recover some missing foreground pixels from the background. Experimental results on the PASCAL VOC 2012 dataset demonstrate our approach can achieve considerable improvements compared with the baseline counterpart, while maintaining the ease of training deep networks. Yu Liu 0012, Michael S. Lew |
ICIP | 1 |
| 2017 | On the Exploration of Convolutional Fusion Networks for Visual Recognition
Yu Liu 0012, Yanming Guo, Michael S. Lew |
MMM (1) | 1 |
| 2017 | What Convnets Make for Image Captioning?
Yu Liu 0012, Yanming Guo, Michael S. Lew |
MMM (1) | 1 |
| 2016 | Learning Relaxed Deep Supervision for Better Edge DetectionabstractWe propose using relaxed deep supervision (RDS) within convolutional neural networks for edge detection. The conventional deep supervision utilizes the general groundtruth to guide intermediate predictions. Instead, we build hierarchical supervisory signals with additional relaxed labels to consider the diversities in deep neural networks. We begin by capturing the relaxed labels from simple detectors (e.g. Canny). Then we merge them with the general groundtruth to generate the RDS. Finally we employ the RDS to supervise the edge network following a coarse-to-fine paradigm. These relaxed labels can be seen as some false positives that are difficult to be classified. Weconsider these false positives in the supervision, and are able to achieve high performance for better edge detection. Wecompensate for the lack of training images by capturing coarse edge annotations from a large dataset of image segmentations to pretrain the model. Extensive experiments demonstrate that our approach achieves state-of-the-art performance on the well-known BSDS500 dataset (ODS F-score of .792) and obtains superior cross-dataset generalization results on NYUD dataset. Yu Liu 0012, Michael S. Lew |
CVPR | 1 |
| 2016 | Deep learning for visual understanding: A review
Yanming Guo, Yu Liu 0012, Ard Oerlemans, Songyang Lao, Song Wu 0003, Michael S. Lew |
Neurocomputing | 2 |
| 2016 | Hierarchical projective invariant contexts for shape recognition
Qi Jia 0001, Xin Fan 0001, Yu Liu 0012, Zhongxuan Luo, He Guo 0001 |
Pattern Recognit. | 3 |
| 2015 | DeepIndex for Accurate and Efficient Image RetrievalabstractIn the well-known Bag-of-Words model, local features, such as the SIFT descriptor, are extracted and quantized into visual words. Then, an index is created to reduce computational burden. However, local clues serve as low-level representations that can not represent high-level semantic concepts. Recently, the success of deep features extracted from convolutional neural networks(CNN) has shown promising results toward bridging the semantic gap. Inspired by this, we attempt to introduce deep features into inverted index based image retrieval and thus propose the DeepIndex framework. Moreover, considering the compensation of different deep features, we incorporate multiple deep features from different fully connected layers, resulting in the multiple DeepIndex. We find the optimal integration of one midlevel deep feature and one high-level deep feature, from two different CNN architectures separately. This can be treated as an attempt to further reduce the semantic gap. Extensive experiments on three benchmark datasets demonstrate that, the proposed DeepIndex method is competitive with the state-of-the-art on Holidays(85:65% mAP), Paris(81:24% mAP), and UKB(3:76 score). In addition, our method is efficient in terms of both memory and time cost. Yu Liu 0012, Yanming Guo, Song Wu 0003, Michael S. Lew |
ICMR | 1 |
| 2014 | A new geometric descriptor for symbols with affine deformations
Qi Jia 0001, Xin Fan 0001, Zhongxuan Luo, Yu Liu 0012, He Guo 0001 |
Pattern Recognit. Lett. | 4 |
| 2013 | A shape matching framework using metric partition constraintabstractThe crucial problem for shape matching is to balance between discrimination power and computation complexity. Popular solutions mainly rely on either global or local information of shape contours, and neglect their intrinsic correlation. But the methods that combine both information may bring high computation complexity. In this paper, we present a shape matching framework, in which a novel shape descriptor named metric partition constraint (MPC) is proposed, and many metric methods can be included. The metric information is used to bridge the local points and the global shape. Meanwhile, we devise a partition smoothing process to improve the robustness to local deformation. Finally, Comprehensive comparisons with the classical shape context and other latest methods on standard datasets show the excellent performance in terms of precision while retaining computational efficiency. Yu Liu 0012, Qi Jia 0001, He Guo 0001, Xin Fan 0001 |
ICIP | 1 |