EDBT 2026 Demo / reviewers in the wild / expert
Wei Ke 0003
dblp:52/7566-3
· DBLP profile ↗
40ranked-venue papers
5as first author
28since 2021 · last 2026
0000-0002-2899-0371ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 28 · 4 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 24 · 3 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Monte Carlo Diffusion for Generalizable Learning-Based RANSACabstractRandom Sample Consensus (RANSAC) is a fundamental approach for robustly estimating parametric models from noisy data. Existing learning-based RANSAC methods utilize deep learning to enhance the robustness of RANSAC against outliers. However, these approaches are trained and tested on the data generated by the same algorithms, leading to limited generalization to out-of-distribution data during inference. Therefore, in this paper, we introduce a novel diffusion-based paradigm that progressively injects noise into ground-truth data, simulating the noisy conditions for training learning-based RANSAC. To enhance data diversity, we incorporate Monte Carlo sampling into the diffusion paradigm, approximating diverse data distributions by introducing different types of randomness at multiple stages. We evaluate our approach in the context of feature matching through comprehensive experiments on the ScanNet and MegaDepth datasets. The experimental results demonstrate that our Monte Carlo diffusion mechanism significantly improves the generalization ability of learning-based RANSAC. We also develop extensive ablation studies that highlight the effectiveness of key components in our framework. Chen Zhao 0025, Wei Ke 0003, Tong Zhang 0023 |
AAAI | 3 |
| 2026 | VulKnow: Enhancing Vulnerability Detection With Structured Knowledge and Large Language ModelsabstractThe rapid proliferation of Internet-of-Things (IoT) systems has led to increasingly heterogeneous, resource-constrained, and security-sensitive software deployments. In this context, detecting vulnerabilities in embedded and system-level IoT code has become particularly critical, as even a single exploitable flaw may compromise entire networks or devices. Vulnerability detection in real-world software continues to pose significant challenges due to the intricate nature of program semantics, the diversity of vulnerability patterns, and the limited explainability offered by current machine learning models. Although Large Language Models (LLMs) have exhibited remarkable capabilities in comprehending and reasoning about source code, their effectiveness is frequently compromised by inadequate domain knowledge and instances of hallucinated outputs. To address these limitations, we propose VulKnow, a retrieval-augmented framework for vulnerability detection that harnesses structured vulnerability knowledge alongside multi-stage prompt-based reasoning. Specifically, VulKnow first constructs a structured knowledge base derived from authoritative sources such as CWE/CVE reports and standard library specifications. Sub-sequently, it conducts context-aware knowledge retrieval and integrates the retrieved items into an LLM-based initial filtering module. To ensure high-confidence and verifiable results, we design a multi-stage reasoning pipeline that progressively validates vulnerabilities through semantic prompts and consistency checks. Experiments conducted on multiple real-world datasets (Linux, Qemu, Big-Vul) demonstrate that VulKnow significantly enhances precision, recall, and explainability compared to existing static analysis tools as well as LLM-only baselines. This positions VulKnow as a reliable solution for practical vulnerability auditing scenarios. Guixiang Liao, Yanli Chen 0001, Wei Ke 0003, Hanzhou Wu, Zhicheng Dong 0003 |
IEEE Internet Things J. | 3 |
| 2026 | Crowd counting with sparse annotation
Shiwei Zhang 0004, Zhengzheng Wang, Qing Liu 0003, Wei Ke 0003, Tong Zhang 0023 |
Pattern Recognit. | 5 |
| 2026 | Semantic Distribution and Authenticity Discrepancy Alignment for AI-Generated Image DetectionabstractGenerative models have achieved remarkable success in producing vivid images. Compared with real images, generated ones still show different semantic structures that features with different semantic classes collapse as a single cluster. Pioneer works leverage the discrepancy of semantic structure in fixed high-level semantic feature space to identify forgery images. Nevertheless, such frozen pre-trained representation models are insensitive to subtle forgery traces. Meanwhile, vanilla fine-tuning methods can distort the pre-trained semantic knowledge and collapse to the real-fake binary distribution, losing generalization capability in newly emerged generative models. In this paper, we propose thesemantic distribution and authenticity discrepancy alignment algorithm (STERM), which learns high-level semantic structures of real-world categories and low-level forgery traces for detecting AI-generated images from unseen generative models and frameworks. Specifically, we first capture semantic features of images by the frozen CLIP and further extract forgery features by a forgery encoder. Then, we propose semantic distribution alignment (SDA) to align the semantic structure of real-world categories by enforcing forgery feature distribution shifting towards the semantic feature space. Next, we introduce authenticity discrepancy alignment (ADA) to minimize the authenticity discrepancy between forgery and semantic features, constraining forgery features from collapsing into the source domain-biased distribution and learning the semantic structure of real-world categories. Extensive experiments on GAN-based and diffusion model-based datasets demonstrate the generalization capability of the proposed method. The source code is publicly available athttps://github.com/freshjh/STERM. Liang Li 0003, Chenggang Yan 0001, Wei Ke 0003, Yihong Gong |
IEEE Trans. Multim. | 4 |
| 2025 | Generating Multimodal Driving Scenes via Next-Scene PredictionabstractGenerative models in Autonomous Driving (AD) enable diverse scenario creation, yet existing methods fall short by only capturing a limited range of modalities, restricting the capability of generating controllable scenes for comprehensive evaluation of AD systems. In this paper, we introduce a multimodal generation framework that incorporates four major data modalities, including a novel addition of the map modality. With tokenized modalities, our scene sequence generation framework autoregressively predicts each scene while managing computational demands through a two-stage approach. The Temporal AutoRegressive (TAR) component captures inter-frame dynamics for each modality, while the Ordered AutoRegressive (OAR) component aligns modalities within each scene by sequentially predicting tokens in a fixed order To maintain coherence between map and ego-action modalities, we introduce the Action-Aware Map Alignment (AMA) module, which applies a transformation based on the ego-action to maintain coherence between these two modalities. Our framework effectively generates complex, realistic driving scenes over extended sequences, ensuring multimodal consistency and offering fine-grained control over scene elements. Project page: https://yanhaowu.github.io/UMGen. Yanhao Wu, Lichao Huang, Shujie Luo, Congpei Qiu, Wei Ke 0003 |
CVPR | 8 |
| 2025 | CountSE: Soft Exemplar Open-Set Object Counting
Shuai Liu 0016, Shiwei Zhang 0004, Wei Ke 0003 |
ICCV | 4 |
| 2025 | Enhancing Zero-Shot Object Counting via Text-Guided Local Ranking and Number-Evoked Global Attention
Shiwei Zhang 0004, Wei Ke 0003 |
ICCV | 3 |
| 2025 | Refining CLIP's Spatial Awareness: A Visual-Centric PerspectiveabstractContrastive Language-Image Pre-training (CLIP) excels in global alignment with language but exhibits limited sensitivity to spatial information, leading to strong performance in zero-shot classification tasks but underperformance in tasks requiring precise spatial understanding. Recent approaches have introduced Region-Language Alignment (RLA) to enhance CLIP's performance in dense multimodal tasks by aligning regional visual representations with corresponding text inputs. However, we find that CLIP ViTs fine-tuned with RLA suffer from notable loss in spatial awareness, which is crucial for dense prediction tasks. To address this, we propose the Spatial Correlation Distillation (SCD) framework, which preserves CLIP's inherent spatial structure and mitigates above degradation. To further enhance spatial correlations, we introduce a lightweight Refiner that extracts refined correlations directly from CLIP before feeding them into SCD, based on an intriguring finding that CLIP naturally capture high-quality dense features. Together, these components form a robust distillation framework that enables CLIP ViTs to integrate both visual-language and visual-centric improvements, achieving state-of-the-art results across various open-vocabulary dense prediction benchmarks. Congpei Qiu, Yanhao Wu, Wei Ke 0003, Xiuxiu Bai, Tong Zhang 0023 |
ICLR | 3 |
| 2025 | Are High-Quality AI-Generated Images More Difficult for Models to Detect?abstractThe remarkable evolution of generative models has enabled the generation of high-quality, visually attractive images, often perceptually indistinguishable from real photographs to human eyes. This has spurred significant attention on AI-generated image (AIGI) detection. Intuitively, higher image quality should increase detection difficulty. However, our systematic study on cutting-edge text-to-image generators reveals a counterintuitive finding: AIGIs with higher quality scores, as assessed by human preference models, tend to be more easily detected by existing models. To investigate this, we examine how the text prompts for generation and image characteristics influence both quality scores and detector accuracy. We observe that images from short prompts tend to achieve higher preference scores while being easier to detect. Furthermore, through clustering and regression analyses, we verify that image characteristics like saturation, contrast, and texture richness collectively impact both image quality and detector accuracy. Finally, we demonstrate that the performance of off-the-shelf detectors can be enhanced across diverse generators and datasets by selecting input patches based on the predicted scores of our regression models, thus substantiating the broader applicability of our findings. Code and data are available at https://github.com/Coxy7/AIGI-Detection-Quality-Paradox. Zijie Cao, ZiYi Dong, Xiangyang Ji, Liang Lin 0004, Wei Ke 0003, Pengxu Wei |
ICML | 9 |
| 2025 | Frequency-aware Correlation Discovering and Spatial Forgery Clue Distilling for Synthetic Image DetectionabstractRecent text-to-image generative models facilitate creating vivid images with arbitrary contents that are indistinguishable from authentic ones by naked eyes. Despite progress in synthetic image detection, detecting the image from new generators remains challenging. Because advanced generators leave fewer visible forgery traces, while different generative frameworks produce varied forgery patterns. We notice that generative models consistently struggle with fine-detailed content generation, creating abnormal spatial dependencies among neighboring pixels in complex texture regions. In this paper, we propose a methodology of gazing local detail of forgery (GLDF) for generator agnostic synthetic image detection, which identifies prominent spatial dependencies to capture subtle forgery. Concretely, we design frequency-aware correlation discovering (FACD) module to learn dynamic filters by instance-adaptive frequency masking block for identifying prominent spatial deficiencies, which distributed in different spatial positions with various patterns. Furthermore, we introduce the spatial forgery clue distilling module (SFCD) to iteratively aggregate and refine spatial dependencies from different positions by spatial aggregating and prototype global interacting blocks. Extensive experiments demonstrate that GLDF outperforms state-of-the-art methods on detecting synthetic images from different generators. Liang Li 0003, Chenggang Yan 0001, Wei Ke 0003, Yihong Gong |
ACM Multimedia | 4 |
| 2025 | A Bayesian dual-pathway network for unsupervised domain adaptation
Yuhang He 0001, Junzhe Chen 0002, Wei Ke 0003, Yihong Gong |
Pattern Recognit. | 4 |
| 2025 | One point is all you need for weakly supervised object detection
Shiwei Zhang 0004, Zhengzheng Wang, Wei Ke 0003 |
Pattern Recognit. | 3 |
| 2025 | Language-Driven Visual Consensus for Zero-Shot Semantic SegmentationabstractThe pre-trained vision-language model, exemplified by CLIP, advances zero-shot semantic segmentation by aligning visual features with class embeddings through a transformer decoder to generate semantic masks. Despite its effectiveness, prevailing methods within this paradigm encounter challenges, including overfitting on seen classes and small fragmentation in segmentation masks. To mitigate these issues, we propose a Language-Driven Visual Consensus (LDVC) approach, fostering improved alignment of linguistic and visual information. Specifically, we leverage class embeddings as anchors due to their discrete and abstract nature, steering visual features toward class embeddings. Moreover, to achieve a more compact visual space, we introduce route attention into the transformer decoder to find visual consensus, thereby enhancing semantic consistency within the same object. Equipped with a vision-language prompting strategy, our approach significantly boosts the generalization capacity of segmentation models for unseen classes. Experimental results underscore the effectiveness of our approach, showcasing mIoU gains of 4.5% on the PASCAL VOC 2012 and 3.6% on the COCO-Stuff 164K for unseen classes compared with the state-of-the-art methods. Wei Ke 0003, Yi Zhu 0004, Xiaodan Liang, Jianzhuang Liu, Qixiang Ye, Tong Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Monocular Depth Estimation on Adverse Weathers With Curriculum Domain Distribution AlignmentabstractDespite the remarkable success of monocular depth estimation, most works focus on ideal experiment conditions, such as favorable weather, where there is few environmental factors impacting the depth estimation system. In practical, when suffering from adverse weather conditions, such as fog and rain, the model trained on favorable weather degrades sharply as the domain shift, caused by the decreasing of visibility. To solve this problem, in this paper, we propose a Curriculum Domain Distribution Alignment (CDA) algorithm to learn the domain-invariant representation, progressively aligning data distributions across favorable weather and adverse weather in the feature space. Concretely, to construct a domain adaptation curriculum, we first separate the target domain into several subsets with increased domain discrepancy based on an optical model. Then, we bridge the distribution discrepancy between domains from easier to harder data by matching the source and target representation subspace. Furthermore, to control the distribution aligning pace, we introduce self-paced learning to learn a dynamic domain adaptation weight, promoting the generalization ability of monocular depth estimation networks against environmental factors. We conduct experiments with six monocular depth estimation frameworks on FoggyCityScapes, RainCityScapes, SnowCityscapes, and All-day Cityscapes, improving RMSE with 8.5 %, 30.5 %, 30.9 %, 20.9 %. The extraordinary performance demonstrates the effectiveness and generalizability of our method under adverse weather conditions. Liang Li 0003, Chenggang Yan 0001, Wei Ke 0003, Yihong Gong |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Frequency-Aware Divide-and-Conquer for Efficient Real Noise RemovalabstractDeep-learning-based approaches have achieved remarkable progress for complex real scenario denoising, yet their accuracy-efficiency tradeoff is still understudied, particularly critical for mobile devices. As real noise is unevenly distributed relative to underlay signals in different frequency bands, we introduce a frequency-aware divide-and-conquer strategy to develop a frequency-aware denoising network (FADN). FADN is materialized by stacking frequency-aware denoising blocks (FADBs), in which a denoised image is progressively predicted by a series of frequency-aware noise dividing and conquering operations. For noise dividing, FADBs decompose the noisy and clean image pairs into low- and high-frequency representations via a wavelet transform (WT) followed by an invertible network and recover the final denoised image by integrating the denoised information from different frequency bands. For noise conquering, the separated low-frequency representation of the noisy image is kept as clean as possible by the supervision of the clean counterpart, while the high-frequency representation combining the estimated residual from the successive FADB is purified under the corresponding accompanied supervision for residual compensation. Since our FADN progressively and pertinently denoises from frequency bands, the accuracy-efficiency tradeoff can be controlled as a requirement by the number of FADBs. Experimental results on the SIDD, DND, and NAM datasets show that our FADN outperforms the state-of-the-art methods by improving the peak signal-to-noise ratio (PSNR) and decreasing the model parameters. The code is released at https://github.com/NekoDaiSiki/FADN. Yunqi Huang, Chang Liu 0047, Wei Ke 0003, Xiaojun Jing |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2024 | Mitigating Object Dependencies: Improving Point Cloud Self-Supervised Learning Through Object ExchangeabstractIn the realm of point cloud scene understanding, particularly in indoor scenes, objects are arranged following human habits, resulting in objects of certain semantics being closely positioned and displaying notable inter-object cor-relations. This can create a tendency for neural networks to exploit these strong dependencies, bypassing the individ-ual object patterns. To address this challenge, we introduce a novel self-supervised learning (SSL) strategy. Our approach leverages both object patterns and contextual cues to produce robust features. It begins with the formulation of an object-exchanging strategy, where pairs of objects with comparable sizes are exchanged across different scenes, effectively disentangling the strong contextual dependencies. Subsequently, we introduce a context-aware feature learning strategy, which encodes object patterns without relying on their specific context by aggregating object features across various scenes. Our extensive experiments demonstrate the superiority of our method over existing SSL techniques, further showing its better robustness to environmental changes. Moreover, we showcase the applicability of our approach by transferring pre-trained models to diverse point cloud datasets.11Our code is available at https:/lgithub.com/YanhaoWu/OESSL Yanhao Wu, Tong Zhang 0023, Wei Ke 0003, Congpei Qiu, Sabine Süsstrunk, Mathieu Salzmann |
CVPR | 3 |
| 2024 | Mind Your Augmentation: The Key to Decoupling Dense Self-Supervised LearningabstractDense Self-Supervised Learning (SSL) creates positive pairs by building positive paired regions or points, thereby aiming to preserve local features, for example of individual objects. However, existing approaches tend to couple objects by leaking information from the neighboring contextual regions when the pairs have a limited overlap. In this paper, we first quantitatively identify and confirm the existence of such a coupling phenomenon. We then address it by developing a remarkably simple yet highly effective solution comprising a novel augmentation method, Region Collaborative Cutout (RCC), and a corresponding decoupling branch. Importantly, our design is versatile and can be seamlessly integrated into existing SSL frameworks, whether based on Convolutional Neural Networks (CNNs) or Vision Transformers (ViTs). We conduct extensive experiments, incorporating our solution into two CNN-based and two ViT-based methods, with results confirming the effectiveness of our approach. Moreover, we provide empirical evidence that our method significantly contributes to the disentanglement of feature representations among objects, both in quantitative and qualitative terms. Congpei Qiu, Tong Zhang 0023, Yanhao Wu, Wei Ke 0003, Mathieu Salzmann, Sabine Süsstrunk |
ICLR | 4 |
| 2024 | Kepler codebookabstractA codebook designed for learning discrete distributions in latent space has demonstrated state-of-the-art results on generation tasks. This inspires us to explore what distribution of codebook is better. Following the spirit of Kepler's Conjecture, we cast the codebook training as solving the sphere packing problem and derive a Kepler codebook with a compact and structured distribution to obtain a codebook for image representations. Furthermore, we implement the Kepler codebook training by simply employing this derived distribution as regularization and using the codebook partition method. We conduct extensive experiments to evaluate our trained codebook for image reconstruction and generation on natural and human face datasets, respectively, achieving significant performance improvement. Besides, our Kepler codebook has demonstrated superior performance when evaluated across datasets and even for reconstructing images with different resolutions. Our trained models and source codes will be publicly released. Junrong Lian, Ziyue Dong, Pengxu Wei, Wei Ke 0003, Chang Liu 0030, Qixiang Ye, Xiangyang Ji, Liang Lin 0004 |
ICML | 4 |
| 2024 | Boosting Semi-supervised Crowd Counting with Scale-based Active LearningabstractThe core of active semi-supervised crowd counting is the sample selection criteria. However, the scale factor has been neglected in active learning approaches despite the fact that the scale of heads varies drastically in the crowd images. In this paper, we propose a simple yet effective active labeling strategy to explicitly select informative unlabeled images, guided by the intra-scale uncertainty and inter-scale inconsistency metrics. The intra-scale uncertainty is quantified through the sum of the query-level entropy of images at different scales. Images are initially ranked based on this uncertainty for preselection. Inter-scale inconsistency is measured by the divergence between the query-level predictions of upscaled and downscaled images, allowing for the identification of the most informative images exhibiting the highest inconsistency. Additionally, we implement a progressive updating scheme for the semi-supervised crowd counting framework, in which the pseudo-labels for unlabeled images are refined iteratively. It further improves the counting accuracy. Through extensive experiments on widely used benchmarks, the proposed approach has demonstrated superior performance compared to previous state-of-the-art semi-supervised and active semi-supervised crowd counting methods. Shiwei Zhang 0004, Wei Ke 0003, Shuai Liu 0016, Xiaopeng Hong, Tong Zhang 0023 |
ACM Multimedia | 2 |
| 2024 | PBT: Progressive Background-Aware Transformer for Infrared Small Target DetectionabstractIn the domain of infrared small target detection (IRSTD), the challenges revolve around detecting small and faint targets from infrared images. These targets lack distinct textures and morphology exist in complex backgrounds with numerous distractions. Current deep-learning methods typically prioritize preserving target features while neglecting the crucial background context, ultimately resulting in false alarms and miss detection. To tackle this issue, we propose a novel approach involving separately focusing on candidate target responses and background context during the encoding stage and aligning them during the decoding stage. Specifically, we introduce the progressive background-aware transformer (PBT) which adopts an asymmetric encoder-decoder architecture. The encoder with task-specific frequency domain priors extracts candidate target responses and background context features separately from shallow and deep blocks, respectively. The following hierarchical decoder progressively refines the candidate target responses under the guidance of rich background context stage by stage, leading to more accurate results. Our experiments demonstrate that PBT surpasses state-of-the-art IRSTD methods across various datasets. The code and dataset are available athttps://github.com/Heron0625/PBT. Huoren Yang, Tingkui Mu, Ziyue Dong, Wei Ke 0003, Qiujie Yang, Zhiping He |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2023 | Spatiotemporal Self-Supervised Learning for Point Clouds in the WildabstractSelf-supervised learning (SSL) has the potential to benefit many applications, particularly those where manually annotating data is cumbersome. One such situation is the semantic segmentation of point clouds. In this context, existing methods employ contrastive learning strategies and define positive pairs by performing various augmentation of point clusters in a single frame. As such, these methods do not exploit the temporal nature of LiDAR data. In this paper, we introduce an SSL strategy that leverages positive pairs in both the spatial and temporal domain. To this end, we design (i) a point-to-cluster learning strategy that aggregates spatial information to distinguish objects; and (ii) a cluster-to-cluster learning strategy based on unsupervised object tracking that exploits temporal correspondences. We demonstrate the benefits of our approach via extensive experiments performed by self-supervised training on two large-scale LiDAR datasets and transferring the resulting models to other point cloud segmentation benchmarks. Our results evidence that our method outperforms the state-of-the-art point cloud SSL methods.11Our code and pretrained models will be found at https://github.com/YanhaoWu/STSSL. Correspondence to Ke Wei. Yanhao Wu, Tong Zhang 0023, Wei Ke 0003, Sabine Süsstrunk, Mathieu Salzmann |
CVPR | 3 |
| 2023 | Automated lesion segmentation in fundus images with many-to-many reassembly of features
Qing Liu 0003, Wei Ke 0003, Yixiong Liang |
Pattern Recognit. | 3 |
| 2022 | Leverage Your Local and Global Representations: A New Self-Supervised Learning StrategyabstractSelf-supervised learning (SSL) methods aim to learn view-invariant representations by maximizing the similar-ity between the features extracted from different crops of the same image regardless of cropping size and content. In essence, this strategy ignores the fact that two crops may truly contain different image information, e.g., background and small objects, and thus tends to restrain the diversity of the learned representations. In this work, we address this issue by introducing a new self-supervised learning strat-egy, LoGo, that explicitly reasons about Local and Global crops. To achieve view invariance, LoGo encourages similarity between global crops from the same image, as well as between a global and a local crop. However, to correctly encode the fact that the content of smaller crops may differ entirely, LoGo promotes two local crops to have dissimi-lar representations, while being close to global crops. Our LoGo strategy can easily be applied to existing SSL meth-ods. Our extensive experiments on a variety of datasets and using different self-supervised learning frameworks vali-date its superiority over existing approaches. Noticeably, we achieve better results than supervised models on trans-fer learning when using only 1/10 of the data.11Our code and pretrained models can be found at https://github.com/ztt1024/LoGo-SSL. Tong Zhang 0023, Congpei Qiu, Wei Ke 0003, Sabine Süsstrunk, Mathieu Salzmann |
CVPR | 3 |
| 2022 | CoupAlign: Coupling Word-Pixel with Sentence-Mask Alignments for Referring Image SegmentationabstractReferring image segmentation aims at localizing all pixels of the visual objects described by a natural language sentence. Previous works learn to straightforwardly align the sentence embedding and pixel-level embedding for highlighting the referred objects, but ignore the semantic consistency of pixels within the same object, leading to incomplete masks and localization errors in predictions. To tackle this problem, we propose CoupAlign, a simple yet effective multi-level visual-semantic alignment method, to couple sentence-mask alignment with word-pixel alignment to enforce object mask constraint for achieving more accurate localization and segmentation. Specifically, the Word-Pixel Alignment (WPA) module performs early fusion of linguistic and pixel-level features in intermediate layers of the vision and language encoders. Based on the word-pixel aligned embedding, a set of mask proposals are generated to hypothesize possible objects. Then in the Sentence-Mask Alignment (SMA) module, the masks are weighted by the sentence embedding to localize the referred object, and finally projected back to aggregate the pixels for the target. To further enhance the learning of the two alignment modules, an auxiliary loss is designed to contrast the foreground and background pixels. By hierarchically aligning pixels and masks with linguistic features, our CoupAlign captures the pixel coherence at both visual and semantic levels, thus generating more accurate predictions. Extensive experiments on popular datasets (e.g., RefCOCO and G-Ref) show that our method achieves consistent improvements over state-of-the-art methods, e.g., about 2% oIoU increase on the validation and testing set of RefCOCO. Especially, CoupAlign has remarkable ability in distinguishing the target from multiple objects of the same class. Code will be available at https://gitee.com/mindspore/models/tree/master/research/cv/CoupAlign. Yi Zhu 0004, Jianzhuang Liu, Xiaodan Liang, Wei Ke 0003 |
NeurIPS | 5 |
| 2022 | Identity-Quantity Harmonic Multi-Object TrackingabstractThe data association problem of multi-object tracking (MOT) aims to assign IDentity (ID) labels to detections and infer a complete trajectory for each target. Most existing methods assume that each detection corresponds to a unique target and thus cannot handle situations when multiple targets occur in a single detection due to detection failure in crowded scenes. To relax this strong assumption for practical applications, we formulate the MOT as a Maximizing An Identity-Quantity Posterior (MAIQP) problem on the basis of associating each detection with an identity and a quantity characteristic and then provide solutions to tackle two key problems arising. Firstly, a local target quantification module is introduced to count the number of targets within one detection. Secondly, we propose an identity-quantity harmony mechanism to reconcile the two characteristics. On this basis, we develop a novel Identity-Quantity HArmonic Tracking (IQHAT) framework that allows assigning multiple ID labels to detections containing several targets. Through extensive experimental evaluations on five benchmark datasets, we demonstrate the superiority of the proposed method. Yuhang He 0001, Xing Wei 0001, Xiaopeng Hong, Wei Ke 0003, Yihong Gong |
IEEE Trans. Image Process. | 4 |
| 2021 | Error-Aware Density Isomorphism Reconstruction for Unsupervised Cross-Domain Crowd CountingabstractThis paper focuses on the unsupervised domain adaptation problem for video-based crowd counting, in which we use labeled data as source domain and unlabelled video data as target domain. It is challenging as there is a huge gap between the source and the target domain and no annotations of samples are available in the target domain. The key issue is how to utilize unlabelled videos in the target domain for knowledge learning and transferring from the source domain. To tackle this problem, we propose a novel Error-aware Density Isomorphism REConstruction Network (EDIREC-Net) for cross-domain crowd counting. EDIREC-Net jointly transfers a pre-trained counting model to target domains using a density isomorphism reconstruction objective and models the reconstruction erroneousness by error reasoning. Specifically, as crowd flows in videos are consecutive, the density maps in adjacent frames turn out to be isomorphic. On this basis, we regard the density isomorphism reconstruction error as a self-supervised signal to transfer the pre-trained counting models to different target domains. Moreover, we leverage an estimation-reconstruction consistency to monitor the density reconstruction erroneousness and suppress unreliable density reconstructions during training. Experimental results on four benchmark datasets demonstrate the superiority of the proposed method and ablation studies investigate the efficiency and robustness. The source code is available at https://github.com/GehenHe/EDIREC-Net. Yuhang He 0001, Zhiheng Ma, Xing Wei 0001, Xiaopeng Hong, Wei Ke 0003, Yihong Gong |
AAAI | 5 |
| 2021 | Kohonen Self-Organizing Map based Route Planning: A RevisitabstractIn this paper, we revisit the long-standing Traveling Salesman Problem (TSP) and focus on the challenging, yet practical route planning problem with limited computational resources. We make contributions to TSP, one of the most famous NP-hard problems by providing a new improved approximate solution, which we term TOpology Preserving Self-Organizing Map (TOPSOM). TOPSOM well preserves the topology of the node map to be traversed by maintaining the continuity of nodes and the distances between them. In addition, to satisfy the requirements of convex hull, we design an elastic competitive Hebbian learning rule. TOPSOM can solve large-scale TSPs with high precision and high efficiency with limited computational costs. Extensive experimental results on mainstream route planning benchmarks including TSPLIB and National TSP’s show that our method consistently outperforms baseline methods, by up to 7.7% in terms of the Percent Deviation of Mean solution to best known solution. Qingshu Guan, Xiaopeng Hong, Wei Ke 0003, Liangfei Zhang, Guanghui Sun, Yihong Gong |
IROS | 3 |
| 2021 | SRN: Side-Output Residual Network for Object Reflection Symmetry Detection and BeyondabstractThis article establishes a baseline for object reflection symmetry detection in natural images by releasing a new benchmark named Sym-PASCAL and proposing an end-to-end deep learning approach for reflection symmetry. Sym-PASCAL spans challenges of multiobjects, object diversity, part invisibility, and clustered backgrounds, which is far beyond those in existing data sets. The end-to-end deep learning approach, referred to as a side-output residual network (SRN), leverages the output residual units (RUs) to fit the errors between the symmetry ground truth and the side outputs of multiple stages of a trunk network. By cascading RUs from deep to shallow, SRN exploits the "flow" of errors along multiple stages to effectively matching object symmetry at different scales and suppress the clustered backgrounds. SRN is interpreted as a boosting-like algorithm, which assembles features using RUs during network forward and backward propagations. SRN is further upgraded to a multitask SRN (MT-SRN) for joint symmetry and edge detection, demonstrating its generality to image-to-mask learning tasks. Experimental results verify that the Sym-PASCAL benchmark is challenging related to real-world images, SRN achieves state-of-the-art performance, and MT-SRN has the capability to simultaneously predict edge and symmetry mask without loss of performance. Wei Ke 0003, Jie Chen 0001, Jianbin Jiao, Guoying Zhao 0001, Qixiang Ye |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2020 | Multiple Anchor Learning for Visual Object DetectionabstractClassification and localization are two pillars of visual object detectors. However, in CNN-based detectors, these two modules are usually optimized under a fixed set of candidate (or anchor) bounding boxes. This configuration significantly limits the possibility to jointly optimize classification and localization. In this paper, we propose a Multiple Instance Learning (MIL) approach that selects anchors and jointly optimizes the two modules of a CNN-based object detector. Our approach, referred to as Multiple Anchor Learning (MAL), constructs anchor bags and selects the most representative anchors from each bag. Such an iterative selection process is potentially NP-hard to optimize. To address this issue, we solve MAL by repetitively depressing the confidence of selected anchors by perturbing their corresponding features. In an adversarial selection-depression manner, MAL not only pursues optimal solutions but also fully leverages multiple anchors/features to learn a detection model. Experiments show that MAL improves the baseline RetinaNet with significant margins on the commonly used MS-COCO object detection benchmark and achieves new state-of-the-art detection performance compared with recent methods. Wei Ke 0003, Tianliang Zhang 0003, Zeyi Huang, Qixiang Ye, Jianzhuang Liu |
CVPR | 1 |
| 2020 | Class-Incremental Learning with Topological Schemas of Memory SpacesabstractClass-incremental learning (CIL) aims to incrementally learn a unified classifier for new classes emerging, which suffers from the catastrophic forgetting problem. To alleviate forgetting and improve the recognition performance, we propose a novel CIL framework, named the topological schemas model (TSM). TSM consists of a Gaussian mixture model arranged on 2D grids (2D-GMM) as the memory of the learned knowledge. To train the 2D-GMM model, we develop a novel competitive expectation-maximization (CEM) method, which contains a global topology embedding step and a local expectation-maximization fine-tuning step. Meanwhile, we choose the image samples of old classes that have the maximum posterior probability with respect to each Gaussian distribution as the episodic points. When finetuning for new classes, we propose the memory preservation loss (MPL) term to ensure episodic points still have maximum probabilities with respect to the corresponding Gaussian distribution. MPL preserves the distribution of 2D-GMM for old knowledge during incremental learning and alleviates catastrophic forgetting. Comprehensive experimental evaluations on two popular CIL benchmarks CIFAR100 and subImageNet demonstrate the superiority of our TSM. Xinyuan Chang, Xiaopeng Hong, Xing Wei 0001, Wei Ke 0003, Yihong Gong |
ICPR | 5 |
| 2020 | Co-Attentive Lifting for Infrared-Visible Person Re-IdentificationabstractInfrared-visible cross-modality person re-identification (IV-ReID) has attracted much attention with the popularity of dual-mode video surveillance systems, where the RGB mode works in the daytime and automatically switches to the infrared mode at night. Despite its significant application value, IV-ReID remains a difficult problem mainly due to two great challenges. First, it is difficult to identify persons in the infrared image, which lacks color and texture clues. Second, there is a significant gap between the infrared and visible modalities where appearances of the same person vary considerably. This paper proposes a novel attention-based approach to handle the two difficulties in a unified framework. 1) We propose an attention lifting mechanism to learn discriminative features in each modality. 2) We propose a co-attentive learning mechanism to bridge the gap between the two modalities. Our method only makes slight modifications of a given backbone network and requires small computation overhead while improving the performance significantly. We conduct extensive experiments to demonstrate the superiority of our proposed method. Xing Wei 0001, Diangang Li, Xiaopeng Hong, Wei Ke 0003, Yihong Gong |
ACM Multimedia | 4 |
| 2020 | Progressive Latent Models for Self-Learning Scene-Specific Pedestrian DetectorsabstractThe performance of offline learned pedestrian detectors significantly drops when they are applied to video scenes of various camera views, occlusions, and background structures. Learning a detector for each video scene can avoid the performance drop but it requires repetitive human effort on data annotation. In this paper, a self-learning approach is proposed, toward specifying a pedestrian detector for each video scene without any human annotation involved. Object locations in video frames are treated as latent variables and a progressive latent model (PLM) is proposed to solve such latent variables. The PLM is deployed as components of object discovery, object enforcement, and label propagation, which are used to learn the object locations in a progressive manner. With the difference of convex (DC) objective functions, PLM is optimized by a concave-convex programming algorithm. With specified network branches and loss functions, PLM is integrated with deep feature learning and optimized in an end-to-end manner. From the perspectives of convex regularization and error rate estimation, detailed optimization analysis and learning stability analysis of the proposed PLM are provided. The extensive experiments demonstrate that even without annotation involved the proposed self-learning approach outperforms weakly supervised learning approaches, while achieving comparable performance with transfer learning approaches. Qixiang Ye, Tianliang Zhang 0003, Wei Ke 0003 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2019 | Orthogonal Decomposition Network for Pixel-Wise Binary ClassificationabstractThe weight sharing scheme and spatial pooling operations in Convolutional Neural Networks (CNNs) introduce semantic correlation to neighboring pixels on feature maps and therefore deteriorate their pixel-wise classification performance. In this paper, we implement an Orthogonal Decomposition Unit (ODU) that transforms a convolutional feature map into orthogonal bases targeting at de-correlating neighboring pixels on convolutional features. In theory, complete orthogonal decomposition produces orthogonal bases which can perfectly reconstruct any binary mask (ground-truth). In practice, we further design incomplete orthogonal decomposition focusing on de-correlating local patches which balances the reconstruction performance and computational cost. Fully Convolutional Networks (FCNs) implemented with ODUs, referred to as Orthogonal Decomposition Networks (ODNs), learn de-correlated and complementary convolutional features and fuse such features in a pixel-wise selective manner. Over pixel-wise binary classification tasks for two-dimensional image processing, specifically skeleton detection, edge detection, and saliency detection, and one-dimensional keypoint detection, specifically S-wave arrival time detection for earthquake localization, ODNs consistently improves the state-of-the-arts with significant margins. Chang Liu 0042, Fang Wan 0001, Wei Ke 0003, Zhuowei Xiao, Xiaosong Zhang 0004, Qixiang Ye |
CVPR | 3 |
| 2019 | C-MIL: Continuation Multiple Instance Learning for Weakly Supervised Object DetectionabstractWeakly supervised object detection (WSOD) is a challenging task when provided with image category supervision but required to simultaneously learn object locations and object detectors. Many WSOD approaches adopt multiple instance learning (MIL) and have non-convex loss functions which are prone to get stuck into local minima (falsely localize object parts) while missing full object extent during training. In this paper, we introduce a continuation optimization method into MIL and thereby creating continuation multiple instance learning (C-MIL), with the intention of alleviating the non-convexity problem in a systematic way. We partition instances into spatially related and class related subsets, and approximate the original loss function with a series of smoothed loss functions defined within the subsets. Optimizing smoothed loss functions prevents the training procedure falling prematurely into local minima and facilitates the discovery of Stable Semantic Extremal Regions (SSERs) which indicate full object extent. On the PASCAL VOC 2007 and 2012 datasets, C-MIL improves the state-of-the-art of weakly supervised object detection and weakly supervised object localization with large margins. Fang Wan 0001, Chang Liu 0042, Wei Ke 0003, Xiangyang Ji, Jianbin Jiao, Qixiang Ye |
CVPR | 3 |
| 2019 | Deep contour and symmetry scored object proposal
Wei Ke 0003, Jie Chen 0001, Qixiang Ye |
Pattern Recognit. Lett. | 1 |
| 2018 | Linear Span Network for Object Skeleton Detection
Chang Liu 0042, Wei Ke 0003, Qixiang Ye |
ECCV (2) | 2 |
| 2017 | SRN: Side-Output Residual Network for Object Symmetry Detection in the WildabstractIn this paper, we establish a baseline for object symmetry detection in complex backgrounds by presenting a new benchmark and an end-to-end deep learning approach, opening up a promising direction for symmetry detection in the wild. The new benchmark, named Sym-PASCAL, spans challenges including object diversity, multi-objects, part-invisibility, and various complex backgrounds that are far beyond those in existing datasets. The proposed symmetry detection approach, named Side-output Residual Network (SRN), leverages output Residual Units (RUs) to fit the errors between the object symmetry ground-truth and the outputs of RUs. By stacking RUs in a deep-to-shallow manner, SRN exploits the flow of errors among multiple scales to ease the problems of fitting complex outputs with limited layers, suppressing the complex backgrounds, and effectively matching object symmetry of different scales. Experimental results validate both the benchmark and its challenging aspects related to real-world images, and the state-of-the-art performance of our symmetry detection approach. The benchmark and the code for SRN are publicly available at https://github.com/KevinKecc/SRN. Wei Ke 0003, Jie Chen 0001, Jianbin Jiao, Guoying Zhao 0001, Qixiang Ye |
CVPR | 1 |
| 2017 | Self-Learning Scene-Specific Pedestrian Detectors Using a Progressive Latent ModelabstractIn this paper, a self-learning approach is proposed towards solving scene-specific pedestrian detection problem without any human annotation involved. The self-learning approach is deployed as progressive steps of object discovery, object enforcement, and label propagation. In the learning procedure, object locations in each frame are treated as latent variables that are solved with a progressive latent model (PLM). Compared with conventional latent models, the proposed PLM incorporates a spatial regularization term to reduce ambiguities in object proposals and to enforce object localization, and also a graph-based label propagation to discover harder instances in adjacent frames. With the difference of convex (DC) objective functions, PLM can be efficiently optimized with a concave-convex programming and thus guaranteeing the stability of self-learning. Extensive experiments demonstrate that even without annotation the proposed self-learning approach outperforms weakly supervised learning approaches, while achieving comparable performance with transfer learning and fully supervised approaches. Qixiang Ye, Tianliang Zhang 0003, Wei Ke 0003, Qiang Qiu 0001, Jie Chen 0001, Guillermo Sapiro, Baochang Zhang 0001 |
CVPR | 3 |
| 2015 | Pedestrian detection via PCA filters based convolutional channel featuresabstractIn this paper, we propose a kind of image representation, named PCA filters based convolutional channel features (PCA-CCF) for pedestrian detection. The motivation is to use the convolutional network architecture with orthogonal PCA filters to enhance the state-of-the-art aggregate channel features (ACF). In PCA-CCF, the convolutional operation improves the feature robustness to pedestrian local deformation. The learned PCA filters reduce the correlations among features of each channel, and therefore, improve feature discrimination capability. With the proposed PCA-CCF features and cascaded AdaBoost classifiers, we develop a coarse-to-fine pedestrian detection approach. Experiments show that such approach achieves 3.04%, 17.87% and 6.28% performance gain on the INRIA, Caltech Reasonable and Caltech Overall pedestrian datasets, respectively. Wei Ke 0003, Pengxu Wei, Qixiang Ye, Jianbin Jiao |
ICASSP | 1 |
| 2013 | Minimum Entropy Models for Laser Line Extraction
Wei Ke 0003, Ce Li 0005, Jianbin Jiao |
CAIP (2) | 3 |