VLDB 2026 Research / reviewers in the wild / expert
Yihong Gong
dblp:62/6520
· DBLP profile ↗
227ranked-venue papers
17as first author
69since 2021 · last 2026
0000-0002-1793-5836ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 135 · 12 first-author · 48 since 2021Artificial intelligence and machine learning · 130 · 7 first-author · 40 since 2021Databases, data management, data science and information retrieval · 23 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 since 2021Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-authorSecurity and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GOAL: Geometrically Optimal Alignment for Continual Generalized Category DiscoveryabstractContinual Generalized Category Discovery (C-GCD) requires identifying novel classes from unlabeled data while retaining knowledge of known classes over time. Existing methods typically update classifier weights dynamically, resulting in forgetting and inconsistent feature alignment. We propose GOAL, a unified framework that introduces a fixed Equiangular Tight Frame (ETF) classifier to impose a consistent geometric structure throughout learning. GOAL conducts supervised alignment for labeled samples and confidence-guided alignment for novel samples, enabling stable integration of new classes without disrupting old ones. Experiments on four benchmarks show that GOAL outperforms prior methods, reducing forgetting by 16.1% and boosting novel class discovery by 3.2%, establishing a strong solution for long-horizon continual discovery. Jizhou Han, Chenhao Ding, Songlin Dong, Yuhang He 0001, Shaokun Wang, Yihong Gong |
AAAI | 7 |
| 2026 | Shared & Domain Self-Adaptive Experts with Frequency-Aware Discrimination for Continual Test-Time AdaptationabstractThis paper focuses on the Continual Test-Time Adaptation (CTTA) task, aiming to enable an agent to continuously adapt to evolving target domains while retaining previously acquired domain knowledge for effective reuse when those domains reappear. Existing shared-parameter paradigms struggle to balance adaptation and forgetting, leading to decreased efficiency and stability. To address this, we propose a frequency-aware shared and self-adaptive expert framework, consisting of two key components: (i) a dual-branch expert architecture that extracts general features and dynamically models domain-specific representations, effectively reducing cross-domain interference and repetitive learning cost; and (ii) an online Frequency-aware Domain Discriminator (FDD), which leverages the robustness of low-frequency image signals for online domain shift detection, guiding dynamic allocation of expert resources for more stable and realistic adaptation. Additionally, we introduce a Continual Repeated Shifts (CRS) benchmark to simulate periodic domain changes for more realistic evaluation. Experimental results show that our method consistently outperforms existing approaches on both classification and segmentation CTTA tasks under standard and CRS settings, with ablations and visualizations confirming its effectiveness and robustness. Jianchao Zhao, Chenhao Ding, Songlin Dong, Jiangyang Li, Yuhang He 0001, Yihong Gong |
AAAI | 7 |
| 2026 | Diversity covariance-aware prompt learning for vision-language models
Zhengdong Zhou, Songlin Dong, Chenhao Ding, Xinyuan Gao, Yuhang He 0001, Yihong Gong |
Pattern Recognit. | 6 |
| 2026 | Unleashing the Potential of All Test Samples: Mean-Shift Guided Test-Time AdaptationabstractVisual-language models (VLMs) like CLIP exhibit strong generalization but struggle with distribution shifts at test time. Existing training-free test-time adaptation (TTA) methods operate strictly within CLIP’s original feature space, relying on high-confidence samples while overlooking the potential of low-confidence ones. We propose MS-TTA, a training-free approach that enhances feature representations beyond CLIP’s space using a single-step k-nearest neighbors (kNN) Mean-Shift. By refining all test samples, MS-TTA improves feature compactness and class separability, leading to more stable adaptation. Additionally, a cache of refined embeddings further enhances inference by providing Mean-Shift-enhanced logits. Extensive evaluations on OOD and Cross-Dataset Benchmarks demonstrate that MS-TTA consistently outperforms state-of-the-art training-free TTA methods, achieving robust adaptation without requiring additional training. Jizhou Han, Chenhao Ding, Songlin Dong, Xinyuan Gao, Yuhang He 0001, Yihong Gong |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2026 | Learn by Reasoning: Analogical Weight Generation for Few-Shot Class-Incremental LearningabstractFew-shot class-incremental Learning (FSCIL) enables models to learn new classes from limited data while retaining performance on previously learned classes. Traditional FSCIL methods often require fine-tuning parameters with limited new class data and suffer from a separation between learning new classes and utilizing old knowledge. Inspired by the analogical learning mechanisms of the human brain, we propose a novel analogical generative method. Our approach includes the Brain-Inspired Analogical Generator (BiAG), which derives new class weights from existing classes without parameter fine-tuning during incremental stages. BiAG consists of three components: Weight Self-Attention Module (WSA), Weight & Prototype Analogical Attention Module (WPAA), and Semantic Conversion Module (SCM). SCM uses Neural Collapse theory for semantic conversion, WSA supplements new class weights, and WPAA computes analogies to generate new class weights. Experiments on miniImageNet, CUB-200, and CIFAR-100 datasets demonstrate that our method achieves higher final and average accuracy compared to SOTA methods. Jizhou Han, Chenhao Ding, Yuhang He 0001, Songlin Dong, Xinyuan Gao, Yihong Gong |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2026 | HemNet: Hemoglobin-Assistant Network for Video-Based Remote Photoplethysmography MeasurementabstractTraditional skin-contact physical sensors typically detect changes of blood volume to predict the periodicity of heartbeat by analyzing the absorption spectra of hemoglobin. However, the contact on human skin may cause uncomfortable feeling and induce difficulty for long-term monitoring. Recently, video-based remote photoplethysmography (rPPG) estimation approaches analyze the periodic facial color changes for matching cardiac cycle in a contactless manner. Nevertheless, the inherent relationship between the changes of facial color and blood volume is not fully exploited. Besides the influence of blood volume (i.e., hemoglobin), there are also other factors such as lighting and reflection that cause the change on facial color. We exploit the physical principles that cause skin color variations to separate the hemoglobin factor driven by blood volume. Based on the physical prior of the reflection of human skin, we introduce an rPPG estimation network assisted by decoupled hemoglobin sequence, named HemNet, which first explicitly leverages hemoglobin to assist rPPG signal estimation. To obtain meaningful hemoglobin from facial video, we design a human skin color disentangler that decouples the facial color variations into four significant features, i.e., hemoglobin, melanin, shading, and specular. We then present a multi-modality rPPG estimator that utilizes cross-covariance attention to extract fused feature from hemoglobin and RGB video inputs. Finally, an adaptive negative Pearson loss is proposed to effectively address phase misalignment between the blood volume in the finger and facial region during the training phase. We evaluate our HemNet on four widely used public benchmark datasets. The superiority of our method is demonstrated in both intra-dataset and cross-dataset test settings. The code is available at https://github.com/jingang-cv/hemnet. Ruize Wu, Jingang Shi, Xin Liu 0012, LinLin Shen, Yihong Gong, Guoying Zhao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2026 | Multi-Task Unified Domain Incremental Learning With Domain Difference AdaptersabstractThis paper focuses on domain incremental learning (DIL) for multiple vision tasks, including object detection, instance segmentation, and image classification. DIL aims to adapt a model to new domains over time without forgetting previously acquired knowledge. Recent DIL methods append learnable prompts to input embeddings of a frozen base model to learn from new domains. However, due to prompts' limited representation ability, they struggle to adapt the feature space to new domain data distributions. To overcome this limitation, we propose a novel DIL method named Domain Difference Adapters (DD-Adapters). Through feature visualization and singular value analysis, we identify the cross-domain clustering ability of the base model and the low-rank property of domain difference. Based on these insights, our method imposes low-rank constraints on the base model to capture the principal components of domain differences, while freezing the base model to maintain its cross-domain clustering ability, thereby adapting to new domains effectively. Additionally, we introduce a prototype-guided domain selector (PDS) to dynamically select the appropriate DD-Adapters during inference, mitigating catastrophic forgetting in DIL. Extensive experimental evaluations on eight benchmark datasets demonstrate the performance superiority of the proposed method on three vision tasks, with minimal extra parameter usage. Xiang Song 0005, Yuhang He 0001, Lin Peng 0003, Yihong Gong |
IEEE Trans. Image Process. | 4 |
| 2026 | Semantic Distribution and Authenticity Discrepancy Alignment for AI-Generated Image DetectionabstractGenerative models have achieved remarkable success in producing vivid images. Compared with real images, generated ones still show different semantic structures that features with different semantic classes collapse as a single cluster. Pioneer works leverage the discrepancy of semantic structure in fixed high-level semantic feature space to identify forgery images. Nevertheless, such frozen pre-trained representation models are insensitive to subtle forgery traces. Meanwhile, vanilla fine-tuning methods can distort the pre-trained semantic knowledge and collapse to the real-fake binary distribution, losing generalization capability in newly emerged generative models. In this paper, we propose thesemantic distribution and authenticity discrepancy alignment algorithm (STERM), which learns high-level semantic structures of real-world categories and low-level forgery traces for detecting AI-generated images from unseen generative models and frameworks. Specifically, we first capture semantic features of images by the frozen CLIP and further extract forgery features by a forgery encoder. Then, we propose semantic distribution alignment (SDA) to align the semantic structure of real-world categories by enforcing forgery feature distribution shifting towards the semantic feature space. Next, we introduce authenticity discrepancy alignment (ADA) to minimize the authenticity discrepancy between forgery and semantic features, constraining forgery features from collapsing into the source domain-biased distribution and learning the semantic structure of real-world categories. Extensive experiments on GAN-based and diffusion model-based datasets demonstrate the generalization capability of the proposed method. The source code is publicly available athttps://github.com/freshjh/STERM. Liang Li 0003, Chenggang Yan 0001, Wei Ke 0003, Yihong Gong |
IEEE Trans. Multim. | 5 |
| 2025 | DualCP: Rehearsal-Free Domain-Incremental Learning via Dual-Level Concept PrototypeabstractDomain-Incremental Learning (DIL) enables vision models to adapt to changing conditions in real-world environments while maintaining the knowledge acquired from previous domains. Given privacy concerns and training time, Rehearsal-Free DIL (RFDIL) is more practical. Inspired by the incremental cognitive process of the human brain, we design Dual-level Concept Prototypes (DualCP) for each class to address the conflict between learning new knowledge and retaining old knowledge in RFDIL. To construct DualCP, we propose a Concept Prototype Generator (CPG) that generates both coarse-grained and fine-grained prototypes for each class. Additionally, we introduce a Coarse-to-Fine calibrator (C2F) to align image features with DualCP. Finally, we propose a Dual Dot-Regression (DDR) loss function to optimize our C2F module. Extensive experiments on the DomainNet, CDDB, and CORe50 datasets demonstrate the effectiveness of our method. Yuhang He 0001, Songlin Dong, Xiang Song 0005, Jizhou Han, Haoyu Luo, Yihong Gong |
AAAI | 7 |
| 2025 | Learning Endogenous Attention for Incremental Object DetectionabstractIn this paper, we focus on a challenging Incremental Object Detection (IOD) problem. Existing IOD methods adopt an image-to-annotation alignment paradigm, which attempts to complete the absent old category annotations and learns both new and old categories concurrently in new tasks. This paradigm inherently introduces missing/redundant/inaccurate annotations of old categories, resulting in a suboptimal performance. Instead, we propose a novel annotation-to-instance alignment IOD paradigm and develop a corresponding method named Learning Endogenous Attention (LEA). Inspired by the human brain, LEA enables the model to focus on annotated task-specific objects, while ignoring irrelevant ones, thus solving the annotation incomplete problem in IOD. Concretely, our LEA consists of Endogenous Attention Modules (EAMs) and an Energy-Based Task Modulator (ETM). During training, we add the dedicated EAMs for each new task and train them to focus on the new categories. During testing, ETM predicts task IDs using energy functions, directing the model to detect task-specific objects. The detection results corresponding to all task IDs are combined as the final output, thereby alleviating the catastrophic forgetting of old knowledge. Extensive experiments on COCO 2017 and Pascal VOC 2007 demonstrate the effectiveness of our method1. Xiang Song 0005, Yuhang He 0001, Yihong Gong |
CVPR | 5 |
| 2025 | Dynamic Integration of Task-Specific Adapters for Class Incremental LearningabstractNon-exemplar Class Incremental Learning (NECIL) enables models to continuously acquire new classes without retraining from scratch and storing old task exemplars, addressing privacy and storage issues. However, the absence of data from earlier tasks exacerbates the challenge of catastrophic forgetting in NECIL. In this paper, we propose a novel framework called Dynamic Integration of task-specific Adapters (DIA), which comprises two key components: Task-Specific Adapter Integration (TSAI) and Patch-Level Model Alignment. TSAI boosts compositionality through a patch-level adapter integration strategy, aggregating richer task-specific information while maintaining low computation costs. Patch-Level Model Alignment maintains feature consistency and accurate decision boundaries via two specialized mechanisms: Patch-Level Distillation Loss (PDL) and Patch-Level Feature Reconstruction (PFR). Specifically, on the one hand, the PDL preserves feature-level consistency between successive models by implementing a distillation loss based on the contributions of patch tokens to new class learning. On the other hand, the PFR promotes classifier alignment by reconstructing old class features from previous tasks that adapt to new task knowledge, thereby preserving well-calibrated decision boundaries. Comprehensive experiments validate the effectiveness of our DIA, revealing significant improvements on NECIL benchmark datasets while maintaining an optimal balance between computational complexity and accuracy. Jiashuo Li, Shaokun Wang, Yuhang He 0001, Xing Wei 0001, Yihong Gong |
CVPR | 7 |
| 2025 | Autoregressive Sequential Pretraining for Visual TrackingabstractRecent advancements in visual object tracking have shifted towards a sequential generation paradigm, where object deformation and motion exhibit strong temporal dependencies. Despite the importance of these dependencies, widely adopted image-level pretrained backbones barely capture the dynamics in the consecutive video, which is the essence of tracking. Thus, we propose AutoRegressive Sequential Pretraining (ARP), an unsupervised spatio-temporal learner, via generating the evolution of object appearance and motion in video sequences. Our method leverages a diffusion model to autoregressively generate the future frame appearance, conditioned on historical embeddings extracted by a general encoder. Furthermore, to ensure trajectory coherence, the same encoder is employed to learn trajectory consistency by generating coordinate sequences in a reverse autoregressive fashion, a process we term backtracking. Further, we integrate the pretrained ARP into AR-TrackV2, creating ARPTrack, which is further fine-tuned for tracking tasks. ARPTrack achieves state-of-the-art performance across multiple benchmarks, becoming the first tracker to surpass 80% AO on GOT-10k, while maintaining high efficiency. These results demonstrate the effectiveness of our approach in capturing temporal dependencies for continuous video tracking. The code will be released soon. Shiyi Liang, Yifan Bai 0001, Yihong Gong, Xing Wei 0001 |
CVPR | 3 |
| 2025 | Boosting Domain Incremental Learning: Selecting the Optimal Parameters is All You NeedabstractDeep neural networks (DNNs) often underperform in real-world, dynamic settings where data distributions change over time. Domain Incremental Learning (DIL) offers a solution by enabling continuous model adaptation, with Parameter-Isolation DIL (PIDIL) emerging as a promising paradigm to reduce knowledge conflicts. However, existing PIDIL methods struggle with parameter selection accuracy, especially as the number of domains and corresponding classes grows. To address this, we propose SOYO, a lightweight framework that improves domain selection in PIDIL. SOYO introduces a Gaussian Mixture Compressor (GMC) and Domain Feature Resampler (DFR) to store and balance prior domain data efficiently, while a Multi-level Domain Feature Fusion Network (MDFN) enhances domain feature extraction. Our framework supports multiple Parameter-Efficient Fine-Tuning (PEFT) methods and is validated across tasks such as image classification, object detection, and speech enhancement. Experimental results on six benchmarks demonstrate SOYO’s consistent superiority over existing baselines, showcasing its robustness and adaptability in complex, evolving environments. Xiang Song 0005, Yuhang He 0001, Jizhou Han, Chenhao Ding, Xinyuan Gao, Yihong Gong |
CVPR | 7 |
| 2025 | Test-Time Adaptation via Distribution-Aware Guidance for Vision-Language Models
Qilong Xue, Chenhao Ding, Zhengdong Zhou, Zeyang Zhao, Yihong Gong |
ICIC (1) | 5 |
| 2025 | CIA: Class- and Instance-aware Adaptation for Vision-Language ModelsabstractFew-shot parameter-efficient tuning methods demonstrate promising potential for Vision-Language (V-L) models in downstream tasks. However, existing approaches primarily focus on class-level alignment between image and text features, overlooking crucial instance-specific semantic information. This limitation leads to suboptimal performance on challenging tasks and restricted generalization capability to unseen data. To address these issues, we propose Class- and Instance-aware Adaptation (CIA), a novel framework that simultaneously optimizes both class-level and instance-level alignments. Specifically, CIA introduces a novel instance encoder that leverages cross-modal self-attention to generate instance-specific text features, accompanied by a carefully designed regularization mechanism to maintain consistency between class-level and instance-level representations. Extensive experiments across 15 benchmark datasets demonstrate that CIA significantly improves the downstream adaptation of V-L models. Lin Peng 0003, Cong Wan, Shaokun Wang, Xiang Song 0005, Yuhang He 0001, Yihong Gong |
ACM Multimedia | 6 |
| 2025 | Frequency-aware Correlation Discovering and Spatial Forgery Clue Distilling for Synthetic Image DetectionabstractRecent text-to-image generative models facilitate creating vivid images with arbitrary contents that are indistinguishable from authentic ones by naked eyes. Despite progress in synthetic image detection, detecting the image from new generators remains challenging. Because advanced generators leave fewer visible forgery traces, while different generative frameworks produce varied forgery patterns. We notice that generative models consistently struggle with fine-detailed content generation, creating abnormal spatial dependencies among neighboring pixels in complex texture regions. In this paper, we propose a methodology of gazing local detail of forgery (GLDF) for generator agnostic synthetic image detection, which identifies prominent spatial dependencies to capture subtle forgery. Concretely, we design frequency-aware correlation discovering (FACD) module to learn dynamic filters by instance-adaptive frequency masking block for identifying prominent spatial deficiencies, which distributed in different spatial positions with various patterns. Furthermore, we introduce the spatial forgery clue distilling module (SFCD) to iteratively aggregate and refine spatial dependencies from different positions by spatial aggregating and prototype global interacting blocks. Extensive experiments demonstrate that GLDF outperforms state-of-the-art methods on detecting synthetic images from different generators. Liang Li 0003, Chenggang Yan 0001, Wei Ke 0003, Yihong Gong |
ACM Multimedia | 5 |
| 2025 | Consistent Supervised-Unsupervised Alignment for Generalized Category DiscoveryabstractGeneralized Category Discovery (GCD) focuses on classifying known categories while simultaneously discovering novel categories from unlabeled data. However, previous GCD methods face challenges due to inconsistent optimization objectives and category confusion. This leads to feature overlap and ultimately hinders performance on novel categories. To address these issues, we propose the Neural Collapse-inspired Generalized Category Discovery (NC-GCD) framework. By pre-assigning and fixing Equiangular Tight Frame (ETF) prototypes, our method ensures an optimal geometric structure and a consistent optimization objective for both known and novel categories. We introduce a Consistent ETF Alignment Loss that unifies supervised and unsupervised ETF alignment and enhances category separability. Additionally, a Semantic Consistency Matcher (SCM) is designed to maintain stable and consistent label assignments across clustering iterations. Our method significantly enhancing novel category accuracy and demonstrating its effectiveness. Jizhou Han, Shaokun Wang, Yuhang He 0001, Chenhao Ding, Xinyuan Gao, Songlin Dong, Yihong Gong |
NeurIPS | 8 |
| 2025 | A Bayesian dual-pathway network for unsupervised domain adaptation
Yuhang He 0001, Junzhe Chen 0002, Wei Ke 0003, Yihong Gong |
Pattern Recognit. | 5 |
| 2025 | CEAT: Continual Expansion and Absorption Transformer for Non-Exemplar Class-Incremental LearningabstractIn dynamic real-world scenarios, continuous learning without forgetting old knowledge is essential, particularly in environments with stricter privacy protection or resource-constrained edge devices where storing old exemplars is infeasible. Therefore, Non-Exemplar Class-Incremental Learning (NECIL) has garnered significant attention. Compared with normal settings, it faces a more severe plasticity-stability dilemma and classifier bias. To address those challenges, we propose a framework based on the vision transformer architecture, called the Continual Expansion and Absorption Transformer (CEAT), which consists of two core components. First, we propose the Continual Expansion and Absorption (CEA) method to alleviate the trade-off between new and old classes by parallelly expanding a set of parameters (i.e. EF layer) on the backbone to learn new tasks, while freezing the backbone to retain old task knowledge. The EF layers can be seamlessly absorbed into the ViT backbone through parameter recombination before inference, mitigating storage and computational burdens. Second, we propose a Dynamic Boundary-Aware (DBA) method to generate dynamic pseudo-features for classifier calibration to address the classifier bias. Extensive experiments demonstrate that our approach achieves state-of-the-art performance, particularly showcasing significant improvements of 4.82% and 5.92% on TinyImageNet and ImageNet-Subset, respectively. Songlin Dong, Xinyuan Gao, Yuhang He 0001, Zhengdong Zhou, Alex Chichung Kot, Yihong Gong |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | Monocular Depth Estimation on Adverse Weathers With Curriculum Domain Distribution AlignmentabstractDespite the remarkable success of monocular depth estimation, most works focus on ideal experiment conditions, such as favorable weather, where there is few environmental factors impacting the depth estimation system. In practical, when suffering from adverse weather conditions, such as fog and rain, the model trained on favorable weather degrades sharply as the domain shift, caused by the decreasing of visibility. To solve this problem, in this paper, we propose a Curriculum Domain Distribution Alignment (CDA) algorithm to learn the domain-invariant representation, progressively aligning data distributions across favorable weather and adverse weather in the feature space. Concretely, to construct a domain adaptation curriculum, we first separate the target domain into several subsets with increased domain discrepancy based on an optical model. Then, we bridge the distribution discrepancy between domains from easier to harder data by matching the source and target representation subspace. Furthermore, to control the distribution aligning pace, we introduce self-paced learning to learn a dynamic domain adaptation weight, promoting the generalization ability of monocular depth estimation networks against environmental factors. We conduct experiments with six monocular depth estimation frameworks on FoggyCityScapes, RainCityScapes, SnowCityscapes, and All-day Cityscapes, improving RMSE with 8.5 %, 30.5 %, 30.9 %, 20.9 %. The extraordinary performance demonstrates the effectiveness and generalizability of our method under adverse weather conditions. Liang Li 0003, Chenggang Yan 0001, Wei Ke 0003, Yihong Gong |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Curriculum Dataset DistillationabstractMost dataset distillation methods struggle to accommodate large-scale datasets due to their substantial computational and memory requirements. Recent research has begun to explore scalable disentanglement methods. However, there are still performance bottlenecks and room for optimization in this direction. In this paper, we present a curriculum-based dataset distillation framework aiming to harmonize performance and scalability. This framework strategically distills synthetic images, adhering to a curriculum that transitions from simple to complex. By incorporating curriculum evaluation, we address the issue of previous methods generating images that tend to be homogeneous and simplistic, doing so at a manageable computational cost. Furthermore, we introduce adversarial optimization towards synthetic images to further improve their representativeness and safeguard against their overfitting to the neural network involved in distilling. This enhances the generalization capability of the distilled images across various neural network architectures and also increases their robustness to noise. Extensive experiments demonstrate that our framework sets new benchmarks in large-scale dataset distillation, achieving substantial improvements of 11.1% on Tiny-ImageNet, 9.0% on ImageNet-1K, and 7.3% on ImageNet-21K. Our distilled datasets and code are available at https://github.com/MIV-XJTU/CUDD. Zhiheng Ma, Anjia Cao, Funing Yang, Yihong Gong, Xing Wei 0001 |
IEEE Trans. Image Process. | 4 |
| 2025 | Cross-Domain Invariant Feature Absorption and Domain-Specific Feature Retention for Domain Incremental Chest X-Ray ClassificationabstractChest X-ray (CXR) images have been widely adopted in clinical care and pathological diagnosis in recent years. Some advanced methods on CXR classification task achieve impressive performance by training the model statically. However, in the real clinical environment, the model needs to learn continually and this can be viewed as a domain incremental learning (DIL) problem. Due to large domain gaps, DIL is faced with catastrophic forgetting. Therefore, in this paper, we propose a Cross-domain invariant feature absorption and Domain-specific feature retention (CaD) framework. To be specific, we adopt a Cross-domain Invariant Feature Absorption (CIFA) module to learn the domain invariant knowledge and a Domain-Specific Feature Retention (DSFR) module to learn the domain-specific knowledge. The CIFA module contains the C(lass)-adapter and an absorbing strategy is used to fuse the common features among different domains. The DSFR module contains the D(omain)-adapter for each domain and it connects to the network in parallel independently to prevent forgetting. A multi-label contrastive loss (MLCL) is used in the training process and improves the class distinctiveness within each domain. We leverage publicly available large-scale datasets to simulate domain incremental learning scenarios, extensive experimental results substantiate the effectiveness of our proposed methods and it has reached state-of-the-art performance. Mengchu Wang, Yuhang He 0001, Lin Peng 0003, Xiang Song 0005, Songlin Dong, Yihong Gong |
IEEE Trans. Medical Imaging | 6 |
| 2025 | Analogical Augmentation and Significance Analysis for Online Task-Free Continual LearningabstractOnline task-free continual learning (OTFCL) is a more challenging variant of continual learning that emphasizes the gradual shift of task boundaries and learning in an online mode. Existing methods rely on a memory buffer of old samples to prevent forgetting. However, the use of memory buffers not only raises privacy concerns but also hinders the efficient learning of new samples. To address this problem, we propose a novel framework called I$^{2}$CANSAY that gets rid of the dependence on memory buffers and efficiently learns the knowledge of new data from one-shot samples. Concretely, our framework comprises two main modules. Firstly, theInter-Class Analogical Augmentation(ICAN) module generates diverse pseudo-features for old classes based on the inter-class analogy of feature distributions for different new classes, serving as a substitute for the memory buffer. Secondly, theIntra-Class Significance Analysis(ISAY) module analyzes the significance of attributes for each class via its distribution standard deviation, and generates an importance vector as a correction bias for the linear classifier, thereby enhancing the capability of learning from new samples. We run our experiments on four popular image classification datasets: CoRe50, CIFAR-10, CIFAR-100, and CUB-200, our approach outperforms the prior state-of-the-art by a large margin. Songlin Dong, Yuhang He 0001, Yuhan Jin, Alex Chichung Kot, Yihong Gong |
IEEE Trans. Multim. | 6 |
| 2024 | Non-exemplar Domain Incremental Object Detection via Learning Domain BiasabstractDomain incremental object detection (DIOD) aims to gradually learn a unified object detection model from a dataset stream composed of different domains, achieving good performance in all encountered domains. The most critical obstacle to this goal is the catastrophic forgetting problem, where the performance of the model improves rapidly in new domains but deteriorates sharply in old ones after a few sessions. To address this problem, we propose a non-exemplar DIOD method named learning domain bias (LDB), which learns domain bias independently at each new session, avoiding saving examples from old domains. Concretely, a base model is first obtained through training during session 1. Then, LDB freezes the weights of the base model and trains individual domain bias for each new incoming domain, adapting the base model to the distribution of new domains. At test time, since the domain ID is unknown, we propose a domain selector based on nearest mean classifier (NMC), which selects the most appropriate domain bias for a test image. Extensive experimental evaluations on two series of datasets demonstrate the effectiveness of the proposed LDB method in achieving high accuracy on new and old domain datasets. The code is available at https://github.com/SONGX1997/LDB. Xiang Song 0005, Yuhang He 0001, Songlin Dong, Yihong Gong |
AAAI | 4 |
| 2024 | Evolving Parameterized Prompt Memory for Continual LearningabstractRecent studies have demonstrated the potency of leveraging prompts in Transformers for continual learning (CL). Nevertheless, employing a discrete key-prompt bottleneck can lead to selection mismatches and inappropriate prompt associations during testing. Furthermore, this approach hinders adaptive prompting due to the lack of shareability among nearly identical instances at more granular level. To address these challenges, we introduce the Evolving Parameterized Prompt Memory (EvoPrompt), a novel method involving adaptive and continuous prompting attached to pre-trained Vision Transformer (ViT), conditioned on specific instance. We formulate a continuous prompt function as a neural bottleneck and encode the collection of prompts on network weights. We establish a paired prompt memory system consisting of a stable reference and a flexible working prompt memory. Inspired by linear mode connectivity, we progressively fuse the working prompt memory and reference prompt memory during inter-task periods, resulting in continually evolved prompt memory. This fusion involves aligning functionally equivalent prompts using optimal transport and aggregating them in parameter space with an adjustable bias based on prompt node attribution. Additionally, to enhance backward compatibility, we propose compositional classifier initialization, which leverages prior prototypes from pre-trained models to guide the initialization of new classifiers in a subspace-aware manner. Comprehensive experiments validate that our approach achieves state-of-the-art performance in both class and domain incremental learning scenarios. Muhammad Rifki Kurniawan, Xiang Song 0005, Zhiheng Ma, Yuhang He 0001, Yihong Gong, Xing Wei 0001 |
AAAI | 5 |
| 2024 | ARTrackV2: Prompting Autoregressive Tracker Where to Look and How to DescribeabstractWe present ARTrackV2, which integrates two pivotal aspects of tracking: determining where to look (localization) and how to describe (appearance analysis) the target object across video frames. Building on the foundation of its predecessor, ARTrackV2 extends the concept by introducing a unified generative framework to “read out” object's trajectory and “retell” its appearance in an autoregressive manner. This approach fosters a time-continuous methodology that models the joint evolution of motion and visual features, guided by previous estimates. Furthermore, ARTrackV2 stands out for its efficiency and simplicity, obviating the less efficient intra-frame autoregression and hand-tuned parameters for appearance updates. Despite its simplicity, ARTrackV2 achieves state-of-the-art performance on prevailing benchmark datasets while demonstrating a remarkable efficiency improvement. In particular, ARTrackV2 achieves an AO score of 79. 5% on GOT-10k and an AUC of 86. 1% on TrackingNet while being 3.6× faster than ARTrack. Yifan Bai 0001, Zeyang Zhao, Yihong Gong, Xing Wei 0001 |
CVPR | 3 |
| 2024 | DYSON: Dynamic Feature Space Self-Organization for Online Task-Free Class Incremental LearningabstractIn this paper, we focus on a challenging Online Task-Free Class Incremental Learning (OTFCIL) problem. Dif-ferent from the existing methods that continuously learn the feature space from data streams, we propose a novel compute-and-align paradigm for the OTFCIL. It first com-putes an optimal geometry, i.e., the class prototype distri-bution, for classifying existing classes and updates it when new classes emerge, and then trains a DNN model by aligning its feature space to the optimal geometry. To this end, we develop a novel Dynamic Neural Collapse (DNC) algorithm to compute and update the optimal geometry. The DNC ex-pands the geometry when new classes emerge without loss of the geometry optimality and guarantees the drift distance of old class prototypes with an explicit upper bound. On this basis, we propose a novel DYnamic feature space Self-OrganizatioN (DYSON) method containing three ma-jor components, including 1) a feature extractor, 2) a Dy-namic Feature-Geometry Alignment (DFGA) module aligning the feature space to the optimal geometry computed by DNC and 3) a training-free class-incremental classifier de-rived from the DNC geometry. Experimental comparison results on four benchmark datasets, including CIFAR10, CI-FAR100, CUB200, and CoRe50, demonstrate the efficiency and superiority of the DYSON method. The source code is released at https://github.com/isCDX2IDYSON. Yuhang He 0001, Yuhan Jin, Songlin Dong, Xing Wei 0001, Yihong Gong |
CVPR | 6 |
| 2024 | Beyond Prompt Learning: Continual Adapter for Efficient Rehearsal-Free Continual Learning
Xinyuan Gao, Songlin Dong, Yuhang He 0001, Yihong Gong |
ECCV (85) | 5 |
| 2024 | Non-exemplar Domain Incremental Learning via Cross-Domain Concept Integration
Yuhang He 0001, Songlin Dong, Xinyuan Gao, Shaokun Wang, Yihong Gong |
ECCV (49) | 6 |
| 2024 | Projecting Points to Axes: Oriented Object Detection via Point-Axis Representation
Zeyang Zhao, Qilong Xue, Yuhang He 0001, Yifan Bai 0001, Xing Wei 0001, Yihong Gong |
ECCV (28) | 6 |
| 2024 | Enhancing Pre-trained ViTs for Downstream Task Adaptation: A Locality-Aware Prompt Learning MethodabstractVision Transformers (ViTs) excel in extracting global information from image patches. However, their inherent limitation lies in effectively extracting information within local regions, hindering their applicability and performance. Particularly, fully supervised pre-trained ViTs, such as Vanilla ViT and CLIP, face the challenge of locality vanishing when adapting to downstream tasks. To address this, we introduce a novel LOcality-aware pRompt lEarning (LORE) method, aiming to improve the adaptation of pre-trained ViTs to downstream tasks. LORE integrates a data-driven Black Box module (i.e., a pre-trained ViT encoder) with a knowledge-driven White Box module. The White Box module is a locality-aware prompt learning mechanism to compensate for ViTs' deficiency in incorporating local information. More specifically, it begins with the design of a Locality Interaction Network (LIN), which treats an image as a neighbor graph and employs graph convolution operations to enhance local relationships among image patches. Subsequently, a Knowledge-Locality Attention (KLA) mechanism is proposed to capture critical local regions from images, learning Knowledge-Locality (K-L) prototypes utilizing relevant semantic knowledge. Afterwards, K-L prototypes guide the training of a Prompt Generator (PG) to generate locality-aware prompts for images. The locality-aware prompts, aggregating crucial local information, serve as additional input for our Black Box module. Combining pre-trained ViTs with our locality-aware prompt learning mechanism, our Black-White Box model enables the capture of both global and local information, facilitating effective downstream task adaptation. Experimental evaluations across four downstream tasks demonstrate the effectiveness and superiority of our LORE. Shaokun Wang, Yuhang He 0001, Yihong Gong |
ACM Multimedia | 4 |
| 2024 | Prompt-Agnostic Adversarial Perturbation for Customized Diffusion ModelsabstractDiffusion models have revolutionized customized text-to-image generation, allowing for efficient synthesis of photos from personal data with textual descriptions. However, these advancements bring forth risks including privacy breaches and unauthorized replication of artworks. Previous researches primarily center around using “prompt-specific methods” to generate adversarial examples to protect personal images, yet the effectiveness of existing methods is hindered by constrained adaptability to different prompts.
In this paper, we introduce a Prompt-Agnostic Adversarial Perturbation (PAP) method for customized diffusion models. PAP first models the prompt distribution using a Laplace Approximation, and then produces prompt-agnostic perturbations by maximizing a disturbance expectation based on the modeled distribution.
This approach effectively tackles the prompt-agnostic attacks, leading to improved defense stability.
Extensive experiments in face privacy and artistic style protection, demonstrate the superior generalization of our method in comparison to existing techniques. Cong Wan, Yuhang He 0001, Xiang Song 0005, Yihong Gong |
NeurIPS | 4 |
| 2024 | Overcoming Catastrophic Forgetting for Multi-Label Class-Incremental LearningabstractDespite the recent progress of class-incremental learning (CIL) methods, their capabilities in real-world scenarios such as multi-label settings remain unexplored. This paper focuses on a more practical CIL problem named multi-label class-incremental learning (MLCIL). MLCIL requires the vision models to overcome catastrophic forgetting of old knowledge while learning new classes from multi-label samples. Direct application of existing CIL methods to MLCIL leads to label absence, representative sample selection, and feature dilution problems. To address these problems, we present a novel AdaPtive Pseudo-Label-drivEn (APPLE) framework consisting of three components. First, the adaptive pseudo-label strategy is proposed to solve the label absence problem, which leverages the old model to annotate old classes for new samples. Second, a cluster sampling strategy is proposed to obtain more diverse samples to alleviate catastrophic forgetting under the MLCIL setting better. Finally, a class attention decoder is designed to mitigate the object feature dilution problem in multi-label samples. The extensive experiments on PASCAL VOC 2007 and MS-COCO demonstrate that our proposed method significantly outperforms other representative state-of-the-art CIL methods. Xiang Song 0005, Kuang Shu, Songlin Dong, Xing Wei 0001, Yihong Gong |
WACV | 6 |
| 2024 | A 10-Gb/s low-power inverter-based optical receiver front-end in 0.13-μm CMOS process
Yihong Gong, Ruiyong Tu, Sini Wu, Qiyan Sun, Jinghu Li 0001, Zhicong Luo |
Integr. | 1 |
| 2024 | A rail-to-rail high speed comparator with LVDS output in 0.18-μm SiGe BiCMOS Technology
Qiyan Sun, Ruiyong Tu, Yihong Gong, Sini Wu, Jinghu Li 0001, Zhicong Luo |
Integr. | 4 |
| 2024 | Coplane-constrained sparse depth sampling and local depth propagation for depth estimation
Zhiwen Yang 0003, Chuqiao Chen, Hongkui Wang, Tingyu Wang 0002, Chenggang Yan 0001, Yihong Gong |
Image Vis. Comput. | 7 |
| 2024 | Few-shot online anomaly detection and segmentation
Shenxing Wei, Xing Wei 0001, Zhiheng Ma, Songlin Dong, Shaochen Zhang, Yihong Gong |
Knowl. Based Syst. | 6 |
| 2024 | Sharpness-Aware Lookahead for Accelerating Convergence and Improving GeneralizationabstractLookahead is a popular stochastic optimizer that can accelerate the training process of deep neural networks. However, the solutions found by Lookahead often generalize worse than those found by its base optimizers, such as SGD and Adam. To address this issue, we propose Sharpness-Aware Lookahead (SALA), a novel optimizer that aims to identify flat minima that generalize well. SALA divides the training process into two stages. In the first stage, the direction towards flat regions is determined by leveraging a quadratic approximation of the optimization trajectory, without incurring any extra computational overhead. In the second stage, however, it is determined by Sharpness-Aware Minimization (SAM), which is particularly effective in improving generalization at the terminal phase of training. In contrast to Lookahead, SALA retains the benefits of accelerated convergence while also enjoying superior generalization performance compared to the base optimizer. Theoretical analysis of the expected excess risk, as well as empirical results on canonical neural network architectures and datasets, demonstrate the advantages of SALA over Lookahead. It is noteworthy that with approximately 25% more computational overhead than the base optimizer, SALA can achieve the same generalization performance as SAM which requires twice the training budget of the base optimizer. Chengli Tan, Jiangshe Zhang 0001, Junmin Liu, Yihong Gong |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | Global self-sustaining and local inheritance for source-free unsupervised domain adaptation
Lin Peng 0003, Yuhang He 0001, Shaokun Wang, Xiang Song 0005, Songlin Dong, Xing Wei 0001, Yihong Gong |
Pattern Recognit. | 7 |
| 2024 | Domain Incremental Object Detection Based on Feature Space Topology Preserving StrategyabstractObject detection with the capacity to incrementally adapt to new domains is a crucial yet relatively under-explored research topic. The catastrophic forgetting problem presents a significant challenge to achieve this goal, where the model’s performance improves quickly in new conditions but deteriorates sharply in old ones after several incremental learning sessions. Drawing on recent discoveries in visual memories of the human brain, we introduce the Topology-Preserving Domain Incremental Object Detection (TP-DIOD) approach, which aims to address the catastrophic forgetting problem by extracting the topological structure of the feature space learned by the Convolutional Neural Network (CNN) model and preserving this topology during the subsequent incremental learning sessions. Specifically, we model the feature space topology using the self-organizing map (SOM) and construct an anchor image set based on the centroid vectors of the SOM nodes to memorize the feature space topology. We then develop the anchor loss function to penalize the topological changes of the feature space during the subsequent incremental learning sessions. Experimental evaluations on two sets of datasets demonstrate the effectiveness of the proposed TP-DIOD method in mitigating the catastrophic forgetting problem and achieving high accuracy on both old and new domain datasets. Xiang Song 0005, Yuhang He 0001, Changxin Wang, Songlin Dong, Xing Wei 0001, Yihong Gong |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2024 | Knowledge Synergy Learning for Multi-Modal TrackingabstractBenefiting from the rich information provided by different modalities, multi-modal tracking has shown significant improvements compared to single-modal tracking. However, in practical applications, multi-modal tracking still faces two major challenges. Firstly, it is crucial to effectively integrate the complementary information from different modalities in order to improve tracking performance. Secondly, as trackers are often deployed in dynamic environments, it is difficult to ensure complete multi-modal data. Thus, handling modal-missing issues is essential to achieve robust and reliable tracking. To address these challenges, this paper proposes a Knowledge Synergy Network (KSNet) that integrates multi-modal features into a comprehensive representation and incorporates a modal compensation mechanism to handle modal-missing issues. With this framework, a multi-modal tracker (KSTrack) is built and trained using multi-modal data. KSTrack is capable of handling both complete and incomplete multi-modal data during inference. Comprehensive experiments on four large-scale RGB-Thermal (RGB-T) and RGB-Depth (RGB-D) benchmarks show that KSTrack surpasses state-of-the-art multi-modal trackers when using multi-modal data and outperforms single-modal trackers by a large margin when using single-modal data. Yuhang He 0001, Zhiheng Ma, Xing Wei 0001, Yihong Gong |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Analogical Learning-Based Few-Shot Class-Incremental LearningabstractFSCIL (Few-shot class-incremental learning) is a prominent research topic in the ML community. It faces two significant challenges: forgetting old class knowledge and overfitting to limited new class training examples. In this paper, we present a novel FSCIL approach inspired by the human brain’s analogical learning mechanism, which enables human beings to form knowledge about a target domain from the knowledge of the source domains that are analogical to the target in some aspects. The proposed analogical learning-based FSCIL (ALFSCIL) method consists of two major components: new class classifier constructor (NCCC) and Meta-Analogical training (MAT). The NCCC module utilizes a multi-head cross-attention transformer to compute analogies between new and old classes, generating new class classifiers by blending old class classifiers based on the computed analogies. The MAT module updates the parameters of the CNN feature extractor, the NCCC module, and the knowledge for each encountered class after each round of the FSCIL session. We turn the optimization process into a bi-level optimization problem(BOP) whose theoretical analysis proves the stability and plasticity of our proposed model. Experimental evaluations reveal that this proposed ALFSCIL method achieves the SOTA performance accuracies on three benchmark datasets: CIFAR100, miniImageNet, and CUB200. Jiashuo Li, Songlin Dong, Yihong Gong, Yuhang He 0001, Xing Wei 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Exploiting Multi-Scale Parallel Self-Attention and Local Variation via Dual-Branch Transformer-CNN Structure for Face Super-ResolutionabstractRecently, deep learning technique has been widely employed to deal with face super-resolution (FSR) problem. It aims to predict the nonlinear relationship between the low-resolution (LR) face images and corresponding high-resolution (HR) ones, which could recover the high-frequency details from the LR degraded textures. However, either CNN-based or Transformer-based approaches mostly enhance the details by exploiting the relationship of local pixels or patches on LR features, the nonlocal features are not fully taken into account for producing high-frequency textures. To improve the above problem, we design a novel dual-branch module which consists of Transformer and CNN respectively. The Transformer branch extracts multiple scale feature embeddings and explores local and nonlocal self-attention simultaneously. Thus, the parallel self-attention mechanism has superior capabilities to capture the local and nonlocal dependencies on face image in the face reconstruction. Furthermore, the traditional CNNs usually extract features by combining pixels in a local convolutional kernel, it may be not effective to recover lost high-frequency details since the variations of local pixels are not well measured, which is important in recovering vivid edges and contours. To this end, we propose the local variation based attention block on the CNN branch, which could enhance the capabilities by directly extracting features from the variation of neighboring pixels. Finally, the Transformer-branch and CNN-branch are combined together by the modulation block to fuse both nonlocal and local advantages from two branches. Experimental results demonstrate the effectiveness of the proposed method when compared with state-of-the-art approaches. Jingang Shi, Yusi Wang, Zitong Yu, Guanxin Li, Xiaopeng Hong, Fei Wang 0037, Yihong Gong |
IEEE Trans. Multim. | 7 |
| 2024 | Brain Cognition-Inspired Dual-Pathway CNN Architecture for Image ClassificationabstractInspired by the global-local information processing mechanism in the human visual system, we propose a novel convolutional neural network (CNN) architecture named cognition-inspired network (CogNet) that consists of a global pathway, a local pathway, and a top-down modulator. We first use a common CNN block to form the local pathway that aims to extract fine local features of the input image. Then, we use a transformer encoder to form the global pathway to capture global structural and contextual information among local parts in the input image. Finally, we construct the learnable top-down modulator where fine local features of the local pathway are modulated by global representations of the global pathway. For ease of use, we encapsulate the dual-pathway computation and modulation process into a building block, called the global-local block (GL block), and a CogNet of any depth can be constructed by stacking a necessary number of GL blocks one after another. Extensive experimental evaluations have revealed that the proposed CogNets have achieved the state-of-the-art performance accuracies on all the six benchmark datasets and are very effective for overcoming the "texture bias" and the "semantic confusion" problems faced by many CNN models. Songlin Dong, Yihong Gong, Jingang Shi, Miao Shang, Xing Wei 0001, Xiaopeng Hong, Tiangang Zhou |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | Deep Class-Incremental Learning From Decentralized DataabstractIn this article, we focus on a new and challenging decentralized machine learning paradigm in which there are continuous inflows of data to be addressed and the data are stored in multiple repositories. We initiate the study of data-decentralized class-incremental learning (DCIL) by making the following contributions. First, we formulate the DCIL problem and develop the experimental protocol. Second, we introduce a paradigm to create a basic decentralized counterpart of typical (centralized) CIL approaches, and as a result, establish a benchmark for the DCIL study. Third, we further propose a decentralized composite knowledge incremental distillation (DCID) framework to transfer knowledge from historical models and multiple local sites to the general model continually. DCID consists of three main components, namely, local CIL, collaborated knowledge distillation (KD) among local models, and aggregated KD from local models to the general one. We comprehensively investigate our DCID framework by using a different implementation of the three components. Extensive experimental results demonstrate the effectiveness of our DCID framework. The source code of the baseline methods and the proposed DCIL is available at https://github.com/Vision-Intelligence-and-Robots-Group/DCIL. Songlin Dong, Jinjie Chen, Qi Tian 0001, Yihong Gong, Xiaopeng Hong |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2023 | DKT: Diverse Knowledge Transfer Transformer for Class Incremental LearningabstractIn the context of incremental class learning, deep neural networks are prone to catastrophic forgetting, where the accuracy of old classes declines substantially as new knowledge is learned. While recent studies have sought to address this issue, most approaches suffer from either the stability-plasticity dilemma or excessive computational and parameter requirements. To tackle these challenges, we propose a novel framework, the Diverse Knowledge Transfer Transformer (DKT), which incorporates two knowledge transfer mechanisms that use attention mechanisms to transfer both task-specific and task-general knowledge to the current task, along with a duplex classifier to address the stability-plasticity dilemma. Additionally, we design a loss function that clusters similar categories and discriminates between old and new tasks in the feature space. The proposed method requires only a small number of extra parameters, which are negligible in comparison to the increasing number of tasks. We perform extensive experiments on CIFAR100, ImageNet100, and ImageNet1000 datasets, which demonstrate that our method outperforms other competitive methods and achieves state-of-the-art performance. Our source code is available at https://github.com/MIVXJTU/DKT. Xinyuan Gao, Yuhang He 0001, Songlin Dong, Xing Wei 0001, Yihong Gong |
CVPR | 6 |
| 2023 | Autoregressive Visual TrackingabstractWe present ARTrack, an autoregressive framework for visual object tracking. ARTrack tackles tracking as a coordinate sequence interpretation task that estimates object trajectories progressively, where the current estimate is induced by previous states and in turn affects subsequences. This time-autoregressive approach models the sequential evolution of trajectories to keep tracing the object across frames, making it superior to existing template matching based trackers that only consider the per-frame localization accuracy. ARTrack is simple and direct, eliminating customized localization heads and post-processings. Despite its simplicity, ARTrack achieves state-of-the-art performance on prevailing benchmark datasets. Source code is available at https://github.com/MIV-XJTU/ARTrack. Xing Wei 0001, Yifan Bai 0001, Yongchao Zheng, Dahu Shi, Yihong Gong |
CVPR | 5 |
| 2023 | Knowledge Restore and Transfer for Multi-Label Class-Incremental LearningabstractCurrent class-incremental learning research mainly focuses on single-label classification tasks while multi-label class-incremental learning (MLCIL) with more practical application scenarios is rarely studied. Although there have been many anti-forgetting methods to solve the problem of catastrophic forgetting in single-label class-incremental learning, these methods have difficulty in solving the MLCIL problem due to label absence and information dilution problems. To solve these problems, we propose a Knowledge Restore and Transfer (KRT) framework containing two key components. First, a dynamic pseudo-label (DPL) module is proposed to solve the label absence problem by restoring the knowledge of old classes to the new data. Second, an incremental cross-attention (ICA) module is designed to maintain and transfer the old knowledge to solve the information dilution problem. Comprehensive experimental results on MS-COCO and PASCAL VOC datasets demonstrate the effectiveness of our method for improving recognition performance and mitigating forgetting on multi-label class-incremental learning tasks. The source code is available at https://gith.ub.com/witdsl/KRT-MLCIL. Songlin Dong, Haoyu Luo, Yuhang He 0001, Xing Wei 0001, Yihong Gong |
ICCV | 6 |
| 2023 | Learning Attention from Attention: Efficient Self-Refinement Transformer for Face Super-ResolutionabstractRecently, Transformer-based architecture has been introduced into face super-resolution task due to its advantage in capturing long-range dependencies. However, these approaches tend to integrate global information in a large searching region, which neglect to focus on the most relevant information and induce blurry effect by the irrelevant textures. Some improved methods simply constrain self-attention in a local window to suppress the useless information. But it also limits the capability of recovering high-frequency details when flat areas dominate the local searching window. To improve the above issues, we propose a novel self-refinement mechanism which could adaptively achieve texture-aware reconstruction in a coarse-to-fine procedure. Generally, the primary self-attention is first conducted to reconstruct the coarse-grained textures and detect the fine-grained regions required further compensation. Then, region selection attention is performed to refine the textures on these key regions. Since self-attention considers the channel information on tokens equally, we employ a dual-branch feature integration module to privilege the important channels in feature extraction. Furthermore, we design the wavelet fusion module which integrate shallow-layer structure and deep-layer detailed feature to recover realistic face images in frequency domain. Extensive experiments demonstrate the effectiveness on a variety of datasets. Guanxin Li, Jingang Shi, Yuan Zong, Fei Wang 0037, Tian Wang 0002, Yihong Gong |
IJCAI | 6 |
| 2023 | Non-Exemplar Class-Incremental Learning via Adaptive Old Class ReconstructionabstractIn the Class-Incremental Learning (CIL) task, rehearsal-based approaches have received a lot of attention recently. However, storing old class samples is often infeasible in application scenarios where device memory is insufficient or data privacy is important. Therefore, it is necessary to rethink Non-Exemplar Class-Incremental Learning (NECIL). In this paper, we propose a novel NECIL method named POLO with an adaPtive Old cLass recOnstruction mechanism, in which a density-based prototype reinforcement method (DBR), a topology-correction prototype adaptation method (TPA), and an adaptive prototype augmentation method (APA) are designed to reconstruct pseudo features of old classes in new incremental sessions. Specifically, the DBR focuses on the low-density features to maintain the model's discriminative ability for old classes. Afterward, the TPA is designed to adapt old class prototypes to new feature spaces in the incremental learning process. Finally, the APA is developed to further adapt pseudo feature spaces of old classes to new feature spaces. Experimental evaluations on four benchmark datasets demonstrate the effectiveness of our proposed method over the state-of-the-art NECIL methods. Shaokun Wang, Weiwei Shi 0003, Yuhang He 0001, Yihong Gong |
ACM Multimedia | 5 |
| 2023 | Topology-preserving transfer learning for weakly-supervised anomaly detection and segmentation
Shenxing Wei, Xing Wei 0001, Muhammad Rifki Kurniawan, Zhiheng Ma, Yihong Gong |
Pattern Recognit. Lett. | 5 |
| 2023 | Semantic Knowledge Guided Class-Incremental LearningabstractDriven by practical needs, research on Class-Incremental Learning (CIL) has received more and more attentions in recent years. A technical challenge to be conquered by CIL methods is the catastrophic forgetting problem, where the model’s performance improves rapidly on new classes while deteriorates drastically on old ones. The main causes behind catastrophic forgetting include network drifts, inter-class confusions, etc. In this paper, we propose a novel CIL method that solves the catastrophic forgetting problem from two aspects. First, to solve the inter-class confusion problem, we propose a novel Semantic knOwledge gUided ciL framework (SOUL) that consists of a CNN feature extractor and a Bi-GCN (Graph Convolutional Network) classifier. In each CIL session, we use the semantic knowledge extracted from the class labels to build two inter-class relation graphs among all the encountered old and new classes. Using these two relation graphs, we develop a Bi-GCN classifier to fuse two kinds of semantic relations in a balanced way, and then to transfer the inter-class relations from semantic modality to image classification weights. The entire SOUL framework is trained end-to-end by the standard BP algorithm, which optimizes the Bi-GCN classifier and the CNN feature extractor jointly. Second, to prevent the network drift, we develop the local topology preserving strategy that divides the global topological structure of the learned feature space into a set of local topological relations, and maintains these local relations at CIL session. Experimental evaluations demonstrate the state-of-the-art performance accuracies on benchmark image classification datasets. Shaokun Wang, Weiwei Shi 0003, Songlin Dong, Xinyuan Gao, Xiang Song 0005, Yihong Gong |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2023 | Semi-Supervised Crowd Counting via Multiple Representation LearningabstractThere has been a growing interest in counting crowds through computer vision and machine learning techniques in recent years. Despite that significant progress has been made, most existing methods heavily rely on fully-supervised learning and require a lot of labeled data. To alleviate the reliance, we focus on the semi-supervised learning paradigm. Usually, crowd counting is converted to a density estimation problem. The model is trained to predict a density map and obtains the total count by accumulating densities over all the locations. In particular, we find that there could be multiple density map representations for a given image in a way that they differ in probability distribution forms but reach a consensus on their total counts. Therefore, we propose multiple representation learning to train several models. Each model focuses on a specific density representation and utilizes the count consistency between models to supervise unlabeled data. To bypass the explicit density regression problem, which makes a strong parametric assumption on the underlying density distribution, we propose an implicit density representation method based on the kernel mean embedding. Extensive experiments demonstrate that our approach outperforms state-of-the-art semi-supervised methods significantly. Xing Wei 0001, Yunfeng Qiu, Zhiheng Ma, Xiaopeng Hong, Yihong Gong |
IEEE Trans. Image Process. | 5 |
| 2023 | Model Behavior Preserving for Class-Incremental LearningabstractDeep models have shown to be vulnerable to catastrophic forgetting, a phenomenon that the recognition performance on old data degrades when a pre-trained model is fine-tuned on new data. Knowledge distillation (KD) is a popular incremental approach to alleviate catastrophic forgetting. However, it usually fixes the absolute values of neural responses for isolated historical instances, without considering the intrinsic structure of the responses by a convolutional neural network (CNN) model. To overcome this limitation, we recognize the importance of the global property of the whole instance set and treat it as a behavior characteristic of a CNN model relevant to model incremental learning. On this basis: 1) we design an instance neighborhood-preserving (INP) loss to maintain the order of pair-wise instance similarities of the old model in the feature space; 2) we devise a label priority-preserving (LPP) loss to preserve the label ranking lists within instance-wise label probability vectors in the output space; and 3) we introduce an efficient derivable ranking algorithm for calculating the two loss functions. Extensive experiments conducted on CIFAR100 and ImageNet show that our approach achieves the state-of-the-art performance. Xiaopeng Hong, Songlin Dong, Jingang Shi, Yihong Gong |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2022 | IDPT: Interconnected Dual Pyramid Transformer for Face Super-ResolutionabstractFace Super-resolution (FSR) task works for generating high-resolution (HR) face images from the corresponding low-resolution (LR) inputs, which has received a lot of attentions because of the wide application prospects. However, due to the diversity of facial texture and the difficulty of reconstructing detailed content from degraded images, FSR technology is still far away from being solved. In this paper, we propose a novel and effective face super-resolution framework based on Transformer, namely Interconnected Dual Pyramid Transformer (IDPT). Instead of straightly stacking cascaded feature reconstruction blocks, the proposed IDPT designs the pyramid encoder/decoder Transformer architecture to extract coarse and detailed facial textures respectively, while the relationship between the dual pyramid Transformers is further explored by a bottom pyramid feature extractor. The pyramid encoder/decoder structure is devised to adapt various characteristics of textures in different spatial spaces hierarchically. A novel fusing modulation module is inserted in each spatial layer to guide the refinement of detailed texture by the corresponding coarse texture, while fusing the shallow-layer coarse feature and corresponding deep-layer detailed feature simultaneously. Extensive experiments and visualizations on various datasets demonstrate the superiority of the proposed method for face super-resolution tasks. Jingang Shi, Yusi Wang, Songlin Dong, Xiaopeng Hong, Zitong Yu, Fei Wang 0037, Changxin Wang, Yihong Gong |
IJCAI | 8 |
| 2022 | Identity-Quantity Harmonic Multi-Object TrackingabstractThe data association problem of multi-object tracking (MOT) aims to assign IDentity (ID) labels to detections and infer a complete trajectory for each target. Most existing methods assume that each detection corresponds to a unique target and thus cannot handle situations when multiple targets occur in a single detection due to detection failure in crowded scenes. To relax this strong assumption for practical applications, we formulate the MOT as a Maximizing An Identity-Quantity Posterior (MAIQP) problem on the basis of associating each detection with an identity and a quantity characteristic and then provide solutions to tackle two key problems arising. Firstly, a local target quantification module is introduced to count the number of targets within one detection. Secondly, we propose an identity-quantity harmony mechanism to reconcile the two characteristics. On this basis, we develop a novel Identity-Quantity HArmonic Tracking (IQHAT) framework that allows assigning multiple ID labels to detections containing several targets. Through extensive experimental evaluations on five benchmark datasets, we demonstrate the superiority of the proposed method. Yuhang He 0001, Xing Wei 0001, Xiaopeng Hong, Wei Ke 0003, Yihong Gong |
IEEE Trans. Image Process. | 5 |
| 2022 | Transductive Semisupervised Deep HashingabstractDeep hashing methods have shown their superiority to traditional ones. However, they usually require a large amount of labeled training data for achieving high retrieval accuracies. We propose a novel transductive semisupervised deep hashing (TSSDH) method which is effective to train deep convolutional neural network (DCNN) models with both labeled and unlabeled training samples. TSSDH method consists of the following four main ingredients. First, we extend the traditional transductive learning (TL) principle to make it applicable to DCNN-based deep hashing. Second, we introduce confidence levels for unlabeled samples to reduce adverse effects from uncertain samples. Third, we employ a Gaussian likelihood loss for hash code learning to sufficiently penalize large Hamming distances for similar sample pairs. Fourth, we design the large-margin feature (LMF) regularization to make the learned features satisfy that the distances of similar sample pairs are minimized and the distances of dissimilar sample pairs are larger than a predefined margin. Comprehensive experiments show that the TSSDH method can produce superior image retrieval accuracies compared to the representative semisupervised deep hashing methods under the same number of labeled training samples. Weiwei Shi 0003, Yihong Gong, Badong Chen, Xinhong Hei 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2021 | Few-Shot Class-Incremental Learning via Relation Knowledge DistillationabstractIn this paper, we focus on the challenging few-shot class incremental learning (FSCIL) problem, which requires to transfer knowledge from old tasks to new ones and solves catastrophic forgetting. We propose the exemplar relation distillation incremental learning framework to balance the tasks of old-knowledge preserving and new-knowledge adaptation. First, we construct an exemplar relation graph to represent the knowledge learned by the original network and update gradually for new tasks learning. Then an exemplar relation loss function for discovering the relation knowledge between different classes is introduced to learn and transfer the structural information in relation graph. A large number of experiments demonstrate that relation knowledge does exist in the exemplars and our approach outperforms other state-of-the-art class-incremental learning methods on the CIFAR100, miniImageNet, and CUB200 datasets. Songlin Dong, Xiaopeng Hong, Xinyuan Chang, Xing Wei 0001, Yihong Gong |
AAAI | 6 |
| 2021 | Error-Aware Density Isomorphism Reconstruction for Unsupervised Cross-Domain Crowd CountingabstractThis paper focuses on the unsupervised domain adaptation problem for video-based crowd counting, in which we use labeled data as source domain and unlabelled video data as target domain. It is challenging as there is a huge gap between the source and the target domain and no annotations of samples are available in the target domain. The key issue is how to utilize unlabelled videos in the target domain for knowledge learning and transferring from the source domain. To tackle this problem, we propose a novel Error-aware Density Isomorphism REConstruction Network (EDIREC-Net) for cross-domain crowd counting. EDIREC-Net jointly transfers a pre-trained counting model to target domains using a density isomorphism reconstruction objective and models the reconstruction erroneousness by error reasoning. Specifically, as crowd flows in videos are consecutive, the density maps in adjacent frames turn out to be isomorphic. On this basis, we regard the density isomorphism reconstruction error as a self-supervised signal to transfer the pre-trained counting models to different target domains. Moreover, we leverage an estimation-reconstruction consistency to monitor the density reconstruction erroneousness and suppress unreliable density reconstructions during training. Experimental results on four benchmark datasets demonstrate the superiority of the proposed method and ablation studies investigate the efficiency and robustness. The source code is available at https://github.com/GehenHe/EDIREC-Net. Yuhang He 0001, Zhiheng Ma, Xing Wei 0001, Xiaopeng Hong, Wei Ke 0003, Yihong Gong |
AAAI | 6 |
| 2021 | Learning to Count via Unbalanced Optimal TransportabstractCounting dense crowds through computer vision technology has attracted widespread attention. Most crowd counting datasets use point annotations. In this paper, we formulate crowd counting as a measure regression problem to minimize the distance between two measures with different supports and unequal total mass. Specifically, we adopt the unbalanced optimal transport distance, which remains stable under spatial perturbations, to quantify the discrepancy between predicted density maps and point annotations. An efficient optimization algorithm based on the regularized semi-dual formulation of UOT is introduced, which alternatively learns the optimal transportation and optimizes the density regressor. The quantitative and qualitative results illustrate that our method achieves state-of-the-art counting and localization performance. Zhiheng Ma, Xing Wei 0001, Xiaopeng Hong, Yunfeng Qiu, Yihong Gong |
AAAI | 6 |
| 2021 | Towards A Universal Model for Cross-Dataset Crowd CountingabstractThis paper proposes to handle the practical problem of learning a universal model for crowd counting across scenes and datasets. We dissect that the crux of this problem is the catastrophic sensitivity of crowd counters to scale shift, which is very common in the real world and caused by factors such as different scene layouts and image resolutions. Therefore it is difficult to train a universal model that can be applied to various scenes. To address this problem, we propose scale alignment as a prime module for establishing a novel crowd counting framework. We derive a closed-form solution to get the optimal image rescaling factors for alignment by minimizing the distances between their scale distributions. A novel neural network together with a loss function based on an efficient sliced Wasserstein distance is also proposed for scale distribution estimation. Benefiting from the proposed method, we have learned a universal model that generally works well on several datasets where can even outperform state-of-the-art models that are particularly fine-tuned for each dataset significantly. Experiments also demonstrate the much better generalizability of our model to unseen scenes. Zhiheng Ma, Xiaopeng Hong, Xing Wei 0001, Yunfeng Qiu, Yihong Gong |
ICCV | 5 |
| 2021 | Anomaly Detection Via Self-Organizing MapabstractAnomaly detection plays a key role in industrial manufacturing for product quality control. Traditional methods for anomaly detection are rule-based with limited generalization ability. Recent methods based on supervised deep learning are more powerful but require large-scale annotated datasets for training. In practice, abnormal products are rare thus it is very difficult to train a deep model in a fully supervised way. In this paper, we propose a novel unsupervised anomaly detection approach based on Self-organizing Map (SOM). Our method, Self-organizing Map for Anomaly Detection (SOMAD) maintains normal characteristics by using topological memory based on multi-scale features. SOMAD achieves state-of-the-art performance on unsupervised anomaly detection and localization on the MVTec dataset. Kaitao Jiang, Zhiheng Ma, Xing Wei 0001, Xiaopeng Hong, Yihong Gong |
ICIP | 6 |
| 2021 | Class Incremental Learning for Video Action ClassificationabstractClass Incremental Learning (CIL) is a hot topic in machine learning for CNN models to learn new classes incrementally. However, most of the CIL studies are for image classification and object recognition tasks and few CIL studies are available for video action classification. To mitigate this problem, in this paper, we present a new Grow When Required network (GWR) based video CIL framework for action classification. GWR learns knowledge incrementally by modeling the manifold of video frames for each encountered action class in feature space. We also introduce a Knowledge Consolidation (KC) method to separate the feature manifolds of old class and new class and introduce an associative matrix for label prediction. Experimental results on KTH and Weizmann demonstrate the effectiveness of the framework. Jiawei Ma, Jianxing Ma, Xiaopeng Hong, Yihong Gong |
ICIP | 5 |
| 2021 | Direct Measure Matching for Crowd CountingabstractTraditional crowd counting approaches usually use Gaussian assumption to generate pseudo density ground truth, which suffers from problems like inaccurate estimation of the Gaussian kernel sizes. In this paper, we propose a new measure-based counting approach to regress the predicted density maps to the scattered point-annotated ground truth directly. First, crowd counting is formulated as a measure matching problem. Second, we derive a semi-balanced form of Sinkhorn divergence, based on which a Sinkhorn counting loss is designed for measure matching. Third, we propose a self-supervised mechanism by devising a Sinkhorn scale consistency loss to resist scale changes. Finally, an efficient optimization method is provided to minimize the overall loss function. Extensive experiments on four challenging crowd counting datasets namely ShanghaiTech, UCF-QNRF, JHU++ and NWPU have validated the proposed method. Xiaopeng Hong, Zhiheng Ma, Xing Wei 0001, Yunfeng Qiu, Yaowei Wang 0001, Yihong Gong |
IJCAI | 7 |
| 2021 | Kohonen Self-Organizing Map based Route Planning: A RevisitabstractIn this paper, we revisit the long-standing Traveling Salesman Problem (TSP) and focus on the challenging, yet practical route planning problem with limited computational resources. We make contributions to TSP, one of the most famous NP-hard problems by providing a new improved approximate solution, which we term TOpology Preserving Self-Organizing Map (TOPSOM). TOPSOM well preserves the topology of the node map to be traversed by maintaining the continuity of nodes and the distances between them. In addition, to satisfy the requirements of convex hull, we design an elastic competitive Hebbian learning rule. TOPSOM can solve large-scale TSPs with high precision and high efficiency with limited computational costs. Extensive experimental results on mainstream route planning benchmarks including TSPLIB and National TSP’s show that our method consistently outperforms baseline methods, by up to 7.7% in terms of the Percent Deviation of Mean solution to best known solution. Qingshu Guan, Xiaopeng Hong, Wei Ke 0003, Liangfei Zhang, Guanghui Sun, Yihong Gong |
IROS | 6 |
| 2021 | Structural Knowledge Organization and Transfer for Class-Incremental LearningabstractDeep models are vulnerable to catastrophic forgetting when fine-tuned on new data. Popular distillation-based methods usually neglect the relations between data samples and may eventually forget essential structural knowledge. To solve these shortcomings, we propose a structural graph knowledge distillation based incremental learning framework to preserve both the positions of samples and their relations. Firstly, a memory knowledge graph (MKG) is generated to fully characterize the structural knowledge of historical tasks. Secondly, we develop a graph interpolation mechanism to enrich the domain of knowledge and alleviate the inter-class sample imbalance issue. Thirdly, we introduce structural graph knowledge distillation to transfer the knowledge of historical tasks. Comprehensive experiments on three datasets validate the proposed method. Xiaopeng Hong, Songlin Dong, Jingang Shi, Yihong Gong |
MMAsia | 6 |
| 2021 | Single-Image super-resolution - When model adaptation matters
Yudong Liang, Radu Timofte, Jinjun Wang, Sanping Zhou, Yihong Gong, Nanning Zheng 0001 |
Pattern Recognit. | 5 |
| 2021 | Beyond Universal Person Re-Identification AttackabstractDeep learning-based person re-identification (Re-ID) has made great progress and achieved high performance recently. In this paper, we make the first attempt to examine the vulnerability of current person Re-ID models against a dangerous attack method, i.e., the universal adversarial perturbation (UAP) attack, which has been shown to fool classification models with a little overhead. We propose a more universal adversarial perturbation (MUAP) method for both image-agnostic and model-insensitive person Re-ID attack. Firstly, we adopt a list-wise attack objective function to disrupt the similarity ranking list directly. Secondly, we propose a model-insensitive mechanism for cross-model attack. Extensive experiments show that the proposed attack approach achieves high attack performance and outperforms other state of the arts by large margin in cross-model scenario. The results also demonstrate the vulnerability of current Re-ID models to MUAP and further suggest the need of designing more robust Re-ID models. Xing Wei 0001, Rongrong Ji, Xiaopeng Hong, Qi Tian 0001, Yihong Gong |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2021 | Analogy-Detail Networks for Object RecognitionabstractThe human visual system can recognize object categories accurately and efficiently and is robust to complex textures and noises. To mimic the analogy-detail dual-pathway human visual cognitive mechanism revealed in recent cognitive science studies, in this article, we propose a novel convolutional neural network (CNN) architecture named analogy-detail networks (ADNets) for accurate object recognition. ADNets disentangle the visual information and process them separately using two pathways: the analogy pathway extracts coarse and global features representing the gist (i.e., shape and topology) of the object, while the detail pathway extracts fine and local features representing the details (i.e., texture and edges) for determining object categories. We modularize the architecture and encapsulate the two pathways into the analogy-detail block as the CNN building block to construct ADNets. For implementation, we propose a general principle that transmutes typical CNN structures into the ADNet architecture and applies the transmutation on representative baseline CNNs. Extensive experiments on CIFAR10, CIFAR100, street view house numbers, and ImageNet data sets demonstrate that ADNets significantly reduce the test error rates of the baseline CNNs by up to 5.76% and outperform other state-of-the-art architectures. Comprehensive analysis and visualizations further demonstrate that ADNets are interpretable and have a better shape-texture tradeoff for recognizing the objects with complex textures. Xiaopeng Hong, Weiwei Shi 0003, Xinyuan Chang, Yihong Gong |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2020 | Infrared-Visible Cross-Modal Person Re-Identification with an X ModalityabstractThis paper focuses on the emerging Infrared-Visible cross-modal person re-identification task (IV-ReID), which takes infrared images as input and matches with visible color images. IV-ReID is important yet challenging, as there is a significant gap between the visible and infrared images. To reduce this ‘gap’, we introduce an auxiliary X modality as an assistant and reformulate infrared-visible dual-mode cross-modal learning as an X-Infrared-Visible three-mode learning problem. The X modality restates from RGB channels to a format with which cross-modal learning can be easily performed. With this idea, we propose an X-Infrared-Visible (XIV) ReID cross-modal learning framework. Firstly, the X modality is generated by a lightweight network, which is learnt in a self-supervised manner with the labels inherited from visible images. Secondly, under the XIV framework, cross-modal learning is guided by a carefully designed modality gap constraint, with information exchanged cross the visible, X, and infrared modalities. Extensive experiments are performed on two challenging datasets SYSU-MM01 and RegDB to evaluate the proposed XIV-ReID approach. Experimental results show that our method considerably achieves an absolute gain of over 7% in terms of rank 1 and mAP even compared with the latest state-of-the-art methods. Diangang Li, Xing Wei 0001, Xiaopeng Hong, Yihong Gong |
AAAI | 4 |
| 2020 | Bi-Objective Continual Learning: Learning 'New' While Consolidating 'Known'abstractIn this paper, we propose a novel single-task continual learning framework named Bi-Objective Continual Learning (BOCL). BOCL aims at both consolidating historical knowledge and learning from new data. On one hand, we propose to preserve the old knowledge using a small set of pillars, and develop the pillar consolidation (PLC) loss to preserve the old knowledge and to alleviate the catastrophic forgetting problem. On the other hand, we develop the contrastive pillar (CPL) loss term to improve the classification performance, and examine several data sampling strategies for efficient onsite learning from ‘new’ with a reasonable amount of computational resources. Comprehensive experiments on CIFAR10/100, CORe50 and a subset of ImageNet validate the BOCL framework. We also reveal the performance accuracy of different sampling strategies when used to finetune a given CNN model. The code will be released. Xiaopeng Hong, Xinyuan Chang, Yihong Gong |
AAAI | 4 |
| 2020 | Superpixel Masking and Inpainting for Self-Supervised Anomaly Detection
Kaitao Jiang, Zhiheng Ma, Xing Wei 0001, Xiaopeng Hong, Yihong Gong |
BMVC | 7 |
| 2020 | Few-Shot Class-Incremental LearningabstractThe ability to incrementally learn new classes is crucial to the development of real-world artificial intelligence systems. In this paper, we focus on a challenging but practical few-shot class-incremental learning (FSCIL) problem. FSCIL requires CNN models to incrementally learn new classes from very few labelled samples, without forgetting the previously learned ones. To address this problem, we represent the knowledge using a neural gas (NG) network, which can learn and preserve the topology of the feature manifold formed by different classes. On this basis, we propose the TOpology-Preserving knowledge InCrementer (TOPIC) framework. TOPIC mitigates the forgetting of the old classes by stabilizing NG's topology and improves the representation learning for few-shot new classes by growing and adapting NG to new training samples. Comprehensive experimental results demonstrate that our proposed method significantly outperforms other state-of-the-art class-incremental learning methods on CIFAR100, miniImageNet, and CUB200 datasets. Xiaopeng Hong, Xinyuan Chang, Songlin Dong, Xing Wei 0001, Yihong Gong |
CVPR | 6 |
| 2020 | Topology-Preserving Class-Incremental Learning
Xinyuan Chang, Xiaopeng Hong, Xing Wei 0001, Yihong Gong |
ECCV (19) | 5 |
| 2020 | Complex Spatial-Temporal Attention Aggregation For Video Person Re-IdentificationabstractVideo-based person re-identification (Re-ID) aims to match pedestrian tracklets of the same identity captured by different cameras. Existing works usually compute the video-level feature representation via simple frame-level feature aggregation, such as average pooling and max pooling. However, the performance of such methods degenerates severely under low signal-noise ratio and partial occlusions. In this paper, we propose a novel Complex Spatial-Temporal Attention Aggregation (CAA), which fully exploits the discriminative information in spatial-temporal dimension via the combination of two aggregation method, namely region-aware aggregation and region-regardless aggregation. We evaluate the proposed method in three widely used video Re-ID datasets, including MARS, iLIDS-VID, and PRID-2011. The experimental results demonstrate that the proposed method outperforms the state of the arts. Xing Wei 0001, Xiaopeng Hong, Yihong Gong |
ICIP | 4 |
| 2020 | Class-Incremental Learning with Topological Schemas of Memory SpacesabstractClass-incremental learning (CIL) aims to incrementally learn a unified classifier for new classes emerging, which suffers from the catastrophic forgetting problem. To alleviate forgetting and improve the recognition performance, we propose a novel CIL framework, named the topological schemas model (TSM). TSM consists of a Gaussian mixture model arranged on 2D grids (2D-GMM) as the memory of the learned knowledge. To train the 2D-GMM model, we develop a novel competitive expectation-maximization (CEM) method, which contains a global topology embedding step and a local expectation-maximization fine-tuning step. Meanwhile, we choose the image samples of old classes that have the maximum posterior probability with respect to each Gaussian distribution as the episodic points. When finetuning for new classes, we propose the memory preservation loss (MPL) term to ensure episodic points still have maximum probabilities with respect to the corresponding Gaussian distribution. MPL preserves the distribution of 2D-GMM for old knowledge during incremental learning and alleviates catastrophic forgetting. Comprehensive experimental evaluations on two popular CIL benchmarks CIFAR100 and subImageNet demonstrate the superiority of our TSM. Xinyuan Chang, Xiaopeng Hong, Xing Wei 0001, Wei Ke 0003, Yihong Gong |
ICPR | 6 |
| 2020 | Polynomial Universal Adversarial Perturbations for Person Re-IdentificationabstractIn this paper, we focus on Universal Adversarial Perturbations (UAP) attack on state-of-the-art person re-identification (Re-ID) methods. Existing UAP methods usually compute a perturbation image and add it to the images of interest. Such a simple constant form greatly limits the attack power. To address this problem, we extend the formulation of UAP to a polynomial form and propose the Polynomial Universal Adversarial Perturbation (PUAP). Unlike traditional UAP methods which only rely on the additive perturbation signal, the proposed PUAP consists of both an additive perturbation and a multiplicative modulation factor. The additive perturbation produces the fundamental component of the signal, while the multiplicative factor modulates the perturbation signal in line with the unit impulse pattern of the input image. Moreover, we introduce a Pearson correlation coefficient loss to generate universal perturbations, for disrupting the outputs of person Re-ID models. Extensive experiments on DukeMTMC-reID, Market-1501, and MARS show that the proposed method can efficiently improve the attack performance, especially when the magnitude of UAP is constrained to a relatively small value. Xing Wei 0001, Rongrong Ji, Xiaopeng Hong, Yihong Gong |
ICPR | 5 |
| 2020 | Multi-Person Action Recognition in Microwave SensorsabstractThe usage of surveillance cameras for video understanding, raises concerns about privacy intrusion recently. This motivates the research community to seek potential alternatives of cameras for emerging multimedia applications. Stepping to this goal, a few researchers have explored the usage of Wi-Fi or Bluetooth sensors to handle action recognition. However, the practical ability of these sensors is limited by their frequency band and deployment inconvenience because of the separate transmitter/receiver architecture. Motivated by the same purpose of reducing privacy issues, we introduce a latest microwave sensor for multi-person action recognition in this paper. The microwave sensor works at 77GHz ~ 80GHz band, and is implemented with both transmitter and receiver inside itself, thus can be easily deployed for action recognition. Although with its advantages, two main challenging issues still remain. One is the difficulty of labelling the invisible signal data with embedding actions. The other is the difficulty of cancelling the environment noise for high-accurate action recognition. To address the challenges, we propose a novel learning framework by designed original loss functions with the considerations on weakly-supervised multi-label learning and attention mechanism to improve the accuracy for action recognition. We build a new microwave sensor data set, and conduct comprehensive experiments to evaluate the recognition accuracy of our proposed framework, and the effectiveness of parameters in each component. The experiment results show that our framework outperforms the state-of-the-art methods up to 14% in terms of mAP. Diangang Li, Jianquan Liu, Shoji Nishimura, Yuka Hayashi, Jun Suzuki 0004, Yihong Gong |
ACM Multimedia | 6 |
| 2020 | Learning Scales from Points: A Scale-aware Probabilistic Model for Crowd CountingabstractCounting people automatically through computer vision technology is a challenging task. Recently, convolution neural network (CNN) based methods have made significant progress. Nonetheless, large scale variations of instances caused by, for example, perspective effects remain unsolved. Moreover, it is problematic to estimate scales with only point annotations. In this paper, we propose a scale-aware probabilistic model to handle this problem. Unlike previous methods that generate a single density map where instances of various scales are processed indiscriminately, we propose a density pyramid network (DPN), where each pyramid level handles instances within a particular scale range. Furthermore, we propose a scale distribution estimator (SDE) to learn scales of people from input data, under the weak supervision of point annotations. Finally, we adopt an instance-level probabilistic scale-aware model (IPSM) to guide the multi-scale training of DPN explicitly. Qualitative and quantitative experimental results demonstrate the effectiveness of the proposed method, which achieves competitive results on four widely used benchmarks. Zhiheng Ma, Xing Wei 0001, Xiaopeng Hong, Yihong Gong |
ACM Multimedia | 4 |
| 2020 | Co-Attentive Lifting for Infrared-Visible Person Re-IdentificationabstractInfrared-visible cross-modality person re-identification (IV-ReID) has attracted much attention with the popularity of dual-mode video surveillance systems, where the RGB mode works in the daytime and automatically switches to the infrared mode at night. Despite its significant application value, IV-ReID remains a difficult problem mainly due to two great challenges. First, it is difficult to identify persons in the infrared image, which lacks color and texture clues. Second, there is a significant gap between the infrared and visible modalities where appearances of the same person vary considerably. This paper proposes a novel attention-based approach to handle the two difficulties in a unified framework. 1) We propose an attention lifting mechanism to learn discriminative features in each modality. 2) We propose a co-attentive learning mechanism to bridge the gap between the two modalities. Our method only makes slight modifications of a given backbone network and requires small computation overhead while improving the performance significantly. We conduct extensive experiments to demonstrate the superiority of our proposed method. Xing Wei 0001, Diangang Li, Xiaopeng Hong, Wei Ke 0003, Yihong Gong |
ACM Multimedia | 5 |
| 2020 | Tracking Persons-of-Interest via Unsupervised Representation Adaptation
Jia-Bin Huang 0001, Jongwoo Lim, Yihong Gong, Jinjun Wang, Narendra Ahuja, Ming-Hsuan Yang 0001 |
Int. J. Comput. Vis. | 4 |
| 2020 | Transductive semi-supervised metric learning for person re-identification
Xinyuan Chang, Zhiheng Ma, Xing Wei 0001, Xiaopeng Hong, Yihong Gong |
Pattern Recognit. | 5 |
| 2020 | Object detection with class aware region proposal network and focused attention objective
Yihong Gong, Weiwei Shi 0003, De Cheng |
Pattern Recognit. Lett. | 2 |
| 2020 | Fusion of Multiple Person Re-id Methods With Model and Data-Aware AbilitiesabstractPerson re-identification (person re-id) has attracted rapidly increasing attention in computer vision and pattern recognition research community in recent years. With the goal of providing match ranking results between each query person image and the gallery ones, the person re-id technique has been widely explored and a large number of person re-id methods have been developed. As these algorithms leverage different kinds of prior assumptions, image features, distance matching functions, et al., each of them has its own strengths and weaknesses. Inspired by these facts, this paper proposes a novel person re-id method based on the idea of inferring superior fusion results from a variety of previous base person re-id algorithms using different methodologies or features. To achieve this goal, we propose a novel framework which mainly consists of two steps: 1) a number of existing person re-id methods are implemented, and the ranking results are obtained in the test datasets. and 2) the robust fusion strategy is applied to obtain better re-ranked matching results by simultaneously considering the recognition abilities of various base re-id methods and the difficulties of different gallery person images to be correctly recognized under the generative model of labels, abilities, and difficulties framework. Comprehensive experiments show the effectiveness of our proposed method, and we have received state-of-the-art results on recent popular person re-id datasets. De Cheng, Zhihui Li 0001, Yihong Gong, Dingwen Zhang |
IEEE Trans. Cybern. | 3 |
| 2020 | Multi-Target Multi-Camera Tracking by Tracklet-to-Target AssignmentabstractThis paper focuses on the Multi-Target Multi-Camera Tracking task (MTMCT), which aims at tracking multiple targets within a multi-camera network. As the trajectory of each target is inherently split into multiple sub-trajectories (namely local tracklets) in a multi-camera network, a major challenge of MTMCT is how to accurately match the local tracklets generated within each camera across different cameras and generate a complete global trajectory for each target, i.e., the cross-camera tracklet matching problem. We solve the cross-camera tracklet matching problem by TRACklet-to-Target Assignment (TRACTA), and propose the Restricted Non-negative Matrix Factorization (RNMF) algorithm to compute the optimal assignment solution that meets a set of constraints, which should be in force in practice. TRACTA can correct the tracking errors caused by occlusions and missed detections in local tracklets, and produce a complete global trajectory for each target across all the cameras. Moreover, we also develop an analytical way of estimating the total number of targets in the camera network, which plays an important role to compute the tracklet-to-target assignment. Experimental evaluations and ablation studies on four MTMCT benchmark datasets show the superiority of the proposed TRACTA method. Yuhang He 0001, Xing Wei 0001, Xiaopeng Hong, Weiwei Shi 0003, Yihong Gong |
IEEE Trans. Image Process. | 5 |
| 2019 | Bayesian Loss for Crowd Count Estimation With Point SupervisionabstractIn crowd counting datasets, each person is annotated by a point, which is usually the center of the head. And the task is to estimate the total count in a crowd scene. Most of the state-of-the-art methods are based on density map estimation, which convert the sparse point annotations into a “ground truth” density map through a Gaussian kernel, and then use it as the learning target to train a density map estimator. However, such a "ground-truth" density map is imperfect due to occlusions, perspective effects, variations in object shapes, etc. On the contrary, we propose Bayesian loss, a novel loss function which constructs a density contribution probability model from the point annotations. Instead of constraining the value at every pixel in the density map, the proposed training loss adopts a more reliable supervision on the count expectation at each annotated point. Without bells and whistles, the loss function makes substantial improvements over the baseline loss on all tested datasets. Moreover, our proposed loss function equipped with a standard backbone network, without using any external detectors or multi-scale architectures, plays favourably against the state of the arts. Our method outperforms previous best approaches by a large margin on the latest and largest UCF-QNRF dataset. Zhiheng Ma, Xing Wei 0001, Xiaopeng Hong, Yihong Gong |
ICCV | 4 |
| 2019 | Consistency-Preserving deep hashing for fast person re-identification
Diangang Li, Yihong Gong, De Cheng, Weiwei Shi 0003, Xinyuan Chang |
Pattern Recognit. | 2 |
| 2019 | Normalized Non-Negative Sparse Encoder for Fast Image RepresentationabstractImage representation based on sparse coding generalizes the bag of words model. Although it reduces the reconstruction error for local features to achieve the state-of-the-art image classification performance, the large computational cost hinders the application of sparse coding-based image features. In this paper, we propose approximating a sparse code using the output of a simple neural network. The resulting parameter learning model for the neural network automatically incorporates non-negative and shift-invariant constraints, leading to an efficient normalized non-negative sparse coding (N3SC) sparse encoder. Without the use of the traditional iterative process to solve the sparse coding objective, the sparse encoder directly “converts” each local feature into a sparse code. We also introduce a method for training the encoder based on the auto-encoder method. In addition, we formally propose the corresponding sparse coding scheme called N3SC, which enforces both the non-negative constraint and the shift-invariant constraint in addition to the traditional sparse coding criteria. As demonstrated by several experiments, the obtained N3SC encoder requires only 3%-10% of the processing time for image feature extraction compared with the standard sparse coding scheme. At the same time, the features extracted using the exact solutions of the N3SC coding scheme and the N3SC encoder offer superior image classification accuracy compared to the accuracy of many existing sparse coding-based representations. Shizhou Zhang, Jinjun Wang, Weiwei Shi 0003, Yihong Gong, Yong Xia 0001, Yanning Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2019 | Discriminative Feature Learning With Foreground Attention for Person Re-IdentificationabstractThe performance of person re-identification (Re-ID) has been seriously affected by the large cross-view appearance variations caused by mutual occlusions and background clutter. Hence, learning a feature representation that can adaptively emphasize the foreground persons becomes very critical to solve the person Re-ID problem. In this paper, we propose a simple yet effective foreground attentive neural network (FANN) to learn a discriminative feature representation for person Re-ID, which can adaptively enhance the positive side of foreground and weaken the negative side of background. Specifically, a novel foreground attentive subnetwork is designed to drive the network’s attention, in which a decoder network is used to reconstruct the binary mask by using a novel local regression loss function, and an encoder network is regularized by the decoder network to focus its attention on the foreground persons. The resulting feature maps of encoder network are further fed into the body part subnetwork and feature fusion subnetwork to learn discriminative features. Besides, a novel symmetric triplet loss function is introduced to supervise feature learning, in which the intra-class distance is minimized and the inter-class distance is maximized in each triplet unit, simultaneously. Training our FANN in a multi-task learning framework, a discriminative feature representation can be learned to find out the matched reference to each probe among various candidates in the gallery. Extensive experimental results on several public benchmark datasets are evaluated, which have shown clear improvements of our method over the state-of-the-art approaches. Sanping Zhou, Jinjun Wang, Deyu Meng, Yudong Liang, Yihong Gong, Nanning Zheng 0001 |
IEEE Trans. Image Process. | 5 |
| 2019 | Fine-Grained Image Classification Using Modified DCNNs Trained by Cascaded Softmax and Generalized Large-Margin LossesabstractWe develop a fine-grained image classifier using a general deep convolutional neural network (DCNN). We improve the fine-grained image classification accuracy of a DCNN model from the following two aspects. First, to better model the h -level hierarchical label structure of the fine-grained image classes contained in the given training data set, we introduce h fully connected (fc) layers to replace the top fc layer of a given DCNN model and train them with the cascaded softmax loss. Second, we propose a novel loss function, namely, generalized large-margin (GLM) loss, to make the given DCNN model explicitly explore the hierarchical label structure and the similarity regularities of the fine-grained image classes. The GLM loss explicitly not only reduces between-class similarity and within-class variance of the learned features by DCNN models but also makes the subclasses belonging to the same coarse class be more similar to each other than those belonging to different coarse classes in the feature space. Moreover, the proposed fine-grained image classification framework is independent and can be applied to any DCNN structures. Comprehensive experimental evaluations of several general DCNN models (AlexNet, GoogLeNet, and VGG) using three benchmark data sets (Stanford car, fine-grained visual classification-aircraft, and CUB-200-2011) for the fine-grained image classification task demonstrate the effectiveness of our method. Weiwei Shi 0003, Yihong Gong, De Cheng, Nanning Zheng 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2018 | Kernelized Subspace Pooling for Deep Local DescriptorsabstractRepresenting local image patches in an invariant and discriminative manner is an active research topic in computer vision. It has recently been demonstrated that local feature learning based on deep Convolutional Neural Network (CNN) can significantly improve the matching performance. Previous works on learning such descriptors have focused on developing various loss functions, regularizations and data mining strategies to learn discriminative CNN representations. Such methods, however, have little analysis on how to increase geometric invariance of their generated descriptors. In this paper, we propose a descriptor that has both highly invariant and discriminative power. The abilities come from a novel pooling method, dubbed Subspace Pooling (SP) which is invariant to a range of geometric deformations. To further increase the discriminative power of our descriptor, we propose a simple distance kernel integrated to the marginal triplet loss that helps to focus on hard examples in CNN training. Finally, we show that by combining SP with the projection distance metric [13], the generated feature descriptor is equivalent to that of the Bilinear CNN model [22], but outperforms the latter with much lower memory and computation consumptions. The proposed method is simple, easy to understand and achieves good performance. Experimental results on several patch matching benchmarks show that our method outperforms the state-of-the-arts significantly. Xing Wei 0001, Yihong Gong, Nanning Zheng 0001 |
CVPR | 3 |
| 2018 | Transductive Semi-Supervised Deep Learning Using Min-Max Features
Weiwei Shi 0003, Yihong Gong, Chris Ding, Zhiheng Ma, Nanning Zheng 0001 |
ECCV (5) | 2 |
| 2018 | Grassmann Pooling as Compact Homogeneous Bilinear Pooling for Fine-Grained Visual Classification
Xing Wei 0001, Yihong Gong, Jiawei Zhang 0002, Nanning Zheng 0001 |
ECCV (3) | 3 |
| 2018 | Specular highlight reduction with known surface geometry
Xing Wei 0001, Xiaobin Xu 0001, Jiawei Zhang 0002, Yihong Gong |
Comput. Vis. Image Underst. | 4 |
| 2018 | Joint Contour Filtering
Xing Wei 0001, Qingxiong Yang, Yihong Gong |
Int. J. Comput. Vis. | 3 |
| 2018 | Pedestrian search in surveillance videos by learning discriminative deep features
Shizhou Zhang, De Cheng, Yihong Gong, Dahu Shi, Xi Qiu, Yong Xia 0001, Yanning Zhang 0001 |
Neurocomputing | 3 |
| 2018 | Person re-identification by the asymmetric triplet and identification loss function
De Cheng, Yihong Gong, Weiwei Shi 0003, Shizhou Zhang |
Multim. Tools Appl. | 2 |
| 2018 | Correction to: Person re-identification by the symmetric triplet and identification loss function
De Cheng, Yihong Gong, Weiwei Shi 0003, Shizhou Zhang |
Multim. Tools Appl. | 2 |
| 2018 | Deep Multimodal Feature Analysis for Action Recognition in RGB+D VideosabstractSingle modality action recognition on RGB or depth sequences has been extensively explored recently. It is generally accepted that each of these two modalities has different strengths and limitations for the task of action recognition. Therefore, analysis of the RGB+D videos can help us to better study the complementary properties of these two types of modalities and achieve higher levels of performance. In this paper, we propose a new deep autoencoder based shared-specific feature factorization network to separate input multimodal signals into a hierarchy of components. Further, based on the structure of the features, a structured sparsity learning machine is proposed which utilizes mixed norms to apply regularization within components and group selection between them for better classification performance. Our experimental results show the effectiveness of our cross-modality feature analysis framework by achieving state-of-the-art accuracy for action classification on five challenging benchmark datasets. Amir Shahroudy, Tian-Tsong Ng, Yihong Gong, Gang Wang 0012 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2018 | Deep feature learning via structured graph Laplacian embedding for person re-identification
De Cheng, Yihong Gong, Xiaojun Chang, Weiwei Shi 0003, Alex Hauptmann 0001, Nanning Zheng 0001 |
Pattern Recognit. | 2 |
| 2018 | Face alignment recurrent network
Qiqi Hou, Jinjun Wang, Ruibin Bai, Sanping Zhou, Yihong Gong |
Pattern Recognit. | 5 |
| 2018 | Entropy and orthogonality based deep discriminative feature learning for object recognition
Weiwei Shi 0003, Yihong Gong, De Cheng, Nanning Zheng 0001 |
Pattern Recognit. | 2 |
| 2018 | Deep self-paced learning for person re-identification
Sanping Zhou, Jinjun Wang, Deyu Meng, Xiaomeng Xin, Yihong Gong, Nanning Zheng 0001 |
Pattern Recognit. | 6 |
| 2018 | Superpixel HierarchyabstractSuperpixel segmentation has been one of the most important tasks in computer vision. In practice, an object can be represented by a number of segments at finer levels with consistent details or included in a surrounding region at coarser levels. Thus, a superpixel segmentation hierarchy is of great importance for applications that require different levels of image details. However, there is no method that can generate all scales of superpixels accurately in real time. In this paper, we propose the superhierarchy algorithm which is able to generate multi-scale superpixels as accurately as the state-of-the-art methods but with one to two orders of magnitude speed-up. The proposed algorithm can be directly integrated with recent efficient edge detectors to significantly outperform the state-of-the-art methods in terms of segmentation accuracy. Quantitative and qualitative evaluations on a number of applications demonstrate that the proposed algorithm is accurate and efficient in generating a hierarchy of superpixels. Xing Wei 0001, Qingxiong Yang, Yihong Gong, Narendra Ahuja, Ming-Hsuan Yang 0001 |
IEEE Trans. Image Process. | 3 |
| 2018 | Large Margin Learning in Set-to-Set Similarity Comparison for Person ReidentificationabstractPerson reidentification aims at matching images of the same person across disjoint camera views, which is a challenging problem in multimedia analysis, multimedia editing, and content-based media retrieval communities. The major challenge lies in how to preserve similarity of the same person across video footages with large appearance variations, while discriminating different individuals. To address this problem, conventional methods usually consider the pairwise similarity between persons by only measuring the point-to-point distance. In this paper, we propose using a deep learning technique to model a novel set-to-set (S2S) distance, in which the underline objective focuses on preserving the compactness of intraclass samples for each camera view, while maximizing the margin between the intraclass set and interclass set. The S2S distance metric consists of three terms, namely, the class-identity term, the relative distance term, and the regularization term. The class-identity term keeps the intraclass samples within each camera view gathering together, the relative distance term maximizes the distance between the intraclass class set and interclass set across different camera views, and the regularization term smoothes the parameters of the deep convolutional neural network. As a result, the final learned deep model can effectively find out the matched target to the probe object among various candidates in the video gallery by learning discriminative and stable feature representations. Using the CUHK01, CUHK03, PRID2011, and Market1501 benchmark datasets, we extensively conducted comparative evaluations to demonstrate the advantages of our method over the state-of-the-art approaches. Sanping Zhou, Jinjun Wang, Qiqi Hou, Yihong Gong, Nanning Zheng 0001 |
IEEE Trans. Multim. | 5 |
| 2018 | Improving CNN Performance Accuracies With Min-Max ObjectiveabstractWe propose a novel method for improving performance accuracies of convolutional neural network (CNN) without the need to increase the network complexity. We accomplish the goal by applying the proposed Min-Max objective to a layer below the output layer of a CNN model in the course of training. The Min-Max objective explicitly ensures that the feature maps learned by a CNN model have the minimum within-manifold distance for each object manifold and the maximum between-manifold distances among different object manifolds. The Min-Max objective is general and able to be applied to different CNNs with insignificant increases in computation cost. Moreover, an incremental minibatch training procedure is also proposed in conjunction with the Min-Max objective to enable the handling of large-scale training data. Comprehensive experimental evaluations on several benchmark data sets with both the image classification and face verification tasks reveal that employing the proposed Min-Max objective in the training process can remarkably improve performance accuracies of a CNN model in comparison with the same model trained without using this objective. Weiwei Shi 0003, Yihong Gong, Jinjun Wang, Nanning Zheng 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2018 | Training DCNN by Combining Max-Margin, Max-Correlation Objectives, and Correntropy Loss for Multilabel Image ClassificationabstractIn this paper, we build a multilabel image classifier using a general deep convolutional neural network (DCNN). We propose a novel objective function that consists of three parts, i.e., max-margin objective, max-correlation objective, and correntropy loss. The max-margin objective explicitly enforces that the minimum score of positive labels must be larger than the maximum score of negative labels by a predefined margin, which not only improves accuracies of the multilabel classifier, but also eases the threshold determination. The max-correlation objective can make the DCNN model learn a latent semantic space, which maximizes the correlations between the feature vectors of the training samples and their corresponding ground-truth label vectors projected into this space. Instead of using the traditional softmax loss, we adopt the correntropy loss from the information theory field to minimize the training errors of the DCNN model. The proposed framework can be end-to-end trained. Comprehensive experimental evaluations on Pascal VOC 2007 and MIR Flickr 25K multilabel benchmark data sets with four DCNN models, i.e., AlexNet, VGG-16, GoogLeNet, and ResNet demonstrate that the proposed objective function can remarkably improve the performance accuracies of a DCNN model for the task of multilabel image classification. Weiwei Shi 0003, Yihong Gong, Nanning Zheng 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2017 | Rank Ordering Constraints Elimination with Application for Kernel Learning
Ying Xie 0002, Chris Ding, Yihong Gong, Zongze Wu 0001 |
AAAI | 3 |
| 2017 | Point to Set Similarity Based Deep Feature Learning for Person Re-IdentificationabstractPerson re-identification (Re-ID) remains a challenging problem due to significant appearance changes caused by variations in view angle, background clutter, illumination condition and mutual occlusion. To address these issues, conventional methods usually focus on proposing robust feature representation or learning metric transformation based on pairwise similarity, using Fisher-type criterion. The recent development in deep learning based approaches address the two processes in a joint fashion and have achieved promising progress. One of the key issues for deep learning based person Re-ID is the selection of proper similarity comparison criteria, and the performance of learned features using existing criterion based on pairwise similarity is still limited, because only P2P distances are mostly considered. In this paper, we present a novel person Re-ID method based on P2S similarity comparison. The P2S metric can jointly minimize the intra-class distance and maximize the inter-class distance, while back-propagating the gradient to optimize parameters of the deep model. By utilizing our proposed P2S metric, the learned deep model can effectively distinguish different persons by learning discriminative and stable feature representations. Comprehensive experimental evaluations on 3DPeS, CUHK01, PRID2011 and Market1501 datasets demonstrate the advantages of our method over the state-of-the-art approaches. Sanping Zhou, Jinjun Wang, Yihong Gong, Nanning Zheng 0001 |
CVPR | 4 |
| 2017 | Discriminative Dictionary Learning With Ranking Metric Embedded for Person Re-IdentificationabstractThe goal of person re-identification (Re-Id) is to match pedestrians captured from multiple non-overlapping cameras. In this paper, we propose a novel dictionary learning based method with the ranking metric embedded, for person Re-Id. A new and essential ranking graph Laplacian term is introduced, which minimizes the intra-personal compactness and maximizes the inter-personal dispersion in the objective. Different from the traditional dictionary learning based approaches and their extensions, which just use the same or not information, our proposed method can explore the ranking relationship among the person images, which is essential for such retrieval related tasks. Simultaneously, one distance measurement has been explicitly learned in the model to further improve the performance. Since we have reformulated these ranking constraints into the graph Laplacian form, the proposed method is easy-to-implement but effective. We conduct extensive experiments on three widely used person Re-Id benchmark datasets, and achieve state-of-the-art performances. De Cheng, Xiaojun Chang, Li Liu 0031, Alex Hauptmann 0001, Yihong Gong, Nanning Zheng 0001 |
IJCAI | 5 |
| 2017 | Video Search via Ranking Network with Very Few Query Exemplars
De Cheng, Lu Jiang 0004, Yihong Gong, Nanning Zheng 0001, Alex Hauptmann 0001 |
MMM (2) | 3 |
| 2017 | Graph Matching via Multiplicative Update AlgorithmabstractAs a fundamental problem in computer vision, graph matching problem can usually be formulated as a Quadratic Programming (QP) problem with doubly stochastic and discrete (integer) constraints. Since it is NP-hard, approximate algorithms are required. In this paper, we present a new algorithm, called Multiplicative Update Graph Matching (MPGM), that develops a multiplicative update technique to solve the QP matching problem. MPGM has three main benefits: (1) theoretically, MPGM solves the general QP problem with doubly stochastic constraint naturally whose convergence and KKT optimality are guaranteed. (2) Em- pirically, MPGM generally returns a sparse solution and thus can also incorporate the discrete constraint approximately. (3) It is efficient and simple to implement. Experimental results show the benefits of MPGM algorithm. Bo Jiang 0002, Jin Tang 0001, Chris Ding, Yihong Gong, Bin Luo 0001 |
NIPS | 4 |
| 2017 | Part-aware trajectories association across non-overlapping uncalibrated cameras
De Cheng, Yihong Gong, Jinjun Wang, Qiqi Hou, Nanning Zheng 0001 |
Neurocomputing | 2 |
| 2017 | Combining local and global hypotheses in deep neural network for multi-label image classification
Qinghua Yu, Jinjun Wang, Shizhou Zhang, Yihong Gong, Jizhong Zhao |
Neurocomputing | 4 |
| 2017 | Correntropy-based level set method for medical image segmentation and bias correction
Sanping Zhou, Jinjun Wang, Yihong Gong |
Neurocomputing | 5 |
| 2017 | A Weight-Adaptive Laplacian Embedding for Graph-Based ClusteringabstractGraph-based clustering methods perform clustering on a fixed input data graph. Thus such clustering results are sensitive to the particular graph construction. If this initial construction is of low quality, the resulting clustering may also be of low quality. We address this drawback by allowing the data graph itself to be adaptively adjusted in the clustering procedure. In particular, our proposed weight adaptive Laplacian (WAL) method learns a new data similarity matrix that can adaptively adjust the initial graph according to the similarity weight in the input data graph. We develop three versions of these methods based on the L2-norm, fuzzy entropy regularizer, and another exponential-based weight strategy, that yield three new graph-based clustering objectives. We derive optimization algorithms to solve these objectives. Experimental results on synthetic data sets and real-world benchmark data sets exhibit the effectiveness of these new graph-based clustering methods. De Cheng, Feiping Nie 0001, Jiande Sun 0001, Yihong Gong |
Neural Comput. | 4 |
| 2017 | Constructing Deep Sparse Coding Network for image classification
Shizhou Zhang, Jinjun Wang, Yihong Gong, Nanning Zheng 0001 |
Pattern Recognit. | 4 |
| 2017 | Balanced Mixture of Deformable Part Models With Automatic Part ConfigurationsabstractThis paper presents a method to improve the traditional mixture of deformable part models (MDPM) method from the learning perspective. First, an object part configuration learning algorithm based on group sparsity constraint is introduced to automatically discover the object part number, size, and location. The algorithm imposes two additional regularization terms in addition to the standard hinge loss function. The first term focuses on automatic part selection and the second term focuses on automatic part placement. Second, this paper introduces an improved MDPM training framework. The framework applies a learned transformation to normalize the prediction score from each individual deformable part model (DPM) into a pseudo probability such that the partition of the entire object appearance feature space becomes less sensitive to the prior distributions of different DPMs. Finally, the two proposed improvements are combined and formulated under the expectation-maximization framework. We evaluate our method mainly using the PASCAL VOC2007 and VOC2010 detection benchmarks and show that the proposed learning algorithms could increase the detection mean AP score by 2.4% and 0.9%, respectively, on these two data sets when using the proposed part selection method and the training algorithm. We also present further in-depth analysis of the proposed algorithm in the experiments. De Cheng, Yihong Gong, Jingjun Wang, Nanning Zheng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2016 | Person Re-identification by Multi-Channel Parts-Based CNN with Improved Triplet Loss FunctionabstractPerson re-identification across cameras remains a very challenging problem, especially when there are no overlapping fields of view between cameras. In this paper, we present a novel multi-channel parts-based convolutional neural network (CNN) model under the triplet framework for person re-identification. Specifically, the proposed CNN model consists of multiple channels to jointly learn both the global full-body and local body-parts features of the input persons. The CNN model is trained by an improved triplet loss function that serves to pull the instances of the same person closer, and at the same time push the instances belonging to different persons farther from each other in the learned feature space. Extensive comparative evaluations demonstrate that our proposed method significantly outperforms many state-of-the-art approaches, including both traditional and deep network-based ones, on the challenging i-LIDS, VIPeR, PRID2011 and CUHK01 datasets. De Cheng, Yihong Gong, Sanping Zhou, Jinjun Wang, Nanning Zheng 0001 |
CVPR | 2 |
| 2016 | Image Quality Assessment Using Similar Scene as Reference
Yudong Liang, Jinjun Wang, Xingyu Wan, Yihong Gong, Nanning Zheng 0001 |
ECCV (5) | 4 |
| 2016 | Tracking Persons-of-Interest via Adaptive Discriminative Features
Yihong Gong, Jia-Bin Huang 0001, Jongwoo Lim, Jinjun Wang, Narendra Ahuja, Ming-Hsuan Yang 0001 |
ECCV (5) | 2 |
| 2016 | Improving CNN Performance with Min-Max Objective
Weiwei Shi 0003, Yihong Gong, Jinjun Wang |
IJCAI | 2 |
| 2016 | Improving DCNN Performance with Sparse Category-Selective Objective Function
Shizhou Zhang, Yihong Gong, Jinjun Wang |
IJCAI | 2 |
| 2016 | Incorporating image priors with deep convolutional neural networks for image super-resolution
Yudong Liang, Jinjun Wang, Sanping Zhou, Yihong Gong, Nanning Zheng 0001 |
Neurocomputing | 4 |
| 2016 | Active contour model based on local and global intensity information for medical image segmentation
Sanping Zhou, Jinjun Wang, Yudong Liang, Yihong Gong |
Neurocomputing | 5 |
| 2015 | Facial landmark detection via cascade multi-channel convolutional neural networkabstractThis paper presents a novel cascade multi-channel convolutional neural networks(CMC-CNN) approach for face alignment. Several CNN are jointly used for the finally output. In our method, each stage CNN takes the local region around the landmarks as input, and each local patches does convolution separately, which can lead network to learn local high-level features. Then a fully connected layer is put to learn global information from these local features. Our methods has achieves the state-of-the-art results when tested on the 300 Face in-the-Wild(300-W) dataset. Qiqi Hou, Jinjun Wang, Lele Cheng, Yihong Gong |
ICIP | 4 |
| 2015 | Incorporating image degeneration modeling with multitask learning for image super-resolutionabstractLearning the non-linear image upscaling process has previously been considered as a simple regression process, where various models have been utilized to describe the correlations between high-resolution (HR) and low-resolution (LR) images/patches. In this paper, we present a multitask learning framework based on deep neural network for image super-resolution, where we jointly consider the image super-resolution process and the image degeneration process. By sharing parameters between the two highly relevant tasks, the proposed framework could effectively improve the obtained neural network based mapping model between HR and LR image patches. Experimental results have demonstrated clear visual improvement and high computational efficiency, especially with large magnification factors. Yudong Liang, Jinjun Wang, Shizhou Zhang, Yihong Gong |
ICIP | 4 |
| 2015 | Multi-cue Normalized Non-Negative Sparse Encoder for image classificationabstractRecently, the sparse coding based image representation has achieved state-of-the-art recognition results on many benchmarks. In this paper, we propose Multi-cue Normalized Non-Negative Sparse Encoder (MN3SE) which enforces both the non-negative constraint and the shift-invariant constraint on top of the traditional sparse coding criteria, and takes multi-cue to further boost the performance. The former constraint reduces information loose by the negative coefficients and improves the coding stability, and the latter allows the sparseness to be self-adaptive to the local feature. The proposed coding scheme is then approximated by an neural network based encoder for speed-up. More importantly, the multi-layer neural network architecture allows us to apply a multi-task learning strategy to fuse information from multi-cue. Specifically, we take one type of descriptor, such as SIFT as the input, and enforce the learned encoder to produce sparse code that can reconstruct not only SIFT but also other types of descriptors such as color moments. In this way, we could achieve not only 10 to 33 times speed up for sparse-coding, the multi-cue enforced learning strategy gives the image feature extracted by MN3SE superior image classification accuracy. Shizhou Zhang, Jinjun Wang, Yudong Liang, Yihong Gong, Nanning Zheng 0001 |
ICME | 4 |
| 2015 | Deep Self-Organizing Map for visual classificationabstractWe proposed a Deep Self-Organizing Map (DSOM) algorithm which is completely different from the existing multi-layers SOM algorithms, such as SOINN. It consists of layers of alternating self-organizing map and sampling operator. The self-organizing layer is made up of certain numbers of SOMs, with each map only looking at a local region block on its input. The winning neuron's index value from every SOM in self-organizing layer is then organized in the sampling layer to generate another 2D map, which could then be fed to a second self-organizing layer. In this way, local information is gathered together, forming more global information in higher layers. The construction method of the DSOM is unique and will be introduced in this paper. Experiments were carried out to discuss how the DSOM architecture parameters affect the performance. We evaluate our proposed DSOM on MNIST and CASIA-HWDB1.1 dataset. Experimental results show that DSOM outperforms the original supervised SOM by 7:17% on MNIST and 7:25% on CASIA-HWDB1.1. Jinjun Wang, Yihong Gong |
IJCNN | 3 |
| 2015 | Robust Deep Auto-encoder for Occluded Face RecognitionabstractOcclusions by sunglasses, scarf, hats, beard, shadow etc, can significantly reduce the performance of face recognition systems. Although there exists a rich literature of researches focusing on face recognition with illuminations, poses and facial expression variations, there is very limited work reported for occlusion robust face recognition. In this paper, we present a method to restore occluded facial regions using deep learning technique to improve face recognition performance. Inspired by SSDA for facial occlusion removal with known occlusion type and explicit occlusion location detection from a preprocessing step, this paper further introduces Double Channel SSDA (DC-SSDA) which requires no prior knowledge of the types and the locations of occlusions. Experimental results based on CMU-PIE face database have showed that, the proposed method is robust to a variety of occlusion types and locations, and the restored faces could yield significant recognition performance improvements over occluded ones. Lele Cheng, Jinjun Wang, Yihong Gong, Qiqi Hou |
ACM Multimedia | 3 |
| 2015 | Training mixture of weighted SVM for object detection using EM algorithm
De Cheng, Jinjun Wang, Xing Wei 0001, Yihong Gong |
Neurocomputing | 4 |
| 2015 | Visual tracking based on online sparse feature learning
Zelun Wang, Jinjun Wang, Yihong Gong |
Image Vis. Comput. | 4 |
| 2015 | Multi-target tracking by learning local-to-global trajectory models
Jinjun Wang, Zelun Wang, Yihong Gong, Yuehu Liu |
Pattern Recognit. | 4 |
| 2015 | Online Multi-Target Tracking With Unified Handling of Complex ScenariosabstractComplex scenarios, including miss detections, occlusions, false detections, and trajectory terminations, make the data association challenging. In this paper, we propose an online tracking-by-detection method to track multiple targets with unified handling of aforementioned complex scenarios, where current detection responses are linked to the previous trajectories. We introduce a dummy node to each trajectory to allow it to temporally disappear. If a trajectory fails to find its matching detection, it will be linked to its corresponding dummy node until the emergence of its matching detection. Source nodes are also incorporated to account for the entrance of new targets. The standard Hungarian algorithm, extended by the dummy nodes, can be exploited to solve the online data association implicitly in a global manner, although it is formulated between two consecutive frames. Moreover, as dummy nodes tend to accumulate in a fake or disappeared trajectory while they only occasionally appear in a real trajectory, we can deal with false detections and trajectory terminations by simply checking the number of consecutive dummy nodes. Our approach works on a single, uncalibrated camera, and requires neither scene prior knowledge nor explicit occlusion reasoning, running at 132 frames/s on the PETS09-S2L1 benchmark sequence. The experimental results validate the effectiveness of the dummy nodes in complex scenarios and show that our proposed approach is robust against false detections and miss detections. Quantitative comparisons with other methods on five benchmark sequences demonstrate that we can achieve comparable results with the most existing offline methods and better results than other online algorithms. Huaizu Jiang, Jinjun Wang, Yihong Gong, Na Rong, Zhenhua Chai, Nanning Zheng 0001 |
IEEE Trans. Image Process. | 3 |
| 2014 | Low Computation Face Verification Using Class Center AnalysisabstractDespite the existence of many state-of-the-art face verification systems, the use of complex features and/or high order recognition models in these systems limits their application in devices with low computation power or low latency requirement. In this paper, we approach the problem by performing verification using simple linear distance model. We introduce a novel probability-based distance metric learning algorithm called Class Center Analysis (CCA) to improve the matching performance in a transformed space. CCA generalizes the classic Neighborhood Components Analysis (NCA) from two aspects. First NCA often leads to distributed clusters, while CCA produces more concentrative clusters, And second, NCA sometimes gives over-fitted distance transformation model, while CCA has better generalization ability. With CCA, our system is able to directly project the difference between face image pair into a real-valued score as their similarity, using only simple matrix-vector operation, and thus consuming very low computation. Our comprehensive experimental evaluation show that, CCA outperforms several other benchmark algorithms in verification accuracy. We have also built the CCA algorithm into a mobile application that uses face image for user authentication. Xinzi Zhang, Jinjun Wang, Yihong Gong, Shizhou Zhang |
ICPR | 3 |
| 2014 | Image parsing by loopy dynamic programming
Shizhou Zhang, Jinjun Wang, Yihong Gong, Xinzi Zhang, Xuguang Lan |
Neurocomputing | 3 |
| 2014 | A Two-Level Topic Model Towards Knowledge Discovery from Citation NetworksabstractKnowledge discovery from scientific articles has received increasing attention recently since huge repositories are made available by the development of the Internet and digital databases. In a corpus of scientific articles such as a digital library, documents are connected by citations and one document plays two different roles in the corpus: document itself and a citation of other documents. In the existing topic models, little effort is made to differentiate these two roles. We believe that the topic distributions of these two roles are different and related in a certain way. In this paper, we propose a Bernoulli process topic (BPT) model which considers the corpus at two levels: document level and citation level. In the BPT model, each document has two different representations in the latent topic space associated with its roles. Moreover, the multi-level hierarchical structure of citation network is captured by a generative process involving a Bernoulli process. The distribution parameters of the BPT model are estimated by a variational approximation approach. An efficient computation algorithm is proposed to overcome the difficulty of matrix inverse operation. In addition to conducting the experimental evaluations on the document modeling and document clustering tasks, we also apply the BPT model to well known corpora to discover the latent topics, recommend important citations, detect the trends of various research areas in computer science between 1991 and 1998, and to investigate the interactions among the research areas. The comparisons against state-of-the-art methods demonstrate a very promising performance. The implementations and the data sets are available online . Zhongfei Zhang, Shenghuo Zhu, Yun Chi, Yihong Gong |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2013 | Comparative Document Summarization via Discriminative Sentence SelectionabstractGiven a collection of document groups, a natural question is to identify the differences among these groups. Although traditional document summarization techniques can summarize the content of the document groups one by one, there exists a great necessity to generate a summary of the differences among the document groups. In this article, we study a novel problem of summarizing the differences between document groups. A discriminative sentence selection method is proposed to extract the most discriminative sentences that represent the specific characteristics of each document group. Experiments and case studies on real-world data sets demonstrate the effectiveness of our proposed method. Dingding Wang 0001, Shenghuo Zhu, Tao Li 0001, Yihong Gong |
ACM Trans. Knowl. Discov. Data | 4 |
| 2012 | Comparative document summarization via discriminative sentence selectionabstractGiven a collection of document groups, a natural question is to identify the differences among them. Although traditional document summarization techniques can summarize the content of the document groups one by one, there exists a great necessity to generate a summary of the differences among the document groups. In this article, we study a novel problem, that of summarizing the differences between document groups. A discriminative sentence selection method is proposed to extract the most discriminative sentences which represent the specific characteristics of each document group. Experiments and case studies on real-world data sets demonstrate the effectiveness of our proposed method. Dingding Wang 0001, Shenghuo Zhu, Tao Li 0001, Yihong Gong |
ACM Trans. Knowl. Discov. Data | 4 |
| 2012 | Discovering Image Semantics in Codebook Derivative SpaceabstractThe sparse coding based approaches for image recognition have recently shown improved performance than traditional bag-of-features technique. Due to high dimensionality of the image descriptor space, existing systems usually require very large codebook size to minimize coding error in order to get satisfactory accuracy. While most research efforts try to address the problem by constructing a relatively smaller codebook with stronger discriminative power, in this paper, we introduce an alternative solution by enhancing the quality of coding. Particularly, we apply the idea similar to Fisher kernel to the coding framework, where we use the image-dependent codebook derivative to represent the image. The proposed idea is generic across multiple coding criteria, and in this paper, it is applied to enhance the locality-constraint linear coding (LLC). Experiments show that, the extracted new feature, called “LLC+,” achieved significantly improved accuracy on several challenging datasets even with a small codebook of 1/20 the reported size used by LLC. This obviously adds to LLC+ the modeling accuracy, processing speed and codebook training advantages. Jinjun Wang, Yihong Gong |
IEEE Trans. Multim. | 2 |
| 2011 | Learning semantic embedding at a large scaleabstractA key problem in image annotation is to learn the underlying semantics. However, finding such semantic embeddings is a challenge task and often requires large amount of tagging information. In this paper, we propose to utilize multi-modality cues by incorporating visual and textual information as embedded objects. The paper further presents a multi-task learning framework that simultaneously learns the approximation of two semantic embeddings with efficient multi-stage convex relaxation technique. The experiments show that the proposed method presents very promising performance in both memory usage and training time for large-scale dataset, as well as image classification accuracy. Min-Hsuan Tsai, Jinjun Wang, Tong Zhang 0005, Yihong Gong, Thomas S. Huang |
ICIP | 4 |
| 2011 | Learning to Search Efficiently in High DimensionsabstractHigh dimensional similarity search in large scale databases becomes an important challenge due to the advent of Internet. For such applications, specialized data structures are required to achieve computational efficiency. Traditional approaches relied on algorithmic constructions that are often data independent (such as Locality Sensitive Hashing) or weakly dependent (such as kd-trees, k-means trees). While supervised learning algorithms have been applied to related problems, those proposed in the literature mainly focused on learning hash codes optimized for compact embedding of the data rather than search efficiency. Consequently such an embedding has to be used with linear scan or another search algorithm. Hence learning to hash does not directly address the search efficiency issue. This paper considers a new framework that applies supervised learning to directly optimize a data structure that supports efficient large scale search. Our approach takes both search quality and computational cost into consideration. Specifically, we learn a boosted search forest that is optimized using pair-wise similarity labeled examples. The output of this search forest can be efficiently converted into an inverted indexing data structure, which can leverage modern text search infrastructure to achieve both scalability and efficiency. Experimental results show that our approach significantly outperforms the start-of-the-art learning to hash methods (such as spectral hashing), as well as state-of-the-art high dimensional search algorithms (such as LSH and k-means trees). Zhen Li 0028, Huazhong Ning, Liangliang Cao, Tong Zhang 0001, Yihong Gong, Thomas S. Huang |
NIPS | 5 |
| 2011 | Detecting communities and their evolutions in dynamic social networks - a Bayesian approach
Tianbao Yang, Yun Chi, Shenghuo Zhu, Yihong Gong, Rong Jin 0001 |
Mach. Learn. | 4 |
| 2011 | Unsupervised Image Categorization by Hypergraph PartitionabstractWe present a framework for unsupervised image categorization in which images containing specific objects are taken as vertices in a hypergraph and the task of image clustering is formulated as the problem of hypergraph partition. First, a novel method is proposed to select the region of interest (ROI) of each image, and then hyperedges are constructed based on shape and appearance features extracted from the ROIs. Each vertex (image) and its k-nearest neighbors (based on shape or appearance descriptors) form two kinds of hyperedges. The weight of a hyperedge is computed as the sum of the pairwise affinities within the hyperedge. Through all of the hyperedges, not only the local grouping relationships among the images are described, but also the merits of the shape and appearance characteristics are integrated together to enhance the clustering performance. Finally, a generalized spectral clustering technique is used to solve the hypergraph partition problem. We compare the proposed method to several methods and its effectiveness is demonstrated by extensive experiments on three image databases. Yuchi Huang, Qingshan Liu 0001, Fengjun Lv, Yihong Gong, Dimitris N. Metaxas |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2011 | Integrating Document Clustering and Multidocument SummarizationabstractDocument understanding techniques such as document clustering and multidocument summarization have been receiving much attention recently. Current document clustering methods usually represent the given collection of documents as a document-term matrix and then conduct the clustering process. Although many of these clustering methods can group the documents effectively, it is still hard for people to capture the meaning of the documents since there is no satisfactory interpretation for each document cluster. A straightforward solution is to first cluster the documents and then summarize each document cluster using summarization methods. However, most of the current summarization methods are solely based on the sentence-term matrix and ignore the context dependence of the sentences. As a result, the generated summaries lack guidance from the document clusters. In this article, we propose a new language model to simultaneously cluster and summarize documents by making use of both the document-term and sentence-term matrices. By utilizing the mutual influence of document clustering and summarization, our method makes; (1) a better document clustering method with more meaningful interpretation; and (2) an effective document summarization method with guidance from document clustering. Experimental results on various document datasets show the effectiveness of our proposed method and the high interpretability of the generated summaries. Dingding Wang 0001, Shenghuo Zhu, Tao Li 0001, Yun Chi, Yihong Gong |
ACM Trans. Knowl. Discov. Data | 5 |
| 2011 | iHelp: An Intelligent Online Helpdesk SystemabstractDue to the importance of high-quality customer service, many companies use intelligent helpdesk systems (e.g., case-based systems) to improve customer service quality. However, these systems face two challenges: 1) Case retrieval measures: most case-based systems use traditional keyword-matching-based ranking schemes for case retrieval and have difficulty to capture the semantic meanings of cases and 2) result representation: most case-based systems return a list of past cases ranked by their relevance to a new request, and customers have to go through the list and examine the cases one by one to identify their desired cases. To address these challenges, we develop iHelp, an intelligent online helpdesk system, to automatically find problem-solution patterns from the past customer-representative interactions. When a new customer request arrives, iHelp searches and ranks the past cases based on their semantic relevance to the request, groups the relevant cases into different clusters using a mixture language model and symmetric matrix factorization, and summarizes each case cluster to generate recommended solutions. Case and user studies have been conducted to show the full functionality and the effectiveness of iHelp. Dingding Wang 0001, Tao Li 0001, Shenghuo Zhu, Yihong Gong |
IEEE Trans. Syst. Man Cybern. Part B | 4 |
| 2010 | A Topic Model for Linked Documents and Update Rules for its EstimationabstractThe latent topic model plays an important role in the unsupervised learning from a corpus, which provides a probabilistic interpretation of the corpus in terms of the latent topic space. An underpinning assumption which most of the topic models are based on is that the documents are assumed to be independent of each other. However, this assumption does not hold true in reality and the relations among the documents are available in different ways, such as the citation relations among the research papers. To address this limitation, in this paper we present a Bernoulli Process Topic (BPT) model, where the interdependence among the documents is modeled by a random Bernoulli process. In the BPT model a document is modeled as a distribution over topics that is a mixture of the distributions associated with the related documents. Although BPT aims at obtaining a better document modeling by incorporating the relations among the documents, it could also be applied to many applications including detecting the topics from corpora and clustering the documents. We apply the BPT model to several document collections and the experimental comparisons against several state-of-the-art approaches demonstrate the promising performance. Shenghuo Zhu, Zhongfei Zhang, Yun Chi, Yihong Gong |
AAAI | 5 |
| 2010 | Locality-constrained Linear Coding for image classificationabstractThe traditional SPM approach based on bag-of-features (BoF) requires nonlinear classifiers to achieve good image classification performance. This paper presents a simple but effective coding scheme called Locality-constrained Linear Coding (LLC) in place of the VQ coding in traditional SPM. LLC utilizes the locality constraints to project each descriptor into its local-coordinate system, and the projected coordinates are integrated by max pooling to generate the final representation. With linear classifier, the proposed approach performs remarkably better than the traditional nonlinear SPM, achieving state-of-the-art performance on several benchmarks. Compared with the sparse coding strategy [22], the objective function used by LLC has an analytical solution. In addition, the paper proposes a fast approximated LLC method by first performing a K-nearest-neighbor search and then solving a constrained least square fitting problem, bearing computational complexity of O(M + K2). Hence even with very large codebooks, our system can still process multiple frames per second. This efficiency significantly adds to the practical values of LLC for real applications. Jinjun Wang, Jianchao Yang, Kai Yu 0001, Fengjun Lv, Thomas S. Huang, Yihong Gong |
CVPR | 6 |
| 2010 | Predicting Facial Beauty without Landmarks
Douglas Gray 0001, Kai Yu 0001, Wei Xu 0007, Yihong Gong |
ECCV (6) | 4 |
| 2010 | Unsupervised Learning from Linked DocumentsabstractDocuments in many corpora, such as digital libraries and webpages, contain both content and link information. In a traditional topic model which plays an important role in the unsupervised learning, the link information is either totally ignored or treated as a feature similar to content. We believe that neither approach is capable of accurately capturing the relations represented by links. To address the limitation of traditional topic models, in this paper we propose a citation-topic (CT) model that explicitly considers the document relations represented by links. In the CT model, instead of being treated as yet another feature, links are used to form the structure of the generative model. As a result, in the CT model a given document is modeled as a mixture of a set of topic distributions, each of which is borrowed (cited) from a document that is related to the given document. We apply the CT model to several document collections and the experimental comparisons against state-of-the-art approaches demonstrate very promising performances. Shenghuo Zhu, Yun Chi, Zhongfei Zhang, Yihong Gong |
ICPR | 5 |
| 2010 | Directed Network Community Detection: A Popularity and Productivity Link ModelabstractIn this paper, we consider the problem of community detection in directed networks by using probabilistic models. Most existing probabilistic models for community detection are either symmetric in which incoming links and outgoing links are treated equally or conditional in which only one type (i.e., either incoming or outgoing) of links is modeled. We present a probabilistic model for directed network community detection that aims to model both incoming links and outgoing links simultaneously and differentially. In particular, we introduce latent variables node productivity and node popularity to explicitly capture outgoing links and incoming links, respectively. We demonstrate the generality of the proposed framework by showing that both symmetric models and conditional models for community detection can be derived from the proposed framework as special cases, leading to better understanding of the existing models. We derive efficient EM algorithms for computing the maximum likelihood solutions to the proposed models. Extensive empirical studies verify the effectiveness of the new models as well as the insights obtained from the unified framework. Tianbao Yang, Yun Chi, Shenghuo Zhu, Yihong Gong, Rong Jin 0001 |
SDM | 4 |
| 2010 | Real-time driving danger-level prediction
Jinjun Wang, Wei Xu 0007, Yihong Gong |
Eng. Appl. Artif. Intell. | 3 |
| 2010 | Incremental spectral clustering by efficiently updating the eigen-system
Huazhong Ning, Wei Xu 0007, Yun Chi, Yihong Gong, Thomas S. Huang |
Pattern Recognit. | 4 |
| 2010 | Resolution enhancement based on learning the sparse association of image patches
Jinjun Wang, Shenghuo Zhu, Yihong Gong |
Pattern Recognit. Lett. | 3 |
| 2010 | Feature Selection for Gene Expression Using Model-Based EntropyabstractGene expression data usually contain a large number of genes but a small number of samples. Feature selection for gene expression data aims at finding a set of genes that best discriminate biological samples of different types. Using machine learning techniques, traditional gene selection based on empirical mutual information suffers the data sparseness issue due to the small number of samples. To overcome the sparseness issue, we propose a model-based approach to estimate the entropy of class variables on the model, instead of on the data themselves. Here, we use multivariate normal distributions to fit the data, because multivariate normal distributions have maximum entropy among all real-valued distributions with a specified mean and standard deviation and are widely used to approximate various distributions. Given that the data follow a multivariate normal distribution, since the conditional distribution of class variables given the selected features is a normal distribution, its entropy can be computed with the log-determinant of its covariance matrix. Because of the large number of genes, the computation of all possible log-determinants is not efficient. We propose several algorithms to largely reduce the computational cost. The experiments on seven gene data sets and the comparison with other five approaches show the accuracy of the multivariate Gaussian generative model for feature selection, and the efficiency of our algorithms. Shenghuo Zhu, Dingding Wang 0001, Kai Yu 0001, Tao Li 0001, Yihong Gong |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2010 | A General Framework to Detect Unsafe System States From Multisensor Data StreamabstractThis paper proposes a general framework for detecting unsafe states of a system whose basic real-time parameters are captured by multiple sensors. Our approach is to learn a danger-level function that can be used to alert the users of dangerous situations in advance so that certain measures can be taken to avoid the collapse. The main challenge to this learning problem is the labeling issue, i.e., it is difficult to assign an objective danger level at each time step to the training data, except at the collapse points, where a definitive penalty can be assigned, and at the successful ends, where a certain reward can be assigned. In this paper, we treat the danger level as an expected future reward (a penalty is regarded as a negative reward) and use temporal difference (TD) learning to learn a function for approximating the expected future reward, given the current and historical sensor readings. The TD learning obtains the approximation by propagating the penalties/rewards observable at collapse points or successful ends to the entire feature space following some constraints. This avoids the labeling issue and naturally allows a general framework to detect unsafe states. Our approach is applied to, but not limited to, the application of monitoring driving safety, and the experimental results demonstrate the effectiveness of the approach. Huazhong Ning, Wei Xu 0007, Yihong Gong, Thomas S. Huang |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2010 | Driving Safety Monitoring Using Semisupervised Learning on Time Series DataabstractThis paper introduces a dangerous-driving warning system that uses statistical modeling to predict driving risks. The major challenge of the research is how to discover the safe/dangerous driving patterns from a sparsely labeled training data set. This paper proposes a semisupervised learning method to utilize both the labeled and the unlabeled data, as well as their interdependence to build a proper danger-level function. In addition, the learned function adopts a continuous parametric form, which is more suitable in modeling the continuous safe/dangerous-driving state transitions in a practical dangerous-driving warning system. Our comprehensive experimental evaluations reveal that, in comparison with driving danger-level estimation using classification-based methods, such as the hidden Markov model (HMM) or the conditional random field algorithm, the proposed method requires less training time and achieved higher prediction accuracy. Jinjun Wang, Shenghuo Zhu, Yihong Gong |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2010 | Human tracking using convolutional neural networksabstractIn this paper, we treat tracking as a learning problem of estimating the location and the scale of an object given its previous location, scale, as well as current and previous image frames. Given a set of examples, we train convolutional neural networks (CNNs) to perform the above estimation task. Different from other learning methods, the CNNs learn both spatial and temporal features jointly from image pairs of two adjacent frames. We introduce multiple path ways in CNN to better fuse local and global information. A creative shift-variant CNN architecture is designed so as to alleviate the drift problem when the distracting objects are similar to the target in cluttered environment. Furthermore, we employ CNNs to estimate the scale through the accurate localization of some key points. These techniques are object-independent so that the proposed method can be applied to track other types of object. The capability of the tracker of handling complex situations is demonstrated in many testing sequences. Jialue Fan, Wei Xu 0007, Ying Wu 0001, Yihong Gong |
IEEE Trans. Neural Networks | 4 |
| 2009 | Comparative document summarization via discriminative sentence selectionabstractGiven a collection of document groups, a quick question is what are the differences in these groups. In this paper, we study a novel problem of summarizing the differences between document groups. A discriminative sentence selection method is proposed to extract the most discriminative sentences which represent the specific characteristics of each document group. Experiments on real world data sets demonstrate the effectiveness of our proposed method. Dingding Wang 0001, Shenghuo Zhu, Tao Li 0001, Yihong Gong |
CIKM | 4 |
| 2009 | Resolution-Invariant Image Representation and its applicationsabstractWe present a resolution-invariant image representation (RIIR) framework in this paper. The RIIR framework includes the methods of building a set of multi-resolution bases from training images, estimating the optimal sparse resolution-invariant representation of any image, and reconstructing the missing patches of any resolution level. As the proposed RIIR framework has many potential resolution enhancement applications, we discuss three novel image magnification applications in this paper. In the first application, we apply the RIIR framework to perform Multi-Scale Image Magnification where we also introduced a training strategy to built a compact RIIR set. In the second application, the RIIR framework is extended to conduct Continuous Image Scaling where a new base at any resolution level can be generated using existing RIIR set on the fly. In the third application, we further apply the RIIR framework onto Content-Base Automatic Zooming applications. The experimental results show that in all these applications, our RIIR based method outperforms existing methods in various aspects. Jinjun Wang, Shenghuo Zhu, Yihong Gong |
CVPR | 3 |
| 2009 | Linear spatial pyramid matching using sparse coding for image classificationabstractRecently SVMs using spatial pyramid matching (SPM) kernel have been highly successful in image classification. Despite its popularity, these nonlinear SVMs have a complexity O(n2∼ n3) in training and O(n) in testing, where n is the training size, implying that it is nontrivial to scaleup the algorithms to handlemore than thousands of training images. In this paper we develop an extension of the SPM method, by generalizing vector quantization to sparse coding followed by multi-scale spatial max pooling, and propose a linear SPM kernel based on SIFT sparse codes. This new approach remarkably reduces the complexity of SVMs to O(n) in training and a constant in testing. In a number of image categorization experiments, we find that, in terms of classification accuracy, the suggested linear SPM based on sparse coding of SIFT descriptors always significantly outperforms the linear SPM kernel on histograms, and is even better than the nonlinear SPM kernels, leading to state-of-the-art performance on several benchmarks by using a single type of descriptors. Jianchao Yang, Kai Yu 0001, Yihong Gong, Thomas S. Huang |
CVPR | 3 |
| 2009 | Action detection in complex scenes with spatial and temporal ambiguitiesabstractIn this paper, we investigate the detection of semantic human actions in complex scenes. Unlike conventional action recognition in well-controlled environments, action detection in complex scenes suffers from cluttered backgrounds, heavy crowds, occluded bodies, and spatial-temporal boundary ambiguities caused by imperfect human detection and tracking. Conventional algorithms are likely to fail with such spatial-temporal ambiguities. In this work, the candidate regions of an action are treated as a bag of instances. Then a novel multiple-instance learning framework, named SMILE-SVM (Simulated annealing Multiple Instance LEarning Support Vector Machines), is presented for learning human action detector based on imprecise action locations. SMILE-SVM is extensively evaluated with satisfactory performances on two tasks: 1) human action detection on a public video action database with cluttered backgrounds, and 2) a real world problem of detecting whether the customers in a shopping mall show an intention to purchase the merchandise on shelf (even if they didn't buy it eventually). In addition, the complementary nature of motion and appearance features in action detection are also validated, demonstrating a boosted performance in our experiments. Yuxiao Hu 0001, Liangliang Cao, Fengjun Lv, Shuicheng Yan, Yihong Gong, Thomas S. Huang |
ICCV | 5 |
| 2009 | Detection driven adaptive multi-cue integration for multiple human trackingabstractIn video surveillance scenarios, appearances of both human and their nearby scenes may experience large variations due to scale and view angle changes, partial occlusions, or interactions of a crowd. These challenges may weaken the effectiveness of a dedicated target observation model even based on multiple cues, which demands for an agile framework to adjust target observation models dynamically to maintain their discriminative power. Towards this end, we propose a new adaptive way to integrate multi-cue in tracking multiple human driven by human detections. Given a human detection can be reliably associated with an existing trajectory, we adapt the way how to combine specifically devised models based on different cues in this tracker so as to enhance the discriminative power of the integrated observation model in its local neighborhood. This is achieved by solving a regression problem efficiently. Specifically, we employ 3 observation models for a single person tracker based on color models of part of torso regions, an elliptical head model, and bags of local features, respectively. Extensive experiments on 3 challenging surveillance datasets demonstrate long-term reliable tracking performance of this method. Ming Yang 0007, Fengjun Lv, Wei Xu 0007, Yihong Gong |
ICCV | 4 |
| 2009 | Knowledge Discovery from Citation NetworksabstractKnowledge discovery from scientific articles has received increasing attentions recently since huge repositories are made available by the development of the Internet and digital databases. In a corpus of scientific articles such as a digital library, documents are connected by citations and one document plays two different roles in the corpus: \emph{document itself} and \emph{a citation of other documents}. In the existing topic models, little effort is made to differentiate these two roles. We believe that the topic distributions of these two roles are different and related in a certain way. In this paper we propose a \emph{Bernoulli Process Topic}~(BPT) model which models the corpus at two levels: \emph{document level} and \emph{citation level}. In the BPT model, each document has two different representations in the latent topic space associated with its roles. Moreover, the multi-level hierarchical structure of the citation network is captured by a generative process involving a Bernoulli process. The distribution parameters of the BPT model are estimated by a variational approximation approach. In addition to conducting the experimental evaluations on the document modeling task, we also apply the BPT model to a well known scientific corpus to discover the latent topics. The comparisons against state-of-the-art methods demonstrate a very promising performance. Zhongfei Zhang, Shenghuo Zhu, Yun Chi, Yihong Gong |
ICDM | 5 |
| 2009 | Mining driving safety pattern using semi-supervised learning on time series dataabstractThis paper introduces a driving danger-level warning system that uses statistical modeling to predict driving risks. The major challenge of the research is how to model the safe/dangerous driving patterns from a sparsely labeled training data set. This paper utilizes both the labeled and the unlabeled data as well as their interdependency to build a proper danger-level function. In addition, the learned function adopts a continuous parametric form, which is more suitable in modeling the continuous safe/dangerous driving state transitions in practical dangerous driving warning system. Our comprehensive experimental evaluations reveal that, in comparison with sequential classification based methods, the proposed method requires less training time and achieved higher prediction accuracy. Yihong Gong |
ICME | 1 |
| 2009 | Normalizing multi-subject variation for drivers' emotion recognitionabstractThe paper attempts the recognition of multiple drivers' emotional state from physiological signals. The major challenge of the research is the severe inter-subject variation such that it is extreme difficult to build a general model for multiple drivers. In this paper, we focus on discovering an optimal feature mapping by utilizing the additional attribute from the drivers. Two models are reported, specifically an auxiliary dimension model and a factorization model. Experimental results show that the proposed method outperform existing algorithms used for emotional state recognition. Jinjun Wang, Yihong Gong |
ICME | 2 |
| 2009 | Resolution-Invariant Image Representation for Content-Based ZoomingabstractThis paper presents a novel Resolution-Invariant Image Representation (RIIR) framework, and applies it for Content-Based Zooming (CBZ) applications. We explain how to generate a multi-resolution bases set, from which the learned image representation can be resolution-invariant. This provides the key technology to support the continues image up-scaling task for the CBZ applications, which existing example-based resolution enhancement approaches cannot handel, or simply 2-D image interpolation algorithm cannot give satisfactory image quality for. We discuss two clustering based methods to construct the bases set. Experimental results show that, both the two methods give good image quality, and the proposed RIIR framework outperforms existing methods in various aspects. Jinjun Wang, Shenghuo Zhu, Yihong Gong |
ICME | 3 |
| 2009 | Large-scale collaborative prediction using a nonparametric random effects modelabstractA nonparametric model is introduced that allows multiple related regression tasks to take inputs from a common data space. Traditional transfer learning models can be inappropriate if the dependence among the outputs cannot be fully resolved by known input-specific and task-specific predictors. The proposed model treats such output responses as conditionally independent, given known predictors and appropriate unobserved random effects. The model is nonparametric in the sense that the dimensionality of random effects is not specified a priori but is instead determined from data. An approach to estimating the model is presented uses an EM algorithm that is efficient on a very large scale collaborative prediction problem. The obtained prediction accuracy is competitive with state-of-the-art results. Kai Yu 0001, John D. Lafferty, Shenghuo Zhu, Yihong Gong |
ICML | 4 |
| 2009 | Detecting video events based on action recognition in complex scenes using spatio-temporal descriptorabstractEvent detection plays an essential role in video content analysis and remains a challenging open problem. In particular, the study on detecting human-related video events in complex scenes with both a crowd of people and dynamic motion is still limited. In this paper, we investigate detecting video events that involve elementary human actions, e.g. making cellphone call, putting an object down, and pointing to something, in complex scenes using a novel spatio-temporal descriptor based approach. A new spatio-temporal descriptor, which temporally integrates the statistics of a set of response maps of low-level features, e.g. image gradients and optical flows, in a space-time cube, is proposed to capture the characteristics of actions in terms of their appearance and motion patterns. Based on this kind of descriptors, the bag-of-words method is utilized to describe a human figure as a concise feature vector. Then, these features are employed to train SVM classifiers at multiple spatial pyramid levels to distinguish different actions. Finally, a Gaussian kernel based temporal filtering is conducted to segment the sequences of events from a video stream taking account of the temporal consistency of actions. The proposed approach is capable of tolerating spatial layout variations and local deformations of human actions due to diverse view angles and rough human figure alignment in complex scenes. Extensive experiments on the 50-hour video dataset of TRECVid 2008 event detection task demonstrate that our approach outperforms the well-known SIFT descriptor based methods and effectively detects video events in challenging real-world conditions. Guangyu Zhu 0002, Ming Yang 0007, Kai Yu 0001, Wei Xu 0007, Yihong Gong |
ACM Multimedia | 5 |
| 2009 | Nonlinear Learning using Local Coordinate CodingabstractThis paper introduces a new method for semi-supervised learning on high dimensional nonlinear manifolds, which includes a phase of unsupervised basis learning and a phase of supervised function learning. The learned bases provide a set of anchor points to form a local coordinate system, such that each data point x on the manifold can be locally approximated by a linear combination of its nearby anchor points, and the linear weights become its local coordinate coding. We show that a high dimensional nonlinear function can be approximated by a global linear function with respect to this coding scheme, and the approximation quality is ensured by the locality of such coding. The method turns a difficult nonlinear learning problem into a simple global linear learning problem, which overcomes some drawbacks of traditional local learning methods. Kai Yu 0001, Tong Zhang 0001, Yihong Gong |
NIPS | 3 |
| 2009 | A Bayesian Approach Toward Finding Communities and Their Evolutions in Dynamic Social NetworksabstractAlthough a large body of work are devoted to finding communities in static social networks, only a few studies examined the dynamics of communities in evolving social networks. In this paper, we propose a dynamic stochastic block model for finding communities and their evolutions in a dynamic social network. The proposed model captures the evolution of communities by explicitly modeling the transition of community memberships for individual nodes in the network. Unlike many existing approaches for modeling social networks that estimate parameters by their most likely values (i.e., point estimation), in this study, we employ a Bayesian treatment for parameter estimation that computes the posterior distributions for all the unknown parameters. This Bayesian treatment allows us to capture the uncertainty in parameter values and therefore is more robust to data noise than point estimation. In addition, an efficient algorithm is developed for Bayesian inference to handle large sparse social networks. Extensive experimental studies based on both synthetic data and real-life data demonstrate that our model achieves higher accuracy and reveals more insights in the data than several state-of-the-art algorithms. Tianbao Yang, Yun Chi, Shenghuo Zhu, Yihong Gong, Rong Jin 0001 |
SDM | 4 |
| 2009 | A latent topic model for linked documentsabstractDocuments in many corpora, such as digital libraries and webpages, contain both content and link information. To explicitly consider the document relations represented by links, in this paper we propose a citation-topic (CT) model which assumes a probabilistic generative process for corpora. In the CT model a given document is modeled as a mixture of a set of topic distributions, each of which is borrowed (cited) from a document that is related to the given document. Moreover, the CT model contains a random process for selecting the related documents according to the structure of the generative model determined by links and therefore, the transitivity of the relations among documents is captured. We apply the CT model on the document clustering task and the experimental comparisons against several state-of-the-art approaches demonstrate very promising performances. Shenghuo Zhu, Yun Chi, Zhongfei Zhang, Yihong Gong |
SIGIR | 5 |
| 2009 | Fast nonparametric matrix factorization for large-scale collaborative filteringabstractWith the sheer growth of online user data, it becomes challenging to develop preference learning algorithms that are sufficiently flexible in modeling but also affordable in computation. In this paper we develop nonparametric matrix factorization methods by allowing the latent factors of two low-rank matrix factorization methods, the singular value decomposition (SVD) and probabilistic principal component analysis (pPCA), to be data-driven, with the dimensionality increasing with data size. We show that the formulations of the two nonparametric models are very similar, and their optimizations share similar procedures. Compared to traditional parametric low-rank methods, nonparametric models are appealing for their flexibility in modeling complex data dependencies. However, this modeling advantage comes at a computational price--it is highly challenging to scale them to large-scale problems, hampering their application to applications such as collaborative filtering. In this paper we introduce novel optimization algorithms, which are simple to implement, which allow learning both nonparametric matrix factorization models to be highly efficient on large-scale problems. Our experiments on EachMovie and Netflix, the two largest public benchmarks to date, demonstrate that the nonparametric models make more accurate predictions of user ratings, and are computationally comparable or sometimes even faster in training, in comparison with previous state-of-the-art parametric matrix factorization models. Kai Yu 0001, Shenghuo Zhu, John D. Lafferty, Yihong Gong |
SIGIR | 4 |
| 2009 | Visual-Quality Optimizing Super ResolutionabstractAbstract In this paper, we propose a robust image super‐resolution (SR) algorithm that aims to maximize the overall visual quality of SR results. We consider a good SR algorithm to be fidelity preserving, image detail enhancing and smooth. Accordingly, we define perception‐based measures for these visual qualities. Based on these quality measures, we formulate image SR as an optimization problem aiming to maximize the overall quality. Since the quality measures are quadratic, the optimization can be solved efficiently. Experiments on a large image set and subjective user study demonstrate the effectiveness of the perception‐based quality measures and the robustness and efficiency of the presented method. Feng Liu 0015, Jinjun Wang, Shenghuo Zhu, Michael Gleicher, Yihong Gong |
Comput. Graph. Forum | 5 |
| 2009 | SoftCuts: A Soft Edge Smoothness Prior for Color Image Super-ResolutionabstractDesigning effective image priors is of great interest to image super-resolution (SR), which is a severely under-determined problem. An edge smoothness prior is favored since it is able to suppress the jagged edge artifact effectively. However, for soft image edges with gradual intensity transitions, it is generally difficult to obtain analytical forms for evaluating their smoothness. This paper characterizes soft edge smoothness based on a novel SoftCuts metric by generalizing the Geocuts method . The proposed soft edge smoothness measure can approximate the average length of all level lines in an intensity image. Thus, the total length of all level lines can be minimized effectively by integrating this new form of prior. In addition, this paper presents a novel combination of this soft edge smoothness prior and the alpha matting technique for color image SR, by adaptively normalizing image edges according to their alpha-channel description. This leads to the adaptive SoftCuts algorithm, which represents a unified treatment of edges with different contrasts and scales. Experimental results are presented which demonstrate the effectiveness of the proposed method. Shengyang Dai, Wei Xu 0007, Ying Wu 0001, Yihong Gong, Aggelos K. Katsaggelos |
IEEE Trans. Image Process. | 5 |
| 2009 | iOLAP: A Framework for Analyzing the Internet, Social Networks, and Other Networked DataabstractAs the amount of noisy, unorganized, linked data on the Internet increases dramatically, how to efficiently analyze such data becomes a challenging research problem. In this paper, we propose a framework, iOLAP, that offers functionalities for analyzing networked data from Internet, social networks, scientific paper citations, etc. We first identify four main data dimensions that are common in most of networked data, namely people, relation, content, and time. Motivated by the fact that various dimensions of data jointly affect each other, we propose a polyadic factorization approach to directly model all the dimensions simultaneously in a unified framework. We provide detailed theoretical analysis of the new modeling framework. In addition to the theoretical framework, we also present an efficient implementation of the algorithm that takes advantage of the sparseness of data and has time complexity linear in the number of data records in a dataset. We then apply the proposed models to analyzing the blogosphere and personalizing recommendation in paper citations. Extensive experimental studies showed that our framework is able to provide deep insights jointed obtained from various dimensions of networked data. Yun Chi, Shenghuo Zhu, Koji Hino, Yihong Gong, Yi Zhang 0001 |
IEEE Trans. Multim. | 4 |
| 2008 | Probabilistic polyadic factorization and its application to personalized recommendationabstractMultiple-dimensional, i.e., polyadic, data exist in many applications, such as personalized recommendation and multiple-dimensional data summarization. Analyzing all the dimensions of polyadic data in a principled way is a challenging research problem. Most existing methods separately analyze the marginal relationships among pairwise dimensions and then combine the results afterwards. Motivated by the fact that various dimensions of polyadic data jointly affect each other, we propose a probabilistic polyadic factorization approach to directly model all the dimensions simultaneously in a unified framework. We then show the connection between the probabilistic polyadic factorization and a non-negative version of the Tucker tensor factorization. We provide detailed theoretical analysis of the new modeling framework, discuss implementation techniques for our models, and propose several extensions to the basic framework. We then apply the proposed models to the application of personalized recommendation. Extensive experiments on a social bookmarking dataset, Delicious, and a paper citation dataset, CiteSeer, demonstrate the effectiveness of the proposed models. Yun Chi, Shenghuo Zhu, Yihong Gong, Yi Zhang 0001 |
CIKM | 3 |
| 2008 | Integrating clustering and multi-document summarization to improve document understandingabstractDocument understanding techniques such as document clustering and multi-document summarization have been receiving much attention in recent years. Current document clustering methods usually represent documents as a term-document matrix and perform clustering algorithms on it. Although these clustering methods can group the documents satisfactorily, it is still hard for people to capture the meanings of the documents since there is no satisfactory interpretation for each document cluster. In this paper, we propose a new language model to simultaneously cluster and summarize the documents. By utilizing the mutual influence of the document clustering and summarization, our method makes (1) a better document clustering method with more meaningful interpretation and (2) a better document summarization method taking the document context information into consideration. Dingding Wang 0001, Shenghuo Zhu, Tao Li 0001, Yun Chi, Yihong Gong |
CIKM | 5 |
| 2008 | Discriminative learning of visual words for 3D human pose estimationabstractThis paper addresses the problem of recovering 3D human pose from a single monocular image, using a discriminative bag-of-words approach. In previous work, the visual words are learned by unsupervised clustering algorithms. They capture the most common patterns and are good features for coarse-grain recognition tasks like object classification. But for those tasks which deal with subtle differences such as pose estimation, such representation may lack the needed discriminative power. In this paper, we propose to jointly learn the visual words and the pose regressors in a supervised manner. More specifically, we learn an individual distance metric for each visual word to optimize the pose estimation performance. The learned metrics rescale the visual words to suppress unimportant dimensions such as those corresponding to background. Another contribution is that we design an Appearance and Position Context (APC) local descriptor that achieves both selectivity and invariance while requiring no background subtraction. We test our approach on both a quasi-synthetic dataset and a real dataset (HumanEva) to verify its effectiveness. Our approach also achieves fast computational speed thanks to the integral histograms used in APC descriptor extraction and fast inference of pose regressors. Huazhong Ning, Wei Xu 0007, Yihong Gong, Thomas S. Huang |
CVPR | 3 |
| 2008 | Training Hierarchical Feed-Forward Visual Recognition Models Using Transfer Learning from Pseudo-Tasks
Amr Ahmed 0001, Kai Yu 0001, Wei Xu 0007, Yihong Gong, Eric P. Xing |
ECCV (3) | 4 |
| 2008 | Latent Pose Estimator for Continuous Action Recognition
Huazhong Ning, Wei Xu 0007, Yihong Gong, Thomas S. Huang |
ECCV (2) | 3 |
| 2008 | Fast image super-resolution using Connected Component enhancementabstractThe paper focuses on reconstructing the discontinuity between homogenous color regions in an interpolated image to improve its perceptual quality. A low-resolution input image is firstly interpolated and then decomposed into several patches. Each patch is then segmented into multiple homogenous regions using connected component analysis technique. Then a spatial-filter is applied to enhance the color/intensity transition between neighboring components. The designed spatial-filter combines the advantages of both bilateral-filtering and unsharp masking methods, with high computational efficiency. The proposed method can be used for image/video super-resolution applications. Experimental results are promising. Jinjun Wang, Yihong Gong |
ICME | 2 |
| 2008 | Temporal difference learning to detect unsafe system statesabstractThis paper proposes a general framework to detect unsafe states of a system whose basic realtime parameters are captured by multi-sensors. Our approach is to learn a danger level function which can be used to alert the users in advance of dangerous situations. The main challenge to this learning problem is the labelling issue, i.e., it is difficult to assign an objective danger level at each time step to the training data, except at the collapse points where a penalty can be assigned and at the successful ends where a certain reward can be assigned. In this paper, we treat the danger level as expected future reward (penalty is regarded as negative reward) and use temporal difference (TD) learning [2] to learn a function to approximate the expected future reward. The TD learning obtains the approximation by propagating the penalty/reward observable at collapse points or successful ends to the entire feature space following some constraints. Our approach is applied to, but not limited to, the application of monitoring of driving safety and the experimental results demonstrate the effectiveness of the approach. Huazhong Ning, Wei Xu 0007, Yihong Gong, Thomas S. Huang |
ICPR | 4 |
| 2008 | Recognition of multiple drivers' emotional stateabstractThe paper attempted the recognition of multiple driverspsila emotional state from physiological signals. The major challenge of the research is due to the severe inter-driver variation such that the features of different emotional state are high correlated, and it is found that simple decorrelation method cannot normalize the features well to achieve acceptable classification accuracy. Hence, in this paper, we propose to apply a latent variable to represent the hidden attribute of individual driver and use statistical training. In addition, we applied temporal constraints for the inference process to improve the recognition accuracy. Experimental results show that the proposed method outperform existing algorithms used for emotional state recognition. Jinjun Wang, Yihong Gong |
ICPR | 2 |
| 2008 | Noisy video super-resolutionabstractLow-quality videos often not only have limited resolution, but also suffer from noise. Directly up-sampling a video without considering noise could deteriorate its visual quality due to magnifying noise. This paper addresses this problem with a unified framework that achieves simultaneous de-noising and super-resolution. This framework formulates noisy video super-resolution as an optimization problem, aiming to maximize the visual quality of the result. We consider a good quality result to be fidelity-preserving, detailpreserving and smooth. Accordingly, we propose measures for these qualities in the scenario of de-noising and superresolution. The experiments on a variety of noisy videos demonstrate the effectiveness of the presented algorithm. Feng Liu 0015, Jinjun Wang, Shenghuo Zhu, Michael Gleicher, Yihong Gong |
ACM Multimedia | 5 |
| 2008 | Deep Learning with Kernel Regularization for Visual RecognitionabstractIn this paper we focus on training deep neural networks for visual recognition tasks. One challenge is the lack of an informative regularization on the network parameters, to imply a meaningful control on the computed function. We propose a training strategy that takes advantage of kernel methods, where an existing kernel function represents useful prior knowledge about the learning task of interest. We derive an efficient algorithm using stochastic gradient descent, and demonstrate very positive results in a wide range of visual recognition tasks. Kai Yu 0001, Wei Xu 0007, Yihong Gong |
NIPS | 3 |
| 2008 | Stochastic Relational Models for Large-scale Dyadic Data using MCMCabstractStochastic relational models provide a rich family of choices for learning and predicting dyadic data between two sets of entities. It generalizes matrix factorization to a supervised learning problem that utilizes attributes of objects in a hierarchical Bayesian framework. Previously empirical Bayesian inference was applied, which is however not scalable when the size of either object sets becomes tens of thousands. In this paper, we introduce a Markov chain Monte Carlo (MCMC) algorithm to scale the model to very large-scale dyadic data. Both superior scalability and predictive accuracy are demonstrated on a collaborative filtering problem, which involves tens of thousands users and a half million items. Shenghuo Zhu, Kai Yu 0001, Yihong Gong |
NIPS | 3 |
| 2008 | Non-greedy active learning for text categorization using convex transductive experimental designabstractIn this paper we propose a non-greedy active learning method for text categorization using least-squares support vector machines (LSSVM). Our work is based on transductive experimental design (TED), an active learning formulation that effectively explores the information of unlabeled data. Despite its appealing properties, the optimization problem is however NP-hard and thus--like most of other active learning methods--a greedy sequential strategy to select one data example after another was suggested to find a suboptimum. In this paper we formulate the problem into a continuous optimization problem and prove its convexity, meaning that a set of data examples can be selected with a guarantee of global optimum. We also develop an iterative algorithm to efficiently solve the optimization problem, which turns out to be very easy-to-implement. Our text categorization experiments on two text corpora empirically demonstrated that the new active learning algorithm outperforms the sequential greedy algorithm, and is promising for active text categorization applications. Kai Yu 0001, Shenghuo Zhu, Wei Xu 0007, Yihong Gong |
SIGIR | 4 |
| 2008 | Dynamic active probing of helpdesk databasesabstractHelpdesk databases are used to store past interactions between customers and companies to improve customer service quality. One common scenario of using helpdesk database is to find whether recommendations exist given a new problem from a customer. However, customers often provide incomplete or even inaccurate information. Manually preparing a list of clarification questions does not work for large databases. This paper investigates the problem of automatic generation of a minimal number of questions to reach an appropriate recommendation. This paper proposes a novel dynamic active probing method. Compared to other alternatives such as decision tree and case-based reasoning, this method has two distinctive features. First, it actively probe the customer to get useful information to reach the recommendation, and the information provided by customer will be immediately used by the method to dynamically generate the next questions to probe. This feature ensures that all available information from the customer is used. Second, this method is based on a probabilistic model, and uses a data augmentation method which avoids overfitting when estimating the probabilities in the model. This feature ensures that the method is robust to databases that are incomplete or contain errors. Experimental results verify the effectiveness of our approach. Shenghuo Zhu, Tao Li 0001, Zhiyuan Chen 0003, Dingding Wang 0001, Yihong Gong |
Proc. VLDB Endow. | 5 |
| 2007 | Soft Edge Smoothness Prior for Alpha Channel Super ResolutionabstractEffective image prior is necessary for image super resolution, due to its severely under-determined nature. Although the edge smoothness prior can be effective, it is generally difficult to have analytical forms to evaluate the edge smoothness, especially for soft edges that exhibit gradual intensity transitions. This paper finds the connection between the soft edge smoothness and a soft cut metric on an image grid by generalizing the Geocuts method (Y. Boykov and V. Kolmogorov, 2003), and proves that the soft edge smoothness measure approximates the average length of all level lines in an intensity image. This new finding not only leads to an analytical characterization of the soft edge smoothness prior, but also gives an intuitive geometric explanation. Regularizing the super resolution problem by this new form of prior can simultaneously minimize the length of all level lines, and thus resulting in visually appealing results. In addition, this paper presents a novel combination of this soft edge smoothness prior and the alpha matting technique for color image super resolution, by normalizing edge segments with their alpha channel description, to achieve a unified treatment of edges with different contrast and scale. Shengyang Dai, Wei Xu 0007, Ying Wu 0001, Yihong Gong |
CVPR | 5 |
| 2007 | Bilateral Back-Projection for Single Image Super ResolutionabstractIn this paper, a novel algorithm for single image super resolution is proposed. Back-projection [1] can minimize the reconstruction error with an efficient iterative procedure. Although it can produce visually appealing result, this method suffers from the chessboard effect and ringing effect, especially along strong edges. The underlining reason is that there is no edge guidance in the error correction process. Bilateral filtering can achieve edge-preserving image smoothing by adding the extra information from the feature domain. The basic idea is to do the smoothing on the pixels which are nearby both in space domain and in feature domain. The proposed bilateral back-projection algorithm strives to integrate the bilateral filtering into the back-projection method. In our approach, the back-projection process can be guided by the edge information to avoid across-edge smoothing, thus the chessboard effect and ringing effect along image edges are removed. Promising results can be obtained by the proposed bilateral back-projection method efficiently. Shengyang Dai, Ying Wu 0001, Yihong Gong |
ICME | 4 |
| 2007 | Efficient Video Object Segmentation by Graph-CutabstractSegmentation of video objects from background is a popular computer vision topic and has many important applications. Most existing methods are either computationally expensive or requiring manual initialization, static cameras, and/or rigid scenes. In a previous work, we proposed a joint spatio-temporal linear regression algorithm to automatically cluster the sparse edge/corner pixels in each video frame and obtain two motion models for the object and background respectively. To label the rest pixels for object segmentation, in this paper, we propose to model the Optical-Flow residual error, color intensity residual error and temporal label consistency features, as well as color/edge orientation consistency constrains, in a graph, and apply the Graph-Cut algorithm to minimize the energy of the graph to obtain an optimal segmentation of the two motion layers boundaries. Finally the object layer is identified from the two using simple heuristics. Experimental segmentation result with videos taken by webcams is promising. Jinjun Wang, Wei Xu 0007, Shenghuo Zhu, Yihong Gong |
ICME | 4 |
| 2007 | Detecting Unsafe Driving Patterns using Discriminative LearningabstractWe propose a discriminative learning approach for fusing multichannel sequential data with application to detect unsafe driving patterns from multi-channel driving recording data. The fusion is performed using a discriminatively trained graphical model -conditional random field (CRF). The proposed approach offers several advantage over existing information fusing approaches. First, it derives its classification power by directly modelling and maximizing the conditional probability. Second, it represents the variable dependency in an undirected graph, which is very efficient in inference. Third, it does not require to label all the training data and utilizes both labelled and unlabelled data efficiently by semi-supervised learning algorithms. The proposed approach is evaluated on driving recording data collected from driving simulator -STISIM. Experiments show it outperforms the simple discriminative classifier (SVM) and generative model (HMM). Wei Xu 0007, Huazhong Ning, Yihong Gong, Thomas S. Huang |
ICME | 4 |
| 2007 | Predictive Matrix-Variate t ModelsabstractIt is becoming increasingly important to learn from a partially-observed random matrix and predict its missing elements. We assume that the entire matrix is a single sample drawn from a matrix-variate t distribution and suggest a matrix-variate t model (MVTM) to predict those missing elements. We show that MVTM generalizes a range of known probabilistic models, and automatically performs model selection to encourage sparse predictive models. Due to the non-conjugacy of its prior, it is difficult to make predictions by computing the mode or mean of the posterior distribution. We suggest an optimization method that sequentially minimizes a convex upper-bound of the log-likelihood, which is very efficient and scalable. The experiments on a toy data and EachMovie dataset show a good predictive accuracy of the model. Shenghuo Zhu, Kai Yu 0001, Yihong Gong |
NIPS | 3 |
| 2007 | Incremental Spectral Clustering With Application to Monitoring of Evolving Blog CommunitiesabstractIn recent years, spectral clustering method has gained attentions because of its superior performance compared to other traditional clustering algorithms such as K-means algorithm. The existing spectral clustering algorithms are all off-line algorithms, i.e., they can not incrementally update the clustering result given a small change of the data set. However, the capability of incrementally updating is essential to some applications such as real time monitoring of the evolving communities of websphere or blogsphere. Unlike traditional stream data, these applications require incremental algorithms to handle not only insertion/deletion of data points but also similarity changes between existing items. This paper extends the standard spectral clustering to such evolving data by introducing the incidence vector/matrix to represent two kinds of dynamics in the same framework and by incrementally updating the eigenvalue system. Our incremental algorithm, initialized by a standard spectral clustering, continuously and efficiently updates the eigenvalue system and generates instant cluster labels, as the data set is evolving. The algorithm is applied to a blog data set. Compared with recomputation of the solution by standard spectral clustering, it achieves similar accuracy but with much lower computational cost. Close inspection into the blog content shows that the incremental approach can discover not only the stable blog communities but also the evolution of the individual multi-topic blogs. Huazhong Ning, Wei Xu 0007, Yun Chi, Yihong Gong, Thomas S. Huang |
SDM | 4 |
| 2007 | Combining content and link for classification using matrix factorizationabstractThe world wide web contains rich textual contents that areinterconnected via complex hyperlinks. This huge database violates the assumption held by most of conventional statistical methods that each web page is considered as an independent and identical sample. It is thus difficult to apply traditional mining or learning methods for solving web mining problems, e.g., web page classification, by exploiting both the content and the link structure. The research in this direction has recently received considerable attention but are still in an early stage. Though a few methods exploit both the link structure or the content information, some of them combine the only authority information with the content information, and the others first decompose the link structure into hub and authority features, then apply them as additional document features. Being practically attractive for its great simplicity, this paper aims to design an algorithm that exploits both the content and linkage information, by carrying out a joint factorization on both the linkage adjacency matrix and the document-term matrix, and derives a new representation for web pages in a low-dimensional factor space, without explicitly separating them as content, hub or authority factors. Further analysis can be performed based on the compact representation of web pages. In the experiments, the proposed method is compared with state-of-the-art methods and demonstrates an excellent accuracy in hypertext classification on the WebKB and Cora benchmarks. Shenghuo Zhu, Kai Yu 0001, Yun Chi, Yihong Gong |
SIGIR | 4 |
| 2007 | Multi-object trajectory tracking
Wei Xu 0007, Yihong Gong |
Mach. Vis. Appl. | 4 |
| 2006 | Video Super-resolution with Scene-specific PriorsabstractIn this paper, we propose a method to improve the spatial resolution of video sequences. Our approach is inspired by previous image hallucination work [12]. There are two main contributions of the proposed method. First, the information from cameras with different spatial-temporal resolutions is combined in our framework. This is achieved by constructing training dictionary using the high resolution images captured by still camera and the low resolution video is enhanced via searching in this scene-specific database. Since the dictionary is customized to a particular scene instead of built from arbitrary images, it has fewer but more representative samples. Second, we enforce the spatio-temporal constraints using the conditional random field (CRF) and the problem of video super-resolution is posed as finding the high resolution video that maximizes the conditional probability. We apply the algorithm to video sequences taken from different scenes and the results demonstrate that our approach can synthesize high quality super-resolution videos. 1 Dan Kong, Wei Xu 0007, Yihong Gong |
BMVC | 5 |
| 2006 | Improving Speaker Diarization by Cross EM RefinementabstractIn this paper, we present a new speaker diarization system that improves the accuracy of traditional hierarchical clustering-based methods with little increase in computational cost. Our contributions are mainly two fold. First, we include a preprocessing called "local clustering" before the hierarchical clustering algorithm to merge very similar adjacent speech segments. This local clustering aims to reduce the number of segments to be clustered by the hierarchical clustering, so as to dramatically increase the processing speed. Second, we perform a postprocessing called "cross EM refinement" to purify the clusters generated by the hierarchical clustering. This algorithm is based on the idea of cross validation and EM algorithm. Our experimental evaluations show that the proposed cross EM refinement approach reduces the speaker diarization error by up to 56%, with an average reduction of 22% compared to the traditional hierarchical clustering method. Huazhong Ning, Wei Xu 0007, Yihong Gong, Thomas S. Huang |
ICME | 3 |
| 2006 | Trend Analysis for Large Document StreamsabstractMore and more powerful computer technology inspires people to investigate information hidden under huge amounts of documents. In this report, we are especially interested in documents with relative time order, which we also call document streams. Examples include TV news, forums, emails of company projects, call center telephone logs, etc. To get an insight into these document streams, first we need to detect the events among the document streams. We use a time-sensitive Dirichlet process mixture model to find the events in the document streams. A time sensitive Dirichlet process mixture model is a generative model, which allows a potentially infinite number of mixture components and uses a Dirichlet compound multinomial model to model the distribution of words in documents. In this report, we consider three different time sensitive Dirichlet process mixture models: an exponential decay kernel model, a polynomial decay function kernel Dirichlet process model and a sliding window kernel model. Experiments on the TDT2 dataset have shown that the time sensitive models perform 18-20% better in terms of accuracy than the Dirichlet process mixture model. The sliding windows kernel and the polynomial kernel are more promising in detecting events. We use ThemeRiver to provide a visualization of the events along the time axis. With the help of ThemeRiver, people can easily get an overall picture of how different events evolve. Besides ThemeRiver, we investigate using top words as a high-level summarization of each event. Experiment results on TDT2 dataset suggests that the sliding window kernel is a better choice both in terms of capturing the trend of the events and expressibility Chengliang Zhang, Shenghuo Zhu, Yihong Gong |
ICMLA | 3 |
| 2006 | Video object segmentation by motion-based sequential feature clusteringabstractSegmentation of video foreground objects from background has many important applications, such as human computer interaction, video compression, multimedia content editing and manipulation. Most existing methods work on image pixels or color segments which are computationally expensive. Some methods require extensive manual inputs, static cameras, and/or rigid scenes. In this paper we propose a fully automatic foreground segmentation method based on sequential clustering of sparse image features. The sparseness makes the method computationally efficient. We use both edge and corner points extracted from each video frame. A joint spatio-temporal linear regression method is developed to compute sparse motion layers of M consecutive frames jointly under the temporal consistency constraint. Once the sparse motion layers have been identified for each frame, the corresponding dense motion layers are created using the Markov Random Field (MRF) model. The MRF model assigns the rest of the image pixels to the motion layers by considering both the color attributes and the spatial relations between each pixel and its surrounding edge/corner points. Experimental evaluations on videos taken by webcams show the effectiveness of the proposed method. Wei Xu 0007, Yihong Gong |
ACM Multimedia | 3 |
| 2005 | Multi-labelled classification using maximum entropy methodabstractMany classification problems require classifiers to assign each single document into more than one category, which is called multi-labelled classification. The categories in such problems usually are neither conditionally independent from each other nor mutually exclusive, therefore it is not trivial to directly employ state-of-the-art classification algorithms without losing information of relation among categories. In this paper, we explore correlations among categories with maximum entropy method and derive a classification algorithm for multi-labelled documents. Our experiments show that this method significantly outperforms the combination of single label approach. Shenghuo Zhu, Wei Xu 0007, Yihong Gong |
SIGIR | 4 |
| 2004 | An Algorithm for Multiple Object Trajectory Tracking
Wei Xu 0007, Yihong Gong |
CVPR (1) | 4 |
| 2004 | A detection-based multiple object tracking methodabstractIn this paper we describe a method for tracking multiple objects whose number is unknown and varies during tracking. Based on preliminary results of object detection in each image which may have missing and/or false detection, the multiple object tracking method keeps a graph structure where it maintains multiple hypotheses about the number and the trajectories of the objects in the video. The image information drives the process of extending and pruning the graph, and determines the best hypothesis to explain the video. While the image-based object detection makes a local decision, the tracking process confirms and validates the detection through time, therefore, it can be regarded as temporal detection which makes a global decision across time. The multiple object tracking method gives feedbacks which are predictions of object locations to the object detection module. Therefore, the method integrates object detection and tracking tightly. The most possible hypothesis provides the multiple object tracking result. The experimental results are presented. Amit Sethi, Yihong Gong |
ICIP | 4 |
| 2004 | Document clustering by concept factorizationabstractIn this paper, we propose a new data clustering method called concept factorization that models each concept as a linear combination of the data points, and each data point as a linear combination of the concepts. With this model, the data clustering task is accomplished by computing the two sets of linear coefficients, and this linear coefficients computation is carried out by finding the non-negative solution that minimizes the reconstruction error of the data points. The cluster label of each data point can be easily derived from the obtained linear coefficients. This method differs from the method of clustering based on non-negative matrix factorization (NMF) \citeXu03 in that it can be applied to data containing negative values and the method can be implemented in the kernel space. Our experimental results show that the proposed data clustering method and its variations performs best among 11 algorithms and their variations that we have evaluated on both TDT2 and Reuters-21578 corpus. In addition to its good performance, the new method also has the merit in its easy and reliable derivation of the clustering results. Wei Xu 0007, Yihong Gong |
SIGIR | 2 |
| 2004 | Maximum entropy model-based baseball highlight detection and classification
Yihong Gong, Wei Xu 0007 |
Comput. Vis. Image Underst. | 1 |
| 2003 | A New Tracking Technique: Object Tracking and Identification from Motion
Terrence Chen, Yihong Gong, Thomas S. Huang |
CAIP | 4 |
| 2003 | Document clustering based on non-negative matrix factorizationabstractIn this paper, we propose a novel document clustering method based on the non-negative factorization of the term-document matrix of the given document corpus. In the latent semantic space derived by the non-negative matrix factorization (NMF), each axis captures the base topic of a particular document cluster, and each document is represented as an additive combination of the base topics. The cluster membership of each document can be easily determined by finding the base topic (the axis) with which the document has the largest projection value. Our experimental evaluations show that the proposed document clustering method surpasses the latent semantic indexing and the spectral clustering methods not only in the easy and reliable derivation of document clustering results, but also in document clustering accuracies. Wei Xu 0007, Xin Liu 0046, Yihong Gong |
SIGIR | 3 |
| 2003 | Video summarization and retrieval using singular value decomposition
Yihong Gong, Xin Liu 0046 |
Multim. Syst. | 1 |
| 2002 | Extract highlights from baseball game video with hidden Markov modelsabstractWe describe a statistical method to detect highlights in a baseball game video. The input video is first segmented into scene shots, within which the camera motion is continuous. Our approach is based on the observations that (1) most highlights in baseball games are composed of certain types of scene shots and (2) those scene shots exhibit special transition context in time. To exploit those two observations, we first build statistical models for each type of scene shots with products of histograms, and then for each type of highlight a hidden Markov model is learned to represent the context of transition in the time domain. A probabilistic model can be obtained by combining the two, which is used for highlight detection and classification. Satisfactory results have been achieved on initial experimental results. Peng Chang 0002, Yihong Gong |
ICIP (1) | 3 |
| 2002 | Creating motion video summaries with partial audio-visual alignmentabstractIn this paper, we propose an audio-visual summarization system which creates an audio and a visual summary of a given video program separately, and then integrates the two summaries with a partial alignment. A bipartite graph-based audio-visual alignment algorithm is developed to efficiently find the best alignment solution that satisfies the predefined alignment requirements. With the proposed system, we strive to produce a motion video summary for the original video that: (1) provides a natural visual and audio content overview; and (2) maximizes the coverage for both audio and visual contents of the original video without having to sacrifice either of them. Such audio-visual summaries dramatically increase the information intensity and depth, and lead to a more effective video content overview. Yihong Gong, Xin Liu 0046 |
ICME (1) | 1 |
| 2002 | Baseball scene classification using multimedia featuresabstractIn this paper, we address the issue of classifying video scenes which is essential in video indexing, archiving and summarization. Compared with previous methods, we emphasize the integration of multimedia features, including image, audio and speech cues. With current state-of-the-art image and audio analysis techniques, most image and audio features we can extract from videos are very low level, therefore, classifying scenes based on features from a single medium yields poor performance. We propose a maximum entropy based method for baseball scene classification in TV broadcast videos. The maximum entropy scheme is chosen because it can automatically select and fuse multimedia features from temporal contexts. Yihong Gong |
ICME (1) | 3 |
| 2002 | An integrated baseball digest system using maximum entropy methodabstractIn this paper, we propose a novel system that is able to automatically detect and classify highlights from baseball game videos in TV broadcast. The digest system gives complete indexes of a baseball game which cover all of the status changes in a game. We achieve this by seamlessly integrating image, audio and speech clues using a maximum entropy based method. What distinguishes our system from previous ones is that we emphasize on the integration of multimedia features and the acquisition of domain knowledge through machine learning process. Integration of multimedia features is important because with the current state-of-the-art image and audio analysis techniques, most image and audio features we can extract from videos are very low level, and detecting/classifying sports game highlights based on features from single medium are doomed to yield poor performances. Acquiring domain knowledge through learning process is preferred over heuristic rules because machine learning process is more powerful for discovering and expressing domain knowledge. We perform extensive experiments on game videos including various stadiums, teams and broadcasted by different TV stations. Wei Xu 0007, Yihong Gong |
ACM Multimedia | 4 |
| 2002 | Document clustering with cluster refinement and model selection capabilitiesabstractIn this paper, we propose a document clustering method that strives to achieve: (1) a high accuracy of document clustering, and (2) the capability of estimating the number of clusters in the document corpus (i.e. the model selection capability). To accurately cluster the given document corpus, we employ a richer feature set to represent each document, and use the Gaussian Mixture Model (GMM) together with the Expectation-Maximization (EM) algorithm to conduct an initial document clustering. From this initial result, we identify a set of discriminative featuresfor each cluster, and refine the initially obtained document clusters by voting on the cluster label of each document using this discriminative feature set. This self-refinement process of discriminative feature identification and cluster label voting is iteratively applied until the convergence of document clusters. On the other hand, the model selection capability is achieved by introducing randomness in the cluster initialization stage, and then discovering a value C for the number of clusters N by which running the document clustering process for a fixed number of times yields sufficiently similar results. Performance evaluations exhibit clear superiority of the proposed method with its improved document clustering and model selection accuracies. The evaluations also demonstrate how each feature as well as the cluster refinement process contribute to the document clustering accuracy. Xin Liu 0046, Yihong Gong, Wei Xu 0007, Shenghuo Zhu |
SIGIR | 2 |
| 2001 | Creating Generic Text SummariesabstractWe propose two generic text summarization methods that create text summaries by ranking and extracting sentences from the original documents. The first method uses standard information retrieval methods to rank sentence relevances, while the second method uses the latent semantic analysis technique to identify semantically important sentences, for summary creations. Both methods strive to select sentences that are highly ranked and different from each other. This is an attempt to create a summary with a wider coverage of the document's main content and less redundancy. Performance evaluations on the two summarization methods are conducted by comparing their summarization outputs with the manual summaries generated by three independent human evaluators. Yihong Gong, Xin Liu 0046 |
ICDAR | 1 |
| 2001 | Video summarization with minimal visual content redundanciesabstractWe propose a video summarization method able to produce a motion video summary that minimizes the visual content redundancy for the input video. We derive the visual content redundancy metric that tells how much redundancy a video contains, and how much the video can be curtailed without losing too much visual content. Then we develop a method that produces a summary of the input video with the minimal redundancy measured. The video summary also contains the original audio segments that are partially synchronized with the summarized visual content. The experimental evaluations show the effectiveness of the redundancy metric and the summarization method. Yihong Gong, Xin Liu 0046 |
ICIP (3) | 1 |
| 2001 | Summarizing Video By Minimizing Visual Content RedundanciesabstractIn this paper, we propose a video summarization method able to produce a motion video summary that minimizes the visual content redundancy for the input video. We derive the visual content redundancy metric that tells how much redundancy a video contains, and how much the video can be curtailed without loosing too much visual content. Then we propose a method that produces a summary of the input video with the minimal redundancy measure. The video summary also contains the original audio segments that are partially synchronized with the summarized visual content. The experimental evaluations show the effectiveness of the redundancy metric and the summarization method. Video summarization examples can be viewed at http://www.ccrl.com/ ygong. Yihong Gong, Xin Liu 0046 |
ICME | 1 |
| 2001 | Generic Text Summarization Using Relevance Measure and Latent Semantic AnalysisabstractIn this paper, we propose two generic text summarization methods that create text summaries by ranking and extracting sentences from the original documents. The first method uses standard IR methods to rank sentence relevances, while the second method uses the latent semantic analysis technique to identify semantically important sentences, for summary creations. Both methods strive to select sentences that are highly ranked and different from each other. This is an attempt to create a summary with a wider coverage of the document's main content and less redundancy. Performance evaluations on the two summarization methods are conducted by comparing their summarization outputs with the manual summaries generated by three independent human evaluators. The evaluations also study the influence of different VSM weighting schemes on the text summarization performances. Finally, the causes of the large disparities in the evaluators' manual summarization results are investigated, and discussions on human text summarization patterns are presented. Yihong Gong, Xin Liu 0046 |
SIGIR | 1 |
| 2000 | Video Summarization Using Singular Value DecompositionabstractThe authors propose a novel technique for video summarization based on singular value decomposition (SVD). For the input video sequence, we create a feature-frame matrix A, and perform the SVD on it. From this SVD, we are able, to not only derive the refined feature space to better cluster visually similar frames, but also define a metric to measure the amount of visual content contained in each frame cluster using its degree of visual changes. Then, in the refined feature space, we find the most static frame cluster, define it as the content unit, and use the context value computed from it as the threshold to cluster the rest of the frames. Based on this clustering result, either the optimal set of keyframes, or a summarized motion video with the user specified time length can be generated to support different user requirements for video browsing and content overview. Our approach ensures that the summarized video representation contains little redundancy, and gives equal attention to the same amount of contents. Yihong Gong, Xin Liu 0046 |
CVPR | 1 |
| 2000 | Video Shot Segmentation and ClassificationabstractWe propose a technique for video shot segmentation and classification based on singular value decomposition (SVD). For the input video sequence, we create a feature-frame matrix A, and perform the SVD on it. From this SVD, we are able to not only derive the refined feature space to better segment the video sequence along the time axis, but also define metrics to enable classifications of the detected video shots. Using these SVD properties, we achieve the two goals of accurate video shot segmentation, and visual content-based shot classification at the same time. Yihong Gong, Xin Liu 0046 |
ICPR | 1 |
| 1999 | A Robust Image Mosaicing Technique Capable of Creating Integrated PanoramasabstractExisting featureless image mosaicing techniques do not pay enough attention to the robustness of the image registration process, and are not able to combine multiple video sequences into an integrated panoramic view. These problems have certainly restricted applications of the existing methods for large-scale panorama composition, video content overview and information visualization. In this paper we propose a method that is able to create an integrated panoramic view for a virtual camera from multiple video sequences which each records a part of a vast scene. The method further enables the user to visualize the integrated panoramic view from an arbitrary viewpoint and orientation by altering the parameters of the virtual camera. To ensure a robust and accurate panoramic view synthesis from long video sequences, we attach a global positioning system (GPS) to the video camera, and utilize its output data to provide initial estimates for the camera's translational parameters, and to prevent the camera parameter recovery process from falling into spurious local minima. Our proposed method is not only suitable for video content overview but also applicable to the areas of information visualization, team collaborations, disastrous rescues, etc. The experimental results demonstrate the effectiveness of the proposed method. Yihong Gong, Guido Proietti, David LaRose |
IV | 1 |
| 1999 | Advancing Content-Based Image Retrieval by Exploiting Image Color and Region Features
Yihong Gong |
Multim. Syst. | 1 |
| 1998 | Image Indexing and Retrieval Based on Human Perceptual Color ClusteringabstractWe propose a new image retrieval method based on human perceptual clustering of color images. This color clustering produces for each image a small set of representative colors which captures the color properties of the image, and a small set of sizable contiguous regions which captures the spatial/geometrical properties of the image. The proposed method outperforms the traditional histogram and its improved methods not only with its richer image retrieval capabilities which cover a wider spectrum of user requirements, but also with its powerful indexing scheme which is essential to cater for large scale image databases. Yihong Gong, Guido Proietti, Christos Faloutsos |
CVPR | 1 |
| 1996 | Image Indexing and Retrieval Based on Color Histograms
Yihong Gong, Chua Hock Chuan, Guo Xiaoyi |
Multim. Tools Appl. | 1 |
| 1995 | Detection of Regions Matching Specified Chromatic Features
Yihong Gong, Masao Sakauchi |
Comput. Vis. Image Underst. | 1 |
| 1995 | Automatic Parsing and Indexing of News Video
HongJiang Zhang, Shuang Yeo Tan, Stephen W. Smoliar, Yihong Gong |
Multim. Syst. | 4 |
| 1992 | A color video image quantization method with stable and efficient color selection capabilityabstractColor video images have become a very important media in communication. This has increased the necessity of displaying and handling color video images on various types of computers. In computerized color image processing, most color images are represented by 24 bits per pixel. Such images usually contain a lot of redundancy and require a large amount of space to be stored. Furthermore, they require expensive full color display devices, so that many general-purpose computers which have only colormap display devices are not capable of displaying them. In order to lower the display and the storage cost, color image quantization algorithms are needed to reduce the number of colors in original images. The authors propose a color quantization method for video images which efficiently generates color quantized video images with stable color allocations.> Yihong Gong, Heitou Zen, Yutaka Ohsawa, Masao Sakauchi |
ICPR (3) | 1 |