EDBT 2026 Demo / reviewers in the wild / expert
Yu Kong 0001
dblp:55/3630 · also Yu Kong Golisano
· DBLP profile ↗
75ranked-venue papers
20as first author
29since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 52 · 11 first-author · 21 since 2021Artificial intelligence and machine learning · 50 · 16 first-author · 23 since 2021Applied, interdisciplinary, general and emerging computing · 3Databases, data management, data science and information retrieval · 2Systems, architecture and hardware · 1Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Open Set Face Forgery Detection via Dual-Level Evidence CollectionabstractThe surge in face forgeries has increasingly undermined confidence in the authenticity of online content. As generation algorithms rapidly evolve, new fake categories will constantly emerge, severely challenging existing face forgery detection methods. Although face forgery detection has recently improved, current techniques remain largely confined to binary Real-vs-Fake classification or the recognition of known fake categories. Moreover, they fail to identify the emergence of entirely new forgery methods. In this work, we study the Open Set Face Forgery Detection (OSFFD) problem, which requires the detection model to identify novel fake categories. To enhance its real-world applicability, we reformulate the OSFFD problem and address it through uncertainty estimation. Specifically, we propose the Dual-Level Evidential face forgery Detection (DLED) approach, which estimates prediction uncertainty by extracting and integrating category-specific evidence on the spatial and frequency levels. Comprehensive experiments across diverse settings demonstrate that our proposed DLED approach achieves state-of-the-art performance. Notably, it surpasses various existing baseline models by a $20\%$ margin on average when identifying forgeries from novel fake categories. Concurrently, our DLED method yields competitive performance on the standard binary Real-versus-Fake face forgery detection task. Zhongyi Cai, Bryce Gernon, Wentao Bao, Matthew Wright 0001, Yu Kong 0001 |
FG | 6 |
| 2026 | Show Me: Unifying Instructional Image and Video Generation with Diffusion ModelsabstractGenerating visual instructions in a given context is essential for developing interactive world simulators. While prior works address this problem through either text-guided image manipulation or video prediction, these tasks are typically treated in isolation. This separation reveals a fundamental issue: image manipulation methods overlook how actions unfold over time, while video prediction models often ignore the intended outcomes. To this end, we propose ShowMe, a unified framework that enables both tasks by selectively activating the spatial and temporal components of video diffusion models. In addition, we introduce structure and motion consistency rewards to improve structural fidelity and temporal coherence. Notably, this unification brings dual benefits: the spatial knowledge gained through video pretraining enhances contextual consistency and realism in non-rigid image edits, while the instruction-guided manipulation stage equips the model with stronger goal-oriented reasoning for video prediction. Experiments on diverse benchmarks demonstrate that our method outperforms expert models in both instructional image and video generation, highlighting the strength of video diffusion models as a unified action-object state transformer. Our code will be available at https://yujiangpu20.github.io/showme/. Yujiang Pu, Zhanbo Huang, Vishnu Naresh Boddeti, Yu Kong 0001 |
WACV | 4 |
| 2025 | IndustryEQA: Pushing the Frontiers of Embodied Question Answering in Industrial ScenariosabstractExisting Embodied Question Answering (EQA) benchmarks primarily focus on household environments, often overlooking safety-critical aspects and reasoning processes pertinent to industrial settings. This drawback limits the evaluation of agent readiness for real-world industrial applications. To bridge this, we introduce IndustryEQA, the first benchmark dedicated to evaluating embodied agent capabilities within safety-critical industrial warehouse scenarios. Built upon the NVIDIA Isaac Sim platform, IndustryEQA provides high-fidelity episodic memory videos featuring diverse industrial assets, dynamic human agents, and carefully designed hazardous situations inspired by real-world safety guidelines. The benchmark includes rich annotations covering six categories: equipment safety, human safety, object recognition, attribute recognition, temporal understanding, and spatial understanding. Besides, it also provides extra reasoning evaluation based on these categories. Specifically, it comprises 971 question-answer pairs generated from small warehouse scenarios and 373 pairs from large ones, incorporating scenarios with and without human. We further propose a comprehensive evaluation framework, including various baseline models, to assess their general perception and reasoning abilities in industrial environments. IndustryEQA aims to steer EQA research towards developing more robust, safety-aware, and practically applicable embodied agents for complex industrial environments. Anh Dao, Lichi Li, Zhongyi Cai, Zhen Tan 0001, Tianlong Chen 0001, Yu Kong 0001 |
NeurIPS | 8 |
| 2025 | Exploiting VLM Localizability and Semantics for Open Vocabulary Action DetectionabstractAction detection aims to detect (recognize and localize) human actions spatially and temporally in videos. Existing approaches focus on the closed-set setting where an action detector is trained and tested on videos from a fixed set of action categories. However, this constrained setting is not viable in an open world where test videos inevitably come beyond the trained action categories. In this paper, we address the practical yet challenging Open-Vocabulary Action Detection (OVAD) problem. It aims to detect any action in test videos while training a model on a fixed set of action categories. To achieve such an open-vocabulary capability, we propose a novel method OpenMixer that exploits the inherent semantics and localizability of large vision-language models (VLM) within the family of query-based detection transformers (DETR). Specifically, the OpenMixer is developed by spatial and temporal OpertMixer blocks (S-OMB and T-OMB), and a dynamically fused alignment (DFA) module. The three components collectively enjoy the merits of strong generalization from pretrained VLMs and end-to-end learning from DETR design. Moreover, we established OVAD benchmarks under various settings, and the experimental results show that the OpenMixer performs the best over baselines for detecting seen and unseen actions. We release the codes, models, and dataset splits at https://github.com/Cogito2012/0penMixer. Wentao Bao, Kai Li 0012, Yuxiao Chen 0002, Deep Patel, Martin Renqiang Min, Yu Kong 0001 |
WACV | 6 |
| 2024 | Prompting Language-Informed Distribution for Compositional Zero-Shot Learning
Wentao Bao, Lichang Chen, Heng Huang 0001, Yu Kong 0001 |
ECCV (14) | 4 |
| 2024 | Learning to Localize Actions in Instructional Videos with LLM-Based Multi-pathway Text-Video Alignment
Yuxiao Chen 0002, Kai Li 0012, Wentao Bao, Deep Patel, Yu Kong 0001, Martin Renqiang Min, Dimitris N. Metaxas |
ECCV (82) | 5 |
| 2024 | SHINE: Saliency-Aware Hierarchical Negative Ranking for Compositional Temporal Grounding
Zixu Cheng, Yujiang Pu, Shaogang Gong, Parisa Kordjamshidi, Yu Kong 0001 |
ECCV (19) | 5 |
| 2024 | Facial Affective Behavior Analysis with Instruction Tuning
Anh Dao, Wentao Bao, Zhen Tan 0001, Tianlong Chen 0001, Huan Liu 0001, Yu Kong 0001 |
ECCV (18) | 7 |
| 2024 | A Survey of Multimodal Sarcasm Detection
Shafkat Farabi, Tharindu Ranasinghe, Diptesh Kanojia, Yu Kong 0001, Marcos Zampieri |
IJCAI | 4 |
| 2023 | Catch Missing Details: Image Reconstruction with Frequency Augmented Variational AutoencoderabstractThe popular VQ-VAE models reconstruct images through learning a discrete codebook but suffer from a significant issue in the rapid quality degradation of image reconstruction as the compression rate rises. One major reason is that a higher compression rate induces more loss of visual signals on the higher frequency spectrum which reflect the details on pixel space. In this paper, a Frequency Complement Module (FCM) architecture is proposed to capture the missing frequency information for enhancing reconstruction quality. The FCM can be easily incorporated into the VQ-VAE structure, and we refer to the new model as Frequancy Augmented VAE (FAVAE). In addition, a Dynamic Spectrum Loss (DSL) is introduced to guide the FCMs to balance between various frequencies dynamically for optimal reconstruction. FA-VAE is further extended to the text-to-image synthesis task, and a Crossattention Autoregressive Transformer (CAT) is proposed to obtain more precise semantic attributes in texts. Extensive reconstruction experiments with different compression rates are conducted on several benchmark datasets, and the results demonstrate that the proposed FA-VAE is able to restore more faithfully the details compared to SOTA methods. CAT also shows improved generation quality with better image-text semantic alignment. Xinmiao Lin, Yikang Li 0001, Jenhao Hsiao, Chiuman Ho, Yu Kong 0001 |
CVPR | 5 |
| 2023 | Uncertainty-aware State Space Transformer for Egocentric 3D Hand Trajectory ForecastingabstractHand trajectory forecasting from egocentric views is crucial for enabling a prompt understanding of human intentions when interacting with AR/VR systems. However, existing methods handle this problem in a 2D image space which is inadequate for 3D real-world applications. In this paper, we set up an egocentric 3D hand trajectory forecasting task that aims to predict hand trajectories in a 3D space from early observed RGB videos in a first-person view. To fulfill this goal, we propose an uncertainty-aware state space Transformer (USST) that takes the merits of the attention mechanism and aleatoric uncertainty within the framework of the classical state-space model. The model can be further enhanced by the velocity constraint and visual prompt tuning (VPT) on large vision transformers. Moreover, we develop an annotation workflow to collect 3D hand trajectories with high quality. Experimental results on H2O and EgoPAT3D datasets demonstrate the superiority of USST for both 2D and 3D trajectory forecasting. The code and datasets are publicly released: https://actionlab-cv.github.io/EgoHandTrajPred. Wentao Bao, Libing Zeng, Zhong Li 0007, Yi Xu 0002, Junsong Yuan 0001, Yu Kong 0001 |
ICCV | 7 |
| 2023 | ATM: Action Temporality Modeling for Video Question AnsweringabstractDespite significant progress in video question answering (VideoQA), existing methods fall short of questions that require causal/temporal reasoning across frames. This can be attributed to imprecise motion representations. We introduce Action Temporality Modeling (ATM) for temporality reasoning via three-fold uniqueness: (1) rethinking the optical flow and realizing that optical flow is effective in capturing the long horizon temporality reasoning; (2) training the visual-text embedding by contrastive learning in an action-centric manner, leading to better action representations in both vision and text modalities; and (3) preventing the model from answering the question given the shuffled video in the fine-tuning stage, to avoid spurious correlation between appearance and motion and hence ensure faithful temporality reasoning. In the experiments, we show that ATM outperforms existing approaches in terms of the accuracy on multiple VideoQAs and exhibits better true temporality reasoning ability. Junwen Chen 0001, Yu Kong 0001 |
ACM Multimedia | 3 |
| 2023 | Ancestor Search: Generalized Open Set Recognition via Hyperbolic Side Information LearningabstractDifferent from the open set recognition, generalized open set recognition learns the most similar known classes for unseen samples using known classes samples and side in-formation of known classes. It is challenging because hierarchically structured side information is distorted when features are embedded in the Euclidean space in existing literature, which incurs the difficulty of identifying the unseen samples. In this paper, we introduce a side information learning algorithm for generalized open set recognition based on the hyperbolic space to alleviate the distortion and accurately identify the unknown samples. Specifically, we propose a hyperbolic side information learning framework to identify the unseen samples and an ancestor search algorithm to search the most similar ancestor from the taxonomy of selected known classes. Experiments on CUB-200 and AWA 2 datasets show that our method improves the performance of generalized open set recognition by a large margin. Xiwen Dengxiong, Yu Kong 0001 |
WACV | 2 |
| 2023 | From Ensemble Clustering to Subspace Clustering: Cluster Structure EncodingabstractIn this study, we propose a novel algorithm to encode the cluster structure by incorporating ensemble clustering (EC) into subspace clustering (SC). First, the low-rank representation (LRR) is learned from a higher order data relationship induced by ensemble K-means coding, which exploits the cluster structure in a co-association matrix of basic partitions (i.e., clustering results). Second, to provide a fast predictive coding mechanism, an encoding function parameterized by neural networks is introduced to predict the LRR derived from partitions. These two steps are jointly proceeded to seamlessly integrate partition information and original features and thus deliver better representations than the ones obtained from each single source. Moreover, an alternating optimization framework is developed to learn the LRR, train the encoding function, and fine-tune the higher order relationship. Extensive experiments on eight benchmark datasets validate the effectiveness of the proposed algorithm on several clustering tasks compared with state-of-the-art EC and SC methods. Zhiqiang Tao, Jun Li 0027, Huazhu Fu, Yu Kong 0001, Yun Fu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2022 | A Dynamic Meta-Learning Model for Time-Sensitive Cold-Start RecommendationsabstractWe present a novel dynamic recommendation model that focuses on users who have interactions in the past but turn relatively inactive recently. Making effective recommendations to these time-sensitive cold-start users is critical to maintain the user base of a recommender system. Due to the sparse recent interactions, it is challenging to capture these users' current preferences precisely. Solely relying on their historical interactions may also lead to outdated recommendations misaligned with their recent interests. The proposed model leverages historical and current user-item interactions and dynamically factorizes a user's (latent) preference into time-specific and time-evolving representations that jointly affect user behaviors. These latent factors further interact with an optimized item embedding to achieve accurate and timely recommendations. Experiments over real-world data help demonstrate the effectiveness of the proposed time-sensitive cold-start recommendation model. Krishna Prasad Neupane, Ervine Zheng, Yu Kong 0001, Qi Yu 0001 |
AAAI | 3 |
| 2022 | OpenTAL: Towards Open Set Temporal Action LocalizationabstractTemporal Action Localization (TAL) has experienced remarkable success under the supervised learning paradigm. However, existing TAL methods are rooted in the closed set assumption, which cannot handle the inevitable unknown actions in open-world scenarios. In this paper, we, for the first time, step toward the Open Set TAL (OSTAL) problem and propose a general framework Open TAL based on Evidential Deep Learning (EDL). Specifically, the OpenTAL consists of uncertainty-aware action classification, actionness prediction, and temporal location regression. With the proposed importance-balanced EDL method, classification uncertainty is learned by collecting categorical evidence majorly from important samples. To distinguish the unknown actions from background video frames, the actionness is learned by the positive-unlabeled learning. The classification uncertainty is further calibrated by leveraging the guidance from the temporal localization quality. The OpenTAL is general to enable existing TAL models for open set scenarios, and experimental results on THUMOS14 and ActivityNet1.3 benchmarks show the effectiveness of our method. The code and pre-trained models are released at https://www.rit.edu/actionlab/opental. Wentao Bao, Qi Yu 0001, Yu Kong 0001 |
CVPR | 3 |
| 2022 | GateHUB: Gated History Unit with Background Suppression for Online Action DetectionabstractOnline action detection is the task of predicting the action as soon as it happens in a streaming video. A major challenge is that the model does not have access to the future and has to solely rely on the history, i.e., the frames observed so far, to make predictions. It is therefore important to accentuate parts of the history that are more informative to the prediction of the current frame. We present GateHUB, Gated History Unit with Background Suppression, that comprises a novel position-guided gated cross attention mechanism to enhance or suppress parts of the history as per how informative they are for current frame prediction. GateHUB further proposes Future-augmented History (FaH) to make history features more informative by using subsequently observed frames when available. In a single unified framework, GateHUB integrates the transformer's ability of long-range temporal modeling and the recurrent model's capacity to selectively encode relevant information. GateHUB also introduces a background suppression objective to further mitigate false positive background frames that closely resemble the action frames. Extensive validation on three benchmark datasets, THUMOS, TVSeries, and HDD, demonstrates that GateHUB significantly outperforms all existing methods and is also more efficient than the existing best work. Furthermore, a flow free version of GateHUB is able to achieve higher or close accuracy at 2.8× higher frame rate compared to all existing methods that require both RGB and optical flow information for prediction. Junwen Chen 0001, Gaurav Mittal, Ye Yu 0003, Yu Kong 0001 |
CVPR | 4 |
| 2022 | Learning of Global Objective for Network Flow in Multi-Object TrackingabstractThis paper concerns the problem of multi-object tracking based on the min-cost flow (MCF) formulation, which is conventionally studied as an instance of linear program. Given its computationally tractable inference, the success of MCF tracking largely relies on the learned cost function of underlying linear program. Most previous studies focus on learning the cost function by only taking into account two frames during training, therefore the learned cost function is sub-optimal for MCF where a multi-frame data association must be considered during inference. In order to address this problem, in this paper we propose a novel differentiable framework that ties training and inference to-gether during learning by solving a bi-level optimization problem, where the lower-level solves a linear program and the upper-level contains a loss function that incorpo-rates global tracking result. By back-propagating the loss through differentiable layers via gradient descent, the glob-ally parameterized cost function is explicitly learned and regularized. With this approach, we are able to learn a better objective for global MCF tracking. As a result, we achieve competitive performances compared to the current state-of-the-art methods on the popular multi-object tracking benchmarks such as MOT16, MOT17 and MOT20. Yu Kong 0001, Seyed Hamid Rezatofighi |
CVPR | 2 |
| 2022 | Universal 3-Dimensional Perturbations for Black-Box Attacks on Video Recognition SystemsabstractWidely deployed deep neural network (DNN) models have been proven to be vulnerable to adversarial perturbations in many applications (e.g., image, audio and text classifications). To date, there are only a few adversarial perturbations proposed to deviate the DNN models in video recognition systems by simply injecting 2D perturbations into video frames. However, such attacks may overly perturb the videos without learning the spatio-temporal features (across temporal frames), which are commonly extracted by DNN models for video recognition. To our best knowledge, we propose the first black-box attack framework that generates universal 3-dimensional (U3D) perturbations to subvert a variety of video recognition systems. U3D has many advantages, such as (1) as the transfer-based attack, U3D can universally attack multiple DNN models for video recognition without accessing to the target DNN model; (2) the high transferability of U3D makes such universal black-box attack easy-to-launch, which can be further enhanced by integrating queries over the target model when necessary; (3) U3D ensures human-imperceptibility; (4) U3D can bypass the existing state-of-the-art defense schemes; (5) U3D can be efficiently generated with a few pre-learned parameters, and then immediately injected to attack real-time DNN-based video recognition systems. We have conducted extensive experiments to evaluate U3D on multiple DNN models and three large-scale video datasets. The experimental results demonstrate its superiority and practicality. Shangyu Xie, Han Wang 0021, Yu Kong 0001, Yuan Hong 0001 |
SP | 3 |
| 2022 | Human Action Recognition and Prediction: A Survey
Yu Kong 0001, Yun Fu 0001 |
Int. J. Comput. Vis. | 1 |
| 2021 | Gradient Frequency Modulation for Visually Explaining Video Understanding Models
Xinmiao Lin, Wentao Bao, Matthew Wright 0001, Yu Kong 0001 |
BMVC | 4 |
| 2021 | DRIVE: Deep Reinforced Accident Anticipation with Visual ExplanationabstractTraffic accident anticipation aims to accurately and promptly predict the occurrence of a future accident from dashcam videos, which is vital for a safety-guaranteed self-driving system. To encourage an early and accurate decision, existing approaches typically focus on capturing the cues of spatial and temporal context before a future accident occurs. However, their decision-making lacks visual explanation and ignores the dynamic interaction with the environment. In this paper, we propose Deep ReInforced accident anticipation with Visual Explanation, named DRIVE. The method simulates both the bottom-up and top-down visual attention mechanism in a dashcam observation environment so that the decision from the pro-posed stochastic multi-task agent can be visually explained by attentive regions. Moreover, the proposed dense anticipation reward and sparse fixation reward are effective in training the DRIVE model with our improved reinforcement learning algorithm. Experimental results show that the DRIVE model achieves state-of-the-art performance on multiple real-world traffic accident datasets. Code and pre-trained model are available at https://www.rit.edu/actionlab/drive. Wentao Bao, Qi Yu 0001, Yu Kong 0001 |
ICCV | 3 |
| 2021 | Evidential Deep Learning for Open Set Action RecognitionabstractIn a real-world scenario, human actions are typically out of the distribution from training data, which requires a model to both recognize the known actions and reject the unknown. Different from image data, video actions are more challenging to be recognized in an open-set setting due to the uncertain temporal dynamics and static bias of human actions. In this paper, we propose a Deep Evidential Action Recognition (DEAR) method to recognize actions in an open testing set. Specifically, we formulate the action recognition problem from the evidential deep learning (EDL) perspective and propose a novel model calibration method to regularize the EDL training. Besides, to mitigate the static bias of video representation, we propose a plug-and-play module to debias the learned representation through contrastive learning. Experimental results show that our DEAR method achieves consistent performance gain on multiple mainstream action recognition models and benchmarks. Code and pre-trained models are available at https://www.rit.edu/actionlab/dear. Wentao Bao, Qi Yu 0001, Yu Kong 0001 |
ICCV | 3 |
| 2021 | Explainable Video Entailment with Grounded Visual EvidenceabstractVideo entailment aims at determining if a hypothesis textual statement is entailed or contradicted by a premise video. The main challenge of video entailment is that it requires fine-grained reasoning to understand the complex and long story-based videos. To this end, we propose to incorporate visual grounding to the entailment by explicitly linking the entities described in the statement to the evidence in the video. If the entities are grounded in the video, we enhance the entailment judgment by focusing on the frames where the entities occur. Besides, in the entailment dataset, the entailed/contradictory (also named as real/fake) statements are formed in pairs with subtle discrepancy, which allows an add-on explanation module to predict which words or phrases make the statement contradictory to the video and regularize the training of the entailment judgment. Experimental results demonstrate that our approach outperforms the state-of-the-art methods. Junwen Chen 0001, Yu Kong 0001 |
ICCV | 2 |
| 2021 | Multiple Instance Relational Learning for Video Anomaly DetectionabstractMost existing video anomaly detection methods are dependent on strong supervision to achieve satisfactory performance, which could be laborious and impractical. Besides, methods using weakly supervised learning less consider the relations among event proposals. To this end, we propose an anomaly event detection method by the instance-based event proposal generation and the proposal relation learning. Specifically, the event proposals are generated by sampling temporal distributions from the multiple instance learning (MIL), while relations among the proposals are captured by graph convolutional network for anomaly localization and classification. The proposed method is free from strong frame-level supervision but only requires video-level annotations. We conduct experiments on four datasets, i.e., UCF crime, UCSD-Peds, UMN, and CityScene, and show state-of-the-art performance on anomaly recognition and detection tasks. Xiwen Dengxiong, Wentao Bao, Yu Kong 0001 |
IJCNN | 3 |
| 2021 | Revealing a history: palimpsest text separation with generative networks
Anna Starynska, David W. Messinger, Yu Kong 0001 |
Int. J. Document Anal. Recognit. | 3 |
| 2021 | Coupling Adversarial Graph Embedding for transductive zero-shot action recognition
Wanru Xu, Yu Kong 0001 |
Neurocomputing | 4 |
| 2021 | Residual Dense Network for Image RestorationabstractRecently, deep convolutional neural network (CNN) has achieved great success for image restoration (IR) and provided hierarchical features at the same time. However, most deep CNN based IR models do not make full use of the hierarchical features from the original low-quality images; thereby, resulting in relatively-low performance. In this work, we propose a novel and efficient residual dense network (RDN) to address this problem in IR, by making a better tradeoff between efficiency and effectiveness in exploiting the hierarchical features from all the convolutional layers. Specifically, we propose residual dense block (RDB) to extract abundant local features via densely connected convolutional layers. RDB further allows direct connections from the state of preceding RDB to all the layers of current RDB, leading to a contiguous memory mechanism. To adaptively learn more effective features from preceding and current local features and stabilize the training of wider network, we proposed local feature fusion in RDB. After fully obtaining dense local features, we use global feature fusion to jointly and adaptively learn global hierarchical features in a holistic way. We demonstrate the effectiveness of RDN with several representative IR applications, single image super-resolution, Gaussian image denoising, image compression artifact reduction, and image deblurring. Experiments on benchmark and real-world datasets show that our RDN achieves favorable performance against state-of-the-art methods for each IR task quantitatively and visually. Yulun Zhang 0001, Yapeng Tian, Yu Kong 0001, Bineng Zhong 0001, Yun Fu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2021 | Accurate and Fast Image Denoising via Attention Guided ScalingabstractImage denoising is a classical topic yet still a challenging problem, especially for reducing noise from the texture information. Feature scaling (e.g., downscale and upscale) is a widely practice in image denoising to enlarge receptive field size and save resources. However, such a common operation would lose some visual informative details. To address those problems, we propose fast and accurate image denoising via attention guided scaling (AGS). We find that the main informative feature channel and visual primitives during the scaling should keep similar. We then propose to extract the global channel-wise attention to maintain main channel information. Moreover, we propose to collect global descriptors by considering the entire spatial feature. And we then distribute the global descriptors to local positions of the scaled feature, based on their specific needs. We further introduce AGS for adversarial training, resulting in a more powerful discriminator. Extensive experiments show the effectiveness of our proposed method, where we clearly surpass all the state-of-the-art methods on most popular synthetic and real-world denoising benchmarks quantitatively and visually. We further show that our network contributes to other high-level vision applications and improves their performances significantly. Yulun Zhang 0001, Kai Li 0012, Gan Sun, Yu Kong 0001, Yun Fu 0001 |
IEEE Trans. Image Process. | 5 |
| 2020 | Group Activity Prediction with Sequential Relational Anticipation Model
Junwen Chen 0001, Wentao Bao, Yu Kong 0001 |
ECCV (21) | 3 |
| 2020 | Publishing Video Data with Indistinguishable Objectsabstractfor all the predefined sensitive objects (e.g., humans and vehicles) in the video, and then propose a video sanitization technique VERRO that randomly generates utility-driven synthetic videos with indistinguishable objects. Therefore, all the objects can be well protected in the generated utility-driven synthetic videos which can be disclosed to any untrusted video recipient. We have conducted extensive experiments on three real videos captured for pedestrians on the streets. The experimental results demonstrate that the generated synthetic videos lie close to the original video for retaining good utility while ensuring rigorous privacy guarantee. Han Wang 0021, Yuan Hong 0001, Yu Kong 0001, Jaideep Vaidya |
EDBT | 3 |
| 2020 | Privacy Attributes-aware Message Passing Neural Network for Visual Privacy Attributes ClassificationabstractVisual Privacy Attribute Classification (VPAC) identifies privacy information leakage via social media images. These images containing privacy attributes such as skin color, face or gender are classified into multiple privacy attribute categories in VPAC. With limited works in this task, current methods often extract features from images and simply classify the extracted feature into multiple privacy attribute classes. The dependencies between privacy attributes, e.g., skin color and face typically coexist in the same image, are usually ignored in classification, which causes performance degradation in VPAC. In this paper, we propose a novel end-to-end Privacy Attributes-aware Message Passing Neural Network (PA-MPNN) to address VPAC. Privacy attributes are considered as nodes on a graph and an MPNN is introduced to model the privacy attribute dependencies. To generate representative features for privacy attribute nodes, a class-wise encoder-decoder is proposed to learn a latent space for each attribute. An attention mechanism with multiple correlation matrices is also introduced in MPNN to learn the privacy attributes graph automatically. Experimental results on the Privacy Attribute Dataset demonstrate that our framework achieves better performance than state-of-the-art methods for visual privacy attributes classification. Hanbin Hong, Wentao Bao, Yuan Hong 0001, Yu Kong 0001 |
ICPR | 4 |
| 2020 | Few-shot Human Motion Prediction via Learning Novel Motion DynamicsabstractHuman motion prediction is a task where we anticipate future motion based on past observation. Previous approaches rely on the access to large datasets of skeleton data, and thus are difficult to be generalized to novel motion dynamics with limited training data. In our work, we propose a novel approach named Motion Prediction Network (MoPredNet) for few-short human motion prediction. MoPredNet can be adapted to predicting new motion dynamics using limited data, and it elegantly captures long-term dependency in motion dynamics. Specifically, MoPredNet dynamically selects the most informative poses in the streaming motion data as masked poses. In addition, MoPredNet improves its encoding capability of motion dynamics by adaptively learning spatio-temporal structure from the observed poses and masked poses. We also propose to adapt MoPredNet to novel motion dynamics based on accumulated motion experiences and limited novel motion dynamics data. Experimental results show that our method achieves better performance over state-of-the-art methods in motion prediction. Chuanqi Zang, Mingtao Pei, Yu Kong 0001 |
IJCAI | 3 |
| 2020 | Object-Aware Centroid Voting for Monocular 3D Object DetectionabstractMonocular 3D object detection aims to detect objects in a 3D physical world from a single camera. However, recent approaches either rely on expensive LiDAR devices, or resort to dense pixel-wise depth estimation that causes prohibitive computational cost. In this paper, we propose an end-to-end trainable monocular 3D object detector without learning the dense depth. Specifically, the grid coordinates of a 2D box are first projected back to 3D space with the pinhole model as 3D centroids proposals. Then, a novel object-aware voting approach is introduced, which considers both the region-wise appearance attention and the geometric projection distribution, to vote the 3D centroid proposals for 3D object localization. With the late fusion and the predicted 3D orientation and dimension, the 3D bounding boxes of objects can be detected from a single RGB image. The method is straightforward yet significantly superior to other monocular-based methods. Extensive experimental results on the challenging KITTI benchmark validate the effectiveness of the proposed method. Wentao Bao, Qi Yu 0001, Yu Kong 0001 |
IROS | 3 |
| 2020 | Uncertainty-based Traffic Accident Anticipation with Spatio-Temporal Relational LearningabstractTraffic accident anticipation aims to predict accidents from dashcam videos as early as possible, which is critical to safety-guaranteed self-driving systems. With cluttered traffic scenes and limited visual cues, it is of great challenge to predict how long there will be an accident from early observed frames. Most existing approaches are developed to learn features of accident-relevant agents for accident anticipation, while ignoring the features of their spatial and temporal relations. Besides, current deterministic deep neural networks could be overconfident in false predictions, leading to high risk of traffic accidents caused by self-driving systems. In this paper, we propose an uncertainty-based accident anticipation model with spatio-temporal relational learning. It sequentially predicts the probability of traffic accident occurrence with dashcam videos. Specifically, we propose to take advantage of graph convolution and recurrent networks for relational feature learning, and leverage Bayesian neural networks to address the intrinsic variability of latent relational representations. The derived uncertainty-based ranking loss is found to significantly boost model performance by improving the quality of relational features. In addition, we collect a new Car Crash Dataset (CCD) for traffic accident anticipation which contains environmental attributes and accident reasons annotations. Experimental results on both public and the newly-compiled datasets show state-of-the-art performance of our model. Our code and CCD dataset are available at https://github.com/Cogito2012/UString. Wentao Bao, Qi Yu 0001, Yu Kong 0001 |
ACM Multimedia | 3 |
| 2020 | Activity-driven Weakly-Supervised Spatio-Temporal Grounding from Untrimmed VideosabstractIn this paper, we study the problem of weakly-supervised spatio-temporal grounding from raw untrimmed video streams. Given a video and its descriptive sentence, spatio-temporal grounding aims at predicting the temporal occurrence and spatial locations of each query object across frames. Our goal is to learn a grounding model in a weakly-supervised fashion, without the supervision of both spatial bounding boxes and temporal occurrences during training. Existing methods have been addressed in trimmed videos, but their reliance on object tracking will easily fail due to frequent camera shot cut in untrimmed videos. To this end, we propose a novel spatio-temporal multiple instance learning framework for untrimmed video grounding. Spatial MIL and temporal MIL are mutually guided to ground each query to specific spatial regions and the occurring frames of a video. Furthermore, an activity described in the sentence is captured to use the informative contextual cues for region proposals refinement and text representation. We conduct extensive evaluation on YouCookII and RoboWatch datasets, and demonstrate our method outperforms state-of-the-art methods. Junwen Chen 0001, Wentao Bao, Yu Kong 0001 |
ACM Multimedia | 3 |
| 2020 | Adversarial Action Prediction NetworksabstractDifferent from after-the-fact action recognition, action prediction task requires action labels to be predicted from partially observed videos containing incomplete action executions. It is challenging because these partial videos have insufficient discriminative information, and their temporal structure is damaged. We study this problem in this paper, and propose an efficient and powerful deep network for learning representative and discriminative features for action prediction. Our approach exploits abundant sequential context information in full videos to enrich the feature representations of partial videos. This information is encoded in latent representations using a variational autoencoder (VAE), which are encouraged to be progress-invariant. Decoding such latent representations using another VAE, we can reconstruct missing information in the features extracted from partial videos. An adversarial learning scheme is adopted to differentiate the reconstructed features from the features directly extracted from full videos in order to well align their distributions. A multi-class classifier is also used to encourage the features to be discriminative. Our network jointly learns features and classifiers, and generates the features particularly optimized for action prediction. Extensive experimental results on UCF101, Sports-1M and BIT datasets demonstrate that our approach remarkably outperforms state-of-the-art methods, and shows significant speedup over these methods. Results also show that actions differ in their prediction characteristics; some actions can be correctly predicted even though only the beginning 10% portion of videos is observed. Yu Kong 0001, Zhiqiang Tao, Yun Fu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2020 | Semi-Supervised Cross-Modality Action Recognition by Latent Tensor Transfer LearningabstractMicrosoft's Kinect sensors are receiving an increasing amount of interests by security researchers since they are cost-effective and can provide both visual and depth modality data at the same time. Unfortunately, depth or RGB modalities are unavailable in training or testing procedures in some realistic scenarios. Therefore, we explore a new problem focusing on the arbitrary absence of modality, which is completely different from the conventional action recognition. The new problem in action recognition aims to deal with cross modality data (e.g., RGB training and depth testing data), “missing” modality data (e.g., RGB training and RGB-D test data), and single-modality data (e.g., RGB/depth in both phases). Accordingly, our method aims to borrow some information (e.g., correlation between two modalities) from the well-established RGB-D dataset and apply it to the existing dataset to recover some latent information to improve the performance of recognition. For instance, a cross-modality regularizer is used to preserve the correlation of RGB and depth modalities. The “missing” knowledge is considered as latent information, which is recovered by low-rank learning in our model. In the real world, the target data are usually sparsely labeled or completely unlabeled; however, we could exploit the pseudolabels of the target as prior knowledge for “supervised” learning in the target domain. Accordingly, we propose a semi-supervised model for transfer learning. The experiments on three widely used RGB-D action datasets show that our method performs better than that of the state-of-the-art transfer learning methods in most cases in terms of accuracy and time efficiency. Chengcheng Jia, Zhengming Ding, Yu Kong 0001, Yun Fu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2020 | Aligned Dynamic-Preserving Embedding for Zero-Shot Action RecognitionabstractZero-shot learning (ZSL) typically explores a shared semantic space in order to recognize novel categories in the absence of any labeled training data. However, the traditional ZSL methods always suffer from serious domain shift problem in human action recognition. This is because: 1) existing ZSL methods are specifically designed for object recognition from static images, which do not capture the temporal dynamics of video sequences, and poor performances are always generated if those methods are directly applied to zero-shot action recognition; 2) these methods always blindly project the target data into a shared space using a semantic mapping obtained by the source data without any adaptation, in which the underlying structures of target data are ignored; and 3) severe inter-class variations exist in various action categories. The traditional ZSL methods do not take relationships across different categories into consideration. In this paper, we propose a novel aligned dynamic-preserving embedding (ADPE) model for zero-shot action recognition in a transductive setting. In our model, an adaptive embedding of target videos is learned, exploring the distributions of both the source and target data. An aligned regularization is further proposed to couple the centers of target semantic representations with their corresponding label prototypes in order to preserve the relationships across different categories. Most significantly, during our embedding, the temporal dynamics of video sequences are simultaneously preserved via exploiting the temporal consistency of video sequences and capturing the temporal evolution of successive segments of actions. Our model can effectively overcome the domain shift problem in zero-shot action recognition. The experiments on Olympic sports, HMDB51, and UCF101 datasets demonstrate the effectiveness of our model. Yu Kong 0001, Qiuqi Ruan, Gaoyun An, Yun Fu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2020 | Visual Object Tracking Via Multi-Stream Deep Similarity Learning NetworksabstractVisual tracking remains a challenging research problem because of appearance variations of the object over time, changing cluttered background and requirement for real-time speed. In this paper, we investigate the problem of real-time accurate tracking in a instance-level tracking-by-verification mechanism. We propose a multi-stream deep similarity learning network to learn a similarity comparison model purely off-line. Our loss function encourages the distance between a positive patch and the background patches to be larger than that between the positive patch and the target template. Then, the learned model is directly used to determine the patch in each frame that is most distinctive to the background context and similar to the target template. Within the learned feature space, even if the distance between positive patches becomes large caused by the interference of background clutter, impact from hard distractors from the same class or the appearance change of the target, our method can still distinguish the target robustly using the relative distance. Besides, we also propose a complete framework considering the recovery from failures and the template updating to further improve the tracking performance without taking too much computing resource. Experiments on visual tracking benchmarks show the effectiveness of the proposed tracker when comparing with several recent real-time-speed trackers as well as trackers already included in the benchmarks. Yu Kong 0001, Yun Fu 0001 |
IEEE Trans. Image Process. | 2 |
| 2019 | Deep Geo-Constrained Auto-Encoder for Non-Landmark GPS EstimationabstractThis paper addresses the problem of geotagging images, i.e., assigning GPS coordinates (i.e., latitude, longitude) to images using image contents. Due to the huge appearance variability of visual features across the world, the images' contents and their GPS coordinates may be inconsistent. This means images captured from geographically close areas may appear visually distinct; and images with visually similar contents may be taken from geographically distant areas. In this paper, we propose a deep Geo-constrained Auto-encoder (DGAE) to solve these inconsistency problems. Using clustered GPS data and visual data, our approach identifies inconsistent data pairs (i.e., image, GPS). We then propose a novel deep learning framework that can learn similar feature representations for geographically close images and distinct feature representations for geographically distant images. We introduce two new constraints: the same-area constraint and the easy-confusing constraint to our feature learning networks. The former one penalizes images from the same area but with very distinct visual features, and the latter one penalizes images from distant areas but with very similar visual features. A deep architecture is developed to further improve learning discriminative features, which can disambiguate different geometric locations. Our approach is extensively evaluated on a newly-compiled large image geotagging dataset from large-scale community-contributed images with 664,720 images and outperforms comparison approaches. Shuhui Jiang, Yu Kong 0001, Yun Fu 0001 |
IEEE Trans. Big Data | 2 |
| 2018 | Action Prediction From Videos via Memorizing Hard-to-Predict SamplesabstractAction prediction based on video is an important problem in computer vision field with many applications, such as preventing accidents and criminal activities. It's challenging to predict actions at the early stage because of the large variations between early observed videos and complete ones. Besides, intra-class variations cause confusions to the predictors as well. In this paper, we propose a mem-LSTM model to predict actions in the early stage, in which a memory module is introduced to record several "hard-to-predict" samples and a variety of early observations. Our method uses Convolution Neural Network (CNN) and Long Short-Term Memory (LSTM) to model partial observed video input. We augment LSTM with a memory module to remember challenging video instances. With the memory module, our mem-LSTM model not only achieves impressive performance in the early stage but also makes predictions without the prior knowledge of observation ratio. Information in future frames is also utilized using a bi-directional layer of LSTM. Experiments on UCF-101 and Sports-1M datasets show that our method outperforms state-of-the-art methods. Yu Kong 0001, Shangqian Gao, Bin Sun 0002, Yun Fu 0001 |
AAAI | 1 |
| 2018 | Residual Dense Network for Image Super-ResolutionabstractA very deep convolutional neural network (CNN) has recently achieved great success for image super-resolution (SR) and offered hierarchical features as well. However, most deep CNN based SR models do not make full use of the hierarchical features from the original low-resolution (LR) images, thereby achieving relatively-low performance. In this paper, we propose a novel residual dense network (RDN) to address this problem in image SR. We fully exploit the hierarchical features from all the convolutional layers. Specifically, we propose residual dense block (RDB) to extract abundant local features via dense connected convolutional layers. RDB further allows direct connections from the state of preceding RDB to all the layers of current RDB, leading to a contiguous memory (CM) mechanism. Local feature fusion in RDB is then used to adaptively learn more effective features from preceding and current local features and stabilizes the training of wider network. After fully obtaining dense local features, we use global feature fusion to jointly and adaptively learn global hierarchical features in a holistic way. Experiments on benchmark datasets with different degradation models show that our RDN achieves favorable performance against state-of-the-art methods. Yulun Zhang 0001, Yapeng Tian, Yu Kong 0001, Bineng Zhong 0001, Yun Fu 0001 |
CVPR | 3 |
| 2018 | Clustered Lifelong Learning Via Representative Task SelectionabstractConsider the lifelong machine learning problem where the objective is to learn new consecutive tasks depending on previously accumulated experiences, i.e., knowledge library. In comparison with most state-of-the-arts which adopt knowledge library with prescribed size, in this paper, we propose a new incremental clustered lifelong learning model with two libraries: feature library and model library, called Clustered Lifelong Learning (CL3), in which the feature library maintains a set of learned features common across all the encountered tasks, and the model library is learned by identifying and adding representative models (clusters). When a new task arrives, the original task model can be firstly reconstructed by representative models measured by capped l2-norm distance, i.e., effectively assigning the new task model to multiple representative models under feature library. Based on this assignment knowledge of new task, the objective of our CL3 model is to transfer the knowledge from both feature library and model library to learn the new task. The new task 1) with a higher outlier probability will then be judged as a new representative, and used to refine both feature library and representative models over time; 2) with lower outlier probability will only update the feature library. For the model optimisation, we cast this problem as an alternating direction minimization problem. To this end, the performance of CL3 is evaluated through comparing with most lifelong learning models, even some batch clustered multi-task learning models. Gan Sun, Yang Cong, Yu Kong 0001, Xiaowei Xu 0001 |
ICDM | 3 |
| 2018 | Hierarchical and Spatio-Temporal Sparse Representation for Human Action RecognitionabstractIn this paper, we present a novel two-layer video representation for human action recognition employing hierarchical group sparse encoding technique and spatio-temporal structure. In the first layer, a new sparse encoding method named locally consistent group sparse coding (LCGSC) is proposed to make full use of motion and appearance information of local features. LCGSC method not only encodes global layouts of features within the same video-level groups, but also captures local correlations between them, which obtains expressive sparse representations of video sequences. Meanwhile, two kinds of efficient location estimation models, namely an absolute location model and a relative location model, are developed to incorporate spatio-temporal structure into LCGSC representations. In the second layer, action-level group is established, where a hierarchical LCGSC encoding scheme is applied to describe videos at different levels of abstractions. On the one hand, the new layer captures higher order dependency between video sequences; on the other hand, it takes label information into consideration to improve discrimination of videos' representations. The superiorities of our hierarchical framework are demonstrated on several challenging datasets. Yu Kong 0001, Qiuqi Ruan, Gaoyun An, Yun Fu 0001 |
IEEE Trans. Image Process. | 2 |
| 2018 | Probabilistic Low-Rank Multitask LearningabstractIn this paper, we consider the problem of learning multiple related tasks simultaneously with the goal of improving the generalization performance of individual tasks. The key challenge is to effectively exploit the shared information across multiple tasks as well as preserve the discriminative information for each individual task. To address this, we propose a novel probabilistic model for multitask learning (MTL) that can automatically balance between low-rank and sparsity constraints. The former assumes a low-rank structure of the underlying predictive hypothesis space to explicitly capture the relationship of different tasks and the latter learns the incoherent sparse patterns private to each task. We derive and perform inference via variational Bayesian methods. Experimental results on both regression and classification tasks on real-world applications demonstrate the effectiveness of the proposed method in dealing with the MTL problems. Yu Kong 0001, Ming Shao, Yun Fu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2017 | Sparse Subspace Clustering by Learning Approximation ℓ0 CodesabstractSubspace clustering has been widely applied to detect meaningful clusters in high-dimensional data spaces. A main challenge in subspace clustering is to quickly calculate a "good" affinity matrix. ℓ0, ℓ1, ℓ2 or nuclear norm regularization is used to construct the affinity matrix in many subspace clustering methods because of their theoretical guarantees and empirical success. However, they suffer from the following problems: (1) ℓ2 and nuclear norm regularization require very strong assumptions to guarantee a subspace-preserving affinity; (2) although ℓ1 regularization can be guaranteed to give a subspace-preserving affinity under certain conditions, it needs more time to solve a large-scale convex optimization problem; (3) ℓ0 regularization can yield a tradeoff between computationally efficient and subspace-preserving affinity by using the orthogonal matching pursuit (OMP) algorithm, but this still takes more time to search the solution in OMP when the number of data points is large. In order to overcome these problems, we first propose a learned OMP (LOMP) algorithm to learn a single hidden neural network (SHNN) to fast approximate the ℓ0code. We then exploit a sparse subspace clustering method based on ℓ0 code which is fast computed by SHNN. Two sufficient conditions are presented to guarantee that our method can give a subspace-preserving affinity. Experiments on handwritten digit and face clustering show that our method not only quickly computes the ℓ0 code, but also outperforms the relevant subspace clustering methods in clustering results. In particular, our method achieves the state-of-the-art clustering accuracy (94.32%) on MNIST. Jun Li 0027, Yu Kong 0001, Yun Fu 0001 |
AAAI | 2 |
| 2017 | Deep Sequential Context Networks for Action PredictionabstractThis paper proposes efficient and powerful deep networks for action prediction from partially observed videos containing temporally incomplete action executions. Different from after-the-fact action recognition, action prediction task requires action labels to be predicted from these partially observed videos. Our approach exploits abundant sequential context information to enrich the feature representations of partial videos. We reconstruct missing information in the features extracted from partial videos by learning from fully observed action videos. The amount of the information is temporally ordered for the purpose of modeling temporal orderings of action segments. Label information is also used to better separate the learned features of different categories. We develop a new learning formulation that enables efficient model training. Extensive experimental results on UCF101, Sports-1M and BIT datasets demonstrate that our approach remarkably outperforms state-of-the-art methods, and is up to 300× faster than these methods. Results also show that actions differ in their prediction characteristics, some actions can be correctly predicted even though only the beginning 10% portion of videos is observed. Yu Kong 0001, Zhiqiang Tao, Yun Fu 0001 |
CVPR | 1 |
| 2017 | Multi-Stream Deep Similarity Learning Networks for Visual TrackingabstractVisual tracking has achieved remarkable success in recent decades, but it remains a challenging problem due to appearance variations over time and complex cluttered background. In this paper, we adopt a tracking-by-verification scheme to overcome these challenges by determining the patch in the subsequent frame that is most similar to the target template and distinctive to the background context. A multi-stream deep similarity learning network is proposed to learn the similarity comparison model. The loss function of our network encourages the distance between a positive patch in the search region and the target template to be smaller than that between positive patch and the background patches. Within the learned feature space, even if the distance between positive patches becomes large caused by the appearance change or interference of background clutter, our method can use the relative distance to distinguish the target robustly. Besides, the learned model is directly used for tracking with no need of model updating, parameter fine-tuning and can run at 45 fps on a single GPU. Our tracker achieves state-of-the-art performance on the visual tracking benchmark compared with other recent real-time-speed trackers, and shows better capability in handling background clutter, occlusion and appearance change. Yu Kong 0001, Yun Fu 0001 |
IJCAI | 2 |
| 2017 | Deep Active Learning Through Cognitive Information ParcelsabstractIn deep learning scenarios, a lot of labeled samples are needed to train the models. However, in practical application fields, since the objects to be recognized are complex and non-uniformly distributed, it is difficult to get enough labeled samples at one time. Active learning can actively improve the accuracy with fewer training labels, which is one of the promising solutions to tackle this problem. Inspired by human being's cognition process to acquire additional knowledge gradually, we propose a novel deep active learning method through Cognitive Information Parcels (CIPs) based on the analysis of model's cognitive errors and expert's instruction. The transformation of the cognitive parcels is defined, and the corresponding representation feature of the objects is obtained to identify the model's cognitive error information. Experiments prove that the samples, selected based on the CIPs, can benefit the target recognition and boost the deep model's performance efficiently. The characterization of cognitive knowledge can avoid the other samples' disturbance to the cognitive property of the model effectively. We believe that our work could provide a trial of thought about the cognitive knowledge used in deep learning field. Wencang Zhao, Yu Kong 0001, Zhengming Ding, Yun Fu 0001 |
ACM Multimedia | 2 |
| 2017 | Max-Margin Heterogeneous Information Machine for RGB-D Action Recognition
Yu Kong 0001, Yun Fu 0001 |
Int. J. Comput. Vis. | 1 |
| 2017 | Deeply Learned View-Invariant Features for Cross-View Action RecognitionabstractClassifying human actions from varied views is challenging due to huge data variations in different views. The key to this problem is to learn discriminative view-invariant features robust to view variations. In this paper, we address this problem by learning view-specific and view-shared features using novel deep models. View-specific features capture unique dynamics of each view while view-shared features encode common patterns across views. A novel sample-affinity matrix is introduced in learning shared features, which accurately balances information transfer within the samples from multiple views and limits the transfer across samples. This allows us to learn more discriminative shared features robust to view variations. In addition, the incoherence between the two types of features is encouraged to reduce information redundancy and exploit discriminative information in them separately. The discriminative power of the learned features is further improved by encouraging features in the same categories to be geometrically closer. Robust view-invariant features are finally learned by stacking several layers of features. Experimental results on three multi-view data sets show that our approaches outperform the state-of-the-art approaches. Yu Kong 0001, Zhengming Ding, Jun Li 0027, Yun Fu 0001 |
IEEE Trans. Image Process. | 1 |
| 2016 | Deep Convolutional Neural Network with Independent Softmax for Large Scale Face RecognitionabstractIn this paper, we present our solution to the MS-Celeb-1M Challenge. This challenge aims to recognize 100k celebrities at the same time. The huge number of celebrities is the bottleneck for training a deep convolutional neural network of which the output is equal to the number of celebrities. To solve this problem, an independent softmax model is proposed to split the single classifier into several small classifiers. Meanwhile, the training data are split into several partitions. This decomposes the large scale training procedure into several medium training procedures which can be solved separately. Besides, a large model is also trained and a simple strategy is introduced to merge the two models. Extensive experiments on the MSR-Celeb-1M dataset demonstrate the superiority of the proposed method. Our solution ranks the first and second in two tracks of the final evaluation. Yue Wu 0008, Jun Li 0027, Yu Kong 0001, Yun Fu 0001 |
ACM Multimedia | 3 |
| 2016 | Learning hierarchical 3D kernel descriptors for RGB-D action recognition
Yu Kong 0001, Behnam Satarboroujeni, Yun Fu 0001 |
Comput. Vis. Image Underst. | 1 |
| 2016 | Max-Margin Action Prediction MachineabstractThe speed with which intelligent systems can react to an action depends on how soon it can be recognized. The ability to recognize ongoing actions is critical in many applications, for example, spotting criminal activity. It is challenging, since decisions have to be made based on partial videos of temporally incomplete action executions. In this paper, we propose a novel discriminative multi-scale kernelized model for predicting the action class from a partially observed video. The proposed model captures temporal dynamics of human actions by explicitly considering all the history of observed features as well as features in smaller temporal segments. A compositional kernel is proposed to hierarchically capture the relationships between partial observations as well as the temporal segments, respectively. We develop a new learning formulation, which elegantly captures the temporal evolution over time, and enforces the label consistency between segments and corresponding partial videos. We prove that the proposed learning formulation minimizes the upper bound of the empirical risk. Experimental results on four public datasets show that the proposed approach outperforms state-of-the-art action prediction methods. Yu Kong 0001, Yun Fu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2016 | Efficient Image Geotagging Using Large DatabasesabstractThere are now billions of images stored on photo sharing websites. These images contain visual cues that reflect the geographical location of where the photograph was taken (e.g., New York City). Linking visual features in images to physical locations has many potential applications, such as tourism recommendation systems. However, the size and nature of these databases pose great challenges. For example, the distribution of images around the world is highly biased towards popular regions. This results in high redundancy in certain locations, while under-representing the features in other regions. Many density estimation methods are unable to handle such datasets. In this paper we employ an on-line unsupervised clustering method, Location Aware Self-Organizing Map (LASOM), to compress a large image database and learn similarity relationships between different geographical locations. Our method achieves promising results when used on a dataset containing approximately 900,000 images. We further show that the learned representation results in minimal information loss as compared to using k-Nearest Neighbor method. The noise reduction property of LASOM allows for superior performance when combining multiple features. The final part of the paper explores clothing as a new information source that may assist in geolocation of images. Dmitry Kit, Yu Kong 0001, Yun Fu 0001 |
IEEE Trans. Big Data | 2 |
| 2016 | Close Human Interaction Recognition Using Patch-Aware ModelsabstractThis paper addresses the problem of recognizing human interactions with close physical contact from videos. Due to ambiguities in feature-to-person assignments and frequent occlusions in close interactions, it is difficult to accurately extract the interacting people. This degrades the recognition performance. We, therefore, propose a hierarchical model, which recognizes close interactions and infers supporting regions for each interacting individual simultaneously. Our model associates a set of hidden variables with spatiotemporal patches and discriminatively infers their states, which indicate the person that the patches belong to. This patch-aware representation explicitly models and accounts for discriminative supporting regions for individuals, and thus overcomes the problem of ambiguities in feature assignments. Moreover, we incorporate the prior for the patches to deal with frequent occlusions during interactions. Using the discriminative supporting regions, our model builds cleaner features for individual action recognition and interaction recognition. Extensive experiments are performed on the BIT-Interaction data set and the UT-Interaction data set set #1 and set #2, and validate the effectiveness of our approach. Yu Kong 0001, Yun Fu 0001 |
IEEE Trans. Image Process. | 1 |
| 2016 | Discriminative Relational Representation Learning for RGB-D Action RecognitionabstractThis paper addresses the problem of recognizing human actions from RGB-D videos. A discriminative relational feature learning method is proposed for fusing heterogeneous RGB and depth modalities, and classifying the actions in RGB-D sequences. Our method factorizes the feature matrix of each modality, and enforces the same semantics for them in order to learn shared features from multimodal data. This allows us to capture the complex correlations between the two modalities. To improve the discriminative power of the relational features, we introduce a hinge loss to measure the classification accuracy when the features are employed for classification. This essentially performs supervised factorization, and learns discriminative features that are optimized for classification. We formulate the recognition task within a maximum margin framework, and solve the formulation using a coordinate descent algorithm. The proposed method is extensively evaluated on two public RGB-D action data sets. We demonstrate that the proposed method can learn extremely low-dimensional features with superior discriminative power, and outperforms the state-of-the-art methods. It also achieves high performance when one modality is missing in testing or training. Yu Kong 0001, Yun Fu 0001 |
IEEE Trans. Image Process. | 1 |
| 2016 | Learning Fast Low-Rank Projection for Image ClassificationabstractRooted in a basic hypothesis that a data matrix is strictly drawn from some independent subspaces, the low-rank representation (LRR) model and its variations have been successfully applied in various image classification tasks. However, this hypothesis is very strict to the LRR model as it cannot always be guaranteed in real images. Moreover, the hypothesis also prevents the sub-dictionaries of different subspaces from collaboratively representing an image. Fortunately, in supervised image classification, low-rank signal can be extracted from the independent label subspaces (ILS) instead of the independent image subspaces (IIS). Therefore, this paper proposes a projective low-rank representation (PLR) model by directly training a projective function to approximate the LRR derived from the labels. To the best of our knowledge, PLR is the first attempt to use the ILS hypothesis to relax the rigorous IIS hypothesis in the LRR models. We further prove a low-rank effect that the representations learned by PLR have high intraclass similarities and large interclass differences, which are beneficial to the classification tasks. The effectiveness of our proposed approach is validated by the experimental results on three databases. Jun Li 0027, Yu Kong 0001, Handong Zhao, Jian Yang 0003, Yun Fu 0001 |
IEEE Trans. Image Process. | 2 |
| 2015 | Bilinear heterogeneous information machine for RGB-D action recognitionabstractThis paper proposes a novel approach to action recognition from RGB-D cameras, in which depth features and RGB visual features are jointly used. Rich heterogeneous RGB and depth data are effectively compressed and projected to a learned shared space, in order to reduce noise and capture useful information for recognition. Knowledge from various sources can then be shared with others in the learned space to learn cross-modal features. This guides the discovery of valuable information for recognition. To capture complex spatiotemporal structural relationships in visual and depth features, we represent both RGB and depth data in a matrix form. We formulate the recognition task as a low-rank bilinear model composed of row and column parameter matrices. The rank of the model parameter is minimized to build a low-rank classifier, which is beneficial for improving the generalization power. The proposed method is extensively evaluated on two public RGB-D action datasets, and achieves state-of-the-art results. It also shows promising results if RGB or depth data are missing in training or testing procedure. Yu Kong 0001, Yun Fu 0001 |
CVPR | 1 |
| 2014 | A Discriminative Model with Multiple Temporal Scales for Action Prediction
Yu Kong 0001, Dmitry Kit, Yun Fu 0001 |
ECCV (5) | 1 |
| 2014 | LASOM: Location Aware Self-Organizing Map for discovering similar and unique visual features of geographical locationsabstractCan a machine tell us if an image was taken in Beijing or New York? Automated identification of the geographical coordinates based on image content is of particular importance to data mining systems, because geolocation provides a large source of context for other useful features of an image. However, successful localization of unannotated images requires a large collection of images that cover all possible locations. Brute-force searches over the entire databases are costly in terms of computation and storage requirements, and achieve limited results. Knowing what visual features make a particular location unique or similar to other locations can be used for choosing a better match between spatially distance locations. However, doing this at global scales is a challenging problem. In this paper we propose an on-line, unsupervised, clustering algorithm called Location Aware Self-Organizing Map (LASOM), for learning the similarity graph between different regions. The goal of LASOM is to select key features in specific locations so as to increase the accuracy in geotagging untagged images, while also reducing computational and storage requirements. Different from other Self-Organizing Map algorithms, LASOM provides the means to learn a conditional distribution of visual features, conditioned on geospatial coordinates. We demonstrate that the generated map not only preserves important visual information, but provides additional context in the form of visual similarity relationships between different geographical areas. We show how this information can be used to improve geotagging results when using large databases. Dmitry Kit, Yu Kong 0001, Yun Fu 0001 |
IJCNN | 2 |
| 2014 | Latent Tensor Transfer Learning for RGB-D Action RecognitionabstractThis paper proposes a method to compensate RGB-D images from the original target RGB images by transferring the depth knowledge of source data. Conventional RGB databases (e.g., UT-Interaction database) do not contain depth information since they are captured by the RGB cameras. Therefore, the methods designed for {RGB} databases cannot take advantage of depth information, which proves useful for simplifying intra-class variations and background subtraction. In this paper, we present a novel transfer learning method that can transfer the knowledge from depth information to the RGB database, and use the additional source information to recognize human actions in RGB videos. Our method takes full advantage of 3D geometric information contained within the learned depth data, thus, can further improve action recognition performance. We treat action data as a fourth-order tensor (row, column, frame and sample), and apply latent low-rank transfer learning to learn shared subspaces of the source and target databases. Moreover, we introduce a novel cross-modality regularizer that plays an important role in finding the correlation between RGB and depth modalities, and then more depth information from the source database can be transferred to that of the target. Our method is extensively evaluated on public by available databases. Results of two action datasets show that our method outperforms existing methods. Chengcheng Jia, Yu Kong 0001, Zhengming Ding, Yun Raymond Fu |
ACM Multimedia | 2 |
| 2014 | Learning a discriminative mid-level feature for action recognition
Cuiwei Liu, Mingtao Pei, Xinxiao Wu, Yu Kong 0001, Yunde Jia |
Sci. China Inf. Sci. | 4 |
| 2014 | Recognising human interaction from videos by a discriminative modelabstractThis study addresses the problem of recognising human interactions between two people. The main difficulties lie in the partial occlusion of body parts and the motion ambiguity in interactions. The authors observed that the interdependencies existing at both the action level and the body part level can greatly help disambiguate similar individual movements and facilitate human interaction recognition. Accordingly, they proposed a novel discriminative method, which model the action of each person by a large‐scale global feature and local body part features, to capture such interdependencies for recognising interaction of two people. A variant of multi‐class Adaboost method is proposed to automatically discover class‐specific discriminative three‐dimensional body parts. The proposed approach is tested on the authors newly introduced BIT‐interaction dataset and the UT‐interaction dataset. The results show that their proposed model is quite effective in recognising human interactions. Yu Kong 0001, Wei Liang 0008, Zhen Dong 0002, Yunde Jia |
IET Comput. Vis. | 1 |
| 2014 | Interactive Phrases: Semantic Descriptionsfor Human Interaction RecognitionabstractThis paper addresses the problem of recognizing human interactions from videos. We propose a novel approach that recognizes human interactions by the learned high-level descriptions, interactive phrases. Interactive phrases describe motion relationships between interacting people. These phrases naturally exploit human knowledge and allow us to construct a more descriptive model for recognizing human interactions. We propose a discriminative model to encode interactive phrases based on the latent SVM formulation. Interactive phrases are treated as latent variables and are used as mid-level features. To complement manually specified interactive phrases, we also discover data-driven phrases from data in order to find potentially useful and discriminative phrases for differentiating human interactions. An information-theoretic approach is employed to learn the data-driven phrases. The interdependencies between interactive phrases are explicitly captured in the model to deal with motion ambiguity and partial occlusion in the interactions. We evaluate our method on the BIT-Interaction data set, UT-Interaction data set, and Collective Activity data set. Experimental results show that our approach achieves superior performance over previous approaches. Yu Kong 0001, Yunde Jia, Yun Fu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2012 | Contour-HOG: A Stub Feature based Level Set Method for Learning Object Contour
Yu Kong 0001, Yun Fu 0001 |
BMVC | 2 |
| 2012 | Learning Human Interaction by Interactive Phrases
Yu Kong 0001, Yunde Jia, Yun Fu 0001 |
ECCV (1) | 1 |
| 2012 | A Hierarchical Model for Human Interaction RecognitionabstractRecognizing human interactions is a challenging task due to partially occluded body parts and motion ambiguities in interactions. We observe that the interdependencies existing at both action level and body part level greatly help disambiguate similar individual movements and facilitate human interaction recognition. In this paper, we propose a novel hierarchical model to capture such interdependencies for recognizing interactions of two persons. We model the action of each person by a large-scale global feature and several body part features. Two types of contextual information are exploited in our model to capture the implicit and complex interdependencies between interaction class, the action classes of two persons and the labels of persons' body parts. We build a challenging human interaction dataset to test our method. Results show that our model is quite effective in recognizing human interactions. Yu Kong 0001, Yunde Jia |
ICME | 1 |
| 2012 | Action recognition with discriminative mid-level features
Cuiwei Liu, Yu Kong 0001, Xinxiao Wu, Yunde Jia |
ICPR | 2 |
| 2012 | Decomposed contour prior for shape recognition
Yu Kong 0001, Yun Fu 0001 |
ICPR | 2 |
| 2011 | Adaptive learning codebook for action recognition
Yu Kong 0001, Xiaoqin Zhang 0002, Weiming Hu 0004, Yunde Jia |
Pattern Recognit. Lett. | 1 |
| 2010 | Compact visual codebook for action recognitionabstractVisual codebook has been popular in object classification as well as action analysis. However, its performance is often sensitive to the codebook size that is usually predefined. Moreover, the codebook generated by unsupervised methods, e.g., K-means, often suffers from the problem of ambiguity and weak efficiency. In other words, the visual codebook contains a lot of noisy and/or ambiguous words. In this paper, we propose a novel method to address these issues by constructing a compact but effective visual codebook using sparse reconstruction. Given a large codebook generated by K-means, we reformulate it in a sparse manner, and learn the weight of each word in the original visual codebook. Since the weights are sparse, they naturally introduce a new compact codebook. We apply this compact codebook to action recognition tasks and verify it on the widely used Weizmann action database. The experimental results show clearly the benefits of the proposed solution. Qingdi Wei, Xiaoqin Zhang 0002, Yu Kong 0001, Weiming Hu 0004, Haibin Ling |
ICIP | 3 |
| 2009 | Learning Group Activity in Soccer Videos from Local Motion
Yu Kong 0001, Weiming Hu 0004, Xiaoqin Zhang 0002, Hanzi Wang, Yunde Jia |
ACCV (1) | 1 |
| 2008 | Group action recognition in soccer videosabstractGroup action recognition in soccer videos is a challenging problem due to the difficulties of group action representation and camera motion estimation. This paper presents a novel approach for recognizing group action with a moving camera. In our approach, ego-motion is estimated by the Kanade-Lucas-Tomasi feature sets on successive frames. The optical flow is then computed on compensated frames. Due to the inaccurate ego-motion estimation, the optical flow can not reflect accurate motion of objects. In this paper, we propose a new motion descriptor which treats the optical flow as spatial patterns and extracts accurate global motion from the noisy optical flow. The latent-dynamic conditional random field model is employed to recognize group action. Experimental results show that our approach is promising. Yu Kong 0001, Xiaoqin Zhang 0002, Qingdi Wei, Weiming Hu 0004, Yunde Jia |
ICPR | 1 |