Kun-Yu Lin

dblp:194/9786 · DBLP profile ↗
← Back
42ranked-venue papers
7as first author
35since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 29 · 5 first-author · 25 since 2021Graphics, computer vision, multimedia, augmented reality and games · 21 · 1 first-author · 21 since 2021Databases, data management, data science and information retrieval · 6 · 2 first-authorSoftware engineering, systems software and programming languages · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 TechCoach: Towards Technical-Point-Aware Descriptive Action Coaching
abstract
To guide a learner in mastering action skills, it is crucial for a coach to 1) reason through the learner's action execution and technical points (TechPoints), and 2) provide detailed, comprehensible feedback on what is done well and what can be improved. However, existing score-based action assessment methods are still far from reaching this practical scenario. To bridge this gap, we investigate a new task termed Descriptive Action Coaching (DescCoach) which requires the model to provide detailed commentary on what is done well and what can be improved beyond a simple quality score for action execution. To this end, we first build a new dataset named EE4D-DescCoach. Through an automatic annotation pipeline, our dataset goes beyond the existing action assessment datasets by providing detailed TechPoint-level commentary. Furthermore, we propose TechCoach, a new framework that explicitly incorporates TechPoint-level reasoning into the DescCoach process. The central to our method lies in the Context-aware TechPoint Reasoner, which enables TechCoach to learn TechPoint-related quality representation by querying visual context under the supervision of TechPoint-level coaching commentary. By leveraging the visual context and the TechPoint-related quality representation, a unified TechPoint-aware Action Assessor is then employed to provide the overall coaching commentary together with the quality score. Combining all of these, we establish a new benchmark for DescCoach and evaluate the effectiveness of our method through extensive experiments.
Yuan-Ming Li, An-Lan Wang, Ling-An Zeng, Kun-Yu Lin, Yu-Ming Tang, Wei-Shi Zheng 0001
AAAI4
2026 Social Incentive Mechanism Based on Data Freshness and Reverse Auction Models for Mobile Crowdsensing
Chih-Lin Hu, Sheng-Min Yuan, Wu-Min Sung, Kun-Yu Lin, Carl K. Chang
COMPSAC4
2025 ParGo: Bridging Vision-Language with Partial and Global Views
abstract
This work presents ParGo, a novel Partial-Global projector designed to connect the vision and language modalities for Multimodal Large Language Models (MLLMs). Unlike previous works that rely on global attention-based projectors, our ParGo bridges the representation gap between the separately pre-trained vision encoders and the LLMs by integrating global and partial views, which alleviates the overemphasis on prominent regions. To facilitate the effective training of ParGo, we collect a large-scale detail-captioned image-text dataset named ParGoCap-1M-PT, consisting of 1 million images paired with high-quality captions. Extensive experiments on several MLLM benchmarks demonstrate the effectiveness of our ParGo, highlighting its superiority in aligning vision and language modalities. Compared to conventional Q-Former projector, our ParGo achieves an improvement of 259.96 in MME benchmark. Furthermore, our experiments reveal that ParGo significantly outperforms other projectors, particularly in tasks that emphasize detail perception ability.
An-Lan Wang, Bin Shan, Kun-Yu Lin, Guozhi Tang, Jingqun Tang, Wei-Shi Zheng 0001
AAAI4
2025 Modeling Multiple Normal Action Representations for Error Detection in Procedural Tasks
abstract
Error detection in procedural activities is essential for consistent and correct outcomes in AR-assisted and robotic systems. Existing methods often focus on temporal ordering errors or rely on static prototypes to represent normal actions. However, these approaches typically overlook the common scenario where multiple, distinct actions are valid following a given sequence of executed actions. This leads to two issues: (1) the model cannot effectively detect errors using static prototypes when the inference environment or action execution distribution differs from training; and (2) the model may also use the wrong prototypes to detect errors if the ongoing action label is not the same as the predicted one. To address this problem, we propose an Adaptive Multiple Normal Action Representation (AMNAR) framework. AMNAR predicts all valid next actions and reconstructs their corresponding normal action representations, which are compared against the ongoing action to detect errors. Extensive experiments demonstrate that AMNAR achieves state-of-the-art performance, highlighting the effectiveness of AMNAR and the importance of modeling multiple valid next actions in error detection. The code is available at https://github.com/iSEE-Laboratory/AMNAR.
Wei-Jin Huang, Yuan-Ming Li, Zhi-Wei Xia, Yu-Ming Tang, Kun-Yu Lin, Jianfang Hu, Wei-Shi Zheng 0001
CVPR5
2025 Person De-reidentification: A Variation-guided Identity Shift Modeling
abstract
Person re-identification (ReID) is to associate images of individuals from different camera views against cross-view variations. Like other surveillance technologies, Re-ID faces serious privacy challenges, particularly the potential for unauthorized tracking. Although various tasks (e.g., face recognition) have developed machine unlearning techniques to address privacy concerns, such methods have not yet been explored within the Re-ID field. In this work, we pioneer the exploration of the person de-reidentification (De-ReID) problem and present its inherent challenges. In the context of ReID, De-ReID is to unlearn the knowledge about accurately matching specific persons so that these "unlearned persons" cannot be re-identified across cameras for privacy guarantee. The primary challenge is to achieve the unlearning without degrading the identity-discriminative feature embeddings to ensure the model’s utility. To address this, we formulate a De-ReID framework that utilizes a labeled dataset of un-learned persons for unlearning and an unlabeled dataset of accessible persons for knowledge preservation. Instead of unlearning based on (pseudo) identity labels, we introduce a variation-guided identity shift mechanism that unlearns the specific persons by fitting the variations in their images while preserving ReID ability on other persons by overcoming the variations in images of accessible persons. As a result, the model shifts the unlearned persons to a feature space that is vulnerable to cross-view variations. Extensive experiments on benchmarks demonstrate the superiority of our method.
Yi-Xing Peng, Yu-Ming Tang, Kun-Yu Lin, Qize Yang, Jingke Meng, Xihan Wei, Wei-Shi Zheng 0001
CVPR3
2025 Mitigating the Human-Robot Domain Discrepancy in Visual Pre-training for Robotic Manipulation
abstract
Learning generalizable visual representations across different embodied environments is essential for effective robotic manipulation in real-world scenarios. However, the limited scale and diversity of robot demonstration data pose a significant challenge. Recent research has explored leveraging large-scale human activity data for pre-training, but the substantial morphological differences between humans and robots introduce a significant human-robot domain discrepancy, hindering the generalization of these models to downstream manipulation tasks. To overcome this, we propose a novel adaptation paradigm that leverages readily available paired human-robot video data to bridge the domain gap. Our method employs a human-robot contrastive alignment loss to align the semantics of human and robot videos, adapting pre-trained models to the robot domain in a parameter-efficient manner. Experiments on 20 simulated tasks across two different benchmarks and five real-world tasks demonstrate significant improvements. These results span both single-task and language-conditioned multi-task settings, evaluated using two different pre-trained models. Compared to existing pre-trained models, our adaptation method improves the average success rate by over 7% across multiple tasks on both simulated benchmarks and real-world evaluations. Project: https://jiaming-zhou.github.io/projects/HumanRobotAlign
Teli Ma, Kun-Yu Lin, Ronghe Qiu, Junwei Liang 0001
CVPR3
2025 Decoupled Distillation to Erase: A General Unlearning Method for Any Class-centric Tasks
abstract
In this work, we present DEcoupLEd Distillation To Erase (DELETE), a general and strong unlearning method for any class-centric tasks. To derive this, we first propose a theoretical framework to analyze the general form of unlearning loss and decompose it into forgetting and retention terms. Through the theoretical framework, we point out that a class of previous methods could be mainly formulated as a loss that implicitly optimizes the forgetting term while lacking supervision for the retention term, disturbing the distribution of pre-trained model and struggling to adequately preserve knowledge of the remaining classes. To address it, we refine the retention term using "dark knowledge" and propose a mask distillation unlearning method. By applying a mask to separate forgetting logits from retention logits, our approach optimizes both the forgetting and refined retention components simultaneously, retaining knowledge of the remaining classes while ensuring thorough forgetting of the target class. Without access to the remaining data or intervention (i.e., used in some works), we achieve state-of-the-art performance across various benchmarks. What’s more, DELETE is a general solution that can be applied to various downstream tasks, including face recognition, backdoor defense, and semantic segmentation with great performance.
Dian Zheng, Qijie Mo, Renjie Lu 0002, Kun-Yu Lin, Wei-Shi Zheng 0001
CVPR5
2025 ViSpeak: Visual Instruction Feedback in Streaming Videos
Shenghao Fu, Qize Yang, Yuan-Ming Li, Yi-Xing Peng, Kun-Yu Lin, Xihan Wei, Jianfang Hu, Xiaohua Xie, Wei-Shi Zheng 0001
ICCV5
2025 ReferDINO: Referring Video Object Segmentation with Visual Grounding Foundations
abstract
Referring video object segmentation (RVOS) aims to segment target objects throughout a video based on a text description. This is challenging as it involves deep vision-language understanding, pixel-level dense prediction and spatiotemporal reasoning. Despite notable progress in recent years, existing methods still exhibit a noticeable gap when considering all these aspects. In this work, we propose \textbf{ReferDINO}, a strong RVOS model that inherits region-level vision-language alignment from foundational visual grounding models, and is further endowed with pixel-level dense perception and cross-modal spatiotemporal reasoning. In detail, ReferDINO integrates two key components: 1) a grounding-guided deformable mask decoder that utilizes location prediction to progressively guide mask prediction through differentiable deformation mechanisms; 2) an object-consistent temporal enhancer that injects pretrained time-varying text features into inter-frame interaction to capture object-aware dynamic changes. Moreover, a confidence-aware query pruning strategy is designed to accelerate object decoding without compromising model performance. Extensive experimental results on five benchmarks demonstrate that our ReferDINO significantly outperforms previous methods (e.g., +3.9% (\mathcal{J}&\mathcal{F}) on Ref-YouTube-VOS) with real-time inference speed (51 FPS).
Tianming Liang, Kun-Yu Lin, Chaolei Tan, Jianguo Zhang 0001, Wei-Shi Zheng 0001, Jianfang Hu
ICCV2
2025 Less Static, More Private: Towards Transferable Privacy-Preserving Action Recognition by Generative Decoupled Learning
Zhi-Wei Xia, Kun-Yu Lin, Yuan-Ming Li, Wei-Jin Huang, Xian-Tuo Tan, Wei-Shi Zheng 0001
ICCV2
2025 Supplementary Material for "NoiseActor: A Noise-Action Collaborative Framework for Privacy-Preserving Action Recognition without Privacy Labels"
abstract
Our supplementary material is organized into the following sections: • Section II provides the details of the LOCATION module. • Section III provides the evaluation protocol for the SBU dataset. • Section IV provides the evaluation protocol for cross-dataset experiment on the UCF101 dataset and the VISPR dataset. • Section V provides implementation details for each datasets. • Section VI provides more anonymized frames on the SBU dataset. • Section VII provides more anonymized frames on the UCF101 dataset. • Section VIII provides details about anonymized video file. • Section IX provides ablation studies of the adapter.
Xiao Li 0074, Xiao-Ming Wu 0002, Delong Zhang, Kun-Yu Lin, Yi-Xing Peng, Ling-An Zeng, Wei-Shi Zheng 0001
ICME4
2025 Task-Oriented 6-DoF Grasp Pose Detection in Clutters
abstract
In general, humans would grasp an object differently for different tasks, e.g., “grasping the handle of a knife to cut” vs. “grasping the blade to hand over”. In the field of robotic grasp pose detection research, some existing works consider this task-oriented grasping and made some progress, but they are generally constrained by low-DoF gripper type or non-cluttered setting, which is not applicable for human assistance in real life. With an aim to get more general and practical grasp models, in this paper, we investigate the problem named Task-Oriented 6-DoF Grasp Pose Detection in Clutters (TO6DGC), which extends the task-oriented problem to a more general 6-DOF Grasp Pose Detection in Cluttered (multi-object) scenario. To this end, we construct a large-scale 6-DoF task-oriented grasping dataset, 6-DoF Task Grasp (6DTG), which features 4391 cluttered scenes with over 2 million 6-DoF grasp poses. Each grasp is annotated with a specific task, involving 6 tasks and 198 objects in total. Moreover, we propose One-Stage TaskGrasp (OSTG), a strong baseline to address the TO6DGC problem. Our OSTG adopts a task-oriented point selection strategy to detect where to grasp, and a task-oriented grasp generation module to decide how to grasp given a specific task. To evaluate the effectiveness of OSTG, extensive experiments are conducted on 6DTG. The results show that our method outperforms various baselines on multiple metrics. Real robot experiments also verify that our OSTG has a better perception of the task-oriented grasp points and 6-DoF grasp poses.
An-Lan Wang, Kun-Yu Lin, Yuan-Ming Li, Wei-Shi Zheng 0001
ICRA3
2025 Panoptic Captioning: An Equivalence Bridge for Image and Text
abstract
This work introduces panoptic captioning, a novel task striving to seek the minimum text equivalent of images, which has broad potential applications. We take the first step towards panoptic captioning by formulating it as a task of generating a comprehensive textual description for an image, which encapsulates all entities, their respective locations and attributes, relationships among entities, as well as global image state. Through an extensive evaluation, our work reveals that state-of-the-art Multi-modal Large Language Models (MLLMs) have limited performance in solving panoptic captioning. To address this, we propose an effective data engine named PancapEngine to produce high-quality data and a novel method named PancapChain to improve panoptic captioning. Specifically, our PancapEngine first detects diverse categories of entities in images by an elaborate detection suite, and then generates required panoptic captions using entity-aware prompts. Additionally, our PancapChain explicitly decouples the challenging panoptic captioning task into multiple stages and generates panoptic captions step by step. More importantly, we contribute a comprehensive metric named PancapScore and a human-curated test set for reliable model evaluation. Experiments show that our PancapChain-13B model can beat state-of-the-art open-source MLLMs like InternVL-2.5-78B and even surpass proprietary models like GPT-4o and Gemini-2.0-Pro, demonstrating the effectiveness of our data engine and method. Project page: https://visual-ai.github.io/pancap/
Kun-Yu Lin, Weining Ren, Kai Han 0001
NeurIPS1
2025 Exploring the Limits of Vision-Language-Action Manipulation in Cross-task Generalization
abstract
The generalization capabilities of vision-language-action (VLA) models to unseen tasks are crucial to achieving general-purpose robotic manipulation in open-world settings. However, the cross-task generalization capabilities of existing VLA models remain significantly underexplored. To address this gap, we introduce **AGNOSTOS**, a novel simulation benchmark designed to rigorously evaluate cross-task zero-shot generalization in manipulation. AGNOSTOS comprises 23 unseen manipulation tasks for test—distinct from common training task distributions—and incorporates two levels of generalization difficulty to assess robustness. Our systematic evaluation reveals that current VLA models, despite being trained on diverse datasets, struggle to generalize effectively to these unseen tasks. To overcome this limitation, we propose **Cross-Task In-Context Manipulation (X-ICM)**, a method that conditions large language models (LLMs) on in-context demonstrations from seen tasks to predict action sequences for unseen tasks. Additionally, we introduce a **dynamics-guided sample selection** strategy that identifies relevant demonstrations by capturing cross-task dynamics. On AGNOSTOS, X-ICM significantly improves cross-task zero-shot generalization performance over leading VLAs, achieving improvements of 6.0\% over $\pi_0$ and 7.9\% over VoxPoser. We believe AGNOSTOS and X-ICM will serve as valuable tools for advancing general-purpose robotic manipulation.
Ke Ye, Teli Ma, Ronghe Qiu, Kun-Yu Lin, Zhi-Lin Zhao 0001, Junwei Liang 0001
NeurIPS7
2025 Human-Centric Transformer for Domain Adaptive Action Recognition
abstract
We study the domain adaptation task for action recognition, namely domain adaptive action recognition, which aims to effectively transfer action recognition power from a label-sufficient source domain to a label-free target domain. Since actions are performed by humans, it is crucial to exploit human cues in videos when recognizing actions across domains. However, existing methods are prone to losing human cues but prefer to exploit the correlation between non-human contexts and associated actions for recognition, and the contexts of interest agnostic to actions would reduce recognition performance in the target domain. To overcome this problem, we focus on uncovering human-centric action cues for domain adaptive action recognition, and our conception is to investigate two aspects of human-centric action cues, namely human cues and human-context interaction cues. Accordingly, our proposed Human-Centric Transformer (HCTransformer) develops a decoupled human-centric learning paradigm to explicitly concentrate on human-centric action cues in domain-variant video feature learning. Our HCTransformer first conducts human-aware temporal modeling by a human encoder, aiming to avoid a loss of human cues during domain-invariant video feature learning. Then, by a Transformer-like architecture, HCTransformer exploits domain-invariant and action-correlated contexts by a context encoder, and further models domain-invariant interaction between humans and action-correlated contexts. We conduct extensive experiments on three benchmarks, namely UCF-HMDB, Kinetics-NecDrone and EPIC-Kitchens-UDA, and the state-of-the-art performance demonstrates the effectiveness of our proposed HCTransformer.
Kun-Yu Lin, Wei-Shi Zheng 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2025 Distilling the Unknown to Unveil Certainty
abstract
Out-of-distribution (OOD) detection is critical for identifying test samples that deviate from in-distribution (ID) data, ensuring network robustness and reliability. This paper presents a flexible framework for OOD knowledge distillation that extracts OOD-sensitive information from a network to develop a binary classifier capable of distinguishing between ID and OOD samples in both scenarios, with and without access to training ID data. To accomplish this, we introduce Confidence Amendment (CA), an innovative methodology that transforms an OOD sample into an ID one while progressively amending prediction confidence derived from the network to enhance OOD sensitivity. This approach enables the simultaneous synthesis of both ID and OOD samples, each accompanied by an adjusted prediction confidence, thereby facilitating the training of a binary classifier sensitive to OOD. Theoretical analysis provides bounds on the generalization error of the binary classifier, demonstrating the pivotal role of confidence amendment in enhancing OOD sensitivity. Extensive experiments spanning various datasets and network architectures confirm the efficacy of the proposed method in detecting OOD samples.
Zhi-Lin Zhao 0001, Longbing Cao, Yixuan Zhang 0006, Kun-Yu Lin, Wei-Shi Zheng 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2025 Weakly-Supervised Temporal Action Localization by Progressive Complementary Learning
abstract
Weakly-Supervised Temporal Action Localization (WSTAL) aims to localize and classify action instances in long untrimmed videos with only video-level category labels as supervision. A critical challenge of WSTAL is the large gap between video-level supervision and unavailable snippet-level supervision. Prevailing methods typically assign pseudo labels to snippets, but these methods suffer from significant noise caused by the pseudo snippet-level labels. In this work, we address the WSTAL from a novel category exclusion perspective, which gradually enhances the snippet-level supervision to bridge the gap. Our proposed Progressive Complementary Learning (ProCL) is inspired by the fact that, video-level labels precisely indicate the categories that all snippets surely do not belong to, which is ignored by previous works. Accordingly, we first exclude these surely non-existent categories by the deterministic complementary learning. And then, we introduce the entropy-based pseudo complementary learning that is able to exclude more categories for snippets of less ambiguity. Furthermore, for the remaining ambiguous snippets, we attempt to reduce the ambiguity by distinguishing foreground actions from the background. Extensive experimental results show that our method achieves new state-of-the-art performance on THUMOS14, ActivityNet1.3, and MultiTHUMOS benchmarks.
Jia-Run Du, Jia-Chang Feng, Kun-Yu Lin, Fa-Ting Hong, Zhongang Qi, Ying Shan, Jianfang Hu, Wei-Shi Zheng 0001
IEEE Trans. Circuits Syst. Video Technol.3
2024 Towards Completeness: A Generalizable Action Proposal Generator for Zero-Shot Temporal Action Localization
Jia-Run Du, Kun-Yu Lin, Jingke Meng, Wei-Shi Zheng 0001
ICPR (16)2
2024 Generalized Intra-Camera Supervised Person Re-Identification
abstract
Person re-identification (Re-ID) is to match the images of the same person from different camera views, which demands a view-invariant feature embedding. Recently, intra-camera supervised (ICS) Re-ID develops the Re-ID models without cross-view annotated data. Existing ICS methods are developed based on the assumptions, such as assuming each person in the training set appears under multiple cameras. However, there is no guarantee that the assumptions are true without cross-view annotations, and their performance degrades when the assumptions are violated. In this work, we generalize the ICS Re-ID and develop an ICS Re-ID model without the assumptions. The absence of prior assumptions and cross-view annotations poses a challenge in exploiting the discriminative information among cross-view images. To this end, we propose to mine the view-invariant relations between cross-view images for Re-ID model to exploit discriminative information and overcome the cross-view variations. Specifically, we learn composited view-aware features by compositing the identity information with different camera view information in the feature composition module. Then, we exploit the composited features to model various view-aware relations between pairwise images. By mining the common patterns among the view-aware relations, we obtain the view-invariant pairwise relation for learning. Besides, leveraging the composited view-aware features, we develop a view-aware marginal constraint for robust cross-view learning. To facilitate learning the feature composition module, we augment an auxiliary network to exploit the camera view information at the feature level. Extensive experimental results show the effectiveness of our method under different scenarios.
Yi-Xing Peng, Yu-Ming Tang, Kun-Yu Lin, Wei-Shi Zheng 0001
IEEE Trans. Circuits Syst. Video Technol.3
2024 TwinFormer: Fine-to-Coarse Temporal Modeling for Long-Term Action Recognition
abstract
The long-term action in untrimmed video generally contains multiple sub-actions, among which various semantic patterns exist (e.g., the co-occurrence or sequentiality between sub-actions). These semantic patterns are temporally coarse, and correlated with multiple local contexts which encode the local temporal evolution of visual elements (e.g., hands, objects) in videos. The local contexts and semantic patterns form the inherent fine-to-coarse temporal structure of long-term actions, which is neglected by existing works. Accordingly, in this work we propose TwinFormer, which exploits a novel fine-to-coarse temporal modeling manner to uncover the temporal structure of long-term actions. The proposed TwinFormer consists of a pair of twin encoders with the same structural design, namely Localcontext Encoder and Semantic-pattern Encoder, and a Temporalbridged Attention to bridge the two twin encoders. The Localcontext Encoder aims to model the local contexts in the longterm action. And the Temporal-bridged Attention is designed to correlate the local contexts with semantic patterns. Furthermore, the Semantic-pattern Encoder reveals the temporal evolution of semantic patterns. Experimental results on three benchmarks demonstrate the effectiveness of the proposed model.
Kun-Yu Lin, Yukun Qiu, Wei-Shi Zheng 0001
IEEE Trans. Multim.2
2024 Out-of-Distribution Detection by Cross-Class Vicinity Distribution of In-Distribution Data
abstract
Deep neural networks for image classification only learn to map in-distribution inputs to their corresponding ground-truth labels in training without differentiating out-of-distribution samples from in-distribution ones. This results from the assumption that all samples are independent and identically distributed (IID) without distributional distinction. Therefore, a pretrained network learned from in-distribution samples treats out-of-distribution samples as in-distribution and makes high-confidence predictions on them in the test phase. To address this issue, we draw out-of-distribution samples from the vicinity distribution of training in-distribution samples for learning to reject the prediction on out-of-distribution inputs. A cross-class vicinity distribution is introduced by assuming that an out-of-distribution sample generated by mixing multiple in-distribution samples does not share the same classes of its constituents. We, thus, improve the discriminability of a pretrained network by finetuning it with out-of-distribution samples drawn from the cross-class vicinity distribution, where each out-of-distribution input corresponds to a complementary label. Experiments on various in-/out-of-distribution datasets show that the proposed method significantly outperforms the existing methods in improving the capacity of discriminating between in- and out-of-distribution samples.
Zhi-Lin Zhao 0001, Longbing Cao, Kun-Yu Lin
IEEE Trans. Neural Networks Learn. Syst.3
2023 Predicting Road Traffic Risks with CNN-and-LSTM Learning Over Spatio-Temporal and Multi-Feature Traffic Data
abstract
Offering traffic safety information to drivers and passengers is one of essential services towards the smart city. Recent research utilizes AI models to analyze the collection of IoT-driven data in transportation environments. Exploring unveiled characteristics of traffic information to improve traffic control and accident prevention on roads, this way becomes plausible. Prior studies exploited various sorts of spatio-temporal traffic data to achieve the traffic prediction using deep learning models. Without understanding the complexity of spatio-temporal data, however, their efforts have not fully shown the effectiveness of deep learning-based traffic prediction and risk presentation. In this paper, our study first applies the Pearson correlation coefficient to clarify that traffic accidents appear in high correlation with time and space patterns. We identify multiple features from traffic domains, and employ CNN first and then LSTM learning techniques on several volumes of spatio-temporal traffic data, including weather, time, traffic flow, and historical traffic accidents and locations, etc. Our study shows that the combination of CNN and LSTM learning on spatio-temporal traffic data is applicable and useful for traffic risk prediction. Under experiments and demonstrations with actual traffic datasets, our proposed traffic risk prediction scheme, called CLwST, can exhibit more accurate results, faster convergence and lower loss in comparison with the two prior studies based on LSTM and ConvLSTM schemes.
Kun-Yu Lin, Pei-Yi Liu, Po-Kai Wang, Chih-Lin Hu, Ying Cai 0001
SSE1
2023 AsyFOD: An Asymmetric Adaptation Paradigm for Few-Shot Domain Adaptive Object Detection
abstract
In this work, we study few-shot domain adaptive object detection (FSDAOD), where only a few target labeled images are available for training in addition to sufficient source labeled images. Critically, in FSDAOD, the data scarcity in the target domain leads to an extreme data imbalance between the source and target domains, which potentially causes over-adaptation in traditional feature alignment. To address the data imbalance problem, we propose an asymmetric adaptation paradigm, namely AsyFOD, which leverages the source and target instances from different perspectives. Specifically, by using target distribution estimation, the AsyFOD first identifies the target-similar source instances, which serves to augment the limited target instances. Then, we conduct asynchronous alignment between target-dissimilar source instances and augmented target instances, which is simple yet effective for alleviating the over-adaptation. Extensive experiments demonstrate that the proposed AsyFOD outperforms all state-of-the-art methods on four FSDAOD benchmarks with various environmental variances, e.g., 3.1% mAP improvement on Cityscapes-to-FoggyCityscapes and 2.9% mAP increase on Sim10k-to-Cityscapes. The code is available at https://github.com/Hlings/AsyFPD.
Yipeng Gao, Kun-Yu Lin, Junkai Yan, Yaowei Wang 0001, Wei-Shi Zheng 0001
CVPR2
2023 Generating Anomalies for Video Anomaly Detection with Prompt-based Feature Mapping
abstract
Anomaly detection in surveillance videos is a challenging computer vision task where only normal videos are available during training. Recent work released the first virtual anomaly detection dataset to assist real-world detection. However, an anomaly gap exists because the anomalies are bounded in the virtual dataset but unbounded in the real world, so it reduces the generalization ability of the virtual dataset. There also exists a scene gap between virtual and real scenarios, including scene-specific anomalies (events that are abnormal in one scene but normal in another) and scene-specific attributes, such as the viewpoint of the surveillance camera. In this paper, we aim to solve the problem of the anomaly gap and scene gap by proposing a prompt-based feature mapping framework (PFMF). The PFMF contains a mapping network guided by an anomaly prompt to generate unseen anomalies with unbounded types in the real scenario, and a mapping adaptation branch to narrow the scene gap by applying domain classifier and anomaly classifier. The proposed framework outperforms the state-of-the-art on three benchmark datasets. Extensive ablation experiments also show the effectiveness of our framework design.
Zuhao Liu 0002, Xiao-Ming Wu 0002, Dian Zheng, Kun-Yu Lin, Wei-Shi Zheng 0001
CVPR4
2023 Event-Guided Procedure Planning from Instructional Videos with Text Supervision
abstract
In this work, we focus on the task of procedure planning from instructional videos with text supervision, where a model aims to predict an action sequence to transform the initial visual state into the goal visual state. A critical challenge of this task is the large semantic gap between observed visual states and unobserved intermediate actions, which is ignored by previous works. Specifically, this semantic gap refers to that the contents in the observed visual states are semantically different from the elements of some action text labels in a procedure. To bridge this semantic gap, we propose a novel event-guided paradigm, which first infers events from the observed states and then plans out actions based on both the states and predicted events. Our inspiration comes from that planning a procedure from an instructional video is to complete a specific event and a specific event usually involves specific actions. Based on the proposed paradigm, we contribute an Event-guided Prompting-based Procedure Planning (E3P) model, which encodes event information into the sequential modeling process to support procedure planning. To further consider the strong action associations within each event, our E3P adopts a mask-and-predict approach for relation mining, incorporating a probabilistic masking scheme for regularization. Extensive experiments on three datasets demonstrate the effectiveness of our proposed model.
An-Lan Wang, Kun-Yu Lin, Jia-Run Du, Jingke Meng, Wei-Shi Zheng 0001
ICCV2
2023 Diversifying Spatial-Temporal Perception for Video Domain Generalization
abstract
Video domain generalization aims to learn generalizable video classification models for unseen target domains by training in a source domain. A critical challenge of video domain generalization is to defend against the heavy reliance on domain-specific cues extracted from the source domain when recognizing target videos. To this end, we propose to perceive diverse spatial-temporal cues in videos, aiming to discover potential domain-invariant cues in addition to domain-specific cues. We contribute a novel model named Spatial-Temporal Diversification Network (STDN), which improves the diversity from both space and time dimensions of video data. First, our STDN proposes to discover various types of spatial cues within individual frames by spatial grouping. Then, our STDN proposes to explicitly model spatial-temporal dependencies between video contents at multiple space-time scales by spatial-temporal relation modeling. Extensive experiments on three benchmarks of different types demonstrate the effectiveness and versatility of our approach.
Kun-Yu Lin, Jia-Run Du, Yipeng Gao, Wei-Shi Zheng 0001
NeurIPS1
2023 Revealing the Distributional Vulnerability of Discriminators by Implicit Generators
abstract
In deep neural learning, a discriminator trained on in-distribution (ID) samples may make high-confidence predictions on out-of-distribution (OOD) samples. This triggers a significant matter for robust, trustworthy and safe deep learning. The issue is primarily caused by the limited ID samples observable in training the discriminator when OOD samples are unavailable. We propose a general approach for fine-tuning discriminators by implicit generators (FIG). FIG is grounded on information theory and applicable to standard discriminators without retraining. It improves the ability of a standard discriminator in distinguishing ID and OOD samples by generating and penalizing its specific OOD samples. According to the Shannon entropy, an energy-based implicit generator is inferred from a discriminator without extra training costs. Then, a Langevin dynamic sampler draws specific OOD samples for the implicit generator. Lastly, we design a regularizer fitting the design principle of the implicit generator to induce high entropy on those generated OOD samples. The experiments on different networks and datasets demonstrate that FIG achieves the state-of-the-art OOD detection performance.
Zhi-Lin Zhao 0001, Longbing Cao, Kun-Yu Lin
IEEE Trans. Pattern Anal. Mach. Intell.3
2023 Supervision Adaptation Balancing In-Distribution Generalization and Out-of-Distribution Detection
abstract
The discrepancy between in-distribution (ID) and out-of-distribution (OOD) samples can lead to distributional vulnerability in deep neural networks, which can subsequently lead to high-confidence predictions for OOD samples. This is mainly due to the absence of OOD samples during training, which fails to constrain the network properly. To tackle this issue, several state-of-the-art methods include adding extra OOD samples to training and assign them with manually-defined labels. However, this practice can introduce unreliable labeling, negatively affecting ID classification. The distributional vulnerability presents a critical challenge for non-IID deep learning, which aims for OOD-tolerant ID classification by balancing ID generalization and OOD detection. In this paper, we introduce a novel supervision adaptation approach to generate adaptive supervision information for OOD samples, making them more compatible with ID samples. First, we measure the dependency between ID samples and their labels using mutual information, revealing that the supervision information can be represented in terms of negative probabilities across all classes. Second, we investigate data correlations between ID and OOD samples by solving a series of binary regression problems, with the goal of refining the supervision information for more distinctly separable ID classes. Our extensive experiments on four advanced network architectures, two ID datasets, and eleven diversified OOD datasets demonstrate the efficacy of our supervision adaptation approach in improving both ID classification and OOD detection capabilities.
Zhi-Lin Zhao 0001, Longbing Cao, Kun-Yu Lin
IEEE Trans. Pattern Anal. Mach. Intell.3
2023 Consistent Intra-Video Contrastive Learning With Asynchronous Long-Term Memory Bank
abstract
Unsupervised representation learning for videos has recently achieved remarkable performance owing to the effectiveness of contrastive learning. Most works on video contrastive learning (VCL) pull all snippets from the same video into the same category, even if some of them are from different actions, leading to temporal collapse, i.e., the snippet representations of a video are invariable with the evolution of time. In this paper, we introduce a novel intra-video contrastive learning (intra-VCL) that further distinguishes intra-video actions to alleviate this issue, which includes an asynchronous long-term memory bank (that caches the representations of all snippets of each video) and mines an extra positive/negative snippet within a video based on the asynchronous long-term memory bank. In addition, since an asynchronous long-term memory bank is required for performing intra-VCL and asynchronous update of the long-term memory leads to inconsistencies when performing contrastive learning, we further propose a consistent contrastive module (CCM) to perform consistent intra-VCL. Specifically, in the CCM, we propose an intra-video self-attention refinement function to reduce the inconsistencies within the asynchronously updated representations (of all snippets of each video) in the long-term memory and an adaptive loss re-weighting to reduce unreliable self-supervision produced by inconsistent contrastive pairs. We call our method as consistent intra-VCL. Extensive experiments demonstrate the effectiveness of the proposed consistent intra-VCL, which achieves state-of-the-art performance on the standard benchmarks of self-supervised action recognition, with top-1 accuracies of 64.2% and 91.0% on HMDB-51 and UCF-101, respectively.
Zelin Chen, Kun-Yu Lin, Wei-Shi Zheng 0001
IEEE Trans. Circuits Syst. Video Technol.2
2023 DilateFormer: Multi-Scale Dilated Transformer for Visual Recognition
abstract
As ade factosolution, the vanilla Vision Transformers (ViTs) are encouraged to model long-range dependencies between arbitrary image patches while the global attended receptive field leads to quadratic computational cost. Another branch of Vision Transformers exploits local attention inspired by CNNs, which only models the interactions between patches in small neighborhoods. Although such a solution reduces the computational cost, it naturally suffers from small attended receptive fields, which may limit the performance. In this work, we explore effective Vision Transformers to pursue a preferable trade-off between the computational complexity and size of the attended receptive field. By analyzing the patch interaction of global attention in ViTs, we observe two key properties in the shallow layers, namely locality and sparsity, indicating the redundancy of global dependency modeling in shallow layers of ViTs. Accordingly, we propose Multi-Scale Dilated Attention (MSDA) to modellocalandsparsepatch interaction within the sliding window. With a pyramid architecture, we construct a Multi-Scale Dilated Transformer (DilateFormer) by stacking MSDA blocks at low-level stages and global multi-head self-attention blocks at high-level stages. Our experiment results show that our DilateFormer achieves state-of-the-art performance on various vision tasks. On ImageNet-1 K classification task, DilateFormer achieves comparable performance with 70% fewer FLOPs compared with existing state-of-the-art models. Our DilateFormer-Base achieves 85.6% top-1 accuracy on ImageNet-1 K classification task, 53.5% box mAP/46.1% mask mAP on COCO object detection/instance segmentation task and 51.1% MS mIoU on ADE20 K semantic segmentation task.
Jiayu Jiao, Yu-Ming Tang, Kun-Yu Lin, Yipeng Gao, Andy Jinhua Ma, Yaowei Wang 0001, Wei-Shi Zheng 0001
IEEE Trans. Multim.3
2023 Incentive Mechanism for Mobile Crowdsensing With Two-Stage Stackelberg Game
abstract
Mobile crowdsensing technologies augment the collective effort on exploiting data from a large crowd of mobile users in ubiquitous environments. When mobile users partake in executing crowdsensing tasks, they can receive rewards and be incentified to stay in virtual teamwork. This paper proposes a game-based incentive mechanism, named Incentive-G, aiming at recruiting mobile users effectively and improving the reliability and quality of sensing data against untrusty or malicious users. The Incentive-G mechanism consists of several design phases, including analyzing sensing data, determining reputations of mobile users, and ensuring data quality and reliability by voting in a task group. This mechanism adopts a two-stage Stackelberg game for analyzing reciprocal relationship between service providers and mobile users, and then optimizes incentive benefits using backward induction. Our analysis shows that the existence and uniqueness of the Stackelberg equilibrium can be validated by identifying the best data-provision strategies for mobile users. In addition, the maximum revenue strategy for a service provider can be found by gathering a sufficient amount of high-quality data from mobile users. Performance results manifest that the Incentive-G mechanism is able to significantly encourage mobile users to contribute their efforts and maximize the revenue for game-based crowdsensing services.
Chih-Lin Hu, Kun-Yu Lin, Carl K. Chang
IEEE Trans. Serv. Comput.2
2022 Adversarial Partial Domain Adaptation by Cycle Inconsistency
Kun-Yu Lin, Yukun Qiu, Wei-Shi Zheng 0001
ECCV (33)1
2022 Intelligent task migration with deep Qlearning in multi-access edge computing
abstract
Abstract Multi‐access edge computing provides computation and network resources in proximity to user applications in mobile environments. Deploying edge servers in network boundary can not only offload the heavy task loading on the cloud, but also alleviate resource‐limited capabilities of mobile devices. Rather than many stand‐alone edge servers, the concept of multi‐server edge computing is recently advocated to contend with the issues of system scalability and service quality against dynamic task workload. This study exploits collaborative computing resources and designs a task migration strategy for multiple edge servers in mobile networks. This study formulates a queueing optimization problem of minimizing the overall service time in a multi‐server system. An intelligent task migration scheme is then developed using the deep reinforcement learning and Q‐learning techniques. With a variety of numerical attributes derived from the queueing model, this intelligent scheme can arrange the task distribution among edge servers to enhance the task processing capability. Simulation‐based results show that the proposed task migration scheme can sustain service efficiency and resource utilization, which is promising as compared with conventional designs without collaborative intelligence in mobile environments.
Sheng-Zhi Huang, Kun-Yu Lin, Chih-Lin Hu
IET Commun.2
2022 Node Pair Information Preserving Network Embedding Based on Adversarial Networks
abstract
Network embedding aims to learn the low-dimensional node representations for networks, which has attracted an increasing amount of attention in recent years. Most existing efforts in this field attempt to embed the network based on node similarity, which generally relies on edge existence statistics of the network. Instead of relying on the global edge existence statistics for every node pair, in this article, we utilize the information between a pair of nodes in a local way and propose a model, called node pair information preserving network embedding (NINE), based on adversarial networks. The main idea lies in preserving the node pair information (NI) by means of adversarial networks. The architecture of the proposed NINE model consists of three main components, namely: 1) NI embedder; 2) NI generator; and 3) NI discriminator. In the NI embedder, to avoid the complicated similarity calculation for a pair of nodes, the original NI vector calculated from the direct neighbor information of the two nodes is adopted as features, and the edge existence information is taken as labels to learn the embedded NI vector in a supervised learning manner. The second component is the NI generator, which takes the original node representation vectors of a node pair as input and outputs the generated NI vector. In order to make the generated NI vector follow the same distribution of the corresponding embedded NI vector, the generative adversarial network (GAN) is adopted, resulting in the third component, called the NI discriminator. Extensive experiments are conducted on seven real-world datasets in three downstream tasks, namely: 1) network reconstruction; 2) link prediction; and 3) node classification. Comparison results with seven state-of-the-art models demonstrate the effectiveness, efficiency, and rationality of our model.
Chang-Dong Wang 0001, Ling Huang 0002, Kun-Yu Lin, Dong Huang 0001, Philip S. Yu
IEEE Trans. Cybern.4
2021 Graph-Based High-Order Relation Modeling for Long-Term Action Recognition
abstract
Long-term actions involve many important visual concepts, e.g., objects, motions, and sub-actions, and there are various relations among these concepts, which we call basic relations. These basic relations will jointly affect each other during the temporal evolution of long-term actions, which forms the high-order relations that are essential for long-term action recognition. In this paper, we propose a Graph-based High-order Relation Modeling (GHRM) module to exploit the high-order relations in the long-term actions for long-term action recognition. In GHRM, each basic relation in the long-term actions will be modeled by a graph, where each node represents a segment in a long video. Moreover, when modeling each basic relation, the information from all the other basic relations will be incorporated by GHRM, and thus the high-order relations in the long-term actions can be well exploited. To better exploit the high-order relations along the time dimension, we design a GHRM-layer consisting of a Temporal-GHRM branch and a Semantic-GHRM branch, which aims to model the local temporal high-order relations and global semantic high-order relations. The experimental results on three long-term action recognition datasets, namely, Breakfast, Charades, and MultiThumos, demonstrate the effectiveness of our model.
Kun-Yu Lin, Haoxin Li, Wei-Shi Zheng 0001
CVPR2
2019 Multi-view Outlier Detection in Deep Intact Space
abstract
Recently, multi-view outlier detection has emerged as a challenging research topic in outlier detection because of complex distributions of data across different views. There are mainly three types of outliers, i.e., attribute outliers, class outliers and class-attribute outliers. Most existing multi-view outlier detection approaches only detect part of the three types of outliers in a pairwise manner across different views, which is not able to accomplish the task of multi-view outlier detection comprehensively and uniformly. Outlier detection in a pairwise manner across different views also leads to time-consuming computation. We propose a new algorithm termed Multi-view Outlier Detection in Deep Intact Space (MODDIS) to find all the three types of outliers simultaneously and avoid comparing different views in a pairwise manner. Rather than leveraging subspace clustering, the performance of which is seriously affected by the dependence of subspaces on most real datasets, neural networks are employed in MODDIS in that neural networks have a stronger representation learning ability. Meanwhile, based on the view insufficiency assumption, a multi-view intact outlierness space assumption is proposed. Based on this assumption, a multi-view latent intact space is constructed to encode outlierness information of all views, where outlierness in any view is a snapshot from some perspective. Finally, an outlier detection measurement is defined in the latent intact space. Experiments are conducted on several UCI datasets and the empirical results demonstrate the effectiveness of our proposed method.
Yu-Xuan Ji, Ling Huang 0002, Heng-Ping He, Chang-Dong Wang 0001, Guangqiang Xie, Kun-Yu Lin
ICDM7
2019 Direction recovery in undirected social networks based on community structure and popularity
Yi-Ming Wen, Ling Huang 0002, Chang-Dong Wang 0001, Kun-Yu Lin
Inf. Sci.4
2018 Multi-view Proximity Learning for Clustering
Kun-Yu Lin, Ling Huang 0002, Chang-Dong Wang 0001, Hongyang Chao
DASFAA (2)1
2018 Direction Recovery in Undirected Social Networks Based on Community Structure and Popularity
Yi-Ming Wen, Chang-Dong Wang 0001, Kun-Yu Lin
DASFAA (1)3
2017 Missing Value Learning
abstract
Missing value is common in many machine learning problems and much effort has been made to handle missing data to improve the performance of the learned model. Sometimes, our task is not to train a model using those unlabeled/labeled data with missing value but process examples according to the values of some specified features. So, there is an urgent need of developing a method to predict those missing values. In this paper, we focus on learning from the known values to learn missing value as close as possible to the true one. It's difficult for us to predict missing value because we do not know the structure of the data matrix and some missing values may relate to some other missing values. We solve the problem by recovering the complete data matrix under the three reasonable constraints: feature relationship, upper recovery error bound and class relationship. The proposed algorithm can deal with both unlabeled and labeled data and generative adversarial idea will be used in labeled data to transfer knowledge. Extensive experiments have been conducted to show the effectiveness of the proposed algorithms.
Zhi-Lin Zhao 0001, Chang-Dong Wang 0001, Kun-Yu Lin, Jian-Huang Lai
CIKM3
2017 Multi-view Unit Intact Space Learning
Kun-Yu Lin, Chang-Dong Wang 0001, Yu-Qin Meng, Zhi-Lin Zhao 0001
KSEM1
2017 CCMS: A nonlinear clustering method based on crowd movement and selection
Kun-Yu Lin, Chang-Dong Wang 0001, Jian-Bo Liu, Dong Huang 0001
Neurocomputing2