Dit-Yan Yeung

dblp:41/5668 · DBLP profile ↗
← Back
203ranked-venue papers
14as first author
53since 2021 · last 2026
0000-0003-3716-8125ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 170 · 13 first-author · 40 since 2021Graphics, computer vision, multimedia, augmented reality and games · 82 · 2 first-author · 26 since 2021Databases, data management, data science and information retrieval · 26 · 2 first-author · 3 since 2021Systems, architecture and hardware · 5 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 since 2021Computer networks · 3 · 1 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 CoherenDream: Boosting Holistic Text Coherence in 3D Generation via Multimodal Large Language Models Feedback
abstract
Score Distillation Sampling (SDS) has achieved remarkable success in text-to-3D content generation. However, SDS-based methods struggle to maintain semantic fidelity for user prompts, particularly when involving multiple objects with intricate interactions. While existing approaches often address 3D consistency through multiview diffusion model fine-tuning on 3D datasets, this strategy inadvertently exacerbates text-3D alignment degradation. The limitation stems from SDS's inherent accumulation of view-independent biases during optimization, which progressively diverges from the ideal text alignment direction. To alleviate this limitation, we propose a novel SDS objective, dubbed as Textual Coherent Score Distillation (TCSD), which integrates alignment feedback from multimodal large language models (MLLMs). Our TCSD leverages cross-modal understanding capabilities of MLLMs to assess and guide the text-3D correspondence during the optimization. We further develop 3DLLaVA-CRITIC - a fine-tuned MLLM specialized for evaluating multiview text alignment in 3D generations. Additionally, we introduce an LLM-layout initialization that significantly accelerates optimization convergence through semantic-aware spatial configuration. Our framework, CoherenDream, achieves consistent improvement across multiple metrics on TIFA subset.As the first study to incorporate MLLMs into SDS optimization, we also conduct extensive ablation studies to explore optimal MLLM adaptations for 3D generation tasks.
Chenhan Jiang, Yihan Zeng, Dit-Yan Yeung
AAAI3
2026 An online forecasting-based fine-tuning pipeline for time-series anomaly prediction
abstract
Time-series anomaly detection is critical for numerous real-world applications and has been extensively studied. However, existing methods are typically designed to identify anomalies within a complete time series. In other words, they rely on access to ground truth data to calculate anomaly scores and distinguish anomalous data from normal patterns. This reliance limits their applicability in scenarios where predicting future anomalies is required, as the ground truth is inherently unavailable. To address this gap, we introduce the concept of Time-Series Anomaly Prediction (TSAP), which focuses on forecasting the occurrence and progression of anomalies in time series simultaneously without relying on ground truth. In this paper, we propose a novel exemplar-based pre-training and fine-tuning pipeline tailored to this task, based on recent achievements in online time-series forecasting techniques. The pipeline begins with an offline pre-training phase, where a deep learning model is trained to capture the underlying temporal correlations in time-series data. During the online fine-tuning stage, a three-step process is employed to predict the timing and evolution of anomalies. This process includes prediction and anomaly detection, motif search for similar patterns, and fine-tuning using exemplars. These steps are repeated as new data arrives. We evaluate the proposed method against state-of-the-art approaches from various relevant categories on both real-world and synthetic datasets. Experimental results show that the proposed method improves anomaly detection accuracy by up to 53.8% in terms of F1 score and enhances time-series forecasting accuracy during and after anomaly periods by up to 82.4% and 49.1% in terms of MSE. Through analysing the results, we prove the proposed method's effectiveness in addressing the new TSAP tasks, which are incapable of being handled by current time-series anomaly detection or online time-series forecasting methods.
Zhou Zhou 0003, Van-Hoan Trinh, Yuet Ming Joyce Yue, Dit-Yan Yeung, Ka-Hing Wong, Wai-Kin Wong
Neural Networks4
2026 Constrained and directional ensemble attention for facial action unit detection
Zhiwen Shao, Bikuan Chen, Yong Zhou 0003, Xuehuai Shi, Canlin Li, Lizhuang Ma, Dit-Yan Yeung
Pattern Recognit.7
2026 Mixture of Cluster-Conditional LoRA Experts for Vision-Language Instruction Tuning
abstract
Instruction tuning of Large Vision-language Models (LVLMs) has revolutionized the development of versatile models with zero-shot generalization across a wide range of downstream vision-language tasks. However, the diversity of different training tasks from various sources and formats would lead to inevitable task conflicts, where different tasks conflict for the same set of model parameters, resulting in sub-optimal instruction-following abilities. To address that, we propose the Mixture of Cluster-conditional LoRA Experts (MoCLE), a novel Mixture of Experts (MoE) architecture designed to activate task-customized model parameters based on instruction clusters. A separate universal expert is further incorporated to improve generalization abilities of MoCLE for novel instructions. Extensive experiments on InstructBLIP and LLaVA demonstrate the effectiveness of MoCLE.
Yunhao Gou, Zhili Liu, Kai Chen 0023, Lanqing Hong, Hang Xu 0004, Zhenguo Li, Dit-Yan Yeung, James T. Kwok, Yu Zhang 0006
IEEE Trans. Image Process.7
2026 TextRSR: Enhanced Arbitrary-Shaped Scene Text Representation via Robust Subspace Recovery
abstract
In recent years, scene text detection research has increasingly focused on arbitrary-shaped texts, where text representation is a fundamental problem. However, most existing methods still struggle to separate adjacent or overlapping texts due to ambiguous spatial positions of points or segmentation masks. Besides, the time efficiency of the entire pipeline is often neglected, resulting in sub-optimal inference speed. To tackle these problems, we first propose a novel text representation method based on robust subspace recovery, which robustly represents complex text shapes by combining orthogonal basis vectors learned from labeled text contours. These basis vectors capture basis contour patterns with distinct information, enabling clearer boundaries even in densely populated text scenarios. Moreover, we propose a dynamic sparse assignment scheme for positive samples that adaptively adjusts their weights during training, which not only accelerates inference speed by eliminating redundant predictions but also enhances feature learning by providing sufficient supervision signals. Building on these innovations, we present TextRSR, an accurate and efficient scene text detection network. Extensive experiments on challenging benchmarks demonstrate the superior accuracy and efficiency of TextRSR compared to state-of-the-art methods. Particularly, TextRSR achieves an F-measure of 88.5% at 37.8 frames per second (FPS) for CTW1500 dataset and an F-measure of 89.1% at 23.1 FPS for Total-Text dataset.
Zhiwen Shao, Shengtian Jiang, Hancheng Zhu, Xuehuai Shi, Canlin Li, Lizhuang Ma, Dit-Yan Yeung
IEEE Trans. Multim.7
2025 G-VEval: A Versatile Metric for Evaluating Image and Video Captions Using GPT-4o
abstract
Evaluation metric of visual captioning is important yet not thoroughly explored. Traditional metrics like BLEU, METEOR, CIDEr, and ROUGE often miss semantic depth, while trained metrics such as CLIP-Score, PAC-S, and Polos are limited in zero-shot scenarios. Advanced Language Model-based metrics also struggle with aligning to nuanced human preferences. To address these issues, we introduce G-VEval, a novel metric inspired by G-Eval and powered by the new GPT-4o. G-VEval uses chain-of-thought reasoning in large multimodal models and supports three modes: reference-free, reference-only, and combined, accommodating both video and image inputs. We also propose MSVD-Eval, a new dataset for video captioning evaluation, to establish a more transparent and consistent framework for both human experts and evaluation metrics. It is designed to address the lack of clear criteria in existing datasets by introducing distinct dimensions of Accuracy, Completeness, Conciseness, and Relevance (ACCR). Extensive results show that G-VEval outperforms existing methods in correlation with human annotations, as measured by Kendall tau-b and Kendall tau-c. This provides a flexible solution for diverse captioning tasks and suggests a straightforward yet effective approach for large language models to understand video content, paving the way for advancements in automated captioning.
Tony Cheng Tong, Zhiwen Shao, Dit-Yan Yeung
AAAI4
2025 Ref-Long: Benchmarking the Long-context Referencing Capability of Long-context Language Models
abstract
Long-context language models (LCLMs) have exhibited impressive capabilities in longcontext understanding tasks.Among these, long-context referencing-a crucial task that requires LCLMs to attribute items of interest to specific parts of long-context data-remains underexplored.To bridge this gap, this paper proposes Referencing Evaluation for Longcontext Language Models (Ref-Long), a novel benchmark designed to assess the long-context referencing capability of LCLMs.Specifically, Ref-Long requires LCLMs to identify the indexes of documents that reference a specific key, emphasizing contextual relationships between the key and the documents over simple retrieval.Based on the task design, we construct three subsets ranging from synthetic to realistic scenarios to form the Ref-Long benchmark.Experimental results of 13 LCLMs reveal significant shortcomings in long-context referencing, even among advanced models like GPT-4o.To further investigate these challenges, we conduct comprehensive analyses, including human evaluations, task format adjustments, fine-tuning experiments, and error analyses, leading to several key insights.Our data and code can be found in https://github. com
Junjie Wu 0007, Gefei Gu, Yanan Zheng, Dit-Yan Yeung, Arman Cohan
ACL (1)4
2025 EMOVA: Empowering Language Models to See, Hear and Speak with Vivid Emotions
abstract
GPT-4o, an omni-modal model that enables vocal conversations with diverse emotions and tones, marks a milestone for omni-modal foundation models. However, empowering Large Language Models to perceive and generate images, texts, and speeches end-to-end with publicly available data remains challenging for the open-source community. Existing vision-language models rely on external tools for speech processing, while speech-language models still suffer from limited or totally without vision-understanding capabilities. To address this gap, we propose the EMOVA (EMotionally Omni-present Voice Assistant), to enable Large Language Models with end-to-end speech abilities while maintaining the leading vision-language performance. With a semantic-acoustic disentangled speech tokenizer, we surprisingly notice that omni-modal alignment can further enhance vision-language and speech abilities compared with the bi-modal aligned counterparts. Moreover, a lightweight style module is introduced for the flexible speech style controls including emotions and pitches. For the first time, EMOVA achieves state-of-the-art performance on both the vision-language and speech benchmarks, and meanwhile, supporting omni-modal spoken dialogue with vivid emotions.
Yunhao Gou, Runhui Huang, Zhili Liu, Daxin Tan, Chunwei Wang, Yihan Zeng, Dingdong Wang, Kun Xiang, Haoli Bai, Jianhua Han, Weike Jin, Nian Xie, James T. Kwok, Hengshuang Zhao, Xiaodan Liang, Dit-Yan Yeung, Zhenguo Li, Qun Liu 0001, Lanqing Hong, Lu Hou 0002
CVPR23
2025 Anyattack: Towards Large-scale Self-supervised Adversarial Attacks on Vision-language Models
abstract
Due to their multimodal capabilities, Vision-Language Models (VLMs) have found numerous impactful applications in real-world scenarios. However, recent studies have revealed that VLMs are vulnerable to image-based adversarial attacks. Traditional targeted adversarial attacks require specific targets and labels, limiting their real-world impact. We present AnyAttack, a self-supervised framework that transcends the limitations of conventional attacks through a novel foundation model approach. By pretraining on the massive LAION-400M dataset without label supervision, AnyAttack achieves unprecedented flexibility - enabling any image to be transformed into an attack vector targeting any desired output across different VLMs. This approach fundamentally changes the threat landscape, making adversarial capabilities accessible at an unprecedented scale. Our extensive validation across five open-source VLMs (CLIP, BLIP, BLIP2, InstructBLIP, and MiniGPT-4) demonstrates AnyAttack’s effectiveness across diverse multimodal tasks. Most concerning, Any-Attack seamlessly transfers to commercial systems including Google Gemini, Claude Sonnet, Microsoft Copilot and OpenAI GPT, revealing a systemic vulnerability requiring immediate attention.
Jiaming Zhang 0006, Junhong Ye, Xingjun Ma, Yige Li, Yunfan Yang, Jitao Sang 0001, Dit-Yan Yeung
CVPR8
2025 Fast and Slow Streams for Online Time Series Forecasting Without Information Leakage
abstract
Current research in online time series forecasting (OTSF) faces two significant issues. The first is information leakage, where models make predictions and are then evaluated on historical time steps that have already been used in backpropagation for parameter updates. The second is practicality: while forecasting in real-world applications typically emphasizes looking ahead and anticipating future uncertainties, prediction sequences in this setting include only one future step with the remaining being observed time points. This necessitates a redefinition of the OTSF setting, focusing on predicting unknown future steps and evaluating unobserved data points. Following this new setting, challenges arise in leveraging incomplete pairs of ground truth and predictions for backpropagation, as well as in generalizing accurate information without overfitting to noise from recent data streams. To address these challenges, we propose a novel dual-stream framework for online forecasting (DSOF): a slow stream that updates with complete data using experience replay, and a fast stream that adapts to recent data through temporal difference learning. This dual-stream approach updates a teacher-student model learned through a residual learning strategy, generating predictions in a coarse-to-fine manner. Extensive experiments demonstrate its improvement in forecasting performance in changing environments. Our code is publicly available at https://github.com/yyalau/iclr2025_dsof.
Ying-yee Ava Lau, Zhiwen Shao, Dit-Yan Yeung
ICLR3
2025 Understanding LLMs' Fluid Intelligence Deficiency: An Analysis of the ARC Task
abstract
Junjie Wu, Mo Yu, Lemao Liu, Dit-Yan Yeung, Jie Zhou. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Junjie Wu 0007, Mo Yu, Lemao Liu, Dit-Yan Yeung, Jie Zhou 0016
NAACL (Long Papers)4
2025 The Stochastic Parrot on LLM's Shoulder: A Summative Assessment of Physical Concept Understanding
abstract
Mo Yu, Lemao Liu, Junjie Wu, Tsz Ting Chung, Shunchi Zhang, Jiangnan Li, Dit-Yan Yeung, Jie Zhou. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Mo Yu, Lemao Liu, Junjie Wu 0007, Tsz Ting Chung, Shunchi Zhang, Dit-Yan Yeung, Jie Zhou 0016
NAACL (Long Papers)7
2025 Learning 3D Persistent Embodied World Models
abstract
The ability to simulate the effects of future actions on the world is a crucial ability of intelligent embodied agents, enabling agents to anticipate the effects of their actions and make plans accordingly. While a large body of existing work has explored how to construct such world models using video models, they are often myopic in nature, without any memory of a scene not captured by currently observed images, preventing agents from making consistent long-horizon plans in complex environments where many parts of the scene are partially observed. We introduce a new persistent embodied world model with an explicit memory of previously generated content, enabling much more consistent long-horizon simulation. During generation time, our video diffusion model predicts RGB-D video of the future observations of the agent. This generation is then aggregated into a persistent 3D map of the environment. By conditioning the video model on this 3D spatial map, we illustrate how this enables video world models to faithfully simulate both seen and unseen parts of the world. Finally, we illustrate the efficacy of such a world model in downstream embodied applications, enabling effective planning and policy learning.
Yilun Du, Yuncong Yang, Peihao Chen, Dit-Yan Yeung, Chuang Gan 0001
NeurIPS6
2025 Automated Evaluation of Large Vision-Language Models on Self-Driving Corner Cases
abstract
Large Vision-Language Models (LVLMs) have received widespread attentions for advancing the interpretable self-driving. Existing evaluations of LVLMs primarily focus on multi-faceted capabilities in natural circumstances, lacking automated and quantifiable assessment for self-driving, let alone the severe road corner cases. In this work, we propose CODA-LM, the very first benchmark for the automatic evaluation of LVLMs for self-driving corner cases. We adopt a hierarchical data structure and prompt powerful LVLMs to analyze complex driving scenes and generate high-quality pre-annotations for the human annotators, while for LVLM evaluation, we show that using the text-only large language models (LLMs) as judges reveals even better alignment with human preferences than the LVLM judges. Moreover, with our CODA-LM, we build CODA-VLM, a new driving LVLM surpassing all open-sourced counterparts on CODA-LM. Our CODA-VLM performs comparably with GPT-4V, even surpassing GPT-4V by +21.42% on the regional perception task. We hope CODA-LM can become the catalyst to promote interpretable self-driving empowered by LVLMs.
Kai Chen 0023, Yanxin Liu, Ruiyuan Gao 0001, Lanqing Hong, Xinhai Zhao, Zhenguo Li, Dit-Yan Yeung, Huchuan Lu, Xu Jia 0012
WACV11
2025 TrackDiffusion: Tracklet-Conditioned Video Generation via Diffusion Models
abstract
Despite remarkable achievements in video synthesis, achieving granular control over complex dynamics, such as nuanced movement among multiple interacting objects, still presents a significant hurdle for dynamic world modeling, compounded by the necessity to manage appearance and disappearance, drastic scale changes, and ensure consistency for instances across frames. These challenges hinder the development of video generation that can faithfully mimic real-world complexity, limiting utility for applications requiring high-level realism and controllability, including advanced scene simulation and training of perception systems. To address that, we propose TrackDiffusion, a novel video generation framework affording fine-grained trajectory-conditioned motion control via diffusion models, which facilitates the precise manipulation of the object trajectories and interactions, overcoming the prevalent limitation of scale and continuity disruptions. A pivotal component of TrackDiffusion is the instance enhancer, which explicitly ensures inter-frame consistency of multiple objects, a critical factor overlooked in the current literature. More-over, we demonstrate that generated video sequences by our TrackDiffusion can be used as training data for visual per-ception models. To the best of our knowledge, this is the first work to apply video diffusion models with tracklet conditions and demonstrate that generated frames can be beneficial for improving the performance of object trackers. 1
Kai Chen 0023, Zhili Liu, Ruiyuan Gao 0001, Lanqing Hong, Dit-Yan Yeung, Huchuan Lu, Xu Jia 0012
WACV6
2025 Graph Neural Networks for multivariate time-series forecasting via learning hierarchical spatiotemporal dependencies
abstract
Multivariate time-series forecasting is one of the essential tasks to draw insights from sequential data. Spatiotemporal Graph Neural Networks (STGNN) have attracted much attention in this field due to their capability to capture the underlying spatiotemporal dependencies. However, current STGNN solutions succumb to a higher degree of error in their predictions due to insufficient modelling of the dependencies and dynamics at different levels. In this paper, a Graph Neural Networks-based model is proposed for multivariate time-series forecasting via learning hierarchical spatiotemporal dependencies (HSDGNN). Specifically, variables are organised as nodes in a graph while each node serves as a subgraph consisting of the attributes of variables. Then two-level convolutions are designed on the hierarchical graph to model the spatial dependencies with different granularities. The changes in graph topologies are also encoded for strengthening dependency modelling across time and spatial dimensions. The proposed model is tested using real-world datasets from different domains, including transportation, electricity, and meteorology. The experimental results demonstrate that HSDGNN can outperform state-of-the-art baselines by up to 15.3% in terms of prediction accuracy, without compromising model scalability. • A new hierarchical spatiotemporal dependency learning-based graph neural network. • The model leverages spatial-, temporal-, and intra-dependency learning processes. • The temporal correlations among dynamic graph topologies are considered. • The model is evaluated on real-world datasets from different engineering domains.
Zhou Zhou 0003, Ronisha Basker, Dit-Yan Yeung
Eng. Appl. Artif. Intell.3
2025 Mirror Detection via Multi-Directional Similarity Perception and Spectral Saliency Enhancement
abstract
Mirror detection is a challenging task, due to the reflective properties of mirrors. Most existing approaches rely on exploiting the relationship between the content inside the mirror and the surrounding environment to aid in locating mirrors. A typical solution is to utilize contextual contrasted features. However, the discontinuity in content at the edges of mirrors may not always be prominent. To overcome this limitation, we propose a novel mirror detection framework called S2MD including two main modules, multi-directional similarity perception module (MSPM) and spectral saliency enhancement decoder module (SSEDM). Specifically, we employ a backbone network to extract multi-scale global information from images using a dual-path approach. Then, we feed these high-level dual-path features into MSPMs to generate direction-sensitive similarity-consistent features. MSPM utilizes active rotating filters and oriented response pooling to model the similarity relations in different orientations. Moreover, the SSEDM is utilized to enhance the spatial contextual contrasted features using feature spectral residuals and fuse the dual-path features to obtain the final predicted mirror mask. Extensive experiments demonstrate that our method achieves state-of-the-art performance on challenging MSD, PMD, and RGBD-Mirror benchmarks. The code is available at https://github.com/RuiChen-stack/M2SD.
Zhiwen Shao, Xuehuai Shi, Bing Liu 0016, Canlin Li, Lizhuang Ma, Dit-Yan Yeung
IEEE Trans. Circuits Syst. Video Technol.7
2025 MF-CLIP: Leveraging CLIP as Surrogate Models for No-Box Adversarial Attacks
Jiaming Zhang 0006, Lingyu Qiu, Qi Yi, Yige Li, Jitao Sang 0001, Changsheng Xu, Dit-Yan Yeung
IEEE Trans. Inf. Forensics Secur.7
2025 Micro-Expression Recognition via Fine-Grained Dynamic Perception
abstract
Facial micro-expression recognition (MER) is a challenging task, due to the transience, subtlety, and dynamics of micro-expressions (MEs). Most existing methods resort to hand-crafted features or deep networks, in which the former often additionally requires key frames, and the latter suffers from small-scale and low-diversity training data. In this article, we develop a novel fine-grained dynamic perception (FDP) framework for MER. We propose to rank frame-level features of a sequence of raw frames in chronological order, in which the rank process encodes the dynamic information of both ME appearances and motions. Specifically, a novel local-global feature-aware transformer is proposed for frame representation learning. A rank scorer is further adopted to calculate rank scores of each frame-level feature. Afterwards, the rank features from rank scorer are pooled in temporal dimension to capture dynamic representation. Finally, the dynamic representation is shared by a MER module and a dynamic image construction module, in which the former predicts the ME category, and the latter uses an encoder-decoder structure to construct the dynamic image. The design of dynamic image construction task is beneficial for capturing facial subtle actions associated with MEs and alleviating the data scarcity issue. Extensive experiments show that our method (i) significantly outperforms the state-of-the-art MER methods, and (ii) works well for dynamic image construction. Particularly, our FDP improves by 4.05%, 2.50%, 7.71%, and 2.11% over the previous best results in terms of F1-score on the CASME II, SAMM, CAS(ME) 2 , and CAS(ME) 3 datasets, respectively. The code is available at https://github.com/CYF-cuber/FDP .
Zhiwen Shao, Xuehuai Shi, Canlin Li, Lizhuang Ma, Dit-Yan Yeung
ACM Trans. Multim. Comput. Commun. Appl.7
2025 High-level LoRA and hierarchical fusion for enhanced micro-expression recognition
Zhiwen Shao, Yong Zhou 0003, Xiang Xiang 0001, Jian Li 0054, Bing Liu 0016, Dit-Yan Yeung
Vis. Comput.7
2024 Gaussian Shell Maps for Efficient 3D Human Generation
abstract
Efficient generation of 3D digital humans is important in several industries, including virtual reality, social media, and cinematic production. 3D generative adversarial net-works (GANs) have demonstrated state-of-the-art (SOTA) quality and diversity for generated assets. Current 3D GAN architectures, however, typically rely on volume representations, which are slow to render, thereby hampering the GAN training and requiring multi- view-inconsistent 2D upsam-plers. Here, we introduce Gaussian Shell Maps (GSMs) as a framework that connects SOTA generator network archi-tectures with emerging 3D Gaussian rendering primitives using an articulable multi shell-based scaffold. In this set-ting, a CNN generates a 3D texture stack with features that are mapped to the shells. The latter represent inflated and deflated versions of a template surface of a digital human in a canonical body pose. Instead of rasterizing the shells directly, we sample 3D Gaussians on the shells whose at-tributes are encoded in the texture features. These Gaus-sians are efficiently and differentiably rendered. The ability to articulate the shells is important during GAN training and, at inference time, to deform a body into arbitrary user-defined poses. Our efficient rendering scheme bypasses the need for view-inconsistent upsamplers and achieves high-quality multi-view consistent renderings at a native resolution of 512 ×512 pixels. We demonstrate that GSMs suc-cessfully generate 3D humans when trained on single-view datasets, including SHHQ and DeepFashion. Project Page: rameenabdal.github.io/GaussianShellMaps
Rameen Abdal, Wang Yifan 0001, Zifan Shi, Yinghao Xu 0001, Ryan Po, Zhengfei Kuang, Qifeng Chen 0001, Dit-Yan Yeung, Gordon Wetzstein
CVPR8
2024 DetDiffusion: Synergizing Generative and Perceptive Models for Enhanced Data Generation and Perception
abstract
Current perceptive models heavily depend on resource-intensive datasets, prompting the need for innovative solutions. Leveraging recent advances in diffusion models, synthetic data, by constructing image inputs from various annotations, proves beneficial for downstream tasks. While prior methods have separately addressed generative and perceptive models, DetDiffusion, for the first time, harmonizes both, tackling the challenges in generating effective data for perceptive models. To enhance image generation with perceptive models, we introduce perception-aware loss (P.A. loss) through segmentation, improving both quality and controllability. To boost the performance of specific perceptive models, our method customizes data augmentation by extracting and utilizing perception-aware attribute (P.A. Attr) during generation. Experimental results from the object detection task highlight DetDiffusion's superior performance, establishing a new state-of-the-art in layout-guided generation. Furthermore, image syntheses from DetDiffusion can effectively augment training data, significantly enhancing downstream detection performance.
Yibo Wang 0039, Ruiyuan Gao 0001, Kai Chen 0023, Kaiqiang Zhou, Yingjie Cai, Lanqing Hong, Zhenguo Li, Lihui Jiang, Dit-Yan Yeung, Qiang Xu 0001, Kai Zhang 0008
CVPR9
2024 Learning High-Resolution Vector Representation from Multi-camera Images for 3D Object Detection
Shuangjie Xu, Maosheng Ye, Zian Qian, Xiaoyi Zou, Dit-Yan Yeung, Qifeng Chen 0001
ECCV (35)6
2024 Eyes Closed, Safety on: Protecting Multimodal LLMs via Image-to-Text Transformation
Yunhao Gou, Kai Chen 0023, Zhili Liu, Lanqing Hong, Hang Xu 0004, Zhenguo Li, Dit-Yan Yeung, James T. Kwok, Yu Zhang 0006
ECCV (17)7
2024 JointDreamer: Ensuring Geometry Consistency and Text Congruence in Text-to-3D Generation via Joint Score Distillation
Chenhan Jiang, Yihan Zeng, Tianyang Hu 0001, Songcun Xu, Wei Zhang 0196, Hang Xu 0004, Dit-Yan Yeung
ECCV (26)7
2024 Implicit Concept Removal of Diffusion Models
Zhili Liu, Kai Chen 0023, Jianhua Han, Lanqing Hong, Hang Xu 0004, Zhenguo Li, Dit-Yan Yeung, James T. Kwok
ECCV (21)8
2024 Rethinking Targeted Adversarial Attacks for Neural Machine Translation
abstract
Targeted adversarial attacks are widely used to evaluate the robustness of neural machine translation systems. Unfortunately, this paper first identifies a critical issue in the existing settings of NMT targeted adversarial attacks, where their attacking results are largely overestimated. To this end, this paper presents a new setting for NMT targeted adversarial attacks that could lead to reliable attacking results. Under the new setting, it then proposes a Targeted Word Gradient adversarial Attack (TWGA) method to craft adversarial examples. Experimental results demonstrate that our proposed setting could provide faithful attacking results for targeted adversarial attacks on NMT systems, and the proposed TWGA method can effectively attack such victim NMT systems. In-depth analyses on a large-scale dataset further illustrate some valuable findings.1Our code and data are available at https://github.com/wujunjie1998/TWGA.
Junjie Wu 0007, Lemao Liu, Wei Bi, Dit-Yan Yeung
ICASSP4
2024 MagicDrive: Street View Generation with Diverse 3D Geometry Control
abstract
Recent advancements in diffusion models have significantly enhanced the data synthesis with 2D control. Yet, precise 3D control in street view generation, crucial for 3D perception tasks, remains elusive. Specifically, utilizing Bird's-Eye View (BEV) as the primary condition often leads to challenges in geometry control (e.g., height), affecting the representation of object shapes, occlusion patterns, and road surface elevations, all of which are essential to perception data synthesis, especially for 3D object detection tasks. In this paper, we introduce MagicDrive, a novel street view generation framework, offering diverse 3D geometry controls including camera poses, road maps, and 3D bounding boxes, together with textual descriptions, achieved through tailored encoding strategies. Besides, our design incorporates a cross-view attention module, ensuring consistency across multiple camera views. With MagicDrive, we achieve high-fidelity street-view image & video synthesis that captures nuanced 3D geometry and various scene descriptions, enhancing tasks like BEV segmentation and 3D object detection. Project Website: https://flymin.github.io/magicdrive
Ruiyuan Gao 0001, Kai Chen 0023, Enze Xie, Lanqing Hong, Zhenguo Li, Dit-Yan Yeung, Qiang Xu 0001
ICLR6
2024 Gaining Wisdom from Setbacks: Aligning Large Language Models via Mistake Analysis
abstract
The rapid development of large language models (LLMs) has not only provided numerous opportunities but also presented significant challenges. This becomes particularly evident when LLMs inadvertently generate harmful or toxic content, either unintentionally or because of intentional inducement. Existing alignment methods usually direct LLMs toward the favorable outcomes by utilizing human-annotated, flawless instruction-response pairs. Conversely, this study proposes a novel alignment technique based on mistake analysis, which deliberately exposes LLMs to erroneous content to learn the reasons for mistakes and how to avoid them. In this case, mistakes are repurposed into valuable data for alignment, effectively helping to avoid the production of erroneous responses. Without external models or human annotations, our method leverages a model's intrinsic ability to discern undesirable mistakes and improves the safety of its generated responses. Experimental results reveal that our method outperforms existing alignment approaches in enhancing model safety while maintaining the overall utility.
Kai Chen 0023, Chunwei Wang, Jianhua Han, Lanqing Hong, Fei Mi, Hang Xu 0004, Zhengying Liu, Wenyong Huang, Zhenguo Li, Dit-Yan Yeung, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001
ICLR11
2024 GeoDiffusion: Text-Prompted Geometric Control for Object Detection Data Generation
abstract
Diffusion models have attracted significant attention due to the remarkable ability to create content and generate data for tasks like image classification. However, the usage of diffusion models to generate the high-quality object detection data remains an underexplored area, where not only image-level perceptual quality but also geometric conditions such as bounding boxes and camera views are essential. Previous studies have utilized either copy-paste synthesis or layout-to-image (L2I) generation with specifically designed modules to encode the semantic layouts. In this paper, we propose the GeoDiffusion, a simple framework that can flexibly translate various geometric conditions into text prompts and empower pre-trained text-to-image (T2I) diffusion models for high-quality detection data generation. Unlike previous L2I methods, our GeoDiffusion is able to encode not only the bounding boxes but also extra geometric conditions such as camera views in self-driving scenes. Extensive experiments demonstrate GeoDiffusion outperforms previous L2I methods while maintaining 4x training time faster. To the best of our knowledge, this is the first work to adopt diffusion models for layout-to-image generation with geometric conditions and demonstrate that L2I-generated images can be beneficial for improving the performance of object detectors.
Kai Chen 0023, Enze Xie, Yibo Wang 0039, Lanqing Hong, Zhenguo Li, Dit-Yan Yeung
ICLR7
2024 RoboDreamer: Learning Compositional World Models for Robot Imagination
abstract
Text-to-video models have demonstrated substantial potential in robotic decision-making, enabling the imagination of realistic plans of future actions as well as accurate environment simulation. However, one major issue in such models is generalization – models are limited to synthesizing videos subject to language instructions similar to those seen at training time. This is heavily limiting in decision-making, where we seek a powerful world model to synthesize plans of unseen combinations of objects and actions in order to solve previously unseen tasks in new environments. To resolve this issue, we introduce RoboDreamer, an innovative approach for learning a compositional world model by factorizing the video generation. We leverage the natural compositionality of language to parse instructions into a set of lower-level primitives, which we condition a set of models on to generate videos. We illustrate how this factorization naturally enables compositional generalization, by allowing us to formulate a new natural language instruction as a combination of previously seen components. We further show how such a factorization enables us to add additional multimodal goals, allowing us to specify a video we wish to generate given both natural language instructions and a goal image. Our approach can successfully synthesize video plans on unseen goals in the RT-X, enables successful robot execution in simulation, and substantially outperforms monolithic baseline approaches to video generation.
Yilun Du, Jiaben Chen, Yandong Li, Dit-Yan Yeung, Chuang Gan 0001
ICML5
2024 Pre-train and Refine: Towards Higher Efficiency in K-Agnostic Community Detection without Quality Degradation
abstract
Community detection (CD) is a classic graph inference task that partitions nodes of a graph into densely connected groups. While many CD methods have been proposed with either impressive quality or efficiency, balancing the two aspects remains a challenge. This study explores the potential of deep graph learning to achieve a better trade-off between the quality and efficiency of K-agnostic CD, where the number of communities K is unknown. We propose PRoCD (Pre-training & Refinement fOr Community Detection), a simple yet effective method that reformulates K-agnostic CD as the binary node pair classification. PRoCD follows a pre-training & refinement paradigm inspired by recent advances in pre-training techniques. We first conduct the offline pre-training of PRoCD on small synthetic graphs covering various topology properties. Based on the inductive inference across graphs, we then generalize the pre-trained model (with frozen parameters) to large real graphs and use the derived CD results as the initialization of an existing efficient CD method (e.g., InfoMap) to further refine the quality of CD results. In addition to benefiting from the transfer ability regarding quality, the online generalization and refinement can also help achieve high inference efficiency, since there is no time-consuming model optimization. Experiments on public datasets with various scales demonstrate that PRoCD can ensure higher efficiency in K-agnostic CD without significant quality degradation.
Meng Qin 0002, Chaorui Zhang, Yu Gao 0041, Weixi Zhang, Dit-Yan Yeung
KDD5
2024 Fourier Amplitude and Correlation Loss: Beyond Using L2 Loss for Skillful Precipitation Nowcasting
abstract
Deep learning approaches have been widely adopted for precipitation nowcasting in recent years. Previous studies mainly focus on proposing new model architectures to improve pixel-wise metrics. However, they frequently result in blurry predictions which provide limited utility to forecasting operations. In this work, we propose a new Fourier Amplitude and Correlation Loss (FACL) which consists of two novel loss terms: Fourier Amplitude Loss (FAL) and Fourier Correlation Loss (FCL). FAL regularizes the Fourier amplitude of the model prediction and FCL complements the missing phase information. The two loss terms work together to replace the traditional L2 losses such as MSE and weighted MSE for the spatiotemporal prediction problem on signal-based data. Our method is generic, parameter-free and efficient. Extensive experiments using one synthetic dataset and three radar echo datasets demonstrate that our method improves perceptual metrics and meteorology skill scores, with a small trade-off to pixel-wise accuracy and structural similarity. Moreover, to improve the error margin in meteorological skill scores such as Critical Success Index (CSI) and Fractions Skill Score (FSS), we propose and adopt the Regional Histogram Divergence (RHD), a distance metric that considers the patch-wise similarity between signal-based imagery patterns with tolerance to local transforms.
Chiu Wai Yan, Shi Quan Foo, Van-Hoan Trinh, Dit-Yan Yeung, Ka-Hing Wong, Wai-Kin Wong
NeurIPS4
2024 A Survey of Automated Data Augmentation for Image Classification: Learning to Compose, Mix, and Generate
abstract
Data augmentation is an effective way to improve the generalization of deep learning models. However, the underlying augmentation methods mainly rely on handcrafted operations, such as flipping and cropping for image data. These augmentation methods are often designed based on human expertise or repeated trials. Meanwhile, automated data augmentation (AutoDA) is a promising research direction that frames the data augmentation process as a learning task and finds the most effective way to augment the data. In this survey, we categorize recent AutoDA methods into the composition-, mixing-, and generation-based approaches and analyze each category in detail. Based on the analysis, we discuss the challenges and future prospects as well as provide guidelines for applying AutoDA methods by considering the dataset, computation effort, and availability of domain-specific transformations. It is hoped that this article can provide a useful list of AutoDA methods and guidelines for data partitioners when deploying AutoDA in practice. The survey can also serve as a reference for further study by researchers in this emerging research area.
Tsz-Him Cheung, Dit-Yan Yeung
IEEE Trans. Neural Networks Learn. Syst.2
2023 Mixed Autoencoder for Self-Supervised Visual Representation Learning
abstract
Masked Autoencoder (MAE) has demonstrated superior performance on various vision tasks via randomly masking image patches and reconstruction. However, effective data augmentation strategies for MAE still remain open questions, different from those in contrastive learning that serve as the most important part. This paper studies the prevailing mixing augmentation for MAE. We first demonstrate that naïve mixing will in contrast degenerate model performance due to the increase of mutual information (MI). To address, we propose homologous recognition, an auxiliary pretext task, not only to alleviate the MI increasement by explicitly requiring each patch to recognize homologous patches, but also to perform object-aware self-supervised pre-training for better downstream dense perception performance. With extensive experiments, we demonstrate that our proposed Mixed Autoencoder (MixedAE) achieves the state-of-the-art transfer results among masked image modeling (MIM) augmentations on different downstream tasks with significant efficiency. Specifically, our MixedAE outperforms MAE by +0.3% accuracy, +1.7 mIoU and +0.9 AP on ImageNet-1K, ADE20K and COCO respectively with a standard ViT-Base. Moreover, MixedAE surpasses iBOT, a strong MIM method combined with instance discrimination, while accelerating training by 2×. To our best knowledge, this is the very first work to consider mixing for MIM from the perspective of pretext task design. Code will be made available.
Kai Chen 0023, Zhili Liu, Lanqing Hong, Hang Xu 0004, Zhenguo Li, Dit-Yan Yeung
CVPR6
2023 Learning 3D-Aware Image Synthesis with Unknown Pose Distribution
abstract
Existing methods for 3D-aware image synthesis largely depend on the 3D pose distribution pre-estimated on the training set. An inaccurate estimation may mislead the model into learning faulty geometry. This work proposes PoF3D that frees generative radiance fields from the requirements of 3D pose priors. We first equip the generator with an efficient pose learner, which is able to infer a pose from a latent code, to approximate the underlying true pose distribution automatically. We then assign the discriminator a task to learn pose distribution under the supervision of the generator and to differentiate real and synthesized images with the predicted pose as the condition. The pose-free generator and the pose-aware discriminator are jointly trained in an adversarial manner. Extensive results on a couple of datasets confirm that the performance of our approach, regarding both image quality and geometry quality, is on par with state of the art. To our best knowledge, PoF3D demonstrates the feasibility of learning high-quality 3D-aware image synthesis without using 3D pose priors for the first time. Project page can be found here.
Zifan Shi, Yujun Shen, Yinghao Xu 0001, Sida Peng, Yiyi Liao, Qifeng Chen 0001, Dit-Yan Yeung
CVPR8
2023 CLIP2: Contrastive Language-Image-Point Pretraining from Real-World Point Cloud Data
abstract
Contrastive Language-Image Pre-training, benefiting from large-scale unlabeled text-image pairs, has demonstrated great performance in open-world vision understanding tasks. However, due to the limited Text-3D data pairs, adapting the success of 2D Vision-Language Models (VLM) to the 3D space remains an open problem. Existing works that leverage VLM for 3D understanding generally resort to constructing intermediate 2D representations for the 3D data, but at the cost of losing 3D geometry information. To take a step toward open-world 3D vision understanding, we propose Contrastive Language-Image-Point Cloud Pretraining (CLIP2) to directly learn the transferable 3D point cloud representation in realistic scenarios with a novel proxy alignment mechanism. Specifically, we exploit naturally-existed correspondences in 2D and 3D scenarios, and build well-aligned and instance-based text-image-point proxies from those complex scenarios. On top of that, we propose a cross-modal contrastive objective to learn semantic and instance-level aligned point cloud representation. Experimental results on both indoor and outdoor scenarios show that our learned 3D representation has great transfer ability in downstream tasks, including zero-shot and few-shot 3D recognition, which boosts the state-of-the-art methods by large margins. Furthermore, we provide analyses of the capability of different representations in real scenarios and present the optional ensemble scheme.
Yihan Zeng, Chenhan Jiang, Jiageng Mao, Jianhua Han, Chaoqiang Ye, Qingqiu Huang, Dit-Yan Yeung, Xiaodan Liang, Hang Xu 0004
CVPR7
2023 SVQNet: Sparse Voxel-Adjacent Query Network for 4D Spatio-Temporal LiDAR Semantic Segmentation
abstract
LiDAR-based semantic perception tasks are critical yet challenging for autonomous driving. Due to the motion of objects and static/dynamic occlusion, temporal information plays an essential role in reinforcing perception by enhancing and completing single-frame knowledge. Previous approaches either directly stack historical frames to the current frame or build a 4D spatio-temporal neighborhood using KNN, which duplicates computation and hinders real-time performance. Based on our observation that stacking all the historical points would damage performance due to a large amount of redundant and misleading information, we propose the Sparse Voxel-Adjacent Query Network (SVQNet) for 4D LiDAR semantic segmentation. To take full advantage of the historical frames high-efficiently, we shunt the historical points into two groups with reference to the current points. One is the Voxel-Adjacent Neighborhood carrying local enhancing knowledge. The other is the Historical Context completing the global knowledge. Then we propose new modules to select and extract the instructive features from the two groups. Our SVQNet achieves state-of-the-art performance in LiDAR semantic segmentation of the SemanticKITTI benchmark and the nuScenes dataset.
Xuechao Chen, Shuangjie Xu, Xiaoyi Zou, Tongyi Cao, Dit-Yan Yeung
ICCV5
2023 ILA-DA: Improving Transferability of Intermediate Level Attack with Data Augmentation
Chiu Wai Yan, Tsz-Him Cheung, Dit-Yan Yeung
ICLR3
2023 Adaptive Online Replanning with Diffusion Models
abstract
Diffusion models have risen a promising approach to data-driven planning, and have demonstrated impressive robotic control, reinforcement learning, and video planning performance. Given an effective planner, an important question to consider is replanning -- when given plans should be regenerated due to both action execution error and external environment changes. Direct plan execution, without replanning, is problematic as errors from individual actions rapidly accumulate and environments are partially observable and stochastic. Simultaneously, replanning at each timestep incurs a substantial computational cost, and may prevent successful task execution, as different generated plans prevent consistent progress to any particular goal. In this paper, we explore how we may effectively replan with diffusion models. We propose a principled approach to determine when to replan, based on the diffusion model's estimated likelihood of existing generated plans. We further present an approach to replan existing trajectories to ensure that new plans follow the same goal state as the original trajectory, which may efficiently bootstrap off previously generated plans. We illustrate how a combination of our proposed additions significantly improves the performance of diffusion planners leading to 38\% gains over past diffusion planning approaches on Maze2D and further enables handling of stochastic and long-horizon robotic control tasks.
Yilun Du, Mengdi Xu, Yikang Shen, Wei Xiao 0003, Dit-Yan Yeung, Chuang Gan 0001
NeurIPS7
2023 Detection Recovery in Online Multi-Object Tracking with Sparse Graph Tracker
abstract
In existing joint detection and tracking methods, pairwise relational features are used to match previous track-lets to current detections. However, the features may not be discriminative enough for a tracker to identify a target from a large number of detections. Selecting only high-scored detections for tracking may lead to missed detections whose confidence score is low. Consequently, in the on-line setting, this results in disconnections of tracklets which cannot be recovered. In this regard, we present Sparse Graph Tracker (SGT), a novel online graph tracker using higher-order relational features which are more discriminative by aggregating the features of neighboring detections and their relations. SGT converts video data into a graph where detections, their connections, and the relational features of two connected nodes are represented by nodes, edges, and edge features, respectively. The strong edge features allow SGT to track targets with tracking candidates selected by top-K scored detections with large K. As a result, even low-scored detections can be tracked, and the missed detections are also recovered. The robustness of K value is shown through the extensive experiments. In the MOT16/17/20 and HiEve Challenge, SGT outperforms the state-of-the-art trackers with real-time inference speed. Especially, a large improvement in MOTA is shown in the MOT20 and HiEve Challenge. Code is available at https://github.com/HYUNJS/SGT.
Jeongseok Hyun, Myunggu Kang, Dongyoon Wee, Dit-Yan Yeung
WACV4
2023 Towards a Better Tradeoff between Quality and Efficiency of Community Detection: An Inductive Embedding Method across Graphs
abstract
Many network applications can be formulated as NP-hard combinatorial optimization problems of community detection (CD) that partitions nodes of a graph into several groups with dense linkage. Most existing CD methods are transductive , which independently optimized their models for each single graph, and can only ensure either high quality or efficiency of CD by respectively using advanced machine learning techniques or fast heuristic approximation. In this study, we consider the CD task and aims to alleviate its NP-hard challenge. Motivated by the efficient inductive inference of graph neural networks (GNNs), we explore the possibility to achieve a better tradeoff between the quality and efficiency of CD via an inductive embedding scheme across multiple graphs of a system and propose a novel inductive community detection (ICD) method. Concretely, ICD first conducts the offline training of an adversarial dual GNN structure on historical graphs to capture key properties of a system. The trained model is then directly generalized to new graphs of the same system for online CD without additional optimization, where a better tradeoff between quality and efficiency can be achieved. Compared with existing inductive approaches, we develop a novel feature extraction module based on graph coarsening, which can efficiently extract informative feature inputs for GNNs. Moreover, our original designs of adversarial dual GNN and clustering regularization loss further enable ICD to capture permutation-invariant community labels in the offline training and help derive community-preserved embedding to support the high-quality online CD. Experiments on a set of benchmarks demonstrate that ICD can achieve a significant tradeoff between quality and efficiency over various baselines.
Meng Qin 0002, Chaorui Zhang, Bo Bai 0001, Gong Zhang 0001, Dit-Yan Yeung
ACM Trans. Knowl. Discov. Data5
2023 High-Quality Temporal Link Prediction for Weighted Dynamic Graphs via Inductive Embedding Aggregation
abstract
Temporal link prediction (TLP) is an inference task on dynamic graphs that predicts future topology using historical graph snapshots. Existing TLP methods are usually designed for unweighted graphs with fixed node sets. Some of them cannot be generalized to the prediction of weighted graphs with non-fixed node sets. Although several methods can still be used to predict weighted graphs, they can only derivelow-qualityprediction snapshots sensitive to large edge weights but fail to distinguish small and zero weights in adjacency matrices. In this study, we consider the challenginghigh-qualityTLP on weighted dynamic graphs and propose a novel inductive dynamic embedding aggregation (IDEA) method, inspired by the high-resolution video prediction. IDEA combines conventional error minimization objectives with a scale difference minimization objective, which can generatehigh-qualityweighted prediction snapshots, distinguishing differences among large, small, and zero weights in adjacency matrices. Since IDEA adopts an inductive dynamic embedding scheme with an attentive node aligning unit and adaptive embedding aggregation module, it can also tackle the TLP on weighted graphs even with non-fixed node sets. Experiments on datasets of various scenarios validate that IDEA can derivehigh-qualityprediction results for weighted dynamic graphs and tackle the variation of node sets.
Meng Qin 0002, Chaorui Zhang, Bo Bai 0001, Gong Zhang 0001, Dit-Yan Yeung
IEEE Trans. Knowl. Data Eng.5
2022 CODA: A Real-World Road Corner Case Dataset for Object Detection in Autonomous Driving
Kaican Li, Kai Chen 0023, Lanqing Hong, Chaoqiang Ye, Jianhua Han, Yukuai Chen, Wei Zhang 0196, Chunjing Xu, Dit-Yan Yeung, Xiaodan Liang, Zhenguo Li, Hang Xu 0004
ECCV (38)10
2022 3D-Aware Indoor Scene Synthesis with Depth Priors
Zifan Shi, Yujun Shen, Jiapeng Zhu 0001, Dit-Yan Yeung, Qifeng Chen 0001
ECCV (16)4
2022 AdaAug: Learning Class- and Instance-adaptive Data Augmentation Policies
Tsz-Him Cheung, Dit-Yan Yeung
ICLR2
2022 Earthformer: Exploring Space-Time Transformers for Earth System Forecasting
abstract
Conventionally, Earth system (e.g., weather and climate) forecasting relies on numerical simulation with complex physical models and hence is both expensive in computation and demanding on domain expertise. With the explosive growth of spatiotemporal Earth observation data in the past decade, data-driven models that apply Deep Learning (DL) are demonstrating impressive potential for various Earth system forecasting tasks. The Transformer as an emerging DL architecture, despite its broad success in other domains, has limited adoption in this area. In this paper, we propose Earthformer, a space-time Transformer for Earth system forecasting. Earthformer is based on a generic, flexible and efficient space-time attention block, named Cuboid Attention. The idea is to decompose the data into cuboids and apply cuboid-level self-attention in parallel. These cuboids are further connected with a collection of global vectors. We conduct experiments on the MovingMNIST dataset and a newly proposed chaotic $N$-body MNIST dataset to verify the effectiveness of cuboid attention and figure out the best design of Earthformer. Experiments on two real-world benchmarks about precipitation nowcasting and El Niño/Southern Oscillation (ENSO) forecasting show that Earthformer achieves state-of-the-art performance.
Zhihan Gao 0001, Xingjian Shi, Hao Wang 0014, Yi Zhu 0001, Yuyang Wang 0001, Mu Li 0003, Dit-Yan Yeung
NeurIPS7
2022 Improving 3D-aware Image Synthesis with A Geometry-aware Discriminator
abstract
3D-aware image synthesis aims at learning a generative model that can render photo-realistic 2D images while capturing decent underlying 3D shapes. A popular solution is to adopt the generative adversarial network (GAN) and replace the generator with a 3D renderer, where volume rendering with neural radiance field (NeRF) is commonly used. Despite the advancement of synthesis quality, existing methods fail to obtain moderate 3D shapes. We argue that, considering the two-player game in the formulation of GANs, only making the generator 3D-aware is not enough. In other words, displacing the generative mechanism only offers the capability, but not the guarantee, of producing 3D-aware images, because the supervision of the generator primarily comes from the discriminator. To address this issue, we propose GeoD through learning a geometry-aware discriminator to improve 3D-aware GANs. Concretely, besides differentiating real and fake samples from the 2D image space, the discriminator is additionally asked to derive the geometry information from the inputs, which is then applied as the guidance of the generator. Such a simple yet effective design facilitates learning substantially more accurate 3D shapes. Extensive experiments on various generator architectures and training datasets verify the superiority of GeoD over state-of-the-art alternatives. Moreover, our approach is registered as a general framework such that a more capable discriminator (i.e., with a third task of novel view synthesis beyond domain classification and geometry extraction) can further assist the generator with a better multi-view consistency. Project page can be found at https://vivianszf.github.io/geod.
Zifan Shi, Yinghao Xu 0001, Yujun Shen, Deli Zhao, Qifeng Chen 0001, Dit-Yan Yeung
NeurIPS6
2021 Probing Toxic Content in Large Pre-Trained Language Models
abstract
Nedjma Ousidhoum, Xinran Zhao, Tianqing Fang, Yangqiu Song, Dit-Yan Yeung. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Nedjma Ousidhoum, Tianqing Fang, Yangqiu Song, Dit-Yan Yeung
ACL/IJCNLP (1)5
2021 MultiSiam: Self-supervised Multi-instance Siamese Representation Learning for Autonomous Driving
abstract
Autonomous driving has attracted much attention over the years but turns out to be harder than expected, probably due to the difficulty of labeled data collection for model training. Self-supervised learning (SSL), which leverages unlabeled data only for representation learning, might be a promising way to improve model performance. Existing SSL methods, however, usually rely on the single-centric-object guarantee, which may not be applicable for multi-instance datasets such as street scenes. To alleviate this limitation, we raise two issues to solve: (1) how to define positive samples for cross-view consistency and (2) how to measure similarity in multi-instance circumstances. We first adopt an IoU threshold during random cropping to transfer global-inconsistency to local-consistency. Then, we propose two feature alignment methods to enable 2D feature maps for multi-instance similarity measurement. Addition-ally, we adopt intra-image clustering with self-attention for further mining intra-image similarity and translation-invariance. Experiments show that, when pre-trained on Waymo dataset, our method called Multi-instance Siamese Network (MultiSiam) remarkably improves generalization ability and achieves state-of-the-art transfer performance on autonomous driving benchmarks, including Cityscapes and BDD100K, while existing SSL counterparts like MoCo, MoCo-v2, and BYOL show significant performance drop. By pre-training on SODA10M, a large-scale autonomous driving dataset, MultiSiam exceeds the ImageNet pre-trained MoCo-v2, demonstrating the potential of domain-specific pre-training. Code will be available at https://github.com/KaiChen1998/MultiSiam.
Kai Chen 0023, Lanqing Hong, Hang Xu 0004, Zhenguo Li, Dit-Yan Yeung
ICCV5
2021 MODALS: Modality-agnostic Automated Data Augmentation in the Latent Space
Tsz-Him Cheung, Dit-Yan Yeung
ICLR2
2021 Stereo Waterdrop Removal with Row-wise Dilated Attention
abstract
Existing vision systems for autonomous driving or robots are sensitive to waterdrops adhered to windows or camera lenses. Most recent waterdrop removal approaches take a single image as input and often fail to recover the missing content behind waterdrops faithfully. Thus, we propose a learning-based model for waterdrop removal with stereo images. To better detect and remove waterdrops from stereo images, we propose a novel row-wise dilated attention module to enlarge attention’s receptive field for effective information propagation between the two stereo images. In addition, we propose an attention consistency loss between the ground-truth disparity map and attention scores to enhance the left-right consistency in stereo images. Because of related datasets’ unavailability, we collect a real-world dataset that contains stereo images with and without waterdrops. Extensive experiments on our dataset suggest that our model outperforms state-of-the-art methods both quantitatively and qualitatively. Our source code and the stereo waterdrop dataset are available at https://github.com/VivianSZF/Stereo-Waterdrop-Removal.
Zifan Shi, Na Fan 0002, Dit-Yan Yeung, Qifeng Chen 0001
IROS3
2021 Clickstream Knowledge Tracing: Modeling How Students Answer Interactive Online Questions
abstract
Knowledge tracing (KT) is a research topic which seeks to model the knowledge acquisition process of students by analyzing their past performance in answering questions, based on which their performance in answering future questions is predicted. However, existing KT models only consider whether a student answers a question correctly when the answer is submitted but not the in-question activities. We argue that the interaction involved in the in-question activities can at least partially reveal the thinking process of the student, and hopefully even the competence of acquiring or understanding each piece of the knowledge required for the question.
Wai-Lun Chan, Dit-Yan Yeung
LAK2
2020 Comparative Evaluation of Label-Agnostic Selection Bias in Multilingual Hate Speech Datasets
abstract
Work on bias in hate speech typically aims to improve classification performance while relatively overlooking the quality of the data.We examine selection bias in hate speech in a language and label independent fashion.We first use topic models to discover latent semantics in eleven hate speech corpora, then, we present two bias evaluation metrics based on the semantic similarity between topics and search words frequently used to build corpora.We discuss the possibility of revising the data collection process by comparing datasets and analyzing contrastive case studies.
Nedjma Ousidhoum, Yangqiu Song, Dit-Yan Yeung
EMNLP (1)3
2019 Multilingual and Multi-Aspect Hate Speech Analysis
abstract
Nedjma Ousidhoum, Zizheng Lin, Hongming Zhang, Yangqiu Song, Dit-Yan Yeung. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Nedjma Ousidhoum, Zizheng Lin, Hongming Zhang 0009, Yangqiu Song, Dit-Yan Yeung
EMNLP/IJCNLP (1)5
2019 Marginalized Average Attentional Network for Weakly-Supervised Learning
Yuan Yuan 0002, Yueming Lyu, Xi Shen 0001, Ivor W. Tsang, Dit-Yan Yeung
ICLR (Poster)5
2019 Semi-semantic Line-Cluster Assisted Monocular SLAM for Indoor Environments
Ting Sun 0001, Dezhen Song, Dit-Yan Yeung, Ming Liu 0001
ICVS3
2019 Effective Feature Learning with Unsupervised Learning for Improving the Predictive Models in Massive Open Online Courses
abstract
The effectiveness of learning in massive open online courses (MOOCs) can be significantly enhanced by introducing personalized intervention schemes which rely on building predictive models of student learning behaviors such as some engagement or performance indicators. A major challenge that has to be addressed when building such models is to design handcrafted features that are effective for the prediction task at hand. In this paper, we make the first attempt to solve the feature learning problem by taking the unsupervised learning approach to learn a compact representation of the raw features with a large degree of redundancy. Specifically, in order to capture the underlying learning patterns in the content domain and the temporal nature of the clickstream data, we train a modified auto-encoder (AE) combined with the long short-term memory (LSTM) network to obtain a fixed-length embedding for each input sequence. When compared with the original features, the new features that correspond to the embedding obtained by the modified LSTM-AE are not only more parsimonious but also more discriminative for our prediction task. Using simple supervised learning models, the learned features can improve the prediction accuracy by up to 17% compared with the supervised neural networks and reduce overfitting to the dominant low-performing group of students, specifically in the task of predicting students' performance. Our approach is generic in the sense that it is not restricted to a specific supervised learning model nor a specific prediction task for MOOC learning analytics.
Mucong Ding, Dit-Yan Yeung, Ting-Chuen Pong
LAK3
2019 Knowledge Query Network for Knowledge Tracing: How Knowledge Interacts with Skills
abstract
Knowledge Tracing (KT) is to trace the knowledge of students as they solve a sequence of problems represented by their related skills. This involves abstract concepts of students' states of knowledge and the interactions between those states and skills. Therefore, a KT model is designed to predict whether students will give correct answers and to describe such abstract concepts. However, existing methods either give relatively low prediction accuracy or fail to explain those concepts intuitively. In this paper, we propose a new model called Knowledge Query Network (KQN) to solve these problems. KQN uses neural networks to encode student learning activities into knowledge state and skill vectors, and models the interactions between the two types of vectors with the dot product. Through this, we introduce a novel concept called probabilistic skill similarity that relates the pairwise cosine and Euclidean distances between skill vectors to the odds ratios of the corresponding skills, which makes KQN interpretable and intuitive.
Dit-Yan Yeung
LAK2
2018 Learning Unmanned Aerial Vehicle Control for Autonomous Target Following
abstract
While deep reinforcement learning (RL) methods have achieved unprecedented successes in a range of challenging problems, their applicability has been mainly limited to simulation or game domains due to the high sample complexity of the trial-and-error learning process. However, real-world robotic applications often need a data-efficient learning process with safety-critical constraints. In this paper, we consider the challenging problem of learning unmanned aerial vehicle (UAV) control for tracking a moving target. To acquire a strategy that combines perception and control, we represent the policy by a convolutional neural network. We develop a hierarchical approach that combines a model-free policy gradient method with a conventional feedback proportional-integral-derivative (PID) controller to enable stable learning without catastrophic failure. The neural network is trained by a combination of supervised learning from raw images and reinforcement learning from games of self-play. We show that the proposed approach can learn a target following policy in a simulator efficiently and the learned behavior can be successfully transferred to the DJI quadrotor platform for real-world UAV control.
Tianbo Liu 0001, Chi Zhang 0067, Dit-Yan Yeung, Shaojie Shen
IJCAI4
2018 Addressing two problems in deep knowledge tracing via prediction-consistent regularization
abstract
Knowledge tracing is one of the key research areas for empowering personalized education. It is a task to model students' mastery level of a knowledge component (KC) based on their historical learning trajectories. In recent years, a recurrent neural network model called deep knowledge tracing (DKT) has been proposed to handle the knowledge tracing task and literature has shown that DKT generally outperforms traditional methods. However, through our extensive experimentation, we have noticed two major problems in the DKT model. The first problem is that the model fails to reconstruct the observed input. As a result, even when a student performs well on a KC, the prediction of that KC's mastery level decreases instead, and vice versa. Second, the predicted performance for KCs across time-steps is not consistent. This is undesirable and unreasonable because student's performance is expected to transit gradually over time. To address these problems, we introduce regularization terms that correspond to reconstruction and waviness to the loss function of the original DKT model to enhance the consistency in prediction. Experiments show that the regularized loss function effectively alleviates the two problems without degrading the original task of DKT.1
Chun-Kit Yeung, Dit-Yan Yeung
L@S2
2018 GaAN: Gated Attention Networks for Learning on Large and Spatiotemporal Graphs
Jiani Zhang 0001, Xingjian Shi, Junyuan Xie, Hao Ma 0001, Irwin King, Dit-Yan Yeung
UAI6
2017 Relational Deep Learning: A Deep Latent Variable Model for Link Prediction
abstract
Link prediction is a fundamental task in such areas as social network analysis, information retrieval, and bioinformatics. Usually link prediction methods use the link structures or node attributes as the sources of information. Recently, the relational topic model (RTM) and its variants have been proposed as hybrid methods that jointly model both sources of information and achieve very promising accuracy. However, the representations (features) learned by them are still not effective enough to represent the nodes (items). To address this problem, we generalize recent advances in deep learning from solely modeling i.i.d. sequences of attributes to jointly modeling graphs and non-i.i.d. sequences of attributes. Specifically, we follow the Bayesian deep learning framework and devise a hierarchical Bayesian model, called relational deep learning (RDL), to jointly model high-dimensional node attributes and link structures with layers of latent variables. Due to the multiple nonlinear transformations in RDL, standard variational inference is not applicable. We propose to utilize the product of Gaussians (PoG) structure in RDL to relate the inferences on different variables and derive a generalized variational inference algorithm for learning the variables and predicting the links. Experiments on three real-world datasets show that RDL works surprisingly well and significantly outperforms the state of the art.
Hao Wang 0014, Xingjian Shi, Dit-Yan Yeung
AAAI3
2017 Sparse Boltzmann Machines with Structure Learning as Applied to Text Analysis
abstract
We are interested in exploring the possibility and benefits of structure learning for deep models. As the first step, this paper investigates the matter for Restricted Boltzmann Machines (RBMs). We conduct the study with Replicated Softmax, a variant of RBMs for unsupervised text analysis. We present a method for learning what we call Sparse Boltzmann Machines, where each hidden unit is connected to a subset of the visible units instead of all of them. Empirical results show that the method yields models with significantly improved model fit and interpretability as compared with RBMs where each hidden unit is connected to all visible units.
Zhourong Chen, Nevin Lianwen Zhang, Dit-Yan Yeung, Peixian Chen
AAAI3
2017 Visual Object Tracking for Unmanned Aerial Vehicles: A Benchmark and New Motion Models
abstract
Despite recent advances in the visual tracking community, most studies so far have focused on the observation model. As another important component in the tracking system, the motion model is much less well-explored especially for some extreme scenarios. In this paper, we consider one such scenario in which the camera is mounted on an unmanned aerial vehicle (UAV) or drone. We build a benchmark dataset of high diversity, consisting of 70 videos captured by drone cameras. To address the challenging issue of severe camera motion, we devise simple baselines to model the camera motion by geometric transformation based on background feature points. An extensive comparison of recent state-of-the-art trackers and their motion model variants on our drone tracking dataset validates both the necessity of the dataset and the effectiveness of the proposed methods. Our aim for this work is to lay the foundation for further research in the UAV tracking area.
Dit-Yan Yeung
AAAI2
2017 Lattice Long Short-Term Memory for Human Action Recognition
abstract
Human actions captured in video sequences are threedimensional signals characterizing visual appearance and motion dynamics. To learn action patterns, existing methods adopt Convolutional and/or Recurrent Neural Networks (CNNs and RNNs). CNN based methods are effective in learning spatial appearances, but are limited in modeling long-term motion dynamics. RNNs, especially Long Short- Term Memory (LSTM), are able to learn temporal motion dynamics. However, naively applying RNNs to video sequences in a convolutional manner implicitly assumes that motions in videos are stationary across different spatial locations. This assumption is valid for short-term motions but invalid when the duration of the motion is long.,,In this work, we propose Lattice-LSTM (L2STM), which extends LSTM by learning independent hidden state transitions of memory cells for individual spatial locations. This method effectively enhances the ability to model dynamics across time and addresses the non-stationary issue of long-term motion dynamics without significantly increasing the model complexity. Additionally, we introduce a novel multi-modal training procedure for training our network. Unlike traditional two-stream architectures which use RGB and optical flow information as input, our two-stream model leverages both modalities to jointly train both input gates and both forget gates in the network rather than treating the two streams as separate entities with no information about the other. We apply this end-to-end system to benchmark datasets (UCF-101 and HMDB-51) of human action recognition. Experiments show that on both datasets, our proposed method outperforms all existing ones that are based on LSTM and/or CNNs of similar model complexities.
Lin Sun 0004, Kui Jia, Kevin Chen 0001, Dit-Yan Yeung, Bertram E. Shi, Silvio Savarese
ICCV4
2017 Spatiotemporal Modeling for Crowd Counting in Videos
abstract
Region of Interest (ROI) crowd counting can be formulated as a regression problem of learning a mapping from an image or a video frame to a crowd density map. Recently, convolutional neural network (CNN) models have achieved promising results for crowd counting. However, even when dealing with video data, CNN-based methods still consider each video frame independently, ignoring the strong temporal correlation between neighboring frames. To exploit the otherwise very useful temporal information in video sequences, we propose a variant of a recent deep learning model called convolutional LSTM (ConvLSTM) for crowd counting. Unlike the previous CNN-based methods, our method fully captures both spatial and temporal dependencies. Furthermore, we extend the ConvLSTM model to a bidirectional ConvLSTM model which can access long-range information in both directions. Extensive experiments using four publicly available datasets demonstrate the reliability of our approach and the effectiveness of incorporating temporal information to boost the accuracy of crowd counting. In addition, we also conduct some transfer learning experiments to show that once our model is trained on one dataset, its learning experience can be transferred easily to a new dataset which consists of only very few video frames for model adaptation.
Xingjian Shi, Dit-Yan Yeung
ICCV3
2017 Temporal Dynamic Graph LSTM for Action-Driven Video Object Detection
abstract
In this paper, we investigate a weakly-supervised object detection framework. Most existing frameworks focus on using static images to learn object detectors. However, these detectors often fail to generalize to videos because of the existing domain shift. Therefore, we investigate learning these detectors directly from boring videos of daily activities. Instead of using bounding boxes, we explore the use of action descriptions as supervision since they are relatively easy to gather. A common issue, however, is that objects of interest that are not involved in human actions are often absent in global action descriptions known as “missing label”. To tackle this problem, we propose a novel temporal dynamic graph Long Short-Term Memory network (TDGraph LSTM). TD-Graph LSTM enables global temporal reasoning by constructing a dynamic graph that is based on temporal correlations of object proposals and spans the entire video. The missing label issue for each individual frame can thus be significantly alleviated by transferring knowledge across correlated objects proposals in the whole video. Extensive evaluations on a large-scale daily-life action dataset (i.e., Charades) demonstrates the superiority of our proposed method. We also release object bounding-box annotations for more than 5,000 frames in Charades. We believe this annotated data can also benefit other research on video-based object recognition in the future.
Yuan Yuan 0002, Xiaodan Liang, Xiaolong Wang 0004, Dit-Yan Yeung, Abhinav Gupta 0001
ICCV4
2017 Gesture-based piloting of an aerial robot using monocular vision
abstract
Aerial robots are becoming popular among general public, and with the development of artificial intelligence (AI), there is a trend to equip aerial robots with a natural user interface (NUI). Hand/arm gestures are an intuitive way to communicate for humans, and various research works have focused on controlling an aerial robot with natural gestures. However, the techniques in this area are still far from mature. Many issues in this area have been poorly addressed, such as the principles of choosing gestures from the design point of view, hardware requirements from an economic point of view, considerations of data availability, and algorithm complexity from a practical perspective. Our work focuses on building an economical monocular system particularly designed for gesture-based piloting of an aerial robot. Natural arm gestures are mapped to rich target directions and convenient fine adjustment is achieved. Practical piloting scenarios, hardware cost and algorithm applicability are jointly considered in our system design. The entire system is successfully implemented in an aerial robot and various properties of the system are tested.
Ting Sun 0001, Shengyi Nie, Dit-Yan Yeung, Shaojie Shen
ICRA3
2017 Deep Learning for Precipitation Nowcasting: A Benchmark and A New Model
abstract
With the goal of making high-resolution forecasts of regional rainfall, precipitation nowcasting has become an important and fundamental technology underlying various public services ranging from rainstorm warnings to flight safety. Recently, the Convolutional LSTM (ConvLSTM) model has been shown to outperform traditional optical flow based methods for precipitation nowcasting, suggesting that deep learning models have a huge potential for solving the problem. However, the convolutional recurrence structure in ConvLSTM-based models is location-invariant while natural motion and transformation (e.g., rotation) are location-variant in general. Furthermore, since deep-learning-based precipitation nowcasting is a newly emerging area, clear evaluation protocols have not yet been established. To address these problems, we propose both a new model and a benchmark for precipitation nowcasting. Specifically, we go beyond ConvLSTM and propose the Trajectory GRU (TrajGRU) model that can actively learn the location-variant structure for recurrent connections. Besides, we provide a benchmark that includes a real-world large-scale dataset from the Hong Kong Observatory, a new training loss, and a comprehensive evaluation protocol to facilitate future research and gauge the state of the art.
Xingjian Shi, Zhihan Gao 0001, Leonard Lausen, Hao Wang 0014, Dit-Yan Yeung, Wai-Kin Wong, Wang-chun Woo
NIPS5
2017 Dynamic Key-Value Memory Networks for Knowledge Tracing
abstract
Knowledge Tracing (KT) is a task of tracing evolving knowledge state of students with respect to one or more concepts as they engage in a sequence of learning activities. One important purpose of KT is to personalize the practice sequence to help students learn knowledge concepts efficiently. However, existing methods such as Bayesian Knowledge Tracing and Deep Knowledge Tracing either model knowledge state for each predefined concept separately or fail to pinpoint exactly which concepts a student is good at or unfamiliar with. To solve these problems, this work introduces a new model called Dynamic Key-Value Memory Networks (DKVMN) that can exploit the relationships between underlying concepts and directly output a student's mastery level of each concept. Unlike standard memory-augmented neural networks that facilitate a single memory matrix or two static memory matrices, our model has one static matrix called key, which stores the knowledge concepts and the other dynamic matrix called value, which stores and updates the mastery levels of corresponding concepts. Experiments show that our model consistently outperforms the state-of-the-art model in a range of KT datasets. Moreover, the DKVMN model can automatically discover underlying concepts of exercises typically performed by human annotations and depict the changing knowledge state of a student.
Jiani Zhang 0001, Xingjian Shi, Irwin King, Dit-Yan Yeung
WWW4
2017 Fine-grained categorization via CNN-based automatic extraction and integration of object-level and part-level features
Ting Sun 0001, Lin Sun 0004, Dit-Yan Yeung
Image Vis. Comput.3
2016 Collaborative Recurrent Autoencoder: Recommend while Learning to Fill in the Blanks
abstract
Hybrid methods that utilize both content and rating information are commonly used in many recommender systems. However, most of them use either handcrafted features or the bag-of-words representation as a surrogate for the content information but they are neither effective nor natural enough. To address this problem, we develop a collaborative recurrent autoencoder (CRAE) which is a denoising recurrent autoencoder (DRAE) that models the generation of content sequences in the collaborative filtering (CF) setting. The model generalizes recent advances in recurrent deep learning from i.i.d. input to non-i.i.d. (CF-based) input and provides a new denoising scheme along with a novel learnable pooling scheme for the recurrent autoencoder. To do this, we first develop a hierarchical Bayesian model for the DRAE and then generalize it to the CF setting. The synergy between denoising and CF enables CRAE to make accurate recommendations while learning to fill in the blanks in sequences. Experiments on real-world datasets from different domains (CiteULike and Netflix) show that, by jointly modeling the order-aware generation of sequences for the content information and performing CF for the ratings, CRAE is able to significantly outperform the state of the art on both the recommendation task based on ratings and the sequence generation task based on content information.
Hao Wang 0014, Xingjian Shi, Dit-Yan Yeung
NIPS3
2016 Natural-Parameter Networks: A Class of Probabilistic Neural Networks
abstract
Neural networks (NN) have achieved state-of-the-art performance in various applications. Unfortunately in applications where training data is insufficient, they are often prone to overfitting. One effective way to alleviate this problem is to exploit the Bayesian approach by using Bayesian neural networks (BNN). Another shortcoming of NN is the lack of flexibility to customize different distributions for the weights and neurons according to the data, as is often done in probabilistic graphical models. To address these problems, we propose a class of probabilistic neural networks, dubbed natural-parameter networks (NPN), as a novel and lightweight Bayesian treatment of NN. NPN allows the usage of arbitrary exponential-family distributions to model the weights and neurons. Different from traditional NN and BNN, NPN takes distributions as input and goes through layers of transformation before producing distributions to match the target output distributions. As a Bayesian treatment, efficient backpropagation (BP) is performed to learn the natural parameters for the distributions over both the weights and neurons. The output distributions of each layer, as byproducts, may be used as second-order representations for the associated tasks such as link prediction. Experiments on real-world datasets show that NPN can achieve state-of-the-art performance.
Hao Wang 0014, Xingjian Shi, Dit-Yan Yeung
NIPS3
2016 Spectral Multimodal Hashing and Its Application to Multimedia Retrieval
abstract
In recent years, multimedia retrieval has sparked much research interest in the multimedia, pattern recognition, and data mining communities. Although some attempts have been made along this direction, performing fast multimodal search at very large scale still remains a major challenge in the area. While hashing-based methods have recently achieved promising successes in speeding-up large-scale similarity search, most existing methods are only designed for uni-modal data, making them unsuitable for multimodal multimedia retrieval. In this paper, we propose a new hashing-based method for fast multimodal multimedia retrieval. The method is based on spectral analysis of the correlation matrix of different modalities. We also develop an efficient algorithm that learns some parameters from the data distribution for obtaining the binary codes. We empirically compare our method with some state-of-the-art methods on two real-world multimedia data sets.
Yi Zhen, Yue Gao 0002, Dit-Yan Yeung, Hongyuan Zha, Xuelong Li 0001
IEEE Trans. Cybern.3
2016 Towards Bayesian Deep Learning: A Framework and Some Existing Methods
abstract
While perception tasks such as visual object recognition and text understanding play an important role in human intelligence, subsequent tasks that involve inference, reasoning, and planning require an even higher level of intelligence. The past few years have seen major advances in many perception tasks using deep learning models. For higher-level inference, however, probabilistic graphical models with their Bayesian nature are still more powerful and flexible. To achieve integrated intelligence that involves both perception and inference, it is naturally desirable to tightly integrate deep learning and Bayesian models within a principled probabilistic framework, which we call Bayesian deep learning. In this unified framework, the perception of text or images using deep learning can boost the performance of higher-level inference and in return, the feedback from the inference process is able to enhance the perception of text or images. This paper proposes a general framework for Bayesian deep learning and reviews its recent applications on recommender systems, topic models, and control. In this paper, we also discuss the relationship and differences between Bayesian deep learning and other related topics such as the Bayesian treatment of neural networks.
Hao Wang 0014, Dit-Yan Yeung
IEEE Trans. Knowl. Data Eng.2
2015 Probabilistic Graphical Models for Boosting Cardinal and Ordinal Peer Grading in MOOCs
abstract
With the enormous scale of massive open online courses (MOOCs), peer grading is vital for addressing the assessment challenge for open-ended assignments or exams while at the same time providing students with an effective learning experience through involvement in the grading process. Most existing MOOC platforms use simple schemes for aggregating peer grades, e.g., taking the median or mean. To enhance these schemes, some recent research attempts have developed machine learning methods under either the cardinal setting (for absolute judgment) or the ordinal setting (for relative judgment). In this paper, we seek to study both cardinal and ordinal aspects of peer grading within a common framework. First, we propose novel extensions to some existing probabilistic graphical models for cardi- nal peer grading. Not only do these extensions give su- perior performance in cardinal evaluation, but they also outperform conventional ordinal models in ordinal eval- uation. Next, we combine cardinal and ordinal models by augmenting ordinal models with cardinal predictions as prior. Such combination can achieve further performance boosts in both cardinal and ordinal evaluations, suggesting a new research direction to pursue for peer grading on MOOCs. Extensive experiments have been conducted using real peer grading data from a course called “Science, Technology, and Society in China I” offered by HKUST on the Coursera platform.
Fei Mi, Dit-Yan Yeung
AAAI2
2015 Relational Stacked Denoising Autoencoder for Tag Recommendation
abstract
Tag recommendation has become one of the most important ways of organizing and indexing online resources like articles, movies, and music. Since tagging information is usually very sparse, effective learning of the content representation for these resources is crucial to accurate tag recommendation. Recently, models proposed for tag recommendation, such as collaborative topic regression and its variants, have demonstrated promising accuracy. However, a limitation of these models is that, by using topic models like latent Dirichlet allocation as the key component, the learned representation may not be compact and effective enough. Moreover, since relational data exist as an auxiliary data source in many applications, it is desirable to incorporate such data into tag recommendation models. In this paper, we start with a deep learning model called stacked denoising autoencoder (SDAE) in an attempt to learn more effective content representation. We propose a probabilistic formulation for SDAE and then extend it to a relational SDAE (RSDAE) model. RSDAE jointly performs deep representation learning and relational learning in a principled way under a probabilistic framework. Experiments conducted on three real datasets show that both learning more effective representation and learning from relational data are beneficial steps to take to advance the state of the art.
Hao Wang 0014, Xingjian Shi, Dit-Yan Yeung
AAAI3
2015 Bayesian adaptive matrix factorization with automatic model selection
abstract
Low-rank matrix factorization has long been recognized as a fundamental problem in many computer vision applications. Nevertheless, the reliability of existing matrix factorization methods is often hard to guarantee due to challenges brought by such model selection issues as selecting the noise model and determining the model capacity. We address these two issues simultaneously in this paper by proposing a robust non-parametric Bayesian adaptive matrix factorization (AMF) model. AMF proposes a new noise model built on the Dirichlet process Gaussian mixture model (DP-GMM) by taking advantage of its high flexibility on component number selection and capability of fitting a wide range of unknown noise. AMF also imposes an automatic relevance determination (ARD) prior on the low-rank factor matrices so that the rank can be determined automatically without the need for enforcing any hard constraint. An efficient variational method is then devised for model inference. We compare AMF with state-of-the-art matrix factorization methods based on data sets ranging from synthetic data to real-world application data. From the results, AMF consistently achieves better or comparable performance.
Peixian Chen, Naiyan Wang, Nevin Lianwen Zhang, Dit-Yan Yeung
CVPR4
2015 DevNet: A Deep Event Network for multimedia event detection and evidence recounting
abstract
In this paper, we focus on complex event detection in internet videos while also providing the key evidences of the detection results. Convolutional Neural Networks (CNNs) have achieved promising performance in image classification and action recognition tasks. However, it remains an open problem how to use CNNs for video event detection and recounting, mainly due to the complexity and diversity of video events. In this work, we propose a flexible deep CNN infrastructure, namely Deep Event Network (DevNet), that simultaneously detects pre-defined events and provides key spatial-temporal evidences. Taking key frames of videos as input, we first detect the event of interest at the video level by aggregating the CNN features of the key frames. The pieces of evidences which recount the detection results, are also automatically localized, both temporally and spatially. The challenge is that we only have video level labels, while the key evidences usually take place at the frame levels. Based on the intrinsic property of CNNs, we first generate a spatial-temporal saliency map by back passing through DevNet, which then can be used to find the key frames which are most indicative to the event, as well as to localize the specific spatial position, usually an object, in the frame of the highly indicative area. Experiments on the large scale TRECVID 2014 MEDTest dataset demonstrate the promising performance of our method, both for event detection and evidence recounting.
Chuang Gan 0001, Naiyan Wang, Yi Yang 0001, Dit-Yan Yeung, Alex Hauptmann 0001
CVPR4
2015 Human Action Recognition Using Factorized Spatio-Temporal Convolutional Networks
abstract
Human actions in video sequences are three-dimensional (3D) spatio-temporal signals characterizing both the visual appearance and motion dynamics of the involved humans and objects. Inspired by the success of convolutional neural networks (CNN) for image classification, recent attempts have been made to learn 3D CNNs for recognizing human actions in videos. However, partly due to the high complexity of training 3D convolution kernels and the need for large quantities of training videos, only limited success has been reported. This has triggered us to investigate in this paper a new deep architecture which can handle 3D signals more effectively. Specifically, we propose factorized spatio-temporal convolutional networks (FstCN) that factorize the original 3D convolution kernel learning as a sequential process of learning 2D spatial kernels in the lower layers (called spatial convolutional layers), followed by learning 1D temporal kernels in the upper layers (called temporal convolutional layers). We introduce a novel transformation and permutation operator to make factorization in FstCN possible. Moreover, to address the issue of sequence alignment, we propose an effective training and inference strategy based on sampling multiple video clips from a given action video sequence. We have tested FstCN on two commonly used benchmark datasets (UCF-101 and HMDB-51). Without using auxiliary training videos to boost the performance, FstCN outperforms existing CNN based methods and achieves comparable performance with a recent method that benefits from using auxiliary training videos.
Lin Sun 0004, Kui Jia, Dit-Yan Yeung, Bertram E. Shi
ICCV3
2015 Understanding and Diagnosing Visual Tracking Systems
abstract
Several benchmark datasets for visual tracking research have been created in recent years. Despite their usefulness, whether they are sufficient for understanding and diagnosing the strengths and weaknesses of different trackers remains questionable. To address this issue, we propose a framework by breaking a tracker down into five constituent parts, namely, motion model, feature extractor, observation model, model updater, and ensemble post-processor. We then conduct ablative experiments on each component to study how it affects the overall result. Surprisingly, our findings are discrepant with some common beliefs in the visual tracking research community. We find that the feature extractor plays the most important role in a tracker. On the other hand, although the observation model is the focus of many studies, we find that it often brings no significant improvement. Moreover, the motion model and model updater contain many details that could affect the result. Also, the ensemble post-processor can improve the result substantially when the constituent trackers have high diversity. Based on our findings, we put together some very elementary building blocks to give a basic tracker which is competitive in performance to the state-of-the-art trackers. We believe our framework can provide a solid baseline when conducting controlled experiments for visual tracking research.
Naiyan Wang, Jianping Shi, Dit-Yan Yeung, Jiaya Jia
ICCV3
2015 Collaborative Deep Learning for Recommender Systems
abstract
Collaborative filtering (CF) is a successful approach commonly used by many recommender systems. Conventional CF-based methods use the ratings given to items by users as the sole source of information for learning to make recommendation. However, the ratings are often very sparse in many applications, causing CF-based methods to degrade significantly in their recommendation performance. To address this sparsity problem, auxiliary information such as item content information may be utilized. Collaborative topic regression (CTR) is an appealing recent method taking this approach which tightly couples the two components that learn from two different sources of information. Nevertheless, the latent representation learned by CTR may not be very effective when the auxiliary information is very sparse. To address this problem, we generalize recently advances in deep learning from i.i.d. input to non-i.i.d. (CF-based) input and propose in this paper a hierarchical Bayesian model called collaborative deep learning (CDL), which jointly performs deep representation learning for the content information and collaborative filtering for the ratings (feedback) matrix. Extensive experiments on three real-world datasets from different domains show that CDL can significantly advance the state of the art.
Hao Wang 0014, Naiyan Wang, Dit-Yan Yeung
KDD3
2015 Convolutional LSTM Network: A Machine Learning Approach for Precipitation Nowcasting
abstract
The goal of precipitation nowcasting is to predict the future rainfall intensity in a local region over a relatively short period of time. Very few previous studies have examined this crucial and challenging weather forecasting problem from the machine learning perspective. In this paper, we formulate precipitation nowcasting as a spatiotemporal sequence forecasting problem in which both the input and the prediction target are spatiotemporal sequences. By extending the fully connected LSTM (FC-LSTM) to have convolutional structures in both the input-to-state and state-to-state transitions, we propose the convolutional LSTM (ConvLSTM) and use it to build an end-to-end trainable model for the precipitation nowcasting problem. Experiments show that our ConvLSTM network captures spatiotemporal correlations better and consistently outperforms FC-LSTM and the state-of-the-art operational ROVER algorithm for precipitation nowcasting.
Xingjian Shi, Zhourong Chen, Hao Wang 0014, Dit-Yan Yeung, Wai-Kin Wong, Wang-chun Woo
NIPS4
2015 Instance-specific canonical correlation analysis
Deming Zhai, Yu Zhang 0006, Dit-Yan Yeung, Hong Chang 0001, Xilin Chen 0001, Wen Gao 0001
Neurocomputing3
2014 Ensemble-Based Tracking: Aggregating Crowdsourced Structured Time Series Data
abstract
We study the problem of aggregating the contributions of multiple contributors in a crowdsourcing setting. The data involved is in a form not typically considered in most crowdsourcing tasks, in that the data is structured and has a temporal dimension. In particular, we study the visual tracking problem in which the unknown data to be estimated is in the form of a sequence of bounding boxes representing the trajectory of the target object being tracked. We propose a factorial hidden Markov model (FHMM) for ensemble-based tracking by learning jointly the unknown trajectory of the target and the reliability of each tracker in the ensemble. For efficient online inference of the FHMM, we devise a conditional particle filter algorithm by exploiting the structure of the joint posterior distribution of the hidden variables. Using the largest open benchmark for visual tracking, we empirically compare two ensemble methods constructed from five state-of-the-art trackers with the individual trackers. The promising experimental results provide empirical evidence for our ensemble approach to "get the best of all worlds".
Naiyan Wang, Dit-Yan Yeung
ICML2
2014 Multicategory large margin classification methods: Hinge losses vs. coherence functions
Zhihua Zhang 0004, Cheng Chen 0015, Guang Dai, Wu-Jun Li, Dit-Yan Yeung
Artif. Intell.5
2013 Online Robust Non-negative Dictionary Learning for Visual Tracking
abstract
This paper studies the visual tracking problem in video sequences and presents a novel robust sparse tracker under the particle filter framework. In particular, we propose an online robust non-negative dictionary learning algorithm for updating the object templates so that each learned template can capture a distinctive aspect of the tracked object. Another appealing property of this approach is that it can automatically detect and reject the occlusion and cluttered background in a principled way. In addition, we propose a new particle representation formulation using the Huber loss function. The advantage is that it can yield robust estimation without using trivial templates adopted by previous sparse trackers, leading to faster computation. We also reveal the equivalence between this new formulation and the previous one which uses trivial templates. The proposed tracker is empirically compared with state-of-the-art trackers on some challenging video sequences. Both quantitative and qualitative comparisons show that our proposed tracker is superior and more stable.
Naiyan Wang, Jingdong Wang 0001, Dit-Yan Yeung
ICCV3
2013 Bayesian Robust Matrix Factorization for Image and Video Processing
abstract
Matrix factorization is a fundamental problem that is often encountered in many computer vision and machine learning tasks. In recent years, enhancing the robustness of matrix factorization methods has attracted much attention in the research community. To benefit from the strengths of full Bayesian treatment over point estimation, we propose here a full Bayesian approach to robust matrix factorization. For the generative process, the model parameters have conjugate priors and the likelihood (or noise model) takes the form of a Laplace mixture. For Bayesian inference, we devise an efficient sampling algorithm by exploiting a hierarchical view of the Laplace distribution. Besides the basic model, we also propose an extension which assumes that the outliers exhibit spatial or temporal proximity as encountered in many computer vision applications. The proposed methods give competitive experimental results when compared with several state-of-the-art methods on some benchmark image and video processing tasks.
Naiyan Wang, Dit-Yan Yeung
ICCV2
2013 SCMF: Sparse Covariance Matrix Factorization for Collaborative Filtering
Jianping Shi, Naiyan Wang, Dit-Yan Yeung, Irwin King, Jiaya Jia
IJCAI4
2013 Learning High-Order Task Relationships in Multi-Task Learning
Yu Zhang 0006, Dit-Yan Yeung
IJCAI2
2013 Learning a Deep Compact Image Representation for Visual Tracking
abstract
In this paper, we study the challenging problem of tracking the trajectory of a moving object in a video with possibly very complex background. In contrast to most existing trackers which only learn the appearance of the tracked object online, we take a different approach, inspired by recent advances in deep learning architectures, by putting more emphasis on the (unsupervised) feature learning problem. Specifically, by using auxiliary natural images, we train a stacked denoising autoencoder offline to learn generic image features that are more robust against variations. This is then followed by knowledge transfer from offline training to the online tracking process. Online tracking involves a classification neural network which is constructed from the encoder part of the trained autoencoder as a feature extractor and an additional classification layer. Both the feature extractor and the classifier can be further tuned to adapt to appearance changes of the moving object. Comparison with the state-of-the-art trackers on some challenging benchmark video sequences shows that our deep learning tracker is very efficient as well as more accurate.
Naiyan Wang, Dit-Yan Yeung
NIPS2
2013 Active hashing and its application to image and text retrieval
Yi Zhen, Dit-Yan Yeung
Data Min. Knowl. Discov.2
2013 Multilabel relationship learning
abstract
Multilabel learning problems are commonly found in many applications. A characteristic shared by many multilabel learning problems is that some labels have significant correlations between them. In this article, we propose a novel multilabel learning method, called MultiLabel Relationship Learning (MLRL), which extends the conventional support vector machine by explicitly learning and utilizing the relationships between labels. Specifically, we model the label relationships using a label covariance matrix and use it to define a new regularization term for the optimization problem. MLRL learns the model parameters and the label covariance matrix simultaneously based on a unified convex formulation. To solve the convex optimization problem, we use an alternating method in which each subproblem can be solved efficiently. The relationship between MLRL and two widely used maximum margin methods for multilabel learning is investigated. Moreover, we also propose a semisupervised extension of MLRL, called SSMLRL, to demonstrate how to make use of unlabeled data to help learn the label covariance matrix. Through experiments conducted on some multilabel applications, we find that MLRL not only gives higher classification accuracy but also has better interpretability as revealed by the label covariance matrix.
Yu Zhang 0006, Dit-Yan Yeung
ACM Trans. Knowl. Discov. Data2
2013 A Regularization Approach to Learning Task Relationships in Multitask Learning
abstract
Multitask learning is a learning paradigm that seeks to improve the generalization performance of a learning task with the help of some other related tasks. In this article, we propose a regularization approach to learning the relationships between tasks in multitask learning. This approach can be viewed as a novel generalization of the regularized formulation for single-task learning. Besides modeling positive task correlation, our approach—multitask relationship learning (MTRL)—can also describe negative task correlation and identify outlier tasks based on the same underlying principle. By utilizing a matrix-variate normal distribution as a prior on the model parameters of all tasks, our MTRL method has a jointly convex objective function. For efficiency, we use an alternating method to learn the optimal model parameters for each task as well as the relationships between tasks. We study MTRL in the symmetric multitask learning setting and then generalize it to the asymmetric setting as well. We also discuss some variants of the regularization approach to demonstrate the use of other matrix-variate priors for learning task relationships. Moreover, to gain more insight into our model, we also study the relationships between MTRL and some existing multitask learning methods. Experiments conducted on a toy problem as well as several benchmark datasets demonstrate the effectiveness of MTRL as well as its high interpretability revealed by the task covariance matrix.
Yu Zhang 0006, Dit-Yan Yeung
ACM Trans. Knowl. Discov. Data2
2012 Sparse Probabilistic Relational Projection
abstract
Probabilistic relational PCA (PRPCA) can learn a projection matrix to perform dimensionality reduction for relational data. However, the results learned by PRPCA lack interpretability because each principal component is a linear combination of all the original variables. In this paper, we propose a novel model, called sparse probabilistic relational projection (SPRP), to learn a sparse projection matrix for relational dimensionality reduction. The sparsity in SPRP is achieved by imposing on the projection matrix a sparsity-inducing prior such as the Laplace prior or Jeffreys prior. We propose an expectation-maximization (EM) algorithm to learn the parameters of SPRP. Compared with PRPCA, the sparsity in SPRP not only makes the results more interpretable but also makes the projection operation much more efficient without compromising its accuracy. All these are verified by experiments conducted on several real applications.
Wu-Jun Li, Dit-Yan Yeung
AAAI2
2012 Supervised Probabilistic Robust Embedding with Sparse Noise
abstract
Many noise models do not faithfully reflect the noise processes introduced during data collection in many real-world applications. In particular, we argue that a type of noise referred to as sparse noise is quite commonly found in many applications and many existing works have been proposed to model such sparse noise. However, all the existing works only focus on unsupervised learning without considering the supervised information, i.e., label information. In this paper, we consider how to model and handle sparse noise in the context of embedding high-dimensional data under a probabilistic formulation for supervised learning. We propose a supervised probabilistic robust embedding (SPRE) model in which data are corrupted either by sparse noise or by a combination of Gaussian and sparse noises. By using the Laplace distribution as a prior to model sparse noise, we devise a two-fold variational EM learning algorithm in which the update of model parameters has analytical solution. We report some classification experiments to compare SPRE with several related models.
Yu Zhang 0006, Dit-Yan Yeung, Eric P. Xing
AAAI2
2012 A Probabilistic Approach to Robust Matrix Factorization
Naiyan Wang, Tiansheng Yao, Jingdong Wang 0001, Dit-Yan Yeung
ECCV (7)4
2012 Overlapping community detection via bounded nonnegative matrix tri-factorization
abstract
Complex networks are ubiquitous in our daily life, with the World Wide Web, social networks, and academic citation networks being some of the common examples. It is well understood that modeling and understanding the network structure is of crucial importance to revealing the network functions. One important problem, known as community detection, is to detect and extract the community structure of networks. More recently, the focus in this research topic has been switched to the detection of overlapping communities. In this paper, based on the matrix factorization approach, we propose a method called bounded nonnegative matrix tri-factorization (BNMTF). Using three factors in the factorization, we can explicitly model and learn the community membership of each node as well as the interaction among communities. Based on a unified formulation for both directed and undirected networks, the optimization problem underlying BNMTF can use either the squared loss or the generalized KL-divergence as its loss function. In addition, to address the sparsity problem as a result of missing edges, we also propose another setting in which the loss function is defined only on the observed edges. We report some experiments on real-world datasets to demonstrate the superiority of BNMTF over other related matrix factorization methods.
Yu Zhang 0006, Dit-Yan Yeung
KDD2
2012 A probabilistic model for multimodal hash function learning
abstract
In recent years, both hashing-based similarity search and multimodal similarity search have aroused much research interest in the data mining and other communities. While hashing-based similarity search seeks to address the scalability issue, multimodal similarity search deals with applications in which data of multiple modalities are available. In this paper, our goal is to address both issues simultaneously. We propose a probabilistic model, called multimodal latent binary embedding (MLBE), to learn hash functions from multimodal data automatically. MLBE regards the binary latent factors as hash codes in a common Hamming space. Given data from multiple modalities, we devise an efficient algorithm for the learning of binary latent factors which corresponds to hash function learning. Experimental validation of MLBE has been conducted using both synthetic data and two realistic data sets. Experimental results show that MLBE compares favorably with two state-of-the-art models.
Yi Zhen, Dit-Yan Yeung
KDD2
2012 Co-Regularized Hashing for Multimodal Data
abstract
Hashing-based methods provide a very promising approach to large-scale similarity search. To obtain compact hash codes, a recent trend seeks to learn the hash functions from data automatically. In this paper, we study hash function learning in the context of multimodal data. We propose a novel multimodal hash function learning method, called Co-Regularized Hashing (CRH), based on a boosted co-regularization framework. The hash functions for each bit of the hash codes are learned by solving DC (difference of convex functions) programs, while the learning for multiple bits proceeds via a boosting procedure so that the bias introduced by the hash functions can be sequentially minimized. We empirically compare CRH with two state-of-the-art multimodal hash function learning methods on two publicly available data sets.
Yi Zhen, Dit-Yan Yeung
NIPS2
2012 Multi-Task Boosting by Exploiting Task Relationships
Yu Zhang 0006, Dit-Yan Yeung
ECML/PKDD (1)2
2012 Transfer Metric Learning with Semi-Supervised Extension
abstract
Distance metric learning plays a very crucial role in many data mining algorithms because the performance of an algorithm relies heavily on choosing a good metric. However, the labeled data available in many applications is scarce, and hence the metrics learned are often unsatisfactory. In this article, we consider a transfer-learning setting in which some related source tasks with labeled data are available to help the learning of the target task. We first propose a convex formulation for multitask metric learning by modeling the task relationships in the form of a task covariance matrix. Then we regard transfer learning as a special case of multitask learning and adapt the formulation of multitask metric learning to the transfer-learning setting for our method, called transfer metric learning (TML). In TML, we learn the metric and the task covariances between the source tasks and the target task under a unified convex formulation. To solve the convex optimization problem, we use an alternating method in which each subproblem has an efficient solution. Moreover, in many applications, some unlabeled data is also available in the target task, and so we propose a semi-supervised extension of TML called STML to further improve the generalization performance by exploiting the unlabeled data based on the manifold assumption. Experimental results on some commonly used transfer-learning applications demonstrate the effectiveness of our method.
Yu Zhang 0006, Dit-Yan Yeung
ACM Trans. Intell. Syst. Technol.2
2012 Solution Path for Manifold Regularized Semisupervised Classification
abstract
Traditional learning algorithms use only labeled data for training. However, labeled examples are often difficult or time consuming to obtain since they require substantial human labeling efforts. On the other hand, unlabeled data are often relatively easy to collect. Semisupervised learning addresses this problem by using large quantities of unlabeled data with labeled data to build better learning algorithms. In this paper, we use the manifold regularization approach to formulate the semisupervised learning problem where a regularization framework which balances a tradeoff between loss and penalty is established. We investigate different implementations of the loss function and identify the methods which have the least computational expense. The regularization hyperparameter, which determines the balance between loss and penalty, is crucial to model selection. Accordingly, we derive an algorithm that can fit the entire path of solutions for every value of the hyperparameter. Its computational complexity after preprocessing is quadratic only in the number of labeled examples rather than the total number of labeled and unlabeled examples.
Gang Wang 0004, Fei Wang 0001, Dit-Yan Yeung, Frederick H. Lochovsky
IEEE Trans. Syst. Man Cybern. Part B4
2011 Social Relations Model for Collaborative Filtering
abstract
We propose a novel probabilistic model for collaborative filtering (CF), called SRMCoFi, which seamlessly integrates both linear and bilinear random effects into a principled framework. The formulation of SRMCoFi is supported by both social psychological experiments and statistical theories. Not only can many existing CF methods be seen as special cases of SRMCoFi, but it also integrates their advantages while simultaneously overcoming their disadvantages. The solid theoretical foundation of SRMCoFi is further supported by promising empirical results obtained in extensive experiments using real CF data sets on movie ratings.
Wu-Jun Li, Dit-Yan Yeung
AAAI2
2011 Multi-Task Learning in Heterogeneous Feature Spaces
abstract
Multi-task learning aims at improving the generalization performance of a learning task with the help of some other related tasks. Although many multi-task learning methods have been proposed, they are all based on the assumption that all tasks share the same data representation. This assumption is too restrictive for general applications. In this paper, we propose a multi-task extension of linear discriminant analysis (LDA), called multi-task discriminant analysis (MTDA), which can deal with learning tasks with different data representations. For each task, MTDA learns a separate transformation which consists of two parts, one specific to the task and one common to all tasks. A by-product of MTDA is that it can alleviate the labeled data deficiency problem of LDA. Moreover, unlike many existing multi-task learning methods, MTDA can handle binary and multi-class problems for each task in a generic way. Experimental results on face recognition show that MTDA consistently outperforms related methods.
Yu Zhang 0006, Dit-Yan Yeung
AAAI2
2011 A Convex Formulation of Modularity Maximization for Community Detection
Emprise Y. K. Chan, Dit-Yan Yeung
IJCAI2
2011 Generalized Latent Factor Models for Social Network Analysis
Wu-Jun Li, Dit-Yan Yeung, Zhihua Zhang 0004
IJCAI2
2011 Discriminative Experimental Design
Yu Zhang 0006, Dit-Yan Yeung
ECML/PKDD (3)2
2011 Semisupervised Generalized Discriminant Analysis
abstract
Generalized discriminant analysis (GDA) is a commonly used method for dimensionality reduction. In its general form, it seeks a nonlinear projection that simultaneously maximizes the between-class dissimilarity and minimizes the within-class dissimilarity to increase class separability. In real-world applications where labeled data are scarce, GDA may not work very well. However, unlabeled data are often available in large quantities at very low cost. In this paper, we propose a novel GDA algorithm which is abbreviated as semisupervised generalized discriminant analysis (SSGDA). We utilize unlabeled data to maximize an optimality criterion of GDA and formulate the problem as an optimization problem that is solved using the constrained concave-convex procedure. The optimization procedure leads to estimation of the class labels for the unlabeled data. We propose a novel confidence measure and a method for selecting those unlabeled data points whose labels are estimated with high confidence. The selected unlabeled data can then be used to augment the original labeled dataset for performing GDA. We also propose a variant of SSGDA, called M-SSGDA, which adopts the manifold assumption to utilize the unlabeled data. Extensive experiments on many benchmark datasets demonstrate the effectiveness of our proposed methods.
Yu Zhang 0006, Dit-Yan Yeung
IEEE Trans. Neural Networks2
2010 Adaptive Transfer Learning
abstract
Transfer learning aims at reusing the knowledge in some source tasks to improve the learning of a target task. Many transfer learning methods assume that the source tasks and the target task be related, even though many tasks are not related in reality. However, when two tasks are unrelated, the knowledge extracted from a source task may not help, and even hurt, the performance of a target task. Thus, how to avoid negative transfer and then ensure a "safe transfer" of knowledge is crucial in transfer learning. In this paper, we propose an Adaptive Transfer learning algorithm based on Gaussian Processes (AT-GP), which can be used to adapt the transfer learning schemes by automatically estimating the similarity between a source and a target task. The main contribution of our work is that we propose a new semi-parametric transfer kernel for transfer learning from a Bayesian perspective, and propose to learn the model with respect to the target task, rather than all tasks as in multi-task learning. We can formulate the transfer learning problem as a unified Gaussian Process (GP) model. The adaptive transfer ability of our approach is verified on both synthetic and real-world datasets.
Bin Cao 0001, Sinno Jialin Pan, Yu Zhang 0006, Dit-Yan Yeung, Qiang Yang 0001
AAAI4
2010 Transductive Learning on Adaptive Graphs
abstract
Graph-based semi-supervised learning methods are based on some smoothness assumption about the data. As a discrete approximation of the data manifold, the graph plays a crucial role in the success of such graph-based methods. In most existing methods, graph construction makes use of a predefined weighting function without utilizing label information even when it is available. In this work, by incorporating label information, we seek to enhance the performance of graph-based semi-supervised learning by learning the graph and label inference simultaneously. In particular, we consider a particular setting of semi-supervised learning called transductive learning. Using the LogDet divergence to define the objective function, we propose an iterative algorithm to solve the optimization problem which has closed-form solution in each step. We perform experiments on both synthetic and real data to demonstrate improvement in the graph and in terms of classification accuracy.
Yan-Ming Zhang 0001, Yu Zhang 0006, Dit-Yan Yeung, Cheng-Lin Liu 0001, Xinwen Hou
AAAI3
2010 Gaussian Process Latent Random Field
abstract
In this paper, we propose a novel supervised extension of GPLVM, called Gaussian process latent random field (GPLRF), by enforcing the latent variables to be a Gaussian Markov random field with respect to a graph constructed from the supervisory information.
Guoqiang Zhong 0001, Wu-Jun Li, Dit-Yan Yeung, Xinwen Hou, Cheng-Lin Liu 0001
AAAI3
2010 Multi-task warped Gaussian process for personalized age estimation
abstract
Automatic age estimation from facial images has aroused research interests in recent years due to its promising potential for some computer vision applications. Among the methods proposed to date, personalized age estimation methods generally outperform global age estimation methods by learning a separate age estimator for each person in the training data set. However, since typical age databases only contain very limited training data for each person, training a separate age estimator using only training data for that person runs a high risk of overfitting the data and hence the prediction performance is limited. In this paper, we propose a novel approach to age estimation by formulating the problem as a multi-task learning problem. Based on a variant of the Gaussian process (GP) called warped Gaussian process (WGP), we propose a multi-task extension called multi-task warped Gaussian process (MTWGP). Age estimation is formulated as a multi-task regression problem in which each learning task refers to estimation of the age function for each person. While MTWGP models common features shared by different tasks (persons), it also allows task-specific (person-specific) features to be learned automatically. Moreover, unlike previous age estimation methods which need to specify the form of the regression functions or determine many parameters in the functions using inefficient methods such as cross validation, the form of the regression functions in MTWGP is implicitly defined by the kernel function and all its model parameters can be learned from data automatically. We have conducted experiments on two publicly available age databases, FG-NET and MORPH. The experimental results are very promising in showing that MTWGP compares favorably with state-of-the-art age estimation methods.
Yu Zhang 0006, Dit-Yan Yeung
CVPR2
2010 Transfer metric learning by learning task relationships
abstract
Distance metric learning plays a very crucial role in many data mining algorithms because the performance of an algorithm relies heavily on choosing a good metric. However, the labeled data available in many applications is scarce and hence the metrics learned are often unsatisfactory. In this paper, we consider a transfer learning setting in which some related source tasks with labeled data are available to help the learning of the target task. We first propose a convex formulation for multi-task metric learning by modeling the task relationships in the form of a task covariance matrix. Then we regard transfer learning as a special case of multi-task learning and adapt the formulation of multi-task metric learning to the transfer learning setting for our method, called transfer metric learning (TML). In TML, we learn the metric and the task covariances between the source tasks and the target task under a unified convex formulation. To solve the convex optimization problem, we use an alternating method in which each subproblem has an efficient solution. Experimental results on some commonly used transfer learning applications demonstrate the effectiveness of our method.
Yu Zhang 0006, Dit-Yan Yeung
KDD2
2010 Worst-Case Linear Discriminant Analysis
abstract
Dimensionality reduction is often needed in many applications due to the high dimensionality of the data involved. In this paper, we first analyze the scatter measures used in the conventional linear discriminant analysis~(LDA) model and note that the formulation is based on the average-case view. Based on this analysis, we then propose a new dimensionality reduction method called worst-case linear discriminant analysis~(WLDA) by defining new between-class and within-class scatter measures. This new model adopts the worst-case view which arguably is more suitable for applications such as classification. When the number of training data points or the number of features is not very large, we relax the optimization problem involved and formulate it as a metric learning problem. Otherwise, we take a greedy approach by finding one direction of the transformation at a time. Moreover, we also analyze a special case of WLDA to show its relationship with conventional LDA. Experiments conducted on several benchmark datasets demonstrate the effectiveness of WLDA when compared with some related dimensionality reduction methods.
Yu Zhang 0006, Dit-Yan Yeung
NIPS2
2010 Probabilistic Multi-Task Feature Selection
abstract
Recently, some variants of the $l_1$ norm, particularly matrix norms such as the $l_{1,2}$ and $l_{1,\infty}$ norms, have been widely used in multi-task learning, compressed sensing and other related areas to enforce sparsity via joint regularization. In this paper, we unify the $l_{1,2}$ and $l_{1,\infty}$ norms by considering a family of $l_{1,q}$ norms for $1 < q\le\infty$ and study the problem of determining the most appropriate sparsity enforcing norm to use in the context of multi-task feature selection. Using the generalized normal distribution, we provide a probabilistic interpretation of the general multi-task feature selection problem using the $l_{1,q}$ norm. Based on this probabilistic interpretation, we develop a probabilistic model using the noninformative Jeffreys prior. We also extend the model to learn and exploit more general types of pairwise relationships between tasks. For both versions of the model, we devise expectation-maximization~(EM) algorithms to learn all model parameters, including $q$, automatically. Experiments have been conducted on two cancer classification applications using microarray gene expression data.
Yu Zhang 0006, Dit-Yan Yeung
NIPS2
2010 SED: supervised experimental design and its application to text classification
abstract
In recent years, active learning methods based on experimental design achieve state-of-the-art performance in text classification applications. Although these methods can exploit the distribution of unlabeled data and support batch selection, they cannot make use of labeled data which often carry useful information for active learning. In this paper, we propose a novel active learning method for text classification, called supervised experimental design (SED), which seamlessly incorporates label information into experimental design. Experimental results show that SED outperforms its counterparts which either discard the label information even when it is available or fail to exploit the distribution of unlabeled data.
Yi Zhen, Dit-Yan Yeung
SIGIR2
2010 Multi-Domain Collaborative Filtering
Yu Zhang 0006, Bin Cao 0001, Dit-Yan Yeung
UAI3
2010 A Convex Formulation for Learning Task Relationships in Multi-Task Learning
Yu Zhang 0006, Dit-Yan Yeung
UAI2
2010 A regularization framework for multiclass classification: A deterministic annealing approach
Zhihua Zhang 0004, Gang Wang 0004, Dit-Yan Yeung, Guang Dai, Frederick H. Lochovsky
Pattern Recognit.3
2010 MILD: Multiple-Instance Learning via Disambiguation
abstract
In multiple-instance learning (MIL), an individual example is called an instance and a bag contains a single or multiple instances. The class labels available in the training set are associated with bags rather than instances. A bag is labeled positive if at least one of its instances is positive; otherwise, the bag is labeled negative. Since a positive bag may contain some negative instances in addition to one or more positive instances, the true labels for the instances in a positive bag may or may not be the same as the corresponding bag label and, consequently, the instance labels are inherently ambiguous. In this paper, we propose a very efficient and robust MIL method, called Multiple-Instance Learning via Disambiguation (MILD), for general MIL problems. First, we propose a novel disambiguation method to identify the true positive instances in the positive bags. Second, we propose two feature representation schemes, one for instance-level classification and the other for bag-level classification, to convert the MIL problem into a standard single-instance learning (SIL) problem that can be solved by well-known SIL algorithms, such as support vector machine. Third, an inductive semi-supervised learning method is proposed for MIL. We evaluate our methods extensively on several challenging MIL applications to demonstrate their promising efficiency, robustness, and accuracy.
Wu-Jun Li, Dit-Yan Yeung
IEEE Trans. Knowl. Data Eng.2
2009 Localized content-based image retrieval through evidence region identification
abstract
Over the past decade, multiple-instance learning (MIL) has been successfully utilized to model the localized content-based image retrieval (CBIR) problem, in which a bag corresponds to an image and an instance corresponds to a region in the image. However, existing feature representation schemes are not effective enough to describe the bags in MIL, which hinders the adaptation of sophisticated single-instance learning (SIL) methods for MIL problems. In this paper, we first propose an evidence region (or evidence instance) identification method to identify the evidence regions supporting the labels of the images (i.e., bags). Then, based on the identified evidence regions, a very effective feature representation scheme, which is also very computationally efficient and robust to labeling noise, is proposed to describe the bags. As a result, the MIL problem is converted into a standard SIL problem and a support vector machine (SVM) can be easily adapted for localized CBIR. Experimental results on two challenging data sets show that our method, called EC-SVM, can outperform the state-of-the-art methods in terms of accuracy, robustness and efficiency.
Wu-Jun Li, Dit-Yan Yeung
CVPR2
2009 Relation Regularized Matrix Factorization
Wu-Jun Li, Dit-Yan Yeung
IJCAI2
2009 Probabilistic Relational PCA
abstract
One crucial assumption made by both principal component analysis (PCA) and probabilistic PCA (PPCA) is that the instances are independent and identically distributed (i.i.d.). However, this common i.i.d. assumption is unreasonable for relational data. In this paper, by explicitly modeling covariance between instances as derived from the relational information, we propose a novel probabilistic dimensionality reduction method, called probabilistic relational PCA (PRPCA), for relational data analysis. Although the i.i.d. assumption is no longer adopted in PRPCA, the learning algorithms for PRPCA can still be devised easily like those for PPCA which makes explicit use of the i.i.d. assumption. Experiments on real-world data sets show that PRPCA can effectively utilize the relational information to dramatically outperform PCA and achieve state-of-the-art performance.
Wu-Jun Li, Dit-Yan Yeung, Zhihua Zhang 0004
NIPS2
2009 Heteroscedastic Probabilistic Linear Discriminant Analysis with Semi-supervised Extension
Yu Zhang 0006, Dit-Yan Yeung
ECML/PKDD (2)2
2009 Semi-Supervised Multi-Task Regression
Yu Zhang 0006, Dit-Yan Yeung
ECML/PKDD (2)2
2009 TagiCoFi: tag informed collaborative filtering
abstract
Besides the rating information, an increasing number of modern recommender systems also allow the users to add personalized tags to the items. Such tagging information may provide very useful information for item recommendation, because the users' interests in items can be implicitly reflected by the tags that they often use. Although some content-based recommender systems have made preliminary attempts recently to utilize tagging information to improve the recommendation performance, few recommender systems based on collaborative filtering (CF) have employed tagging information to help the item recommendation procedure. In this paper, we propose a novel framework, called tag informed collaborative filtering (TagiCoFi), to seamlessly integrate tagging information into the CF procedure. Experimental results demonstrate that TagiCoFi outperforms its counterpart which discards the tagging information even when it is available, and achieves state-of-the-art performance.
Yi Zhen, Wu-Jun Li, Dit-Yan Yeung
RecSys3
2008 Human action recognition using Local Spatio-Temporal Discriminant Embedding
abstract
Human action video sequences can be considered as nonlinear dynamic shape manifolds in the space of image frames. In this paper, we address learning and classifying human actions on embedded low-dimensional manifolds. We propose a novel manifold embedding method, called Local Spatio-Temporal Discriminant Embedding (LSTDE). The discriminating capabilities of the proposed method are two-fold: (1) for local spatial discrimination, LSTDE projects data points (silhouette-based image frames of human action sequences) in a local neighborhood into the embedding space where data points of the same action class are close while those of different classes are far apart; (2) in such a local neighborhood, each data point has an associated short video segment, which forms a local temporal subspace on the embedded manifold. LSTDE finds an optimal embedding which maximizes the principal angles between those temporal subspaces associated with data points of different classes. Benefiting from the joint spatio-temporal discriminant embedding, our method is potentially more powerful for classifying human actions with similar space-time shapes, and is able to perform recognition on a frame-by-frame or short video segment basis. Experimental results demonstrate that our method can accurately recognize human actions, and can improve the recognition performance over some representative manifold embedding methods, especially on highly confusing human action types.
Kui Jia, Dit-Yan Yeung
CVPR2
2008 Semi-Supervised Discriminant Analysis using robust path-based similarity
abstract
Linear discriminant analysis (LDA), which works by maximizing the within-class similarity and minimizing the between-class similarity simultaneously, is a popular dimensionality reduction technique in pattern recognition and machine learning. In real-world applications when labeled data are limited, LDA does not work well. Under many situations, however, it is easy to obtain unlabeled data in large quantities. In this paper, we propose a novel dimensionality reduction method, called semi-supervised discriminant analysis (SSDA), which can utilize both labeled and unlabeled data to perform dimensionality reduction in the semi-supervised setting. Our method uses a robust path-based similarity measure to capture the manifold structure of the data and then uses the obtained similarity to maximize the separability between different classes. A kernel extension of the proposed method for nonlinear dimensionality reduction in the semi-supervised setting is also presented. Experiments on face recognition demonstrate the effectiveness of the proposed method.
Yu Zhang 0006, Dit-Yan Yeung
CVPR2
2008 Learning Two-View Stereo Matching
Jianxiong Xiao, Jingni Chen, Dit-Yan Yeung, Long Quan
ECCV (3)3
2008 Structuring Visual Words in 3D for Arbitrary-View Object Localization
Jianxiong Xiao, Jingni Chen, Dit-Yan Yeung, Long Quan
ECCV (3)3
2008 Posterior Consistency of the Silverman g-prior in Bayesian Model Choice
abstract
Kernel supervised learning methods can be unified by utilizing the tools from regularization theory. The duality between regularization and prior leads to interpreting regularization methods in terms of maximum a posteriori estimation and has motivated Bayesian interpretations of kernel methods. In this paper we pursue a Bayesian interpretation of sparsity in the kernel setting by making use of a mixture of a point-mass distribution and prior that we refer to as ``Silverman's g-prior.'' We provide a theoretical analysis of the posterior consistency of a Bayesian model choice procedure based on this prior. We also establish the asymptotic relationship between this procedure and the Bayesian information criterion.
Zhihua Zhang 0004, Michael I. Jordan, Dit-Yan Yeung
NIPS3
2008 Semi-supervised Discriminant Analysis Via CCCP
Yu Zhang 0006, Dit-Yan Yeung
ECML/PKDD (2)2
2008 A Scalable Kernel-Based Semisupervised Metric Learning Algorithm with Out-of-Sample Generalization Ability
abstract
In recent years, metric learning in the semisupervised setting has aroused a lot of research interest. One type of semisupervised metric learning utilizes supervisory information in the form of pairwise similarity or dissimilarity constraints. However, most methods proposed so far are either limited to linear metric learning or unable to scale well with the data set size. In this letter, we propose a nonlinear metric learning method based on the kernel approach. By applying low-rank approximation to the kernel matrix, our method can handle significantly larger data sets. Moreover, our low-rank approximation scheme can naturally lead to out-of-sample generalization. Experiments performed on both artificial and real-world data show very promising results.
Dit-Yan Yeung, Hong Chang 0001, Guang Dai
Neural Comput.1
2008 Robust path-based spectral clustering
Hong Chang 0001, Dit-Yan Yeung
Pattern Recognit.2
2008 A New Solution Path Algorithm in Support Vector Regression
abstract
In this paper, regularization path algorithms were proposed as a novel approach to the model selection problem by exploring the path of possibly all solutions with respect to some regularization hyperparameter in an efficient way. This approach was later extended to a support vector regression (SVR) model called epsilon-SVR. However, the method requires that the error parameter epsilon be set a priori. This is only possible if the desired accuracy of the approximation can be specified in advance. In this paper, we analyze the solution space for epsilon-SVR and propose a new solution path algorithm, called epsilon-path algorithm, which traces the solution path with respect to the hyperparameter epsilon rather than lambda. Although both two solution path algorithms possess the desirable piecewise linearity property, our epsilon-path algorithm overcomes some limitations of the original lambda-path algorithm and has more advantages. It is thus more appealing for practical use.
Gang Wang 0004, Dit-Yan Yeung, Frederick H. Lochovsky
IEEE Trans. Neural Networks2
2007 Image Hallucination Using Neighbor Embedding over Visual Primitive Manifolds
abstract
In this paper, we propose a novel learning-based method for image hallucination, with image super-resolution being a specific application that we focus on here. Given a low-resolution image, its underlying higher-resolution details are synthesized based on a set of training images. In order to build a compact yet descriptive training set, we investigate the characteristic local structures contained in large volumes of small image patches. Inspired by progress in manifold learning research, we take the assumption that small image patches in the low-resolution and high-resolution images form manifolds with similar local geometry in the corresponding image feature spaces. This assumption leads to a super-resolution approach which reconstructs the feature vector corresponding to an image patch by its neighbors in the feature space. In addition, the residual errors associated with the reconstructed image patches are also estimated to compensate for the information loss in the local averaging process. Experimental results show that our hallucination method can synthesize higher-quality images compared with other methods.
Dit-Yan Yeung
CVPR2
2007 Locally Smooth Metric Learning with Application to Image Retrieval
abstract
In this paper, we propose a novel metric learning method based on regularized moving least squares. Unlike most previous metric learning methods which learn a global Mahalanobis distance, we define locally smooth metrics using local affine transformations which are more flexible. The data set after metric learning can preserve the original topological structures. Moreover, our method is fairly efficient and may be used as a preprocessing step for various subsequent learning tasks, including classification, clustering, and nonlinear dimensionality reduction. In particular, we demonstrate that our method can boost the performance of content-based image retrieval (CBIR) tasks. Experimental results provide empirical evidence for the effectiveness of our approach.
Dit-Yan Yeung, Hong Chang 0001
ICCV1
2007 Kernel selection forl semi-supervised kernel machines
abstract
Existing semi-supervised learning methods are mostly based on either the cluster assumption or the manifold assumption. In this paper, we propose an integrated regularization framework for semi-supervised kernel machines by incorporating both the cluster assumption and the manifold assumption. Moreover, it supports kernel learning in the form of kernel selection. The optimization problem involves joint optimization over all the labeled and unlabeled data points, a convex set of basic kernels, and a discrete space of unknown labels for the unlabeled data. When the manifold assumption is incorporated, graph Laplacian kernels are used as the basic kernels for learning an optimal convex combination of graph Laplacian kernels. Comparison with related methods on the USPS data set shows very promising results.
Guang Dai, Dit-Yan Yeung
ICML2
2007 A kernel path algorithm for support vector machines
abstract
The choice of the kernel function which determines the mapping between the input space and the feature space is of crucial importance to kernel methods. The past few years have seen many efforts in learning either the kernel function or the kernel matrix. In this paper, we address this model selection issue by learning the hyperparameter of the kernel function for a support vector machine (SVM). We trace the solution path with respect to the kernel hyperparameter without having to train the model multiple times. Given a kernel hyperparameter value and the optimal solution obtained for that value, we find that the solutions of the neighborhood hyperparameters can be calculated exactly. However, the solution path does not exhibit piecewise linearity and extends nonlinearly. As a result, the breakpoints cannot be computed in advance. We propose a method to approximate the breakpoints. Our method is both efficient and general in the sense that it can be applied to many kernel functions in common use.
Gang Wang 0004, Dit-Yan Yeung, Frederick H. Lochovsky
ICML2
2007 Boosting Kernel Discriminant Analysis and Its Application on Tissue Classification of Gene Expression Data
Guang Dai, Dit-Yan Yeung
IJCAI2
2007 A Scalable Kernel-Based Algorithm for Semi-Supervised Metric Learning
Dit-Yan Yeung, Hong Chang 0001, Guang Dai
IJCAI1
2007 Semi-supervised Cast Indexing for Feature-Length Films
Tao Wang 0003, Jean-Yves Bouguet, Wei Hu 0002, Yimin Zhang 0002, Dit-Yan Yeung
MMM (1)6
2007 Kernel-based distance metric learning for content-based image retrieval
Hong Chang 0001, Dit-Yan Yeung
Image Vis. Comput.2
2007 Surrogate maximization/minimization algorithms and extensions
Zhihua Zhang 0004, James T. Kwok, Dit-Yan Yeung
Mach. Learn.3
2007 Face recognition using a kernel fractional-step discriminant analysis algorithm
Guang Dai, Dit-Yan Yeung, Yuntao Qian
Pattern Recognit.2
2007 Learning the kernel matrix by maximizing a KFD-based class separability criterion
Dit-Yan Yeung, Hong Chang 0001, Guang Dai
Pattern Recognit.1
2007 A Kernel Approach for Semisupervised Metric Learning
abstract
While distance function learning for supervised learning tasks has a long history, extending it to learning tasks with weaker supervisory information has only been studied recently. In particular, some methods have been proposed for semisupervised metric learning based on pairwise similarity or dissimilarity information. In this paper, we propose a kernel approach for semisupervised metric learning and present in detail two special cases of this kernel approach. The metric learning problem is thus formulated as an optimization problem for kernel learning. An attractive property of the optimization problem is that it is convex and, hence, has no local optima. While a closed-form solution exists for the first special case, the second case is solved using an iterative majorization procedure to estimate the optimal solution asymptotically. Experimental results based on both synthetic and real-world data show that this new kernel approach is promising for nonlinear metric learning.
Dit-Yan Yeung, Hong Chang 0001
IEEE Trans. Neural Networks1
2006 Tensor Embedding Methods
Guang Dai, Dit-Yan Yeung
AAAI2
2006 A Manifold Regularization Approach to Calibration Reduction for Sensor-Network Based Tracking
Jeffrey Junfeng Pan, Qiang Yang 0001, Hong Chang 0001, Dit-Yan Yeung
AAAI4
2006 Graph Laplacian Kernels for Object Classification from a Single Example
abstract
Classification with only one labeled example per class is a challenging problem in machine learning and pattern recognition. While there have been some attempts to address this problem in the context of specific applications, very little work has been done so far on the problem under more general object classification settings. In this paper, we propose a graph-based approach to the problem. Based on a robust path-based similarity measure proposed recently, we construct a weighted graph using the robust path-based similarities as edge weights. A kernel matrix, called graph Laplacian kernel, is then defined based on the graph Laplacian. With the kernel matrix, in principle any kernel-based classifier can be used for classification. In particular, we demonstrate the use of a kernel nearest neighbor classifier on some synthetic data and real-world image sets, showing that our method can successfully solve some difficult classification tasks with only very few labeled examples.
Hong Chang 0001, Dit-Yan Yeung
CVPR (2)2
2006 Locally Linear Models on Face Appearance Manifolds with Application to Dual-Subspace Based Classification
abstract
Recently, there has been a flurry of research on face recognition based on multiple images or shots from either a video sequence or an image set. This paper is also such an attempt in multiple-shot face recognition. Specifically, we propose a novel nonparametric method that first extracts discriminating local models via clustering. We apply a hierarchical distance-based clustering procedure according to some distance measure on the appearance manifold to cluster similar face images together. Based on the local models extracted, we then construct the intrapersonal and extrapersonal subspaces. Given a new test image, the angle between the projections of the image onto the two subspaces is used as a distance measure for classification. Since a test example contains multiple face images in multiple-shot face recognition, the final classification combines the classification decisions of all individual test images via a majority voting scheme. We compare our method empirically with some previous methods based on a database of video sequences of human faces, showing that out method significantly outperforms other methods.
Dit-Yan Yeung
CVPR (2)2
2006 Extending Kernel Fisher Discriminant Analysis with the Weighted Pairwise Chernoff Criterion
Guang Dai, Dit-Yan Yeung, Hong Chang 0001
ECCV (4)2
2006 A Coordinated Detection and Response Scheme for Distributed Denial-of-Service Attacks
abstract
Distributed denial-of-service (DDoS) attacks present serious threats to servers in the Internet. They can exhaust critical resources at a target host with the help of a large number of compromised Internet hosts and hence deny services to legitimate clients. This paper studies some existing schemes for the detection and defense against TCP-based DDoS attacks. We propose a distributed scheme that can mitigate the damage caused by DDoS through a coordinated detection and response framework. This proposed scheme composes of a number of heterogeneous defense systems which cooperate with each other in protecting Internet servers. We have set up a network testbed for carrying out extensive experiments using real server machines, routers and software attack tools. Experimental results show that, compared to existing schemes, our proposed scheme can greatly improve the throughput of legitimate traffic and reduce the attack traffic during DDoS attacks. To investigate the scale-up behavior of our scheme, we have also developed a software simulator for larger-scale experiments. Simulation results show that our scheme performs consistently well even in networks with more than 3000 nodes and under high traffic load.
Ho-Yu Lam, Chi-Pan Li, Samuel T. Chanson, Dit-Yan Yeung
ICC4
2006 Solution Path for Semi-Supervised Classification with Manifold Regularization
abstract
With very low extra computational cost, the entire solution path can be computed for various learning algorithms like support vector classification (SVC) and support vector regression (SVR). In this paper, we extend this promising approach to semi-supervised learning algorithms. In particular, we consider finding the solution path for the Laplacian support vector machine (LapSVM) which is a semi-supervised classification model based on manifold regularization. One advantage of the this algorithm is that the coefficient path is piecewise linear with respect to the regularization parameter, hence its computational complexity is quadratic in the number of labeled examples.
Gang Wang 0004, Dit-Yan Yeung, Frederick H. Lochovsky
ICDM3
2006 Facial Expression Recognition using Advanced Local Binary Patterns, Tsallis Entropies and Global Appearance Features
abstract
This paper proposes a novel facial expression recognition approach based on two sets of features extracted from the face images: texture features and global appearance features. The first set is obtained by using the extended local binary patterns in both intensity and gradient maps and computing the Tsallis entropy of the Gabor filtered responses. The second set of features is obtained by performing null-space based linear discriminant analysis on the training face images. The proposed method is evaluated by extensive experiments on the JAFFE database, and compared with two widely used facial expression recognition approaches. Experimental results show that the proposed approach maintains high recognition rate in a wide range of resolution levels and outperforms the other alternative methods.
Shu Liao, Albert C. S. Chung, Dit-Yan Yeung
ICIP4
2006 Local Discriminant Embedding with Tensor Representation
abstract
We present a subspace learning method, called local discriminant embedding with tensor representation (LDET), that addresses simultaneously the generalization and data representation problems in subspace learning. LDET learns multiple interrelated subspaces for obtaining a lower-dimensional embedding by incorporating both class label information and neighborhood information. By encoding each object as a second- or higher-order tensor, LDET can capture higher-order structures in the data without requiring a large sample size. Extensive empirical studies have been performed to compare LDET with a second- or third-order tensor representation and the original LDE on their face recognition performance. Not only does LDET have a lower computational complexity than LDE, but LDET is also superior to LDE in terms of its recognition accuracy.
Jian Xia, Dit-Yan Yeung, Guang Dai
ICIP2
2006 Two-dimensional solution path for support vector regression
abstract
Recently, a very appealing approach was proposed to compute the entire solution path for support vector classification (SVC) with very low extra computational cost. This approach was later extended to a support vector regression (SVR) model called ε-SVR. However, the method requires that the error parameter ε be set a priori, which is only possible if the desired accuracy of the approximation can be specified in advance. In this paper, we show that the solution path for ε-SVR is also piecewise linear with respect to ε. We further propose an efficient algorithm for exploring the two-dimensional solution space defined by the regularization and error parameters. As opposed to the algorithm for SVC, our proposed algorithm for ε-SVR initializes the number of support vectors to zero and then increases it gradually as the algorithm proceeds. As such, a good regression function possessing the sparseness property can be obtained after only a few iterations.
Gang Wang 0004, Dit-Yan Yeung, Frederick H. Lochovsky
ICML2
2006 Intrusion Detection Routers: Design, Implementation and Evaluation Using an Experimental Testbed
abstract
In this paper, we present the design, the implementation details, and the evaluation results of an intrusion detection and defense system for distributed denial-of-service (DDoS) attack. The evaluation is conducted using an experimental testbed. The system, known as intrusion detection router (IDR), is deployed on network routers to perform online detection on any DDoS attack event, and then react with defense mechanisms to mitigate the attack. The testbed is built up by a cluster of sufficient number of Linux machines to mimic a portion of the Internet. Using the testbed, we conduct real experiments to evaluate the IDR system and demonstrate that IDR is effective in protecting the network from various DDoS attacks.
Eric Ying Kwong Chan, H. W. Chan, K. M. Chan, Vivien P. S. Chan, Samuel T. Chanson, Matthew M. H. Cheung, C. F. Chong, Kam-Pui Chow, Albert K. T. Hui, Lucas C. K. Hui, S. K. Ip, Luke C. K. Lam, W. C. Lau, Kevin K. H. Pun, Anthony Y. F. Tsang, Wai Wan Tsang, Sam C. W. Tso, Dit-Yan Yeung, Siu-Ming Yiu, Kwun Yin Yu, W. Ju
IEEE J. Sel. Areas Commun.18
2006 Model-based transductive learning of the kernel matrix
Zhihua Zhang 0004, James T. Kwok, Dit-Yan Yeung
Mach. Learn.3
2006 Robust locally linear embedding
Hong Chang 0001, Dit-Yan Yeung
Pattern Recognit.2
2006 Locally linear metric adaptation with application to semi-supervised clustering and image retrieval
Hong Chang 0001, Dit-Yan Yeung
Pattern Recognit.2
2006 Relaxational metric adaptation and its application to semi-supervised clustering and content-based image retrieval
Hong Chang 0001, Dit-Yan Yeung, William Kwok-Wai Cheung
Pattern Recognit.2
2006 Extending the relevant component analysis algorithm for metric learning using both positive and negative equivalence constraints
Dit-Yan Yeung, Hong Chang 0001
Pattern Recognit.1
2005 Stepwise Metric Adaptation Based on Semi-Supervised Learning for Boosting Image Retrieval Performance
abstract
For a specific set of features chosen for representing images, the performance of a content-based image retrieval (CBIR) system depends critically on the similarity measure used. Based on a recently proposed semisupervised metric learning method called locally linear metric adaptation (LLMA), we propose in this paper a stepwise LLMA algorithm for boosting the retrieval performance of CBIR systems by incorporating relevance feedback from users collected over multiple query sessions. Unlike most existing metric learning methods which learn a global Mahalanobis metric, the transformation performed by LLMA is more general in that it is linear locally but nonlinear globally. Moreover, the efficiency problem is well addressed by the stepwise LLMA algorithm. We also report experimental results performed on a real-world color image database to demonstrate the effectiveness of our method. 1
Hong Chang 0001, Dit-Yan Yeung
BMVC2
2005 Robust Path-Based Spectral Clustering with Application to Image Segmentation
abstract
Spectral clustering and path-based clustering are two recently developed clustering approaches that have delivered impressive results in a number of challenging clustering tasks. However, they are not robust enough against noise and outliers in the data. In this paper, based on M-estimation from robust statistics, we develop a robust path-based spectral clustering method by defining a robust path-based similarity measure for spectral clustering. Our method is significantly more robust than spectral clustering and path-based clustering. We have performed experiments based on both synthetic and real-world data, comparing our method with some other methods. In particular, color images from the Berkeley segmentation dataset and benchmark are used in the image segmentation experiments. Experimental results show that our method consistently outperforms other methods due to its higher robustness.
Hong Chang 0001, Dit-Yan Yeung
ICCV2
2005 Nonlinear dimensionality reduction for classification using kernel weighted subspace method
abstract
We study the use of kernel subspace methods that learn low-dimensional subspace representations for classification tasks. In particular, we propose a new method called kernel weighted nonlinear discriminant analysis (KWNDA) which possesses several appealing properties. First, like all kernel methods, it handles nonlinearity in a disciplined manner that is also computationally attractive. Second, by introducing weighting functions into the discriminant criterion, it outperforms existing kernel discriminant analysis methods in terms of the classification accuracy. Moreover, it also effectively deals with the small sample size problem. We empirically compare different subspace methods with respect to their classification performance of facial images based on the simple nearest neighbor rule. Experimental results show that KWNDA substantially outperforms competing linear as well as nonlinear subspace methods.
Guang Dai, Dit-Yan Yeung
ICIP (2)2
2004 Bayesian Inference on Principal Component Analysis Using Reversible Jump Markov Chain Monte Carlo
Zhihua Zhang 0004, Kap Luk Chan, James T. Kwok, Dit-Yan Yeung
AAAI4
2004 Super-Resolution through Neighbor Embedding
Hong Chang 0001, Dit-Yan Yeung, Yimin Xiong
CVPR (1)2
2004 Locally linear metric adaptation for semi-supervised clustering
abstract
Many supervised and unsupervised learning algorithms are very sensitive to the choice of an appropriate distance metric. While classification tasks can make use of class label information for metric learning, such information is generally unavailable in conventional clustering tasks. Some recent research sought to address a variant of the conventional clustering problem called semi-supervised clustering, which performs clustering in the presence of some background knowledge or supervisory information expressed as pairwise similarity or dissimilarity constraints. However, existing metric learning methods for semi-supervised clustering mostly perform global metric learning through a linear transformation. In this paper, we propose a new metric learning method which performs nonlinear transformation globally but linear transformation locally. In particular, we formulate the learning problem as an optimization problem and present two methods for solving it. Through some toy data sets, we show empirically that our locally linear metric adaptation (LLMA) method can handle some difficult cases that cannot be handled satisfactorily by previous methods. We also demonstrate the effectiveness of our method on some real data sets.
Hong Chang 0001, Dit-Yan Yeung
ICML2
2004 Surrogate maximization/minimization algorithms for AdaBoost and the logistic regression model
abstract
Surrogate maximization (or minimization) (SM) algorithms are a family of algorithm that can be regarded as a generalization of expectation-maximization (EM) algorithms. There are three major approaches to the construction of surrogate function, all relying on the convexity of some function. In this paper, we solve the boosting problem by proposing SM algorithms for the corresponding optimization problem. Specifically, for AdaBoost, we derive an SM algorithm that can be shown to be identical to the algorithm proposed by Collins et al. (2002) based on Bregman distance. More importantly, for LogitBoost (or logistic boosting), we use several methods to construct different surrogate functions which result in different SM algorithms. By combining multiple methods, we are able to derive an SM algorithm that is also the same as an algorithm derived by Collins et al. (2002). Our approach based on SM algorithms is much simpler and convergence results follow naturally.
Zhihua Zhang 0004, James T. Kwok, Dit-Yan Yeung
ICML3
2004 Bayesian inference for transductive learning of kernel matrix using the Tanner-Wong data augmentation algorithm
abstract
In kernel methods, an interesting recent development seeks to learn a good kernel from empirical data automatically. In this paper, by regarding the transductive learning of the kernel matrix as a missing data problem, we propose a Bayesian hierarchical model for the problem and devise the Tanner-Wong data augmentation algorithm for making inference on the model. The Tanner-Wong algorithm is closely related to Gibbs sampling, and it also bears a strong resemblance to the expectation-maximization (EM) algorithm. For an efficient implementation, we propose a simplified Bayesian hierarchical model and the corresponding Tanner-Wong algorithm. We express the relationship between the kernel on the input space and the kernel on the output space as a symmetric-definite generalized eigenproblem. Based on this eigenproblem, an efficient approach to choosing the base kernel matrices is presented. The effectiveness of our Bayesian model with the Tanner-Wong algorithm is demonstrated through some classification experiments showing promising results.
Zhihua Zhang 0004, Dit-Yan Yeung, James T. Kwok
ICML2
2004 Time series clustering with ARMA mixtures
Yimin Xiong, Dit-Yan Yeung
Pattern Recognit.2
2003 Parametric Distance Metric Learning with Label Information
Zhihua Zhang 0004, James T. Kwok, Dit-Yan Yeung
IJCAI3
2003 Host-based intrusion detection using dynamic and static behavioral models
Dit-Yan Yeung
Pattern Recognit.1
2002 Mixtures of ARMA Models for Model-Based Time Series Clustering
abstract
Clustering problems are central to many knowledge discovery and data mining tasks. However, most existing clustering methods can only work with fixed-dimensional representations of data patterns. In this paper we study the clustering of data patterns that are represented as sequences or time series possibly of different lengths. We propose a model-based approach to this problem using mixtures of autoregressive moving average (ARMA) models. We derive an expectation-maximization (EM) algorithm for learning the mixing coefficients as well as the parameters of component models. Experiments were conducted on simulated and real datasets. Results show that our method compares favorably with another method recently proposed by others for similar time series clustering problems.
Yimin Xiong, Dit-Yan Yeung
ICDM2
2002 User Profiling for Intrusion Detection Using Dynamic and Static Behavioral Models
Dit-Yan Yeung
PAKDD1
2002 Bidirectional Deformable Matching with Application to Handwritten Character Extraction
abstract
To achieve integrated segmentation and recognition in complex scenes, the model-based approach has widely been accepted as a promising paradigm. However, the performance is still far from satisfactory when the target object is highly deformed and the level of outlier contamination is high. In this paper, we first describe two Bayesian frameworks, one for classifying input patterns and another for detecting target patterns in complex scenes using deformable models. Then, we show that the two frameworks are similar to the forward-reverse setting of Hausdorff matching and that their matching and discriminating properties are complementary to each other. By properly combining the two frameworks, we propose a new matching scheme called bidirectional matching. This combined approach inherits the advantages of the two Bayesian frameworks. In particular, we have obtained encouraging empirical results on shape-based pattern extraction, using a subset of the CEDAR handwriting database containing handwritten words of highly varying shape.
William Kwok-Wai Cheung, Dit-Yan Yeung, Roland T. Chin
IEEE Trans. Pattern Anal. Mach. Intell.2
2002 On deformable models for visual pattern recognition
William Kwok-Wai Cheung, Dit-Yan Yeung, Roland T. Chin
Pattern Recognit.2
2001 PenCalc: A Novel Application of On-Line Mathematical Expression Recognition Technology
abstract
Most of the calculator programs found in existing pen-based mobile computing devices, such as personal digital assistants (PDA) and other handheld devices, do not take full advantages of the pen technology offered by these devices. Instead, input of expressions is still done through a virtual keypad shown on the screen, and the stylus (i.e., electronic pen) is simply used as a pointing device. In this paper we propose an intelligent handwriting-based calculator program with which the user can enter expressions simply by writing them on the screen using a stylus. In addition, variables can be defined to store intermediate results for subsequent calculations, as in ordinary algebraic calculations. The proposed software is the result of a novel application of on-line mathematical expression recognition technology which has mostly been used by others only for some mathematical expression editor programs.
Kam-Fai Chan, Dit-Yan Yeung
ICDAR2
2001 Error detection, error correction and performance evaluation in on-line mathematical expression recognition
Kam-Fai Chan, Dit-Yan Yeung
Pattern Recognit.2
2000 Mathematical expression recognition: a survey
Kam-Fai Chan, Dit-Yan Yeung
Int. J. Document Anal. Recognit.2
2000 An efficient syntactic approach to structural analysis of on-line handwritten mathematical expressions
Kam-Fai Chan, Dit-Yan Yeung
Pattern Recognit.2
1999 A Bidirectional Matching Algorithm for Deformable Pattern Detection with Application to Handwritten Word Retrieval
abstract
A Bayesian framework for deformable pattern classification was proposed by K.W. Cheung et al. (1998), with promising results for isolated handwritten character recognition. Its performance, however degrades significantly when it is applied to detect deformable patterns in complex scenes, where the amount of outliers due to other neighboring objects or the background is usually large. Also, the fact that the associated evidence measure does not penalize models resting on white space results in a high false alarm rate. Another Bayesian framework for deformable pattern detection is proposed. The framework possesses the intrinsic property of matching with only part of an image (segmentation) and its associated evidence measure can penalize white space implicitly. However, limited data exploration capability is the major trade-off. By properly combining the two frameworks, a new matching algorithm called bidirectional matching is proposed. This combined approach possesses the advantages of the two frameworks and gives robust results for non-rigid shape extraction. To evaluate the performance of the proposed approach, we have applied it to shape-based handwritten word retrieval. Using a subset of the bb dataset in the CEDAR database, we can achieve a recall rate of 59% and a precision rate of 43%.
William Kwok-Wai Cheung, Dit-Yan Yeung, Roland T. Chin
ICCV2
1999 An Environment Model for Nonstationary Reinforcement Learning
Samuel P. M. Choi, Dit-Yan Yeung, Nevin Lianwen Zhang
NIPS2
1999 Recognizing on-line handwritten alphanumeric characters through flexible structural matching
Kam-Fai Chan, Dit-Yan Yeung
Pattern Recognit.2
1998 Elastic structural matching for online handwritten alphanumeric character recognition
abstract
We propose a simple yet robust structural approach for recognizing online handwriting. Our approach is designed to achieve reasonable speed, fairly high accuracy and sufficient tolerance to variations. Experimental results show that the recognition rates are 98.60% for digits, 98.49% for uppercase letters, 97.44% for lowercase letters, and 97.40% for the combined set. When the rejected cases are excluded from the calculation, the rates can be increased to 99.93%, 99.53%, 98.55% and 98.07%, respectively. On the average, the recognition speed is about 7.5 characters per second running in Prolog on a Sun SPARC 10 Unix workstation and the memory requirement is reasonably low.
Kam-Fai Chan, Dit-Yan Yeung
ICPR2
1998 A Bayesian Framework for Deformable Pattern Recognition With Application to Handwritten Character Recognition
abstract
Deformable models have recently been proposed for many pattern recognition applications due to their ability to handle large shape variations. These proposed approaches represent patterns or shapes as deformable models, which deform themselves to match with the input image, and subsequently feed the extracted information into a classifier. The three components-modeling, matching, and classification-are often treated as independent tasks. In this paper, we study how to integrate deformable models into a Bayesian framework as a unified approach for modeling, matching, and classifying shapes. Handwritten character recognition serves as a testbed for evaluating the approach. With the use of our system, recognition is invariant to affine transformation as well as other handwriting variations. In addition, no preprocessing or manual setting of hyperparameters (e.g., regularization parameter and character width) is required. Besides, issues on the incorporation of constraints on model flexibility, detection of subparts, and speed-up are investigated. Using a model set with only 23 prototypes without any discriminative training, we can achieve an accuracy of 94.7 percent with no rejection on a subset (11,791 images by 100 writers) of handwritten digits from the NIST SD-1 dataset.
William Kwok-Wai Cheung, Dit-Yan Yeung, Roland T. Chin
IEEE Trans. Pattern Anal. Mach. Intell.2
1997 On-line handwritten alphanumeric character recognition using dominant points in strokes
Xiaolin Li 0006, Dit-Yan Yeung
Pattern Recognit.2
1997 Constructive algorithms for structure learning in feedforward neural networks for regression problems
abstract
In this survey paper, we review the constructive algorithms for structure learning in feedforward neural networks for regression problems. The basic idea is to start with a small network, then add hidden units and weights incrementally until a satisfactory solution is found. By formulating the whole problem as a state-space search, we first describe the general issues in constructive algorithms, with special emphasis on the search strategy. A taxonomy, based on the differences in the state transition mapping, the training algorithm, and the network architecture, is then presented.
James T. Kwok, Dit-Yan Yeung
IEEE Trans. Neural Networks2
1997 Objective functions for training new hidden units in constructive neural networks
abstract
In this paper, we study a number of objective functions for training new hidden units in constructive algorithms for multilayer feedforward networks. The aim is to derive a class of objective functions the computation of which and the corresponding weight updates can be done in O(N) time, where N is the number of training patterns. Moreover, even though input weight freezing is applied during the process for computational efficiency, the convergence property of the constructive algorithms using these objective functions is still preserved. We also propose a few computational tricks that can be used to improve the optimization of the objective functions under practical situations. Their relative performance in a set of two-dimensional regression problems is also discussed.
James T. Kwok, Dit-Yan Yeung
IEEE Trans. Neural Networks2
1996 Competitive Mixture of Deformable Models for Pattern Classification
abstract
Following the success of applying deformable models to feature extraction, a natural next step is to apply such models to pattern classification. Recently, we have cast a deformable model under a Bayesian framework for classification, giving promising results. However, deformable model methods are computationally expensive due to the required iterative optimization process. The problem is even more severe when there are a large number of models (e.g., for character recognition), because each of them has to deform and match with the input data before a final classification can be derived. In this paper, we propose to combine the deformable models into a mixture, in which the individual models compete with each other to survive the matching process during classification. Models that do not compete well are eliminated early, thus allowing substantial savings in computation. This process of competition-elimination has been applied to handwritten digit recognition in which significant speedup can be achieved without sacrificing recognition accuracy.
William Kwok-Wai Cheung, Dit-Yan Yeung, Roland T. Chin
CVPR2
1996 Bayesian Regularization in Constructive Neural Networks
James T. Kwok, Dit-Yan Yeung
ICANN2
1996 Use of bias term in projection pursuit learning improves approximation and convergence properties
abstract
In a regression problem, one is given a multidimensional random vector X, the components of which are called predictor variables, and a random variable, Y, called response. A regression surface describes a general relationship between X and Y. A nonparametric regression technique that has been successfully applied to high-dimensional data is projection pursuit regression (PPR). The regression surface is approximated by a sum of empirically determined univariate functions of linear combinations of the predictors. Projection pursuit learning (PPL) formulates PPR using a 2-layer feedforward neural network. The smoothers in PPR are nonparametric, whereas those in PPL are based on Hermite functions of some predefined highest order R. We demonstrate that PPL networks in the original form do not have the universal approximation property for any finite R, and thus cannot converge to the desired function even with an arbitrarily large number of hidden units. But, by including a bias term in each linear projection of the predictor variables, PPL networks can regain these capabilities, independent of the exact choice of R. Experimentally, it is shown in this paper that this modification increases the rate of convergence with respect to the number of hidden units, improves the generalization performance, and makes it less sensitive to the setting of R. Finally, we apply PPL to chaotic time series prediction, and obtain superior results compared with the cascade-correlation architecture.
James T. Kwok, Dit-Yan Yeung
IEEE Trans. Neural Networks2
1995 A grammatical inference approach to on-line handwriting modeling and recognition: a pilot study
abstract
In this paper, we present a grammar-based approach to the modeling and recognition of temporal sequences. Unlike hidden Markov models which require humans to determine in advance the appropriate model architecture to work on, our approach does not rely on prior knowledge about the topology of the underlying grammars. In particular, a discrete-time recurrent neural network model is proposed to learn separately the dynamics of each embedded subgrammar (or subpattern) class. These subgrammar network models are trained using an unsupervised learning paradigm called auto-associative (or self-supervised) learning. In this pilot study, some issues of this new approach to temporal sequence processing are investigated in the domain of on-line handwriting modeling and recognition. Some possible future research directions are also discussed.
Dit-Yan Yeung
ICDAR1
1995 Predictive Q-Routing: A Memory-based Reinforcement Learning Approach to Adaptive Traffic Control
Samuel P. M. Choi, Dit-Yan Yeung
NIPS2
1995 Improving the approximation and convergence capabilities of projection pursuit learning
James T. Kwok, Dit-Yan Yeung
Neural Process. Lett.2
1993 Constructive neural networks as estimators of bayesian discriminant functions
Dit-Yan Yeung
Pattern Recognit.1
1993 On reducing learning time in context-dependent mappings
abstract
An approach to overcoming the slow convergence problems often associated with learning complex nonlinear mappings is presented. The mappings are learned in a context-dependent manner so that complex problems are decomposed into simpler subproblems corresponding to different contexts. While no general conditions for determining applicability the method have been found, its power is illustrated through experiments in controlling simulated robot manipulators in two and three degrees of freedom. The experiments also indicate that the method shows promising scale-up properties.
Dit-Yan Yeung, George A. Bekey
IEEE Trans. Neural Networks1
1991 A Neural Network Approach to Constructive Induction
Dit-Yan Yeung
ML1
1989 Using a context-sensitive learning for robot arm control
abstract
A class of networks called context-sensitive learning networks is proposed for use in the learning of complex nonlinear mappings. Particular attention is given to a network architecture for learning to control a robot arm by learning independently the different entries of the inverse Jacobian matrix. Computer simulation results show that the network is able to learn the inverse Jacobian of the PUMA 560 arm for inverse kinematic control. The network also generalized well when unseen testing examples are presented to it.>
Dit-Yan Yeung, George A. Gekey
ICRA1
1987 A decentralized approach to the motion planning problem for multiple mobile robots
abstract
In this paper the motion planning problem for multiple mobile robots is addressed. Conventional methods of planning the motion for a single moving object are based on the assumption of a static environment, and so they cannot be used here because each of the robots is in a dynamic environment consisting of other moving robots. Centralized approaches to the multiple moving objects problem were shown to be intractable. In order to find a practical solution for the problem, it is necessary to reduce the complexity of it by decomposing the problem and introducing various heuristic techniques. We are proposing here a decentralized approach which is based on the decomposition of the problem into two subproblems: the global path planning problem and the local path replanning problem. This approach is based on a framework of problem solving using a group of intelligent agents.
Dit-Yan Yeung, George A. Bekey
ICRA1