EDBT 2026 Demo / reviewers in the wild / expert
Francis E. H. Tay
dblp:61/569 · also Eng Hock Tay, Francis Eng Hock Tay
· DBLP profile ↗
38ranked-venue papers
6as first author
7since 2021 · last 2026
0000-0001-7189-2825ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 4 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 4 since 2021Software engineering, systems software and programming languages · 10 · 1 first-authorSystems, architecture and hardware · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3Databases, data management, data science and information retrieval · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
11 papers |
Generative modeling · 36% 3D vision · 21% Robot manipulation · 10% | |
| Computer graphics and multimedia
2 papers |
Visual content generation and editing · 70% Multimedia analysis and retrieval · 30% | |
| Databases, data mining, and information retrieval
2 papers |
Information retrieval · 89% Data mining · 7% Data integration and cleaning · 4% |
Topics — the 30 heaviest of 37, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling › autoregressive model
autoregressive image generation |
1.0 | 1 | 2026 | Next Patch Prediction for AutoRegressive Visual Generation · AAAI 2026 |
Machine learning › Generative modeling
autoregressive model |
1.0 | 1 | 2026 | Next Patch Prediction for AutoRegressive Visual Generation · AAAI 2026 |
Machine learning › Generative modeling
image generation |
1.0 | 1 | 2026 | Next Patch Prediction for AutoRegressive Visual Generation · AAAI 2026 |
Machine learning › Generative modeling › autoregressive model
next-token prediction |
1.0 | 1 | 2026 | Next Patch Prediction for AutoRegressive Visual Generation · AAAI 2026 |
Machine learning › Generative modeling
diffusion model |
0.9 | 1 | 2025 | DreamDance: Animating Human Images by Enriching 3D Geometry Cues from 2D Poses · ICCV 2025 |
Visual content generation and editing › image animation
human image animation |
0.9 | 1 | 2025 | DreamDance: Animating Human Images by Enriching 3D Geometry Cues from 2D Poses · ICCV 2025 |
Computer vision › 3D vision
3d scene reconstruction |
0.8 | 1 | 2024 | You Only Scan Once: A Dynamic Scene Reconstruction Pipeline for 6-DoF Robotic Grasping of Novel Objects · ICRA 2024 |
Robotics › Robot manipulation › grasping
6-dof grasping |
0.8 | 1 | 2024 | You Only Scan Once: A Dynamic Scene Reconstruction Pipeline for 6-DoF Robotic Grasping of Novel Objects · ICRA 2024 |
Computer vision › 3D vision › 3d scene reconstruction
dynamic scene reconstruction |
0.8 | 1 | 2024 | You Only Scan Once: A Dynamic Scene Reconstruction Pipeline for 6-DoF Robotic Grasping of Novel Objects · ICRA 2024 |
Robotics › Robot manipulation
grasping |
0.8 | 1 | 2024 | You Only Scan Once: A Dynamic Scene Reconstruction Pipeline for 6-DoF Robotic Grasping of Novel Objects · ICRA 2024 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning
masked autoencoder |
0.6 | 1 | 2022 | Masked Autoencoders for Point Cloud Self-supervised Learning · ECCV (2) 2022 |
Computer vision › 3D vision
point cloud analysis |
0.6 | 1 | 2022 | Masked Autoencoders for Point Cloud Self-supervised Learning · ECCV (2) 2022 |
Computer vision › Image recognition and object detection
image classification |
0.5 | 1 | 2021 | Tokens-to-Token ViT: Training Vision Transformers from Scratch on ImageNet · ICCV 2021 |
Machine learning › Deep learning architectures and training › transformer
vision transformer |
0.5 | 1 | 2021 | Tokens-to-Token ViT: Training Vision Transformers from Scratch on ImageNet · ICCV 2021 |
Computer vision › Face, body and person analysis › human pose estimation
human pose tracking |
0.4 | 1 | 2020 | A Simple Baseline for Pose Tracking in Videos of Crowed Scenes · ACM Multimedia 2020 |
Machine learning › Efficient and distributed learning › model compression
knowledge distillation |
0.4 | 1 | 2020 | Revisiting Knowledge Distillation via Label Smoothing Regularization · CVPR 2020 |
Machine learning › Deep learning architectures and training › regularization
label smoothing |
0.4 | 1 | 2020 | Revisiting Knowledge Distillation via Label Smoothing Regularization · CVPR 2020 |
Computer vision › Video understanding and tracking › video summarization
unsupervised video summarization |
0.4 | 1 | 2020 | Unsupervised Video Summarization With Cycle-Consistent Adversarial LSTM Networks · IEEE Trans. Multim. 2020 |
Computer vision › Video understanding and tracking
video summarization |
0.4 | 1 | 2020 | Unsupervised Video Summarization With Cycle-Consistent Adversarial LSTM Networks · IEEE Trans. Multim. 2020 |
Information retrieval › hashing
hashing-based retrieval |
0.4 | 1 | 2020 | Central Similarity Quantization for Efficient Image and Video Retrieval · CVPR 2020 |
Multimedia analysis and retrieval
video summarization |
0.4 | 1 | 2019 | Cycle-SUM: Cycle-Consistent Adversarial LSTM Networks for Unsupervised Video Summarization · AAAI 2019 |
Computer vision › Image recognition and object detection › object detection
object proposal generation |
0.2 | 1 | 2016 | Scale-Aware Pixelwise Object Proposal Networks · IEEE Trans. Image Process. 2016 |
Computer vision › Image recognition and object detection › object detection
small object detection |
0.2 | 1 | 2016 | Scale-Aware Pixelwise Object Proposal Networks · IEEE Trans. Image Process. 2016 |
Computer vision › 3D vision › object pose estimation
object pose tracking |
0.2 | 1 | 2024 | You Only Scan Once: A Dynamic Scene Reconstruction Pipeline for 6-DoF Robotic Grasping of Novel Objects · ICRA 2024 |
Computer vision › Video understanding and tracking
frame selection |
0.1 | 1 | 2020 | Unsupervised Video Summarization With Cycle-Consistent Adversarial LSTM Networks · IEEE Trans. Multim. 2020 |
Machine learning › Representation and self-supervised learning › representation learning › embedding learning
hash code learning |
0.1 | 1 | 2020 | Central Similarity Quantization for Efficient Image and Video Retrieval · CVPR 2020 |
Computer vision › Video understanding and tracking
multi-object tracking |
0.1 | 1 | 2020 | A Simple Baseline for Pose Tracking in Videos of Crowed Scenes · ACM Multimedia 2020 |
Machine learning › Generative modeling
generative adversarial network |
0.1 | 1 | 2019 | Cycle-SUM: Cycle-Consistent Adversarial LSTM Networks for Unsupervised Video Summarization · AAAI 2019 |
Machine learning › Deep learning architectures and training › convolutional neural network › convolutional neural network architecture
fully convolutional network |
0.1 | 1 | 2016 | Scale-Aware Pixelwise Object Proposal Networks · IEEE Trans. Image Process. 2016 |
Data integration and cleaning › data preprocessing
discretization |
0.0 | 1 | 2002 | A Modified Chi2 Algorithm for Discretization · IEEE Trans. Knowl. Data Eng. 2002 |
Methods — techniques the papers use, named apart from their topics
geometry diffusion · 1.7diffusion model · 1.7cross-domain controller · 1.7cycle consistency · 1.2LSTM · 1.2multi-scale coarse-to-fine patch grouping · 1.0point cloud registration · 0.8mesh reconstruction · 0.8self-supervised learning · 0.6deep-narrow backbone · 0.5hadamard matrix · 0.4central similarity · 0.4bernoulli distribution · 0.4adversarial learning · 0.4chi-square merging · 0.0c4.5 · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Next Patch Prediction for AutoRegressive Visual GenerationabstractAutoregressive models, built based on the Next Token Prediction (NTP) paradigm, show great potential in developing a unified framework that integrates both language and vision tasks. Pioneering works introduce NTP to autoregressive visual generation tasks. In this work, we rethink the NTP for autoregressive image generation and extend it to a novel Next Patch Prediction (NPP) paradigm. Our key idea is to group and aggregate image tokens into patch tokens with higher information density. By using patch tokens as a more compact input sequence, the autoregressive model is trained to predict the next patch, significantly reducing computational costs. To further exploit the natural hierarchical structure of image data, we propose a multi-scale coarse-to-fine patch grouping strategy. With this strategy, the training process begins with a large patch size and ends with vanilla NTP where the patch size is 1x1, thus maintaining the original inference process without modifications. Extensive experiments across a diverse range of model sizes demonstrate that NPP could reduce the training cost to around 0.6 times while improving image generation quality by up to 1.0 FID score on the ImageNet 256x256 generation benchmark. Notably, our method retains the original autoregressive model architecture without introducing additional trainable parameters or specifically designing a custom image tokenizer, offering a flexible and plug-and-play solution for enhancing autoregressive visual generation. Yatian Pang, Peng Jin 0001, Bin Lin 0014, Chaoran Feng 0001, Zhenyu Tang 0004, Liuhan Chen, Francis E. H. Tay, Ser-Nam Lim, Harry Yang, Li Yuan 0007 |
AAAI | 9 |
| 2025 | DreamDance: Animating Human Images by Enriching 3D Geometry Cues from 2D PosesabstractIn this work, we present DreamDance, a novel method for animating human images using only skeleton pose sequences as conditional inputs. Existing approaches struggle with generating coherent, high-quality content in an efficient and user-friendly manner. Concretely, baseline methods relying on only 2D pose guidance lack the cues of 3D information, leading to suboptimal results, while methods using 3D representation as guidance achieve higher quality but involve a cumbersome and time-intensive process. To address these limitations, DreamDance enriches 3D geometry cues from 2D poses by introducing an efficient diffusion model, enabling high-quality human image animation with various guidance. Our key insight is that human images naturally exhibit multiple levels of correlation, progressing from coarse skeleton poses to fine-grained geometry cues, and further from these geometry cues to explicit appearance details. Capturing such correlations could enrich the guidance signals, facilitating intra-frame coherency and inter-frame consistency. Specifically, we construct the TikTok-Dance5K dataset, comprising 5K high-quality dance videos with detailed frame annotations, including human pose, depth, and normal maps. Next, we introduce a Mutually Aligned Geometry Diffusion Model to generate fine-grained depth and normal maps for enriched guidance. Finally, a Cross-domain Controller incorporates multi-level guidance to animate human images effectively with a video diffusion model. Extensive experiments demonstrate that our method achieves state-of-the-art performance in animating human images. Yatian Pang, Bin Bin, Mingzhe Zheng, Francis E. H. Tay, Ser-Nam Lim, Harry Yang, Li Yuan 0007 |
ICCV | 5 |
| 2024 | You Only Scan Once: A Dynamic Scene Reconstruction Pipeline for 6-DoF Robotic Grasping of Novel ObjectsabstractIn the realm of robotic grasping, achieving accurate and reliable interactions with the environment is a pivotal challenge. Traditional methods of grasp planning methods utilizing partial point clouds derived from depth image often suffer from reduced scene understanding due to occlusion, ultimately impeding their grasping accuracy. Furthermore, scene reconstruction methods have primarily relied upon static techniques, which are susceptible to environment change during manipulation process limits their efficacy in real-time grasping tasks. To address these limitations, this paper introduces a novel two-stage pipeline for dynamic scene reconstruction. In the first stage, our approach takes scene scanning as input to register each target object with mesh reconstruction and novel object pose tracking. In the second stage, pose tracking is still performed to provide object poses in real-time, enabling our approach to transform the reconstructed object point clouds back into the scene. Unlike conventional methodologies, which rely on static scene snapshots, our method continuously captures the evolving scene geometry, resulting in a comprehensive and up-to-date point cloud representation. By circumventing the constraints posed by occlusion, our method enhances the overall grasp planning process and empowers state-of-the-art 6-DoF robotic grasping algorithms to exhibit markedly improved accuracy. Haozhe Wang 0003, Zhengshen Zhang, Francis E. H. Tay, Marcelo H. Ang |
ICRA | 5 |
| 2024 | 3D Affordance Keypoint Detection for Robotic ManipulationabstractThis paper presents a novel approach for affordance-informed robotic manipulation by introducing 3D keypoints to enhance the understanding of object parts’ functionality. The proposed approach provides direct information about what the potential use of objects is, as well as guidance on where and how a manipulator should engage, whereas conventional methods treat affordance detection as a semantic segmentation task, focusing solely on answering the what question. To address this gap, we propose a Fusion-based Affordance Keypoint Network (FAKP-Net) by introducing 3D keypoint quadruplet that harnesses the synergistic potential of RGB and Depth image to provide information on execution position, direction, and extent. Benchmark testing demonstrates that FAKP-Net outperforms existing models by significant margins in affordance segmentation task and keypoint detection task. Real-world experiments also showcase the reliability of our method in accomplishing manipulation tasks with previously unseen objects. Our source code and video demo will be public. Ruiteng Zhao, Chengran Yuan, Yuwei Wu 0002, Zhengshen Zhang, Marcelo H. Ang, Francis E. H. Tay |
IROS | 10 |
| 2024 | Learnable Central Similarity Quantization for Efficient Image and Video RetrievalabstractData-dependent hashing methods aim to learn hash functions from the pairwise or triplet relationships among the data, which often lead to low efficiency and low collision rate by only capturing the local distribution of the data. To solve the limitation, we propose central similarity, in which the hash codes of similar data pairs are encouraged to approach a common center and those of dissimilar pairs to converge to different centers. As a new global similarity metric, central similarity can improve the efficiency and retrieval accuracy of hash learning. By introducing a new concept, hash centers, we principally formulate the computation of the proposed central similarity metric, in which the hash centers refer to a set of points scattered in the Hamming space with a sufficient mutual distance between each other. To construct well-separated hash centers, we provide two efficient methods: 1) leveraging the Hadamard matrix and Bernoulli distributions to generate data-independent hash centers and 2) learning data-dependent hash centers from data representations. Based on the proposed similarity metric and hash centers, we propose central similarity quantization (CSQ) that optimizes the central similarity between data points with respect to their hash centers instead of optimizing the local similarity to generate a high-quality deep hash function. We also further improve the CSQ with data-dependent hash centers, dubbed as CSQ with learnable center (CSQLC). The proposed CSQ and CSQLC are generic and applicable to image and video hashing scenarios. We conduct extensive experiments on large-scale image and video retrieval tasks, and the proposed CSQ yields noticeably boosted retrieval performance, i.e., 3%-20% in mean average precision (mAP) over the previous state-of-the-art methods, which also demonstrates that our methods can generate cohesive hash codes for similar data pairs and dispersed hash codes for dissimilar pairs. Li Yuan 0007, Tao Wang 0053, Xiaopeng Zhang 0008, Francis E. H. Tay, Zequn Jie, Yonghong Tian 0001, Wei Liu 0005, Jiashi Feng |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2022 | Masked Autoencoders for Point Cloud Self-supervised Learning
Yatian Pang, Wenxiao Wang 0001, Francis E. H. Tay, Wei Liu 0005, Yonghong Tian 0001, Li Yuan 0007 |
ECCV (2) | 3 |
| 2021 | Tokens-to-Token ViT: Training Vision Transformers from Scratch on ImageNetabstractTransformers, which are popular for language modeling, have been explored for solving vision tasks recently, e.g., the Vision Transformer (ViT) for image classification. The ViT model splits each image into a sequence of tokens with fixed length and then applies multiple Transformer layers to model their global relation for classification. However, ViT achieves inferior performance to CNNs when trained from scratch on a midsize dataset like ImageNet. We find it is because: 1) the simple tokenization of input images fails to model the important local structure such as edges and lines among neighboring pixels, leading to low training sample efficiency; 2) the redundant attention backbone design of ViT leads to limited feature richness for fixed computation budgets and limited training samples. To overcome such limitations, we propose a new Tokens-To-Token Vision Transformer (T2T-VTT), which incorporates 1) a layer-wise Tokens-to-Token (T2T) transformation to progressively structurize the image to tokens by recursively aggregating neighboring Tokens into one Token (Tokens-to-Token), such that local structure represented by surrounding tokens can be modeled and tokens length can be reduced; 2) an efficient backbone with a deep-narrow structure for vision transformer motivated by CNN architecture design after empirical study. Notably, T2T-ViT reduces the parameter count and MACs of vanilla ViT by half, while achieving more than 3.0% improvement when trained from scratch on ImageNet. It also outperforms ResNets and achieves comparable performance with MobileNets by directly training on ImageNet. For example, T2T-ViT with comparable size to ResNet50 (21.5M parameters) can achieve 83.3% top1 accuracy in image resolution 384x384 on ImageNet.1 Li Yuan 0007, Yunpeng Chen, Tao Wang 0053, Weihao Yu 0001, Yujun Shi, Zihang Jiang, Francis E. H. Tay, Jiashi Feng, Shuicheng Yan |
ICCV | 7 |
| 2020 | Revisiting Knowledge Distillation via Label Smoothing RegularizationabstractKnowledge Distillation (KD) aims to distill the knowledge of a cumbersome teacher model into a lightweight student model. Its success is generally attributed to the privileged information on similarities among categories provided by the teacher model, and in this sense, only strong teacher models are deployed to teach weaker students in practice. In this work, we challenge this common belief by following experimental observations: 1) beyond the acknowledgment that the teacher can improve the student, the student can also enhance the teacher significantly by reversing the KD procedure; 2) a poorly-trained teacher with much lower accuracy than the student can still improve the latter significantly. To explain these observations, we provide a theoretical analysis of the relationships between KD and label smoothing regularization. We prove that 1) KD is a type of learned label smoothing regularization and 2) label smoothing regularization provides a virtual teacher model for KD. From these results, we argue that the success of KD is not fully due to the similarity information between categories from teachers, but also to the regularization of soft targets, which is equally or even more important. Based on these analyses, we further propose a novel Teacher-free Knowledge Distillation (Tf-KD) framework, where a student model learns from itself or manually-designed regularization distribution. The Tf-KD achieves comparable performance with normal KD from a superior teacher, which is well applied when a stronger teacher model is unavailable. Meanwhile, Tf-KD is generic and can be directly deployed for training deep neural networks. Without any extra computation cost, Tf-KD achieves up to 0.65\% improvement on ImageNet over well-established baseline models, which is superior to label smoothing regularization. Li Yuan 0007, Francis E. H. Tay, Guilin Li 0001, Tao Wang 0053, Jiashi Feng |
CVPR | 2 |
| 2020 | Central Similarity Quantization for Efficient Image and Video RetrievalabstractExisting data-dependent hashing methods usually learn hash functions from pairwise or triplet data relationships, which only capture the data similarity locally, and often suffer from low learning efficiency and low collision rate. In this work, we propose a new \emph{global} similarity metric, termed as \emph{central similarity}, with which the hash codes of similar data pairs are encouraged to approach a common center and those for dissimilar pairs to converge to different centers, to improve hash learning efficiency and retrieval accuracy. We principally formulate the computation of the proposed central similarity metric by introducing a new concept, i.e., \emph{hash center} that refers to a set of data points scattered in the Hamming space with a sufficient mutual distance between each other. We then provide an efficient method to construct well separated hash centers by leveraging the Hadamard matrix and Bernoulli distributions. Finally, we propose the Central Similarity Quantization (CSQ) that optimizes the central similarity between data points w.r.t.\ their hash centers instead of optimizing the local similarity. CSQ is generic and applicable to both image and video hashing scenarios. Extensive experiments on large-scale image and video retrieval tasks demonstrate that CSQ can generate cohesive hash codes for similar data pairs and dispersed hash codes for dissimilar pairs, achieving a noticeable boost in retrieval performance, i.e. 3\%-20\% in mAP over the previous state-of-the-arts. Li Yuan 0007, Tao Wang 0053, Xiaopeng Zhang 0008, Francis E. H. Tay, Zequn Jie, Wei Liu 0005, Jiashi Feng |
CVPR | 4 |
| 2020 | A Simple Baseline for Pose Tracking in Videos of Crowed ScenesabstractThis paper presents our solution to ACM MM challenge: Large-scale Human-centric Video Analysis in Complex Events[13]; specifically, here we focus on Track3: Crowd Pose Tracking in Complex Events. Remarkable progress has been made in multi-pose training in recent years. However, how to track the human pose in crowded and complex environments has not been well addressed. We formulate the problem as several subproblems to be solved. First, we use a multi-object tracking method to assign human ID to each bounding box generated by the detection model. After that, a pose is generated to each bounding box with ID. At last, optical flow is used to take advantage of the temporal information in the videos and generate the final pose tracking result. Li Yuan 0007, Shuning Chang, Ziyuan Huang 0003, Xuecheng Nie, Francis E. H. Tay, Jiashi Feng, Shuicheng Yan |
ACM Multimedia | 7 |
| 2020 | Unsupervised Video Summarization With Cycle-Consistent Adversarial LSTM NetworksabstractVideo summarization is an important technique to browse, manage and retrieve a large amount of videos efficiently. The main objective of video summarization is to minimize the information loss when selecting a subset of video frames from the original video, hence the summary video can faithfully represent the overall story of the original video. Recently developed unsupervised video summarization approaches are free of requiring tedious annotation on important frames to train a video summarization model and thus are practically attractive. However, their performance is still limited due to the difficulty of minimizing information loss between the summary and original videos. In this paper, we address unsupervised video summarization by developing a novel Cycle-consistent Adversarial LSTM architecture to effectively reduce the information loss in the summary video. The proposed model, named Cycle-SUM, consists of a frame selector and a cycle-consistent learning based evaluator. The selector is a bi-directional LSTM network to capture the long-range relationship between video frames. To overcome the difficulty of specifying a suitable information preserving metric between original video and summary video, the evaluator is introduced to “supervise” selector to improve the video summarization quality. Specifically, the evaluator is composed of two generative adversarial networks (GANs), in which the forward GAN component is learned to reconstruct the original video from summary video, while the backward GAN learns to invert the process. We establish the relation between mutual information maximization and such cycle learning procedure and further introduce cycle-consistent loss to regularize the summarization. Extensive experiments on three video summarization benchmark datasets demonstrate a state-of-the-art performance, and show the superiority of the Cycle-SUM model compared with other unsupervised approaches. Li Yuan 0007, Francis E. H. Tay, Ping Li 0006, Jiashi Feng |
IEEE Trans. Multim. | 2 |
| 2019 | Cycle-SUM: Cycle-Consistent Adversarial LSTM Networks for Unsupervised Video SummarizationabstractIn this paper, we present a novel unsupervised video summarization model that requires no manual annotation. The proposed model termed Cycle-SUM adopts a new cycleconsistent adversarial LSTM architecture that can effectively maximize the information preserving and compactness of the summary video. It consists of a frame selector and a cycle-consistent learning based evaluator. The selector is a bi-direction LSTM network that learns video representations that embed the long-range relationships among video frames. The evaluator defines a learnable information preserving metric between original video and summary video and “supervises” the selector to identify the most informative frames to form the summary video. In particular, the evaluator is composed of two generative adversarial networks (GANs), in which the forward GAN is learned to reconstruct original video from summary video while the backward GAN learns to invert the processing. The consistency between the output of such cycle learning is adopted as the information preserving metric for video summarization. We demonstrate the close relation between mutual information maximization and such cycle learning procedure. Experiments on two video summarization benchmark datasets validate the state-of-theart performance and superiority of the Cycle-SUM model over previous baselines. Li Yuan 0007, Francis E. H. Tay, Ping Li 0006, Jiashi Feng |
AAAI | 2 |
| 2018 | Object Proposal Generation With Fully Convolutional NetworksabstractObject proposal generation, as a preprocessing technique, has been widely used in current object detection pipelines to guide the search of objects and avoid exhaustive sliding window search across images. Current object proposals are mostly based on low-level image cues, such as edges and saliency. However, objectness is possibly a high-level semantic concept showing whether one region contains objects. This paper presents a framework utilizing fully convolutional networks (FCNs) to produce object proposal positions and bounding box location refinement with Support Vector Machine (SVM) to further improve proposal localization. Experiments on the PASCAL VOC 2007 show that using high-level semantic object proposals obtained by FCN, the object recall can be improved. An improvement in detection mean average precision is also seen when using our proposals in the Fast R-convolutional neural network framework. In addition, we also demonstrate that our method shows stronger robustness when introduced to image perturbations, e.g., blurring, JPEG compression, and salt and pepper noise. Finally, the generalization capability of our model (trained on the PASCAL VOC 2007) is evaluated and validated by testing on PASCAL VOC 2012 validation set, ILSVRC 2013 validation set, and MS COCO 2014 validation set. Zequn Jie, Wen Feng Lu, Siavash Sakhavi, Yunchao Wei, Francis E. H. Tay, Shuicheng Yan |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2016 | Scale-Aware Pixelwise Object Proposal NetworksabstractObject proposal is essential for current state-of-the-art object detection pipelines. However, the existing proposal methods generally fail in producing results with satisfying localization accuracy. The case is even worse for small objects, which, however, are quite common in practice. In this paper, we propose a novel scale-aware pixelwise object proposal network (SPOP-net) to tackle the challenges. The SPOP-net can generate proposals with high recall rate and average best overlap, even for small objects. In particular, in order to improve the localization accuracy, a fully convolutional network is employed which predicts locations of object proposals for each pixel. The produced ensemble of pixelwise object proposals enhances the chance of hitting the object significantly without incurring heavy extra computational cost. To solve the challenge of localizing objects at small scale, two localization networks, which are specialized for localizing objects with different scales are introduced, following the divide-and-conquer philosophy. Location outputs of these two networks are then adaptively combined to generate the final proposals by a large-/small-size weighting network. Extensive evaluations on PASCAL VOC 2007 and COCO 2014 show the SPOP network is superior over the state-of-the-art models. The high-quality proposals from SPOP-net also significantly improve the mean average precision of object detection with Fast-Regions with CNN features framework. Finally, the SPOP-net (trained on PASCAL VOC) shows great generalization performance when testing it on ILSVRC 2013 validation set. Zequn Jie, Xiaodan Liang, Jiashi Feng, Wen Feng Lu, Francis E. H. Tay, Shuicheng Yan |
IEEE Trans. Image Process. | 5 |
| 2008 | A Long-term Wearable Vital Signs Monitoring System using BSNabstractWearable medical devices provide great convenience to the elderly for the monitoring of vital signs at home. This paper describes a novel long-term wearable vital signs monitoring system which can real-time measure physiological signs such as ECG, SpO2 (Saturation of Arterial Oxygen) and systolic blood pressure. The proposed device is easy to wear, convenient to use and has almost no impact on the quality of life. A miniaturized PCB integrated with the ECG analog front-end and SpO2 LED drive circuit is attached onto the BSN node. By simply wearing the chest band embedded with ECG micromachined electrodes and the ear-probe which incorporates LEDs and photodetector, the measured physiological signals can be wirelessly transmitted to a hand-held device, such as PDA phone, where the heart beat rate, SpO2 and systolic blood pressure are calculated and displayed. Ding G. Guo, Francis E. H. Tay, L. M. Yu, Myo Naing Nyan, F. W. Chong, K. L. Yap |
DSD | 2 |
| 2005 | Dynamic Behaviors of High-G Mems Accelerometer Incorporated with Novel Micro-FlexuresabstractIn this work, a new MEMS accelerometer with large detectable range of more than 50 g is designed. Three types of flexure designs were studied: the conventional straight flexure, newly proposed interlapped-L flexure and rectangular flexure. Their dimensions were optimized to achieve the desired requirements of the accelerometer. Capacitive sensing method and electrostatic actuation are selected to be the sensing and actuation methods. Governing equations derived from this model are used to compare with that of a second order spring-mass-damper system. These mathematical models are then used to formulate the various types of sensitivities. Finite Element Analysis software ANSYS is used in the design stage to simulate its dynamic behavior. The accelerometers with interlapped-L flexure and rectangular flexure present very large detectable range of 60 g and 80 g, sensitivity of 23.1 and 17.3 fF/g, with the noise floor of 17.9 and 18.2 μ g/(Hz) 1/2 in atmosphere. Bangtao Chen, Jianmin Miao, Chunkiat Lim, Francis E. H. Tay, Ciprian Iliescu |
Int. J. Softw. Eng. Knowl. Eng. | 4 |
| 2005 | Design Analysis of The Propulsion and Control System of an Underactuated Remotely Operated Vehicle Using Axiomatic Design Theory - Part 1abstractIn this paper, the system design issues of the Propulsion and Control System of the ROV II are analyzed and addressed. The design concept, some of the upgraded features of the ROV II in comparison to ROV I and the unified pilot training and control system developed, will also be briefly discussed in this paper. Teck Hong Koh, Francis E. H. Tay, Michael Wai Shing Lau, Eicher Low, Gerald Seet |
Int. J. Softw. Eng. Knowl. Eng. | 2 |
| 2005 | Design Analysis of the Propulsion and Control System of an Underactuated Remotely Operated Vehicle Using Axiomatic Design Theory - Part 2abstractIn this paper, Axiomatic Design (AD) theory was adopted for the design analysis of an underactuated Remotely Operated Vehicle (ROV) system and its subcomponents. The system design issues of the Propulsion and Control System of the ROV II are analyzed and addressed based on the Independence Axiom methodology. The top-level Functional Requirements (FRs) for the thruster design and configuration are identified and its corresponding Design Parameters (DPs) are also presented. Teck Hong Koh, Francis E. H. Tay, Michael Wai Shing Lau, Eicher Low, Gerald Seet |
Int. J. Softw. Eng. Knowl. Eng. | 2 |
| 2005 | Pareto Simulated Annealing (Sa)-Based Multi-Objective Optimization for Mems Design and ApplicationabstractIn this paper we present a global optimization method for multiple objective functions using the Pareto Simulated Annealing (SA). This novel optimization method is very useful and promising for design and application in the field of Micro-Electro-Mechanical Systems (MEMS). Previously published global optimization method has been reported by us for only single objective function. The proposed method automatically assigns different objective weights to each objective functions so that it can generate multiple solutions simultaneously. It also offers the trade-off between the objective functions so that we will be able to select the most suitable solution for MEMS design and applications. Based on the global Pareto ranking of the solutions, the optimization method can provide the best solution (the first Pareto ranking) as well. Andojo Ongkodjojo Ong, Francis E. H. Tay |
Int. J. Softw. Eng. Knowl. Eng. | 2 |
| 2005 | Modeling and Analysis of Nanotips for Thermoelectric CoolersabstractThermoelectric cooling system can be used in many applications. The properties of structured point contact at the cold end affect the performance of thermoelectric cooling. The performance of metal point junctions as a function of tip radius show that the smaller the tip radius, the better the figure of merit. This paper presents the modeling and analysis of the nanotips. Each nanotip will have to carry some fraction of the weight of the thermoelectric material and the silicon slabs. The small apical area of each nanotip causes a very high pressure build up. The device may be loaded externally to ensure proper nanotip and thermoelectric material contact. A deformation of approximately 10 to 15 nm would be ideal. Finite Element Analysis (FEA) needed to be applied to determine how tall each nanotip must be to maximize soft and hard tip contacts. FEA was also needed to approximate the amount of external loading that will be necessary. ANSYS software was used in the modeling and analysis of the nanotips. Kwong-Luck Tan, Prashant Padmanabhan, Ciprian Iliescu, Francis E. H. Tay |
Int. J. Softw. Eng. Knowl. Eng. | 4 |
| 2005 | Sleep Monitoring Devices Using Electric Field (E-Field) Mattress for Children with EczemaabstractThis paper presents the concept of implementing different dielectric properties of surroundings to monitor the sleep of eczema patients through the use of an electric field mattress. Different sensor methods, as well as different arrangements of electrodes in the E-field mattress, are discussed to find out which would achieve optimum results. This paper also explains how readings taken from the electrodes can be used to analyze the sleep pattern of eczema patients. Kwong-Luck Tan, Francis E. H. Tay, Hugo Van Bever |
Int. J. Softw. Eng. Knowl. Eng. | 2 |
| 2005 | Smart Shirt That Can Call for Help After A FallabstractA smart shirt was developed for geriatric healthcare purposes. It can summon medical assistances via SMS and e-mail when its bearer has fallen. Single axis accelerometers are arranged in medio-lateral, vertical and antero-posterior direction on the shoulder part of the shirt. Using Bluetooth™ transmitter/receiver, signals are transmitted and processed on a Personal Computer (PC) for fall-detection. Upon detection of falls, SMS (Short Messaging Service) and e-mail are sent through GSM (Global System for Mobile Communication) network supported by a local telecommunication service provider. The primary advantage of our detection system compared to other researchers' systems is that fall notification can be sent to individuals and medical health care unit simultaneously to lead to a shortened interval of the arrival of assistance. Francis E. H. Tay, Myo Naing Nyan, Teck Hong Koh, K. H. W. Seah, Yih Yiow Sitoh |
Int. J. Softw. Eng. Knowl. Eng. | 1 |
| 2005 | Multi-Channel Biotelemetry System Using Microcontroller with Uhf Transmit FunctionabstractThis paper presents the development of a three-channel telemetry system using Microchip rfPIC12F675 as transmitter and using rfRXD0420 as receiver. The highly integrated radio frequency IC chip features the system which is simple in structure, compact in size and low in power consumption. Using 3V power supply for the transmitter, the system is able to transmit low amplitude signal from 8 m to more than 50 m with power consumption as low as 10 mW. The biotelemetry system can be used in human or animal bio-signal acquisition and monitoring. Guolin Xu, Francis E. H. Tay, Ciprian Iliescu, Victor Samper |
Int. J. Softw. Eng. Knowl. Eng. | 2 |
| 2005 | Theoretical Analysis and Experiment of A Novel Dep Chip With 3-D Silicon ElectrodesabstractThis paper presents a novel dielectrophoresis (DEP) device where the DEP electrodes define the channel walls. This is achieved by fabricating microfluidic channel walls from highly doped silicon so that they can also function as DEP electrodes. Compared with planar electrodes, this device increases the exhibited dielectrophoretic force on the particle, therefore decreases the applied potential and reduces the heating of the solution. A DEP device with triangle electrodes has been designed and fabricated. Compared with the other two configurations, semi-circular and square, triangle electrode presents an increased force, which can decrease the applied voltage and reduce the Joule effect. Yeast cells have been used to for testing the performance of the device. Liming Yu, Francis E. H. Tay, Guolin Xu, Ciprian Iliescu, Marioara Avram |
Int. J. Softw. Eng. Knowl. Eng. | 2 |
| 2005 | Control-Oriented Modeling of 2d Torsional MicromirrorabstractThis paper presents a simplified model of 2D torsional micromirror for control design. The micromirror system is modeled as a two inputs two outputs (TITO) system with nonlinear electrostatic torques, which are coupled with the two inputs and the two tilt angles. The coupling property of the electrostatic torques makes it difficult to design the control law. To surmount the problem, a new model based on extracting the dominant nonlinear terms and Taylor's expansion is proposed. The micromirror system together with the simplified model is coded in Matlab™ S-function. The simulation result shows that the model is feasible for control design. Francis E. H. Tay, Siong Chau Fook, Guangya Zhou |
Int. J. Softw. Eng. Knowl. Eng. | 2 |
| 2003 | Support vector machine with adaptive parameters in financial time series forecastingabstractA novel type of learning machine called support vector machine (SVM) has been receiving increasing interest in areas ranging from its original application in pattern recognition to other applications such as regression estimation due to its remarkable generalization performance. This paper deals with the application of SVM in financial time series forecasting. The feasibility of applying SVM in financial forecasting is first examined by comparing it with the multilayer back-propagation (BP) neural network and the regularized radial basis function (RBF) neural network. The variability in performance of SVM with respect to the free parameters is investigated experimentally. Adaptive parameters are then proposed by incorporating the nonstationarity of financial time series into SVM. Five real futures contracts collated from the Chicago Mercantile Market are used as the data sets. The simulation shows that among the three methods, SVM outperforms the BP neural network in financial forecasting, and there are comparable generalization performance between SVM and the regularized RBF neural network. Furthermore, the free parameters of SVM have a great effect on the generalization performance. SVM with adaptive parameters can both achieve higher generalization performance and use fewer support vectors than the standard SVM in financial forecasting. Lijuan Cao, Francis E. H. Tay |
IEEE Trans. Neural Networks | 2 |
| 2002 | Modified support vector machines in financial time series forecasting
Francis E. H. Tay, Lijuan Cao |
Neurocomputing | 1 |
| 2002 | ε Descending Support Vector Machines for Financial Time Series Forecasting
Francis E. H. Tay, Lijuan Cao |
Neural Process. Lett. | 1 |
| 2002 | A Modified Chi2 Algorithm for DiscretizationabstractSince the ChiMerge algorithm was first proposed by Kerber (1992), it has become a widely used and discussed discretization method. The Chi2 algorithm is a modification to the ChiMerge method. It automates the discretization process by introducing an inconsistency rate as the stopping criterion and it automatically selects the significance value. In addition, it adds a finer phase aimed at feature selection to broaden the applications of the ChiMerge algorithm. However, the Chi2 algorithm does not consider the inaccuracy inherent in ChiMerge's merging criterion. The user-defined inconsistency rate also brings about inaccuracy to the discretization process. These two drawbacks are first discussed in this paper and modifications to overcome them are then proposed. By comparison, results with the original Chi2 algorithm using C4.5, the modified Chi2 algorithm, performs better than the original Chi2 algorithm. It becomes a completely automatic discretization method. Francis E. H. Tay, Lixiang Shen |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2001 | A comparative study of saliency analysis and genetic algorithm for feature selection in support vector machines
Francis E. H. Tay, Lijuan Cao |
Intell. Data Anal. | 1 |
| 2001 | Improved financial time series forecasting by combining Support Vector Machines with self-organizing feature map
Francis E. H. Tay, Lijuan Cao |
Intell. Data Anal. | 1 |
| 2001 | A discretization method for rough sets theory
Lixiang Shen, Francis E. H. Tay |
Intell. Data Anal. | 2 |
| 2001 | Financial Forecasting Using Support Vector Machines
Lijuan Cao, Francis E. H. Tay |
Neural Comput. Appl. | 2 |
| 2000 | Feature Selection for Support Vector Machines
Lijuan Cao, Francis E. H. Tay |
IDEAL | 2 |
| 2000 | epsilon-Descending Support Vector Machines for Financial Time Series Forecasting
Lijuan Cao, Francis E. H. Tay |
IDEAL | 2 |
| 2000 | Classifying Market States with WARS
Lixiang Shen, Francis E. H. Tay |
IDEAL | 2 |
| 1999 | Neuro-Genetic Based Method to the Classification of Acupuncture Needle: A Case Study
Lijuan Cao, Francis E. H. Tay |
GECCO | 2 |
| 1999 | Classification of the Market States Using Neural Network
Lijuan Cao, Francis E. H. Tay, Lawrence Ma, Wai Cheong Yeong |
GECCO | 2 |