VLDB 2026 Research / reviewers in the wild / expert
Yu-Chuan Su
dblp:53/6299
· DBLP profile ↗
28ranked-venue papers
15as first author
10since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 20 · 11 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 10 first-author · 6 since 2021Systems, architecture and hardware · 3 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Scaling Inference Time Compute for Diffusion ModelsabstractGenerative models have made significant impacts across various domains, largely due to their ability to scale during training by increasing data, computational resources, and model size, a phenomenon characterized by the scaling laws. Recent research has begun to explore inference-time scaling behavior in Large Language Models (LLMs), revealing how performance can further improve with additional computation during inference. Unlike LLMs, diffusion models inherently possess the flexibility to adjust inference-time computation via the number of denoising steps, although the performance gains typically flatten after a few dozen. In this work, we explore the inference-time scaling behavior of diffusion models beyond increasing denoising steps and investigate how the generation performance can further improve with increased computation. Specifically, we consider a search problem aimed at identifying better noises for the diffusion sampling process. We structure the design space along two axes: the verifiers used to provide feedback, and the algorithms used to find better noise candidates. Through extensive experiments on class-conditioned and text-conditioned image generation benchmarks, our findings reveal that increasing inference-time compute leads to substantial improvements in the quality of samples generated by diffusion models, and with the complicated nature of images, combinations of the components in the framework can be specifically chosen to conform with different application scenario. Nanye Ma, Shangyuan Tong, Hexiang Hu, Yu-Chuan Su, Yandong Li, Tommi S. Jaakkola, Xuhui Jia, Saining Xie |
CVPR | 5 |
| 2025 | Generating Long-Take Videos via Effective Keyframes and GuidanceabstractWe tackle the challenge of generating long-take videos encompassing multiple non-repetitive yet coherent events. Existing approaches generate long videos conditioned on single input guidance, often leading to repetitive content. To address this problem, we develop a framework that uses multiple guidance sources to enhance long video generation. The main idea of our approach is to decouple video generation into keyframe generation and frame interpolation. In this process, keyframe generation focuses on cre-ating multiple coherent events, while the frame interpolation stage generates smooth intermediate frames between keyframes using existing video generation models. A novel mask attention module is further introduced to improve co-herence and efficiency. Experiments on challenging real-world videos demonstrate that the proposed method outper-forms prior methods by up to 9.5% in objective metrics. Hsin-Ping Huang, Yu-Chuan Su, Ming-Hsuan Yang 0001 |
WACV | 2 |
| 2025 | Fine-grained Controllable Video Generation via Object Appearance and ContextabstractWhile text-to-video generation shows state-of-the-art results, fine-grained output control remains challenging for users relying solely on natural language prompts. In this work, we present FACTOR for fine-grained controllable video generation. FACTOR provides an intuitive interface where users can manipulate the trajectory and appearance of individual objects in conjunction with a text prompt. We propose a unified framework to integrate these control signals into an existing text-to-video model. Our approach involves a multimodal condition module with a joint encoder, control-attention layers, and an appearance augmentation mechanism. This design enables FACTOR to generate videos that closely align with detailed user specifications. Extensive experiments on standard benchmarks and user-provided inputs demonstrate a notable improvement in controllability by FACTOR over competitive baselines. Hsin-Ping Huang, Yu-Chuan Su, Deqing Sun, Lu Jiang 0004, Xuhui Jia, Yukun Zhu, Ming-Hsuan Yang 0001 |
WACV | 2 |
| 2024 | Instruct-Imagen: Image Generation with Multi-modal InstructionabstractThis paper presents Instruct-Imagen, a model that tackles heterogeneous image generation tasks and generalizes across unseen tasks. We introduce multi-modal in-struction for image generation, a task representation artic-ulating a range of generation intents with precision. It uses natural language to amalgamate disparate modalities (e.g., text, edge, style, subject, etc.), such that abundant generation intents can be standardized in a uniform format. We then build Instruct - Imagen by fine-tuning a pre-trained text-to-image diffusion model with two stages. First, we adapt the model using the retrieval-augmented training, to enhance model's capabilities to ground its generation on external multi-modal context. Subsequently, we fine-tune the adapted model on diverse image generation tasks that requires vision-language understanding (e.g., subject-driven generation, etc.), each paired with a multi-modal instruction encapsulating the task's essence. Human evaluation on various image generation datasets re-veals that Instruct-Imagen matches or surpasses prior task-specific models in-domain and demonstrates promising generalization to unseen and more complex tasks. Our evaluation suite will be made publicly available. Hexiang Hu, Kelvin C. K. Chan, Yu-Chuan Su, Wenhu Chen, Yandong Li, Kihyuk Sohn, Xue Ben, Boqing Gong, William W. Cohen, Ming-Wei Chang, Xuhui Jia |
CVPR | 3 |
| 2024 | A Versatile Diffusion Transformer with Mixture of Noise Levels for Audiovisual GenerationabstractTraining diffusion models for audiovisual sequences allows for a range of generation tasks by learning conditional distributions of various input-output combinations of the two modalities. Nevertheless, this strategy often requires training a separate model for each task which is expensive. Here, we propose a novel training approach to effectively learn arbitrary conditional distributions in the audiovisual space. Our key contribution lies in how we parameterize the diffusion timestep in the forward diffusion process. Instead of the standard fixed diffusion timestep, we propose applying variable diffusion timesteps across the temporal dimension and across modalities of the inputs. This formulation offers flexibility to introduce variable noise levels for various portions of the input, hence the term mixture of noise levels. We propose a transformer-based audiovisual latent diffusion model and show that it can be trained in a task-agnostic fashion using our approach to enable a variety of audiovisual generation tasks at inference time. Experiments demonstrate the versatility of our method in tackling cross-modal and multimodal interpolation tasks in the audiovisual space. Notably, our proposed approach surpasses baselines in generating temporally and perceptually consistent samples conditioned on the input. Project page: neurips13025.github.io Gwanghyun Kim, Alonso Martinez, Yu-Chuan Su, Brendan Jou, José Lezama, Agrim Gupta, Lijun Yu, Lu Jiang 0004, Aren Jansen, Jacob Walker, Krishna Somandepalli |
NeurIPS | 3 |
| 2023 | Towards Authentic Face Restoration with Iterative Diffusion Models and BeyondabstractAn authentic face restoration system is becoming increasingly demanding in many computer vision applications, e.g., image enhancement, video communication, and taking portrait. Most of the advanced face restoration models can recover high-quality faces from low-quality ones but usually fail to faithfully generate realistic and high-frequency details that are favored by users. To achieve authentic restoration, we propose IDM, an Iteratively learned face restoration system based on denoising Diffusion Models (DDMs). We define the criterion of an authentic face restoration system, and argue that denoising diffusion models are naturally endowed with this property from two aspects: intrinsic iterative refinement and extrinsic iterative enhancement. Intrinsic learning can preserve the content well and gradually refine the high-quality details, while extrinsic enhancement helps clean the data and improve the restoration task one step further. We demonstrate superior performance on blind face restoration tasks. Beyond restoration, we find the authentically cleaned data by the proposed restoration system is also helpful to image generation tasks in terms of training stabilization and sample quality. Without modifying the models, we achieve better quality than state-of-the-art on FFHQ and ImageNet generation using either GANs or diffusion models. Tingbo Hou, Yu-Chuan Su, Xuhui Jia, Yandong Li, Matthias Grundmann 0002 |
ICCV | 3 |
| 2022 | Rethinking Deep Face RestorationabstractA model that can authentically restore a low-quality face image to a high-quality one can benefit many applications. While existing approaches for face restoration make significant progress in generating high-quality faces, they often fail to preserve facial features that compromise the authenticity of reconstructed faces. Because the human visual system is very sensitive to faces, even minor changes may significantly degrade the perceptual quality. In this work, we argue that the problems of existing models can be traced down to the two sub-tasks of the face restoration problem, i.e. face generation and face reconstruction, and the fragile balance between them. Based on the observation, we propose a new face restoration model that improves both generation and reconstruction. Besides the model improvement, we also introduce a new evaluation metric for measuring models' ability to preserve the identity in the restored faces. Extensive experiments demonstrate that our model achieves state-of-the-art performance on multiple face restoration benchmarks, and the proposed metric has a higher correlation with user preference. The user study shows that our model produces higher quality faces while better preserving the identity 86.4% of the time compared with state-of-the-art methods. Yu-Chuan Su, Chun-Te Chu, Yandong Li, Marius Renn, Yukun Zhu, Changyou Chen, Xuhui Jia |
CVPR | 2 |
| 2022 | 2.5D visual relationship detection
Yu-Chuan Su, Soravit Changpinyo, Xiangning Chen, Sathish Thoppay, Cho-Jui Hsieh, Lior Shapira, Radu Soricut, Hartwig Adam, Matthew Brown 0001, Ming-Hsuan Yang 0001, Boqing Gong |
Comput. Vis. Image Underst. | 1 |
| 2022 | Learning Spherical Convolution for $360^{\circ }$360∘ RecognitionabstractWhile 360 cameras offer tremendous new possibilities in vision, graphics, and augmented reality, the spherical images they produce make visual recognition non-trivial. Ideally, 360 imagery could inherit the convolutional neural networks (CNNs) trained with great success on perspective projection images. However, existing methods to transfer CNNs from perspective to spherical images introduce significant computational costs and/or degradations in accuracy. We propose to learn a Spherical Convolution Network (SphConv) that translates a planar CNN to the equirectangular projection of 360 images. Given a source CNN for perspective images, SphConv learns to reproduce the flat filter outputs on 360 data. The key benefits are 1) efficient and accurate recognition for 360 images, and 2) the ability to leverage pre-trained networks for perspective images. We propose two instantiations of SphConv---Spherical Kernel, which learns location dependent kernels on the sphere, and Kernel Transformer Network, which learns a functional transformation that generates SphConv from the source CNN. Validating our approach with multiple source CNNs and datasets, we show that it successfully preserves the source CNN's accuracy, while offering efficiency, transferability, and scalability to typical image resolutions. We further introduce a spherical Faster R-CNN based on SphConv and show that we can learn a spherical object detector without any object annotations in 360 images. Yu-Chuan Su, Kristen Grauman |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2021 | Learning Compressible 360$^{\circ }$∘ Video IsomersabstractStandard video encoders developed for conventional narrow field-of-view video are widely applied to 360°video as well, with reasonable results. However, while this approach commits arbitrarily to a projection of the spherical frames, we observe that some orientations of a 360°video, once projected, are more compressible than others. We introduce an approach to predict the sphere rotation that will yield the maximal compression rate. Given video clips in their original encoding, a convolutional neural network learns the association between a clip's visual content and its compressibility at different rotations of a cubemap projection. Given a novel video, our learning-based approach efficiently infers the most compressible direction in one shot, without repeated rendering and compression of the source video. We validate our idea on thousands of video clips and multiple popular video codecs. The results show that this untapped dimension of 360°compression has substantial potential-“good” rotations are typically 8-18 percent more compressible than bad ones, and our learning approach can predict them reliably 78 percent of the time. Yu-Chuan Su, Kristen Grauman |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2019 | Kernel Transformer Networks for Compact Spherical ConvolutionabstractIdeally, 360° imagery could inherit the deep convolutional neural networks (CNNs) already trained with great success on perspective projection images. However, existing methods to transfer CNNs from perspective to spherical images introduce significant computational costs and/or degradations in accuracy. We present the Kernel Transformer Network (KTN) to efficiently transfer convolution kernels from perspective images to the equirectangular projection of 360° images. Given a source CNN for perspective images as input, the KTN produces a function parameterized by a polar angle and kernel as output. Given a novel 360° image, that function in turn can compute convolutions for arbitrary layers and kernels as would the source CNN on the corresponding tangent plane projections. Distinct from all existing methods, KTNs allow model transfer: the same model can be applied to different source CNNs with the same base architecture. This enables application to multiple recognition tasks without re-training the KTN. Validating our approach with multiple source CNNs and datasets, we show that KTNs improve the state of the art for spherical convolution. KTNs successfully preserve the source CNN's accuracy, while offering transferability, scalability to typical image resolutions, and, in many cases, a substantially lower memory footprint. Yu-Chuan Su, Kristen Grauman |
CVPR | 1 |
| 2018 | Learning Compressible 360° Video IsomersabstractStandard video encoders developed for conventional narrow field-of-view video are widely applied to 360° video as well, with reasonable results. However, while this approach commits arbitrarily to a projection of the spherical frames, we observe that some orientations of a 360° video, once projected, are more compressible than others. We introduce an approach to predict the sphere rotation that will yield the maximal compression rate. Given video clips in their original encoding, a convolutional neural network learns the association between a clip's visual content and its compressibility at different rotations of a cubemap projection. Given a novel video, our learning-based approach efficiently infers the most compressible direction in one shot, without repeated rendering and compression of the source video. We validate our idea on thousands of video clips and multiple popular video codecs. The results show that this untapped dimension of 360° compression has substantial potential-"good" rotations are typically 8-10% more compressible than bad ones, and our learning approach can predict them reliably 82% of the time. Yu-Chuan Su, Kristen Grauman |
CVPR | 1 |
| 2017 | Making 360° Video Watchable in 2D: Learning Videography for Click Free Viewingabstract360° Video requires human viewers to actively control where to look while watching the video. Although it provides a more immersive experience of the visual content, it also introduces additional burden for viewers, awkward interfaces to navigate the video lead to suboptimal viewing experiences. Virtual cinematography is an appealing direction to remedy these problems, but conventional methods are limited to virtual environments or rely on hand-crafted heuristics. We propose a new algorithm for virtual cinematography that automatically controls a virtual camera within a 360° video. Compared to the state of the art, our algorithm allows more general camera control, avoids redundant outputs, and extracts its output videos substantially more efficiently. Experimental results on over 7 hours of real in the wild video show that our generalized camera control is crucial for viewing 360° video, while the proposed efficient algorithm is essential for making the generalized control computationally tractable. Yu-Chuan Su, Kristen Grauman |
CVPR | 1 |
| 2017 | Learning Spherical Convolution for Fast Features from 360° ImageryabstractWhile 360° cameras offer tremendous new possibilities in vision, graphics, and augmented reality, the spherical images they produce make core feature extraction non-trivial. Convolutional neural networks (CNNs) trained on images from perspective cameras yield “flat" filters, yet 360° images cannot be projected to a single plane without significant distortion. A naive solution that repeatedly projects the viewing sphere to all tangent planes is accurate, but much too computationally intensive for real problems. We propose to learn a spherical convolutional network that translates a planar CNN to process 360° imagery directly in its equirectangular projection. Our approach learns to reproduce the flat filter outputs on 360° data, sensitive to the varying distortion effects across the viewing sphere. The key benefits are 1) efficient feature extraction for 360° images and video, and 2) the ability to leverage powerful pre-trained networks researchers have carefully honed (together with massive labeled image training sets) for perspective images. We validate our approach compared to several alternative methods in terms of both raw CNN output accuracy as well as applying a state-of-the-art “flat" object detector to 360° data. Our method yields the most accurate results while saving orders of magnitude in computation versus the existing exact reprojection solution. Yu-Chuan Su, Kristen Grauman |
NIPS | 1 |
| 2016 | Pano2Vid: Automatic Cinematography for Watching 360° Videos
Yu-Chuan Su, Dinesh Jayaraman, Kristen Grauman |
ACCV (4) | 1 |
| 2016 | Detecting Engagement in Egocentric Video
Yu-Chuan Su, Kristen Grauman |
ECCV (5) | 1 |
| 2016 | Leaving Some Stones Unturned: Dynamic Feature Prioritization for Activity Detection in Streaming Video
Yu-Chuan Su, Kristen Grauman |
ECCV (7) | 1 |
| 2015 | Combination of feature engineering and ranking models for paper-author identification in KDD cup 2013
Chun-Liang Li, Yu-Chuan Su, Ting-Wei Lin, Cheng-Hao Tsai, Wei-Cheng Chang, Kuan-Hao Huang, Tzu-Ming Kuo, Shan-Wei Lin, Young-San Lin, Yu-Chen Lu, Chun-Pai Yang, Cheng-Xia Chang, Wei-Sheng Chin, Yu-Chin Juan, Hsiao-Yu Fish Tung, Jui-Pin Wang, Cheng-Kuang Wei, Felix Wu, Tu-Chun Yin, Tong Yu 0001, Yong Zhuang, Shou-De Lin, Hsuan-Tien Lin, Chih-Jen Lin |
J. Mach. Learn. Res. | 2 |
| 2014 | Effective string processing and matching for author disambiguation
Wei-Sheng Chin, Yong Zhuang, Yu-Chin Juan, Felix Wu, Hsiao-Yu Fish Tung, Tong Yu 0001, Jui-Pin Wang, Cheng-Xia Chang, Chun-Pai Yang, Wei-Cheng Chang, Kuan-Hao Huang, Tzu-Ming Kuo, Shan-Wei Lin, Young-San Lin, Yu-Chen Lu, Yu-Chuan Su, Cheng-Kuang Wei, Tu-Chun Yin, Chun-Liang Li, Ting-Wei Lin, Cheng-Hao Tsai, Shou-De Lin, Hsuan-Tien Lin, Chih-Jen Lin |
J. Mach. Learn. Res. | 16 |
| 2014 | Scalable Mobile Visual Classification by Kernel Preserving Projection Over High-Dimensional FeaturesabstractScalable mobile visual classification-classifying images/videos in a large semantic space on mobile devices in real-time-is an emerging problem as observing the paradigm shift towards mobile platforms and the explosive growth of visual data. Though seeing the advances in detecting thousands of concepts in the servers, the scalability is handicapped in mobile devices due to the severe resource constraints within. However, certain emerging applications require such scalable visual classification with prompt response for detecting local contexts (e.g., Google Glass) or ensuring user satisfaction. In this work, we point out the ignored challenges for scalable mobile visual classification and provide a feasible solution. To overcome the limitations of mobile visual classification, we propose an unsupervised linear dimension reduction algorithm, kernel preserving projection (KPP), which approximates the kernel matrix of high dimensional features with low dimensional linear embedding. We further introduce sparsity to the projection matrix to ensure its compliance with mobile computing (with merely 12% non-zero entries). By inspecting the similarity of linear dimension reduction with low-rank linear distance metric and Taylor expansion of RBF kernel, we justified the feasibility for the proposed KPP method over high-dimensional features. Experimental results on three public datasets confirm that the proposed method outperforms existing dimension reduction methods. What is even more, we can greatly reduce the storage consumption and efficiently compute the classification results on the mobile devices. Yu-Chuan Su, Tzu-Hsuan Chiu, Yin-Hsi Kuo, Chun-Yen Yeh, Winston H. Hsu |
IEEE Trans. Multim. | 1 |
| 2013 | Enabling low bitrate mobile visual recognition: a performance versus bandwidth evaluationabstractThe rapid development of technologies in both hardware and software have made content-based multimedia services feasible on mobile devices such as smartphones and tablets; and the strong needs for mobile visual search and recognition have been emerging. While many real applications of visual recognition require a large scale recognition systems, the same technologies that support server-based scalable visual recognition may not be feasible on mobile devices due to the resource constraints. Although the client-server framework ensures the scalability, the real-time response subjects to the limitation on network bandwidth. Therefore, the main challenge for mobile visual recognition system should be the recognition bitrate, which is the amount of data transmission under the same recognition performance. For this work, we exploit and compare various strategies such as compact features, feature compression, feature signatures by hashing, image scaling, etc., to enable low bitrate mobile visual recognition. We argue that thumbnail image is a competitive candidate for low bitrate visual recognition because it carries multiple features at once and multi-feature fusion is important as the size of semantic space increases. Our evaluations on two subsets of ImageNet, both contain more than 10,000 images with 19 and 137 categories, verify the efficacy of thumbnail images. We further suggest a new strategy that combines single (local) feature signature and the thumbnail image, which achieves significant bitrate reduction from (average) 102,570 to 4,661 bytes with merely (overall) 10% performance degradation. Yu-Chuan Su, Tzu-Hsuan Chiu, Yan-Ying Chen, Chun-Yen Yeh, Winston H. Hsu |
ACM Multimedia | 1 |
| 2013 | Flickr-tag prediction using multi-modal fusion and meta informationabstractWe present our evaluation and analysis on Yahoo! Large-scale Flickr-tag Image Classification dataset. Our evaluations show that combining multi-features and different classification models, the MAP of tag prediction can be significantly improve over ordinary linear classification. Further analysis shows that some tags are given not because of the visual content but the meta information of images. Our experiments show that we can make more accurate prediction on certain tags using meta information without any training process, compared with visual content based classifiers. Combine the meta information, multi-features and multi-models fusion, we achieve significantly better performance than simple linear classification. We also evaluate the performance of various mid-level feature, and the results suggest that "Concept Bank" feature may be a promising direction for the task. Yu-Chuan Su, Tzu-Hsuan Chiu, Guan-Long Wu, Chun-Yen Yeh, Felix Wu, Winston H. Hsu |
ACM Multimedia | 1 |
| 2012 | Evaluating Gaussian Like Image Representations over Local FeaturesabstractRecently, several Gaussian like image representations are proposed as an alternative of the bag-of-word representation over local features. These representations are proposed to overcome the quantization error problem faced in bag-of-word representation. They are shown to be effective in different applications, the Extended Hierarchical Gaussianization reached excellent performance using single feature in VOC2009, Vector of Locally Aggregated Descriptors and Fisher Kernel reached excellent performance using only signature like representation on Holiday dataset. Despite their success and similarity, no comparative study about these representations has been made. In this paper, we perform a systematic comparison about three emerging different gaussian like representations: Extended Hierarchical Gaussianization, Fisher Kernel and Vector of Locally Aggregated Descriptors. We evaluate the performance and the influence of feature and parameters of these representations on Holiday and CC_Web_Video datasets, and several important properties about these representations have been observed during our investigation. This study provides better understanding about these gaussian like image representations that are believed to be promising in various applications. Yu-Chuan Su, Guan-Long Wu, Tzu-Hsuan Chiu, Winston H. Hsu, Kuo-Wei Chang |
ICME | 1 |
| 2012 | Sharing the trees among random forests for effective and efficient concept detectionabstractIn this paper, we focus on the random forest based concept detection system, and we intend to improve the efficiency of the system in testing phase and to save memory and storage usages by reducing the total number of trees (classifiers). However, reducing the tree number often results in poor performance. In this article, we proposed a method called tree-sharing to cope with this issue. Unlike the traditional method that treats each concept independently, our work shares the trees among concepts, and leave the most important ones from the view of whole system. Experiments on different concept sets show tree-sharing can greatly reduce the number of total trees while the performance decreases slightly. Even in the worst case, we achieve 80% of original performance with only 5% of trees. Tzu-Hsuan Chiu, Guan-Long Wu, Yu-Chuan Su, Winston H. Hsu |
MMSP | 3 |
| 2011 | Scalable mobile video question-answering system with locally aggregated descriptors and random projectionabstractWe present a scalable mobile video Question-Answering system with locally aggregated descriptors and random projection using user-generated videos all around the world for ACM Multimedia 2011 Technicolor challenge: "precise event recognition and description from video excerpts." Our proposed system takes a video excerpt as a query and explores its canonical semantics of that. We collect a public events video dataset containing 7 topics with 1963 YouTube videos to evaluate our proposed system. The experiment results show that our proposed video feature representation outperforms a state-of-the-art near-duplicate retrieval based on color histogram. The signature generated by random projection not only ensures real-time efficiency but also achieves a competitive MAP performance with original feature. Guan-Long Wu, Yu-Chuan Su, Tzu-Hsuan Chiu, Liang-Chi Hsieh, Winston H. Hsu |
ACM Multimedia | 2 |
| 2008 | Development of a dual-stage virtual metrology architecture for TFT-LCD manufacturingabstractProcessing quality of thin film transistor-liquid crystal display (TFT-LCD) manufacturing is a key factor for production yield. In general, the processing quality of production equipment is not only related to its own manufacturing process but also affected by the process result of the previous equipment. Current proposed virtual metrology (VM) architectures are all applied to conjecture processing quality of a single stage. These single-stage VM architectures lack the ability of detecting the processing drift occur between different stages. This paper proposes a novel dual-stage VM architecture for quality conjecturing that involves two pieces of equipment. The architecture consists of two stages. In Stage I, the fore-equipment VM model is established as usual and then the rear-equipment VM model that depends on stage-I VM output is built in Stage II. An example of the proposed dual-stage VM architecture applied to two sets of TFT-LCD chemical vapor deposition (CVD) equipment is presented. Experimental results demonstrate that this dual-stage VM architecture is promising for TFT-LCD manufacturing. Yu-Chuan Su, Wen-Huang Tsai, Fan-Tien Cheng, Wei-Ming Wu |
ICRA | 1 |
| 2007 | Method for Evaluating Reliance Level of a Virtual Metrology SystemabstractA method for evaluating reliance level of a virtual metrology system (VMS) is proposed. This method calculates a reliance index (RI) value between 0 and 1 by analyzing the process data of production equipment to decide if the virtual metrology result is reliable. A RI threshold is also defined in this method. If a RI value is higher than the threshold, the conjecture result is reliant; otherwise, the conjecture result needs to be further examined. In addition to the RI, the method also proposes process data similarity indices (SIs). The SIs are defined to evaluate the degree of similarity between the input set of process data and those historical sets of process data used to establish the conjecture model. Two kinds of SIs are included in the method: global similarity index (GSI) and individual similarity index (ISI). Both the GSI and ISI are applied to assist the RI in gauging the reliance level and locating the key parameter(s) that cause major deviation, hence the VMS manufacturability problem is resolved. An illustrative example with 300-mm semiconductor foundry production equipment in Taiwan is demonstrated in this work. The real experimental results show that this method is applicable to the VMS of (such as semiconductor and TFT-LCD) production equipment. Fan-Tien Cheng, Yeh-Tung Chen, Yu-Chuan Su, Deng-Lin Zeng |
ICRA | 3 |
| 2003 | Development of a generic tester for distributed object-oriented systemsabstractWith the popularity of PC and the coming of Internet era, more and more software systems are constructed under a distributed object-oriented environment. Software systems constructed with the distributed object-oriented technology have the advantages of openness, modularization, flexibility and easy maintenance. Also, a distributed software system can efficiently solve the problems that need complicated computing. However, testing and debugging in the later stage of software development demand great deal of resources, and open-environment architecture of general distributed systems tends to increase the complexity of testing. To resolve the problems mentioned above, this work proposed a Generic Tester for distributed object-oriented systems. This Generic Tester is applicable to the tests of an individual component or module and even the whole distributed object-oriented system as long as the functions and operations of the components or system can be presented with only class diagrams (as well as interface definitions) and sequence diagrams generated by the tools used during software development. Research results indicate that this Generic Tester enables an integrated planning of software development and testing, reduces testing cost, and improves overall development efficiency. Fan-Tien Cheng, Chin-Hui Wang, Yu-Chuan Su |
ICRA | 3 |