VLDB 2026 Research / reviewers in the wild / expert
Ming-Ching Chang
dblp:21/4361
· DBLP profile ↗
82ranked-venue papers
12as first author
44since 2021 · last 2026
0000-0001-9325-5341ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 64 · 7 first-author · 34 since 2021Artificial intelligence and machine learning · 38 · 7 first-author · 18 since 2021Systems, architecture and hardware · 3 · 2 since 2021Computer networks · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | How Bias Binds: Measuring Hidden Associations for Bias Control in Text-to-Image CompositionsabstractText-to-image generative models often exhibit bias related to sensitive attributes. However, current research tends to focus narrowly on single-object prompts with limited contextual diversity. In reality, each object or attribute within a prompt can contribute to bias. For example, the prompt ``an assistant wearing a pink hat'' may reflect female-inclined biases associated with a pink hat. Neglecting joint semantic bindings in prompts leads to significant failures of current debiasing methods. We present a preliminary investigation into how bias manifests under semantic binding, where contextual associations between objects and attributes affect generative outcomes. We demonstrate that the underlying bias distribution can be amplified based on these associations. Therefore, we introduce a bias adherence score that quantifies how specific object-attribute bindings activate bias. To delve deeper, we develop a training-free context-bias control framework to explore how token decoupling can facilitate the debiasing of semantic bindings. This framework achieves over 10% debiasing improvement in compositional generation tasks. Our analysis of bias scores across various attribute-object bindings and token decorrelation highlights a fundamental challenge: reducing bias without disrupting essential semantic relationships. These findings expose critical limitations in current debiasing approaches when applied to semantically bound contexts, underscoring the need to reassess prevailing bias mitigation strategies. Jeng-Lin Li, Ming-Ching Chang, Wei-Chao Chen |
AAAI | 2 |
| 2026 | PreClaim-GCT: Self-supervised Graph-Transformer Pretraining for Incident Outcome Prediction from Healthcare Claims
Ming-Ching Chang, Emily Leckman-Westin |
AIME (2) | 2 |
| 2026 | LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval
Weiheng Lu, Jian Li 0062, An Yu, Ming-Ching Chang |
ICPR (5) | 4 |
| 2026 | PatchEAD: Unifying Industrial Visual Prompting Frameworks for Patch-Exclusive Anomaly Detection
Jeng-Lin Li, Po-Hsuan Huang, Ming-Ching Chang, Wei-Chao Chen |
WACV | 4 |
| 2025 | Training-Free Image Manipulation Localization Using Diffusion ModelsabstractImage manipulation localization (IML) is a critical technique in media forensics, focusing on identifying tampered regions within manipulated images. Most existing IML methods require extensive training on labeled datasets with both image-level and pixel-level annotations. These methods often struggle with new manipulation types and exhibit low generalizability. In this work, we propose a training-free IML approach using diffusion models. Our method adaptively selects an appropriate number of diffusion timesteps for each input image in the forward process and performs both conditional and unconditional reconstructions in the backward process without relying on external conditions. By comparing these reconstructions, we generate a localization map highlighting regions of manipulation based on inconsistencies. Extensive experiments were conducted using sixteen state-of-the-art (SoTA) methods across six IML datasets. The results demonstrate that our training-free method outperforms SoTA unsupervised and weakly-supervised techniques. Furthermore, our method competes effectively against fully-supervised methods on novel (unseen) manipulation types. Zhenfei Zhang, Ming-Ching Chang, Xin Li 0005 |
AAAI | 2 |
| 2025 | E4: Energy-Efficient DNN Inference for Edge Video Analytics via Early Exiting and DVFSabstractDeep neural network (DNN) models are increasingly popular in edge video analytic applications. However, the computeintensive nature of DNN models pose challenges for energyefficient inference on resource-constrained edge devices. Most existing solutions focus on optimizing DNN inference latency and accuracy, often overlooking energy efficiency. They also fail to account for the varying complexity of video frames, leading to sub-optimal performance in edge video analytics. In this paper, we propose an EnergyEfficient Early-Exit (E4) framework that enhances DNN inference efficiency for edge video analytics by integrating a novel early-exit mechanism with dynamic voltage and frequency scaling (DVFS) governors. It employs an attentionbased cascade module to analyze video frame diversity and automatically determine optimal DNN exit points. Additionally, E4 features a just-in-time (JIT) profiler that uses coordinate descent search to co-optimize CPU and GPU clock frequencies for each layer before the DNN exit points. Extensive evaluations demonstrate that E4 outperforms current state-of-the-art methods, achieving up to 2.8× speedup and 26% average energy saving while maintaining high accuracy. Yang Zhao 0020, Ming-Ching Chang, Changyao Lin, Jie Liu 0001 |
AAAI | 3 |
| 2025 | Latent Orthogonal Perturbation: A Data Poisoning Approach to Prevent Unauthorized LoRA Fine-TuningabstractGenerative AI (GenAI) enables creators to produce high-quality images without advanced artistic skills, but also raises concerns over the unauthorized replication of distinctive styles. With the rise of fine-tuning techniques like LoRA, mimicking specific artists’ styles has become increasingly accessible. To address this, we propose Latent Orthogonal Perturbation (LO-Perturbation), a data poisoning method that subtly injects orthogonal noise into the latent space of artworks. These perturbations introduce only subtle visual changes but significantly degrade the model’s ability to learn fine-grained stylistic features during LoRA fine-tuning. We evaluate our method using five standard image quality metrics. Results show that, while Glaze and Nightshade prioritize preserving visual fidelity, LO-Perturbation leads to significantly greater degradation in image quality and alignment under LoRA fine-tuning—achieving a 56.92% lower Q-Align score and a 46.18% higher FID, along with the lowest CLIP-IQA and LIQE scores. Our method directly disrupts the fine-tuning process, a threat that is less emphasized in existing protection techniques, thus offering an additional layer of protection for artists. Hong Bin Tan, Ching-Yun Fu, Ming-Ching Chang, Wei-Chao Chen |
AVSS | 3 |
| 2025 | Enhancing LLM Question Answering with RAG through Dense Vector Search and Re-RankingabstractRetrieval-Augmented Generation (RAG) has emerged as a powerful framework for enhancing Large Language Models (LLMs) by incorporating external knowledge through information retrieval (IR) techniques. However, in question-answering tasks, RAG often retrieves documents that are only semantically similar to the query, which may not provide the most relevant information for generating accurate responses. To address this limitation, we propose an improved retrieval pipeline that combines dense vector search with a re-ranking mechanism to more effectively identify and extract highly relevant knowledge from the retrieved content. We evaluated our approach on two Chinese datasets, TTQA and TMMLU+, using 17 different LLMs. Experimental results show that our method improves performance by up to 21.24% over baseline approaches, particularly on two finance-related subsets, after incorporating domain-specific financial regulations to enhance the knowledge base used in the TMMLU+ dataset. Te-Lun Yang, Jyi-Shane Liu, Yuen-Hsien Tseng, Jyh-Shing Roger Jang, Ming-Ching Chang, Wei-Chao Chen |
AVSS | 5 |
| 2025 | BF-YOLOv7: Enhancing Helmet Rule Violation Detection
Chun-Ming Tsai, Jun-Wei Hsieh, Ming-Ching Chang |
IEA/AIE (2) | 3 |
| 2025 | RLMiniStyler: Light-weight RL Style Agent for Arbitrary Sequential Neural Style GenerationabstractArbitrary style transfer aims to apply the style of any given artistic image to another content image. Still, existing deep learning-based methods often require significant computational costs to generate diverse stylized results. Motivated by this, we propose a novel reinforcement learning-based framework for arbitrary style transfer RLMiniStyler. This framework leverages a unified reinforcement learning policy to iteratively guide the style transfer process by exploring and exploiting stylization feedback, generating smooth sequences of stylized results while achieving model lightweight. Furthermore, we introduce an uncertainty-aware multi-task learning strategy that automatically adjusts loss weights to adapt to the content and style balance requirements at different training stages, thereby accelerating model convergence. Through a series of experiments across image various resolutions, we have validated the advantages of RLMiniStyler over other state-of-the-art methods in generating high-quality, diverse artistic image sequences at a lower cost. Codes are available at https://github.com/fengxiaoming520/RLMiniStyler. Jing Hu 0009, Chengming Feng, Shu Hu 0001, Ming-Ching Chang, Xin Li 0005, Xi Wu 0004, Xin Wang 0045 |
IJCAI | 4 |
| 2025 | A Comprehensive Evaluation of Encrypted DNN Inference MethodsabstractFully Homomorphic Encryption (FHE) in the realm of deep neural network (DNN) encrypted inference represents a pivotal advancement in privacy-preserving machine learning. This technology allows users to securely access DNN inference services hosted on remote servers without compromising their personal privacy. Given its wide range of potential applications, FHE has garnered significant research attention. Despite its rapid progress, FHE still faces considerable challenges, particularly the high computational resource demands. Moreover, varying configurations of FHE schemes can lead to notable differences in performance, whether in terms of efficiency or inference accuracy, making it difficult to strike an optimal balance tailored to specific application requirements.To tackle this challenge, we introduce a new approach that simulates FHE-induced errors to assess the impact of different FHE architectures and parameter configurations on encrypted inference during the model testing phase. Our simulation framework enables efficient approximation of homomorphic computation outcomes on a DNN model, specifically for third-generation FHEs, without the need for executing the complete homomorphic process. This method significantly streamlines the research process for optimizing parameters based on specific application needs. We validate the effectiveness of our approach through extensive performance benchmarking across a variety of experimental settings. Yu-Te Ku, Ming-Chien Ho, Feng-Hao Liu, Chih-Fan Hsu, Ming-Ching Chang, Shih-Hao Hung, Wei-Chao Chen |
ISCAS | 6 |
| 2025 | Physics-Guided Exposure Parameter Estimation for Image Metadata VerificationabstractThe rapid rise of AI-generated and edited images has made it increasingly difficult to distinguish authentic photographs from manipulated content. Digital image forensics addresses this challenge by detecting inconsistencies between image content and metadata. Among various forensic cues, the exposure triangle parameters namely ISO Speed Ratings (ISO), aperture (F-number), and shutter speed offer a physically grounded reference for verifying authenticity. We propose a physics-guided latent triad regression framework that predicts these parameters directly from image pixel content while enforcing exposure value ($E V$) consistency through the exposure equation. Our model predicts$I S O, F$-number, and$E V$in log space, deriving shutter speed to ensure physically coherent and non-redundant predictions. Trained on RAISE-2K, it achieves strong correlations with ground-truth parameters ($R^{2} \approx 0.32$to 0.35) and high$\text{E V}$consistency ($R^{2}=0.69$). By embedding physical exposure laws into learning, the framework produces interpretable, exposure-consistent predictions that enhance metadata verification, camera provenance analysis, and image authenticity assessment. Sharmilee Rajkumar Rajan, Ming-Ching Chang, Pradeep K. Atrey |
ISM | 2 |
| 2025 | A Semantically Impactful Image Manipulation Dataset: Characterizing Image Manipulations Using Semantic SignificanceabstractWe investigate how to characterize semantic significance (SS) in detecting image manipulations (IMD) for media forensics. We introduce the Characterization of Seman-tic Impact for IMD (CSI-IMD) dataset, which focuses on localizing and evaluating the semantic impact of image manipulations to counter advanced generative techniques. Our evaluation of 10 state-of-the-art IMD and localization methods on CSI-IMD reveals key insights. Unlike existing datasets, CSI-IMD provides detailed semantic annotations beyond traditional manipulation masks, aiding in the development of new defensive strategies. The dataset features manipulations from advanced generation methods, offering various levels of semantic significance. It is divided into two parts: a gold-standard set of 1,000 manu-ally annotated manipulations with high-quality control, and an extended set of 500,000 automated manipulations for large-scale training and analysis. We also propose a new SS-focused task to assess the impact of semantically targeted manipulations. Our experiments show that current IMD methods struggle with manipulations created using stable diffusion, with TruFor and Cat-Net performing the best among those tested. The CSI-IMD dataset will become available at https://github.com/csiimd/csiimd. Ming-Ching Chang, Matthias Kirchner, Zhenfei Zhang, Xin Li 0005, Arslan Basharat, Anthony Hoogs |
WACV | 2 |
| 2025 | Who Brings the Frisbee: Probing Hidden Hallucination Factors in Large Vision-Language Model via Causality AnalysisabstractRecent advancements in large vision-language models (LVLM) have significantly enhanced their ability to comprehend visual inputs alongside natural language. How-ever, a major challenge in their real-world application is hallucination, where LVLMs generate non-existent visual elements, eroding user trust. The underlying mechanism driving this multimodal hallucination is poorly understood. Minimal research has illuminated whether contexts such as sky, tree, or grass field involve the LVLM in hallucinating a frisbee. We hypothesize that hiddenfactors, such as objects, contexts, and semantic foreground-background structures, induce hallucination. This study proposes a novel causal approach: a hallucination probing system to identify these hidden factors. By analyzing the causality between images, text prompts, and network saliency, we systematically ex-plore interventions to block these factors. Our experimen-tal findings show that a straightforward technique based on our analysis can significantly reduce hallucinations. Additionally, our analyses indicate the potential to edit network internals to minimize hallucinated outputs. Po-Hsuan Huang, Jeng-Lin Li, Chin-Po Chen, Ming-Ching Chang, Wei-Chao Chen |
WACV | 4 |
| 2025 | Unsupervised domain adaptation for cross-style, cross-year land use understanding from historical mapsabstractDigitizing historical topographic maps is essential for spatial analysis in GIS; however, conventional methods for digitizing these maps are labor-intensive and challenging due to non-explicit boundaries and inconsistent map styles. We address these challenges by proposing a new Map Style Segmentation (MapStyleSeg) method that employs unsupervised domain adaptation (UDA) from deep learning (DL) to enhance cross-style, cross-year automatic map segmentation and conversion. Our method, MapStyleSeg, is exemplified by training on a fully annotated topographic map of Taiwan in 2017 and applying it to a 2001 topographic map without annotations. We also evaluated different encoder-decoder architectures and loss functions. Our results show that using the ResNet-101 backbone with the SegFormer decoder and a mix of focal and Dice loss yields the best performance: 94.94% overall accuracy (Acc), 81.8% mean Intersection over Union (mIoU), outperforming standard U-Net models without UDA (88.23% Acc, 49.3% mIoU). Our approach addresses the challenges of digitizing historical maps with varying styles, further advancing GIS digitization of historical maps, and offering useful information for urban planning, environmental monitoring, and decision-making processes. This work highlights the novel use of DL algorithms to automate complex GIS data processing that transforms historical maps into spatial datasets. Jun-Hua Wang, Andy Da-Yu Wang, Hsiung-Ming Liao, Ming-Ching Chang, Richard Tzong-Han Tsai |
Int. J. Geogr. Inf. Sci. | 4 |
| 2025 | Optimizing Encrypted Neural Networks: Model Design, Quantization and Fine-Tuning Using FHEW/TFHEabstractThird-generation Fully Homomorphic Encryption (FHE), particularly the FHEW/TFHE schemes, is recognized for its balanced security requirements, small parameters, and low memory usage, though the current methods in the scenarios of Deep Neural Network (DNN) inference still have high computational costs, limiting the practical applicability. This work demonstrates how to improve practicality of the third-generation technologies for DNN tasks while preserving its key advantages. Our work focuses on two main contributions. First, we developed a computational architecture called FHE-Neuron, which reconfigures the parameters and bootstrapping structure of traditional FHEW/TFHE Boolean operations. This architecture significantly reducing the cost of encrypted DNN inference by dynamically switching the precision of encrypted data during computation—using high precision for cost-effective linear operations and low precision for computationally expensive nonlinear operations. Second, we introduced an FHE-aware Quantization and Fine-tuning framework that optimizes model parameters to align with FHE-Neuron’s constraints, ensuring high accuracy in encrypted inference. We validate our approach on various neural network models across several computing platforms. In our experiments, our method achieves one-image inference time on average 4.5 milliseconds for MNIST and 17 milliseconds for Fashion MNIST, achieving accuracy rates of 96.52% and 88.57% respectively. For the CIFAR-10 dataset, our system completes one image inference in 30 seconds with a 90.5% accuracy rate. Yu-Te Ku, Feng-Hao Liu, Chih-Fan Hsu, Ming-Ching Chang, Shih-Hao Hung, I-Ping Tu, Wei-Chao Chen |
Proc. Priv. Enhancing Technol. | 4 |
| 2025 | Scale-Aware Crowd Counting Network With Annotation Error ModelingabstractTraditional crowd-counting networks suffer from information loss when feature maps are reduced by pooling layers, leading to inaccuracies in counting crowds at a distance. Existing methods often assume correct annotations during training, disregarding the impact of noisy annotations, especially in crowded scenes. Furthermore, using a fixed Gaussian density model does not account for the varying pixel distribution of the camera distance. To overcome these challenges, we propose a Scale-Aware Crowd Counting Network (SACC-Net) that introduces a scale-aware loss function with error-compensation capabilities of noisy annotations. For the first time, we simultaneously model labeling errors (mean) and scale variations (variance) by spatially varying Gaussian distributions to produce fine-grained density maps for crowd counting. Furthermore, the proposed scale-aware Gaussian density model can be dynamically approximated with a low-rank approximation, leading to improved convergence efficiency with comparable accuracy. To create a smoother scale-aware feature space, this paper proposes a novel Synthetic Fusion Module (SFM) and an Intra-block Fusion Module (IFM) to generate fine-grained heat maps for better crowd counting. The lightweight version of our model, named SACC-LW, enhances the computational efficiency while retaining accuracy. The superiority and generalization properties of scale-aware loss function are extensively evaluated for different backbone architectures and performance metrics on six public datasets: UCF-QNRF, UCF CC 50, NWPU, ShanghaiTech A, ShanghaiTech B, and JHU. Experimental results also demonstrate that SACC-Net outperforms all state-of-the-art methods, validating its effectiveness in achieving superior crowd-counting accuracy. The source code is available at https://github.com/Naughty725. Yi-Kuan Hsieh, Jun-Wei Hsieh, Xin Li 0005, Yu-Ming Zhang, Yu-Chee Tseng, Ming-Ching Chang |
IEEE Trans. Image Process. | 6 |
| 2025 | Open-Set Occluded Person Identification With mmWave RadarabstractRadio frequency sensors can penetrate non-metal objects and provide complementary information to vision sensors for person identification (PID) purposes. However, there is a lack of research on millimeter wave (mmWave) radar for PID under occlusions, particularly in addressing the open-set recognition problem. Thus, we propose an open-set occluded PID (OSO-PID) framework that can deal with various obstacle and occlusion scenarios with open-set recognition capability. We first introduce a new dataset, mmWave-ocPID, comprising mmWave radar measurements and RGB-depth images, collected from 23 human subjects. We next design a novel neural network, mm-PIDNet, for occluded person identification using mmWave radar measurements. mm-PIDNet incorporates a transformer encoder, a bidirectional long short-term memory module, and a novel supervised contrastive learning module to improve PID performance. For open-set recognition, we enhance the mmWave radar-based PID method by integrating supervised contrastive learning with the Weibull models, which can identify out-of-distribution samples. We perform extensive indoor experiments with a variety of obstacles and occlusion scenarios. Our experimental results show that mm-PIDNet achieves an F1-score of 0.93 on average, outperforming state-of-the-art methods by up to 13.41% for occluded cases. For open-set PID, the OSO-PID framework achieves an F1-score above 0.8 when the openness is less than 14.36%. Tao Wang 0118, Yang Zhao 0020, Ming-Ching Chang, Jie Liu 0001 |
IEEE Trans. Mob. Comput. | 3 |
| 2025 | Spotting the Fakes: A Deep Dive into GAN-Generated Face DetectionabstractGenerative Adversarial Networks (GANs) have enabled the creation of highly authentic facial images, which are increasingly used in deceptive social media profiles and other forms of disinformation, resulting in serious consequences. Significant progress has been made in developing GAN-generated face detection systems to identify these fake images. This study offers a comprehensive review of recent advancements in GAN-generated face detection, focusing on techniques that detect facial images generated by GAN models. We categorize detection methods into three groups: (1) deep learning-based approaches, (2) physics-based methods, and (3) physiology-based methods. We summarize key concepts in each category, connecting them to relevant implementations, datasets, and evaluation metrics. Additionally, we provide a comparative analysis between automated detection and human visual performance to highlight the strengths and weaknesses of both approaches. Furthermore, we review related surveys, including detecting morphed faces, manipulated faces, DeepFake, and faces generated by diffusion models. Finally, we discuss unresolved challenges and suggest potential directions for future research. Xin Wang 0045, Ting Yu Tsai, Shu Hu 0001, Ming-Ching Chang, Pradeep K. Atrey, Siwei Lyu |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2024 | Pushing the Limit of Fine-Tuning for Few-Shot Learning: Where Feature Reusing Meets Cross-Scale AttentionabstractDue to the scarcity of training samples, Few-Shot Learning (FSL) poses a significant challenge to capture discriminative object features effectively. The combination of transfer learning and meta-learning has recently been explored by pre-training the backbone features using labeled base data and subsequently fine-tuning the model with target data. However, existing meta-learning methods, which use embedding networks, suffer from scaling limitations when dealing with a few labeled samples, resulting in suboptimal results. Inspired by the latest advances in FSL, we further advance the approach of fine-tuning a pre-trained architecture by a strengthened hierarchical feature representation. The technical contributions of this work include: 1) a hybrid design named Intra-Block Fusion (IBF) to strengthen the extracted features within each convolution block; and 2) a novel Cross-Scale Attention (CSA) module to mitigate the scaling inconsistencies arising from the limited training samples, especially for cross-domain tasks. We conducted comprehensive evaluations on standard benchmarks, including three in-domain tasks (miniImageNet, CIFAR-FS, and FC100), as well as two cross-domain tasks (CDFSL and Meta-Dataset). The results have improved significantly over existing state-of-the-art approaches on all benchmark datasets. In particular, the FSL performance on the in-domain FC100 dataset is more than three points better than the latest PMF (Hu et al. 2022). Ying-Yu Chen, Jun-Wei Hsieh, Xin Li 0005, Ming-Ching Chang |
AAAI | 4 |
| 2024 | SMILEtrack: SiMIlarity LEarning for Occlusion-Aware Multiple Object TrackingabstractDespite recent progress in Multiple Object Tracking (MOT), several obstacles such as occlusions, similar objects, and complex scenes remain an open challenge. Meanwhile, a systematic study of the cost-performance tradeoff for the popular tracking-by-detection paradigm is still lacking. This paper introduces SMILEtrack, an innovative object tracker that effectively addresses these challenges by integrating an efficient object detector with a Siamese network-based Similarity Learning Module (SLM). The technical contributions of SMILETrack are twofold. First, we propose an SLM that calculates the appearance similarity between two objects, overcoming the limitations of feature descriptors in Separate Detection and Embedding (SDE) models. The SLM incorporates a Patch Self-Attention (PSA) block inspired by the vision Transformer, which generates reliable features for accurate similarity matching. Second, we develop a Similarity Matching Cascade (SMC) module with a novel GATE function for robust object matching across consecutive video frames, further enhancing MOT performance. Together, these innovations help SMILETrack achieve an improved trade-off between the cost (e.g., running speed) and performance (e.g., tracking accuracy) over several existing state-of-the-art benchmarks, including the popular BYTETrack method. SMILETrack outperforms BYTETrack by 0.4-0.8 MOTA and 2.1-2.2 HOTA points on MOT17 and MOT20 datasets. Code is available at http://github.com/pingyang1117/SMILEtrack_official. Yu-Hsiang Wang, Jun-Wei Hsieh, Ping-Yang Chen, Ming-Ching Chang, Hung-Hin So, Xin Li 0005 |
AAAI | 4 |
| 2024 | A New Benchmark and Model for Challenging Image Manipulation DetectionabstractThe ability to detect manipulation in multimedia data is vital in digital forensics. Existing Image Manipulation Detection (IMD) methods are mainly based on detecting anomalous features arisen from image editing or double compression artifacts. All existing IMD techniques encounter challenges when it comes to detecting small tampered regions from a large image. Moreover, compression-based IMD approaches face difficulties in cases of double compression of identical quality factors. To investigate the State-of-The-Art (SoTA) IMD methods in those challenging conditions, we introduce a new Challenging Image Manipulation Detection (CIMD) benchmark dataset, which consists of two subsets, for evaluating editing-based and compression-based IMD methods, respectively. The dataset images were manually taken and tampered with high-quality annotations. In addition, we propose a new two-branch network model based on HRNet that can better detect both the image-editing and compression artifacts in those challenging conditions. Extensive experiments on the CIMD benchmark show that our model significantly outperforms SoTA IMD methods on CIMD. The dataset is available at: https://github.com/ZhenfeiZ/CIMD. Zhenfei Zhang, Mingyang Li 0007, Ming-Ching Chang |
AAAI | 3 |
| 2024 | Image Manipulation Detection with Implicit Neural Representation and Limited Supervision
Zhenfei Zhang, Mingyang Li 0007, Xin Li 0005, Ming-Ching Chang, Jun-Wei Hsieh |
ECCV (88) | 4 |
| 2024 | Improving Limited Supervised Foot Ulcer Segmentation Using Cross-Domain Augmentation StrategiesabstractDiabetic foot ulcers pose health risks, including higher morbidity, mortality, and amputation rates. Monitoring wound areas is crucial for proper care, but manual segmentation is subjective due to complex wound features and background variation. Expert annotations are costly and time-intensive, thus hampering large dataset creation. Existing segmentation models relying on extensive annotations are impractical in real-world scenarios with limited annotated data. In this paper, we propose a cross-domain augmentation method named TransMix that combines Augmented Global Pre-training (AGP) and Localized CutMix Fine-tuning (LCF) to enrich wound segmentation data for model learning. TransMix can effectively improve the foot ulcer segmentation model training by leveraging other dermatology datasets not on ulcer skins or wounds. AGP effectively increases the overall image variability, while LCF increases the diversity of wound regions. Experimental results show that TransMix increases the variability of wound regions and substantially improves the Dice score for models trained with only 40 annotated images under various proportions. Shang-Jui Kuo, Chia-Ching Lin, Jeng-Lin Li, Ming-Ching Chang |
ICASSP | 5 |
| 2024 | Invited Paper: Efficient Design of FHEW/TFHE Bootstrapping Implementation with Scalable ParametersabstractFully Homomorphic Encryption (FHE) is vital for computing over encrypted data, thereby enabling numerous privacy-preserving applications. This work focuses on the third generation FHE schemes (e.g., FHEW and TFHE), known for their fast bootstrapping, small FHE parameters, and robust security built on milder assumptions. Ming-Chien Ho, Yu-Te Ku, Feng-Hao Liu, Chih-Fan Hsu, Ming-Ching Chang, Shih-Hao Hung, Wei-Chao Chen |
ICCAD | 6 |
| 2024 | Learning With Instance-Dependent Noisy Labels By Anchor Hallucination And Hard Sample Label CorrectionabstractLearning from noisy-labeled data is crucial for real-world applications. Traditional Noisy-Label Learning (NLL) methods categorize training data into clean and noisy sets based on the loss distribution of training samples. However, they often neglect that clean samples, especially those with intricate visual patterns, may also yield substantial losses. This oversight is particularly significant in datasets with Instance-Dependent Noise (IDN), where mislabeling probabilities correlate with visual appearance. Our approach explicitly distinguishes between clean $v s$. noisy and easy $v s$. hard samples. We identify training samples with small losses, assuming they have simple patterns and correct labels. Utilizing these easy samples, we hallucinate multiple anchors to select hard samples for label correction. Corrected hard samples, along with the easy samples, are used as labeled data in subsequent semi-supervised training. Experiments on synthetic and real-world IDN datasets demonstrate the superior performance of our method over other state-of-the-art NLL methods. Po-Hsuan Huang, Chia-Ching Lin, Chih-Fan Hsu, Ming-Ching Chang, Wei-Chao Chen |
ICIP | 4 |
| 2024 | Set-Nas: Sample-Efficient Training For Neural Architecture Search With Strong Predictor And Stratified SamplingabstractSample-efficient neural architecture search (NAS) techniques have advanced rapidly. Two lines of methods, namely neural predictor and sequential search, have shown promising performance in improving the sample efficiency of NAS. However, as far as we know, little attention has been paid to the middle ground between these two lines. Inspired by the analogy between NAS and evolutionary optimization, we propose a new Sample-Efficient Training for NAS (SETNAS) based on strategies that improve fitness scores and sampling mechanisms. We develop a strong neural predictor called the Fully Bidirectional Graph Convolutional Network evolutionary (Fully-BiGCN) that significantly enhances the predictor capability of the features in each layer. The developed predictor is embedded into an iterative stratified sampling process to retain only a subset of best-fit architectures using the same training budget. SET-NAS achieves remarkable results compared to the state-of-the-art in predictor-based NAS. Using NASBench-201 as the benchmark, SET-NAS takes only $27.1 \%$ (CIFAR-10), $49.0 \%$ (CIFAR-100), and $51.75 \%$ (ImageNet-16) of training cost of other state-of-the-art predictor-based methods to find the promising network architecture. Yu-Ming Zhang, Jun-Wei Hsieh, Yu-Hsiu Chang, Xin Li 0005, Ming-Ching Chang, Chun-Chieh Lee, Kuo-Chin Fan |
ICIP | 5 |
| 2024 | MOTE-NAS: Multi-Objective Training-based Estimate for Efficient Neural Architecture SearchabstractNeural Architecture Search (NAS) methods seek effective optimization toward performance metrics regarding model accuracy and generalization while facing challenges regarding search costs and GPU resources. Recent Neural Tangent Kernel (NTK) NAS methods achieve remarkable search efficiency based on a training-free model estimate; however, they overlook the non-convex nature of the DNNs in the search process. In this paper, we develop Multi-Objective Training-based Estimate (MOTE) for efficient NAS, retaining search effectiveness and achieving the new state-of-the-art in the accuracy and cost trade-off. To improve NTK and inspired by the Training Speed Estimation (TSE) method, MOTE is designed to model the actual performance of DNNs from macro to micro perspective by draw loss landscape and convergence speed simultaneously. Using two reduction strategies, the MOTE is generated based on a reduced architecture and a reduced dataset. Inspired by evolutionary search, our iterative ranking-based, coarse-to-fine architecture search is highly effective. Experiments on NASBench-201 show MOTE-NAS achieves 94.32% accuracy on CIFAR-10, 72.81% on CIFAR-100, and 46.38% on ImageNet-16-120, outperforming NTK-based NAS approaches. An evaluation-free (EF) version of MOTE-NAS delivers high efficiency in only 5 minutes, delivering a model more accurate than KNAS. Jun-Wei Hsieh, Xin Li 0005, Ming-Ching Chang, Chun-Chieh Lee, Kuo-Chin Fan |
NeurIPS | 4 |
| 2023 | GAN-Generated Faces Detection: A Survey and New PerspectivesabstractGenerative Adversarial Networks (GAN) have led to the generation of very realistic face images, which have been used in fake social media accounts and other disinformation matters that can generate profound impacts. Therefore, the corresponding GAN-face detection techniques are under active development that can examine and expose such fake faces. In this work, we aim to provide a comprehensive review of recent progress in GAN-face detection. We focus on methods that can detect face images that are generated or synthesized from GAN models. We classify the existing detection works into four categories: (1) deep learning-based, (2) physical-based, (3) physiological-based methods, and (4) evaluation and comparison against human visual performance. For each category, we summarize the key ideas and connect them with method implementations. We also discuss open problems and suggest future research directions. Xin Wang 0045, Shu Hu 0001, Ming-Ching Chang, Siwei Lyu |
ECAI | 4 |
| 2023 | Fisheye Multiple Object Tracking by Learning Distortions Without DewarpingabstractWe develop a new Multiple Object Tracking (MOT) scheme for fisheye cameras that can directly perform vehicle detection, re-identification, and tracking under fisheye distortions without explicit dewarping. Fisheye cameras provide omnidirectional coverage that is wider than traditional cameras, reducing fewer need of cameras to monitor road intersections. However, the problem of distorted views introduces new challenges for fisheye MOT. In this paper, we propose a Fish-Eye Multiple Object Tracking (FEMOT) approach with two novelties. We develop the Distorted Fisheye Image Augmentation (DFIA) method to improve object detection and re-identification on fisheye cameras, where fisheye model training can be performed on existing datasets of traditional cameras via fisheye data synthesis and augmentation. We also develop the Hybrid Data Association (HDA) method to perform tracking directly on fisheye views, without the need of de-warping. The developed FEMOT framework provides practical design and advancement that enables large-scale use of fisheye cameras in smart city and surveillance applications. Ping-Yang Chen, Jun-Wei Hsieh, Ming-Ching Chang, Munkhjargal Gochoo, Fang-Pang Lin, Yong-Sheng Chen |
ICIP | 3 |
| 2022 | Instance Contour Adjustment via Structure-Driven CNN
Shuchen Weng, Ming-Ching Chang, Boxin Shi |
ECCV (7) | 3 |
| 2022 | Vocbench: A Neural Vocoder Benchmark for Speech SynthesisabstractNeural vocoders, used for converting the spectral representations of an audio signal to the waveforms, are a commonly used component in speech synthesis pipelines. It focuses on synthesizing waveforms from low-dimensional representation, such as Mel-Spectrograms. In recent years, different approaches have been introduced to develop such vocoders. However, it becomes more challenging to assess these new vocoders and compare their performance to previous ones. To address this problem, we present VocBench, a framework that benchmark the performance of state-of-the-art neural vocoders. VocBench uses a systematic study to evaluate different neural vocoders in a shared environment that enables a fair comparison between them. In our experiments, we use the same setup for datasets, training pipeline, and evaluation metrics for all neural vocoders. We perform a subjective and objective evaluation to compare the performance of each vocoder along a different axis. Our results demonstrate that the framework can show competitive efficacy and quality of the synthesized samples for each vocoder. VocBench framework is available at https://github.com/facebookresearch/vocoder-benchmark. Ehab A. AlBadawy, Andrew Gibiansky, Jilong Wu, Ming-Ching Chang, Siwei Lyu |
ICASSP | 5 |
| 2022 | Eyes Tell All: Irregular Pupil Shapes Reveal GAN-Generated FacesabstractGenerative adversarial network (GAN) generated high-realistic human faces are visually challenging to discern from real ones. They have been used as profile images for fake social media accounts, which leads to high negative social impacts. In this work, we show that GAN-generated faces can be exposed via irregular pupil shapes. This phenomenon is caused by the lack of physiological constraints in the GAN models. We demonstrate that such artifacts exist widely in high-quality GAN-generated faces. We design an automatic method to segment the pupils from the eyes and analyze their shapes to distinguish GAN-generated faces from real ones. Qualitative and quantitative evaluations of our method on the Flickr-Faces-HQ dataset and a StyleGAN2 generated face dataset demonstrate the effectiveness and simplicity of our method. Shu Hu 0001, Xin Wang 0045, Ming-Ching Chang, Siwei Lyu |
ICASSP | 4 |
| 2022 | Text-Image De-Contextualization Detection Using Vision-Language ModelsabstractText-image de-contextualization, which uses inconsistent image-text pairs, is an emerging form of misinformation and drawing increasing attention due to the great threat to information authenticity. With real content but semantic mismatch in multiple modalities, the detection of de-contextualization is a challenging problem in media forensics. Inspired by the recent advances in vision-language models with powerful relationship learning between images and texts, we leverage the vision-language models to the media de-contextualization detection task. Two popular models, namely CLIP and VinVL, are evaluated and compared on several news and social media datasets to show their performance in detecting image-text inconsistency in de-contextualization. We also summarize interesting observations and shed lights to the use of vision-language models in de-contextualization detection. Mingzhen Huang, Shan Jia, Ming-Ching Chang, Siwei Lyu |
ICASSP | 3 |
| 2022 | Improving Class Activation Map for Weakly Supervised Object LocalizationabstractWe propose a Weakly Supervised Object Localization (WSOL) method that can locate an object within a given image using a pre-trained network learned with only class labels without location annotations. Most existing WSOL methods rely on thresholding a Class Activation Map (CAM) generated by the pre-trained network to highlight and localize the object. However such approaches often produce incomplete object bounding boxes, as only the discriminative parts of the object are selected during thresholding. We revisit current CAM-based WSOL approaches and propose a pipeline to: (1) refine the CAM map using Weighted Global Average Pooling (WGAP), (2) recombine weights to make use of the negative features, (3) adaptively select a suitable threshold to achieve better object localization. Our method does not require additional learning or hyperparameter tuning. We show that our simple approach can achieve competitive results when evaluated on the CUB-200-2011 and ILSVRC 2016 datasets against other state-of-the-art methods. Zhenfei Zhang, Ming-Ching Chang, Tien D. Bui |
ICASSP | 2 |
| 2022 | Simultaneous multi-person tracking and activity recognition based on cohesive cluster search
Wenbo Li 0001, Yi Wei 0006, Siwei Lyu, Ming-Ching Chang |
Comput. Vis. Image Underst. | 4 |
| 2022 | DetPoseNet: Improving Multi-Person Pose Estimation via Coarse-Pose FilteringabstractHuman detection and pose estimation are essential for understanding human activities in images and videos. Mainstream multi-human pose estimation methods take a top-down approach, where human detection is first performed, then each detected person bounding box is fed into a pose estimation network. This top-down approach suffers from the early commitment of initial detections in crowded scenes and other cases with ambiguities or occlusions, leading to pose estimation failures. We propose the DetPoseNet, an end-to-end multi-human detection and pose estimation framework in a unified three-stage network. Our method consists of a coarse-pose proposal extraction sub-net, a coarse-pose based proposal filtering module, and a multi-scale pose refinement sub-net. The coarse-pose proposal sub-net extracts whole-body bounding boxes and body keypoint proposals in a single shot. The coarse-pose filtering step based on the person and keypoint proposals can effectively rule out unlikely detections, thus improving subsequent processing. The pose refinement sub-net performs cascaded pose estimation on each refined proposal region. Multi-scale supervision and multi-scale regression are used in the pose refinement sub-net to simultaneously strengthen context feature learning. Structure-aware loss and keypoint masking are applied to further improve the pose refinement robustness. Our framework is flexible to accept most existing top-down pose estimators as the role of the pose refinement sub-net in our approach. Experiments on COCO and OCHuman datasets demonstrate the effectiveness of the proposed framework. The proposed method is computationally efficient (5-6x speedup) in estimating multi-person poses with refined bounding boxes in sub-seconds. Lipeng Ke, Ming-Ching Chang, Honggang Qi, Siwei Lyu |
IEEE Trans. Image Process. | 2 |
| 2021 | FlagDetSeg: Multi-Nation Flag Detection and Segmentation in the WildabstractWe present a simple and effective flag detection approach for multi-nation flag instance segmentation in-the-wild based on data augmentation and Mask-RCNN PointRend. To the best of our knowledge, this is the first multi-nation flag detection work incorporating recent deep object detection with code and dataset that will be released for public use. Flag images with binary segmentation are collected from public domain including the Open Image V6 and annotated for up to 225 countries. Additional flag images are generated from template flag images with cropping, warping, masking, and color adaption to hallucinate realistic-looking flag images for training and testing. Data augmentation is performed by fusing and transforming the segmented flags on top of natural image backgrounds to synthesize new images. To cope with the large variability of flags with the lack of authentic annotated flags, we combine the trained binary Mask-RCNN segmentation weights with the new multi-nation classifier for fine-tuning. For evaluation, the proposed model is compared with other popular detectors and instance segmentation methods including YOLACT++. Results show the efficacy of the proposed approach. Shou-Fang Wu, Ming-Ching Chang, Siwei Lyu, Cheng-Shih Wong, Abhineet Kumar Pandey, Po-Chi Su |
AVSS | 2 |
| 2021 | A Video Analytic System for Rail Crossing Point ProtectionabstractWith the rise of AI deep learning, video surveillance based on deep neural networks can provide real-time detection and tracking of vehicles and pedestrians. We present a video analytic system for monitoring railway crossing and providing security protection for rail intersections. Our system can automatically determine the rail-crossing gate status via visual detection and analyze traffic by detecting and tracking passing vehicles, thus to oversee a set of rail-transportation related safety events. Assuming a fixed camera view, each gate RoI can be manually annotated once for each site during system setup, and then gate status can be automatically detected afterwards. Vehicles are detected using YOLOv4 and multi-target tracking is performed using DeepSORT. Safety-related events including trespassing are continuously monitored using rule-based triggering. Experimental evaluation is performed on a Youtube rail crossing dataset as well as a private dataset. On the private dataset of 76 total minutes from 38 videos, our system can successfully detect all 56 events out of 58 annotated events. On the public dataset of 14.21 hrs of videos, it detects 58 out of 62 events. Guangliang Zhao, Abhineet Kumar Pandey, Ming-Ching Chang, Siwei Lyu |
AVSS | 3 |
| 2021 | Multi-Teacher Single-Student Visual Transformer with Multi-Level Attention for Face Spoofing Detection
Yao-Hui Huang, Jun-Wei Hsieh, Ming-Ching Chang, Lipeng Ke, Siwei Lyu, Arpita Samanta Santra |
BMVC | 3 |
| 2021 | Learnable Discrete Wavelet Pooling (LDW-Pooling) for Convolutional Networks
Bor-Shiun Wang, Jun-Wei Hsieh, Ping-Yang Chen, Ming-Ching Chang, Lipeng Ke, Siwei Lyu |
BMVC | 4 |
| 2021 | TransRPN: Towards the Transferable Adversarial Perturbations using Region Proposal Networks and Beyond
Yuezun Li, Ming-Ching Chang, Pu Sun 0001, Honggang Qi, Junyu Dong, Siwei Lyu |
Comput. Vis. Image Underst. | 2 |
| 2021 | Fast Online Video Pose Estimation by Dynamic Bayesian Modeling of Mode TransitionsabstractWe propose a fast online video pose estimation method to detect and track human upper-body poses based on a conditional dynamic Bayesian modeling of pose modes without referring to future frames. The estimation of human body poses from videos is an important task with many applications. Our method extends fast image-based pose estimation to live video streams by leveraging the temporal correlation of articulated poses between frames. Video pose estimation is inferred over a time window using a conditional dynamic Bayesian network (CDBN), which we term time-windowed CDBN. Specifically, latent pose modes and their transitions are modeled and co-determined from the combination of three modules: 1) inference based on current observations; 2) the modeling of mode-to-mode transitions as a probabilistic prior; and 3) the modeling of state-to-mode transitions using a multimode softmax regression. Given the predicted pose modes, the body poses in terms of arm joint locations can then be determined more accurately and robustly. Our method is suitable to investigate high frame rate (HFR) scenarios, where pose mode transitions can effectively capture action-related temporal information to boost performance. We evaluate our method on a newly collected HFR-Pose dataset and four major video pose datasets (VideoPose2, TUM Kitchen, FLIC, and Penn_Action). Our method achieves improvements in both accuracy and efficiency over existing online video pose estimation methods. Ming-Ching Chang, Lipeng Ke, Honggang Qi, Longyin Wen, Siwei Lyu |
IEEE Trans. Cybern. | 1 |
| 2021 | Parallel Residual Bi-Fusion Feature Pyramid Network for Accurate Single-Shot Object DetectionabstractThis paper proposes the Parallel Residual Bi-Fusion Feature Pyramid Network (PRB-FPN) for fast and accurate single-shot object detection. Feature Pyramid (FP) is widely used in recent visual detection, however the top-down pathway of FP cannot preserve accurate localization due to pooling shifting. The advantage of FP is weakened as deeper backbones with more layers are used. In addition, it cannot keep up accurate detection of both small and large objects at the same time. To address these issues, we propose a new parallel FP structure with bi-directional (top-down and bottom-up) fusion and associated improvements to retain high-quality features for accurate localization. We provide the following design improvements: (1) A parallel bifusion FP structure with a bottom-up fusion module (BFM) to detect both small and large objects at once with high accuracy. (2) A concatenation and re-organization (CORE) module provides a bottom-up pathway for feature fusion, which leads to the bi-directional fusion FP that can recover lost information from lower-layer feature maps. (3) The CORE feature is further purified to retain richer contextual information. Such CORE purification in both top-down and bottom-up pathways can be finished in only a few iterations. (4) The adding of a residual design to CORE leads to a new Re-CORE module that enables easy training and integration with a wide range of deeper or lighter backbones. The proposed network achieves state-of-the-art performance on the UAVDT17 and MS COCO datasets. Code is available at https://github.com/pingyang1117/PRBNet_PyTorch. Ping-Yang Chen, Ming-Ching Chang, Jun-Wei Hsieh, Yong-Sheng Chen |
IEEE Trans. Image Process. | 2 |
| 2020 | 3D Single-Person Concurrent Activity Detection Using Stacked Relation NetworkabstractWe aim to detect real-world concurrent activities performed by a single person from a streaming 3D skeleton sequence. Different from most existing works that deal with concurrent activities performed by multiple persons that are seldom correlated, we focus on concurrent activities that are spatio-temporally or causally correlated and performed by a single person. For the sake of generalization, we propose an approach based on a decompositional design to learn a dedicated feature representation for each activity class. To address the scalability issue, we further extend the class-level decompositional design to the postural-primitive level, such that each class-wise representation does not need to be extracted by independent backbones, but through a dedicated weighted aggregation of a shared pool of postural primitives. There are multiple interdependent instances deriving from each decomposition. Thus, we propose Stacked Relation Networks (SRN), with a specialized relation network for each decomposition, so as to enhance the expressiveness of instance-wise representations via the inter-instance relationship modeling. SRN achieves state-of-the-art performance on a public dataset and a newly collected dataset. The relation weights within SRN are interpretable among the activity contexts. The new dataset and code are available at https://github.com/weiyi1991/UA_Concurrent/ Yi Wei 0006, Wenbo Li 0001, Yanbo Fan, Linghan Xu, Ming-Ching Chang, Siwei Lyu |
AAAI | 5 |
| 2020 | MagGAN: High-Resolution Face Attribute Editing with Mask-Guided Generative Adversarial Network
Yi Wei 0006, Zhe Gan, Wenbo Li 0001, Siwei Lyu, Ming-Ching Chang, Lei Zhang 0001, Jianfeng Gao 0001, Pengchuan Zhang |
ACCV (4) | 5 |
| 2020 | Railcar Detection, Identification and Tracking for Rail Yard ManagementabstractWe present a video analytics system combining railcar detection, classification, Federal Railroad Admin. (FRA) text identification, and logo detection into a system for locomotive transportation and yard management. Existing RFID-based systems are limited by sensor deployment and cannot visually identify railcars when they are away. As there are typically tens of tracks and hundreds of railcars in a yard, an automatic vision system is desirable. The proposed AI system is developed for autonomous yard inventory checking, such that the arrival, departure, and movement of individual railcars can be automatically monitored and managed in the facility. Our system consists of multiple cameras with edge computing devices installed at check points (track entrances and branches), such that visual detection and tracking of railcars can be performed and meta-data can be exchanged. After knowing the railcar locations and types, scene text detection is performed to search and recognize FRA ID markings and logos that can uniquely identify each railcar. Information fusion a database in the central hub can further improve railcar identification and reduce errors. Early results on real-world field collected data demonstrate the efficacy of the proposed approach. Ming-Ching Chang, Guangliang Zhao, Abhineet Kumar Pandey, Andrew Pulver, Peter H. Tu |
ICIP | 1 |
| 2020 | Drone-Based Vehicle Flow Estimation and its Application to Traffic Conflict Hotspot Detection at IntersectionsabstractDrones can provide a wider field of view, high mobility and flexibility for monitoring and analyzing traffic flows and safety conditions. In case of a perpendicular viewing angle to the ground, there will be a very less occlusion that can occur and make vehicle tracking be easier. Thus, a drone-based solution will be better for traffic conflict hotspot detection at an interaction. However, due to its observation far from the ground, limited battery time, and bandwidth, this solution should be edge-based and have a good recognition rate in small object detection. However, current edge-based SoTA (state-of-the-art) methods are weak in a small object detection. We propose CoBiF net (Concatenated Bi-Fusion feature pyramid network), a one-stage object detection model for a real-time small object detection, which consists of SPP (spatial pyramid pooling), FE (Feature Extractor), CF (Concatenated Feature) block, and BFM (Bottom-up Fusion Module). CoBiF net is memory-and-bandwidth saving for the most edge devices. Extensive experiments on UA VDT benchmark show the proposed method achieved the SoTA results for the small object detection task in terms of accuracy and efficiency. Ping-Yang Chen, Jun-Wei Hsieh, Munkhjargal Gochoo, Ming-Ching Chang, Chien-Yao Wang, Yong-Sheng Chen, Hong-Yuan Mark Liao |
ICIP | 4 |
| 2020 | Driver License Field Detection Using Real-Time Deep Networks
Chun-Ming Tsai, Jun-Wei Hsieh, Ming-Ching Chang |
IEA/AIE | 3 |
| 2020 | Explainable and Efficient Sequential Correlation Network for 3D Single Person Concurrent Activity DetectionabstractWe present the sequential correlation network (SCN) to improve concurrent activity detection. SCN combines a recurrent neural network and a correlation model hierarchically to model the complex correlations and temporal dynamics of concurrent activities. SCN has several advantages that enable effective learning even from a small dataset for real-world deployment. Unlike the majority of approaches assuming that each subject performs one activity at a time, SCN is end-to- end trainable, i.e., it can automatically learn the inclusive or exclusive relations of concurrent activities. SCN is lightweight in design using only a small set of learnable parameters to model the spatio-temporal correlations of activities. This also enhances the explainability of the learned parameters. Furthermore, the learning of SCN can benefit from the initialization using semantically meaningful priors. We evaluate the proposed method against the state-of-the-art method on two benchmark datasets with human skeletal data, SCN achieves comparable performance to the SOTA but with much faster inference speed and less memory usage. Yi Wei 0006, Wenbo Li 0001, Ming-Ching Chang, Hongxia Jin, Siwei Lyu |
IROS | 3 |
| 2020 | UA-DETRAC: A new benchmark and protocol for multi-object detection and tracking
Longyin Wen, Dawei Du, Zhaowei Cai, Zhen Lei 0001, Ming-Ching Chang, Honggang Qi, Jongwoo Lim, Ming-Hsuan Yang 0001, Siwei Lyu |
Comput. Vis. Image Underst. | 5 |
| 2019 | Graph-to-Graph Energy Minimization for Video Object SegmentationabstractWe describe a new unsupervised video object segmentation (VOS) method based on the graph-to-graph energy minimization, which focuses on exploiting the mutual bootstrapping information between bottom-up (i.e., using pixel/superpixel attributes) and top-down (i.e., using learned appearance and motion cues) processes in a unified framework. Specifically, we construct a graph-to-graph energy function to encode the spatial similarities among superpixels (superpixel-graph) and temporal consistency among regions (region-graph). An efficient heuristic iterative algorithm is used to minimize the energy function to get the optimal assignment of superpixel and region labels to complete the VOS task. Experiments on two challenging benchmarks (i.e., SegTrack v2 and DAVIS) show that the proposed method achieves favorable performance against the state-of-the-art unsupervised VOS methods and comparable performance with the state-of-the-art semi-supervised methods. Yuezun Li, Longyin Wen, Ming-Ching Chang, Siwei Lyu |
AVSS | 3 |
| 2019 | Exploring the Vulnerability of Single Shot Module in Object Detectors via Imperceptible Background Patches
Yuezun Li, Xiao Bian, Ming-Ching Chang, Siwei Lyu |
BMVC | 3 |
| 2019 | Efficient algorithms for graph regularized PLSA for probabilistic topic modeling
Xin Wang 0045, Ming-Ching Chang, Siwei Lyu |
Pattern Recognit. | 2 |
| 2018 | Pixel Offset Regression (POR) for Single-shot Instance SegmentationabstractState-of-the-art instance segmentation methods including Mask-RCNN and MNC are multi-shot, as multiple region of interest (ROI) forward passes are required to distinguish candidate regions. Multi-shot architectures usually achieve good performance on public benchmarks. However, hundreds of ROI forward passes in sequel limits their running efficiency, which is a critical point in several utilities such as vehicle surveillance. As such, we arrange our focus on seeking a well trade-off between performance and efficiency. In this paper, we introduce a novel Pixel Offset Regression (POR) scheme which can simply extend single-shot object detector to single-shot instance segmentation system, i.e., segmenting all instances in a single pass. Our framework is based on VGG161with following four parts: (1) a single-shot detection branch to generate object detections, (2) a segmentation branch to estimate foreground masks, (3) a pixel offset regression branch to effectively estimate the distance and orientation from each pixel to the respective object center and (4) a merging process combining output of each branch to obtain instances. Our framework is evaluated on Berkeley-BDD, KITTI and PASCAL VOC2012 validation set, with comparison against several VGG16 based multi-shot methods. Without whistles and bells, our framework exhibits decent performance, which shows good potential for fast speed required applications. Yuezun Li, Xiao Bian, Ming-Ching Chang, Longyin Wen, Siwei Lyu |
AVSS | 3 |
| 2018 | UA-DETRAC 2018: Report of AVSS2018 & IWT4S Challenge on Advanced Traffic MonitoringabstractA desirable smart traffic-monitoring and street-safety system can elicit and support the intervention of law enforcement agencies or medical staff. Recently, there has been a dramatically higher demand for such smart systems. To this end, the International Workshop on Traffic and Street Surveillance for Safety and Security (IWT4S) was organized in conjunction with the 15th IEEE International Conference on Advanced Video and Signal-based Surveillance (AVSS 2018). Our goal is to advance the state-of-the-art detection and tracking algorithms and provide a comprehensive performance evaluation for them. We evaluate 5 submitted detection and 7 submitted tracking methods on the large-scale UA-DETRAC benchmark, and the results are shared publicly on the website http://detrac-db. rit.albany.edu. We expect this challenge to advance the research and development of new detection and tracking methods for transportation applications. Siwei Lyu, Ming-Ching Chang, Dawei Du, Wenbo Li 0001, Yi Wei 0006, Marco Del Coco, Pierluigi Carcagnì, Arne Schumann, Bharti Munjal, Dinh-Quoc-Trung Dang, Doo-Hyun Choi, Erik Bochinski, Fabio Galasso, Filiz Bunyak, Guna Seetharaman, Jang-Woon Baek, Jong Taek Lee, Kannappan Palaniappan, Kil-Taek Lim, Kiyoung Moon, Kwang-Ju Kim, Lars Wilko Sommer, Meltem Brandlmaier, Minsung Kang, Moongu Jeon, Noor Al-Shakarji, Oliver Acatay, Pyong-Kun Kim, Sikandar Amin, Thomas Sikora, Tien Ba Dinh, Tobias Senst, Vu-Gia-Hy Che, Young-Chul Lim, Yun-Su Chung |
AVSS | 2 |
| 2018 | Robust Adversarial Perturbation on Deep Proposal-based Models
Yuezun Li, Daniel Tian, Ming-Ching Chang, Xiao Bian, Siwei Lyu |
BMVC | 3 |
| 2018 | Multi-Scale Structure-Aware Network for Human Pose Estimation
Lipeng Ke, Ming-Ching Chang, Honggang Qi, Siwei Lyu |
ECCV (2) | 2 |
| 2018 | Multi-Scale Supervised Network for Human Pose EstimationabstractHuman pose estimation is an important topic in computer vision with many applications including gesture and activity recognition. However, pose estimation from image is challenging due to appearance variations, occlusions, clutter background, and complex activities. To alleviate these problems, we develop a robust pose estimation method based on the recent deep conv-deconv modules with two improvements: (1) multi -scale supervision of body keypoints, and (2) a global regression to improve structural consistency of keypoints. We refine keypoint detection heatmaps using layer-wise multi-scale supervision to better capture local contexts. Pose inference via keypoint association is optimized globally using a regression network at the end. Our method can effectively disambiguate keypoint matches in close proximity including the mismatch of left-right body parts, and better infer occluded parts. Experimental results show that our method achieves competitive performance among state-of-the-art methods on the MPII and FLIC datasets. Lipeng Ke, Honggang Qi, Ming-Ching Chang, Siwei Lyu |
ICIP | 3 |
| 2018 | Explain Black-box Image Classifications Using Superpixel-based InterpretationabstractHow to best understand and interpret the decisions of deep neural networks is a crucial topic, as the impact of intelligent deep network systems is prevalent in many applications. We propose a superpixel based method to interpret and explain the results of black-box deep networks in the widely-applied image classification tasks. We perform probabilistic prediction difference analysis upon one or more superpixels clustered from image pixels. Our method generates a superpixel score map visualization that can provide rich interpretation regarding image components. Such interpretation provides supportive/unsupportive likelihood of image regions upon the decisions performed by the black-box classifier. We compare our method against state-of-art pixelwise interpretation methods over the latest deep neural network classifiers on the ImageNet dataset. Results show that our method produces more consistent interpretations in less computation time. Our method also supports interactive interpretation, where users can acquire explanations on specified regions through a convenient interface for a prompt reaction. Yi Wei 0006, Ming-Ching Chang, Yiming Ying, Ser-Nam Lim, Siwei Lyu |
ICPR | 2 |
| 2018 | Multimodal Sensor System for Pressure Ulcer Wound Assessment and CareabstractWe present a multimodal sensor system for wound assessment and pressure ulcer care. Multiple imaging modalities including RGB, three- dimensional (3-D) depth, thermal, multispectral, and chemical sensing are integrated into a portable hand-held probe for real-time wound assessment. Analytic and quantitative algorithms for various assessments including tissue composition, wound measurement in 3-D, temperature profiling, spectral, and chemical vapor analysis are developed. After each assessment scan, 3-D models of the wound are generated on the fly for geometric measurement, while multimodal observations are analyzed to estimate healing progress. Collaboration between developers and clinical practitioners was conducted at the Charlie Norwood VA Medical Center for in-field data collection and experimental evaluation. A total of 133 assessment sessions from 23 enrolled subjects were collected, on which the multimodal data were analyzed and validated with respect to clinical notes associated with each subject. The system can be operated by nontechnical caregivers on a regular basis to aid wound assessment and care. A web portal front-end was developed for clinical decision and telehealth support, where all historical patient data including wound measurements and analysis can be organized online. Ming-Ching Chang, Ting Yu 0003, Jiajia Luo, Kun Duan, Peter H. Tu, Yang Zhao 0020, Nandini Nagraj, Vrinda Rajiv, Michael Priebe, Elena A. Wood, Maximillian Stachura |
IEEE Trans. Ind. Informatics | 1 |
| 2017 | UA-DETRAC 2017: Report of AVSS2017 & IWT4S Challenge on Advanced Traffic MonitoringabstractThe rapid advances of transportation infrastructure have led to a dramatic increase in the demand for smart systems capable of monitoring traffic and street safety. Fundamental to these applications are a community-based evaluation platform and benchmark for object detection and multi-object tracking. To this end, we organize the AVSS2017 Challenge on Advanced Traffic Monitoring, in conjunction with the International Workshop on Traffic and Street Surveillance for Safety and Security (IWT4S), to evaluate the state-of-the-art object detection and multi-object tracking algorithms in the relevance of traffic surveillance. Submitted algorithms are evaluated using the large-scale UA-DETRAC benchmark and evaluation protocol. The benchmark, the evaluation toolkit and the algorithm performance are publicly available from the website http://detrac-db.rit.albany.edu. Siwei Lyu, Ming-Ching Chang, Dawei Du, Longyin Wen, Honggang Qi, Yuezun Li, Yi Wei 0006, Lipeng Ke, Tao Hu 0011, Marco Del Coco, Pierluigi Carcagnì, Dmitriy Anisimov, Erik Bochinski, Fabio Galasso, Filiz Bunyak, Hao Ye 0005, Hong Wang 0014, Kannappan Palaniappan, Koray Ozcan, Li Wang 0033, Liang Wang 0001, Martin Lauer, Nattachai Watcharapinchai, Nenghui Song, Noor Al-Shakarji, Sikandar Amin, Sitapa Watcharapinchai, Tatiana Khanova, Thomas Sikora, Tino Kutschbach, Volker Eiselein, Wei Tian 0001, Xiangyang Xue 0001, Xiaoyi Yu, Yao Lu 0028, Yingbin Zheng, Yongzhen Huang, Yuqi Zhang 0001 |
AVSS | 2 |
| 2017 | Adaptive RNN Tree for Large-Scale Human Action RecognitionabstractIn this work, we present the RNN Tree (RNN-T), an adaptive learning framework for skeleton based human action recognition. Our method categorizes action classes and uses multiple Recurrent Neural Networks (RNNs) in a treelike hierarchy. The RNNs in RNN-T are co-trained with the action category hierarchy, which determines the structure of RNN-T. Actions in skeletal representations are recognized via a hierarchical inference process, during which individual RNNs differentiate finer-grained action classes with increasing confidence. Inference in RNN-T ends when any RNN in the tree recognizes the action with high confidence, or a leaf node is reached. RNN-T effectively addresses two main challenges of large-scale action recognition: (i) able to distinguish fine-grained action classes that are intractable using a single network, and (ii) adaptive to new action classes by augmenting an existing model. We demonstrate the effectiveness of RNN-T/ACH method and compare it with the state-of-the-art methods on a large-scale dataset and several existing benchmarks. Wenbo Li 0001, Longyin Wen, Ming-Ching Chang, Ser-Nam Lim, Siwei Lyu |
ICCV | 3 |
| 2017 | In-bed patient motion and pose analysis using depth videos for pressure ulcer preventionabstractWe present a real-time depth based computer vision system for pressure ulcer prevention, in-bed patient care and monitoring. Our system can effectively determine whether or not a mobility-compromised patient has been correctly repo-sitioned at the required frequency. A depth sensor is used to detect and recognize patient movements, motion patterns, and pose positions. If the patient has stayed in an unchanged pose for too long and needs pressure releasing movements, our system can notify caregivers for repositioning or assistance. Privacy concerns are mitigated by removing the RGB components of the video stream from the camera capturing, and only processing depth measurements. We collaborated with clinical practitioners at the Charlie Norwood VA Medical Center for in-field data collection and experimental evaluation. A web portal front-end is developed such that all historical patient movements, pose positions, and repositioning data can be organized to support telehealth applications. Ming-Ching Chang, Ting Yi, Kun Duan, Jiajia Luo, Peter H. Tu, Michael Priebe, Elena A. Wood, Max E. Stachura |
ICIP | 1 |
| 2017 | Cloud tracking for solar irradiance predictionabstractWe propose a video analytic system to segment and track clouds for the purpose of solar irradiance prediction. Ground-based imaging sensors are used to monitor potential sun occlusions for solar irradiance drop prediction, which can be used to assist power ramp control strategies. Sky images are first rectified to remove fisheye artifacts. Cloud pixels are segmented and classified into low-to-high transparencies. Evolution of cloud boundaries are tracked using optical flow. Tracking extrapolation predicts future cloud movements and deformations as well as potential sun occlusions. To accurately estimate solar irradiance drop, we propose a novel scheme based on back-projecting the predicted sun occlusion onto the evolving cloud boundary to count for cloud transparency and irradiance drop estimation. Experimental results show that short-term solar irradiance drop is predictable with reasonable accuracy. Ming-Ching Chang, Peter H. Tu |
ICIP | 1 |
| 2017 | Hybrid structure hypergraph for online deformable object trackingabstractRecent advances in visual tracking field design part-based model to handle the deformation and occlusion challenges. Previous methods only consider the sole degree of dependencies (e.g., pairwise or high-order dependencies) between object parts in consecutive frames. However, the degree of dependencies of different object parts in consecutive frames are not consistent, especially when large deformation and occlusion happen. To that end, we design a hybrid structure hypergraph based tracker, which use a non-uniform hypergraph to model the dependencies among object parts. The tracking task is further formulated as the dense structures extracting problem on the non-uniform hypergraph, which is solved by an approximate algorithm efficiently. Several experiments are carried out on publicly available online deformable object tracking dataset, i.e., Deform-SOT dataset, to demonstrate the favorable performance of the proposed method against the state-of-the-art online tracking methods. Shengkun Li, Dawei Du, Longyin Wen, Ming-Ching Chang, Siwei Lyu |
ICIP | 4 |
| 2017 | Multi-Camera Multi-Target Tracking with Space-Time-View Hyper-graph
Longyin Wen, Zhen Lei 0001, Ming-Ching Chang, Honggang Qi, Siwei Lyu |
Int. J. Comput. Vis. | 3 |
| 2016 | Co-Regularized PLSA for Multi-Modal LearningabstractMany learning problems in real world applications involve rich datasets comprising multiple information modalities. In this work, we study co-regularized PLSA(coPLSA) as an efficient solution to probabilistic topic analysis of multi-modal data. In coPLSA, similarities between topic compositions of a data entity across different data modalities are measured with divergences between discrete probabilities, which are incorporated as a co-regularizer to augment individual PLSA models over each data modality. We derive efficient iterative learning algorithms for coPLSA with symmetric KL, L2 and L1 divergences as co-regularizers, in each case the essential optimization problem affords simple numerical solutions that entail only matrix arithmetic operations and numerical solution of 1D nonlinear equations. We evaluate the performance of the coPLSA algorithms on text/image cross-modal retrieval tasks, on which they show competitive performance with state-of-the-art methods. Xin Wang 0045, Ming-Ching Chang, Yiming Ying, Siwei Lyu |
AAAI | 2 |
| 2016 | Efficient large-scale photometric reconstruction using Divide-Recon-Fuse 3D Structure from MotionabstractWe propose an efficient framework for large-scale 3D reconstruction from a large set of photos following the Structure-from-Motion (SfM) paradigm with divide-conquer and fusion. Our main novelty is to ensure commonality from overlaps between image sets corresponding to their reconstructions, which facilitates effective stitching and fusion. Specifically, such commonality is ensured by selecting a set of duplicated images (which are termed anchor images) in adjacent image sets prior to the 3D reconstruction. The anchor images can assist accurate fusion of the 3D point clouds. We describe an efficient RANSAC scheme for pairwise stitching. Our method is intuitively scalable to large site reconstruction via subdivision and fusion following a graph construct. We further describe another RANSAC algorithm to improve loop closure in our anchor image approach. Experimental results on reconstructing a large portion of a university campus demonstrate the efficacy of our method. Yueming Yang, Ming-Ching Chang, Longyin Wen, Peter H. Tu, Honggang Qi, Siwei Lyu |
AVSS | 2 |
| 2015 | Fast Online Upper Body Pose Estimation from VideoabstractEstimation of human body poses from video is an important problem in computer vision with many applications. Most existing methods for video pose estimation are offline in nature, where all frames in the video are used in the process to estimate the body pose in each frame. In this work, we describe a fast online video upper body pose estimation method (CDBN-MODEC) that is based on a conditional dynamic Bayesian network model, which predicts upper body pose in a frame without using information from future frames. Our method combines fast single image based pose estimation methods with the temporal correlation of poses between frames. We collect a new high frame rate upper body pose dataset that better reflects practical scenarios calling for fast online video pose estimation. When evaluated on this dataset and the VideoPose2 benchmark dataset, CDBN-MODEC achieves improvements in both performance and running efficiency over several state-of-art online video pose estimation methods. Ming-Ching Chang, Honggang Qi, Xin Wang 0045, Hong Cheng 0002, Siwei Lyu |
BMVC | 1 |
| 2015 | Bridging computer vision and social science: A multi-camera vision system for social interaction training analysisabstractWe investigate the use of a vision-based system capable of estimating social states such as rapport and hostility. We study the correlation between interpretations automatically generated by our system and those reported by social scientists. Our multi-camera vision system collects visual cues including location (proximity), motion, pose, gaze, and facial expressions in real-time from multiple subjects moving freely in an unconstrained environment. We performed experiments on 80+ subjects. Preliminary regression analysis suggests high correlation between machine distilled time series signals and assessments made by human experts. Jixu Chen, Ming-Ching Chang, Tai-Peng Tian, Ting Yu 0003, Peter H. Tu |
ICIP | 2 |
| 2015 | Seeing as it happens: Real time 3D video event visualizationabstractWe present a video event visualization system that can render steerable 3D views of tracked targets onto a reconstructed 3D site representation. The framework takes object tracking meta-data generated from a multi-camera event tracking system as input and produces an immersive 3D playback as a representation of the observation. This 3D representation can provide users seeing as it happens of the events for surveillance applications. Our system can further “animate” the virtual viewing camera and generate a first-person immersive playback of the event, either from the trajectory of a specified real-world target or from a virtual avatar. Such synthetic view can provide additional insights for event recognition for on-line monitoring, investigation, and forensic applications. Yueming Yang, Ming-Ching Chang, Peter H. Tu, Siwei Lyu |
ICIP | 2 |
| 2012 | Spatio-Temporal Phrases for Activity Recognition
Xiaoming Liu 0002, Ming-Ching Chang, Weina Ge, Tsuhan Chen |
ECCV (3) | 3 |
| 2012 | Group context learning for event recognitionabstractWe address the problem of group-level event recognition from videos. The events of interest are defined based on the motion and interaction of members in a group over time. Example events include group formation, dispersion, following, chasing, flanking, and fighting. To recognize these complex group events, we propose a novel approach that learns the group-level scenario context from automatically extracted individual trajectories. We first perform a group structure analysis to produce a weighted graph that represents the probabilistic group membership of the individuals. We then extract features from this graph to capture the motion and action contexts among the groups. The features are represented using the “bag-of-words” scheme. Finally, our method uses the learned Support Vector Machine (SVM) to classify a video segment into the six event categories. Our implementation builds upon a mature multi-camera multi-target tracking system that recognizes the group-level events involving up to 20 individuals in real-time. Weina Ge, Ming-Ching Chang, Xiaoming Liu 0002 |
WACV | 3 |
| 2011 | Gaze and body pose estimation from a distanceabstractWe present a comprehensive approach to track gaze by estimating location, body pose, and head pose direction of multiple individuals in unconstrained environments. The approach combines person detections from fixed cameras with directional face detections obtained from actively controlled pan tilt zoom (PTZ) cameras. The main contribution of this work is to estimate both body pose and head pose (gaze) direction independently from motion direction, using a combination of sequential Monte Carlo Filtering and MCMC sampling. There are numerous benefits in tracking body pose and gaze in surveillance. It allows to track people's focus of attention, can optimize the control of active cameras for biometric face capture, and can provide better interaction metrics between pairs of people. The availability of gaze and face detection information also improves localization and data association for tracking in crowded environments. The performance of the system will be demonstrated on data captured at a real-time surveillance site. Nils Krahnstoever, Ming-Ching Chang, Weina Ge |
AVSS | 2 |
| 2011 | Probabilistic group-level motion analysis and scenario recognitionabstractThis paper addresses the challenge of recognizing behavior of groups of individuals in unconstraint surveillance environments. As opposed to approaches that rely on agglomerative or decisive hierarchical clustering techniques, we propose to recognize group interactions without making hard decisions about the underlying group structure. Instead we use a probabilistic grouping strategy evaluated from the pairwise spatial-temporal tracking information. A path-based grouping scheme determines a soft segmentation of groups and produces a weighted connection graph where its edges express the probability of individuals belonging to a group. Without further segmenting this graph, we show how a large number of low- and high-level behavior recognition tasks can be performed. Our work builds on a mature multi-camera multi-target person tracking system that operates in real-time. We derive probabilistic models to analyze individual track motion as well as group interactions. We show that the soft grouping can combine with motion analysis elegantly to robustly detect and predict group-level activities. Experimental results demonstrate the efficacy of our approach. Ming-Ching Chang, Nils Krahnstoever, Weina Ge |
ICCV | 1 |
| 2011 | Tracking gaze direction from far-field surveillance camerasabstractWe present a real-time approach to estimating the gaze direction of multiple individuals using a network of far-field surveillance cameras. This work is part of a larger surveillance system that utilizes a network of fixed cameras as well as PTZ cameras to perform site-wide tracking of individuals. Based on the tracking information, one or more PTZ cameras are cooperatively controlled to obtain close-up facial images of individuals. Within these close-up shots, face detection and head pose estimation are performed and the results are provided back to the tracking system to track the individual gazes. A new cost metric based on location and gaze orientation is proposed to robustly associate head observations with tracker states. The tracking system can thus leverage the newly obtained gaze information for two purposes: (i) improve the localization of individuals in crowded settings, and (ii) aid high-level surveillance tasks such as understanding gesturing, interactions between individuals, and finding the object-of-interest that people are looking at. In security application, our system can detect if a subject is looking at the security cameras or guard posts. Karthik Sankaranarayanan, Ming-Ching Chang, Nils Krahnstoever |
WACV | 2 |
| 2011 | Measuring 3D shape similarity by graph-based matching of the medial scaffolds
Ming-Ching Chang, Benjamin B. Kimia |
Comput. Vis. Image Underst. | 1 |
| 2010 | Group Level Activity Recognition in Crowded Environments across Multiple CamerasabstractEnvironments such as schools, public parks and prisons and others that contain a large number of people are typically characterized by frequent and complex social interactions. In order to identify activities and behaviors in such environments, it is necessary to understand the interactions that take place at a group level. To this end, this paper addresses the problem of detecting and predicting suspicious and in particular aggressive behaviors between groups of individuals such as gangs in prison yards. The work builds on a mature multi-camera multi-target person tracking system that operates in real-time and has the ability to handle crowded conditions. We consider two approaches for grouping individuals: (i) agglomerative clustering favored by the computer vision community, as well as (ii) decisive clustering based on the concept of modularity, which is favored by the social network analysis community. We show the utility of such grouping analysis towards the detection of group activities of interest. The presented algorithm is integrated with a system operating in real-time to successfully detect highly realistic aggressive behaviors enacted by correctional officers in a simulated prison environment. We present results from these enactments that demonstrate the efficacy of our approach. Ming-Ching Chang, Nils Krahnstoever, Ser-Nam Lim, Ting Yu 0003 |
AVSS | 1 |
| 2009 | Surface reconstruction from point clouds by transforming the medial scaffold
Ming-Ching Chang, Frederic Fol Leymarie, Benjamin B. Kimia |
Comput. Vis. Image Underst. | 1 |
| 2008 | Regularizing 3D medial axis using medial scaffold transformsabstractThis paper addresses a key bottleneck in the use of the 3D medial axis (MA) representation, namely, how the complex MA structure can be regularized so that similar, within-category 3D shapes yield similar 3D MA that are distinct from the non-category shapes. We rely on previous work which (i) constructs a hierarchical MA hypergraph, the medial scaffold (MS), and (ii) the theoretical classification of the instabilities of this structure, or transitions (sudden topological changes due to a small perturbation). The shapes at transition point are degenerate. Our approach is to recognize the transitions which are close-by to a given shape and transform the shape to this transition point, and repeat until no close-by transitions exists. This move towards degeneracy is the basis of simplification of shape. We derive 11 transforms from 7 transitions and follow a greedy scheme in applying the transform. The results show that the simplified MA preserves with-in-category similarity, thus indicating its potential use in various applications including shape analysis, manipulation, and matching. Ming-Ching Chang, Benjamin B. Kimia |
CVPR | 1 |
| 2001 | Fast Search Algorithms for Industrial InspectionabstractThis paper presents an efficient general purpose search algorithm for alignment and an applied procedure for IC print mark quality inspection. The search algorithm is based on normalized cross-correlation and enhances it with a hierarchical resolution pyramid, dynamic programming, and pixel over-sampling to achieve subpixel accuracy on one or more targets. The general purpose search procedure is robust with respect to linear change of image intensity and thus can be applied to general industrial visual inspection. Accuracy, speed, reliability, and repeatability are all critical for the industrial use. After proper optimization, the proposed procedure was tested on the IC inspection platforms in the Mechanical Industry Research Laboratories (MIRL), Industrial Technology Research Institute (ITRI), Taiwan. The proposed method meets all these criteria and has worked well in field tests on various IC products. Ming-Ching Chang, Chiou-Shann Fuh, Hsien-Yei Chen |
Int. J. Pattern Recognit. Artif. Intell. | 1 |