VLDB 2026 Research / reviewers in the wild / expert
Chun-Fu Chen 0001
dblp:48/915 · also Chun-Fu (Richard) Chen, Chun-Fu Richard Chen, Richard Chen 0002
· DBLP profile ↗
44ranked-venue papers
7as first author
28since 2021 · last 2026
0000-0002-5912-5620ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 26 · 4 first-author · 23 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 4 first-author · 10 since 2021Systems, architecture and hardware · 6Computer networks · 3 · 3 since 2021Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Physical Self-Supervised Learning: IMU Sensing without Manual LabelsabstractDeep neural networks have become a promising approach for IMU-based sensing, but their scalability is fundamentally limited by costly labeled data and poor robustness to heterogeneous devices, placements, and users. Existing unsupervised and self-supervised methods reduce but do not remove this dependence, still requiring labeled data for domain adaptation and largely ignoring known physical structure. We propose physical self-supervised learning, an autoencoder-style paradigm for label-free IMU sensing. We replace the conventional neural decoder with an auto-adaptive physics decoder—a learnable family of kinematic equations that enforces explicit physical structure while adapting across environments—and adopt a hybrid two-stage IMU encoder with reconstruction in a structured latent space to mitigate sensor noise. Our framework further introduces probabilistic frequency-spatial constraints to disentangle sensor and object motion, a multi-view kinematic tree to exploit sparse physical self-supervised signals, and an uncertainty-aware formulation to handle the inherent ambiguity of IMU inference. Evaluated on inertial tracking and full-body motion capture over public datasets and realistic deployments, physical self-supervised learning reduces errors by up to 5× for tracking and 4× for motion capture in challenging generalization scenarios, consistently outperforming state-of-the-art supervised and self-supervised baselines without any labels. Yuyang Leng, Renyuan Liu, Shaohan Hu, Peijun Zhao, Chun-Fu Chen 0001, Songqing Chen, Shuochao Yao |
MobiSys | 5 |
| 2025 | Probing LLM Hallucination from Within: Perturbation-Driven Method via Internal Knowledge
Seongmin Lee 0007, Hsiang Hsu, Chun-Fu Chen 0001, Polo Chau |
IEEE Big Data | 3 |
| 2025 | PaLD: Detection of Text Partially Written by Large Language ModelsabstractAdvances in large language models (LLM) have produced text that appears increasingly human-like and difficult to detect with the human eye. In order to mitigate the impact of misusing LLM-generated texts, e.g., copyright infringement, fair student assessment, fraud, and other societally harmful LLM usage, a line of work on detecting human and LLM-written text has been explored. While recent work has focused on classifying entire text samples (e.g., paragraphs) as human or LLM-written, this paper investigates a more realistic setting of mixed-text, where the text's individual segments (e.g., sentences) could each be written by either a human or an LLM. A text encountered in practical usage cannot generally be assumed to be fully human or fully LLM-written; simply predicting whether it is human or LLM-written is insufficient as it does not provide the user with full context on its origins, such as the amount of LLM-written text, or locating the LLM-written parts. Therefore, we study two relevant problems in the mixed-text setting: (i) estimating the percentage of a text that was LLM-written, and (ii) determining which segments were LLM-written. To this end, we propose Partial-LLM Detector (PaLD), a black-box method that leverages the scores of text classifiers. Experimentally, we demonstrate the effectiveness of PaLD compared to baseline methods that build on existing LLM text detectors. Eric Lei, Hsiang Hsu, Chun-Fu Chen 0001 |
ICLR | 3 |
| 2025 | PASS: Private Attributes Protection with Stochastic Data SubstitutionabstractThe growing Machine Learning (ML) services require extensive collections of user data, which may inadvertently include people’s private information irrelevant to the services. Various studies have been proposed to protect private attributes by removing them from the data while maintaining the utilities of the data for downstream tasks. Nevertheless, as we theoretically and empirically show in the paper, these methods reveal severe vulnerability because of a common weakness rooted in their adversarial training based strategies. To overcome this limitation, we propose a novel approach, PASS, designed to stochastically substitute the original sample with another one according to certain probabilities, which is trained with a novel loss function soundly derived from information-theoretic objective defined for utility-preserving private attributes protection. The comprehensive evaluation of PASS on various datasets of different modalities, including facial images, human activity sensory signals, and voice recording datasets, substantiates PASS’s effectiveness and generalizability. Yizhuo Chen, Chun-Fu Chen 0001, Hsiang Hsu, Shaohan Hu, Tarek F. Abdelzaher |
ICML | 2 |
| 2025 | DAF: An Efficient End-to-End Dynamic Activation Framework for on-Device DNN TrainingabstractRecent advancements in on-device training for deep neural networks have underscored the critical need for efficient activation compression to overcome the memory constraints of mobile and edge devices. As activations dominate memory usage during training and are essential for gradient computation, compressing them without compromising accuracy remains a key research challenge. While existing methods for dynamic activation quantization promise theoretical memory savings, their practical deployment is impeded by system-level challenges such as computational overhead and memory fragmentation. Renyuan Liu, Yuyang Leng, Kaiyan Liu, Shaohan Hu, Chun-Fu Chen 0001, Peijun Zhao, Heechul Yun, Shuochao Yao |
MobiSys | 5 |
| 2025 | The Unseen Threat: Residual Knowledge in Machine Unlearning under Perturbed SamplesabstractMachine unlearning offers a practical alternative to avoid full model re-training by approximately removing the influence of specific user data. While existing methods certify unlearning via statistical indistinguishability from re-trained models, these guarantees do not naturally extend to model outputs when inputs are adversarially perturbed. In particular, slight perturbations of forget samples may still be correctly recognized by the unlearned model---even when a re-trained model fails to do so---revealing a novel privacy risk: information about the forget samples may persist in their local neighborhood. In this work, we formalize this vulnerability as residual knowledge and show that it is inevitable in high-dimensional settings. To mitigate this risk, we propose a fine-tuning strategy, named RURK, that penalizes the model’s ability to re-recognize perturbed forget samples. Experiments on vision benchmarks with deep neural networks demonstrate that residual knowledge is prevalent across existing unlearning methods and that our approach effectively prevents residual knowledge. Hsiang Hsu, Pradeep Niroula, Zichang He, Ivan Brugere, Freddy Lécué, Chun-Fu Chen 0001 |
NeurIPS | 6 |
| 2025 | HeavyWater and SimplexWater: Distortion-free LLM Watermarks for Low-Entropy DistributionsabstractLarge language model (LLM) watermarks enable authentication of text provenance, curb misuse of machine-generated text, and promote trust in AI systems. Current watermarks operate by changing the next-token predictions output by an LLM. The updated (i.e., watermarked) predictions depend on random side information produced, for example, by hashing previously generated tokens. LLM watermarking is particularly challenging in low-entropy generation tasks -- such as coding -- where next-token predictions are near-deterministic. In this paper, we propose an optimization framework for watermark design. Our goal is to understand how to most effectively use random side information in order to maximize the likelihood of watermark detection and minimize the distortion of generated text. Our analysis informs the design of two new watermarks: HeavyWater and SimplexWater. Both watermarks are tunable, gracefully trading-off between detection accuracy and text distortion. They can also be applied to any LLM and are agnostic to side information generation. We examine the performance of HeavyWater and SimplexWater through several benchmarks, demonstrating that they can achieve high watermark detection accuracy with minimal compromise of text generation quality, particularly in the low-entropy regime. Our theoretical analysis also reveals surprising new connections between LLM watermarking and coding theory. Dor Tsur, Carol Xuan Long, Claudio Mayrink Verdun, Sajani Vithana, Hsiang Hsu, Chun-Fu Chen 0001, Haim H. Permuter, Flávio P. Calmon |
NeurIPS | 6 |
| 2024 | Dropout-Based Rashomon Set Exploration for Efficient Predictive Multiplicity EstimationabstractPredictive multiplicity refers to the phenomenon in which classification tasks may admit multiple competing models that achieve almost-equally-optimal performance, yet generate conflicting outputs for individual samples.
This presents significant concerns, as it can potentially result in systemic exclusion, inexplicable discrimination, and unfairness in practical applications.
Measuring and mitigating predictive multiplicity, however, is computationally challenging due to the need to explore all such almost-equally-optimal models, known as the Rashomon set, in potentially huge hypothesis spaces.
To address this challenge, we propose a novel framework that utilizes dropout techniques for exploring models in the Rashomon set.
We provide rigorous theoretical derivations to connect the dropout parameters to properties of the Rashomon set, and empirically evaluate our framework through extensive experimentation.
Numerical results show that our technique consistently outperforms baselines in terms of the effectiveness of predictive multiplicity metric estimation, with runtime speedup up to $20\times \sim 5000\times$.
With efficient Rashomon set exploration and metric estimation, mitigation of predictive multiplicity is then achieved through dropout ensemble and model selection. Hsiang Hsu, Guihong Li, Shaohan Hu, Chun-Fu Chen 0001 |
ICLR | 4 |
| 2024 | OVOR: OnePrompt with Virtual Outlier Regularization for Rehearsal-Free Class-Incremental LearningabstractRecent works have shown that by using large pre-trained models along with learnable prompts, rehearsal-free methods
for class-incremental learning (CIL) settings can achieve superior performance to prominent rehearsal-based ones.
Rehearsal-free CIL methods struggle with distinguishing classes from different tasks, as those are not trained together.
In this work we propose a regularization method based on virtual outliers to tighten decision boundaries of the classifier,
such that confusion of classes among different tasks is mitigated.
Recent prompt-based methods often require a pool of task-specific prompts, in order to prevent overwriting knowledge
of previous tasks with that of the new task, leading to extra computation in querying and composing an
appropriate prompt from the pool.
This additional cost can be eliminated, without sacrificing accuracy, as we reveal in the paper.
We illustrate that a simplified prompt-based method can achieve results comparable to
previous state-of-the-art (SOTA) methods equipped with a prompt pool, using much less learnable parameters and lower inference cost.
Our regularization method has demonstrated its compatibility with different prompt-based methods, boosting
those previous SOTA rehearsal-free CIL methods' accuracy on the ImageNet-R and CIFAR-100 benchmarks. Our source code is available at https://github.com/jpmorganchase/ovor. Chun-Fu Chen 0001, Hsiang Hsu |
ICLR | 2 |
| 2024 | Machine Unlearning for Image-to-Image Generative ModelsabstractMachine unlearning has emerged as a new paradigm to deliberately forget data samples from a given model in order to adhere to stringent regulations.
However, existing machine unlearning methods have been primarily focused on classification models, leaving the landscape of unlearning for generative models relatively unexplored.
This paper serves as a bridge, addressing the gap by providing a unifying framework of machine unlearning for image-to-image generative models.
Within this framework, we propose a computationally-efficient algorithm, underpinned by rigorous theoretical analysis, that demonstrates negligible performance degradation on the retain samples, while effectively removing the information from the forget samples.
Empirical studies on two large-scale datasets, ImageNet-1K and Places-365, further show that our algorithm does not rely on the availability of the retain samples, which further complies with data retention policy.
To our best knowledge, this work is the first that represents systemic, theoretical, empirical explorations of machine unlearning specifically tailored for image-to-image generative models. Guihong Li, Hsiang Hsu, Chun-Fu Chen 0001, Radu Marculescu |
ICLR | 3 |
| 2024 | MaSS: Multi-attribute Selective Suppression for Utility-preserving Data Transformation from an Information-theoretic PerspectiveabstractThe growing richness of large-scale datasets has been crucial in driving the rapid advancement and wide adoption of machine learning technologies. The massive collection and usage of data, however, pose an increasing risk for people’s private and sensitive information due to either inadvertent mishandling or malicious exploitation. Besides legislative solutions, many technical approaches have been proposed towards data privacy protection. However, they bear various limitations such as leading to degraded data availability and utility, or relying on heuristics and lacking solid theoretical bases. To overcome these limitations, we propose a formal information-theoretic definition for this utility-preserving privacy protection problem, and design a data-driven learnable data transformation framework that is capable of selectively suppressing sensitive attributes from target datasets while preserving the other useful attributes, regardless of whether or not they are known in advance or explicitly annotated for preservation. We provide rigorous theoretical analyses on the operational bounds for our framework, and carry out comprehensive experimental evaluations using datasets of a variety of modalities, including facial images, voice audio clips, and human activity motion sensor signals. Results demonstrate the effectiveness and generalizability of our method under various configurations on a multitude of tasks. Our source code is available at this URL. Yizhuo Chen, Chun-Fu Chen 0001, Hsiang Hsu, Shaohan Hu, Marco Pistoia, Tarek F. Abdelzaher |
ICML | 2 |
| 2024 | RashomonGB: Analyzing the Rashomon Effect and Mitigating Predictive Multiplicity in Gradient BoostingabstractThe Rashomon effect is a mixed blessing in responsible machine learning. It enhances the prospects of finding models that perform well in accuracy while adhering to ethical standards, such as fairness or interpretability. Conversely, it poses a risk to the credibility of machine decisions through predictive multiplicity. While recent studies have explored the Rashomon effect across various machine learning algorithms, its impact on gradient boosting---an algorithm widely applied to tabular datasets---remains unclear. This paper addresses this gap by systematically analyzing the Rashomon effect and predictive multiplicity in gradient boosting algorithms. We provide rigorous theoretical derivations to examine the Rashomon effect in the context of gradient boosting and offer an information-theoretic characterization of the Rashomon set. Additionally, we introduce a novel inference technique called RashomonGB to efficiently inspect the Rashomon effect in practice. On more than 20 datasets, our empirical results show that RashomonGB outperforms existing baselines in terms of improving the estimation of predictive multiplicity metrics and model selection with group fairness constraints. Lastly, we propose a framework to mitigate predictive multiplicity in gradient boosting and empirically demonstrate its effectiveness. Hsiang Hsu, Ivan Brugere, Freddy Lécué, Chun-Fu Chen 0001 |
NeurIPS | 5 |
| 2024 | DynaSpa: Exploiting Spatial Sparsity for Efficient Dynamic DNN Inference on DevicesabstractRecent advancements in exploring machine learning models' dynamic spatial sparsity have demonstrated great potential for superior efficiency and adaptability without compromising accuracy when compared to conventional static-and-dense DNNs. However, realizing theoretical inference acceleration under practical deployment environments is still faced with significant system challenges. Current vendor libraries and tensor compilers fall short due to their extra data copy operations or insufficient computation schemes, especially for DNN operators with dynamic spatial sparsity. Renyuan Liu, Yuyang Leng, Shilei Tian, Shaohan Hu, Chun-Fu Chen 0001, Shuochao Yao |
SenSys | 5 |
| 2024 | Improved Techniques for Quantizing Deep Networks with Adaptive Bit-WidthsabstractQuantizing deep networks with adaptive bit-widths is a promising technique for efficient inference across many devices and resource constraints. In contrast to static methods that repeat the quantization process and train different models for different constraints, adaptive quantization enables us to flexibly adjust the bit-widths of a single deep network during inference for instant adaptation in different scenarios. While existing research shows encouraging results on common image classification benchmarks, this paper investigates how to train such adaptive networks more effectively. Specifically, we present two novel techniques for quantizing deep neural networks with adaptive bit-widths of weights and activations. First, we propose a collaborative strategy to choose a high-precision "teacher" for transferring knowledge to the low-precision "student" while jointly optimizing the model with all bit-widths. Second, to effectively transfer knowledge, we develop a dynamic block swapping method by randomly replacing the blocks in the lower-precision student network with the corresponding blocks in the higher-precision teacher network. Extensive experiments on multiple image and video classification datasets, well demonstrate the efficacy of our approach over state-of-the-art methods. Ximeng Sun, Rameswar Panda, Chun-Fu Chen 0001, Naigang Wang, Bowen Pan, Aude Oliva, Rogério Feris, Kate Saenko |
WACV | 3 |
| 2022 | VALHALLA: Visual Hallucination for Machine TranslationabstractDesigning better machine translation systems by considering auxiliary inputs such as images has attracted much attention in recent years. While existing methods show promising performance over the conventional text-only translation systems, they typically require paired text and image as input during inference, which limits their applicability to real-world scenarios. In this paper, we introduce a visual hallucination framework, called VALHALLA, which requires only source sentences at inference time and instead uses hallucinated visual representations for multi-modal machine translation. In particular, given a source sentence an autoregressive hallucination transformer is used to predict a discrete visual representation from the input text, and the combined text and hallucinated representations are utilized to obtain the target translation. We train the hallucination transformer jointly with the translation transformer using standard backpropagation with crossentropy losses while being guided by an additional loss that encourages consistency between predictions using either groundtruth or hallucinated visual representations. Extensive experiments on three standard translation datasets with a diverse set of language pairs demonstrate the effectiveness of our approach over both text-only baselines and state-of-the-art methods. Project page: http://www.svcl.ucsd.jects/valhalla.edu/pro. Yi Li 0051, Rameswar Panda, Chun-Fu Chen 0001, Rogério Feris, David D. Cox, Nuno Vasconcelos |
CVPR | 4 |
| 2022 | Task2Sim: Towards Effective Pre-training and Transfer from Synthetic DataabstractPre-training models on Imagenet or other massive datasets of real images has led to major advances in Computer vision, albeit accompanied with shortcomings related to curation cost, privacy, usage rights, and ethical issues. In this paper, for the first time, we study the transferability of pre-trained models based on synthetic data generated by graphics simulators to downstream tasks from very different domains. In using such synthetic data for pre-training, we find that downstream performance on different tasks are fa-vored by different configurations of simulation parameters (e.g. lighting, object pose, backgrounds, etc.), and that there is no one-size-fits-all solution. It is thus better to tailor syn-thetic pre-training data to a specific downstream task, for best performance. We introduce Task2Sim, a unified model mapping downstream task representations to optimal sim-ulation parameters to generate synthetic pre-training data for them. Task2Sim learns this mapping by training to find the set of best parameters on a set of “seen” tasks. Once trained, it can then be used to predict best simulation pa-rameters for novel “unseen” tasks in one shot, without re-quiring additional training. Given a budget in number of images per class, our extensive experiments with 20 di-verse downstream tasks show Task2Sim's task-adaptive pre-training data results in significantly better downstream per-formance than non-adaptively choosing simulation param-eters on both seen and unseen tasks. It is even competitive with pre-training on real images from Imagenet. Samarth Mishra, Rameswar Panda, Cheng Perng Phoo, Chun-Fu Chen 0001, Leonid Karlinsky, Kate Saenko, Venkatesh Saligrama, Rogério Feris |
CVPR | 4 |
| 2022 | Utility-Preserving Biometric Information Anonymization
Bill Moriarty, Chun-Fu Chen 0001, Shaohan Hu, Sean J. Moran, Marco Pistoia, Vincenzo Piuri, Pierangela Samarati |
ESORICS (2) | 2 |
| 2022 | RegionViT: Regional-to-Local Attention for Vision Transformers
Chun-Fu Chen 0001, Rameswar Panda, Quanfu Fan |
ICLR | 1 |
| 2022 | Can an Image Classifier Suffice For Action Recognition?
Quanfu Fan, Chun-Fu Chen 0001, Rameswar Panda |
ICLR | 2 |
| 2022 | Procedural Image Programs for Representation LearningabstractLearning image representations using synthetic data allows training neural networks without some of the concerns associated with real images, such as privacy and bias. Existing work focuses on a handful of curated generative processes which require expert knowledge to design, making it hard to scale up. To overcome this, we propose training with a large dataset of twenty-one thousand programs, each one generating a diverse set of synthetic images. These programs are short code snippets, which are easy to modify and fast to execute using OpenGL. The proposed dataset can be used for both supervised and unsupervised representation learning, and reduces the gap between pre-training with real and procedurally generated images by 38%. Manel Baradad Jurjo, Chun-Fu Chen 0001, Jonas Wulff, Tongzhou Wang 0001, Rogério Feris, Antonio Torralba 0001, Phillip Isola |
NeurIPS | 2 |
| 2021 | NASTransfer: Analyzing Architecture Transferability in Large Scale Neural Architecture SearchabstractNeural Architecture Search (NAS) is an open and challenging problem in machine learning. While NAS offers great promise, the prohibitive computational demand of most of the existing NAS methods makes it difficult to directly search the architectures on large-scale tasks. The typical way of conducting large scale NAS is to search for an architectural building block on a small dataset (either using a proxy set from the large dataset or a completely different small scale dataset) and then transfer the block to a larger dataset. Despite a number of recent results that show the promise of transfer from proxy datasets, a comprehensive evaluation of different NAS methods studying the impact of different source datasets has not yet been addressed. In this work, we propose to analyze the architecture transferability of different NAS methods by performing a series of experiments on large scale benchmarks such as ImageNet1K and ImageNet22K. We find that: (i) The size and domain of the proxy set does not seem to influence architecture performance on the target dataset. On average, transfer performance of architectures searched using completely different small datasets (e.g., CIFAR10) perform similarly to the architectures searched directly on proxy target datasets. However, design of proxy sets has considerable impact on rankings of different NAS methods. (ii) While different NAS methods show similar performance on a source dataset (e.g., CIFAR10), they significantly differ on the transfer performance to a large dataset (e.g., ImageNet1K). (iii) Even on large datasets, random sampling baseline is very competitive, but the choice of the appropriate combination of proxy set and search strategy can provide significant improvement over it. We believe that our extensive empirical analysis will prove useful for future design of NAS algorithms. Rameswar Panda, Michele Merler, Mayoore S. Jaiswal, Hui Wu 0009, Kandan Ramakrishnan, Ulrich Finkler, Chun-Fu Chen 0001, Minsik Cho, Rogério Feris, David S. Kung 0001, Bishwaranjan Bhattacharjee |
AAAI | 7 |
| 2021 | Deep Analysis of CNN-Based Spatio-Temporal Representations for Action RecognitionabstractIn recent years, a number of approaches based on 2D or 3D convolutional neural networks (CNN) have emerged for video action recognition, achieving state-of-the-art results on several large-scale benchmark datasets. In this paper, we carry out in-depth comparative analysis to better understand the differences between these approaches and the progress made by them. To this end, we develop an unified framework for both 2D-CNN and 3D-CNN action models, which enables us to remove bells and whistles and provides a common ground for fair comparison. We then conduct an effort towards a large-scale analysis involving over 300 action recognition models. Our comprehensive analysis reveals that a) a significant leap is made in efficiency for action recognition, but not in accuracy; b) 2D-CNN and 3D-CNN models behave similarly in terms of spatio-temporal representation abilities and transferability. Our codes are available at https://github.com/IBM/action-recognition-pytorch. Chun-Fu Chen 0001, Rameswar Panda, Kandan Ramakrishnan, Rogério Feris, John Cohn, Aude Oliva, Quanfu Fan |
CVPR | 1 |
| 2021 | Detector-Free Weakly Supervised Grounding by SeparationabstractNowadays, there is an abundance of data involving images and surrounding free-form text weakly corresponding to those images. Weakly Supervised phrase-Grounding (WSG) deals with the task of using this data to learn to localize (or to ground) arbitrary text phrases in images without any additional annotations. However, most recent SotA methods for WSG assume an existence of a pre-trained object detector, relying on it to produce the ROIs for localization. In this work, we focus on the task of Detector-Free WSG (DF-WSG) to solve WSG without relying on a pre-trained detector. The key idea behind our proposed Grounding by Separation (GbS) method is synthesizing ‘text to image-regions’ associations by random alpha-blending of arbitrary image pairs and using the corresponding texts of the pair as conditions to recover the alpha map from the blended image via a segmentation network. At test time, this allows using the query phrase as a condition for a non-blended query image, thus interpreting the test image as a composition of a region corresponding to the phrase and the complement region. Our GbS shows an 8.5% accuracy improvement over previous DF-WSG SotA, for a range of benchmarks including Flickr30K, Visual Genome, and ReferIt, as well as a complementary improvement (above 7%) over the detector-based approaches for WSG. Assaf Arbelle, Sivan Doveh, Amit Alfassy, Joseph Shtok, Guy Lev, Eli Schwartz, Hilde Kuehne, Hila Levi, Prasanna Sattigeri, Rameswar Panda, Chun-Fu Chen 0001, Alexander M. Bronstein, Kate Saenko, Shimon Ullman, Raja Giryes, Rogério Feris, Leonid Karlinsky |
ICCV | 11 |
| 2021 | CrossViT: Cross-Attention Multi-Scale Vision Transformer for Image ClassificationabstractThe recently developed vision transformer (ViT) has achieved promising results on image classification compared to convolutional neural networks. Inspired by this, in this paper, we study how to learn multi-scale feature representations in transformer models for image classification. To this end, we propose a dual-branch transformer to com-bine image patches (i.e., tokens in a transformer) of different sizes to produce stronger image features. Our approach processes small-patch and large-patch tokens with two separate branches of different computational complexity and these tokens are then fused purely by attention multiple times to complement each other. Furthermore, to reduce computation, we develop a simple yet effective token fusion module based on cross attention, which uses a single token for each branch as a query to exchange information with other branches. Our proposed cross-attention only requires linear time for both computational and memory complexity instead of quadratic time otherwise. Extensive experiments demonstrate that our approach performs better than or on par with several concurrent works on vision transformer, in addition to efficient CNN models. For example, on the ImageNet1K dataset, with some architectural changes, our approach outperforms the recent DeiT by a large margin of 2% with a small to moderate increase in FLOPs and model parameters. Our source codes and models are available at https://github.com/IBM/CrossViT. Chun-Fu Chen 0001, Quanfu Fan, Rameswar Panda |
ICCV | 1 |
| 2021 | A Broad Study on the Transferability of Visual Representations with Contrastive LearningabstractTremendous progress has been made in visual representation learning, notably with the recent success of self-supervised contrastive learning methods. Supervised contrastive learning has also been shown to outperform its cross-entropy counterparts by leveraging labels for choosing where to contrast. However, there has been little work to explore the transfer capability of contrastive learning to a different domain. In this paper, we conduct a comprehensive study on the transferability of learned representations of different contrastive approaches for linear evaluation, full-network transfer, and few-shot recognition on 12 downstream datasets from different domains, and object detection tasks on MSCOCO and VOC0712. The results show that the contrastive approaches learn representations that are easily transferable to a different downstream task. We further observe that the joint objective of self-supervised contrastive loss with cross-entropy/supervised-contrastive loss leads to better transferability of these models over their supervised counterparts. Our analysis reveals that the representations learned from the contrastive approaches contain more low/mid-level semantics than cross-entropy models, which enables them to quickly adapt to a new task. Our codes and models will be publicly available to facilitate future research on transferability of visual representations.1 Ashraful Islam, Chun-Fu Chen 0001, Rameswar Panda, Leonid Karlinsky, Richard J. Radke, Rogério Feris |
ICCV | 2 |
| 2021 | AdaMML: Adaptive Multi-Modal Learning for Efficient Video RecognitionabstractMulti-modal learning, which focuses on utilizing various modalities to improve the performance of a model, is widely used in video recognition. While traditional multi-modal learning offers excellent recognition results, its computational expense limits its impact for many real-world applications. In this paper, we propose an adaptive multi-modal learning framework, called AdaMML, that selects on-the-fly the optimal modalities for each segment conditioned on the input for efficient video recognition. Specifically, given a video segment, a multi-modal policy net-work is used to decide what modalities should be used for processing by the recognition model, with the goal of improving both accuracy and efficiency. We efficiently train the policy network jointly with the recognition model using standard back-propagation. Extensive experiments on four challenging diverse datasets demonstrate that our proposed adaptive approach yields 35% − 55% reduction in computation when compared to the traditional baseline that simply uses all the modalities irrespective of the in-put, while also achieving consistent improvements in accuracy over the state-of-the-art methods. Project page: https://rpand002.github.io/adamml.html. Rameswar Panda, Chun-Fu Chen 0001, Quanfu Fan, Ximeng Sun, Kate Saenko, Aude Oliva, Rogério Feris |
ICCV | 2 |
| 2021 | Dynamic Network Quantization for Efficient Video InferenceabstractDeep convolutional networks have recently achieved great success in video recognition, yet their practical realization remains a challenge due to the large amount of computational resources required to achieve robust recognition. Motivated by the effectiveness of quantization for boosting efficiency, in this paper, we propose a dynamic network quantization framework, that selects optimal precision for each frame conditioned on the input for efficient video recognition. Specifically, given a video clip, we train a very lightweight network in parallel with the recognition network, to produce a dynamic policy indicating which numerical precision to be used per frame in recognizing videos. We train both networks effectively using standard backpropagation with a loss to achieve both competitive performance and resource efficiency required for video recognition. Extensive experiments on four challenging diverse benchmark datasets demonstrate that our proposed approach provides significant savings in computation and memory usage while outperforming the existing state-of-the-art methods. Project page: https://cs-people.bu.edu/sunxm/VideoIQ/project.html. Ximeng Sun, Rameswar Panda, Chun-Fu Chen 0001, Aude Oliva, Rogério Feris, Kate Saenko |
ICCV | 3 |
| 2021 | Dynamic Distillation Network for Cross-Domain Few-Shot Recognition with Unlabeled DataabstractMost existing works in few-shot learning rely on meta-learning the network on a large base dataset which is typically from the same domain as the target dataset. We tackle the problem of cross-domain few-shot learning where there is a large shift between the base and target domain. The problem of cross-domain few-shot recognition with unlabeled target data is largely unaddressed in the literature. STARTUP was the first method that tackles this problem using self-training. However, it uses a fixed teacher pretrained on a labeled base dataset to create soft labels for the unlabeled target samples. As the base dataset and unlabeled dataset are from different domains, projecting the target images in the class-domain of the base dataset with a fixed pretrained model might be sub-optimal. We propose a simple dynamic distillation-based approach to facilitate unlabeled images from the novel/base dataset. We impose consistency regularization by calculating predictions from the weakly-augmented versions of the unlabeled images from a teacher network and matching it with the strongly augmented versions of the same images from a student network. The parameters of the teacher network are updated as exponential moving average of the parameters of the student network. We show that the proposed network learns representation that can be easily adapted to the target domain even though it has not been trained with target-specific classes during the pretraining phase. Our model outperforms the current state-of-the art method by 4.4% for 1-shot and 3.6% for 5-shot classification in the BSCD-FSL benchmark, and also shows competitive performance on traditional in-domain few-shot learning task. Ashraful Islam, Chun-Fu Chen 0001, Rameswar Panda, Leonid Karlinsky, Rogério Feris, Richard J. Radke |
NeurIPS | 2 |
| 2020 | Parallelization of Classical Numerical optimization in Quantum Variational AlgorithmsabstractNumerical optimization has been extensively used in many real-world applications related to Scientific Computing, Artificial Intelligence and, more recently, Quantum Computing. However, existing optimizers conduct their internal computations sequentially, which affects their performance. We observed a general pattern that enabled us to parallelize such internal computations and achieve significant speedup. We designed a novel parallelization algorithm for optimizers, which consists of pattern detection, prediction, precomputation, and caching. Importantly, our design does not require any change to the optimizers. Instead, it simply modifies the function to be optimized, thereby leading to several engineering advantages, including simplicity, modularity and portability. We implemented this solution and included it in the Qiskit Aqua open-source project. In this paper, we present an evaluation on both standard benchmarks and real-world quantum-computing applications. The evaluation results confirm that our approach (1) incurs negligible overhead, (2) effectively speeds up optimization, and (3) does not affect the accuracy of the results or the convergence of the optimizers. Marco Pistoia, Peng Liu 0010, Chun-Fu Chen 0001, Shaohan Hu, Stephen P. Wood |
ICST | 3 |
| 2019 | Big-Little Net: An Efficient Multi-Scale Feature Representation for Visual and Speech Recognition
Chun-Fu Chen 0001, Quanfu Fan, Neil Mallinar, Tom Sercu, Rogério Feris |
ICLR (Poster) | 1 |
| 2019 | More Is Less: Learning Efficient Video Representations by Big-Little Network and Depthwise Temporal AggregationabstractCurrent state-of-the-art models for video action recognition are mostly based on expensive 3D ConvNets. This results in a need for large GPU clusters to train and evaluate such architectures. To address this problem, we present an lightweight and memory-friendly architecture for action recognition that performs on par with or better than current architectures by using only a fraction of resources. The proposed architecture is based on a combination of a deep subnet operating on low-resolution frames with a compact subnet operating on high-resolution frames, allowing for high efficiency and accuracy at the same time. We demonstrate that our approach achieves a reduction by 3~4 times in FLOPs and ~2 times in memory usage compared to the baseline. This enables training deeper models with more input frames under the same computational budget. To further obviate the need for large-scale 3D convolutions, a temporal aggregation module is proposed to model temporal dependencies in a video at very small additional computational costs. Our models achieve strong performance on several action recognition benchmarks including Kinetics, Something-Something and Moments-in-time. The code and models are available at \url{https://github.com/IBM/bLVNet-TAM}. Quanfu Fan, Chun-Fu Chen 0001, Hilde Kuehne, Marco Pistoia, David D. Cox |
NeurIPS | 2 |
| 2018 | NISP: Pruning Networks Using Neuron Importance Score PropagationabstractTo reduce the significant redundancy in deep Convolutional Neural Networks (CNNs), most existing methods prune neurons by only considering the statistics of an individual layer or two consecutive layers (e.g., prune one layer to minimize the reconstruction error of the next layer), ignoring the effect of error propagation in deep networks. In contrast, we argue that for a pruned network to retain its predictive power, it is essential to prune neurons in the entire neuron network jointly based on a unified goal: minimizing the reconstruction error of important responses in the "final response layer" (FRL), which is the second-to-last layer before classification. Specifically, we apply feature ranking techniques to measure the importance of each neuron in the FRL, formulate network pruning as a binary integer optimization problem, and derive a closed-form solution to it for pruning neurons in earlier layers. Based on our theoretical analysis, we propose the Neuron Importance Score Propagation (NISP) algorithm to propagate the importance scores of final responses to every neuron in the network. The CNN is pruned by removing neurons with least importance, and it is then fine-tuned to recover its predictive power. NISP is evaluated on several datasets with multiple CNN models and demonstrated to achieve significant acceleration and compression with negligible accuracy loss. Ruichi Yu, Ang Li 0001, Chun-Fu Chen 0001, Jui-Hsin Lai, Vlad I. Morariu, Xintong Han, Mingfei Gao, Ching-Yung Lin, Larry Davis 0001 |
CVPR | 3 |
| 2018 | SC-Conv: Sparse-Complementary Convolution for Efficient Model Utilization on CNNsabstractWe propose sparse-complementary convolution (SC-Conv) to improve model utilization of convolution neural networks (CNNs). The networks with SC-Conv achieve better accuracy than the regular convolution under similar computations and parameters. The proposed SC-Conv is paired with two deterministic sparse kernels, and one of kernels is complementary to the other one at in either spatial or channel domain or both; the deterministic sparsity increases the computational speed theoretically and practically; furthermore, by having the complementary characteristic, SC-Conv retains the same receptive field to the conventional convolution. This insightful but straightforward SC-Conv reuses of modern network architectures (ResNet and DenseNet), and at the same FLOPs and parameters, SC-Conv improves top-1 classification accuracy on ImageNet by 0.6 points for ResNet-101 and keep the same model complexity. Furthermore, SC-Conv also outperforms recent sparse networks by 1.3 points at top-1 accuracy for ImageNet, and after integrating SC-Conv with the sparse network, we further improve another 1.8 points accuracy at similar FLOPs and parameters. Chun-Fu Chen 0001, Jinwook Oh, Quanfu Fan, Marco Pistoia |
ISM | 1 |
| 2017 | UI X-Ray: Interactive Mobile UI Testing Based on Computer VisionabstractUser Interface/eXperience (UI/UX) significantly affects the lifetime of any software program, particularly mobile apps. A bad UX can undermine the success of a mobile app even if that app enables sophisticated capabilities. A good UX, however, needs to be supported of a highly functional and user friendly UI design. In spite of the importance of building mobile apps based on solid UI designs, UI discrepancies---inconsistencies between UI design and implementation---are among the most numerous and expensive defects encountered during testing. This paper presents UI X-Ray, an interactive UI testing system that integrates computer-vision methods to facilitate the correction of UI discrepancies---such as inconsistent positions, sizes and colors of objects and fonts. Using UI X-Ray does not require any programming experience; therefore, UI X-Ray can be used even by non-programmers---particularly designers---which significantly reduces the overhead involved in writing tests. With the feature of interactive interface, UI testers can quickly generate defect reports and revision instructions---which would otherwise be done manually. We verified our UI X-Ray on 4 developed mobile apps of which the entire development history was saved. UI X-Ray achieved a 99.03% true-positive rate, which significantly surpassed the 20.92% true-positive rate obtained via manual analysis. Furthermore, evaluating the results of our automated analysis can be completed quickly (< 1 minute per view on average) compared to hours of manual work required by UI testers. On the other hand, UI X-Ray received the appreciations from skilled designers and UI X-Ray improves their current work flow to generate UI defect reports and revision instructions. The proposed system, UI X-Ray, presented in this paper has recently become part of a commercial product. Chun-Fu Chen 0001, Marco Pistoia, Conglei Shi, Paolo Girolami, Joe W. Ligman, Yong Wang 0021 |
IUI | 1 |
| 2016 | Efficient nuclei segmentation based on spectral graph partitioningabstractBiomedical image processing that offers computer-aided diagnosis is much more popular due to the availability of high quality and large quantity of medical data. Our well-developed biomedical image computing system, which automatically extracts and segments the nucleus and cytoplasm of cell in medical images, is no doubt following this idea. Nonetheless, even though previous system provide good algorithmic performance, its throughput is limited by high computation load and data dependency. Therefore, we deploy spectral graph partitioning to improve computation speed of the most complex module, maker-controlled watershed transform for nuclei detection. By modeling our problem as a graph and embedding architectural costs as the attributes in vertices and edges, we equally distribute workload among processors and reduce overhead in data transfer rate. We deploy the proposed approach on Intel Core i7-930 CPU with four cores and eight threads and test 153 medical images; as a consequence, we achieve less data transfer and better load balance as compared to conventional workload distribution through clustering and other graph partitioning methods. Gwo Giun Lee, Shi-Yu Hung, Tai-Ping Wang, Chun-Fu Chen 0001, Chi-Kuang Sun, Yi-Hua Liao |
ISCAS | 4 |
| 2015 | Implementation of Gabor feature extraction algorithm for electrocardiogram on FPGAabstractThis paper implements an electrocardiogram (ECG) feature extraction system onto a Field Programmable Gate Array (FPGA). The algorithm extracts ECG features including R peaks, QRS onsets, Q peaks, S peaks, QRS offsets, P onsets, P peaks, P offsets, T onsets, T peaks, and T offsets based on multi-scale analysis in Gabor Wavelet Transform (GWT); and estimates the amplitudes of P, R, T peaks and Q, S depths. Subjects' ECG signals are acquired through the Texas Instrument (TI) ADS1298R-ECGPDK ECG signal acquisition, and then transferred to Xilinx XUPV5-LX110T Evaluation Platform for ECG feature extraction via Serial Peripheral Interface (SPI). Consequently, the results that are obtained by the feature extraction process will be stored in the CompactFlash Card of the evaluation platform and displayed on the monitor via the National Instrument (NI) data acquisition. The experimental results show that the proposed system that achieves low comparative error surpasses the performance in other related works. Gwo Giun Lee, Zuo-Jheng Huang, Chih-Yuan Chen, Chun-Fu Chen 0001 |
ISCAS | 4 |
| 2015 | Efficient Multi-training Framework of Image Deep Learning on GPU ClusterabstractIn this paper, we develop a pipelining schema for image deep learning on GPU cluster to leverage heavy workload of training procedure. In addition, it is usually necessary to train multiple models to obtain a good deep learning model due to the limited a priori knowledge on deep neural network structure. Therefore, adopting parallel and distributed computing appears is an obvious path forward, but the mileage varies depending on how amenable a deep network can be parallelized and the availability of rapid prototyping capabilities with low cost of entry. In this work, we propose a framework to organize the training procedures of multiple deep learning models into a pipeline on a GPU cluster, where each stage is handled by a particular GPU with a partition of the training dataset. Instead of frequently migrating data among the disks, CPUs, and GPUs, our framework only moves partially trained models to reduce bandwidth consumption and to leverage the full computation capability of the cluster. In this paper, we deploy the proposed framework on popular image recognition tasks using deep learning, and the experiments show that the proposed method reduces overall training time up to dozens of hours compared to the baseline method. Chun-Fu Chen 0001, Gwo Giun Lee, Yinglong Xia, Wan-Yi Sabrina Lin, Toyotaro Suzumura, Ching-Yung Lin |
ISM | 1 |
| 2015 | Multi-modality Mobile Image Recognition Based on Thermal and Visual CamerasabstractThe advances of mobile computing and sensor technology have turned the mobile devices into powerful instruments. The integration of thermal and visual cameras extends the capability of computer vision, due to the fact that both images reveal different characteristics in images, however, image alignment is a challenge. This paper proposes an effective approach to align image pairs for event detection on mobile through image recognition. We leverage thermal and visual cameras as multi-modality sources for image recognition. By analyzing the heat pattern, the proposed APP can identify the heating sources and help users inspect their house heating system, on the other hand, with applying image recognition, the proposed APP furthermore can help field workers identify the asset condition and provide the guidance to solve their issues. Jui-Hsin Larry Lai, Chung-Ching Lin, Chun-Fu Chen 0001, Ching-Yung Lin |
ISM | 3 |
| 2014 | Content-adaptive depth map enhancement based on motion distributionabstractThis paper provides a motion-based content-adaptive depth map enhancement algorithm to enhance the quality of the depth map and reduce the artifacts in the synthesized views. The proposed algorithm extracts depth cues from the motion distribution at the specific scenario of camera movement to align the distribution of depth and motion. In real world scenarios, when the camera is panning in horizontal direction, the nearer distance between the object and the camera, the larger motion will be, and vice versa; therefore, we could interpret the depth from motion in this. Moreover, in the scenario of fixed camera, the depth cue from motion could be derived in the same approach, and the depth variation within one moving object shall be small. Hence, the depth values of moving object should not be rapidly changing. In addition, this paper also employs the bi-directional motion-compensated infinite impulse response low-pass filter to stabilize the consistency of depth maps over time. As a consequence, the algorithm so introduced not only aligns the depth map to depth cues from motion but also enhance stability and consistency of depth maps in the spatial-temporal domain. Experiment results via enhanced depth maps show that the synthesized results would be better in both objective and subjective measurement in comparison with the results using original depth maps and the state-of-the-art depth enhancement algorithms. Gwo Giun Lee, Bo-Syun Li, Chun-Fu Chen 0001 |
VCIP | 3 |
| 2013 | Depth map enhancement based on Z-displacement of objectsabstractA depth map enhancement algorithm based on Z-displacement of objects is presented in this paper. Most depth map enhancement algorithms utilize either the spatial or temporal information only. In this paper, information in another dimension, Z-displacement, is utilized to enhance the depth map of videos by estimating the scale change of objects. Our experimental results show that the proposed scale changing detection method has higher accuracy and lower complexity compared with the conventional scale changing detection approaches such as normalized gradient correlation. Based on the estimated scale change, the proposed depth map enhancement corrects the depth map by estimating the Z-displacement of objects in the 3D space. Experimental results demonstrate that better depth maps and visual quality of view synthesis are achieved by the proposed method as compared to the state-of-arts. Gwo Giun Lee, Ciao-Siang Siao, Chunhui Cui, Chun-Fu Chen 0001, Huan-Hsiang Lin |
ISCAS | 4 |
| 2013 | Motion-based depth estimation for 2D-to-3D video conversionabstractThis paper presents a motion-based depth estimation algorithm for automatic 2D-to-3D video conversion algorithm by employing the co-occurrence matrix of motion vectors (MVCM). Video scenes possess distinct signatures of MVCM, which enables exploiting the corresponding motion-depth relation for depth generation. The subsequent motion-compensated depth updating scheme provides stable and comfort 3D visual quality as synthesized by depth-image-based rendering. The simulation results of several high-definition image sequences indicate that the proposed algorithm produces better and more reasonable depth than two motion-based depth estimation algorithms. With the adaptive depth estimation scheme using MVCM, the proposed 2D-to-3D video conversion algorithm can accommodate a great variety of visual contents. It thus provides an efficient and reliable solution towards the problem of automatic 3D video content creation. Ming-Jiun Wang, Chun-Fu Chen 0001, Gwo Giun Lee |
VCIP | 2 |
| 2012 | Cell segmentation and NC ratio analysis of third harmonic generation virtual biopsy images based on marker-controlled gradient watershed algorithmabstractTraditional biopsy procedure requires invasive tissue removal from a living subject followed by time-consuming complicatedly processing, so noninvasive in vivo virtual biopsy is a highly desired technique which own ability to obtain exhaustive tissue images without removing tissues from subjects. Some sets of in vivo virtual biopsy images provided by some healthy volunteers are processed by our cell segmentation approach based on marker-controlled gradient watershed algorithm to isolate the nuclei and cytoplasm and also evaluate their Nuclear-to-Cytoplasmic (NC) ratio. From our experimental results, our algorithm has significant potential for in vivo cell segmentation and NC ratio analysis to identify or detect the early symptoms of some skin diseases with abnormal NC ratios, such as skin cancers in clinical diagnosis. Huan-Hsiang Lin, Ming-Rung Tsai, Chun-Fu Chen 0001, Szu-Yu Chen, Yi-Hua Liao, Gwo Giun Lee, Chi-Kuang Sun |
ISCAS | 3 |
| 2012 | Quantifying Intrinsic Parallelism Using Linear Algebra for Algorithm/Architecture CoexplorationabstractDegree of parallelism (DoP) is an essential complexity metric that characterizes the number of independent operation sets (IOSs) that can be concurrently executed within an algorithm. This paper presents a generic framework to identify IOSs and to quantify the DoP based on rank theorem in linear algebra. This framework is applied to extract algorithmic parallelisms at various granularities, namely, multigrain parallelism. Our parallelism is intrinsic and platform independent and can provide insights into architectural information, thus facilitating mapping onto generic platforms and early back annotation for modifying algorithms. It plays a significant role in the concurrent optimization of both algorithms and architectures, referred to as Algorithm/Architecture Coexploration (AAC), by trading off between the DoP and the number of operations (NoO). This paper reports three case studies for AAC. The case study on an IDCT reveals that our framework accurately quantifies the parallelism for mapping the algorithm onto generic platforms, including FPGA and multicore systems. The IDCT parallelized by our technique surpasses a conventional spectral parallelization. By exploiting fine-grain parallelism, this paper presents a better porting of a discrete wavelet transform (DWT) onto single instruction multiple data (SIMD) machines compared with a commercial compiler. A high-quality deinterlacer is implemented on a low-cost multicore platform for real-time high-definition applications by analyzing the multigrain parallelism. These case studies reveal the effectiveness of our parallel analysis framework which is applicable to generic systems. Compared with traditional graph traversal techniques, our linear algebraic approach impressively features low complexity and is practical for complicated algorithms. Gwo Giun Lee, He-Yuan Lin, Chun-Fu Chen 0001, Tsung-Yuan Huang |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2011 | Reconfigurable inverse transform architecture for multiple purpose video codingabstractIn this paper, an area efficient reconfigurable inverse transformation architecture for multiple standards is proposed. We present a top-down design methodology with complexity analysis, commonalities extraction, and dataflow modeling to systematically design reconfigurable architecture. By exporting and sharing the commonalities, the adder usage of the proposed reconfigurable inverse transform processing element can be reduced 44% compared with the total amount of adders in performing target inverse transform types. Then, the reconfigurable architecture is synthesized using TSMC 0.18 urn library. The working frequency is 108Mhz, which is derived from the dataflow scheduling. The area synthesis result is 32k gates, which indicates that the proposed design has more efficient area than other documented design in VLSI implementation. In addition, the proposed architecture also satisfies the accuracy requirement. Therefore, the proposed design have lower cost and enough flexibility for multi-standard purposes with 1920×1088 resolution and 64 frames per second and the color format is 4:2:0 for real time processing. Tsung-Yuan Huang, He-Yuan Lin, Chun-Fu Chen 0001, Gwo Giun Lee |
ISCAS | 3 |