EDBT 2026 Demo / reviewers in the wild / expert
Cheng-Yu Hsieh
dblp:40/4421
· DBLP profile ↗
26ranked-venue papers
7as first author
17since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 20 · 4 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 6 since 2021Computer networks · 3 · 2 first-authorDatabases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multiattribute Decision Making System Based on Interval-Valued Intuitionistic Fuzzy Weights of AttributesabstractIn this paper, we propose a new multiattribute decision making (MADM) approach in interval-valued intuitionistic fuzzy (IVIF) environments, where the attributes’ weights are represented by interval-valued intuitionistic fuzzy values (IVIFVs). First, we use the decision matrix (DM) offered by the decision maker (DK) and our proposed score function (SF) to construct a score matrix (SM). Then, a new nonlinear programming (NLP) model is developed using the hesitancy degree of each IVIFV shown in the DM, Mishra et al.’s Hellinger distance measure between interval-valued intuitionistic fuzzy sets (IVIFSs), and the IVIF weights of the attributes offered by the DK. Then, our proposed NLP model is solved to obtain the optimal weights (OPWs) of the attributes. We utilize the obtained OPWs of attributes and the obtained SM to compute the weighted scores of alternatives, where the alternatives are ranked using the weighted scores. The existing MADM approaches have drawbacks that they get unreasonable preference order (POs) of alternatives or cannot obtain the POs of alternatives in some situations. The experimental results show that proposed MADM approach gets more reasonable POs of alternatives to overcome the drawbacks of the existing MADM approaches in IVIF environments. Shyi-Ming Chen, Cheng-Yu Hsieh, Xin-Yao Zou |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2025 | Perception Tokens Enhance Visual Reasoning in Multimodal Language ModelsabstractMultimodal language models (MLMs) still face challenges in fundamental visual perception tasks where specialized models excel. Tasks requiring reasoning about 3D structures benefit from depth estimation, and reasoning about 2D object instances benefits from object detection. Yet, MLMs can not produce intermediate depth or boxes to reason over. Fine-tuning MLMs on relevant data doesn’t generalize well and outsourcing computation to specialized vision tools is too compute-intensive and memory-inefficient. To address this, we introduce Perception Tokens, intrinsic image representations designed to assist reasoning tasks where language is insufficient. Perception tokens act as auxiliary reasoning tokens, akin to chain-of-thought prompts in language models. For example, in a depth-related task, an MLM augmented with perception tokens can reason by generating a depth map as tokens, enabling it to solve the problem effectively. We propose Aurora, a training method that augments MLMs with perception tokens for improved reasoning over visual inputs. AURORA leverages a VQVAE to transform intermediate image representations, such as depth maps into a tokenized format and bounding box tokens, which are then used in a multi-task training framework. AURORA achieves notable improvements across counting benchmarks: +10.8% on BLINK, +11.3% on CVBench, and +8.3% on SEED-Bench, outperforming fine-tuning approaches in generalization across datasets. It also improves on relative depth: over +6% on BLINK. With perception tokens, Aurora expands the scope of MLMs beyond language-based reasoning, paving the way for more effective visual reasoning capabilities. Code and data will be released at the project page. Mahtab Bigverdi, Zelun Luo, Cheng-Yu Hsieh, Ethan Shen, Dongping Chen, Linda G. Shapiro, Ranjay Krishna |
CVPR | 3 |
| 2025 | NVILA: Efficient Frontier Visual Language ModelsabstractVisual language models (VLMs) have made significant advances in accuracy in recent years. However, their efficiency has received much less attention. This paper introduces NVILA, a family of open VLMs designed to optimize both efficiency and accuracy. Building on top of VILA, we improve its model architecture by first scaling up the spatial and temporal resolutions, and then compressing visual tokens. This "scale-then-compress" approach enables NVILA to efficiently process high-resolution images and long videos. We also conduct a systematic investigation to enhance the efficiency of NVILA throughout its entire lifecycle, from training to deployment. NVILA matches or surpasses the accuracy of many leading open and proprietary VLMs across a wide range of image and video benchmarks. At the same time, it reduces training costs by 1.9-5.1×, prefilling latency by 1.6-2.2×, and decoding latency by 1.2-2.8×. Ligeng Zhu, Baifeng Shi, Zhuoyang Zhang, Yuming Lou, Shang Yang, Haocheng Xi, Shiyi Cao, Yuxian Gu, Dacheng Li, Xiuyu Li, Haotian Tang, Yunhao Fang, Yukang Chen, Cheng-Yu Hsieh, De-An Huang, An-Chieh Cheng, Jinyi Hu, Sifei Liu, Ranjay Krishna, Pavlo Molchanov 0001, Jan Kautz, Hongxu Yin, Song Han 0003, Yao Lu 0006 |
CVPR | 15 |
| 2025 | Synthetic Visual GenomeabstractReasoning over visual relationships—spatial, functional, interactional, social, etc.—is considered to be a fundamental component of human cognition. Yet, despite the major advances in visual comprehension in multimodal language models (MLMs), precise reasoning over relationships and their generations remains a challenge. We introduce Robin: an MLM instruction-tuned with densely annotated relationships capable of constructing high-quality dense scene graphs at scale. To train Robin, we curate SVG1, a synthetic scene graph dataset by completing the missing relations of selected objects in existing scene graphs using a teacher MLM and a carefully designed filtering process to ensure high-quality. To generate more accurate and rich scene graphs at scale for any image, we introduce SG-Edit: a self-distillation framework where GPT-4o further refines Robin’s predicted scene graphs by removing unlikely relations and/or suggesting relevant ones. In total, our dataset contains 146K images and 5.6M relationships for 2.6M objects. Results show that our Robin-3B model, despite being trained on less than 3 million instances, outperforms similar-size models trained on over 300 million instances on relationship understanding benchmarks, and even surpasses larger models up to 13B parameters. Notably, it achieves state-of-the-art performance in referring expression comprehension with a score of 88.2, surpassing the previous best of 87.4. Our results suggest that training on the refined scene graph data is crucial to maintaining high performance across diverse visual reasoning tasks2. Zixian Ma, Chenhao Zheng, Cheng-Yu Hsieh, Ximing Lu, Khyathi Raghavi Chandu, Quan Kong, Norimasa Kobori, Ali Farhadi, Yejin Choi 0001, Ranjay Krishna |
CVPR | 5 |
| 2025 | RealEdit: Reddit Edits As a Large-scale Empirical Dataset for Image TransformationsabstractExisting image editing models struggle to meet real-world demands; despite excelling in academic benchmarks, we are yet to see them adopted to solve real user needs. The datasets that power these models use artificial edits, lacking the scale and ecological validity necessary to address the true diversity of user requests. In response, we introduce RealEdit, a large-scale image editing dataset with authentic user requests and human-made edits sourced from Reddit. RealEdit contains a test set of 9.3K examples the community can use to evaluate models on real user requests. Our results show that existing models fall short on these tasks, implying a need for realistic training data. So, we introduce 48K training examples, with which we train our RealEdit model. Our model achieves substantial gains—outperforming competitors by up to 165 Elo points in human judgment and 92% relative improvement on the automated VIEScore metric on our test set. We deploy our model back on Reddit, testing it on new requests, and receive positive feedback. Beyond image editing, we explore RealEdit ’s potential in detecting edited images by partnering with a deepfake detection non-profit. Finetuning their model on RealEdit data improves its F1-score by 14 percentage points, underscoring the dataset’s value for broad, impactful applications.1 Peter V. Sushko, Ayana Bharadwaj, Zhi Yang Lim, Vasily Ilin, Ben Caffee, Dongping Chen, Mohammadreza Salehi, Cheng-Yu Hsieh, Ranjay Krishna |
CVPR | 8 |
| 2025 | A Novel Score Function for Ranking Interval-Valued Intuitionistic Fuzzy Values
Shyi-Ming Chen, Cheng-Yu Hsieh, Xin-Yao Zou |
IEA/AIE (2) | 2 |
| 2024 | The Hard Positive Truth About Vision-Language Compositionality
Amita Kamath, Cheng-Yu Hsieh, Kai-Wei Chang 0001, Ranjay Krishna |
ECCV (14) | 2 |
| 2024 | Is C4 Dataset Optimal for Pruning? An Investigation of Calibration Data for LLM PruningabstractAbhinav Bandari, Lu Yin, Cheng-Yu Hsieh, Ajay Kumar Jaiswal, Tianlong Chen, Li Shen, Ranjay Krishna, Shiwei Liu. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Abhinav Bandari, Lu Yin 0006, Cheng-Yu Hsieh, Ajay Jaiswal, Tianlong Chen 0001, Li Shen 0008, Ranjay Krishna, Shiwei Liu 0003 |
EMNLP | 3 |
| 2024 | Lookback Lens: Detecting and Mitigating Contextual Hallucinations in Large Language Models Using Only Attention MapsabstractWhen asked to summarize articles or answer questions given a passage, large language models (LLMs) can hallucinate details and respond with unsubstantiated answers that are inaccurate with respect to the input context.This paper describes a simple approach for detecting such contextual hallucinations.We hypothesize that contextual hallucinations are related to the extent to which an LLM attends to information in the provided context versus its own generations.Based on this intuition, we propose a simple hallucination detection model whose input features are given by the ratio of attention weights on the context versus newly generated tokens (for each attention head).We find that a linear classifier based on these lookback ratio features is as effective as a richer detector that utilizes the entire hidden states of an LLM or a text-based entailment model.The lookback ratio-based detector-Lookback Lens-is found to transfer across tasks and even models, allowing a detector that is trained on a 7B model to be applied (without retraining) to a larger 13B model.We further apply this detector to mitigate contextual hallucinations, and find that a simple classifier-guided decoding approach is able to reduce the amount of hallucination, for example by 9.6% in the XSum summarization task. 1 Yung-Sung Chuang, Linlu Qiu, Cheng-Yu Hsieh, Ranjay Krishna, James R. Glass |
EMNLP | 3 |
| 2024 | Outlier Weighed Layerwise Sparsity (OWL): A Missing Secret Sauce for Pruning LLMs to High SparsityabstractLarge Language Models (LLMs), renowned for their remarkable performance across diverse domains, present a challenge due to their colossal model size when it comes to practical deployment. In response to this challenge, efforts have been directed toward the application of traditional network pruning techniques to LLMs, uncovering a massive number of parameters can be pruned in one-shot without hurting performance. Building upon insights gained from pre-LLM models, particularly BERT-level language models, prevailing LLM pruning strategies have consistently adhered to the practice of uniformly pruning all layers at equivalent sparsity levels, resulting in robust performance. However, this observation stands in contrast to the prevailing trends observed in the field of vision models, where non-uniform layerwise sparsity typically yields substantially improved results. To elucidate the underlying reasons for this disparity, we conduct a comprehensive analysis of the distribution of token features within LLMs. In doing so, we discover a strong correlation with the emergence of outliers, defined as features exhibiting significantly greater magnitudes compared to their counterparts in feature dimensions. Inspired by this finding, we introduce a novel LLM pruning methodology that incorporates a tailored set of **non-uniform layerwise sparsity ratios** specifically designed for LLM pruning, termed as **O**utlier **W**eighed **L**ayerwise sparsity (**OWL**). The sparsity ratio of OWL is directly proportional to the outlier ratio observed within each layer, facilitating a more effective alignment between layerwise weight sparsity and outlier ratios. Our empirical evaluation, conducted across the LLaMA-V1/V2, Vicuna, OPT, and Mistral, spanning various benchmarks, demonstrates the distinct advantages offered by OWL over previous methods. For instance, OWL exhibits a remarkable performance gain, surpassing the state-of-the-art Wanda and SparseGPT by **61.22** and **6.80** perplexity at a high sparsity level of 70%, respectively, while delivering **2.6$\times$** end-to-end inference speed-up in the DeepSparse inference engine. Code is available at https://github.com/luuyin/OWL.git. Lu Yin 0006, You Wu 0001, Zhenyu Zhang 0015, Cheng-Yu Hsieh, Yaqing Wang 0007, Yiling Jia, Gen Li 0012, Ajay Jaiswal, Mykola Pechenizkiy, Michael Bendersky, Zhangyang Wang, Shiwei Liu 0003 |
ICML | 4 |
| 2024 | The Unmet Promise of Synthetic Training Images: Using Retrieved Real Images Performs BetterabstractGenerative text-to-image models enable us to synthesize unlimited amounts of images in a controllable manner, spurring many recent efforts to train vision models with synthetic data. However, every synthetic image ultimately originates from the upstream data used to train the generator. Does the intermediate generator provide additional information over directly training on relevant parts of the upstream data?
Grounding this question in the setting of image classification, we compare finetuning on task-relevant, targeted synthetic data generated by Stable Diffusion---a generative model trained on the LAION-2B dataset---against finetuning on targeted real images retrieved directly from LAION-2B. We show that while synthetic data can benefit some downstream tasks, it is universally matched or outperformed by real data from the simple retrieval baseline. Our analysis suggests that this underperformance is partially due to generator artifacts and inaccurate task-relevant visual details in the synthetic images. Overall, we argue that targeted retrieval is a critical baseline to consider when training with synthetic data---a baseline that current methods do not yet surpass. We release code, data, and models at [https://github.com/scottgeng00/unmet-promise/](https://github.com/scottgeng00/unmet-promise). Scott Geng, Cheng-Yu Hsieh, Vivek Ramanujan, Matthew Wallingford, Chun-Liang Li, Pang Wei Koh, Ranjay Krishna |
NeurIPS | 2 |
| 2024 | DataComp-LM: In search of the next generation of training sets for language modelsabstractWe introduce DataComp for Language Models, a testbed for controlled dataset experiments with the goal of improving language models.As part of DCLM, we provide a standardized corpus of 240T tokens extracted from Common Crawl, effective pretraining recipes based on the OpenLM framework, and a broad suite of 53 downstream evaluations.Participants in the DCLM benchmark can experiment with data curation strategies such as deduplication, filtering, and data mixing atmodel scales ranging from 412M to 7B parameters.As a baseline for DCLM, we conduct extensive experiments and find that model-based filtering is key to assembling a high-quality training set.The resulting dataset, DCLM-Baseline, enables training a 7B parameter language model from scratch to 63% 5-shot accuracy on MMLU with 2T training tokens.Compared to MAP-Neo, the previous state-of-the-art in open-data language models, DCLM-Baseline represents a 6 percentage point improvement on MMLU while being trained with half the compute.Our results highlight the importance of dataset design for training language models and offer a starting point for further research on data curation. We release the \dclm benchmark, framework, models, and datasets at https://www.datacomp.ai/dclm/ Jeffrey Li, Alex Fang, Georgios Smyrnis, Maor Ivgi, Matt Jordan, Samir Yitzhak Gadre, Hritik Bansal, Etash Kumar Guha, Sedrick Keh, Kushal Arora, Niklas Muennighoff, Reinhard Heckel, Jean Mercat, Mayee F. Chen, Suchin Gururangan, Mitchell Wortsman, Alon Albalak, Yonatan Bitton, Marianna Nezhurina, Amro Abbas, Cheng-Yu Hsieh, Dhruba Ghosh, Josh Gardner 0001, Maciej Kilian, Hanlin Zhang 0002, Rulin Shao, Sarah M. Pratt, Sunny Sanyal, Gabriel Ilharco, Giannis Daras, Kalyani Marathe, Aaron Gokaslan, Jieyu Zhang 0001, Khyathi Raghavi Chandu, Igor Vasiljevic, Sham M. Kakade, Shuran Song, Sujay Sanghavi, Fartash Faghri, Sewoong Oh, Luke Zettlemoyer, Kyle Lo, Alaaeldin El-Nouby, Hadi Pouransari, Alexander Toshev, Stephanie Wang, Dirk Groeneveld, Luca Soldaini, Pang Wei Koh, Jenia Jitsev, Thomas Kollar, Alexandros G. Dimakis, Yair Carmon, Achal Dave, Ludwig Schmidt, Vaishaal Shankar |
NeurIPS | 23 |
| 2023 | SugarCrepe: Fixing Hackable Benchmarks for Vision-Language CompositionalityabstractIn the last year alone, a surge of new benchmarks to measure $\textit{compositional}$ understanding of vision-language models have permeated the machine learning ecosystem.Given an image, these benchmarks probe a model's ability to identify its associated caption amongst a set of compositional distractors.Surprisingly, we find significant biases in $\textit{all}$ these benchmarks rendering them hackable. This hackability is so dire that blind models with no access to the image outperform state-of-the-art vision-language models.To remedy this rampant vulnerability, we introduce $\textit{SugarCrepe}$, a new benchmark for vision-language compositionality evaluation.We employ large language models, instead of rule-based templates used in previous benchmarks, to generate fluent and sensical hard negatives, and utilize an adversarial refinement mechanism to maximally reduce biases. We re-evaluate state-of-the-art models and recently proposed compositionality inducing strategies, and find that their improvements were hugely overestimated, suggesting that more innovation is needed in this important direction.We release $\textit{SugarCrepe}$ and the code for evaluation at: https://github.com/RAIVNLab/sugar-crepe. Cheng-Yu Hsieh, Jieyu Zhang 0001, Zixian Ma, Aniruddha Kembhavi, Ranjay Krishna |
NeurIPS | 1 |
| 2022 | Summit-assisted Evolutionary MultitaskingabstractEvolutionary computation has served as a blooming research area for decades, in which evolutionary algorithms are inspired from mechanisms of evolution as well as cognitive and social behaviors of creatures in nature as searching processes. An emerging branch of evolutionary computation proposed in recent year is the evolutionary multitasking, which is to manipulate evolutionary algorithms for tackling multitask optimization problems. Recent studies of evolutionary multitasking have drawn much attention on designing effective knowledge transfer mechanisms. This study proposes a novel evolutionary multi-tasking method named summit-assisted evolutionary multitasking (SaEMT) by integrating summit-based knowledge transfer into the general multi-population evolutionary multitasking method. There are two main features in the proposed method, including the summit-based recombination, and the dynamic control of transfer rate. Empirical results show that the proposed method can outperform classical and advanced evolutionary multitasking methods in terms of solution quality and convergence speed. Experimental results also discover that the SaEMT is able to complete in acceptable running time. Cheng-Yu Hsieh, Rung-Tzuo Liaw |
CEC | 1 |
| 2022 | Understanding Programmatic Weak Supervision via Source-aware Influence FunctionabstractProgrammatic Weak Supervision (PWS) aggregates the source votes of multiple weak supervision sources into probabilistic training labels, which are in turn used to train an end model. With its increasing popularity, it is critical to have some tool for users to understand the influence of each component (\eg, the source vote or training data) in the pipeline and interpret the end model behavior. To achieve this, we build on Influence Function (IF) and propose source-aware IF, which leverages the generation process of the probabilistic labels to decompose the end model's training objective and then calculate the influence associated with each (data, source, class) tuple. These primitive influence score can then be used to estimate the influence of individual component of PWS, such as source vote, supervision source, and training data. On datasets of diverse domains, we demonstrate multiple use cases: (1) interpreting incorrect predictions from multiple angles that reveals insights for debugging the PWS pipeline, (2) identifying mislabeling of sources with a gain of 9\%-37\% over baselines, and (3) improving the end model's generalization performance by removing harmful components in the training objective (13\%-24\% better than ordinary IF). Jieyu Zhang 0001, Cheng-Yu Hsieh, Alexander Ratner |
NeurIPS | 3 |
| 2022 | Nemo: Guiding and Contextualizing Weak Supervision for Interactive Data ProgrammingabstractWeak Supervision (WS) techniques allow users to efficiently create large training datasets by programmatically labeling data with heuristic sources of supervision. While the success of WS relies heavily on the provided labeling heuristics, the process of how these heuristics are created in practice has remained under-explored. In this work, we formalize the development process of labeling heuristics as an interactive procedure, built around the existing workflow where users draw ideas from a selected set of development data for designing the heuristic sources. With the formalism, shown in Figure 1, we study two core problems of (1) how to strategically select the development data to guide users in efficiently creating informative heuristics, and (2) how to exploit the information within the development process to contextualize and better learn from the resultant heuristics. Building upon two novel methodologies that effectively tackle the respective problems considered, we present Nemo, an end-to-end interactive system that improves the overall productivity of WS learning pipeline by an average 20% (and up to 47% in one task) compared to the prevailing WS approach. Cheng-Yu Hsieh, Jieyu Zhang 0001, Alexander Ratner |
Proc. VLDB Endow. | 1 |
| 2021 | Evaluations and Methods for Explanation through Robustness Analysis
Cheng-Yu Hsieh, Chih-Kuan Yeh, Xuanqing Liu, Pradeep Ravikumar, Seungyeon Kim 0001, Sanjiv Kumar, Cho-Jui Hsieh |
ICLR | 1 |
| 2019 | On the (In)fidelity and Sensitivity of ExplanationsabstractWe consider objective evaluation measures of saliency explanations for complex black-box machine learning models. We propose simple robust variants of two notions that have been considered in recent literature: (in)fidelity, and sensitivity. We analyze optimal explanations with respect to both these measures, and while the optimal explanation for sensitivity is a vacuous constant explanation, the optimal explanation for infidelity is a novel combination of two popular explanation methods. By varying the perturbation distribution that defines infidelity, we obtain novel explanations by optimizing infidelity, which we show to out-perform existing explanations in both quantitative and qualitative measurements. Another salient question given these measures is how to modify any given explanation to have better values with respect to these measures. We propose a simple modification based on lowering sensitivity, and moreover show that when done appropriately, we could simultaneously improve both sensitivity as well as fidelity. Chih-Kuan Yeh, Cheng-Yu Hsieh, Arun Suggala, David I. Inouye, Pradeep Ravikumar |
NeurIPS | 2 |
| 2018 | A Deep Model With Local Surrogate Loss for General Cost-Sensitive Multi-Label LearningabstractMulti-label learning is an important machine learning problem with a wide range of applications. The variety of criteria for satisfying different application needs calls for cost-sensitive algorithms, which can adapt to different criteria easily. Nevertheless, because of the sophisticated nature of the criteria for multi-label learning, cost-sensitive algorithms for general criteria are hard to design, and current cost-sensitive algorithms can at most deal with some special types of criteria. In this work, we propose a novel cost-sensitive multi-label learning model for any general criteria. Our key idea within the model is to iteratively estimate a surrogate loss that approximates the sophisticated criterion of interest near some local neighborhood, and use the estimate to decide a descent direction for optimization. The key idea is then coupled with deep learning to form our proposed model. Experimental results validate that our proposed model is superior to existing cost-sensitive algorithms and existing deep learning models across different criteria. Cheng-Yu Hsieh, Yi-An Lin, Hsuan-Tien Lin |
AAAI | 1 |
| 2018 | Automatic Bridge Bidding Using Deep Reinforcement LearningabstractBridge is among the zero-sum games for which artificial intelligence has not yet outperformed expert human players. The main difficulty lies in the bidding phase of bridge, which requires cooperative decision making with partial information. Existing artificial intelligence systems for bridge bidding rely on, and are thus restricted by, human-designed bidding systems or features. In this work, we propose a flexible and pioneering bridge-bidding system, which can learn either with or without the aid of human domain knowledge. The system is based on a novel deep reinforcement learning model, which extracts sophisticated features and learns to bid automatically based on raw card data. The model includes an upper-confidence-bound algorithm and additional techniques to achieve a balance between exploration and exploitation. We further study how different pieces of human knowledge can be exploited to assist the model. Our experiments demonstrate the promising performance of our proposed model. In particular, the model can advance from having no knowledge on bidding to achieving a superior performance compared with a champion-winning computer bridge program that implements a human-designed bidding system. In addition, further synergies can be extracted by incorporating expert knowledge into the proposed model. Chih-Kuan Yeh, Cheng-Yu Hsieh, Hsuan-Tien Lin |
IEEE Trans. Games | 2 |
| 2013 | Efficient video segment matching for detecting temporal-based video copies
Chih-Yi Chiu, Cheng-Yu Hsieh |
Neurocomputing | 3 |
| 2013 | Efficient Video Stream Monitoring for Near-Duplicate Detection and Localization in a Large-Scale RepositoryabstractIn this article, we study the efficiency problem of video stream near-duplicate monitoring in a large-scale repository. Existing stream monitoring methods are mainly designed for a short video to scan over a query stream; they have difficulty being scalable for a large number of long videos. We present a simple but effective algorithm called incremental similarity update to address the problem. That is, a similarity upper bound between two videos can be calculated incrementally by leveraging the prior knowledge of the previous calculation. The similarity upper bound takes a lightweight computation to filter out unnecessary time-consuming computation for the actual similarity between two videos, making the search process more efficient. We integrate the algorithm with inverted indexing to obtain a candidate list from the repository for the given query stream. Meanwhile, the algorithm is applied to scan each candidate for locating exact near-duplicate subsequences. We implement several state-of-the-art methods for comparison in terms of accuracy, execution time, and memory consumption. Experimental results demonstrate the proposed algorithm yields comparable accuracy, compact memory size, and more efficient execution time. Chih-Yi Chiu, Guei-Wun Han, Cheng-Yu Hsieh, Sheng-Yang Li |
ACM Trans. Inf. Syst. | 4 |
| 2012 | Video Query Reformulation for Near-Duplicate DetectionabstractIn this paper, we present a novel near-duplicate video detection approach based on video query reformulation to expedite the video subsequence search process. The proposed video query reformulation method addresses two key issues: 1) how to efficiently skip unnecessary subsequence matches and 2) how to effectively increase the skip probability. First, we present an incremental update mechanism that rapidly estimates the similarity between two video subsequences to skip unnecessary matches. Second, we formulate an optimization problem of subsequence partition to increase the skip probability; a trust-region-based gradient descent algorithm is applied to solve the optimization problem. Extensive experiments cover various feature representations, subsequence granularities, and baseline methods; the results demonstrate that the proposed query reformulation method is robust and efficient to deal with a variety of near-duplicates in a large-scale video dataset. Chih-Yi Chiu, Sheng-Yang Li, Cheng-Yu Hsieh |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2011 | Design and implementation of dynamic charging plan for IMS-based multicast servicesabstract3GPP has standardized Multimedia Broadcast Multicast Service (MBMS) to deliver multimedia services with high bandwidth efficiency through IP multicasting technology. In this paper, we present the implementation details of an MBMS test-bed with the standardized Policy and Charging Control (PCC) to realize the dynamic charging plan for multicast services. Specifically, we illustrate the detailed message flow and the design block diagram with 3GPP policy and charging control. The corresponding performance metrics (e.g., signaling delays) have been measured through our test-bed. Sok-Ian Sou, Cheng-Yu Hsieh, Fen-Yen Lee, Yu-Fu Lin, Jeu-Yih Jeng, Chien-Wei Cheng |
APNOMS | 2 |
| 2007 | All-optical multicast routing in sparse splitting WDM networksabstractThis paper studies all-optical multicast routing in wavelength-routed optical networks with sparse light splitting. In a sparse splitting network, only a small percentage of nodes is capable of light splitting, i.e., multicast capable, and most of the nodes are multicast incapable. The typical approach to this problem is combining an existing Steiner tree heuristics with some rerouting procedure to refine the trees. Therefore, the cost in terms of the total number of wavelengths used for all tree links (referred to as wavelength channel cost) is very high. In this paper, we propose a new mechanism that constructs light-trees for sparse splitting optical networks without additional rerouting. We design two efficient schemes to build a light-tree for any given multicast session. We then extend our mechanism to support dynamic group membership. The simulation results show that our mechanism can build light-trees with the least wavelength channel cost and with the smallest number of wavelengths used per link. Cheng-Yu Hsieh, Wanjiun Liao |
IEEE J. Sel. Areas Commun. | 1 |
| 2003 | All Optical Multicast Routing in Sparse-Splitting Optical NetworksabstractThis paper studies all-optical multicast routing in wavelength-routed optical networks with sparse light splitting. In a sparse splitting network, only a small percentage of nodes are capable of light splitting, i.e., multicast capable. The typical solutions of existing multicast routing algorithms for sparse splitting networks combine an existing Steiner tree heuristic with some rerouting procedures to refine the trees. The resulting tree cost in terms of the total number of wavelengths used on all tree links is then very expensive. In this paper, we propose a new mechanism that constructs all-optical multicast trees for sparse splitting networks without an additional rerouting procedure in the tree construction. Two efficient approaches are suggested and evaluated by simulations. The results show that our mechanism builds light-trees with the least wavelength channel cost and with the smallest number of wavelengths used per link. Cheng-Yu Hsieh, Wanjiun Liao |
LCN | 1 |