VLDB 2026 Research / reviewers in the wild / expert
I-Hong Jhuo
dblp:80/7642
· DBLP profile ↗
24ranked-venue papers
11as first author
7since 2021 · last 2026
0009-0009-3893-3758ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 21 · 10 first-author · 6 since 2021Artificial intelligence and machine learning · 11 · 8 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GOT-JEPA: Generic Object Tracking With Model Adaptation and Occlusion Handling Using Joint-Embedding Predictive ArchitectureabstractThe human visual system tracks objects by integrating current observations with previously observed information, adapting to target and scene changes, and reasoning about occlusion at fine granularity. In contrast, recent generic object trackers are often optimized for training targets, which limits robustness and generalization in unseen scenarios, and their occlusion reasoning remains coarse, lacking detailed modeling of occlusion patterns. To address these limitations in generalization and occlusion perception, we propose GOT-JEPA, a model-predictive pretraining framework that extends JEPA from predicting image features to predicting tracking models. Given identical historical information, a teacher predictor generates pseudo-tracking models from a clean current frame, and a student predictor learns to predict the same pseudo-tracking models from a corrupted version of the current frame. This design provides stable pseudo supervision and explicitly trains the predictor to produce reliable tracking models under occlusions, distractors, and other adverse observations, improving generalization to dynamic environments. Building on GOT-JEPA, we further propose OccuSolver to enhance occlusion perception for object tracking. OccuSolver adapts a point-centric point tracker for object-aware visibility estimation and detailed occlusion-pattern capture. Conditioned on object priors iteratively generated by the tracker, OccuSolver incrementally refines visibility states, strengthens occlusion handling, and produces higher-quality reference labels that progressively improve subsequent model predictions. Extensive evaluations on seven benchmarks show that our method effectively enhances tracker generalization and robustness. The code will be available at https://github.com/chenshihfang/GOT. Shih-Fang Chen, Jun-Cheng Chen, I-Hong Jhuo, Yen-Yu Lin |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Consistent View Synthesis with Bidirectional Epipolar Attention and ReconstructionabstractNovel view synthesis from a single image aims to generate novel scene views given a reference image and a sequence of camera poses. Its primary difficulty lies in effectively leveraging a generative model to achieve high-quality image generation while simultaneously ensuring consistency and faithfulness across synthesized views. In this paper, we propose a novel approach to address the consistency and faithfulness issues in view synthesis. Specifically, we develop a new attention layer, termed bidirectional epipolar attention, which utilizes a pair of complementary epipolar lines to guide the associations between features from different viewpoints. Each bidirectional epipolar layer calculates forward and backward epipolar lines, enabling geometrically constrained attention that improves cross-view consistency. To ensure faithful synthesis, we introduce an epipolar-aware reconstruction module that prevents creating novel content in regions where the newly generated image overlaps with existing ones. Extensive experimental results demonstrate that our method outperforms previous approaches to novel view synthesis, achieving superior performance in both image quality and consistency. The source code is available at https://github.com/fallantbell/Bidirectional-Epipolar-Synthesis. I-Chung Chiu, Jun-Cheng Chen, I-Hong Jhuo, Yen-Yu Lin |
ICIP | 3 |
| 2025 | Generation and Comprehension Hand-in-Hand: Vision-guided Expression Diffusion for Boosting Referring Expression Generation and ComprehensionabstractReferring expression generation (REG) and comprehension (REC) are vital and complementary in joint visual and textual reasoning. Existing REC datasets typically contain insufficient image-expression pairs for training, hindering the generalization of REC models to unseen referring expressions. Moreover, REG methods frequently struggle to bridge the visual and textual domains due to the limited capacity, leading to low-quality and restricted diversity in expression generation. To address these issues, we propose a novel VIsion-guided Expression Diffusion Model (VIE-DM) for the REG task, where diverse synonymous expressions adhering to both image and text contexts of the target object are generated to augment REC datasets. VIE-DM consists of a vision-text condition (VTC) module and a transformer decoder. Our VTC and token selection design effectively addresses the feature discrepancy problem prevalent in existing REG methods. This enables us to generate high-quality, diverse synonymous expressions that can serve as augmented data for REC model learning. Extensive experiments on five datasets demonstrate the high quality and large diversity of our generated expressions. Furthermore, the augmented image-expression pairs consistently enhance the performance of existing REC models, achieving state-of-the-art results. Jingcheng Ke, Jun-Cheng Chen, I-Hong Jhuo, Chia-Wen Lin, Yen-Yu Lin |
ICLR | 3 |
| 2025 | Improving Visual Object Tracking Through Visual PromptingabstractLearning a discriminative model to distinguish a target from its surrounding distractors is essential to generic visual object tracking. Dynamic target representation adaptation against distractors is challenging due to the limited discriminative capabilities of prevailing trackers. We present a new visual Prompting mechanism for generic Visual Object Tracking (PiVOT) to address this issue. PiVOT proposes a prompt generation network with the pre-trained foundation model CLIP to automatically generate and refine visual prompts, enabling the transfer of foundation model knowledge for tracking. While CLIP offers broad category-level knowledge, the tracker, trained on instance-specific data, excels at recognizing unique object instances. Thus, PiVOT first compiles a visual prompt highlighting potential target locations. To transfer the knowledge of CLIP to the tracker, PiVOT leverages CLIP to refine the visual prompt based on the similarities between candidate objects and the reference templates across potential targets. Once the visual prompt is refined, it can better highlight potential target locations, thereby reducing irrelevant prompt information. With the proposed prompting mechanism, the tracker can generate improved instance-aware feature maps through the guidance of the visual prompt, thus effectively reducing distractors. The proposed method does not involve CLIP during training, thereby keeping the same training complexity and preserving the generalization capability of the pretrained foundation model. Extensive experiments across multiple benchmarks indicate that PiVOT, using the proposed prompting method can suppress distracting objects and enhance the tracker. Shih-Fang Chen, Jun-Cheng Chen, I-Hong Jhuo, Yen-Yu Lin |
IEEE Trans. Multim. | 3 |
| 2025 | Make Graph-Based Referring Expression Comprehension Great Again Through Expression-Guided Dynamic Gating and RegressionabstractOne common belief is that with complex models and pre-training on large-scale datasets, transformer-based methods for referring expression comprehension (REC) perform much better than existing graph-based methods. We observe that since most graph-based methods adopt an off-the-shelf detector to locate candidate objects (i.e., regions detected by the object detector), they face two challenges that result in subpar performance: (1) the presence of significant noise caused by numerous irrelevant objects during reasoning, and (2) inaccurate localization outcomes attributed to the provided detector. To address these issues, we introduce a plug-and-adapt module guided by sub-expressions, called dynamic gate constraint (DGC), which can adaptively disable irrelevant proposals and their connections in graphs during reasoning. We further introduce an expression-guided regression strategy (EGR) to refine location prediction. Extensive experimental results on the RefCOCO, RefCOCO+, RefCOCOg, Flickr30 K, RefClef, and Ref-reasoning datasets demonstrate the effectiveness of the DGC module and the EGR strategy in consistently boosting the performances of various graph-based REC methods. Without any pretaining, the proposed graph-based method achieves better performance than the state-of-the-art (SOTA) transformer-based methods. Jingcheng Ke, Dele Wang, Jun-Cheng Chen, I-Hong Jhuo, Chia-Wen Lin, Yen-Yu Lin |
IEEE Trans. Multim. | 4 |
| 2024 | CLIPREC: Graph-Based Domain Adaptive Network for Zero-Shot Referring Expression ComprehensionabstractReferring expression comprehension (REC) is a cross-modal matching task that aims to localize the target object in an image specified by a text description. Most existing approaches for this task focus on identifying only objects whose categories are covered by training data. This restricts their generalization to unseen categories and practical usage. To address this issue, we propose a domain adaptive network called CLIPREC for zero-shot REC, which integrates the Contrastive Language-Image Pretraining (CLIP) model for graph-based REC. The proposed CLIPREC is composed of a graph collaborative attention module with two directed graphs: one for objects in an image and the other for their corresponding categorical labels. To carry out zero-shot REC, we leverage the strong common image-text feature space from the CLIP model to correlate the two graphs. Furthermore, a multilayer perceptron is introduced to enable feature alignment so that the CLIP model is adapted to the expression representation from the language parser, resulting in effective reasoning from expressions involving both seen and unseen object categories. Extensive experimental and ablation results on several widely-adopted benchmarks show that the proposed approach performs favorably against state-of-the-art approaches for zero-shot REC. Jingcheng Ke, Jia Wang 0020, Jun-Cheng Chen, I-Hong Jhuo, Chia-Wen Lin, Yen-Yu Lin |
IEEE Trans. Multim. | 4 |
| 2023 | Task-Adaptive Feature Matching Loss for Image DeblurringabstractImage deblurring is a highly challenging and ill-posed image restoration problem. Contemporary deep learning-based approaches usually tackle this problem by exploiting the encoder-decoder-based models trained by the commonly used mean squared error loss with the feature matching loss as a regularization to obtain perceptual consistent restored results as the ground truths. We argue that since the general backbone models for computing feature matching loss are usually not trained on the image deblurring task, the loss lacks specific knowledge of blur and usually leads to suboptimal performance. To address this issue, we propose a task-adaptive feature matching loss for image deblurring where we synthesize blurred images in different blur extents and employ triplet loss to finetune the backbone model for learning specific blur priors. Then, we leverage the finetuned backbone to compute feature matching loss which can greatly enhance the existing image deblurring models for better perceptual results. With extensive experiments on the GoPro and RealBlur datasets, both qualitative and quantitative results show that the SOTA deblurring models trained with the proposed loss can effectively obtain better and sharper restored images in terms of various perceptual image quality metrics than the original models while maintaining comparable PSNR and SSIM performances. Chiao-Chang Chang, Bo-Cheng Yang, Yi-Ting Liu, Jun-Cheng Chen, I-Hong Jhuo, Yen-Yu Lin |
ICIP | 5 |
| 2020 | Trajectory Prediction in Heterogeneous Environment via Attended Ecology EmbeddingabstractTrajectory prediction is a highly desirable feature for safe navigation or autonomous vehicle in complex traffic. In this paper, we consider the practical environment of predicting trajectory in the heterogeneous traffic ecology. The proposed method has various applications in trajectory prediction problems and also in applied fields beyond tracking. One challenge stands out of the trajectory prediction-heterogeneous environment. Particularly, many factors should be considered in the environments, i.e., multiple types of road-agents, social interactions and terrains. The information is complicated and large that may result in inaccurate trajectory prediction. We propose two social and visual enforced attention modules to circumvent the problem and a variant of an Info-GAN structure to predict the trajectory with multi-modal behaviors. Experimental results show that the proposed method significantly outperforms state-of-the-art methods in both heterogeneous and homogeneous real environments. Wei-Cheng Lai, Zi-Xiang Xia, Hao-Siang Lin, Lien-Feng Hsu, Hong-Han Shuai, I-Hong Jhuo, Wen-Huang Cheng |
ACM Multimedia | 6 |
| 2019 | Supervised Set-to-Set Hashing in Visual RecognitionabstractVisual data, such as an image or a sequence of video frames, is often naturally represented as a point set. In this paper, we consider the fundamental problem of finding a nearest set from a collection of sets, to a query set. This problem has obvious applications in large-scale visual retrieval and recognition, and also in applied fields beyond computer vision. One challenge stands out in solving the problem---set representation and measure of similarity. Particularly, the query set and the sets in dataset collection can have varying cardinalities. The training collection is large enough such that linear scan is impractical. We propose a simple representation scheme that encodes both statistical and structural information of the sets. The derived representations are integrated in a kernel framework for flexible similarity measurement. For the query set process, we adopt a learning-to-hash pipeline that turns the kernel representations into hash bits based on simple learners, using multiple kernel learning. Experiments on two visual retrieval datasets show unambiguously that our set-to-set hashing framework outperforms prior methods that do not take the set-to-set search setting. I-Hong Jhuo |
IJCAI | 1 |
| 2016 | A feature fusion framework for hashingabstractA hash algorithm converts data into compact strings. In the multimedia domain, effective hashing is the key to large-scale similarity search in high-dimensional feature space. A limit of existing hashing techniques is that they typically use single features. In order to improve search performance, it is necessary to utilize multiple features. Due to the compactness requirement, concatenation of hash values from different features is not an optimal solution. Thus a fusion process is desired. In this paper, we solve the multiple feature fusion problem by a hash bit selection framework. Given multiple features, we derive an n-bit hash value of improved performance compared with hash values of the same length computed from each individual feature. The framework utilizes a feature-independent hash algorithm to generate a sufficient number of bits from each feature, and selects n bits from the hash bit pool by leveraging pair-wise label information. The metric bit reliability is used for ranking the bits. It is estimated by bit-level hypothesis testing. In addition, we also take into account the dependence among bits. A weighted graph is constructed for refined bit selection, where the bit reliability is used as vertex weights and the mutual information among hash bits is used as edge weights. We demonstrate our framework with LSH. Extensive experiments confirm that our method is effective, and outperforms several state-of-the-art methods. I-Hong Jhuo, Li Weng, Wen-Huang Cheng, D. T. Lee |
ICPR | 1 |
| 2015 | Supervised Multi-scale Locality Sensitive HashingabstractLSH is a popular framework to generate compact representations of multimedia data, which can be used for content based search. However, the performance of LSH is limited by its unsupervised nature and the underlying feature scale. In this work, we propose to improve LSH by incorporating two elements - supervised hash bit selection and multi-scale feature representation. First, a feature vector is represented by multiple scales. At each scale, the feature vector is divided into segments. The size of a segment is decreased gradually to make the representation correspond to a coarse-to-fine view of the feature. Then each segment is hashed to generate more bits than the target hash length. Finally the best ones are selected from the hash bit pool according to the notion of bit reliability, which is estimated by bit-level hypothesis testing. Li Weng, I-Hong Jhuo, Miaojing Shi, Meng Sun 0001, Wen-Huang Cheng, Laurent Amsaleg |
ICMR | 2 |
| 2014 | Unsupervised Feature Learning for RGB-D Image Classification
I-Hong Jhuo, Shenghua Gao, Liansheng Zhuang, D. T. Lee, Yi Ma 0001 |
ACCV (1) | 1 |
| 2014 | Model reference adaptive iterative learning control for nonlinear systems using observer designabstractIn this paper, we propose an observer based model reference adaptive iterative learning control (MRAILC) using model reference adaptive control strategy for more general class of uncertain nonlinear systems with non-canonical form and iteration-varying reference trajectories. Due to the system state vector is assumed to be unmeasurable, a state tracking error observer is applied for state tracking error estimation. Based on the state tracking error observer and a mixed time-domain and s-domain technique, a relative degree one output observation error model whose inputs are some uncertain nonlinearities and filtered signals which is derived to solve the relative degree problem caused by the system states are not measurable. Besides, we also apply some auxiliary signals and an averaging filter to transfer the original output observation error to a new formulation so that we can implement the AILC without using differentiators. The filtered fuzzy neural network (filtered-FNN) using the system state estimation vector as the input vector is applied for approximation of the unknown plant nonlinearities. In order to overcome the lumped uncertainties associated with function approximation error and state estimation error, a normalization signal is applied as a bounding function for designing a robust AILC. The stabilization learning component is used to guarantee the boundedness of internal signals. Based on a Lyapunov like analysis, we show that all the adjustable parameters as well as internal signals remain bounded for all iterations and the norm of output tracking error will asymptotically converge to a tunable residual set. Ying-Chung Wang, Chiang-Ju Chien, I-Hong Jhuo |
FUZZ-IEEE | 3 |
| 2014 | Image auto-annotation by exploiting web informationabstractWe consider the image auto-annotation problem by exploiting information from Internet. Given a collection of semantically similar images and a keyword that accurately describes these images, our goal is to find a set of popular tags to annotate each image, conforming to those used for similar images found on the web. We propose a novel framework to exploit classification based learning and bipartitioning clustering algorithms for extracting meaningful tags from semantical images on the web. Specifically, we adopt multiple kernel learning (MKL) to first select relevant images with their associated tags, which are obtained from the web based on keyword search, and then build a bipartite graph to model the relation between related tags and images. Finally, we partition over the bipartite graph to produce a set of significant tags for each image. We evaluate our proposed method by using the colorful Natural Scene and Events datasets to generate related images and tags from the Flickr website. The experimental results show that our proposed method has superior performance compared with baseline methods. I-Hong Jhuo, Li Weng |
ICIP | 1 |
| 2014 | Video Event Detection via Multi-modality Deep LearningabstractDetecting complex video events based on audio and visual modalities is still a largely unresolved issue. While the conventional video representation methods extract each modality ineffectively, we propose a regularized multi-modality deep learning for video event detection. We first build an auto-encoder based on unconstrained minimization and adopt the conjugate gradient method with linear search for optimization. The learned auto-encoder can capture the relationship between the audio and visual modality corresponding to the same video event at each layer of the network. To make the network robust to local variance, we adopt the commonly used local contrast normalization and spatial maximum pooling to each modality for video representation. Compared with traditional methods using manually designed features, our method is more efficient. Experimental results on publicly available video event detection datasets demonstrate that the proposed method consistently outperforms the state-of-the-art video representation methods. I-Hong Jhuo, D. T. Lee |
ICPR | 1 |
| 2014 | Discovering joint audio-visual codewords for video event detection
I-Hong Jhuo, Guangnan Ye, Shenghua Gao, Dong Liu 0001, Yu-Gang Jiang 0001, D. T. Lee, Shih-Fu Chang |
Mach. Vis. Appl. | 1 |
| 2012 | Robust visual domain adaptation with low-rank reconstructionabstractVisual domain adaptation addresses the problem of adapting the sample distribution of the source domain to the target domain, where the recognition task is intended but the data distributions are different. In this paper, we present a low-rank reconstruction method to reduce the domain distribution disparity. Specifically, we transform the visual samples in the source domain into an intermediate representation such that each transformed source sample can be linearly reconstructed by the samples of the target domain. Unlike the existing work, our method captures the intrinsic relatedness of the source samples during the adaptation process while uncovering the noises and outliers in the source domain that cannot be adapted, making it more robust than previous methods. We formulate our problem as a constrained nuclear norm and ℓ2, 1norm minimization objective and then adopt the Augmented Lagrange Multiplier (ALM) method for the optimization. Extensive experiments on various visual adaptation tasks show that the proposed method consistently and significantly beats the state-of-the-art domain adaptation methods. I-Hong Jhuo, Dong Liu 0001, D. T. Lee, Shih-Fu Chang |
CVPR | 1 |
| 2012 | Robust late fusion with rank minimizationabstractIn this paper, we propose a rank minimization method to fuse the predicted confidence scores of multiple models, each of which is obtained based on a certain kind of feature. Specifically, we convert each confidence score vector obtained from one model into a pairwise relationship matrix, in which each entry characterizes the comparative relationship of scores of two test samples. Our hypothesis is that the relative score relations are consistent among component models up to certain sparse deviations, despite the large variations that may exist in the absolute values of the raw scores. Then we formulate the score fusion problem as seeking a shared rank-2 pairwise relationship matrix based on which each original score matrix from individual model can be decomposed into the common rank-2 matrix and sparse deviation errors. A robust score vector is then extracted to fit the recovered low rank score relation matrix. We formulate the problem as a nuclear norm and ℓ1norm optimization objective function and employ the Augmented Lagrange Multiplier (ALM) method for the optimization. Our method is isotonic (i.e., scale invariant) to the numeric scales of the scores originated from different models. We experimentally show that the proposed method achieves significant performance gains on various tasks including object categorization and video event detection. Guangnan Ye, Dong Liu 0001, I-Hong Jhuo, Shih-Fu Chang |
CVPR | 3 |
| 2012 | Joint audio-visual bi-modal codewords for video event detectionabstractJoint audio-visual patterns often exist in videos and provide strong multi-modal cues for detecting multimedia events. However, conventional methods generally fuse the visual and audio information only at a superficial level, without adequately exploring deep intrinsic joint patterns. In this paper, we propose a joint audio-visual bi-modal representation, called bi-modal words. We first build a bipartite graph to model relation across the quantized words extracted from the visual and audio modalities. Partitioning over the bipartite graph is then applied to construct the bi-modal words that reveal the joint patterns across modalities. Finally, different pooling strategies are employed to re-quantize the visual and audio words into the bi-modal words and form bi-modal Bag-of-Words representations that are fed to subsequent multimedia event classifiers. We experimentally show that the proposed multi-modal feature achieves statistically significant performance gains over methods using individual visual and audio features alone and alternative multi-modal fusion methods. Moreover, we found that average pooling is the most suitable strategy for bi-modal feature generation. Guangnan Ye, I-Hong Jhuo, Dong Liu 0001, Yu-Gang Jiang 0001, D. T. Lee, Shih-Fu Chang |
ICMR | 2 |
| 2012 | The acousticvisual emotion guassians model for automatic generation of music videoabstractThis paper presents a novel content-based system that utilizes the perceived emotion of multimedia content as a bridge to connect music and video. Specifically, we propose a novel machine learning framework, called Acousticvisual Emotion Guassians (AVEG), to jointly learn the tripartite relationship among music, video, and emotion from an emotion-annotated corpus of music videos. For a music piece (or a video sequence), the AVEG model is applied to predict its emotion distribution in a stochastic emotion space from the corresponding low-level acoustic (resp. visual) features. Finally, music and video are matched by measuring the similarity between the two corresponding emotion distributions, based on a distance measure such as KL divergence. Ju-Chiang Wang, Yi-Hsuan Yang, I-Hong Jhuo, Yen-Yu Lin, Hsin-Min Wang |
ACM Multimedia | 3 |
| 2011 | Multiple-Instance Learning: Multiple Feature Selection on Instance RepresentationabstractIn multiple-Instance Learning (MIL), training class labels are attached to sets of bags composed of unlabeled instances, and the goal is to deal with classification of bags. Most previous MIL algorithms, which tackle classification problems, consider each instance as a represented feature. Although the algorithms work well in some prediction problems, considering diverse features to represent an instance may provide more significant information for learning task. Moreover, since each instance may be mapped into diverse feature spaces, encountering a large number of irrelevant or redundant features is inevitable. In this paper, we propose a method to select relevant instances and concurrently consider multiple features for each instance, which is termed as MIL-MFS. MIL-MFS is based on multiple kernel learning (MKL), and it iteratively selects the fusing multiple features for classifier training. Experimental results show that the MIL-MFS combined with multiple kernel learning can significantly improve the classification performance. I-Hong Jhuo, D. T. Lee |
AAAI | 1 |
| 2010 | Boosted Multiple Kernel Learning for Scene Category RecognitionabstractScene images typically include diverse and distinctive properties. It is reasonable to consider different features in establishing a scene category recognition system with a promising performance. We propose an adaptive model to represent various features in a unified domain, i.e., a set of kernels, and transform the discriminant information contained in each kernel into a set of weak learners, called dyadic hyper cuts. Based on this model, we present a novel approach to carrying out incremental multiple kernel learning for feature fusion by applying AdaBoost to the union of the sets of weak learners. We further evaluate the performance of this approach by a benchmark dataset for scene category recognition. Experimental results show a significantly improved performance in both accuracy and efficiency. I-Hong Jhuo, D. T. Lee |
ICPR | 1 |
| 2010 | Boosting-based multiple kernel learning for image re-rankingabstractRe-ranking the returned images from a query relies on two important steps to improve its effectiveness: the estimation of the image relevance and the enhancement of the similarity function. However, attaining an effective visual similarity and an accurate re-ranking are quite challenging. We address these issues by first evaluating the image relevance to the query from the dataset according to the visual features and the co-occurrence of local patches of images. Then we boost the visual similarity measure associated with image relevance, and propose an enhancement algorithm, called Boosting-MKL, which not only incrementally learns the feature fusion but also generally preserves the initial local ranking. Specifically, we perform a random walk over a similarity graph for re-ranking. The experimental results demonstrate that our proposed approach significantly improves the effectiveness of visual similarity measure and the performance of image reranking. I-Hong Jhuo, D. T. Lee |
ACM Multimedia | 1 |
| 2010 | Scene Location Guide by Image-Based Retrieval
I-Hong Jhuo, Tsuhan Chen, D. T. Lee |
MMM | 1 |