Guangnan Ye

dblp:20/9998 · DBLP profile ↗
← Back
30ranked-venue papers
4as first author
19since 2021 · last 2026
0009-0007-4973-7942ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 18 · 4 first-author · 8 since 2021Artificial intelligence and machine learning · 14 · 2 first-author · 9 since 2021Databases, data management, data science and information retrieval · 7 · 1 first-author · 4 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Rethinking the Role of Entropy in Optimizing Tool-Use Behaviors for Large Language Model Agents
abstract
Zeping Li, Hongru Wang, Yiwen Zhao, Guanhua Chen, Yixia Li, Keyang Chen, Yixin Cao, Guangnan Ye, Hongfeng Chai, Zhenfei Yin. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Zeping Li, Hongru Wang 0003, Guanhua Chen 0001, Yixia Li, Keyang Chen, Yixin Cao 0002, Guangnan Ye, Hongfeng Chai, Zhenfei Yin
ACL (1)8
2026 GQLBench: A Large-Scale Cross-Domain, Cross-Dialect Benchmark for NL2GQL
abstract
Despite growing interest in NL2GQL, benchmarking progress has been constrained by the lack of resources that are simultaneously largescale, cross-domain, and cross-dialect.To address this gap, we present GQLBench, a new benchmark built through an automated and scalable framework that integrates NL2SQL-to-NL2GQL conversion with graph-native data generation.GQLBench supports executionbased evaluation on both Cypher and ISO GQL, covering hundreds of graph databases and over 20k natural language questions for each dialect.By combining converted data from mature NL2SQL resources with synthetic graphspecific queries, it captures both schema diversity from real-world relational sources and graph-native reasoning challenges, including long paths and cycles.Beyond overall performance comparison, GQLBench also enables fine-grained evaluation across dialects, graph patterns, and query complexity.Experiments on advanced LLMs show that even strong proprietary models struggle on GQLBench, with gemini-3-flash achieving only 35.40% average execution accuracy across the two dialects.Our data and code are available at https://github.com/qxssadf/GQLBench.
Yanning Su, Guangnan Ye, Hongfeng Chai
ACL (1)5
2026 PRISM: Probabilistic Reward Model with Inherent Structural Modeling
abstract
Yuhang Zhou, Yixin Cao, Yuchen Ni, Shihan Dou, Xutian Chen, Ge Zhang, Xiang Liu, Guangnan Ye. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yixin Cao 0002, Yuchen Ni, Shihan Dou, Xutian Chen, Ge Zhang 0009, Guangnan Ye
ACL (1)8
2026 TabLoft: Tabular Data Generation Based on LLM with Ordered Features
Luyu Chen, Changhao Wu, Guangnan Ye, Hongfeng Chai
ICDE5
2026 Dynamic graph learning for integrating temporal relationships in stock prediction
Ziyue Dai, Qianru Zeng, Nianwang Lin, Hongjie Xia, Keyu Zhao, Sen Liu 0002, Guangnan Ye, Jie Wu 0003, Hongfeng Chai
Expert Syst. Appl.9
2026 MasterKey: A multi-target backdoor attack in federated learning
Haohe Jia, Hongbin Zhu, Guangnan Ye, Hongfeng Chai
Knowl. Based Syst.4
2026 Graph Self-Supervised Learning via Learnable View Augmentation for Recommender System
abstract
In the field of recommender systems, graph neural networks (GNNs) have been extensively applied to collaborative filtering to generate personalized recommendations for users. To solve the problem of lack of observed data and contrasting interactions during representation learning, graph contrastive learning as an effective self-supervised learning (SSL) technique is presented to obtain augmented user and item representations. Nevertheless, most self-supervised approaches to generate recommendation either disrupt the graph structure or node embeddings through random augmentations or introduce augmented SSL information from biased data through heuristic methods. To overcome these challenges, we propose a learnable view augmentation model for collaborative filtering (LACF). Specifically, our framework embeds parameterized learnable view generators layer by layer into the automatic augmentation strategy, thus dynamically optimizing the adaptive augmented views of users and items through the backpropagation of weight gradients. In addition, LACF introduces a multiscale learning strategy that guides the view generator with layer-wise aware optimization and graph-level adaptive augmentation, enabling joint learning of representations with topological heterogeneity and semantic similarity from integrated viewpoint, achieving superior view augmentation. Extensive experiments on realworld datasets demonstrate that our LACF outperforms state-of-the-art baselines. In-depth analysis confirms the advantages of LACF in resistance against noise disturbances, alleviating data sparsity, and improving training efficiency.
Hengjing Xiang, Yanfeng Xu, Sen Liu 0002, Zhi Liu 0011, Guangnan Ye
IEEE Trans. Ind. Informatics7
2025 StageDesigner: Artistic Stage Generation for Scenography via Theater Scripts
abstract
In this work, we introduce StageDesigner, the first comprehensive framework for artistic stage generation using large language models combined with layout-controlled diffusion models. Given the professional requirements of stage scenography, StageDesigner simulates the workflows of seasoned artists to generate immersive 3D stage scenes. Specifically, our approach is divided into three primary modules: Script Analysis, which extracts thematic and spatial cues from input scripts; Foreground Generation, which constructs and arranges essential 3D objects; and Background Generation, which produces a harmonious background aligned with the narrative atmosphere and maintains spatial coherence by managing occlusions between foreground and background elements. Furthermore, we introduce the StagePro-V1 dataset, a dedicated dataset with 276 unique stage scenes spanning different historical styles and annotated with scripts, images, and detailed 3D layouts, specifically tailored for this task. Finally, evaluations using both standard and newly proposed metrics, along with extensive user studies, demonstrate the effectiveness of StageDesigner, showcasing its ability to produce visually and thematically cohesive stages that meet both artistic and spatial coherence standards. Project can be found at: https://deadsmither5.github.io/2025/01/03/StageDesigner/
Zhaoxing Gan, Mengtian Li 0002, Ruhua Chen, Zhongxia Ji, Sichen Guo, Huanling Hu, Guangnan Ye, Zuo Hu
CVPR7
2025 Towards Million-Scale Adversarial Robustness Evaluation With Stronger Individual Attacks
abstract
As deep learning models are increasingly deployed in safety-critical applications, evaluating their vulnerabilities to adversarial perturbations is essential for ensuring their reliability and trustworthiness. Over the past decade, a large number of white-box adversarial robustness evaluation methods (i.e., attacks) have been proposed, ranging from single-step to multi-step methods and from individual to ensemble methods. Despite these advances, challenges remain in conducting meaningful and comprehensive robustness evaluations, particularly when it comes to large-scale testing and ensuring evaluations reflect real-world adversarial risks. In this work, we focus on image classification models and propose a novel individual attack method, Probability Margin Attack (PMA), which defines the adversarial margin in the probability space rather than the logits space. We analyze the relationship between PMA and existing cross-entropy or logits-margin-based attacks, and show that PMA can outperform the current state-of-the-art individual methods. Building on PMA, we propose two types of ensemble attacks that balance effectiveness and efficiency. Furthermore, we create a million-scale dataset, CC1M, derived from the existing CC3M dataset, and use it to conduct the first million-scale white-box adversarial robustness evaluation of adversarially-trained ImageNet models. Our findings provide valuable insights into the robustness gaps between individual versus ensemble attacks and small-scale versus million-scale evaluations.
Hanxun Huang, Guangnan Ye, Xingjun Ma
CVPR4
2025 Enhancing Federated Knowledge Distillation in Heterogeneous and Non-IID Scenarios
abstract
Federated Learning (FL) allows multiple participants to train models together while keeping their data private. Some FL frameworks use Knowledge Distillation to address model heterogenity, but many struggle in non-IID and heterogeneous environments, making it hard for clients to learn from each other. In this work, we show that the entropy of the softmax-averaged logits from clients reflects the model’s convergence. Based on this, we propose a new loss function, Sharpened Symmetric KL Divergence Loss (SSKL), which combines KL and Reverse KL Divergence with Label Sharpening to reduce the impact of non-IID data. Experiments demonstrate that our approach improves performance and reduces accuracy decline in non-IID and heterogeneous settings.
Wenjie Lv, Xingjun Ma, Guangnan Ye, Hongfeng Chai
ICASSP6
2025 FedCAda: Adaptive Client-Side Optimization for Accelerated and Stable Federated Learning
abstract
Federated learning (FL) enables collaborative model training across distributed clients while preserving data privacy. However, achieving both acceleration and stability, particularly on the client side, remains a challenge. In this paper, we introduce FedCAda, an adaptive algorithm that leverages an Adam-like approach to adjust first and second moment estimates on the client side while aggregating adaptive parameters on the server side. This design aims to accelerate convergence without compromising stability and performance. We also explore several algorithms with different adjustment functions and find that stronger constraints on adaptive parameters are necessary in the early stages of FL, when information from other clients is limited. Experiments on public datasets demonstrate that FedCAda surpasses state-of-the-art methods in adaptability, convergence, and stability, advancing adaptive algorithms for FL.
Liuzhi Zhou, Kun Zhai, Xingjun Ma, Guangnan Ye, Hongfeng Chai
ICASSP7
2025 FSPFL: Mitigating Communication Gap in Personalization Federated Learning Through Flexible Sparsity Allocation
abstract
Federated learning has emerged as a promising distributed learning paradigm that enables model training across decentralized devices while preserving data privacy. However, two critical challenges hinder its practical deployment: model performance degradation due to client drift in high data heterogeneity scenarios and communication gap due to varying client communication capabilities. In this paper, we provide a comprehensive analysis of these challenges and propose Flexible Sparse Personalized Federated Learning (FSPFL), a novel framework that jointly optimizes model personalization and communication efficiency. FSPFL adaptively allocates local model sparsity while incorporating personalization mechanisms to trade off model performance and communication efficiency. Extensive experiments demonstrate that FSPFL significantly mitigates the communication gap, and outperforms existing methods. Our results show that FSPFL improves communication efficiency by up to$4.1 \times$than the baselines while maintaining similar model accuracy in heterogeneous scenarios where clients have diverse data distributions and communication capabilities.
Liuzhi Zhou, Nianwang Lin, Sen Liu 0002, Guangnan Ye, Yimin Yu, Hongfeng Chai
IWQoS6
2025 ArtNVG: Content-Style Separated Artistic Neighboring-View Gaussian Stylization
abstract
As demand from the film and gaming industries for 3D scenes with target styles grows, the importance of advanced 3D stylization techniques increases. However, recent methods often struggle to maintain local consistency in color and texture throughout stylized scenes, which is essential for maintaining aesthetic coherence. To solve this problem, this paper introduces ArtNVG, an innovative 3D stylization framework that efficiently generates stylized 3D scenes by leveraging reference style images. Built on 3D Gaussian Splatting (3DGS), ArtNVG achieves rapid optimization and rendering while upholding high reconstruction quality. Our framework realizes high-quality 3D stylization by incorporating two pivotal techniques: Content-Style Separated Control and Attention-based Neighboring-View Alignment. Content-Style Separated Control uses the CSGO model and the Tile ControlNet to decouple the content and style control, reducing risks of information leakage. Concurrently, Attention-based Neighboring-View Alignment ensures consistency of local colors and textures across neighboring views, significantly improving visual quality. Extensive experiments validate that ArtNVG surpasses existing SOTA methods, delivering superior results in content preservation, style alignment, and local consistency.
Zixiao Gu, Zhenye Zhang, Mengtian Li 0002, Zhongxia Ji, Ruhua Chen, Zuo Hu, Guangnan Ye
ICMR7
2025 EditMaster: Bridging Text instruction and Visual Example for Multimodal guided Image Editing
abstract
Recent advances in image editing systems reveal critical limitations in handling complex real-world scenarios requiring multimodal condition controls. While text instructions enable broad semantic guidance, visual examples provide precise visual reference in specific scenarios, existing unimodal approaches fail to synergize these complementary modalities effectively. We propose EditMaster, a unified framework that integrates text and visual controls through multimodal instruction learning, enabling precise image manipulation with bidirectional consistency. Our framework introduces three core innovations: A multimodal large language model enhanced with extended visual tokens replaces CLIP text encoders, generating pre-edited visual guidance that aligns textual commands with visual examples to guide diffusion model toward high-quality outputs; The Mask-Based Decoupled Residual Exemplar-Attention module preserves unedited regions through spatial masking while integrating visual details via residual pathways; A systematic data construction method converts unimodal editing datasets into a task-specific multimodal dataset, eliminating the need for de novo data construction. Experiments show that our approach outperforms unimodal baselines and excels in complex multimodal instruction editing, setting a new benchmark for this field.
Mengtian Li 0002, Jiewei Tang, Junyu Deng, Guangnan Ye, Yu-Gang Jiang 0001
ACM Multimedia8
2025 FinIR: The 2nd Workshop on Financial Information Retrieval in the Era of Generative AI
abstract
Recent advancements in Generative AI, such as Large Language Models (LLMs), have demonstrated remarkable success across various general tasks. Extensive studies have explored leveraging generative models in finance, but significant challenges persist. This half-day workshop explores potential approaches and research directions to address these challenges by equipping generative models with advanced Information Retrieval (IR) models. Specifically, this workshop seeks to provide a platform for discussing innovative ideas that facilitate the advancement of IR technology to enrich generative models in finance from four key perspectives: (i) financial IR techniques (ii) financial IR benchmarking and evaluation (iii) financial systems and agents/assistants (iv) and trustworthiness, privacy and security when applying financial IR and generative models. This workshop aims to deepen understanding, accelerate progress, and support the advancement of IR technology to enhance generative models to address financial challenges.
Fengbin Zhu, Yunshan Ma 0002, Fuli Feng, Chao Wang 0049, Huan-Bo Luan, Guangnan Ye, Shuo Zhang 0006, Dhagash Mehta, Pingping Chen 0004, Bing Xiang, Tat-Seng Chua
SIGIR6
2025 MultiHGPT: Multi-task heterogeneous graph prompt tuning
Yixin Cao 0002, Zeping Li, Guangnan Ye, Hongfeng Chai
Inf. Process. Manag.5
2024 CI-STHPAN: Pre-trained Attention Network for Stock Selection with Channel-Independent Spatio-Temporal Hypergraph
abstract
Quantitative stock selection is one of the most challenging FinTech tasks due to the non-stationary dynamics and complex market dependencies. Existing studies rely on channel mixing methods, exacerbating the issue of distribution shift in financial time series. Additionally, complex model structures they build make it difficult to handle very long sequences. Furthermore, most of them are based on predefined stock relationships thus making it difficult to capture the dynamic and highly volatile stock markets. To address the above issues, in this paper, we propose Channel-Independent based Spatio-Temporal Hypergraph Pre-trained Attention Networks (CI-STHPAN), a two-stage framework for stock selection, involving Transformer and HGAT based stock time series self-supervised pre-training and stock-ranking based downstream task fine-tuning. We calculate the similarity of stock time series of different channel in dynamic intervals based on Dynamic Time Warping (DTW), and further construct channel-independent stock dynamic hypergraph based on the similarity. Experiments with NASDAQ and NYSE markets data over five years show that our framework outperforms SOTA approaches in terms of investment return ratio (IRR) and Sharpe ratio (SR). Additionally, we find that even without introducing graph information, self-supervised learning based on the vanilla Transformer Encoder also surpasses SOTA results. Notable improvements are gained on the NYSE market. It is mainly attributed to the improvement of fine-tuning approach on Information Coefficient (IC) and Information Ratio based IC (ICIR), indicating that the fine-tuning method enhances the accuracy and stability of the model prediction.
Hongjie Xia, Huijie Ao, Guangnan Ye, Hongfeng Chai
AAAI6
2024 White-box Multimodal Jailbreaks Against Large Vision-Language Models
abstract
Recent advancements in Large Vision-Language Models (VLMs) have underscored their superiority in various multimodal tasks. However, the adversarial robustness of VLMs has not been fully explored. Existing methods mainly assess robustness through unimodal adversarial attacks that perturb images, while assuming inherent resilience against text-based attacks. Different from existing attacks, in this work we propose a more comprehensive strategy that jointly attacks both text and image modalities to exploit a broader spectrum of vulnerability within VLMs. Specifically, we propose a dual optimization objective aimed at guiding the model to generate highly toxic affirmative responses. Our attack method begins by optimizing an adversarial image prefix from random noise to generate diverse harmful responses in the absence of text input, thus imbuing the image with toxic semantics. Subsequently, an adversarial text suffix is integrated and co-optimized with the adversarial image prefix to maximize the probability of eliciting affirmative responses to various harmful instructions. The discovered adversarial image prefix and text suffix are collectively denoted as a Universal Master Key (UMK). When integrated into various malicious queries, UMK can circumvent the alignment defenses of VLMs and lead to the generation of objectionable content, known as jailbreaks. The experimental results demonstrate that our universal attack strategy can effectively jailbreak MiniGPT-4 with a 96% success rate, highlighting the fragility of VLMs and the exigency for new alignment strategies. Codes are available at https://github.com/roywang021/UMK. Disclaimer: This paper contains potentially disturbing and offensive content.
Xingjun Ma, Hanxu Zhou, Chuanjun Ji, Guangnan Ye, Yu-Gang Jiang 0001
ACM Multimedia5
2021 On Sample Based Explanation Methods for NLP: Faithfulness, Efficiency and Semantic Evaluation
abstract
Wei Zhang, Ziming Huang, Yada Zhu, Guangnan Ye, Xiaodong Cui, Fan Zhang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Wei Zhang 0057, Ziming Huang, Yada Zhu, Guangnan Ye, Fan Zhang 0003
ACL/IJCNLP (1)4
2015 Large Video Event Ontology Browsing, Search and Tagging (EventNet Demo)
abstract
In this demo we present PITAGORA\footnote{Demo video available at http://bit.ly/1GgtUrN}: a mobile web contextual social network designed for the check-in area of an airport. The app provides recommendation of potential friends, local experts and targeted services. Recommendation is hybrid and combines social media analysis and collaborative filtering techniques. Users' recommendation has been evaluated through a user study with good results.
Hongliang Xu, Guangnan Ye, Dong Liu 0001, Shih-Fu Chang
ACM Multimedia2
2015 EventNet: A Large Scale Structured Concept Library for Complex Event Detection in Video
abstract
Event-specific concepts are the semantic concepts specifically designed for the events of interest, which can be used as a mid-level representation of complex events in videos. Existing methods only focus on defining event-specific concepts for a small number of pre-defined events, but cannot handle novel unseen events. This motivates us to build a large scale event-specific concept library that covers as many real-world events and their concepts as possible. Specifically, we choose WikiHow, an online forum containing a large number of how-to articles on human daily life events. We perform a coarse-to-fine event discovery process and discover 500 events from WikiHow articles. Then we use each event name as query to search YouTube and discover event-specific concepts from the tags of returned videos. After an automatic filter process, we end up with 95,321 videos and 4,490 concepts. We train a Convolutional Neural Network (CNN) model on the 95,321 videos over the 500 events, and use the model to extract deep learning feature from video content. With the learned deep learning feature, we train 4,490 binary SVM classifiers as the event-specific concept library. The concepts and events are further organized in a hierarchical structure defined by WikiHow, and the resultant concept library is called EventNet. Finally, the EventNet concept library is used to generate concept based representation of event videos. To the best of our knowledge, EventNet represents the first video event ontology that organizes events and their concepts into a semantic structure. It offers great potential for event retrieval and browsing. Extensive experiments over the zero-shot event retrieval task when no training samples are available show that the proposed EventNet concept library consistently and significantly outperforms the state-of-the-art (such as the 20K ImageNet concepts trained with CNN) by a large margin up to 207%. We will also show that EventNet structure can help users find relevant concepts for novel event queries that cannot be well handled by conventional text based semantic analysis alone. The unique two-step approach of first applying event detection models followed by detection of event-specific concepts also provides great potential to improve the efficiency and accuracy of Event Recounting since only a very small number of event-specific concept classifiers need to be fired after event detection.
Guangnan Ye, Hongliang Xu, Dong Liu 0001, Shih-Fu Chang
ACM Multimedia1
2014 Event-Driven Semantic Concept Discovery by Exploiting Weakly Tagged Internet Images
abstract
Analysis and detection of complex events in videos require a semantic representation of the video content. Existing video semantic representation methods typically require users to pre-define an exhaustive concept lexicon and manually annotate the presence of the concepts in each video, which is infeasible for real-world video event detection problems. In this paper, we propose an automatic semantic concept discovery scheme by exploiting Internet images and their associated tags. Given a target event and its textual descriptions, we crawl a collection of images and their associated tags by performing text based image search using the noun and verb pairs extracted from the event textual descriptions. The system first identifies the candidate concepts for an event by measuring whether a tag is a meaningful word and visually detectable. Then a concept visual model is built for each candidate concept using a SVM classifier with probabilistic output. Finally, the concept models are applied to generate concept based video representations. We use the TRECVID Multimedia Event Detection (MED) 2013 as our video test set and crawl 400K Flickr images to automatically discover 2, 000 visual concepts. We show significant performance gains of the proposed concept discovery method over different video event detection tasks including supervised event modeling over concept space and semantic based zero-shot retrieval without training examples. Importantly, we show the proposed method of automatic concept discovery outperforms other well-known concept library construction approaches such as Classemes and ImageNet by a large margin (228%) in zero-shot event retrieval. Finally, subjective evaluation by humans also confirms clear superiority of the proposed method in discovering concepts for event representation.
Yin Cui, Guangnan Ye, Dong Liu 0001, Shih-Fu Chang
ICMR3
2014 Discovering joint audio-visual codewords for video event detection
I-Hong Jhuo, Guangnan Ye, Shenghua Gao, Dong Liu 0001, Yu-Gang Jiang 0001, D. T. Lee, Shih-Fu Chang
Mach. Vis. Appl.2
2013 Sample-Specific Late Fusion for Visual Category Recognition
abstract
Late fusion addresses the problem of combining the prediction scores of multiple classifiers, in which each score is predicted by a classifier trained with a specific feature. However, the existing methods generally use a fixed fusion weight for all the scores of a classifier, and thus fail to optimally determine the fusion weight for the individual samples. In this paper, we propose a sample-specific late fusion method to address this issue. Specifically, we cast the problem into an information propagation process which propagates the fusion weights learned on the labeled samples to individual unlabeled samples, while enforcing that positive samples have higher fusion scores than negative samples. In this process, we identify the optimal fusion weights for each sample and push positive samples to top positions in the fusion score rank list. We formulate our problem as a L∞norm constrained optimization problem and apply the Alternating Direction Method of Multipliers for the optimization. Extensive experiment results on various visual categorization tasks show that the proposed method consistently and significantly beats the state-of-the-art late fusion methods. To the best knowledge, this is the first method supporting sample-specific fusion weight learning.
Dong Liu 0001, Kuan-Ting Lai, Guangnan Ye, Ming-Syan Chen, Shih-Fu Chang
CVPR3
2013 Large-Scale Video Hashing via Structure Learning
abstract
Recently, learning based hashing methods have become popular for indexing large-scale media data. Hashing methods map high-dimensional features to compact binary codes that are efficient to match and robust in preserving original similarity. However, most of the existing hashing methods treat videos as a simple aggregation of independent frames and index each video through combining the indexes of frames. The structure information of videos, e.g., discriminative local visual commonality and temporal consistency, is often neglected in the design of hash functions. In this paper, we propose a supervised method that explores the structure learning techniques to design efficient hash functions. The proposed video hashing method formulates a minimization problem over a structure-regularized empirical loss. In particular, the structure regularization exploits the common local visual patterns occurring in video frames that are associated with the same semantic class, and simultaneously preserves the temporal consistency over successive frames from the same video. We show that the minimization objective can be efficiently solved by an Accelerated Proximal Gradient (APG) method. Extensive experiments on two large video benchmark datasets (up to around 150K video clips with over 12 million frames) show that the proposed method significantly outperforms the state-of-the-art hashing methods.
Guangnan Ye, Dong Liu 0001, Jun Wang 0006, Shih-Fu Chang
ICCV1
2012 Robust late fusion with rank minimization
abstract
In this paper, we propose a rank minimization method to fuse the predicted confidence scores of multiple models, each of which is obtained based on a certain kind of feature. Specifically, we convert each confidence score vector obtained from one model into a pairwise relationship matrix, in which each entry characterizes the comparative relationship of scores of two test samples. Our hypothesis is that the relative score relations are consistent among component models up to certain sparse deviations, despite the large variations that may exist in the absolute values of the raw scores. Then we formulate the score fusion problem as seeking a shared rank-2 pairwise relationship matrix based on which each original score matrix from individual model can be decomposed into the common rank-2 matrix and sparse deviation errors. A robust score vector is then extracted to fit the recovered low rank score relation matrix. We formulate the problem as a nuclear norm and ℓ1norm optimization objective function and employ the Augmented Lagrange Multiplier (ALM) method for the optimization. Our method is isotonic (i.e., scale invariant) to the numeric scales of the scores originated from different models. We experimentally show that the proposed method achieves significant performance gains on various tasks including object categorization and video event detection.
Guangnan Ye, Dong Liu 0001, I-Hong Jhuo, Shih-Fu Chang
CVPR1
2012 Weak attributes for large-scale image retrieval
abstract
Attribute-based query offers an intuitive way of image retrieval, in which users can describe the intended search targets with understandable attributes. In this paper, we develop a general and powerful framework to solve this problem by leveraging a large pool of weak attributes comprised of automatic classifier scores or other mid-level representations that can be easily acquired with little or no human labor. We extend the existing retrieval model of modeling dependency within query attributes to modeling dependency of query attributes on a large pool of weak attributes, which is more expressive and scalable. To efficiently learn such a large dependency model without overfitting, we further propose a semi-supervised graphical model to map each multiattribute query to a subset of weak attributes. Through extensive experiments over several attribute benchmarks, we demonstrate consistent and significant performance improvements over the state-of-the-art techniques. In addition, we compile the largest multi-attribute image retrieval dateset to date, including 126 fully labeled query attributes and 6,000 weak attributes of 0.26 million images.
Felix X. Yu, Rongrong Ji, Ming-Hen Tsai, Guangnan Ye, Shih-Fu Chang
CVPR4
2012 Joint audio-visual bi-modal codewords for video event detection
abstract
Joint audio-visual patterns often exist in videos and provide strong multi-modal cues for detecting multimedia events. However, conventional methods generally fuse the visual and audio information only at a superficial level, without adequately exploring deep intrinsic joint patterns. In this paper, we propose a joint audio-visual bi-modal representation, called bi-modal words. We first build a bipartite graph to model relation across the quantized words extracted from the visual and audio modalities. Partitioning over the bipartite graph is then applied to construct the bi-modal words that reveal the joint patterns across modalities. Finally, different pooling strategies are employed to re-quantize the visual and audio words into the bi-modal words and form bi-modal Bag-of-Words representations that are fed to subsequent multimedia event classifiers. We experimentally show that the proposed multi-modal feature achieves statistically significant performance gains over methods using individual visual and audio features alone and alternative multi-modal fusion methods. Moreover, we found that average pooling is the most suitable strategy for bi-modal feature generation.
Guangnan Ye, I-Hong Jhuo, Dong Liu 0001, Yu-Gang Jiang 0001, D. T. Lee, Shih-Fu Chang
ICMR1
2012 Hybrid social media network
abstract
Analysis and recommendation of multimedia information can be greatly improved if we know the interactions between the content, user, and concept, which can be easily observed from the social media networks. However, there are many heterogeneous entities and relations in such networks, making it difficult to fully represent and exploit the diverse array of information. In this paper, we develop a hybrid social media network, through which the heterogeneous entities and relations are seamlessly integrated and a joint inference procedure across the heterogeneous entities and relations can be developed. The network can be used to generate personalized information recommendation in response to specific targets of interests, e.g., personalized multimedia albums, target advertisement and friend/topic recommendation. In the proposed network, each node denotes an entity and the multiple edges between nodes characterize the diverse relations between the entities (e.g., friends, similar contents, related concepts, favorites, tags, etc). Given a query from a user indicating his/her information needs, a propagation over the hybrid social media network is employed to infer the utility scores of all the entities in the network while learning the edge selection function to activate only a sparse subset of relevant edges, such that the query information can be best propagated along the activated paths. Driven by the intuition that much redundancy exists among the diverse relations, we have developed a robust optimization framework based on several sparsity principles. We show significant performance gains of the proposed method over the state of the art in multimedia retrieval and recommendation using data crawled from social media sites. To the best of our knowledge, this is the first model supporting not only aggregation but also judicious selection of heterogeneous relations in the social media networks.
Dong Liu 0001, Guangnan Ye, Ching-Ting Chen, Shuicheng Yan, Shih-Fu Chang
ACM Multimedia2
2011 Consumer video understanding: a benchmark database and an evaluation of human and machine performance
abstract
Recognizing visual content in unconstrained videos has become a very important problem for many applications. Existing corpora for video analysis lack scale and/or content diversity, and thus limited the needed progress in this critical area. In this paper, we describe and release a new database called CCV, containing 9,317 web videos over 20 semantic categories, including events like "baseball" and "parade", scenes like "beach", and objects like "cat". The database was collected with extra care to ensure relevance to consumer interest and originality of video content without post-editing. Such videos typically have very little textual annotation and thus can benefit from the development of automatic content analysis techniques.
Yu-Gang Jiang 0001, Guangnan Ye, Shih-Fu Chang, Daniel P. W. Ellis, Alexander C. Loui
ICMR2