VLDB 2026 Research / reviewers in the wild / expert
Xiaoyong Wei
dblp:20/1707 · also Xiao-Yong Wei
· DBLP profile ↗
51ranked-venue papers
10as first author
35since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 26 · 8 first-author · 14 since 2021Artificial intelligence and machine learning · 16 · 14 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 2 first-author · 7 since 2021Databases, data management, data science and information retrieval · 6 · 5 since 2021Computer networks · 2 · 2 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multi-Agent Undercover Gaming: Hallucination Removal Through Counterfactual Test for Multimodal ReasoningabstractHallucination continues to pose a major obstacle in the reasoning capabilities of large language models (LLMs). Although the Multi-Agent Debate (MAD) paradigm offers a promising solution by promoting consensus among multiple agents to enhance reliability, it relies on the unrealistic assumption that all debaters are rational and reflective, which is a condition that may not hold when agents themselves are prone to hallucinations. To address this gap, we introduce the Multi-agent Undercover Gaming (MUG) protocol, inspired by social deduction games like ''Who is Undercover?''. MUG reframes MAD as a process of detecting ''undercover'' agents (those suffering from hallucinations) by employing multimodal counterfactual tests. Specifically, we modify reference images to introduce counterfactual evidence and observe whether agents can accurately identify these changes, providing ground-truth for identifying hallucinating agents and enabling robust, crowd-powered multimodal reasoning. MUG advances MAD protocols along three key dimensions: (1) enabling factual verification beyond statistical consensus through counterfactual testing; (2) introducing cross-evidence reasoning via dynamically modified evidence sources instead of relying on static inputs; and (3) fostering active reasoning, where agents engage in probing discussions rather than passively answering questions. Collectively, these innovations offer a more reliable and effective framework for multimodal reasoning in LLMs. Dayong Liang, Xiaoyong Wei, Changmeng Zheng |
AAAI | 2 |
| 2026 | OptScale: Probabilistic Optimality for Inference-time ScalingabstractInference-time scaling has emerged as a powerful technique for enhancing the reasoning performance of Large Language Models (LLMs). However, existing approaches often rely on heuristic strategies for parallel sampling, lacking a principled foundation. To address this gap, we propose a probabilistic framework that formalizes the optimality of inference-time scaling under the assumption that parallel samples are independently and identically distributed (i.i.d.), and where the Best-of-N selection strategy follows a probability distribution that can be estimated. Within this framework, we derive a theoretical lower bound on the required number of samples to achieve a target performance level, providing the first principled guidance for compute-efficient scaling. Leveraging this insight, we develop OptScale, a practical algorithm that dynamically determines the optimal number of sampled responses. OptScale employs a language model-based predictor to estimate probabilistic prior parameters, enabling the decision of the minimal number of samples needed that satisfy predefined performance thresholds and confidence levels. Extensive experiments on representative reasoning benchmarks (including MATH-500, GSM8K, AIME, and AMC) demonstrate that OptScale significantly reduces sampling overhead while remaining better or on par with state-of-the-art reasoning performance. Our work offers both a theoretical foundation and a practical solution for principled inference-time scaling, addressing a critical gap in the efficient deployment of LLMs for complex reasoning. Youkang Wang, Jian Wang 0054, Rubing Chen, Xiaoyong Wei |
AAAI | 4 |
| 2026 | Think in Latent Thoughts: A New Paradigm for Gloss-Free Sign Language TranslationabstractMany Sign language translation (SLT) systems quietly assume that brief chunks of signing map directly to spoken-language words.That assumption breaks down because signers often create meaning on the fly using context, space, and movement.We revisit SLT and argue that it is mainly a cross-modal reasoning task, not just a straightforward video-to-text conversion.We thus introduce a reasoning-driven SLT framework that uses an ordered sequence of latent thoughts as an explicit middle layer between the video and the generated text.These latent thoughts gradually extract and organize meaning over time.On top of this, we use a planthen-ground decoding method: the model first decides what it wants to say, and then looks back at the video to find the evidence.This separation improves coherence and faithfulness.We also built and released a new large-scale gloss-free SLT dataset with stronger context dependencies and more realistic meanings.Experiments across several benchmarks show consistent gains over existing gloss-free methods. Xiaoyong Wei, Li Qing |
ACL (1) | 3 |
| 2026 | METER: Evaluating Multi-Level Contextual Causal Reasoning in Large Language ModelsabstractPengfeng Li, Chen Huang, Chaoqun Hao, Hongyao Chen, Xiao-Yong Wei, Wenqiang Lei, See-Kiong Ng. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Pengfeng Li, Chen Huang 0006, Chaoqun Hao, Hongyao Chen, Xiaoyong Wei, Wenqiang Lei, See-Kiong Ng |
ACL (1) | 5 |
| 2026 | OsmT: Bridging Openstreetmap Queries and Natural Language With Open-Source Tag-Aware Language ModelsabstractBridging natural language and structured query languages is a long-standing challenge in the database community. While recent advances in language models have shown promise in this direction, existing solutions often rely on large-scale closed-source models that suffer from high inference costs, limited transparency, and lack of adaptability for lightweight deployment. In this paper, we present OsmT, an open-source tag-aware language model specifically designed to bridge natural language and Overpass Query Language (OverpassQL), a structured query language for accessing large-scale OpenStreetMap (OSM) data. To enhance the accuracy and structural validity of generated queries, we introduce a Tag Retrieval Augmentation (TRA) mechanism that incorporates contextually relevant tag knowledge into the generation process. This mechanism is designed to capture the hierarchical and relational dependencies present in the OSM database, addressing the topological complexity inherent in geospatial query formulation. In addition, we define a reverse task, OverpassQL-to-Text, which translates structured queries into natural language explanations to support query interpretation and improve user accessibility. We evaluate OsmT on a public benchmark against strong baselines and observe consistent improvements in both query generation and interpretation. Despite using significantly fewer parameters, our model achieves competitive accuracy, demonstrating the effectiveness of open-source pre-trained language models in bridging natural language and structured query languages within schema-rich geospatial environments. Zhuoyue Wan, Chen Zhang 0013, Yuanfeng Song, Shuaimin Li, Ruiqiang Xiao, Xiaoyong Wei, Raymond Chi-Wing Wong |
ICDE | 7 |
| 2026 | Adaptive Multi-Agent Reasoning for Text-to-Video RetrievalabstractThe rise of short-form video platforms and the emergence of multimodal large language models (MLLMs) have amplified the need for scalable, effective, zero-shot text-to-video retrieval systems. While recent advances in large-scale pretraining have improved zero-shot cross-modal alignment, existing methods still struggle with query-dependent temporal reasoning, limiting their effectiveness on complex queries involving temporal, logical, or causal relationships. To address these limitations, we propose an adaptive multi-agent retrieval framework that dynamically orchestrates specialized agents over multiple reasoning iterations based on the demands of each query. The framework includes: (1) a retrieval agent for scalable retrieval over large video corpora, (2) a reasoning agent for zero-shot contextual temporal reasoning, and (3) a query reformulation agent for refining ambiguous queries and recovering performance for those that degrade over iterations. These agents are dynamically coordinated by an orchestration agent, which leverages intermediate feedback and reasoning outcomes to guide execution. We also introduce a novel communication mechanism that incorporates retrieval-performance memory and historical reasoning traces to improve coordination and decision-making. Experiments on three TRECVid benchmarks spanning eight years show that our framework achieves a twofold improvement over CLIP4Clip and significantly outperforms state-of-the-art methods by a large margin. The code is available at https://github.com/nikkiwoo-gh/multi-agent-retrieval. Jiaxin Wu 0001, Xiaoyong Wei, Qing Li 0001 |
ICMR | 2 |
| 2026 | RelAgent: a multi-agent solution for molecular relationship groundingabstractMOTIVATION: Molecular captions, patents, and medicinal-chemistry notes describe substructures and their relations in natural language, whereas computational models operate on formal representations such as SMILES. Bridging this semantic-structural gap is important for patent interpretation, structural relationship analysis, and controllable molecular editing, yet current large language models struggle to ground textual references to precise molecular components. RESULTS: We propose RelAgent, a cooperative multi-agent framework for molecular relationship grounding. RelAgent decomposes the task into three interpretable stages: entity extraction, substructure localization, and ontology-guided relationship reasoning, and then uses verifier agents to rank structurally plausible candidates. This design supports fine-grained reasoning over molecular substructure and substantially improves performance on the MolGround benchmark. RelAgent achieves 81.4% entity-extraction F1, 56.0% exact-match localization F1, and 54.6% relationship F1 on an open-source LLaMA3.1-8B model, improving the REL F1 from 0.1% to 54.6% and exceeding the vanilla Gemini-3.1-Pro baseline in our experiments. These results indicate that agentic, structure-aware reasoning is a practical direction for interpretable molecular understanding in bioinformatics. AVAILABILITY AND IMPLEMENTATION: The source code for RelAgent is available at https://github.com/Anya-RB-Chen/RelAgent. Rubing Chen, Jiaxin Wu 0001, Chen Zhang 0013, Xiaoyong Wei |
Bioinform. | 4 |
| 2026 | Self-Paced Learning for Images of Antinuclear AntibodiesabstractAntinuclear antibody (ANA) testing is a critical method for diagnosing autoimmune disorders such as Lupus, Sjögren's syndrome, and scleroderma. Despite its importance, manual ANA detection is slow, labor-intensive, and demands years of training. ANA detection is complicated by over 100 coexisting antibody types, resulting in vast fluorescent pattern combinations. Although machine learning and deep learning have enabled automation, ANA detection in real-world clinical settings presents unique challenges as it involves multi-instance, multi-label (MIML) learning. In this paper, a novel framework for ANA detection is proposed that handles the complexities of MIML tasks using unaltered microscope images without manual preprocessing. Inspired by human labeling logic, it identifies consistent ANA sub-regions and assigns aggregated labels accordingly. These steps are implemented using three task-specific components: an instance sampler, a probabilistic pseudo-label dispatcher, and self-paced weight learning rate coefficients. The instance sampler suppresses low-confidence instances by modeling pattern confidence, while the dispatcher adaptively assigns labels based on instance distinguishability. Self-paced learning adjusts training according to empirical label observations. Our framework overcomes limitations of traditional MIML methods and supports end-to-end optimization. Extensive experiments on one ANA dataset and three public medical MIML benchmarks demonstrate the superiority of our framework. On the ANA dataset, our model achieves up to +7.0% F1-Macro and +12.6% mAP gains over the best prior method, setting new state-of-the-art results. It also ranks top-2 across all key metrics on public datasets, reducing Hamming loss and one-error by up to 18.2% and 26.9%, respectively. The source code can be accessed at https://github.com/fletcherjiang/ANA-SelfPacedLearning. Guangwu Qian, Jiaxin Wu 0001, Qing Li 0001, Yongkang Wu, Xiaoyong Wei |
IEEE Trans. Medical Imaging | 7 |
| 2025 | Removal of Hallucination on Hallucination: Debate-Augmented RAGabstractRetrieval-Augmented Generation (RAG) enhances factual accuracy by integrating external knowledge, yet it introduces a critical issue: erroneous or biased retrieval can mislead generation, compounding hallucinations, a phenomenon we term Hallucination on Hallucination.To address this, we propose Debate-Augmented RAG (DRAG), a training-free framework that integrates Multi-Agent Debate (MAD) mechanisms into both retrieval and generation stages.In retrieval, DRAG employs structured debates among proponents, opponents, and judges to refine retrieval quality and ensure factual reliability.In generation, DRAG introduces asymmetric information roles and adversarial debates, enhancing reasoning robustness and mitigating factual inconsistencies.Evaluations across multiple tasks demonstrate that DRAG improves retrieval reliability, reduces RAG-induced hallucinations, and significantly enhances overall factual accuracy.Our code is available at https://github.com/ Huenao/Debate- Wengyu Zhang, Chen Zhang 0013, Xiaoyong Wei, Qing Li 0001 |
ACL (1) | 5 |
| 2025 | ProtoD2C2: Prototype-Based Dual-Supervision, Dual-Consistency, and Dual-Contrastive Learning for OCT Fluid SegmentationabstractMixed supervision strikes a balance between performance and annotation efficiency by using a small subset of fully pixel-annotated data alongside sparsely annotated data. The key challenge lies in effectively utilizing sparse labels, as conventional pseudo-labeling relies mainly on prediction probabilities or lowlevel similarities, neglecting intra-class compactness and interclass separability. To address this issue, we proposed ProtoD2C2, a novel prototype-based framework for OCT fluid segmentation under mixed supervision, which includes three key components: (1) dual-supervision scheme that combines direct supervision with prototype-generated pseudo-labels; (2) dual-consistency at both the prediction level and prototype level; (3) hierarchical dual-contrastive learning strategy that first enforces anatomical structure priors, then promotes fine-grained discrimination among fluid subtypes. To the best of our knowledge, we are the first to introduce prototype learning for OCT fluid segmentation under mixed supervision. We evaluated our method on two public OCT fluid segmentation datasets RETOUCH and AROI, demonstrating superior performance compared to state-of-theart annotation-efficient segmentation approaches. Meixia Zhang, Qingqing Tang, Xiaoyong Wei, Haixian Zhang |
BIBM | 5 |
| 2025 | Sound Bridge: Associating Egocentric and Exocentric Videos via Audio CuesabstractUnderstanding human behavior and environmental information in egocentric videos is very challenging due to the invisibility of some actions (e.g., laughing and sneezing) and the local nature of the first-person view. Leveraging the corresponding exocentric video to provide global context has shown promising results. However, existing visual-to-visual and visual-to-textual Ego-Exo video alignment methods struggle with the issue that some activities may have non-visual overlap. To address this, we propose using sound as a bridge, as audio is often consistent across Ego-Exo videos. However, direct audio-to-audio alignment lacks context. Thus, we introduce two context-aware sound modules: one aligns audio with vision via a visual-audio cross-attention module, and another aligns text with sound closed caption generated by LLM. Experimental results on two Ego-Exo video association benchmarks show that each of the proposed modules enhances the state-of-the-art methods. Moreover, the proposed sound-aware egocentric or exocentric representation boosts the performance of downstream tasks, such as action recognition of exocentric videos and scene recognition of egocentric videos. The code and models can be accessed at https://github.com/shhuangcoder/SoundBridge. Sihong Huang, Jiaxin Wu 0001, Xiaoyong Wei, Yi Cai 0001, Dongmei Jiang, Yaowei Wang 0001 |
CVPR | 3 |
| 2025 | A Survey of WebAgents: Towards Next-Generation AI Agents for Web Automation with Large Foundation ModelsabstractWith the advancement of web techniques, they have significantly revolutionized various aspects of people's lives. Despite the importance of the web, many tasks performed on it are repetitive and time-consuming, negatively impacting the overall quality of life. To efficiently handle these tedious daily tasks, one of the most promising approaches is to advance autonomous agents to incorporate human-like intelligence based on Artificial Intelligence (AI) techniques, referred to as AI Agents. AI Agents offer significant advantages in handling such tasks since they can operate continuously without fatigue or performance degradation. Therefore, leveraging AI Agents - termed WebAgents in the context of web - to automatically assist people in handling tedious daily tasks can dramatically enhance productivity and efficiency. Recently, Large Foundation Models (LFMs) containing billions of parameters have exhibited human-like language understanding and reasoning capabilities, showing proficiency in performing various complex tasks. This naturally raises the question: 'Can LFMs be utilized to develop powerful AI Agents that automatically handle web tasks, providing significant convenience to users?' To fully explore the potential of LFMs, extensive research has emerged on WebAgents designed to complete daily web tasks according to user instructions, significantly enhancing the convenience of daily human life. In this survey, we comprehensively review existing research studies on WebAgents across three key aspects: architectures, training, and trustworthiness. Additionally, several promising directions for future research are explored to provide deeper insights. Liang-Bo Ning 0001, Ziran Liang, Zhuohang Jiang, Haohao Qu, Yujuan Ding, Wenqi Fan, Xiaoyong Wei, Shanru Lin, Hui Liu 0031, Philip S. Yu, Qing Li 0001 |
KDD (2) | 7 |
| 2025 | Interactive Video Search with Multi-modal LLM Video Captioning
Yu-Tong Cheng, Jiaxin Wu 0001, Zhixin Ma 0001, Jiangshan He, Xiaoyong Wei, Chong-Wah Ngo |
MMM (5) | 5 |
| 2025 | GraphATC: advancing multilevel and multi-label anatomical therapeutic chemical classification via atom-level graph learningabstractThe accurate categorization of compounds within the anatomical therapeutic chemical (ATC) system is fundamental for drug development and fundamental research. Although this area has garnered significant research focus for over a decade, the majority of prior studies have concentrated solely on the Level 1 labels defined by the World Health Organization (WHO), neglecting the labels of the remaining four levels. This narrow focus fails to address the true nature of the task as a multilevel, multi-label classification challenge. Moreover, existing benchmarks like Chen-2012 and ATC-SMILES have become outdated, lacking the incorporation of new drugs or updated properties of existing ones that have emerged in recent years and have been integrated into the WHO ATC system. To tackle these shortcomings, we present a comprehensive approach in this paper. Firstly, we systematically cleanse and enhance the drug dataset, expanding it to encompass all five levels through a rigorous cross-resource validation process involving KEGG, PubChem, ChEMBL, ChemSpider, and ChemicalBook. This effort culminates in the creation of a novel benchmark termed ATC-GRAPH. Secondly, we extend the classification task to encompass Level 2 and introduce graph-based learning techniques to provide more accurate representations of drug molecular structures. This approach not only facilitates the modeling of Polymers, Macromolecules, and Multi-Component drugs more precisely but also enhances the overall fidelity of the classification process. The efficacy of our proposed framework is validated through extensive experiments, establishing a new state-of-the-art methodology. To facilitate the replication of this study, we have made the benchmark dataset, source code, and web server openly accessible. Wengyu Zhang, Qi Tian 0001, Wenqi Fan, Dongmei Jiang, Yaowei Wang 0001, Qing Li 0001, Xiaoyong Wei |
Briefings Bioinform. | 8 |
| 2025 | Hierarchical and Heterogeneous Federated Learning via a Learning-on-Model ParadigmabstractFederated Learning (FL) collaboratively trains a shared global model without exposing clients' private data. In practical FL systems, clients (e.g., smartphones and wearables) typically have disparate system resources. Traditional FL, however, adopts a one-size-fits-all solution, where a homogeneous large model is sent to and trained on each client. This method results in an overwhelming workload for less capable clients and starvation for others. To tackle this, we proposeFedConv, a client-friendly FL framework, minimizing the system overhead on resource-constrained clients by providing heterogeneous customized sub-models.FedConvfeatures a novellearning-on-modelparadigm that learns the parameters of heterogeneous sub-models viaconvolutional compression. To aggregate heterogeneous sub-models, we proposetransposed convolutional dilationto convert them back to large models with a unified size while retaining personalized information. The compression and dilation processes, transparent to clients, are tuned on the server using a small public dataset. We further propose ahierarchical and clustering-based local trainingstrategy for enhanced performance. Extensive experiments on six datasets show thatFedConvoutperforms state-of-the-art FL systems in terms of model accuracy (by more than 35% on average), computation and communication overhead (with 33% and 25% reduction, respectively). Leming Shen, Qiang Yang 0018, Kaiyan Cui, Yuanqing Zheng, Xiaoyong Wei, Jianwei Liu 0008, Jinsong Han |
IEEE Trans. Mob. Comput. | 5 |
| 2025 | Multi-Agent Attacks for Black-Box Social RecommendationsabstractThe rise of online social networks has facilitated the evolution of social recommender systems, which incorporate social relations to enhance users’ decision-making process. With the great success of Graph Neural Networks (GNNs) in learning node representations, GNN-based social recommendations have been widely studied to model user-item interactions and user-user social relations simultaneously. Despite their great successes, recent studies have shown that these advanced recommender systems are highly vulnerable to adversarial attacks, in which attackers can inject well-designed fake user profiles to disrupt recommendation performances. While most existing studies mainly focus on targeted attacks to promote target items on vanilla recommender systems, untargeted attacks to degrade the overall prediction performance are less explored on social recommendations under a black-box scenario. To perform untargeted attacks on social recommender systems, attackers can construct malicious social relationships for fake users to enhance the attack performance. However, the coordination of social relations and item profiles is challenging for attacking black-box social recommendations. To address this limitation, we first conduct several preliminary studies to demonstrate the effectiveness of cross-community connections and cold-start items in degrading recommendations performance. Specifically, we propose a novel framework MultiAttack based on multi-agent reinforcement learning to coordinate the generation of cold-start item profiles and cross-community social relations for conducting untargeted attacks on black-box social recommendations. Comprehensive experiments on various real-world datasets demonstrate the effectiveness of our proposed attacking framework under the black-box setting. Shijie Wang 0002, Wenqi Fan, Xiaoyong Wei, Xiaowei Mei, Shanru Lin, Qing Li 0001 |
ACM Trans. Inf. Syst. | 3 |
| 2024 | Compositional Inversion for Stable Diffusion ModelsabstractInversion methods, such as Textual Inversion, generate personalized images by incorporating concepts of interest provided by user images. However, existing methods often suffer from overfitting issues, where the dominant presence of inverted concepts leads to the absence of other desired concepts. It stems from the fact that during inversion, the irrelevant semantics in the user images are also encoded, forcing the inverted concepts to occupy locations far from the core distribution in the embedding space. To address this issue, we propose a method that guides the inversion process towards the core distribution for compositional embeddings. Additionally, we introduce a spatial regularization approach to balance the attention on the concepts being composed. Our method is designed as a post-training approach and can be seamlessly integrated with other inversion methods. Experimental results demonstrate the effectiveness of our proposed approach in mitigating the overfitting problem and generating more diverse and balanced compositions of concepts in the synthesized images. The source code is available at https://github.com/zhangxulu1996/Compositional-Inversion. Xulu Zhang, Xiaoyong Wei, Jinlin Wu, Zhaoxiang Zhang 0001, Zhen Lei 0001, Qing Li 0001 |
AAAI | 2 |
| 2024 | Instruct Once, Chat Consistently in Multiple Rounds: An Efficient Tuning Framework for DialogueabstractTuning language models for dialogue generation has been a prevalent paradigm for building capable dialogue agents.Yet, traditional tuning narrowly views dialogue generation as resembling other language generation tasks, ignoring the role disparities between two speakers and the multi-round interactive process that dialogues ought to be.Such a manner often leads to unsatisfactory chat consistency for the built agent.In this work, we emphasize the interactive, communicative nature of dialogue and argue that it is more feasible to model the speaker roles of agent and user separately, enabling the agent to adhere to its role consistently.With this in mind, we propose an efficient Multi-round Interactive Dialogue Tuning (MIDI-Tuning) framework 1 .It models the agent and user individually with two adapters built upon large language models.The adapters make use of respective utterances round by round in alternating order and they are tuned via a round-level memory caching mechanism.Extensive experiments demonstrate that, our framework performs superior to traditional finetuning and harbors the tremendous potential for improving dialogue consistency. Jian Wang 0054, Chak Tou Leong, Jiashuo Wang, Dongding Lin, Wenjie Li 0002, Xiaoyong Wei |
ACL (1) | 6 |
| 2024 | Prior Knowledge Integration via LLM Encoding and Pseudo Event Regulation for Video Moment RetrievalabstractIn this paper, we explore the use of large language models (LLMs) to enhance video moment retrieval (VMR) by integrating general knowledge and pseudo-events as priors. We address the limitations of LLMs in generating continuous outputs, such as salience scores and inter-frame embeddings, which are critical for capturing inter-frame relations. To address these limitations, we propose using LLM encoders, which refine inter-concept relations in multimodal embeddings effectively, even without textual training. Our feasibility study shows that this capability extends to other embeddings like BLIP and T5 when they exhibit similar patterns to CLIP embeddings. We present a general framework for integrating LLM encoders into existing VMR architectures, specifically within the fusion module. The LLM encoder's ability to refine concept relation can help the model to achieve a balanced understanding of the foreground concepts (e.g., persons, faces) and background concepts (e.g., street, mountains) rather focusing only on the visually dominant foreground concepts. Additionally, we utilize pseudo-events, identified via event detection, to guide accurate moment prediction within event boundaries, reducing distractions from adjacent moments. Our plug-in approach for semantic refinement and pseudo-event regulation demonstrates state-of-the-art VMR performance through experimental validation. The source code can be accessed at https://github.com/fletcherjiang/LLMEPET. Wengyu Zhang, Xulu Zhang, Xiaoyong Wei, Chang Wen Chen, Qing Li 0001 |
ACM Multimedia | 4 |
| 2024 | Generative Active Learning for Image Synthesis PersonalizationabstractThis paper presents a pilot study that explores the application of active learning, traditionally studied in the context of discriminative models, to generative models. We specifically focus on image synthesis personalization tasks. The primary challenge in conducting active learning on generative models lies in the open-ended nature of querying, which differs from the closed form of querying in discriminative models that typically target a single concept. We introduce the concept of anchor directions to transform the querying process into a semi-open problem. We propose a direction-based uncertainty sampling strategy to enable generative active learning and tackle the exploitation-exploration dilemma. Extensive experiments are conducted to validate the effectiveness of our approach, demonstrating that an open-source model can achieve superior performance compared to closed-source models developed by large companies, such as Google's StyleDrop. The source code is available at https://github.com/zhangxulu1996/GAL4Personalization. Xulu Zhang, Wengyu Zhang, Xiaoyong Wei, Jinlin Wu, Zhaoxiang Zhang 0001, Zhen Lei 0001, Qing Li 0001 |
ACM Multimedia | 3 |
| 2024 | A Picture Is Worth a Graph: A Blueprint Debate Paradigm for Multimodal ReasoningabstractThis paper presents a pilot study aimed at introducing multi-agent debate into multimodal reasoning. The study addresses two key challenges: the trivialization of opinions resulting from excessive summarization and the diversion of focus caused by distractor concepts introduced from images. These challenges stem from the inductive (bottom-up) nature of existing debating schemes. To address the issue, we propose a deductive (top-down) debating approach called Blueprint Debate on Graphs (BDoG). In BDoG, debates are confined to a blueprint graph to prevent opinion trivialization through world-level summarization. Moreover, by storing evidence in branches within the graph, BDoG mitigates distractions caused by frequent but irrelevant concepts. Extensive experiments validate that BDoG is able to achieve state-of-the-art results in ScienceQA and MMBench with significant improvements over previous methods. The source code can be accessed at https://github.com/thecharm/BDoG. Changmeng Zheng, Dayong Liang, Wengyu Zhang, Xiaoyong Wei, Tat-Seng Chua, Qing Li 0001 |
ACM Multimedia | 4 |
| 2024 | FedConv: A Learning-on-Model Paradigm for Heterogeneous Federated ClientsabstractFederated Learning (FL) facilitates collaborative training of a shared global model without exposing clients' private data. In practical FL systems, clients (e.g., edge servers, smartphones, and wearables) typically have disparate system resources. Conventional FL, however, adopts a one-size-fits-all solution, where a homogeneous large global model is transmitted to and trained on each client, resulting in an overwhelming workload for less capable clients and starvation for other clients. To address this issue, we propose FedConv, a client-friendly FL framework, which minimizes the computation and memory burden on resource-constrained clients by providing heterogeneous customized sub-models. FedConv features a novel learning-on-model paradigm that learns the parameters of the heterogeneous sub-models via convolutional compression. Unlike traditional compression methods, the compressed models in FedConv can be directly trained on clients without decompression. To aggregate the heterogeneous sub-models, we propose transposed convolutional dilation to convert them back to large models with a unified size while retaining personalized information from clients. The compression and dilation processes, transparent to clients, are optimized on the server leveraging a small public dataset. Extensive experiments on six datasets demonstrate that FedConv outperforms state-of-the-art FL systems in terms of model accuracy (by more than 35% on average), computation and communication overhead (with 33% and 25% reduction, respectively). Leming Shen, Qiang Yang 0018, Kaiyan Cui, Yuanqing Zheng, Xiaoyong Wei, Jianwei Liu 0008, Jinsong Han |
MobiSys | 5 |
| 2024 | Towards Bridged Vision and Language: Learning Cross-Modal Knowledge Representation for Relation ExtractionabstractIn natural language processing, relation extraction (RE) is to detect and classify the semantic relationship of two given entities within a sentence. Previous RE methods consider only the textual contents and suffer performance decline in social media when texts lack contexts. Incorporating text-related visual information can supplement the missing semantics for relation extraction in social media posts. However, textual relations are usually abstract and of high-level semantics, which causes the semantic gap between visual contents and textual expressions. In this paper, we propose RECK - a neural network for relation extraction with cross-modal knowledge representations. Different from previous multimodal methods training a common subspace for all modalities, we bridge the semantic gaps by explicitly selecting knowledge paths from external knowledge through the cross-modal object-entity pairs. We further extend the paths into a knowledge graph, and adopt a graph attention network to capture the multi-grained relevant concepts which can provide higher level and key semantics information from external knowledge. Besides, we employ a cross-modal attention mechanism to align and fuse the multimodal information. Experimental results on a multimodal RE dataset show that our model achieves new state-of-the-art performance with knowledge evidence. Guohua Wang 0003, Changmeng Zheng, Yi Cai 0001, Ze Fu, Yaowei Wang 0001, Xiaoyong Wei, Qing Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2024 | Empowering Molecule Discovery for Molecule-Caption Translation With Large Language Models: A ChatGPT PerspectiveabstractMolecule discovery plays a crucial role in various scientific fields, advancing the design of tailored materials and drugs, which contributes to the development of society and human well-being. Specifically, molecule-caption translation is an important task for molecule discovery, aligning human understanding with molecular space. However, most of the existing methods heavily rely on domain experts, require excessive computational cost, or suffer from sub-optimal performance. On the other hand, Large Language Models (LLMs), like ChatGPT, have shown remarkable performance in various cross-modal tasks due to their powerful capabilities in natural language understanding, generalization, and in-context learning (ICL), which provides unprecedented opportunities to advance molecule discovery. Despite several previous works trying to apply LLMs in this task, the lack of domain-specific corpus and difficulties in training specialized LLMs still remain challenges. In this work, we propose a novel LLM-based framework (MolReGPT) for molecule-caption translation, where an In-Context Few-Shot Molecule Learning paradigm is introduced to empower molecule discovery with LLMs like ChatGPT to perform their in-context learning capability without domain-specific pre-training and fine-tuning. MolReGPT leverages the principle of molecular similarity to retrieve similar molecules and their text descriptions from a local database to enable LLMs to learn the task knowledge from context examples. We evaluate the effectiveness of MolReGPT on molecule-caption translation, including molecule understanding and text-based molecule generation. Experimental results show that compared to fine-tuned models, MolReGPT outperforms MolT5-base and is comparable to MolT5-large without additional training. To the best of our knowledge, MolReGPT is the first work to leverage LLMs via in-context learning in molecule-caption translation for advancing molecule discovery. Our work expands the scope of LLM applications, as well as providing a new paradigm for molecule discovery and design. Jiatong Li 0003, Wenqi Fan, Xiaoyong Wei, Hui Liu 0031, Jiliang Tang, Qing Li 0001 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2023 | Rethinking Multimodal Entity and Relation Extraction from a Translation Point of ViewabstractWe revisit the multimodal entity and relation extraction from a translation point of view.Special attention is paid on the misalignment issue in text-image datasets which may mislead the learning.We are motivated by the fact that the cross-modal misalignment is a similar problem of cross-lingual divergence issue in machine translation.The problem can then be transformed and existing solutions can be borrowed by treating a text and its paired image as the translation to each other.We implement a multimodal back-translation using diffusionbased generative models for pseudo-paralleled pairs and a divergence estimator by constructing a high-resource corpora as a bridge for low-resource learners.Fine-grained confidence scores are generated to indicate both types and degrees of alignments with which better representations are obtained.The method has been validated in the experiments by outperforming 14 state-of-the-art methods in both entity and relation extraction tasks.The source code is available at https://github.com/thecharm/TMR. Changmeng Zheng, Yi Cai 0001, Xiaoyong Wei, Qing Li 0001 |
ACL (1) | 4 |
| 2023 | Open-Scenario Domain Adaptive Object Detection in Autonomous DrivingabstractExisting domain adaptive object detection algorithms (DAOD) have demonstrated their effectiveness in discriminating and localizing objects across scenarios. However, these algorithms typically assume a single source and target domain for adaptation, which is not representative of the more complex data distributions in practice. To address this issue, we propose a novel Open-Scenario Domain Adaptive Object Detection (OSDA), which leverages multiple source and target domains for more practical and effective domain adaptation. We are the first to increase the granularity of the background category by building the foundation model using contrastive vision-language pre-training in an open-scenario setting for better distinguishing foreground and background, which is under-explored in previous studies. The performance gains by introducing the pre-training have been observed and have validated the model's ability to detect objects across domains. To further fine-tune the model for domain-specific object detection, we propose a hierarchical feature alignment strategy to obtain a better common feature space among the various source and target domains. In the case of multi-source domains, the cross-reconstruction framework is introduced for learning more domain invariances. The proposed method is able to alleviate knowledge forgetting without any additional computational costs. Extensive experiments across different scenarios demonstrate the effectiveness of the proposed model. Zeyu Ma 0002, Ziqiang Zheng, Jiwei Wei, Xiaoyong Wei, Yang Yang 0002, Heng Tao Shen |
ACM Multimedia | 4 |
| 2023 | Entity-Graph Enhanced Cross-Modal Pretraining for Instance-Level Product RetrievalabstractOur goal in this research is to study a more realistic environment in which we can conduct weakly-supervised multi-modal instance-level product retrieval for fine-grained product categories. We first contribute the Product1M datasets and define two real practical instance-level retrieval tasks that enable evaluations on price comparison and personalized recommendations. For both instance-level tasks, accurately identifying the intended product target mentioned in visual-linguistic data and mitigating the impact of irrelevant content are quite challenging. To address this, we devise a more effective cross-modal pretraining model capable of adaptively incorporating key concept information from multi-modal data. This is accomplished by utilizing an entity graph, where nodes represented entities and edges denoted the similarity relations between them. Specifically, a novel Entity-Graph Enhanced Cross-Modal Pretraining (EGE-CMP) model is proposed for instance-level commodity retrieval, which explicitly injects entity knowledge in both node-based and subgraph-based ways into the multi-modal networks via a self-supervised hybrid-stream transformer. This could reduce the confusion between different object contents, thereby effectively guiding the network to focus on entities with real semantics. Experimental results sufficiently verify the efficacy and generalizability of our EGE-CMP, outperforming several SOTA cross-modal baselines like CLIP Radford et al. 2021, UNITER Chen et al. 2020 and CAPTURE Zhan et al. 2021. Xunlin Zhan, Yunchao Wei, Xiaoyong Wei, Yaowei Wang 0001, Minlong Lu, Xiaochun Cao, Xiaodan Liang |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Region Attentive Action Unit Intensity Estimation With Uncertainty Weighted Multi-Task LearningabstractFacial action units (AUs) refer to a comprehensive set of atomic facial muscle movements. Recent works have focused on exploring complementary information by learning the relationships among AUs. Most existing approaches process AU co-occurrence and enhance AU recognition by learning the dependencies among AUs from labels, however, the complementary information among features of different AUs are ignored. Moreover, ground truth annotations suffer from a large intra-class variance and their associated intensity levels may vary depending on the annotators’ experience. In this paper, we propose the Region Attentive AU intensity estimation method with Uncertainty Weighted Multi-task Learning (RA-UWML). A RoI-Net is first used to extract features from the pre-defined facial patches where the AUs locate. Then, we use the co-occurrence of AUs using both within patch and between patches representation learning. Within a given patch, we propose sharing representation learning in a multi-task manner. To achieve complementarity and avoid redundancy between different image patches, we propose to use a multi-head self-attention mechanism to adaptively and attentively encode each patch specific representation. Moreover, the AU intensity is represented as a Gaussian distribution, instead of a single value, where the mean value indicates the most likely AU intensity and the variance indicates the uncertainty of the estimated AU intensity. The estimated variances are leveraged to automatically weight the loss of each AU in the multitask learning model. In extensive experiments on the Disfa, Fera2015 and Feafa benchmarks, it is shown that the proposed AU intensity estimation model achieves better results compared to the state-of-the-art models. Dongmei Jiang, Xiaoyong Wei, Ke Lu 0002, Hichem Sahli |
IEEE Trans. Affect. Comput. | 4 |
| 2023 | DRAKE: Deep Pair-Wise Relation Alignment for Knowledge-Enhanced Multimodal Scene Graph Generation in Social Media PostsabstractScene Graph Generation (SGG) is a typical computer vision task that detects objects and corresponding predicates in an image. Existing SGG methods focus on modeling visual contexts to generate scene graphs and are conducted on well-annotated datasets with high-quality images. However, the quality is unguaranteed for images in social media posts, so that some images may be incomplete or occluded by some obstacles, hence might not provide sufficient visual context for SGG. Therefore, previous methods might result in missing or false visual relationship detection due to lacking visual contexts. To effectively generate the scene graphs in social media, we study multimodal scene graph generation (MSG) in this paper. MSG aims to develop visual scene graphs from images in social media posts with the support of text sentences. However, leveraging textual contents by simple multimodal alignment such as object-level alignment neglects the inherent pair-wise mapping between multimodal object pairs. To address the limitations, we propose a method named Deep pair-wise Relation Alignment for Knowledge-Enhanced (DRAKE) multimodal scene graph generation. The model supplements the missing visual contexts with well-aligned textual knowledge. It first represents the textual information into object-aware knowledge representation with the help of vision data. Furthermore, our proposed DRAKE facilitates the interaction of the info between multimodal pair-wise representations. A multimodal context enhancement layer can be devised to help the model generate the scene graph. To evaluate the model performance of SGG on social media images, we propose a social media SGG dataset called MSG. We comprehensively analyze the effectiveness of our proposed method on the MSG dataset. The experimental results on the MSG dataset indicate that our model outperforms the previous methods. To fairly compare our method with other SGG models, we also conduct experiments on the Visual Genome dataset for more analysis The MSG dataset is released onhttps://github.com/FuZe4ever/MSG. Ze Fu, Changmeng Zheng, Yi Cai 0001, Xiaoyong Wei, Yaowei Wang 0001, Qing Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | M5Product: Self-harmonized Contrastive Learning for E-commercial Multi-modal PretrainingabstractDespite the potential of multi-modal pre-training to learn highly discriminative feature representations from complementary data modalities, current progress is being slowed by the lack of large-scale modality-diverse datasets. By leveraging the natural suitability of E-commerce, where different modalities capture complementary semantic information, we contribute a large-scale multi-modal pretraining dataset M5Product. The dataset comprises 5 modalities (image, text, table, video, and audio), covers over 6,000 categories and 5,000 attributes, and is 500× larger than the largest publicly available dataset with a similar number of modalities. Furthermore, M5Product contains incomplete modality pairs and noise while also having a long-tailed distribution, resembling most real-world problems. We further propose Self-harmonized ContrAstive LEarning (SCALE), a novel pretraining framework that integrates the different modalities into a unified model through an adaptive feature fusion mechanism, where the importance of each modality is learned directly from the modality embeddings and impacts the inter-modality contrastive learning and masked tasks within a multi-modal transformer model. We evaluate the current multi-modal pre-training state-of-the-art approaches and benchmark their ability to learn from unlabeled data when faced with the large number of modalities in the M5Product dataset. We conduct extensive experiments on four downstream tasks and demonstrate the superiority of our SCALE model, providing insights into the importance of dataset scale and diversity. Dataset and codes are available at11https://xiaodongsuper.github.io/M5Product_dataset/. Xunlin Zhan, Yangxin Wu, Yunchao Wei, Michael Kampffmeyer, Xiaoyong Wei, Minlong Lu, Yaowei Wang 0001, Xiaodan Liang |
CVPR | 6 |
| 2022 | Identifying the kind behind SMILES - anatomical therapeutic chemical classification using structure-only representationsabstractAnatomical Therapeutic Chemical (ATC) classification for compounds/drugs plays an important role in drug development and basic research. However, previous methods depend on interactions extracted from STITCH dataset which may make it depend on lab experiments. We present a pilot study to explore the possibility of conducting the ATC prediction solely based on the molecular structures. The motivation is to eliminate the reliance on the costly lab experiments so that the characteristics of a drug can be pre-assessed for better decision-making and effort-saving before the actual development. To this end, we construct a new benchmark consisting of 4545 compounds which is with larger scale than the one used in previous study. A light-weight prediction model is proposed. The model is with better explainability in the sense that it is consists of a straightforward tokenization that extracts and embeds statistically and physicochemically meaningful tokens, and a deep network backed by a set of pyramid kernels to capture multi-resolution chemical structural characteristics. Its efficacy has been validated in the experiments where it outperforms the state-of-the-art methods by 15.53% in accuracy and by 69.66% in terms of efficiency. We make the benchmark dataset, source code and web server open to ease the reproduction of this study. Zhen-Qun Yang, Xulu Zhang, Wenqi Fan, Yaowei Wang 0001, Qing Li 0001, Xiaoyong Wei |
Briefings Bioinform. | 9 |
| 2022 | Deep learning-based person re-identification methods: A survey and outlook of recent works
Zhangqiang Ming, Min Zhu 0005, Xiangkun Wang, Jiamin Zhu, Junlong Cheng, Chengrui Gao, Xiaoyong Wei |
Image Vis. Comput. | 8 |
| 2022 | A multi-scale multi-attention network for dynamic facial expression recognition
Xiaohan Xia, Le Yang 0009, Xiaoyong Wei, Hichem Sahli, Dongmei Jiang |
Multim. Syst. | 3 |
| 2021 | MDA-GCNFTG: identifying miRNA-disease associations based on graph convolutional networks via graph sampling through the feature and topology graphabstractAccurate identification of the miRNA-disease associations (MDAs) helps to understand the etiology and mechanisms of various diseases. However, the experimental methods are costly and time-consuming. Thus, it is urgent to develop computational methods towards the prediction of MDAs. Based on the graph theory, the MDA prediction is regarded as a node classification task in the present study. To solve this task, we propose a novel method MDA-GCNFTG, which predicts MDAs based on Graph Convolutional Networks (GCNs) via graph sampling through the Feature and Topology Graph to improve the training efficiency and accuracy. This method models both the potential connections of feature space and the structural relationships of MDA data. The nodes of the graphs are represented by the disease semantic similarity, miRNA functional similarity and Gaussian interaction profile kernel similarity. Moreover, we considered six tasks simultaneously on the MDA prediction problem at the first time, which ensure that under both balanced and unbalanced sample distribution, MDA-GCNFTG can predict not only new MDAs but also new diseases without known related miRNAs and new miRNAs without known related diseases. The results of 5-fold cross-validation show that the MDA-GCNFTG method has achieved satisfactory performance on all six tasks and is significantly superior to the classic machine learning methods and the state-of-the-art MDA prediction methods. Moreover, the effectiveness of GCNs via the graph sampling strategy and the feature and topology graph in MDA-GCNFTG has also been demonstrated. More importantly, case studies for two diseases and three miRNAs are conducted and achieved satisfactory performance. Yanyi Chu, Xuhong Wang, Qiuying Dai, Yanjing Wang 0003, Shaoliang Peng, Xiaoyong Wei, Jingfei Qiu, Dennis R. Salahub, Yi Xiong 0002 |
Briefings Bioinform. | 7 |
| 2021 | Deep Collocative Learning for Immunofixation Electrophoresis Image AnalysisabstractImmunofixation Electrophoresis (IFE) analysis is of great importance to the diagnosis of Multiple Myeloma, which is among the top-9 cancer killers in the United States, but has rarely been studied in the context of deep learning. Two possible reasons are: 1) the recognition of IFE patterns is dependent on the co-location of bands that forms a binary relation, different from the unary relation (visual features to label) that deep learning is good at modeling; 2) deep classification models may perform with high accuracy for IFE recognition but is not able to provide firm evidence (where the co-location patterns are) for its predictions, rendering difficulty for technicians to validate the results. We propose to address these issues with collocative learning, in which a collocative tensor has been constructed to transform the binary relations into unary relations that are compatible with conventional deep networks, and a location-label-free method that utilizes the Grad-CAM saliency map for evidence backtracking has been proposed for accurate localization. In addition, we have proposed Coached Attention Gates that can regulate the inference of the learning to be more consistent with human logic and thus support the evidence backtracking. The experimental results show that the proposed method has obtained a performance gain over its base model ResNet18 by 741.30% in IoU and also outperformed popular deep networks of DenseNet, CBAM, and Inception-v3. Xiaoyong Wei, Zhen-Qun Yang, Xulu Zhang, Ga Liao, Ailin Sheng, Shaohua Kevin Zhou, Yongkang Wu |
IEEE Trans. Medical Imaging | 1 |
| 2020 | Multi-View Weighted Feature Fusion Using CNN for Pneumonia Detection on Chest X-RaysabstractChest X-ray is still the most common and important method for diagnosing pneumonia. However, the analysis of chest radiographs requires professional radiologists, and overreliance on radiologists may lead to erroneous diagnosis or missed diagnosis. Using convolutional neural networks(CNNs) for diagnosis chest diseases on chest X-ray has achieved better results, but most of the previous models are only trained by frontal-view X-ray images. Unique from past studies, in this paper, we proposed a model that can learn multi-view semantic information from chest X-rays to detect pneumonia. Our model includes two stages of feature extraction and feature fusion, and is trained on MIMIC-CXR-JPG dataset, currently the largest publicly available chest x-ray dataset, containing 377,110 JPG format images. We demonstrate that such multi-view weighted feature fusion model outperforms the models that use features only from one view. Our results are better than previous models for pneumonia detection. Shaoliang Peng, Xiongjun Zhao, Xiaoyong Wei, Donqing Wei, Yuehua Peng |
HealthCom | 3 |
| 2020 | Exploring Entity-Level Spatial Relationships for Image-Text MatchingabstractExploring the entity-level (i.e., objects in an image, words in a text) spatial relationship contributes to understanding multimedia content precisely. The ignorance of spatial information in previous works probably leads to misunderstandings of image contents. For instance, sentences `Boats are on the water' and `Boats are under the water' describe the same objects, but correspond to different sceneries. To this end, we utilize the relative position of objects to capture entity-level spatial relationships for image-text matching. Specifically, we fuse semantic and spatial relationships of image objects in a visual intra-modal relation module. The module performs promisingly to understand image contents and improve object representation learning. It contributes to capturing entity-level latent correspondence of image-text pairs. Then the query (text) plays a role of textual context to refine the interpretable alignments of image-text pairs in the inter-modal relation module. Our proposed method achieves state-of-the-art results on MSCOCO and Flickr30K datasets. Yaxian Xia, Lun Huang, Wenmin Wang 0001, Xiaoyong Wei, Jie Chen 0001 |
ICASSP | 4 |
| 2019 | Attention on Attention for Image CaptioningabstractAttention mechanisms are widely used in current encoder/decoder frameworks of image captioning, where a weighted average on encoded vectors is generated at each time step to guide the caption decoding process. However, the decoder has little idea of whether or how well the attended vector and the given attention query are related, which could make the decoder give misled results. In this paper, we propose an Attention on Attention (AoA) module, which extends the conventional attention mechanisms to determine the relevance between attention results and queries. AoA first generates an information vector and an attention gate using the attention result and the current context, then adds another attention by applying element-wise multiplication to them and finally obtains the attended information, the expected useful knowledge. We apply AoA to both the encoder and the decoder of our image captioning model, which we name as AoA Network (AoANet). Experiments show that AoANet outperforms all previously published methods and achieves a new state-of-the-art performance of 129.8 CIDEr-D score on MS COCO Karpathy offline test split and 129.6 CIDEr-D (C40) score on the official online testing server. Code is available at https://github.com/husthuaan/AoANet. Lun Huang, Wenmin Wang 0001, Jie Chen 0001, Xiaoyong Wei |
ICCV | 4 |
| 2017 | Contextual Noise Reduction for Domain Adaptive Near-Duplicate Retrieval on Merchandize ImagesabstractIn this paper, we have proposed a novel method which utilizes the contextual relationship among visual words for reducing the Quantization errors in near-duplicate image retrieval (NDR). Instead of following the track of conventional NDR techniques which usually search new solutions by borrowing ideas from the text domain, we propose to model the problem back to image domain, which results in a more natural way of solution search. The idea of the proposed method is to construct a context graph that encapsulates the contextual relationship within an image and treat the graph as a pseudo-image, so that classical image filters can be adopted to reduce the mismapped visual words which are contextually inconsistent with others.With these contextual noises reduced, the method provides purified inputs to the subsequent processes in NDR, and improves the overall accuracy. More importantly, the purification further increases the sparsity of the image feature vectors, which thus speeds up the conventional methods by 1662% times and makes NDR practical to online applications on merchandize images where the requirement of response time is critical. The way of considering contextual noise reduction in image domain also makes the problem open to all sophisticated filters. Our study shows the classic anisotropic diffusion filter can be employed to address the cross-domain issue, resulting in the superiority of the method to conventional ones in both effectiveness and efficiency. Zhen-Qun Yang, Xiaoyong Wei, Zhang Yi 0001, Gerald Friedland |
IEEE Trans. Image Process. | 2 |
| 2014 | Collaborative error reduction for hierarchical classification
Shiai Zhu, Xiaoyong Wei, Chong-Wah Ngo |
Comput. Vis. Image Underst. | 2 |
| 2014 | Visual Typo Correction by Collocative Optimization: A Case Study on Merchandize ImagesabstractNear-duplicate retrieval (NDR) in merchandize images is of great importance to a lot of online applications on e-Commerce websites. In those applications where the requirement of response time is critical, however, the conventional techniques developed for a general purpose NDR are limited, because expensive post-processing like spatial verification or hashing is usually employed to compromise the quantization errors among the visual words used for the images. In this paper, we argue that most of the errors are introduced because of the quantization process where the visual words are considered individually, which has ignored the contextual relations among words. We propose a "spelling or phrase correction" like process for NDR, which extends the concept of collocations to visual domain for modeling the contextual relations. Binary quadratic programming is used to enforce the contextual consistency of words selected for an image, so that the errors (typos) are eliminated and the quality of the quantization process is improved. The experimental results show that the proposed method can improve the efficiency of NDR by reducing vocabulary size by 1000% times, and under the scenario of merchandize image NDR, the expensive local interest point feature used in conventional approaches can be replaced by color-moment feature, which reduces the time cost by 9202% while maintaining comparable performance to the state-of-the-art methods. Xiaoyong Wei, Zhen-Qun Yang, Chong-Wah Ngo, Wei Zhang 0031 |
IEEE Trans. Image Process. | 1 |
| 2013 | Error recovered hierarchical classificationabstractHierarchical classification (HC) is a popular and efficient way for detecting the semantic concepts from the images. However, the conventional HC, which always selects the branch with the highest classification response to go on, has the risk of propagating serious errors from higher levels of the hierarchy to the lower levels. We argue that the highest-response-first strategy is too arbitrary, because the candidate nodes are considered individually which ignores the semantic relationship among them. In this paper, we propose a novel method for HC, which is able to utilize the semantic relationship among candidate nodes and their children to recover the responses of unreliable classifiers of the candidate nodes, with the hope of providing the branch selection a more globally valid and semantically consistent view. The experimental results show that the proposed method outperforms the conventional HC methods and achieves a satisfactory balance between the accuracy and efficiency. Shiai Zhu, Xiaoyong Wei, Chong-Wah Ngo |
ACM Multimedia | 2 |
| 2013 | Free-gram phrase identification for modeling Chinese text
Xi Peng 0001, Zhang Yi 0001, Xiaoyong Wei, Dezhong Peng, Yongsheng Sang |
Inf. Process. Lett. | 3 |
| 2013 | Coaching the Exploration and Exploitation in Active Learning for Interactive Video RetrievalabstractConventional active learning approaches for interactive video/image retrieval usually assume the query distribution is unknown, as it is difficult to estimate with only a limited number of labeled instances available. Thus, it is easy to put the system in a dilemma whether to explore the feature space in uncertain areas for a better understanding of the query distribution or to harvest in certain areas for more relevant instances. In this paper, we propose a novel approach called coached active learning that makes the query distribution predictable through training and, therefore, avoids the risk of searching on a completely unknown space. The estimated distribution, which provides a more global view of the feature space, can be used to schedule not only the timing but also the step sizes of the exploration and the exploitation in a principled way. The results of the experiments on a large-scale data set from TRECVID 2005-2009 validate the efficiency and effectiveness of our approach, which demonstrates an encouraging performance when facing domain-shift, outperforms eight conventional active learning methods, and shows superiority to six state-of-the-art interactive video retrieval systems. Xiaoyong Wei, Zhen-Qun Yang |
IEEE Trans. Image Process. | 1 |
| 2012 | Mining in-class social networks for large-scale pedagogical analysisabstractModeling the in-class student social networks is a highly desired goal in educational literature. However, due to the difficulty to collect social data, most of the conventional studies can only be conducted in a qualitative way on a small-scale of dataset obtained through questionnaires or interviews. We propose to solve the problems of data collection, social network construction and analysis with multimedia technology, in the way that we can automatically recognize the positions and identities of the students in classroom and construct the in-class social networks accordingly. With the social networks and the statistics on a large-scale dataset, we have demonstrated that the pedagogical analysis for investigating the co-learning patterns among the students can be conducted in a quantitative way, which provides the statistical clues about why prior studies reach conflicting conclusions on the relation between the students' positions in social networks and their academic performances. The experimental results have validated the effectiveness of the proposed approaches in both technical and pedagogical senses. Xiaoyong Wei, Zhen-Qun Yang |
ACM Multimedia | 1 |
| 2011 | Coached active learning for interactive video searchabstractActive learning with uncertainty sampling has been popularly employed in implementing interactive video search, due to its promise to reduce labeling efforts. However, since the ultimate goal of interactive search is to find as many relevant shots as possible, the purely explorative learning strategy always places conventional active learning in a dilemma whether to explore uncertain areas for a better understanding of query distribution or to harvest in certain areas for more relevant instances. In this paper, we propose a novel paradigm of active learning, where a coaching process is introduced to guide the leaner by jointly consulting an estimated prior query distribution and a posterior query distribution indicated by current classifier outcomes. To bypass the difficulty of estimating the prior query distribution from a limited number of labeled relevant instances, we propose to estimate the distribution using a set of semantic distributions which are statistically from the same distributions as the labeled relevant instances. With the coaching of both prior and posterior query distributions, the learning can be conducted and scheduled with a global perspective, and thus can explicitly balance the trade-off between exploitation and exploration. The results of the experiments on TRECVID 2005--2009 datasets validate the efficiency and effectiveness of our approach, which outperforms the conventional active learning methods with uncertainty sampling and also shows superiority to several state-of-the art interactive video search systems. Xiaoyong Wei, Zhen-Qun Yang |
ACM Multimedia | 1 |
| 2011 | Concept-Driven Multi-Modality Fusion for Video SearchabstractAs it is true for human perception that we gather information from different sources in natural and multi-modality forms, learning from multi-modalities has become an effective scheme for various information retrieval problems. In this paper, we propose a novel multi-modality fusion approach for video search, where the search modalities are derived from a diverse set of knowledge sources, such as text transcript from speech recognition, low-level visual features from video frames, and high-level semantic visual concepts from supervised learning. Since the effectiveness of each search modality greatly depends on specific user queries, prompt determination of the importance of a modality to a user query is a critical issue in multi-modality search. Our proposed approach, named concept-driven multi-modality fusion (CDMF), explores a large set of predefined semantic concepts for computing multi-modality fusion weights in a novel way. Specifically, in CDMF, we decompose the query-modality relationship into two components that are much easier to compute: query-concept relatedness and concept-modality relevancy. The former can be efficiently estimated online using semantic and visual mapping techniques, while the latter can be computed offline based on concept detection accuracy of each modality. Such a decomposition facilitates the need of adaptive learning of fusion weights for each user query on-the-fly, in contrast to the existing approaches which mostly adopted predefined query classes and/or modality weights. Experimental results on TREC video-retrieval evaluation 2005-2008 dataset validate the effectiveness of our approach, which outperforms the existing multi-modality fusion methods and achieves near-optimal performance (from oracle fusion) for many test queries. Xiaoyong Wei, Yu-Gang Jiang 0001, Chong-Wah Ngo |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2008 | Fusing semantics, observability, reliability and diversity of concept detectors for video searchabstractEffective utilization of semantic concept detectors for large-scale video search has recently become a topic of intensive studies. One of main challenges is the selection and fusion of appropriate detectors, which considers not only semantics but also the reliability of detectors, observability and diversity of detectors in target video domains. In this paper, we present a novel fusion technique which considers different aspects of detectors for query answering. In addition to utilizing detectors for bridging the semantic gap of user queries and multimedia data, we also address the issue of "observability gap" among detectors which could not be directly inferred from semantic reasoning such as using ontology. To facilitate the selection of detectors, we propose the building of two vector spaces: semantic space (SS) and observability space (OS). We categorize the set of detectors selected separately from SS and OS into four types: anchor, bridge, positive and negative concepts. A multi-level fusion strategy is proposed to novelly combine detectors, allowing the enhancement of detector reliability while enabling the observability, semantics and diversity of concepts being utilized for query answering. By experimenting the proposed approach on TRECVID 2005-2007 datasets and queries, we demonstrate the significance of considering observability, reliability and diversity, in addition to the semantics of detectors to queries. Xiaoyong Wei, Chong-Wah Ngo |
ACM Multimedia | 1 |
| 2008 | Selection of Concept Detectors for Video Search by Ontology-Enriched Semantic SpacesabstractThis paper describes the construction and utilization of two novel semantic spaces, namely ontology-enriched semantic space (OSS) and ontology-enriched orthogonal semantic space (OS2), to facilitate the selection of concept detectors for video search. These two semantic spaces are enriched with ontology knowledge, while emphasizing consistent and uniform comparison of ontological relatedness among concepts for query-to-concept mapping. OS2, in addition to being a linear space like OSS, also guarantees orthogonality of the semantic space. Compared with other ontology reasoning measures, both spaces are capable of providing platforms that offer a global view of concept inter-relatedness, by allowing evaluation of concept similarity in metric spaces. We simulate OSS and OS2by using LSCOM concepts and experiment search effectiveness with VIREO-374 concept detectors. Empirical observations indicate that the proposed semantic spaces enable more effective selection of concept detectors than eight other existing ontology measures. OS2, in particular, is better in providing a viable and reasonable solution for fusion of multiple concept detectors. Xiaoyong Wei, Chong-Wah Ngo, Yu-Gang Jiang 0001 |
IEEE Trans. Multim. | 1 |
| 2007 | Ontology-enriched semantic space for video searchabstractMultimedia-based ontology construction and reasoning have recently been recognized as two important issues in video search, particularly for bridging semantic gap. The lack of coincidence between low-level features and user expectation makes concept-based ontology reasoning an attractive mid-level framework for interpreting high-level semantics. In this paper, we propose a novel model, namely ontology-enriched semantic space (OSS), to provide a computable platform for modeling and reasoning concepts in a linear space. OSS enlightens the possibility of answering conceptual questions such as a high coverage of semantic space with minimal set of concepts, and the set of concepts to be developed for video search. More importantly, the query-to-concept mapping can be more reasonably conducted by guaranteeing the uniform and consistent comparison of concept scores for video search. We explore OSS for several tasks including concept-based video search, word sense disambiguation and multi-modality fusion. Our empirical findings show that OSS is a feasible solution to timely issues such as the measurement of concept combination and query-concept dependent fusion. Xiaoyong Wei, Chong-Wah Ngo |
ACM Multimedia | 1 |
| 2005 | Authorization Based on Palmprint
Xiaoyong Wei, Dan Xu 0001 |
ICIC (1) | 1 |