Yujiu Yang 0001

dblp:30/3847 · also Yu-Jiu Yang 0001 · DBLP profile ↗
← Back
153ranked-venue papers
3as first author
123since 2021 · last 2026
0000-0002-6427-1024ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 121 · 2 first-author · 106 since 2021Graphics, computer vision, multimedia, augmented reality and games · 63 · 52 since 2021Databases, data management, data science and information retrieval · 15 · 2 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Systems, architecture and hardware · 1Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 O-DisCo-Edit: Object Distortion Control for Unified Realistic Video Editing
abstract
Diffusion models have recently advanced video editing, yet controllable editing remains challenging due to the need for precise manipulation of diverse object properties. Current methods require different control signal for diverse editing tasks, which complicates model design and demands significant training resources. To address this, we propose O-DisCo-Edit, a unified framework that incorporates a novel object distortion control (O-DisCo). This signal, based on random and adaptive noise, flexibly encapsulates a wide range of editing cues within a single representation. Paired with a “copy-form” preservation module for preserving non-edited regions, O-DisCo-Edit enables efficient, high-fidelity editing through an effective training paradigm. Extensive experiments and comprehensive human evaluations consistently demonstrate that O-DisCo-Edit surpasses both specialized and multitask state-of-the-art methods across various video editing tasks.
Junjie Wang 0012, Lin Liu 0016, Ruihang Chu, Xiaopeng Zhang 0008, Qi Tian 0001, Yujiu Yang 0001
AAAI7
2026 Probing the Safety Robustness of LLMs in Latent Space
abstract
Tianle Gu, Kexin Huang, Zongqi Wang, Yixu Wang, Jie Li, Xin Wang, Yang Yao, Yujiu Yang, Yan Teng, Yingchun Wang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Tianle Gu, Zongqi Wang, Yixu Wang, Jie Li 0052, Xin Wang 0119, Yujiu Yang 0001, Yan Teng 0002, Yingchun Wang 0004
ACL (1)8
2026 Distributional Clarity: The Hidden Driver of RL-Friendliness in Large Language Models
abstract
Shaoning Sun, Mingzhu Cai, Huang He, Bingjin Chen, Siqi Bao, Yujiu Yang, Hua Wu, Haifeng Wang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Shaoning Sun, Mingzhu Cai, Huang He, Bingjin Chen, Siqi Bao, Yujiu Yang 0001, Hua Wu 0003, Haifeng Wang 0001
ACL (1)6
2026 SCAN: Structured Capability Assessment and Navigation for LLMs
abstract
Evaluating Large Language Models (LLMs) has become increasingly important, with automatic evaluation benchmarks gaining prominence as alternatives to human evaluation.While existing research has focused on approximating model rankings, such benchmarks fail to provide users and developers with a comprehensive and fine-grained understanding of a specific model's capabilities.To fill this gap, we propose SCAN (Structured Capability Assessment and Navigation), a practical framework that enables detailed characterization of LLM capabilities through comprehensive and fine-grained evaluation.SCAN incorporates four key components: (1) TaxBuilder, which extracts capability-indicating tags from extensive queries to construct a hierarchical taxonomy automatically; (2) RealMix, a query synthesis and filtering mechanism that ensures sufficient evaluation data for each capability tag; (3) a suite of visualization and analysis tools that facilitate efficient navigation and analysis of model capabilities; and (4) a PC 2based (Pre-Comparison-derived Criteria) LLMas-a-Judge approach that achieves significantly higher accuracy (see the definition of accuracy in § D) compared to classic LLM-as-a-Judge method.Using SCAN, we conduct a comprehensive evaluation of 21 mainstream LLMs.Our detailed analysis of the GPT-OSS family reveals substantial performance variations, even within sub-capabilities belonging to the same category of capability.This finding highlights the importance of fine-grained evaluation in accurately understanding LLM behavior.Project homepage and resources are available at https://github.com/liudan193/SCAN.
Zongqi Wang, Tianle Gu, Siqi Bao, Yujiu Yang 0001
ACL (1)6
2026 Towards Token-Level Text Anomaly Detection
abstract
Despite significant progress in text anomaly detection for web applications such as spam filtering and fake news detection, existing methods are fundamentally limited to document-level analysis, unable to identify which specific parts of a text are anomalous. We introduce token-level anomaly detection, a novel paradigm that enables fine-grained localization of anomalies within text. We formally define text anomalies at both document and token-levels, and propose a unified detection framework that operates across multiple levels. To facilitate research in this direction, we collect and annotate three benchmark datasets spanning spam, reviews and grammar errors with token-level labels. Experimental results demonstrate that our framework achieves better performance than other 6 baselines, opening new possibilities for precise anomaly localization in text. All the codes and data are publicly available on https://github.com/charles-cao/TokenCore.
Yang Cao 0019, Bicheng Yu, Sikun Yang, Ming Liu 0028, Yujiu Yang 0001
WWW5
2026 Semantic-Assisted Object Clustering for Multi-Modal Referring Video Segmentation
abstract
This paper concentrates on Multi-modal Referring Video Segmentation task, where a well optimized model is able to recognize and segment the target objects referred by the given guidance signals, e.g., language description. Early approaches model this task as a sequence prediction problem. The lack of a global view of video content leads to difficulties in effectively utilizing inter-frame relationships. Some recent works propose to perform temporal modeling with vanilla attention mechanism. However, the condensed visual representation tends to be messy about target information due to occlusion or motion blur. Unlimited non-local operation would spread such noise to all the sequences and interfere with the extraction of global representations. To address the above issue, we present Semantic-assisted Object Cluster network (SOC) and the improved SOC++ in this paper. Our method unifies temporally selective interaction and cross-modal alignment to achieve video-level understanding. In SOC++, a proxy-assisted multi-modal fusion module is introduced to perform preliminary bidirectional activation. Then a semantic integration module with progressive frame-to-video structure facilitates joint space learning across modalities and time steps. Considering that potential noisy visual embeddings would impair the overall representation of target objects in unconstrained inter-frame interactions, we propose to perform tendentious video aggregation through emphasizing the indicative role of the informative frames with lower entropy in this part. A multi-modal query contrastive supervision is also utilized to help construct well-aligned joint space at the video level. Moreover, to integrate the advantage of high-level video information and the low-level details of each frame, we introduce a dynamic query fusion module that performs joint updating of these embeddings. We conduct extensive experiments on popular referring video segmentation benchmarks, and our method outperforms state-of-the-art competitors on all benchmarks by a remarkable margin. Besides, the emphasis on temporal coherence enhances the segmentation stability and adaptability of our method in processing text expressions with temporal variations..
Yong Liu 0033, Zhuoyan Luo, Yicheng Xiao, Shuyan Li, Xiu Li 0001, Yujiu Yang 0001, Yansong Tang
IEEE Trans. Pattern Anal. Mach. Intell.7
2025 MorphMark: Flexible Adaptive Watermarking for Large Language Models
abstract
Watermarking by altering token sampling probabilities based on red-green list is a promising method for tracing the origin of text generated by large language models (LLMs).However, existing watermark methods often struggle with a fundamental dilemma: improving watermark effectiveness (the detectability of the watermark) often comes at the cost of reduced text quality.This trade-off limits their practical application.To address this challenge, we first formalize the problem within a multiobjective trade-off analysis framework.Within this framework, we identify a key factor that influences the dilemma.Unlike existing methods, where watermark strength is typically treated as a fixed hyperparameter, our theoretical insights lead to the development of MorphMark-a method that adaptively adjusts the watermark strength in response to changes in the identified factor, thereby achieving an effective resolution of the dilemma.In addition, MorphMark also prioritizes flexibility since it is an modelagnostic and model-free watermark method, thereby offering a practical solution for realworld deployment, particularly in light of the rapid evolution of AI models.Extensive experiments demonstrate that MorphMark achieves a superior resolution of the effectiveness-quality dilemma, while also offering greater flexibility and time and space efficiency.
Zongqi Wang, Tianle Gu, Baoyuan Wu, Yujiu Yang 0001
ACL (1)4
2025 Chain-of-Reasoning: Towards Unified Mathematical Reasoning in Large Language Models via a Multi-Paradigm Perspective
abstract
Yiyao Yu, Yuxiang Zhang, Dongdong Zhang, Xiao Liang, Hengyuan Zhang, Xingxing Zhang, Mahmoud Khademi, Hany Hassan Awadalla, Junjie Wang, Yujiu Yang, Furu Wei. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Yiyao Yu, Dongdong Zhang 0001, Xingxing Zhang 0002, Mahmoud Khademi, Hany Hassan, Junjie Wang 0011, Yujiu Yang 0001, Furu Wei
ACL (1)10
2025 ShifCon: Enhancing Non-Dominant Language Capabilities with a Shift-based Multilingual Contrastive Framework
abstract
Hengyuan Zhang, Chenming Shang, Sizhe Wang, Dongdong Zhang, Yiyao Yu, Feng Yao, Renliang Sun, Yujiu Yang, Furu Wei. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Chenming Shang, Dongdong Zhang 0001, Yiyao Yu, Renliang Sun, Yujiu Yang 0001, Furu Wei
ACL (1)8
2025 IG-Diff: Complex Night Scene Restoration with Illumination-Guided Diffusion Model
Yifan Chen 0001, Chunle Guo, Chongyi Li, Yujiu Yang 0001
CGI (3)5
2025 Edge-free but Structure-aware: Prototype-Guided Knowledge Distillation from GNNs to MLPs
abstract
Distilling high-accuracy Graph Neural Networks (GNNs) to low-latency multilayer perceptrons (MLPs) on graph tasks has become a hot research topic. However, conventional MLP learning relies almost exclusively on graph nodes and fails to effectively capture the graph structural information. Previous methods address this issue by processing graph edges into extra inputs for MLPs, but such graph structures may be unavailable for various scenarios. To this end, we propose Prototype-Guided Knowledge Distillation (PGKD), which does not require graph edges (edge-free setting) yet learns structure-aware MLPs. Our insight is to distill graph structural information from GNNs. Specifically, we first employ the class prototypes to analyze the impact of graph structures on GNN teachers, and then design two losses to distill such information from GNNs to MLPs. Experimental results on popular graph benchmarks demonstrate the effectiveness and robustness of the proposed PGKD.
Taiqiang Wu, Zhe Zhao 0006, Jiahao Wang 0005, Xingyu Bai, Ngai Wong 0001, Yujiu Yang 0001
COLING7
2025 ProReflow: Progressive Reflow with Decomposed Velocity
abstract
Diffusion models have achieved significant progress in both image and video generation while still suffering from huge computation costs. As an effective solution, rectified flow aims to rectify the diffusion process of diffusion models into a straight line for few-step and even one-step generation. However, in this paper, we suggest that the original training pipeline of reflow is not optimal and introduce two techniques to improve it. Firstly, we introduce progressive reflow, which progressively reflows the diffusion models in local timesteps until the whole diffusion progresses, reducing the difficulty of flow matching. Second, we introduce aligned v-prediction, which highlights the importance of direction matching in flow matching over magnitude matching. Experimental results on SDv1.5 and SDXL demonstrate the effectiveness of our method, for example, conducting on SDv1.5 achieves an FID of 10.70 on MSCOCO2014 validation set with only 4 sampling steps, close to our teacher model (32 DDIM steps, FID = 10.05). Our codes will be released at Github.
Lei Ke, Haohang Xu, Xuefei Ning, Yu Li 0022, Haoling Li, Dongsheng Jiang, Yujiu Yang 0001, Linfeng Zhang 0001
CVPR9
2025 HyperSeg: Hybrid Segmentation Assistant with Fine-grained Visual Perceiver
abstract
This paper aims to address universal segmentation for image and video perception with the strong reasoning ability empowered by Visual Large Language Models (VLLMs). Despite significant progress in current unified segmentation methods, limitations in adaptation to both image and video scenarios, as well as the complex reasoning segmentation, make it difficult for them to handle various challenging instructions and achieve an accurate understanding of fine-grained vision-language correlations. We propose HyperSeg, the first VLLM-based universal segmentation model for pixel-level image and video perception, en compassing generic segmentation tasks and more complex reasoning perception tasks requiring powerful reasoning abilities and world knowledge. Besides, to fully leverage the recognition capabilities of VLLMs and the fine-grained visual information, HyperSeg incorporates hybrid entity recognition and fine-grained visual perceiver modules for various segmentation tasks. Combined with the temporal adapter, HyperSeg achieves a comprehensive understanding of temporal information. Experimental results validate the effectiveness of our insights in resolving universal image and video segmentation tasks, including the more complex reasoning perception tasks. Our code is available at https://github.com/congvvc/HyperSeg.
Haoxian Tan, Yong Liu 0033, Dengjie Li, Yujiu Yang 0001
CVPR8
2025 DnLUT: Ultra-Efficient Color Image Denoising via Channel-Aware Lookup Tables
abstract
While deep neural networks have revolutionized image de-noising capabilities, their deployment on edge devices remains challenging due to substantial computational and memory requirements. To this end, we present DnLUT, an ultra-efficient lookup table-based framework that achieves high-quality color image denoising with minimal resource consumption. Our key innovation lies in two complementary components: a Pairwise Channel Mixer (PCM) that effectively captures inter-channel correlations and spatial dependencies in parallel, and a novel L-shaped convolution design that maximizes receptive field coverage while minimizing storage overhead. By converting these components into optimized lookup tables post-training, DnLUT achieves remarkable efficiency - requiring only 500KB storage and 0.1% energy consumption compared to its CNN contestant DnCNN, while delivering 20× faster inference. Extensive experiments demonstrate that DnLUT outperforms all existing LUT-based methods by over 1dB in PSNR, establishing a new state-of-the-art in resource-efficient color image de-noising. The project is available at https://github.com/Stephen0808/DnLUT.
Sidi Yang, Binxiao Huang, Yulun Zhang 0001, Dahai Yu 0001, Yujiu Yang 0001, Ngai Wong 0001
CVPR5
2025 IDOL: Instant Photorealistic 3D Human Creation from a Single Image
abstract
Creating a high-fidelity, animatable 3D full-body avatar from a single image is a challenging task due to the diverse appearance and poses of humans and the limited availability of high-quality training data. To achieve fast and high-quality human reconstruction, this work rethinks the task from the perspectives of dataset, model, and representation. First, we introduce a large-scale HUman-centric GEnerated dataset, HuGe100K, consisting of 100K diverse, photorealistic sets of human images. Each set contains 24-view frames in specific human poses, generated using a pose-controllable image-to-multi-view model. Next, leveraging the diversity in views, poses, and appearances within HuGe100K, we develop a scalable feed-forward transformer model to predict a 3D human Gaussian representation in a uniform space from a given human image. This model is trained to disentangle human pose, body shape, clothing geometry, and texture. The estimated Gaussians can be animated without post-processing. We conduct comprehensive experiments to validate the effectiveness of the proposed dataset and method. Our model demonstrates the ability to efficiently reconstruct photorealistic humans at 1K resolution from a single input image using a single GPU instantly. Additionally, it seamlessly supports various applications, as well as shape and texture editing tasks.
Yiyu Zhuang, Jiaxi Lv, Hao Wen 0005, Qing Shuai, Ailing Zeng, Hao Zhu 0004, Shifeng Chen, Yujiu Yang 0001, Xun Cao, Wei Liu 0005
CVPR8
2025 Invisible Entropy: Towards Safe and Efficient Low-Entropy LLM Watermarking
abstract
Logit-based LLM watermarking traces and verifies AI-generated content by maintaining green and red token lists and increasing the likelihood of green tokens during generation.However, it fails in low-entropy scenarios, where predictable outputs make green token selection difficult without disrupting natural text flow.Existing approaches address this by assuming access to the original LLM to calculate entropy and selectively watermark high-entropy tokens.However, these methods face two major challenges:(1) high computational costs and detection delays due to reliance on the original LLM, and (2) potential risks of model leakage.To address these limitations, we propose Invisible Entropy (IE), a watermarking paradigm designed to enhance both safety and efficiency.Instead of relying on the original LLM, IE introduces a lightweight feature extractor and an entropy tagger to predict whether the entropy of the next token is high or low.Furthermore, based on theoretical analysis, we develop a threshold navigator that adaptively sets entropy thresholds.It identifies a threshold where the watermark ratio decreases as the green token count increases, enhancing the naturalness of the watermarked text and improving detection robustness.Experiments on HumanEval and MBPP datasets demonstrate that IE reduces parameter size by 99% while achieving performance on par with state-of-the-art methods.Our work introduces a safe and efficient paradigm for low-entropy watermarking.We release both our standalone implementation IE-official-repo and an integration into the existing package MarkLLM.
Tianle Gu, Zongqi Wang, Yuanqi Yao, Xiangliang Zhang 0001, Yujiu Yang 0001, Xiuying Chen
EMNLP6
2025 ToM: Leveraging Tree-oriented MapReduce for Long-Context Reasoning in Large Language Models
abstract
Large Language Models (LLMs), constrained by limited context windows, often face significant performance degradation when reasoning over long contexts.To address this, Retrieval-Augmented Generation (RAG) retrieves and reasons over chunks but frequently sacrifices logical coherence due to its reliance on similarity-based rankings.Similarly, divideand-conquer frameworks (DCF) split documents into small chunks for independent reasoning and aggregation.While effective for local reasoning, DCF struggles to capture longrange dependencies and risks inducing conflicts by processing chunks in isolation.To overcome these limitations, we propose ToM, a novel Tree-oriented MapReduce framework for long-context reasoning.ToM leverages the inherent hierarchical structure of long documents (e.g., main headings and subheadings) by constructing a DocTree through hierarchical semantic parsing and performing bottom-up aggregation.Using a Tree MapReduce approach, ToM enables recursive reasoning: in the Map step, rationales are generated at child nodes; in the Reduce step, these rationales are aggregated across sibling nodes to resolve conflicts or reach consensus at parent nodes.Experimental results on 70B+ LLMs show that ToM significantly outperforms existing divide-andconquer frameworks and retrieval-augmented generation methods, achieving better logical coherence and long-context reasoning.Our code is available at https://github.com/gjn12- 31/ToM.
Jiani Guo, Zuchao Li, Jie Wu 0001, Qianren Wang, Yun Li 0011, Lefei Zhang, Hai Zhao 0001, Yujiu Yang 0001
EMNLP8
2025 Teaching Your Models to Understand Code via Focal Preference Alignment
abstract
Jie Wu, Haoling Li, Xin Zhang, Xiao Liu, Yangyu Huang, Jianwen Luo, Yizhen Zhang, Zuchao Li, Ruihang Chu, Yujiu Yang, Scarlett Li. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Jie Wu 0001, Haoling Li, Xin Zhang 0099, Xiao Liu 0029, Yangyu Huang, Zuchao Li, Ruihang Chu, Yujiu Yang 0001, Scarlett Li
EMNLP10
2025 CoHD: A Counting-Aware Hierarchical Decoding Framework for Generalized Referring Expression Segmentation
Zhuoyan Luo, Yinghao Wu, Tianheng Cheng, Yong Liu 0033, Yicheng Xiao, Hongfa Wang, Yujiu Yang 0001
ICCV8
2025 Scalable Image Tokenization with Index Backpropagation Quantization
Fengyuan Shi 0001, Zhuoyan Luo, Yixiao Ge, Yujiu Yang 0001, Ying Shan, Limin Wang 0002
ICCV4
2025 Instructseg: Unifying Instructed Visual Segmentation with Multi-Modal Large Language Models
abstract
Boosted by Multi-modal Large Language Models (MLLMs), text-guided universal segmentation models for the image and video domains have made rapid progress recently. However, these methods are often developed separately for specific domains, overlooking the similarities in task settings and solutions across these two areas. In this paper, we define the union of referring segmentation and reasoning segmentation at both the image and video levels as Instructed Visual Segmentation (IVS). Correspondingly, we propose InstructSeg, an end-to-end segmentation pipeline equipped with MLLMs for IVS. Specifically, we employ an object-aware video perceiver to extract temporal and object information from reference frames, facilitating comprehensive video understanding. Additionally, we introduce vision-guided multi-granularity text fusion to better integrate global and detailed text information with fine-grained visual guidance. By leveraging multi-task and end-to-end training, InstructSeg demonstrates superior performance across diverse image and video segmentation tasks, surpassing both segmentation specialists and MLLM-based methods with a single model. Our code is available at https://github.com/congvvc/InstructSeg.
Haoxian Tan, Yingsen Zeng, Yong Liu 0033, Hongfa Wang, Yujiu Yang 0001
ICCV7
2025 Advancing Visual Large Language Model for Multi-Granular Versatile Perception
abstract
Perception is a fundamental task in the field of computer vision, encompassing a diverse set of subtasks that can be systematically categorized into four distinct groups based on two dimensions: prediction type and instruction type. Notably, existing researches often focus solely on a limited subset of these potential combinations, which constrains their applicability and versatility across various contexts. In response to this challenge, we present MVP-LM, a Multi-granular and Versatile Perception framework incorporating Visual Large Language Model. Our framework is designed to integrate both word-based and sentence-based perception tasks alongside box and mask predictions within a single architecture. MVP-LM features an innovative multi-granularity decoder in conjunction with a CoT-inspired dataset unification strategy, enabling seamless supervised fine-tuning across a wide spectrum of tasks, including but not limited to panoptic segmentation, detection, grounding, and referring expression segmentation. Furthermore, we introduce a query enhancement strategy aimed at harnessing the decoding and generative capabilities inherent in VLLMs. Extensive experiments conducted across a range of benchmarks in both word-based and sentence-based perception tasks substantiate the efficacy of our framework. The code will be available at https://github.com/xiangwentao666/MVP-LM.
Wentao Xiang, Haoxian Tan, Dengjie Li, Yujiu Yang 0001
ICCV6
2025 ChartMimic: Evaluating LMM's Cross-Modal Reasoning Capability via Chart-to-Code Generation
abstract
We introduce a new benchmark, ChartMimic, aimed at assessing the visually-grounded code generation capabilities of large multimodal models (LMMs). ChartMimic utilizes information-intensive visual charts and textual instructions as inputs, requiring LMMs to generate the corresponding code for chart rendering. ChartMimic includes $4,800$ human-curated (figure, instruction, code) triplets, which represent the authentic chart use cases found in scientific papers across various domains (e.g., Physics, Computer Science, Economics, etc). These charts span $18$ regular types and $4$ advanced types, diversifying into $201$ subcategories. Furthermore, we propose multi-level evaluation metrics to provide an automatic and thorough assessment of the output code and the rendered charts. Unlike existing code generation benchmarks, ChartMimic places emphasis on evaluating LMMs' capacity to harmonize a blend of cognitive capabilities, encompassing visual understanding, code generation, and cross-modal reasoning. The evaluation of $3$ proprietary models and $14$ open-weight models highlights the substantial challenges posed by ChartMimic. Even the advanced GPT-4o, InternVL2-Llama3-76B only achieved an average score across Direct Mimic and Customized Mimic tasks of $82.2$ and $61.6$, respectively, indicating significant room for improvement. We anticipate that ChartMimic will inspire the development of LMMs, advancing the pursuit of artificial general intelligence.
Cheng Yang 0002, Chufan Shi, Bo Shui, Junjie Wang 0011, Mohan Jing, Linran Xu, Siheng Li, Gongye Liu, Xiaomei Nie, Deng Cai 0002, Yujiu Yang 0001
ICLR14
2025 IDA-VLM: Towards Movie Understanding via ID-Aware Large Vision-Language Model
abstract
The rapid advancement of Large Vision-Language models (LVLMs) has demonstrated a spectrum of emergent capabilities. Nevertheless, current models only focus on the visual content of a single scenario, while their ability to associate instances across different scenes has not yet been explored, which is essential for understanding complex visual content, such as movies with multiple characters and intricate plots. Towards movie understanding, a critical initial step for LVLMs is to unleash the potential of character identities memory and recognition across multiple visual scenarios. To achieve the goal, we propose visual instruction tuning with ID reference and develop an ID-Aware Large Vision-Language Model, IDA-VLM. Furthermore, our research introduces a novel benchmark MM-ID, to examine LVLMs on instance IDs memory and recognition across four dimensions: matching, location, question-answering, and captioning. Our findings highlight the limitations of existing LVLMs in recognizing and associating instance identities with ID reference. This paper paves the way for future artificial intelligence systems to possess multi-identity visual inputs, thereby facilitating the comprehension of complex visual narratives like movies.
Yatai Ji, Jie Wu 0001, Peize Sun, Xuefeng Xiao 0001, Sidi Yang, Yujiu Yang 0001, Ping Luo 0002
ICLR8
2025 IterComp: Iterative Composition-Aware Feedback Learning from Model Gallery for Text-to-Image Generation
abstract
Advanced diffusion models like Stable Diffusion 3, Omost, and FLUX have made notable strides in compositional text-to-image generation. However, these methods typically exhibit distinct strengths for compositional generation, with some excelling in handling attribute binding and others in spatial relationships. This disparity highlights the need for an approach that can leverage the complementary strengths of various models to comprehensively improve the composition capability. To this end, we introduce IterComp, a novel framework that aggregates composition-aware model preferences from multiple models and employs an iterative feedback learning approach to enhance compositional generation. Specifically, we curate a gallery of six powerful open-source diffusion models and evaluate their three key compositional metrics: attribute binding, spatial relationships, and non-spatial relationships. Based on these metrics, we develop a composition-aware model preference dataset comprising numerous image-rank pairs to train composition-aware reward models. Then, we propose an iterative feedback learning method to enhance compositionality in a closed-loop manner, enabling the progressive self-refinement of both the base diffusion model and reward models over multiple iterations. Detailed theoretical proof demonstrates the effectiveness of this method. Extensive experiments demonstrate our significant superiority over previous methods, particularly in multi-category object composition and complex semantic alignment. IterComp opens new research avenues in reward feedback learning for diffusion models and compositional generation. Code: https://github.com/YangLing0818/IterComp
Ling Yang 0006, Guohao Li 0001, Yaqi Cai, Jiake Xie, Yujiu Yang 0001, Mengdi Wang 0001, Bin Cui 0001
ICLR7
2025 Critical Tokens Matter: Token-Level Contrastive Estimation Enhances LLM's Reasoning Capability
abstract
Mathematical reasoning tasks pose significant challenges for large language models (LLMs) because they require precise logical deduction and sequence analysis. In this work, we introduce the concept of critical tokens – elements within reasoning trajectories that significantly influence incorrect outcomes. We present a novel framework for identifying these tokens through rollout sampling and demonstrate their substantial divergence from traditional error tokens. Through extensive experiments on datasets such as GSM8K and MATH500, we show that identifying and replacing critical tokens significantly improves model accuracy. We propose an efficient methodology for pinpointing these tokens in large-scale datasets using contrastive estimation and extend this framework to enhance model training processes with direct preference optimization (DPO). Experimental results on GSM8K and MATH500 benchmarks with the widely used models Llama-3 (8B and 70B) and Deepseek-math (7B) demonstrate the effectiveness of the proposed approach, cDPO. Our results underscore the potential of leveraging critical tokens to reduce errors in reasoning tasks, advancing the development of AI systems capable of robust logical deduction.
Zicheng Lin, Qiuzhi Liu, Xing Wang 0007, Ruilin Luo, Chufan Shi, Siheng Li, Yujiu Yang 0001, Zhaopeng Tu
ICML9
2025 EpiCoder: Encompassing Diversity and Complexity in Code Generation
abstract
Existing methods for code generation use code snippets as seed data, restricting the complexity and diversity of the synthesized data. In this paper, we introduce a novel feature tree-based synthesis framework, which revolves around hierarchical code features derived from high-level abstractions of code. The feature tree is constructed from raw data and refined iteratively to increase the quantity and diversity of the extracted features, which captures and recognizes more complex patterns and relationships within the code. By adjusting the depth and breadth of the sampled subtrees, our framework provides precise control over the complexity of the generated code, enabling functionalities that range from function-level operations to multi-file scenarios. We fine-tuned widely-used base models to obtain EpiCoder series, achieving state-of-the-art performance on multiple benchmarks at both the function and file levels. In particular, empirical evidence indicates that our approach shows significant potential in the synthesizing of repository-level code data. Our code and data are publicly available.
Yaoxiang Wang, Haoling Li, Xin Zhang 0099, Jie Wu 0001, Xiao Liu 0029, Wenxiang Hu, Zhongxin Guo, Yangyu Huang, Yujiu Yang 0001, Jinsong Su, Qi Chen 0009, Scarlett Li
ICML10
2025 Wan-Move: Motion-controllable Video Generation via Latent Trajectory Guidance
abstract
We present Wan-Move, a simple and scalable framework that brings motion control to video generative models. Existing motion-controllable methods typically suffer from coarse control granularity and limited scalability, leaving their outputs insufficient for practical use. We narrow this gap by achieving precise and high-quality motion control. Our core idea is to directly make the original condition features motion-aware for guiding video synthesis. To this end, we first represent object motions with dense point trajectories, allowing fine-grained control over the scene. We then project these trajectories into latent space and propagate the first frame's features along each trajectory, producing an aligned spatiotemporal feature map that tells how each scene element should move. This feature map serves as the updated latent condition, which is naturally integrated into the off-the-shelf image-to-video model, e.g., Wan-I2V-14B, as motion guidance without any architecture change. It removes the need for auxiliary motion encoders and makes fine-tuning base models easily scalable. Through scaled training, Wan-Move generates 5-second, 480p videos whose motion controllability rivals Kling 1.5 Pro's commercial Motion Brush, as indicated by user studies. To support comprehensive evaluation, we further design MoveBench, a rigorously curated benchmark featuring diverse content categories and hybrid-verified annotations. It is distinguished by larger data volume, longer video durations, and high-quality motion annotations. Extensive experiments on MoveBench and the public dataset consistently show Wan-Move's superior motion quality. Code, models, and benchmark data are made available.
Ruihang Chu, Yefei He, Zhekai Chen, Shiwei Zhang 0001, Xiaogang Xu 0002, Dingdong Wang, Hongwei Yi, Xihui Liu, Hengshuang Zhao, Yu Liu 0063, Yingya Zhang, Yujiu Yang 0001
NeurIPS13
2025 Improving Video Generation with Human Feedback
abstract
Video generation has achieved significant advances through rectified flow techniques, but issues like unsmooth motion and misalignment between videos and prompts persist. In this work, we develop a systematic pipeline that harnesses human feedback to mitigate these problems and refine the video generation model. Specifically, we begin by constructing a large-scale human preference dataset focused on modern video generation models, incorporating pairwise annotations across multi-dimensions. We then introduce VideoReward, a multi-dimensional video reward model, and examine how annotations and various design choices impact its rewarding efficacy. From a unified reinforcement learning perspective aimed at maximizing reward with KL regularization, we introduce three alignment algorithms for flow-based models. These include two training-time strategies: direct preference optimization for flow (Flow-DPO) and reward weighted regression for flow (Flow-RWR), and an inference-time technique, Flow-NRG, which applies reward guidance directly to noisy videos. Experimental results indicate that VideoReward significantly outperforms existing reward models, and Flow-DPO demonstrates superior performance compared to both Flow-RWR and supervised fine-tuning methods. Additionally, Flow-NRG lets users assign custom weights to multiple objectives during inference, meeting personalized video quality needs.
Jie Liu 0047, Gongye Liu, Jiajun Liang, Ziyang Yuan, Xiaokun Liu, Mingwu Zheng, Xiele Wu, Qiulin Wang, Menghan Xia, Xintao Wang 0002, Xiaohong Liu 0001, Pengfei Wan 0001, Di Zhang 0026, Kun Gai, Yujiu Yang 0001, Wanli Ouyang
NeurIPS16
2025 Unlocking Multimodal Mathematical Reasoning via Process Reward Model
abstract
Process Reward Models (PRMs) have shown promise in enhancing the mathematical reasoning capabilities of Large Language Models (LLMs) through Test-Time Scaling (TTS). However, their integration into multimodal reasoning remains largely unexplored. In this work, we take the first step toward unlocking the potential of PRMs in multimodal mathematical reasoning. We identify three key challenges: (i) the scarcity of high-quality reasoning data constrains the capabilities of foundation Multimodal Large Language Models (MLLMs), which imposes further limitations on the upper bounds of TTS and reinforcement learning (RL); (ii) a lack of automated methods for process labeling within multimodal contexts persists; (iii) the employment of process rewards in unimodal RL faces issues like reward hacking, which may extend to multimodal scenarios. To address these issues, we introduce URSA, a three-stage Unfolding multimodal pRocess-Supervision Aided training framework. We first construct MMathCoT-1M, a high-quality large-scale multimodal Chain-of-Thought (CoT) reasoning dataset, to build a stronger math reasoning foundation MLLM, URSA-8B. Subsequently, we go through an automatic process to synthesize process supervision data, which emphasizes both logical correctness and perceptual consistency. We introduce DualMath-1.1M to facilitate the training of URSA-8B-RM. Finally, we propose Process-Supervised Group-Relative-Policy-Optimization (PS-GRPO), pioneering a multimodal PRM-aided online RL method that outperforms vanilla GRPO. With PS-GRPO application, URSA-8B-PS-GRPO outperforms Gemma3-12B and GPT-4o by 8.4% and 2.7% on average across 6 benchmarks.
Ruilin Luo, Zhuofan Zheng, Xinzhe Ni, Zicheng Lin, Songtao Jiang, Yiyao Yu, Chufan Shi, Ruihang Chu, Yujiu Yang 0001
NeurIPS12
2025 PocketSR: The Super-Resolution Expert in Your Pocket Mobiles
abstract
Real-world image super-resolution (RealSR) aims to enhance the visual quality of in-the-wild images, such as those captured by mobile phones. While existing methods leveraging large generative models demonstrate impressive results, the high computational cost and latency make them impractical for edge deployment. In this paper, we introduce PocketSR, an ultra-lightweight, single-step model that brings generative modeling capabilities to RealSR while maintaining high fidelity. To achieve this, we design LiteED, a highly efficient alternative to the original computationally intensive VAE in SD, reducing parameters by 97.5\% while preserving high-quality encoding and decoding. Additionally, we propose online annealing pruning for the U-Net, which progressively shifts generative priors from heavy modules to lightweight counterparts, ensuring effective knowledge transfer and further optimizing efficiency. To mitigate the loss of prior knowledge during pruning, we incorporate a multi-layer feature distillation loss. Through an in-depth analysis of each design component, we provide valuable insights for future research. PocketSR, with a model size of 146M parameters, processes 4K images in just 0.8 seconds, achieving a remarkable speedup over previous methods. Notably, it delivers performance on par with state-of-the-art single-step and even multi-step RealSR models, making it a highly practical solution for edge-device applications.
Haoze Sun, Linfeng Jiang, Renjing Pei, Zhixin Wang, Haoyu Chen 0003, Fenglong Song, Yujiu Yang 0001, Wenbo Li 0002
NeurIPS11
2025 PeRL: Permutation-Enhanced Reinforcement Learning for Interleaved Vision-Language Reasoning
abstract
Inspired by the impressive reasoning capabilities demonstrated by reinforcement learning approaches like DeepSeek-R1, recent emerging research has begun exploring the use of reinforcement learning (RL) to enhance vision-language models (VLMs) for multimodal reasoning tasks. However, most existing multimodal reinforcement learning approaches remain limited to spatial reasoning within single-image contexts, yet still struggle to generalize to more complex and real-world scenarios involving multi-image positional reasoning, where understanding the relationships across images is crucial. To address this challenge, we propose a general reinforcement learning approach PeRL tailored for interleaved multimodal tasks, and a multi-stage strategy designed to enhance the exploration-exploitation trade-off, thereby improving learning efficiency and task performance. Specifically, we introduce permutation of image sequences to simulate varied positional relationships to explore more spatial and positional diversity. Furthermore, we design a rollout filtering mechanism for resampling to focus on trajectories that contribute most to learning optimal behaviors to exploit learned policies effectively. We evaluate our model on 5 widely-used multi-image benchmarks and 3 single-image benchmarks. Our experiments confirm that PeRL trained model consistently surpasses R1-related and interleaved VLM baselines by a large margin, achieving state-of-the-art performance on multi-image benchmarks, while preserving comparable performance on single-image tasks.
Shuoshuo Zhang, Haoling Li, Zhongzhi Li, Jie Wu 0001, Lei Ji 0001, Yeyun Gong, Yelong Shen, Yujiu Yang 0001
NeurIPS12
2025 Improving CTR Prediction with Graph-Enhanced Interest Networks for Sparse Behavior Sequences
abstract
Predicting click-through rates is crucial in various fields, including online advertising and recommendation systems. The key to improving the performance of CTR prediction lies in learning a robust user representation, particularly by analyzing their historical behaviors. Previous studies usually model behavior sequences through attention-based sequence models or graph-based methods, which usually struggle to explore diverse latent interests or accurately model user behaviors. Moreover, this challenge is exacerbated when users' historical behaviors are sparse, a common issue in real-world business-to-business (B2B) e-commerce scenarios. In this paper, we propose a novel Graph-Enhanced Interest Network (GEIN) to capture users' latent intents and facilitate the sequential learning of sparse behavior sequences. Specifically, we first construct a hierarchical item-intent heterogeneous graph to enrich the representation of sparse behaviors using diverse information from graphs. Next, we build a user-level behavior interest factor graph to accurately capture user interests. Additionally, a contrastive learning mechanism is incorporated to mitigate the negative robustness impacts caused by sparsity. Extensive experiments on real-world datasets demonstrate that our proposed GEIN outperforms a wide range of state-of-the-art methods. Furthermore, online A/B testing also confirms the superiority of GEIN over competing baselines in a real-world production environment.
Xuanzhou Liu, Zhibo Xiao, Luwei Yang, Hansheng Xue, Jianxing Ma, Yujiu Yang 0001
WSDM6
2025 Learning High-Quality Dynamic Memory for Video Object Segmentation
abstract
Recently, several spatial-temporal memory-based methods have verified that storing intermediate frames with masks as memory helps segment target objects in videos. However, they mainly focus on better matching between the current frame and memory frames without paying attention to the quality of the memory. Consequently, frames with poor segmentation masks may be memorized, leading to error accumulation problems. Besides, the linear increase of memory frames with the growth of frame numbers limits the ability of the models to handle long videos. To this end, we propose a Quality-aware Dynamic Memory Network (QDMN) to evaluate the segmentation quality of each frame, allowing the memory bank to selectively store accurately segmented frames and prevent error accumulation. Then, we combine the segmentation quality with temporal consistency to dynamically update the memory bank and make the models can handle videos of arbitrary length. The above operation ensures the reliability of memory frames and improves the quality of memory at the frame level. Moreover, we observe that the memory features extracted from reliable frames still contain noise and have limited representation capabilities. To address this problem, we propose to perform memory enhancement and anchoring on the basis of QDMN to improve the quality of memory from the feature level, resulting in a more robust and effective network QDMN++. Our method achieves state-of-the-art performance on all popular benchmarks. Moreover, extensive experiments demonstrate that the proposed memory screening mechanism can be applied to any memory-based methods as generic plugins.
Yong Liu 0033, Wei Zhao 0013, Weihao Xia 0001, Jiahao Wang 0005, Yansong Tang, Yujiu Yang 0001
IEEE Trans. Pattern Anal. Mach. Intell.10
2025 Global and Local Semantic Completion Learning for Vision-Language Pre-Training
abstract
Cross-modal alignment plays a crucial role in vision-language pre-training (VLP) models, enabling them to capture meaningful associations across different modalities. For this purpose, inspired by the success of masked language modeling (MLM) tasks in the NLP pre-training area, numerous masked modeling tasks have been proposed for VLP to further promote cross-modal interactions. The core idea of previous masked modeling tasks is to focus on reconstructing the masked tokens based on visible context for learning local-local alignment, i.e., associations between image patches and text tokens. However, most of them pay little attention to the global semantic features generated for the masked data, resulting in a limited cross-modal alignment ability of global representations to local features of the other modality. Therefore, in this paper, we propose a novel Global and Local Semantic Completion Learning (GLSCL) task to facilitate global-local alignment and local-local alignment simultaneously. Specifically, the GLSCL task complements the missing semantics of masked data and recovers global and local features by cross-modal interactions. Our GLSCL consists of masked global semantic completion (MGSC) and masked local token completion (MLTC). MGSC promotes learning more representative global features, which have a great impact on the performance of downstream tasks, while MLTC reconstructs modal-fusion local tokens, further enhancing accurate comprehension of multimodal data. To evaluate the proposed approaches on cross-modal alignment, we develop a validation benchmark called ALIGN-BENCH. Moreover, we present a flexible vision encoder, enabling our model to simultaneously perform image-text and video-text multimodal tasks. Experimental results show that our proposed method obtains state-of-the-art performance on various vision-language benchmarks, such as visual question answering, image-text retrieval, and video-text retrieval.
Rongcheng Tu, Yatai Ji, Jie Jiang 0015, Weijie Kong, Chengfei Cai, Hongfa Wang, Yujiu Yang 0001, Wei Liu 0005
IEEE Trans. Pattern Anal. Mach. Intell.8
2025 Efficient Text-Guided 3D-Aware Generation With Score Distillation on 3D Distribution
abstract
Text-to-3D generation enables the creation of 3D content with infinite possibilities. Existing methods typically involve training 3D generative models, which suffer from poor semantic alignment due to the scarcity of paired 3D data, or optimizing a 3D representation with 2D diffusion guidance, resulting in slow inference, low diversity, and Janus problems. In this paper, we introduce InstantDreamer, a model designed for text-guided 3D-aware generation in a single forward pass without requiring paired training datasets, thereby enhancing efficiency. To accomplish this, we extend score distillation to learn a 3D-aware semantics distribution. We distill priors from diffusion models into a 3D-aware generator, amortizing the optimization time required for new prompts and eliminating the necessity of paired training data. We equip the generator with hierarchical semantics conditioning, explicitly allowing the model to perceive the correspondence between the text distribution and the 3D latent space. Our elaborate designs empower our 3D generative model with multi-view semantic consistency and feed-forward 3D generation capabilities, thus eliminating the need for score distillation-based optimization for each prompt. Both quantitative and qualitative results on the mainstream benchmarks demonstrate that our InstantDreamer generates competitive multi-view semantic consistent 3D assets compared with state-of-the-art methods. Our method outperforms previous approaches in terms of CLIP R-Precision (66.31) and FID (28.47) while also exhibiting a significant boost in generation speed.
Yiji Cheng, Xiaoke Huang 0001, Jiaxiang Liu 0004, Shikun Feng, Yujiu Yang 0001, Yansong Tang
IEEE Trans. Circuits Syst. Video Technol.7
2024 Understanding Multimodal Deep Neural Networks: A Concept Selection View
Chenming Shang, Hao Wen 0005, Yujiu Yang 0001
CogSci4
2024 Prior Relational Schema Assists Effective Contrastive Learning for Inductive Knowledge Graph Completion
abstract
Knowledge Graph Completion (KGC) is a task aimed at uncovering the inherent relationships among known knowledge triplets in a Knowledge Graph (KG) and subsequently predicting missing links. Presently, there is a rising interest in inductive knowledge graph completion, where missing links may pertain to previously unobserved entities. Previous inductive KGC methods mainly rely on descriptive information of entities to improve the representation of unseen entities, neglecting to provide effective prior knowledge for relation modeling. To tackle this challenge, we capture prior schema-level interactions related to relations by leveraging entity type information, thereby furnishing effective prior constraints when reasoning with newly introduced entities. Moreover, We employ normal in-batch negatives and introduce schema-guided negatives to bolster the efficiency of normal contrastive representation learning. Experimental results demonstrate that our approach consistently achieves state-of-the-art performance on various established metrics across multiple benchmark datasets for link prediction. Notably, our method achieves a 20.5% relative increase in Hits@1 on the HumanWiki-Ind dataset.
Ruilin Luo, Jiayi Li 0002, Jianghangfan Zhang, Jing Xiao 0006, Yujiu Yang 0001
LREC/COLING5
2024 Rolling Shutter Correction with Intermediate Distortion Flow Estimation
abstract
This paper proposes to correct the rolling shutter (RS) distorted images by estimating the distortion flow from the global shutter (GS) to RS directly. Existing methods usually perform correction using the undistortion flow from the RS to GS. They initially predict the flow from consecutive RS frames, subsequently rescaling it as the displacement fields from the RS frame to the underlying GS image using time-dependent scaling factors. Following this, RS-aware forward warping is employed to convert the RS image into its GS counterpart. Nevertheless, this strategy is prone to two shortcomings. First, the undistortion flow estimation is rendered inaccurate by merely linear scaling the flow, due to the complex non-linear motion nature. Second, RS-aware forward warping often results in unavoidable artifacts. To address these limitations, we introduce a new framework that directly estimates the distortion flow and rectifies the RS image with the backward warping operation. More specifically, we first propose a global correlation-based flow attention mechanism to estimate the initial distortion flow and GS feature jointly, which are then refined by the following coarse-to-fine decoder layers. Additionally, a multi-distortion flow prediction strategy is integrated to mitigate the issue of inaccurate flow estimation further. Experimental results validate the effectiveness of the proposed method, which outperforms state-of-the-art approaches on various benchmarks while maintaining high efficiency. The project is available at https://github.com/ljzycmd/DFRSC.
Mingdeng Cao, Sidi Yang, Yujiu Yang 0001, Yinqiang Zheng
CVPR3
2024 Universal Segmentation at Arbitrary Granularity with Language Instruction
abstract
This paper aims to achieve universal segmentation of arbitrary semantic level. Despite significant progress in recent years, specialist segmentation approaches are limited to specific tasks and data distribution. Retraining a new model for adaptation to new scenarios or settings takes expensive computation and time cost, which raises the demand for versatile and universal segmentation model that can cater to various granularity. Although some attempts have been made for unifying different segmentation tasks or generalization to various scenarios, limitations in the definition of paradigms and input-output spaces make it difficult for them to achieve accurate understanding of content at arbitrary granularity. To this end, we present UniLSeg, a universal segmentation model that can perform segmentation at any semantic level with the guidance of language instructions. For training UniLSeg, we reorganize a group of tasks from original diverse distributions into a unified data format, where images with texts describing segmentation targets as input and corresponding masks are output. Combined with a automatic annotation engine for utilizing numerous unlabeled data, UniLSeg achieves excellent performance on various tasks and settings, surpassing both specialist and unified segmentation models. Code is available here.
Yong Liu 0033, Cairong Zhang, Jiahao Wang 0005, Yujiu Yang 0001, Yansong Tang
CVPR5
2024 Incremental Residual Concept Bottleneck Models
abstract
Concept Bottleneck Models (CBMs) map the black-box visual representations extracted by deep neural networks onto a set of interpretable concepts and use the concepts to make predictions, enhancing the transparency of the decision-making process. Multimodal pre-trained models can match visual representations with textual concept embeddings, allowing for obtaining the interpretable concept bottleneck without the expertise concept annotations. Recent research has focused on the concept bank establishment and the high-quality concept selection. However, it is challenging to construct a comprehensive concept bank through humans or large language models, which severely limits the performance of CBMs. In this work, we propose the Incremental Residual Concept Bottleneck Model (Res-CBM) to address the challenge of concept completeness. Specifically, the residual concept bottleneck model employs a set of optimizable vectors to complete missing concepts, then the incremental concept discovery module converts the complemented vectors with unclear meanings into potential concepts in the candidate concept bank. Our approach can be applied to any user-defined concept bank, as a post-hoc processing method to enhance the performance of any CBMs. Furthermore, to measure the descriptive efficiency of CBMs, the Concept Utilization Efficiency (CUE) metric is proposed. Experiments show that the Res-CBM outperforms the current state-of-the-art methods in terms of both accuracy and efficiency and achieves comparable performance to black-box models across multiple datasets.
Chenming Shang, Shiji Zhou, Xinzhe Ni, Yujiu Yang 0001, Yuwang Wang
CVPR5
2024 CoSeR: Bridging Image and Language for Cognitive Super-Resolution
abstract
Existing super-resolution (SR) models primarily focus on restoring local texture details, often neglecting the global semantic information within the scene. This oversight can lead to the omission of crucial semantic details or the intro-duction of inaccurate textures during the recovery process. In our work, we introduce the Cognitive Super-Resolution (CoSeR) framework, empowering SR models with the ca-pacity to comprehend low-resolution images. We achieve this by marrying image appearance and language under-standing to generate a cognitive embedding, which not only activates prior information from large text-to-image diffusion models but also facilitates the generation of high-quality reference images to optimize the SR process. To fur-ther improve image fidelity, we propose a novel condition injection scheme called “Ali-in-Attention ”, consolidating all conditional information into a single module. Conse-quently, our method successfully restores semantically cor-rect and photorealistic details, demonstrating state-of-the-art performance across multiple benchmarks. Project page: https://coser-main.github.io/
Haoze Sun, Wenbo Li 0002, Jianzhuang Liu, Haoyu Chen 0003, Renjing Pei, Xueyi Zou, Youliang Yan, Yujiu Yang 0001
CVPR8
2024 Bridging the Gap: A Unified Video Comprehension Framework for Moment Retrieval and Highlight Detection
abstract
Video Moment Retrieval (MR) and Highlight Detection (HD) have attracted significant attention due to the growing demand for video analysis. Recent approaches treat MR and HD as similar video grounding problems and address them together with transformer-based architecture. However, we observe that the emphasis of MR and HD differs, with one necessitating the perception of local relationships and the other prioritizing the understanding of global contexts. Consequently, the lack of task-specific design will inevitably lead to limitations in associating the intrinsic specialty of two tasks. To tackle the issue, we propose a Unified Video COMprehension framework (UVCOM) to bridge the gap and jointly solve MR and HD effectively. By performing progressive integration on intra and inter-modality across multi-granularity, UVCOM achieves the comprehensive understanding in processing a video. Moreover, we present multi-aspect contrastive learning to consolidate the local relation modeling and global knowledge accumulation via well aligned multi-modal space. Extensive experiments on QVHighlights, Charades-STA, TACoS, YouTube Highlights and TVSum datasets demonstrate the effectiveness and rationality of UVCOM which outperforms the state-of-the-art methods by a remarkable margin. Code is available at https://github.com/EasonXiao-888/UVCOM.
Yicheng Xiao, Zhuoyan Luo, Yong Liu 0033, Yue Ma 0016, Hengwei Bian, Yatai Ji, Yujiu Yang 0001, Xiu Li 0001
CVPR7
2024 LoCa: Logit Calibration for Knowledge Distillation
abstract
Knowledge Distillation (KD), aiming to train a better student model by mimicking the teacher model, plays an important role in model compression. One typical way is to align the output logits. However, we find a common issue named mis-instruction, that the student would be misled when the predictions based on teacher logits do not follow the labels. Meanwhile, there is other useful dark knowledge in the logits such as the class discriminability, which is vital for distillation. In this paper, we propose a simple yet effective Logit Calibration (LoCa) method, which calibrates the logits from the teacher model based on the ground-truth labels. The key insight is to correct the prediction (to address the mis-instruction issue) and maintain useful dark knowledge simultaneously. Our proposed LoCa does not require any additional parameters. Empirical results on image classification and text generation tasks demonstrate that LoCa can effectively improve the performance of baselines.
Runming Yang, Taiqiang Wu, Yujiu Yang 0001
ECAI3
2024 A Comprehensive Study of Multimodal Large Language Models for Image Quality Assessment
Tianhe Wu, Kede Ma, Jie Liang 0007, Yujiu Yang 0001, Lei Zhang 0006
ECCV (74)4
2024 Taming Lookup Tables for Efficient Image Retouching
Sidi Yang, Binxiao Huang, Mingdeng Cao, Yatai Ji, Hanzhong Guo, Ngai Wong 0001, Yujiu Yang 0001
ECCV (58)7
2024 Encouraging Divergent Thinking in Large Language Models through Multi-Agent Debate
abstract
Modern large language models (LLMs) like ChatGPT have shown remarkable performance on general language tasks but still struggle on complex reasoning tasks, which drives the research on cognitive behaviors of LLMs to explore human-like problem-solving strategies.Along this direction, one representative strategy is self-reflection, which asks an LLM to refine the solution with the feedback generated by itself iteratively.However, our study shows that such reflection-style methods suffer from the Degeneration-of-Thought (DoT) problem: once the LLM has established confidence in its solutions, it is unable to generate novel thoughts later through reflection even if its initial stance is incorrect.To address the DoT problem, we propose a Multi-Agent Debate (MAD) framework, in which multiple agents express their arguments in the state of "tit for tat" and a judge manages the debate process to obtain a final solution.Clearly, our MAD framework encourages divergent thinking in LLMs which would be helpful for tasks that require deep levels of contemplation.Experiment results on two challenging datasets, commonsense machine translation and counterintuitive arithmetic reasoning, demonstrate the effectiveness of our MAD framework.Extensive analyses suggest that the adaptive break of debate and the modest level of "tit for tat" state are required for MAD to obtain good performance.Moreover, we find that LLMs might not be a fair judge if different LLMs are used for agents.Code is available at https://github. com/Skytliang/Multi-Agents-Debate.
Zhiwei He 0002, Wenxiang Jiao, Xing Wang 0007, Yan Wang 0060, Rui Wang 0015, Yujiu Yang 0001, Shuming Shi 0001, Zhaopeng Tu
EMNLP7
2024 PTD-SQL: Partitioning and Targeted Drilling with LLMs in Text-to-SQL
abstract
Large Language Models (LLMs) have emerged as powerful tools for Text-to-SQL tasks, exhibiting remarkable reasoning capabilities.Different from tasks such as math word problems and commonsense reasoning, SQL solutions have a relatively fixed pattern.This facilitates the investigation of whether LLMs can benefit from categorical thinking, mirroring how humans acquire knowledge through inductive reasoning based on comparable examples.In this study, we propose that employing query group partitioning allows LLMs to focus on learning the thought processes specific to a single problem type, consequently enhancing their reasoning abilities across diverse difficulty levels and problem categories.Our experiments reveal that multiple advanced LLMs, when equipped with PTD-SQL, can either surpass or match previous state-of-theart (SOTA) methods on the Spider and BIRD datasets.Intriguingly, models with varying initial performances have exhibited significant improvements, mainly at the boundary of their capabilities after targeted drilling, suggesting a parallel with human progress.Code is available at https://github.com/lrlbbzl/PTD-SQL.
Ruilin Luo, Binghuai Lin, Zicheng Lin, Yujiu Yang 0001
EMNLP5
2024 SciAgent: Tool-augmented Language Models for Scientific Reasoning
abstract
Yubo Ma, Zhibin Gou, Junheng Hao, Ruochen Xu, Shuohang Wang, Liangming Pan, Yujiu Yang, Yixin Cao, Aixin Sun. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Yubo Ma, Zhibin Gou, Junheng Hao, Ruochen Xu, Shuohang Wang, Liangming Pan, Yujiu Yang 0001, Yixin Cao 0002, Aixin Sun
EMNLP7
2024 A Thorough Examination of Decoding Methods in the Era of LLMs
abstract
Decoding methods play an indispensable role in converting language models from next-token predictors into practical task solvers.Prior research on decoding methods, primarily focusing on task-specific models, may not extend to the current era of general-purpose large language models (LLMs).Moreover, the recent influx of decoding strategies has further complicated this landscape.This paper provides a comprehensive and multifaceted analysis of various decoding methods within the context of LLMs, evaluating their performance, robustness to hyperparameter changes, and decoding speeds across a wide range of tasks, models, and deployment environments.Our findings reveal that decoding method performance is notably task-dependent and influenced by factors such as alignment, model size, and quantization.Intriguingly, sensitivity analysis exposes that certain methods achieve superior performance at the cost of extensive hyperparameter tuning, highlighting the trade-off between attaining optimal results and the practicality of implementation in varying contexts.
Chufan Shi, Deng Cai 0002, Zhisong Zhang, Yujiu Yang 0001, Wai Lam
EMNLP6
2024 ToolBeHonest: A Multi-level Hallucination Diagnostic Benchmark for Tool-Augmented Large Language Models
abstract
Yuxiang Zhang, Jing Chen, Junjie Wang, Yaxin Liu, Cheng Yang, Chufan Shi, Xinyu Zhu, Zihao Lin, Hanwen Wan, Yujiu Yang, Tetsuya Sakai, Tian Feng, Hayato Yamana. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Junjie Wang 0011, Cheng Yang 0007, Chufan Shi, Hanwen Wan, Yujiu Yang 0001, Tetsuya Sakai, Tian Feng 0001, Hayato Yamana
EMNLP10
2024 Hint-Enhanced In-Context Learning Wakes Large Language Models Up For Knowledge-Intensive Tasks
abstract
In-context learning (ICL) ability has emerged with the increasing scale of large language models (LLMs), enabling them to learn input-label mappings from demonstrations and perform well on downstream tasks. However, under the standard ICL setting, LLMs may sometimes neglect query-related information in demonstrations, leading to incorrect predictions. To address this limitation, we propose a new paradigm called Hint-enhanced In-Context Learning (HICL) to explore the power of ICL in open-domain question answering, an important form in knowledge-intensive tasks. HICL leverages LLMs’ reasoning ability to extract query-related knowledge from demonstrations, then concatenates the knowledge to prompt LLMs in a more explicit way. Furthermore, we track the source of this knowledge to identify specific examples, and introduce a Hint-related Example Retriever (HER) to select informative examples for enhanced demonstrations. We evaluate HICL with HER on 3 open-domain QA benchmarks, and observe average performance gains of 2.89 EM score and 2.52 F1 score on gpt-3.5-turbo, 7.62 EM score and 7.27 F1 score on LLaMA-2-Chat-7B compared with standard setting.
Qingyan Guo, Xinzhe Ni, Chufan Shi, Lemao Liu, Haiyun Jiang, Yujiu Yang 0001
ICASSP7
2024 CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing
abstract
Recent developments in large language models (LLMs) have been impressive. However, these models sometimes show inconsistencies and problematic behavior, such as hallucinating facts, generating flawed code, or creating offensive and toxic content. Unlike these models, humans typically utilize external tools to cross-check and refine their initial content, like using a search engine for fact-checking, or a code interpreter for debugging. Inspired by this observation, we introduce a framework called CRITIC that allows LLMs, which are essentially “black boxes” to validate and progressively amend their own outputs in a manner similar to human interaction with tools. More specifically, starting with an initial output, CRITIC interacts with appropriate tools to evaluate certain aspects of the text, and then revises the output based on the feedback obtained during this validation process. Comprehensive evaluations involving free-form question answering, mathematical program synthesis, and toxicity reduction demonstrate that CRITIC consistently enhances the performance of LLMs. Meanwhile, our research highlights the crucial importance of external feedback in promoting the ongoing self-improvement of LLMs.
Zhibin Gou, Zhihong Shao, Yeyun Gong, Yelong Shen, Yujiu Yang 0001, Nan Duan 0001, Weizhu Chen
ICLR5
2024 ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving
abstract
Large language models have made significant progress in various language tasks, yet they still struggle with complex mathematics. In this paper, we propose ToRA a series of Tool-integrated Reasoning Agents designed to solve challenging mathematical problems by seamlessly integrating natural language reasoning with the utilization of external tools (e.g., computation libraries and symbolic solvers), thereby amalgamating the analytical prowess of language and the computational efficiency of tools. To train ToRA, we curate interactive tool-use trajectories on mathematical datasets, apply imitation learning on the annotations, and propose output space shaping to further refine models' reasoning behavior. As a result, ToRA models significantly outperform open-source models on 10 mathematical reasoning datasets across all scales with 13%-19% absolute improvements on average. Notably, ToRA-7B reaches 44.6% on the competition-level dataset MATH, surpassing the best open-source model WizardMath-70B by 22% absolute. ToRA-34B is also the first open-source model that achieves an accuracy exceeding 50% on MATH, which significantly outperforms GPT-4's CoT result, and is competitive with GPT-4 solving problems with programs. Additionally, we conduct a comprehensive analysis of the benefits and remaining challenges of tool interaction for mathematical reasoning, providing valuable insights for future research.
Zhibin Gou, Zhihong Shao, Yeyun Gong, Yelong Shen, Yujiu Yang 0001, Minlie Huang, Nan Duan 0001, Weizhu Chen
ICLR5
2024 Connecting Large Language Models with Evolutionary Algorithms Yields Powerful Prompt Optimizers
abstract
Large Language Models (LLMs) excel in various tasks, but they rely on carefully crafted prompts that often demand substantial human effort. To automate this process, in this paper, we propose a novel framework for discrete prompt optimization, called EvoPrompt, which borrows the idea of evolutionary algorithms (EAs) as they exhibit good performance and fast convergence. To enable EAs to work on discrete prompts, which are natural language expressions that need to be coherent and human-readable, we connect LLMs with EAs. This approach allows us to simultaneously leverage the powerful language processing capabilities of LLMs and the efficient optimization performance of EAs. Specifically, abstaining from any gradients or parameters, EvoPrompt starts from a population of prompts and iteratively generates new prompts with LLMs based on the evolutionary operators, improving the population based on the development set. We optimize prompts for both closed- and open-source LLMs including GPT-3.5 and Alpaca, on 31 datasets covering language understanding, generation tasks, as well as BIG-Bench Hard (BBH) tasks. EvoPrompt significantly outperforms human-engineered prompts and existing methods for automatic prompt generation (e.g., up to 25% on BBH). Furthermore, EvoPrompt demonstrates that connecting LLMs with EAs creates synergies, which could inspire further research on the combination of LLMs and conventional algorithms.
Qingyan Guo, Rui Wang 0028, Junliang Guo, Kaitao Song, Xu Tan 0003, Jiang Bian 0002, Yujiu Yang 0001
ICLR9
2024 Spurious Feature Diversification Improves Out-of-distribution Generalization
abstract
Generalization to out-of-distribution (OOD) data is a critical challenge in machine learning. Ensemble-based methods, like weight space ensembles that interpolate model parameters, have been shown to achieve superior OOD performance. However, the underlying mechanism for their effectiveness remains unclear. In this study, we closely examine WiSE-FT, a popular weight space ensemble method that interpolates between a pre-trained and a fine-tuned model. We observe an unexpected ``FalseFalseTrue" phenomenon, in which WiSE-FT successfully corrects many cases where each individual model makes incorrect predictions, which contributes significantly to its OOD effectiveness. To gain further insights, we conduct theoretical analysis in a multi-class setting with a large number of spurious features. Our analysis predicts the above phenomenon and it further shows that ensemble-based models reduce prediction errors in the OOD settings by utilizing a more diverse set of spurious features. Contrary to the conventional wisdom that focuses on learning invariant features for better OOD performance, our findings suggest that incorporating a large number of diverse spurious features weakens their individual contributions, leading to improved overall OOD generalization performance. Additionally, our findings provide the first explanation for the mysterious phenomenon of weight space ensembles outperforming output space ensembles in OOD. Empirically we demonstrate the effectiveness of utilizing diverse spurious features on a MultiColorMNIST dataset, and our experimental results are consistent with the theoretical analysis. Building upon the new theoretical insights into the efficacy of ensemble methods, we further identify an issue of WiSE-FT caused by the overconfidence of fine-tuned models in OOD situations. This overconfidence magnifies the fine-tuned model's incorrect prediction, leading to deteriorated OOD ensemble performance. To remedy this problem, we propose a novel method called BAlaNced averaGing (BANG) to mitigate the overconfidence problem, which significantly enhances the OOD performance of WiSE-FT.
Yifan Hao 0002, Honam Wong, Hanze Dong, Yujiu Yang 0001, Tong Zhang 0001
ICLR7
2024 Continuous Invariance Learning
abstract
Invariance learning methods aim to learn invariant features in the hope that they generalize under distributional shift. Although many tasks are naturally characterized by continuous domains, current invariance learning techniques generally assume categorically indexed domains. For example, auto-scaling in cloud computing often needs a CPU utilization prediction model that generalizes across different times (e.g., time of a day and date of a year), where `time' is a continuous domain index. In this paper, we start by theoretically showing that existing invariance learning methods can fail for continuous domain problems. Specifically, the naive solution of splitting continuous domains into discrete ones ignores the underlying relationship among domains, and therefore potentially leads to suboptimal performance. To address this challenge, we then propose Continuous Invariance Learning (CIL), which extracts invariant features across continuously indexed domains. CIL is a novel adversarial procedure which measures and controls the conditional independence between the labels and continuous domain indices given the extracted features. Our theoretical analysis demonstrates that CIL learns features that satisfy the invariant constraint with infinite samples. Empirical results on both synthetic and real-world datasets (including data collected from production systems) show that CIL consistently outperforms strong baselines among all the tasks.
Lin Yong, Fan Zhou 0012, Lintao Ma, Jianmeng Liu, Yansu He, Yu Liu 0071, James Y. Zhang, Yujiu Yang 0001, Hao Wang 0014
ICLR10
2024 Accelerating Diffusion Models for Inverse Problems through Shortcut Sampling
Gongye Liu, Haoze Sun, Jiayi Li 0002, Yujiu Yang 0001
IJCAI5
2024 Multimodal Prototype-Enhanced Network for Few-Shot Action Recognition
abstract
Current methods for few-shot action recognition mainly fall into the metric learning framework following ProtoNet, which demonstrates the importance of prototypes. Although they achieve relatively good performance, the effect of multimodal information is ignored, e.g. label texts. In this work, we propose a novel MultimOdal PRototype-ENhanced Network (MORN), which uses the semantic information of label texts as multimodal information to enhance prototypes. A CLIP visual encoder and a frozen CLIP text encoder are introduced to obtain features with good multimodal initialization. Then in the visual flow, visual prototypes are computed by a visual prototype-computed module. In the text flow, a semantic-enhanced (SE) module and an inflating operation are used to obtain text prototypes. The final multimodal prototypes are then computed by a multimodal prototype-enhanced (MPE) module. Besides, we define a PRototype SImilarity DiffErence (PRIDE) to evaluate the quality of prototypes, which is used to verify our improvement on the prototype level and effectiveness of MORN. We conduct extensive experiments on four popular few-shot action recognition datasets: HMDB51, UCF101, Kinetics and SSv2, and MORN achieves state-of-the-art results. When plugging PRIDE into the training stage, the performance can be further improved.
Xinzhe Ni, Yong Liu 0033, Hao Wen 0005, Yatai Ji, Jing Xiao 0006, Yujiu Yang 0001
ICMR6
2024 InsCL: A Data-efficient Continual Learning Paradigm for Fine-tuning Large Language Models with Instructions
abstract
Yifan Wang, Yafei Liu, Chufan Shi, Haoling Li, Chen Chen, Haonan Lu, Yujiu Yang. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Chufan Shi, Haoling Li, Chen Chen 0015, Haonan Lu, Yujiu Yang 0001
NAACL-HLT7
2024 MLLMGuard: A Multi-dimensional Safety Evaluation Suite for Multimodal Large Language Models
abstract
Powered by remarkable advancements in Large Language Models (LLMs), Multimodal Large Language Models (MLLMs) demonstrate impressive capabilities in manifold tasks.However, the practical application scenarios of MLLMs are intricate, exposing them to potential malicious instructions and thereby posing safety risks.While current benchmarks do incorporate certain safety considerations, they often lack comprehensive coverage and fail to exhibit the necessary rigor and robustness.For instance, the common practice of employing GPT-4V as both the evaluator and a model to be evaluated lacks credibility, as it tends to exhibit a bias toward its own responses.In this paper, we present MLLMGuard, a multi-dimensional safety evaluation suite for MLLMs, including a bilingual image-text evaluation dataset, inference utilities, and a lightweight evaluator.MLLMGuard's assessment comprehensively covers two languages (English and Chinese) and five important safety dimensions (Privacy, Bias, Toxicity, Truthfulness, and Legality), each with corresponding rich subtasks.Focusing on these dimensions, our evaluation dataset is primarily sourced from platforms such as social media, and it integrates text-based and image-based red teaming techniques with meticulous annotation by human experts.This can prevent inaccurate evaluation caused by data leakage when using open-source datasets and ensures the quality and challenging nature of our benchmark.Additionally, a fully automated lightweight evaluator termed GuardRank is developed, which achieves significantly higher evaluation accuracy than GPT-4.Our evaluation results across 13 advanced models indicate that MLLMs still have a substantial journey ahead before they can be considered safe and responsible.
Tianle Gu, Dandan Liang, Yixu Wang, Haiquan Zhao 0002, Yuanqi Yao, Xingge Qiao, Keqing Wang, Yujiu Yang 0001, Yan Teng 0002, Yu Qiao 0001, Yingchun Wang 0004
NeurIPS10
2024 Not All Tokens Are What You Need for Pretraining
abstract
Previous language model pre-training methods have uniformly applied a next-token prediction loss to all training tokens. Challenging this norm, we posit that ''Not all tokens in a corpus are equally important for language model training''. Our initial analysis examines token-level training dynamics of language model, revealing distinct loss patterns for different tokens. Leveraging these insights, we introduce a new language model called Rho-1. Unlike traditional LMs that learn to predict every next token in a corpus, Rho-1 employs Selective Language Modeling (SLM), which selectively trains on useful tokens that aligned with the desired distribution. This approach involves scoring training tokens using a reference model, and then training the language model with a focused loss on tokens with higher scores. When continual continual pretraining on 15B OpenWebMath corpus, Rho-1 yields an absolute improvement in few-shot accuracy of up to 30% in 9 math tasks. After fine-tuning, Rho-1-1B and 7B achieved state-of-the-art results of 40.6% and 51.8% on MATH dataset, respectively - matching DeepSeekMath with only 3% of the pretraining tokens. Furthermore, when continual pretraining on 80B general tokens, Rho-1 achieves 6.8% average enhancement across 15 diverse tasks, increasing both data efficiency and performance of the language model pre-training.
Zhenghao Lin, Zhibin Gou, Yeyun Gong, Xiao Liu 0029, Yelong Shen, Ruochen Xu, Chen Lin 0001, Yujiu Yang 0001, Jian Jiao 0007, Nan Duan 0001, Weizhu Chen
NeurIPS8
2024 AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents
abstract
Evaluating large language models (LLMs) as general-purpose agents is essential for understanding their capabilities and facilitating their integration into practical applications. However, the evaluation process presents substantial challenges. A primary obstacle is the benchmarking of agent performance across diverse scenarios within a unified framework, especially in maintaining partially-observable environments and ensuring multi-round interactions. Moreover, current evaluation frameworks mostly focus on the final success rate, revealing few insights during the process and failing to provide a deep understanding of the model abilities. To address these challenges, we introduce AgentBoard, a pioneering comprehensive benchmark and accompanied open-source evaluation framework tailored to analytical evaluation of LLM agents. AgentBoard offers a fine-grained progress rate metric that captures incremental advancements as well as a comprehensive evaluation toolkit that features easy assessment of agents for multi-faceted analysis through interactive visualization. This not only sheds light on the capabilities and limitations of LLM agents but also propels the interpretability of their performance to the forefront. Ultimately, AgentBoard serves as a significant step towards demystifying agent behaviors and accelerating the development of stronger LLM agents.
Junlei Zhang, Cheng Yang 0007, Yujiu Yang 0001, Yaohui Jin, Zhen-Zhong Lan, Lingpeng Kong, Junxian He
NeurIPS5
2024 Unchosen Experts Can Contribute Too: Unleashing MoE Models' Power by Self-Contrast
abstract
Mixture-of-Experts (MoE) has emerged as a prominent architecture for scaling model size while maintaining computational efficiency. In MoE, each token in the input sequence activates a different subset of experts determined by a routing mechanism. However, the unchosen experts in MoE models do not contribute to the output, potentially leading to underutilization of the model's capacity. In this work, we first conduct exploratory studies to demonstrate that increasing the number of activated experts does not necessarily improve and can even degrade the output quality. Then, we show that output distributions from an MoE model using different routing strategies substantially differ, indicating that different experts do not always act synergistically. Motivated by these findings, we propose **S**elf-**C**ontrast **M**ixture-**o**f-**E**xperts (SCMoE), a training-free strategy that utilizes unchosen experts in a self-contrast manner during inference. In SCMoE, the next-token probabilities are determined by contrasting the outputs from strong and weak activation using the same MoE model. Our method is conceptually simple and computationally lightweight, as it incurs minimal latency compared to greedy decoding. Experiments on several benchmarks (GSM8K, StrategyQA, MBPP and HumanEval) demonstrate that SCMoE can consistently enhance Mixtral 8x7B’s reasoning capability across various domains. For example, it improves the accuracy on GSM8K from 61.79 to 66.94. Moreover, combining SCMoE with self-consistency yields additional gains, increasing major@20 accuracy from 75.59 to 78.31.
Chufan Shi, Cheng Yang 0002, Jiahao Wang 0005, Taiqiang Wu, Siheng Li, Deng Cai 0002, Yujiu Yang 0001, Yu Meng 0001
NeurIPS8
2024 RealCompo: Balancing Realism and Compositionality Improves Text-to-Image Diffusion Models
abstract
Diffusion models have achieved remarkable advancements in text-to-image generation. However, existing models still have many difficulties when faced with multiple-object compositional generation. In this paper, we propose ***RealCompo***, a new *training-free* and *transferred-friendly* text-to-image generation framework, which aims to leverage the respective advantages of text-to-image models and spatial-aware image diffusion models (e.g., layout, keypoints and segmentation maps) to enhance both realism and compositionality of the generated images. An intuitive and novel *balancer* is proposed to dynamically balance the strengths of the two models in denoising process, allowing plug-and-play use of any model without extra training. Extensive experiments show that our RealCompo consistently outperforms state-of-the-art text-to-image models and spatial-aware image diffusion models in multiple-object compositional generation while keeping satisfactory realism and compositionality of the generated images. Notably, our RealCompo can be seamlessly extended with a wide range of spatial-aware image diffusion models and stylized diffusion models. Code is available at: https://github.com/YangLing0818/RealCompo
Ling Yang 0006, Yaqi Cai, Zhaochen Yu, Kai-Ni Wang, Jiake Xie, Minkai Xu, Yujiu Yang 0001, Bin Cui 0001
NeurIPS10
2024 Prior Bilinear-Based Models for Knowledge Graph Completion
Jiayi Li 0002, Ruilin Luo, Jing Xiao 0006, Yujiu Yang 0001
ECML/PKDD (3)5
2024 Deep Evolutional Instant Interest Network for CTR Prediction in Trigger-Induced Recommendation
abstract
The recommendation has been playing a key role in many industries, e.g., e-commerce, streaming media, social media, etc. Recently, a new recommendation scenario, called Trigger-Induced Recommendation (TIR), where users are able to explicitly express their instant interests via trigger items, is emerging as an essential role in many e-commerce platforms, e.g., Alibaba.com and Amazon. Without explicitly modeling the user's instant interest, traditional recommendation methods usually obtain sub-optimal results in TIR. Even though there are a few methods considering the trigger and target items simultaneously to solve this problem, they still haven't taken into account temporal information of user behaviors, the dynamic change of user instant interest when the user scrolls down and the interactions between the trigger and target items. To tackle these problems, we propose a novel method -- Deep Evolutional Instant Interest Network (DEI2N), for click-through rate prediction in TIR scenarios. Specifically, we design a User Instant Interest Modeling Layer to predict the dynamic change of the intensity of instant interest when the user scrolls down. Temporal information is utilized in user behavior modeling. Moreover, an Interaction Layer is introduced to learn better interactions between the trigger and target items. We evaluate our method on several offline and real-world industrial datasets. Experimental results show that our proposed DEI2N outperforms state-of-the-art baselines. In addition, online A/B testing demonstrates the superiority over the existing baseline in real-world production environments.
Zhibo Xiao, Luwei Yang, Tao Zhang 0124, Wei Ning, Yujiu Yang 0001
WSDM6
2024 Generalizable Black-Box Adversarial Attack With Meta Learning
abstract
In the scenario of black-box adversarial attack, the target model's parameters are unknown, and the attacker aims to find a successful adversarial perturbation based on query feedback under a query budget. Due to the limited feedback information, existing query-based black-box attack methods often require many queries for attacking each benign example. To reduce query cost, we propose to utilize the feedback information across historical attacks, dubbed example-level adversarial transferability. Specifically, by treating the attack on each benign example as one task, we develop a meta-learning framework by training a meta generator to produce perturbations conditioned on benign examples. When attacking a new benign example, the meta generator can be quickly fine-tuned based on the feedback information of the new task as well as a few historical attacks to produce effective perturbations. Moreover, since the meta-train procedure consumes many queries to learn a generalizable generator, we utilize model-level adversarial transferability to train the meta generator on a white-box surrogate model, then transfer it to help the attack against the target model. The proposed framework with the two types of adversarial transferability can be naturally combined with any off-the-shelf query-based attack methods to boost their performance, which is verified by extensive experiments. The source code is available at https://github.com/SCLBD/MCG-Blackbox.
Yong Zhang 0034, Baoyuan Wu, Jingyi Zhang 0005, Yanbo Fan, Yujiu Yang 0001
IEEE Trans. Pattern Anal. Mach. Intell.7
2024 Exploring Human-Like Translation Strategy with Large Language Models
abstract
Abstract Large language models (LLMs) have demonstrated impressive capabilities in general scenarios, exhibiting a level of aptitude that approaches, in some aspects even surpasses, human-level intelligence. Among their numerous skills, the translation abilities of LLMs have received considerable attention. Compared to typical machine translation that focuses solely on source-to-target mapping, LLM-based translation can potentially mimic the human translation process, which might take preparatory steps to ensure high-quality translation. This work explores this possibility by proposing the MAPS framework, which stands for Multi-Aspect Prompting and Selection. Specifically, we enable LLMs first to analyze the given source sentence and induce three aspects of translation-related knowledge (keywords, topics, and relevant demonstrations) to guide the final translation process. Moreover, we employ a selection mechanism based on quality estimation to filter out noisy and unhelpful knowledge. Both automatic (3 LLMs × 11 directions × 2 automatic metrics) and human evaluation (preference study and MQM) demonstrate the effectiveness of MAPS. Further analysis shows that by mimicking the human translation process, MAPS reduces various translation errors such as hallucination, ambiguity, mistranslation, awkward style, untranslated text, and omission. Source code is available at https://github.com/zwhe99/MAPS-mt.
Zhiwei He 0002, Wenxiang Jiao, Zhuosheng Zhang 0001, Yujiu Yang 0001, Rui Wang 0015, Zhaopeng Tu, Shuming Shi 0001, Xing Wang 0007
Trans. Assoc. Comput. Linguistics5
2024 An Energy-based Model for Word-level AutoCompletion in Computer-aided Translation
abstract
Abstract Word-level AutoCompletion (WLAC) is a rewarding yet challenging task in Computer-aided Translation. Existing work addresses this task through a classification model based on a neural network that maps the hidden vector of the input context into its corresponding label (i.e., the candidate target word is treated as a label). Since the context hidden vector itself does not take the label into account and it is projected to the label through a linear classifier, the model cannot sufficiently leverage valuable information from the source sentence as verified in our experiments, which eventually hinders its overall performance. To alleviate this issue, this work proposes an energy-based model for WLAC, which enables the context hidden vector to capture crucial information from the source sentence. Unfortunately, training and inference suffer from efficiency and effectiveness challenges, therefore we employ three simple yet effective strategies to put our model into practice. Experiments on four standard benchmarks demonstrate that our reranking-based approach achieves substantial improvements (about 6.07%) over the previous state-of-the-art model. Further analyses show that each strategy of our approach contributes to the final performance.1
Cheng Yang 0007, Guoping Huang, Mo Yu, Zhirui Zhang, Siheng Li, Shuming Shi 0001, Yujiu Yang 0001, Lemao Liu
Trans. Assoc. Comput. Linguistics8
2024 StyleCrafter: Taming Artistic Video Diffusion with Reference-Augmented Adapter Learning
abstract
Text-to-video (T2V) models have shown remarkable capabilities in generating diverse videos. However, they struggle to produce user-desired artistic videos due to (i) text's inherent clumsiness in expressing specific styles and (ii) the generally degraded style fidelity. To address these challenges, we introduce StyleCrafter, a generic method that enhances pretrained T2V models with a style control adapter, allowing video generation in any style by feeding a reference image. Considering the scarcity of artistic video data, we propose to first train a style control adapter using style-rich image datasets, then transfer the learned stylization ability to video generation through a tailor-made finetuning paradigm. To promote content-style disentanglement, we employ carefully designed data augmentation strategies to enhance decoupled learning. Additionally, we propose a scale-adaptive fusion module to balance the influences of text-based content features and image-based style features, which helps generalization across various text and style combinations. StyleCrafter efficiently generates high-quality stylized videos that align with the content of the texts and resemble the style of the reference images. Experiments demonstrate that our approach is more flexible and efficient than existing competitors. Project page: https://gongyeliu.github.io/StyleCrafter.github.io/
Gongye Liu, Menghan Xia, Yong Zhang 0034, Haoxin Chen, Jinbo Xing, Yibo Wang 0039, Xintao Wang 0002, Ying Shan, Yujiu Yang 0001
ACM Trans. Graph.9
2023 MvP: Multi-view Prompting Improves Aspect Sentiment Tuple Prediction
abstract
Generative methods greatly promote aspectbased sentiment analysis via generating a sequence of sentiment elements in a specified format.However, existing studies usually predict sentiment elements in a fixed order, which ignores the effect of the interdependence of the elements in a sentiment tuple and the diversity of language expression on the results.In this work, we propose Multi-view Prompting (MVP) that aggregates sentiment elements generated in different orders, leveraging the intuition of human-like problem-solving processes from different views.Specifically, MVP introduces element order prompts to guide the language model to generate multiple sentiment tuples, each with a different element order, and then selects the most reasonable tuples by voting.MVP can naturally model multi-view and multi-task as permutations and combinations of elements, respectively, outperforming previous task-specific designed methods on multiple ABSA tasks with a single model.Extensive experiments show that MVP significantly advances the state-of-the-art performance on 10 datasets of 4 benchmark tasks, and performs quite effectively in low-resource settings.Detailed evaluation verified the effectiveness, flexibility, and cross-task transferability of MVP. 1 * Equal contribution.
Zhibin Gou, Qingyan Guo, Yujiu Yang 0001
ACL (1)3
2023 Solving Math Word Problems via Cooperative Reasoning induced Language Models
abstract
Xinyu Zhu, Junjie Wang, Lin Zhang, Yuxiang Zhang, Yongfeng Huang, Ruyi Gan, Jiaxing Zhang, Yujiu Yang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Junjie Wang 0011, Yongfeng Huang 0001, Ruyi Gan, Jiaxing Zhang 0001, Yujiu Yang 0001
ACL (1)8
2023 GLeaD: Improving GANs with A Generator-Leading Task
abstract
Generative adversarial network (GAN) is formulated as a two-player game between a generator (G) and a discriminator (D), where D is asked to differentiate whether an image comes from real data or is produced by G. Under such a formulation, D plays as the rule maker and hence tends to dominate the competition. Towards a fairer game in GANs, we propose a new paradigm for adversarial training, which makes G assign a task to D as well. Specifically, given an image, we expect D to extract representative features that can be adequately decoded by G to reconstruct the input. That way, instead of learning freely, D is urged to align with the view of G for domain classification. Experimental results on various datasets demonstrate the substantial superiority of our approach over the baselines. For instance, we improve the FID of StyleGAN2 from 4.30 to 2.55 on LSUN Bedroom and from 4.04 to 2.82 on LSUN Church. We believe that the pioneering attempt present in this work could inspire the community with better designed generator-leading tasks for GAN improvement. Project page is at https://ezioby.github.io/glead/.
Qingyan Bai, Ceyuan Yang, Yinghao Xu 0001, Xihui Liu, Yujiu Yang 0001, Yujun Shen
CVPR5
2023 Seeing What You Miss: Vision-Language Pre-training with Semantic Completion Learning
abstract
Cross-modal alignment is essential for vision-language pre-training (VLP) models to learn the correct corresponding information across different modalities. For this purpose, inspired by the success of masked language modeling (MLM) tasks in the NLP pre-training area, numerous masked modeling tasks have been proposed for VLP to further promote cross-modal interactions. The core idea of previous masked modeling tasks is to focus on reconstructing the masked tokens based on visible context for learning local-to-local alignment. However, most of them pay little attention to the global semantic features generated for the masked data, resulting in a limited cross-modal alignment ability of global representations. Therefore, in this paper, we propose a novel Semantic Completion Learning (SCL) task, complementary to existing masked modeling tasks, to facilitate global-to-local alignment. Specifically, the SCL task complements the missing semantics of masked data by capturing the corresponding information from the other modality, promoting learning more representative global features which have a great impact on the performance of downstream tasks. Moreover, we present a flexible vision encoder, which enables our model to perform image-text and video-text multimodal tasks simultaneously. Experimental results show that our proposed method obtains state-of-the-art performance on various vision-language benchmarks, such as visual question answering, image-text retrieval, and video-text retrieval.
Yatai Ji, Rongcheng Tu, Jie Jiang 0015, Weijie Kong, Chengfei Cai, Hongfa Wang, Yujiu Yang 0001, Wei Liu 0005
CVPR8
2023 MAP: Multimodal Uncertainty-Aware Vision-Language Pre-training Model
abstract
Multimodal semantic understanding often has to deal with uncertainty, which means the obtained messages tend to refer to multiple targets. Such uncertainty is problematic for our interpretation, including inter- and intra-modal uncertainty. Little effort has studied the modeling of this uncertainty, particularly in pretraining on unlabeled datasets and fine-tuning in task-specific downstream datasets. In this paper, we project the representations of all modalities as probabilistic distributions via a Probability Distribution Encoder (PDE) by utilizing sequence-level interactions. Compared to the existing deterministic methods, such uncertainty modeling can convey richer multimodal semantic information and more complex relationships. Furthermore, we integrate uncertainty modeling with popular pretraining frameworks and propose suitable pretraining tasks: Distribution-based Vision-Language Contrastive learning (D-VLC), Distribution-based Masked Language Modeling (D-MLM), and Distribution-based Image-Text Matching (D-ITM). The fine-tuned models are applied to challenging downstream tasks, including image-text retrieval, visual question answering, visual reasoning, and visual entailment, and achieve state-of-the-art results.
Yatai Ji, Junjie Wang 0011, Yuan Gong 0002, Yanru Zhu, Hongfa Wang, Jiaxing Zhang 0001, Tetsuya Sakai, Yujiu Yang 0001
CVPR9
2023 RIFormer: Keep Your Vision Backbone Effective But Removing Token Mixer
abstract
This paper studies how to keep a vision backbone effective while removing token mixers in its basic building blocks. Token mixers, as self-attention for vision transformers (ViTs), are intended to perform information communication between different spatial tokens but suffer from considerable computational cost and latency. However, directly removing them will lead to an incomplete model structure prior, and thus brings a significant accuracy drop. To this end, we first develop an RepIdentityFormer base on the re-parameterizing idea, to study the token mixer free model architecture. And we then explore the improved learning paradigm to break the limitation of simple token mixer free backbone, and summarize the empirical practice into 5 guidelines. Equipped with the proposed optimization strategy, we are able to build an extremely simple vision backbone with encouraging performance, while enjoying the high efficiency during inference. Extensive experiments and ablative analysis also demonstrate that the inductive bias of network architecture, can be incorporated into simple network structure with appropriate optimization strategy. We hope this work can serve as a starting point for the exploration of optimization-driven efficient network design.
Jiahao Wang 0005, Songyang Zhang 0001, Yong Liu 0033, Taiqiang Wu, Yujiu Yang 0001, Xihui Liu, Kai Chen 0026, Ping Luo 0002, Dahua Lin
CVPR5
2023 3D GAN Inversion with Facial Symmetry Prior
abstract
Recently, a surge of high-quality 3D-aware GANs have been proposed, which leverage the generative power of neural rendering. It is natural to associate 3D GANs with GAN inversion methods to project a real image into the generator's latent space, allowing free-view consistent synthesis and editing, referred as 3D GAN inversion. Although with the facial prior preserved in pre-trained 3D GANs, reconstructing a 3D portrait with only one monocular image is still an ill-pose problem. The straightforward application of 2D GAN inversion methods focuses on texture similarity only while ignoring the correctness of 3D geometry shapes. It may raise geometry collapse effects, especially when reconstructing a side face under an extreme pose. Besides, the synthetic results in novel views are prone to be blurry. In this work, we propose a novel method to promote 3D GAN inversion by introducing facial symmetry prior. We design a pipeline and constraints to make full use of the pseudo auxiliary view obtained via image flipping, which helps obtain a view-consistent and well-structured geometry shape during the inversion process. To enhance texture fidelity in unobserved viewpoints, pseudo labels from depth-guided 3D warping can provide extra supervision. We design constraints to filter out conflict areas for optimization in asymmetric situations. Comprehensive quantitative and qualitative evaluations on image reconstruction and editing demonstrate the superiority of our method.
Yong Zhang 0034, Xuan Wang 0009, Tengfei Wang 0002, Xiaoyu Li 0002, Yuan Gong 0002, Yanbo Fan, Xiaodong Cun, Ying Shan, A. Cengiz Öztireli, Yujiu Yang 0001
CVPR11
2023 Specialist or Generalist? Instruction Tuning for Specific NLP Tasks
abstract
The potential of large language models (LLMs) to simultaneously perform a wide range of natural language processing (NLP) tasks has been the subject of extensive research.Although instruction tuning has proven to be a data-efficient method for transforming LLMs into such generalist models, their performance still lags behind specialist models trained exclusively for specific tasks.In this paper, we investigate whether incorporating broadcoverage generalist instruction tuning can contribute to building a specialist model.We hypothesize that its efficacy depends on task specificity and skill requirements.Our experiments assess four target tasks with distinct coverage levels, revealing that integrating generalist instruction tuning consistently enhances model performance when the task coverage is broad.The effect is particularly pronounced when the amount of task-specific training data is limited.Further investigation into three target tasks focusing on different capabilities demonstrates that generalist instruction tuning improves understanding and reasoning abilities.However, for tasks requiring factual knowledge, generalist data containing hallucinatory information may negatively affect the model's performance.Overall, our work provides a systematic guide for developing specialist models with general instruction tuning.Our code and other related resources can be found at https://github.com/DavidFanzz/ Generalist_or_Specialist.
Chufan Shi, Yixuan Su, Cheng Yang 0002, Yujiu Yang 0001, Deng Cai 0002
EMNLP4
2023 Question Answering as Programming for Solving Time-Sensitive Questions
abstract
Question answering plays a pivotal role in human daily life because it involves our acquisition of knowledge about the world.However, due to the dynamic and ever-changing nature of real-world facts, the answer can be completely different when the time constraint in the question changes.Recently, Large Language Models (LLMs) have shown remarkable intelligence in question answering, while our experiments reveal that the aforementioned problems still pose a significant challenge to existing LLMs.This can be attributed to the LLMs' inability to perform rigorous reasoning based on surfacelevel text semantics.To overcome this limitation, rather than requiring LLMs to directly answer the question, we propose a novel approach where we reframe the Question Answering task as Programming (QAaP).Concretely, by leveraging modern LLMs' superior capability in understanding both natural language and programming language, we endeavor to harness LLMs to represent diversely expressed text as wellstructured code and select the best matching answer from multiple candidates through programming.We evaluate our QAaP framework on several time-sensitive question answering datasets and achieve decent improvement, up to 14.5% over strong baselines.1
Cheng Yang 0007, Bei Chen 0008, Siheng Li, Jian-Guang Lou, Yujiu Yang 0001
EMNLP6
2023 Recouple Event Field via Probabilistic Bias for Event Extraction
abstract
Event Extraction (EE), aiming to identify and classify event triggers and arguments from event mentions, has benefited from pre-trained language models (PLMs). However, existing PLM-based methods ignore the information of trigger/argument fields, which is crucial for understanding event schemas. To this end, we propose a Probabilistic reCoupling model enhanced Event extraction framework (ProCE). Specifically, we first model the syntactic-related event fields as probabilistic biases, to clarify the event fields from ambiguous entanglement. Furthermore, considering multiple occurrences of the same triggers/arguments in EE, we explore probabilistic interaction strategies among multiple fields of the same triggers/arguments, to recouple the corresponding clarified distributions and capture more latent information fields. Experiments on EE datasets demonstrate the effectiveness and generalization of our proposed approach.
Xingyu Bai, Taiqiang Wu, Zhe Zhao 0006, Xuefeng Yang, Jiayi Li 0002, Weijie Liu 0002, Qi Ju 0002, Weigang Guo, Yujiu Yang 0001
ICASSP10
2023 A Two-Branch Network for Video Anomaly Detection with Spatio-Temporal Feature Learning
abstract
Video anomaly detection is very challenging, as most anomalies are rare and inconclusive. Previous weakly supervised learning approaches utilize the classifier trained with video-level labels to locate anomalous clips from the video. However, the anomalous clips often contain both anomalies and numerous irrelevant background behaviors, increasing the difficulty of localization. In this work, we propose a two-branch network to obtain the global and each local object’s action information of the clip respectively, where the local objects are extracted by a pre-trained object detector. This local-cum-global perception highlights the anomalous features from the background noise. We further propose a spatio-temporal relationship network, which is based on the attention mechanism to model the spatial relations of different objects and the temporal correlations among different clips to efficiently capture the spatio-temporal distribution of anomalies in the video. Extensive experiments on two benchmarks show that our method achieves significant performance gains.
Guoqiu Li, Shengjie Chen, Yujiu Yang 0001, Zhenhua Guo 0001
ICASSP3
2023 Syngen: A Syntactic Plug-And-Play Module for Generative Aspect-Based Sentiment Analysis
abstract
Aspect-based Sentiment Analysis (ABSA) is a sentiment analysis task at fine-grained level. Recently, generative frameworks have attracted increasing attention in ABSA due to their ability to unify subtasks and their continuity to upstream pre-training tasks. However, these generative models suffer from the neighboring dependency problem that induces neighboring words to get higher attention. In this paper, we propose SynGen, a plug-and-play syntactic information aware module. As a plug-in module, our SynGen can be easily applied to any generative framework backbones. The key insight of our module is to add syntactic inductive bias to attention assignment and thus direct attention to the correct target words. To the best of our knowledge, we are the first ones to introduce syntactic information to generative ABSA frameworks. Our module design is based on two main principles: (1) maintaining the structural integrity of backbone PLMs and (2) disentangling the added syntactic information and original semantic information. Empirical results on four popular ABSA datasets demonstrate that Syn-Gen enhanced model achieves a comparable performance to the state-of-the-art model with relaxed labeling specification and less training consumption.
Chengze Yu, Taiqiang Wu, Jiayi Li 0002, Xingyu Bai, Yujiu Yang 0001
ICASSP5
2023 ToonTalker: Cross-Domain Face Reenactment
abstract
We target cross-domain face reenactment in this paper, i.e., driving a cartoon image with the video of a real person and vice versa. Recently, many works have focused on one-shot talking face generation to drive a portrait with a real video, i.e., within-domain reenactment. Straightforwardly applying those methods to cross-domain animation will cause inaccurate expression transfer, blur effects, and even apparent artifacts due to the domain shift between cartoon and real faces. Only a few works attempt to settle cross-domain face reenactment. The most related work AnimeCeleb [13] requires constructing a dataset with pose vector and cartoon image pairs by animating 3D characters, which makes it inapplicable anymore if no paired data is available. In this paper, we propose a novel method for cross-domain reenactment without paired data. Specifically, we propose a transformer-based framework to align the motions from different domains into a common latent space where motion transfer is conducted via latent code addition. Two domain-specific motion encoders and two learnable motion base memories are used to capture domain properties. A source query transformer and a driving one are exploited to project domain-specific motion to the canonical space. The edited motion is projected back to the domain of the source with a transformer. Moreover, since no paired data is provided, we propose a novel cross-domain training scheme using data from two domains with the designed analogy constraint. Besides, we contribute a cartoon dataset in Disney style. Extensive evaluations demonstrate the superiority of our method over competing methods.
Yuan Gong 0002, Yong Zhang 0034, Xiaodong Cun, Yanbo Fan, Xuan Wang 0009, Baoyuan Wu, Yujiu Yang 0001
ICCV8
2023 Global Knowledge Calibration for Fast Open-Vocabulary Segmentation
abstract
Recent advancements in pre-trained vision-language models, such as CLIP, have enabled the segmentation of arbitrary concepts solely from textual inputs, a process commonly referred to as open-vocabulary semantic segmentation (OVS). However, existing OVS techniques confront a fundamental challenge: the trained classifier tends to over-fit on the base classes observed during training, resulting in suboptimal generalization performance to unseen classes. To mitigate this issue, recent studies have proposed the use of an additional frozen pre-trained CLIP for classification. Nonetheless, this approach incurs heavy computational overheads as the CLIP vision encoder must be repeatedly forward-passed for each mask, rendering it impractical for real-world applications. To address this challenge, our objective is to develop a fast OVS model that can perform comparably or better without the extra computational burden of the CLIP image encoder during inference. To this end, we propose a core idea of preserving the generalizable representation when fine-tuning on known classes. Specifically, we introduce a text diversification strategy that generates a set of synonyms for each training category, which prevents the learned representation from collapsing onto specific known category names. Additionally, we employ a text-guided knowledge distillation method to preserve the generalizable knowledge of CLIP. Extensive experiments demonstrate that our proposed model achieves robust generalization performance across various datasets. Furthermore, we perform a preliminary exploration of open-vocabulary video segmentation and present a benchmark that can facilitate future open-vocabulary research in the video domain.
Kunyang Han, Yong Liu 0033, Jun Hao Liew, Henghui Ding, Yansong Tang, Yujiu Yang 0001, Jiashi Feng, Yao Zhao 0001, Yunchao Wei
ICCV8
2023 UniKD: Universal Knowledge Distillation for Mimicking Homogeneous or Heterogeneous Object Detectors
abstract
Knowledge distillation (KD) has become a standard method to boost the performance of lightweight object detectors. Most previous works are feature-based, where students mimic the features of homogeneous teacher detectors. However, distilling the knowledge from the heterogeneous teacher fails in this manner due to the serious semantic gap, which greatly limits the flexibility of KD in practical applications. Bridging this semantic gap now requires case-by-case algorithm design which is time-consuming and heavily relies on experienced adjustment. To alleviate this problem, we propose Universal Knowledge Distillation (UniKD), introducing additional decoder heads with deformable cross-attention called Adaptive Knowledge Extractor (AKE). In UniKD, AKEs are first pretrained on the teacher’s output to infuse the teacher’s content and positional knowledge into a fixed-number set of knowledge embeddings. The fixed AKEs are then attached to the student’s backbone to encourage the student to absorb the teacher’s knowledge in these knowledge embeddings. In this query-based distillation paradigm, detection-relevant information can be dynamically aggregated into a knowledge embedding set and transferred between different detectors. When the teacher model is too large for online inference, its output can be stored on disk in advance to save the computation overhead, which is more storage efficient than feature-based methods. Extensive experiments demonstrate that our UniKD can plug and play in any homogeneous or heterogeneous teacher-student pairs and significantly outperforms conventional feature-based KD.
Shanshan Lao, Guanglu Song, Boxiao Liu, Yu Liu 0015, Yujiu Yang 0001
ICCV5
2023 Masked Autoencoders Are Stronger Knowledge Distillers
abstract
Knowledge distillation (KD) has shown great success in improving student’s performance by mimicking the intermediate output of the high-capacity teacher in fine-grained visual tasks, e.g. object detection. This paper proposes a technique called Masked Knowledge Distillation (MKD) that enhances this process using a masked autoencoding scheme. In MKD, random patches of the input image are masked, and the corresponding missing feature is recovered by forcing it to imitate the output of the teacher. MKD is based on two core designs. First, using the student as the encoder, we develop an adaptive decoder architecture, which includes a spatial alignment module that operates on the multi-scale features in the feature pyramid network (FPN) [20], a simple decoder, and a spatial recovery module that mimics the teacher’s output from the latent representation and mask tokens. Second, we introduce the masked convolution in each convolution block to keep the masked patches unaffected by others. By coupling these two designs, we can further improve the completeness and effectiveness of teacher knowledge learning. We conduct extensive experiments on different architectures with object detection and semantic segmentation. The results show that all the students can achieve further improvements compared to the conventional KD. Notably, we establish the new state-of-the-art results by boosting RetinaNet ResNet-18, and ResNet-50 from 33.4 to 37.5 mAP, and 37.4 to 41.5 mAP, respectively.
Shanshan Lao, Guanglu Song, Boxiao Liu, Yu Liu 0015, Yujiu Yang 0001
ICCV5
2023 D2Match: Leveraging Deep Learning and Degeneracy for Subgraph Matching
abstract
Subgraph matching is a fundamental building block for graph-based applications and is challenging due to its high-order combinatorial nature. Existing studies usually tackle it by combinatorial optimization or learning-based methods. However, they suffer from exponential computational costs or searching the matching without theoretical guarantees. In this paper, we develop $D^2$Match by leveraging the efficiency of Deep learning and Degeneracy for subgraph matching. More specifically, we first prove that subgraph matching can degenerate to subtree matching, and subsequently is equivalent to finding a perfect matching on a bipartite graph. We can then yield an implementation of linear time complexity by the built-in tree-structured aggregation mechanism on graph neural networks. Moreover, circle structures and node attributes can be easily incorporated in $D^2$Match to boost the matching performance. Finally, we conduct extensive experiments to show the superior performance of our $D^2$Match and confirm that our $D^2$Match indeed exploits the subtrees and differs from existing GNNs-based subgraph matching methods that depend on memorizing the data distribution divergence.
Xuanzhou Liu, Yujiu Yang 0001, Haiqin Yang
ICML4
2023 Feature Expansion for Graph Neural Networks
abstract
Graph neural networks aim to learn representations for graph-structured data and show impressive performance in node classification. Recently, many methods have studied the representations of GNNs from the perspective of optimization goals and spectral graph theory. However, the feature space that dominates representation learning has not been systematically studied in graph neural networks. In this paper, we propose to fill this gap by analyzing the feature space of both spatial and spectral models. We decompose graph neural networks into determined feature spaces and trainable weights, providing the convenience of studying the feature space explicitly using matrix space analysis. In particular, we find theoretically that the feature space tends to be linearly correlated due to repeated aggregations. In this case, the feature space is bounded by the poor representation of shared weights or the limited dimensionality of node attributes in existing models, leading to poor performance. Motivated by these findings, we propose 1) feature subspaces flattening and 2) structural principal components to expand the feature space. Extensive experiments verify the effectiveness of our proposed more comprehensive feature space, with comparable inference time to the baseline, and demonstrate its efficient convergence capability.
Guangyi Chen 0002, Kun Zhang 0001, Yujiu Yang 0001
ICML6
2023 SOC: Semantic-Assisted Object Cluster for Referring Video Object Segmentation
abstract
This paper studies referring video object segmentation (RVOS) by boosting video-level visual-linguistic alignment. Recent approaches model the RVOS task as a sequence prediction problem and perform multi-modal interaction as well as segmentation for each frame separately. However, the lack of a global view of video content leads to difficulties in effectively utilizing inter-frame relationships and understanding textual descriptions of object temporal variations. To address this issue, we propose Semantic-assisted Object Cluster (SOC), which aggregates video content and textual guidance for unified temporal modeling and cross-modal alignment. By associating a group of frame-level object embeddings with language tokens, SOC facilitates joint space learning across modalities and time steps. Moreover, we present multi-modal contrastive supervision to help construct well-aligned joint space at the video level. We conduct extensive experiments on popular RVOS benchmarks, and our method outperforms state-of-the-art competitors on all benchmarks by a remarkable margin. Besides, the emphasis on temporal coherence enhances the segmentation stability and adaptability of our method in processing text expressions with temporal variations. Code is available at https://github.com/RobertLuo1/NeurIPS2023_SOC.
Zhuoyan Luo, Yicheng Xiao, Yong Liu 0033, Shuyan Li, Yansong Tang, Xiu Li 0001, Yujiu Yang 0001
NeurIPS8
2023 Assessor360: Multi-sequence Network for Blind Omnidirectional Image Quality Assessment
abstract
Blind Omnidirectional Image Quality Assessment (BOIQA) aims to objectively assess the human perceptual quality of omnidirectional images (ODIs) without relying on pristine-quality image information. It is becoming more significant with the increasing advancement of virtual reality (VR) technology. However, the quality assessment of ODIs is severely hampered by the fact that the existing BOIQA pipeline lacks the modeling of the observer's browsing process. To tackle this issue, we propose a novel multi-sequence network for BOIQA called Assessor360, which is derived from the realistic multi-assessor ODI quality assessment procedure. Specifically, we propose a generalized Recursive Probability Sampling (RPS) method for the BOIQA task, combining content and details information to generate multiple pseudo viewport sequences from a given starting point. Additionally, we design a Multi-scale Feature Aggregation (MFA) module with a Distortion-aware Block (DAB) to fuse distorted and semantic features of each viewport. We also devise Temporal Modeling Module (TMM) to learn the viewport transition in the temporal domain. Extensive experimental results demonstrate that Assessor360 outperforms state-of-the-art methods on multiple OIQA datasets. The code and models are available at https://github.com/TianheWu/Assessor360.
Tianhe Wu, Shuwei Shi, Haoming Cai, Mingdeng Cao, Jing Xiao 0006, Yinqiang Zheng, Yujiu Yang 0001
NeurIPS7
2023 Interactive Story Visualization with Multiple Characters
abstract
Accurate Story visualization requires several necessary elements, such as identity consistency across frames, the alignment between plain text and visual content, and a reasonable layout of objects in images. Most previous works endeavor to meet these requirements by fitting a text-to-image (T2I) model on a set of videos in the same style and with the same characters, e.g., the FlintstonesSV dataset. However, the learned T2I models typically struggle to adapt to new characters, scenes, and styles, and often lack the flexibility to revise the layout of the synthesized images. This paper proposes a system for generic interactive story visualization, capable of handling multiple novel characters and supporting the editing of layout and local structure. It is developed by leveraging the prior knowledge of large language and T2I models, trained on massive corpora. The system comprises four interconnected components: story-to-prompt generation (S2P), text-to-layout generation (T2L), controllable text-to-image generation (C-T2I), and image-to-video animation (I2V). First, the S2P module converts concise story information into detailed prompts required for subsequent stages. Next, T2L generates diverse and reasonable layouts based on the prompts, offering users the ability to adjust and refine the layout to their preferences. The core component, C-T2I, enables the creation of images guided by layouts, sketches, and actor-specific identifiers to maintain consistency and detail across visualizations. Finally, I2V enriches the visualization process by animating the generated images. Extensive experiments and a user study are conducted to validate the effectiveness and flexibility of interactive editing of the proposed system.
Yuan Gong 0002, Youxin Pang, Xiaodong Cun, Menghan Xia, Yingqing He, Haoxin Chen, Longyue Wang, Yong Zhang 0034, Xintao Wang 0002, Ying Shan, Yujiu Yang 0001
SIGGRAPH Asia11
2023 Multi-Granularity Interest Learning for Click-Through Rate Prediction
abstract
Embedding&MLP paradigm based on deep learning has been widely used in click-through rate prediction tasks. Such approaches compress user features into a fixed-length vector, which causes the bottleneck of interest learning and makes it difficult to express diverse interests of users. To improve the accuracy and diversity of recommendation, more comprehensive user interest learning is required. We present Multi-Granularity Interest Learning(MGIL)—combining the memorization of user's historical preferences with generalization of user's interests. Our fine-grained interest module learns interests at the granularity of items, preserving the memory of users' historical preferences as much as possible. The coarse-grained interest module, on the other hand, extracts multiple high-level interests from user behavior sequences to characterize multiple aspects of user interest, which are more generalized and abstract. The two modules complement each other from different granularities. We conduct extensive experiments on three real-world datasets, Movielens, Taobao and Amazon. Experimental results demonstrate that our method significantly improves the accuracy and diversity of recommendation compared to state-of-the-art models.
Chang Niu, Yujiu Yang 0001
SMC3
2023 Modeling Fine-grained Information via Knowledge-aware Hierarchical Graph for Zero-shot Entity Retrieval
abstract
Zero-shot entity retrieval, aiming to link mentions to candidate entities under the zero-shot setting, is vital for many tasks in Natural Language Processing. Most existing methods represent mentions/entities via the sentence embeddings of corresponding context from the Pre-trained Language Model. However, we argue that such coarse-grained sentence embeddings can not fully model the mentions/entities, especially when the attention scores towards mentions/entities are relatively low. In this work, we propose GER, a Graph enhanced Entity Retrieval framework, to capture more fine-grained information as complementary to sentence embeddings. We extract the knowledge units from the corresponding context and then construct a mention/entity centralized graph. Hence, we can learn the fine-grained information about mention/entity by aggregating information from these knowledge units. To avoid the graph bottleneck for the central mention/entity node, we construct a hierarchical graph and design a novel Hierarchical Graph Attention Network~(HGAN). Experimental results on popular benchmarks demonstrate that our proposed GER framework performs better than previous state-of-the-art models.
Taiqiang Wu, Xingyu Bai, Weigang Guo, Weijie Liu 0002, Siheng Li, Yujiu Yang 0001
WSDM6
2023 GAN Inversion: A Survey
abstract
GAN inversion aims to invert a given image back into the latent space of a pretrained GAN model so that the image can be faithfully reconstructed from the inverted code by the generator. As an emerging technique to bridge the real and fake image domains, GAN inversion plays an essential role in enabling pretrained GAN models, such as StyleGAN and BigGAN, for applications of real image editing. Moreover, GAN inversion interprets GAN's latent space and examines how realistic images can be generated. In this paper, we provide a survey of GAN inversion with a focus on its representative algorithms and its applications in image restoration and image manipulation. We further discuss the trends and challenges for future research. A curated list of GAN inversion methods, datasets, and other related information can be found at https://github.com/weihaox/awesome-gan-inversion.
Weihao Xia 0001, Yulun Zhang 0001, Yujiu Yang 0001, Jing-Hao Xue, Bolei Zhou, Ming-Hsuan Yang 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2023 A simple and effective patch-Based method for frame-level face anti-spoofing
Shengjie Chen, Yujiu Yang 0001, Zhenhua Guo 0001
Pattern Recognit. Lett.3
2023 VDTR: Video Deblurring With Transformer
abstract
Video deblurring is still an unsolved problem due to the challenging spatio-temporal modeling process. While existing convolutional neural network (CNN)-based methods show a limited capacity of effective spatial and temporal modeling for video deblurring. This paper presents VDTR, an effective Transformer-based model that makes the first attempt to adapt pure Transformer for video deblurring. VDTR exploits the superior long-range and relation modeling capabilities of Transformer for both spatial and temporal modeling. However, it is challenging to design an appropriate Transformer-based model for video deblurring due to the complicated non-uniform blurs, misalignment across multiple frames and the high computational costs for high-resolution spatial modeling. To address these problems, VDTR advocates performing attention within non-overlapping windows and exploiting the hierarchical structure for long-range dependencies modeling. For frame-level spatial modeling, we propose an encoder-decoder Transformer that utilizes multi-scale features for deblurring. For multi-frame temporal modeling, we adapt Transformer to fuse multiple spatial features efficiently. Compared with CNN-based methods, the proposed method achieves highly competitive results on both synthetic and real-world video deblurring benchmarks, including DVD, GOPRO, REDS and BSD. We hope such a Transformer-based architecture can serve as a powerful alternative baseline for video deblurring and other video restoration tasks. The source code will be available athttps://github.com/ljzycmd/VDTR.
Mingdeng Cao, Yanbo Fan, Yong Zhang 0034, Jue Wang 0001, Yujiu Yang 0001
IEEE Trans. Circuits Syst. Video Technol.5
2022 MCSCSet: A Specialist-annotated Dataset for Medical-domain Chinese Spelling Correction
abstract
Chinese Spelling Correction (CSC) is gaining increasing attention in recent years. Despite its extensive use in many applications, such as search engine and optical character recognition system, little has been explored in medical scenarios in which complex and uncommon medical entities are easily misspelled. Correcting the misspellings of medical entities is arguably more difficult than those in the open domain due to its requirements of specific domain knowledge. In this work, we define the task of Medical-domain Chinese Spelling Correction (MCSC) and propose MCSCSet, a large-scale specialist-annotated dataset that contains about 200k samples. In contrast to existing open-domain CSC datasets, MCSCSet involves: i) extensive real-world medical queries collected from Tencent Yidian, ii) corresponding misspelled sentences manually annotated by medical specialists. Our work further offers a medical-domain confusion set consisting of the common error-prone characters in medicine and their corresponding misspellings. Extensive empirical studies have shown significant gaps between the open-domain and medical-domain spelling correction, highlighting the need to develop high-quality datasets that allow for CSC in specific domains. Moreover, our work benchmarks several representative methods, establishing baselines for future work.
Wangjie Jiang, Zhihao Ye, Zijing Ou, Ruihui Zhao, Jianguang Zheng, Yi Liu 0057, Bang Liu 0003, Siheng Li, Yujiu Yang 0001, Yefeng Zheng 0001
CIKM9
2022 Learning Adaptive Warping for RealWorld Rolling Shutter Correction
abstract
This paper proposes the first real-world rolling shutter (RS) correction dataset, BS-RSC, and a corresponding model to correct the RS frames in a distorted video. Mobile devices in the consumer market with CMOS-based sensors for video capture often result in rolling shutter effects when relative movements occur during the video acquisition process, calling for RS effect removal techniques. However, current state-of-the-art RS correction methods often fail to remove RS effects in real scenarios since the motions are various and hard to model. To address this issue, we propose a real-world RS correction dataset BS-RSC. Real distorted videos with corresponding ground truth are recorded simultaneously via a well-designed beam-splitter-based acquisition system. BS-RSC contains various motions of both camera and objects in dynamic scenes. Further, an RS correction model with adaptive warping is proposed. Our model can warp the learned RS features into global shutter counterparts adaptively with predicted multiple displacement fields. These warped features are aggregated and then reconstructed into high-quality global shutter frames in a coarse-to-fine strategy. Experimental results demonstrate the effectiveness of the proposed method, and our dataset can improve the model's ability to remove the RS effects in the real world. The project is available at https://github.com/ljzycmd/BSRSC.
Mingdeng Cao, Zhihang Zhong, Jiahao Wang 0005, Yinqiang Zheng, Yujiu Yang 0001
CVPR5
2022 High-Fidelity GAN Inversion with Padding Space
Qingyan Bai, Yinghao Xu 0001, Jiapeng Zhu 0001, Weihao Xia 0001, Yujiu Yang 0001, Yujun Shen
ECCV (15)5
2022 Global Spectral Filter Memory Network for Video Object Segmentation
Yong Liu 0033, Jiahao Wang 0005, Yansong Tang, Yujiu Yang 0001
ECCV (29)7
2022 Learning Quality-aware Dynamic Memory for Video Object Segmentation
Yong Liu 0033, Wei Zhao 0013, Weihao Xia 0001, Yujiu Yang 0001
ECCV (29)7
2022 StyleHEAT: One-Shot High-Resolution Editable Talking Face Generation via Pre-trained StyleGAN
Yong Zhang 0034, Xiaodong Cun, Mingdeng Cao, Yanbo Fan, Xuan Wang 0009, Qingyan Bai, Baoyuan Wu, Jue Wang 0001, Yujiu Yang 0001
ECCV (17)10
2022 Multi-Turn Incomplete Utterance Restoration As Object Detection
abstract
In this paper, we investigate the task of multi-turn incomplete utterance restoration to tackle the issue of frequent coreference and information omission in multi-turn dialogues. Recent works mainly focus on edit-based approaches which have been proven to outperform traditional generation-based models in terms of accuracy and efficiency. However, they only model token-level edit relationships while ignoring span-level edit relationships. Our experiments find this breaks the semantic integrity of edit span, which causes inaccurate edit span prediction and disfluent utterance restoration. To address the problem, we propose a novel approach to directly model span-level edit relationships between the incomplete utterance and context. Specifically, we build an edit matrix in which each rectangular region represents a span-level edit operation. Then, we detect the region with a well-designed dual-branch detection module inspired by object detection. Empirical results demonstrate that our method outperforms state-of-the-art methods significantly on two public datasets. In addition, further studies verify that our method is capable of preserving the semantic integrity of edit span.
Wangjie Jiang, Siheng Li, Jiayi Li 0002, Yujiu Yang 0001
ICASSP4
2022 Attention Probe: Vision Transformer Distillation in the Wild
abstract
Vision transformers (ViTs) require intensive computational resources to achieve high performance, which usually makes them not suitable for mobile devices. A feasible strategy is to compress them using the original training data, which may be not accessible due to privacy limitations or transmission restrictions. In this case, utilizing the massive unlabeled data in the wild is an alternative paradigm, which has been proved effective for compressing convolutional neural networks (CNNs). However, due to the significant differences in model structure and computation mechanism between CNNs and ViTs, it is still an open issue that whether the similar paradigm is suitable for ViTs. In this work, we propose to effectively compress ViTs using the unlabeled data in the wild, consisting of two stages. First, we design an effective tool in selecting valuable data from the wild, dubbed Attention Probe. Second, based on the selected data, we develop a probe knowledge distillation algorithm to train a lightweight student transformer, through maximizing the similarities on both the outputs and intermediate features, between the heavy teacher and the lightweight student models. Extensive experimental results on several benchmarks demonstrate that the student transformer obtained by the proposed method can achieve comparable performance with the baseline that requires the original training data. Code is available at: https://github.com/IIGROUP/AttentionProbe.
Jiahao Wang 0005, Mingdeng Cao, Shuwei Shi, Baoyuan Wu, Yujiu Yang 0001
ICASSP5
2022 Identity-Guided Face Generation with Multi-Modal Contour Conditions
abstract
Recent face generation methods have tried to synthesize faces based on the given contour condition, like a low-resolution image or sketch. However, the problem of identity ambiguity remains unsolved, which usually occurs when the contour is too vague to provide reliable identity information (e.g., when its resolution is extremely low). Thus feasible solutions of image restoration could be infinite. In this work, we propose a novel framework that takes the contour and an extra image specifying the identity as the inputs, where the contour can be of various modalities, including the low-resolution image, sketch, and semantic label map. Concretely, we propose a novel dual-encoder architecture, in which an identity encoder extracts the identity-related feature, accompanied by a main encoder to obtain the rough contour information and further fuse all the information together. The encoder output is iteratively fed into a pre-trained StyleGAN generator until getting a satisfying result. To the best of our knowledge, this is the first work that achieves identity-guided face generation conditioned on multi-modal contour images. Moreover, our method can produce photo-realistic results with 1024×1024 resolution.
Qingyan Bai, Weihao Xia 0001, Yujiu Yang 0001
ICIP4
2022 AACP: Model Compression by Accurate and Automatic Channel Pruning
abstract
Channel pruning is formulated as a neural architecture search (NAS) problem recently, which achieves impressive performance in model compression. However, prior arts only considered one kind of constraint (FLOPs, inference latency or model size) when pruning a neural network. This will lead to an unbalanced-pruning problem, where the FLOPs of a pruned network are under budget but the inference latency and model size are still unaffordable. Another challenge is that the supernet training process of typical NAS-based channel pruning methods is computationally expensive. To address these problems, we propose a novel Accurate and Automatic Channel Pruning (AACP) method. Firstly, we impose multiple constraints on channel pruning to address the unbalanced-pruning problem. To solve this complicated multi-objective problem, AACP proposes Improved Differential Evolution (IDE) algorithm which is more effective than typical evolutionary algorithms in searching for optimal architectures. Secondly, AACP proposes a Pruned Structure Accuracy Estimator (PSAE) which can estimate the performance of sub-networks without training a supernet and speeds up the performance estimation process. Our method achieves state-of-the-art performance on several benchmarks. On CIFAR10, our method reduces 65% FLOPs of ResNet110 with an improvement of 0.26% top-1 accuracy. On ImageNet, we reduce 42% FLOPs of ResNet50 with a small loss of 0.06% top-1 accuracy. The code is available at https://github.com/linlb11/AACP.
Lanbo Lin, Shengjie Chen, Yujiu Yang 0001, Zhenhua Guo 0001
ICPR3
2022 Augmenting Anchors by the Detector Itself
abstract
Usually, it is difficult to determine the scale and aspect ratio of anchors for anchor-based object detection methods. Current state-of-the-art object detectors either determine anchor parameters according to objects' shape and scale in a dataset, or avoid this problem by utilizing anchor-free methods, however, the former scheme is dataset-specific and the latter methods could not get better performance than the former ones. In this paper, we propose a novel anchor augmentation method named AADI, which means Augmenting Anchors by the Detector Itself. AADI is not an anchor-free method, instead, it can convert the scale and aspect ratio of anchors from a continuous space to a discrete space, which greatly alleviates the problem of anchors' designation. Furthermore, AADI is a learning-based anchor augmentation method, but it does not add any parameters or hyper-parameters, which is beneficial for research and downstream tasks. Extensive experiments on COCO dataset demonstrate the effectiveness of AADI, specifically, AADI achieves significant performance boosts on many state-of-the-art object detectors (eg. at least +2.4 box AP on Faster R-CNN, +2.2 box AP on Mask R-CNN, and +0.9 box AP on Cascade Mask R-CNN). We hope that this simple and cost-efficient method can be widely used in object detection. Code and models are available at https://github.com/WanXiaopei/aadi.
Xiaopei Wan, Guoqiu Li, Yujiu Yang 0001, Zhenhua Guo 0001
IJCAI3
2022 EmpHi: Generating Empathetic Responses with Human-like Intents
abstract
In empathetic conversations, humans express their empathy to others with empathetic intents.However, most existing empathetic conversational methods suffer from a lack of empathetic intents, which leads to monotonous empathy.To address the bias of the empathetic intents distribution between empathetic dialogue models and humans, we propose a novel model to generate empathetic responses with humanconsistent empathetic intents, EmpHi for short.Precisely, EmpHi learns the distribution of potential empathetic intents with a discrete latent variable, then combines both implicit and explicit intent representation to generate responses with various empathetic intents.Experiments show that EmpHi outperforms state-ofthe-art models in terms of empathy, relevance, and diversity on both automatic and human evaluation.Moreover, the case studies demonstrate the high interpretability and outstanding performance of our model.Our code are avaliable at https://github.com/mattc95/EmpHi.
Mao Yan Chen, Siheng Li, Yujiu Yang 0001
NAACL-HLT3
2022 Rethinking Alignment in Video Super-Resolution Transformers
abstract
The alignment of adjacent frames is considered an essential operation in video super-resolution (VSR). Advanced VSR models, including the latest VSR Transformers, are generally equipped with well-designed alignment modules. However, the progress of the self-attention mechanism may violate this common sense. In this paper, we rethink the role of alignment in VSR Transformers and make several counter-intuitive observations. Our experiments show that: (i) VSR Transformers can directly utilize multi-frame information from unaligned videos, and (ii) existing alignment methods are sometimes harmful to VSR Transformers. These observations indicate that we can further improve the performance of VSR Transformers simply by removing the alignment module and adopting a larger attention window. Nevertheless, such designs will dramatically increase the computational burden, and cannot deal with large motions. Therefore, we propose a new and efficient alignment method called patch alignment, which aligns image patches instead of pixels. VSR Transformers equipped with patch alignment could demonstrate state-of-the-art performance on multiple benchmarks. Our work provides valuable insights on how multi-frame information is used in VSR and how to select alignment methods for different networks/datasets. Codes and models will be released at https://github.com/XPixelGroup/RethinkVSRAlignment.
Shuwei Shi, Jinjin Gu, Liangbin Xie, Xintao Wang 0002, Yujiu Yang 0001, Chao Dong 0005
NeurIPS5
2022 Dual Contrastive Learning for Unsupervised Knowledge Selection
abstract
Although dialogue systems based on the Seq2Seq model have achieved success, they suffer from tending to generate general responses.Recent works have shown that selecting external knowledge is helpful for dialogue systems to generate informative and diverse responses.However, selecting appropriate knowledge from an unlabeled knowledge set, which is referred to as unsupervised knowledge selection, remains a tricky challenge.Therefore, we propose a dual contrastive method, which utilizes two source-target pairs which are based on the same knowledge set to construct dual contrasts.Specifically, for a source utterance, we consider its paired and unpaired target response as a positive and negative sample, then obtain the positive and negative posterior distribution over the knowledge candidates set, respectively.Then we lead the prior distribution to be close to the positive posterior distribution and distant from the negative one.Similarly, the posterior distribution is treated with the same criterion.Experimental results show that our method improves generated responses in terms of BLUE, DISTINCT, and knowledge utilization.Our codes are available at https://github.com/CaoXiang1997/DualCL4UKS.
Yujiu Yang 0001
SEKE3
2022 A novel 2D contactless fingerprint matching method
Hao Gui, Yujiu Yang 0001, Zhenhua Guo 0001
Neurocomputing4
2022 Real-time human-centric segmentation for complex video scenes
Weihao Xia 0001, Yujiu Yang 0001
Image Vis. Comput.6
2021 TediGAN: Text-Guided Diverse Face Image Generation and Manipulation
abstract
In this work, we propose TediGAN, a novel framework for multi-modal image generation and manipulation with textual descriptions. The proposed method consists of three components: StyleGAN inversion module, visual-linguistic similarity learning, and instance-level optimization. The inversion module maps real images to the latent space of a well-trained StyleGAN. The visual-linguistic similarity learns the text-image matching by mapping the image and text into a common embedding space. The instancelevel optimization is for identity preservation in manipulation. Our model can produce diverse and high-quality images with an unprecedented resolution at 10242. Using a control mechanism based on style-mixing, our TediGAN inherently supports image synthesis with multi-modal inputs, such as sketches or semantic labels, with or without instance guidance. To facilitate text-guided multi-modal synthesis, we propose the Multi-Modal CelebA-HQ, a large-scale dataset consisting of real face images and corresponding semantic segmentation map, sketch, and textual descriptions. Extensive experiments on the introduced dataset demonstrate the superior performance of our proposed method. Code and data are available at https://github.com/weihaox/TediGAN.
Weihao Xia 0001, Yujiu Yang 0001, Jing-Hao Xue, Baoyuan Wu
CVPR2
2021 Probabilistic Modeling of Semantic Ambiguity for Scene Graph Generation
abstract
To generate "accurate" scene graphs, almost all existing methods predict pairwise relationships in a deterministic manner. However, we argue that visual relationships are often semantically ambiguous. Specifically, inspired by linguistic knowledge, we classify the ambiguity into three types: Synonymy Ambiguity, Hyponymy Ambiguity, and Multi-view Ambiguity. The ambiguity naturally leads to the issue of implicit multi-label, motivating the need for diverse predictions. In this work, we propose a novel plug-and-play Probabilistic Uncertainty Modeling (PUM) module. It models each union region as a Gaussian distribution, whose variance measures the uncertainty of the corresponding visual content. Compared to the conventional deterministic methods, such uncertainty modeling brings stochasticity of feature representation, which naturally enables diverse predictions. As a byproduct, PUM also manages to cover more fine-grained relationships and thus alleviates the issue of bias towards frequent relationships. Extensive experiments on the large-scale Visual Genome benchmark show that combining PUM with newly proposed ResCAGCN can achieve state-of-the-art performances, especially under the mean recall metric. Furthermore, we show the universal effectiveness of PUM by plugging it into some existing models and provide insightful analysis of its ability to generate diverse yet plausible visual relationships.
Gengcong Yang, Jingyi Zhang 0005, Yong Zhang 0034, Baoyuan Wu, Yujiu Yang 0001
CVPR5
2021 PoseDet: Fast Multi-Person Pose Estimation Using Pose Embedding
abstract
Current methods of multi-person pose estimation typically treat the localization and the association of body joints separately. It is convenient but inefficient, leading to additional computation and a waste of time. This paper, however, presents a novel framework PoseDet (Estimating Pose by Detection) to localize and associate body joints simultaneously at higher inference speed. Moreover, we propose the keypoint-aware pose embedding to represent an object in terms of the locations of its keypoints. The proposed pose embedding contains semantic and geometric information, allowing us to efficiently access discriminative and informative features. It is utilized for candidate classification and body joint localization in PoseDet, leading to robust predictions of various poses. This simple framework achieves an unprecedented speed and a competitive accuracy on the COCO benchmark compared with state-of-the-art methods. Extensive experiments on the CrowdPose benchmark show the robustness in the crowd scenes. Code is available at https://github.com/IIGROUP/PoseDet.
Weihao Xia 0001, Haoqian Wang, Yujiu Yang 0001
FG6
2021 More: A Metric Learning Based Framework for Open-Domain Relation Extraction
abstract
Open relation extraction (OpenRE) is the task of extracting relation schemes from open-domain corpora. Most existing OpenRE methods either do not fully benefit from high-quality labeled corpora or can not learn semantic representation directly, affecting downstream clustering efficiency. To address these problems, in this work, we propose a novel learning framework named MORE (Metric learning-based Open Relation Extraction). The framework utilizes deep metric learning to obtain rich supervision signals from labeled data and drive the neural model to learn semantic relational representation directly. Experiments result in two real-world datasets show that our method outperforms other state-of-the-art baselines. Our source code is available on Github1.
Renze Lou, Kai Zhang 0033, Mao Yan Chen, Yujiu Yang 0001
ICASSP5
2021 Adder Attention for Vision Transformer
abstract
Transformer is a new kind of calculation paradigm for deep learning which has shown strong performance on a large variety of computer vision tasks. However, compared with conventional deep models (e.g., convolutional neural networks), vision transformers require more computational resources which cannot be easily deployed on mobile devices. To this end, we present to reduce the energy consumptions using adder neural network (AdderNet). We first theoretically analyze the mechanism of self-attention and the difficulty for applying adder operation into this module. Specifically, the feature diversity, i.e., the rank of attention map using only additions cannot be well preserved. Thus, we develop an adder attention layer that includes an additional identity mapping. With the new operation, vision transformers constructed using additions can also provide powerful feature representations. Experimental results on several benchmarks demonstrate that the proposed approach can achieve highly competitive performance to that of the baselines while achieving an about 2~3× reduction on the energy consumption.
Han Shu, Jiahao Wang 0005, Hanting Chen, Yujiu Yang 0001, Yunhe Wang 0001
NeurIPS5
2021 Overview of the NLPCC 2021 Shared Task: AutoIE2
Weigang Guo, Xuefeng Yang, Xingyu Bai, Taiqiang Wu, Weijie Liu 0002, Zhe Zhao 0006, Qi Ju 0002, Yujiu Yang 0001
NLPCC (2)8
2021 HSCJN: A holistic semantic constraint joint network for diverse response generation
Pengda Si, Zeyang Lei, Guangxu Xun, Yujiu Yang 0001
Comput. Speech Lang.5
2021 Cali-sketch: Stroke calibration and completion for high-quality face image generation from human-like sketches
Weihao Xia 0001, Yujiu Yang 0001, Jing-Hao Xue
Neurocomputing2
2021 Guest Editorial: Special issue on deep learning with small samples
Jing-Hao Xue, Jufeng Yang, Yan Yan 0001, Yujiu Yang 0001, Zongqing Lu 0001, Zhanyu Ma
Neurocomputing5
2021 Domain Fingerprints for No-Reference Image Quality Assessment
abstract
Human fingerprints are detailed and nearly unique markers of human identity. Such a unique and stable fingerprint is also left on each acquired image. It can reveal how an image was degraded during the image acquisition procedure and thus is closely related to the quality of an image. In this work, we propose a new no-reference image quality assessment (NR-IQA) approach called domain-aware IQA (DA-IQA), which for the first time introduces the concept of domain fingerprint to the NR-IQA field. The domain fingerprint of an image is learned from image collections of different degradations and then used as the unique characteristics to identify the degradation sources and assess the quality of the image. To this end, we design a new domain-aware architecture, which enables simultaneous determination of both the distortion sources and the quality of an image. With the distortion in an image better characterized, the image quality can be more accurately assessed, as verified by extensive experiments, which show that the proposed DA-IQA performs better than almost all the compared state-of-the-art NR-IQA methods.
Weihao Xia 0001, Yujiu Yang 0001, Jing-Hao Xue, Jing Xiao 0006
IEEE Trans. Circuits Syst. Video Technol.2
2020 Topic Enhanced Controllable CVAE for Dialogue Generation (Student Abstract)
abstract
Neural generation models have shown great potential in conversation generation recently. However, these methods tend to generate uninformative or irrelevant responses. In this paper, we present a novel topic-enhanced controllable CVAE (TEC-CVAE) model to address this issue. On the one hand, the model learns the context-interactive topic knowledge through a novel multi-hop hybrid attention in the encoder. On the other hand, we design a topic-aware controllable decoder to constrain the expression of the stochastic latent variable in the CVAE to reduce irrelevant responses. Experimental results on two public datasets show that the two mechanisms synchronize to improve both relevance and diversity, and the proposed model outperforms other competitive methods.
Pengda Si, Zeyang Lei, Yujiu Yang 0001
AAAI4
2020 STCN: A Lightweight Sleep Staging Model with Multiple Channels
abstract
Sleep staging is an important means of diagnosing sleep disorders and monitoring sleep quality. Concurrently, the RNN (Recurrent Neural Network) models commonly used in this field limits the overall calculation efficiency and performance. Specifically, we utilize the TCN (Temporal Convolutional Network) model advanced in the audio field as an alternative for RNN models to improve temporal information collection capability. Leverage the SE (squeeze-and-extraction) module to merge the multi-lead information flexibly. Implement the feature extraction layer by the design of CNN to reduce computational complexity. The proposed model, named STCN, achieved 85.01% accuracy, 78.80% F1 value on the Sleep-EDFx dataset, and achieved 84.10% accuracy, 81.27% F1 value on the Physionet 2018 dataset. Only 1.6 million parameters are obtained in STCN, which is significantly smaller than the parameters of Deepsleepnet. The proposed model is compatible with scenarios from different datasets, achieves an adaptive fusion of each lead information and supplies a more reliable prediction for sleep quality assessment to provide a foundation for convenient applications in the future.
Yui Lo, Yujiu Yang 0001
BIBM3
2020 DT-QDC: A Dataset for Question Comprehension in Online Test
abstract
With the transformation of education from the traditional classroom environment to online education and assessment, it is more and more important to accurately assess the difficulty of questions than ever. As teachers may not be able to follow the student’s performance and learning behavior closely, a well-defined method to measure the difficulty of questions to guide learning is necessary. In this paper, we explore the concept of question difficulty and provide our new Chinese DT-QDC dataset. This is currently the largest and only Chinese question dataset, and it also has enriched attributes and difficulty labels. Additional attributes such as keywords, chapter, and question type would allow models to understand questions more precisely. We proposed the MTMS-BERT and ORMS-BERT, which can improve the judgment of difficulty from different views. The proposed methods outperforms different baselines by 7.79% on F1-score and 15.92% on MAE, 28.26% on MSE on the new DT-QDC dataset, laying the foundation for the question difficulty comprehension task.
Sijin Wu, Yujiu Yang 0001, Nicholas Yung, Zhengchen Shen, Zeyang Lei
COLING2
2020 Sparse Adversarial Attack via Perturbation Factorization
Yanbo Fan, Baoyuan Wu, Tuanhui Li, Yong Zhang 0034, Zhifeng Li 0001, Yujiu Yang 0001
ECCV (22)7
2020 Match4Rec: A Novel Recommendation Algorithm Based on Bidirectional Encoder Representation with the Matching Task
Lingxiao Zhang, Jiangpeng Yan, Yujiu Yang 0001, Xiu Li 0001
ICONIP (3)3
2020 Stock-UniBERT: A News-based Cost-sensitive Ensemble BERT Model for Stock Trading
abstract
Financial news plays an important role in investors' decisions and then influences stock markets. Previous studies mainly focus on establishing sentiment index from financial text and then making stock return prediction and trading strategy based on the index. This procedure demands costly manual label and may not directly correspond to actual stock market reaction. This paper solves this problem by using labels of stocks' residual return as sentiment labels for BERT model training. Distinct from ordinary task, buying or selling action will be taken after judgement of the stock news' sentiment. Hence, weighted cross-entropy loss and cost-sensitive accuracy are used to reveal influence and cost of judgement. Different settings of weighted cross-entropy loss are applied to learn self-adaptively and a selection method is designed to seek capable base classifiers for ensemble learning. This paper then develops a stock trading strategy based on the ensemble BERT model. Experiments and ablation study show the robust effectiveness of our strategy.
Xiliu Man, Jianwu Lin, Yujiu Yang 0001
INDIN3
2020 Cognitive Representation Learning of Self-Media Online Article Quality
abstract
The automatic quality assessment of self-media online articles is an urgent and new issue, which is of great value to the online recommendation and search. Different from traditional and well-formed articles, self-media online articles are mainly created by users, which have the appearance characteristics of different text levels and multi-modal hybrid editing, along with the potential characteristics of diverse content, different styles, large semantic spans and good interactive experience requirements. To solve these challenges, we establish a joint model CoQAN in combination with the layout organization, writing characteristics and text semantics, designing different representation learning subnetworks, especially for the feature learning process and interactive reading habits on mobile terminals. It is more consistent with the cognitive style of expressing an expert's evaluation of articles. We have also constructed a large scale real-world assessment dataset. Extensive experimental results show that the proposed framework significantly outperforms state-of-the-art methods, and effectively learns and integrates different factors of the online article quality assessment.
Shen Huang, Gongfu Li, Qiang Deng, Dongliang Liao, Pengda Si, Yujiu Yang 0001
ACM Multimedia7
2020 Controllable Continuous Gaze Redirection
abstract
In this work, we present interpGaze, a novel framework for controllable gaze redirection that achieves both precise redirection and continuous interpolation. Given two gaze images with different attributes, our goal is to redirect the eye gaze of one person into any gaze direction depicted in the reference image or to generate continuous intermediate results. To accomplish this, we design a model including three cooperative components: an encoder, a controller and a decoder. The encoder maps images into a well-disentangled and hierarchically-organized latent space. The controller adjusts the magnitudes of latent vectors to the desired strength of corresponding attributes by altering a control vector. The decoder converts the desired representations from the attribute space to the image space. To facilitate covering the full space of gaze directions, we introduce a high-quality gaze image dataset with a large range of directions, which also benefits researchers in related areas. Extensive experimental validation and comparisons to several baseline methods show that the proposed interpGaze outperforms state-of-the-art methods in terms of image quality and redirection precision.
Weihao Xia 0001, Yujiu Yang 0001, Jing-Hao Xue, Wensen Feng
ACM Multimedia2
2020 G2T: Generating Fluent Descriptions for Knowledge Graph
abstract
Generating natural language descriptions for knowledge graph (KG) is an important category for intelligent writing. Recent models on this task substitute the sequence encoder in a commonly used encoder-decoder framework with a graph encoder. However, these models suffer from entity missing and repetition. In this paper, we propose a novel end-to-end generation model named G2T, which integrates a novel Graph Structure Enhanced Mechanism (GSEM) and a Copy Coverage Loss (CCL). Instead of just considering graph structure in the encoding phase in most existing methods, our GSEM fully utilizes graph structure in the decoding phase and helps to mitigate entity missing problem. Moreover, our CCL can further improve performance by avoiding generating repeated entities. With their help, our model is capable of generating fluent description for KG. The results of automatic and human evaluations show that our model outperforms the state-of-the-art models.
Yunzhou Shi, Zhiling Luo, Haiqing Chen, Yujiu Yang 0001
SIGIR7
2020 Unsupervised multi-domain multimodal image-to-image translation with explicit domain-constrained disentanglement
Weihao Xia 0001, Yujiu Yang 0001, Jing-Hao Xue
Neural Networks2
2020 FAT-RE: A faster dependency-free model for relation extraction
Lifang Ding, Zeyang Lei, Guangxu Xun, Yujiu Yang 0001
J. Web Semant.4
2019 A Human-Like Semantic Cognition Network for Aspect-Level Sentiment Classification
abstract
In this paper, we propose a novel Human-like Semantic Cognition Network (HSCN) for aspect-level sentiment classification, motivated by the principles of human beings’ reading cognitive process (pre-reading, active reading, post-reading). We first design a word-level interactive perception module to capture the correlation between context words and the given target words, which can be regarded as pre-reading. Second, to mimic the process of active reading, we propose a targetaware semantic distillation module to produce the targetspecific context representation for aspect-level sentiment prediction. Third, we further devise a semantic deviation metric module to measure the semantic deviation between the targetspecific context representation and the given target, which evaluates the degree we understand the target-specific context semantics. The measured semantic deviation is then used to fine-tune the above active reading process in a feedback regulation way. To verify the effectiveness of our approach, we conduct extensive experiments on three widely used datasets. The experiments demonstrate that HSCN achieves impressive results compared to other strong competitors.
Zeyang Lei, Yujiu Yang 0001, Min Yang 0007, Wei Zhao 0013, Jun Guo 0008, Yi Liu 0021
AAAI2
2019 Compressing Convolutional Neural Networks via Factorized Convolutional Filters
abstract
This work studies the model compression for deep convolutional neural networks (CNNs) via filter pruning. The workflow of a traditional pruning consists of three sequential stages: pre-training the original model, selecting the pre-trained filters via ranking according to a manually designed criterion (e.g., the norm of filters), and learning the remained filters via fine-tuning. Most existing works follow this pipeline and focus on designing different ranking criteria for filter selection. However, it is difficult to control the performance due to the separation of filter selection and filter learning. In this work, we propose to conduct filter selection and filter learning simultaneously, in a unified model. To this end, we define a factorized convolutional filter (FCF), consisting of a standard real-valued convolutional filter and a binary scalar, as well as a dot-product operator between them. We train a CNN model with factorized convolutional filters (CNN-FCF) by updating the standard filter using back-propagation, while updating the binary scalar using the alternating direction method of multipliers (ADMM) based optimization method. With this trained CNN-FCF model, we only keep the standard filters corresponding to the 1-valued scalars, while all other filters and all binary scalars are discarded, to obtain a compact CNN model. Extensive experiments on CIFAR-10 and ImageNet demonstrate the superiority of the proposed method over state-of-the-art filter pruning methods.
Tuanhui Li, Baoyuan Wu, Yujiu Yang 0001, Yanbo Fan, Yong Zhang 0034, Wei Liu 0005
CVPR3
2019 Residual Dilated Network with Attention for Image Blind Denoising
abstract
Image denoising has recently witnessed substantial progress. However, many existing methods remain suboptimal for texture restoration due to treating different image regions and channels indiscriminately. Also they need to specify the noise level in advance, which largely hinders their use in blind denoising. Therefore, we introduce both attention mechanism and automatic noise level estimation into image denoising. Specifically, we propose a new, effective end-to-end attention-embedded neural network for image denoising, named as Residual Dilated Attention Network (RDAN). Our RDAN is composed of a series of tailored Residual Dilated Attention Blocks (RDAB) and Residual Conv Attention Blocks (RCAB). The RDAB and RCAB incorporates both non-local and local operations, which enable a comprehensive capture of structural information. In addition, we incorporate a Gaussian-based noise level estimation into RDAN to accomplish blind denoising. Experimental results have demonstrated that our RDAN can substantially outperforms the state-of-the-art denoising methods as well as promisingly preserve image texture.
Guanqun Hou, Yujiu Yang 0001, Jing-Hao Xue
ICME2
2019 Self-supervised Feature Learning for 3D Medical Images by Playing a Rubik's Cube
Xinrui Zhuang, Yuexiang Li, Kai Ma 0002, Yujiu Yang 0001, Yefeng Zheng 0001
MICCAI (4)5
2018 Sentiment Lexicon Enhanced Attention-Based LSTM for Sentiment Classification
abstract
Deep neural networks have gained great success recently for sentiment classification. However, these approaches do not fully exploit the linguistic knowledge. In this paper, we propose a novel sentiment lexicon enhanced attention-based LSTM (SLEA-LSTM) model to improve the performance of sentence-level sentiment classification. Our method successfully integrates sentiment lexicon into deep neural networks via single-head or multi-head attention mechanisms. We conduct extensive experiments on MR and SST datasets. The experimental results show that our model achieved comparable or better performance than the state-of-the-art methods.
Zeyang Lei, Yujiu Yang 0001, Min Yang 0007
AAAI2
2018 SAAN: A Sentiment-Aware Attention Network for Sentiment Analysis
abstract
Analyzing public opinions towards products, services and social events is an important but challenging task. Despite the remarkable successes of deep neural networks in sentiment analysis, these approaches do not make full use of the prior sentiment knowledge (e.g., sentiment lexicon, negation words, intensity words). In this paper, we propose a Sentiment-Aware Attention Network (SAAN) to boost the performance of sentiment analysis, which adopts a three-step strategy to learn the sentiment-specific sentence representation. First, we employ a word-level mutual attention mechanism to model word-level correlation. Next, a phrase-level convolutional attention is designed to obtain phrase-level correlation. Finally, a sentence-level multi-head attention mechanism is proposed to capture various sentimental information from different subspaces. The experiments on Movie Review (MR) and Stanford Sentiment Treebank (SST) show that our model consistently outperform the previous methods for sentiment analysis.
Zeyang Lei, Yujiu Yang 0001, Min Yang 0007
SIGIR2
2016 Combining user-based and global lexicon features for sentiment analysis in twitter
abstract
Generally speaking, sentiment lexicons employed in the majority of current sentiment analysis systems are trained globally from public data stream source or other large independent corpus. However, sentiments are rather subjective and personal states of mind that the individuality and diversity of characteristics, particular writing habit and idiolect could play a crucial role in the judgment of sentiment expressed by a specific user. In this paper, we present a novel feature construction method to combine user-based and global lexicon features in sentiment analysis for short social media text. After the creation of user-based sentiment lexicons from user-timeline corpus, a rule-based fusing approach is adopted subsequently to generate user-based lexicon features in combination with general lexicon features. Experiments show that user-based features may capture potential user preferences hence adjusting the bias caused by representing an individual's sentiment with an averaged lexicon score, and our proposed method yield better results in comparison with some of the state-of-the-art sentiment analysis systems in twitter.
Yujiu Yang 0001, Xianyu Bao, Biqing Huang
IJCNN2
2015 Within-class penalty based multi-class support vector machine
abstract
Support vector machine (SVM) is a widely used maximum margin classifier, but the classification performance is largely affected by outliers. In this paper, we propose a novel multi-class SVM method to reduce the influence of outliers on the classification performance. Our proposed method includes an efficient optimization model via considering the within-class scatter and an optimization way. Specifically, the method is based on one assumption that penalizing the within-class scatter can reduce the number of misclassified outliers near the decision boundary, because data points of each class could be compacted by the within-class penalty. Experiments on benchmark databases demonstrate the effectiveness of the assumption and the proposed method.
Xiaoshuang Shi, Zhenhua Guo 0001, Yujiu Yang 0001, Lin Yang 0002
ICIP3
2015 A Comprehensive Survey of Recommendation System Based on Taxi GPS Trajectory
abstract
In recent years, the service system based on taxi GPS trajectory gradually becomes a hot research topic. In this paper we first give some descriptions and definitions of the taxi GPS trajectory problems. Different from the traditional recommendation system, the service system based on taxi GPS trajectory will lead to some special challenges and we will give corresponding solutions to solve these problems. Therefore, we propose a recommendation system framework for this issue via the emphasis on temporal and spatial information mining. Then, we discuss the different classification method by different points of views including the statistics of spatial information, the modeling of time information, mining methods and knowledge discovery models. Finally, we point out the promising directions in this field.
Yuanhang Hu, Yujiu Yang 0001, Biqing Huang
ICSS2
2015 A Framework of Joint Graph Embedding and Sparse Regression for Dimensionality Reduction
abstract
Over the past few decades, a large number of algorithms have been developed for dimensionality reduction. Despite the different motivations of these algorithms, they can be interpreted by a common framework known as graph embedding. In order to explore the significant features of data, some sparse regression algorithms have been proposed based on graph embedding. However, the problem is that these algorithms include two separate steps: (1) embedding learning and (2) sparse regression. Thus their performance is largely determined by the effectiveness of the constructed graph. In this paper, we present a framework by combining the objective functions of graph embedding and sparse regression so that embedding learning and sparse regression can be jointly implemented and optimized, instead of simply using the graph spectral for sparse regression. By the proposed framework, supervised, semisupervised, and unsupervised learning algorithms could be unified. Furthermore, we analyze two situations of the optimization problem for the proposed framework. By adopting an ℓ2,1-norm regularization for the proposed framework, it can perform feature selection and subspace learning simultaneously. Experiments on seven standard databases demonstrate that joint graph embedding and sparse regression method can significantly improve the recognition performance and consistently outperform the sparse regression method.
Xiaoshuang Shi, Zhenhua Guo 0001, Zhihui Lai 0001, Yujiu Yang 0001, Zhifeng Bao, David Zhang 0001
IEEE Trans. Image Process.4
2014 Face recognition by sparse discriminant analysis via joint L2, 1-norm minimization
Xiaoshuang Shi, Yujiu Yang 0001, Zhenhua Guo 0001, Zhihui Lai 0001
Pattern Recognit.2
2013 Node Classification in Social Network via a Factor Graph Model
Yujiu Yang 0001, Wenhuang Liu
PAKDD (1)2
2010 WeightLOFCC: A Heuristic Weight-Setting Strategy of LOF Applied to Outlier Detection in Time Series Data
Hongrui Xie, Yujiu Yang 0001, Wenhuang Liu
ADMA (1)2
2010 A Probability Click Tracking Model Analysis of Web Search Results
Yujiu Yang 0001, Xinyi Shu, Wenhuang Liu
ICONIP (1)1
2008 Sparse Kernel-Based Feature Weighting
Shuang-Hong Yang, Yujiu Yang 0001, Bao-Gang Hu
PAKDD2
2008 Fighting WebSpam: Detecting Spam on the Graph Via Content and Link Features
Yujiu Yang 0001, Shuang-Hong Yang, Bao-Gang Hu
PAKDD1
2007 Pairwise Constraints-Guided Non-negative Matrix Factorization for Document Clustering
abstract
Nonnegative Matrix Factorization (NMF) has been proven to be effective in text mining. However, since NMF is a well-known unsupervised components analysis technique, the existing NMF method can not deal with prior constraints, which are beneficial to clustering or classification tasks. In this paper, we address the text clustering problem via a novel strategy, called Pairwise Constraintsguided Non-negative Matrix Factorization (PCNMF for short). Differing from the traditional NMF method, the proposed method can capture the available abundance prior constraints in original space, which result in more effective for clustering or information retrieval. Therefore, PCNMF enforces the discriminative capability in the reduced space. Utilizing the appropriate transformation, PCNMF represents as a new optimization problem, which can be efficiently solved by an iterative approach. The cluster membership of each document can be easily determined as the standard NMF. Empirical studies based on Benchmark document corpus demonstrate appealing results.
Yujiu Yang 0001, Bao-Gang Hu
Web Intelligence1
2005 Sparse Kernel Fisher Discriminant Analysis
Hong-Jie Xing, Yujiu Yang 0001, Bao-Gang Hu
ISNN (1)2
2004 Geometric Interpretation of Nonlinear Approximation Capability for Feedforward Neural Networks
Bao-Gang Hu, Hong-Jie Xing, Yujiu Yang 0001
ISNN (1)3