VLDB 2026 Research / reviewers in the wild / expert
Yu Cheng 0001
dblp:96/3060-1
· DBLP profile ↗
145ranked-venue papers
12as first author
86since 2021 · last 2026
0000-0002-2315-5641ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 124 · 7 first-author · 78 since 2021Graphics, computer vision, multimedia, augmented reality and games · 55 · 5 first-author · 32 since 2021Databases, data management, data science and information retrieval · 16 · 7 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 4Computer networks · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Less Is More: Vision Representation Compression for Efficient Video Generation with Large Language ModelsabstractVideo generation using Large Language Models (LLMs) has shown promising potential, effectively leveraging the extensive LLM infrastructure to provide a unified framework for multimodal understanding and content generation. However, these methods face critical challenges, i.e., token redundancy and inefficiencies arising from long sequences, which constrain their performance and efficiency compared to diffusion-based approaches. In this study, we investigate the impact of token redundancy in LLM-based video generation by information-theoretic analysis and propose Vision Representation Compression (VRC), a novel framework designed to achieve more in both performance and efficiency with less video token representations. VRC introduces learnable representation compressor and decompressor to compress video token representations, enabling autoregressive next-sequence prediction in a compact latent space. Our approach reduces redundancy, shortens token sequences, and improves model's ability to capture underlying video structures. Our experiments demonstrate that VRC reduces token sequence lengths by a factor of 4, achieving more than 9~14x acceleration in inference while maintaining performance comparable to state-of-the-art video generation models. VRC not only accelerates the inference but also significantly reduces memory requirements during both model training and inference. Yucheng Zhou 0001, Jihai Zhang 0002, Guanjie Chen, Jianbing Shen, Yu Cheng 0001 |
AAAI | 5 |
| 2026 | Native Hybrid Attention for Efficient Sequence ModelingabstractTransformers excel at sequence modeling but face quadratic complexity, while linear attention offers improved efficiency but often compromises recall accuracy over long contexts.In this work, we introduce Native Hybrid Attention (NHA), a novel hybrid architecture of linear and full attention that integrates both intra & inter-layer hybridization into a unified layer design.NHA maintains longterm context in key-value slots updated by a linear RNN, and augments them with shortterm tokens from a sliding window.A single softmax attention operation is then applied over all keys and values, enabling pertoken and per-head context-dependent weighting without requiring additional fusion parameters.The inter-layer behavior is controlled through a single hyperparameter, the sliding window size, which allows smooth adjustment between purely linear and full attention while keeping all layers structurally uniform.Experimental results show that NHA surpasses Transformers and other hybrid baselines on recall-intensive and commonsense reasoning tasks.Furthermore, pretrained LLMs can be structurally hybridized with NHA, achieving competitive accuracy while delivering significant efficiency gains.Code is available at https://github.com/JusenD/NHA. Jusen Du, Jiaxi Hu, Zhang Tao, Weigao Sun, Yu Cheng 0001 |
ACL (1) | 5 |
| 2026 | Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning ModelsabstractInstruction-following is essential for aligning large language models (LLMs) with user intent.Yet recent reasoning-oriented models, despite their strong performance on complex mathematical problems, often fail to comply with simple natural language directives.In this work, we analyze the interaction between reasoning ability and instruction adherence in large reasoning models (LRMs).Using a controlled evaluation framework (MathIF), we uncover a persistent trade-off: as models scale reasoning capacity through long chains-of-thought or reinforcement learning on reasoning traces, their obedience to instructions degrades, particularly when generation length grows.We further show that interventions such as constraining or repeating instructions can partially restore compliance, but typically at the expense of reasoning performance.Taken together, our findings expose a dilemma between intelligence and obedience in current training paradigms and underscore the need for instruction-aware approaches to developing controllable reasoning models. Tingchen Fu, Yafu Li, Jiawei Gu, Xiaoye Qu, Yu Cheng 0001 |
ACL (1) | 5 |
| 2026 | Nirvana: A Specialized Generalist Model With Task-Aware Memory MechanismabstractYuhua Jiang, Shuang Cheng, Yihao Liu, Ermo Hua, Che Jiang, Weigao Sun, Yu Cheng, Feifei Gao, Biqing Qi, Bowen Zhou. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yuhua Jiang, Shuang Cheng, Yihao Liu 0008, Ermo Hua, Che Jiang, Weigao Sun, Yu Cheng 0001, Biqing Qi, Bowen Zhou 0002 |
ACL (1) | 7 |
| 2026 | One Refiner to Unlock Them All: Inference-Time Reasoning Elicitation via Reinforcement Query RefinementabstractLarge Language Models (LLMs) often fail to utilize their latent reasoning capabilities due to a distributional mismatch between ambiguous human inquiries and the structured logic required for machine activation.Existing alignment methods either incur prohibitive O(N ) costs by fine-tuning each model individually or rely on static prompts that fail to resolve query-level structural complexity.In this paper, we propose ReQueR (Reinforcement Query Refinement), a modular framework that treats reasoning elicitation as an inference-time alignment task.We train a specialized Refiner policy via Reinforcement Learning to rewrite raw queries into explicit logical decompositions, treating frozen LLMs as the environment.Rooted in the classical Zone of Proximal Development from educational psychology, we introduce the Adaptive Solver Hierarchy, a curriculum mechanism that stabilizes training by dynamically aligning environmental difficulty with the Refiner's evolving competence.ReQueR yields consistent absolute gains of 1.7%-7.2%across diverse architectures and benchmarks, outperforming strong baselines by 2.1% on average.Crucially, it provides a promising paradigm for one-to-many inference-time reasoning elicitation, enabling a single Refiner trained on a small set of models to effectively unlock reasoning in diverse unseen models. Yixiao Zhou 0001, Dongzhou Cheng, Zhiliang Wu, Yi Yang 0001, Yu Cheng 0001, Hehe Fan |
ACL (1) | 5 |
| 2026 | Each Rank Could be an Expert: Single-Ranked Mixture of Experts LoRA for Multi-task LearningabstractLow-Rank Adaptation (LoRA) is widely used for adapting large language models (LLMs) to specific domains due to its efficiency and modularity. However, vanilla LoRA struggles with task conflicts in multi-task scenarios. Recent works adopt Mixture of Experts (MoE) by treating each LoRA module as an expert, thereby mitigating task interference through multiple specialized LoRA modules. While effective, these methods often isolate knowledge within individual tasks, failing to fully exploit the shared knowledge across related tasks. In this paper, we establish a connection between single LoRA and multi-LoRA MoE, integrating them into a unified framework. We demonstrate that the dynamic routing of multiple LoRAs is functionally equivalent to rank partitioning and block-level activation within a single LoRA. To systematically study the role of expert granularity in multi-task learning, we conduct an in-depth investigation within our unified framework. Our empirical results show that a finer-grained expert partitioning not only yields significant performance gains but also captures more diverse parameter patterns. These empirical findings are supported by our theoretical analysis, which proves that finer granularity expands parameter space diversity and tightens the model's error bound. Building on these findings, we propose Single-ranked Mixture of Experts LoRA (SMoRA ), which embeds MoE into LoRA by treating each rank as an independent expert. With a dynamic rank-wise activation mechanism, SMoRA facilitates a flexible composition of knowledge, enabling the model to learn deeper and more diverse features while mitigating task conflicts. Experiments demonstrate that SMoRA activates fewer parameters yet achieves better performance in multi-task scenarios. Ziyu Zhao 0001, Yixiao Zhou 0001, Zhi Zhang 0005, Didi Zhu, Tao Shen 0002, Zexi Li 0001, Jinluan Yang, Xuwu Wang, Jing Su 0005, Kun Kuang 0001, Zhongyu Wei, Fei Wu 0001, Yu Cheng 0001 |
KDD (1) | 14 |
| 2026 | Mitigating Multilingual Hallucination in Large Vision-Language ModelsabstractWhile Large Vision-Language Models (LVLMs) have exhibited remarkable capabilities across a wide range of tasks, they suffer from hallucination problems, where models generate plausible yet incorrect answers given the input image-query pair. This hallucination phenomenon is even more severe when querying the image in non-English languages, while existing methods for mitigating hallucinations in LVLMs only consider the English scenarios. In this article, we make the first attempt to mitigate this important multilingual hallucination in LVLMs. With thorough experimental analysis, we found that multilingual hallucination in LVLMs is a systemic problem that could arise from deficiencies in multilingual capabilities or inadequate multimodal abilities. To this end, we propose a two-stage Multilingual Hallucination Removal (MHR) framework for LVLMs, aiming to improve resistance to hallucination for both high-resource and low-resource languages. Specifically, in the first stage, considering that most non-English languages cannot follow instructions well and output non-sense answers given the input image, we boost multilingual instruction-following ability with a multilingual supervised fine-tuning. The second phase is aimed at enhancing the LVLM’s ability to diminish multilingual hallucinations. Instead of relying on the intricate manual annotations of multilingual resources, we fully leverage the inherent capabilities of the LVLM and propose a novel cross-lingual alignment method, which generates multiple responses for each image-query input and then identifies the hallucination-aware pairs for each language. These data pairs are finally used for direct preference optimization to prompt the LVLMs to favor non-hallucinating responses. Experimental results show that our MHR achieves a substantial reduction in hallucination generation for LVLMs. Our code and model weights are available at https://github.com/ssmisya/MHR . Xiaoye Qu, Wei Wei 0002, Daizong Liu, Jianfeng Dong, Yu Cheng 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2025 | Cooperative or Competitive? Understanding the Interaction between Attention Heads From A Game Theory PerspectiveabstractDespite the remarkable success of attentionbased large language models (LLMs), the precise interaction mechanisms between attention heads remain poorly understood.In contrast to prevalent methods that focus on individual head contributions, we rigorously analyze the intricate interplay among attention heads through a novel framework based on the Harsanyi dividend, a concept from cooperative game theory.Our analysis reveals that significant positive Harsanyi dividends are sparsely distributed across head combinations, indicating that most heads do not contribute cooperatively.Moreover, certain head combinations exhibit negative dividends, indicating implicit competitive relationships.To further optimize the interactions among attention heads, we propose a training-free Game-theoretic Attention Calibration (GAC) method.Specifically, GAC selectively retains heads demonstrating significant cooperative gains and applies fine-grained distributional adjustments to the remaining heads.Comprehensive experiments across 17 benchmarks demonstrate the effectiveness of our proposed GAC and its superior generalization capabilities across diverse model families, scales, and modalities.The source code is available at Xiaoye Qu, Zengqi Yu, Dongrui Liu, Wei Wei 0002, Daizong Liu, Jianfeng Dong, Yu Cheng 0001 |
ACL (1) | 7 |
| 2025 | PRMBench: A Fine-grained and Challenging Benchmark for Process-Level Reward ModelsabstractProcess-level Reward Models (PRMs) are crucial for complex reasoning and decisionmaking tasks, where each intermediate step plays an important role in the reasoning process.Since large language models (LLMs) suffer from various types of errors during the reasoning process, PRMs are required to possess nuanced capabilities for detecting various implicit error types in real-world scenarios.However, current benchmarks primarily focus on step correctness, failing to evaluate PRMs' performance systematically.To address this gap, we introduce PRMBENCH, a processlevel benchmark specifically designed to assess the fine-grained error detection capabilities of PRMs.PRMBENCH comprises 6,216 carefully designed problems and 83,456 step-level labels, evaluating models across multiple dimensions, including simplicity, soundness, and sensitivity.In our experiments on 25 models, spanning across both open-source PRMs and LLMs prompted as critic models, we uncover significant weaknesses in current PRMs.These findings reveal the challenges inherent in processlevel evaluation and highlight key directions for future research, establishing PRMBENCH as a robust testbed for advancing research on PRM evaluation and development. Zhaochen Su, Xiaoye Qu, Yu Cheng 0001 |
ACL (1) | 5 |
| 2025 | Unveiling Attractor Cycles in Large Language Models: A Dynamical Systems View of Successive ParaphrasingabstractDynamical systems theory provides a framework for analyzing iterative processes and evolution over time.Within such systems, repetitive transformations can lead to stable configurations, known as attractors, including fixed points and limit cycles.Applying this perspective to large language models (LLMs), which iteratively map input text to output text, provides a principled approach to characterizing long-term behaviors.Successive paraphrasing serves as a compelling testbed for exploring such dynamics, as paraphrases re-express the same underlying meaning with linguistic variation.Although LLMs are expected to explore a diverse set of paraphrases in the text space, our study reveals that successive paraphrasing converges to stable periodic states, such as 2period attractor cycles, limiting linguistic diversity.This phenomenon is attributed to the selfreinforcing nature of LLMs, as they iteratively favour and amplify certain textual forms over others.This pattern persists with increasing generation randomness or alternating prompts and LLMs.These findings underscore inherent constraints in LLM generative capability, while offering a novel dynamical systems perspective for studying their expressive potential.Our code is available here. Zhilin Wang, Yafu Li, Jianhao Yan, Yu Cheng 0001, Yue Zhang 0004 |
ACL (1) | 4 |
| 2025 | Look, Compare, Decide: Alleviating Hallucination in Large Vision-Language Models via Multi-View Multi-Path ReasoningabstractRecently, Large Vision-Language Models (LVLMs) have demonstrated impressive capabilities in multi-modal context comprehension. However, they still suffer from hallucination problems referring to generating inconsistent outputs with the image content. To mitigate hallucinations, previous studies mainly focus on retraining LVLMs with custom datasets. Although effective, they inherently come with additional computational costs. In this paper, we propose a training-free framework, MVP, that aims to reduce hallucinations by making the most of the innate capabilities of the LVLMs via Multi-View Multi-Path Reasoning. Specifically, we first devise a multi-view information-seeking strategy to thoroughly perceive the comprehensive information in the image, which enriches the general global information captured by the original vision encoder in LVLMs. Furthermore, during the answer decoding, we propose multi-path reasoning for each information view to quantify and aggregate the certainty scores for each potential answer among multiple decoding paths and finally decide the output answer. By fully grasping the information in the image and carefully considering the certainty of the potential answers when decoding, our MVP can effectively reduce hallucinations in LVLMs. The extensive experiments verify that our proposed MVP significantly mitigates the hallucination problem across four well-known LVLMs. Xiaoye Qu, Jiashuo Sun, Wei Wei 0002, Daizong Liu, Jianfeng Dong, Yu Cheng 0001 |
COLING | 6 |
| 2025 | From Head to Tail: Towards Balanced Representation in Large Vision-Language Models through Adaptive Data CalibrationabstractLarge Vision-Language Models (LVLMs) have achieved significant progress in combining visual comprehension with language generation. Despite this success, the training data of LVLMs still suffers from Long-Tail (LT) problems, where the data distribution is highly imbalanced. Previous works have mainly focused on traditional VLM architectures, i.e., CLIP or ViT, and specific tasks such as recognition and classification. Nevertheless, the exploration of LVLM (e.g. LLaVA) and more general tasks (e.g. Visual Question Answering and Visual Reasoning) remains under-explored. In this paper, we first conduct an in-depth analysis of the LT issues in LVLMs and identify two core causes: the overrepresentation of head concepts and the underrepresentation of tail concepts. Based on the above observation, we propose an Adaptive Data Refinement Framework (ADR), which consists of two stages: Data Rebalancing (DR) and Data Synthesis (DS). In the DR stage, we adaptively rebalance the redundant data based on entity distributions, while in the DS stage, we leverage Denoising Diffusion Probabilistic Models (DDPMs) and scarce images to supplement under-represented portions. Through comprehensive evaluations across eleven benchmarks, our proposed ADR effectively mitigates the long-tail problem in the training data, improving the average performance of LLaVA 1.5 relatively by 4.36%, without increasing the training data volume. Xiaoye Qu, Yu Cheng 0001 |
CVPR | 4 |
| 2025 | Extrapolating and Decoupling Image-to-Video Generation Models: Motion Modeling is Easier Than You ThinkabstractImage-to-Video (I2V) generation aims to synthesize a video clip according to a given image and condition (e.g., text). The key challenge of this task lies in simultaneously generating natural motions while preserving the original appearance of the images. However, current I2V diffusion models (I2V-DMs) often produce videos with limited motion degrees or exhibit uncontrollable motion that conflicts with the textual condition. To address these limitations, we propose a novel Extrapolating and Decoupling framework, which introduces model merging techniques to the I2V domain for the first time. Specifically, our framework consists of three separate stages: (1) Starting with a base I2V-DM, we explicitly inject the textual condition into the temporal module using a lightweight, learnable adapter and fine-tune the integrated model to improve motion controllability. (2) We introduce a training-free extrapolation strategy to amplify the dynamic range of the motion, effectively reversing the fine-tuning process to enhance the motion degree significantly. (3) With the above two-stage models excelling in motion controllability and degree, we decouple the relevant parameters associated with each type of motion ability and inject them into the base I2V-DM. Since the I2V-DM handles different levels of motion controllability and dynamics at various denoising time steps, we adjust the motion-aware parameters accordingly over time. Extensive qualitative and quantitative experiments have been conducted to demonstrate the superiority of our framework over existing methods. Code is available at https://github.com/Chuge0335/EDG Xiaoye Qu, Zhenyi Lu, Wei Wei 0002, Yu Cheng 0001 |
CVPR | 6 |
| 2025 | Bit-Flip Error Resilience in LLMs: A Comprehensive Analysis and Defense FrameworkabstractYuhang Chen, Zhen Tan, Ajay Kumar Jaiswal, Huaizhi Qu, Xinyu Zhao, Qi Lin, Yu Cheng, Andrew Kwong, Zhichao Cao, Tianlong Chen. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Zhen Tan 0001, Ajay Jaiswal, Huaizhi Qu, Yu Cheng 0001, Andrew Kwong, Zhichao Cao 0002, Tianlong Chen 0001 |
EMNLP | 7 |
| 2025 | CLIP-MoE: Towards Building Mixture of Experts for CLIP with Diversified Multiplet UpcyclingabstractContrastive Language-Image Pre-training (CLIP) has become a cornerstone in multimodal intelligence.However, recent studies discovered that CLIP can only encode one aspect of the feature space, leading to substantial information loss and indistinctive features.To mitigate this issue, this paper introduces a novel strategy that fine-tunes a series of complementary CLIP models and transforms them into a CLIP-MoE.Specifically, we propose a model-agnostic Diversified Multiplet Upcycling (DMU) framework for CLIP.Instead of training multiple CLIP models from scratch, DMU leverages a pre-trained CLIP and fine-tunes it into a diverse set with highly cost-effective multistage contrastive learning, thus capturing distinct feature subspaces efficiently.To fully exploit these fine-tuned models while minimizing computational overhead, we transform them into a CLIP-MoE, which dynamically activates a subset of CLIP experts, achieving an effective balance between model capacity and computational cost.Comprehensive experiments demonstrate the superior performance of CLIP-MoE across various zero-shot retrieval, zero-shot image classification tasks, and downstream Multimodal Large Language Model (MLLM) benchmarks when used as a vision encoder.Code is available at https: //github.com/OpenSparseLLMs/CLIP-MoE. Jihai Zhang 0002, Xiaoye Qu, Tong Zhu 0002, Yu Cheng 0001 |
EMNLP | 4 |
| 2025 | Towards Stabilized and Efficient Diffusion Transformers Through Long-Skip-Connections With Spectral Constraints
Guanjie Chen, Yucheng Zhou 0001, Xiaoye Qu, Tianlong Chen 0001, Yu Cheng 0001 |
ICCV | 6 |
| 2025 | LangBridge: Interpreting Image as a Combination of Language EmbeddingsabstractRecent years have witnessed remarkable advances in Large Vision-Language Models (LVLMs), which have achieved human-level performance across various complex vision-language tasks. Following LLaVA's paradigm, mainstream LVLMs typically employ a shallow MLP for visual-language alignment through a two-stage training process: pretraining for cross-modal alignment followed by instruction tuning. While this approach has proven effective, the underlying mechanisms of how MLPs bridge the modality gap remain poorly understood. Although some research has explored how LLMs process transformed visual tokens, few studies have investigated the fundamental alignment mechanism. Furthermore, the MLP adapter requires retraining whenever switching LLM backbones. To address these limitations, we first investigate the working principles of MLP adapters and discover that they learn to project visual embeddings into subspaces spanned by corresponding text embeddings progressively. Based on this insight, we propose LangBridge, a novel adapter that explicitly maps visual tokens to linear combinations of LLM vocabulary embeddings. This innovative design enables pretraining-free adapter transfer across different LLMs while maintaining performance. Our experimental results demonstrate that a LangBridge adapter pre-trained on Qwen2-0.5B can be directly applied to larger models such as LLaMA3-8B or Qwen2.5-14B while maintaining competitive performance. Overall, LangBridge enables interpretable vision-language alignment by grounding visual representations in LLM vocab embedding, while its plug-and-play design ensures efficient reuse across multiple LLMs with nearly no performance degradation. See our project page at https://curryx-001.github.io/LangBridge.github.io/ Jiaqi Liao, Yuwei Niu, Fanqing Meng, Hao Li 0069, Changyao Tian, Yinuo Du, Yuwen Xiong, Dianqi Li, Xizhou Zhu, Jifeng Dai, Yu Cheng 0001 |
ICCV | 12 |
| 2025 | ImageGen-CoT: Enhancing Text-to-Image in-context Learning with Chain-of-Thought Reasoning
Jiaqi Liao, Zhengyuan Yang, Dianqi Li, Yu Cheng 0001 |
ICCV | 6 |
| 2025 | Weak to Strong Generalization for Large Language Models with Multi-capabilitiesabstractAs large language models (LLMs) grow in sophistication, some of their capabilities surpass human abilities, making it essential to ensure their alignment with human values and intentions, i.e., Superalignment. This superalignment challenge is particularly critical for complex tasks, as annotations provided by humans, as weak supervisors, may be overly simplistic, incomplete, or incorrect. Previous work has demonstrated the potential of training a strong model using the weak dataset generated by a weak model as weak supervision. However, these studies have been limited to a single capability. In this work, we conduct extensive experiments to investigate weak to strong generalization for LLMs with multi-capabilities. The experiments reveal that different capabilities tend to remain relatively independent in this generalization, and the effectiveness of weak supervision is significantly impacted by the quality and diversity of the weak datasets. Moreover, the self-bootstrapping of the strong model leads to performance degradation due to its overconfidence and the limited diversity of its generated dataset. To address these issues, we proposed a novel training framework using reward models to select valuable data, thereby providing weak supervision for strong model training. In addition, we propose a two-stage training method on both weak and selected datasets to train the strong model. Experimental results demonstrate our method significantly improves the weak to strong generalization with multi-capabilities. Yucheng Zhou 0001, Jianbing Shen, Yu Cheng 0001 |
ICLR | 3 |
| 2025 | Diving into Self-Evolving Training for Multimodal ReasoningabstractSelf-evolving training—where models iteratively learn from their own outputs—has emerged as a key approach for complex reasoning tasks, addressing the scarcity of high-quality chain-of-thought data. However, its effectiveness in multimodal reasoning, a domain more intricate than text-only reasoning, remains underexplored, and the understanding of critical factors in this training paradigm remains limited. Furthermore, a central challenge for this training method is performance saturation, which impedes further improvements and scalability. Inspired by reinforcement learning (RL), in this paper, we reframe self-evolving training for multimodal reasoning through the lens of RL, identifying three pivotal factors: $\textit{Training Method}$, $\textit{Reward Model}$, and $\textit{Prompt Variation}$. Through systematic analysis, we establish relatively optimal design principles that significantly enhance multimodal reasoning capabilities. Moreover, delving deeper into training dynamics, we uncover the roots of saturation and propose a new automatic balancing mechanism to mitigate this limitation. Building on these insights, we propose M-STaR (**M**ultimodal **S**elf-evolving **T**r**a**ining for **R**easoning), a framework that achieves consistent performance gains across models of varying sizes and diverse benchmarks. All resources will be made publicly available. Wei Liu 0131, Yu Cheng 0001, Junxian He |
ICML | 5 |
| 2025 | Make LoRA Great Again: Boosting LoRA with Adaptive Singular Values and Mixture-of-Experts Optimization AlignmentabstractWhile Low-Rank Adaptation (LoRA) enables parameter-efficient fine-tuning for Large Language Models (LLMs), its performance often falls short of Full Fine-Tuning (Full FT). Current methods optimize LoRA by initializing with static singular value decomposition (SVD) subsets, leading to suboptimal leveraging of pre-trained knowledge. Another path for improving LoRA is incorporating a Mixture-of-Experts (MoE) architecture. However, weight misalignment and complex gradient dynamics make it challenging to adopt SVD prior to the LoRA MoE architecture. To mitigate these issues, we propose Great LoRA Mixture-of-Expert (GOAT), a framework that (1) adaptively integrates relevant priors using an SVD-structured MoE, and (2) aligns optimization with full fine-tuned MoE by deriving a theoretical scaling factor. We demonstrate that proper scaling, without modifying the architecture or training algorithms, boosts LoRA MoE’s efficiency and performance. Experiments across 25 datasets, including natural language understanding, commonsense reasoning, image classification, and natural language generation, demonstrate GOAT’s state-of-the-art performance, closing the gap with Full FT. Our code is available at: https://github.com/Facico/GOAT-PEFT Chenghao Fan, Zhenyi Lu, Chengfeng Gu, Xiaoye Qu, Wei Wei 0002, Yu Cheng 0001 |
ICML | 7 |
| 2025 | Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning BenchmarkabstractThe ability to organically reason over and with both text and images is a pillar of human intelligence, yet the ability of Multimodal Large Language Models (MLLMs) to perform such multimodal reasoning remains under-explored. Existing benchmarks often emphasize text-dominant reasoning or rely on shallow visual cues, failing to adequately assess integrated visual and textual reasoning. We introduce EMMA (Enhanced MultiModal reAsoning), a benchmark targeting organic multimodal reasoning across mathematics, physics, chemistry, and coding. EMMA tasks demand advanced cross-modal reasoning that cannot be addressed by reasoning independently in each modality, offering an enhanced test suite for MLLMs' reasoning capabilities. Our evaluation of state-of-the-art MLLMs on EMMA reveals significant limitations in handling complex multimodal and multi-step reasoning tasks, even with advanced techniques like Chain-of-Thought prompting and test-time compute scaling underperforming. These findings underscore the need for improved multimodal architectures and training paradigms to close the gap between human and model reasoning in multimodality. Yunzhuo Hao, Jiawei Gu, Huichen Will Wang, Zhengyuan Yang, Yu Cheng 0001 |
ICML | 7 |
| 2025 | Divide and Conquer: Grounding LLMs as Efficient Decision-Making Agents via Offline Hierarchical Reinforcement LearningabstractWhile showing sophisticated reasoning abilities, large language models (LLMs) still struggle with long-horizon decision-making tasks due to deficient exploration and long-term credit assignment, especially in sparse-reward scenarios. Inspired by the divide-and-conquer principle, we propose an innovative framework GLIDER (Grounding Language Models as EffIcient Decision-Making Agents via Offline HiErarchical Reinforcement Learning) that introduces a parameter-efficient and generally applicable hierarchy to LLM policies. We develop a scheme where the low-level controller is supervised with abstract, step-by-step plans that are learned and instructed by the high-level policy. This design decomposes complicated problems into a series of coherent chain-of-thought reasoning sub-tasks, providing flexible temporal abstraction to significantly enhance exploration and learning for long-horizon tasks. Furthermore, GLIDER facilitates fast online adaptation to non-stationary environments owing to the strong transferability of its task-agnostic low-level skills. Experiments on ScienceWorld and ALFWorld benchmarks show that GLIDER achieves consistent performance gains, along with enhanced generalization capabilities. Zican Hu, Wei Liu 0131, Xiaoye Qu, Xiangyu Yue 0001, Chunlin Chen 0001, Zhi Wang 0001, Yu Cheng 0001 |
ICML | 7 |
| 2025 | Liger: Linearizing Large Language Models to Gated Recurrent StructuresabstractTransformers with linear recurrent modeling offer linear-time training and constant-memory inference. Despite their demonstrated efficiency and performance, pretraining such non-standard architectures from scratch remains costly and risky. The linearization of large language models (LLMs) transforms pretrained standard models into linear recurrent structures, enabling more efficient deployment. However, current linearization methods typically introduce additional feature map modules that require extensive fine-tuning and overlook the gating mechanisms used in state-of-the-art linear recurrent models. To address these issues, this paper presents Liger, short for Linearizing LLMs to gated recurrent structures. Liger is a novel approach for converting pretrained LLMs into gated linear recurrent models without adding extra parameters. It repurposes the pretrained key matrix weights to construct diverse gating mechanisms, facilitating the formation of various gated recurrent structures while avoiding the need to train additional components from scratch. Using lightweight fine-tuning with Low-Rank Adaptation (LoRA), Liger restores the performance of the linearized gated recurrent models to match that of the original LLMs. Additionally, we introduce Liger Attention, an intra-layer hybrid attention mechanism, which significantly recovers 93% of the Transformer-based LLM performance at 0.02% pre-training tokens during the linearization process, achieving competitive results across multiple benchmarks, as validated on models ranging from 1B to 8B parameters. Disen Lan, Weigao Sun, Jiaxi Hu, Jusen Du, Yu Cheng 0001 |
ICML | 5 |
| 2025 | Test-Time Preference Optimization: On-the-Fly Alignment via Iterative Textual FeedbackabstractLarge language models (LLMs) have presented impressive performance but often lack the flexibility to adapt to human preferences quickly without retraining. Inspired by the recent efforts on test-time scaling, we make the first attempt to propose Test-time Preference Optimization (TPO), a framework that aligns LLM outputs with human preferences during inference, eliminating the need to update model parameters. Instead of relying on purely numerical rewards, TPO translates reward signals into \emph{textual} critiques and uses them as textual rewards to iteratively refine its response. Evaluations on benchmarks covering instruction following, preference alignment, safety, and mathematics reveal that TPO progressively improves alignment with human preferences. Notably, after only a few TPO steps, the initially unaligned Llama-3.1-70B-SFT model can surpass the aligned counterpart, Llama-3.1-70B-Instruct. Furthermore, TPO scales efficiently with both the search width and depth of the inference process. Through case studies, we illustrate how TPO exploits the innate capacity of LLM to interpret and act upon reward signals. Our findings establish TPO as a practical, lightweight alternative for test-time preference optimization, achieving alignment on the fly. Yafu Li, Xuyang Hu, Xiaoye Qu, Yu Cheng 0001 |
ICML | 5 |
| 2025 | Occult: Optimizing Collaborative Communications across Experts for Accelerated Parallel MoE Training and InferenceabstractMixture-of-experts (MoE) architectures could achieve impressive computational efficiency with expert parallelism, which relies heavily on all-to-all communication across devices. Unfortunately, such communication overhead typically constitutes a significant portion of the total runtime, hampering the scalability of distributed training and inference for modern MoE models (consuming over 40% runtime in large-scale training). In this paper, we first define $\textit{collaborative communication}$ to illustrate this intrinsic limitation, and then propose system- and algorithm-level innovations to reduce communication costs. Specifically, given a pair of experts co-activated by one token, we call them as $\textit{collaborated}$, which comprises $2$ cases as $\textit{intra-}$ and $\textit{inter-collaboration}$, depending on whether they are kept on the same device. Our pilot investigations reveal that augmenting the proportion of intra-collaboration can accelerate expert parallel at scale. It motivates us to strategically $\underline{\texttt{o}}$ptimize $\underline{\texttt{c}}$ollaborative $\underline{\texttt{c}}$omm$\underline{\texttt{u}}$nication for acce$\underline{\texttt{l}}$era$\underline{\texttt{t}}$ed MoE training and inference, dubbed $\textbf{\texttt{Occult}}$. Our designs are capable of $\underline{either}$ delivering exact results with reduced communication cost, $\underline{or}$ controllably minimizing the cost with collaboration pruning, materialized by modified fine-tuning. Comprehensive experiments on various MoE-LLMs demonstrate that $\texttt{Occult}$ can be faster than popular state-of-the-art inference or training frameworks (over 50% speed up across multiple tasks and models) with comparable or superior quality compared to the standard fine-tuning. Codes will be available upon acceptance. Shuqing Luo, Pingzhi Li, Jie Peng 0002, Yang Zhao 0013, Yu Cao 0001, Yu Cheng 0001, Tianlong Chen 0001 |
ICML | 6 |
| 2025 | Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video GenerationabstractText-to-video (T2V) models like Sora have made significant strides in visualizing complex prompts, which is increasingly viewed as a promising path towards constructing the universal world simulator. Cognitive psychologists believe that the foundation for achieving this goal is the ability to understand intuitive physics. However, the capacity of these models to accurately represent intuitive physics remains largely unexplored. To bridge this gap, we introduce PhyGenBench, a comprehensive \textbf{Phy}sics \textbf{Gen}eration \textbf{Ben}chmark designed to evaluate physical commonsense correctness in T2V generation. PhyGenBench comprises 160 carefully crafted prompts across 27 distinct physical laws, spanning four fundamental domains, which could comprehensively assesses models' understanding of physical commonsense. Alongside PhyGenBench, we propose a novel evaluation framework called PhyGenEval. This framework employs a hierarchical evaluation structure utilizing appropriate advanced vision-language models and large language models to assess physical commonsense. Through PhyGenBench and PhyGenEval, we can conduct large-scale automated assessments of T2V models' understanding of physical commonsense, which align closely with human feedback. Our evaluation results and in-depth analysis demonstrate that current models struggle to generate videos that comply with physical commonsense. Moreover, simply scaling up models or employing prompt engineering techniques is insufficient to fully address the challenges presented by PhyGenBench (e.g., dynamic scenarios). We hope this study will inspire the community to prioritize the learning of physical commonsense in these models beyond entertainment applications. We will release the data and codes at https://github.com/OpenGVLab/PhyGenBench Fanqing Meng, Jiaqi Liao, Quanfeng Lu, Wenqi Shao, Kaipeng Zhang, Yu Cheng 0001, Dianqi Li, Ping Luo 0002 |
ICML | 7 |
| 2025 | Dynamic Data Mixing Maximizes Instruction Tuning for Mixture-of-ExpertsabstractTong Zhu, Daize Dong, Xiaoye Qu, Jiacheng Ruan, Wenliang Chen, Yu Cheng. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Tong Zhu 0002, Daize Dong, Xiaoye Qu, Jiacheng Ruan, Wenliang Chen, Yu Cheng 0001 |
NAACL (Long Papers) | 6 |
| 2025 | Learning to Reason under Off-Policy GuidanceabstractRecent advances in large reasoning models (LRMs) demonstrate that sophisticated behaviors such as multi-step reasoning and self-reflection can emerge via reinforcement learning with verifiable rewards~(RLVR).
However, existing RLVR approaches are inherently ``on-policy'', limiting learning to a model's own outputs and failing to acquire reasoning abilities beyond its initial capabilities.
To address this issue, we introduce LUFFY (Learning to reason Under oFF-policY guidance), a framework that augments RLVR with off-policy reasoning traces.
LUFFY dynamically balances imitation and exploration by combining off-policy demonstrations with on-policy rollouts during training.
Specifically, LUFFY combines the Mixed-Policy GRPO framework, which has a theoretically guaranteed convergence rate, alongside policy shaping via regularized importance sampling to avoid superficial and rigid imitation during mixed-policy training.
Compared with previous RLVR methods, LUFFY achieves an over +6.4 average gain across six math benchmarks and an advantage of over +6.2 points in out-of-distribution tasks.
Most significantly, we show that LUFFY successfully trains weak models in scenarios where on-policy RLVR completely fails. These results provide compelling evidence that LUFFY transcends the fundamental limitations of on-policy RLVR and demonstrates the great potential of utilizing off-policy guidance in RLVR. Jianhao Yan, Yafu Li, Zican Hu, Zhi Wang 0001, Ganqu Cui, Xiaoye Qu, Yu Cheng 0001, Yue Zhang 0004 |
NeurIPS | 7 |
| 2025 | Text-to-Decision Agent: Offline Meta-Reinforcement Learning from Natural Language SupervisionabstractOffline meta-RL usually tackles generalization by inferring task beliefs from high-quality samples or warmup explorations. The restricted form limits their generality and usability since these supervision signals are expensive and even infeasible to acquire in advance for unseen tasks. Learning directly from the raw text about decision tasks is a promising alternative to leverage a much broader source of supervision. In the paper, we propose **T**ext-to-**D**ecision **A**gent (**T2DA**), a simple and scalable framework that supervises offline meta-RL with natural language. We first introduce a generalized world model to encode multi-task decision data into a dynamics-aware embedding space. Then, inspired by CLIP, we predict which textual description goes with which decision embedding, effectively bridging their semantic gap via contrastive language-decision pre-training and aligning the text embeddings to comprehend the environment dynamics. After training the text-conditioned generalist policy, the agent can directly realize zero-shot text-to-decision generation in response to language instructions. Comprehensive experiments on MuJoCo and Meta-World benchmarks show that T2DA facilitates high-capacity zero-shot generalization and outperforms various types of baselines. Our code is available at [https://github.com/NJU-RL/T2DA](https://github.com/NJU-RL/T2DA). Zican Hu, Jianxiang Tang, Chunlin Chen 0001, Daoyi Dong, Yu Cheng 0001, Zhenhong Sun, Zhi Wang 0001 |
NeurIPS | 8 |
| 2025 | VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation ModelsabstractRecent advancements in text-to-video (T2V) diffusion models have enabled high-fidelity and realistic video synthesis. However, current T2V models often struggle to generate physically plausible content due to their limited inherent ability to accurately understand physics. We found that while the representations within T2V models possess some capacity for physics understanding, they lag significantly behind those from recent video self-supervised learning methods. To this end, we propose a novel framework called {VideoREPA}, which distills physics understanding capability from video understanding foundation models into T2V models by aligning token-level relations. This closes the physics understanding gap and enables more physics-plausible generation. Specifically, we introduce the {Token Relation Distillation (TRD) loss}, leveraging spatio-temporal alignment to provide soft guidance suitable for finetuning powerful pre-trained T2V models—a critical departure from prior representation alignment (REPA) methods. To our knowledge, VideoREPA is the first REPA method designed for finetuning T2V models and specifically for injecting physical knowledge. Empirical evaluations show that VideoREPA substantially enhances the physics commonsense of baseline method, CogVideoX, achieving significant improvement on relevant benchmarks and demonstrating a strong capacity for generating videos consistent with intuitive physics. Code and more video results are available at https://videorepa.github.io/. Jiaqi Liao, Shaofeng Zhang, Fanqing Meng, Xiangpeng Wan, Junchi Yan, Yu Cheng 0001 |
NeurIPS | 7 |
| 2025 | Scaling Physical Reasoning with the PHYSICS DatasetabstractLarge Language Models (LLMs) have achieved remarkable progress on advanced reasoning tasks such as mathematics and coding competitions. Meanwhile, physics, despite being both reasoning-intensive and essential to real-world understanding, received limited academic and industrial attention. This paper introduces PHYSICS, a dataset containing 16,568 high-quality physics problems spanning subjects and difficulty levels, to facilitate this issue. Specifically, PHYSICS is curated with exercises from over 100 textbooks through a carefully designed pipeline for quality control. It covers five major physics domains: Mechanics, Electromagnetism, Thermodynamics, Optics, and Modern Physics. It also spans a wide range of difficulty levels, from high school to graduate-level physics courses. To utilize the data for improving and evaluating the model's physical reasoning capabilities, we split the dataset into training and test sets, and provide reasoning paths generated by powerful reasoning models for the training data to facilitate model training. In addition, for the evaluation part, we find that existing evaluation frameworks exhibit biases in aspects such as units, simplification, and precision in physics domain. To balance efficiency and accuracy, we introduce a Rule+Model evaluation framework tailored to physics problems. Our evaluations on current state-of-the-art open-source and proprietary models highlight the limitations of current models in handling physics-related tasks. We hope that our dataset and evaluation methodology will jointly advance the development of LLMs in the field of physics. The code and data can be found at: https://github.com/Zhengsh123/PHYSICS. Shenghe Zheng, Qianjia Cheng, Junchi Yao, Mengsong Wu, Ning Ding 0002, Yu Cheng 0001, Shuyue Hu, Lei Bai 0001, Dongzhan Zhou, Ganqu Cui, Peng Ye 0006 |
NeurIPS | 7 |
| 2025 | A Survey of Attacks on Large Vision-Language Models: Resources, Advances, and Future TrendsabstractWith the significant development of large models in recent years, large vision-language models (LVLMs) have demonstrated remarkable capabilities across a wide range of multimodal understanding and reasoning tasks. Compared with traditional large language models (LLMs), LVLMs present great potential and challenges due to their closer proximity to the multiresource real-world applications and the complexity of multimodal processing. However, the vulnerability of LVLMs is relatively underexplored, posing potential security risks in the daily use of LVLM applications. In this article, we provide a comprehensive review of the various forms of existing LVLM attacks. Specifically, we first introduce the background of attacks targeting LVLMs, including the attack preliminary, attack challenges, and attack resources. Then, we systematically review the development of LVLM attack methods, such as adversarial attacks that manipulate model outputs, jailbreak attacks that exploit model vulnerabilities for unauthorized actions, prompt injection attacks that engineer the prompt type and pattern, and data poisoning that affects model training. Finally, we discuss promising future research directions in LVLM attacks. We believe that our survey provides insights into the current landscape of LVLM vulnerabilities, inspiring more researchers to explore and mitigate potential safety issues in LVLM developments. Daizong Liu, Xiaoye Qu, Pan Zhou 0001, Yu Cheng 0001, Wei Hu 0003 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | Enhancing Low-Resource Relation Representations through Multi-View DecouplingabstractRecently, prompt-tuning with pre-trained language models (PLMs) has demonstrated the significantly enhancing ability of relation extraction (RE) tasks. However, in low-resource scenarios, where the available training data is scarce, previous prompt-based methods may still perform poorly for prompt-based representation learning due to a superficial understanding of the relation. To this end, we highlight the importance of learning high-quality relation representation in low-resource scenarios for RE, and propose a novel prompt-based relation representation method, named MVRE (Multi-View Relation Extraction), to better leverage the capacity of PLMs to improve the performance of RE within the low-resource prompt-tuning paradigm. Specifically, MVRE decouples each relation into different perspectives to encompass multi-view relation representations for maximizing the likelihood during relation inference. Furthermore, we also design a Global-Local loss and a Dynamic-Initialization method for better alignment of the multi-view relation-representing virtual words, containing the semantics of relation labels during the optimization learning process and initialization. Extensive experiments on three benchmark datasets show that our method can achieve state-of-the-art in low-resource settings. Chenghao Fan, Wei Wei 0002, Xiaoye Qu, Zhenyi Lu, Wenfeng Xie, Yu Cheng 0001, Dangyang Chen |
AAAI | 6 |
| 2024 | Unsupervised Domain Adaptative Temporal Sentence Localization with Mutual Information MaximizationabstractTemporal sentence localization (TSL) aims to localize a target segment in a video according to a given sentence query. Though respectable works have made decent achievements in this task, they severely rely on abundant yet expensive manual annotations for training. Moreover, these trained data-dependent models usually can not generalize well to unseen scenarios because of the inherent domain shift. To facilitate this issue, in this paper, we target another more practical but challenging setting: unsupervised domain adaptative temporal sentence localization (UDA-TSL), which explores whether the localization knowledge can be transferred from a fully-annotated data domain (source domain) to a new unannotated data domain (target domain). Particularly, we propose an effective and novel baseline for UDA-TSL to bridge the multi-modal gap across different domains and learn the potential correspondence between the video-query pairs in target domain. We first develop separate modality-specific domain adaptation modules to smoothly balance the minimization of the domain shifts in cross-dataset video and query domains. Then, to fully exploit the semantic correspondence of both modalities in target domain for unsupervised localization, we devise a mutual information learning module to adaptively align the video-query pairs which are more likely to be relevant in target domain, leading to more truly aligned target pairs and ensuring the discriminability of target features. In this way, our model can learn domain-invariant and semantic-aligned cross-modal representations. Three sets of migration experiments show that our model achieves competitive performance compared to existing methods. Daizong Liu, Xiaoye Qu, Jianfeng Dong, Yang Yang 0002, Pan Zhou 0001, Yu Cheng 0001 |
AAAI | 8 |
| 2024 | Confidence is not Timeless: Modeling Temporal Validity for Rule-based Temporal Knowledge Graph ForecastingabstractRecently, Temporal Knowledge Graph Forecasting (TKGF) has emerged as a pivotal domain for forecasting future events.Unlike black-box neural network methods, rule-based approaches are lauded for their efficiency and interpretability.For this line of work, it is crucial to correctly estimate the predictive effectiveness of the rules, i.e., the confidence.However, the existing literature lacks in-depth investigation into how confidence evolves with time.Moreover, inaccurate and heuristic confidence estimation limits the performance of rule-based methods.To alleviate such issues, we propose a framework named TempValid to explicitly model the temporal validity of rules for TKGF.Specifically, we design a time function to model the interaction between temporal information with confidence.TempValid conceptualizes confidence and other coefficients as learnable parameters to avoid inaccurate estimation and combinatorial explosion.Furthermore, we introduce a rule-adversarial negative sampling and a time-aware negative sampling strategies to facilitate TempValid learning.Extensive experiments show that TempValid significantly outperforms previous state-of-theart (SOTA) rule-based methods on six TKGF datasets.Moreover, it exhibits substantial advancements in cross-domain and resourceconstrained rule learning scenarios. Rikui Huang, Wei Wei 0002, Xiaoye Qu, Shengzhe Zhang, Dangyang Chen, Yu Cheng 0001 |
ACL (1) | 6 |
| 2024 | Living in the Moment: Can Large Language Models Grasp Co-Temporal Reasoning?abstractZhaochen Su, Juntao Li, Jun Zhang, Tong Zhu, Xiaoye Qu, Pan Zhou, Yan Bowen, Yu Cheng, Min Zhang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Zhaochen Su, Juntao Li 0005, Jun Zhang 0069, Tong Zhu 0002, Xiaoye Qu, Pan Zhou 0001, Yan Bowen, Yu Cheng 0001, Min Zhang 0005 |
ACL (1) | 8 |
| 2024 | Towards Robust Temporal Activity Localization Learning with Noisy LabelsabstractThis paper addresses the task of temporal activity localization (TAL). Although recent works have made significant progress in TAL research, almost all of them implicitly assume that the dense frame-level correspondences in each video-query pair are correctly annotated. However, in reality, such an assumption is extremely expensive and even impossible to satisfy due to subjective labeling. To alleviate this issue, in this paper, we explore a new TAL setting termed Noisy Temporal activity localization (NTAL), where a TAL model should be robust to the mixed training data with noisy moment boundaries. Inspired by the memorization effect of neural networks, we propose a novel method called Co-Teaching Regularizer (CTR) for NTAL. Specifically, we first learn a Gaussian Mixture Model to divide the mixed training data into preliminary clean and noisy subsets. Subsequently, we refine the labels of the two subsets by an adaptive prediction function so that their true positive and false positive samples could be identified. To avoid single model being prone to its mistakes learned by the mixed data, we adopt a co-teaching paradigm, which utilizes two models sharing the same framework to teach each other for robust learning. A curriculum strategy is further introduced to gradually learn the moment confidence from easy to hard. Experiments on three datasets demonstrate that our CTR is significantly more robust to the noisy training data compared to the existing methods. Daizong Liu, Xiaoye Qu, Jianfeng Dong, Pan Zhou 0001, Guoshun Nan, Keke Tang, Wanlong Fang, Yu Cheng 0001 |
LREC/COLING | 9 |
| 2024 | Rethinking Weakly-Supervised Video Temporal Grounding From a Game Perspective
Zeyu Xiong, Wanlong Fang, Xiaoye Qu, Chen Chen 0006, Jianfeng Dong, Keke Tang, Pan Zhou 0001, Yu Cheng 0001, Daizong Liu |
ECCV (45) | 9 |
| 2024 | Learning the Unlearned: Mitigating Feature Suppression in Contrastive Learning
Jihai Zhang 0002, Xiang Lan 0004, Xiaoye Qu, Yu Cheng 0001, Mengling Feng, Bryan Hooi |
ECCV (83) | 4 |
| 2024 | LLaMA-MoE: Building Mixture-of-Experts from LLaMA with Continual Pre-TrainingabstractMixture-of-Experts (MoE) has gained increasing popularity as a promising framework for scaling up large language models (LLMs).However, training MoE from scratch in a largescale setting still suffers from data-hungry and instability problems.Motivated by this limit, we investigate building MoE models from existing dense large language models.Specifically, based on the well-known LLaMA-2 7B model, we obtain an MoE model by: (1) Expert Construction, which partitions the parameters of original Feed-Forward Networks (FFNs) into multiple experts; (2) Continual pretraining, which further trains the transformed MoE model and additional gate networks.In this paper, we comprehensively explore different methods for expert construction and various data sampling strategies for continual pretraining.After these stages, our LLaMA-MoE models could maintain language abilities and route the input tokens to specific experts with part of the parameters activated.Empirically, by training 200B tokens, LLaMA-MoE-3.5Bmodels significantly outperform dense models that contain similar activation parameters. Tong Zhu 0002, Xiaoye Qu, Daize Dong, Jiacheng Ruan, Jingqi Tong, Conghui He, Yu Cheng 0001 |
EMNLP | 7 |
| 2024 | On the Universal Truthfulness Hyperplane Inside LLMsabstractWhile large language models (LLMs) have demonstrated remarkable abilities across various fields, hallucination remains a significant challenge.Recent studies have explored hallucinations through the lens of internal representations, proposing mechanisms to decipher LLMs' adherence to facts.However, these approaches often fail to generalize to out-of-distribution data, leading to concerns about whether internal representation patterns reflect fundamental factual awareness, or only overfit spurious correlations on the specific datasets.In this work, we investigate whether a universal truthfulness hyperplane that distinguishes the model's factually correct and incorrect outputs exists within the model.To this end, we scale up the number of training datasets and conduct an extensive evaluation -we train the truthfulness hyperplane on a diverse collection of over 40 datasets and examine its cross-task, cross-domain, and in-domain generalization.Our results indicate that increasing the diversity of the training datasets significantly enhances the performance in all scenarios, while the volume of data samples plays a less critical role.This finding supports the optimistic hypothesis that a universal truthfulness hyperplane may indeed exist within the model, offering promising directions for future research.Code is publicly available at https://github.com/hkust-nlp/ Universal_Truthfulness_Hyperplane.Tend to overfit Junteng Liu, Shiqi Chen 0002, Yu Cheng 0001, Junxian He |
EMNLP | 3 |
| 2024 | SURf: Teaching Large Vision-Language Models to Selectively Utilize Retrieved InformationabstractLarge Vision-Language Models (LVLMs) have become pivotal at the intersection of computer vision and natural language processing.However, the full potential of LVLMs' Retrieval-Augmented Generation (RAG) capabilities remains underutilized.Existing works either focus solely on the text modality or are limited to specific tasks.Moreover, most LVLMs struggle to selectively utilize retrieved information and are sensitive to irrelevant or misleading references.To address these challenges, we propose a self-refinement framework designed to teach LVLMs to Selectively Utilize Retrieved Information (SURf).Specifically, when given questions that are incorrectly answered by the LVLM backbone, we obtain references that help correct the answers (positive references) and those that do not (negative references).We then fine-tune the LVLM backbone using a combination of these positive and negative references.Our experiments across three tasks and seven datasets demonstrate that our framework significantly enhances LVLMs' ability to effectively utilize retrieved multimodal references and improves their robustness against irrelevant or misleading information.The source code is available at https://github.com/GasolSun36/SURf. * Work done during internship at Shanghai AI Laboratory.† Both are corresponding authors.How many apples in the images? VQAVanilla: There are three apples.The image depicting five apples on a tree...The picture shows 7 apples .... leaves...Ours: There are four apples.Describe this image in details. Captioning Vanilla: A person walking in snow.The image depicting a...the skier is in a crouched position... The image captures a dynamic scene ..a skier dressed in a ... Jiashuo Sun, Jihai Zhang 0002, Yucheng Zhou 0001, Zhaochen Su, Xiaoye Qu, Yu Cheng 0001 |
EMNLP | 6 |
| 2024 | Merge, Then Compress: Demystify Efficient SMoE with Hints from Its Routing PolicyabstractSparsely activated Mixture-of-Experts (SMoE) has shown promise to scale up the learning capacity of neural networks, however, they have issues like: ($a$) $\textit{High Memory Usage,}$ due to duplication of the network layers into multiple copies as experts; and ($b$) $\textit{Redundancy in Experts,}$ as common learning-based routing policies suffer from representational collapse. Therefore, vanilla SMoE models are memory inefficient and non-scalable, especially for resource-constrained downstream scenarios. In this paper, we ask: Can we craft a compact SMoE model by consolidating expert information? What is the best recipe to merge multiple experts into fewer but more knowledgeable experts? Our pilot investigation reveals that conventional model merging methods fail to be effective in such expert merging for SMoE. The potential reasons are: ($1$) redundant information overshadows critical experts; ($2$) appropriate neuron permutation for each expert is missing to bring all of them in alignment. To address these challenges, we propose a novel merging algorithm for SMoE, $\textit{i.e.}$, $\texttt{M-SMoE}$, which leverages routing statistics to guide expert merging. Specifically, it starts with neuron permutation alignment for experts; then, dominant experts and their "group members" are formed based on routing policies; lastly, every expert group is merged into a single expert by utilizing each expert's activation frequency as their weight for merging, thus diminishing the impact of insignificant experts. Moreover, we draw an interesting observation that our proposed merging promotes a low dimensionality in the merged expert's weight space, naturally paving the way for additional compression. Hence, our final method, $\texttt{MC-SMoE}$ ($\textit{i.e.}$, Merge, then Compress SMoE), further decomposes the merged experts into low-rank and structural sparse alternatives. Extensive experiments across $8$ benchmarks validate the effectiveness of our proposals. For instance, our $\texttt{MC-SMoE}$ achieves up to $80\%$ memory and a $20\%$ FLOPs reduction, with virtually no loss in performance. Our code is provided as supplementary material. Pingzhi Li, Zhenyu Zhang 0015, Prateek Yadav, Yi-Lin Sung, Yu Cheng 0001, Mohit Bansal, Tianlong Chen 0001 |
ICLR | 5 |
| 2024 | Sparse MoE with Language Guided Routing for Multilingual Machine TranslationabstractSparse Mixture-of-Experts (SMoE) has gained increasing popularity as a promising framework for scaling up multilingual machine translation (MMT) models with negligible extra computational overheads. However, current SMoE solutions neglect the intrinsic structures of the MMT problem: ($a$) $\textit{Linguistics Hierarchy.}$ Languages are naturally grouped according to their lingual properties like genetic families, phonological characteristics, etc; ($b$) $\textit{Language Complexity.}$ The learning difficulties are varied for diverse languages due to their grammar complexity, available resources, etc. Therefore, routing a fixed number of experts (e.g., $1$ or $2$ experts in usual) only at the word level leads to inferior performance. To fill in the missing puzzle, we propose $\textbf{\texttt{Lingual-SMoE}}$ by equipping the SMoE with adaptive and linguistic-guided routing policies. Specifically, it ($1$) extracts language representations to incorporate linguistic knowledge and uses them to allocate experts into different groups; ($2$) determines the number of activated experts for each target language in an adaptive and automatic manner, according to their translation difficulties, which aims to mitigate the potential over-/under-fitting issues of learning simple/challenges translations. Sufficient experimental studies on MMT benchmarks with {$16$, $50$, $100$} language pairs and various network architectures, consistently validate the superior performance of our proposals. For instance, $\texttt{Lingual-SMoE}$ outperforms its dense counterpart by over $5\%$ BLEU scores on $\texttt{OPUS-100}$ dataset. Xuxi Chen, Yu Cheng 0001, Tianlong Chen 0001 |
ICLR | 3 |
| 2024 | MoE-RBench: Towards Building Reliable Language Models with Sparse Mixture-of-Experts
Guanjie Chen, Tianlong Chen 0001, Yu Cheng 0001 |
ICML | 4 |
| 2024 | Not All Inputs Are Valid: Towards Open-Set Video Moment Retrieval using LanguageabstractVideo Moment Retrieval (VMR) targets to retrieve the specific moment corresponding to a sentence query from an untrimmed video. Although recent respectable works have made remarkable progress in this task, they implicitly are rooted in the closed-set assumption that all the given queries as video-relevant. Given an OOD query in open-set scenarios, they still utilize it for wrong retrieval, which might lead to irrecoverable losses in high-risk scenarios, e.g., criminal activity detection. To this end, we creatively explore a brand-new VMR setting termed Open-Set Video Moment Retrieval (OS-VMR), where we should not only retrieve the precise moments based on ID query, but also reject OOD queries. In this paper, we make the first attempt to step toward OS-VMR and propose a novel model OpenVMR, which first distinguishes ID and OOD queries based on the normalizing flow technology, and then conducts moment retrieval based on ID queries. Specifically, we first learn the ID distribution by constructing a normalizing flow, and assume the ID query distribution obeys the multi-variate Gaussian distribution. Then, we introduce an uncertainty score to search the ID-OOD separating boundary. After that, we refine the ID-OOD boundary by pulling together ID query features. Besides, video-query matching and frame-query matching are designed for coarse-grained and fine-grained cross-modal interaction, respectively. Finally, a positive-unlabeled learning module is introduced for moment retrieval. Experimental results on three VMR datasets show the effectiveness of our OpenVMR. Wanlong Fang, Daizong Liu, Xiaoye Qu, Jianfeng Dong, Pan Zhou 0001, Renfu Li, Zichuan Xu, Lixing Chen, Panpan Zheng, Yu Cheng 0001 |
ACM Multimedia | 11 |
| 2024 | MAGIS: LLM-Based Multi-Agent Framework for GitHub Issue ResolutionabstractIn software development, resolving the emergent issues within GitHub repositories is a complex challenge that involves not only the incorporation of new code but also the maintenance of existing code.
Large Language Models (LLMs) have shown promise in code generation but face difficulties in resolving Github issues, particularly at the repository level.
To overcome this challenge, we empirically study the reason why LLMs fail to resolve GitHub issues and analyze the major factors.
Motivated by the empirical findings, we propose a novel LLM-based **M**ulti-**A**gent framework for **G**itHub **I**ssue re**S**olution, **MAGIS**, consisting of four agents customized for software evolution: Manager, Repository Custodian, Developer, and Quality Assurance Engineer agents.
This framework leverages the collaboration of various agents in the planning and coding process to unlock the potential of LLMs to resolve GitHub issues.
In experiments, we employ the SWE-bench benchmark to compare MAGIS with popular LLMs, including GPT-3.5, GPT-4, and Claude-2.
MAGIS can resolve **13.94%** GitHub issues, significantly outperforming the baselines.
Specifically, MAGIS achieves an eight-fold increase in resolved ratio over the direct application of GPT-4, the advanced LLM. Wei Tao 0003, Yucheng Zhou 0001, Yanlin Wang 0001, Hongyu Zhang 0002, Yu Cheng 0001 |
NeurIPS | 6 |
| 2024 | On Giant's Shoulders: Effortless Weak to Strong by Dynamic Logits FusionabstractEfficient fine-tuning of large language models for task-specific applications is imperative, yet the vast number of parameters in these models makes their training increasingly challenging.
Despite numerous proposals for effective methods, a substantial memory overhead remains for gradient computations during updates. \thm{Can we fine-tune a series of task-specific small models and transfer their knowledge directly to a much larger model without additional training?}
In this paper, we explore weak-to-strong specialization using logit arithmetic, facilitating a direct answer to this question.
Existing weak-to-strong methods often employ a static knowledge transfer ratio and a single small model for transferring complex knowledge, which leads to suboptimal performance.
To surmount these limitations,
we propose a dynamic logit fusion approach that works with a series of task-specific small models, each specialized in a different task.
This method adaptively allocates weights among these models at each decoding step,
learning the weights through Kullback-Leibler divergence constrained optimization problems.
We conduct extensive experiments across various benchmarks in both single-task and multi-task settings, achieving leading results.
By transferring expertise from the 7B model to the 13B model, our method closes the performance gap by 96.4\% in single-task scenarios and by 86.3\% in multi-task scenarios compared to full fine-tuning of the 13B model. Notably, we achieve surpassing performance on unseen tasks. Moreover, we further demonstrate that our method can effortlessly integrate in-context learning for single tasks and task arithmetic for multi-task scenarios. Chenghao Fan, Zhenyi Lu, Wei Wei 0002, Xiaoye Qu, Dangyang Chen, Yu Cheng 0001 |
NeurIPS | 7 |
| 2024 | Twin-Merging: Dynamic Integration of Modular Expertise in Model MergingabstractIn the era of large language models, model merging is a promising way to combine multiple task-specific models into a single multitask model without extra training.
However, two challenges remain: (a) interference between different models and (b) heterogeneous data during testing. Traditional model merging methods often show significant performance gaps compared to fine-tuned models due to these issues.
Additionally, a one-size-fits-all model lacks flexibility for diverse test data, leading to performance degradation.
We show that both shared and exclusive task-specific knowledge are crucial for merging performance, but directly merging exclusive knowledge hinders overall performance.
In view of this, we propose Twin-Merging, a method that encompasses two principal stages:
(1) modularizing knowledge into shared and exclusive components, with compression to reduce redundancy and enhance efficiency;
(2) dynamically merging shared and task-specific knowledge based on the input.
This approach narrows the performance gap between merged and fine-tuned models and improves adaptability to heterogeneous data.
Extensive experiments on $20$ datasets for both language and vision tasks demonstrate the effectiveness of our method, showing an average improvement of $28.34\%$ in absolute normalized score for discriminative tasks and even surpassing the fine-tuned upper bound on the generative tasks. Zhenyi Lu, Chenghao Fan, Wei Wei 0002, Xiaoye Qu, Dangyang Chen, Yu Cheng 0001 |
NeurIPS | 6 |
| 2024 | ConflictBank: A Benchmark for Evaluating the Influence of Knowledge Conflicts in LLMs
Zhaochen Su, Jun Zhang 0069, Xiaoye Qu, Tong Zhu 0002, Yanshu Li, Jiashuo Sun, Juntao Li 0005, Min Zhang 0005, Yu Cheng 0001 |
NeurIPS | 9 |
| 2024 | ProS: Facial Omni-Representation Learning via Prototype-based Self-DistillationabstractThis paper presents a novel approach, called Prototype-based Self-Distillation (ProS), for unsupervised face representation learning. The existing supervised methods heavily rely on a large amount of annotated training facial data, which poses challenges in terms of data collection and privacy concerns. To address these issues, we propose ProS, which leverages a vast collection of unlabeled face images to learn a comprehensive facial omni-representation. In particular, ProS consists of two vision-transformers (teacher and student models) that are trained with different augmented images (cropping, blurring, coloring, etc.). Besides, we build a face-aware retrieval system along with augmentations to obtain the curated images comprising predominantly facial areas. To enhance the discrimination of learned features, we introduce a prototype-based matching loss that aligns the similarity distributions between features (teacher or student) and a set of learnable prototypes. After pre-training, the teacher vision transformer serves as a backbone for downstream tasks, including attribute estimation, expression recognition, and landmark alignment, achieved through simple fine-tuning with additional layers. Extensive experiments demonstrate that our method achieves state-of-the-art performance on various tasks, both in full and few-shot settings. Further, we investigate pre-training with synthetic face images, and ProS exhibits promising performance in this scenario as well. Xing Di, Yiyu Zheng, Xiaoming Liu 0002, Yu Cheng 0001 |
WACV | 4 |
| 2024 | Transform-Equivariant Consistency Learning for Temporal Sentence GroundingabstractThis paper addresses the temporal sentence grounding (TSG). Although existing methods have made decent achievements in this task, they not only severely rely on abundant video-query paired data for training, but also easily fail into the dataset distribution bias. To alleviate these limitations, we introduce a novel Equivariant Consistency Regulation Learning (ECRL) framework to learn more discriminative query-related frame-wise representations for each video, in a self-supervised manner. Our motivation comes from that the temporal boundary of the query-guided activity should be consistently predicted under various video-level transformations. Concretely, we first design a series of spatio-temporal augmentations on both foreground and background video segments to generate a set of synthetic video samples. In particular, we devise a self-refine module to enhance the completeness and smoothness of the augmented video. Then, we present a novel self-supervised consistency loss (SSCL) applied on the original and augmented videos to capture their invariant query-related semantic by minimizing the KL-divergence between the sequence similarity of two videos and a prior Gaussian distribution of timestamp distance. At last, a shared grounding head is introduced to predict the transform-equivariant query-guided segment boundaries for both the original and augmented videos. Extensive experiments on three challenging datasets (ActivityNet, TACoS, and Charades-STA) demonstrate both effectiveness and efficiency of our proposed ECRL framework. Daizong Liu, Xiaoye Qu, Jianfeng Dong, Pan Zhou 0001, Zichuan Xu, Haozhao Wang, Xing Di, Weining Lu, Yu Cheng 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 9 |
| 2023 | Frido: Feature Pyramid Diffusion for Complex Scene Image SynthesisabstractDiffusion models (DMs) have shown great potential for high-quality image synthesis. However, when it comes to producing images with complex scenes, how to properly describe both image global structures and object details remains a challenging task. In this paper, we present Frido, a Feature Pyramid Diffusion model performing a multi-scale coarse-to-fine denoising process for image synthesis. Our model decomposes an input image into scale-dependent vector quantized features, followed by a coarse-to-fine gating for producing image output. During the above multi-scale representation learning stage, additional input conditions like text, scene graph, or image layout can be further exploited. Thus, Frido can be also applied for conditional or cross-modality image synthesis. We conduct extensive experiments over various unconditioned and conditional image generation tasks, ranging from text-to-image synthesis, layout-to-image, scene-graph-to-image, to label-to-image. More specifically, we achieved state-of-the-art FID scores on five benchmarks, namely layout-to-image on COCO and OpenImages, scene-graph-to-image on COCO and Visual Genome, and label-to-image on COCO. Wan-Cyuan Fan, Yen-Chun Chen 0001, Dongdong Chen 0001, Yu Cheng 0001, Lu Yuan 0001, Yu-Chiang Frank Wang |
AAAI | 4 |
| 2023 | Hypotheses Tree Building for One-Shot Temporal Sentence LocalizationabstractGiven an untrimmed video, temporal sentence localization (TSL) aims to localize a specific segment according to a given sentence query. Though respectable works have made decent achievements in this task, they severely rely on dense video frame annotations, which require a tremendous amount of human effort to collect. In this paper, we target another more practical and challenging setting: one-shot temporal sentence localization (one-shot TSL), which learns to retrieve the query information among the entire video with only one annotated frame. Particularly, we propose an effective and novel tree-structure baseline for one-shot TSL, called Multiple Hypotheses Segment Tree (MHST), to capture the query-aware discriminative frame-wise information under the insufficient annotations. Each video frame is taken as the leaf-node, and the adjacent frames sharing the same visual-linguistic semantics will be merged into the upper non-leaf node for tree building. At last, each root node is an individual segment hypothesis containing the consecutive frames of its leaf-nodes. During the tree construction, we also introduce a pruning strategy to eliminate the interference of query-irrelevant nodes. With our designed self-supervised loss functions, our MHST is able to generate high-quality segment hypotheses for ranking and selection with the query. Experiments on two challenging datasets demonstrate that MHST achieves competitive performance compared to existing methods. Daizong Liu, Pan Zhou 0001, Xing Di, Weining Lu, Yu Cheng 0001 |
AAAI | 6 |
| 2023 | DSEE: Dually Sparsity-embedded Efficient Tuning of Pre-trained Language ModelsabstractXuxi Chen, Tianlong Chen, Weizhu Chen, Ahmed Hassan Awadallah, Zhangyang Wang, Yu Cheng. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Xuxi Chen, Tianlong Chen 0001, Weizhu Chen, Ahmed Awadallah 0001, Zhangyang Wang, Yu Cheng 0001 |
ACL (1) | 6 |
| 2023 | You Are Catching My Attention: Are Vision Transformers Bad Learners under Backdoor Attacks?abstractVision Transformers (ViTs), which made a splash in the field of computer vision (CV), have shaken the dominance of convolutional neural networks (CNNs). However, in the process of industrializing ViTs, backdoor attacks have brought severe challenges to security. The success of ViTs benefits from the self-attention mechanism. However, compared with CNNs, we find that this mechanism of capturing global information within patches makes ViTs more sensitive to patch-wise triggers. Under such observations, we delicately design a novel backdoor attack framework for ViTs, dubbed BadViT, which utilizes a universal patch-wise trigger to catch the model's attention from patches beneficial for classification to those with triggers, thereby manipulating the mechanism on which ViTs survive to confuse itself. Furthermore, we propose invisible variants of BadViT to increase the stealth of the attack by limiting the strength of the trigger perturbation. Through a large number of experiments, it is proved that BadViT is an efficient backdoor attack method against ViTs, which is less dependent on the number of poisons, with satisfactory convergence, and is transferable for downstream tasks. Furthermore, the risks inside of ViTs to backdoor attacks are also explored from the perspective of existing advanced defense schemes. Zenghui Yuan, Pan Zhou 0001, Yu Cheng 0001 |
CVPR | 4 |
| 2023 | Robustness Challenges in Model Distillation and Pruning for Natural Language UnderstandingabstractMengnan Du, Subhabrata Mukherjee, Yu Cheng, Milad Shokouhi, Xia Hu, Ahmed Hassan Awadallah. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023. Mengnan Du, Subhabrata Mukherjee, Yu Cheng 0001, Milad Shokouhi, Xia Ben Hu, Ahmed Awadallah 0001 |
EACL | 3 |
| 2023 | Adaptive Budget Allocation for Parameter-Efficient Fine-Tuning
Qingru Zhang, Minshuo Chen, Alexander Bukharin, Yu Cheng 0001, Weizhu Chen, Tuo Zhao |
ICLR | 5 |
| 2023 | Filling the Information Gap between Video and Query for Language-Driven Moment RetrievalabstractThis paper addresses the challenging task of language-driven moment retrieval. Previous methods are typically trained to localize the target moment corresponding to a single sentence query in a complicated video. However, this specific moment generally delivers richer contents than the query, i.e., the semantics of one query may miss certain object details or actions in the complex foreground-background visual contents. Such information imbalance between two modalities makes it difficult to finely align their representations. To this end, instead of training with a single query, we propose to utilize the diversity and complementarity among different queries corresponding to the same video moment for enriching the textual semantics. Specifically, we develop a Teacher-Student Moment Retrieval (TSMR) framework to fill this cross-modal information gap. A teacher model is trained to not only encode a certain query but also capture extra complementary queries to aggregate contextual semantics for obtaining more comprehensive moment-related query representations. Since the additional queries are inaccessible during inference, we further introduce an adaptive knowledge distillation mechanism to train a student model with a single query input by selectively absorbing the knowledge from the teacher model. In this manner, the student model is more robust to the cross-modal information gap during the moment retrieval guided by a single query. Experimental results on two benchmarks demonstrate the effectiveness of our proposed method. Daizong Liu, Xiaoye Qu, Jianfeng Dong, Guoshun Nan, Pan Zhou 0001, Zichuan Xu, Lixing Chen, Yu Cheng 0001 |
ACM Multimedia | 9 |
| 2023 | DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT ModelsabstractGenerative Pre-trained Transformer (GPT) models have exhibited exciting progress in capabilities, capturing the interest of practitioners and the public alike. Yet, while the literature on the trustworthiness of GPT models remains limited, practitioners have proposed employing capable GPT models for sensitive applications to healthcare and finance – where mistakes can be costly. To this end, this work proposes a comprehensive trustworthiness evaluation for large language models with a focus on GPT-4 and GPT-3.5, considering diverse perspectives – including toxicity, stereotype bias, adversarial robustness, out-of-distribution robustness, robustness on adversarial demonstrations, privacy, machine ethics, and fairness. Based on our evaluations, we discover previously unpublished vulnerabilities to trustworthiness threats. For instance, we find that GPT models can be easily misled to generate toxic and biased outputs and leak private information in both training data and conversation history. We also find that although GPT-4 is usually more trustworthy than GPT-3.5 on standard benchmarks, GPT-4 is more vulnerable given jailbreaking system or user prompts, potentially due to the reason that GPT-4 follows the (misleading) instructions more precisely. Our work illustrates a comprehensive trustworthiness evaluation of GPT models and sheds light on the trustworthiness gaps. Our benchmark is publicly available at https://decodingtrust.github.io/. Boxin Wang, Hengzhi Pei, Chulin Xie, Mintong Kang, Chejian Xu, Zidi Xiong, Ritik Dutta, Rylan Schaeffer, Sang T. Truong, Simran Arora, Mantas Mazeika, Dan Hendrycks, Zinan Lin 0001, Yu Cheng 0001, Oluwasanmi Koyejo, Dawn Song, Bo Li 0026 |
NeurIPS | 16 |
| 2023 | Learning Deep Generative Clustering via Mutual Information MaximizationabstractDeep clustering refers to joint representation learning and clustering using deep neural networks. Existing methods can be mainly categorized into two types: discriminative and generative methods. The former learns representations for clustering with discriminative mechanisms directly, and the latter estimate the latent distribution of each cluster for generating data points and then infers cluster assignments. Although generative methods have the advantage of estimating the latent distributions of clusters, their performances still significantly fall behind discriminative methods. In this work, we argue that this performance gap might be partly due to the overlap of data distribution of different clusters. In fact, there is little guarantee of generative methods to separate the distributions of different clusters in the data space. To tackle these problems, we theoretically prove that mutual information maximization promotes the separation of different clusters in the data space, which provides a theoretical justification for deep generative clustering with mutual information maximization. Our theoretical analysis directly leads to a model which integrates a hierarchical generative adversarial network and mutual information maximization. Moreover, we further propose three techniques and empirically show their effects to stabilize and enhance the model. The proposed approach notably outperforms other generative models for deep clustering on public benchmarks. Xiaojiang Yang, Junchi Yan, Yu Cheng 0001, Yizhe Zhang 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2022 | Efficient Robust Training via Backward SmoothingabstractAdversarial training is so far the most effective strategy in defending against adversarial examples. However, it suffers from high computational costs due to the iterative adversarial attacks in each training step. Recent studies show that it is possible to achieve fast Adversarial Training by performing a single-step attack with random initialization. However, such an approach still lags behind state-of-the-art adversarial training algorithms on both stability and model robustness. In this work, we develop a new understanding towards Fast Adversarial Training, by viewing random initialization as performing randomized smoothing for better optimization of the inner maximization problem. Following this new perspective, we also propose a new initialization strategy, backward smoothing, to further improve the stability and model robustness over single-step robust training methods. Experiments on multiple benchmarks demonstrate that our method achieves similar model robustness as the original TRADES method while using much less training time (~3x improvement with the same training schedule). Yu Cheng 0001, Zhe Gan, Quanquan Gu, Jingjing Liu 0001 |
AAAI | 2 |
| 2022 | Playing Lottery Tickets with Vision and LanguageabstractLarge-scale pre-training has recently revolutionized vision-and-language (VL) research. Models such as LXMERT and UNITER have significantly lifted the state of the art over a wide range of VL tasks. However, the large number of parameters in such models hinders their application in practice. In parallel, work on the lottery ticket hypothesis (LTH) has shown that deep neural networks contain small matching subnetworks that can achieve on par or even better performance than the dense networks when trained in isolation. In this work, we perform the first empirical study to assess whether such trainable subnetworks also exist in pre-trained VL models. We use UNITER as the main testbed (also test on LXMERT and ViLT), and consolidate 7 representative VL tasks for experiments, including visual question answering, visual commonsense reasoning, visual entailment, referring expression comprehension, image-text retrieval, GQA, and NLVR2. Through comprehensive analysis, we summarize our main findings as follows. (i) It is difficult to find subnetworks that strictly match the performance of the full model. However, we can find relaxed winning tickets at 50%-70% sparsity that maintain 99% of the full accuracy. (ii) Subnetworks found by task-specific pruning transfer reasonably well to the other tasks, while those found on the pre-training tasks at 60%/70% sparsity transfer universally, matching 98%/96% of the full accuracy on average over all the tasks. (iii) Besides UNITER, other models such as LXMERT and ViLT can also play lottery tickets. However, the highest sparsity we can achieve for ViLT is far lower than LXMERT and UNITER (30% vs. 70%). (iv) LTH also remains relevant when using other training methods (e.g., adversarial training). Zhe Gan, Yen-Chun Chen 0001, Tianlong Chen 0001, Yu Cheng 0001, Shuohang Wang, Jingjing Liu 0001, Zicheng Liu 0001 |
AAAI | 5 |
| 2022 | Memory-Guided Semantic Learning Network for Temporal Sentence GroundingabstractTemporal sentence grounding (TSG) is crucial and fundamental for video understanding. Although existing methods train well-designed deep networks with large amount of data, we find that they can easily forget the rarely appeared cases during training due to the off-balance data distribution, which influences the model generalization and leads to unsatisfactory performance. To tackle this issue, we propose a memory-augmented network, called Memory-Guided Semantic Learning Network (MGSL-Net), that learns and memorizes the rarely appeared content in TSG task. Specifically, our proposed model consists of three main parts: cross-modal interaction module, memory augmentation module, and heterogeneous attention module. We first align the given video-query pair by a cross-modal graph convolutional network, and then utilize memory module to record the cross-modal shared semantic features in the domain-specific persistent memory. During training, the memory slots are dynamically associated with both common and rare cases, alleviating the forgetting issue. In testing, the rare cases can thus be enhanced by retrieving the stored memories, leading to better generalization. At last, the heterogeneous attention module is utilized to integrate the enhanced multi-modal features in both video and query domains. Experimental results on three benchmarks show the superiority of our method on both effectiveness and efficiency, which substantially improves the accuracy not only on the entire dataset but also on the rare cases. Daizong Liu, Xiaoye Qu, Xing Di, Yu Cheng 0001, Zichuan Xu, Pan Zhou 0001 |
AAAI | 4 |
| 2022 | Unsupervised Temporal Video Grounding with Deep Semantic ClusteringabstractTemporal video grounding (TVG) aims to localize a target segment in a video according to a given sentence query. Though respectable works have made decent achievements in this task, they severely rely on abundant video-query paired data, which is expensive to collect in real-world scenarios. In this paper, we explore whether a video grounding model can be learned without any paired annotations. To the best of our knowledge, this paper is the first work trying to address TVG in an unsupervised setting. Considering there is no paired supervision, we propose a novel Deep Semantic Clustering Network (DSCNet) to leverage all semantic information from the whole query set to compose the possible activity in each video for grounding. Specifically, we first develop a language semantic mining module, which extracts implicit semantic features from the whole query set. Then, these language semantic features serve as the guidance to compose the activity in video via a video-based semantic aggregation module. Finally, we utilize a foreground attention branch to filter out the redundant background activities and refine the grounding results. To validate the effectiveness of our DSCNet, we conduct experiments on both ActivityNet Captions and Charades-STA datasets. The results demonstrate that our DSCNet achieves competitive performance, and even outperforms most weakly-supervised approaches. Daizong Liu, Xiaoye Qu, Yinzhen Wang, Xing Di, Yu Cheng 0001, Zichuan Xu, Pan Zhou 0001 |
AAAI | 6 |
| 2022 | A Good Prompt Is Worth Millions of Parameters: Low-resource Prompt-based Learning for Vision-Language ModelsabstractLarge pre-trained vision-language (VL) models can learn a new task with a handful of examples and generalize to a new task without fine-tuning.However, these VL models are hard to deploy for real-world applications due to their impractically huge sizes and slow inference speed.To solve this limitation, we study prompt-based low-resource learning of VL tasks with our proposed method, FEWVLM, relatively smaller than recent fewshot learners.For FEWVLM, we pre-train a sequence-to-sequence transformer model with prefix language modeling (PrefixLM) and masked language modeling (MaskedLM).Furthermore, we analyze the effect of diverse prompts for few-shot tasks.Experimental results on VQA show that FEWVLM with prompt-based learning outperforms Frozen (Tsimpoukelli et al., 2021) which is 31× larger than FEWVLM by 18.2% point and achieves comparable results to a 246× larger model, PICa (Yang et al., 2021).In our analysis, we observe that (1) prompts significantly affect zero-shot performance but marginally affect few-shot performance, (2) models with noisy prompts learn as quickly as hand-crafted prompts given larger training data, and (3) MaskedLM helps VQA tasks while PrefixLM boosts captioning performance.Our code is publicly available at https://github. com/woojeongjin/FewVLM * Work was mainly done while Woojeong Jin 0001, Yu Cheng 0001, Yelong Shen, Weizhu Chen, Xiang Ren 0001 |
ACL (1) | 2 |
| 2022 | The Principle of Diversity: Training Stronger Vision Transformers Calls for Reducing All Levels of RedundancyabstractVision transformers (ViTs) have gained increasing popularity as they are commonly believed to own higher mod-eling capacity and representation flexibility, than traditional convolutional networks. However, it is questionable whether such potential has been fully unleashed in prac-tice, as the learned ViTs often suffer from over-smoothening, yielding likely redundant models. Recent works made pre-liminary attempts to identify and alleviate such redundancy, e.g., via regularizing embedding similarity or re-injecting convolution-like structures. However, a “head-to-toe as-sessment” regarding the extent of redundancy in ViTs, and how much we could gain by thoroughly mitigating such, has been absent for this field. This paper, for the first time, systematically studies the ubiquitous existence of re-dundancy at all three levels: patch embedding, attention map, and weight space. In view of them, we advocate a principle of diversity for training ViTs, by presenting cor-responding regularizers that encourage the representation diversity and coverage at each of those levels, that enabling capturing more discriminative information. Extensive ex-periments on ImageNet with a number of ViT backbones validate the effectiveness of our proposals, largely eliminating the observed ViT redundancy and significantly boosting the model generalization. For example, our diversified DeiT obtains 0.70% ~ 1.76% accuracy boosts on ImageNet with highly reduced similarity. Our codes are fully available in https://github.com/VITA-Group/Diverse-ViT. Tianlong Chen 0001, Zhenyu Zhang 0015, Yu Cheng 0001, Ahmed Awadallah 0001, Zhangyang Wang |
CVPR | 3 |
| 2022 | Scalable Learning to Optimize: A Learned Optimizer Can Train Big Models
Xuxi Chen, Tianlong Chen 0001, Yu Cheng 0001, Weizhu Chen, Ahmed Awadallah 0001, Zhangyang Wang |
ECCV (23) | 3 |
| 2022 | DnA: Improving Few-Shot Transfer Learning with Low-Rank Decomposition and Alignment
Ziyu Jiang, Tianlong Chen 0001, Xuxi Chen, Yu Cheng 0001, Luowei Zhou, Lu Yuan 0001, Ahmed Awadallah 0001, Zhangyang Wang |
ECCV (20) | 4 |
| 2022 | Point Cloud Domain Adaptation via Masked Local 3D Structure Prediction
Hanxue Liang, Hehe Fan, Zhiwen Fan, Yi Wang 0076, Tianlong Chen 0001, Yu Cheng 0001, Zhangyang Wang |
ECCV (3) | 6 |
| 2022 | Learning Visual Representation from Modality-Shared Contrastive Language-Image Pre-training
Haoxuan You, Luowei Zhou, Bin Xiao 0004, Noel Codella, Yu Cheng 0001, Ruochen Xu, Shih-Fu Chang, Lu Yuan 0001 |
ECCV (27) | 5 |
| 2022 | Backdoor Attacks on Crowd CountingabstractCrowd counting is a regression task that estimates the number of people in a scene image, which plays a vital role in a range of safety-critical applications, such as video surveillance, traffic monitoring and flow control. In this paper, we investigate the vulnerability of deep learning based crowd counting models to backdoor attacks, a major security threat to deep learning. A backdoor attack implants a backdoor trigger into a target model via data poisoning so as to control the model's predictions at test time. Different from image classification models on which most of existing backdoor attacks have been developed and tested, crowd counting models are regression models that output multi-dimensional density maps, thus requiring different techniques to manipulate. In this paper, we propose two novel Density Manipulation Backdoor Attacks (DMBA- and DMBA+) to attack the model to produce arbitrarily large or small density estimations. Experimental results demonstrate the effectiveness of our DMBA attacks on five classic crowd counting models and four types of datasets. We also provide an in-depth analysis of the unique challenges of backdooring crowd counting models and reveal two key elements of effective attacks: 1) full and dense triggers and 2) manipulation of the ground truth counts or density maps. Our work could help evaluate the vulnerability of crowd counting models to potential backdoor attacks. Tailai Zhang, Xingjun Ma, Pan Zhou 0001, Jian Lou 0001, Zichuan Xu, Xing Di, Yu Cheng 0001, Lichao Sun 0001 |
ACM Multimedia | 8 |
| 2022 | M³ViT: Mixture-of-Experts Vision Transformer for Efficient Multi-task Learning with Model-Accelerator Co-designabstractMulti-task learning (MTL) encapsulates multiple learned tasks in a single model and often lets those tasks learn better jointly. Multi-tasking models have become successful and often essential for many sophisticated systems such as autonomous driving and indoor robots. However, when deploying MTL onto those real-world systems that are often resource-constrained or latency-sensitive, two prominent challenges arise: (i) during training, simultaneously optimizing all tasks is often difficult due to gradient conflicts across tasks, and the challenge is amplified when a growing number of tasks have to be squeezed into one compact model; (ii) at inference, current MTL regimes have to activate nearly the entire model even to just execute a single task. Yet most real systems demand only one or two tasks at each moment, while flexibly switching between tasks per need: therefore such “all tasks activated” inference is also highly inefficient and non-scalable in practice. In this paper, we present a model-accelerator co-design framework to enable efficient on-device MTL, that tackles both training and inference bottlenecks. Our framework, dubbed M³ViT, customizes mixture-of-experts (MoE) layers into a vision transformer (ViT) backbone for MTL, and sparsely activates task-specific experts during training, which effectively disentangles the parameter spaces to avoid different tasks’ training conflicts. Then at inference with any task of interest, the same design allows for activating only the task-corresponding sparse “expert” pathway, instead of the full model. Our new model design is further enhanced by hardware-level innovations, in particular, a novel computation reordering scheme tailored for memory-constrained MTL that achieves zero-overhead switching between tasks and can scale to any number of experts. Extensive experiments on PASCAL-Context and NYUD-v2 datasets at both software and hardware levels are conducted to demonstrate the effectiveness of the proposed design. When executing the practical scenario of single-task inference, M³ViT achieves higher accuracies than encoder-focused MTL methods, while significantly reducing 88% inference FLOPs. When implemented on a hardware platform of one Xilinx ZCU104 FPGA, our co-design framework reduces the memory requirement by 2.40×, while achieving energy efficiency (as the product of latency and power) up to 9.23× times higher than a comparable FPGA baseline. Hanxue Liang, Zhiwen Fan, Rishov Sarkar, Ziyu Jiang, Tianlong Chen 0001, Yu Cheng 0001, Cong Hao, Zhangyang Wang |
NeurIPS | 7 |
| 2021 | EarlyBERT: Efficient BERT Training via Early-bird Lottery TicketsabstractXiaohan Chen, Yu Cheng, Shuohang Wang, Zhe Gan, Zhangyang Wang, Jingjing Liu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Xiaohan Chen 0001, Yu Cheng 0001, Shuohang Wang, Zhe Gan, Zhangyang Wang, Jingjing Liu 0001 |
ACL/IJCNLP (1) | 2 |
| 2021 | Context-Aware Biaffine Localizing Network for Temporal Sentence GroundingabstractThis paper addresses the problem of temporal sentence grounding (TSG), which aims to identify the temporal boundary of a specific segment from an untrimmed video by a sentence query. Previous works either compare pre-defined candidate segments with the query and select the best one by ranking, or directly regress the boundary timestamps of the target segment. In this paper, we propose a novel localization framework that scores all pairs of start and end indices within the video simultaneously with a biaffine mechanism. In particular, we present a Context-aware Biaffine Localizing Network (CBLN) which incorporates both local and global contexts into features of each start/end position for biaffine-based localization. The local contexts from the adjacent frames help distinguish the visually similar appearance, and the global contexts from the entire video contribute to reasoning the temporal relation. Besides, we also develop a multi-modal self-attention module to provide fine-grained query-guided video representation for this biaffine strategy. Extensive experiments show that our CBLN significantly outperforms state-of-thearts on three public datasets (ActivityNet Captions, TACoS, and Charades-STA), demonstrating the effectiveness of the proposed localization framework. The code is available at https://github.com/liudaizong/CBLN. Daizong Liu, Xiaoye Qu, Jianfeng Dong, Pan Zhou 0001, Yu Cheng 0001, Wei Wei 0002, Zichuan Xu, Yulai Xie 0002 |
CVPR | 5 |
| 2021 | UC2: Universal Cross-Lingual Cross-Modal Vision-and-Language Pre-TrainingabstractVision-and-language pre-training has achieved impressive success in learning multimodal representations between vision and language. To generalize this success to non-English languages, we introduce UC2, the first machine translation-augmented framework for cross-lingual cross-modal representation learning. To tackle the scarcity problem of multilingual captions for image datasets, we first augment existing English-only datasets with other languages via machine translation (MT). Then we extend the standard Masked Language Modeling and Image-Text Matching training objectives to multilingual setting, where alignment between different languages is captured through shared visual context (i.e., using image as pivot). To facilitate the learning of a joint embedding space of images and all languages of interest, we further propose two novel pre-training tasks, namely Masked Region-to-Token Modeling (MRTM) and Visual Translation Language Modeling (VTLM), leveraging MT-enhanced translated data. Evaluation on multilingual image-text retrieval and multilingual visual question answering benchmarks demonstrates that our proposed framework achieves new state of the art on diverse non-English benchmarks while maintaining comparable performance to monolingual pre-trained models on English tasks. Mingyang Zhou 0004, Luowei Zhou, Shuohang Wang, Yu Cheng 0001, Zhou Yu 0005, Jingjing Liu 0001 |
CVPR | 4 |
| 2021 | InfoBERT: Improving Robustness of Language Models from An Information Theoretic Perspective
Boxin Wang, Shuohang Wang, Yu Cheng 0001, Zhe Gan, Ruoxi Jia 0001, Bo Li 0026, Jingjing Liu 0001 |
ICLR | 3 |
| 2021 | APo-VAE: Text Generation in Hyperbolic SpaceabstractShuyang Dai, Zhe Gan, Yu Cheng, Chenyang Tao, Lawrence Carin, Jingjing Liu. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Shuyang Dai, Zhe Gan, Yu Cheng 0001, Chenyang Tao, Lawrence Carin, Jingjing Liu 0001 |
NAACL-HLT | 3 |
| 2021 | Data-Efficient GAN Training Beyond (Just) Augmentations: A Lottery Ticket PerspectiveabstractTraining generative adversarial networks (GANs) with limited real image data generally results in deteriorated performance and collapsed models. To conquer this challenge, we are inspired by the latest observation, that one can discover independently trainable and highly sparse subnetworks (a.k.a., lottery tickets) from GANs. Treating this as an inductive prior, we suggest a brand-new angle towards data-efficient GAN training: by first identifying the lottery ticket from the original GAN using the small training set of real images; and then focusing on training that sparse subnetwork by re-using the same set. We find our coordinated framework to offer orthogonal gains to existing real image data augmentation methods, and we additionally present a new feature-level augmentation that can be applied together with them. Comprehensive experiments endorse the effectiveness of our proposed framework, across various GAN architectures (SNGAN, BigGAN, and StyleGAN-V2) and diverse datasets (CIFAR-10, CIFAR-100, Tiny-ImageNet, ImageNet, and multiple few-shot generation datasets). Codes are available at: https://github.com/VITA-Group/Ultra-Data-Efficient-GAN-Training. Tianlong Chen 0001, Yu Cheng 0001, Zhe Gan, Jingjing Liu 0001, Zhangyang Wang |
NeurIPS | 2 |
| 2021 | Chasing Sparsity in Vision Transformers: An End-to-End ExplorationabstractVision transformers (ViTs) have recently received explosive popularity, but their enormous model sizes and training costs remain daunting. Conventional post-training pruning often incurs higher training budgets. In contrast, this paper aims to trim down both the training memory overhead and the inference complexity, without sacrificing the achievable accuracy. We carry out the first-of-its-kind comprehensive exploration, on taking a unified approach of integrating sparsity in ViTs "from end to end''. Specifically, instead of training full ViTs, we dynamically extract and train sparse subnetworks, while sticking to a fixed small parameter budget. Our approach jointly optimizes model parameters and explores connectivity throughout training, ending up with one sparse network as the final output. The approach is seamlessly extended from unstructured to structured sparsity, the latter by considering to guide the prune-and-grow of self-attention heads inside ViTs. We further co-explore data and architecture sparsity for additional efficiency gains by plugging in a novel learnable token selector to adaptively determine the currently most vital patches. Extensive results on ImageNet with diverse ViT backbones validate the effectiveness of our proposals which obtain significantly reduced computational cost and almost unimpaired generalization. Perhaps most surprisingly, we find that the proposed sparse (co-)training can sometimes \textit{improve the ViT accuracy} rather than compromising it, making sparsity a tantalizing "free lunch''. For example, our sparsified DeiT-Small at ($5\%$, $50\%$) sparsity for (data, architecture), improves $\mathbf{0.28\%}$ top-1 accuracy, and meanwhile enjoys $\mathbf{49.32\%}$ FLOPs and $\mathbf{4.40\%}$ running time savings. Our codes are available at https://github.com/VITA-Group/SViTE. Tianlong Chen 0001, Yu Cheng 0001, Zhe Gan, Lu Yuan 0001, Lei Zhang 0001, Zhangyang Wang |
NeurIPS | 2 |
| 2021 | The Elastic Lottery Ticket HypothesisabstractLottery Ticket Hypothesis (LTH) raises keen attention to identifying sparse trainable subnetworks, or winning tickets, which can be trained in isolation to achieve similar or even better performance compared to the full models. Despite many efforts being made, the most effective method to identify such winning tickets is still Iterative Magnitude-based Pruning (IMP), which is computationally expensive and has to be run thoroughly for every different network. A natural question that comes in is: can we “transform” the winning ticket found in one network to another with a different architecture, yielding a winning ticket for the latter at the beginning, without re-doing the expensive IMP? Answering this question is not only practically relevant for efficient “once-for-all” winning ticket finding, but also theoretically appealing for uncovering inherently scalable sparse patterns in networks. We conduct extensive experiments on CIFAR-10 and ImageNet, and propose a variety of strategies to tweak the winning tickets found from different networks of the same model family (e.g., ResNets). Based on these results, we articulate the Elastic Lottery Ticket Hypothesis (E-LTH): by mindfully replicating (or dropping) and re-ordering layers for one network, its corresponding winning ticket could be stretched (or squeezed) into a subnetwork for another deeper (or shallower) network from the same family, whose performance is nearly the same competitive as the latter’s winning ticket directly found by IMP. We have also extensively compared E-LTH with pruning-at-initialization and dynamic sparse training methods, as well as discussed the generalizability of E-LTH to different model families, layer types, and across datasets. Code is available at https://github.com/VITA-Group/ElasticLTH. Xiaohan Chen 0001, Yu Cheng 0001, Shuohang Wang, Zhe Gan, Jingjing Liu 0001, Zhangyang Wang |
NeurIPS | 2 |
| 2021 | MaxVA: Fast Adaptation of Step Sizes by Maximizing Observed Variance of Gradients
Chen Zhu 0001, Yu Cheng 0001, Zhe Gan, Furong Huang, Jingjing Liu 0001, Tom Goldstein |
ECML/PKDD (3) | 2 |
| 2021 | Meta Module Network for Compositional Visual ReasoningabstractNeural Module Network (NMN) exhibits strong interpretability and compositionality thanks to its handcrafted neural modules with explicit multi-hop reasoning capability. However, most NMNs suffer from two critical draw-backs: 1) scalability: customized module for specific function renders it impractical when scaling up to a larger set of functions in complex tasks; 2) generalizability: rigid pre-defined module inventory makes it difficult to generalize to unseen functions in new tasks/domains. To design a more powerful NMN architecture for practical use, we propose Meta Module Network (MMN) centered on a novel meta module, which can take in function recipes and morph into diverse instance modules dynamically. The instance modules are then woven into an execution graph for complex visual reasoning, inheriting the strong explainability and compositionality of NMN. With such a flexible instantiation mechanism, the parameters of instance modules are inherited from the central meta module, retaining the same model complexity as the function set grows, which promises better scalability. Meanwhile, as functions are encoded into the embedding space, unseen functions can be readily represented based on its structural similarity with previously observed ones, which ensures better generalizability. Experiments on GQA and CLEVR datasets validate the superiority of MMN over state-of-the-art NMN designs. Synthetic experiments on held-out unseen functions from GQA dataset also demonstrate the strong generalizability of MMN. Our code and model are released in Github1. Wenhu Chen, Zhe Gan, Yu Cheng 0001, William Yang Wang, Jingjing Liu 0001 |
WACV | 4 |
| 2021 | EnlightenGAN: Deep Light Enhancement Without Paired SupervisionabstractDeep learning-based methods have achieved remarkable success in image restoration and enhancement, but are they still competitive when there is a lack of paired training data? As one such example, this paper explores the low-light image enhancement problem, where in practice it is extremely challenging to simultaneously take a low-light and a normal-light photo of the same visual scene. We propose a highly effective unsupervised generative adversarial network, dubbed EnlightenGAN, that can be trained without low/normal-light image pairs, yet proves to generalize very well on various real-world test images. Instead of supervising the learning using ground truth data, we propose to regularize the unpaired training using the information extracted from the input itself, and benchmark a series of innovations for the low-light image enhancement problem, including a global-local discriminator structure, a self-regularized perceptual loss fusion, and the attention mechanism. Through extensive experiments, our proposed approach outperforms recent methods under a variety of metrics in terms of visual quality and subjective user study. Thanks to the great flexibility brought by unpaired training, EnlightenGAN is demonstrated to be easily adaptable to enhancing real-world images from various domains. Our codes and pre-trained models are available at: https://github.com/VITA-Group/EnlightenGAN. Yifan Jiang 0001, Xinyu Gong, Ding Liu 0001, Yu Cheng 0001, Xiaohui Shen, Jianchao Yang, Pan Zhou 0001, Zhangyang Wang |
IEEE Trans. Image Process. | 4 |
| 2021 | Bayesian Cycle-Consistent Generative Adversarial Networks via Marginalizing Latent SamplingabstractRecent techniques built on generative adversarial networks (GANs), such as cycle-consistent GANs, are able to learn mappings among different domains built from unpaired data sets, through min-max optimization games between generators and discriminators. However, it remains challenging to stabilize the training process and thus cyclic models fall into mode collapse accompanied by the success of discriminator. To address this problem, we propose an novel Bayesian cyclic model and an integrated cyclic framework for interdomain mappings. The proposed method motivated by Bayesian GAN explores the full posteriors of cyclic model via sampling latent variables and optimizes the model with maximum a posteriori (MAP) estimation. Hence, we name it Bayesian CycleGAN. In addition, original CycleGAN cannot generate diversified results. But it is feasible for Bayesian framework to diversify generated images by replacing restricted latent variables in inference process. We evaluate the proposed Bayesian CycleGAN on multiple benchmark data sets, including Cityscapes, Maps, and Monet2photo. The proposed method improve the per-pixel accuracy by 15% for the Cityscapes semantic segmentation task within origin framework and improve 20% within the proposed integrated framework, showing better resilience to imbalance confrontation. The diversified results of Monet2Photo style transfer also demonstrate its superiority over original cyclic model. We provide codes for all of our experiments in https://github.com/ranery/Bayesian-CycleGAN. Haoran You, Yu Cheng 0001, Tianheng Cheng, Chun-Liang Li, Pan Zhou 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2020 | What Makes A Good Story? Designing Composite Rewards for Visual StorytellingabstractPrevious storytelling approaches mostly focused on optimizing traditional metrics such as BLEU, ROUGE and CIDEr. In this paper, we re-examine this problem from a different angle, by looking deep into what defines a natural and topically-coherent story. To this end, we propose three assessment criteria: relevance, coherence and expressiveness, which we observe through empirical analysis could constitute a “high-quality” story to the human eye. We further propose a reinforcement learning framework, ReCo-RL, with reward functions designed to capture the essence of these quality criteria. Experiments on the Visual Storytelling Dataset (VIST) with both automatic and human evaluation demonstrate that our ReCo-RL model achieves better performance than state-of-the-art baselines on both traditional metrics and the proposed new criteria. Junjie Hu 0001, Yu Cheng 0001, Zhe Gan, Jingjing Liu 0001, Jianfeng Gao 0001, Graham Neubig |
AAAI | 2 |
| 2020 | Contrastively Smoothed Class Alignment for Unsupervised Domain Adaptation
Shuyang Dai, Yu Cheng 0001, Yizhe Zhang 0002, Zhe Gan, Jingjing Liu 0001, Lawrence Carin |
ACCV (4) | 2 |
| 2020 | Distilling Knowledge Learned in BERT for Text GenerationabstractLarge-scale pre-trained language model such as BERT has achieved great success in language understanding tasks.However, it remains an open question how to utilize BERT for language generation.In this paper, we present a novel approach, Conditional Masked Language Modeling (C-MLM), to enable the finetuning of BERT on target generation tasks.The finetuned BERT (teacher) is exploited as extra supervision to improve conventional Seq2Seq models (student) for better text generation performance.By leveraging BERT's idiosyncratic bidirectional nature, distilling knowledge learned in BERT can encourage auto-regressive Seq2Seq models to plan ahead, imposing global sequence-level supervision for coherent text generation.Experiments show that the proposed approach significantly outperforms strong Transformer baselines on multiple language generation tasks such as machine translation and text summarization.Our proposed model also achieves new state of the art on IWSLT German-English and English-Vietnamese MT datasets.1 Yen-Chun Chen 0001, Zhe Gan, Yu Cheng 0001, Jingzhou Liu, Jingjing Liu 0001 |
ACL | 3 |
| 2020 | INSET: Sentence Infilling with INter-SEntential TransformerabstractMissing sentence generation (or sentence infilling) fosters a wide range of applications in natural language generation, such as document auto-completion and meeting note expansion.This task asks the model to generate intermediate missing sentences that can syntactically and semantically bridge the surrounding context.Solving the sentence infilling task requires techniques in natural language processing ranging from understanding to discourselevel planning to generation.In this paper, we propose a framework to decouple the challenge and address these three aspects respectively, leveraging the power of existing largescale pre-trained models such as BERT and GPT-2.We empirically demonstrate the effectiveness of our model in learning a sentence representation for generation and further generating a missing sentence that fits the context. Yizhe Zhang 0002, Oussama Elachqar, Yu Cheng 0001 |
ACL | 4 |
| 2020 | Discourse-Aware Neural Extractive Text SummarizationabstractRecently BERT has been adopted for document encoding in state-of-the-art text summarization models.However, sentence-based extractive models often result in redundant or uninformative phrases in the extracted summaries.Also, long-range dependencies throughout a document are not well captured by BERT, which is pre-trained on sentence pairs instead of documents.To address these issues, we present a discourse-aware neural summarization model -DISCOBERT 1 .DISCOBERT extracts sub-sentential discourse units (instead of sentences) as candidates for extractive selection on a finer granularity.To capture the long-range dependencies among discourse units, structural discourse graphs are constructed based on RST trees and coreference mentions, encoded with Graph Convolutional Networks.Experiments show that the proposed model outperforms state-of-the-art methods by a significant margin on popular summarization benchmarks compared to other BERT-base models. Zhe Gan, Yu Cheng 0001, Jingjing Liu 0001 |
ACL | 3 |
| 2020 | Adversarial Robustness: From Self-Supervised Pre-Training to Fine-TuningabstractPretrained models from self-supervision are prevalently used in fine-tuning downstream tasks faster or for better accuracy. However, gaining robustness from pretraining is left unexplored. We introduce adversarial training into self-supervision, to provide general-purpose robust pretrained models for the first time. We find these robust pretrained models can benefit the subsequent fine-tuning in two ways: i) boosting final model robustness; ii) saving the computation cost, if proceeding towards adversarial fine-tuning. We conduct extensive experiments to demonstrate that the proposed framework achieves large performance margins (eg, 3.83% on robust accuracy and 1.3% on standard accuracy, on the CIFAR-10 dataset), compared with the conventional end-to-end adversarial training baseline. Moreover, we find that different self-supervised pretrained models have diverse adversarial vulnerability. It inspires us to ensemble several pretraining tasks, which boosts robustness more. Our ensemble strategy contributes to a further improvement of 3.59% on robust accuracy, while maintaining a slightly higher standard accuracy on CIFAR-10. Our codes are available at https://github.com/TAMU-VITA/Adv-SS-Pretraining. Tianlong Chen 0001, Sijia Liu 0001, Shiyu Chang, Yu Cheng 0001, Lisa Amini, Zhangyang Wang |
CVPR | 4 |
| 2020 | BachGAN: High-Resolution Image Synthesis From Salient Object LayoutabstractWe propose a new task towards more practical applications for image generation - high-quality image synthesis from salient object layout. This new setting requires users to provide only the layout of salient objects (i.e., foreground bounding boxes and categories) and lets the model complete the drawing with an invented background and a matching foreground. Two main challenges spring from this new task: (i) how to generate fine-grained details and realistic textures without segmentation map input; and (ii) how to create and weave a background into standalone objects in a seamless way. To tackle this, we propose Background Hallucination Generative Adversarial Network (BachGAN), which leverages a background retrieval module to first select a set of segmentation maps from a large candidate pool, then encodes these candidate layouts via a background fusion module to hallucinate a suitable background for the given objects. By generating the hallucinated background representation dynamically, our model can synthesize high-resolution images with both photo-realistic foreground and integral background. Experiments on Cityscapes and ADE20K datasets demonstrate the advantage of BachGAN over existing approaches, measured on both visual fidelity of generated images and visual alignment between output images and input layouts. Yandong Li, Yu Cheng 0001, Zhe Gan, Licheng Yu, Liqiang Wang 0001, Jingjing Liu 0001 |
CVPR | 2 |
| 2020 | Violin: A Large-Scale Dataset for Video-and-Language InferenceabstractWe introduce a new task, Video-and-Language Inference, for joint multimodal understanding of video and text. Given a video clip with aligned subtitles as premise, paired with a natural language hypothesis based on the video content, a model needs to infer whether the hypothesis is entailed or contradicted by the given video clip. A new large-scale dataset, named Violin (VIdeO-and-Language INference), is introduced for this task, which consists of 95,322 video-hypothesis pairs from 15,887 video clips, spanning over 582 hours of video. These video clips contain rich content with diverse temporal dynamics, event shifts, and people interactions, collected from two sources: (i) popular TV shows, and (ii) movie clips from YouTube channels. In order to address our new multimodal inference task, a model is required to possess sophisticated reasoning skills, from surface-level grounding (e.g., identifying objects and characters in the video) to in-depth commonsense reasoning (e.g., inferring causal relations of events in the video). We present a detailed analysis of the dataset and an extensive evaluation over many strong baselines, providing valuable insights on the challenges of this new task. Jingzhou Liu, Wenhu Chen, Yu Cheng 0001, Zhe Gan, Licheng Yu, Yiming Yang 0002, Jingjing Liu 0001 |
CVPR | 3 |
| 2020 | Behind the Scene: Revealing the Secrets of Pre-trained Vision-and-Language Models
Jize Cao, Zhe Gan, Yu Cheng 0001, Licheng Yu, Yen-Chun Chen 0001, Jingjing Liu 0001 |
ECCV (6) | 3 |
| 2020 | UNITER: UNiversal Image-TExt Representation Learning
Yen-Chun Chen 0001, Licheng Yu, Ahmed El Kholy, Faisal Ahmed 0001, Zhe Gan, Yu Cheng 0001, Jingjing Liu 0001 |
ECCV (30) | 7 |
| 2020 | Multi-Fact Correction in Abstractive Text SummarizationabstractPre-trained neural abstractive summarization systems have dominated extractive strategies on news summarization performance, at least in terms of ROUGE.However, systemgenerated abstractive summaries often face the pitfall of factual inconsistency: generating incorrect facts with respect to the source text.To address this challenge, we propose Span-Fact, a suite of two factual correction models that leverages knowledge learned from question answering models to make corrections in system-generated summaries via span selection.Our models employ single or multimasking strategies to either iteratively or autoregressively replace entities in order to ensure semantic consistency w.r.t. the source text, while retaining the syntactic structure of summaries generated by abstractive summarization models.Experiments show that our models significantly boost the factual consistency of system-generated summaries without sacrificing summary quality in terms of both automatic metrics and human evaluation.* *Most of this work was done when the first author was an intern at Microsoft.CNNDM Source (CNN) About a quarter of a million Australian homes and businesses have no power after a "once in a decade" storm battered Sydney and nearby areas.About 4,500 people Yue Dong 0002, Shuohang Wang, Zhe Gan, Yu Cheng 0001, Jackie Chi Kit Cheung, Jingjing Liu 0001 |
EMNLP (1) | 4 |
| 2020 | HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-trainingabstractWe present HERO, a novel framework for large-scale video+language omnirepresentation learning.HERO encodes multimodal inputs in a hierarchical structure, where local context of a video frame is captured by a Cross-modal Transformer via multimodal fusion, and global video context is captured by a Temporal Transformer.In addition to standard Masked Language Modeling (MLM) and Masked Frame Modeling (MFM) objectives, we design two new pre-training tasks: (i) Video-Subtitle Matching (VSM), where the model predicts both global and local temporal alignment; and (ii) Frame Order Modeling (FOM), where the model predicts the right order of shuffled video frames.HERO is jointly trained on HowTo100M and large-scale TV datasets to gain deep understanding of complex social dynamics with multi-character interactions.Comprehensive experiments demonstrate that HERO achieves new state of the art on multiple benchmarks over Text-based Video/Video-moment Retrieval, Video Question Answering (QA), Video-and-language Inference and Video Captioning tasks across different domains.We also introduce two new challenging benchmarks How2QA and How2R for Video QA and Retrieval, collected from diverse video content over multimodalities. 1 Yen-Chun Chen 0001, Yu Cheng 0001, Zhe Gan, Licheng Yu, Jingjing Liu 0001 |
EMNLP (1) | 3 |
| 2020 | Contrastive Distillation on Intermediate Representations for Language Model CompressionabstractExisting language model compression methods mostly use a simple L 2 loss to distill knowledge in the intermediate representations of a large BERT model to a smaller one.Although widely used, this objective by design assumes that all the dimensions of hidden representations are independent, failing to capture important structural knowledge in the intermediate layers of the teacher network.To achieve better distillation efficacy, we propose Contrastive Distillation on Intermediate Representations (CODIR), a principled knowledge distillation framework where the student is trained to distill knowledge through intermediate layers of the teacher via a contrastive objective.By learning to distinguish positive sample from a large set of negative samples, CoDIR facilitates the student's exploitation of rich information in teacher's hidden layers.CoDIR can be readily applied to compress large-scale language models in both pretraining and finetuning stages, and achieves superb performance on the GLUE benchmark, outperforming state-of-the-art compression methods. 1 Zhe Gan, Yuwei Fang, Yu Cheng 0001, Shuohang Wang, Jingjing Liu 0001 |
EMNLP (1) | 4 |
| 2020 | Cross-Thought for Sentence Encoder Pre-trainingabstractIn this paper, we propose Cross-Thought, a novel approach to pre-training sequence encoder, which is instrumental in building reusable sequence embeddings for large-scale NLP tasks such as question answering.Instead of using the original signals of full sentences, we train a Transformer-based sequence encoder over a large set of short sequences, which allows the model to automatically select the most useful information for predicting masked words.Experiments on question answering and textual entailment tasks demonstrate that our pre-trained encoder can outperform state-of-the-art encoders trained with continuous sentence signals as well as traditional masked language modeling baselines.Our proposed approach also achieves new state of the art on HotpotQA (full-wiki setting) by improving intermediate information retrieval performance.1 Shuohang Wang, Yuwei Fang, Zhe Gan, Yu Cheng 0001, Jingjing Liu 0001, Jing Jiang 0001 |
EMNLP (1) | 5 |
| 2020 | FreeLB: Enhanced Adversarial Training for Natural Language Understanding
Chen Zhu 0001, Yu Cheng 0001, Zhe Gan, Tom Goldstein, Jingjing Liu 0001 |
ICLR | 2 |
| 2020 | Graph Optimal Transport for Cross-Domain AlignmentabstractCross-domain alignment between two sets of entities (e.g., objects in an image, words in a sentence) is fundamental to both computer vision and natural language processing. Existing methods mainly focus on designing advanced attention mechanisms to simulate soft alignment, where no training signals are provided to explicitly encourage alignment. Plus, the learned attention matrices are often dense and difficult to interpret. We propose Graph Optimal Transport (GOT), a principled framework that builds upon recent advances in Optimal Transport (OT). In GOT, cross-domain alignment is formulated as a graph matching problem, by representing entities as a dynamically-constructed graph. Two types of OT distances are considered: (i) Wasserstein distance (WD) for node (entity) matching; and (ii) Gromov-Wasserstein distance (GWD) for edge (structure) matching. Both WD and GWD can be incorporated into existing neural network models, effectively acting as a drop-in regularizer. The inferred transport plan also yields sparse and self-normalized alignment, enhancing the interpretability of the learned model. Experiments show consistent outperformance of GOT over baselines across a wide range of tasks, including image-text retrieval, visual question answering, image captioning, machine translation, and text summarization. Liqun Chen 0001, Zhe Gan, Yu Cheng 0001, Lawrence Carin, Jingjing Liu 0001 |
ICML | 3 |
| 2020 | Sequential Attention GAN for Interactive Image EditingabstractMost existing text-to-image synthesis tasks are static single-turn generation, based on pre-defined textual descriptions of images. To explore more practical and interactive real-life applications, we introduce a new task - Interactive Image Editing, where users can guide an agent to edit images via multi-turn textual commands on-the-fly. In each session, the agent takes a natural language description from the user as the input, and modifies the image generated in previous turn to a new design, following the user description. The main challenges in this sequential and interactive image generation task are two-fold: 1) contextual consistency between a generated image and the provided textual description; 2) step-by-step region-level modification to maintain visual consistency across the generated image sequence in each session. To address these challenges, we propose a novel Sequential Attention Generative Adversarial Network (SeqAttnGAN), which applies a neural state tracker to encode the previous image and the textual description in each turn of the sequence, and uses a GAN framework to generate a modified version of the image that is consistent with the preceding images and coherent with the description. To achieve better region-specific refinement, we also introduce a sequential attention mechanism into the model. To benchmark on the new task, we introduce two new datasets, Zap-Seq and DeepFashion-Seq, which contain multi-turn sessions with image-description sequences in the fashion domain. Experiments on both datasets show that the proposed SeqAttnGAN model outperforms state-of-the-art approaches on the interactive image editing task across all evaluation metrics including visual quality, image sequence coherence and text-image consistency. Yu Cheng 0001, Zhe Gan, Yitong Li 0001, Jingjing Liu 0001, Jianfeng Gao 0001 |
ACM Multimedia | 1 |
| 2020 | Fine-grained Iterative Attention Network for Temporal Language Localization in VideosabstractTemporal language localization in videos aims to ground one video segment in an untrimmed video based on a given sentence query. To tackle this task, designing an effective model to extract ground-ing information from both visual and textual modalities is crucial. However, most previous attempts in this field only focus on unidirectional interactions from video to query, which emphasizes which words to listen and attends to sentence information via vanilla soft attention, but clues from query-by-video interactions implying where to look are not taken into consideration. In this paper, we propose a Fine-grained Iterative Attention Network (FIAN) that consists of an iterative attention module for bilateral query-video in-formation extraction. Specifically, in the iterative attention module, each word in the query is first enhanced by attending to each frame in the video through fine-grained attention, then video iteratively attends to the integrated query. Finally, both video and query information is utilized to provide robust cross-modal representation for further moment localization. In addition, to better predict the target segment, we propose a content-oriented localization strategy instead of applying recent anchor-based localization. We evaluate the proposed method on three challenging public benchmarks: ActivityNet Captions, TACoS, and Charades-STA. FIAN significantly outperforms the state-of-the-art approaches. Xiaoye Qu, Pengwei Tang, Zhikang Zou, Yu Cheng 0001, Jianfeng Dong, Pan Zhou 0001, Zichuan Xu |
ACM Multimedia | 4 |
| 2020 | Large-Scale Adversarial Training for Vision-and-Language Representation LearningabstractWe present VILLA, the first known effort on large-scale adversarial training for vision-and-language (V+L) representation learning. VILLA consists of two training stages: (i) task-agnostic adversarial pre-training; followed by (ii) task-specific adversarial finetuning. Instead of adding adversarial perturbations on image pixels and textual tokens, we propose to perform adversarial training in the embedding space of each modality. To enable large-scale training, we adopt the ``free'' adversarial training strategy, and combine it with KL-divergence-based regularization to promote higher invariance in the embedding space. We apply VILLA to current best-performing V+L models, and achieve new state of the art on a wide range of tasks, including Visual Question Answering, Visual Commonsense Reasoning, Image-Text Retrieval, Referring Expression Comprehension, Visual Entailment, and NLVR2. Zhe Gan, Yen-Chun Chen 0001, Chen Zhu 0001, Yu Cheng 0001, Jingjing Liu 0001 |
NeurIPS | 5 |
| 2019 | Multi-step Reasoning via Recurrent Dual Attention for Visual DialogabstractThis paper presents a new model for visual dialog, Recurrent Dual Attention Network (ReDAN), using multi-step reasoning to answer a series of questions about an image. In each question-answering turn of a dialog, ReDAN infers the answer progressively through multiple reasoning steps. In each step of the reasoning process, the semantic representation of the question is updated based on the image and the previous dialog history, and the recurrently-refined representation is used for further reasoning in the subsequent step. On the VisDial v1.0 dataset, the proposed ReDAN model achieves a new state-of-the-art of 64.47% NDCG score. Visualization on the reasoning process further demonstrates that ReDAN can locate context-relevant visual and textual clues via iterative refinement, which can lead to the correct answer step-by-step. Zhe Gan, Yu Cheng 0001, Ahmed El Kholy, Jingjing Liu 0001, Jianfeng Gao 0001 |
ACL (1) | 2 |
| 2019 | StoryGAN: A Sequential Conditional GAN for Story VisualizationabstractIn this work, we propose a new task called Story Visualization. Given a multi-sentence paragraph, the story is visualized by generating a sequence of images, one for each sentence. In contrast to video generation, story visualization focuses less on the continuity in generated images (frames), but more on the global consistency across dynamic scenes and characters -- a challenge that has not been addressed by any single-image or video generation methods. Therefore, we propose a new story-to-image-sequence generation model, StoryGAN, based on the sequential conditional GAN framework. Our model is unique in that it consists of a deep Context Encoder that dynamically tracks the story flow, and two discriminators at the story and image levels, to enhance the image quality and the consistency of the generated sequences. To evaluate the model, we modified existing datasets to create the CLEVR-SV and Pororo-SV datasets. Empirically, StoryGAN outperformed state-of-the-art models in image quality, contextual consistency metrics, and human evaluation. Yitong Li 0001, Zhe Gan, Yelong Shen, Jingjing Liu 0001, Yu Cheng 0001, Yuexin Wu, Lawrence Carin, David E. Carlson, Jianfeng Gao 0001 |
CVPR | 5 |
| 2019 | Domain Adaptive Text Style TransferabstractDianqi Li, Yizhe Zhang, Zhe Gan, Yu Cheng, Chris Brockett, Bill Dolan, Ming-Ting Sun. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Dianqi Li, Yizhe Zhang 0002, Zhe Gan, Yu Cheng 0001, Chris Brockett, William B. Dolan, Ming-Ting Sun |
EMNLP/IJCNLP (1) | 4 |
| 2019 | Patient Knowledge Distillation for BERT Model CompressionabstractSiqi Sun, Yu Cheng, Zhe Gan, Jingjing Liu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Yu Cheng 0001, Zhe Gan, Jingjing Liu 0001 |
EMNLP/IJCNLP (1) | 2 |
| 2019 | Relation-Aware Graph Attention Network for Visual Question AnsweringabstractIn order to answer semantically-complicated questions about an image, a Visual Question Answering (VQA) model needs to fully understand the visual scene in the image, especially the interactive dynamics between different objects. We propose a Relation-aware Graph Attention Network (ReGAT), which encodes each image into a graph and models multi-type inter-object relations via a graph attention mechanism, to learn question-adaptive relation representations. Two types of visual object relations are explored: (i) Explicit Relations that represent geometric positions and semantic interactions between objects; and (ii) Implicit Relations that capture the hidden dynamics between image regions. Experiments demonstrate that ReGAT outperforms prior state-of-the-art approaches on both VQA 2.0 and VQA-CP v2 datasets. We further show that ReGAT is compatible to existing VQA architectures, and can be used as a generic relation encoder to boost the model performance for VQA. Zhe Gan, Yu Cheng 0001, Jingjing Liu 0001 |
ICCV | 3 |
| 2019 | Mixed-Supervised Dual-Network for Medical Image Segmentation
Nir Ben-Shlomo, C. Eduardo Corrales, Yu Cheng 0001, Tao Zhang 0006, Jayender Jagadeesan |
MICCAI (2) | 5 |
| 2019 | A hybrid approach with optimization-based and metric-based meta-learner for few-shot learning
Yu Cheng 0001, Mo Yu, Tao Zhang 0006 |
Neurocomputing | 2 |
| 2019 | Attend to count: Crowd counting with adaptive capacity multi-scale CNNs
Zhikang Zou, Yu Cheng 0001, Xiaoye Qu, Shouling Ji, Pan Zhou 0001 |
Neurocomputing | 2 |
| 2018 | Diverse Few-Shot Text Classification with Multiple MetricsabstractMo Yu, Xiaoxiao Guo, Jinfeng Yi, Shiyu Chang, Saloni Potdar, Yu Cheng, Gerald Tesauro, Haoyu Wang, Bowen Zhou. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Mo Yu, Jinfeng Yi, Shiyu Chang, Saloni Potdar, Yu Cheng 0001, Gerald Tesauro, Haoyu Wang 0002 |
NAACL-HLT | 6 |
| 2018 | Dialog-based Interactive Image RetrievalabstractExisting methods for interactive image retrieval have demonstrated the merit of integrating user feedback, improving retrieval results. However, most current systems rely on restricted forms of user feedback, such as binary relevance responses, or feedback based on a fixed set of relative attributes, which limits their impact. In this paper, we introduce a new approach to interactive image search that enables users to provide feedback via natural language, allowing for more natural and effective interaction. We formulate the task of dialog-based interactive image retrieval as a reinforcement learning problem, and reward the dialog system for improving the rank of the target image during each dialog turn. To mitigate the cumbersome and costly process of collecting human-machine conversations as the dialog system learns, we train our system with a user simulator, which is itself trained to describe the differences between target and candidate images. The efficacy of our approach is demonstrated in a footwear retrieval application. Experiments on both simulated and real-world data show that 1) our proposed learning framework achieves better accuracy than other supervised and reinforcement learning baselines and 2) user feedback based on natural language rather than pre-specified attributes leads to more effective retrieval results, and a more natural and expressive communication interface. Hui Wu 0009, Yu Cheng 0001, Steven Rennie, Gerald Tesauro, Rogério Feris |
NeurIPS | 3 |
| 2017 | Fully-Adaptive Feature Sharing in Multi-Task Networks with Applications in Person Attribute ClassificationabstractMulti-task learning aims to improve generalization performance of multiple prediction tasks by appropriately sharing relevant information across them. In the context of deep neural networks, this idea is often realized by hand-designed network architectures with layers that are shared across tasks and branches that encode task-specific features. However, the space of possible multi-task deep architectures is combinatorially large and often the final architecture is arrived at by manual exploration of this space, which can be both error-prone and tedious. We propose an automatic approach for designing compact multi-task deep learning architectures. Our approach starts with a thin multi-layer network and dynamically widens it in a greedy manner during training. By doing so iteratively, it creates a tree-like deep architecture, on which similar tasks reside in the same branch until at the top layers. Evaluation on person attributes classification tasks involving facial and clothing attributes suggests that the models produced by the proposed method are fast, compact and can closely match or exceed the state-of-the-art accuracy from strong baselines by much more expensive models. Yongxi Lu, Abhishek Kumar 0001, Shuangfei Zhai, Yu Cheng 0001, Tara Javidi, Rogério Feris |
CVPR | 4 |
| 2017 | S3Pool: Pooling with Stochastic Spatial Sampling
Shuangfei Zhai, Hui Wu 0009, Abhishek Kumar 0001, Yu Cheng 0001, Yongxi Lu, Zhongfei Zhang, Rogério Feris |
CVPR | 4 |
| 2017 | Jointly Attentive Spatial-Temporal Pooling Networks for Video-Based Person Re-identificationabstractPerson Re-Identification (person re-id) is a crucial task as its applications in visual surveillance and human-computer interaction. In this work, we present a novel joint Spatial and Temporal Attention Pooling Network (ASTPN) for video-based person re-identification, which enables the feature extractor to be aware of the current input video sequences, in a way that interdependency from the matching items can directly influence the computation of each other's representation. Specifically, the spatial pooling layer is able to select regions from each frame, while the attention temporal pooling performed can select informative frames over the sequence, both pooling guided by the information from distance matching. Experiments are conduced on the iLIDS-VID, PRID-2011 and MARS datasets and the results demonstrate that this approach outperforms existing state-of-art methods. We also analyze how the joint pooling in both dimensions can boost the person re-id performance more effectively than using either of them separately 1. Shuangjie Xu, Yu Cheng 0001, Kang Gu, Yang Yang 0002, Shiyu Chang, Pan Zhou 0001 |
ICCV | 2 |
| 2017 | Boosting Deep Learning Risk Prediction with Generative Adversarial Networks for Electronic Health RecordsabstractThe rapid growth of Electronic Health Records (EHRs), as well as the accompanied opportunities in Data-Driven Healthcare (DDH), has been attracting widespread interests and attentions. Recent progress in the design and applications of deep learning methods has shown promising results and is forcing massive changes in healthcare academia and industry, but most of these methods rely on massive labeled data. In this work, we propose a general deep learning framework which is able to boost risk prediction performance with limited EHR data. Our model takes a modified generative adversarial network namely ehrGAN, which can provide plausible labeled EHR data by mimicking real patient records, to augment the training dataset in a semi-supervised learning manner. We use this generative model together with a convolutional neural network (CNN) based prediction model to improve the onset prediction performance. Experiments on two real healthcare datasets demonstrate that our proposed framework produces realistic data samples and achieves significant improvements on classification tasks with the generated data over several stat-of-the-art baselines. Zhengping Che, Yu Cheng 0001, Shuangfei Zhai, Zhaonan Sun, Yan Liu 0002 |
ICDM | 2 |
| 2017 | MMD GAN: Towards Deeper Understanding of Moment Matching NetworkabstractGenerative moment matching network (GMMN) is a deep generative model that differs from Generative Adversarial Network (GAN) by replacing the discriminator in GAN with a two-sample test based on kernel maximum mean discrepancy (MMD). Although some theoretical guarantees of MMD have been studied, the empirical performance of GMMN is still not as competitive as that of GAN on challenging and large benchmark datasets. The computational efficiency of GMMN is also less desirable in comparison with GAN, partially due to its requirement for a rather large batch size during the training. In this paper, we propose to improve both the model expressiveness of GMMN and its computational efficiency by introducing {\it adversarial kernel learning} techniques, as the replacement of a fixed Gaussian kernel in the original GMMN. The new approach combines the key ideas in both GMMN and GAN, hence we name it MMD-GAN. The new distance measure in MMD-GAN is a meaningful loss that enjoys the advantage of weak$^*$ topology and can be optimized via gradient descent with relatively small batch sizes. In our evaluation on multiple benchmark datasets, including MNIST, CIFAR-10, CelebA and LSUN, the performance of MMD-GAN significantly outperforms GMMN, and is competitive with other representative GAN works. Chun-Liang Li, Wei-Cheng Chang, Yu Cheng 0001, Yiming Yang 0002, Barnabás Póczos |
NIPS | 3 |
| 2017 | CHI: A contemporaneous health index for degenerative disease monitoring using longitudinal measurements
Yijun Huang, Heather L. Evans, William B. Lober, Yu Cheng 0001, Xiaoning Qian, Ji Liu 0002, Shuai Huang 0001 |
J. Biomed. Informatics | 5 |
| 2017 | Robust extrinsic calibration from pedestrians
Yu Cheng 0001, Tao Zhang 0006 |
Signal Process. Image Commun. | 2 |
| 2017 | Unsupervised Sequential Outlier Detection With Deep ArchitecturesabstractUnsupervised outlier detection is a vital task and has high impact on a wide variety of applications domains, such as image analysis and video surveillance. It also gains long-standing attentions and has been extensively studied in multiple research areas. Detecting and taking action on outliers as quickly as possible are imperative in order to protect network and related stakeholders or to maintain the reliability of critical systems. However, outlier detection is difficult due to the one class nature and challenges in feature construction. Sequential anomaly detection is even harder with more challenges from temporal correlation in data, as well as the presence of noise and high dimensionality. In this paper, we introduce a novel deep structured framework to solve the challenging sequential outlier detection problem. We use autoencoder models to capture the intrinsic difference between outliers and normal instances and integrate the models to recurrent neural networks that allow the learning to make use of previous context as well as make the learners more robust to warp along the time axis. Furthermore, we propose to use a layerwise training procedure, which significantly simplifies the training procedure and hence helps achieve efficient and scalable training. In addition, we investigate a fine-tuning step to update all parameters set by incorporating the temporal correlation in the sequence. We further apply our proposed models to conduct systematic experiments on five real-world benchmark data sets. Experimental results demonstrate the effectiveness of our model, compared with other state-of-the-art approaches. Weining Lu, Yu Cheng 0001, Cao Xiao, Shiyu Chang, Shuai Huang 0001, Bin Liang 0001, Thomas S. Huang |
IEEE Trans. Image Process. | 2 |
| 2016 | Walk and Learn: Facial Attribute Representation Learning from Egocentric Video and Contextual DataabstractThe way people look in terms of facial attributes (ethnicity, hair color, facial hair, etc.) and the clothes or accessories they wear (sunglasses, hat, hoodies, etc.) is highly dependent on geo-location and weather condition, respectively. This work explores, for the first time, the use of this contextual information, as people with wearable cameras walk across different neighborhoods of a city, in order to learn a rich feature representation for facial attribute classification, without the costly manual annotation required by previous methods. By tracking the faces of casual walkers on more than 40 hours of egocentric video, we are able to cover tens of thousands of different identities and automatically extract nearly 5 million pairs of images connected by or from different face tracks, along with their weather and location context, under pose and lighting variations. These image pairs are then fed into a deep network that preserves similarity of images connected by the same track, in order to capture identity-related attribute features, and optimizes for location and weather prediction to capture additional facial attribute features. Finally, the network is fine-tuned with manually annotated samples. We perform an extensive experimental analysis on wearable data and two standard benchmark datasets based on web images (LFWA and CelebA). Our method outperforms by a large margin a network trained from scratch. Moreover, even without using manually annotated identity labels for pre-training as in previous methods, our approach achieves results that are better than the state of the art. Jing Wang 0069, Yu Cheng 0001, Rogério Feris |
CVPR | 2 |
| 2016 | Outlier faces detector via efficient cohesive subgraph identificationabstractA personal or enterprise collection of a large set of face images may contain many types of tags used for querying the collection. Often the tags have many irrelevant content that may not reflect the image content in terms of the facial characteristics. In this paper, we propose a data curation method to filter out the irrelevant face images using a face recognition based subgraph identification. Results on retrievals from the Internet using popular celebrities show the efficacy of our approach after we cleanse the images collection retrieved and applying our algorithm to the collection. Yu Cheng 0001, Nalini K. Ratha, Sharath Pankanti |
ICIP | 1 |
| 2016 | Deep Structured Energy Based Models for Anomaly DetectionabstractIn this paper, we attack the anomaly detection problem by directly modeling the data distribution with deep architectures. We hence propose deep structured energy based models (DSEBMs), where the energy function is the output of a deterministic deep neural network with structure. We develop novel model architectures to integrate EBMs with different types of data such as static data, sequential data, and spatial data, and apply appropriate model architectures to adapt to the data structure. Our training algorithm is built upon the recent development of score matching (Hyvarinen, 2005), which connects an EBM with a regularized autoencoder, eliminating the need for complicated sampling method. Statistically sound decision criterion can be derived for anomaly detection purpose from the perspective of the energy landscape of the data distribution. We investigate two decision criteria for performing anomaly detection: the energy score and the reconstruction error. Extensive empirical studies on benchmark anomaly detection tasks demonstrate that our proposed model consistently matches or outperforms all the competing methods. Shuangfei Zhai, Yu Cheng 0001, Weining Lu, Zhongfei Zhang |
ICML | 2 |
| 2016 | Doubly Convolutional Neural NetworksabstractBuilding large models with parameter sharing accounts for most of the success of deep convolutional neural networks (CNNs). In this paper, we propose doubly convolutional neural networks (DCNNs), which significantly improve the performance of CNNs by further exploring this idea. In stead of allocating a set of convolutional filters that are independently learned, a DCNN maintains groups of filters where filters within each group are translated versions of each other. Practically, a DCNN can be easily implemented by a two-step convolution procedure, which is supported by most modern deep learning libraries. We perform extensive experiments on three image classification benchmarks: CIFAR-10, CIFAR-100 and ImageNet, and show that DCNNs consistently outperform other competing architectures. We have also verified that replacing a convolutional layer with a doubly convolutional layer at any depth of a CNN can improve its performance. Moreover, various design choices of DCNNs are demonstrated, which shows that DCNN can serve the dual purpose of building more accurate models and/or reducing the memory footprint without sacrificing the accuracy. Shuangfei Zhai, Yu Cheng 0001, Zhongfei Zhang, Weining Lu |
NIPS | 2 |
| 2015 | An Exploration of Parameter Redundancy in Deep Networks with Circulant ProjectionsabstractWe explore the redundancy of parameters in deep neural networks by replacing the conventional linear projection in fully-connected layers with the circulant projection. The circulant structure substantially reduces memory footprint and enables the use of the Fast Fourier Transform to speed up the computation. Considering a fully-connected neural network layer with d input nodes, and d output nodes, this method improves the time complexity from O(d2) to O(dlogd) and space complexity from O(d2) to O(d). The space savings are particularly important for modern deep convolutional neural network architectures, where fully-connected layers typically contain more than 90% of the network parameters. We further show that the gradient computation and optimization of the circulant projections can be performed very efficiently. Our experiments on three standard datasets show that the proposed approach achieves this significant gain in storage and efficiency with minimal increase in error rate compared to neural networks with unstructured projections. Yu Cheng 0001, Felix X. Yu, Rogério Feris, Sanjiv Kumar, Alok N. Choudhary, Shih-Fu Chang |
ICCV | 1 |
| 2015 | Legislative Prediction with Dual Uncertainty Minimization from Heterogeneous InformationabstractVoting on legislative bills to form new laws serves as a key function of most legislature. Predicting the votes of such deliberative bodies leads to better understanding of government policies and generates actionable strategies for social good. In this paper, we present a novel prediction model that maximizes the usage of publicly accessible heterogeneous data, i.e., bill text and lawmakers' profile data, to carry out effective legislative prediction. In particular, we propose to design a probabilistic prediction model which achieves high consistency with past vote records while ensuring the minimum uncertainty of the vote prediction reflecting the firm legal ground often held by the lawmakers. In addition, the proposed legislative prediction model enjoys the following properties: inductive and analytical solution, abilities to deal with the prediction on new bills and new legislators, and robustness to the missing vote issue. We conduct extensive empirical study using the real legislative data and compare with other representative methods in both quantitative political science and data mining communities. The experimental results clearly corroborate that the proposed method provides superior prediction accuracy with visible performance gain. Yu Cheng 0001, Ankit Agrawal 0001, Huan Liu 0001, Alok N. Choudhary |
SDM | 1 |
| 2014 | Temporal Sequence Modeling for Video Event DetectionabstractWe present a novel approach for event detection in video by temporal sequence modeling. Exploiting temporal information has lain at the core of many approaches for video analysis (i.e., action, activity and event recognition). Unlike previous works doing temporal modeling at semantic event level, we propose to model temporal dependencies in the data at sub-event level without using event annotations. This frees our model from ground truth and addresses several limitations in previous work on temporal modeling. Based on this idea, we represent a video by a sequence of visual words learnt from the video, and apply the Sequence Memoizer [21] to capture long-range dependencies in a temporal context in the visual sequence. This data-driven temporal model is further integrated with event classification for jointly performing segmentation and classification of events in a video. We demonstrate the efficacy of our approach on two challenging datasets for visual recognition. Yu Cheng 0001, Quanfu Fan, Sharath Pankanti, Alok N. Choudhary |
CVPR | 1 |
| 2014 | Social Role Identification via Dual Uncertainty Minimization RegularizationabstractIn this paper, we study a challenging problem of inferring individuals' role and statuses in a professional social network, which is of central importance in workforce optimization and human capital management. Realizing the natural setting of social nodes associated with dual view information, i.e., The local node characteristics and the global network influence, we present a novel model that explores graph regularization techniques and integrates such information to achieve improved prediction performance. In particular, our prediction model is built upon the graph transductive learning framework that encodes an uncertainty regularization term in the conventional empirical risk minimization principle. Through taking advantage of the information from both the local profile and the global network characteristics, the final inference of the role or statues achieves minimum an empirical loss on the labeled set, as well as a minimum uncertainty on the unlabeled social nodes. We perform extensive empirical study using real-world data and compare with representative peer approaches. The experimental results on three real social network data sets show that the proposed model greatly outperforms a number of baseline models and is able to effectively infer in a wide range of scenarios. Yu Cheng 0001, Ankit Agrawal 0001, Alok N. Choudhary, Huan Liu 0001, Tao Zhang 0006 |
ICDM | 1 |
| 2014 | RiskWheel: Interactive visual analytics for surveillance event detectionabstractDetecting human behaviors in vast amounts of video is a challenging task in a variety of real-world applications. Thus an interactive tool designed to support this task with human in the loop is of significance in various domains including public safety and security. In this paper, we design and develop an interactive visual analytics system, RiskWheel, that enables effective analysis of detection results and utilization of user feedback to improve surveillance event detection. In particular, we propose 1) an interactive approach to visualize data with temporal relations and 2) a novel risk ranking method to differentiate detection results and present more informative ones to the user for better interaction. In our experiments, we demonstrate RiskWheel through a case study on TRECVID Surveillance Event Detection (SED) task [1]. The experimental results quantitatively show that RiskWheel outperforms multiple baselines, demonstrating the power of the risk ranking technique. Yu Cheng 0001, Lisa M. Brown, Qibin Fan, Rogério Feris, Sharath Pankanti, Tao Zhang 0006 |
ICME | 1 |
| 2014 | Batch Mode Active Learning with Hierarchical-Structured Embedded VarianceabstractWe consider the problem of active learning when the categories are represented as a tree with leaf nodes as outputs and internal nodes as clusters of the outputs at multiple granularity. Recent work has improved the traditional techniques by moving beyond “flat” structure through incorporation of the label hierarchy into the uncertainty measure. However, these methods have two major limitations when used. First, these methods roughly use the information in the label structure but do not take into account the training samples, which may lead to a sampling bias due to their crude approximation of the class relations. Second, none of these methods can work in a batch mode to reduce the computational time of training. We propose a batch mode active learning scheme that exploits both the hierarchical structure of the labels and the characteristics of the training data to select the most informative data for human labeling. We achieve this goal by first using an approach based on graph embedding that embeds the relationships between the labels and data points in a transformed low-dimensional space. Then, we compute uncertainty by calculating the variance among the points and the labels in the embedding space. Finally, the selection criterion is designed to construct batches and incorporate a diversity measure. Experimental results indicate that our technique achieves a notable improvement in performance over the state-of-the-art approaches. Yu Cheng 0001, Zhengzhang Chen, Hongliang Fei, Alok N. Choudhary |
SDM | 1 |
| 2013 | A probabilistic graphical model for brand reputation assessment in social networksabstractSocial media has become a popular platform that connects people who share information, in particular personal opinions. Through such a fast information exchange mechanism, reputation of individuals, consumer products, or business companies can be quickly built up within a social network. Recently, applications mining social network data start emerging to find the communities sharing the same interests for marketing purposes. Knowing the reputation of social network entities, such as celebrities or business companies, can help develop better strategies for election campaigns or new product advertisements. In this paper, we propose a probabilistic graphical model to collectively measure reputations of entities in social networks. By collecting and analyzing large amount of user activities on Facebook, our model can effectively and efficiently rank entities, such as presidential candidates, professional sport teams, musician bands, and companies, based on their social reputation. The proposed model produces results largely consistent with the two publicly available systems - movie ranking in Internet Movie Database and business school ranking by the US news & World Report - with the correlation coefficients of 0.75 and -0.71, respectively. Kunpeng Zhang 0001, Doug Downey, Zhengzhang Chen, Yusheng Xie, Yu Cheng 0001, Ankit Agrawal 0001, Wei-keng Liao, Alok N. Choudhary |
ASONAM | 5 |
| 2013 | Elver: Recommending Facebook pages in cold start situation without content featuresabstractRecommender systems are vital to the success of online retailers and content providers. One particular challenge in recommender systems is the “cold start” problem. The word “cold” refers to the items that are not yet rated by any user or the users who have not yet rated any items. We propose Elver to recommend and optimize page-interest targeting on Facebook. Existing techniques for cold recommendation mostly rely on content features in the event of lacking user ratings. Since it is very hard to construct universally meaningful features for the millions of Facebook pages, Elver makes minimal assumption of content features. Elver employs iterative matrix completion technology and nonnegative factorization procedure to work with meagre content inklings. Experiments on Facebook data shows the effectiveness of Elver at different levels of sparsity. Yusheng Xie, Zhengzhang Chen, Kunpeng Zhang 0001, Yu Cheng 0001, Ankit Agrawal 0001, Alok N. Choudhary |
IEEE BigData | 5 |
| 2013 | Feedback-driven multiclass active learning for data streamsabstractActive learning is a promising way to efficiently build up training sets with minimal supervision. Most existing methods consider the learning problem in a pool-based setting. However, in a lot of real-world learning tasks, such as crowdsourcing, the unlabeled samples, arrive sequentially in the form of continuous rapid streams. Thus, preparing a pool of unlabeled data for active learning is impractical. Moreover, performing exhaustive search in a data pool is expensive, and therefore unsuitable for supporting on-the-fly interactive learning in large scale data. In this paper, we present a systematic framework for stream-based multi-class active learning. Following the reinforcement learning framework, we propose a feedback-driven active learning approach by adaptively combining different criteria in a time-varying manner. Our method is able to balance exploration and exploitation during the learning process. Extensive evaluation on various benchmark and real-world datasets demonstrates the superiority of our framework over existing methods. Yu Cheng 0001, Zhengzhang Chen, Lu Liu 0005, Ankit Agrawal 0001, Alok N. Choudhary |
CIKM | 1 |
| 2013 | Bootstrapping active name disambiguation with crowdsourcingabstractName disambiguation is a challenging and important problem in many domains, such as digital libraries, social media management and people search systems. Traditional methods, based on direct assignment using supervised machine learning techniques, seem to be the most effective, but their performances are highly dependent on the amount of training data, while large data annotation can be expensive and time-consuming requiring hours of manual inspection by a domain expert. To efficiently acquire labeled data, we propose a bootstrapping algorithm for the name disambiguation task based on active learning and crowdsourced labeling. We show that the proposed method can leverage the advantages of exploration and exploitation by combining two strategies, thereby improving the overall quality of the training data at minimal expense. The experimental results on two datasets DBLP and ArnetMiner demonstrate the superiority of our framework over existing methods. Yu Cheng 0001, Zhengzhang Chen, Ankit Agrawal 0001, Alok N. Choudhary |
CIKM | 1 |
| 2013 | Mining diabetes complication and treatment patterns for clinical decision supportabstractThe fast development of hospital information systems (HIS) produces a large volume of electronic medical records, which provides a comprehensive source for exploratory analysis and statistics to support clinical decision-making. In this paper, we investigate how to utilize the heterogeneous medical records to aid the clinical treatments of diabetes mellitus. Diabetes mellitus, simply diabetes, is a group of metabolic diseases, which is often accompanied with many complications. We propose a Symptom-Diagnosis-Treatment model to mine the diabetes complication patterns and to unveil the latent association mechanism between treatments and symptoms from large volume of electronic medical records. Furthermore, we study the demographic statistics of patient population w.r.t. complication patterns in real data and observe several interesting phenomena. The discovered complication and treatment patterns can help physicians better understand their specialty and learn previous experiences. Our experiments on a collection of one-year diabetes clinical records from a famous geriatric hospital demonstrate the effectiveness of our approaches. Lu Liu 0005, Jie Tang 0001, Yu Cheng 0001, Ankit Agrawal 0001, Wei-keng Liao, Alok N. Choudhary |
CIKM | 3 |
| 2013 | Forecast Oriented Classification of Spatio-Temporal Extreme Events
Zhengzhang Chen, Yusheng Xie, Yu Cheng 0001, Kunpeng Zhang 0001, Ankit Agrawal 0001, Wei-keng Liao, Nagiza F. Samatova, Alok N. Choudhary |
IJCAI | 3 |
| 2013 | JobMiner: a real-time system for mining job-related patterns from social mediaabstractThe various kinds of booming social media not only provide a platform where people can communicate with each other, but also spread useful domain information, such as career and job market information. For example, LinkedIn publishes a large amount of messages either about people who want to seek jobs or companies who want to recruit new members. By collecting information, we can have a better understanding of the job market and provide insights to job-seekers, companies and even decision makers. In this paper, we analyze the job information from the social network point of view. We first collect the job-related information from various social media sources. Then we construct an inter-company job-hopping network, with the vertices denoting companies and the edges denoting flow of personnel between companies. We subsequently employ graphmining techniques to mine influential companies and related company groups based on the job-hopping network model. Demonstration on LinkedIn data shows that our system JobMiner can provide a better understanding of the dynamic processes and a more accurate identification of important entities in the job market. Yu Cheng 0001, Yusheng Xie, Zhengzhang Chen, Ankit Agrawal 0001, Alok N. Choudhary, Songtao Guo |
KDD | 1 |
| 2013 | Graphical Modeling of Macro Behavioral Targeting in Social NetworksabstractWe investigate a class of emerging online marketing challenges in social networks; macro behavioral targeting (MBT) is introduced as non-personalized broadcasting efforts to massive populations. We propose a new probabilistic graphical model for MBT. Further, a linear-time approximation method is proposed to circumvent an intractable parametric representation of user behaviors. We compare the proposed model with the existing state-of-the-art method on real datasets from social networks. Our model outperforms in all categories by comfortable margins. Ankit Agrawal 0001, Zhengzhang Chen, Yu Cheng 0001, Alok N. Choudhary, Md. Mostofa Ali Patwary, Yusheng Xie, Kunpeng Zhang 0001 |
SDM | 3 |
| 2012 | On active learning in hierarchical classificationabstractMost of the existing active learning algorithms assume all the category labels as independent or consider them in a "flat" structure. However, in reality, there are many applications in which the set of possible labels are often organized in a hierarchical structure. In this paper, we consider the problem of active learning when the categories are represented as a tree. Our goal is to exploit the structure information of the label tree in active learning to select the most informative samples to be labeled. We propose an algorithm that estimates the semantic space, embedding the category hierarchy. In this space, each category label is represented as a prototype and the uncertainty is measured using a variance-based fashion. We also demonstrate notable performance improvement with the proposed approach on synthetic and real datasets. Yu Cheng 0001, Kunpeng Zhang 0001, Yusheng Xie, Ankit Agrawal 0001, Alok N. Choudhary |
CIKM | 1 |
| 2012 | VOXSUP: a social engagement frameworkabstractSocial media websites are currently central hubs on the Internet. Major online social media platforms are not only places for individual users to socialize but are increasingly more important as channels for companies to advertise, public figures to engage, etc. In order to optimize such advertising and engaging efforts, there is an emerging challenge for knowledge discovery on today's Internet. The goal of knowledge discovery is to understand the entire online social landscape instead of merely summarizing the statistics. To answer this challenge, we have created VOXSUP as a unified social engagement framework. Unlike most existing tools, VOXSUP not only aggregates and filters social data from the Internet, but also provides what we call Voxsupian Knowledge Discovery (VKD). VKD consists of an almost human-level understanding of social conversations at any level of granularity from a single comment sentiment to multi-lingual inter-platform user demographics. Here we describe the technologies that are crucial to VKD, and subsequently go beyond experimental verification and present case studies from our live VOXSUP system. Yusheng Xie, Daniel Honbo, Alok N. Choudhary, Kunpeng Zhang 0001, Yu Cheng 0001, Ankit Agrawal 0001 |
KDD | 5 |
| 2012 | Sentiment identification by incorporating syntax, semantics and context informationabstractThis paper proposes a method based on conditional random fields to incorporate sentence structure (syntax and semantics) and context information to identify sentiments of sentences within a document. It also proposes and evaluates two different active learning strategies for labeling sentiment data. The experiments with the proposed approach demonstrate a 5-15% improvement in accuracy on Amazon customer reviews compared to existing supervised learning and rule-based methods. Kunpeng Zhang 0001, Yusheng Xie, Yu Cheng 0001, Daniel Honbo, Doug Downey, Ankit Agrawal 0001, Wei-keng Liao, Alok N. Choudhary |
SIGIR | 3 |
| 2011 | Mining millions of reviews: a technique to rank products based on importance of reviewsabstractAs online shopping becomes increasingly more popular, many shopping web sites encourage existing customers to add reviews of products purchased. These reviews make an impact on the purchasing decisions of potential customers. At Amazon.com for instance, some products receive hundreds of reviews. It is overwhelming and time restrictive for most customers to read, comprehend and make decisions based on all of these reviews. Customers most likely end up reading only a small fraction of the reviews usually in the order which they are presented on the product page. Incorporating various product review factors, such as: content related to product quality, time of the review, content related to product durability and historically older positive customer reviews will have different impacts on the products rankings. Thus, the automated mining of product reviews and opinions to produce a re-calculated product ranking score is a valuable tool which would allow potential customers to make more informed decisions. In this paper, we present a product ranking model that applies weights to product review factors to calculate a products ranking score. Our experiments use the customer reviews from Amazon.com as input to our product ranking model which produces product ranking results that closely relate to the products sales ranking as reported by the retailer. Kunpeng Zhang 0001, Yu Cheng 0001, Wei-keng Liao, Alok N. Choudhary |
ICEC | 2 |