VLDB 2026 Research / reviewers in the wild / expert
Zheng Zhang 0001
dblp:z/ZhengZhang
· DBLP profile ↗
101ranked-venue papers
5as first author
43since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 49 · 39 since 2021Systems, architecture and hardware · 25 · 3 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 7 since 2021Software engineering, systems software and programming languages · 11 · 1 first-authorDatabases, data management, data science and information retrieval · 10 · 3 since 2021Computer networks · 5Security and privacy · 4 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AuthGuard: Generalizable Deepfake Detection via Language GuidanceabstractExisting deepfake detection techniques struggle to keep-up with the ever-evolving novel, unseen forgeries methods. This limitation stems from their reliance on statistical artifacts learned during training, which are often tied to specific generation processes that may not be representative of samples from new, unseen deepfake generation methods encountered at test time. We propose that incorporating language guidance can improve deepfake detection generalization by integrating human-like commonsense reasoning – such as recognizing logical inconsistencies and perceptual anomalies – alongside statistical cues. To achieve this, we train an expert deepfake vision encoder by combining discriminative classification with image-text contrastive learning, where the text is generated by generalist MLLMs using few-shot prompting. This allows the encoder to extract both language-describable, commonsense deepfake artifacts and statistical forgery artifacts from pixel-level distributions. To further enhance robustness, we integrate data uncertainty learning into vision-language contrastive learning, mitigating noise in image-text supervision. Our expert vision encoder seamlessly interfaces with an LLM, further enabling more generalized and interpretable deepfake detection while also boosting accuracy. The resulting framework, AuthGuard, achieves state-of-the-art deepfake detection accuracy in both in-distribution and out-of-distribution settings, achieving AUC gains of 6.15% on the DFDC dataset and 16.68% on the DF40 dataset. Additionally, AuthGuard significantly enhances deepfake reasoning, improving performance by 24.69% on the DDVQA dataset. Guangyu Shen, Tianchen Zhao, Zheng Zhang 0001, Dongsheng An, Zhuowen Tu, Yifan Xing |
WACV | 5 |
| 2025 | PerSphere: A Comprehensive Framework for Multi-Faceted Perspective Retrieval and SummarizationabstractAs online platforms and recommendation algorithms evolve, people are increasingly trapped in echo chambers, leading to biased understandings of various issues. To combat this issue, we have introduced PerSphere, a benchmark designed to facilitate multi-faceted perspective retrieval and summarization, thus breaking free from these information silos. For each query within PerSphere, there are two opposing claims, each supported by distinct, non-overlapping perspectives drawn from one or more documents. Our goal is to accurately summarize these documents, aligning the summaries with the respective claims and their underlying perspectives. This task is structured as a two-step end-to-end pipeline that includes comprehensive document retrieval and multi-faceted summarization. Furthermore, we propose a set of metrics to evaluate the comprehensiveness of the retrieval and summarization content. Experimental results on various counterparts for the pipeline show that recent models struggle with such a complex task. Analysis shows that the main challenge lies in long context and perspective extraction, and we propose a simple but effective multi-agent summarization system, offering a promising solution to enhance performance on PerSphere. Yingjie Li 0008, Xiangkun Hu, Qinglin Qi, Qipeng Guo, Zheng Zhang 0001, Yue Zhang 0004 |
ACL (1) | 7 |
| 2025 | Bridging Information Asymmetry in Text-video Retrieval: A Data-centric ApproachabstractAs online video content rapidly grows, the task of text-video retrieval (TVR) becomes increasingly important. A key challenge in TVR is the information asymmetry between video and text: videos are inherently richer in information, while their textual descriptions often capture only fragments of this complexity. This paper introduces a novel, data-centric framework to bridge this gap by enriching textual representations to better match the richness of video content. During training, videos are segmented into event-level clips and captioned to ensure comprehensive coverage. During retrieval, a large language model (LLM) generates semantically diverse queries to capture a broader range of possible matches. To enhance retrieval efficiency, we propose a query selection mechanism that identifies the most relevant and diverse queries, reducing computational cost while improving accuracy. Our method achieves state-of-the-art results across multiple benchmarks, demonstrating the power of data-centric approaches in addressing information asymmetry in TVR. This work paves the way for new research focused on leveraging data to improve cross-modal retrieval. Zechen Bai, Tianjun Xiao, Tong He 0002, Pichao Wang, Zheng Zhang 0001, Thomas Brox, Zheng Shou 0001 |
ICLR | 5 |
| 2025 | NovelQA: Benchmarking Question Answering on Documents Exceeding 200K TokensabstractRecent advancements in Large Language Models (LLMs) have pushed the boundaries of natural language processing, especially in long-context understanding. However, the evaluation of these models' long-context abilities remains a challenge due to the limitations of current benchmarks. To address this gap, we introduce NovelQA, a benchmark tailored for evaluating LLMs with complex, extended narratives. NovelQA, constructed from English novels, offers a unique blend of complexity, length, and narrative coherence, making it an ideal tool for assessing deep textual understanding in LLMs. This paper details the design and construction of NovelQA, focusing on its comprehensive manual annotation process and the variety of question types aimed at evaluating nuanced comprehension. Our evaluation of long-context LLMs on NovelQA reveals significant insights into their strengths and weaknesses. Notably, the models struggle with multi-hop reasoning, detail-oriented questions, and handling extremely long inputs, averaging over 200,000 tokens. Results highlight the need for substantial advancements in LLMs to enhance their long-context comprehension and contribute effectively to computational literary analysis. Cunxiang Wang, Ruoxi Ning, Boqi Pan, Tonghui Wu, Qipeng Guo, Cheng Deng 0001, Guangsheng Bao, Xiangkun Hu, Zheng Zhang 0001, Yue Zhang 0004 |
ICLR | 9 |
| 2025 | VideoSAM: Open-World Video SegmentationabstractVideo segmentation is essential for advancing robotics and autonomous driving, particularly in open-world settings where continuous perception and object association across video frames are critical. While the Segment Anything Model (SAM) has excelled in static image segmentation, extending its capabilities to video segmentation poses significant challenges. We tackle two major hurdles: a) SAM's embedding limitations in associating objects across frames, and b) granularity inconsistencies in object segmentation. To this end, we introduce VideoSAM, an end-to-end framework designed to address these challenges by improving object tracking and segmentation consistency in dynamic environments. VideoSAM integrates an agglomerated backbone, RADIO, enabling object association through similarity metrics and introduces Cycle-ack-Pairs Propagation with a memory mechanism for stable object tracking. Additionally, we incorporate an autoregressive object-token mechanism within the SAM decoder to maintain consistent granularity across frames. Our method is extensively evaluated on the UVO and BURST benchmarks, and robotic videos from RoboTAP, demonstrating its effectiveness and robustness in real-world scenarios. All codes will be available. Pinxue Guo, Jianxiong Gao, Chongruo Wu, Tong He 0002, Zheng Zhang 0001, Tianjun Xiao |
ICRA | 6 |
| 2025 | ShortListing Model: A Streamlined Simplex Diffusion for Discrete Variable GenerationabstractGenerative modeling of discrete variables is challenging yet crucial for applications in natural language processing and biological sequence design. We introduce the Shortlisting Model (SLM), a novel simplex-based diffusion model inspired by progressive candidate pruning. SLM operates on simplex centroids, reducing generation complexity and enhancing scalability. Additionally, SLM incorporates a flexible implementation of classifier-free guidance, enhancing unconditional generation performance. Extensive experiments on DNA promoter and enhancer design, protein design, character-level and large-vocabulary language modeling demonstrate the competitive performance and strong potential of SLM. Our code can be found at https://github.com/GenSI-THUAIR/SLM. Yuxuan Song 0002, Jingjing Gong, Qiying Yu, Zheng Zhang 0001, Mingxuan Wang, Hao Zhou 0012, Wei-Ying Ma |
NeurIPS | 6 |
| 2025 | DAPO: An Open-Source LLM Reinforcement Learning System at ScaleabstractInference scaling empowers LLMs with unprecedented reasoning ability, with reinforcement learning as the core technique to elicit complex reasoning. However, key technical details of state-of-the-art reasoning LLMs are concealed (such as in OpenAI o1 blog and DeepSeek R1 technical report), thus the community still struggles to reproduce their RL training results. We propose the **D**ecoupled Clip and **D**ynamic s**A**mpling **P**olicy **O**ptimization (**DAPO**) algorithm, and fully open-source a state-of-the-art large-scale RL system that achieves 50 points on AIME 2024 using Qwen2.5-32B base model. Unlike previous works that withhold training details, we introduce four key techniques of our algorithm that make large-scale LLM RL a success. In addition, we open-source our training code, which is built on the verl framework, along with a carefully curated and processed dataset. These components of our open-source system enhance reproducibility and support future research in large-scale LLM RL. Qiying Yu, Zheng Zhang 0001, Ruofei Zhu, Yufeng Yuan, Xiaochen Zuo, Yu Yue, Weinan Dai, Tiantian Fan, Gaohong Liu, Juncai Liu, Lingjun Liu, Xin Liu 0039, Haibin Lin, Bole Ma, Guangming Sheng, Yuxuan Tong, Chi Zhang 0022, Mofan Zhang, Ru Zhang 0006, Wang Zhang 0017, Jiaze Chen, Jiangjie Chen, Hongli Yu, Yuxuan Song 0002, Xiangpeng Wei, Hao Zhou 0012, Wei-Ying Ma, Ya-Qin Zhang, Mingxuan Wang |
NeurIPS | 2 |
| 2024 | EventGround: Narrative Reasoning by Grounding to Eventuality-centric Knowledge GraphsabstractNarrative reasoning relies on the understanding of eventualities in story contexts, which requires a wealth of background world knowledge. To help machines leverage such knowledge, existing solutions can be categorized into two groups. Some focus on implicitly modeling eventuality knowledge by pretraining language models (LMs) with eventuality-aware objectives. However, this approach breaks down knowledge structures and lacks interpretability. Others explicitly collect world knowledge of eventualities into structured eventuality-centric knowledge graphs (KGs). However, existing research on leveraging these knowledge sources for free-texts is limited. In this work, we propose an initial comprehensive framework called EventGround, which aims to tackle the problem of grounding free-texts to eventuality-centric KGs for contextualized narrative reasoning. We identify two critical problems in this direction: the event representation and sparsity problems. We provide simple yet effective parsing and partial information extraction methods to tackle these problems. Experimental results demonstrate that our approach consistently outperforms baseline models when combined with graph neural network (GNN) or large language model (LLM) based graph reasoning models. Our framework, incorporating grounded knowledge, achieves state-of-the-art performance while providing interpretable evidence. Cheng Jiayang, Chunkit Chan, Xin Liu 0039, Yangqiu Song, Zheng Zhang 0001 |
LREC/COLING | 6 |
| 2024 | Adaptive Slot Attention: Object Discovery with Dynamic Slot NumberabstractObject-centric learning (OCL) extracts the representation of objects with slots, offering an exceptional blend of flexibility and interpretability for abstracting low-level perceptual features. A widely adopted method within OCL is slot attention, which utilizes attention mechanisms to iteratively refine slot representations. However, a major draw-back of most object-centric models, including slot attention, is their reliance on predefining the number of slots. This not only necessitates prior knowledge of the dataset but also overlooks the inherent variability in the number of objects present in each instance. To overcome this fundamental limitation, we present a novel complexity-aware object auto-encoder framework. Within this framework, we introduce an adaptive slot attention (AdaSlot) mecha-nism that dynamically determines the optimal number of slots based on the content of the data. This is achieved by proposing a discrete slot sampling module that is responsible for selecting an appropriate number of slots from a candidate list. Furthermore, we introduce a masked slot decoder that suppresses unselected slots during the decoding process. Our framework, tested extensively on object discovery tasks with various datasets, shows performance matching or exceeding top fixed-slot models. Moreover, our analysis substantiates that our method exhibits the capability to dynamically adapt the slot number according to each instance's complexity, offering the potential for further exploration in slot attention research. Project will be available at https://kfan21.github.io/AdaSlot/ Zechen Bai, Tianjun Xiao, Tong He 0002, Max Horn, Yanwei Fu 0001, Francesco Locatello, Zheng Zhang 0001 |
CVPR | 8 |
| 2024 | ECON: On the Detection and Resolution of Evidence ConflictsabstractCheng Jiayang, Chunkit Chan, Qianqian Zhuang, Lin Qiu, Tianhang Zhang, Tengxiao Liu, Yangqiu Song, Yue Zhang, Pengfei Liu, Zheng Zhang. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Cheng Jiayang, Chunkit Chan, Qianqian Zhuang, Tianhang Zhang, Tengxiao Liu, Yangqiu Song, Yue Zhang 0004, Pengfei Liu 0003, Zheng Zhang 0001 |
EMNLP | 10 |
| 2024 | One Token to Seg Them All: Language Instructed Reasoning Segmentation in VideosabstractWe introduce VideoLISA, a video-based multimodal large language model designed to tackle the problem of language-instructed reasoning segmentation in videos. Leveraging the reasoning capabilities and world knowledge of large language models, and augmented by the Segment Anything Model, VideoLISA generates temporally consistent segmentation masks in videos based on language instructions. Existing image-based methods, such as LISA, struggle with video tasks due to the additional temporal dimension, which requires temporal dynamic understanding and consistent segmentation across frames. VideoLISA addresses these challenges by integrating a Sparse Dense Sampling strategy into the video-LLM, which balances temporal context and spatial detail within computational constraints. Additionally, we propose a One-Token-Seg-All approach using a specially designed <TRK> token, enabling the model to segment and track objects across multiple frames. Extensive evaluations on diverse benchmarks, including our newly introduced ReasonVOS benchmark, demonstrate VideoLISA's superior performance in video object segmentation tasks involving complex reasoning, temporal understanding, and object tracking. While optimized for videos, VideoLISA also shows promising generalization to image segmentation, revealing its potential as a unified foundation model for language-instructed object segmentation. Code and model will be available at: https://github.com/showlab/VideoLISA. Zechen Bai, Tong He 0002, Haiyang Mei, Pichao Wang, Ziteng Gao, Joya Chen, Zheng Zhang 0001, Zheng Shou 0001 |
NeurIPS | 8 |
| 2024 | Rethinking The Training And Evaluation of Rich-Context Layout-to-Image GenerationabstractRecent advancements in generative models have significantly enhanced their capacity for image generation, enabling a wide range of applications such as image editing, completion and video editing. A specialized area within generative modeling is layout-to-image (L2I) generation, where predefined layouts of objects guide the generative process. In this study, we introduce a novel regional cross-attention module tailored to enrich layout-to-image generation. This module notably improves the representation of layout regions, particularly in scenarios where existing methods struggle with highly complex and detailed textual descriptions. Moreover, while current open-vocabulary L2I methods are trained in an open-set setting, their evaluations often occur in closed-set environments. To bridge this gap, we propose two metrics to assess L2I performance in open-vocabulary scenarios. Additionally, we conduct a comprehensive user study to validate the consistency of these metrics with human preferences. Jiaxin Cheng, Tong He 0002, Tianjun Xiao, Zheng Zhang 0001, Yicong Zhou |
NeurIPS | 5 |
| 2024 | Unified Lexical Representation for Interpretable Visual-Language AlignmentabstractVisual-Language Alignment (VLA) has gained a lot of attention since CLIP's groundbreaking work.
Although CLIP performs well, the typical direct latent feature alignment lacks clarity in its representation and similarity scores.
On the other hand, lexical representation, a vector whose element represents the similarity between the sample and a word from the vocabulary, is a natural sparse representation and interpretable, providing exact matches for individual words.
However, lexical representations are difficult to learn due to no ground-truth supervision and false-discovery issues, and thus requires complex design to train effectively.
In this paper, we introduce LexVLA, a more interpretable VLA framework by learning a unified lexical representation for both modalities without complex design.
We use DINOv2 as our visual model for its local-inclined features and Llama 2, a generative language model, to leverage its in-context lexical prediction ability.
To avoid the false discovery, we propose an overuse penalty to refrain the lexical representation from falsely frequently activating meaningless words.
We demonstrate that these two pre-trained uni-modal models can be well-aligned by fine-tuning on the modest multi-modal dataset and avoid intricate training configurations.
On cross-modal retrieval benchmarks, LexVLA, trained on the CC-12M multi-modal dataset, outperforms baselines fine-tuned on larger datasets (e.g., YFCC15M) and those trained from scratch on even bigger datasets (e.g., 1.1B data, including CC-12M).
We conduct extensive experiments to analyze LexVLA.
Codes are available at https://github.com/Clementine24/LexVLA. Yikai Wang 0002, Yanwei Fu 0001, Dongyu Ru, Zheng Zhang 0001, Tong He 0002 |
NeurIPS | 5 |
| 2024 | Can Language Models Learn to Skip Steps?abstractTrained on vast corpora of human language, language models demonstrate emergent human-like reasoning abilities. Yet they are still far from true intelligence, which opens up intriguing opportunities to explore the parallels of humans and model behaviors. In this work, we study the ability to skip steps in reasoning—a hallmark of human expertise developed through practice. Unlike humans, who may skip steps to enhance efficiency or to reduce cognitive load, models do not inherently possess such motivations to minimize reasoning steps. To address this, we introduce a controlled framework that stimulates step-skipping behavior by iteratively refining models to generate shorter and accurate reasoning paths. Empirical results indicate that models can develop the step skipping ability under our guidance. Moreover, after fine-tuning on expanded datasets that include both complete and skipped reasoning sequences, the models can not only resolve tasks with increased efficiency without sacrificing accuracy, but also exhibit comparable and even enhanced generalization capabilities in out-of-domain scenarios. Our work presents the first exploration into human-like step-skipping ability and provides fresh perspectives on how such cognitive abilities can benefit AI models. Tengxiao Liu, Qipeng Guo, Xiangkun Hu, Cheng Jiayang, Yue Zhang 0004, Xipeng Qiu, Zheng Zhang 0001 |
NeurIPS | 7 |
| 2024 | RAGChecker: A Fine-grained Framework for Diagnosing Retrieval-Augmented GenerationabstractDespite Retrieval-Augmented Generation (RAG) has shown promising capability in leveraging external knowledge, a comprehensive evaluation of RAG systems is still challenging due to the modular nature of RAG, evaluation of long-form responses and reliability of measurements. In this paper, we propose a fine-grained evaluation framework, RAGChecker, that incorporates a suite of diagnostic metrics for both the retrieval and generation modules. Meta evaluation verifies that RAGChecker has significantly better correlations with human judgments than other evaluation metrics. Using RAGChecker, we evaluate 8 RAG systems and conduct an in-depth analysis of their performance, revealing insightful patterns and trade-offs in the design choices of RAG architectures. The metrics of RAGChecker can guide researchers and practitioners in developing more effective RAG systems. Dongyu Ru, Xiangkun Hu, Tianhang Zhang, Peng Shi 0010, Shuaichen Chang, Cheng Jiayang, Cunxiang Wang, Shichao Sun, Huanyu Li 0010, Binjie Wang, Jiarong Jiang, Tong He 0002, Zhiguo Wang 0006, Pengfei Liu 0003, Yue Zhang 0004, Zheng Zhang 0001 |
NeurIPS | 18 |
| 2024 | 4DBInfer: A 4D Benchmarking Toolbox for Graph-Centric Predictive Modeling on RDBsabstractGiven a relational database (RDB), how can we predict missing column values in some target table of interest? Although RDBs store vast amounts of rich, informative data spread across interconnected tables, the progress of predictive machine learning models as applied to such tasks arguably falls well behind advances in other domains such as computer vision or natural language processing. This deficit stems, at least in part, from the lack of established/public RDB benchmarks as needed for training and evaluation purposes. As a result, related model development thus far often defaults to tabular approaches trained on ubiquitous single-table benchmarks, or on the relational side, graph-based alternatives such as GNNs applied to a completely different set of graph datasets devoid of tabular characteristics. To more precisely target RDBs lying at the nexus of these two complementary regimes, we explore a broad class of baseline models predicated on: (i) converting multi-table datasets into graphs using various strategies equipped with efficient subsampling, while preserving tabular characteristics; and (ii) trainable models with well-matched inductive biases that output predictions based on these input subgraphs. Then, to address the dearth of suitable public benchmarks and reduce siloed comparisons, we assemble a diverse collection of (i) large-scale RDB datasets and (ii) coincident predictive tasks. From a delivery standpoint, we operationalize the above four dimensions (4D) of exploration within a unified, scalable open-source toolbox called 4DBInfer; please see https://github.com/awslabs/multi-table-benchmark . David P. Wipf, Zheng Zhang 0001, Christos Faloutsos, Weinan Zhang 0001, Muhan Zhang, Zhenkun Cai, Jiahang Li 0002, Zunyao Mao, Yakun Song, Yanlin Zhang, Chuan Lei, Xiao Qin 0003, Ning Li 0029, Han Zhang 0057 |
NeurIPS | 4 |
| 2024 | FreshGNN: Reducing Memory Access via Stable Historical Embeddings for Graph Neural Network TrainingabstractA key performance bottleneck when training graph neural network (GNN) models on large, real-world graphs is loading node features onto a GPU. Due to limited GPU memory, expensive data movement is necessary to facilitate the storage of these features on alternative devices with slower access (e.g. CPU memory). Moreover, the irregularity of graph structures contributes to poor data locality which further exacerbates the problem. Consequently, existing frameworks capable of efficiently training large GNN models usually incur a significant accuracy degradation because of the currently-available shortcuts involved. To address these limitations, we instead propose FreshGNN, a general-purpose GNN mini-batch training framework that leverages a historical cache for storing and reusing GNN node embeddings instead of re-computing them through fetching raw features at every iteration. Critical to its success, the corresponding cache policy is designed, using a combination of gradient-based and staleness criteria, to selectively screen those embeddings which are relatively stable and can be cached, from those that need to be re-computed to reduce estimation errors and subsequent downstream accuracy loss. When paired with complementary system enhancements to support this selective historical cache, FreshGNN is able to accelerate the training speed on large graph datasets such as ogbn-papers100M and MAG240M by 3.4× up to 20.5× and reduce the memory access by 59%, with less than 1% influence on test accuracy. Kezhao Huang, Haitian Jiang, Guangxuan Xiao, David P. Wipf, Xiang Song 0003, Zengfeng Huang, Jidong Zhai, Zheng Zhang 0001 |
Proc. VLDB Endow. | 10 |
| 2023 | An AMR-based Link Prediction Approach for Document-level Event Argument ExtractionabstractRecent works have introduced Abstract Meaning Representation (AMR) for Document-level Event Argument Extraction (Doc-level EAE), since AMR provides a useful interpretation of complex semantic structures and helps to capture long-distance dependency.However, in these works AMR is used only implicitly, for instance, as additional features or training signals.Motivated by the fact that all event structures can be inferred from AMR, this work reformulates EAE as a link prediction problem on AMR graphs.Since AMR is a generic structure and does not perfectly suit EAE, we propose a novel graph structure, Tailored AMR Graph (TAG), which compresses less informative subgraphs and edge types, integrates span information, and highlights surrounding events in the same document.With TAG, we further propose a novel method using graph neural networks as a link prediction model to find event arguments.Our extensive experiments on WikiEvents and RAMS show that this simpler approach outperforms the state-of-the-art models by 3.63pt and 2.33pt F1, respectively, and do so with reduced 56% inference time.The code is available at https://github.com/ayyyq/TARA. Yuqing Yang 0004, Qipeng Guo, Xiangkun Hu, Yue Zhang 0004, Xipeng Qiu, Zheng Zhang 0001 |
ACL (1) | 6 |
| 2023 | Dual Cache for Long Document Neural Coreference ResolutionabstractRecent works show the effectiveness of cachebased neural coreference resolution models on long documents.These models incrementally process a long document from left to right and extract relations between mentions and entities in a cache, resulting in much lower memory and computation cost compared to computing all mentions in parallel.However, they do not handle cache misses when high-quality entities are purged from the cache, which causes wrong assignments and leads to prediction errors.We propose a new hybrid cache that integrates two eviction policies to capture global and local entities separately, and effectively reduces the aggregated cache misses up to half as before, while improving F1 score of coreference by 0.7 ∼ 5.7pt.As such, the hybrid policy can accelerate existing cache-based models and offer a new long document coreference resolution solution.Results show that our method outperforms existing methods on four benchmarks while saving up to 83% of inference time against non-cache-based models.Further, we achieve a new state-of-the-art on a long document coreference benchmark, LitBank. Qipeng Guo, Xiangkun Hu, Yue Zhang 0004, Xipeng Qiu, Zheng Zhang 0001 |
ACL (1) | 5 |
| 2023 | Distributed Marker Representation for Ambiguous Discourse Markers and Entangled RelationsabstractDiscourse analysis is an important task because it models intrinsic semantic structures between sentences in a document.Discourse markers are natural representations of discourse in our daily language.One challenge is that the markers as well as pre-defined and human-labeled discourse relations can be ambiguous when describing the semantics between sentences.We believe that a better approach is to use a contextual-dependent distribution over the markers to express discourse information.In this work, we propose to learn a Distributed Marker Representation (DMR) by utilizing the (potentially) unlimited discourse marker data with a latent discourse sense, thereby bridging markers with sentence pairs.Such representations can be learned automatically from data without supervision, and in turn provide insights into the data itself.Experiments show the SOTA performance of our DMR on the implicit discourse relation recognition task and strong interpretability.Our method also offers a valuable tool to understand complex ambiguity and entanglement among discourse markers and manually defined discourse relations. Dongyu Ru, Xipeng Qiu, Yue Zhang 0004, Zheng Zhang 0001 |
ACL (1) | 5 |
| 2023 | StoryAnalogy: Deriving Story-level Analogies from Large Language Models to Unlock Analogical UnderstandingabstractCheng Jiayang, Lin Qiu, Tsz Chan, Tianqing Fang, Weiqi Wang, Chunkit Chan, Dongyu Ru, Qipeng Guo, Hongming Zhang, Yangqiu Song, Yue Zhang, Zheng Zhang. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Cheng Jiayang, Tsz Ho Chan, Tianqing Fang, Weiqi Wang 0001, Chunkit Chan, Dongyu Ru, Qipeng Guo, Hongming Zhang 0009, Yangqiu Song, Yue Zhang 0004, Zheng Zhang 0001 |
EMNLP | 12 |
| 2023 | Plan, Verify and Switch: Integrated Reasoning with Diverse X-of-ThoughtsabstractAs large language models (LLMs) have shown effectiveness with different prompting methods, such as Chain of Thought, Program of Thought, we find that these methods have formed a great complementarity to each other on math reasoning tasks.In this work, we propose XoT, an integrated problem solving framework by prompting LLMs with diverse reasoning thoughts.For each question, XoT always begins with selecting the most suitable method then executes each method iteratively.Within each iteration, XoT actively checks the validity of the generated answer and incorporates the feedback from external executors, allowing it to dynamically switch among different prompting methods.Through extensive experiments on 10 popular math reasoning datasets, we demonstrate the effectiveness of our proposed approach and thoroughly analyze the strengths of each module.Moreover, empirical results suggest that our framework is orthogonal to recent work that makes improvements on single reasoning methods and can further generalise to logical reasoning domain.By allowing method switching, XoT provides a fresh perspective on the collaborative integration of diverse reasoning thoughts in a unified framework. Tengxiao Liu, Qipeng Guo, Yuqing Yang 0004, Xiangkun Hu, Yue Zhang 0004, Xipeng Qiu, Zheng Zhang 0001 |
EMNLP | 7 |
| 2023 | Enhancing Uncertainty-Based Hallucination Detection with Stronger FocusabstractTianhang Zhang, Lin Qiu, Qipeng Guo, Cheng Deng, Yue Zhang, Zheng Zhang, Chenghu Zhou, Xinbing Wang, Luoyi Fu. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Tianhang Zhang, Qipeng Guo, Cheng Deng 0001, Yue Zhang 0004, Zheng Zhang 0001, Chenghu Zhou, Xinbing Wang, Luoyi Fu |
EMNLP | 6 |
| 2023 | Unsupervised Open-Vocabulary Object Localization in VideosabstractIn this paper, we show that recent advances in video representation learning and pre-trained vision-language models allow for substantial improvements in self-supervised video object localization. We propose a method that first localizes objects in videos via a slot attention approach and then assigns text to the obtained slots. The latter is achieved by an unsupervised way to read localized semantic information from the pre-trained CLIP model. The resulting video object localization is entirely unsupervised apart from the implicit annotation contained in CLIP, and it is effectively the first unsupervised approach that yields good results on regular video benchmarks. Zechen Bai, Tianjun Xiao, Dominik Zietlow, Max Horn, Carl-Johann Simon-Gabriel, Zheng Shou 0001, Francesco Locatello, Bernt Schiele, Thomas Brox, Zheng Zhang 0001, Yanwei Fu 0001, Tong He 0002 |
ICCV | 12 |
| 2023 | Rethinking Amodal Video Segmentation from Learning Supervised Signals with Object-centric RepresentationabstractVideo amodal segmentation is a particularly challenging task in computer vision, which requires to deduce the full shape of an object from the visible parts of it. Recently, some studies have achieved promising performance by using motion flow to integrate information across frames under a self-supervised setting. However, motion flow has a clear limitation by the two factors of moving cameras and object deformation. This paper presents a rethinking to previous works. We particularly leverage the supervised signals with object-centric representation in real-world scenarios. The underlying idea is the supervision signal of the specific object and the features from different views can mutually benefit the deduction of the full mask in any specific frame. We thus propose an Efficient object-centric Representation amodal Segmentation (EoRaS). Specially, beyond solely relying on supervision signals, we design a translation module to project image features into the Bird’s-Eye View (BEV), which introduces 3D information to improve current feature quality. Furthermore, we propose a multi-view fusion layer based temporal module which is equipped with a set of object slots and interacts with features from different views by attention mechanism to fulfill sufficient object representation completion. As a result, the full mask of the object can be decoded from image features updated by object slots. Extensive experiments on both real-world and synthetic benchmarks demonstrate the superiority of our proposed method, achieving state-of-the-art performance. Our code will be released at https://github.com/kfan21/EoRaS. Jingshi Lei, Xuelin Qian, Miaopeng Yu, Tianjun Xiao, Tong He 0002, Zheng Zhang 0001, Yanwei Fu 0001 |
ICCV | 7 |
| 2023 | Coarse-to-Fine Amodal Segmentation with Shape PriorabstractAmodal object segmentation is a challenging task that involves segmenting both visible and occluded parts of an object. In this paper, we propose a novel approach, called Coarse-to-Fine Segmentation (C2F-Seg), that addresses this problem by progressively modeling the amodal segmentation. C2F-Seg initially reduces the learning space from the pixel-level image space to the vector-quantized latent space. This enables us to better handle long-range dependencies and learn a coarse-grained amodal segment from visual features and visible segments. However, this latent space lacks detailed information about the object, which makes it difficult to provide a precise segmentation directly. To address this issue, we propose a convolution refine module to inject fine-grained information and provide a more precise amodal object segmentation based on visual features and coarse-predicted segmentation. To help the studies of amodal object segmentation, we create a synthetic amodal dataset, named as MOViD-Amodal (MOViD-A), which can be used for both image and video amodal object segmentation. We extensively evaluate our model on two benchmark datasets: KINS and COCO-A. Our empirical results demonstrate the superiority of C2F-Seg. Moreover, we exhibit the potential of our approach for video amodal object segmentation tasks on FISHBOWL and our proposed MOViD-A. Project page at: https://jianxgao.github.io/C2F-Seg. Jianxiong Gao, Xuelin Qian, Yikai Wang 0002, Tianjun Xiao, Tong He 0002, Zheng Zhang 0001, Yanwei Fu 0001 |
ICCV | 6 |
| 2023 | Object-Centric Multiple Object TrackingabstractUnsupervised object-centric learning methods allow the partitioning of scenes into entities without additional localization information and are excellent candidates for reducing the annotation burden of multiple-object tracking (MOT) pipelines. Unfortunately, they lack two key properties: objects are often split into parts and are not consistently tracked over time. In fact, state-of-the-art models achieve pixel-level accuracy and temporal consistency by relying on supervised object detection with additional ID labels for the association through time. This paper proposes a video object-centric model for MOT. It consists of an index-merge module that adapts the object-centric slots into detection outputs and an object memory module that builds complete object prototypes to handle occlusions. Benefited from object-centric learning, we only require sparse detection labels (0%-6.25%) for object localization and feature binding. Relying on our self-supervised Expectation-Maximization-inspired loss for object association, our approach requires no ID labels. Our experiments significantly narrow the gap between the existing object-centric model and the fully supervised state-of-the-art and outperform several unsupervised trackers. Code is available at https://github.com/amazon-science/object-centric-multiple-object-tracking. Max Horn, Yizhuo Ding, Tong He 0002, Zechen Bai, Dominik Zietlow, Carl-Johann Simon-Gabriel, Bing Shuai, Zhuowen Tu, Thomas Brox, Bernt Schiele, Yanwei Fu 0001, Francesco Locatello, Zheng Zhang 0001, Tianjun Xiao |
ICCV | 15 |
| 2023 | Bridging the Gap to Real-World Object-Centric Learning
Maximilian Seitzer, Max Horn, Andrii Zadaianchuk, Dominik Zietlow, Tianjun Xiao, Carl-Johann Simon-Gabriel, Tong He 0002, Zheng Zhang 0001, Bernhard Schölkopf, Thomas Brox, Francesco Locatello |
ICLR | 8 |
| 2022 | RLET: A Reinforcement Learning Based Approach for Explainable QA with Entailment TreesabstractInterpreting the reasoning process from questions to answers poses a challenge in approaching explainable QA.A recently proposed structured reasoning format, entailment tree, manages to offer explicit logical deductions with entailment steps in a tree structure.To generate entailment trees, prior single pass sequence-tosequence models lack visible internal decision probability, while stepwise approaches are supervised with extracted single step data and cannot model the tree as a whole.In this work, we propose RLET, a Reinforcement Learning based Entailment Tree generation framework, which is trained utilising the cumulative signals across the whole tree.RLET iteratively performs single step reasoning with sentence selection and deduction generation modules, from which the training signal is accumulated across the tree with elaborately designed aligned reward function that is consistent with the evaluation.To the best of our knowledge, we are the first to introduce RL into the entailment tree generation task.Experiments on three settings of the EntailmentBank dataset demonstrate the strength of using RL framework. Tengxiao Liu, Qipeng Guo, Xiangkun Hu, Yue Zhang 0004, Xipeng Qiu, Zheng Zhang 0001 |
EMNLP | 6 |
| 2022 | Inductive Relation Prediction Using Analogy Subgraph Embeddings
Jiarui Jin, Yangkun Wang, Kounianhua Du, Weinan Zhang 0001, Zheng Zhang 0001, David P. Wipf, Yong Yu 0001 |
ICLR | 5 |
| 2022 | Why Propagate Alone? Parallel Use of Labels and Features on Graphs
Yangkun Wang, Jiarui Jin, Weinan Zhang 0001, Yongyi Yang, Jiuhai Chen, Yong Yu 0001, Zheng Zhang 0001, Zengfeng Huang, David P. Wipf |
ICLR | 8 |
| 2022 | Learning Enhanced Representation for Tabular Data via Neighborhood PropagationabstractPrediction over tabular data is an essential and fundamental problem in many important downstream tasks. However, existing methods either take a data instance of the table independently as input or do not fully utilize the multi-row features and labels to directly change and enhance the target data representations. In this paper, we propose to 1) construct a hypergraph from relevant data instance retrieval to model the cross-row and cross-column patterns of those instances, and 2) perform message Propagation to Enhance the target data instance representation for Tabular prediction tasks. Specifically, our specially-designed message propagation step benefits from 1) the fusion of label and features during propagation, and 2) locality-aware multiplicative high-order interaction between features. Experiments on two important tabular prediction tasks validate the superiority of the proposed PET model against other baselines. Additionally, we demonstrate the effectiveness of the model components and the feature enhancement ability of PET via various ablation studies and visualizations. The code is available at https://github.com/KounianhuaDu/PET. Kounianhua Du, Weinan Zhang 0001, Ruiwen Zhou, Yangkun Wang, Xilong Zhao, Jiarui Jin, Zheng Zhang 0001, David P. Wipf |
NeurIPS | 8 |
| 2022 | Self-supervised Amodal Video Object SegmentationabstractAmodal perception requires inferring the full shape of an object that is partially occluded. This task is particularly challenging on two levels: (1) it requires more information than what is contained in the instant retina or imaging sensor, (2) it is difficult to obtain enough well-annotated amodal labels for supervision. To this end, this paper develops a new framework of Self-supervised amodal Video object segmentation (SaVos). Our method efficiently leverages the visual information of video temporal sequences to infer the amodal mask of objects. The key intuition is that the occluded part of an object can be explained away if that part is visible in other frames, possibly deformed as long as the deformation can be reasonably learned. Accordingly, we derive a novel self-supervised learning paradigm that efficiently utilizes the visible object parts as the supervision to guide the training on videos. In addition to learning type prior to complete masks for known types, SaVos also learns the spatiotemporal prior, which is also useful for the amodal task and could generalize to unseen types. The proposed framework achieves the state-of-the-art performance on the synthetic amodal segmentation benchmark FISHBOWL and the real world benchmark KINS-Video-Car. Further, it lends itself well to being transferred to novel distributions using test-time adaptation, outperforming existing models even after the transfer to a new distribution. Jian Yao 0004, Yuxin Hong, Chiyu Wang, Tianjun Xiao, Tong He 0002, Francesco Locatello, David P. Wipf, Yanwei Fu 0001, Zheng Zhang 0001 |
NeurIPS | 9 |
| 2022 | Learning Cognitive Map Representations for Navigation by Sensory-Motor IntegrationabstractHow to transform a mixed flow of sensory and motor information into memory state of self-location and to build map representations of the environment are central questions in the navigation research. Studies in neuroscience have shown that place cells in the hippocampus of the rodent brains form dynamic cognitive representations of locations in the environment. We propose a neural-network model called sensory-motor integration network model (SeMINet) to learn cognitive map representations by integrating sensory and motor information while an agent is exploring a virtual environment. This biologically inspired model consists of a deep neural network representing visual features of the environment, a recurrent network of place units encoding spatial information by sensorimotor integration, and a secondary network to decode the locations of the agent from spatial representations. The recurrent connections between the place units sustain an activity bump in the network without the need of sensory inputs, and the asymmetry in the connections propagates the activity bump in the network, forming a dynamic memory state which matches the motion of the agent. A competitive learning process establishes the association between the sensory representations and the memory state of the place units, and is able to correct the cumulative path-integration errors. The simulation results demonstrate that the network forms neural codes that convey location information of the agent independent of its head direction. The decoding network reliably predicts the location even when the movement is subject to noise. The proposed SeMINet thus provides a brain-inspired neural-network model for cognitive map updated by both self-motion cues and visual cues. Dongye Zhao, Zheng Zhang 0001, Hong Lu 0001, Sen Cheng, Bailu Si, Xisheng Feng |
IEEE Trans. Cybern. | 2 |
| 2022 | GraphHINGE: Learning Interaction Models of Structured Neighborhood on Heterogeneous Information NetworkabstractHeterogeneous information network (HIN) has been widely used to characterize entities of various types and their complex relations. Recent attempts either rely on explicit path reachability to leverage path-based semantic relatedness or graph neighborhood to learn heterogeneous network representations before predictions. These weakly coupled manners overlook the rich interactions among neighbor nodes, which introduces an early summarization issue. In this article, we propose GraphHINGE ( H eterogeneous IN teract and aggre G at E ), which captures and aggregates the interactive patterns between each pair of nodes through their structured neighborhoods. Specifically, we first introduce Neighborhood-based Interaction (NI) module to model the interactive patterns under the same metapaths, and then extend it to Cross Neighborhood-based Interaction (CNI) module to deal with different metapaths. Next, in order to address the complexity issue on large-scale networks, we formulate the interaction modules via a convolutional framework and learn the parameters efficiently with fast Fourier transform. Furthermore, we design a novel neighborhood-based selection (NS) mechanism, a sampling strategy, to filter high-order neighborhood information based on their low-order performance. The extensive experiments on six different types of heterogeneous graphs demonstrate the performance gains by comparing with state-of-the-arts in both click-through rate prediction and top-N recommendation tasks. Jiarui Jin, Kounianhua Du, Weinan Zhang 0001, Jiarui Qin, Yong Yu 0001, Zheng Zhang 0001, Alexander J. Smola |
ACM Trans. Inf. Syst. | 7 |
| 2021 | A Unified Generative Framework for Aspect-based Sentiment AnalysisabstractHang Yan, Junqi Dai, Tuo Ji, Xipeng Qiu, Zheng Zhang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Hang Yan 0001, Junqi Dai, Tuo Ji, Xipeng Qiu, Zheng Zhang 0001 |
ACL/IJCNLP (1) | 5 |
| 2021 | A Unified Generative Framework for Various NER SubtasksabstractHang Yan, Tao Gui, Junqi Dai, Qipeng Guo, Zheng Zhang, Xipeng Qiu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Hang Yan 0001, Tao Gui, Junqi Dai, Qipeng Guo, Zheng Zhang 0001, Xipeng Qiu |
ACL/IJCNLP (1) | 5 |
| 2021 | Fork or Fail: Cycle-Consistent Training with Many-to-One MappingsabstractCycle-consistent training is widely used for jointly learning a forward and inverse mapping between two domains of interest without the cumbersome requirement of collecting matched pairs within each domain. In this regard, the implicit assumption is that there exists (at least approximately) a ground-truth bijection such that a given input from either domain can be accurately reconstructed from successive application of the respective mappings. But in many applications no such bijection can be expected to exist and large reconstruction errors can compromise the success of cycle-consistent training. As one important instance of this limitation, we consider practically-relevant situations where there exists a many-to-one or surjective mapping between domains. To address this regime, we develop a conditional variational autoencoder (CVAE) approach that can be viewed as converting surjective mappings to implicit bijections whereby reconstruction errors in both directions can be minimized, and as a natural byproduct, realistic output diversity can be obtained in the one-to-many direction. As theoretical motivation, we analyze a simplified scenario whereby minima of the proposed CVAE-based energy function align with the recovery of ground-truth surjective mappings. On the empirical side, we consider a synthetic image dataset with known ground-truth, as well as a real-world application involving natural language generation from knowledge graphs and vice versa, a prototypical surjective case. For the latter, our CVAE pipeline can capture such many-to-one mappings during cycle training while promoting textural diversity for graph-to-text tasks. Qipeng Guo, Zhijing Jin 0001, Ziyu Wang 0006, Xipeng Qiu, Weinan Zhang 0001, Jun Zhu 0001, Zheng Zhang 0001, David P. Wipf |
AISTATS | 7 |
| 2021 | Learning Hierarchical Graph Neural Networks for Image ClusteringabstractWe propose a hierarchical graph neural network (GNN) model that learns how to cluster a set of images into an unknown number of identities using a training set of images annotated with labels belonging to a disjoint set of identities. Our hierarchical GNN uses a novel approach to merge connected components predicted at each level of the hierarchy to form a new graph at the next level. Unlike fully unsupervised hierarchical clustering, the choice of grouping and complexity criteria stems naturally from supervision in the training set. The resulting method, Hi-LANDER, achieves an average of 49% improvement in F-score and 7% increase in Normalized Mutual Information (NMI) relative to current GNN-based clustering algorithms. Additionally, state-of-the-art GNN-based methods rely on separate models to predict linkage probabilities and node densities as intermediate steps of the clustering process. In contrast, our unified framework achieves a three-fold decrease in computational cost. Our training and inference code are released1. Yifan Xing, Tong He 0002, Tianjun Xiao, Yuanjun Xiong, Wei Xia 0009, David P. Wipf, Zheng Zhang 0001, Stefano Soatto |
ICCV | 8 |
| 2021 | Graph Neural Networks Inspired by Classical Iterative AlgorithmsabstractDespite the recent success of graph neural networks (GNN), common architectures often exhibit significant limitations, including sensitivity to oversmoothing, long-range dependencies, and spurious edges, e.g., as can occur as a result of graph heterophily or adversarial attacks. To at least partially address these issues within a simple transparent framework, we consider a new family of GNN layers designed to mimic and integrate the update rules of two classical iterative algorithms, namely, proximal gradient descent and iterative reweighted least squares (IRLS). The former defines an extensible base GNN architecture that is immune to oversmoothing while nonetheless capturing long-range dependencies by allowing arbitrary propagation steps. In contrast, the latter produces a novel attention mechanism that is explicitly anchored to an underlying end-to-end energy function, contributing stability with respect to edge uncertainty. When combined we obtain an extremely simple yet robust model that we evaluate across disparate scenarios including standardized benchmarks, adversarially-perturbated graphs, graphs with heterophily, and graphs involving long-range dependencies. In doing so, we compare against SOTA GNN approaches that have been explicitly designed for the respective task, achieving competitive or superior node classification accuracy. Our code is available at https://github.com/FFTYYY/TWIRLS. And for an extended version of this work, please see https://arxiv.org/abs/2103.06064. Yongyi Yang, Yangkun Wang, Jinjing Zhou, Zhewei Wei, Zheng Zhang 0001, Zengfeng Huang, David P. Wipf |
ICML | 7 |
| 2021 | GRIN: Generative Relation and Intention Network for Multi-agent Trajectory PredictionabstractLearning the distribution of future trajectories conditioned on the past is a crucial problem for understanding multi-agent systems. This is challenging because humans make decisions based on complex social relations and personal intents, resulting in highly complex uncertainties over trajectories. To address this problem, we propose a conditional deep generative model that combines advances in graph neural networks. The prior and recognition model encodes two types of latent codes for each agent: an inter-agent latent code to represent social relations and an intra-agent latent code to represent agent intentions. The decoder is carefully devised to leverage the codes in a disentangled way to predict multi-modal future trajectory distribution. Specifically, a graph attention network built upon inter-agent latent code is used to learn continuous pair-wise relations, and an agent's motion is controlled by its latent intents and its observations of all other agents. Through experiments on both synthetic and real-world datasets, we show that our model outperforms previous work in multiple performance metrics. We also show that our model generates realistic multi-modal trajectories. Longyuan Li, Jian Yao 0004, Li Kevin Wenliang, Tong He 0002, Tianjun Xiao, Junchi Yan, David P. Wipf, Zheng Zhang 0001 |
NeurIPS | 8 |
| 2021 | Scalable Graph Neural Networks with Deep Graph LibraryabstractLearning from graph and relational data plays a major role in many applications including social network analysis, marketing, e-commerce, information retrieval, knowledge modeling, medical and biological sciences, engineering, and others. Recently, Graph Neural Networks (GNNs) have emerged as a promising new learning framework capable of bringing the power of deep representation learning to graph and relational data. This ever-growing body of research has shown that GNNs achieve state-of-the-art performance for problems such as link prediction, fraud detection, target-ligand binding activity prediction, knowledge-graph completion, and product recommendations. In practice, many of the real-world graphs are very large. It is urgent to have scalable solutions to train GNN on large graphs efficiently. Da Zheng 0004, Xiang Song 0003, Zheng Zhang 0001, George Karypis |
WSDM | 5 |
| 2021 | Syntax-guided text generation via graph neural network
Qipeng Guo, Xipeng Qiu, Xiangyang Xue 0001, Zheng Zhang 0001 |
Sci. China Inf. Sci. | 4 |
| 2020 | Multi-Scale Self-Attention for Text ClassificationabstractIn this paper, we introduce the prior knowledge, multi-scale structure, into self-attention modules. We propose a Multi-Scale Transformer which uses multi-scale multi-head self-attention to capture features from different scales. Based on the linguistic perspective and the analysis of pre-trained Transformer (BERT) on a huge corpus, we further design a strategy to control the scale distribution for each layer. Results of three different kinds of tasks (21 datasets) show our Multi-Scale Transformer outperforms the standard Transformer consistently and significantly on small and moderate size datasets. Qipeng Guo, Xipeng Qiu, Pengfei Liu 0003, Xiangyang Xue 0001, Zheng Zhang 0001 |
AAAI | 5 |
| 2020 | GenWiki: A Dataset of 1.3 Million Content-Sharing Text and Graphs for Unsupervised Graph-to-Text GenerationabstractData collection for the knowledge graph-to-text generation is expensive.As a result, research on unsupervised models has emerged as an active field recently.However, most unsupervised models have to use non-parallel versions of existing small supervised datasets, which largely constrain their potential.In this paper, we propose a large-scale, general-domain dataset, GenWiki.Our unsupervised dataset has 1.3M text and graph examples, respectively.With a human-annotated test set, we provide this new benchmark dataset for future research on unsupervised text generation from knowledge graphs. 1 Zhijing Jin 0001, Qipeng Guo, Xipeng Qiu, Zheng Zhang 0001 |
COLING | 4 |
| 2020 | CoLAKE: Contextualized Language and Knowledge EmbeddingabstractWith the emerging branch of incorporating factual knowledge into pre-trained language models such as BERT, most existing models consider shallow, static, and separately pre-trained entity embeddings, which limits the performance gains of these models.Few works explore the potential of deep contextualized knowledge representation when injecting knowledge.In this paper, we propose the Contextualized Language and Knowledge Embedding (CoLAKE), which jointly learns contextualized representation for both language and knowledge with the extended MLM objective.Instead of injecting only entity embeddings, CoLAKE extracts the knowledge context of an entity from large-scale knowledge bases.To handle the heterogeneity of knowledge context and language context, we integrate them in a unified data structure, word-knowledge graph (WK graph).CoLAKE is pre-trained on large-scale WK graphs with the modified Transformer encoder.We conduct experiments on knowledge-driven tasks, knowledge probing tasks, and language understanding tasks.Experimental results show that CoLAKE outperforms previous counterparts on most of the tasks.Besides, CoLAKE achieves surprisingly high performance on our synthetic task called word-knowledge graph completion, which shows the superiority of simultaneously contextualizing language and knowledge representation. 1 Tianxiang Sun, Yunfan Shao, Xipeng Qiu, Qipeng Guo, Yaru Hu, Xuanjing Huang 0001, Zheng Zhang 0001 |
COLING | 7 |
| 2020 | An Efficient Neighborhood-based Interaction Model for Recommendation on Heterogeneous GraphabstractThere is an influx of heterogeneous information network (HIN) based recommender systems in recent years since HIN is capable of characterizing complex graphs and contains rich semantics. Although the existing approaches have achieved performance improvement, while practical, they still face the following problems. On one hand, most existing HIN-based methods rely on explicit path reachability to leverage path-based semantic relatedness between users and items, e.g., metapath-based similarities. These methods are hard to use and integrate since path connections are sparse or noisy, and are often of different lengths. On the other hand, other graph-based methods aim to learn effective heterogeneous network representations by compressing node together with its neighborhood information into single embedding before prediction. This weakly coupled manner in modeling overlooks the rich interactions among nodes, which introduces an early summarization issue. In this paper, we propose an end-to-end Neighborhood-based Interaction Model for Recommendation (NIRec) to address above problems. Specifically, we first analyze the significance of learning interactions in HINs and then propose a novel formulation to capture the interactive patterns between each pair of nodes through their metapath-guided neighborhoods. Then, to explore complex interactions between metapaths and deal with the learning complexity on large-scale networks, we formulate interaction in a convolutional way and learn efficiently with fast Fourier transform. The extensive experiments on four different types of heterogeneous graphs demonstrate the performance gains of NIRec comparing with state-of-the-arts. To the best of our knowledge, this is the first work providing an efficient neighborhood-based interaction model in the HIN-based recommendations. Jiarui Jin, Jiarui Qin, Kounianhua Du, Weinan Zhang 0001, Yong Yu 0001, Zheng Zhang 0001, Alexander J. Smola |
KDD | 7 |
| 2020 | Scalable Graph Neural Networks with Deep Graph LibraryabstractLearning from graph and relational data plays a major role in many applications including social network analysis, marketing, e-commerce, information retrieval, knowledge modeling, medical and biological sciences, engineering, and others. In the last few years, Graph Neural Networks (GNNs) have emerged as a promising new supervised learning framework capable of bringing the power of deep representation learning to graph and relational data. This ever-growing body of research has shown that GNNs achieve state-of-the-art performance for problems such as link prediction, fraud detection, target-ligand binding activity prediction, knowledge-graph completion, and product recommendations. In practice, many of the real-world graphs are very large. It is urgent to have scalable solutions to train GNN on large graphs efficiently. Da Zheng 0004, Zheng Zhang 0001, George Karypis |
KDD | 4 |
| 2020 | FeatGraph: a flexible and efficient backend for graph neural network systemsabstractGraph neural networks (GNNs) are gaining popularity as a promising approach to machine learning on graphs. Unlike traditional graph workloads where each vertex/edge is associated with a scalar, GNNs attach a feature tensor to each vertex/edge. This additional feature dimension, along with consequently more complex vertex- and edge-wise computations, has enormous implications on locality and parallelism, which existing graph processing systems fail to exploit. This paper proposes FeatGraph to accelerate GNN workloads by co-optimizing graph traversal and feature dimension computation. FeatGraph provides a flexible programming interface to express diverse GNN models by composing coarse-grained sparse templates with fine-grained user-defined functions (UDFs) on each vertex/edge. FeatGraph incorporates optimizations for graph traversal into the sparse templates and allows users to specify optimizations for UDFs with a feature dimension schedule (FDS). FeatGraph speeds up end-to-end GNN training and inference by up to 32× on CPU and 7× on GPU. Zihao Ye 0001, Da Zheng 0004, Mu Li 0003, Zheng Zhang 0001, Zhiru Zhang, Yida Wang 0003 |
SC | 7 |
| 2020 | DGL-KE: Training Knowledge Graph Embeddings at ScaleabstractKnowledge graphs have emerged as a key abstraction for organizing information in diverse domains and their embeddings are increasingly used to harness their information in various information retrieval and machine learning tasks. However, the ever growing size of knowledge graphs requires computationally efficient algorithms capable of scaling to graphs with millions of nodes and billions of edges. This paper presents DGL-KE, an open-source package to efficiently compute knowledge graph embeddings. DGL-KE introduces various novel optimizations that accelerate training on knowledge graphs with millions of nodes and billions of edges using multi-processing, multi-GPU, and distributed parallelism. These optimizations are designed to increase data locality, reduce communication overhead, overlap computations with memory accesses, and achieve high operation efficiency. Experiments on knowledge graphs consisting of over 86M nodes and 338M edges show that DGL-KE can compute embeddings in 100 minutes on an EC2 instance with 8 GPUs and 30 minutes on an EC2 cluster with 4 machines with 48 cores/machine. These results represent a 2× ~ 5× speedup over the best competing approaches. DGL-KE is available on https://github.com/awslabs/dgl-ke. Da Zheng 0004, Xiang Song 0003, Chao Ma 0025, Zeyuan Tan, Zihao Ye 0001, Zheng Zhang 0001, George Karypis |
SIGIR | 8 |
| 2019 | Low-Rank and Locality Constrained Self-Attention for Sequence ModelingabstractSelf-attention mechanism becomes more and more popular in natural language processing (NLP) applications. Recent studies show the Transformer architecture which relies mainly on the attention mechanism achieves much success on large datasets. But a raised problem is its generalization ability is weaker than CNN and RNN on many moderate-sized datasets. We think the reason can be attributed to its unsuitable inductive bias of the self-attention structure. In this paper, we regard the self-attention as matrix decomposition problem and propose an improved self-attention module by introducing two linguistic constraints: low-rank and locality. We further develop the low-rank attention and band attention to parameterize the self-attention mechanism under the low-rank and locality constraints. Experiments on several real NLP tasks show our model outperforms the vanilla Transformer and other self-attention models on moderate size datasets. Additionally, evaluation on a synthetic task gives us a more detailed understanding of working mechanisms of different architectures. Qipeng Guo, Xipeng Qiu, Xiangyang Xue 0001, Zheng Zhang 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2018 | Loss Functions for Multiset PredictionabstractWe study the problem of multiset prediction. The goal of multiset prediction is to train a predictor that maps an input to a multiset consisting of multiple items. Unlike existing problems in supervised learning, such as classification, ranking and sequence generation, there is no known order among items in a target multiset, and each item in the multiset may appear more than once, making this problem extremely challenging. In this paper, we propose a novel multiset loss function by viewing this problem from the perspective of sequential decision making. The proposed multiset loss function is empirically evaluated on two families of datasets, one synthetic and the other real, with varying levels of difficulty, against various baseline loss functions including reinforcement learning, sequence, and aggregated distribution matching loss functions. The experiments reveal the effectiveness of the proposed loss function over the others. Sean Welleck, Zixin Yao, Yu Gai, Jialin Mao, Zheng Zhang 0001, Kyunghyun Cho |
NeurIPS | 5 |
| 2017 | Saliency-based Sequential Image Attention with Multiset PredictionabstractHumans process visual scenes selectively and sequentially using attention. Central to models of human visual attention is the saliency map. We propose a hierarchical visual architecture that operates on a saliency map and uses a novel attention mechanism to sequentially focus on salient regions and take additional glimpses within those regions. The architecture is motivated by human visual attention, and is used for multi-label image classification on a novel multiset task, demonstrating that it achieves high precision and recall while localizing objects with its attention. Unlike conventional multi-label image classification models, the model supports multiset prediction due to a reinforcement-learning based training process that allows for arbitrary label permutation and multiple instances per label. Sean Welleck, Jialin Mao, Kyunghyun Cho, Zheng Zhang 0001 |
NIPS | 4 |
| 2015 | The application of two-level attention models in deep convolutional neural network for fine-grained image classificationabstractFine-grained classification is challenging because categories can only be discriminated by subtle and local differences. Variances in the pose, scale or rotation usually make the problem more difficult. Most fine-grained classification systems follow the pipeline of finding foreground object or object parts (where) to extract discriminative features (what). In this paper, we propose to apply visual attention to fine-grained classification task using deep neural network. Our pipeline integrates three types of attention: the bottom-up attention that propose candidate patches, the object-level top-down attention that selects relevant patches to a certain object, and the part-level top-down attention that localizes discriminative parts. We combine these attentions to train domain-specific deep nets, then use it to improve both the what and where aspects. Importantly, we avoid using expensive annotations like bounding box or part information from end-to-end. The weak supervision constraint makes our work easier to generalize. We have verified the effectiveness of the method on the subsets of ILSVRC2012 dataset and CUB200 2011 dataset. Our pipeline delivered significant improvements and achieved the best accuracy under the weakest supervision condition. The performance is competitive against other methods that rely on additional annotations. Tianjun Xiao, Yichong Xu, Kuiyuan Yang, Yuxin Peng 0001, Zheng Zhang 0001 |
CVPR | 6 |
| 2015 | Multiple Granularity Descriptors for Fine-Grained CategorizationabstractFine-grained categorization, which aims to distinguish subordinate-level categories such as bird species or dog breeds, is an extremely challenging task. This is due to two main issues: how to localize discriminative regions for recognition and how to learn sophisticated features for representation. Neither of them is easy to handle if there is insufficient labeled data. We leverage the fact that a subordinate-level object already has other labels in its ontology tree. These "free" labels can be used to train a series of CNN-based classifiers, each specialized at one grain level. The internal representations of these networks have different region of interests, allowing the construction of multi-grained descriptors that encode informative and discriminative features covering all the grain levels. Our multiple granularity framework can be learned with the weakest supervision, requiring only image-level label and avoiding the use of labor-intensive bounding box or part annotations. Experimental results on three challenging fine-grained image datasets demonstrate that our approach outperforms state-of-the-art algorithms, including those requiring strong labels. Dequan Wang, Jie Shao 0006, Wei Zhang 0016, Xiangyang Xue 0001, Zheng Zhang 0001 |
ICCV | 6 |
| 2015 | Distributed Outlier Detection using Compressive SensingabstractComputing outliers and related statistical aggregation functions from large-scale big data sources is a critical operation in many cloud computing scenarios, e.g. service quality assurance, fraud detection, or novelty discovery. Such problems commonly have to be solved in a distributed environment where each node only has a local slice of the entirety of the data. To process a query on the global data, each node must transmit its local slice of data or an aggregated subset thereof to a global aggregator node, which can then compute the desired statistical aggregation function. In this context, reducing the total communication cost is often critical to the overall efficiency. Ying Yan 0006, Bojun Huang, Xuzhan Sun, Jiaqi Mu, Zheng Zhang 0001, Thomas Moscibroda |
SIGMOD Conference | 6 |
| 2014 | Error-Driven Incremental Learning in Deep Convolutional Neural Network for Large-Scale Image ClassificationabstractSupervised learning using deep convolutional neural network has shown its promise in large-scale image classification task. As a building block, it is now well positioned to be part of a larger system that tackles real-life multimedia tasks. An unresolved issue is that such model is trained on a static snapshot of data. Instead, this paper positions the training as a continuous learning process as new classes of data arrive. A system with such capability is useful in practical scenarios, as it gradually expands its capacity to predict increasing number of new classes. It is also our attempt to address the more fundamental issue: a good learning system must deal with new knowledge that it is exposed to, much as how human do. Tianjun Xiao, Kuiyuan Yang, Yuxin Peng 0001, Zheng Zhang 0001 |
ACM Multimedia | 5 |
| 2014 | Error-bounded Sampling for Analytics on Big Sparse DataabstractAggregation queries are at the core of business intelligence and data analytics. In the big data era, many scalable shared-nothing systems have been developed to process aggregation queries over massive amount of data. Microsoft's SCOPE is a well-known instance in this category. Nevertheless, aggregation queries are still expensive, because query processing needs to consume the entire data set, which is often hundreds of terabytes. Data sampling is a technique that samples a small portion of data to process and returns an approximate result with an error bound, thereby reducing the query's execution time. While similar problems were studied in the database literature, we encountered new challenges that disable most of prior efforts: (1) error bounds are dictated by end users and cannot be compromised, (2) data is sparse , meaning data has a limited population but a wide range. For such cases, conventional uniform sampling often yield high sampling rates and thus deliver limited or no performance gains. In this paper, we propose error-bounded stratified sampling to reduce sample size. The technique relies on the insight that we may only reduce the sampling rate with the knowledge of data distributions. The technique has been implemented into Microsoft internal search query platform. Results show that the proposed approach can reduce up to 99% sample size comparing with uniform sampling, and its performance is robust against data volume and other key performance metrics. Ying Yan 0006, Liang Jeff Chen, Zheng Zhang 0001 |
Proc. VLDB Endow. | 3 |
| 2013 | TimeStream: reliable stream computation in the cloudabstractTimeStream is a distributed system designed specifically for low-latency continuous processing of big streaming data on a large cluster of commodity machines. The unique characteristics of this emerging application domain have led to a significantly different design from the popular MapReduce-style batch data processing. In particular, we advocate a powerful new abstraction called resilient substitution that caters to the specific needs in this new computation model to handle failure recovery and dynamic reconfiguration in response to load changes. Several real-world applications running on our prototype have been shown to scale robustly with low latency while at the same time maintaining the simple and concise declarative programming model. TimeStream handles an on-line advertising aggregation pipeline at a rate of 700,000 URLs per second with a 2-second delay, while performing sentiment analysis of Twitter data at a peak rate close to 10,000 tweets per second, with approximately 2-second delay. Zhengping Qian, Chunzhi Su, Zhuojie Wu, Hongyu Zhu 0003, Taizhi Zhang, Lidong Zhou, Zheng Zhang 0001 |
EuroSys | 9 |
| 2013 | Failure Recovery: When the Cure Is Worse Than the Disease
Sean McDirmid, Mao Yang 0004, Li Zhuang, Yingwei Luo, Tom Bergan, Madan Musuvathi, Zheng Zhang 0001, Lidong Zhou |
HotOS | 9 |
| 2012 | MadLINQ: large-scale distributed matrix computation for the cloudabstractThe computation core of many data-intensive applications can be best expressed as matrix computations. The MadLINQ project addresses the following two important research problems: the need for a highly scalable, efficient and fault-tolerant matrix computation system that is also easy to program, and the seamless integration of such specialized execution engines in a general purpose data-parallel computing system. Zhengping Qian, Xiuwei Chen, Nanxi Kang, Mingcheng Chen, Thomas Moscibroda, Zheng Zhang 0001 |
EuroSys | 7 |
| 2012 | A uniform solution to the independent set problem through tissue P systems with cell separation
Xingyi Zhang 0001, Xiangxiang Zeng, Bin Luo 0001, Zheng Zhang 0001 |
Frontiers Comput. Sci. | 4 |
| 2010 | Language-based replay via data flow cutabstractA replay tool aiming to reproduce a program's execution interposes itself at an appropriate replay interface between the program and the environment. During recording, it logs all non-deterministic side effects passing through the interface from the environment and feeds them back during replay. The replay interface is critical for correctness and recording overhead of replay tools. Ming Wu 0007, Fan Long, Xi Wang 0005, Zhilei Xu, Haoxiang Lin, Xuezheng Liu, Huayang Guo, Lidong Zhou, Zheng Zhang 0001 |
SIGSOFT FSE | 10 |
| 2009 | MPIWiz: subgroup reproducible replay of mpi applicationsabstractMessage Passing Interface (MPI) is a widely used standard for managing coarse-grained concurrency on distributed computers. Debugging parallel MPI applications, however, has always been a particularly challenging task due to their high degree of concurrent execution and non-deterministic behavior. Deterministic replay is a potentially powerful technique for addressing these challenges, with existing MPI replay tools adopting either data-replay or order-replay approaches. Unfortunately, each approach has its tradeoffs. Data-replay generates substantial log sizes by recording every communication message. Order-replay generates small logs, but requires all processes to be replayed together. We believe that these drawbacks are the primary reasons that inhibit the wide adoption of deterministic replay as the critical enabler of cyclic debugging of MPI applications. Ruini Xue, Xuezheng Liu, Ming Wu 0007, Zheng Zhang 0001, Geoffrey M. Voelker |
PPoPP | 7 |
| 2008 | Towards cinematic internet video-on-demandabstractVideo-on-demand (VoD) is increasingly popular with Internet users. It gives users greater choice and more control than live streaming or file downloading. Systems such as MSN Video and YouTube deliver content at low bitrates. This may suit short clips, but great films and 5-minute bloopers are as different as symphonies and jingles. For cinema, poor quality and high jitter are less acceptable. Combining user control with high bitrate is compelling, but technically challenging. Bin Cheng 0001, Lex Stein, Hai Jin 0001, Zheng Zhang 0001 |
EuroSys | 4 |
| 2008 | Hang analysis: fighting responsiveness bugsabstractSoft hang is an action that was expected to respond instantly but instead drives an application into a coma. While the application usually responds eventually, users cannot issue other requests while waiting. Such hang problems are widespread in productivity tools such as desktop applications; similar issues arise in server programs as well. Hang problems arise because the software contains blocking or time-consuming operations in graphical user interface (GUI) and other time-critical call paths that should not. Xi Wang 0005, Xuezheng Liu, Zhilei Xu, Haoxiang Lin, Xiaoge Wang, Zheng Zhang 0001 |
EuroSys | 7 |
| 2008 | A framework for lazy replication in P2P VoDabstractVideo-on-Demand (VoD) is a compelling application, but costly due to the load it places on servers. Peer-to-peer (P2P) techniques hold the potential to reduce centralized costs by sharing data between peers. There are many diffi-cult design issues associated with P2P for VoD. Viewing the problem as designing a large distributed cache, many of the issues can be expressed in terms of caching algorithms. In an earlier paper [6], we studied the performance of Grid-Cast, a P2P VoD system deployed on CERNET. From sys-tem traces, we found that departure misses are the major cause of server load. Motivated by this finding, this paper examines how to use replication to decrease departure misses and thereby further reduce server load. This paper proposes and evaluates a framework for lazy replication. Lazy replication postpones replication, trying to make efficient use of bandwidth. In our framework, two predictors are plugged in to create the working replication algorithm. Lazy replication with several predictors is com-pared with a näıve eager replication algorithm. We find that lazy replication is more efficient than eager replication, even when using two simple predictors. With these two simple predictors, lazy replication can decrease server load by 15% from multivideo caching with only a minor increase in net-work traffic. 1. Bin Cheng 0001, Lex Stein, Hai Jin 0001, Zheng Zhang 0001 |
NOSSDAV | 4 |
| 2008 | D3S: Debugging Deployed Distributed Systems
Xuezheng Liu, Xi Wang 0005, Feibo Chen, Xiaochen Lian, Ming Wu 0007, M. Frans Kaashoek, Zheng Zhang 0001 |
NSDI | 9 |
| 2008 | Corey: An Operating System for Many Cores
Silas Boyd-Wickizer, Haibo Chen 0001, Rong Chen 0001, Yandong Mao, M. Frans Kaashoek, Robert Morris 0005, Aleksey Pesterev, Lex Stein, Ming Wu 0007, Yue-hua Dai, Zheng Zhang 0001 |
OSDI | 12 |
| 2008 | R2: An Application-Level Kernel for Record and Replay
Xi Wang 0005, Xuezheng Liu, Zhilei Xu, Ming Wu 0007, M. Frans Kaashoek, Zheng Zhang 0001 |
OSDI | 8 |
| 2008 | Conditional correlation analysis for safe region-based memory managementabstractRegion-based memory management is a popular scheme in systems software for better organization and performance. In the scheme, a developer constructs a hierarchy of regions of different lifetimes and allocates objects in regions. When the developer deletes a region, the runtime will recursively delete all its subregions and simultaneously reclaim objects in the regions. The developer must construct a consistent placement of objects in regions; otherwise, if a region that contains pointers to other regions is not always deleted before pointees, an inconsistency will surface and cause dangling pointers, which may lead to either crashes or leaks. Xi Wang 0005, Zhilei Xu, Xuezheng Liu, Xiaoge Wang, Zheng Zhang 0001 |
PLDI | 6 |
| 2008 | Robust incentives via multi-level Tit-for-TatabstractAbstract Much work has been done to address the need for incentive models in real deployed peer‐to‐peer networks. In this paper, we discuss problems found with the incentive model in a large, deployed peer‐to‐peer network, Maze. We evaluate several alternatives, and propose an incentive system that generates preferences for well‐behaved nodes while correctly punishing colluders. We discuss our proposal as a hybrid between Tit‐for‐Tat and EigenTrust, and show its effectiveness through simulation of real traces of the Maze system. Copyright © 2007 John Wiley & Sons, Ltd. Qiao Lian, Mao Yang 0004, Zheng Zhang 0001, Yafei Dai, Xiaoming Li 0001 |
Concurr. Comput. Pract. Exp. | 4 |
| 2008 | Evaluation and optimization of a peer-to-peer video-on-demand system
Bin Cheng 0001, Xiuzheng Liu, Zheng Zhang 0001, Hai Jin 0001, Lex Stein, Xiaofei Liao |
J. Syst. Archit. | 3 |
| 2008 | GridCast: Improving peer sharing for P2P VoDabstractVideo-on-Demand (VoD) is a compelling application, but costly. VoD is costly due to the load it places on video source servers. Many have proposed using peer-to-peer (P2P) techniques to shift load from servers to peers. Yet, nobody has implemented and deployed a system to openly and systematically evaluate how these techniques work. This article describes the design, implementation and evaluation of GridCast, a real deployed P2P VoD system. GridCast has been live on CERNET since May of 2006. It provides seek, pause, and play operations, and employs peer sharing to improve system scalability. In peak months, GridCast has served videos to 23,000 unique users. From the first deployment, we have gathered information to understand the system and evaluate how to further improve peer sharing through caching and replication. We first show that GridCast with single video caching (SVC) can decrease load on source servers by an average of 22% from a client-server architecture. We analyze the net effect on system resources and determine that peer upload is largely idle. This leads us to changing the caching algorithm to cache multiple videos (MVC). MVC decreases source load by an average of 51% over the client-server. The improvement is greater as user load increases. This bodes well for peer-assistance at larger scales. A detailed analysis of MVC shows that departure misses become a major issue in a P2P VoD system with caching optimization. Motivated by this observation, we examine how to use replication to eliminate departure misses and further reduce server load. A framework for lazy replication is presented and evaluated in this article. In this framework, two predictors are plugged in to create the working replication algorithm. With these two simple predictors, lazy replication can decrease server load by 15% from MVC with only a minor increase in network traffic. Bin Cheng 0001, Lex Stein, Hai Jin 0001, Xiaofei Liao, Zheng Zhang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2007 | An Empirical Study of Collusion Behavior in the Maze P2P File-Sharing SystemabstractPeer-to-peer networks often use incentive policies to encourage cooperation between nodes. Such systems are generally susceptible to collusion by groups of users in order to gain unfair advantages over others. While techniques have been proposed to combat Web spam collusion, there are few measurements of real collusion in deployed systems. In this paper, we report analysis and measurement results of user collusion in Maze, a large-scale peer-to-peer file sharing system with a non-net-zero point-based incentive policy. We search for colluding behavior by examining complete user logs, and incrementally refine a set of collusion detectors to identify common collusion patterns. We find collusion patterns similar to those found in Web spamming. We evaluate how proposed reputation systems would perform on the Maze system. Our results can help guide the design of more robust incentive schemes. Qiao Lian, Zheng Zhang 0001, Mao Yang 0004, Ben Y. Zhao, Yafei Dai, Xiaoming Li 0001 |
ICDCS | 2 |
| 2007 | Machine Bank: Own Your Virtual Personal ComputerabstractIn this paper, we report the design, implementation and experimental results of Machine Bank, a system engineered towards the popular shared-lab scenario, where users outnumber available PCs and may get different PCs in different sessions. Machine Bank allows users to preserve their entire working environment across sessions. Each client runs virtual machine, which is saved to and reinstantiated from a content-addressable backend storage. We carefully designed lightweight hooks at client side that implements caching and tracking logics to improve reinstantiation speed as well as to remove unnecessary network and disk traffic. Our detailed evaluation demonstrates that these techniques are effective, and the overall performance fits well with the shared-lab usage. Zheng Zhang 0001 |
IPDPS | 3 |
| 2007 | WiDS Checker: Combating Bugs in Distributed Systems
Xuezheng Liu, Aimin Pan, Zheng Zhang 0001 |
NSDI | 4 |
| 2007 | An Analytical Framework and Its Applications for Studying Brick Storage ReliabilityabstractThe reliability of a large-scale storage system is influenced by a complex set of inter-dependent factors. This paper presents a comprehensive and extensible analytical framework that offers quantitative answers to many design tradeoffs. We apply the framework to a number of important design strategies that a designer and/or administrator must face in reality, including topology-aware replica placement, proactive replication that uses small background network bandwidth and unused disk space to create additional copies. We also quantify the impact of slow (but potentially more accurate) failure detection and lazy replacement of failed disks. We use detailed simulation to verify and refine our analytical model. These results demonstrate the versatility of the framework and serve as a solid step towards more quantitative studies of fundamental system tradeoffs between reliability, performance, and cost in large-scale distributed storage systems. Ming Chen 0004, Wei Chen 0013, Likun Liu, Zheng Zhang 0001 |
SRDS | 4 |
| 2006 | Automated known problem diagnosis with event tracesabstractComputer problem diagnosis remains a serious challenge to users and support professionals. Traditional troubleshooting methods relying heavily on human intervention make the process inefficient and the results inaccurate even for solved problems, which contribute significantly to user's dissatisfaction. We propose to use system behavior information such as system event traces to build correlations with solved problems, instead of using only vague text descriptions as in existing practices. The goal is to enable automatic identification of the root cause of a problem if it is a known one, which would further lead to its resolution. By applying statistical learning techniques to classifying system call sequences, we show our approach can achieve considerable accuracy of root cause recognition by studying four case examples. Chun Yuan 0004, Ni Lao, Ji-Rong Wen, Zheng Zhang 0001, Yi-Min Wang, Wei-Ying Ma |
EuroSys | 5 |
| 2005 | WiDS: An Integrated Toolkit for Distributed System Development
Shiding Lin, Aimin Pan, Zheng Zhang 0001 |
HotOS | 3 |
| 2005 | On the Impact of Replica Placement to the Reliability of Distributed Brick Storage SystemsabstractData reliability of distributed brick storage systems critically depends on the replica placement policy, and the two governing forces are repair speed and sensitivity to multiple concurrent failures. In this paper, we provide an analytical framework to reason and quantify the impact of replica placement policy to system reliability. The novelty of the framework is its consideration of the bounded network bandwidth for data maintenance. We apply the framework to two popular schemes, namely sequential placement and random placement, and show that both have drawbacks that significantly degrade data reliability. We then propose the stripe placement scheme and find the near-optimal configuration parameter such that it provides much better reliability. We further discuss the possibility of addressing the problem of correlated brick failures in our analytical framework. Qiao Lian, Wei Chen 0013, Zheng Zhang 0001 |
ICDCS | 3 |
| 2005 | Z-Ring: Fast Prefix Routing via a Low Maintenance Membership ProtocolabstractIn this paper, we introduce Z-ring, a fast prefix routing protocol for peer-to-peer overlay networks. Z-ring incorporates cost-efficient membership protocol to achieve fast routing with small maintenance cost. Z-ring achieves routing in logGN steps, where N is the network size and G is the size of a group that can be maintained by a membership protocol with low cost. With G=4096, it translates to one-hop routing for intranet environments (N<4096), two-hop routing for mid-scale internet applications (N<16 million), and three-hop routing for ultra-large Internet applications (N<64 billion). Z-ring maintains good routing success rate under churn and low maintenance cost even at large network size. Its modularized use of the membership protocol also makes it adaptive to dynamic and wide-range network size changes. Qiao Lian, Wei Chen 0013, Zheng Zhang 0001, Shaomei Wu, Ben Y. Zhao |
ICNP | 3 |
| 2005 | Simulating Large-Scale P2P Systems with the WiDS ToolkitabstractCurrent simulation technologies support at most hundreds of thousands of nodes, and fall short on the emerging large-scale networking systems that usually involve millions of nodes. We meet this challenge with our distributed simulation engine that is able to run millions of instances and is tested with a production P2P protocol, using commodity PC clusters. This simulation engine is part of the WiDS toolkit, which takes a holistic approach to the research and development of distributed systems. We also propose a critical optimization, called slow message relaxation (SMR), to trade simulation accuracy for performance. By taking advantage of the fact that distributed protocols are resilient to network fluctuation, SMR executes events in a logical time window much wider than the conventional look ahead scheme allows. We analyze and bound the potential effect of the distortion on application logic and other general metrics. Our experiments demonstrate that the simulation engine is able to achieve order of a magnitude speedup with statistically accurate simulation results. Shiding Lin, Aimin Pan, Zheng Zhang 0001 |
MASCOTS | 4 |
| 2005 | Sigma: A Fault-Tolerant Mutual Exclusion Algorithm in Dynamic Distributed Systems Subject to Process Crashes and Memory LossesabstractThis paper introduces the Sigma algorithm that solves fault-tolerant mutual exclusion problem in dynamic systems where the set of processes may be large and change dynamically, processes may crash, and the recovery or replacement of crashed processes may lose all state information (memory losses). Sigma algorithm includes new messaging mechanisms to tolerate process crashes and memory losses. It does not require any extra cost for process recovery. The paper also shows that the threshold used by the Sigma algorithm is necessary for systems with process crashes and memory losses. Wei Chen 0013, Shiding Lin, Qiao Lian, Zheng Zhang 0001 |
PRDC | 4 |
| 2005 | Using model checker and replay facility to debug complex distributed systemabstractA correct system is only derived from a correct implementation of a correct specification. Unfortunately, this imposes a heavy burden in the development process, especially for complex, distributed system ranging from machine room computing and storage services as well as large-scale P2P applications. A specification, if authored in formal language such as TLA+, Spec#, SPIN etc., is ready for model checking. The state explosion problem, however, prohibits all specification states to be thoroughly traversed. Often ad hoc heuristics are applied to drastically reduce the scale so as to make the model checking phase tractable. A correct implementation can be even more challenging, especially when we encounter non-deterministic bugs that are hard to reproduce. The gap between spec and implementation often leaves one to wonder whether the implementation or the spec is faulty, or even both. Motivated by our experiences in developing several complete large scale distributed systems, we are designing and implementing a suite of testing and debugging facility on top of our previously developed WiDS platform. Xuezheng Liu, Aimin Pan, Zheng Zhang 0001 |
SOSP | 4 |
| 2004 | Strider: a black-box, state-based approach to change and configuration management and support
Yi-Min Wang, Chad Verbowski, John Dunagan, Helen J. Wang, Chun Yuan 0004, Zheng Zhang 0001 |
Sci. Comput. Program. | 7 |
| 2004 | Evaluation of Edge Caching/Offloading for Dynamic Content DeliveryabstractAs dynamic content becomes increasingly dominant, it becomes an important research topic as how the edge resources such as client-side proxies, which are otherwise underutilized for such content, can be put into use. However, it is unclear what will be the best strategy, and the design/deployment trade offs lie therein. In this paper, using one representative e-commerce benchmark, we report our experience of an extensive investigation of different offloading and caching options. Our results point out that, while great benefits can be reached in general, advanced offloading strategies can be overly complex and even counterproductive. In contrast, simple augmentation at proxies to enable fragment caching and page composition achieves most of the benefit without compromising important considerations such as security. We also present proxy+ architecture which supports such capabilities for existing Web applications with minimal reengineering effort. Chun Yuan 0004, Zheng Zhang 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2003 | Building Topology-Aware Overlays Using Global Soft-StatabstractDistributed hash table (DHT) based overlay networks offer an administration-free and fault-tolerant storage space that maps "keys" to "values". For these systems to function efficiently, their structures must fit that of the underlying network. Existing techniques for discovering network proximity information, such as landmark clustering and expanding-ring search are either inaccurate or expensive. The lack of global proximity information in overlay construction and maintenance can result in bad proximity approximation or excessive communication. To address these problems, we propose the following: (1) Combining landmark clustering and round-trip time (RTT) measurements to generate proximity information, achieving both efficiency and accuracy. (2) Controlled placement of global proximity information on the system itself as soft-state, such that nodes can independently access relevant information efficiently. (3) Publish/subscribe functionality that allows nodes to subscribe to the relevant soft-state and get notified as the state changes necessitate overlay restructuring. Zhichen Xu, Chunqiang Tang, Zheng Zhang 0001 |
ICDCS | 3 |
| 2003 | Trading Replication Consistency for Performance and Availability: an Adaptive ApproachabstractReplication system is one of the most fundamental building blocks of wide-area applications. Due to the inevitable dependencies on wide-area communication, trade-off between performance, availability and replication consistency is often a necessity. While a number of proposals have been made to provide a tunable consistency bound between strong and weak extremes, many of them rely on a statically specified enforcement across replicas. This approach, while easy to implement, neglects the dynamic contexts within which replicas are operating, delivering sub-optimal performance and/or system availability. In this paper we analyze the problem of optimal performance/availability for a given consistency level under heterogeneous workload and network condition. We prove several optimization rules for different goals. Based on these results, we developed an adaptive update window protocol in which consistency enforcement across replicas is self-tuned to achieve optimal performance/availability. A prototype system, FRACS, is built and evaluated in this paper. The experiment results demonstrate significant advantages of adaptation over static approach for a variety of workloads. Chi Zhang 0070, Zheng Zhang 0001 |
ICDCS | 2 |
| 2003 | STRIDER: A Black-box, State-based Approach to Change and Configuration Management and Support
Yi-Min Wang, Chad Verbowski, John Dunagan, Helen J. Wang, Chun Yuan 0004, Zheng Zhang 0001 |
LISA | 7 |
| 2003 | Evaluation of edge caching/offloading for dynamic content deliveryabstractAbstract—As dynamic content becomes increasingly dominant, it becomes an important research topic as how the edge resources such as client-side proxies, which are otherwise underutilized for such content, can be put into use. However, it is unclear what will be the best strategy, and the design/deployment trade offs lie therein. In this paper, using one representative e-commerce benchmark, we report our experience of an extensive investigation of different offloading and caching options. Our results point out that, while great benefits can be reached in general, advanced offloading strategies can be overly complex and even counterproductive. In contrast, simple augmentation at proxies to enable fragment caching and page composition achieves most of the benefit without compromising important considerations such as security. We also present Proxy+ architecture which supports such capabilities for existing Web applications with minimal reengineering effort. Index Terms—Edge caching, offloading, dynamic content, fragment caching, page composition. 1 Chun Yuan 0004, Zheng Zhang 0001 |
WWW | 3 |
| 2002 | ReVive: Cost-Effective Architectural Support for Rollback Recovery in Shared-Memory MultiprocessorsabstractThis paper presents ReVive, a novel general-purpose rollback recovery mechanism for shared-memory multiprocessors. ReVive carefully balances the conflicting requirements of availability, performance, and hardware cost. ReVive performs checkpointing, logging, and distributed parity protection, all memory-based. It enables recovery from a wide class of errors, including the permanent loss of an entire node. To maintain high performance, ReVive includes specialized hardware that performs frequent operations in the background, such as log and parity updates. To keep the cost low, more complex checkpointing and recovery functions are performed in software, while the hardware modifications are limited to the directory controllers of the machine. Our simulation results on a 16-processor system indicate that the average error-free execution time overhead of using ReVive is only 6.3%, while the achieved availability is better than 99.999% even when the errors occur as often as once per day. Milos Prvulovic, Josep Torrellas, Zheng Zhang 0001 |
ISCA | 3 |
| 2002 | Reperasure: Replication Protocol Using Erasure-Code in Peer-to-Peer Storage NetworkabstractPeer-to-peer overlay networks offer a convenient way to host an infrastructure that can scale to the size of the Internet and yet stay manageable. These overlays are essentially self-organizing distributed hash tables (DHT). The dynamic nature of the system, however, poses serious challenges of data reliability. Furthermore, in order to see wider adoption, it is time to design support for generic replication mechanisms capable of handling arbitrary update requests - most of the existing proposals are deep archival systems in nature. Utilizing the fact that DHT can function as a super-reliable and high performance disk when data stored inside are erasure coded, we believe practical and simple protocols can be designed. In this paper, we introduce the reperasure protocol, a layer on top of the basic DHT, which efficiently supports strong consistency semantic with high availability guarantee. By relieving the DHT layer out of replication duo, inside, a cleaner overall architecture can be derived because of the clear division of responsibility. Zheng Zhang 0001, Qiao Lian |
SRDS | 1 |
| 2001 | Designing a Robust Namespace for Distributed File ServicesabstractA number of ongoing research projects follow a partition-based approach to provide highly scalable distributed storage services. These systems maintain namespaces that reference objects distributed across multiple locations in the system. Typically, atomic commitment protocols, such as 2-phase commit, are used for updating the namespace, in order to guarantee its consistency even in the presence of failures. Atomic commitment protocols are known to impose a high overhead to failure-free execution. Furthermore, they use conservative recovery procedures and may considerably restrict the concurrency of overlapping operations in the system. This paper proposes a set of new protocols implementing the fundamental operations in a distributed namespace. The protocols impose a minimal overhead to failure-free execution. They are robust against both communication and host failures, and use aggressive recovery procedures to re-execute incomplete operations. The proposed protocols are compared with their 2-phase commit counterparts and are shown to outperform them in all critical performance factors: communication round-trips, synchronous I/O, operation concurrency. Zheng Zhang 0001, Christos T. Karamanolis |
SRDS | 1 |
| 1999 | Excel-NUMA: Toward Programmability, Simplicity, and High PerformanceabstractWhile hardware-coherent scalable shared-memory multiprocessors are relatively easy to program, they still require substantial programming effort to deliver high performance. Specifically, to minimize remote accesses, data must be carefully laid out in memory for locality and application working sets carefully tuned for caches. It has been claimed that this programming effort is less necessary in hardware COMA machines like Flat-COMA thanks to automatic line-based data migration. Unfortunately, Flat-COMA is complex to design. Consequently, we would like a machine as programmable as Flat-COMA, as simple as plain CC-NUMA, and that outperforms both. This paper presents our proposal: Excel-NUMA (EX-NUMA). The idea is to exploit the fact that, after a memory line is written and cached, the storage that kept the line in memory is unutilized. We use that storage to temporarily hold remote data displaced from the local caches. This enables automatic data migration, like in Flat-COMA, enhancing programmability. The hardware required to manage the system is a simple, local module added to a CC-NUMA; the global cache coherence protocol is not changed. Simulations of Splash2 applications show that EX-NUMA outperforms CC-NUMA and Flat-COMA in every single application and eliminates most of the conflict misses. Zheng Zhang 0001, Marcelo H. Cintra, Josep Torrellas |
IEEE Trans. Computers | 1 |
| 1997 | The Memory Performance of DSS Commercial Workloads in Shared-Memory MultiprocessorsabstractAlthough cache-coherent shared-memory multiprocessors are often used to run commercial workloads, little work has been done to characterize how well these machines support such workloads. In particular, we do not have much insight into the demands of commercial workloads on the memory subsystem of these machines. In this paper, we analyze in detail the memory access patterns of several queries that are representative of Decision Support System (DSS) databases. Our analysis shows that the memory use of queries differs largely depending on how the queries access the database data, namely via indices or by sequentially scanning the records. The former queries, which we call Index queries, suffer most of their shared-data misses on indices and on lock-related metadata structures. The latter queries, which we call Sequential queries, suffer most of their shared-data misses on the database records as they are scanned. An analysis of the data locality in the queries shows that both Index and Sequential queries exhibit spatial locality and, therefore, can benefit from relatively long cache lines. Interestingly, shared data is reused very little inside queries. However, there is data reuse across Sequential queries. Finally, we show that the performance of Sequential queries can be improved moderately with data prefetching. Pedro Trancoso, Josep Lluís Larriba-Pey, Zheng Zhang 0001, Josep Torrellas |
HPCA | 3 |
| 1997 | Reducing Remote Conflict Misses: NUMA with Remote Cache versus COMAabstractMany future applications for scalable shared-memory multiprocessors are likely to have large working sets that overflow secondary or tertiary caches. Two possible solutions to this problem are to add a very large cache called remote cache that caches remote data (NUMA-RC), or organize the machine as a cache-only memory architecture (COMA). This paper tries to determine which solution is best. To compare the performance of the two organizations for the same amount of total memory, we introduce a model of data sharing. The model uses three data sharing patterns: replication, read-mostly migration, and read-writs migration. Replication data is accessed in read-mostly mode by several processors, while migration data is accessed largely by one processor at a time. For large working sets, the weight of the migration data largely determines whether COMA outperforms NUMA-RC. Ideally, COMA only needs to fit the replication data in its extra memory; the migration data will simply be swapped between attraction memories. The remote cache of NUMA-RC, instead, needs to house both the replication and the migration data. However, simulations of seven Splash2 applications show that COMA does not outperform NUMA-RC. This is due to two reasons. First, the extra memory added has more associativity in NUMA-RC than in COMA and, therefore, can be utilized better by the working set in NUMA-RC. Second, COMA memory accesses are more expensive. Of course, our results are affected by the applications used, which have been optimized for a cache-coherent NUMA machine. Overall, since NUMA-RC is cheaper, NUMA-RC is more cost-effective for these applications. Zheng Zhang 0001, Josep Torrellas |
HPCA | 1 |
| 1997 | The Performance of the Cedar Multistage Switching NetworkabstractWhile multistage switching networks for vector multiprocessors have been studied extensively, detailed evaluations of their performance are rare. Indeed, analytical models, simulations with pseudosynthetic loads, studies focused on average-value parameters, and measurements of networks disconnected from the machine, all provide limited information. In this paper, instead, we present an in-depth empirical analysis of a multistage switching network in a realistic setting: We use hardware probes to examine the performance of the omega network of the Cedar shared-memory machine executing real applications. The machine is configured with 16 vector processors. The analysis suggests that the performance of multistage switching networks is limited by traffic nonuniformities. We identify two major nonuniformities that degrade Cedar's performance and are likely to slow down other networks too. The first one is the contention caused by the return messages in a vector access as they converge from the memories to one processor port. This traffic convergence penalizes vector reads and, more importantly, causes tree saturation. The second nonuniformity is the uneven contention delays induced by a relatively fair scheme to resolve message collisions. Based on our observations, we argue that intuitive optimizations for multistage switching networks may not be the most cost-effective ones. Instead, we suggest changes to increase the network bandwidth at the root of the traffic convergence tree and to delay traffic convergence up until the final stages of the network. Josep Torrellas, Zheng Zhang 0001 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 1996 | Distance-Adaptive Update Protocols for Scalable Shared-Memory MultiprocessorsabstractWhile update protocols generally induce lower miss rates than invalidate protocols, they tend to generate much traffic. This is one of the reasons why they are considered less cost-effectively scalable than invalidate protocols and, as a result, are avoided in most existing designs of scalable shared-memory multiprocessors. However, given the increasing relative cost of cache misses, update protocols are becoming more worthy of exploration. In this paper, we present a model of sharing that is key to investigating the performance of optimized update protocols: the update distance model. The model gives insight into the update patterns that optimized protocols need to handle. Using this model, we design a new family of protocols that we call distance-adaptive protocols. In these schemes, the directory records the update patterns observed and then uses them to selectively send updates and invalidations to processors. As a result, traffic and miss rates are kept low. We present an implementation of these protocols based on a dynamic pointer scheme. A performance comparison between one of these protocols and efficient invalidate and delayed competitive-update protocols over five applications shows that the new protocol decreases the execution time by an average of 15% and 10% respectively. Alain Raynaud, Zheng Zhang 0001, Josep Torrellas |
HPCA | 2 |
| 1995 | Speeding Up Irregular Applications in Shared-Memory Multiprocessors: Memory Binding and Group PrefetchingabstractWhile many parallel applications exhibit good spatial locality, other important codes in areas like graph problem-solving or CAD do not. Often, these irregular codes contain small records accessed via pointers. Consequently, while the former applications benefit from long cache lines, the latter prefer short lines. One good solution is to combine short lines with prefetching. In this way, each application can exploit the amount of spatial locality that it has. However, prefetching, if provided, should also work for the irregular codes. This paper presents a new prefetching scheme that, while usable by regular applications, is specifically targeted to irregular ones: Memory Binding and Group Prefetching.The idea is to hardware-bind and prefetch together groups of data that the programmer suggests are strongly related to each other. Examples are the different fields in a record or two records linked by a permanent pointer. This prefetching scheme, combined with short cache lines, results in a memory hierarchy design that can be exploited by both regular and irregular applications. Overall, it is better to use a system with short lines (16-32 bytes) and our prefetching than a system with long lines (128 bytes) with or without our prefetching. The former system runs 6 out of 7 Splash-class applications faster. In particular, some of the most irregular applications run 25-40% faster. Zheng Zhang 0001, Josep Torrellas |
ISCA | 1 |
| 1994 | The performance of the Cedar multistage switching networkabstractWhile multistage switching networks for vector multiprocessors have been studied extensively, detailed evaluations of their performance are rare. Indeed, analytical models, simulations with pseudo-synthetic loads, studies focused on average-value parameters, and measurements of networks disconnected from the machine all provide limited information. In this paper, instead, we present an in-depth empirical analysis of a multistage switching network in a realistic setting: we use hardware probes to examine the performance of the omega network of the Cedar shared-memory machine executing real applications. The machine is configured with 16 vector processors. The analysis suggests that the performance of multistage switching networks is limited by traffic non-uniformities. We identify two major non-uniformities that degrade Cedar's performance and are likely to slow down other networks too. The first one is the contention caused by the return messages in a vector access as they converge from the memories to one processor port. This traffic convergence penalizes vector reads and, more importantly, causes tree saturation. The second non-uniformity is the uneven contention delays induced by even a relatively fair scheme to resolve message collisions. Based on our observations, we argue that intuitive optimizations for multistage switching networks may not be cost-effective. Instead, we suggest changes to increase the network bandwidth at the root of the traffic convergence tree and to delay traffic convergence up until the final stages of the network.> Josep Torrellas, Zheng Zhang 0001 |
SC | 2 |