VLDB 2026 Research / reviewers in the wild / expert
Jiaxin Shi
dblp:151/7509
· DBLP profile ↗
62ranked-venue papers
16as first author
33since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 46 · 15 first-author · 29 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 4 first-author · 5 since 2021Systems, architecture and hardware · 5 · 1 since 2021Software engineering, systems software and programming languages · 2 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Neural Eigenfunctions are Structured Representation LearnersabstractThis paper revisits the canonical concept of learning structured representations without label supervision by eigendecomposition. Yet, unlike prior spectral methods such as Laplacian Eigenmap which operate in a nonparametric manner, we aim to parametrically model the principal eigenfunctions of an integral operator defined by a kernel and a data distribution using a neural network for enhanced scalability and reasonable out-of-sample generalization. To achieve this goal, we first present a new series of objective functions that generalize the EigenGame Gemp et al. 2020 to function space for learning neural eigenfunctions. We then show that, when the similarity metric is derived from positive relations in a data augmentation setup, a representation learning objective function that resembles those of popular self-supervised learning methods emerges, with an additional symmetry-breaking property for producing structured representations where features are ordered by importance. We call such a structured, adaptive-length deep representation Neural Eigenmap. We demonstrate using Neural Eigenmap as adaptive-length codes in image retrieval systems. By truncation according to feature importance, our method requires up to $16\times$16× shorter representation length than leading self-supervised learning ones to achieve similar retrieval performance. We further apply our method to graph data and report strong results on a node representation learning benchmark with more than one million nodes. Zhijie Deng, Jiaxin Shi, Hao Zhang 0025, Peng Cui 0007, Cewu Lu, Jun Zhu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2026 | LoRA of Change: Learning to Generate LoRA for the Editing Instruction From a Single Before-After Image PairabstractIn this paper, we propose the LoRA of Change (LoC) framework for image editing with visual instructions, i.e., before-after image pairs. Compared to the ambiguities, insufficient specificity, and diverse interpretations of natural language, visual instructions can accurately reflect users' intent. Building on the success of LoRA in text-based image editing and generation, we dynamically learn an instruction-specific LoRA to encode the "change" in a before-after image pair, enhancing the interpretability and reusability of our model. Furthermore, generalizable models for image editing with visual instructions typically require quad data, i.e., a before-after image pair, along with query and target images. Due to the scarcity of such quad data, existing models are limited to a narrow range of visual instructions. To overcome this limitation, we introduce the LoRA Reverse optimization technique, enabling large-scale training with paired data alone. Extensive qualitative and quantitative experiments demonstrate that our model produces high-quality images that align with user intent and support a broad spectrum of real-world visual instructions. Jiequan Cui, Hanwang Zhang, Jiaxin Shi, Jingjing Chen 0001, Chi Zhang 0007, Yu-Gang Jiang 0001 |
IEEE Trans. Image Process. | 4 |
| 2025 | Incremental Transformer: Efficient Encoder for Incremented Text Over MRC and Conversation TasksabstractSome encoder inputs such as conversation histories are frequently extended with short additional inputs like new responses. However, to obtain the real-time encoding of the extended input, existing Transformer-based encoders like BERT have to encode the whole extended input again without utilizing the existing encoding of the original input, which may be prohibitively slow for real-time applications. In this paper, we introduce Incremental Transformer, an efficient encoder dedicated for faster encoding of incremented input. It takes only added input as input but attends to cached representations of original input in lower layers for better performance. By treating questions as additional inputs of a passage, Incremental Transformer can also be applied to accelerate MRC tasks. Experimental results show tiny decline in effectiveness but significant speedup against traditional full encoder across various MRC and multi-turn conversational question answering tasks. With the help from simple distillation-like auxiliary losses, Incremental Transformer achieves a speedup of 6.2x, with a mere 2.2 point accuracy reduction in comparison to RoBERTa-Large on SQuADV1.1. Yuechen Wang, Jiaxin Shi, Wengang Zhou 0001, Qi Tian 0001, Houqiang Li |
COLING | 3 |
| 2025 | Chain of Semantics Programming in 3D Gaussian Splatting Representation for 3D Vision Groundingabstract3D Vision Grounding (3DVG) is a fundamental research area that enables agents to perceive and interact with the 3D world. The challenge of the 3DVG task lies in understanding fine-grained semantics and spatial relationships within both the utterance and 3D scene. To address this challenge, we propose a zero-shot neuro-symbolic framework that utilizes a large language model (LLM) as neuro-symbolic functions to ground the object within the 3D Gaussian Splatting (3DGS) representation. By utilizing 3DGS representation, we can dynamically render high-quality 2D images from various viewpoints to enrich the semantic information. Given the complexity of spatial relationships, we construct a relationship graph and chain of semantics that decouple spatial relationships and facilitate step-by-step reasoning within 3DGS representation. Additionally, we employ a grounded-aware self-check mechanism to enable the LLM to reflect on its responses and mitigate the effects of ambiguity in spatial reasoning. We evaluate our method using two publicly available datasets, Nr3D and Sr3D, achieving accuracies of 60.8% and 91.4%, respectively. Notably, our method surpasses current state-of-the-art zero-shot methods on the Nr3D dataset. In addition, it outperforms the recent supervised models on the Sr3D dataset. Jiaxin Shi, Mingyue Xiang, Zhi Weng |
CVPR | 1 |
| 2025 | Learning-Order Autoregressive Models with Application to Molecular Graph GenerationabstractAutoregressive models (ARMs) have become the workhorse for sequence generation tasks, since many problems can be modeled as next-token prediction. While there appears to be a natural ordering for text (i.e., left-to-right), for many data types, such as graphs, the canonical ordering is less obvious. To address this problem, we introduce a variant of ARM that generates high-dimensional data using a probabilistic ordering that is sequentially inferred from data. This model incorporates a trainable probability distribution, referred to as an order-policy, that dynamically decides the autoregressive order in a state-dependent manner. To train the model, we introduce a variational lower bound on the exact log-likelihood, which we optimize with stochastic gradient estimation. We demonstrate experimentally that our method can learn meaningful autoregressive orderings in image and graph generation. On the challenging domain of molecular graph generation, we achieve state-of-the-art results on the QM9 and ZINC250k benchmarks, evaluated using the Fréchet ChemNet Distance (FCD), Synthetic Accessibility Score (SAS), Quantitative Estimate of Drug-likeness (QED). Zhe Wang 0055, Jiaxin Shi, Nicolas Heess, Arthur Gretton, Michalis K. Titsias |
ICML | 2 |
| 2025 | Ca2-VDM: Efficient Autoregressive Video Diffusion Model with Causal Generation and Cache SharingabstractWith the advance of diffusion models, today's video generation has achieved impressive quality. To extend the generation length and facilitate real-world applications, a majority of video diffusion models (VDMs) generate videos in an autoregressive manner, i.e., generating subsequent clips conditioned on the last frame(s) of the previous clip. However, existing autoregressive VDMs are highly inefficient and redundant: The model must re-compute all the conditional frames that are overlapped between adjacent clips. This issue is exacerbated when the conditional frames are extended autoregressively to provide the model with long-term context. In such cases, the computational demands increase significantly (i.e., with a quadratic complexity w.r.t. the autoregression step). In this paper, we propose **Ca2-VDM**, an efficient autoregressive VDM with **Ca**usal generation and **Ca**che sharing. For **causal generation**, it introduces unidirectional feature computation, which ensures that the cache of conditional frames can be precomputed in previous autoregression steps and reused in every subsequent step, eliminating redundant computations. For **cache sharing**, it shares the cache across all denoising steps to avoid the huge cache storage cost. Extensive experiments demonstrated that our Ca2-VDM achieves state-of-the-art quantitative and qualitative video generation results and significantly improves the generation speed. Code is available: https://github.com/Dawn-LX/CausalCache-VDM Kaifeng Gao, Jiaxin Shi, Hanwang Zhang, Chunping Wang 0001, Jun Xiao 0001, Long Chen 0016 |
ICML | 2 |
| 2025 | Slide-CLIP: A Simple and Effective Pruning Method for CLIPabstractPre-trained vision-language models have achieved great zero-shot performance in various downstream tasks. With the rapid development of vision-language models, many task-specific Contrastive Language-Image Pre-training (CLIP) models are proposed and utilized in diverse domains. However, their large model size hinders their utilization on platforms with limited hardware resources. Recently, several CLIP pruning methods have been proposed, we find they are effective but resources-consuming, costing a great amount of GPU hours at pruning stage. At the same time, we notice the emergence of fast large language model (LLM) pruning method such as SparseGPT which is simple and effective. All these lead to the question: Does there exist a simple but effective pruning method for CLIP? In this paper, we propose Slide-CLIP as an answer, which incorporates (i) applying a fast pruning method such as SparseGPT to obtain a specific layer with desired sparsity; (ii) applying Layer-wise Sliding Distillation (LSD) at the next layer of the pruned layer to minimize the mean squared error (MSE) of the next layer output before pruning and after pruning; (iii) iterating the pruning routine of (i) and (ii) in a slide manner until the last layer of model is reached; and (iv) applying a progressive process on top of (iii) at sparsity above 50% to minimize the performance drop. Extensive experiments on various CLIP models demonstrate the effectiveness of the proposed Slide-CLIP pruning method. The code is publicly available at https://github.com/bill426/Slide-CLIP Jiaxin Shi, Shaohua Kevin Zhou |
IJCNN | 1 |
| 2025 | Generating Creative Chess PuzzlesabstractWhile Generative AI rapidly advances in various domains, generating truly creative, aesthetic, and counter-intuitive outputs remains a challenge. This paper presents an approach to tackle these difficulties in the domain of chess puzzles. We start by benchmarking Generative AI architectures, and then introduce an RL framework with novel rewards based on chess engine search statistics to overcome some of those shortcomings. The rewards are designed to enhance a puzzle's uniqueness, counter-intuitiveness, diversity, and realism. Our RL approach dramatically increases counter-intuitive puzzle generation by 10x, from 0.22\% (supervised) to 2.5\%, surpassing existing dataset rates (2.1\%) and the best Lichess-trained model (0.4\%). Our puzzles meet novelty and diversity benchmarks, retain aesthetic themes, and are rated by human experts as more creative, enjoyable, and counter-intuitive than composed book puzzles, even approaching classic compositions. Our final outcome is a curated booklet of these novel AI-generated puzzles, which is acknowledged for creativity by three world-renowned experts. Xidong Feng, Vivek Veeriah, Marcus Chiam, Michael Dennis 0001, Federico Barbero, Johan S. Obando-Ceron, Jiaxin Shi, Satinder Singh 0001, Shaobo Hou, Nenad Tomasev, Tom Zahavy |
NeurIPS | 7 |
| 2025 | Informed Correctors for Discrete Diffusion ModelsabstractDiscrete diffusion has emerged as a powerful framework for generative modeling in discrete domains, yet efficiently sampling from these models remains challenging. Existing sampling strategies often struggle to balance computation and sample quality when the number of sampling steps is reduced, even when the model has learned the data distribution well. To address these limitations, we propose a predictor-corrector sampling scheme where the corrector is informed by the diffusion model to more reliably counter the accumulating approximation errors. To further enhance the effectiveness of our informed corrector, we introduce complementary architectural modifications based on hollow transformers and a simple tailored training objective that leverages more training signal. We use a synthetic example to illustrate the failure modes of existing samplers and show how informed correctors alleviate these problems. On the Text8 dataset, the informed corrector improves sample quality by generating text with significantly fewer errors than the baselines. On tokenized ImageNet 256x256, this approach consistently produces superior samples with fewer steps, achieving improved FID scores for discrete diffusion models. These results underscore the potential of informed correctors for fast and high-fidelity generation using discrete diffusion. Yixiu Zhao, Jiaxin Shi, Feng Chen 0046, Shaul Druckmann, Lester Mackey, Scott W. Linderman |
NeurIPS | 2 |
| 2025 | DPUSwap: building an infinite swap with DPU for cloud computing system
Shiqiang Nie, Jianqiang Ma, Jiaxin Shi, Weiguo Wu |
J. Supercomput. | 5 |
| 2024 | Preparing Lessons for Progressive Training on Language ModelsabstractThe rapid progress of Transformers in artificial intelligence has come at the cost of increased resource consumption and greenhouse gas emissions due to growing model sizes. Prior work suggests using pretrained small models to improve training efficiency, but this approach may not be suitable for new model structures. On the other hand, training from scratch can be slow, and progressively stacking layers often fails to achieve significant acceleration. To address these challenges, we propose a novel method called Apollo, which prepares lessons for expanding operations by learning high-layer functionality during training of low layers. Our approach involves low-value-prioritized sampling (LVPS) to train different depths and weight sharing to facilitate efficient expansion. We also introduce an interpolation method for stable model depth extension. Experiments demonstrate that Apollo achieves state-of-the-art acceleration ratios, even rivaling methods using pretrained models, making it a universal and efficient solution for training deep models while reducing time, financial, and environmental costs. Yu Pan 0005, Ye Yuan 0016, Yichun Yin, Jiaxin Shi, Zenglin Xu, Ming Zhang 0004, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001 |
AAAI | 4 |
| 2024 | Leveraging Machine-Generated Rationales to Facilitate Social Meaning Detection in ConversationsabstractWe present a generalizable classification approach that leverages Large Language Models (LLMs) to facilitate the detection of implicitly encoded social meaning in conversations.We design a multi-faceted prompt to extract a textual explanation of the reasoning that connects visible cues to underlying social meanings.These extracted explanations or rationales serve as augmentations to the conversational text to facilitate dialogue understanding and transfer.Our empirical results over 2,340 experimental settings demonstrate the significant positive impact of adding these rationales.Our findings hold true for in-domain classification, zero-shot, and few-shot domain transfer for two different social meaning detection tasks, each spanning two different corpora. Ritam Dutt, Jiaxin Shi, Divyanshu Sheth, Prakhar Gupta, Carolyn P. Rosé |
ACL (1) | 3 |
| 2024 | Cross-Lingual Transfer for Natural Language Inference via Multilingual Prompt TranslatorabstractBased on multilingual pre-trained models, cross-lingual transfer with prompt learning has shown promising effectiveness, where soft prompt learned in a source language is transferred to target languages for downstream tasks, particularly in the low-resource scenario. To efficiently transfer soft prompt, we propose a novel framework, Multilingual Prompt Translator (MPT), where a multilingual prompt translator is introduced to properly process crucial knowledge embedded in prompt by changing language knowledge while retaining task knowledge. More concretely, we first train prompt in source language and employ translator to translate it into target prompt. Besides, we extend an external corpus as auxiliary data, on which an alignment task for predicted answer probability is designed to convert language knowledge, thereby equipping target prompt with multilingual knowledge. In few-shot settings on XNLI, MPT demonstrates superiority over baselines by remarkable improvements. MPT is more prominent compared with vanilla prompting when transferring to languages quite distinct from source language. Code is available at https://github.com/qiuxiaoyu9954/MPT. Xiaoyu Qiu, Yuechen Wang, Jiaxin Shi, Wengang Zhou 0001, Houqiang Li |
ICME | 3 |
| 2024 | Non-confusing Generation of Customized Concepts in Diffusion ModelsabstractWe tackle the common challenge of inter-concept visual confusion in compositional concept generation using text-guided diffusion models (TGDMs). It becomes even more pronounced in the generation of customized concepts, due to the scarcity of user-provided concept visual examples. By revisiting the two major stages leading to the success of TGDMs---1) contrastive image-language pre-training (CLIP) for text encoder that encodes visual semantics, and 2) training TGDM that decodes the textual embeddings into pixels---we point that existing customized generation methods only focus on fine-tuning the second stage while overlooking the first one. To this end, we propose a simple yet effective solution called CLIF: contrastive image-language fine-tuning. Specifically, given a few samples of customized concepts, we obtain non-confusing textual embeddings of a concept by fine-tuning CLIP via contrasting a concept and the over-segmented visual regions of other concepts. Experimental results demonstrate the effectiveness of CLIF in preventing the confusion of multi-customized concept generation. Project page: https://clif-official.github.io/clif. Jingyuan Chen 0003, Jiaxin Shi, Junzhong Miao, Tao Jin 0004, Zhou Zhao 0001, Fei Wu 0001, Shuicheng Yan, Hanwang Zhang |
ICML | 3 |
| 2024 | Action Imitation in Common Action Space for Customized Action Image SynthesisabstractWe propose a novel method, \textbf{TwinAct}, to tackle the challenge of decoupling actions and actors in order to customize the text-guided diffusion models (TGDMs) for few-shot action image generation. TwinAct addresses the limitations of existing methods that struggle to decouple actions from other semantics (e.g., the actor's appearance) due to the lack of an effective inductive bias with few exemplar images. Our approach introduces a common action space, which is a textual embedding space focused solely on actions, enabling precise customization without actor-related details. Specifically, TwinAct involves three key steps: 1) Building common action space based on a set of representative action phrases; 2) Imitating the customized action within the action space; and 3) Generating highly adaptable customized action images in diverse contexts with action similarity loss. To comprehensively evaluate TwinAct, we construct a novel benchmark, which provides sample images with various forms of actions. Extensive experiments demonstrate TwinAct's superiority in generating accurate, context-independent customized actions while maintaining the identity consistency of different subjects, including animals, humans, and even customized actors. Jingyuan Chen 0003, Jiaxin Shi, Zirun Guo, Zehan Wang 0001, Tao Jin 0004, Zhou Zhao 0001, Fei Wu 0001, Shuicheng Yan, Hanwang Zhang |
NeurIPS | 3 |
| 2024 | Simplified and Generalized Masked Diffusion for Discrete DataabstractMasked (or absorbing) diffusion is actively explored as an alternative to autoregressive models for generative modeling of discrete data. However, existing work in this area has been hindered by unnecessarily complex model formulations and unclear relationships between different perspectives, leading to suboptimal parameterization, training objectives, and ad hoc adjustments to counteract these issues. In this work, we aim to provide a simple and general framework that unlocks the full potential of masked diffusion models. We show that the continuous-time variational objective of masked diffusion models is a simple weighted integral of cross-entropy losses. Our framework also enables training generalized masked diffusion models with state-dependent masking schedules. When evaluated by perplexity, our models trained on OpenWebText surpass prior diffusion language models at GPT-2 scale and demonstrate superior performance on 4 out of 5 zero-shot language modeling tasks. Furthermore, our models vastly outperform previous discrete diffusion models on pixel-level image modeling, achieving 2.75 (CIFAR-10) and 3.40 (ImageNet 64x64) bits per dimension that are better than autoregressive models of similar sizes. Jiaxin Shi, Kehang Han, Arnaud Doucet, Michalis K. Titsias |
NeurIPS | 1 |
| 2024 | Human Motion Generation: A SurveyabstractHuman motion generation aims to generate natural human pose sequences and shows immense potential for real-world applications. Substantial progress has been made recently in motion data collection technologies and generation methods, laying the foundation for increasing interest in human motion generation. Most research within this field focuses on generating human motions based on conditional signals, such as text, audio, and scene contexts. While significant advancements have been made in recent years, the task continues to pose challenges due to the intricate nature of human motion and its implicit relationship with conditional signals. In this survey, we present a comprehensive literature review of human motion generation, which, to the best of our knowledge, is the first of its kind in this field. We begin by introducing the background of human motion and generative models, followed by an examination of representative methods for three mainstream sub-tasks: text-conditioned, audio-conditioned, and scene-conditioned human motion generation. Additionally, we provide an overview of common datasets and evaluation metrics. Lastly, we discuss open problems and outline potential future research directions. We hope that this survey could provide the community with a comprehensive glimpse of this rapidly evolving field and inspire novel ideas that address the outstanding challenges. Wentao Zhu 0004, Xiaoxuan Ma 0001, Dongwoo Ro, Hai Ci, Jinlu Zhang 0001, Jiaxin Shi, Feng Gao 0014, Qi Tian 0001, Yizhou Wang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2023 | BEST: BERT Pre-training for Sign Language Recognition with Coupling TokenizationabstractIn this work, we are dedicated to leveraging the BERT pre-training success and modeling the domain-specific statistics to fertilize the sign language recognition~(SLR) model. Considering the dominance of hand and body in sign language expression, we organize them as pose triplet units and feed them into the Transformer backbone in a frame-wise manner. Pre-training is performed via reconstructing the masked triplet unit from the corrupted input sequence, which learns the hierarchical correlation context cues among internal and external triplet units. Notably, different from the highly semantic word token in BERT, the pose unit is a low-level signal originally locating in continuous space, which prevents the direct adoption of the BERT cross entropy objective. To this end, we bridge this semantic gap via coupling tokenization of the triplet unit. It adaptively extracts the discrete pseudo label from the pose triplet unit, which represents the semantic gesture / body state. After pre-training, we fine-tune the pre-trained encoder on the downstream SLR task, jointly with the newly added task-specific layer. Extensive experiments are conducted to validate the effectiveness of our proposed method, achieving new state-of-the-art performance on all four benchmarks with a notable gain. Weichao Zhao, Hezhen Hu, Wengang Zhou 0001, Jiaxin Shi, Houqiang Li |
AAAI | 4 |
| 2023 | Reasoning over Hierarchical Question Decomposition Tree for Explainable Question AnsweringabstractJiajie Zhang, Shulin Cao, Tingjian Zhang, Xin Lv, Juanzi Li, Lei Hou, Jiaxin Shi, Qi Tian. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Shulin Cao, Tingjian Zhang, Juan-Zi Li, Lei Hou 0001, Jiaxin Shi |
ACL (1) | 7 |
| 2023 | Sequence Modeling with Multiresolution Convolutional MemoryabstractEfficiently capturing the long-range patterns in sequential data sources salient to a given task---such as classification and generative modeling---poses a fundamental challenge. Popular approaches in the space tradeoff between the memory burden of brute-force enumeration and comparison, as in transformers, the computational burden of complicated sequential dependencies, as in recurrent neural networks, or the parameter burden of convolutional networks with many or large filters. We instead take inspiration from wavelet-based multiresolution analysis to define a new building block for sequence modeling, which we call a MultiresLayer. The key component of our model is the multiresolution convolution, capturing multiscale trends in the input sequence. Our MultiresConv can be implemented with shared filters across a dilated causal convolution tree. Thus it garners the computational advantages of convolutional networks and the principled theoretical motivation of wavelet decompositions. Our MultiresLayer is straightforward to implement, requires significantly fewer parameters, and maintains at most a $O(N \log N)$ memory footprint for a length $N$ sequence. Yet, by stacking such layers, our model yields state-of-the-art performance on a number of sequence classification and autoregressive density estimation tasks using CIFAR-10, ListOps, and PTB-XL datasets. Jiaxin Shi, Ke Alexander Wang, Emily B. Fox |
ICML | 1 |
| 2023 | A Finite-Particle Convergence Rate for Stein Variational Gradient DescentabstractWe provide the first finite-particle convergence rate for Stein variational gradient descent (SVGD), a popular algorithm for approximating a probability distribution with a collection of particles. Specifically, whenever the target distribution is sub-Gaussian with a Lipschitz score, SVGD with $n$ particles and an appropriate step size sequence drives the kernel Stein discrepancy to zero at an order ${1/}{\sqrt{\log\log n}}$ rate. We suspect that the dependence on $n$ can be improved, and we hope that our explicit, non-asymptotic proof strategy will serve as a template for future refinements. Jiaxin Shi, Lester Mackey |
NeurIPS | 1 |
| 2022 | Program Transfer for Answering Complex Questions over Knowledge BasesabstractShulin Cao, Jiaxin Shi, Zijun Yao, Xin Lv, Jifan Yu, Lei Hou, Juanzi Li, Zhiyuan Liu, Jinghui Xiao. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Shulin Cao, Jiaxin Shi, Zijun Yao 0002, Jifan Yu, Lei Hou 0001, Juan-Zi Li, Zhiyuan Liu 0001, Jinghui Xiao |
ACL (1) | 2 |
| 2022 | KQA Pro: A Dataset with Explicit Compositional Programs for Complex Question Answering over Knowledge BaseabstractShulin Cao, Jiaxin Shi, Liangming Pan, Lunyiu Nie, Yutong Xiang, Lei Hou, Juanzi Li, Bin He, Hanwang Zhang. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Shulin Cao, Jiaxin Shi, Liangming Pan, Lunyiu Nie, Yutong Xiang, Lei Hou 0001, Juan-Zi Li, Hanwang Zhang |
ACL (1) | 2 |
| 2022 | Double Control Variates for Gradient Estimation in Discrete Latent Variable ModelsabstractStochastic gradient-based optimisation for discrete latent variable models is challenging due to the high variance of gradients. We introduce a variance reduction technique for score function estimators that makes use of double control variates. These control variates act on top of a main control variate, and try to further reduce the variance of the overall estimator. We develop a double control variate for the REINFORCE leave-one-out estimator using Taylor expansions. For training discrete latent variable models, such as variational autoencoders with binary latent variables, our approach adds no extra computational cost compared to standard training with the REINFORCE leave-one-out estimator. We apply our method to challenging high-dimensional toy examples and for training variational autoencoders with binary latent variables. We show that our estimator can have lower variance compared to other state-of-the-art estimators. Michalis K. Titsias, Jiaxin Shi |
AISTATS | 2 |
| 2022 | Triple-as-Node Knowledge Graph and Its Embeddings
Jiaxin Shi, Shulin Cao, Lei Hou 0001, Juan-Zi Li |
DASFAA (1) | 2 |
| 2022 | GraphQ IR: Unifying the Semantic Parsing of Graph Query Languages with One Intermediate RepresentationabstractSubject to the huge semantic gap between natural and formal languages, neural semantic parsing is typically bottlenecked by its complexity of dealing with both input semantics and output syntax.Recent works have proposed several forms of supplementary supervision but none is generalized across multiple formal languages.This paper proposes a unified intermediate representation (IR) for graph query languages, named GraphQ IR.It has a natural-language-like expression that bridges the semantic gap and formally defined syntax that maintains the graph structure.Therefore, a neural semantic parser can more precisely convert user queries into GraphQ IR, which can be later losslessly compiled into various downstream graph query languages.Extensive experiments on several benchmarks including KQA PRO, OVERNIGHT, GRAILQA and METAQA-Cypher under standard i.i.d., out-of-distribution and low-resource settings validate GraphQ IR's superiority over the previous state-of-the-arts with a maximum 11% accuracy improvement. Lunyiu Nie, Shulin Cao, Jiaxin Shi, Jiuding Sun, Lei Hou 0001, Juan-Zi Li, Jidong Zhai |
EMNLP | 3 |
| 2022 | G-MAP: General Memory-Augmented Pre-trained Language Model for Domain TasksabstractRecently, domain-specific PLMs have been proposed to boost the task performance of specific domains (e.g., biomedical and computer science) by continuing to pre-train general PLMs with domain-specific corpora.However, this Domain-Adaptive Pre-Training (DAPT; Gururangan et al. ( 2020)) tends to forget the previous general knowledge acquired by general PLMs, which leads to a catastrophic forgetting phenomenon and sub-optimal performance.To alleviate this problem, we propose a new framework of General Memory-Augmented Pre-trained Language Model (G-MAP), which augments the domain-specific PLM by a memory representation built from the frozen general PLM without losing any general knowledge.Specifically, we propose a new memory-augmented layer, and based on it, different augmented strategies are explored to build the memory representation and then adaptively fuse it into the domain-specific PLM.We demonstrate the effectiveness of G-MAP on various domains (biomedical and computer science publications, news, and reviews) and different kinds (text classification, QA, NER) of tasks, and the extensive results show that the proposed G-MAP 1 can achieve SOTA results on all tasks. Zhongwei Wan, Yichun Yin, Wei Zhang 0196, Jiaxin Shi, Lifeng Shang, Guangyong Chen, Xin Jiang 0002, Qun Liu 0001 |
EMNLP | 4 |
| 2022 | Sampling with Mirrored Stein Operators
Jiaxin Shi, Chang Liu 0030, Lester Mackey |
ICLR | 1 |
| 2022 | NeuralEF: Deconstructing Kernels by Deep Neural NetworksabstractLearning the principal eigenfunctions of an integral operator defined by a kernel and a data distribution is at the core of many machine learning problems. Traditional nonparametric solutions based on the Nystrom formula suffer from scalability issues. Recent work has resorted to a parametric approach, i.e., training neural networks to approximate the eigenfunctions. However, the existing method relies on an expensive orthogonalization step and is difficult to implement. We show that these problems can be fixed by using a new series of objective functions that generalizes the EigenGame to function space. We test our method on a variety of supervised and unsupervised learning problems and show it provides accurate approximations to the eigenfunctions of polynomial, radial basis, neural network Gaussian process, and neural tangent kernels. Finally, we demonstrate our method can scale up linearised Laplace approximation of deep neural networks to modern image classification datasets through approximating the Gauss-Newton matrix. Code is available at https://github.com/thudzj/neuraleigenfunction. Zhijie Deng, Jiaxin Shi, Jun Zhu 0001 |
ICML | 2 |
| 2022 | Gradient Estimation with Discrete Stein OperatorsabstractGradient estimation---approximating the gradient of an expectation with respect to the parameters of a distribution---is central to the solution of many machine learning problems. However, when the distribution is discrete, most common gradient estimators suffer from excessive variance. To improve the quality of gradient estimation, we introduce a variance reduction technique based on Stein operators for discrete distributions. We then use this technique to build flexible control variates for the REINFORCE leave-one-out estimator. Our control variates can be adapted online to minimize variance and do not require extra evaluations of the target function. In benchmark generative modeling tasks such as training binary variational autoencoders, our gradient estimator achieves substantially lower variance than state-of-the-art estimators with the same number of function evaluations. Jiaxin Shi, Jessica Hwang, Michalis K. Titsias, Lester Mackey |
NeurIPS | 1 |
| 2021 | TWAG: A Topic-Guided Wikipedia Abstract GeneratorabstractFangwei Zhu, Shangqing Tu, Jiaxin Shi, Juanzi Li, Lei Hou, Tong Cui. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Fangwei Zhu, Shangqing Tu, Jiaxin Shi, Juan-Zi Li, Lei Hou 0001, Tong Cui |
ACL/IJCNLP (1) | 3 |
| 2021 | TransferNet: An Effective and Transparent Framework for Multi-hop Question Answering over Relation GraphabstractMulti-hop Question Answering (QA) is a challenging task because it requires precise reasoning with entity relations at every step towards the answer.The relations can be represented in terms of labels in knowledge graph (e.g., spouse) or text in text corpus (e.g., they have been married for 26 years).Existing models usually infer the answer by predicting the sequential relation path or aggregating the hidden graph features.The former is hard to optimize, and the latter lacks interpretability.In this paper, we propose Trans-ferNet, an effective and transparent model for multi-hop QA, which supports both label and text relations in a unified framework.Trans-ferNet jumps across entities at multiple steps.At each step, it attends to different parts of the question, computes activated scores for relations, and then transfer the previous entity scores along activated relations in a differentiable way.We carry out extensive experiments on three datasets and demonstrate that TransferNet surpasses the state-of-the-art models by a large margin.In particular, on MetaQA, it achieves 100% accuracy in 2-hop and 3-hop questions.By qualitative analysis, we show that TransferNet has transparent and interpretable intermediate results. Jiaxin Shi, Shulin Cao, Lei Hou 0001, Juan-Zi Li, Hanwang Zhang |
EMNLP (1) | 1 |
| 2021 | Scalable Variational Gaussian Processes via Harmonic Kernel DecompositionabstractWe introduce a new scalable variational Gaussian process approximation which provides a high fidelity approximation while retaining general applicability. We propose the harmonic kernel decomposition (HKD), which uses Fourier series to decompose a kernel as a sum of orthogonal kernels. Our variational approximation exploits this orthogonality to enable a large number of inducing points at a low computational cost. We demonstrate that, on a range of regression and classification problems, our approach can exploit input space symmetries such as translations and reflections, and it significantly outperforms standard variational methods in scalability and accuracy. Notably, our approach achieves state-of-the-art results on CIFAR-10 among pure GP models. Shengyang Sun, Jiaxin Shi, Andrew Gordon Wilson, Roger B. Grosse |
ICML | 2 |
| 2020 | Sparse Orthogonal Variational Inference for Gaussian ProcessesabstractWe introduce a new interpretation of sparse variational approximations for Gaussian processes using inducing points, which can lead to more scalable algorithms than previous methods. It is based on decomposing a Gaussian process as a sum of two independent processes: one spanned by a finite basis of inducing points and the other capturing the remaining variation. We show that this formulation recovers existing approximations and at the same time allows to obtain tighter lower bounds on the marginal likelihood and new stochastic variational inference algorithms. We demonstrate the efficiency of these algorithms in several Gaussian process models ranging from standard regression to multi-class classification using (deep) convolutional Gaussian processes and report state-of-the-art results on CIFAR-10 among purely GP-based models. Jiaxin Shi, Michalis K. Titsias, Andriy Mnih |
AISTATS | 1 |
| 2020 | Unbiased Scene Graph Generation From Biased TrainingabstractToday's scene graph generation (SGG) task is still far from practical, mainly due to the severe training bias, e.g., collapsing diverse "human walk on / sit on / lay on beach" into "human on beach". Given such SGG, the down-stream tasks such as VQA can hardly infer better scene structures than merely a bag of objects. However, debiasing in SGG is not trivial because traditional debiasing methods cannot distinguish between the good and bad bias, e.g., good context prior (e.g., "person read book" rather than "eat") and bad long-tailed bias (e.g., "near" dominating "behind / in front of"). In this paper, we present a novel SGG framework based on causal inference but not the conventional likelihood. We first build a causal graph for SGG, and perform traditional biased training with the graph. Then, we propose to draw the counterfactual causality from the trained graph to infer the effect from the bad bias, which should be removed. In particular, we use Total Direct Effect (TDE) as the proposed final predicate score for unbiased SGG. Note that our framework is agnostic to any SGG model and thus can be widely applied in the community who seeks unbiased predictions. By using the proposed Scene Graph Diagnosis toolkit on the SGG benchmark Visual Genome and several prevailing models, we observed significant improvements over the previous state-of-the-art methods. Kaihua Tang, Yulei Niu, Jianqiang Huang 0001, Jiaxin Shi, Hanwang Zhang |
CVPR | 4 |
| 2020 | Baidu Kunlun An AI processor for diversified workloadsabstractThis article consists only of a collection of slides from the author's conference presentation. Jian Ouyang, Mijung Noh, Yin Ma, Canghai Gu, SoonGon Kim, Ki-il Hong, Wang-Keun Bae, Zhibiao Zhao, Xiaozhang Gong, Jiaxin Shi, Hefei Zhu, Xueliang Du |
Hot Chips Symposium | 14 |
| 2020 | Nonparametric Score EstimatorsabstractEstimating the score, i.e., the gradient of log density function, from a set of samples generated by an unknown distribution is a fundamental task in inference and learning of probabilistic models that involve flexible yet intractable densities. Kernel estimators based on Stein’s methods or score matching have shown promise, however their theoretical properties and relationships have not been fully-understood. We provide a unifying view of these estimators under the framework of regularized nonparametric regression. It allows us to analyse existing estimators and construct new ones with desirable properties by choosing different hypothesis spaces and regularizers. A unified convergence analysis is provided for such estimators. Finally, we propose score estimators based on iterative regularization that enjoy computational benefits from curl-free kernels and fast convergence. Jiaxin Shi, Jun Zhu 0001 |
ICML | 2 |
| 2020 | Calibrating nested sensor arrays for DOA estimation utilizing continuous multiplication operator
Ye Tian 0014, Jiaxin Shi, Hong Yue, Xiaoliu Rong |
Signal Process. | 2 |
| 2019 | Learning to Embed Sentences Using Attentive Recursive Trees
Jiaxin Shi, Lei Hou 0001, Juan-Zi Li, Zhiyuan Liu 0001, Hanwang Zhang |
AAAI | 1 |
| 2019 | DeepChannel: Salience Estimation by Contrastive Learning for Extractive Document SummarizationabstractWe propose DeepChannel, a robust, data-efficient, and interpretable neural model for extractive document summarization. Given any document-summary pair, we estimate a salience score, which is modeled using an attention-based deep neural network, to represent the salience degree of the summary for yielding the document. We devise a contrastive training strategy to learn the salience estimation network, and then use the learned salience score as a guide and iteratively extract the most salient sentences from the document as our generated summary. In experiments, our model not only achieves state-of-the-art ROUGE scores on CNN/Daily Mail dataset, but also shows strong robustness in the out-of-domain test on DUC2007 test set. Moreover, our model reaches a ROUGE-1 F-1 score of 39.41 on CNN/Daily Mail test set with merely 1/100 training set, demonstrating a tremendous data efficiency. Jiaxin Shi, Lei Hou 0001, Juan-Zi Li, Zhiyuan Liu 0001, Hanwang Zhang |
AAAI | 1 |
| 2019 | Explainable and Explicit Visual Reasoning Over Scene GraphsabstractWe aim to dismantle the prevalent black-box neural architectures used in complex visual reasoning tasks, into the proposed eXplainable and eXplicit Neural Modules (XNMs), which advance beyond existing neural module networks towards using scene graphs - objects as nodes and the pairwise relationships as edges - for explainable and explicit reasoning with structured knowledge. XNMs allow us to pay more attention to teach machines how to "think'', regardless of what they "look''. As we will show in the paper, by using scene graphs as an inductive bias, 1) we can design XNMs in a concise and flexible fashion, i.e., XNMs merely consist of 4 meta-types, which significantly reduce the number of parameters by 10 to 100 times, and 2) we can explicitly trace the reasoning-flow in terms of graph attentions. XNMs are so generic that they support a wide range of scene graph implementations with various qualities. For example, when the graphs are detected perfectly, XNMs achieve 100% accuracy on both CLEVR and CLEVR CoGenT, establishing an empirical performance upper-bound for visual reasoning; when the graphs are noisily detected from real-world images, XNMs are still robust to achieve a competitive 67.5% accuracy on VQAv2.0, surpassing the popular bag-of-objects attention models without graph structures. Jiaxin Shi, Hanwang Zhang, Juan-Zi Li |
CVPR | 1 |
| 2019 | Semi-supervised Entity Alignment via Joint Knowledge Embedding Model and Cross-graph ModelabstractChengjiang Li, Yixin Cao, Lei Hou, Jiaxin Shi, Juanzi Li, Tat-Seng Chua. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Chengjiang Li, Yixin Cao 0002, Lei Hou 0001, Jiaxin Shi, Juan-Zi Li, Tat-Seng Chua |
EMNLP/IJCNLP (1) | 4 |
| 2019 | Functional variational Bayesian Neural Networks
Shengyang Sun, Guodong Zhang 0006, Jiaxin Shi, Roger B. Grosse |
ICLR (Poster) | 3 |
| 2019 | Scalable Training of Inference Networks for Gaussian-Process ModelsabstractInference in Gaussian process (GP) models is computationally challenging for large data, and often difficult to approximate with a small number of inducing points. We explore an alternative approximation that employs stochastic inference networks for a flexible inference. Unfortunately, for such networks, minibatch training is difficult to be able to learn meaningful correlations over function outputs for a large dataset. We propose an algorithm that enables such training by tracking a stochastic, functional mirror-descent algorithm. At each iteration, this only requires considering a finite number of input locations, resulting in a scalable and easy-to-implement algorithm. Empirical results show comparable and, sometimes, superior performance to existing sparse variational GP methods. Jiaxin Shi, Mohammad Emtiyaz Khan, Jun Zhu 0001 |
ICML | 1 |
| 2019 | Sliced Score Matching: A Scalable Approach to Density and Score Estimation
Yang Song 0011, Sahaj Garg, Jiaxin Shi, Stefano Ermon |
UAI | 3 |
| 2019 | Non-coherent direction of arrival estimation utilizing linear model approximation
Ye Tian 0014, Jiaxin Shi, Qiusheng Lian |
Signal Process. | 2 |
| 2018 | Kernel Implicit Variational Inference
Jiaxin Shi, Shengyang Sun, Jun Zhu 0001 |
ICLR (Poster) | 1 |
| 2018 | A Spectral Approach to Gradient Estimation for Implicit DistributionsabstractRecently there have been increasing interests in learning and inference with implicit distributions (i.e., distributions without tractable densities). To this end, we develop a gradient estimator for implicit distributions based on Stein’s identity and a spectral decomposition of kernel operators, where the eigenfunctions are approximated by the Nystr{ö}m method. Unlike the previous works that only provide estimates at the sample points, our approach directly estimates the gradient function, thus allows for a simple and principled out-of-sample extension. We provide theoretical results on the error bound of the estimator and discuss the bias-variance tradeoff in practice. The effectiveness of our method is demonstrated by applications to gradient-free Hamiltonian Monte Carlo and variational inference with implicit distributions. Finally, we discuss the intuition behind the estimator by drawing connections between the Nystr{ö}m method and kernel PCA, which indicates that the estimator can automatically adapt to the geometry of the underlying distribution. Jiaxin Shi, Shengyang Sun, Jun Zhu 0001 |
ICML | 1 |
| 2018 | Message Passing Stein Variational Gradient DescentabstractStein variational gradient descent (SVGD) is a recently proposed particle-based Bayesian inference method, which has attracted a lot of interest due to its remarkable approximation ability and particle efficiency compared to traditional variational inference and Markov Chain Monte Carlo methods. However, we observed that particles of SVGD tend to collapse to modes of the target distribution, and this particle degeneracy phenomenon becomes more severe with higher dimensions. Our theoretical analysis finds out that there exists a negative correlation between the dimensionality and the repulsive force of SVGD which should be blamed for this phenomenon. We propose Message Passing SVGD (MP-SVGD) to solve this problem. By leveraging the conditional independence structure of probabilistic graphical models (PGMs), MP-SVGD converts the original high-dimensional global inference problem into a set of local ones over the Markov blanket with lower dimensions. Experimental results show its advantages of preventing vanishing repulsive force in high-dimensional space over SVGD, and its particle efficiency and approximation flexibility over other inference methods on graphical models. Jingwei Zhuo, Chang Liu 0030, Jiaxin Shi, Jun Zhu 0001, Ning Chen 0002, Bo Zhang 0010 |
ICML | 3 |
| 2018 | Semi-crowdsourced Clustering with Deep Generative ModelsabstractWe consider the semi-supervised clustering problem where crowdsourcing provides noisy information about the pairwise comparisons on a small subset of data, i.e., whether a sample pair is in the same cluster. We propose a new approach that includes a deep generative model (DGM) to characterize low-level features of the data, and a statistical relational model for noisy pairwise annotations on its subset. The two parts share the latent variables. To make the model automatically trade-off between its complexity and fitting data, we also develop its fully Bayesian variant. The challenge of inference is addressed by fast (natural-gradient) stochastic variational inference algorithms, where we effectively combine variational message passing for the relational part and amortized learning of the DGM under a unified framework. Empirical results on synthetic and real-world datasets show that our model outperforms previous crowdsourced clustering methods. Yucen Luo, Tian Tian 0001, Jiaxin Shi, Jun Zhu 0001, Bo Zhang 0010 |
NeurIPS | 3 |
| 2018 | Non-fragile chaotic synchronization for discontinuous neural networks with time-varying delays and random feedback gain uncertainties
Xiao Peng 0003, Huaiqin Wu, Ka Song, Jiaxin Shi |
Neurocomputing | 4 |
| 2018 | Analyzing the Training Processes of Deep Generative ModelsabstractAmong the many types of deep models, deep generative models (DGMs) provide a solution to the important problem of unsupervised and semi-supervised learning. However, training DGMs requires more skill, experience, and know-how because their training is more complex than other types of deep models such as convolutional neural networks (CNNs). We develop a visual analytics approach for better understanding and diagnosing the training process of a DGM. To help experts understand the overall training process, we first extract a large amount of time series data that represents training dynamics (e.g., activation changes over time). A blue-noise polyline sampling scheme is then introduced to select time series samples, which can both preserve outliers and reduce visual clutter. To further investigate the root cause of a failed training process, we propose a credit assignment algorithm that indicates how other neurons contribute to the output of the neuron causing the training failure. Two case studies are conducted with machine learning experts to demonstrate how our approach helps understand and diagnose the training processes of DGMs. We also show how our approach can be directly used to analyze other types of deep models, such as CNNs. Mengchen Liu, Jiaxin Shi, Kelei Cao, Jun Zhu 0001, Shixia Liu |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2017 | On Modeling Sense Relatedness in Multi-prototype Word EmbeddingabstractTo enhance the expression ability of distributional word representation learning model, many researchers tend to induce word senses through clustering, and learn multiple embedding vectors for each word, namely multi-prototype word embedding model. However, most related work ignores the relatedness among word senses which actually plays an important role. In this paper, we propose a novel approach to capture word sense relatedness in multi-prototype word embedding model. Particularly, we differentiate the original sense and extended senses of a word by introducing their global occurrence information and model their relatedness through the local textual context information. Based on the idea of fuzzy clustering, we introduce a random process to integrate these two types of senses and design two non-parametric methods for word sense induction. To make our model more scalable and efficient, we use an online joint learning framework extended from the Skip-gram model. The experimental results demonstrate that our model outperforms both conventional single-prototype embedding models and other multi-prototype embedding models, and achieves more stable performance when trained on smaller data. Yixin Cao 0002, Jiaxin Shi, Juan-Zi Li, Zhiyuan Liu 0001, Chengjiang Li |
IJCNLP(1) | 2 |
| 2017 | Global synchronization in finite time for fractional-order neural networks with discontinuous activations and time delays
Xiao Peng 0003, Huaiqin Wu, Ka Song, Jiaxin Shi |
Neural Networks | 4 |
| 2017 | Fast In-Memory Transaction Processing Using RDMA and HTMabstractDrTM is a fast in-memory transaction processing system that exploits advanced hardware features such as remote direct memory access (RDMA) and hardware transactional memory (HTM). To achieve high efficiency, it mostly offloads concurrency control such as tracking read/write accesses and conflict detection into HTM in a local machine and leverages the strong consistency between RDMA and HTM to ensure serializability among concurrent transactions across machines. To mitigate the high probability of HTM aborts for large transactions, we design and implement an optimized transaction chopping algorithm to decompose a set of large transactions into smaller pieces such that HTM is only required to protect each piece. We further build an efficient hash table for DrTM by leveraging HTM and RDMA to simplify the design and notably improve the performance. We describe how DrTM supports common database features like read-only transactions and logging for durability. Evaluation using typical OLTP workloads including TPC-C and SmallBank shows that DrTM has better single-node efficiency and scales well on a six-node cluster; it achieves greater than 1.51, 34 and 5.24, 138 million transactions per second for TPC-C and SmallBank on a single node and the cluster, respectively. Such numbers outperform a state-of-the-art single-node system (i.e., Silo) and a distributed transaction system (i.e., Calvin) by at least 1.9X and 29.6X for TPC-C. Haibo Chen 0001, Rong Chen 0001, Xingda Wei, Jiaxin Shi, Yanzhe Chen, Binyu Zang, Haibing Guan |
ACM Trans. Comput. Syst. | 4 |
| 2017 | Towards Better Analysis of Deep Convolutional Neural NetworksabstractDeep convolutional neural networks (CNNs) have achieved breakthrough performance in many pattern recognition tasks such as image classification. However, the development of high-quality deep models typically relies on a substantial amount of trial-and-error, as there is still no clear understanding of when and why a deep model works. In this paper, we present a visual analytics approach for better understanding, diagnosing, and refining deep CNNs. We formulate a deep CNN as a directed acyclic graph. Based on this formulation, a hybrid visualization is developed to disclose the multiple facets of each neuron and the interactions between them. In particular, we introduce a hierarchical rectangle packing algorithm and a matrix reordering algorithm to show the derived features of a neuron cluster. We also propose a biclustering-based edge bundling method to reduce visual clutter caused by a large number of connections between neurons. We evaluated our method on a set of CNNs and the results are generally favorable. Mengchen Liu, Jiaxin Shi, Zhen Li 0044, Chongxuan Li, Jun Zhu 0001, Shixia Liu |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2017 | PlenoPatch: Patch-Based Plenoptic Image ManipulationabstractPatch-based image synthesis methods have been successfully applied for various editing tasks on still images, videos and stereo pairs. In this work we extend patch-based synthesis to plenoptic images captured by consumer-level lenselet-based devices for interactive, efficient light field editing. In our method the light field is represented as a set of images captured from different viewpoints. We decompose the central view into different depth layers, and present it to the user for specifying the editing goals. Given an editing task, our method performs patch-based image synthesis on all affected layers of the central view, and then propagates the edits to all other views. Interaction is done through a conventional 2D image editing user interface that is familiar to novice users. Our method correctly handles object boundary occlusion with semi-transparency, thus can generate more realistic results than previous methods. We demonstrate compelling results on a wide range of applications such as hole-filling, object reshuffling and resizing, changing object depth, light field upscaling and parallax magnification. Jue Wang 0001, Eli Shechtman, Zi-Ye Zhou, Jiaxin Shi, Shi-Min Hu 0001 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2016 | Fast and general distributed transactions using RDMA and HTMabstractRecent transaction processing systems attempt to leverage advanced hardware features like RDMA and HTM to significantly boost performance, which, however, pose several limitations like requiring priori knowledge of read/write sets of transactions and providing no availability support. In this paper, we present DrTM+R, a fast in-memory transaction processing system that retains the performance benefit from advanced hardware features, while supporting general transactional workloads and high availability through replication. DrTM+R addresses the generality issue by designing a hybrid OCC and locking scheme, which leverages the strong atomicity of HTM and the strong consistency of RDMA to preserve strict serializability with high performance. To resolve the race condition between the immediate visibility of records updated by HTM transactions and the unready replication of such records, DrTM+R leverages an optimistic replication scheme that uses seqlock-like versioning to distinguish the visibility of tuples and the readiness of record replication. Evaluation using typical OLTP workloads like TPC-C and SmallBank shows that DrTM+R scales well on a 6-node cluster and achieves over 5.69 and 94 million transactions per second without replication for TPC-C and SmallBank respectively. Enabling 3-way replication on DrTM+R only incurs at most 41% overhead before reaching network bottleneck, and is still an order-of-magnitude faster than a state-of-the-art distributed transaction system (Calvin). Yanzhe Chen, Xingda Wei, Jiaxin Shi, Rong Chen 0001, Haibo Chen 0001 |
EuroSys | 3 |
| 2016 | Fast and Concurrent RDF Queries with RDMA-Based Distributed Graph Exploration
Jiaxin Shi, Youyang Yao, Rong Chen 0001, Haibo Chen 0001, Feifei Li 0001 |
OSDI | 1 |
| 2015 | PowerLyra: differentiated graph computation and partitioning on skewed graphsabstractNatural graphs with skewed distribution raise unique challenges to graph computation and partitioning. Existing graph-parallel systems usually use a "one size fits all" design that uniformly processes all vertices, which either suffer from notable load imbalance and high contention for high-degree vertices (e.g., Pregel and GraphLab), or incur high communication cost and memory consumption even for low-degree vertices (e.g., PowerGraph and GraphX). Rong Chen 0001, Jiaxin Shi, Yanzhe Chen, Haibo Chen 0001 |
EuroSys | 2 |
| 2015 | Fast in-memory transaction processing using RDMA and HTMabstractWe present DrTM, a fast in-memory transaction processing system that exploits advanced hardware features (i.e., RDMA and HTM) to improve latency and throughput by over one order of magnitude compared to state-of-the-art distributed transaction systems. The high performance of DrTM are enabled by mostly offloading concurrency control within a local machine into HTM and leveraging the strong consistency between RDMA and HTM to ensure serializability among concurrent transactions across machines. We further build an efficient hash table for DrTM by leveraging HTM and RDMA to simplify the design and notably improve the performance. We describe how DrTM supports common database features like read-only transactions and logging for durability. Evaluation using typical OLTP workloads including TPC-C and SmallBank show that DrTM scales well on a 6-node cluster and achieves over 5.52 and 138 million transactions per second for TPC-C and SmallBank Respectively. This number outperforms a state-of-the-art distributed transaction system (namely Calvin) by at least 17.9X for TPC-C. Xingda Wei, Jiaxin Shi, Yanzhe Chen, Rong Chen 0001, Haibo Chen 0001 |
SOSP | 2 |
| 2015 | Bipartite-Oriented Distributed Graph Partitioning for Big Learning
Rong Chen 0001, Jiaxin Shi, Haibo Chen 0001, Binyu Zang |
J. Comput. Sci. Technol. | 2 |