Jindong Chen

dblp:07/5066 · DBLP profile ↗
← Back
28ranked-venue papers
7as first author
14since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 18 · 4 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 1 since 2021Theory of computation · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-author
YearPublicationVenuePosition
2026 Optimizing Generative Ranking Relevance via Reinforcement Learning in Xiaohongshu Search
abstract
Ranking relevance is a fundamental task in search engines, aiming to identify the items most relevant to a given user query. Traditional relevance models typically produce scalar scores or directly predict relevance labels, limiting both interpretability and the modeling of complex relevance signals. Inspired by recent advances in Chain-of-Thought (CoT) reasoning for complex tasks, we investigate whether explicit reasoning can enhance both interpretability and performance in relevance modeling. However, existing reasoning-based Generative Relevance Models (GRMs) primarily rely on supervised fine-tuning on large amounts of human-annotated or synthetic CoT data, which often leads to limited generalization. Moreover, domain-agnostic, free-form reasoning tends to be overly generic and insufficiently grounded, limiting its potential to handle the diverse and ambiguous cases prevalent in open-domain search. In this work, we formulate relevance modeling in Xiaohongshu search as a reasoning task and introduce a Reinforcement Learning (RL)-based training framework to enhance the grounded reasoning capabilities of GRMs. Specifically, we incorporate practical business-specific relevance criteria into the multi-step reasoning prompt design and propose Stepwise Advantage Masking (SAM), a lightweight process-supervision strategy which facilitates effective learning of these criteria through improved credit assignment. To enable industrial deployment, we further distill the large-scale RL-tuned model to a lightweight version suitable for real-world search systems. Extensive offline evaluations and online A/B tests demonstrate that our approach consistently delivers significant improvements across key relevance and business metrics, validating its effectiveness, robustness, and practicality for large-scale industrial search systems.
Ziyang Zeng, Heming Jing, Jindong Chen, Yige Sun, Zheyong Xie, Shaosheng Cao, Yao Hu 0002
KDD (1)3
2026 Fraud detection based on GNNs with local augmentation and adaptive relation aggregation
Zhou Mengzhe, Jindong Chen, Zhang Wen, Zhihua Yan
Expert Syst. Appl.2
2026 HAM-VAE: A hierarchical adaptive modulation VAE for data augmentation in fake review detection
Jindong Chen, Zhihua Yan, Zhichao Zheng 0010
Knowl. Based Syst.1
2025 Accelerating Inference of Retrieval-Augmented Generation via Sparse Context Selection
abstract
Large language models (LLMs) augmented with retrieval exhibit robust performance and extensive versatility by incorporating external contexts. However, the input length grows linearly in the number of retrieved documents, causing a dramatic increase in latency. In this paper, we propose a novel paradigm named Sparse RAG, which seeks to cut computation costs through sparsity. Specifically, Sparse RAG encodes retrieved documents in parallel, which eliminates latency introduced by long-range attention of retrieved documents. Then, LLMs selectively decode the output by only attending to highly relevant caches auto-regressively, which are chosen via prompting LLMs with special control tokens. It is notable that Sparse RAG combines the assessment of each individual document and the generation of the response into a single process. The designed sparse mechanism in a RAG system can facilitate the reduction of the number of documents loaded during decoding for accelerating the inference of the RAG system. Additionally, filtering out undesirable contexts enhances the model’s focus on relevant context, inherently improving its generation quality. Evaluation results on four datasets show that Sparse RAG can be used to strike an optimal balance between generation quality and computational efficiency, demonstrating its generalizability across tasks.
Jia-Chen Gu, Caitlin Sikora, Ho Ko, Yinxiao Liu, Chu-Cheng Lin, Lei Shu 0004, Liangchen Luo, Lei Meng 0008, Jindong Chen
ICLR11
2025 ScreenQA: Large-Scale Question-Answer Pairs Over Mobile App Screenshots
abstract
Yu-Chung Hsiao, Fedir Zubach, Gilles Baechler, Srinivas Sunkara, Victor Carbune, Jason Lin, Maria Wang, Yun Zhu, Jindong Chen. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Yu-Chung Hsiao, Fedir Zubach, Gilles Baechler, Srinivas Sunkara, Victor Carbune, Maria Wang, Jindong Chen
NAACL (Long Papers)9
2024 RewriteLM: An Instruction-Tuned Large Language Model for Text Rewriting
abstract
Large Language Models (LLMs) have demonstrated impressive capabilities in creative tasks such as storytelling and E-mail generation. However, as LLMs are primarily trained on final text results rather than intermediate revisions, it might be challenging for them to perform text rewriting tasks. Most studies in the rewriting tasks focus on a particular transformation type within the boundaries of single sentences. In this work, we develop new strategies for instruction tuning and reinforcement learning to better align LLMs for cross-sentence rewriting tasks using diverse wording and structures expressed through natural languages including 1) generating rewriting instruction data from Wiki edits and public corpus through instruction generation and chain-of-thought prompting; 2) collecting comparison data for reward model training through a new ranking function. To facilitate this research, we introduce OpenRewriteEval, a novel benchmark covers a wide variety of rewriting types expressed through natural language instructions. Our results show significant improvements over a variety of baselines.
Lei Shu 0004, Liangchen Luo, Jayakumar Hoskere, Yinxiao Liu, Simon Tong, Jindong Chen, Lei Meng 0008
AAAI7
2024 ScreenAI: A Vision-Language Model for UI and Infographics Understanding
Gilles Baechler, Srinivas Sunkara, Maria Wang, Fedir Zubach, Hassan Mansoor, Vincent Etter, Victor Carbune, Jindong Chen, Abhanshu Sharma
IJCAI9
2023 Cappy: Outperforming and Boosting Large Multi-Task LMs with a Small Scorer
abstract
Large language models (LLMs) such as T0, FLAN, and OPT-IML excel in multi-tasking under a unified instruction-following paradigm, where they also exhibit remarkable generalization abilities to unseen tasks. Despite their impressive performance, these LLMs, with sizes ranging from several billion to hundreds of billions of parameters, demand substantial computational resources, making their training and inference expensive and inefficient. Furthermore, adapting these models to downstream applications, particularly complex tasks, is often unfeasible due to the extensive hardware requirements for finetuning, even when utilizing parameter-efficient approaches such as prompt tuning. Additionally, the most powerful multi-task LLMs, such as OPT-IML-175B and FLAN-PaLM-540B, are not publicly accessible, severely limiting their customization potential. To address these challenges, we introduce a pretrained small scorer, \textit{Cappy}, designed to enhance the performance and efficiency of multi-task LLMs. With merely 360 million parameters, Cappy functions either independently on classification tasks or serve as an auxiliary component for LLMs, boosting their performance. Moreover, Cappy enables efficiently integrating downstream supervision without requiring LLM finetuning nor the access to their parameters. Our experiments demonstrate that, when working independently on 11 language understanding tasks from PromptSource, Cappy outperforms LLMs that are several orders of magnitude larger. Besides, on 45 complex tasks from BIG-Bench, Cappy boosts the performance of the advanced multi-task LLM, FLAN-T5, by a large margin. Furthermore, Cappy is flexible to cooperate with other LLM adaptations, including finetuning and in-context learning, offering additional performance enhancement.
Bowen Tan, Eric P. Xing, Zhiting Hu, Jindong Chen
NeurIPS6
2022 Towards Better Semantic Understanding of Mobile Interfaces
abstract
Improving the accessibility and automation capabilities of mobile devices can have a significant positive impact on the daily lives of countless users. To stimulate research in this direction, we release a human-annotated dataset with approximately 500k unique annotations aimed at increasing the understanding of the functionality of UI elements. This dataset augments images and view hierarchies from RICO, a large dataset of mobile UIs, with annotations for icons based on their shapes and semantics, and associations between different elements and their corresponding text labels, resulting in a significant increase in the number of UI elements and the categories assigned to them. We also release models using image-only and multimodal inputs; we experiment with various architectures and study the benefits of using multimodal inputs on the new dataset. Our models demonstrate strong performance on an evaluation set of unseen apps, indicating their generalizability to newer screens. These models, combined with the new dataset, can enable innovative functionalities like referring to UI elements by their labels, improved coverage and better semantics for icons etc., which would go a long way in making UIs more usable for everyone.
Srinivas Sunkara, Maria Wang, Gilles Baechler, Yu-Chung Hsiao, Jindong Chen, Abhanshu Sharma, James W. Stout
COLING6
2021 ActionBert: Leveraging User Actions for Semantic Understanding of User Interfaces
abstract
As mobile devices are becoming ubiquitous, regularly interacting with a variety of user interfaces (UIs) is a common aspect of daily life for many people. To improve the accessibility of these devices and to enable their usage in a variety of settings, building models that can assist users and accomplish tasks through the UI is vitally important. However, there are several challenges to achieve this. First, UI components of similar appearance can have different functionalities, making understanding their function more important than just analyzing their appearance. Second, domain-specific features like Document Object Model (DOM) in web pages and View Hierarchy (VH) in mobile applications provide important signals about the semantics of UI elements, but these features are not in a natural language format. Third, owing to a large diversity in UIs and absence of standard DOM or VH representations, building a UI understanding model with high coverage requires large amounts of training data. Inspired by the success of pre-training based approaches in NLP for tackling a variety of problems in a data-efficient way, we introduce a new pre-trained UI representation model called ActionBert. Our methodology is designed to leverage visual, linguistic and domain-specific features in user interaction traces to pre-train generic feature representations of UIs and their components. Our key intuition is that user actions, e.g., a sequence of clicks on different UI components, reveals important information about their functionality. We evaluate the proposed model on a wide variety of downstream tasks, ranging from icon classification to UI component retrieval based on its natural language description. Experiments show that the proposed ActionBert model outperforms multi-modal baselines across all downstream tasks by up to 15.5%.
Zecheng He, Srinivas Sunkara, Xiaoxue Zang, Nevan Wichers, Gabriel Schubiner, Ruby B. Lee, Jindong Chen
AAAI9
2021 PhotoChat: A Human-Human Dialogue Dataset With Photo Sharing Behavior For Joint Image-Text Modeling
abstract
Xiaoxue Zang, Lijuan Liu, Maria Wang, Yang Song, Hao Zhang, Jindong Chen. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Xiaoxue Zang, Maria Wang, Yang Song 0008, Jindong Chen
ACL/IJCNLP (1)6
2021 Bootstrapping Information Extraction via Conceptualization
abstract
Bootstrapping enables us to use existing knowledge to find patterns and extract new knowledge from free texts, from which more patterns can be found. Due to its minimally supervised, domain-independent, and language-independent nature, it has been widely adopted in real-world applications. However, as iterations go on, semantic drift may happen. The extraction may shift from the target class to other classes and result in errors, which propagate in the succeeding iterations and hurt the performance significantly. Existing solutions simply throw away bad patterns, sacrificing recall to ensure high precision. However, we argue that most of these patterns and instances can be kept as long as being applied selectively, guided by prior knowledge. In this paper, we propose a pattern-based extraction framework with three distinguished features: (1) it uses conceptual taxonomies to guide the extraction to reduce semantic drift; (2) it uses the knowledge of existing triples to improve the precision; (3) it integrates all patterns to form a generalized pattern set with quantified confidence measurement. The proposed solution is applied on enriching two real-world knowledge bases and achieves higher precision and recall compared to existing solutions.
Jiaqing Liang, Suo Feng, Chenhao Xie 0002, Yanghua Xiao, Jindong Chen, Seung-won Hwang
ICDE5
2021 UIBert: Learning Generic Multimodal Representations for UI Understanding
abstract
To improve the accessibility of smart devices and to simplify their usage, building models which understand user interfaces (UIs) and assist users to complete their tasks is critical. However, unique challenges are proposed by UI-specific characteristics, such as how to effectively leverage multimodal UI features that involve image, text, and structural metadata and how to achieve good performance when high-quality labeled data is unavailable. To address such challenges we introduce UIBert, a transformer-based joint image-text model trained through novel pre-training tasks on large-scale unlabeled UI data to learn generic feature representations for a UI and its components. Our key intuition is that the heterogeneous features in a UI are self-aligned, i.e., the image and text features of UI components, are predictive of each other. We propose five pretraining tasks utilizing this self-alignment among different features of a UI component and across various components in the same UI. We evaluate our method on nine real-world downstream UI tasks where UIBert outperforms strong multimodal baselines by up to 9.26% accuracy.
Chongyang Bai, Xiaoxue Zang, Srinivas Sunkara, Abhinav Rastogi, Jindong Chen, Blaise Agüera y Arcas
IJCAI6
2021 Multimodal Icon Annotation For Mobile Applications
abstract
Annotating user interfaces (UIs) that involves localization and classification of meaningful UI elements on a screen is a critical step for many mobile applications such as screen readers and voice control of devices. Annotating object icons, such as menu, search, and arrow backward, is especially challenging due to the lack of explicit labels on screens, their similarity to pictures, and their diverse shapes. Existing studies either use view hierarchy or pixel based methods to tackle the task. Pixel based approaches are more popular as view hierarchy features on mobile platforms are often incomplete or inaccurate, however it leaves out instructional information in the view hierarchy such as resource-ids or content descriptions. We propose a novel deep learning based multi-modal approach that combines the benefits of both pixel and view hierarchy features as well as leverages the state-of-the-art object detection techniques. In order to demonstrate the utility provided, we create a high quality UI dataset by manually annotating the most commonly used 29 icons in Rico, a large scale mobile design dataset consisting of 72k UI screenshots. The experimental results indicate the effectiveness of our multi-modal approach. Our model not only outperforms a widely used object classification baseline but also pixel based object detection models. Our study sheds light on how to combine view hierarchy with pixel features for annotating UI elements.
Xiaoxue Zang, Jindong Chen
MobileHCI3
2020 An artificial bee colony-based kernel ridge regression for automobile insurance fraud identification
Chun Yan, Wei Liu 0051, Maozhen Li 0001, Jindong Chen
Neurocomputing5
2019 Deep Short Text Classification with Knowledge Powered Attention
abstract
Short text classification is one of important tasks in Natural Language Processing (NLP). Unlike paragraphs or documents, short texts are more ambiguous since they have not enough contextual information, which poses a great challenge for classification. In this paper, we retrieve knowledge from external knowledge source to enhance the semantic representation of short texts. We take conceptual information as a kind of knowledge and incorporate it into deep neural networks. For the purpose of measuring the importance of knowledge, we introduce attention mechanisms and propose deep Short Text Classification with Knowledge powered Attention (STCKA). We utilize Concept towards Short Text (CST) attention and Concept towards Concept Set (C-CS) attention to acquire the weight of concepts from two aspects. And we classify a short text with the help of conceptual information. Unlike traditional approaches, our model acts like a human being who has intrinsic ability to make decisions based on observation (i.e., training data for machines) and pays more attention to important knowledge. We also conduct extensive experiments on four public datasets for different tasks. The experimental results and case studies show that our model outperforms the state-of-the-art methods, justifying the effectiveness of knowledge powered attention.
Jindong Chen, Yizhou Hu, Yanghua Xiao, Haiyun Jiang
AAAI1
2019 CN-Probase: A Data-Driven Approach for Large-Scale Chinese Taxonomy Construction
abstract
Taxonomies play an important role in machine intelligence. However, most well-known taxonomies are in English, and non-English taxonomies, especially Chinese ones, are still very rare. In this paper, we focus on automatic Chinese taxonomy construction and propose an effective generation and verification framework to build a large-scale and high-quality Chinese taxonomy. In the generation module, we extract isA relations from multiple sources of Chinese encyclopedia, which ensures the coverage. To further improve the precision of taxonomy, we apply three heuristic approaches in verification module. As a result, we construct the largest Chinese taxonomy with high precision about 95% called CN-Probase. Our taxonomy has been deployed on Aliyun, with over 82 million API calls in six months.
Jindong Chen, Jiangjie Chen, Yanghua Xiao, Zhendong Chu, Jiaqing Liang, Wei Wang 0009
ICDE1
2019 Relation Extraction Using Supervision from Topic Knowledge of Relation Labels
abstract
Explicitly exploring the semantics of a relation is significant for high-accuracy relation extraction, which is, however, not fully studied in previous work. In this paper, we mine the topic knowledge of a relation to explicitly represent the semantics of this relation, and model relation extraction as a matching problem. That is, the matching score between a sentence and a candidate relation is predicted for an entity pair. To this end, we propose a deep matching network to precisely model the semantic similarity between a sentence-relation pair. Besides, the topic knowledge also allows us to derive the importance information of samples as well as two knowledge-guided negative sampling strategies in the training process. We conduct extensive experiments to evaluate the proposed framework and observe improvements in AUC of 11.5% and max F1 of 5.4% over the baselines with state-of-the-art performance.
Haiyun Jiang, Deqing Yang, Jindong Chen, Jiaqing Liang, Chao Wang 0095, Yanghua Xiao, Wei Wang 0009
IJCAI5
2019 Learning Question-Guided Video Representation for Multi-Turn Video Question Answering
abstract
Understanding and conversing about dynamic scenes is one of the key capabilities of AI agents that navigate the environment and convey useful information to humans.Video question answering is a specific scenario of such AI-human interaction where an agent generates a natural language response to a question regarding the video of a dynamic scene.Incorporating features from multiple modalities, which often provide supplementary information, is one of the challenging aspects of video question answering.Furthermore, a question often concerns only a small segment of the video, hence encoding the entire video sequence using a recurrent neural network is not computationally efficient.Our proposed question-guided video representation module efficiently generates the token-level video summary guided by each word in the question.The learned representations are then fused with the question to generate the answer.Through empirical evaluation on the Audio Visual Scene-aware Dialog (AVSD) dataset (Alamri et al., 2019a), our proposed models in single-turn and multiturn question answering achieve state-of-theart performance on several automatic natural language generation evaluation metrics.
Guan-Lin Chao, Abhinav Rastogi, Semih Yavuz, Dilek Hakkani-Tür, Jindong Chen, Ian Lane
SIGdial5
2018 Resolving Referring Expressions in Images with Labeled Elements
abstract
Images may have elements containing text and a bounding box associated with them, for example, text identified via optical character recognition on a computer screen image, or a natural image with labeled objects. We present an end-to-end trainable architecture to incorporate the information from these elements and the image to segment/identify the part of the image a natural language expression is referring to. We calculate an embedding for each element and then project it onto the corresponding location (i.e., the associated bounding box) of the image feature map. We show that this architecture gives an improvement in resolving referring expressions, over only using the image, and other methods that incorporate the element information. We demonstrate experimental results on the referring expression datasets based on COCO and on a webpage image referring expression dataset that we developed.
Nevan Wichers, Dilek Hakkani-Tür, Jindong Chen
SLT3
2015 Ensemble of SVM Classifiers with Different Representations for Societal Risk Classification
abstract
Using the posts of Tianya Forum as the data source and adopting the societal risk indicators from socio psychology, we conduct document-level multiple societal risk classification of BBS posts. Two kinds of models are applied to generate the representations of posts respectively: Bag-of-Words focuses on extracting the occurrence information of words in posts, and a deep learning model as Post Vector is designed to capture the semantics and word order of posts. Based on the different post representations, two types of support vector machine (SVM) classifiers are developed and compared in the societal risk classification of the posts. Furthermore, as the complementary information contained in the two different post representations, several SVM ensemble methods at the decision score level of the two SVM classifiers are proposed to improve the performance of societal risk classification. The experimental results reveal that the SVM ensemble method achieves better results in document-level societal risk classification than SVM based on single representation.
Jindong Chen, Xijin Tang 0001
KSEM1
2015 The Challenges and Feasibility of Societal Risk Classification Based on Deep Learning of Representations
abstract
Using the posts of Tianya Forum as the data source and adopting the socio psychology study results on societal risks perception, we analyze the challenges and feasibility of the document-level multiple societal risk classification of BBS posts. To effectively capture the semantics and word order of documents, a deep learning model as Post Vector is applied to realize the distributed vector representations of the posts in the vector space. Based on the distributed vector representations, cross-validated classification of the posts labeled by different annotators with KNN method and pair wise similarities comparisons of the posts between risk categories are implemented. The big variance of the results of cross validation shows the differences of individual risk perceptions, which reflects the challenges of societal risk classification. Furthermore, the higher similarities of posts in same societal risk category manifest the feasibility of the classification of societal risks, and indicate the possibility to improve the performance of the societal risk classification of BBS posts.
Jindong Chen, Xijin Tang 0001
SMC1
2008 Image objects and multi-scale features for annotation detection
abstract
This paper investigates several issues in the problem of detecting handwritten markings, or annotations, on printed documents. One issue is to define the appropriate units over which to perform feature measurements and assign type labels. We propose an alpha-shape tree that operates across multiple scales. A second issue is to devise image features that offer inferential power for machine learning algorithms. We report on a feature that measures edge turn statistics. A third issue is how to combine local and neighborhood evidence. We exploit the alpha shape tree in a direct inference architecture. Information propagation schemes such as Markov random fields may be readily layered on top of our output.
Jindong Chen, Eric Saund, Yizhou Wang 0001
ICPR1
2005 Avoiding Local Optima in Single Particle Reconstruction
Marshall W. Bern, Jindong Chen, Hao Chi Wong
RECOMB2
1999 Energy formulations of A-splines
Chandrajit L. Bajaj, Jindong Chen, Robert J. Holt, Arun N. Netravali
Comput. Aided Geom. Des.2
1997 A Triangulation-Based Object Reconstruction Method
abstract
Reconstructing the shape of a 3D object from a digital scan of its surface has a range of applications, such asreverse engineering, authoring 3D synthetic worlds, shape analysis, 3D faxing and tailor-fit modeling. Input data
Fausto Bernardini, Chandrajit L. Bajaj, Jindong Chen, Daniel Schikore
SCG3
1995 Modeling with Cubic A-Patches
abstract
We present a sufficient criterion for the Bernstein Bezier (BB) form of a trivariate polynomial within a tetrahedron, such that the real zero contour of the polynomial defines a smooth and single-sheeted algebraic surface patch, We call this an A-patch.We present algorithms to build a mesh of cubic A-patches to interpolate a given set of scattered point data in three dimensions, respecting tbe topology of any surface triangulation T of the given point set.In these algorithms we first specify "normals" an the data points, then build a simplicial hull consisting of tetrahedral surrounding the surface triangulation 2', and finally construct cubic A-patches within each tetrahedron.The resulting surface constructed is C' (tangent plane) continuous and single sheeted in each of the tetrahedral.We also show how to adjust the free parameters of the A-patches to achieve both local and global shape control.
Chandrajit L. Bajaj, Jindong Chen
ACM Trans. Graph.2
1990 Shortest Paths on a Polyhedron
abstract
We present an algorithm for determining the shortest path between a source point and any destination point along the surface of a polyhedron (need not be convex). Our algorithm uses a new approach which deviates from the conventional “continuous Dijkstra” technique. It takes Ο(n2) time and ⊖(n) space to determine the shortest path and to compute the inward layout which can be used to construct a structure for processing queries of shortest path from the source point to any destination point.
Jindong Chen, Yijie Han
SCG1