EDBT 2026 Demo / reviewers in the wild / expert
Qianying Wang 0002
dblp:86/11012-2 · also QianYing Wang 0002
· DBLP profile ↗
47ranked-venue papers
0as first author
45since 2021 · last 2026
0000-0003-2058-2282ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 30 · 28 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 17 since 2021Human-computer interaction and ubiquitous computing · 8 · 8 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ST-SAM: Multimodal Scene Text Segmentation with Dense Visual and Sparse Textual Prompts via SAMabstractScene text segmentation is a critical preprocessing step in various text-based applications. Specialist text segmentation methods, often relying on a detect-then-segment paradigm, tend to exhibit reduced robustness and can lead to cascading errors. The introduction of the Segment Anything Model (SAM) has revolutionized general segmentation by leveraging vision foundation models. However, SAM still falls short when applied to domain-specific tasks such as scene text segmentation. To bridge this gap between SAM and specialized scene text segmentation approaches, we propose ST-SAM (Scene Text SAM), a parameter-efficient fine-tuning framework tailored to adapt SAM for high-quality scene text segmentation without relying on explicit text detection. ST-SAM incorporates a multimodal prompting mechanism: a lightweight visual encoder generates multi-scale spatial features to provide precise visual context; and textual prompts generated by a large language model offer high-level semantic guidance. We demonstrate the advantages of the proposed ST-SAM as follows: (1) ST-SAM achieves new state-of-the-art performance on multiple scene text segmentation benchmarks, including 85.30% fgIoU on Total-Text and 91.03% fgIoU on TextSeg, outperforming both specialist and generalist models. (2) ST-SAM enables effective domain adaptation by flexibly adapting the general SAM architecture to the domain of scene text. (3) By discarding the detect-then-segment pipeline, ST-SAM simplifies the inference process while still achieving robust performance on complex text cases. Yaqiang Wu, Jiayi Yan, Yu Zhou 0015, Lingling Zhang 0005, Qianying Wang 0002 |
AAAI | 8 |
| 2026 | GoT-CQA: Graph-of-Thought guided compositional reasoning for chart question answering
Lingling Zhang 0005, Muye Huang, Qianying Wang 0002, Yaxian Wang, Ziqi He, Jun Liu 0002 |
Comput. Vis. Image Underst. | 3 |
| 2026 | Memory-enriched thought-by-thought framework for complex Diagram Question Answering
Xinyu Zhang 0021, Lingling Zhang 0005, Yanrui Wu, Muye Huang, Qianying Wang 0002, Jun Liu 0002 |
Comput. Vis. Image Underst. | 7 |
| 2025 | Unleashing the Potential of Model Bias for Generalized Category DiscoveryabstractGeneralized Category Discovery is a significant and complex task that aims to identify both known and undefined novel categories from a set of unlabeled data, leveraging another labeled dataset containing only known categories. The primary challenges stem from model bias induced by pre-training on only known categories and the lack of precise supervision for novel ones, leading to category bias towards known categories and category confusion among different novel categories, which hinders models' ability to identify novel categories effectively. To address these challenges, we propose a novel framework named Self-Debiasing Calibration (SDC). Unlike prior methods that regard model bias towards known categories as an obstacle to novel category identification, SDC provides a novel insight into unleashing the potential of the bias to facilitate novel category learning. Specifically, we utilize the biased pre-trained model to guide the subsequent learning process on unlabeled data. The output of the biased model serves two key purposes. First, it provides an accurate modeling of category bias, which can be utilized to measure the degree of bias and debias the output of the current training model. Second, it offers valuable insights for distinguishing different novel categories by transferring knowledge between similar categories. Based on these insights, SDC dynamically adjusts the output logits of the current training model using the output of the biased model. This approach produces less biased logits to effectively address the issue of category bias towards known categories, and generates more accurate pseudo labels for unlabeled data, thereby mitigating category confusion for novel categories. Experiments on three benchmark datasets show that SDC outperforms SOTA methods, especially in the identification of novel categories. Wenbin An, Haonan Lin, Jiahao Nie 0002, Feng Tian 0002, Wenkai Shi, Yaqiang Wu, Qianying Wang 0002, Ping Chen 0001 |
AAAI | 7 |
| 2025 | Chorus of the Past: Toward Designing a Multi-agent Conversational Reminiscence System with Digital Artifacts for Older AdultsabstractReminiscence has been shown to provide benefits for older adults, but traditionally relies on personal photos as memory cues and interactions with real people who may not always be available. We present ReminiBuddy, a novel LLM-powered multi-agent conversational system, which allows older adults to engage with two distinct agents - one embodying an older identity and the other a younger identity - while using not only personal photos but also 3D models of generic nostalgic objects as memory cues. Our study, with older adult participants, found that the conversational approach both enjoyable and beneficial for reminiscence. While the younger agent was perceived as more emotionally engaging, the older one fostered greater resonance in content. Personal photos prompted autobiographical memories, whereas 3D generic nostalgic objects evoked shared memories of an era, contributing to a more multifaceted reminiscence experience. We further present design implications for better supporting older adults in reminiscing with LLM-powered conversational agents. Jingwei Sun 0005, Nianlong Li, Zhangwei Lu, Liuxin Zhang, Yu Zhang 0124, Qianying Wang 0002, Mingming Fan 0001 |
CHI | 9 |
| 2025 | Mitigating Object Hallucinations in Large Vision-Language Models with Assembly of Global and Local AttentionabstractDespite great success across various multimodal tasks, Large Vision-Language Models (LVLMs) often encounter object hallucinations with generated textual responses being inconsistent with the actual objects in images. We examine different LVLMs and pinpoint that one root cause of object hallucinations lies with deficient attention on discriminative image features. Specifically, LVLMs often predominantly attend to prompt-irrelevant global features instead of prompt-relevant local features, undermining their visual grounding capacity and leading to object hallucinations. We propose Assembly of Global and Local Attention (AGLA), a training-free and plug-and-play approach that mitigates hallucinations by assembling global features for response generation and local features for visual discrimination simultaneously. Specifically, we introduce an image-prompt matching scheme that captures prompt-relevant local features from images, leading to an augmented view of the input image where prompt-relevant content is highlighted while irrelevant distractions are suppressed. Hallucinations can thus be mitigated with a calibrated logit distribution that is from generative global features of the original image and discriminative local features of the augmented image. Extensive experiments show the superiority of AGLA in LVLM hallucination mitigation, demonstrating its wide applicability across both discriminative and generative tasks. Our code is available at https://github.com/Lackel/AGLA. Wenbin An, Feng Tian 0002, Sicong Leng, Jiahao Nie 0002, Haonan Lin, Qianying Wang 0002, Ping Chen 0001, Shijian Lu |
CVPR | 6 |
| 2025 | GBC-Splat: Generalizable Gaussian-Based Clothed Human Digitalization under Sparse RGB CamerasabstractWe present an efficient approach for generalizable clothed human digitalization, termed GBC-Splat. Unlike previous methods that necessitate per-subject optimizations or discount watertight geometry, the proposed method is dedicated to reconstructing complete human shapes and Gaussian Splatting via sparse view RGB inputs in a feed-forward manner. We first extract a fine-grained mesh using a combination of implicit occupancy field regression and explicit disparity estimation between views. The reconstructed high-quality geometry allows us to easily anchor Gaussian primitives to mesh surface according to surface normal and texture, which allows 6-DoF photorealistic novel view synthesis. In addition, we introduce a simple yet effective algorithm to subdivide Gaussian primitives in high-frequency areas to further enhance the visual quality. Without the assistance of human parametric models, our method can tackle loose garments, such as dresses and costumes. Our method outperforms state-of-the-art methods in terms of novel view synthesis while keeping high efficiency, enabling the potential of deployment in real-time applications. Hanzhang Tu, Zhanfeng Liao, Boyao Zhou, Shunyuan Zheng, Liuxin Zhang, Qianying Wang 0002, Yebin Liu |
CVPR | 7 |
| 2025 | Boosting Knowledge Utilization in Multimodal Large Language Models via Adaptive Logits Fusion and Attention ReallocationabstractDespite their recent progress, Multimodal Large Language Models (MLLMs) often struggle in knowledge-intensive tasks due to the limited and outdated parametric knowledge acquired during training. Multimodal Retrieval Augmented Generation addresses this issue by retrieving contextual knowledge from external databases, thereby enhancing MLLMs with expanded knowledge sources.
However, existing MLLMs often fail to fully leverage the retrieved contextual knowledge for response generation. We examine representative MLLMs and identify two major causes, namely, attention bias toward different tokens and knowledge conflicts between parametric and contextual knowledge. To this end, we design Adaptive Logits Fusion and Attention Reallocation (ALFAR), a training-free and plug-and-play approach that improves MLLM responses by maximizing the utility of the retrieved knowledge. Specifically, ALFAR tackles the challenges from two perspectives. First, it alleviates attention bias by adaptively shifting attention from visual tokens to relevant context tokens according to query-context relevance. Second, it decouples and weights parametric and contextual knowledge at output logits, mitigating conflicts between the two types of knowledge. As a plug-and-play method, ALFAR achieves superior performance across diverse datasets without requiring additional training or external tools. Extensive experiments over multiple MLLMs and benchmarks show that ALFAR consistently outperforms the state-of-the-art by large margins. Our code and data are available at https://github.com/Lackel/ALFAR. Wenbin An, Jiahao Nie 0002, Feng Tian 0002, Haonan Lin, Mingxiang Cai, Yaqiang Wu, Qianying Wang 0002, Shijian Lu |
NeurIPS | 7 |
| 2025 | Causal-R: A Causal-Reasoning Geometry Problem Solver for Optimized Solution ExplorationabstractThe task of geometry problem solving has been a long-standing focus in the automated mathematics community and draws growing attention due to its complexity for both symbolic and neural models. Although prior studies have explored various effective approaches for enhancing problem solving performances, two fundamental challenges remain unaddressed, which are essential to the application in practical scenarios. First, the multi-step reasoning gap between the initial geometric conditions and ultimate problem goal leads to a great search space for solution exploration. Second, obtaining multiple interpretable and shorter solutions remains an open problem. In this work, we introduce the Causal-Reasoning Geometry Problem Solver to overcome these challenges. Specifically, the Causal Graph Reasoning theory is proposed to perform symbolic reasoning before problem solving. Several causal graphs are constructed according to predefined rule base, where each graph is composed of primitive nodes, causal edges and prerequisite edges. By applying causal graph deduction from initial conditions, the reachability status of nodes are iteratively conveyed by causal edges until reaching the target nodes, representing feasible causal deduction paths. In this way, the search space of solutions is compressed from the beginning, the end and intermediate reasoning paths, while ensuring the interpretability and variety of solutions. To achieve this, we further propose Forward Matrix Deduction which transforms the causal graphs into matrices and vectors, and applies matrix operations to update the status value of reachable nodes in iterations. Finally, multiple solutions can be generated by tracing back from the target nodes after validation. Experiments demonstrate the effectiveness of our method to obtain multiple shorter and interpretable solutions. Code is available after acceptance. Lingling Zhang 0005, Muye Huang, Qianying Wang 0002, Jun Liu 0002 |
NeurIPS | 5 |
| 2025 | A Dual-Stick Controller for Enhancing Raycasting Interactions with Virtual ObjectsabstractThis work presents Dual-Stick, a novel controller with two sticks connected at the end that innovates a Dual-Ray interaction paradigm to enrich raycasting input in Virtual Reality (VR). Dual-Stick leverages the inherent human dexterity in using everyday tools such as clamps and tweezers to adjust the relative angle between two sticks. This design supports Dual-Ray interactions that provide with a heuristics-based enhanced mechanism. It also offers more flexible manipulation by taking advantages of additional degrees of freedom provided by clamping angle. We conducted two studies to evaluate the effectiveness of Dual-Ray in target selection and manipulation tasks. The results indicated that Dual-Ray significantly improved efficiency in target selection compared to single-ray input but did not outperform the enhanced single-ray technique. In terms of manipulation, Dual-Ray effectively reduced completion time and mode switching compared to single-ray input. Nianlong Li, Zhenxuan He, Luyao Shen, Tianren Luo, Teng Han, Boyu Gao 0003, Yu Zhang 0199, Liuxin Zhang, Feng Tian 0001, Qianying Wang 0002 |
VR | 11 |
| 2025 | TDGI: Translation-Guided Double-Graph Inference for Document-Level Relation ExtractionabstractDocument-level relation extraction (DocRE) aims at predicting relations of all entity pairs in one document, which plays an important role in information extraction. DocRE is more challenging than previous sentence-level relation extraction, as it often requires coreference and logical reasoning across multiple sentences. Graph-based methods are the mainstream solution to this complex reasoning in DocRE. They generally construct the heterogeneous graphs with entities, mentions, and sentences as nodes, co-occurrence and co-reference relations as edges. Their performance is difficult to further break through because the semantics and direction of the relation are not jointly considered in graph inference process. To this end, we propose a novel translation-guided double-graph inference network named TDGI for DocRE. On one hand, TDGI includes two relation semantics-aware and direction-aware reasoning graphs, i.e., mention graph and entity graph, to mine relations among long-distance entities more explicitly. Each graph consists of three elements: vectorized nodes, edges, and direction weights. On the other hand, we devise an interesting translation-based graph updating strategy that guides the embeddings of mention/entity nodes, relation edges, and direction weights following the specific translation algebraic structure, thereby to enhance the reasoning skills of TDGI. In the training procedure of TDGI, we minimize the relation multi-classification loss and triple contrastive loss together to guarantee the model's stability and robustness. Comprehensive experiments on three widely-used datasets show that TDGI achieves outstanding performance comparing with state-of-the-art baselines. Lingling Zhang 0005, Jun Liu 0002, Qianying Wang 0002, Jiaxin Wang 0002, Xiaojun Chang |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | Hierarchy-Based Diagram-Sentence Matching on Dual-Modal Graphs
Lingling Zhang 0005, Jun Liu 0002, Jiaxin Wang 0002, Qianying Wang 0002 |
Pattern Recognit. | 7 |
| 2024 | Transfer and Alignment Network for Generalized Category DiscoveryabstractGeneralized Category Discovery (GCD) is a crucial real-world task that aims to recognize both known and novel categories from an unlabeled dataset by leveraging another labeled dataset with only known categories. Despite the improved performance on known categories, current methods perform poorly on novel categories. We attribute the poor performance to two reasons: biased knowledge transfer between labeled and unlabeled data and noisy representation learning on the unlabeled data. The former leads to unreliable estimation of learning targets for novel categories and the latter hinders models from learning discriminative features. To mitigate these two issues, we propose a Transfer and Alignment Network (TAN), which incorporates two knowledge transfer mechanisms to calibrate the biased knowledge and two feature alignment mechanisms to learn discriminative features. Specifically, we model different categories with prototypes and transfer the prototypes in labeled data to correct model bias towards known categories. On the one hand, we pull instances with known categories in unlabeled data closer to these prototypes to form more compact clusters and avoid boundary overlap between known and novel categories. On the other hand, we use these prototypes to calibrate noisy prototypes estimated from unlabeled data based on category similarities, which allows for more accurate estimation of prototypes for novel categories that can be used as reliable learning targets later. After knowledge transfer, we further propose two feature alignment mechanisms to acquire both instance- and category-level knowledge from unlabeled data by aligning instance features with both augmented features and the calibrated prototypes, which can boost model performance on both known and novel categories with less noise. Experiments on three benchmark datasets show that our model outperforms SOTA methods, especially on novel categories. Theoretical analysis is provided for an in-depth understanding of our model in general. Our code and data are available at https://github.com/Lackel/TAN. Wenbin An, Feng Tian 0002, Wenkai Shi, Yan Chen 0031, Yaqiang Wu, Qianying Wang 0002, Ping Chen 0001 |
AAAI | 6 |
| 2024 | A Unified Knowledge Transfer Network for Generalized Category DiscoveryabstractGeneralized Category Discovery (GCD) aims to recognize both known and novel categories in an unlabeled dataset by leveraging another labeled dataset with only known categories. Without considering knowledge transfer from known to novel categories, current methods usually perform poorly on novel categories due to the lack of corresponding supervision. To mitigate this issue, we propose a unified Knowledge Transfer Network (KTN), which solves two obstacles to knowledge transfer in GCD. First, the mixture of known and novel categories in unlabeled data makes it difficult to identify transfer candidates (i.e., samples with novel categories). For this, we propose an entropy-based method that leverages knowledge in the pre-trained classifier to differentiate known and novel categories without requiring extra data or parameters. Second, the lack of prior knowledge of novel categories presents challenges in quantifying semantic relationships between categories to decide the transfer weights. For this, we model different categories with prototypes and treat their similarities as transfer weights to measure the semantic similarities between categories. On the basis of two treatments, we transfer knowledge from known to novel categories by conducting pre-adjustment of logits and post-adjustment of labels for transfer candidates based on the transfer weights between different categories. With the weighted adjustment, KTN can generate more accurate pseudo-labels for unlabeled data, which helps to learn more discriminative features and boost model performance on novel categories. Extensive experiments show that our method outperforms state-of-the-art models on all evaluation metrics across multiple benchmark datasets. Furthermore, different from previous clustering-based methods that can only work offline with abundant data, KTN can be deployed online conveniently with faster inference speed. Code and data are available at https://github.com/yibai-shi/KTN. Wenkai Shi, Wenbin An, Feng Tian 0002, Yan Chen 0031, Yaqiang Wu, Qianying Wang 0002, Ping Chen 0001 |
AAAI | 6 |
| 2024 | See Widely, Think Wisely: Toward Designing a Generative Multi-agent System to Burst Filter BubblesabstractThe proliferation of AI-powered search and recommendation systems has accelerated the formation of “filter bubbles” that reinforce people’s biases and narrow their perspectives. Previous research has attempted to address this issue by increasing the diversity of information exposure, which is often hindered by a lack of user motivation to engage with. In this study, we took a human-centered approach to explore how Large Language Models (LLMs) could assist users in embracing more diverse perspectives. We developed a prototype featuring LLM-powered multi-agent characters that users could interact with while reading social media content. We conducted a participatory design study with 18 participants and found that multi-agent dialogues with gamification incentives could motivate users to engage with opposing viewpoints. Additionally, progressive interactions with assessment tasks could promote thoughtful consideration. Based on these findings, we provided design implications with future work outlooks for leveraging LLMs to help users burst their filter bubbles. Yu Zhang 0124, Jingwei Sun 0005, Cen Yao, Mingming Fan 0001, Liuxin Zhang, Qianying Wang 0002, Xin Geng 0001, Yong Rui |
CHI | 7 |
| 2024 | E-GPS: Explainable Geometry Problem Solving via Top-Down Solver and Bottom-Up GeneratorabstractGeometry Problem Solving has drawn growing attention recently due to its application prospects in intelligent ed-ucation field. However, existing methods are still inade-quate to meet the needs of practical application, suffering from the following limitations: 1) explainability is not en-sured which is essential in real teaching scenarios; 2) the small scale and incomplete annotation of existing datasets make it hard for model to comprehend geometric knowl-edge. To tackle the above problems, we propose a novel method called Explainable Geometry Problem Solving (E-GPS). E-GPS first parses the geometric diagram and prob-lem text into unified formal language representations. Then, the answer and explainable reasoning and solving steps are obtained by a Top-Down Problem Solver (TD-PS), which innovatively solves the problem from the target and focuses on what is needed. To alleviate the data issues, a Bottom-Up Problem Generator (BU-PG) is devised to augment the data set with various well-annotated constructed geome-try problems. It enables us to train an enhanced theorem predictor with a better grasp of theorem knowledge, which further improves the efficiency ofTD-PS. Extensive experi-ments demonstrate that E-GPS maintains comparable solving performances with fewer steps and provides outstanding explainability. Lingling Zhang 0005, Jun Liu 0002, Yaxian Wang, Qianying Wang 0002 |
CVPR | 7 |
| 2024 | A Multi-Scale Bimodal Fusion Network for Robust and Accurate Online Handwriting RecognitionabstractOnline handwriting recognition based on sensor trajectory information faces several unresolved challenges: 1) sensor signals lack sufficient global spatial context; 2) different recognition tasks have inconsistent requirements for feature receptive fields. This is due to the inconsistent scales of the input sequences and the different semantic complexity of different language units. In this paper, we propose an online handwritten text recognition method based on multi-scale bimodal feature fusion to address these challenges. First, we employ sequence-generated pseudo-images to supplement the two-dimensional spatial information, and then extract multi-scale features from both trajectories and images simultaneously. Subsequently, our designed bimodal embedding learning module jointly learns feature embeddings for trajectories and images at different scales. These embeddings are then fed into a novel position-aware multi-scale fusion module to extract features for text prediction. The proposed modules effectively mitigate the issues of scales and semantics misalignment. Experimental results demonstrate significant performance improvements on various handwriting recognition datasets using our approach. Yaqiang Wu, Wanjun Lv, Qianying Wang 0002 |
ICASSP | 7 |
| 2024 | A Tri-Branch Network with Prototype-aware Matching for Universal Category DiscoveryabstractIn this paper, we propose a novel task, Universal Category Discovery (UCD), to address the challenge of partial overlap between source and target domain categories. Different from previous tasks that assume all known categories exist in the target domain, UCD introduces "private-known" categories that only exist in the source domain and aims to classify unlabeled data as "common" or "novel" categories while avoiding misclassifying them into "private-known" categories. For this task, we propose a Tri-branch network with bidirectional Prototype-aware Matching (TriPM). TriPM effectively transfers knowledge from labeled to unlabeled data by bidirectionally matching similar data pairs, while a prototype matching strategy reduces the negative transfer risk from "private-known" categories. Finally, we propose a tri-branch network to decouple knowledge acquisition from labeled data, unlabeled data, and their interactions, which can avoid knowledge forgetting, explore novel patterns, and transfer common knowledge, respectively. Experiments demonstrate our model’s superiority over SOTA methods. Haonan Lin, Wenbin An, Yan Chen 0031, Feng Tian 0002, Yuzhe Yao, Wei Ding 0003, Qianying Wang 0002, Ping Chen 0001 |
ICME | 7 |
| 2024 | Schedule Your Edit: A Simple yet Effective Diffusion Noise Schedule for Image EditingabstractText-guided diffusion models have significantly advanced image editing, enabling high-quality and diverse modifications driven by text prompts. However, effective editing requires inverting the source image into a latent space, a process often hindered by prediction errors inherent in DDIM inversion.
These errors accumulate during the diffusion process, resulting in inferior content preservation and edit fidelity, especially with conditional inputs.
We address these challenges by investigating the primary contributors to error accumulation in DDIM inversion and identify the singularity problem in traditional noise schedules as a key issue.
To resolve this, we introduce the *Logistic Schedule*, a novel noise schedule designed to eliminate singularities, improve inversion stability, and provide a better noise space for image editing. This schedule reduces noise prediction errors, enabling more faithful editing that preserves the original content of the source image. Our approach requires no additional retraining and is compatible with various existing editing methods.
Experiments across eight editing tasks demonstrate the Logistic Schedule's superior performance in content preservation and edit fidelity compared to traditional noise schedules, highlighting its adaptability and effectiveness.
The project page is available at https://lonelvino.github.io/SYE/. Haonan Lin, Yan Chen 0031, Jiahao Wang 0004, Wenbin An, Mengmeng Wang 0005, Feng Tian 0002, Yong Liu 0007, Guang Dai, Jingdong Wang 0001, Qianying Wang 0002 |
NeurIPS | 10 |
| 2024 | Flipped Classroom: Aligning Teacher Attention with Student in Generalized Category DiscoveryabstractRecent advancements have shown promise in applying traditional Semi-Supervised Learning strategies to the task of Generalized Category Discovery (GCD). Typically, this involves a teacher-student framework in which the teacher imparts knowledge to the student to classify categories, even in the absence of explicit labels. Nevertheless, GCD presents unique challenges, particularly the absence of priors for new classes, which can lead to the teacher's misguidance and unsynchronized learning with the student, culminating in suboptimal outcomes. In our work, we delve into why traditional teacher-student designs falter in generalized category discovery as compared to their success in closed-world semi-supervised learning. We identify inconsistent pattern learning as the crux of this issue and introduce FlipClass—a method that dynamically updates the teacher to align with the student's attention, instead of maintaining a static teacher reference. Our teacher-attention-update strategy refines the teacher's focus based on student feedback, promoting consistent pattern recognition and synchronized learning across old and new classes. Extensive experiments on a spectrum of benchmarks affirm that FlipClass significantly surpasses contemporary GCD methods, establishing new standards for the field. Haonan Lin, Wenbin An, Jiahao Wang 0004, Yan Chen 0031, Feng Tian 0002, Mengmeng Wang 0005, Qianying Wang 0002, Guang Dai, Jingdong Wang 0001 |
NeurIPS | 7 |
| 2024 | Multi-view cognition with path search for one-shot part labeling
Lingling Zhang 0005, Tao Qin 0002, Jun Liu 0002, Yifei Li 0006, Qianying Wang 0002 |
Comput. Vis. Image Underst. | 6 |
| 2024 | Towards Workplace Metaverse: A Human-Centered Approach for Designing and Evaluating XR Virtual DisplaysabstractWork is becoming more and more flexible nowadays. It can take place in the office, at home, or even on the go; tasks may encompass activities in the form of documents, multimedia, or 3D models in virtual space. Under such circumstances, personal computers (PC), the most widely used productivity devices for work today, cannot well address the diverse needs such as screen size, privacy, and flexibility to display diverse content formats (e.g., 2D to 3D), due to their fixed hardware specs. We believe the solution lies in Extended Reality (XR). In this article, we explored how XR glasses can be used for PC’s virtual extended displays, and conducted user interviews and usability tests to propose a systematic user experience design and evaluation framework. We discovered that the design space encompasses four dimensions general placement, display specs, operating system integration, and interaction behaviors) and summarized users’ corresponding preferences. We proposed a quality-of-experience (QoE) evaluation framework for XR virtual displays consisting of visual quality, visual fatigue and discomfort, as well as immersiveness, and identified clarity as the most significant factor that affects user satisfaction. Our design and evaluation frameworks could serve as a resource for both practitioners and scholars with an interest in the design and evaluation of virtual displays. Yu Zhang 0124, Jingwei Sun 0005, Qicheng Ding, Liuxin Zhang, Qianying Wang 0002, Xin Geng 0001, Yong Rui |
Int. J. Hum. Comput. Interact. | 5 |
| 2024 | Multi-focus image fusion with parameter adaptive dual channel dynamic threshold neural P systems
Lingling Zhang 0005, Jun Liu 0002, Qianying Wang 0002 |
Neural Networks | 5 |
| 2024 | Alignment Relation is What You Need for Diagram ParsingabstractAs a knowledge carrier, the diagram is widely distributed in many aspects of human life, such as textbooks, architectural drawings, and documents. Different from natural images, representations of visual elements in the diagram are sparser, and similar visual representations can reflect dissimilar semantics. Thus, current methods fail to capture the visual elements with precise semantics. To address this issue, regarding the aligned visual and textual elements as pairs is the way to assign the precise semantics of textual elements to visual elements. We build the first diagram dataset named align diagram element (ADE), which includes annotations for alignment relations between visual and textual elements. And we propose a visual-textual alignment model (VTAM) including graph construction and optimal aligning phases. In the graph construction phase, the relational graphs are constructed between different elements with four relational operators. The relational operators are designed to measure the relations between different elements, according to distance, connection line, inclusion, and feature similarity. In the optimal aligning phase, the representation at each visual-textual pair is improved as a weighted sum of the representations on all relational graphs. Experimental results show that our VTAM achieves a significant improvement of 10.9% on mean test folds of the ADE dataset than the current best competitor. In order to explore the role of alignment relations in diagram parsing, we introduce VTAM to diagram-related tasks, such as diagram question answering (DQA). And we achieve 2.8% to 5.9% and 4.6% to 5.1% improvements on AI2D and Foodwebs after adding VTAM. Our dataset and code are released at: https://github.com/ADE-dataset/ADE-dataset. Xinyu Zhang 0021, Lingling Zhang 0005, Jun Liu 0002, Qianying Wang 0002 |
IEEE Trans. Image Process. | 6 |
| 2024 | Context-Aware Commonsense Knowledge Graph Reasoning With Path-Guided ExplanationsabstractCommonsense knowledge graphs (CKGs) store massive commonsense knowledge as triples whose nodes consist of free-form texts. CKG reasoning aims to predict missing nodes in incomplete commonsense triples, which is challenging as it requires more accurate embeddings for reasoning. Compared to conventional knowledge graphs (KGs), CKGs have deficient structural information due to their sparsity and contain nodes indistinguishable due to the conceptual diversity. These issues limit the performance of previous reasoning methods, because they face difficulties obtaining precise CKG representations. To address these issues, we propose a context-aware CKG reasoning framework with path-guided explanations, named CoRPe. Firstly, CoRPe constructs context sentences based on the target commonsense triple using designed templates. The context captures reasoning paths instantiated from the first-order logic. Secondly, to improve CKG representations, CoRPe injects context semantics and employs a context-augmented tuning strategy on a pre-trained language model (PLM) via a synergistic optimization. Finally, CoRPe embeds structural information using a graph convolutional network (GCN) and associates the textual semantics for joint scoring. Extensive experiments on two CKGs show that CoRPe outperforms state-of-the-art KG and CKG reasoning baselines in terms of embedding and reasoning performance. Furthermore, the interpretability of CoRPe is reflected in the implicit logic during reasoning. Yudai Pan, Jun Liu 0002, Tianzhe Zhao, Lingling Zhang 0005, Qianying Wang 0002 |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2024 | FPrompt-PLM: Flexible-Prompt on Pretrained Language Model for Continual Few-Shot Relation ExtractionabstractRelation extraction (RE) aims to identify the relation between two entities within a sentence, which plays a crucial role in information extraction. Traditional supervised setting on RE does not fit the actual scenario, due to the continuous emergence of new relations and the unavailability of massive labeled examples. Continual few-shot relation extraction (CFS-RE) is proposed as a potential solution to the above situation, which requires the model to learn new relations sequentially from a few examples. Apparently, CFS-RE is more challenging than previous RE, as the catastrophic forgetting of old knowledge and few-shot overfitting on a handful of examples. To this end, we propose a novel flexible-prompt framework on pretrained language model named FPrompt-PLM for CFS-RE, which includes flexible-prompt embedding, pretrained-language understanding, and nearest-prototype learning modules. Note that two pools in FPrompt-PLM, i.e., prompt and prototype pools, are continual updated and applied for prediction of all seen relations at current time-step. The former pool records the distinctive prompt embedding in each time period, and the latter records all learned relation prototypes. Besides, three progressive stages are introduced to learn FPrompt-PLM's parameters and apply this model for CFS-RE testing, which includes meta-training, continual meta-finetuning, and testing stages. And we improve the CFS-RE loss by incorporating multiple distillation losses as well as a novel prototype-diversity loss in these stages to alleviate the catastrophic forgetting and few-shot overfitting problems. Comprehensive experiments on two widely-used datasets show that FPrompt-PLM achieves significant performance improvements over the SOTA baselines. Lingling Zhang 0005, Yifei Li 0006, Qianying Wang 0002, Hang Yan 0010, Jiaxin Wang 0002, Jun Liu 0002 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2023 | Generalized Category Discovery with Decoupled Prototypical NetworkabstractGeneralized Category Discovery (GCD) aims to recognize both known and novel categories from a set of unlabeled data, based on another dataset labeled with only known categories. Without considering differences between known and novel categories, current methods learn about them in a coupled manner, which can hurt model's generalization and discriminative ability. Furthermore, the coupled training approach prevents these models transferring category-specific knowledge explicitly from labeled data to unlabeled data, which can lose high-level semantic information and impair model performance. To mitigate above limitations, we present a novel model called Decoupled Prototypical Network (DPN). By formulating a bipartite matching problem for category prototypes, DPN can not only decouple known and novel categories to achieve different training targets effectively, but also align known categories in labeled and unlabeled data to transfer category-specific knowledge explicitly and capture high-level semantics. Furthermore, DPN can learn more discriminative features for both known and novel categories through our proposed Semantic-aware Prototypical Learning (SPL). Besides capturing meaningful semantic information, SPL can also alleviate the noise of hard pseudo labels through semantic-weighted soft assignment. Extensive experiments show that DPN outperforms state-of-the-art models by a large margin on all evaluation metrics across multiple benchmark datasets. Code and data are available at https://github.com/Lackel/DPN. Wenbin An, Feng Tian 0002, Wei Ding 0003, Qianying Wang 0002, Ping Chen 0001 |
AAAI | 5 |
| 2023 | DNA: Denoised Neighborhood Aggregation for Fine-grained Category DiscoveryabstractDiscovering fine-grained categories from coarsely labeled data is a practical and challenging task, which can bridge the gap between the demand for fine-grained analysis and the high annotation cost.Previous works mainly focus on instance-level discrimination to learn low-level features, but ignore semantic similarities between data, which may prevent these models learning compact cluster representations.In this paper, we propose Denoised Neighborhood Aggregation (DNA), a self-supervised framework that encodes semantic structures of data into the embedding space.Specifically, we retrieve k-nearest neighbors of a query as its positive keys to capture semantic similarities between data and then aggregate information from the neighbors to learn compact cluster representations, which can make fine-grained categories more separatable.However, the retrieved neighbors can be noisy and contain many false-positive keys, which can degrade the quality of learned embeddings.To cope with this challenge, we propose three principles to filter out these false neighbors for better representation learning.Furthermore, we theoretically justify that the learning objective of our framework is equivalent to a clustering loss, which can capture semantic similarities between data to form compact fine-grained clusters.Extensive experiments on three benchmark datasets show that our method can retrieve more accurate neighbors (21.31% accuracy improvement) and outperform state-of-the-art models by a large margin (average 9.96% improvement on three metrics).Our code and data are available at https://github.com/Lackel/DNA. Wenbin An, Feng Tian 0002, Wenkai Shi, Yan Chen 0031, Qianying Wang 0002, Ping Chen 0001 |
EMNLP | 6 |
| 2023 | A Diffusion Weighted Graph Framework for New Intent DiscoveryabstractNew Intent Discovery (NID) aims to recognize both new and known intents from unlabeled data with the aid of limited labeled data containing only known intents.Without considering structure relationships between samples, previous methods generate noisy supervisory signals which cannot strike a balance between quantity and quality, hindering the formation of new intent clusters and effective transfer of the pre-training knowledge.To mitigate this limitation, we propose a novel Diffusion Weighted Graph Framework (DWGF) to capture both semantic similarities and structure relationships inherent in data, enabling more sufficient and reliable supervisory signals.Specifically, for each sample, we diffuse neighborhood relationships along semantic paths guided by the nearest neighbors for multiple hops to characterize its local structure discriminately.Then, we sample its positive keys and weigh them based on semantic similarities and local structures for contrastive learning.During inference, we further propose Graph Smoothing Filter (GSF) to explicitly utilize the structure relationships to filter high-frequency noise embodied in semantically ambiguous samples on the cluster boundary.Extensive experiments show that our method outperforms state-of-the-art models on all evaluation metrics across multiple benchmark datasets. Wenkai Shi, Wenbin An, Feng Tian 0002, Qianying Wang 0002, Ping Chen 0001 |
EMNLP | 5 |
| 2023 | SHGAE: Social Hypergraph AutoEncoder for Friendship Inference
Yan Chen 0031, Tianliang Qi, Feng Tian 0002, Yaqiang Wu, Qianying Wang 0002 |
ICANN (6) | 6 |
| 2023 | Diagram Visual Grounding: Learning to See with Gestalt-Perceptual AttentionabstractDiagram visual grounding aims to capture the correlation between language expression and local objects in the diagram, and plays an important role in the applications like textbook question answering and cross-modal retrieval. Most diagrams consist of several colors and simple geometries. This results in sparse low-level visual features, which further aggravates the gap between low-level visual and high-level semantic features of diagrams. The phenomenon brings challenges to the diagram visual grounding. To solve the above issues, we propose a gestalt-perceptual attention model to align the diagram objects and language expressions. For low-level visual features, inspired by the gestalt that simulates human visual system, we build a gestalt-perception graph network to make up the features learned by the traditional backbone network. For high-level semantic features, we design a multi-modal context attention mechanism to facilitate the interaction between diagrams and language expressions, so as to enhance the semantics of diagrams. Finally, guided by diagram features and linguistic embedding, the target query is gradually decoded to generate the coordinates of the referred object. By conducting comprehensive experiments on diagrams and natural images, we demonstrate that the proposed model achieves superior performance over the competitors. Our code will be released at https://github.com/AIProCode/GPA. Lingling Zhang 0005, Jun Liu 0002, Xinyu Zhang 0021, Qianying Wang 0002 |
IJCAI | 6 |
| 2023 | A prediction model of student performance based on self-attention mechanism
Yan Chen 0031, Ganglin Wei, Yunwei Chen, Feng Tian 0002, Qianying Wang 0002, Yaqiang Wu |
Knowl. Inf. Syst. | 8 |
| 2023 | Disentangling interest and conformity for eliminating popularity bias in session-based recommendation
Qidong Liu 0002, Feng Tian 0002, Qianying Wang 0002 |
Knowl. Inf. Syst. | 4 |
| 2023 | Side-by-Side vs Face-to-Face: Evaluating Colocated Collaboration via a Transparent Wall-sized DisplayabstractTraditional wall-sized displays mostly only support side-by-side co-located collaboration, while transparent displays naturally support face-to-face interaction. Many previous works assume transparent displays support collaboration. Yet it is unknown how exactly its afforded face-to-face interaction can support loose or close collaboration, especially compared to the side-by-side configuration offered by traditional large displays. In this paper, we used an established experimental task that operationalizes different collaboration coupling and layout locality, to compare pairs of participants collaborating side-by-side versus face-to-face in each collaborative situation. We compared quantitative measures and collected interview and observation data to further illustrate and explain our observed user behavior patterns. The results showed that the unique face-to-face collaboration brought by transparent display can result in more efficient task performance, different territorial behavior, and both positive and negative collaborative factors. Our findings provided empirical understanding about the collaborative experience supported by wall-sized transparent displays and shed light on its future design. Jiangtao Gong, Mengdi Chu, Minghao Luo, Liuxin Zhang, Yaqiang Wu, Qianying Wang 0002, Can Liu 0003 |
Proc. ACM Hum. Comput. Interact. | 9 |
| 2023 | MoCA: Incorporating domain pretraining and cross attention for textbook question answering
Fangzhi Xu, Qika Lin, Jun Liu 0002, Lingling Zhang 0005, Tianzhe Zhao, Qi Chai, Yudai Pan, Yi Huang 0017, Qianying Wang 0002 |
Pattern Recognit. | 9 |
| 2023 | RPMG-FSS: Robust Prior Mask Guided Few-Shot Semantic SegmentationabstractFew-shot semantic segmentation (FSS) has been developed to perform pixel-level segmentation with only a few dense labeled examples for training, which relieves the expensive annotation problem in traditional segmentation models. Current researches on FSS generally act the labeled masks on the corresponding support images to obtain the class-specific embeddings, and predict the pixel-level masks for query images by matching their pixels to these class-specific embeddings. Their performance is difficult to further break through because of the limited supervision from single-view support images and the neglect of position information from similar pixels between query and support images. To solve these issues, we propose a novel robust prior mask guided model named RPMG-FSS for the challenging FSS task. The core of RPMG-FSS is to produce a robust prior mask with good generalization ability on novel classes to better assist the following query mask prediction. Note that each element in the prior mask corresponds to one pixel in query image. It not only considers the interaction within one view and between multiple views of the support image, but also fuses the top-$k$similarity values to all support pixels and these pixels’ position information. The parameters in RPMG-FSS are optimized with the combination of segmentation loss and multi-view contrastive loss. Comprehensive experiments on two datasets show that our RPMG-FSS achieves outstanding performance comparing with the current popular baselines. The code is released onhttps://github.com/dxzxy12138/RPMG-FSS/tree/master Lingling Zhang 0005, Xinyu Zhang 0021, Qianying Wang 0002, Xiaojun Chang, Jun Liu 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | DisAVR: Disentangled Adaptive Visual Reasoning Network for Diagram Question AnsweringabstractDiagram Question Answering (DQA) aims to correctly answer questions about given diagrams, which demands an interplay of good diagram understanding and effective reasoning. However, the same appearance of objects in diagrams can express different semantics. This kind of visual semantic ambiguity problem makes it challenging to represent diagrams sufficiently for better understanding. Moreover, since there are questions about diagrams from different perspectives, it is also crucial to perform flexible and adaptive reasoning on content-rich diagrams. In this paper, we propose a Disentangled Adaptive Visual Reasoning Network for DQA, named DisAVR, to jointly optimize the dual-process of representation and reasoning. DisAVR mainly comprises three modules: improved region feature learning, question parsing, and disentangled adaptive reasoning. Specifically, the improved region feature learning module is designed to first learn robust diagram representation by integrating detail-aware patch features and semantically-explicit text features with region features. Subsequently, the question parsing module decomposes the question into three types of question guidance including region, spatial relation and semantic relation guidance to dynamically guide subsequent reasoning. Next, the disentangled adaptive reasoning module decomposes the whole reasoning process by employing three visual reasoning cells to construct a soft fully-connected multi-layer stacked routing space. These three cells in each layer reason over object regions, semantic and spatial relations in the diagram under the corresponding question guidance. Moreover, an adaptive routing mechanism is designed to flexibly explore more optimal reasoning paths for specific diagram-question pairs. Extensive experiments on three DQA datasets demonstrate the superiority of our DisAVR. Yaxian Wang, Bifan Wei, Jun Liu 0002, Lingling Zhang 0005, Jiaxin Wang 0002, Qianying Wang 0002 |
IEEE Trans. Image Process. | 6 |
| 2022 | Remote Co-teaching in Rural Classroom: Current Practices, Impacts, and ChallengesabstractThe shortage of high-quality teachers is one of the biggest educational problems faced by underdeveloped areas. With the development of information and communication technologies (ICTs), China has begun a remote co-teaching intervention program using ICTs for rural classes, forming a unique “co-teaching classroom”. We conducted semi-structured interviews with nine remote urban teachers and twelve local rural teachers. We identified the remote co-teaching classes’ standard practices and co-teachers’ collaborative work process. We also found that remote teachers’ high-quality class directly impacted local teachers and students. Furthermore, interestingly, local teachers were also actively involved in making indirect impacts on their students by deeply coordinating with remote teachers and adapting the resources offered by the remote teachers. We conclude by summarizing and discussing the challenges faced by teachers, lessons learned from the current program, and related design implications to achieve a more adaptive and sustainable ICT4D program design. Siling Guo, Tianchen Sun, Jiangtao Gong, Zhicong Lu, Liuxin Zhang, Qianying Wang 0002 |
CHI | 6 |
| 2022 | Fine-grained Category Discovery under Coarse-grained supervision with Hierarchical Weighted Self-contrastive LearningabstractNovel category discovery aims at adapting models trained on known categories to novel categories.Previous works only focus on the scenario where known and novel categories are of the same granularity.In this paper, we investigate a new practical scenario called Fine-grained Category Discovery under Coarsegrained supervision (FCDC).FCDC aims at discovering fine-grained categories with only coarse-grained labeled data, which can adapt models to categories of different granularity from known ones and reduce significant labeling cost.It is also a challenging task since supervised training on coarse-grained categories tends to focus on inter-class distance (distance between coarse-grained classes) but ignore intra-class distance (distance between fine-grained sub-classes) which is essential for separating fine-grained categories.Considering most current methods cannot transfer knowledge from coarse-grained level to fine-grained level, we propose a hierarchical weighted self-contrastive network by building a novel weighted self-contrastive module and combining it with supervised learning in a hierarchical manner.Extensive experiments on public datasets show both effectiveness and efficiency of our model over compared methods.Code and data are available at https://github.com/Lackel/ Hierarchical_Weighted_SCL. Wenbin An, Feng Tian 0002, Ping Chen 0001, Siliang Tang, Qianying Wang 0002 |
EMNLP | 6 |
| 2022 | Inductive Relation Prediction with Logical Reasoning Using Contrastive RepresentationsabstractRelation prediction in knowledge graphs (KGs) aims at predicting missing relations in incomplete triples, whereas the dominant embedding paradigm has a restriction on handling unseen entities during testing.In the realworld scenario, the inductive setting is more common because entities in the training process are finite.Previous methods capture an inductive ability by implicit logic in KGs.However, it would be challenging to preciously acquire entity-independent relational semantics of compositional logic rules and to deal with the deficient supervision of logic caused by the scarcity of relational semantics.To this end, we propose a novel graph convolutional network (GCN)-based model LogCo with logical reasoning by contrastive representations.LogCo firstly extracts enclosing subgraphs and relational paths between two entities to supply the entity-independence.Then a contrastive strategy for relational path instances and the subgraph is proposed for the issue of deficient supervision.The contrastive representations are learned for a joint training regime.Finally, prediction results and logic rules for reasoning are attained.Comprehensive experiments on twelve inductive datasets show that LogCo achieves outstanding performance comparing with SOTA inductive baselines. Yudai Pan, Jun Liu 0002, Lingling Zhang 0005, Tianzhe Zhao, Qika Lin, Qianying Wang 0002 |
EMNLP | 7 |
| 2021 | Towards Robust Visual Information Extraction in Real World: New Dataset and Novel SolutionabstractVisual Information Extraction (VIE) has attracted considerable attention recently owing to its various advanced applications such as document understanding, automatic marking and intelligent education. Most existing works decoupled this problem into several independent sub-tasks of text spotting (text detection and recognition) and information extraction, which completely ignored the high correlation among them during optimization. In this paper, we propose a robust Visual Information Extraction System (VIES) towards real-world scenarios, which is an unified end-to-end trainable framework for simultaneous text detection, recognition and information extraction by taking a single document image as input and outputting the structured information. Specifically, the information extraction branch collects abundant visual and semantic representations from text spotting for multimodal feature fusion and conversely, provides higher-level semantic clues to contribute to the optimization of text spotting. Moreover, regarding the shortage of public benchmarks, we construct a fully-annotated dataset called EPHOIE (https://github.com/HCIILAB/EPHOIE), which is the first Chinese benchmark for both text spotting and visual information extraction. EPHOIE consists of 1,494 images of examination paper head with complex layouts and background, including a total of 15,771 Chinese handwritten or printed text instances. Compared with the state-of-the-art methods, our VIES shows significant superior performance on the EPHOIE dataset and achieves a 9.01% F-score gain on the widely used SROIE dataset under the end-to-end scenario. Chongyu Liu, Guozhi Tang, Jiaxin Zhang 0003, Shuaitao Zhang, Qianying Wang 0002, Yaqiang Wu, Mingxiang Cai |
AAAI | 7 |
| 2021 | All in One Group: Current Practices, Lessons and Challenges of Chinese Home-School Communication in IM Group ChatabstractWhen schools and families form a good partnership, children benefit. With the recent flourishing of communication apps, families and schools in China have shifted their primary communication channels to chat groups hosted on popular instant-messenger(IM) tools such as WeChat and QQ. With an interview study consisting of 18 parents and 9 teachers, followed by a survey study with 210 teachers, we found that IM group chat has become the most popular way that the majority of parents and teachers communicate, from among the many different channels available. While there are definite advantages to this kind of group chat, we also found a number of problematic issues, including a lack of privacy and repeated negative feedback shared by both parents and teachers. We discuss our results on how IM-based group chat could affect Chinese teachers’ authoritative figures, affect Chinese teacher’s work-life balance and potentially compromise Chinese students’ privacy. Jiangtao Gong, Zhicong Lu, Qicheng Ding, Yu Zhang 0124, Liuxin Zhang, Qianying Wang 0002 |
CHI | 7 |
| 2021 | MatchVIE: Exploiting Match Relevancy between Entities for Visual Information ExtractionabstractVisual Information Extraction (VIE) task aims to extract key information from multifarious document images (e.g., invoices and purchase receipts). Most previous methods treat the VIE task simply as a sequence labeling problem or classification problem, which requires models to carefully identify each kind of semantics by introducing multimodal features, such as font, color, layout. But simply introducing multimodal features can't work well when faced with numeric semantic categories or some ambiguous texts. To address this issue, in this paper we propose a novel key-value matching model based on a graph neural network for VIE (MatchVIE). Through key-value matching based on relevancy evaluation, the proposed MatchVIE can bypass the recognitions to various semantics, and simply focuses on the strong relevancy between entities. Besides, we introduce a simple but effective operation, Num2Vec, to tackle the instability of encoded values, which helps model converge more smoothly. Comprehensive experiments demonstrate that the proposed MatchVIE can significantly outperform previous methods. Notably, to the best of our knowledge, MatchVIE may be the first attempt to tackle the VIE task by modeling the relevancy between keys and values and it is a good complement to the existing methods. Guozhi Tang, Lele Xie, Jingdong Chen, Qianying Wang 0002, Yaqiang Wu |
IJCAI | 7 |
| 2021 | HoloBoard: a Large-format Immersive Teaching Board based on pseudo HoloGraphicsabstractIn this paper, we present HoloBoard, an interactive large-format pseduo-holographic display system for lecture based classes. With its unique properties of immersive visual display and transparent screen, we designed and implemented a rich set of novel interaction techniques like immersive presentation, role-play, and lecturing behind the scene that are potentially valuable for lecturing in class. We conducted a controlled experimental study to compare a HoloBoard class with a normal class through measuring students’ learning outcomes and three dimensions of engagement (i.e., behavioral, emotional, and cognitive engagement). We used pre-/post- knowledge tests and multimodal learning analytics to measure students’ learning outcomes and learning experiences. Results indicated that the lecture-based class utilizing HoloBoard lead to slightly better learning outcomes and a significantly higher level of student engagement. Given the results, we discussed the impact of HoloBoard as an immersive media in the classroom setting and suggest several design implications for deploying HoloBoard in immersive teaching practices. Jiangtao Gong, Teng Han, Siling Guo, Jiannan Li, Siyu Zha, Liuxin Zhang, Feng Tian 0001, Qianying Wang 0002, Yong Rui |
UIST | 8 |
| 2021 | Grabbing the Long Tail: A data normalization method for diverse and informative dialogue generation
Zhiqiang Zhan, Yang Zhang 0002, Jiangtao Gong, Qianying Wang 0002, Liuxin Zhang |
Neurocomputing | 5 |
| 2020 | Decoupled Attention Network for Text RecognitionabstractText recognition has attracted considerable research interests because of its various applications. The cutting-edge text recognition methods are based on attention mechanisms. However, most of attention methods usually suffer from serious alignment problem due to its recurrency alignment operation, where the alignment relies on historical decoding results. To remedy this issue, we propose a decoupled attention network (DAN), which decouples the alignment operation from using historical decoding results. DAN is an effective, flexible and robust end-to-end text recognizer, which consists of three components: 1) a feature encoder that extracts visual features from the input image; 2) a convolutional alignment module that performs the alignment operation based on visual features from the encoder; and 3) a decoupled text decoder that makes final prediction by jointly using the feature map and attention maps. Experimental results show that DAN achieves state-of-the-art performance on multiple text recognition tasks, including offline handwritten text recognition and regular/irregular scene text recognition. Codes will be released.1 Canjie Luo, Xiaoxue Chen, Yaqiang Wu, Qianying Wang 0002, Mingxiang Cai |
AAAI | 7 |
| 2019 | Modeling Human Intelligence in Customer-Agent Conversation Using Fine-Grained Dialogue Acts
Qicheng Ding, Guoguang Zhao, Penghui Xu, Yucheng Jin 0001, Yu Zhang 0124, Changjian Hu, Qianying Wang 0002 |
NLPCC (2) | 7 |