EDBT 2026 Demo / reviewers in the wild / expert
Jingyan Qin
dblp:153/1300
· DBLP profile ↗
31ranked-venue papers
0as first author
21since 2021 · last 2026
0000-0002-4101-4316ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 10 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 4 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Dual-Geometry Graph Network: Unifying Local and Global Priors for Few-Shot LearningabstractIn few-shot learning, utilizing local and global geometric priors to capture both subtle local class metrics and coarse global structures within the meta-task are important to obtain discriminative embeddings. However, existing graph-based and curvature-based few-shot approaches only focus on either one kind of geometric prior but neglect the other. To effectively utilize the pros of these two paradigms, we propose a novel Dual-Geometry Graph Network (DGGN) to adaptively integrate the local and global geometric priors via two key pathways. Specifically, the local-wise metric modeling pathway utilizes Ollivier-Ricci curvature to capture task-specific local class metrics among the instances, and the global-wise connectivity modeling pathway utilizes resistive embedding to capture global instance distributions and connectivity patterns of the entire meta-task. In addition, we introduce two new regularization loss functions to explicitly enhance the geometric representation ability of the local and global pathways respectively. We validate that DGGN's superior performance stems from its adaptively topological refinements by measuring the graph edit distance, demonstrating its ability to match the underlying data distribution. Extensive experiments show that DGGN sets a new state-of-the-art on standard, cross-domain, and semi-supervised few-shot benchmarks. Xiaobin Zhu 0001, Jingyan Qin, Xu-Cheng Yin |
AAAI | 4 |
| 2026 | Degradation decomposition learning for self-supervised blind image super-resolution
Xiaobin Zhu 0001, Liuling Chen, Jingyan Qin, Xu-Cheng Yin |
Pattern Recognit. | 5 |
| 2025 | Aligning enhanced feature representation for generalized zero-shot learning
Zhiyu Fang, Xiaobin Zhu 0001, Jingyan Qin, Xu-Cheng Yin |
Sci. China Inf. Sci. | 5 |
| 2025 | Multi-Scale Texture Fusion for Reference-Based Image Super-Resolution: New Dataset and Solution
Xiaobin Zhu 0001, Jingyan Qin, Roberto Marcondes Cesar Junior, Xu-Cheng Yin |
Int. J. Comput. Vis. | 3 |
| 2025 | Decoupling and Interaction: task coordination in single-stage object detection
Jia-Wei Ma, Shu Tian, Haixia Man, Song-Lu Chen, Jingyan Qin, Xu-Cheng Yin |
Multim. Tools Appl. | 5 |
| 2024 | Arbitrary Time Information Modeling via Polynomial Approximation for Temporal Knowledge Graph EmbeddingabstractDistinguished from traditional knowledge graphs (KGs), temporal knowledge graphs (TKGs) must explore and reason over temporally evolving facts adequately. However, existing TKG approaches still face two main challenges, i.e., the limited capability to model arbitrary timestamps continuously and the lack of rich inference patterns under temporal constraints. In this paper, we propose an innovative TKGE method (PTBox) via polynomial decomposition-based temporal representation and box embedding-based entity representation to tackle the above-mentioned problems. Specifically, we decompose time information by polynomials and then enhance the model’s capability to represent arbitrary timestamps flexibly by incorporating the learnable temporal basis tensor. In addition, we model every entity as a hyperrectangle box and define each relation as a transformation on the head and tail entity boxes. The entity boxes can capture complex geometric structures and learn robust representations, improving the model’s inductive capability for rich inference patterns. Theoretically, our PTBox can encode arbitrary time information or even unseen timestamps while capturing rich inference patterns and higher-arity relations of the knowledge base. Extensive experiments on real-world datasets demonstrate the effectiveness of our method. Zhiyu Fang, Jingyan Qin, Xiaobin Zhu 0001, Xu-Cheng Yin |
LREC/COLING | 2 |
| 2024 | LayoutFormer: Hierarchical Text Detection Towards Scene Text UnderstandingabstractExisting scene text detectors generally focus on accu-rately detecting single-level (i.e., word-level, line-level, or paragraph-level) text entities without exploring the relationships among different levels of text entities. To comprehensively understand scene texts, detecting multi-level texts while exploring their contextual information is criti-cal. To this end, we propose a unified framework (dubbed LayoutFormer) for hierarchical text detection, which simultaneously conducts multi-level text detection and predicts the geometric layouts for promoting scene text understanding. In LayoutFormer, WordDecoder, LineDecoder, and Pa- raDecoder are proposed to be responsible for word-level text prediction, line-level text prediction, and paragraph- level text prediction, respectively. Meanwhile, WordDe-coder and ParaDecoder adaptively learn word-line and line-paragraph relationships, respectively. In addition, we propose a Prior Location Sampler to be used on multi-scale features to adaptively select a few representative foreground features for updating text queries. It can improve hierar- chical detection performance while significantly reducing the computational cost. Comprehensive experiments verify that our method achieves state-of-the-art performance on single-level and hierarchical text detection. Jia-Wei Ma, Xiaobin Zhu 0001, Jingyan Qin, Xu-Cheng Yin |
CVPR | 4 |
| 2024 | Attention Decoupling for Query-Based Object DetectionabstractBenefiting from attention mechanisms, query-based detectors have a strong model capacity. They predict classification and regression by utilizing their shared queries and features in the decoder. Inter-task biases cause multi-directional gradients that disturb each other to limit model optimization. In this work, we introduce an attention decoupling (AD) for query-based detectors to explicitly align multi-task features. Specifically, AD consists of a Dense-to-Sparse Query Generator (DSQG) and a Split Cross-Attention (SCA), enabling query and feature decoupling respectively in decoding phase. Then, we propose a task consistency loss (TCL) which integrates a novel task alignment metric to classification loss to further improve task consistency across multiple decoding stages. Thus, AD effectively mitigates query-based detectors’ task misalignment problem and inspires subsequent multi-task paradigms. Moreover, extensive experiments on COCO dataset demonstrate that the proposed AD can enhance a variety of representative detectors. Remarkably, AD-DINO achieves the state-of-the-art performance. Jia-Wei Ma, Haixia Man, Shu Tian, Jingyan Qin, Xu-Cheng Yin |
ICASSP | 5 |
| 2024 | Exploring Stable Meta-Optimization Patterns via Differentiable Reinforcement Learning for Few-Shot ClassificationabstractExisting few-shot learning methods generally focus on designing exquisite structures of meta-learners for learning task-specific prior to improve the discriminative ability of global embeddings. However, they often ignore the importance of learning stability in meta-training, making it difficult to obtain a relatively optimal model. From this key observation, we propose an innovative generic differentiable Reinforcement Learning (RL) strategy for few-shot classification. It aims to explore stable meta-optimization patterns in meta-training by learning generalizable optimizations for producing task-adaptive embeddings. Accordingly, our differentiable RL strategy models the embedding procedure of feature transformation layers in meta-learner to optimize the gradient flow implicitly. Also, we propose a memory module to associate historical and current task states and actions for exploring inter-task similarity. Notably, our RL-based strategy can be easily extended to various backbones. In addition, we propose a novel task state encoder to encode task representation, which fully explores inner-task similarities between support set and query set. Extensive experiments verify that our approach can improve the performance of different backbones and achieve promising results against state-of-the-art methods in few-shot classification. Xiaobin Zhu 0001, Jingyan Qin, Xu-Cheng Yin |
ACM Multimedia | 5 |
| 2024 | Transformer-based Reasoning for Learning Evolutionary Chain of Events on Temporal Knowledge GraphabstractTemporal Knowledge Graph (TKG) reasoning often involves completing missing factual elements along the timeline. Although existing methods can learn good embeddings for each factual element in quadruples by integrating temporal information, they often fail to infer the evolution of temporal facts. This is mainly because of (1) insufficiently exploring the internal structure and semantic relationships within individual quadruples and (2) inadequately learning a unified representation of the contextual and temporal correlations among different quadruples. To overcome these limitations, we propose a novel Transformer-based reasoning model (dubbed ECEformer) for TKG to learn the Evolutionary Chain of Events (ECE). Specifically, we unfold the neighborhood subgraph of an entity node in chronological order, forming an evolutionary chain of events as the input for our model. Subsequently, we utilize a Transformer encoder to learn the embeddings of intra-quadruples for ECE. We then craft a mixed-context reasoning module based on the multi-layer perceptron (MLP) to learn the unified representations of inter-quadruples for ECE while accomplishing temporal knowledge reasoning. In addition, to enhance the timeliness of the events, we devise an additional time prediction task to complete effective temporal information within the learned unified representation. Extensive experiments on six benchmark datasets verify the state-of-the-art performance and the effectiveness of our method. Zhiyu Fang, Shuai-Long Lei, Xiaobin Zhu 0001, Shi-Xue Zhang, Xu-Cheng Yin, Jingyan Qin |
SIGIR | 7 |
| 2024 | Semi-supervised domain adaptation via subspace explorationabstractAbstract Recent methods of learning latent representations in Domain Adaptation (DA) often entangle the learning of features and exploration of latent space into a unified process. However, these methods can cause a false alignment problem and do not generalise well to the alignment of distributions with large discrepancy. In this study, the authors propose to explore a robust subspace for Semi‐Supervised Domain Adaptation (SSDA) explicitly. To be concrete, for disentangling the intricate relationship between feature learning and subspace exploration, the authors iterate and optimise them in two steps: in the first step, the authors aim to learn well‐clustered latent representations by aggregating the target feature around the estimated class‐wise prototypes; in the second step, the authors adaptively explore a subspace of an autoencoder for robust SSDA. Specially, a novel denoising strategy via class‐agnostic disturbance to improve the discriminative ability of subspace is adopted. Extensive experiments on publicly available datasets verify the promising and competitive performance of our approach against state‐of‐the‐art methods. Xiaobin Zhu 0001, Zhiyu Fang, Jingyan Qin, Xu-Cheng Yin |
IET Comput. Vis. | 5 |
| 2024 | Sample Weighting with Hierarchical Equalization Loss for Dense Object DetectionabstractLabel assignment (LA) is one of the essential phases in the object detection paradigm and aims to classify samples as foreground or background. Current LA strategies generally discriminate samples by explicit thresholds and then calculate weighted losses based on their significances. However, existing methods mostly neglect to consider the importance of samples comprehensively due to the uneven distribution of objects and the limitations of detector structures. In this paper, we propose a hierarchical equalization loss (HEL) by reconsidering the underlying factors affecting sample weights. First, we mitigate sample imbalance at three progressive levels. (1) Task level. We propose task-reconciled weights (TRW) to overcome the effects caused by inter-task inconsistencies (i.e., the inherent differences of classification and localization). (2) Instance level. We propose instance-aware normalization (IAN) for reconstructing the distribution of sample weights within an instance to suppress environmental noise. (3) Pyramid level. We propose hierarchical modulation (HM) to alleviate the unbalanced distribution of multi-scale objects on feature pyramids. Then, we stack the above three mechanisms and formulate the effective weighted loss. Moreover, we propose a staggered candidate bag construction (SCBC) mechanism to further improve the robustness of our method. Without adding any extra overhead, HEL can improve the performance of representative detectors by an impressive margin. Equipped with HEL, a single “ResNet-50+FPN+Head” detector can achieve a performance of 41.9 AP on COCO under 1× schedule, outperforming other existing LA methods. Extensive experiments conducted on multiple backbones and datasets demonstrate the effectiveness of our method. Jia-Wei Ma, Lei Chen 0069, Shu Tian, Song-Lu Chen, Jingyan Qin, Xu-Cheng Yin |
IEEE Trans. Multim. | 6 |
| 2023 | Learning Correction Filter via Degradation-Adaptive Regression for Blind Single Image Super-ResolutionabstractAlthough existing image deep learning super-resolution (SR) methods achieve promising performance on benchmark datasets, they still suffer from severe performance drops when the degradation of the low-resolution (LR) input is not covered in training. To address the problem, we propose an innovative unsupervised method of Learning Correction Filter via Degradation-Adaptive Regression for Blind Single Image Super-Resolution. Highly inspired by the generalized sampling theory, our method aims to enhance the strength of off-the-shelf SR methods trained on known degradations and adapt to unknown complex degradations to generate improved results. Specifically, we first conduct degradation estimation for each local image region by learning the internal distribution in an unsupervised manner via GAN. Instead of assuming degradation are spatially invariant across the whole image, we learn correction filters to adjust degradations to known degradations in a spatially variant way by a novel linearly-assembled pixel degradation-adaptive regression module (DARM). DARM is lightweight and easy to optimize on a dictionary of multiple pre-defined filter bases. Extensive experiments on synthetic and real-world datasets verify the effectiveness of our method both qualitatively and quantitatively. Code can be available at: https://github.com/edbca/DARSR. Xiaobin Zhu 0001, Jianqing Zhu, Shi-Xue Zhang, Jingyan Qin, Xu-Cheng Yin |
ICCV | 6 |
| 2023 | HFENet: Hybrid Feature Enhancement Network for Detecting Texts in Scenes and Traffic PanelsabstractText detection in complex scene images is a challenging task for intelligent transportation. Existing scene text detection methods often adopt multi-scale feature learning strategies to extract informative feature representations for covering objects of various sizes. However, the sampling operation inherent in multi-scale feature generation can easily impair high-frequency details (e.g., textures and boundaries), which are critical for text detection. In this work, we propose an innovative Hybrid Feature Enhancement Network (dubbed HFENet) to explicitly improve the quality of high-frequency information for detecting texts in scenes and traffic panels. To be concrete, we propose a simple yet effective self-guided feature enhancement module (SFEM) for globally lifting feature representations to highly discriminative and high-frequency abundant ones. Notably, our SFEM is pluggable and will be removed after training without introducing extra computational costs. In addition, due to the challenge and importance of accurately predicting boundaries for text detection, we propose a novel boundary enhancement module (BEM) to explicitly strengthen local feature representations in the guidance of boundary annotation for accurate localization. Extensive experiments on multiple publicly available datasets (i.e., MSRA-TD500, CTW1500, Total-Text, Traffic Guide Panel Dataset, Chinese Road Plate Dataset, and ASAYAR_TXT) verify the state-of-the-art performance of our method. Xiaobin Zhu 0001, Jingyan Qin, Xu-Cheng Yin |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2023 | RUArt: A Novel Text-Centered Solution for Text-Based Visual Question AnsweringabstractText-based visual question answering (VQA) requires to read and understand text in an image to correctly answer a given question. However, most current methods simply add optical character recognition (OCR) tokens extracted from the image into the VQA model without considering contextual information of OCR tokens and mining the relationships between OCR tokens and scene objects. In this paper, we propose a novel text-centered method called RUArt (Reading, Understanding and Answering the Related Text) for text-based VQA. Taking an image and a question as input, RUArt first reads the image and obtains text and scene objects. Then, it understands the question, OCRed text and objects in the context of the scene, and further mines the relationships among them. Finally, it answers the related text for the given question through text semantic matching and reasoning. We evaluate our RUArt on two text-based VQA benchmarks (ST-VQA and TextVQA) and conduct extensive ablation studies for exploring the reasons behind RUArt’s effectiveness. Experimental results demonstrate that our method can effectively explore the contextual information of the text and mine the stable relationships between the text and objects. Zanxia Jin, Heran Wu, Jingyan Qin, Lei Xiao 0001, Xu-Cheng Yin |
IEEE Trans. Multim. | 5 |
| 2022 | Learning Aligned Cross-Modal Representation for Generalized Zero-Shot ClassificationabstractLearning a common latent embedding by aligning the latent spaces of cross-modal autoencoders is an effective strategy for Generalized Zero-Shot Classification (GZSC). However, due to the lack of fine-grained instance-wise annotations, it still easily suffer from the domain shift problem for the discrepancy between the visual representation of diversified images and the semantic representation of fixed attributes. In this paper, we propose an innovative autoencoder network by learning Aligned Cross-Modal Representations (dubbed ACMR) for GZSC. Specifically, we propose a novel Vision-Semantic Alignment (VSA) method to strengthen the alignment of cross-modal latent features on the latent subspaces guided by a learned classifier. In addition, we propose a novel Information Enhancement Module (IEM) to reduce the possibility of latent variables collapse meanwhile encouraging the discriminative ability of latent variables. Extensive experiments on publicly available datasets demonstrate the state-of-the-art performance of our method. Zhiyu Fang, Xiaobin Zhu 0001, Jingyan Qin, Xu-Cheng Yin |
AAAI | 5 |
| 2022 | From Token to Word: OCR Token Evolution via Contrastive Learning and Semantic Matching for Text-VQAabstractText-based Visual Question Answering (Text-VQA) is a question-answering task to understand scene text, where the text is usually recognized by Optical Character Recognition (OCR) systems. However, the text from OCR systems often includes spelling errors, such as "pepsi" being recognized as "peosi". These OCR errors are one of the major challenges for Text-VQA systems. To address this, we propose a novel Text-VQA method to alleviate OCR errors via OCR token evolution. First, we artificially create the misspelled OCR tokens in the training time, and make the system more robust to the OCR errors. To be specific, we propose an OCR Token-Word Contrastive (TWC) learning task, which pre-trains word representation by augmenting OCR tokens via the Levenshtein distance between the OCR tokens and words in a dictionary. Second, by assuming that the majority of characters in misspelled OCR tokens are still correct, a multimodal transformer is proposed and fine-tuned to predict the answer using character-based word embedding. Specifically, we introduce a vocabulary predictor with character-level semantic matching, which enables the model to recover the correct word from the vocabulary even with misspelled OCR tokens. A variety of experimental evaluations show that our method outperforms the state-of-the-art methods on both TextVQA and ST-VQA datasets. The code will be released at https://github.com/xiaojino/TWA. Zanxia Jin, Zheng Shou 0001, Satoshi Tsutsui, Jingyan Qin, Xu-Cheng Yin |
ACM Multimedia | 5 |
| 2022 | Scene text detection via decoupled feature pyramid networks
Jie-Bo Hou, Xiaobin Zhu 0001, Jingyan Qin |
Int. J. Document Anal. Recognit. | 5 |
| 2022 | Depth-Guided Progressive Network for Object DetectionabstractMulti-scale object detection in natural scenes is still challenging. To enhance the multi-scale perception capability, some algorithms combine the lower-level and higher-level information via multi-scale feature fusion strategies. However, the inherent spatial properties among instances and relations between foreground and background are ignored. In addition, the human-defined “center-based” regression quality evaluation strategy, predicting a high-to-low score based on a linear relationship with the distance to the center of ground-truth box, is not robust to scale-variant objects. In this work, we propose a Depth-Guided Progressive Network (DGPNet) for multi-scale object detection. Specifically, besides the prediction of classification and localization, the depth is estimated and used to guide the image features in a weighted manner to obtain a better spatial representation. Therefore, depth estimation and 2D object detection are simultaneously learned via a unified network, where the depth features are merged as auxiliary information into the detection branch to enhance the discrimination among multi-scale objects. Moreover, to overcome the difficulty of empirically fitting the localization quality function, high-quality predicted boxes on scale-variant objects are more adaptively obtained by an IoU-aware progressive sampling strategy. We divide the sampling process into two stages, i.e., “statistical-aware” and “IoU-aware”. The former selects thresholds for positive samples based on statistical characteristics of multi-scale instances, and the latter further selects high-quality samples by IoU on the basis of the former. Therefore, the final ranking scores better reflect the quality of localization. Experiments verify that our method outperforms state-of-the-art methods on the KINS and Cityscapes dataset. Jia-Wei Ma, Song-Lu Chen, Feng Chen 0040, Shu Tian, Jingyan Qin, Xu-Cheng Yin |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2021 | Enhancing the generalization of feature construction using genetic programming for imbalanced data with augmented non-overlap degreeabstractGenetic programming (GP) has a significant achievement in feature construction and non-overlap degree can help to improve the generalization ability of GP based feature construction. However, the non-overlap degree is biased towards the majority class. In this paper, a novel GP based feature construction method with augmented non-overlap degree is proposed to enhance the generalization ability for imbalanced data. And the constructed features are evaluated by a novel function based on the combination of the area under the ROC curve metric and the augmented non-overlap degree. The generalization performance is evaluated not only by a particular classification algorithm, but also by six widely used classification algorithms. The experiments conducted on five imbalanced biomedical datasets with different imbalance rates show that the proposed GP-AANO method can achieve superior generalization performance for classification. Jingyan Qin, Haiyan Gong, Xiaotong Zhang 0002, Yadong Wan |
BIBM | 2 |
| 2021 | Multi-orientation scene text detection with scale-guided regression
Jie-Bo Hou, Xiaobin Zhu 0001, Jingyan Qin, Xu-Cheng Yin |
Neurocomputing | 5 |
| 2020 | Toward high accuracy and visualization: An interpretable feature extraction method based on genetic programming and non-overlap degreeabstractGenetic programming (GP) has shown promising results in interpretable feature extraction, but few works considered both classification accuracy and data visualization as objectives. Evaluating the extracted features based on the combination of accuracy measures and visualization measures can help to achieve the two objectives simultaneously. However, the exploitation of improper visualization measures and combination methods will decrease the classification accuracy. In this paper, a novel feature extraction method based on GP and non-overlap degree is proposed to extract interpretable features for high accuracy and visualization. And a novel function that maximizes the product of the accuracy of a linear classifier and the non-overlap degree is proposed to evaluate the extracted features. The proposed method, named GP-ANO, is compared with other methods on five medical datasets by six common machine learning methods. The experimental results demonstrate that the GP-ANO method outperforms other compared methods in terms of both classification accuracy and data visualization. Jie He 0001, Xiaotong Zhang 0002, Huadong Fu, Jingyan Qin |
BIBM | 5 |
| 2020 | Draw Portraits by Music: A Music based Image Style Transformationabstract"Draw portraits by music", an interactive work of art. Compared with music visualization and image style conversion, it's AI's imitation of human synaesthetic. New portraits gradually appear on the screen and are synchronized with music in real-time. Users select music and images as the main interactive contents, the parameters of the music are used as the dynamic expression of human emotions, and the new pixel generation process of the image is regarded as the result of emotions affecting humans. Jingyan Qin, Wenfa Li |
ACM Multimedia | 2 |
| 2020 | Little World: Virtual Humans Accompany Children on Dramatic PerformanceabstractEvery child is the leading actor in her/his unique world. To help them achieve performance, an interactive art called 'little world' is proposed to let virtual humans accompany children on drama performance. Theatrical adaptation rewrites the novel to an interactive drama suitable for children. Little world builds drama scenes and virtual humans for characters and lets children interact with them by speech and actions. Xiaohui Wang 0004, Xiaoxue Ding, Jingyan Qin |
ACM Multimedia | 4 |
| 2020 | First Impression: AI Understands PersonalityabstractWhen you first encounter a person, a mental image of that person is formed. First impression, an interactive art, is proposed to let AI understand human personality at first glance. The mental image is demonstrated by Beijing opera facial makeups, which shows the character personality with a combination of realism and symbolism. We build Beijing opera facial makeup dataset and semantic dataset of facial features to establish relationships among real faces, personalities and facial makeups. First impression detects faces, recognizes personality from facial appearance and finds the matching Beijing opera facial makeup. Finally, the morphing process from real face to facial makeup is shown to let users enjoy the process of AI understanding personality. Xiaohui Wang 0004, Xia Liang, Miao Lu, Jingyan Qin |
ACM Multimedia | 4 |
| 2020 | Ranking via partial ordering for answer selection
Zanxia Jin, Bowen Zhang 0011, Jingyan Qin, Xu-Cheng Yin |
Inf. Sci. | 4 |
| 2020 | A reformative teaching-learning-based optimization algorithm for solving numerical and engineering design optimization problems
Xiaotong Zhang 0002, Jingyan Qin, Jie He 0001 |
Soft Comput. | 3 |
| 2019 | Identification of usability problems and requirements of elderly Chinese users for smart TV interactionsabstractWith the development of Information and Communications Technology (ICT), smart TV is gradually becoming universal and penetrating daily life. Smart TV has a wide range of user groups, and the elderly is an important group. The special physiological and psychological characteristics of Chinese elderly highlight the usability problems of smart TV interactions for them. The main functions of smart TVs were selected and operated by the elderly in this study. With the help of physiological measurements, behaviour analyses and interviews during natural usage scenarios, we determined the usability issues and user requirements of Chinese elderly for interacting with smart TVs. The research shows that the elderly in China have the intention to use smart TV products, but the usability of interactive systems affects the experience of using a smart TV. The results indicate that different content search methods result in different user experiences. Pinyin search is difficult to operate using a remote control, and some influencing factors are revealed. Voice search results in the best user experience for the elderly, but the recognition accuracy is easily affected by factors such as user accents and environmental noise. Hierarchical search is easy to operate, but it often takes a long time to finish an involved task. Other usability issues of some main functions, i.e. screen mirroring, shopping, playing games, system settings and application downloads, were also obtained, and the requirements of the elderly were well understood. These findings have implications for interaction design and implementation of smart TV service systems. Jinhua Dou, Jingyan Qin, Qingju Wang 0002, Qichao Zhao |
Behav. Inf. Technol. | 2 |
| 2017 | Data augmentation for unbalanced face recognition training sets
Biao Leng, Kai Yu 0003, Jingyan Qin |
Neurocomputing | 3 |
| 2016 | Cascade shallow CNN structure for face verification and identification
Biao Leng, Yu Liu 0015, Kai Yu 0003, Songting Xu, Ziqing Yuan, Jingyan Qin |
Neurocomputing | 6 |
| 2016 | Artistic Coloring: Color Transfer from PaintingabstractColor transfer is to alter an image’s color composition by reference to the color characteristics of another image. In this paper, we build a system called artistic coloring that realizes automatic color transfer from famous paintings. It properly extracts the wonderful color characteristics of famous paintings and applies them to color transfer. Specially, we investigate the traditional color theme extraction methods and find their deficiencies. Based on this, we quantify the processing of human painting to find the main colors in the color palette and propose an artistic balanced color theme extraction algorithm aimed specially at paintings. In the experiments, a user study is carried out to evaluate the artistic balanced extraction method. Our proposed method achieves the highest score 3.98 in a 5-point evaluation, which is 41% higher than traditional methods at most. We have successfully tested the artistic coloring system on lots of images with different painting styles. The results are natural and have the similar color characteristics with their corresponding reference paintings. Xiaohui Wang 0004, Jingyan Qin, Yujiao Gao |
Int. J. Pattern Recognit. Artif. Intell. | 2 |