EDBT 2026 Demo / reviewers in the wild / expert
Bo Sun 0006
dblp:35/892-6
· DBLP profile ↗
38ranked-venue papers
12as first author
22since 2021 · last 2026
0000-0003-1168-1051ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 3 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 4 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 7 · 4 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Computer networks · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A hippocampal-PFC inspired neuro-symbolic architecture for contextually anchored question chain generation
Wang Ruan, Bo Sun 0006, Jun He 0009, Tengda Qi, Guomin Zheng, Tore Hoel |
Neurocomputing | 2 |
| 2026 | Multi-scale steering of large vision language models via visual information intervention
Dongliang Zhao, Bo Sun 0006, Jun He 0009, Yinghui Zhang 0004 |
Neurocomputing | 2 |
| 2025 | Automated Coding Utterances Toward Chinese Course Core Competence with Large Language Models
Mingyang Yue 0001, Tengda Qi, Mingxuan Wang, Wang Ruan, Jun He 0009, Bo Sun 0006, Guomin Zheng |
ICIC (23) | 6 |
| 2025 | TAD-IVR: Enhancing Temporal Action Detection via Instrumental Variable RegressionabstractTemporal action detection is a key task in video understanding, with one major challenge being the handling of confounders. Confounders include both observed factors (e.g., temporal order, co-occurrence patterns of actions) and unobserved factors (e.g., lighting, individual states), which can introduce bias and affect predictions. While causal inference methods have been introduced, they often rely on fixed representations of confounders, limiting their adaptability to dynamic contexts, particularly with unobserved confounders. To address this, we propose TAD-IVR, which combines Transformer with instrumental variable (IV) regression. Transformer flexibly captures the temporal dependencies of actions, improving the representation of confounders, while IV regression uses exogenous variables to eliminate the influence of unobserved confounders, thus reducing prediction bias. Additionally, we introduce mutual information constraints and zero-sum optimization strategies to enforce more informative and accurate feature representations. Experimental results show that TAD-IVR effectively mitigates confounding effects and improves detection accuracy. Minglin Hong, Bo Sun 0006, Jun He 0009, Yinghui Zhang 0004 |
ICME | 2 |
| 2025 | Backward Design-Driven Modular Retrieval-Augmented Generation Framework for Automated Instructional Design in Chinese Language EducationabstractInstructional Design spans multiple disciplines, making its automation a key challenge in modern education. Instructional design generation based on a large language model faces significant challenges in term of reliability and long-distant textual logic. To Address those issues, this study presents a novel framework that integrates backward design theory with a modular Retrieval-Augmented Generation (RAG) approach. It ensures factual accuracy, goal alignment and efficiency by grounding content in reliable resources through retrieval-enhanced generation, aligning activities with learning outcomes using backward design principles, and optimizes scalability with modular RAG methods. A six-dimensional evaluation ensures the generation of high-quality and relevant content. Experimental results comparing with ChatGPT-4o highlight the framework’s effectiveness, especially in resource-constrained settings such as Chinese language education. The proposed framework offers a scalable and precise solution for automated Instructional Design, aligning closely with educational objectives. Meanwhile, it also provides potential solution for similar long textual and goal-oriented generation tasks. Wang Ruan, Tengda Qi, Jun He 0009, Bo Sun 0006, Guomin Zheng |
IJCNN | 4 |
| 2025 | CDE-CCL: A Classroom Dialogue Evaluation Framework for Chinese Core Literacy via LLMabstractClassroom dialogue evaluation is a crucial component of the teaching process assessment. It not only enhances the quality of classroom interaction and student engagement but also fosters the development of students’ core literacy. By leveraging classroom dialogue evaluation, teachers can implement personalized teaching strategies, overcoming the problem of unequal distribution of educational resources caused by regional disparities. However, current classroom dialogue evaluation frameworks are manually implemented, which results in low automation, high time consumption, and a lack of systematic assessment related to core literacy. To address these limitations, we propose the Classroom Dialogue Evaluation Framework for Chinese Core Literacy (CDE-CCL). This framework integrates Chinese Core Literacy, dialogue round segmentation, prompt engineering, and LoRA fine-tuning techniques to encode complete classroom dialogues, thereby enabling automated evaluation of classroom dialogues. Experimental results demonstrate the superiority of CDE-CCL in both tasks of classroom dialogue segmentation and encoding over existing methods across various aspects and educational levels. Mingxuan Wang, Mingyang Yue 0001, Guomin Zheng, Jun He 0009, Bo Sun 0006 |
IJCNN | 6 |
| 2025 | Mutual introspective distillation for unbiased scene graph generation
Bo Sun 0006, Zhuo Hao, Lejun Yu, Jun He 0009 |
J. Supercomput. | 1 |
| 2024 | Multi-level neural prompt for zero-shot weakly supervised group activity recognition
Yinghui Zhang 0004, Bo Sun 0006, Jun He 0009, Lejun Yu, Xiaochong Zhao |
Neurocomputing | 2 |
| 2024 | Unbiased scene graph generation using the self-distillation method
Bo Sun 0006, Zhuo Hao, Lejun Yu, Jun He 0009 |
Vis. Comput. | 1 |
| 2023 | Data Augmentation Ensemble Module based on Natural Guidance for X-ray Prohibited Items DetectionabstractAutomatic prohibited items detection plays an important role in protecting public security. Till now, object detection powered by deep learning provides a promising solution to automatic security inspection. However, in the one hand, according to the imaging principle of X-ray images, texture information would be lost, and in the other hand, different prohibited items with the same material are easily confused for the similar imaging color, leading to poor detection performance. Thus, to improve the detection performance of the basic object detection models on prohibited items detection task, we first propose the Data Augmentation Ensemble Module (DAEM) based on Nature Guidance for more accurate prohibited items detection. Specifically, inspired by the fact that inspectors detect items based on the characteristics of prohibited items in nature, we introduce natural images as prior knowledge to build X-ray security image - natural image sample pairs for supervising the model training. Besides, we adopt data augmentation strategies to enhance the diversity of the X-ray images, and then we combine the predictions from different data augmentation methods by ensemble learning to yield more accurate results. We verify the DAEM's performance by plug it into three different object detection models, and the experiments demonstrate that our framework can significantly improve the performance compared with the SOTA method on the PIDray dataset. Jun He 0009, Yangcai Zhong, Bo Sun 0006, Yinghui Zhang 0004 |
IJCNN | 3 |
| 2023 | Knowledge tracing based on multi-feature fusion
Yongkang Xiao, Rong Xiao 0002, Bo Sun 0006 |
Neural Comput. Appl. | 6 |
| 2023 | CoConGAN: Cooperative contrastive learning for few-shot cross-domain heterogeneous face translation
Yinghui Zhang 0004, Wansong Hu, Bo Sun 0006, Jun He 0009, Lejun Yu |
Neural Comput. Appl. | 3 |
| 2023 | Supervised Contrastive Learned Deep Model for Question Continuation EvaluationabstractQuestion continuation evaluation (QCE) is a branch task of dialogue act prediction (DAP) in the natural language processing area, which is aimed at predicting whether each question in a dialogue is worthy of being followed-up under a specific context. QCE is important for communication, education, and even entertainment. Regrettably, QCE has always been disregarded as an auxiliary task for conversational machine reading comprehension. QCE involves more information and relationships than the original DAP task, making it more complex. Moreover, the classification of QCE inherently renders the samples confusing. In this article, a transformer long short-term memory (LSTM)-based supervised contrastive learned model for QCE is proposed to automatically distribute QCE labels. This model is mainly constructed with transformer encoder blocks and LSTM modules, and supervised contrastive learning (SCL) is innovatively introduced to the training process. This model is good at extracting both information about corpora and the relationships among corpora, and SCL alleviates any confusion. With the only applicable dataset, i.e., Question Answering in Context (QuAC), experiments are conducted. This model is proven to perform well and is robust to missing data. The performance is 2.3% (accuracy) and 12.2% (macro-F1 score) higher than baselines from QuAC and only decreases by approximately 2.3% when 10% data remain. Bo Sun 0006, Jun He 0009, Yinghui Zhang 0004 |
IEEE Trans. Hum. Mach. Syst. | 1 |
| 2023 | Cross-language multimodal scene semantic guidance and leap sampling for video captioning
Bo Sun 0006, Yijia Zhao, Zhuo Hao, Lejun Yu, Jun He 0009 |
Vis. Comput. | 1 |
| 2022 | ENG-Face: cross-domain heterogeneous face synthesis with enhanced asymmetric CycleGAN
Yinghui Zhang 0004, Lejun Yu, Bo Sun 0006, Jun He 0009 |
Appl. Intell. | 3 |
| 2022 | Machine Reading Comprehension with Rich KnowledgeabstractMachine reading comprehension (MRC) is a crucial and challenging task in natural language processing (NLP). With the development of deep learning, language models have achieved excellent results. However, these models still cannot answer complex questions. Currently, researchers often utilize structured knowledge, such as knowledge bases (KBs), as external knowledge by directly extracting triples to enhance the results of machine reading. Although they can support certain background knowledge, the triples are limited to the interrelationships among entities or words. Unlike structured knowledge, unstructured knowledge is rich and extensive. However, these methods ignore unstructured knowledge resources, such as Wikipedia. In addition, the effect of combining the two types of knowledge is still not known. In this study, we first attempt to explore the usefulness of combining them. We introduce a fusion mechanism into a rich knowledge fusion layer (RKF) to obtain more useful and relevant knowledge from different external knowledge resources. Further to promote interaction among different types of knowledge, a bi-matching layer is added. We propose the RKF-NET framework based on BERT, and our experimental results demonstrate the effectiveness of two classic datasets: SQuAD1.1 and the Easy-Challenge (ARC). Jun He 0009, Yinghui Zhang 0004, Bo Sun 0006, Rong Xiao 0002, Yongkang Xiao |
Int. J. Pattern Recognit. Artif. Intell. | 4 |
| 2022 | Dynamic Micro-Expression Recognition Using Knowledge DistillationabstractMicro-expression is a spontaneous expression that occurs when a person tries to mask his or her inner emotion, and can neither be forged nor suppressed. It is a kind of short-duration, low-intensity, and usually local-motion facial expression. However, owing to these characteristics of micro-expression, it is difficult to obtain micro-expression data, which is the bottleneck of applying deep learning methods to micro-expression recognition. In addition, micro-expression is still a type of expression, and it can also be encoded by the facial action coding system. Therefore, there is a certain correlation between action unit recognition and micro-expression recognition. Addressing those, we propose a novel knowledge transfer technique distills and transfers knowledge from action unit for micro-expression recognition, where knowledge from a pre-trained deep teacher neural network is distilled and transferred to a shallow student neural network. Specifically, a teacher-student correlative framework is designed with a novel objective function. And features extracted from the teacher network is used as prior knowledge to guide the student part to efficiently learning from the target micro-expression dataset. Experiments are conducted on four available published micro-expression datasets (SMIC2, CASME, CASME II, and SAMM). The experimental results show that our model outperforms the state-of-the-art systems. Bo Sun 0006, Siming Cao, Dongliang Li, Jun He 0009, Lejun Yu |
IEEE Trans. Affect. Comput. | 1 |
| 2021 | Automatically Predict Question Difficulty for Reading Comprehension ExercisesabstractQuestion difficulty is a critical indicator for educational examination and personalized learning resources recommendation. Its evaluation mainly based on experts’ experience, which is both subjective and labor intensive. Recently, many studies pay more attention to using neural network for question difficulty prediction (QDP). Though these methods have improved efficiency of difficulty prediction, they only regarded the difficulty prediction task as a simple classification or prediction task, which ignored the influence of the input text relation on it, such as confusion relation of multiple choice for English reading comprehension items. Therefore, in this paper, we proposed a Convolutional Neural Network with Multi-view attention (MACNN) to extract different relation from multiple-parts text in reading comprehension exercises with multiple-choice for automatically predicting question difficulty. Our experimental results demonstrate the effectiveness of the proposed framework on a real-world dataset. Besides, we give interpretable insights to analyze the effect of three module we designed with attention mechanism by attention weights visualization experiment. Jun He 0009, Bo Sun 0006, Lejun Yu, Yinghui Zhang 0004 |
ICTAI | 3 |
| 2021 | Dual Multi-Task Network with Bridge-Temporal-Attention for Student Emotion Recognition via Classroom VideoabstractEmotion recognition is one of the most significant technologies for building a smart education environment. The quantitative analysis of students' emotion in classroom is helpful to improve teaching effect. Though temporal features of emotion generation and disappearance process has been demonstrated to be of great benefit for emotion recognition, it has received little attention. Recently, a large of data has been accumulated in education, like classroom video. Considering the practical scenarios, it is more challenge to obtain the temporal information of video emotion. Therefore, this paper firstly uses the deep learning and Attention Mechanism methods to automatically modelling the temporal process. Besides, we make use of multitask learning to build a DMTN-BTA model for students' emotion recognition through classroom videos. While, temporal segment labels are unnecessary in the model. The method consists of the CNN for spatio-temporal features extraction on emotion recognition and temporal segmentation tasks and BLSTM-RNN with a novel ‘Bridge-Temporal-Attention’ for the emotion recognition task. Experiment results show that our model outperforms the prior single-task learning methods on the BNU-LSVED 2.0 and the current state-of-the-art methods on the official FABO. Jun He 0009, Bo Sun 0006, Lejun Yu |
IJCNN | 3 |
| 2021 | NIR-VIS Heterogeneous Face Synthesis via Enhanced Asymmetric CycleGANabstractHeterogeneous face image involves many fields, among which the near infrared (NIR) and visible (VIS) face image recognition is a hot field. NIR image has excellent imaging effect under extreme illumination conditions. Although few NIR images are available, unpaired image-to-image translation provides a solution for converting NIR-VIS to each other. However, NIR image contains relatively poor information, while VIS image contains relatively rich information, which results in asymmetry between NIR and VIS, and not suitable for unpaired translation based on two domains of the same complexity. In this paper, an enhanced asymmetric CycleGAN(EN-ASGAN) with edge retention module, auxiliary encoder module and generators equipped with different sizes is used to convert NIR-VIS face images. To validate the effective of EN-ASGAN, we conduct experiments on CASIA NIR-VIS 2.0 dataset. The experimental results are evaluated qualitatively and quantitatively and show that EN-ASGAN is effective for heterogeneous face generation in unpaired NIR-VIS images. Yinghui Zhang 0004, Bo Sun 0006, Jun He 0009, Lejun Yu |
IJCNN | 3 |
| 2021 | Analyses and Benchmark of a Spontaneous Student Affect DatabaseabstractBNU-LSVED2.0 is the first large-scale spontaneous and multi-modal student affect database adapted to the classroom environment. On the basis of BNU-LSVED2.0, BNU-SDED is created. This study analyzes BNU-LSVED2.0 and BNU-SDED and develops a benchmark for BNU-SDED, which can provide some insights for future research on education improvement and emotion recognition. Firstly, student personality traits are analyzed to ensure the validity of data. Secondly, the relationship between PAD emotional dimensions and students’ emotional states in learning is analyzed. Based on these analyses, a set of mapping labels between categorical emotions and PAD dimensions is presented. Thirdly, BNU-SDED is used to perform basic baseline experiments. The results demonstrate the effectiveness of the data and reveal the trends behind students’ affects in learning. These analyses can benefit both developers and users of student affect databases. Bo Sun 0006, Sixu Lu, Jun He 0009, Lejun Yu |
ISM | 1 |
| 2021 | Student Class Behavior Dataset: a video dataset for recognizing, detecting, and captioning students' behaviors in classroom scenes
Bo Sun 0006, Kaijie Zhao, Jun He 0009, Lejun Yu, Huanqing Yan, Ao Luo |
Neural Comput. Appl. | 1 |
| 2020 | Feedback evaluations to promote image captioningabstractImage captioning can be treated as a policy gradient problem. A retrieval model to obtain the discriminability score to distinguish between two images, given the caption for one of them, has been proposed previously; the discriminability score and one of the image captioning evaluation metrics were optimised using policy gradient. Based on this, two methods to evaluate the caption and caption‐generating process, referred to as feedback evaluations, are proposed in this study. The results of the evaluations were used to improve the model. First, an auxiliary retrieval loss (ARL) is introduced to evaluate the generated caption to improve the discriminability of the model. ARL has been utilised as a feedback evaluation method because it calculates similarity between the generated caption and convolutional neural network features. With ARL, a higher similarity and better discriminability were achieved. Second, an evaluation reward is introduced to evaluate the captioning process. With ER, the overall evaluation metrics can be improved. A policy gradient was used, and a captioning model could be trained by jointly adjusting the captioning process and captioning itself. The attention long short‐term memory network was trained with ARL and ER successively and it demonstrated state‐of‐the‐art performance on the COCO database. Jun He 0009, Yijia Zhao, Bo Sun 0006, Lejun Yu |
IET Image Process. | 3 |
| 2020 | Local relation network with multilevel attention for visual question answering
Bo Sun 0006, Zeng Yao, Yinghui Zhang 0004, Lejun Yu |
J. Vis. Commun. Image Represent. | 1 |
| 2019 | Feature augmentation for imbalanced classification with conditional mixture WGANs
Yinghui Zhang 0004, Bo Sun 0006, Yongkang Xiao, Rong Xiao 0002, Yungang Wei |
Signal Process. Image Commun. | 2 |
| 2018 | Edge Convolutional Network for Facial Action Intensity EstimationabstractIn this paper, we propose a novel convolutional neural architecture for facial action unit intensity estimation. While Convolutional Neural Networks (CNNs) have shown great promise in a wide range of computer vision tasks, these achievements have not translated as well to facial expression analysis, with hand crafted features (e.g. the Histogram of Orientated Gradient) still being very competitive. We introduce a novel Edge Convolutional Network (ECN) that is able to capture subtle changes in facial appearance. Our model is able to learn edge-like detectors that can capture subtle wrinkles and facial muscle contours at multiple orientations and frequencies. The core novelty of our ECN model is in its first layer which integrates three main components: an edge filter generator, a receptive gate and a filter rotator. All the components are differentiable and our ECN model is end-to-end trainable and learns the important edge detectors for facial expression analysis. Experiments on two facial action unit datasets show that the proposed ECN outperforms state-of-the-art methods for both AU intensity estimation tasks. Liandong Li, Tadas Baltrusaitis, Bo Sun 0006, Louis-Philippe Morency |
FG | 3 |
| 2018 | Affect recognition from facial movements and body gestures by hierarchical deep spatio-temporal features and fusion strategy
Bo Sun 0006, Siming Cao, Jun He 0009, Lejun Yu |
Neural Networks | 1 |
| 2018 | Content-Based Efficient Messages Transmission in WSNsabstractThe WSNs are mainly to monitor various types of sensing content. In the automated control system, administration centers (ACs) often send notifications to a set of nodes meeting given content ranges to implement specific actions. For instance, we notify the sensing devices with sensing temperatures greater than 35 to turn on the cooling system. At present, the transmission of notification messages is mostly based on node identification, which is isolated from sensing contents and unable to accurately locate nodes meeting the sensing content requirements. In this paper, we generalize two types of sensing contents: the continuous values and the discrete values, and a content‐based efficient message transmission mechanism (CEMT) is proposed for the typical tree‐like topologies in WSNs. Thus, a highly effective notifying method for content‐based multicast and anycast messages is designed, which accurately sends notification messages to nodes whose sensing values belong to the given range. Sometimes, multiple sensing types of nodes are mixed together to construct a net topology, offering underlying transport services to each other. At this point, CEMT builds a specialized logical tree for the same type of nodes, and notification messages are transmitted along the logical tree. CEMT reduces the transmission time, bandwidth, and number of processing nodes when a content‐based message is sent, and it takes less storage of content routing entries in nodes. Rong Xiao 0002, Bo Sun 0006, Yongkang Xiao, Yungang Wei |
Wirel. Commun. Mob. Comput. | 2 |
| 2017 | Multi View Facial Action Unit Detection Based on CNN and BLSTM-RNNabstractThis paper presents our work in the FG 2017 Facial Expression Recognition and Analysis challenge (FERA 2017) and we participate in the AU occurrence sub-challenge. Our work of AU occurrence recognition is based on deep learning, and we design convolution neural network (CNN) models for two types of work: facial view recognition and AU occurrence recognition. For facial view recognition, our model could achieve 97.7% accuracy on validation dataset about 9 facial views. For AU occurrence recognition, we use both visual features and temporal information of dataset. We use CNN models to get deep visual feature and then use BLSTM-RNN to learn the high-level feature in the time domain. When training models, we divide dataset into 9 parts based on 9 facial views, and each model is trained in a specific view. When recognizing AUs, we recognize facial view first and then choose the corresponding model for AU occurrence recognition. Finally, our method shows good performance, the F1 score of test data is 0.507 and the accuracy is 0.735. Jun He 0009, Dongliang Li, Siming Cao, Bo Sun 0006, Lejun Yu |
FG | 5 |
| 2017 | A new deep-learning framework for group emotion recognitionabstractIn this paper, we target the Group-level emotion recognition sub-challenge of the fifth Emotion Recognition in the Wild (EmotiW 2017) Challenge, which is based on the Group Affect Database 2.0 containing images of groups of people in a wide variety of social events. We use Seetaface to detect and align the faces in the group images and extract two kinds of face-image visual features: VGGFace-lstm, DCNN-lstm. As group image features, we propose using Pyramid Histogram of Oriented Gradients (PHOG), CENTRIST, DCNN features, VGG features. To the testing group images on which the faces have been detected, the final emotion is estimated using group image features and face-level visual features. While to the testing group images on which the faces cannot be detected, the face-level visual features are fused for final recognition. The final achievements we have gained are 79.78% accuracy on the Group Affect Database 2.0 testing set, which is much higher than the corresponding baseline results 53.62%. Qinglan Wei, Yijia Zhao, Qihua Xu, Liandong Li, Jun He 0009, Lejun Yu, Bo Sun 0006 |
ICMI | 7 |
| 2017 | BNU-LSVED 2.0: Spontaneous multimodal student affect database with multi-dimensional labels
Qinglan Wei, Bo Sun 0006, Jun He 0009, Lejun Yu |
Signal Process. Image Commun. | 2 |
| 2016 | LSTM for dynamic emotion and group emotion recognition in the wildabstractIn this paper, we describe our work in the fourth Emotion Recognition in the Wild (EmotiW 2016) Challenge. For video based emotion recognition sub-challenge, we extract acoustic features, LBPTOP, Dense SIFT and CNN-LSTM features to recognize the emotions of film characters. For group level emotion recognition sub-challenge, we use LSTM and GEM model. We train linear SVM classifiers for these kinds of features on the AFEW6.0 and HAPPEI dataset, and use a fusion network we proposed to combine all the extracted features at the decision level. The final achievements we have gained are 51.54% accuracy on the AFEW testing set and 0.836 RMSE on the HAPPEI testing set. Bo Sun 0006, Qinglan Wei, Liandong Li, Qihua Xu, Jun He 0009, Lejun Yu |
ICMI | 1 |
| 2015 | The Design of a Visual Tool for the Quick Customization of Virtual Characters in OSSLabstractThis paper implements a visual tool for the quick customization of virtual characters in 3D virtual worlds with a three-layer structure based on the Open Sim and Second Life platforms and "Knowledge" theory. "Knowledge" refers to specific controlling conditions and the corresponding behaviors of virtual characters. The three-layer structure is divided into the data layer, entity layer, and implementation layer. The data layer is used to store the data needed by the virtual characters' login and the behavioral control, the entity layer enables the exchange of data with the data layer and passes parameters to the implementation layer, and the implementation layer allows for realizing the knowledge of virtual characters with the support of the API in the Open Metaverse open source framework. Moreover, visualization of the tool is implemented through the interfaces provided by the entity layer. With this tool, users can generate a group of virtual characters quickly and conveniently at any place in a certain virtual world and select a variety of knowledge for diverse virtual characters. Yungang Wei, Xiaoran Qin, Xiaoye Tan, Xiaohang Yu, Bo Sun 0006 |
CW | 5 |
| 2015 | Regularized single-image super-resolution based on progressive gradient estimationabstractGradient domain optimization is widely used in regularized image super-resolution, in which the gradient of high resolution (HR) is estimated for calculating the regularization energy. In this paper, a progressive gradient estimation (PGE) is proposed. In PGE, the gradient of the reconstructed HR image in the previous round of optimization is taken as the estimated gradient in the current round. Then, the estimated image gradient is progressively improved. When the estimated image gradient converges, a high quality HR image can be reconstructed. Experimental results show that the reconstructed HR images by PGE have good qualitative and quantitative performances. Lejun Yu, Feng-Xiang Ge, Bo Sun 0006, Jun He 0009, Robert Sablatnig |
ICIP | 4 |
| 2015 | Combining Multimodal Features within a Fusion Network for Emotion Recognition in the WildabstractIn this paper, we describe our work in the third Emotion Recognition in the Wild (EmotiW 2015) Challenge. For each video clip, we extract MSDF, LBP-TOP, HOG, LPQ-TOP and acoustic features to recognize the emotions of film characters. For the static facial expression recognition based on video frame, we extract MSDF, DCNN and RCNN features. We train linear SVM classifiers for these kinds of features on the AFEW and SFEW dataset, and we propose a novel fusion network to combine all the extracted features at decision level. The final achievement we gained is 51.02% on the AFEW testing set and 51.08% on the SFEW testing set, which are much better than the baseline recognition rate of 39.33% and 39.13%. Bo Sun 0006, Liandong Li, Guoyan Zhou, Xuewen Wu, Jun He 0009, Lejun Yu, Qinglan Wei |
ICMI | 1 |
| 2014 | Exploring the Use of a 3D Virtual Environment in Chinese Cultural TransmissionabstractThrough scene construction and intelligent interaction in a 3D virtual world environment, we developed the project "Confucius' Journey". Considering the problems in such applications, such as the lack of interaction and reduced effectiveness in representing the application purpose, we explored interactive objects and virtual human technology. In addition, we can verify the advantage of using the 3D platform via the experimental results. Yungang Wei, Xiaoye Tan, Xiaoran Qin, Xiaohang Yu, Bo Sun 0006 |
CW | 5 |
| 2014 | Combining Multimodal Features with Hierarchical Classifier Fusion for Emotion Recognition in the WildabstractEmotion recognition in the wild is a very challenging task. In this paper, we investigate a variety of different multimodal features from video and audio to evaluate their discriminative ability to human emotion analysis. For each clip, we extract SIFT, LBP-TOP, PHOG, LPQ-TOP and audio features. We train different classifiers for every kind of features on the dataset from EmotiW 2014 Challenge, and we propose a novel hierarchical classifier fusion method for all the extracted features. The final achievement we gained on the test set is 47.17% which is much better than the best baseline recognition rate of 33.7%. Bo Sun 0006, Liandong Li, Tian Zuo, Guoyan Zhou, Xuewen Wu |
ICMI | 1 |
| 2009 | An Improved Cloud Rendering MethodabstractClouds play an important role for creating realistic images of outdoor scenes. Many methods have been proposed for modeling and rendering clouds. Among these methods, Dobashi et al. proposed a simple but efficient cloud rendering algorithm based on splatting algorithm and take into account single scattering in couds. This algorithm has been used in many papers to generate realistic clouds images. However, in images generated by using Dobashi's method, lattice like artifact can be easily perceived. This paper describes a method to improve the quality of cloud image generated by dobashi method. We perturb the positions of cloud particles and resample the density at perturbed position within neighboring meatballs to remove the regular pattern from regular space lattice structure and adding noise to texture of billboard used in the original method. Bo Sun 0006, Jun He 0009, Yongkang Xiao, Rong Xiao 0002 |
ICIG | 2 |