Chengbin Quan

dblp:176/1175 · DBLP profile ↗
← Back
16ranked-venue papers
1as first author
11since 2021 · last 2026
0009-0007-5450-6567ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021Human-computer interaction and ubiquitous computing · 5 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Computer networks · 2Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Hardware-Accelerated Algorithm for Complex Function Roots Density Graph Plotting
abstract
Solving and visualizing the potential roots of complex functions is essential in both theoretical and applied domains, yet often computationally intensive. We present a hardware-accelerated algorithm for complex function roots density graph plotting by approximating functions with polynomials and solving their roots using single-shift QR iteration. By leveraging the Hessenberg structure of companion matrices and optimizing QR decomposition with Givens rotations, we design a pipelined FPGA architecture capable of processing a large amount of polynomials with high throughput. Our implementation achieves up to 65× higher energy efficiency than CPU-based approaches, and while it trails modern GPUs in performance. Compared with state-of-the-art QR decomposition solutions, our design specificly optimize QR decomposition for complex-valued Hessenberg matrices up to size 6x6, exhibiting a moderate throughput of 16.5M QR decompositions per second, while prior works have predominantly focused on 4x4 general matrices.
Ruibai Tang, Chengbin Quan
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2024 Language-aware Visual Semantic Distillation for Video Question Answering
abstract
Significant progress in video question answering (VideoQA) have been made thanks to thriving large image-language pretraining frameworks. Although image-language models can efficiently represent both video and language branches, they typically employ goal-free vision perception and do not interact vision with language well during the answer generation, thus omitting crucial visual cues. In this paper, we are inspired by the human recognition and learning pattern and propose VideoDistill, a framework with language-aware (i.e., goal-driven) behavior in both vision perception and answer generation. VideoDistill generates answers only from question-related visual embeddings and follows a thinking-observing-answering approach that closely resembles human behavior, distinguishing it from previous research. Specifically, we develop a language-aware gating mechanism to replace the standard cross-attention, avoiding language's direct fusion into visual representations. We incorporate this mechanism into two key components of the entire framework. The first component is a differentiable sparse sampling module, which selects frames containing the necessary dynamics and semantics relevant to the questions. The second component is a vision refinement module that merges existing spatial-temporal attention layers to ensure extracting multi-grained visual semantics associated with the questions. We conduct evaluations on various challenging video question-answering benchmarks, and VideoDistill achieves state-of-the-art performance in both general and long-form VideoQA datasets. In Addition, we verify that VideoDistill can effectively alleviate the utilization of language shortcut solutions in the EgoTaskQA dataset.
Chao Yang 0026, Yu Qiao 0001, Chengbin Quan, Youjian Zhao
CVPR4
2024 Teeth-SEG: An Efficient Instance Segmentation Framework for Orthodontic Treatment Based on Multi-Scale Aggregation and Anthropic Prior Knowledge
abstract
Teeth localization, segmentation, and labeling in 2D images have great potential in modern dentistry to enhance dental diagnostics, treatment planning, and population-based studies on oral health. However, general instance segmentation frameworks are incompetent due to 1) the sub-tle differences between some teeth’ shapes (e.g., maxillary first premolar and second premolar), 2) the teeth's position and shape variation across subjects, and 3) the presence of abnormalities in the dentition (e.g., caries and edentulism). To address these problems, we propose a ViT-based frame-work named TeethSEG, which consists of stacked Multi-Scale Aggregation (MSA) blocks and an Anthropic Prior Knowledge (APK) layer. Specifically, to compose the two modules, we design a unique permutation-based upscaler to ensure high efficiency while establishing clear segmentation boundaries with multi-head selflcross-gating layers to emphasize particular semantics meanwhile maintaining the divergence between token embeddings. Besides, we collect the first open-sourced intraoral image dataset IO150K, which comprises over 150k intraoral photos, and all photos are annotated by orthodontists using a human-machine hybrid algorithm. Experiments on IO150K demonstrate that our TeethSEG outperforms the state-of-the-art segmentation models on dental image segmentation.
Shaofeng Wang, Gaoyue Sun, Feifei Zuo, Chengbin Quan, Youjian Zhao
CVPR7
2024 LLaMA-Excitor: General Instruction Tuning via Indirect Feature Interaction
abstract
Existing methods to fine-tune LLMs, like Adapter, Prefix-tuning, and LoRA, which introduce extra modules or ad-ditional input sequences to inject new skills or knowledge, may compromise the innate abilities of LLMs. In this paper, we propose LLaMA-Excitor, a lightweight method that stimulates the LLMs' potential to better follow instructions by gradually paying more attention to worthwhile information. Specifically, LLaMA-Excitor does not directly change the intermediate hidden state during the self-attention calculation. We designed the Excitor block as a bypass module that reconstructs Keys and changes the importance of Values in self-attention using learnable prompts. LLaMA-Excitor ensures a self-adaptive allocation of additional attention to input instructions, thus effectively preserving LLMs' pre-trained knowledge when fine-tuning LLMs on low-quality instruction-following datasets. Furthermore, we unify the modeling of multi-modal and language-only tuning, extending LLaMA-Excitor to a powerful visual instruction follower without the need for complex multi-modal alignment. Our approach is evaluated in language-only and multi-modal scenarios. Compared with the original LLaMA-7B, LLaMA-Excitor is the only PEFT method that maintains basic capabilities and achieves +3.12% relative improvement on the MMLU benchmark. In the visual instruction tuning, we achieve a new state-of-the-art image captioning performance on MSCOCO (157.5 CIDEr), and a comparable performance on ScienceQA (88.39%) to cutting-edge models with more parameters and extensive vision-language pertaining. The code will be available at https://zoubo9034.github.io/Excitor/.
Chao Yang 0026, Yu Qiao 0001, Chengbin Quan, Youjian Zhao
CVPR4
2023 Diversified Teaching Methods of Hardware Experiment
abstract
This article mainly takes the digital circuit experiment course as the background, and introduces the difficulties of hardware experiment course due to lack of flexibility during COVID-19. Through the continuous analysis and adjustment of teaching methods and contents, we summarize the advantages and disadvantages of online teaching and traditional experimental teaching. At the same time, we also propose hybrid hardware experiment teaching method, which has higher flexibility, and the effect is consistent with the traditional experiment teaching effect, which can ensure synchronization and meet the individual needs. In the future, both online teaching and traditional teaching will play an important role, and will no longer be two completely independent teaching methods. The hybrid hardware experiment teaching method is a new experimental teaching methods worth trying and spreading.
Yuchao Gao, Ninghan Zheng, Chengbin Quan
FIE4
2023 DFCP: Few-Shot DeepFake Detection via Contrastive Pretraining
abstract
Abuses of forgery techniques have created a considerable problem of misinformation on social media. Although scholars devote many efforts to face forgery detection (a.k.a DeepFake detection) and achieve some results, two issues still hinder the practical application. 1) Most detectors do not generalize well to unseen datasets. 2) In a supervised manner, most previous works require a considerable amount of manually labeled data. To address these problems, we propose a simple contrastive pertaining framework for DeepFake detection (DFCP), which works in a finetuning-after-pretraining manner, and requires only a few labels (5%). Specifically, we design a two-stream framework to simultaneously learn high-frequency texture features and high-level semantics information during pretraining. In addition, a video-based frame sampling strategy is proposed to mitigate potential noise data in the instance-discriminative contrastive learning to achieve better performance. Experimental results on several downstream datasets show the state-of-the-art performance of the proposed DFCP, which works at frame-level (w/o temporal reasoning) with high efficiency but outperforms video-level methods.
Jiazhi Guan, Chengbin Quan, Youjian Zhao
ICME4
2023 Dual-Modality Co-Learning for Unveiling Deepfake in Spatio-Temporal Space
abstract
The emergence of photo-realistic deepfakes on a large scale has become a significant societal concern, which has garnered considerable attention from the research community. Several recent studies have identified the critical issue of “temporal inconsistency” resulting from the frame reassembling process of deepfake generation techniques. However, due to the lack of task-specific design, the spatio-temporal modeling of current methods remains insufficient in three critical aspects: 1) inapparent temporal changes are prone to be undermined compared to abundant spatial cues; 2) minor inconsistent regions are often concealed by motions with greater amplitude during downsampling; 3) capturing both transient inconsistencies and persistent motions simultaneously remains a significant challenge. In this paper, we propose a novel Dual-Modality Co-Learning framework tailored for these characteristics, which achieves more effectual deepfake detection with complementary information from RGB and optical flow modalities. In particular, we designed a Multi-Scale Motion Regularization module to encourage the network to equally prioritize both the significant spatial cues and the subtle temporal facial motion cues. Additionally, we developed a Multi-Span Cross-Attention module to effectively integrate the information from both RGB and optical flow modalities and improve the detection accuracy with multi-span predictions. Extensive experiments validate the effectiveness our ideas and demonstrate the superior performance of our approach.
Jiazhi Guan, Hang Zhou 0009, Zhizhi Guo, Tianshu Hu, Lirui Deng 0001, Chengbin Quan, Youjian Zhao
ICMR6
2023 SpaceCLIP: A Vision-Language Pretraining Framework With Spatial Reconstruction On Text
abstract
The tremendous progress of vision-to-language retrieval over these years is fueled by contrastive vision-language pretraining (VLP), such as CLIP. Although, contrastive methods do not exhibit the same level of performance on other downstream tasks (e.g., video question answering and natural language grounding). One possible reason is they ignore the misalignment between vision and language, especially the absence of spatial information in language. To mitigate this issue, We start from a new perspective and propose a contrastive VLP framework with spatial reconstruction on text (SpaceCLIP). Specifically, we introduce a unique reconstruction method to assign text representations into the same spatial structure with images or videos and a pretraining objective, SpatialNCE, to reduce the computational overhead and ensure performance on downstream tasks. Empirically, we show SpaceCLIP outperforms other methods with performance gains ranging from 2.1% up to 9.0% on MSRVTT and EgoCLIP multiple-choice questions answering, 2.5% up to 11.0% on EPIC-KITCHENS-100 and MSRVTT multi-instance retrieval, and 0.31% up to 7.2% on Ego4D natural language query benchmark.
Chao Yang 0026, Chengbin Quan, Youjian Zhao
ACM Multimedia3
2022 Identity-Referenced Deepfake Detection with Contrastive Learning
abstract
With current advancements in deep learning technology, it is becoming easier to create high-quality face forgery videos, causing concerns about the misuse of deepfake technology. In recent years, research on deepfake detection has become a popular topic. Many detection methods have been proposed, most of which focus on exploiting image artifacts or frequency domain features for detection. In this work, we propose using real images of the same identity as a reference to improve detection performance. Specifically, a real image of the same identity is used as a reference image and input into the model together with the image to be tested to learn the distinguishable identity representation, which is achieved by contrastive learning. Our method achieves superior performance on both FaceForensics++ and Celeb-DF with relatively little training data, and also achieves very competitive results on cross-manipulation and cross-dataset evaluations, demonstrating the effectiveness of our solution.
Dongyao Shen, Youjian Zhao, Chengbin Quan
IH&MMSec3
2022 Delving into Sequential Patches for Deepfake Detection
abstract
Recent advances in face forgery techniques produce nearly visually untraceable deepfake videos, which could be leveraged with malicious intentions. As a result, researchers have been devoted to deepfake detection. Previous studies have identified the importance of local low-level cues and temporal information in pursuit to generalize well across deepfake methods, however, they still suffer from robustness problem against post-processings. In this work, we propose the Local- & Temporal-aware Transformer-based Deepfake Detection (LTTD) framework, which adopts a local-to-global learning protocol with a particular focus on the valuable temporal information within local sequences. Specifically, we propose a Local Sequence Transformer (LST), which models the temporal consistency on sequences of restricted spatial regions, where low-level information is hierarchically enhanced with shallow layers of learned 3D filters. Based on the local temporal embeddings, we then achieve the final classification in a global contrastive way. Extensive experiments on popular datasets validate that our approach effectively spots local forgery cues and achieves state-of-the-art performance.
Jiazhi Guan, Hang Zhou 0009, Zhibin Hong, Errui Ding, Jingdong Wang 0001, Chengbin Quan, Youjian Zhao
NeurIPS6
2021 Make experimental device to be a toy - the experimental teaching in Digital Logic Circuit course
abstract
This Research-to-Practice Work in Progress paper presents a type of Digital Logic Circuit experiment device like toys to attract the students' interest for the experiments. Digital Logic Circuit course is a course that emphasizes the actual hand-on experiment. Because it is fundamental course, its experiment maybe monotonous and boring for the students. The traditional experiment is hard to attract students. In this paper, we tried to turn the experimental device to the digital toys. We divided the whole experimental device into several kinds of modular blocks, which can perfectly cover the contents of the traditional experiments and greatly improve the fun of students' experiment. The experiment integrated the traditional chip connection type experiments and the programmable device experiments. The experiment blocks were colored, like the toy blocks, and it is easy to connect, so the experimental process is very interesting. The students can do the experiment like play toys, to build a simple and interesting digital circuit. We had used this device in the Digital Logic Lab course in a small scale, the students were more willing to do the experiment than others not using this device. The results proved that this kind of experiment is beneficial to train students in improving their hardware application ability, not only in mastering the knowledge effectively, but also in strengthening the ability of practice.
Ninghan Zheng, Yuchao Gao, Chengbin Quan, Weidong Liu 0001
FIE4
2019 The organization and Evaluation of Digital Logic Projects
abstract
This Research to Practice WIP presents the analysis of the project organization and evaluation in Digital Logic course. How to reasonably organize the team, analyze the feasibility of the task, and correctly evaluate each member in the team is the problem in the teaching process. The project is aimed to mobilize the team members, let them play their roles in the project, and master the curriculum knowledge. In this paper, we use different team organization, feasibility analysis of the task, and use comprehensive evaluation for the project members in the team. Through the improvement of the organization, feasibility analysis of every projects and evaluation of the team in project, students reflect well and the teaching efficiency is improved.
Chengbin Quan, Ninghan Zheng
FIE2
2017 Training students' practical and innovation ability in hardware experiment
abstract
How to effectively cultivate students' practical ability and innovative spirit is the subject of Computer Science in Colleges and Universities, especially for the first-year or second-year undergraduate students. This paper introduces the experimental teaching reform trial of the Digital Logic courses, and sums up the experience of how to stimulate students' awareness of innovation in the hardware experiment teaching and how to improve the students' practical ability. This paper proposes that we should start the student independent innovation experiment as soon as possible at the university stage. We design the independent innovation experiment in digital logic courses that experiment is an open-minded experiment. After years of experiments carried out, the students deepened understanding of the knowledge of theory course, improve the interest in the design of hardware, understand the basic processes of the design of electronic products, improve the ability of practical, and establish the consciousness of innovation and practice. Our trial has proved that it is very meaningful and feasible to enhance the ability of innovation practice in the low grade students of computer major.
Chengbin Quan
FIE3
2016 Exploration of the Computer Hardware Experiment teaching method based on the cloud platform
abstract
This paper presents a new method for Computer Hardware Experiment teaching. Previous hardware experiments of computer is built on a computer connected to a hardware equipment locally, and then student operate the hardware devices on the switches, buttons, etc. to carry out experiments and verify the developed hardware code is correct. Such experimental methods have lots of significant limitations. To solve the problems, we have developed a cloud-based real hardware experiment system platform. Students can use the experiment board anywhere internet available at any time. They can land at any place where they can access the cloud platform server, that will help them to get a hardware device, and then they can operate, and get results the same as local. We set up a cloud, and many sub-clouds in different cities. Each sub-clouds have a certain number of the special hardware experiment devices. The hardware device is assigned automatically, especially it can solve the problem of insufficient experiment hardware within a university. The platform has been set up with an experimental database, to analysis what is the difficult points for students and make a big data to help for improving computer hardware experiment educational methods. It provided a good experimental support for the online hardware experiment teaching content of MOOC.
Chengbin Quan, Youjian Zhao
FIE1
2015 Narrowing down the debugging space of slow search response time
abstract
When using search engines, users often care about search response time (SRT) in addition to result accuracy. It is thus the operators' responsibility to closely monitor and improving SRT. The first critical step of improving SRT is to pinpoint the root causes of slow SRT. However, this task is very challenging because SRT can be impacted by many factors, e.g., networks, data centers, browsers, and the page content. In this paper, we propose FOCUS, a systematic framework to narrow down the debugging space of slow SRT by identifying the bottleneck of slow SRT regarding various factors. The bottleneck provides operators more specific direction for further investigation. We deployed FOCUS in a global top search engine. Based on the output of FOCUS, operators successfully identified four potential causes which would not have been easy to find without FOCUS. Our what-if simulation analysis shows that, the proposed solutions, focusing on these bottlenecks, can improve SRT significantly, and they are more effective than some ad hoc solutions.
Youjian Zhao, Dan Pei, Chengbin Quan, Qingqian Tao, Xiyang Chen, Dai Tan, Xiaowei Jing, Mei Feng
IPCCC4
2015 Learning thresholds for PV change detection from operators' labels
abstract
Page Views (PVs) are very crucial for search engines due to their close relationship to the revenue. When PVs change significantly, operators must be informed so that they can diagnose and fix the problem quickly, and prevent further loss. In reality, PVs can be counted in many ways (e.g., PVs originated from different ISPs), and different PVs are of different interest to operators (e.g., the PVs of a larger ISP is more important). As a result, different PVs often require different detection standards, or thresholds. However, attempts to tune a number of thresholds have been hampered by the cost of the manual effort involved. To address the above problem, we propose a practical framework, called PTL (practical threshold learning). Operators only need to provide a few simple labels about the detection results, then PTL will automatically tune the thresholds for different PVs. Using 4-month PVs from a global top search engine, our evaluation demonstrates that PTL can improve the accuracy of detection dramatically. More importantly, it introduces very little labeling overhead for operators. For example, when detecting the PVs of 103 ISPs, PTL can reduce the overall false negative rate from 96% to 9% using only 29 labels per week on average.
Youjian Zhao, Kaixin Sui, Shiwen Cheng, Dan Pei, Chengbin Quan, Jiao Luo, Xiaowei Jing, Mei Feng
IPCCC6