AI-Guided Learning: Research on Knowledge and Skill Acquisition Support Methods Using Deep Learning Audio-Video Processing Techniques

📅 2026-08-09
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses two key challenges in audiovisual learning: excessive content length and the lack of scalable feedback for skill imitation. To tackle these issues, the authors propose the first end-to-end AI-guided learning framework encompassing three stages—consumption, comprehension, and imitation—driven entirely by deep learning. The framework integrates three core systems: AIxSpeed, which enables phoneme-level variable playback speed based on speech recognition confidence; FastPerson, a multimodal video summarization module; and Profy, an unsupervised model for speech proficiency assessment with acoustic distance visualization. Experimental results demonstrate that AIxSpeed enhances subjective viewing experience at 1.3× playback speed, FastPerson reduces viewing time by 53% without compromising quiz performance, and Profy significantly improves pronunciation intelligibility, collectively validating the feasibility of efficient, scalable learning support using unlabeled data.
📝 Abstract
Audio and video have become major learning media, but learners face two persistent challenges: the time cost of consuming long-form content sequentially and the lack of scalable feedback for imitation-based skill acquisition. This dissertation proposes an AI-guided learning framework that supports three interconnected stages: Consume, Understand, and Imitate. It develops and evaluates three systems. AIxSpeed dynamically adjusts audio playback speed at the phoneme level using speech-recognition-model confidence as a proxy for listening difficulty. FastPerson generates multimodal video summaries that preserve visual and auditory information and lets learners switch between summarized and full versions by chapter. Profy learns proficiency from largely unannotated speech data and visualizes classifier-relevant regions and model-derived acoustic distances to support pronunciation practice. Technical and user evaluations show that AIxSpeed achieved average playback factors of 1.30x on LibriSpeech and 1.29x on UME-ERJ and received higher mean opinion scores than matched constant-speed playback; FastPerson reduced viewing time by 53% with no statistically significant difference in quiz scores compared with normal playback; and Profy showed an observed improvement in pronunciation intelligibility, with non-overlapping pre- and post-practice confidence intervals. Together, these systems demonstrate how deep learning can support efficient content consumption, multimodal understanding, and repeated skill practice while retaining learner access to the original material.
Problem

Research questions and friction points this paper is trying to address.

audio-video learning
time efficiency
skill acquisition
scalable feedback
imitation-based learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

AI-guided learning
deep learning
multimodal summarization
dynamic audio speed adaptation
proficiency modeling
🔎 Similar Papers
No similar papers found.