Video summarization using deep image captioning models
Abstract
This research presents a novel approach for video summarization by leveraging deep image captioning models. A pretrained image captioning model, namely Salesforce's bootstrapped language image pretraining (BLIP), is used to extract keyframes from a video at regular intervals and produce natural language descriptions. These textual descriptions are then filtered for non-repetition and concatenated into a coherent summary, allowing users to understand the video’s content without viewing it in full. The proposed framework aims to improve video browsing, indexing, and retrieval efficiency, particularly for big datasets. The proposed method achieves a significant reduction in redundancy by 40% compared to raw captioning sequences. Evaluation using semantic consistency checks demonstrates that the BLIP-based framework maintains high descriptive accuracy even in complex scenes, providing a scalable solution for large-scale video indexing.
Keywords
BLIP; Deep learning; Image captioning; Keyframe extraction; Semantic analysis; Transformers; Video summarization
Full Text:
PDFDOI: http://doi.org/10.11591/ijeecs.v43.i2.pp547-554
Refbacks
- There are currently no refbacks.

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
Indonesian Journal of Electrical Engineering and Computer Science (IJEECS)
p-ISSN: 2502-4752, e-ISSN: 2502-4760
This journal is published by the Institute of Advanced Engineering and Science (IAES).