科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ International Journal of Computer Information Systems and Industrial Management Applications2026-08-19· Closed captioning

Novel Method for Caption Based Description in Image using Machine Learning

Rachita Dubey, Vivek Shukla

原始摘要(英文原文)· Original abstract
The challenging research area of image captioning aims to produce written descriptions that faithfully capture the atmosphere and events captured in a photograph. In order to recognize objects interpret changing conditions and understand the semantic connections within the image advanced machine learning algorithms are required. This study builds an image captioning system using CNN, RNN and ResNet architectures. Here the MSCOCO benchmark dataset is used to enable precise scenario inference with CNN acting as the encoder and RNN as the decoder. By utilizing skip connections and avoiding multiple convolutional layers the system uses ResNet which efficiently utilizes its layers and reduces computation time. This approach resolves the gradient explosion problem and enhances model performance. The proposed model performs better than previous implementations in several evaluation criteria such as BLEU, METEOR, CIDEr and ROUGE. Using the Pillow library the system preprocesses images to enhance brightness and minimize size for optimal training. Additionally the models predicted accuracy is raised by utilizing the TorchVision library. With these enhancements the system can now generate high-quality captions with precise semantic understanding and contextual value.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Novel Method for Caption Based Description in Image using Machine Learning — 科研速览 Science Skim