科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ International Journal of Research and Development in Engineering Sciences2026-03-06· Computer science

A Privacy-Preserving Universal Multimodal Framework for Real-Time Any-to-Any Transformation

Ramakrishna Kolikipogu, Aruna Devarakonda, Swathi Palamakula

原始摘要(英文原文)· Original abstract
We propose a unified multimodal framework, Universal Multi-Modal Generation Enabling Any-to-Any Transformation, that enables seamless transformation between text, audio, and image inputs and outputs. The system integrates three core capabilities: speech understanding using Whisper, visual understanding through LLaVA, and speech synthesis via PyTorch-based text-to-speech models. All modules are deployed on-premise using Docker, providing a privacy-centric execution environment and reducing operational overhead associated with cloud processing. The framework supports advanced workflows including document/PDF-to-text extraction, text-to-speech conversion, and image-driven description generation, thereby enabling accessible and interactive multimodal content pipelines. The implementation emphasizes efficient orchestration and inference to meet real-time constraints. Experimental results across multiple cross-modal tasks demonstrate robust accuracy and consistently low latency, suggesting that local, containerized multimodal systems can deliver scalable performance for practical applications. The proposed approach is particularly relevant to accessibility, education, and content creation, where rapid modality conversion and data privacy are essential.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

A Privacy-Preserving Universal Multimodal Framework for Real-Time Any-to-Any Transformation — 科研速览 Science Skim