PRODUCTION MLOPS ARCHITECTURE FOR SCALABLE DEPLOYMENT OF GENERATIVE AI IN HEALTHCARE PAYER SYSTEMS
Abstract
Healthcare payer organizations increasingly deploy Generative Artificial Intelligence to automate claims adjudication, prior authorization, policy interpretation, customer support, and revenue cycle management. However, production deployment of Generative AI requires robust MLOps infrastructure capable of ensuring scalability, reliability, security, governance, and continuous model lifecycle management. Conventional AI deployment strategies frequently encounter challenges related to model drift, governance, monitoring, deployment automation, and regulatory compliance. This paper proposes a Production MLOps architecture integrating cloud-native deployment, Retrieval-Augmented Generation, enterprise knowledge management, automated model governance, continuous integration and deployment, monitoring, and scalable inference services for healthcare payer systems. The proposed framework enables reliable deployment, version management, model monitoring, prompt management, security governance, and automated optimization throughout the AI lifecycle. Experimental evaluation demonstrates improvements in deployment reliability, operational scalability, inference performance, governance effectiveness, regulatory compliance, and administrative productivity. The proposed architecture provides a secure, explainable, and production-ready foundation for enterprise Generative AI deployment in healthcare payer organizations. Keywords— MLOps, Generative Artificial Intelligence, Healthcare Payer Systems, Cloud Computing, Retrieval-Augmented Generation, Model Governance, Enterprise AI, Healthcare Claims