Author(s):
Maksym Nenashev
ABSTRACT
This paper presents the design and implementation of an AI-based multimodal content moderation system deployed in a real-world production environment. The study addresses the scalability and effectiveness challenges associated with moderating large volumes of user-generated text and image content. The proposed system integrates natural language processing, convolutional neural networks, and embedding-based facial recognition within an ensemble of eight parallel models to improve classification accuracy, robustness, and redundancy. Experimental results demonstrate strong performance, with text-classification accuracy exceeding 90% and image-detection accuracy reaching 92%, while maintaining end-to-end latency below one second. These findings confirm the practical value of multimodel architectures for large-scale digital platforms and highlight their potential to improve platform safety, operational efficiency, and broader socio-economic outcomes.
Keywords:
Artificial Intelligence, Content Moderation, Multi-model Architecture, NLP, Computer Vision, Face Recognition, MLOps, Risk Control
Pages:
105-111
UDK:
004.896:351.824.11(4)