Design and implementation of an Intelligent Unified Social Media Monitoring Framework: A Multimodal Generative AI Framework

Authors

  • Hammad Ali Huboweb Technologies Private Limited, Tele Tower, Model Town, Lahore 54000, Pakistan Author
  • Muhammad Khurram Zahur Bajwa Department of Management and Innovation Systems, University of Salerno, 84084, Fisciano, Italy, Author
  • Nasir Ayub Engineering Calrom Limited, M16EG, United Kingdom Author
  • Umair Ghafoor Engineering Calrom Limited, M16EG, United Kingdom Author
  • Zainab Shehzadi FAST National University of Computer and Emerging Sciences, Lahore, Pakistan Author
  • Muhammad Zulkifl Hasan Faculty of Information Technology, University of Central Punjab, Lahore, Pakistan Author
  • Muhammad Zunnurain Hussain Bahria University Lahore Campus Author
  • Ali Irfan Department of Computer Science, Faculty of Computer Science & IT, Superior University Lahore, 54000, Pakistan Author
  • Sehar Fayyaz Department of Computer Science, Faculty of Computer Science & IT, Superior University Lahore, 54000, Pakistan Author

DOI:

https://doi.org/10.5281/zenodo.20717170

Keywords:

 Multimodal Generation, Style Control, Social Media Automation, Generative AI, Content Personalisation

Abstract

With the increasing popularity of social media, however, there's an absolute lack of preference for the creation of the personalised version, something that is very hard to create efficiently across media (multimodal channels) in the same style. Often dealing with urgent problems within this context, automatic creation of texts with correct mastery of their stylistic features adapted to the various cultures of platforms, this study complements the urgent problem mentioned above. This proposed research is novel in the sense that it introduces a novel "Multimodal Generative Framework (MMGF)" to facilitate the synchronised text-image-video generation process while providing adaptivity, adaptable style variation, and modulation based on LLM, diffusion model and reinforcement learning. This method includes a style embedding process (dynamic approach) to guarantee consistency across several modalities, and a reward-based optimisation process to optimise content against participation associated with the platform, measured by metrics. The system runs on the AWS cloud platform, and is tested with multiple large-scale experiments with over one million labelled posts from Instagram, TikTok and Twitter. The results show an improvement of 15 times the production speed in relation to the creation done manually and a style preservation accuracy of 89% for both modalities. 83% of all content that is generated is deemed suitable for direct publishing in user studies, whereas the traditional way of generating content in the pipeline seems poor. The effort would encompass not only the general technical aspects of modularisation, but also basic practical aspects, which would decrease the barriers for small content creators. The features are also shared with the research community, which helps to initiate valuable discussions on ethical issues related to synthetic media on the open-source SocialGen-AI platform. The functions are also made available to the research community, enabling informed conversations around the ethical implications of synthetic media on the open-source SocialGen-AI platform. The results set a new standard for AI-assisted content brands on both a local and cultural level.

The system is based on a dual-layer prompting approach: it normalises the user intent with a large instruction following model, and generates Text, Image, Video materials using both Imagen and Phenaki. The perception module, which is based on the YOLOv8, ensures quality assurance by segmenting and editing the images as well as detecting errors, while human-in-the-loop systems will correct errors like fabrications and inconsistencies in terms of styling the images. To predict engagement, we tested a few regression techniques with the public Instagram Reach Analysis Data, which are available on Kaggle, including the following: Impressions, reach, likes, comments, shares, and profile visits. The tests performed indicated that the data were linear, as Linear Regression had the highest accuracy (R2 = 99.96 %). The other classifiers gave good results: Random Forest: R2 value 89.72, Neural Networks: R2 value 90.06 and KNN: satisfactory results - R2 value 72.82, Decision Tree: poor results - R2 value 44.34. The comparison of the actual engagements with the projected engagements is reinforced with the charts. This all-in-one solution will save time in creating content and planning, and deliver reliable engagement predictions that will turn into a valuable, scalable and brand-safe answer for businesses and small companies.

 

 

Downloads

Download data is not yet available.

Downloads

Published

2026-06-14

How to Cite

Design and implementation of an Intelligent Unified Social Media Monitoring Framework: A Multimodal Generative AI Framework. (2026). Annual Methodological Archive Research Review, 4(6), 138-178. https://doi.org/10.5281/zenodo.20717170

Similar Articles

141-150 of 1712

You may also start an advanced similarity search for this article.

Most read articles by the same author(s)

1 2 3 > >>