An Intelligent Real-Time Fake News Detection Framework Using Machine Learning and Textual Feature Engineering

Authors

  • Hamid Ghous Department of Computer Science & Information Technology, University of Southern Punjab (USP) Multan, Punjab, Pakistan Author
  • Syed Zulfiqar Haider Bukhari Directorate of Information Technology, Islamia University Bahawalpur (IUB), Punjab, Pakistan Author
  • Muhammad Talah Zubair Department of Computer NCBA&E Lahore Sub Campus Multan, Punjab, Pakistan Author
  • Muhammad Akmal Shahzad Department of Computer Science & Information Technology, University of Southern Punjab (USP) Multan, Punjab, Pakistan Author
  • Ghulam Muhy Ud Deen Raee Department of Computer Science & Information Technology, University of Southern Punjab (USP) Multan, Punjab, Pakistan Author
  • Wajiha Qavi Department of Computer Science & Information Technology, University of Southern Punjab (USP) Multan, Punjab, Pakistan Author

DOI:

https://doi.org/10.5281/zenodo.20814258

Keywords:

Fake News Detection, Machine Learning, Probability Calibration, Real-Time Verification, Deep Text Analytics

Abstract

The rapid dissemination of misinformation via digital platforms has created an urgent need for automated detection systems that can operate at scale. This paper presents an intelligent real-time fake news detection framework that integrates classical machine learning with optimized textual feature engineering. By focusing on efficient architectures, the proposed system addresses the limitations of manual fact-checking, which is increasingly inadequate in the face of modern misinformation campaigns.

The methodology involves a comprehensive text preprocessing pipeline—including conversion to lowercase, removal of URLs, numbers, and punctuation, as well as tokenization, stopword removal, and lemmatization. For feature representation, the system utilizes Term Frequency-Inverse Document Frequency with a bigram configuration, capturing up to 50,000 unique features to maintain contextual depth. The classification phase evaluates multiple models, specifically Logistic Regression, Support Vector Machines, Naive Bayes, Random Forest, and the SGD Classifier.

A key innovation of this framework is the application of probability calibration to mitigate the issue of overconfident misclassifications commonly found in standard NLP models. By applying techniques such as Platt Scaling, the system transforms raw scores into reliable posterior probabilities. Experimental evaluations show that the calibrated Linear SVM is the most effective model, achieving a peak accuracy of 97.95% and an ROC-AUC of 99.75%, surpassing traditional uncalibrated benchmarks.

The final system is deployed using a Streamlit-based web interface, providing real-time predictions with confidence scores and interactive visualizations powered by Plotly and Seaborn. The framework demonstrates that optimized traditional machine learning models can outperform more complex architectures in efficiency and reliability. These results provide a scalable solution for news platforms to verify content instantly, bridging the gap between theoretical research and practical application.

 

Downloads

Download data is not yet available.

Downloads

Published

2026-06-22

How to Cite

An Intelligent Real-Time Fake News Detection Framework Using Machine Learning and Textual Feature Engineering. (2026). Annual Methodological Archive Research Review, 4(6), 230-241. https://doi.org/10.5281/zenodo.20814258

Similar Articles

1-10 of 963

You may also start an advanced similarity search for this article.