An Intelligent Real-Time Fake News Detection Framework Using Machine Learning and Textual Feature Engineering
DOI:
https://doi.org/10.5281/zenodo.20814258Keywords:
Fake News Detection, Machine Learning, Probability Calibration, Real-Time Verification, Deep Text AnalyticsAbstract
The rapid dissemination of misinformation via digital platforms has created an urgent need for automated detection systems that can operate at scale. This paper presents an intelligent real-time fake news detection framework that integrates classical machine learning with optimized textual feature engineering. By focusing on efficient architectures, the proposed system addresses the limitations of manual fact-checking, which is increasingly inadequate in the face of modern misinformation campaigns.
The methodology involves a comprehensive text preprocessing pipeline—including conversion to lowercase, removal of URLs, numbers, and punctuation, as well as tokenization, stopword removal, and lemmatization. For feature representation, the system utilizes Term Frequency-Inverse Document Frequency with a bigram configuration, capturing up to 50,000 unique features to maintain contextual depth. The classification phase evaluates multiple models, specifically Logistic Regression, Support Vector Machines, Naive Bayes, Random Forest, and the SGD Classifier.
A key innovation of this framework is the application of probability calibration to mitigate the issue of overconfident misclassifications commonly found in standard NLP models. By applying techniques such as Platt Scaling, the system transforms raw scores into reliable posterior probabilities. Experimental evaluations show that the calibrated Linear SVM is the most effective model, achieving a peak accuracy of 97.95% and an ROC-AUC of 99.75%, surpassing traditional uncalibrated benchmarks.
The final system is deployed using a Streamlit-based web interface, providing real-time predictions with confidence scores and interactive visualizations powered by Plotly and Seaborn. The framework demonstrates that optimized traditional machine learning models can outperform more complex architectures in efficiency and reliability. These results provide a scalable solution for news platforms to verify content instantly, bridging the gap between theoretical research and practical application.