Abstract
The rampant
spread of fake news in digital media poses a serious threat
to public trust,
democratic stability, and societal well-being. This study focuses on developing an automated
fake news detection system using
Python and Natural Language Processing (NLP) techniques. Employing benchmark
datasets such as the Fake and Real News Dataset and the LIAR dataset, the
research explores both traditional machine learning classifiers (Logistic
Regression, Naïve Bayes, SVM, Random Forest, XGBoost) and advanced deep
learning models (LSTM, BiLSTM, and BERT). Text preprocessing techniques like
tokenization, stopword removal, lemmatization, and vectorization (TF-IDF,
Word2Vec, BERT embeddings) are applied to convert raw news text into
machine-readable features. Performance evaluation using metrics
such as accuracy,
precision, recall, F1-score, and ROC-AUC indicates
that transformer-based BERT significantly outperforms other models,
achieving an accuracy of 96.2% and an AUC of 0.973. The findings highlight the
potential of contextual embeddings in fake news detection and reinforce the
effectiveness of Python's NLP ecosystem in automating content credibility assessment. This research contributes to combating misinformation by offering scalable and accurate solutions for
real-world deployment in digital platforms.