🎯 Research Overview
This research investigates the application of Natural Language Processing (NLP) and Text Mining techniques for detecting fraudulent activities on social media platforms. The study focuses on three specific fraud types: cyberbullying, fake reviews, and misinformation.
Author: Santosh Poudel
Degree: MSc Information Technology
Institution: University of the West of Scotland
Completion Date: December 2023
📊 Abstract
This research addresses the significant challenges in identifying fraudulent activities in social media environments. The project explores the effectiveness of NLP techniques, including Naive Bayes classification, TF-IDF, and N-gram modelling, across three distinct datasets. The study demonstrates adaptive fraud detection capabilities with accuracy rates up to 86% on fake review detection.
🚀 Key Features
- Multi-dataset analysis: Cyberbullying, Fake Reviews, and Misinformation datasets
- Advanced NLP techniques: TF-IDF, Unigram/Bigram features, Naive Bayes
- Comprehensive evaluation: Precision, Recall, F1-Score metrics
- Real-world application: Social media fraud detection framework
📈 Key Results
| Dataset | Best Model | Accuracy | Precision | Recall | F1-Score |
|---|---|---|---|---|---|
| Cyberbullying | Unigram Features | 73% | 71.8% | 81.4% | 75.8% |
| Fake Reviews | Combined Features | 86% | 82% | 93% | 87% |
| Misinformation | Combined Features | 81% | 78% | 91% | 84% |
🛠️ Technical Implementation
Algorithms & Techniques
- Machine Learning: Naive Bayes Classifier
- Feature Extraction: TF-IDF, Unigram, Bigram, Combined features
- NLP Processing: Tokenization, Stemming, Lemmatization, Stop-word removal
- Evaluation Metrics: Confusion Matrix, Accuracy, Precision, Recall, F1-Score
Tools & Technologies
- Programming: Python 3.12.0
- Libraries: NLTK, Scikit-learn, Pandas, NumPy, Matplotlib, Seaborn
- Platform: Google Colaboratory
- Data Source: Kaggle datasets
