An enterprise-grade, multi-layered verification pipeline combining automated web ingestion, deep NLP claim extraction, live evidence retrieval, and LIME interpretability models.
Validated performance using cross-domain datasets containing over 39,000 distinct articles.
"Seamlessly bridges the gap between raw web text and explainable fact-checking intelligence."
— Core Systems Engineering SpecificationBuilt upon high-performance open-source AI frameworks
Comprehensive Capabilities
Multi-method extraction utilizing TLS impersonation (`curl_cffi`), `newspaper3k`, and BeautifulSoup with seamless manual text pasting fallbacks for secure paywalls.
Blends rapid TF-IDF + Logistic Regression baselines with fine-tuned DistilBERT models through a meta-weighting layer to optimize overall classification accuracy.
Isolates factual assertions via spaCy sentence segmentation, querying live search sources and ranking evidence by semantic similarity and source credibility.
Generates granular plain-language breakdowns and feature-weight lists illustrating precisely which words tipped the model's confidence scores.
Exposes raw scoring arrays, source-quality metrics, and technical feature weights to empower advanced fact-checkers and academic auditors.
One-click export capabilities delivering thorough documentation of extracted claims, evidence sources, numeric consistency checks, and disclaimers.
Cleaned Training Samples
Logistic Regression Baseline
DistilBERT Model Accuracy
Total Pipeline Latency
System Design & Control Flow
The system receives URLs or raw text. The scraper service handles TLS fingerprinting and fallbacks to ensure raw content retrieval.
Loaded models evaluate textual vectors concurrently through both baseline TF-IDF Logistic Regression and fine-tuned DistilBERT models.
spaCy segments sentences into atomic factual assertions, which are cross-referenced against high-quality scraped search snippets.
Results aggregate into a unified credibility tier with LIME insights, interactive metric panels, and report downloads.
| Model Architecture | Test Accuracy | Precision (Real) | Recall (Real) | F1-Score (Real) |
|---|---|---|---|---|
| Logistic Regression (TF-IDF Baseline) | 98.94% | 0.99 | 0.99 | 0.99 |
| DistilBERT (Fine-Tuned Transformer) | 99.91% | 0.999 | 0.999 | 0.999 |
| Ensemble Meta-Model (Optimized) | 99.92% | 0.999 | 0.999 | 0.999 |
Deployment Manual
1. Clone repository and set up a virtual environment:
git clone https://github.com/your-username/news-verifier.git
cd news-verifier
python -m venv venv
source venv/bin/activate
2. Install Python dependencies and language data:
pip install -r requirements.txt
python -m spacy download en_core_web_sm
3. Place pre-trained model files inside local directory structure:
models/
├── tfidf_vectorizer.pkl
├── logistic_regression.pkl
├── ensemble_meta_model.pkl
└── transformer_model/
4. Run the Streamlit web server dashboard:
streamlit run app.py