File size: 2,917 Bytes
41279c3
 
 
 
 
 
 
 
 
924fd67
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
---
title: News Digest
emoji: πŸ“°
colorFrom: blue
colorTo: purple
sdk: docker
pinned: false
---

# OneNews

**⚠️ Experimental / Educational project β€” built to learn how scrapers and AI work from scratch. Not production-ready. Not intended to be.**

**Your daily news, scraped, scored, sorted, and served in tabs. No AI slop, no crypto bros, no "you won't believe what happens next."**

OneNews pulls from ~40 RSS feeds across 9 categories β€” Geopolitical, World Health, Tech, Cybersecurity, Funny/Weird, Gaming, Movies, Arab World, and Tunisia β€” runs them through a trust/recency/coverage scoring engine, and displays everything in a clean tabbed web UI. All written from scratch to understand every moving part.

## What makes it special?

- **Built to learn.** Everything is hand-rolled β€” no black boxes, no copy-paste agency code.
- **Zero ads, zero trackers.** Just headlines, summaries, and a relevance score.
- **Keyword-based topic detection** with regex patterns (not an LLM hallucinating categories).
- **Source-domain fallback** β€” when keywords fail, the domain decides (onion.com β†’ Funny/Weird, variety.com β†’ Movies, etc.)
- **Trust scoring** based on source rep, article length, factual language, clickbait penalties, and opinion markers.
- **Date-grouped UI** with daily headers. News doesn't care about your feed order.
- **Scoring model you can tune** β€” every weight lives in one config file.

## Quick Start

```bash
python webapp.py
```

Opens at `http://localhost:5050`. First load takes ~5 min while 40 RSS feeds get fetched and scored. Go make tea.

## Stack

| Layer | Tech |
|-------|------|
| Backend | Flask |
| Scraping | feedparser + requests + trafilatura |
| NLP | regex patterns (no transformers β€” runs on a toaster) |
| Scoring | weighted: pop 0.20 Γ— trust 0.35 Γ— coverage 0.30 Γ— recency 0.15 |
| UI | vanilla HTML/CSS/JS, zero frameworks |

## Categories

Geopolitical | World Health | Tech | Cybersecurity | Funny/Weird | Gaming | Movies | Arab World | Tunisia

## What I learned building this

- How RSS feeds actually work under the hood (you'd be surprised how many are broken)
- Trafilatura for article extraction (beats rolling your own readability)
- The surprisingly hard problem of classifying news with just regex
- Why domain-level fallbacks matter when keyword matching comes up empty
- That Flask is perfectly fine for small stuff despite what Twitter says
- Scoring is 90% tuning weights and 10% math

## FAQ

**Q: Will it scale?**  
A: It scales to your sofa. It's a single Flask process running on whatever laptop you own.

**Q: Is this production-ready?**  
A: No. Read the first sentence of this README again.

**Q: Night mode?**  
A: The UI is already dark. You're welcome.

**Q: Why 4 AM Tunisian time for refresh?**  
A: That's when the coffee kicks in.

**Q: Can I contribute?**  
A: This is a learning project. Fork it and make your own.