will-it-bang
Can you predict a tweet's engagement from only what's knowable before it's posted?
This project says yes β and shows how. We train a LightGBM model on ~110k tweets from the Twitter Community Archive to forecast the exposure-adjusted engagement rate of a tweet using:
- The tweet text itself
- Structural features (length, URLs, media, mentions, emoji, etc.)
- The author's pre-tweet history of replies received, quotes received, and mentions received (never dump-time totals)
- Reply/quote parent context from both CA and the X API
- Calendar / platform-era signals
- A semantic embedding of the text (MiniLM β PCA + residual engagement CAV)
- Lexical features (TF-IDF β SVD)
- An account-level prior capturing the author's typical engagement rate
Crucially, nothing leaks the future. No dump-time favorite/retweet counts as features. No current follower counts (they're snapshots, not histories). No undated likes. The model sees only what was knowable at posting time.
Scoreboard
The primary metric is within-account Spearman by stratum β given an author's tweets, how well does the model rank them by eventual engagement?
| Metric | Value |
|---|---|
| Within-account Spearman (non-replies, weighted mean) | 0.246 |
| Within-account Spearman (replies, weighted mean) | 0.307 |
| Global Spearman (non-replies) | 0.616 |
| Rate RMSE | 0.588 |
| Training rows / accounts | 88,500 / 868 |
| Validation rows / accounts | 22,125 / 594 |
| Features | 326 |
Release-v1 checkpoint trained on X-complete full data (~110k / 934 accounts).
What makes this hard?
1. Leakage is everywhere. The naive approach uses a tweet's eventual favorites/retweets as features for that same tweet. Even "just the author's average engagement on previous tweets" is dangerous if that average includes dump-time totals that post-date the candidate. We ban all of that.
2. The signal is sparse. Most tweets get little engagement. The model must pick out the ~10% that will outperform an author's own median.
3. Reply vs. standalone tweets behave differently. A thoughtful reply to a hot thread has a very different engagement profile than an original post. We model these separately and finally blend them with a within-author ranker.
How it works
Training pipeline
enriched_tweets.parquet (CA dump, ~900MB)
β
βββββββ΄βββββββββββ
β prepare.py β βββ Feature engineering
β β - Text β MiniLM β PCA
β β - TF-IDF β SVD
β β - Ego history (replies/quotes/mentions received)
β β - Parent context (CA + X API fallback)
β β - Calendar / platform eras
βββββββ¬βββββββββββ
β
βββββββ΄βββββββββββ
β dataset.parquet β
βββββββ¬βββββββββββ
β
βββββββ΄βββββββββββ
β train.py β βββ Training
β β - Exposure-rate target (maturity-corrected)
β β - Account target encoding (LOO + shrinkage)
β β - Residual embedding CAV
β β - Asymmetric MSE loss
β β - Hybrid level + within-author LambdaRank
βββββββ¬βββββββββββ
β
βββββββ΄βββββββββββ
β model.txt β βββ LightGBM booster
β + transforms β + PCA, TF-IDFβSVD, residualizer
ββββββββββββββββββ
Feature groups
| Group | Description | Count |
|---|---|---|
| Structure | text_len, has_url, has_media, n_mentions, caps_ratio, is_reply, etc. | ~20 |
| Calendar | hour, day-of-week, month, year, platform era (pre-timeline β musk_late) | ~12 |
| Ego history | inbound replies/quotes/mentions (last 7d, 30d, all-time, per-tweet rates), crosses with media/reply flags | ~32 |
| Parent context | reply/quote parent MiniLM cosine, parent age, parent's own pre-T replies, same-author flag | ~16 |
| URL structure | text urls, domain-level features (twitter, youtube, github, etc.) | ~8 |
| Account prior | account_mean_rate (shrunken LOO prior, caller-supplied) + log_account_n |
2 |
| Residual CAV | projection of MiniLM residual onto engagement direction (after regressing out baseline) | 2 |
| Embedding PCA (MiniLM) | 96-d PCA of all-MiniLM-L6-v2 |
96 |
| Text SVD | 128-d TF-IDF β TruncatedSVD | 128 |
The target
y = log1p(eng / (min(age_days, 30) + 1))
where eng = favorite_count + retweet_count
This is an exposure rate β the intensity of engagement per unit tweet lifetime, capped at 30 days so dump maturity doesn't distort scores. Retweets (RT @...) are excluded from training.
Leakage rules (non-negotiable)
| Allowed | Banned |
|---|---|
| Pre-T reply/quote/mention arrivals to the author | Dump-time favorite_count / retweet_count of any tweet |
| Parent tweet content + pre-T parent reply count | Current all_account.num_followers (snapshot, leaks future) |
| Calendar / platform era | account_display_name (mutable, no as-of-T history) |
| Text and its structure | Undated likes |
| Account-level shrunken mean of prior tweets' rates | Username fed into an LM for "reputation" scoring |
What's in this repository
Inference bundle
| File | Role |
|---|---|
model_level.txt |
LightGBM level booster (exposure-rate target) |
model_rank_nr.txt / model_rank_rp.txt |
Hybrid stratum rankers (non-reply / reply) |
emb_pca.joblib |
MiniLM β PCA transform |
tfidf.joblib + text_svd.joblib |
TF-IDF β TruncatedSVD |
res_emb_cav_direction.npy |
Residual engagement CAV direction (384-d) |
res_emb_baseline_coef.npy + res_emb_baselines.json |
OLS residualizer coefficients |
feature_schema.json |
Ordered features, dtypes, defaults, groups |
constants.json |
Global priors, encoder ID, exposure parameters |
metrics.json |
Full holdout evaluation |
predict.py |
Reference scorer (JSON in β score out) |
The bundle excludes: training data, per-account TE maps, emb-kNN train index, CA graph caches. Those are private to the research pipeline.
Research pipeline (not in this repo)
prepare.pyβ Full feature engineering from CA parquet dumptrain.pyβ Model training with all ablation knobsanalyze.pyβ Detailed evaluation with per-stratum breakdownsexperiments.mdβ Complete experiment log (50+ runs with outcomes)AGENTS.mdβ Technical guidance for the data and leakage rules
Quick start: scoring a tweet
# Install dependencies
pip install lightgbm scikit-learn numpy pandas joblib pyarrow sentence-transformers
# Download the bundle from Hugging Face
# (you're looking at it β clone or download the files above)
# Score a single tweet
cat > tweet.json << 'EOF'
{
"tweet_id": "123456789",
"full_text": "just shipped something i've been working on for months. feels good.",
"is_reply": false,
"has_media": false,
"has_url": false,
"n_mentions": 0,
"n_hashtags": 0,
"account_mean_rate": 0.15,
"log_account_n": 3.5
}
EOF
python predict.py --bundle . --input tweet.json
For full context (parent tweets, ego history), populate the remaining schema fields (see feature_schema.json). Missing fields default to 0 or the global mean.
The account_mean_rate recipe
This is the most important caller-supplied value. Given the author's earlier tweets with known engagement:
- For each tweet in the author's history, compute the same exposure rate
y_i - Let
n= number of historical tweets,y_bar= their mean rate - Shrink toward the global prior using
m = 20:
account_mean_rate = (n * y_bar + m * global_mean) / (n + m)
log_account_n = log1p(n)
If n = 0: use global_mean β 0.151, log_account_n = 0.
Research process
This project ran over 50 experiments across 18 months. The full log is in experiments.md, but the headline lessons are:
- Signal > technique. The biggest lifts came from denser leakage-safe history (inbound replies/quotes, mentions, parent context from X API) β not from embedding discovery, LM labels, or modeling tricks.
- Metric selection matters. Within-author ranking (does the model rank an author's own tweets correctly?) is harder and more useful than cross-author level accuracy.
- Failure modes are open. Viral non-replies are still under-predicted. Media/image understanding is weak. Reply dyad features don't yet pierce the noise floor.
License
Apache 2.0
Citation
@misc{will-it-bang,
author = {Chris Carroll Smith},
title = {will-it-bang: Predicting Tweet Engagement from Pre-Publication Features},
year = {2026},
howpublished = {\url{https://huggingface.co/chriscarrollsmith/will-it-bang}}
}
Built with the Twitter Community Archive data, LightGBM, and sentence-transformers/all-MiniLM-L6-v2.