Sid-GPT: a 92M-parameter GPT built from scratch to write stronger PM resume bullets
Sid-GPT takes a weak, responsibility-style resume bullet and rewrites it the way a resume that got selected for a product role would: strong action verb, clear scope, and a quantified result.
"Responsible for managing the product roadmap for the mobile app." → "Managed the product roadmap for the mobile app, increasing the user engagement by 20%."
Everything here was built end to end by one person on a single consumer GPU (RTX 3060, 12 GB): the tokenizer, the transformer, the pretraining loop, the fine-tuning pipeline, the evaluation, and the demo. No pretrained weights, no AutoModel, no borrowed checkpoints.
Highlights
| Parameters | 91.6M (12 layers, 12 heads, 768 hidden, 512 context) |
| Tokenizer | Custom byte-level BPE, 8,000 vocab, trained on the corpus |
| Pretraining | 8,000 steps over a ~47M-token resume corpus (36.6M public tokens plus consented real resumes), from random initialisation (≈77 minutes on one RTX 3060) |
| Fine-tuning | 1,104 instruction pairs: 904 from consented candidate resumes + 200 hand-written coverage pairs |
| Quantified rewrites | 0% → 93% of unseen weak bullets come back with a metric-bearing result |
| Held-out resume perplexity | 21.1 (down from the 14M-parameter baseline's 33) |
| Scale-up | 14M → 92M parameters cut pretraining validation loss from 3.94 to 3.57 |
Quick examples
Settings: temperature=0.5, top_k=20.
| Weak bullet | Sid-GPT rewrite |
|---|---|
| Responsible for managing the product roadmap for the mobile app. | Managed the product roadmap for the mobile app, increasing the user engagement by 20%. |
| Responsible for writing user stories and working with developers. | Created user stories and worked with developers and designers to improve user experience by 25%. |
| Responsible for analyzing user behavior data with SQL. | Analyzed user behavior patterns in SQL to identify a drop-off point affecting 30% of trial users via funnel analysis in Excel. |
| Responsible for building dashboards for the business team. | Built dashboards for the business team, improving the visibility of the business by 30%. |
| Worked on an automation project to reduce manual work. | Rebuilt a self-serve workflow that eliminated manual review time by 40% and improved team productivity. |
A larger, unfiltered batch of 30 generations is in samples.json.
Try it
In the browser: open Sid_GPT_Demo.ipynb in Google Colab (File → Upload notebook) and run all cells for a shareable web app in about two minutes. The project's landing page lives in the Sid-GPT Space.
In Python:
pip install torch tokenizers huggingface_hub
huggingface-cli download Sidhartha-Rajput/sid-gpt-resume-llm --local-dir sid-gpt
cd sid-gpt
python generate.py --prompt "Responsible for running A/B tests on the checkout page."
The prompt template used in fine-tuning is
Turn this into a strong, quantified PM resume bullet:\n<your bullet>, which generate.py applies for you. Sampling at --temperature 0.4–0.6 --top_k 20 gives the cleanest phrasing.
How it was built
- Tokenizer. Byte-level BPE (8k vocab) trained directly on resume text, so common resume vocabulary like "stakeholders", "KPIs", "A/B" and "cross-functional" becomes compact tokens.
- Architecture. A decoder-only GPT written in plain PyTorch: token and position embeddings, causal multi-head self-attention, GELU MLP blocks, pre-LayerNorm, weight-tied output head. Source in
model.py. - Pretraining. Next-token prediction from random weights on a public resume corpus (InferencePrince555/Resume-Dataset) plus a set of consented real resumes. Mixed precision, AdamW, validation split held out from the public data.
- Supervised fine-tuning. Instruction/response pairs with loss computed on the response tokens only. Weak-to-strong pairs come from bullets in resumes of candidates who were selected for APM, PM, and BA roles (anonymised, with names and contact details removed), plus 200 hand-written pairs spanning growth, analytics, B2B, fintech, e-commerce, edtech, AI, and operations (
data/synthetic_pairs.json). - Evaluation. A held-out set of 99 pairs from unseen candidates, a held-out resume perplexity measurement, a result-orientation check on 30 unseen weak bullets, and a generation scan confirming no verbatim reproduction of training text and no contact details in outputs.
Training curves: pretrain_loss.png. Numbers: training_summary.json and results.json.
Data and privacy
- Candidates contributed their resumes with permission. Names, emails, phone numbers, and profile links were removed before any training text was produced, and candidates are identified only by a one-way hash.
- The raw resumes and the real fine-tuning pairs are not included in this repository. Only the model, tokenizer, code, and the 200 hand-written pairs are published.
- A scan of 30 generations found no verbatim 10-word span from the training resumes and no email, phone, or URL patterns.
What it is best at
- Turning "Responsible for X" into an action-led, outcome-oriented line.
- Giving you a structure to react to: verb → scope → measurable result.
- Producing quick drafts that you then correct with your real context.
Use it as a drafting partner. The figures it writes are illustrative placeholders that show where a metric belongs and what kind of metric fits; swap in your own real numbers before using any line on a resume. Prompts closest to its fine-tuning distribution (product, analytics, growth, operations bullets of one to two sentences) work best.
Roadmap
- Scale to a ~350M-parameter model with a larger consented corpus.
- Metric placeholders (
[X%]) mode so users fill in their own numbers. - Role-aware rewriting (APM vs PM vs BA vs data analyst).
- Preference tuning on rewrites that real hiring managers rank higher.
safetensorsexport and atransformers-compatible wrapper (export_safetensors.pyis included).
Files
| File | Purpose |
|---|---|
sft.pt |
Fine-tuned weights and config |
model.py |
The GPT architecture |
generate.py |
Standalone inference script |
tokenizer/ |
BPE vocab, merges, and tokenizer.json |
config.json |
Model hyperparameters |
data/synthetic_pairs.json |
The 200 hand-written weak→strong pairs |
samples.json |
30 generations at temperature=0.5, top_k=20 |
results.json, training_summary.json |
Metrics |
pretrain_loss.png |
Pretraining curve |
export_safetensors.py |
Optional weights export |
Citation
@misc{sidgpt2026,
title = {Sid-GPT: a from-scratch GPT for product-management resume bullets},
author = {Sidhartha},
year = {2026},
url = {https://huggingface.co/Sidhartha-Rajput/sid-gpt-resume-llm}
}
Released under the MIT license.
- Downloads last month
- 39