Sid-GPT: a 92M-parameter GPT built from scratch to write stronger PM resume bullets

Sid-GPT takes a weak, responsibility-style resume bullet and rewrites it the way a resume that got selected for a product role would: strong action verb, clear scope, and a quantified result.

"Responsible for managing the product roadmap for the mobile app." → "Managed the product roadmap for the mobile app, increasing the user engagement by 20%."

Everything here was built end to end by one person on a single consumer GPU (RTX 3060, 12 GB): the tokenizer, the transformer, the pretraining loop, the fine-tuning pipeline, the evaluation, and the demo. No pretrained weights, no AutoModel, no borrowed checkpoints.

Highlights

Parameters 91.6M (12 layers, 12 heads, 768 hidden, 512 context)
Tokenizer Custom byte-level BPE, 8,000 vocab, trained on the corpus
Pretraining 8,000 steps over a ~47M-token resume corpus (36.6M public tokens plus consented real resumes), from random initialisation (≈77 minutes on one RTX 3060)
Fine-tuning 1,104 instruction pairs: 904 from consented candidate resumes + 200 hand-written coverage pairs
Quantified rewrites 0% → 93% of unseen weak bullets come back with a metric-bearing result
Held-out resume perplexity 21.1 (down from the 14M-parameter baseline's 33)
Scale-up 14M → 92M parameters cut pretraining validation loss from 3.94 to 3.57

Quick examples

Settings: temperature=0.5, top_k=20.

Weak bullet Sid-GPT rewrite
Responsible for managing the product roadmap for the mobile app. Managed the product roadmap for the mobile app, increasing the user engagement by 20%.
Responsible for writing user stories and working with developers. Created user stories and worked with developers and designers to improve user experience by 25%.
Responsible for analyzing user behavior data with SQL. Analyzed user behavior patterns in SQL to identify a drop-off point affecting 30% of trial users via funnel analysis in Excel.
Responsible for building dashboards for the business team. Built dashboards for the business team, improving the visibility of the business by 30%.
Worked on an automation project to reduce manual work. Rebuilt a self-serve workflow that eliminated manual review time by 40% and improved team productivity.

A larger, unfiltered batch of 30 generations is in samples.json.

Try it

In the browser: open Sid_GPT_Demo.ipynb in Google Colab (File → Upload notebook) and run all cells for a shareable web app in about two minutes. The project's landing page lives in the Sid-GPT Space.

In Python:

pip install torch tokenizers huggingface_hub
huggingface-cli download Sidhartha-Rajput/sid-gpt-resume-llm --local-dir sid-gpt
cd sid-gpt
python generate.py --prompt "Responsible for running A/B tests on the checkout page."

The prompt template used in fine-tuning is Turn this into a strong, quantified PM resume bullet:\n<your bullet>, which generate.py applies for you. Sampling at --temperature 0.4–0.6 --top_k 20 gives the cleanest phrasing.

How it was built

  1. Tokenizer. Byte-level BPE (8k vocab) trained directly on resume text, so common resume vocabulary like "stakeholders", "KPIs", "A/B" and "cross-functional" becomes compact tokens.
  2. Architecture. A decoder-only GPT written in plain PyTorch: token and position embeddings, causal multi-head self-attention, GELU MLP blocks, pre-LayerNorm, weight-tied output head. Source in model.py.
  3. Pretraining. Next-token prediction from random weights on a public resume corpus (InferencePrince555/Resume-Dataset) plus a set of consented real resumes. Mixed precision, AdamW, validation split held out from the public data.
  4. Supervised fine-tuning. Instruction/response pairs with loss computed on the response tokens only. Weak-to-strong pairs come from bullets in resumes of candidates who were selected for APM, PM, and BA roles (anonymised, with names and contact details removed), plus 200 hand-written pairs spanning growth, analytics, B2B, fintech, e-commerce, edtech, AI, and operations (data/synthetic_pairs.json).
  5. Evaluation. A held-out set of 99 pairs from unseen candidates, a held-out resume perplexity measurement, a result-orientation check on 30 unseen weak bullets, and a generation scan confirming no verbatim reproduction of training text and no contact details in outputs.

Training curves: pretrain_loss.png. Numbers: training_summary.json and results.json.

Data and privacy

  • Candidates contributed their resumes with permission. Names, emails, phone numbers, and profile links were removed before any training text was produced, and candidates are identified only by a one-way hash.
  • The raw resumes and the real fine-tuning pairs are not included in this repository. Only the model, tokenizer, code, and the 200 hand-written pairs are published.
  • A scan of 30 generations found no verbatim 10-word span from the training resumes and no email, phone, or URL patterns.

What it is best at

  • Turning "Responsible for X" into an action-led, outcome-oriented line.
  • Giving you a structure to react to: verb → scope → measurable result.
  • Producing quick drafts that you then correct with your real context.

Use it as a drafting partner. The figures it writes are illustrative placeholders that show where a metric belongs and what kind of metric fits; swap in your own real numbers before using any line on a resume. Prompts closest to its fine-tuning distribution (product, analytics, growth, operations bullets of one to two sentences) work best.

Roadmap

  • Scale to a ~350M-parameter model with a larger consented corpus.
  • Metric placeholders ([X%]) mode so users fill in their own numbers.
  • Role-aware rewriting (APM vs PM vs BA vs data analyst).
  • Preference tuning on rewrites that real hiring managers rank higher.
  • safetensors export and a transformers-compatible wrapper (export_safetensors.py is included).

Files

File Purpose
sft.pt Fine-tuned weights and config
model.py The GPT architecture
generate.py Standalone inference script
tokenizer/ BPE vocab, merges, and tokenizer.json
config.json Model hyperparameters
data/synthetic_pairs.json The 200 hand-written weak→strong pairs
samples.json 30 generations at temperature=0.5, top_k=20
results.json, training_summary.json Metrics
pretrain_loss.png Pretraining curve
export_safetensors.py Optional weights export

Citation

@misc{sidgpt2026,
  title  = {Sid-GPT: a from-scratch GPT for product-management resume bullets},
  author = {Sidhartha},
  year   = {2026},
  url    = {https://huggingface.co/Sidhartha-Rajput/sid-gpt-resume-llm}
}

Released under the MIT license.

Downloads last month
39
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train Sidhartha-Rajput/sid-gpt-resume-llm

Space using Sidhartha-Rajput/sid-gpt-resume-llm 1