Buckets:
| Name | Size | Uploaded | Xet hash |
|---|---|---|---|
| code | 21 items | ||
| processed | 930 items | ||
| raw | 2,337 items | ||
| README.md | 5.97 kB xet | 2ae9d8ef | |
| STATUS.md | 2.04 kB xet | 6925b158 |
SkillsStorage — agent-skills corpus (SKILL.md bundles)
Corpus of agent skills for training a skill-retrieval embedding model. Built from the
"Skills Data Ledger" (2026-09-27). Bucket: hf://buckets/Mercity/SkillsStorage.
Live progress: STATUS.md (regenerated every 30 min while jobs run).
A skill is a bundle, not a file: a folder holding SKILL.md plus any scripts/, references/,
assets/, templates, evals, etc. Every source is normalised to two tables that join on skill_uid:
skills (one row per bundle, with the SKILL.md text) and files (every other file of the bundle).
Layout
raw/ untouched source data (re-process from here)
hf/<owner>__<dataset>/ server-side copies of 20 Hugging Face datasets (GitSkills, ClawHub dump, other skill sets, benchmarks)
vendor/ 31 official/curated GitHub repos, one tar.gz each (whole repo minus .git) + manifest.json
clawhub_api/ ClawHub public API: listing.jsonl (all public skills) + zips/ (bundles newer than the HF dump)
github_delta/<run>/ GitHub crawl shards (one row per file) + repos.jsonl (per-repo outcome) + done.txt
seeds/ (local) repo seed lists used by the crawl (data/seeds in the code repo)
processed/
gitskills/skills/ 1,877,981 rows — every distinct SKILL.md content of the Jul-2026 GitHub census + repo metadata
gitskills/files/ 5,858,945 rows — bundle files from GitSkills' artifact_siblings (text when GitSkills fetched it)
gitskills/occurrences/ 3,797,117 rows — every SKILL.md occurrence incl. verbatim copies (repo, path, file_sha)
clawhub/skills.parquet 80,961 ClawHub skills (HF dump 2026-09-28) + ClawScan / VirusTotal / SkillSpector verdicts
clawhub/files.parquet 291,125 bundle files of those skills
clawhub/api_listing.parquet 81,082 public ClawHub skills: downloads, installs, stars, versions, topics
clawhub/api_delta_*.parquet 1,659 skills created / re-versioned after the dump (+21,810 bundle files)
github_delta/<run>/{skills,files}/ GitHub crawl, same layout; runs below
code/ every script that produced the above (see "Pipeline")
GitHub crawl runs (processed/github_delta/<run>):
| run | what | repos |
|---|---|---|
pass1_sourcegraph |
popular repos (Sourcegraph slice) with SKILL.md that are not in GitSkills | 8,386 (complete) |
pass2_new |
repos created since July 2026 (GitHub topic search) + skills.sh / claudemarketplaces / Anthropic-marketplace repos | 29,296 (complete) |
backfill_gitskills |
GitSkills repos with multi-file skills, re-fetched so their bundles are complete (GitSkills' own folder listings are truncated for 236k skills and ~1/3 of its sibling files lack text) | 136,218 (in progress) |
Schemas
skills (all sources): skill_uid, source, repo / owner+slug, skill_dir, skill_md (full SKILL.md text),
name, description, frontmatter_valid, body_chars, cjk_ratio, bundle counts/bytes, plus source extras:
- gitskills:
file_sha(git blob SHA),location_class,has_scripts,has_references,sibling_count,composition_truncated, commit dates,repo_license,repo_stars,repo_forks,repo_is_fork,repo_language, … - github_delta:
commit,file_sha(git blob SHA),n_bundle_files_with_content,fetched_at,via(git | codeload) - clawhub:
version,license(MIT-0), scanner verdicts (clawscan_verdict,virustotal_*,skillspector_*, …)
files (all sources): skill_uid, rel_path (path inside the skill folder), size, blob_sha/sha256, content (text),
content_bin (small binaries), has_content, skipped (why content is absent: bundle_cap >500 files or >8 MB per bundle,
repo_cap >100 MB per repo, too_large >1 MB file, binary_large binary >256 KB, not_fetched_by_gitskills, …).
Nothing is dropped silently: files without content are still listed with path, size and reason.
skill_uid: gh:{repo}:{skill_dir} (GitSkills, Jul 2026), gh:{repo}:{skill_dir}@{commit12} (our crawl),
clawhub:{owner}/{slug}@{version}. A SKILL.md at the repo root has skill_dir = "." (bundle = whole repo, capped).
file_sha/blob_sha are git blob SHAs in both GitSkills and our crawl, so exact copies dedup across sources.
Sources and decisions
- Bucket A (no account, bulk): GitSkills, ClawHub dump + API delta, other HF skill sets (raw only; mostly overlap GitHub), benchmarks, vendor repos.
- Bucket B (no account, crawl): GitHub passes above. Anonymous git is throttled at times; the crawler falls back to codeload tarballs (results verified identical to git, incl. blob SHAs).
- Out of scope by decision: Chinese-only registries (Tencent SkillHub, ModelScope). CJK-heavy skills elsewhere are kept and flagged with
cjk_ratio.
License cautions
- GitHub content keeps its origin repo's license (~55% of repos have none): filter on
repo_license/ the repo. anthropics/skillsdocx/pdf/pptx/xlsx skills are source-available, not OSS;trailofbits/skillsis CC-BY-SA-4.0.- ClawHub is MIT-0 throughout. Scanner-flagged (malicious/suspicious) skills are kept but flagged.
Pipeline (code/)
All long jobs run as systemd services via run_job.sh (memory caps, auto-restart, resumable); watchdog.sh relaunches
dead jobs and writes STATUS.md; memguard.sh freezes low-priority jobs when RAM is low.
mirror_hf.py (bucket A copies) · vendor_repos.py · clawhub_api.py + build_clawhub_api.py · build_clawhub.py ·
build_gitskills.py · github_topic_seeds.py + build_pass2_seeds.py · github_bundles.py (crawler) ·
build_github_delta.py + ship_delta.py (convert → upload → delete) · crawl_chain.sh · status.py.
Secrets
The HF token lives only in a local, git-ignored .env. Nothing in this bucket contains it.
- Total size
- 63.7 GB
- Files
- 3,288
- Last updated
- Oct 2
- Pre-warmed CDN
- US EU US EU