Buckets:

63.7 GB
3,288 files
Updated less than a minute ago
Name
Size
code
processed
raw
README.md5.97 kB
xet
STATUS.md2.04 kB
xet
README.md

SkillsStorage — agent-skills corpus (SKILL.md bundles)

Corpus of agent skills for training a skill-retrieval embedding model. Built from the "Skills Data Ledger" (2026-09-27). Bucket: hf://buckets/Mercity/SkillsStorage. Live progress: STATUS.md (regenerated every 30 min while jobs run).

A skill is a bundle, not a file: a folder holding SKILL.md plus any scripts/, references/, assets/, templates, evals, etc. Every source is normalised to two tables that join on skill_uid: skills (one row per bundle, with the SKILL.md text) and files (every other file of the bundle).

Layout

raw/                               untouched source data (re-process from here)
  hf/<owner>__<dataset>/           server-side copies of 20 Hugging Face datasets (GitSkills, ClawHub dump, other skill sets, benchmarks)
  vendor/                          31 official/curated GitHub repos, one tar.gz each (whole repo minus .git) + manifest.json
  clawhub_api/                     ClawHub public API: listing.jsonl (all public skills) + zips/ (bundles newer than the HF dump)
  github_delta/<run>/              GitHub crawl shards (one row per file) + repos.jsonl (per-repo outcome) + done.txt
seeds/ (local)                     repo seed lists used by the crawl (data/seeds in the code repo)
processed/
  gitskills/skills/                1,877,981 rows — every distinct SKILL.md content of the Jul-2026 GitHub census + repo metadata
  gitskills/files/                 5,858,945 rows — bundle files from GitSkills' artifact_siblings (text when GitSkills fetched it)
  gitskills/occurrences/           3,797,117 rows — every SKILL.md occurrence incl. verbatim copies (repo, path, file_sha)
  clawhub/skills.parquet           80,961 ClawHub skills (HF dump 2026-09-28) + ClawScan / VirusTotal / SkillSpector verdicts
  clawhub/files.parquet            291,125 bundle files of those skills
  clawhub/api_listing.parquet      81,082 public ClawHub skills: downloads, installs, stars, versions, topics
  clawhub/api_delta_*.parquet      1,659 skills created / re-versioned after the dump (+21,810 bundle files)
  github_delta/<run>/{skills,files}/   GitHub crawl, same layout; runs below
code/                              every script that produced the above (see "Pipeline")

GitHub crawl runs (processed/github_delta/<run>):

run what repos
pass1_sourcegraph popular repos (Sourcegraph slice) with SKILL.md that are not in GitSkills 8,386 (complete)
pass2_new repos created since July 2026 (GitHub topic search) + skills.sh / claudemarketplaces / Anthropic-marketplace repos 29,296 (complete)
backfill_gitskills GitSkills repos with multi-file skills, re-fetched so their bundles are complete (GitSkills' own folder listings are truncated for 236k skills and ~1/3 of its sibling files lack text) 136,218 (in progress)

Schemas

skills (all sources): skill_uid, source, repo / owner+slug, skill_dir, skill_md (full SKILL.md text), name, description, frontmatter_valid, body_chars, cjk_ratio, bundle counts/bytes, plus source extras:

  • gitskills: file_sha (git blob SHA), location_class, has_scripts, has_references, sibling_count, composition_truncated, commit dates, repo_license, repo_stars, repo_forks, repo_is_fork, repo_language, …
  • github_delta: commit, file_sha (git blob SHA), n_bundle_files_with_content, fetched_at, via (git | codeload)
  • clawhub: version, license (MIT-0), scanner verdicts (clawscan_verdict, virustotal_*, skillspector_*, …)

files (all sources): skill_uid, rel_path (path inside the skill folder), size, blob_sha/sha256, content (text), content_bin (small binaries), has_content, skipped (why content is absent: bundle_cap >500 files or >8 MB per bundle, repo_cap >100 MB per repo, too_large >1 MB file, binary_large binary >256 KB, not_fetched_by_gitskills, …). Nothing is dropped silently: files without content are still listed with path, size and reason.

skill_uid: gh:{repo}:{skill_dir} (GitSkills, Jul 2026), gh:{repo}:{skill_dir}@{commit12} (our crawl), clawhub:{owner}/{slug}@{version}. A SKILL.md at the repo root has skill_dir = "." (bundle = whole repo, capped). file_sha/blob_sha are git blob SHAs in both GitSkills and our crawl, so exact copies dedup across sources.

Sources and decisions

  • Bucket A (no account, bulk): GitSkills, ClawHub dump + API delta, other HF skill sets (raw only; mostly overlap GitHub), benchmarks, vendor repos.
  • Bucket B (no account, crawl): GitHub passes above. Anonymous git is throttled at times; the crawler falls back to codeload tarballs (results verified identical to git, incl. blob SHAs).
  • Out of scope by decision: Chinese-only registries (Tencent SkillHub, ModelScope). CJK-heavy skills elsewhere are kept and flagged with cjk_ratio.

License cautions

  • GitHub content keeps its origin repo's license (~55% of repos have none): filter on repo_license / the repo.
  • anthropics/skills docx/pdf/pptx/xlsx skills are source-available, not OSS; trailofbits/skills is CC-BY-SA-4.0.
  • ClawHub is MIT-0 throughout. Scanner-flagged (malicious/suspicious) skills are kept but flagged.

Pipeline (code/)

All long jobs run as systemd services via run_job.sh (memory caps, auto-restart, resumable); watchdog.sh relaunches dead jobs and writes STATUS.md; memguard.sh freezes low-priority jobs when RAM is low. mirror_hf.py (bucket A copies) · vendor_repos.py · clawhub_api.py + build_clawhub_api.py · build_clawhub.py · build_gitskills.py · github_topic_seeds.py + build_pass2_seeds.py · github_bundles.py (crawler) · build_github_delta.py + ship_delta.py (convert → upload → delete) · crawl_chain.sh · status.py.

Secrets

The HF token lives only in a local, git-ignored .env. Nothing in this bucket contains it.

Total size
63.7 GB
Files
3,288
Last updated
Oct 2
Pre-warmed CDN
US EU US EU

Contributors