Buckets:
| """Pass-2 seed list: repos from GitHub topic search (created since July) + skills.sh / claudemarketplaces / | |
| Anthropic-marketplace sitemaps, minus repos already in GitSkills or the Sourcegraph pass. Forks dropped; | |
| ordered by stars.""" | |
| import json, duckdb | |
| c = duckdb.connect(); c.execute("SET memory_limit='200MB'; SET threads=1") | |
| c.execute("create table gs as select lower(full_name) n from 'work/gitskills/data/repos/*.parquet'") | |
| c.execute("create table sg as select lower(column0) n from read_csv('data/seeds/sourcegraph_repos.txt', header=false)") | |
| c.execute("""create table tp as select full_name, bool_or(fork) fork, max(stars) stars | |
| from read_json('data/seeds/github_topics.jsonl') where full_name is not null group by 1""") | |
| sm = json.load(open("data/seeds/sitemap_repos.json")) | |
| c.execute("create table sm as select unnest(?) full_name", [list(sm)]) | |
| rows = c.execute("""select full_name from ( | |
| select full_name, stars from tp where not fork | |
| union all select full_name, 0 from sm where lower(full_name) not in (select lower(full_name) from tp)) | |
| where lower(full_name) not in (select n from gs) and lower(full_name) not in (select n from sg) | |
| order by stars desc, full_name""").fetchall() | |
| open("data/seeds/pass2.txt", "w").write("\n".join(r[0] for r in rows) + "\n") | |
| print("pass2 seeds:", len(rows), flush=True) | |
Xet Storage Details
- Size:
- 1.35 kB
- Xet hash:
- 57914820d23762085e71318b1b768b9a047da9af19b86005199faffe2e687a76
·
Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.