api / README.md
lljz66's picture
assets-only root README
96a28db verified
|
Raw
History Blame Contribute Delete
2.17 kB
metadata
title: Asset Extractor API
emoji: 🐍
colorFrom: blue
colorTo: green
sdk: docker
pinned: false
app_port: 7860

Asset Extractor API

FastAPI service that extracts assets from any URL, grouped by type: image · video · audio · svg/icon · font.

Usage

POST /assets

Give it a URL, get back the page's assets grouped by category.

curl -X POST https://lljz66-api.hf.space/assets \
  -H 'Content-Type: application/json' \
  -d '{"url": "https://example.com"}'
{
  "url": "https://example.com",
  "title": "Example Domain",
  "total": 3,
  "assets": {
    "image": ["https://example.com/logo.png"],
    "video": [],
    "audio": [],
    "svg_icon": ["https://example.com/favicon.ico"],
    "font": ["https://fonts.gstatic.com/..."]
  }
}

Request fields

field type default meaning
url string page to extract from (required)
include_data_uris bool false keep data: URIs
groups list[string] all subset: image, video, audio, svg_icon, font

Response fields

field type meaning
url string final URL after redirects
title string | null page title
total int total asset count
assets object grouped asset URL lists

Static HTML only — JS-rendered assets are not captured.

Structure (monolith-modular)

app/
├── main.py                 # FastAPI app assembly
├── config.py               # settings (fetch limits, SSRF block-list)
├── schemas.py              # pydantic request/response models
├── routers/
│   └── assets.py           # POST /assets (HTTP layer)
└── services/
    ├── fetcher.py          # safe async page fetch (SSRF/size/timeout guards)
    ├── resolver.py         # absolutize asset references
    ├── asset_service.py    # orchestrator: fetch -> parse -> extract -> group
    └── extractors/         # images, videos, audios, icons_svg, fonts

Local dev

pip install -r requirements.txt
uvicorn app.main:app --reload --port 7860