Tenor / README.md
CJHauser's picture
Update README.md
460a071 verified
|
Raw
History Blame Contribute Delete
5.07 kB
---
title: Personal Tenor Scraper API
emoji: 🐈
colorFrom: yellow
colorTo: pink
sdk: docker
app_port: 7860
pinned: false
---
# Personal Tenor Scraper API
Unofficial personal replacement for the Tenor public API, which Google
shut down on **June 30, 2026**. Tenor.com itself is still online for
humans, so this scrapes the live site instead of calling the (now dead)
`api.tenor.com` endpoints.
**Heads up**: this depends on Tenor's page markup not changing. It uses
two independent extraction strategies (embedded JSON state blob, then a
raw-HTML regex fallback) so it should survive minor site tweaks, but a
big redesign on Tenor's end could break it. Check `scraper.py` if search
results suddenly come back empty.
## Deploying on Hugging Face Spaces
1. Create a new Space, SDK = **Docker**.
2. Push these files (`app.py`, `scraper.py`, `cache.py`, `requirements.txt`, `Dockerfile`, this `README.md`).
3. (Optional but recommended once it's public) Set a Space secret `API_KEY` to require auth β€” see below.
4. (Optional) Enable persistent storage on the Space so the SQLite cache/favorites survive restarts. Without it, everything resets when the Space sleeps/restarts.
5. Space builds and starts listening on port 7860 automatically.
## Auth
If you set the `API_KEY` environment variable/secret on the Space,
every endpoint requires it via either:
- Header: `X-API-Key: your-key-here`
- Query param: `?api_key=your-key-here`
Leave `API_KEY` unset for open access (fine if you're the only one who
knows the Space URL, less fine once people share it around).
## Endpoints
### `GET /search`
```
/search?q=cats&limit=10&format=full
```
Params:
- `q` (required) β€” search term
- `limit` β€” max results, default 20, capped at 50
- `page` β€” pagination, default 1
- `format` β€” `full` (default, full metadata), `urls` (flat list of direct URLs), or `discord` (list of `{url, embed_url, title}`)
### `GET /random`
```
/random?q=thumbs+up&limit=1&format=urls
```
Same params as `/search` minus `page`, returns shuffled picks.
### `GET /direct`
```
/direct?q=high+five
```
No JSON β€” just **302 redirects straight to a gif URL**. Great for
sticking directly into a Discord message as a link, or hitting from a
webhook/bot that just wants a URL back with zero parsing:
```
Just paste https://your-space.hf.space/direct?q=high+five directly in Discord chat and it unfurls as the gif.
```
### `POST /favorites`
Save a gif you like for quick reuse later (e.g. your most-used reaction gifs).
```json
{
"gif_url": "https://media.tenor.com/xyz.gif",
"mp4_url": "https://media.tenor.com/xyz.mp4",
"title": "thumbs up guy",
"tag": "reactions"
}
```
### `GET /favorites?tag=reactions`
List saved favorites, optionally filtered by tag.
### `DELETE /favorites/<id>`
Remove a saved favorite by its numeric id.
### `GET /favorites/random?tag=reactions`
Random pick from your saved favorites β€” good for a Discord bot's
"random reaction gif" command without re-scraping every time.
### `GET /health`
Basic liveness check, returns `{"status": "ok", "time": ...}`.
## Using it from a Discord bot later
Since you said the bot wiring comes later β€” the API is shaped so a
future bot command can just do:
```python
import requests
resp = requests.get(
"https://your-space.hf.space/random",
params={"q": query, "format": "discord", "limit": 1},
)
gif = resp.json()["results"][0]
await ctx.send(gif["embed_url"]) # Discord auto-embeds the URL
```
## Rate limiting
In-memory sliding window, default 60 requests / 60 seconds per API
key (or per IP if no key is set). Adjust via `RATE_LIMIT_MAX` and
`RATE_LIMIT_WINDOW` env vars. This is per-instance, not distributed β€”
fine for a single free-tier Space.
## Caching
Search results are cached in SQLite for `CACHE_TTL_SECONDS` (default
3600 = 1 hour) so repeated searches for the same term don't re-scrape
Tenor every time. Cache path defaults to `/data/tenor_cache.db` if
persistent storage is enabled on the Space, otherwise falls back to a
local file that resets on restart.
## Env vars summary
| Var | Default | Purpose |
|---|---|---|
| `API_KEY` | unset (auth disabled) | require `X-API-Key` header/`api_key` param |
| `RATE_LIMIT_MAX` | 60 | max requests per window |
| `RATE_LIMIT_WINDOW` | 60 | window length in seconds |
| `CACHE_TTL_SECONDS` | 3600 | how long cached search results stay fresh |
| `TENOR_CACHE_PATH` | `/data/tenor_cache.db` | SQLite path when persistent storage is enabled |
| `PORT` | 7860 | server port (HF Spaces expects 7860) |
## Known limitations
- Scraping HTML is inherently more fragile than a real API β€” expect occasional breakage if Tenor changes their frontend.
- No official support/guarantee from Tenor; this is purely a personal workaround now that the API is gone.
- Rate limiting is per-instance/in-memory, resets on restart.
- If Tenor ever moves to fully client-rendered results with nothing in the initial HTML, you'd need to add a headless-browser fallback (see the `render_with_browser()` stub in `scraper.py` for notes on wiring up Playwright).