--- title: Personal Tenor Scraper API emoji: 🐈 colorFrom: yellow colorTo: pink sdk: docker app_port: 7860 pinned: false --- # Personal Tenor Scraper API Unofficial personal replacement for the Tenor public API, which Google shut down on **June 30, 2026**. Tenor.com itself is still online for humans, so this scrapes the live site instead of calling the (now dead) `api.tenor.com` endpoints. **Heads up**: this depends on Tenor's page markup not changing. It uses two independent extraction strategies (embedded JSON state blob, then a raw-HTML regex fallback) so it should survive minor site tweaks, but a big redesign on Tenor's end could break it. Check `scraper.py` if search results suddenly come back empty. ## Deploying on Hugging Face Spaces 1. Create a new Space, SDK = **Docker**. 2. Push these files (`app.py`, `scraper.py`, `cache.py`, `requirements.txt`, `Dockerfile`, this `README.md`). 3. (Optional but recommended once it's public) Set a Space secret `API_KEY` to require auth — see below. 4. (Optional) Enable persistent storage on the Space so the SQLite cache/favorites survive restarts. Without it, everything resets when the Space sleeps/restarts. 5. Space builds and starts listening on port 7860 automatically. ## Auth If you set the `API_KEY` environment variable/secret on the Space, every endpoint requires it via either: - Header: `X-API-Key: your-key-here` - Query param: `?api_key=your-key-here` Leave `API_KEY` unset for open access (fine if you're the only one who knows the Space URL, less fine once people share it around). ## Endpoints ### `GET /search` ``` /search?q=cats&limit=10&format=full ``` Params: - `q` (required) — search term - `limit` — max results, default 20, capped at 50 - `page` — pagination, default 1 - `format` — `full` (default, full metadata), `urls` (flat list of direct URLs), or `discord` (list of `{url, embed_url, title}`) ### `GET /random` ``` /random?q=thumbs+up&limit=1&format=urls ``` Same params as `/search` minus `page`, returns shuffled picks. ### `GET /direct` ``` /direct?q=high+five ``` No JSON — just **302 redirects straight to a gif URL**. Great for sticking directly into a Discord message as a link, or hitting from a webhook/bot that just wants a URL back with zero parsing: ``` Just paste https://your-space.hf.space/direct?q=high+five directly in Discord chat and it unfurls as the gif. ``` ### `POST /favorites` Save a gif you like for quick reuse later (e.g. your most-used reaction gifs). ```json { "gif_url": "https://media.tenor.com/xyz.gif", "mp4_url": "https://media.tenor.com/xyz.mp4", "title": "thumbs up guy", "tag": "reactions" } ``` ### `GET /favorites?tag=reactions` List saved favorites, optionally filtered by tag. ### `DELETE /favorites/` Remove a saved favorite by its numeric id. ### `GET /favorites/random?tag=reactions` Random pick from your saved favorites — good for a Discord bot's "random reaction gif" command without re-scraping every time. ### `GET /health` Basic liveness check, returns `{"status": "ok", "time": ...}`. ## Using it from a Discord bot later Since you said the bot wiring comes later — the API is shaped so a future bot command can just do: ```python import requests resp = requests.get( "https://your-space.hf.space/random", params={"q": query, "format": "discord", "limit": 1}, ) gif = resp.json()["results"][0] await ctx.send(gif["embed_url"]) # Discord auto-embeds the URL ``` ## Rate limiting In-memory sliding window, default 60 requests / 60 seconds per API key (or per IP if no key is set). Adjust via `RATE_LIMIT_MAX` and `RATE_LIMIT_WINDOW` env vars. This is per-instance, not distributed — fine for a single free-tier Space. ## Caching Search results are cached in SQLite for `CACHE_TTL_SECONDS` (default 3600 = 1 hour) so repeated searches for the same term don't re-scrape Tenor every time. Cache path defaults to `/data/tenor_cache.db` if persistent storage is enabled on the Space, otherwise falls back to a local file that resets on restart. ## Env vars summary | Var | Default | Purpose | |---|---|---| | `API_KEY` | unset (auth disabled) | require `X-API-Key` header/`api_key` param | | `RATE_LIMIT_MAX` | 60 | max requests per window | | `RATE_LIMIT_WINDOW` | 60 | window length in seconds | | `CACHE_TTL_SECONDS` | 3600 | how long cached search results stay fresh | | `TENOR_CACHE_PATH` | `/data/tenor_cache.db` | SQLite path when persistent storage is enabled | | `PORT` | 7860 | server port (HF Spaces expects 7860) | ## Known limitations - Scraping HTML is inherently more fragile than a real API — expect occasional breakage if Tenor changes their frontend. - No official support/guarantee from Tenor; this is purely a personal workaround now that the API is gone. - Rate limiting is per-instance/in-memory, resets on restart. - If Tenor ever moves to fully client-rendered results with nothing in the initial HTML, you'd need to add a headless-browser fallback (see the `render_with_browser()` stub in `scraper.py` for notes on wiring up Playwright).