opticparse-python / launch-posts /opticparse-devto-article.md
Nanny7's picture
initial deploy
bcf46c3
|
Raw
History Blame Contribute Delete
4.11 kB
metadata
title: >-
  I Built a Free API That Scrapes Any Website Using Plain English β€” No CSS
  Selectors
published: true
description: >-
  How I built Opticparse: a vision-based web scraper that uses AI to extract
  data from any webpage using natural language, powered by Playwright + Vision
  AI (Groq LLaMA Vision + Gemini).
tags: webdev, python, ai, opensource
cover_image: null

I've wasted days of my life maintaining CSS selectors.

You know the drill β€” you write the perfect scraper, it works great for a week, then the site does a frontend redesign, your selectors break, and you spend another afternoon hunting through the DOM again.

So I built Opticparse β€” a completely different approach.

How It Works

Instead of selectors, Opticparse:

  1. Opens a real Chromium browser (via Playwright)
  2. Navigates to your URL and waits for JavaScript to load
  3. Screenshots the page
  4. Sends the screenshot to a vision AI model
  5. Returns structured JSON based on your natural language query
curl -X POST https://opticparse.onrender.com/api/vision-scrape \
  -H "X-API-Key: YOUR_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "target_url": "https://news.ycombinator.com",
    "extraction_query": "Extract all story titles and upvote counts as a JSON array"
  }'
[
  {"title": "Show HN: I built a vision web scraper", "upvotes": 342},
  {"title": "The future of AI agents", "upvotes": 187}
]

No selectors. No XPath. No DOM inspection. The AI figures out where everything is from the screenshot.

The Technical Stack

Request
  β”‚
  β–Ό
FastAPI (Python)
  β”‚
  β–Ό
Playwright Chromium (headless, stealth mode)
  β”œβ”€β”€ navigator.webdriver = undefined (anti-bot)
  β”œβ”€β”€ Route interception (blocks media/fonts β†’ 60% less bandwidth)
  └── Screenshots full page as PNG
  β”‚
  β–Ό
AI Provider Rotation (tries in order until success):
  1. Groq β€” llama-3.2-11b-vision (< 1s, free)
  2. GitHub Models β€” Vision AI (Groq LLaMA Vision + Gemini) (free 150 req/day)
  3. OpenRouter β€” Vision AI (Groq LLaMA Vision + Gemini) (free credits)
  β”‚
  β–Ό
5-minute response cache (MD5 hash of url + query)
  β”‚
  β–Ό
Cleaned JSON response

The Stealth Mode

Cloudflare and other WAFs detect headless browsers by checking navigator.webdriver. I added a simple init script to neutralize this:

await context.add_init_script(
    "Object.defineProperty(navigator, 'webdriver', {get: () => undefined});"
)

Combined with a real Chrome user agent and Accept-Language headers, this bypasses most basic bot detection.

The AI Provider Rotation

The smartest part: if one AI provider rate-limits, the next one kicks in automatically. Zero downtime, effectively unlimited free capacity.

AI_PROVIDERS = [
    {"name": "Groq", "model": "llama-3.2-11b-vision-preview"},  # Fastest
    {"name": "GitHub Models", "model": "Vision AI (Groq LLaMA Vision + Gemini)"},               # Fallback
    {"name": "OpenRouter", "model": "openai/Vision AI (Groq LLaMA Vision + Gemini)"},           # Fallback
]

for provider in AI_PROVIDERS:
    try:
        result = await call_provider(provider, screenshot, query)
        return result  # Success β€” stop here
    except RateLimitError:
        continue  # Try next provider

The 5-Minute Cache

Same URL + same query = instant response from cache. Especially useful for dashboard-style apps that poll the same pages repeatedly.

def get_cache(url, query):
    key = hashlib.md5(f"{url}|{query}".encode()).hexdigest()
    entry = _cache.get(key)
    if entry and time.time() - entry["ts"] < 300:
        return entry["data"]
    return None

Try It Free

Available on opticparse.com with a free tier β€” no credit card: Opticparse on opticparse.com

Full source code: GitHub (MIT license)


What websites have you tried to scrape that kept breaking? Let me know in the comments β€” I'll test it against Opticparse.