content_pipeline / README.md
AI Engineer
Deploy Streamlit Content Agent
0b29030
|
Raw
History Blame Contribute Delete
5.47 kB

A newer version of the Streamlit SDK is available: 1.62.0

Upgrade
metadata
title: Content Generation Agent
emoji: ✍️
colorFrom: indigo
colorTo: purple
sdk: streamlit
sdk_version: 1.38.0
app_file: app.py
pinned: false

✍️ Content Generation Agent (MVP)

A production-shaped AI marketing agent that turns a founder's raw product update note into polished, grounded, and guardrailed marketing copy for LinkedIn, Instagram, and WhatsApp, along with a downloadable branded PowerPoint deck.

Built as a Streamlit app. Run it fully offline in Mock Mode for a zero-cost interactive demo, or connect a Groq / OpenRouter API key to crawl and generate copy for your live product.


🚀 Live Demo Quick Start

1. Local Setup

Ensure you have Python 3.10+ installed. Run the following commands in your terminal:

# Create a virtual environment
python -m venv venv

# Activate the virtual environment
# On Windows (PowerShell):
.\venv\Scripts\activate
# On macOS/Linux:
source venv/bin/activate

# Install dependencies
pip install -r requirements.txt

2. Run Tests

Verify the pipeline works with the offline mock runner:

python -m tests.test_pipeline

3. Launch UI

Launch the Streamlit web application:

streamlit run app.py

This will automatically open the application at http://localhost:8501 in your browser.


🛠️ Offline Mock Mode vs. Live Mode

The application supports two execution modes, switchable dynamically from the sidebar settings:

🟢 Mock Mode (Default)

  • API Key Required: None ($0 cost).
  • Behavior: Uses deterministic mock responses to simulate the pipeline.
  • Use Case: Perfect for testing the UI, inspecting formatting, and running demo sessions offline.

🔵 Live Mode

  • API Key Required: A valid Groq or OpenRouter key (pasted into the sidebar or set via .env).
  • Behavior: Active crawler scrapes website content, and the selected LLM (e.g., Llama 3.3, Claude 3.5 Sonnet, Gemini Flash) generates live copy.

📐 Pipeline Architecture

The agent is designed as a structured pipeline to prevent hallucinations, enforce constraints, and rank generated copy:

graph TD
    A[Product Website] -->|Crawler| B(Crawled Pages)
    B -->|Brand Brain LLM| C[🧠 Brand Brain]
    C -->|Prompt Suggestions| D[Founder Input]
    D -->|Brief Builder LLM| E[📋 Structured Brief]
    C & E -->|Fact Pack Builder LLM| F[📦 Fact Pack]
    F -->|LinkedIn Writer LLM| G1[LinkedIn Copy]
    F -->|Instagram Writer LLM| G2[Instagram Copy]
    F -->|WhatsApp Writer LLM| G3[WhatsApp Copy]
    G1 & G2 & G3 -->|Guardrail Layer| H[Claim Checking + Lints + Policy]
    H -->|Ranker Heuristics| I[🏆 Scored & Badged Output]
    F & E -->|python-pptx| J[📊 Branded PPT Deck]

🔒 Grounded Guardrails & Robustness

To ensure enterprise-level reliability, the agent implements several production safety features:

  1. Untrusted Data Handling: Scraped website content is treated strictly as raw context/data. The system prompts instruct LLM writers to ignore any instructions embedded in crawled HTML (preventing prompt injection).
  2. Fact Pack Grounding: Writers never see raw crawled HTML; they only write copy using the curated Fact Pack approved facts to prevent hallucinations.
  3. Deterministic Guardrails:
    • Claim Checker: Flags numerical values, compliance assertions, and superlatives (like "guaranteed") that cannot be verified in the Fact Pack.
    • Lints: Enforces strict word count, character, and hashtag limits per channel.
    • Policy Filter: Flags offensive, inappropriate, or out-of-bounds language.
  4. Resilient Website Crawler:
    • Sitemap-First BFS: Prioritizes crawling specific entered seed URLs before falling back to sitemap links.
    • Browser Emulation: Bypasses bot protections using standard Chrome headers.
    • SSL Verification Fallback: Retries unverified connections, allowing development and local test sites to be crawled.
  5. Unicode Surrogate Repair: Automatically sanitizes emoji surrogate code points from API models to prevent web client encoding crashes.

📂 Codebase Layout

content_agent/
├── app.py                 # Streamlit UI & Interactive Interface
├── core/
│   ├── config.py          # Environment, Settings & Model Registry
│   ├── crawler.py         # Resilient Web Scraper (requests + bs4)
│   ├── brand_brain.py     # Brand Brain builder from scraped corpus
│   ├── brief.py           # Structurer for the founder brief
│   ├── fact_pack.py       # Grounded Fact Pack builder
│   ├── writers.py         # Multi-channel copywriting templates
│   ├── guardrails.py      # Claim, policy, and formatting validation
│   ├── ranker.py          # Score evaluation for variants
│   ├── ppt.py             # PowerPoint rendering engine (python-pptx)
│   ├── llm.py             # OpenAI-compatible API client & parser
│   ├── mock_backend.py    # Offline simulation database
│   └── schemas.py         # Strict dataclasses for pipeline state
├── tests/
│   └── test_pipeline.py   # Complete integration & test suite
└── Dockerfile             # Deployment config

🐳 Docker Deployment

To deploy in any cloud environment:

docker compose up --build

The app will bind to port 8501.