lecturelens-api / README.md
aliSaac510's picture
chore: add Hugging Face Spaces configuration metadata to README
bed214b
|
Raw
History Blame Contribute Delete
9.99 kB
metadata
title: LectureLens API
emoji: πŸŽ₯
colorFrom: blue
colorTo: indigo
sdk: docker
pinned: false
license: mit

LectureLens API

LectureLens FastAPI Python Docker HuggingFace

Technical audio & video quality analysis for Zoom lecture recordings.
Part of the automated lecture QC pipeline alongside Speechmatics and n8n.


What It Does

LectureLens receives a single audio or video file per request and returns:

Output Description
KPI metrics Loudness (LUFS), SNR, sharpness, brightness, clipping, freeze detection…
Alerts Structured list with severity, kpi, message, suggested_fix, and optional timestamp_range
Composite score A single 0–1 quality score derived from weighted KPIs

It does not evaluate content organisation, transcript quality, or pedagogy β€” that's handled by a separate LLM step in n8n.


Architecture Position

Zoom recording.completed webhook
        β”‚
        β–Ό
  n8n Switch β†’ verify signature β†’ download video + per-speaker audio
        β”‚
        β”œβ”€β”€β–Ί LectureLens API  ◄── this service
        β”‚       one call per file (video once, audio per speaker)
        β”‚
        β”œβ”€β”€β–Ί Speechmatics β†’ transcript with timestamps
        β”‚
        β–Ό
  n8n Code node β†’ merge: LectureLens results + transcript
        β”‚
        β–Ό
  Unified lecture quality report

Tech Stack

Tool Purpose
FastAPI + Uvicorn API framework with auto /docs
FFmpeg (system binary) loudnorm, silencedetect, freezedetect, blackdetect, ffprobe
pyloudnorm ITU-R BS.1770-4 loudness (cross-check)
librosa SNR / noise floor from raw waveform
DNSMOS (Microsoft ONNX) Non-reference speech quality score (1–5)
OpenCV Frame sampling, brightness histogram, Laplacian sharpness
MUSIQ (IQA-PyTorch) Perceptual video quality score (0–1)
PyYAML thresholds.yaml β€” edit thresholds without code changes

All tools are free, open-source, CPU-only β€” no GPU required.


Quick Start (Local)

1. Prerequisites

  • Python 3.11+
  • FFmpeg installed and on PATH
    • macOS: brew install ffmpeg
    • Ubuntu: sudo apt install ffmpeg
    • Windows: download from ffmpeg.org and add to PATH

2. Clone & Install

git clone https://github.com/YOUR_USERNAME/lecturelens-api.git
cd lecturelens-api
python -m venv .venv
source .venv/bin/activate  # Windows: .venv\Scripts\activate

# CPU-only PyTorch first (saves ~1GB vs full install)
pip install torch==2.2.2+cpu torchvision==0.17.2+cpu \
    --index-url https://download.pytorch.org/whl/cpu

pip install -r requirements.txt

3. Download DNSMOS Model

python scripts/download_dnsmos.py

This downloads Microsoft's sig_bak_ovr.onnx (~8 MB) into dnsmos/.

4. Configure

cp .env.example .env
# Edit .env and set LECTURELENS_API_KEY to a strong secret

5. Run

uvicorn app.main:app --reload --port 7860

Open http://localhost:7860/docs for the interactive Swagger UI.


API Reference

Authentication

All /analyze requests require the header:

X-API-Key: <your-secret-key>

GET /health

No authentication required.

{ "status": "ok", "version": "1.0.0" }

POST /analyze

Request (multipart/form-data):

Field Type Required Description
file binary βœ… MP4/MOV (video) or M4A/WAV/MP3 (audio)
media_type "audio" | "video" βœ… Selects the analysis pipeline
participant_label string β€” Speaker name β€” echoed back in the response
language string β€” Default ar (reserved for future use)

Audio response example:

{
  "status": "completed",
  "media_type": "audio",
  "participant_label": "ali saad",
  "processing_time_seconds": 5.8,
  "metrics": {
    "integrated_loudness_lufs": -19.8,
    "true_peak_dbtp": -3.2,
    "clipped_samples_count": 0,
    "snr_db": 24.1,
    "loudness_range_lu": 6.4,
    "silence_segments": [{ "start": 12.4, "end": 14.1 }],
    "dnsmos_ovrl": 3.6
  },
  "alerts": [
    {
      "severity": "warning",
      "category": "audio",
      "kpi": "integrated_loudness_lufs",
      "message": "Integrated loudness is -19.8 LUFS β€” 3.8 LU below the target of -14 LUFS.",
      "suggested_fix": "Adjust microphone gain or apply loudness normalisation before uploading.",
      "timestamp_range": null
    }
  ],
  "overall_audio_score": 0.71,
  "overall_video_score": null
}

Video response example:

{
  "status": "completed",
  "media_type": "video",
  "participant_label": null,
  "processing_time_seconds": 22.3,
  "metrics": {
    "resolution": "1920x1080",
    "resolution_consistent": true,
    "frame_rate_fps": 30.0,
    "dropped_frames_ratio": 0.002,
    "avg_brightness": 132.4,
    "avg_sharpness_laplacian": 184.7,
    "frozen_segments": [],
    "black_segments": [],
    "compression_artifact_score": 0.12,
    "musiq_quality_score": 0.78
  },
  "alerts": [],
  "overall_audio_score": null,
  "overall_video_score": 0.83
}

Error Codes

HTTP error_code Cause
400 INVALID_MEDIA_TYPE Unsupported file extension
401 UNAUTHORIZED Missing or wrong X-API-Key
413 FILE_TOO_LARGE Exceeds MAX_FILE_SIZE_MB
422 CORRUPTED_FILE File could not be opened/decoded
500 PROCESSING_FAILED Unexpected internal error

cURL Examples

# Health check
curl https://<space>.hf.space/health

# Analyze audio
curl -X POST https://<space>.hf.space/analyze \
  -H "X-API-Key: your-secret-key" \
  -F "file=@speaker_ali.m4a" \
  -F "media_type=audio" \
  -F "participant_label=ali saad"

# Analyze video
curl -X POST https://<space>.hf.space/analyze \
  -H "X-API-Key: your-secret-key" \
  -F "file=@lecture.mp4" \
  -F "media_type=video"

KPI Reference

Audio

KPI Tool Target
integrated_loudness_lufs ffmpeg loudnorm βˆ’14 LUFS Β± 2 LU
true_peak_dbtp ffmpeg loudnorm < βˆ’1.0 dBTP
clipped_samples_count numpy 0
snr_db librosa > 20 dB
loudness_range_lu ffmpeg loudnorm 4–15 LU
silence_segments ffmpeg silencedetect No segment > 3s
dnsmos_ovrl DNSMOS ONNX > 3.0 / 5

Video

KPI Tool Target
resolution_consistent ffprobe true
dropped_frames_ratio ffprobe < 1%
avg_brightness OpenCV 80–180
avg_sharpness_laplacian OpenCV Laplacian > 100
frozen_segments ffmpeg freezedetect No segment > 2s
black_segments ffmpeg blackdetect None
compression_artifact_score Gradient heuristic < 0.3
musiq_quality_score MUSIQ (IQA-PyTorch) > 0.6

Thresholds

All alert thresholds live in thresholds.yaml.
Edit and redeploy β€” no code changes needed.


Deploy to Hugging Face Spaces

  1. Create a new Space β†’ Docker type
  2. Push this repo to the Space
  3. In Settings β†’ Variables and Secrets, add:
    • LECTURELENS_API_KEY β†’ your secret key
    • MAX_FILE_SIZE_MB β†’ e.g. 500
  4. The Space will build automatically (~5–10 min first build)
  5. Test with GET /health

Tip: HF Spaces (free tier) sleep after inactivity.
Always call GET /health first from n8n to wake the Space before sending large files.


n8n Integration

See the full n8n node configuration in the API documentation or refer to Section 6 of the original design specification.

Quick reference:

Existing node New node to add
download video Analyze Video Quality (POST /analyze, media_type=video)
download audeo Analyze Audio Quality (POST /analyze, media_type=audio)
Merge Speakers Aggregate Final Report (Code node)

Running Tests

pytest tests/ -v

Project Structure

lecturelens/
β”œβ”€β”€ app/
β”‚   β”œβ”€β”€ main.py              # FastAPI app + routes
β”‚   β”œβ”€β”€ config.py            # Settings (env vars)
β”‚   β”œβ”€β”€ auth.py              # X-API-Key middleware
β”‚   β”œβ”€β”€ schemas.py           # Pydantic models
β”‚   β”œβ”€β”€ alert_engine.py      # Alert generation + scoring
β”‚   β”œβ”€β”€ utils.py             # FFmpeg helpers + temp files
β”‚   └── analyzers/
β”‚       β”œβ”€β”€ audio_analyzer.py
β”‚       └── video_analyzer.py
β”œβ”€β”€ dnsmos/
β”‚   └── sig_bak_ovr.onnx    # DNSMOS model (download via script)
β”œβ”€β”€ scripts/
β”‚   └── download_dnsmos.py
β”œβ”€β”€ tests/
β”‚   β”œβ”€β”€ test_audio.py
β”‚   β”œβ”€β”€ test_video.py
β”‚   └── test_api.py
β”œβ”€β”€ thresholds.yaml
β”œβ”€β”€ Dockerfile
β”œβ”€β”€ requirements.txt
└── .env.example

Roadmap

  • v1.0 β€” Audio pipeline (Loudness, SNR, Clipping, Silence, DNSMOS)
  • v1.0 β€” Video pipeline (Brightness, Sharpness, Freeze, Black, MUSIQ)
  • v1.0 β€” Alert engine with thresholds.yaml
  • v1.0 β€” Hugging Face Docker deployment
  • Phase 2 β€” Lip-sync offset detection (SyncNet)
  • Phase 2 β€” Deeper compression artefact analysis
  • Phase 2 β€” Unified report merging LectureLens + LLM content evaluation

License

MIT β€” see LICENSE.