--- title: LectureLens API emoji: πŸŽ₯ colorFrom: blue colorTo: indigo sdk: docker pinned: false license: mit --- # LectureLens API
![LectureLens](https://img.shields.io/badge/LectureLens-v1.0.0-6366f1?style=for-the-badge) ![FastAPI](https://img.shields.io/badge/FastAPI-0.111-009688?style=for-the-badge&logo=fastapi) ![Python](https://img.shields.io/badge/Python-3.11-3776AB?style=for-the-badge&logo=python) ![Docker](https://img.shields.io/badge/Docker-ready-2496ED?style=for-the-badge&logo=docker) ![HuggingFace](https://img.shields.io/badge/HuggingFace-Spaces-FFD21E?style=for-the-badge) **Technical audio & video quality analysis for Zoom lecture recordings.** Part of the automated lecture QC pipeline alongside Speechmatics and n8n.
--- ## What It Does LectureLens receives a single audio or video file per request and returns: | Output | Description | |--------|-------------| | **KPI metrics** | Loudness (LUFS), SNR, sharpness, brightness, clipping, freeze detection… | | **Alerts** | Structured list with `severity`, `kpi`, `message`, `suggested_fix`, and optional `timestamp_range` | | **Composite score** | A single 0–1 quality score derived from weighted KPIs | It does **not** evaluate content organisation, transcript quality, or pedagogy β€” that's handled by a separate LLM step in n8n. --- ## Architecture Position ``` Zoom recording.completed webhook β”‚ β–Ό n8n Switch β†’ verify signature β†’ download video + per-speaker audio β”‚ β”œβ”€β”€β–Ί LectureLens API ◄── this service β”‚ one call per file (video once, audio per speaker) β”‚ β”œβ”€β”€β–Ί Speechmatics β†’ transcript with timestamps β”‚ β–Ό n8n Code node β†’ merge: LectureLens results + transcript β”‚ β–Ό Unified lecture quality report ``` --- ## Tech Stack | Tool | Purpose | |------|---------| | **FastAPI** + Uvicorn | API framework with auto `/docs` | | **FFmpeg** (system binary) | `loudnorm`, `silencedetect`, `freezedetect`, `blackdetect`, `ffprobe` | | **pyloudnorm** | ITU-R BS.1770-4 loudness (cross-check) | | **librosa** | SNR / noise floor from raw waveform | | **DNSMOS** (Microsoft ONNX) | Non-reference speech quality score (1–5) | | **OpenCV** | Frame sampling, brightness histogram, Laplacian sharpness | | **MUSIQ** (IQA-PyTorch) | Perceptual video quality score (0–1) | | **PyYAML** | `thresholds.yaml` β€” edit thresholds without code changes | All tools are **free, open-source, CPU-only** β€” no GPU required. --- ## Quick Start (Local) ### 1. Prerequisites - Python 3.11+ - FFmpeg installed and on `PATH` - macOS: `brew install ffmpeg` - Ubuntu: `sudo apt install ffmpeg` - Windows: download from [ffmpeg.org](https://ffmpeg.org/download.html) and add to PATH ### 2. Clone & Install ```bash git clone https://github.com/YOUR_USERNAME/lecturelens-api.git cd lecturelens-api python -m venv .venv source .venv/bin/activate # Windows: .venv\Scripts\activate # CPU-only PyTorch first (saves ~1GB vs full install) pip install torch==2.2.2+cpu torchvision==0.17.2+cpu \ --index-url https://download.pytorch.org/whl/cpu pip install -r requirements.txt ``` ### 3. Download DNSMOS Model ```bash python scripts/download_dnsmos.py ``` This downloads Microsoft's `sig_bak_ovr.onnx` (~8 MB) into `dnsmos/`. ### 4. Configure ```bash cp .env.example .env # Edit .env and set LECTURELENS_API_KEY to a strong secret ``` ### 5. Run ```bash uvicorn app.main:app --reload --port 7860 ``` Open **http://localhost:7860/docs** for the interactive Swagger UI. --- ## API Reference ### Authentication All `/analyze` requests require the header: ``` X-API-Key: ``` ### `GET /health` No authentication required. ```json { "status": "ok", "version": "1.0.0" } ``` ### `POST /analyze` **Request** (`multipart/form-data`): | Field | Type | Required | Description | |-------|------|----------|-------------| | `file` | binary | βœ… | MP4/MOV (video) or M4A/WAV/MP3 (audio) | | `media_type` | `"audio"` \| `"video"` | βœ… | Selects the analysis pipeline | | `participant_label` | string | β€” | Speaker name β€” echoed back in the response | | `language` | string | β€” | Default `ar` (reserved for future use) | **Audio response example:** ```json { "status": "completed", "media_type": "audio", "participant_label": "ali saad", "processing_time_seconds": 5.8, "metrics": { "integrated_loudness_lufs": -19.8, "true_peak_dbtp": -3.2, "clipped_samples_count": 0, "snr_db": 24.1, "loudness_range_lu": 6.4, "silence_segments": [{ "start": 12.4, "end": 14.1 }], "dnsmos_ovrl": 3.6 }, "alerts": [ { "severity": "warning", "category": "audio", "kpi": "integrated_loudness_lufs", "message": "Integrated loudness is -19.8 LUFS β€” 3.8 LU below the target of -14 LUFS.", "suggested_fix": "Adjust microphone gain or apply loudness normalisation before uploading.", "timestamp_range": null } ], "overall_audio_score": 0.71, "overall_video_score": null } ``` **Video response example:** ```json { "status": "completed", "media_type": "video", "participant_label": null, "processing_time_seconds": 22.3, "metrics": { "resolution": "1920x1080", "resolution_consistent": true, "frame_rate_fps": 30.0, "dropped_frames_ratio": 0.002, "avg_brightness": 132.4, "avg_sharpness_laplacian": 184.7, "frozen_segments": [], "black_segments": [], "compression_artifact_score": 0.12, "musiq_quality_score": 0.78 }, "alerts": [], "overall_audio_score": null, "overall_video_score": 0.83 } ``` ### Error Codes | HTTP | `error_code` | Cause | |------|-------------|-------| | 400 | `INVALID_MEDIA_TYPE` | Unsupported file extension | | 401 | `UNAUTHORIZED` | Missing or wrong `X-API-Key` | | 413 | `FILE_TOO_LARGE` | Exceeds `MAX_FILE_SIZE_MB` | | 422 | `CORRUPTED_FILE` | File could not be opened/decoded | | 500 | `PROCESSING_FAILED` | Unexpected internal error | --- ## cURL Examples ```bash # Health check curl https://.hf.space/health # Analyze audio curl -X POST https://.hf.space/analyze \ -H "X-API-Key: your-secret-key" \ -F "file=@speaker_ali.m4a" \ -F "media_type=audio" \ -F "participant_label=ali saad" # Analyze video curl -X POST https://.hf.space/analyze \ -H "X-API-Key: your-secret-key" \ -F "file=@lecture.mp4" \ -F "media_type=video" ``` --- ## KPI Reference ### Audio | KPI | Tool | Target | |-----|------|--------| | `integrated_loudness_lufs` | ffmpeg loudnorm | βˆ’14 LUFS Β± 2 LU | | `true_peak_dbtp` | ffmpeg loudnorm | < βˆ’1.0 dBTP | | `clipped_samples_count` | numpy | 0 | | `snr_db` | librosa | > 20 dB | | `loudness_range_lu` | ffmpeg loudnorm | 4–15 LU | | `silence_segments` | ffmpeg silencedetect | No segment > 3s | | `dnsmos_ovrl` | DNSMOS ONNX | > 3.0 / 5 | ### Video | KPI | Tool | Target | |-----|------|--------| | `resolution_consistent` | ffprobe | `true` | | `dropped_frames_ratio` | ffprobe | < 1% | | `avg_brightness` | OpenCV | 80–180 | | `avg_sharpness_laplacian` | OpenCV Laplacian | > 100 | | `frozen_segments` | ffmpeg freezedetect | No segment > 2s | | `black_segments` | ffmpeg blackdetect | None | | `compression_artifact_score` | Gradient heuristic | < 0.3 | | `musiq_quality_score` | MUSIQ (IQA-PyTorch) | > 0.6 | --- ## Thresholds All alert thresholds live in [`thresholds.yaml`](thresholds.yaml). Edit and redeploy β€” no code changes needed. --- ## Deploy to Hugging Face Spaces 1. Create a new Space β†’ **Docker** type 2. Push this repo to the Space 3. In **Settings β†’ Variables and Secrets**, add: - `LECTURELENS_API_KEY` β†’ your secret key - `MAX_FILE_SIZE_MB` β†’ e.g. `500` 4. The Space will build automatically (~5–10 min first build) 5. Test with `GET /health` > **Tip:** HF Spaces (free tier) sleep after inactivity. > Always call `GET /health` first from n8n to wake the Space before sending large files. --- ## n8n Integration See the full n8n node configuration in the [API documentation](docs/n8n_integration.md) or refer to Section 6 of the original design specification. **Quick reference:** | Existing node | New node to add | |---------------|----------------| | `download video` | `Analyze Video Quality` (POST /analyze, media_type=video) | | `download audeo` | `Analyze Audio Quality` (POST /analyze, media_type=audio) | | `Merge Speakers` | `Aggregate Final Report` (Code node) | --- ## Running Tests ```bash pytest tests/ -v ``` --- ## Project Structure ``` lecturelens/ β”œβ”€β”€ app/ β”‚ β”œβ”€β”€ main.py # FastAPI app + routes β”‚ β”œβ”€β”€ config.py # Settings (env vars) β”‚ β”œβ”€β”€ auth.py # X-API-Key middleware β”‚ β”œβ”€β”€ schemas.py # Pydantic models β”‚ β”œβ”€β”€ alert_engine.py # Alert generation + scoring β”‚ β”œβ”€β”€ utils.py # FFmpeg helpers + temp files β”‚ └── analyzers/ β”‚ β”œβ”€β”€ audio_analyzer.py β”‚ └── video_analyzer.py β”œβ”€β”€ dnsmos/ β”‚ └── sig_bak_ovr.onnx # DNSMOS model (download via script) β”œβ”€β”€ scripts/ β”‚ └── download_dnsmos.py β”œβ”€β”€ tests/ β”‚ β”œβ”€β”€ test_audio.py β”‚ β”œβ”€β”€ test_video.py β”‚ └── test_api.py β”œβ”€β”€ thresholds.yaml β”œβ”€β”€ Dockerfile β”œβ”€β”€ requirements.txt └── .env.example ``` --- ## Roadmap - [x] v1.0 β€” Audio pipeline (Loudness, SNR, Clipping, Silence, DNSMOS) - [x] v1.0 β€” Video pipeline (Brightness, Sharpness, Freeze, Black, MUSIQ) - [x] v1.0 β€” Alert engine with `thresholds.yaml` - [x] v1.0 β€” Hugging Face Docker deployment - [ ] Phase 2 β€” Lip-sync offset detection (SyncNet) - [ ] Phase 2 β€” Deeper compression artefact analysis - [ ] Phase 2 β€” Unified report merging LectureLens + LLM content evaluation --- ## License MIT β€” see [LICENSE](LICENSE).