Spaces:
Sleeping
Sleeping
Heebin Moon Claude Opus 4.6 commited on
Commit ·
a2749f3
0
Parent(s):
Initial commit: YouTube transcript formatter
Browse filesYouTube 영상의 자막을 추출하고 포맷팅하는 도구
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- .env.example +17 -0
- .gitignore +27 -0
- README.md +116 -0
- cli.py +111 -0
- formatter.py +687 -0
- requirements.txt +15 -0
- run +68 -0
- templates/index.html +1448 -0
- transcriber.py +565 -0
- web.py +237 -0
.env.example
ADDED
|
@@ -0,0 +1,17 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# OpenAI API 키 (https://platform.openai.com/api-keys 에서 발급)
|
| 2 |
+
OPENAI_API_KEY=sk-proj-여기에_키를_입력하세요
|
| 3 |
+
|
| 4 |
+
# 기본 변환 모드 (api 또는 local)
|
| 5 |
+
DEFAULT_MODE=api
|
| 6 |
+
|
| 7 |
+
# 로컬 Whisper 모델 크기 (tiny, base, small, medium, large)
|
| 8 |
+
DEFAULT_MODEL=base
|
| 9 |
+
|
| 10 |
+
# MD 포맷팅용 API 키 (사용하는 LLM만 설정하면 됩니다)
|
| 11 |
+
# Google Gemini (https://aistudio.google.com/app/apikey)
|
| 12 |
+
GOOGLE_API_KEY=여기에_키를_입력하세요
|
| 13 |
+
# Anthropic Claude (https://console.anthropic.com/)
|
| 14 |
+
ANTHROPIC_API_KEY=여기에_키를_입력하세요
|
| 15 |
+
|
| 16 |
+
# 웹 서버 포트
|
| 17 |
+
PORT=8000
|
.gitignore
ADDED
|
@@ -0,0 +1,27 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# 환경 변수 (API 키 포함)
|
| 2 |
+
.env
|
| 3 |
+
|
| 4 |
+
# 출력물
|
| 5 |
+
output/
|
| 6 |
+
|
| 7 |
+
# Python
|
| 8 |
+
__pycache__/
|
| 9 |
+
*.pyc
|
| 10 |
+
*.pyo
|
| 11 |
+
*.egg-info/
|
| 12 |
+
dist/
|
| 13 |
+
build/
|
| 14 |
+
|
| 15 |
+
# Claude Code 로컬 설정
|
| 16 |
+
.claude/
|
| 17 |
+
|
| 18 |
+
# OS
|
| 19 |
+
.DS_Store
|
| 20 |
+
Thumbs.db
|
| 21 |
+
|
| 22 |
+
# Whisper 모델 캐시
|
| 23 |
+
*.pt
|
| 24 |
+
|
| 25 |
+
# 임시 파일
|
| 26 |
+
*.tmp
|
| 27 |
+
temp/
|
README.md
ADDED
|
@@ -0,0 +1,116 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# 🎬 YouTube Script Extractor
|
| 2 |
+
|
| 3 |
+
YouTube 영상에서 스크립트를 추출하고, AI로 깔끔하게 정리하는 도구입니다.
|
| 4 |
+
|
| 5 |
+
## ✨ 주요 기능
|
| 6 |
+
|
| 7 |
+
- **자동 스크립트 추출** — YouTube URL만 입력하면 음성을 텍스트로 변환
|
| 8 |
+
- **AI 마크다운 정리** — LLM이 스크립트를 구조화된 마크다운으로 정리
|
| 9 |
+
- **다국어 번역** — 14개 언어 지원 (한국어, 영어, 일본어 등)
|
| 10 |
+
- **키프레임 분석 (CV)** — 영상의 핵심 장면을 추출하고 Vision AI로 분석
|
| 11 |
+
- **Smart Mix** — 파이프라인 단계별로 다른 LLM을 지정 가능
|
| 12 |
+
- **웹 UI + CLI** — 브라우저 기반 UI와 커맨드라인 모두 지원
|
| 13 |
+
|
| 14 |
+
## 🤖 지원 LLM
|
| 15 |
+
|
| 16 |
+
| 모델 | 가격 | Vision | 비고 |
|
| 17 |
+
|------|------|--------|------|
|
| 18 |
+
| Ollama (로컬) | 무료 | ❌ | 인터넷 불필요 |
|
| 19 |
+
| Gemini 2.5 Flash-Lite | $0.10/1M토큰 | ✅ | ⭐ 최저가 |
|
| 20 |
+
| Gemini 3.1 Flash-Lite | $0.25/1M토큰 | ✅ | 최신 모델 |
|
| 21 |
+
| GPT-4o-mini | $0.15/1M토큰 | ✅ | OpenAI |
|
| 22 |
+
| Claude Sonnet 4.6 | $3/1M토큰 | ✅ | 고품질 |
|
| 23 |
+
| Claude Opus 4.6 | $5/1M토큰 | ✅ | 최고 품질 |
|
| 24 |
+
|
| 25 |
+
## 🚀 빠른 시작
|
| 26 |
+
|
| 27 |
+
### 1. 설치
|
| 28 |
+
|
| 29 |
+
```bash
|
| 30 |
+
git clone https://huggingface.co/spaces/YOUR_USERNAME/youtube-script-extractor
|
| 31 |
+
cd youtube-script-extractor
|
| 32 |
+
|
| 33 |
+
# ffmpeg 설치 (macOS)
|
| 34 |
+
brew install ffmpeg
|
| 35 |
+
|
| 36 |
+
# Python 의존성 설치
|
| 37 |
+
pip install -r requirements.txt
|
| 38 |
+
```
|
| 39 |
+
|
| 40 |
+
### 2. API 키 설정
|
| 41 |
+
|
| 42 |
+
```bash
|
| 43 |
+
cp .env.example .env
|
| 44 |
+
# .env 파일을 열어 사용할 LLM의 API 키를 입력하세요
|
| 45 |
+
```
|
| 46 |
+
|
| 47 |
+
### 3. 실행
|
| 48 |
+
|
| 49 |
+
```bash
|
| 50 |
+
# 웹 UI (추천)
|
| 51 |
+
chmod +x run
|
| 52 |
+
./run
|
| 53 |
+
|
| 54 |
+
# 또는 직접 실행
|
| 55 |
+
python3 web.py
|
| 56 |
+
```
|
| 57 |
+
|
| 58 |
+
브라우저에서 `http://localhost:8000` 접속
|
| 59 |
+
|
| 60 |
+
### CLI 모드
|
| 61 |
+
|
| 62 |
+
```bash
|
| 63 |
+
./run "https://www.youtube.com/watch?v=VIDEO_ID"
|
| 64 |
+
./run "https://www.youtube.com/watch?v=VIDEO_ID" --mode local --format srt
|
| 65 |
+
```
|
| 66 |
+
|
| 67 |
+
## ⚙️ 설정
|
| 68 |
+
|
| 69 |
+
`.env` 파일에서 설정합니다:
|
| 70 |
+
|
| 71 |
+
```env
|
| 72 |
+
# 필수: Whisper API용
|
| 73 |
+
OPENAI_API_KEY=sk-proj-...
|
| 74 |
+
|
| 75 |
+
# 선택: 사용하는 LLM만 설정
|
| 76 |
+
GOOGLE_API_KEY=... # Gemini
|
| 77 |
+
ANTHROPIC_API_KEY=... # Claude
|
| 78 |
+
|
| 79 |
+
# 기본 설정
|
| 80 |
+
DEFAULT_MODE=api # api 또는 local
|
| 81 |
+
DEFAULT_MODEL=base # Whisper 모델 (tiny/base/small/medium/large)
|
| 82 |
+
PORT=8000
|
| 83 |
+
```
|
| 84 |
+
|
| 85 |
+
## 📁 프로젝트 구조
|
| 86 |
+
|
| 87 |
+
```
|
| 88 |
+
├── run # 실행 스크립트 (설치 + 실행)
|
| 89 |
+
├── web.py # FastAPI 웹 서버
|
| 90 |
+
├── cli.py # CLI 인터페이스
|
| 91 |
+
├── transcriber.py # 음성 추출 + STT + 키프레임 추출
|
| 92 |
+
├── formatter.py # LLM 호출 + 마크다운 정리 + 번역
|
| 93 |
+
├── templates/
|
| 94 |
+
│ └── index.html # 웹 UI
|
| 95 |
+
├── requirements.txt # Python 의존성
|
| 96 |
+
├── .env.example # 환경 변수 템플릿
|
| 97 |
+
└── output/ # 추출 결과물 (git 제외)
|
| 98 |
+
```
|
| 99 |
+
|
| 100 |
+
## 🔧 Smart Mix (단계별 LLM 선택)
|
| 101 |
+
|
| 102 |
+
웹 UI의 ⚙ 버튼으로 파이프라인 단계별 LLM을 설정할 수 있습니다:
|
| 103 |
+
|
| 104 |
+
- **MD 정리** — 스크립트를 마크다운으로 구조화 (저렴한 모델 OK)
|
| 105 |
+
- **번역** — 다국어 번역 (저렴한 모델 OK)
|
| 106 |
+
- **키프레임 분석** — Vision AI로 화면 분석 (Vision 지원 모델 필요)
|
| 107 |
+
|
| 108 |
+
## 📋 요구사항
|
| 109 |
+
|
| 110 |
+
- Python 3.9+
|
| 111 |
+
- ffmpeg
|
| 112 |
+
- API 키 (사용하려는 LLM에 따라)
|
| 113 |
+
|
| 114 |
+
## 📄 License
|
| 115 |
+
|
| 116 |
+
MIT
|
cli.py
ADDED
|
@@ -0,0 +1,111 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
#!/usr/bin/env python3
|
| 2 |
+
"""YouTube Script Extractor - CLI."""
|
| 3 |
+
|
| 4 |
+
import argparse
|
| 5 |
+
import os
|
| 6 |
+
import sys
|
| 7 |
+
|
| 8 |
+
from dotenv import load_dotenv
|
| 9 |
+
|
| 10 |
+
load_dotenv()
|
| 11 |
+
|
| 12 |
+
from transcriber import process_video
|
| 13 |
+
|
| 14 |
+
|
| 15 |
+
def main():
|
| 16 |
+
parser = argparse.ArgumentParser(
|
| 17 |
+
description="YouTube 영상에서 음성인식으로 스크립트를 추출합니다.",
|
| 18 |
+
formatter_class=argparse.RawDescriptionHelpFormatter,
|
| 19 |
+
epilog="""
|
| 20 |
+
사용 예시:
|
| 21 |
+
# Whisper API 사용 (빠르고 정확)
|
| 22 |
+
python cli.py "https://www.youtube.com/watch?v=..." --mode api
|
| 23 |
+
|
| 24 |
+
# 로컬 Whisper 사용 (무료)
|
| 25 |
+
python cli.py "https://www.youtube.com/watch?v=..." --mode local --model base
|
| 26 |
+
|
| 27 |
+
# SRT 자막만 생성
|
| 28 |
+
python cli.py "https://www.youtube.com/watch?v=..." --format srt
|
| 29 |
+
|
| 30 |
+
# 출력 디렉토리 지정
|
| 31 |
+
python cli.py "https://www.youtube.com/watch?v=..." --output ./my_scripts
|
| 32 |
+
""",
|
| 33 |
+
)
|
| 34 |
+
|
| 35 |
+
parser.add_argument("url", help="YouTube 영상 URL")
|
| 36 |
+
parser.add_argument(
|
| 37 |
+
"--mode",
|
| 38 |
+
choices=["api", "local"],
|
| 39 |
+
default=os.environ.get("DEFAULT_MODE", "api"),
|
| 40 |
+
help="변환 모드 (기본: api)",
|
| 41 |
+
)
|
| 42 |
+
parser.add_argument(
|
| 43 |
+
"--api-key",
|
| 44 |
+
default=None,
|
| 45 |
+
help="OpenAI API 키 (미지정 시 OPENAI_API_KEY 환경변수 사용)",
|
| 46 |
+
)
|
| 47 |
+
parser.add_argument(
|
| 48 |
+
"--model",
|
| 49 |
+
default=os.environ.get("DEFAULT_MODEL", "base"),
|
| 50 |
+
choices=["tiny", "base", "small", "medium", "large"],
|
| 51 |
+
help="로컬 Whisper 모델 크기 (기본: base)",
|
| 52 |
+
)
|
| 53 |
+
parser.add_argument(
|
| 54 |
+
"--format",
|
| 55 |
+
default="txt,srt",
|
| 56 |
+
help="출력 형식, 쉼표로 구분 (기본: txt,srt)",
|
| 57 |
+
)
|
| 58 |
+
parser.add_argument(
|
| 59 |
+
"--output",
|
| 60 |
+
default="./output",
|
| 61 |
+
help="출력 디렉토리 (기본: ./output)",
|
| 62 |
+
)
|
| 63 |
+
|
| 64 |
+
args = parser.parse_args()
|
| 65 |
+
|
| 66 |
+
# API 키 처리
|
| 67 |
+
api_key = args.api_key or os.environ.get("OPENAI_API_KEY")
|
| 68 |
+
if args.mode == "api" and not api_key:
|
| 69 |
+
print("오류: API 모드에서는 OpenAI API 키가 필요합니다.")
|
| 70 |
+
print(" --api-key 인자를 사용하거나 OPENAI_API_KEY 환경변수를 설정하세요.")
|
| 71 |
+
sys.exit(1)
|
| 72 |
+
|
| 73 |
+
formats = [f.strip() for f in args.format.split(",")]
|
| 74 |
+
|
| 75 |
+
def on_progress(msg: str):
|
| 76 |
+
print(f" → {msg}")
|
| 77 |
+
|
| 78 |
+
print(f"YouTube Script Extractor")
|
| 79 |
+
print(f" URL: {args.url}")
|
| 80 |
+
print(f" 모드: {'Whisper API' if args.mode == 'api' else f'로컬 Whisper ({args.model})'}")
|
| 81 |
+
print(f" 출력: {', '.join(formats)}")
|
| 82 |
+
print()
|
| 83 |
+
|
| 84 |
+
try:
|
| 85 |
+
result = process_video(
|
| 86 |
+
url=args.url,
|
| 87 |
+
mode=args.mode,
|
| 88 |
+
api_key=api_key,
|
| 89 |
+
model_size=args.model,
|
| 90 |
+
output_dir=args.output,
|
| 91 |
+
formats=formats,
|
| 92 |
+
on_progress=on_progress,
|
| 93 |
+
)
|
| 94 |
+
|
| 95 |
+
print()
|
| 96 |
+
print(f"제목: {result['title']}")
|
| 97 |
+
print(f"감지된 언어: {result['language']}")
|
| 98 |
+
print(f"생성된 파일:")
|
| 99 |
+
for fmt, path in result["files"].items():
|
| 100 |
+
print(f" [{fmt.upper()}] {path}")
|
| 101 |
+
|
| 102 |
+
except KeyboardInterrupt:
|
| 103 |
+
print("\n중단되었습니다.")
|
| 104 |
+
sys.exit(130)
|
| 105 |
+
except Exception as e:
|
| 106 |
+
print(f"\n오류 발생: {e}")
|
| 107 |
+
sys.exit(1)
|
| 108 |
+
|
| 109 |
+
|
| 110 |
+
if __name__ == "__main__":
|
| 111 |
+
main()
|
formatter.py
ADDED
|
@@ -0,0 +1,687 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""LLM을 사용하여 Whisper 추출 텍스트를 마크다운 정리/번역하는 모듈."""
|
| 2 |
+
|
| 3 |
+
from __future__ import annotations
|
| 4 |
+
|
| 5 |
+
import base64
|
| 6 |
+
import json
|
| 7 |
+
import os
|
| 8 |
+
import re
|
| 9 |
+
import time
|
| 10 |
+
from typing import Callable, List, Optional
|
| 11 |
+
|
| 12 |
+
# ── LLM 정보 (가격순 정렬) ──────────────────────────────────
|
| 13 |
+
LLM_MODELS = {
|
| 14 |
+
"ollama": {
|
| 15 |
+
"name": "Ollama (로컬, 무료)",
|
| 16 |
+
"model": "llama3.2",
|
| 17 |
+
"version": "llama3.2",
|
| 18 |
+
"price_rank": 0,
|
| 19 |
+
"quality_rank": 5,
|
| 20 |
+
"needs_key": False,
|
| 21 |
+
"env_key": None,
|
| 22 |
+
"supports_vision": False,
|
| 23 |
+
"description": "로컬 실행, 무료, 인터넷 불필요",
|
| 24 |
+
},
|
| 25 |
+
"gemini-flash-lite": {
|
| 26 |
+
"name": "Gemini 2.5 Flash-Lite",
|
| 27 |
+
"model": "gemini-2.5-flash-lite",
|
| 28 |
+
"version": "2.5-flash-lite",
|
| 29 |
+
"price_rank": 1,
|
| 30 |
+
"quality_rank": 4,
|
| 31 |
+
"needs_key": True,
|
| 32 |
+
"env_key": "GOOGLE_API_KEY",
|
| 33 |
+
"supports_vision": True,
|
| 34 |
+
"description": "⭐ 최저가, $0.10/1M토큰, 빠르고 안정적",
|
| 35 |
+
},
|
| 36 |
+
"gemini-flash": {
|
| 37 |
+
"name": "Gemini 3.1 Flash-Lite (Preview)",
|
| 38 |
+
"model": "gemini-3.1-flash-lite-preview",
|
| 39 |
+
"version": "3.1-flash-lite",
|
| 40 |
+
"price_rank": 2,
|
| 41 |
+
"quality_rank": 3,
|
| 42 |
+
"needs_key": True,
|
| 43 |
+
"env_key": "GOOGLE_API_KEY",
|
| 44 |
+
"supports_vision": True,
|
| 45 |
+
"description": "최신 모델, $0.25/1M토큰, 향상된 성능",
|
| 46 |
+
},
|
| 47 |
+
"gpt-4o-mini": {
|
| 48 |
+
"name": "GPT-4o-mini (OpenAI)",
|
| 49 |
+
"model": "gpt-4o-mini",
|
| 50 |
+
"version": "4o-mini",
|
| 51 |
+
"price_rank": 3,
|
| 52 |
+
"quality_rank": 2,
|
| 53 |
+
"needs_key": True,
|
| 54 |
+
"env_key": "OPENAI_API_KEY",
|
| 55 |
+
"supports_vision": True,
|
| 56 |
+
"description": "빠르고 저렴, $0.15/1M토큰",
|
| 57 |
+
},
|
| 58 |
+
"claude-sonnet": {
|
| 59 |
+
"name": "Claude Sonnet 4.6 (Anthropic)",
|
| 60 |
+
"model": "claude-sonnet-4-6",
|
| 61 |
+
"version": "sonnet-4.6",
|
| 62 |
+
"price_rank": 4,
|
| 63 |
+
"quality_rank": 1,
|
| 64 |
+
"needs_key": True,
|
| 65 |
+
"env_key": "ANTHROPIC_API_KEY",
|
| 66 |
+
"supports_vision": True,
|
| 67 |
+
"description": "Anthropic, 빠르고 고품질 ($3/1M토큰)",
|
| 68 |
+
},
|
| 69 |
+
"claude-opus": {
|
| 70 |
+
"name": "Claude Opus 4.6 (Anthropic)",
|
| 71 |
+
"model": "claude-opus-4-6",
|
| 72 |
+
"version": "opus-4.6",
|
| 73 |
+
"price_rank": 5,
|
| 74 |
+
"quality_rank": 0,
|
| 75 |
+
"needs_key": True,
|
| 76 |
+
"env_key": "ANTHROPIC_API_KEY",
|
| 77 |
+
"supports_vision": True,
|
| 78 |
+
"description": "Anthropic 최고 모델, 구조화 능력 최고 ($5/1M토큰)",
|
| 79 |
+
},
|
| 80 |
+
}
|
| 81 |
+
|
| 82 |
+
# ── 번역 지원 언어 ────────────────────────────────────────
|
| 83 |
+
LANGUAGES = {
|
| 84 |
+
"ko": "한국어",
|
| 85 |
+
"en": "English",
|
| 86 |
+
"ja": "日本語",
|
| 87 |
+
"zh-CN": "中文(简体)",
|
| 88 |
+
"zh-TW": "中文(繁體)",
|
| 89 |
+
"es": "Español",
|
| 90 |
+
"fr": "Français",
|
| 91 |
+
"de": "Deutsch",
|
| 92 |
+
"pt": "Português",
|
| 93 |
+
"ru": "Русский",
|
| 94 |
+
"vi": "Tiếng Việt",
|
| 95 |
+
"th": "ภาษาไทย",
|
| 96 |
+
"ar": "العربية",
|
| 97 |
+
"hi": "हिन्दी",
|
| 98 |
+
}
|
| 99 |
+
|
| 100 |
+
|
| 101 |
+
def get_models_sorted(sort_by: str = "price") -> list:
|
| 102 |
+
"""정렬된 LLM 모델 리스트를 반환한다."""
|
| 103 |
+
key = "price_rank" if sort_by == "price" else "quality_rank"
|
| 104 |
+
return sorted(
|
| 105 |
+
[{"id": k, **v} for k, v in LLM_MODELS.items()],
|
| 106 |
+
key=lambda x: x[key],
|
| 107 |
+
)
|
| 108 |
+
|
| 109 |
+
|
| 110 |
+
def get_languages() -> list:
|
| 111 |
+
"""지원되는 번역 언어 목록을 반환한다."""
|
| 112 |
+
return [{"code": k, "name": v} for k, v in LANGUAGES.items()]
|
| 113 |
+
|
| 114 |
+
|
| 115 |
+
def make_llm_config(
|
| 116 |
+
global_llm: Optional[str] = None,
|
| 117 |
+
global_api_key: Optional[str] = None,
|
| 118 |
+
global_ollama_model: str = "llama3.2",
|
| 119 |
+
format_llm: Optional[str] = None,
|
| 120 |
+
format_api_key: Optional[str] = None,
|
| 121 |
+
translate_llm: Optional[str] = None,
|
| 122 |
+
translate_api_key: Optional[str] = None,
|
| 123 |
+
keyframe_llm: Optional[str] = None,
|
| 124 |
+
keyframe_api_key: Optional[str] = None,
|
| 125 |
+
) -> dict:
|
| 126 |
+
"""단계별 LLM 설정 딕셔너리를 생성한다.
|
| 127 |
+
|
| 128 |
+
per-step 값이 없으면 global 값으로 fallback.
|
| 129 |
+
반환: {"format": {...}, "translate": {...}, "keyframe": {...}}
|
| 130 |
+
"""
|
| 131 |
+
def _resolve(step_llm, step_key):
|
| 132 |
+
llm = step_llm or global_llm
|
| 133 |
+
api_key = step_key or global_api_key
|
| 134 |
+
return {
|
| 135 |
+
"llm": llm,
|
| 136 |
+
"api_key": api_key,
|
| 137 |
+
"ollama_model": global_ollama_model,
|
| 138 |
+
}
|
| 139 |
+
|
| 140 |
+
return {
|
| 141 |
+
"format": _resolve(format_llm, format_api_key),
|
| 142 |
+
"translate": _resolve(translate_llm, translate_api_key),
|
| 143 |
+
"keyframe": _resolve(
|
| 144 |
+
keyframe_llm or "gemini-flash-lite",
|
| 145 |
+
keyframe_api_key,
|
| 146 |
+
),
|
| 147 |
+
}
|
| 148 |
+
|
| 149 |
+
|
| 150 |
+
# ── 프롬프트 ──────────────────────────────────────────────
|
| 151 |
+
|
| 152 |
+
FORMAT_SYSTEM_PROMPT = """You are an expert document formatter. Your task is to transform raw speech-to-text transcriptions into clean, well-structured Markdown documents.
|
| 153 |
+
|
| 154 |
+
Rules:
|
| 155 |
+
- Detect the content's language and write the output in the SAME language
|
| 156 |
+
- Add a clear title as # heading
|
| 157 |
+
- Write a brief 2-3 sentence summary at the top
|
| 158 |
+
- Divide content into logical sections with ## headings
|
| 159 |
+
- Clean up filler words, repetitions, and stutters
|
| 160 |
+
- Fix obvious grammar/punctuation errors
|
| 161 |
+
- Keep the original meaning and tone intact
|
| 162 |
+
- Use bullet points or numbered lists where appropriate
|
| 163 |
+
- Add --- horizontal rules between major sections
|
| 164 |
+
- Do NOT add information that wasn't in the original text
|
| 165 |
+
- Output ONLY the formatted Markdown, no explanations"""
|
| 166 |
+
|
| 167 |
+
FORMAT_USER_TEMPLATE = """Here is a raw speech-to-text transcription from a YouTube video titled "{title}".
|
| 168 |
+
Please format it into a clean, readable Markdown document.
|
| 169 |
+
|
| 170 |
+
---
|
| 171 |
+
{text}
|
| 172 |
+
---"""
|
| 173 |
+
|
| 174 |
+
FORMAT_SYSTEM_PROMPT_WITH_KEYFRAMES = """You are an expert document formatter. Your task is to transform raw speech-to-text transcriptions into clean, well-structured Markdown documents, enhanced with visual context from video keyframes.
|
| 175 |
+
|
| 176 |
+
Rules:
|
| 177 |
+
- Detect the content's language and write the output in the SAME language
|
| 178 |
+
- Add a clear title as # heading
|
| 179 |
+
- Write a brief 2-3 sentence summary at the top
|
| 180 |
+
- Divide content into logical sections with ## headings
|
| 181 |
+
- Clean up filler words, repetitions, and stutters
|
| 182 |
+
- Fix obvious grammar/punctuation errors
|
| 183 |
+
- Keep the original meaning and tone intact
|
| 184 |
+
- Use bullet points or numbered lists where appropriate
|
| 185 |
+
- Add --- horizontal rules between major sections
|
| 186 |
+
- KEYFRAME CONTEXT: You are also given timestamped descriptions of visual keyframes.
|
| 187 |
+
Insert relevant visual descriptions as blockquotes (> 🖼 [timestamp] description) at
|
| 188 |
+
appropriate positions in the transcript where they add context.
|
| 189 |
+
- Only include keyframe descriptions that add meaningful value (skip redundant ones)
|
| 190 |
+
- Do NOT add information that wasn't in the original text or keyframes
|
| 191 |
+
- Output ONLY the formatted Markdown, no explanations"""
|
| 192 |
+
|
| 193 |
+
FORMAT_USER_TEMPLATE_WITH_KEYFRAMES = """Here is a raw speech-to-text transcription from a YouTube video titled "{title}".
|
| 194 |
+
Please format it into a clean, readable Markdown document.
|
| 195 |
+
|
| 196 |
+
---
|
| 197 |
+
TRANSCRIPT:
|
| 198 |
+
{text}
|
| 199 |
+
---
|
| 200 |
+
|
| 201 |
+
VISUAL KEYFRAME DESCRIPTIONS (timestamped):
|
| 202 |
+
{keyframe_descriptions}
|
| 203 |
+
---"""
|
| 204 |
+
|
| 205 |
+
KEYFRAME_ANALYSIS_SYSTEM_PROMPT = """You are a visual content analyst. Analyze the provided video keyframes and describe what is shown.
|
| 206 |
+
|
| 207 |
+
Rules:
|
| 208 |
+
- Describe each frame concisely (1-3 sentences)
|
| 209 |
+
- Include any visible text (OCR) exactly as shown
|
| 210 |
+
- Note visual elements: diagrams, charts, code, slides, people, scenes
|
| 211 |
+
- Focus on informational content, not aesthetic quality
|
| 212 |
+
- For each frame, output one line in the format: [MM:SS] description
|
| 213 |
+
- Keep descriptions factual and relevant to the video content
|
| 214 |
+
- Output ONLY the descriptions, no extra commentary"""
|
| 215 |
+
|
| 216 |
+
KEYFRAME_ANALYSIS_USER_TEMPLATE = """Analyze these keyframes from a video. For each image, describe what is shown and extract any visible text.
|
| 217 |
+
The timestamps for each frame are provided as labels."""
|
| 218 |
+
|
| 219 |
+
TRANSLATE_SYSTEM_PROMPT = """You are a professional translator. Translate the given text accurately into {target_lang}.
|
| 220 |
+
|
| 221 |
+
Rules:
|
| 222 |
+
- Maintain the original meaning, tone, and nuance
|
| 223 |
+
- If the text uses Markdown formatting, preserve the Markdown structure
|
| 224 |
+
- Translate naturally and idiomatically, not word-by-word
|
| 225 |
+
- Keep proper nouns, brand names, and technical terms appropriately
|
| 226 |
+
- Do NOT add explanations, notes, or commentary
|
| 227 |
+
- Output ONLY the translated text"""
|
| 228 |
+
|
| 229 |
+
TRANSLATE_USER_TEMPLATE = """Translate the following text into {target_lang}:
|
| 230 |
+
|
| 231 |
+
---
|
| 232 |
+
{text}
|
| 233 |
+
---"""
|
| 234 |
+
|
| 235 |
+
|
| 236 |
+
# ── LLM 호출 함수들 (범용) ────────────────────────────────
|
| 237 |
+
|
| 238 |
+
def _call_openai(system_prompt: str, user_prompt: str, api_key: str) -> str:
|
| 239 |
+
"""OpenAI GPT-4o-mini 호출."""
|
| 240 |
+
from openai import OpenAI
|
| 241 |
+
|
| 242 |
+
client = OpenAI(api_key=api_key)
|
| 243 |
+
response = client.chat.completions.create(
|
| 244 |
+
model="gpt-4o-mini",
|
| 245 |
+
messages=[
|
| 246 |
+
{"role": "system", "content": system_prompt},
|
| 247 |
+
{"role": "user", "content": user_prompt},
|
| 248 |
+
],
|
| 249 |
+
temperature=0.3,
|
| 250 |
+
)
|
| 251 |
+
return response.choices[0].message.content
|
| 252 |
+
|
| 253 |
+
|
| 254 |
+
def _call_gemini(system_prompt: str, user_prompt: str, api_key: str,
|
| 255 |
+
model: str = "gemini-2.5-flash-lite") -> str:
|
| 256 |
+
"""Google Gemini 호출."""
|
| 257 |
+
import google.generativeai as genai
|
| 258 |
+
|
| 259 |
+
genai.configure(api_key=api_key)
|
| 260 |
+
gmodel = genai.GenerativeModel(model)
|
| 261 |
+
|
| 262 |
+
prompt = system_prompt + "\n\n" + user_prompt
|
| 263 |
+
response = gmodel.generate_content(prompt)
|
| 264 |
+
return response.text
|
| 265 |
+
|
| 266 |
+
|
| 267 |
+
def _call_ollama(system_prompt: str, user_prompt: str, model_name: str = "llama3.2") -> str:
|
| 268 |
+
"""Ollama 로컬 모델 호출."""
|
| 269 |
+
import urllib.request
|
| 270 |
+
|
| 271 |
+
payload = json.dumps({
|
| 272 |
+
"model": model_name,
|
| 273 |
+
"messages": [
|
| 274 |
+
{"role": "system", "content": system_prompt},
|
| 275 |
+
{"role": "user", "content": user_prompt},
|
| 276 |
+
],
|
| 277 |
+
"stream": False,
|
| 278 |
+
"options": {"temperature": 0.3},
|
| 279 |
+
}).encode("utf-8")
|
| 280 |
+
|
| 281 |
+
req = urllib.request.Request(
|
| 282 |
+
"http://localhost:11434/api/chat",
|
| 283 |
+
data=payload,
|
| 284 |
+
headers={"Content-Type": "application/json"},
|
| 285 |
+
)
|
| 286 |
+
|
| 287 |
+
try:
|
| 288 |
+
with urllib.request.urlopen(req, timeout=300) as resp:
|
| 289 |
+
data = json.loads(resp.read().decode("utf-8"))
|
| 290 |
+
return data["message"]["content"]
|
| 291 |
+
except Exception as e:
|
| 292 |
+
if "Connection refused" in str(e):
|
| 293 |
+
raise RuntimeError(
|
| 294 |
+
"Ollama가 실행 중이 아닙니다. "
|
| 295 |
+
"'ollama serve' 명령으로 먼저 시작해주세요. "
|
| 296 |
+
"(설치: brew install ollama && ollama pull llama3.2)"
|
| 297 |
+
)
|
| 298 |
+
raise
|
| 299 |
+
|
| 300 |
+
|
| 301 |
+
def _call_claude(system_prompt: str, user_prompt: str, api_key: str, model: str = "claude-sonnet-4-6") -> str:
|
| 302 |
+
"""Anthropic Claude 호출."""
|
| 303 |
+
import anthropic
|
| 304 |
+
|
| 305 |
+
client = anthropic.Anthropic(api_key=api_key)
|
| 306 |
+
response = client.messages.create(
|
| 307 |
+
model=model,
|
| 308 |
+
max_tokens=8192,
|
| 309 |
+
system=system_prompt,
|
| 310 |
+
messages=[
|
| 311 |
+
{"role": "user", "content": user_prompt},
|
| 312 |
+
],
|
| 313 |
+
)
|
| 314 |
+
return response.content[0].text
|
| 315 |
+
|
| 316 |
+
|
| 317 |
+
# ── Vision LLM 호출 함수들 ───────────────────────────────
|
| 318 |
+
|
| 319 |
+
def _call_gemini_vision(
|
| 320 |
+
system_prompt: str,
|
| 321 |
+
user_prompt: str,
|
| 322 |
+
images: List[dict],
|
| 323 |
+
api_key: str,
|
| 324 |
+
model: str = "gemini-2.5-flash-lite",
|
| 325 |
+
) -> str:
|
| 326 |
+
"""Google Gemini Vision 호출."""
|
| 327 |
+
import google.generativeai as genai
|
| 328 |
+
from PIL import Image
|
| 329 |
+
|
| 330 |
+
genai.configure(api_key=api_key)
|
| 331 |
+
gmodel = genai.GenerativeModel(model)
|
| 332 |
+
|
| 333 |
+
parts = [system_prompt + "\n\n" + user_prompt]
|
| 334 |
+
for img_info in images:
|
| 335 |
+
img = Image.open(img_info["path"])
|
| 336 |
+
parts.append(img)
|
| 337 |
+
parts.append(f"[Timestamp: {img_info['timestamp']}]")
|
| 338 |
+
|
| 339 |
+
response = gmodel.generate_content(parts)
|
| 340 |
+
return response.text
|
| 341 |
+
|
| 342 |
+
|
| 343 |
+
def _call_openai_vision(
|
| 344 |
+
system_prompt: str,
|
| 345 |
+
user_prompt: str,
|
| 346 |
+
images: List[dict],
|
| 347 |
+
api_key: str,
|
| 348 |
+
) -> str:
|
| 349 |
+
"""OpenAI GPT-4o-mini Vision 호출."""
|
| 350 |
+
from openai import OpenAI
|
| 351 |
+
|
| 352 |
+
client = OpenAI(api_key=api_key)
|
| 353 |
+
content = [{"type": "text", "text": user_prompt}]
|
| 354 |
+
for img_info in images:
|
| 355 |
+
with open(img_info["path"], "rb") as f:
|
| 356 |
+
b64 = base64.b64encode(f.read()).decode()
|
| 357 |
+
content.append({
|
| 358 |
+
"type": "image_url",
|
| 359 |
+
"image_url": {"url": f"data:image/jpeg;base64,{b64}"},
|
| 360 |
+
})
|
| 361 |
+
content.append({"type": "text", "text": f"[Timestamp: {img_info['timestamp']}]"})
|
| 362 |
+
|
| 363 |
+
response = client.chat.completions.create(
|
| 364 |
+
model="gpt-4o-mini",
|
| 365 |
+
messages=[
|
| 366 |
+
{"role": "system", "content": system_prompt},
|
| 367 |
+
{"role": "user", "content": content},
|
| 368 |
+
],
|
| 369 |
+
temperature=0.3,
|
| 370 |
+
)
|
| 371 |
+
return response.choices[0].message.content
|
| 372 |
+
|
| 373 |
+
|
| 374 |
+
def _call_claude_vision(
|
| 375 |
+
system_prompt: str,
|
| 376 |
+
user_prompt: str,
|
| 377 |
+
images: List[dict],
|
| 378 |
+
api_key: str,
|
| 379 |
+
model: str = "claude-sonnet-4-6",
|
| 380 |
+
) -> str:
|
| 381 |
+
"""Anthropic Claude Vision 호출."""
|
| 382 |
+
import anthropic
|
| 383 |
+
|
| 384 |
+
client = anthropic.Anthropic(api_key=api_key)
|
| 385 |
+
content = []
|
| 386 |
+
for img_info in images:
|
| 387 |
+
with open(img_info["path"], "rb") as f:
|
| 388 |
+
b64 = base64.b64encode(f.read()).decode()
|
| 389 |
+
content.append({
|
| 390 |
+
"type": "image",
|
| 391 |
+
"source": {"type": "base64", "media_type": "image/jpeg", "data": b64},
|
| 392 |
+
})
|
| 393 |
+
content.append({"type": "text", "text": f"[Timestamp: {img_info['timestamp']}]"})
|
| 394 |
+
content.append({"type": "text", "text": user_prompt})
|
| 395 |
+
|
| 396 |
+
response = client.messages.create(
|
| 397 |
+
model=model,
|
| 398 |
+
max_tokens=8192,
|
| 399 |
+
system=system_prompt,
|
| 400 |
+
messages=[{"role": "user", "content": content}],
|
| 401 |
+
)
|
| 402 |
+
return response.content[0].text
|
| 403 |
+
|
| 404 |
+
|
| 405 |
+
def _call_vision_llm(
|
| 406 |
+
system_prompt: str,
|
| 407 |
+
user_prompt: str,
|
| 408 |
+
images: List[dict],
|
| 409 |
+
llm_provider: str,
|
| 410 |
+
api_key: Optional[str] = None,
|
| 411 |
+
) -> str:
|
| 412 |
+
"""Vision LLM 호출 디스패처 (Rate limit 자동 재시도 포함)."""
|
| 413 |
+
def _do_call():
|
| 414 |
+
if llm_provider in ("gemini-flash", "gemini-flash-lite"):
|
| 415 |
+
key = api_key or os.environ.get("GOOGLE_API_KEY", "")
|
| 416 |
+
if not key:
|
| 417 |
+
raise ValueError("Google API 키가 필요합니다. (GOOGLE_API_KEY)")
|
| 418 |
+
model_id = LLM_MODELS[llm_provider]["model"]
|
| 419 |
+
return _call_gemini_vision(system_prompt, user_prompt, images, key, model=model_id)
|
| 420 |
+
|
| 421 |
+
elif llm_provider == "gpt-4o-mini":
|
| 422 |
+
key = api_key or os.environ.get("OPENAI_API_KEY", "")
|
| 423 |
+
if not key:
|
| 424 |
+
raise ValueError("OpenAI API 키가 필요합니다.")
|
| 425 |
+
return _call_openai_vision(system_prompt, user_prompt, images, key)
|
| 426 |
+
|
| 427 |
+
elif llm_provider in ("claude-sonnet", "claude-opus"):
|
| 428 |
+
key = api_key or os.environ.get("ANTHROPIC_API_KEY", "")
|
| 429 |
+
if not key:
|
| 430 |
+
raise ValueError("Anthropic API 키가 필요합니다. (ANTHROPIC_API_KEY)")
|
| 431 |
+
model_id = LLM_MODELS[llm_provider]["model"]
|
| 432 |
+
return _call_claude_vision(system_prompt, user_prompt, images, key, model=model_id)
|
| 433 |
+
|
| 434 |
+
elif llm_provider == "ollama":
|
| 435 |
+
raise ValueError("Ollama는 Vision(이미지 분석)을 지원하지 않습니다.")
|
| 436 |
+
|
| 437 |
+
else:
|
| 438 |
+
raise ValueError(f"지원하지 않는 Vision LLM: {llm_provider}")
|
| 439 |
+
|
| 440 |
+
return _retry_on_rate_limit(_do_call)
|
| 441 |
+
|
| 442 |
+
|
| 443 |
+
# ── Rate Limit 재시도 로직 ────────────────────────────────
|
| 444 |
+
|
| 445 |
+
def _is_rate_limit_error(error: Exception) -> bool:
|
| 446 |
+
"""429 Rate Limit 에러인지 확인한다."""
|
| 447 |
+
err_str = str(error).lower()
|
| 448 |
+
err_type = type(error).__name__
|
| 449 |
+
return (
|
| 450 |
+
"rate_limit" in err_str
|
| 451 |
+
or "rate limit" in err_str
|
| 452 |
+
or "429" in err_str
|
| 453 |
+
or "resource_exhausted" in err_str
|
| 454 |
+
or "quota" in err_str
|
| 455 |
+
or err_type == "RateLimitError"
|
| 456 |
+
)
|
| 457 |
+
|
| 458 |
+
|
| 459 |
+
def _parse_retry_after(error: Exception) -> float:
|
| 460 |
+
"""에러 메시지에서 대기 시간(초)을 추출한다."""
|
| 461 |
+
err_str = str(error)
|
| 462 |
+
# "Please try again in 2.129s" 같은 패턴
|
| 463 |
+
match = re.search(r"try again in (\d+\.?\d*)s", err_str)
|
| 464 |
+
if match:
|
| 465 |
+
return float(match.group(1))
|
| 466 |
+
# "Retry-After: 5" 헤더 패턴
|
| 467 |
+
match = re.search(r"retry.?after:?\s*(\d+)", err_str, re.IGNORECASE)
|
| 468 |
+
if match:
|
| 469 |
+
return float(match.group(1))
|
| 470 |
+
return 0.0
|
| 471 |
+
|
| 472 |
+
|
| 473 |
+
def _retry_on_rate_limit(func, *args, max_retries: int = 3, **kwargs):
|
| 474 |
+
"""Rate limit 에러 시 exponential backoff으로 재시도한다."""
|
| 475 |
+
for attempt in range(max_retries + 1):
|
| 476 |
+
try:
|
| 477 |
+
return func(*args, **kwargs)
|
| 478 |
+
except Exception as e:
|
| 479 |
+
if not _is_rate_limit_error(e) or attempt >= max_retries:
|
| 480 |
+
raise
|
| 481 |
+
# 에러에서 대기 시간 추출, 없으면 exponential backoff
|
| 482 |
+
wait = _parse_retry_after(e)
|
| 483 |
+
if wait <= 0:
|
| 484 |
+
wait = (2 ** attempt) * 2 # 2초, 4초, 8초
|
| 485 |
+
wait = min(wait + 0.5, 60) # 여유 0.5초 추가, 최대 60초
|
| 486 |
+
print(f"⏳ Rate limit 초과, {wait:.1f}초 후 재시도... ({attempt + 1}/{max_retries})")
|
| 487 |
+
time.sleep(wait)
|
| 488 |
+
|
| 489 |
+
|
| 490 |
+
# ── 텍스트 LLM 호출 디스패처 ─────────────────────────────
|
| 491 |
+
|
| 492 |
+
def _call_llm(
|
| 493 |
+
system_prompt: str,
|
| 494 |
+
user_prompt: str,
|
| 495 |
+
llm_provider: str,
|
| 496 |
+
api_key: Optional[str] = None,
|
| 497 |
+
ollama_model: str = "llama3.2",
|
| 498 |
+
) -> str:
|
| 499 |
+
"""텍스트 LLM 호출 디스패처 (Rate limit 자동 재시도 포함)."""
|
| 500 |
+
def _do_call():
|
| 501 |
+
if llm_provider == "gpt-4o-mini":
|
| 502 |
+
key = api_key or os.environ.get("OPENAI_API_KEY", "")
|
| 503 |
+
if not key:
|
| 504 |
+
raise ValueError("OpenAI API 키가 필요합니다.")
|
| 505 |
+
return _call_openai(system_prompt, user_prompt, key)
|
| 506 |
+
|
| 507 |
+
elif llm_provider in ("gemini-flash", "gemini-flash-lite"):
|
| 508 |
+
key = api_key or os.environ.get("GOOGLE_API_KEY", "")
|
| 509 |
+
if not key:
|
| 510 |
+
raise ValueError("Google API 키가 필요합니다. (GOOGLE_API_KEY)")
|
| 511 |
+
model_id = LLM_MODELS[llm_provider]["model"]
|
| 512 |
+
return _call_gemini(system_prompt, user_prompt, key, model=model_id)
|
| 513 |
+
|
| 514 |
+
elif llm_provider == "ollama":
|
| 515 |
+
return _call_ollama(system_prompt, user_prompt, ollama_model)
|
| 516 |
+
|
| 517 |
+
elif llm_provider in ("claude-sonnet", "claude-opus"):
|
| 518 |
+
key = api_key or os.environ.get("ANTHROPIC_API_KEY", "")
|
| 519 |
+
if not key:
|
| 520 |
+
raise ValueError("Anthropic API 키가 필요합니다. (ANTHROPIC_API_KEY)")
|
| 521 |
+
model_id = LLM_MODELS[llm_provider]["model"]
|
| 522 |
+
return _call_claude(system_prompt, user_prompt, key, model=model_id)
|
| 523 |
+
|
| 524 |
+
else:
|
| 525 |
+
raise ValueError(f"지원하지 않는 LLM: {llm_provider}")
|
| 526 |
+
|
| 527 |
+
return _retry_on_rate_limit(_do_call)
|
| 528 |
+
|
| 529 |
+
|
| 530 |
+
# ── 키프레임 분석 ─────────────────────────────────────────
|
| 531 |
+
|
| 532 |
+
def analyze_keyframes(
|
| 533 |
+
keyframe_paths: List[dict],
|
| 534 |
+
llm_provider: str = "gemini-flash-lite",
|
| 535 |
+
api_key: Optional[str] = None,
|
| 536 |
+
on_progress: Optional[Callable] = None,
|
| 537 |
+
batch_size: int = 10,
|
| 538 |
+
) -> str:
|
| 539 |
+
"""Vision LLM으로 키프레임 이미지를 분석한다.
|
| 540 |
+
|
| 541 |
+
keyframe_paths: [{"path": str, "timestamp": str}, ...]
|
| 542 |
+
반환: 타임스탬프별 설명 텍스트
|
| 543 |
+
"""
|
| 544 |
+
def _notify(percent: int, detail: str):
|
| 545 |
+
if on_progress:
|
| 546 |
+
on_progress({"step": "keyframe_analysis", "percent": percent, "detail": detail})
|
| 547 |
+
|
| 548 |
+
model_info = LLM_MODELS.get(llm_provider, {})
|
| 549 |
+
model_name = model_info.get("name", llm_provider)
|
| 550 |
+
|
| 551 |
+
total = len(keyframe_paths)
|
| 552 |
+
_notify(5, f"{total}개 키프레임을 {model_name}으로 분석 준비 중...")
|
| 553 |
+
|
| 554 |
+
all_descriptions = []
|
| 555 |
+
batches = [keyframe_paths[i:i + batch_size] for i in range(0, total, batch_size)]
|
| 556 |
+
|
| 557 |
+
for idx, batch in enumerate(batches):
|
| 558 |
+
pct = int(10 + (idx / len(batches)) * 80)
|
| 559 |
+
_notify(pct, f"배치 {idx + 1}/{len(batches)} 분석 중 ({len(batch)}프레임)...")
|
| 560 |
+
|
| 561 |
+
try:
|
| 562 |
+
result = _call_vision_llm(
|
| 563 |
+
system_prompt=KEYFRAME_ANALYSIS_SYSTEM_PROMPT,
|
| 564 |
+
user_prompt=KEYFRAME_ANALYSIS_USER_TEMPLATE,
|
| 565 |
+
images=batch,
|
| 566 |
+
llm_provider=llm_provider,
|
| 567 |
+
api_key=api_key,
|
| 568 |
+
)
|
| 569 |
+
all_descriptions.append(result.strip())
|
| 570 |
+
except Exception as e:
|
| 571 |
+
_notify(pct, f"배치 {idx + 1} 분석 오류: {str(e)}")
|
| 572 |
+
raise
|
| 573 |
+
|
| 574 |
+
_notify(100, f"{total}개 키프레임 분석 완료")
|
| 575 |
+
return "\n".join(all_descriptions)
|
| 576 |
+
|
| 577 |
+
|
| 578 |
+
# ── 메인 함수들 ───────────────────────────────────────────
|
| 579 |
+
|
| 580 |
+
def _truncate_text(text: str, max_chars: int = 100000) -> tuple:
|
| 581 |
+
"""텍스트가 너무 길면 잘라낸다. (text, truncated) 반환."""
|
| 582 |
+
if len(text) > max_chars:
|
| 583 |
+
return text[:max_chars], True
|
| 584 |
+
return text, False
|
| 585 |
+
|
| 586 |
+
|
| 587 |
+
def format_as_markdown(
|
| 588 |
+
text: str,
|
| 589 |
+
title: str,
|
| 590 |
+
llm_provider: str = "gpt-4o-mini",
|
| 591 |
+
api_key: Optional[str] = None,
|
| 592 |
+
ollama_model: str = "llama3.2",
|
| 593 |
+
on_progress: Optional[Callable] = None,
|
| 594 |
+
keyframe_descriptions: Optional[str] = None,
|
| 595 |
+
) -> str:
|
| 596 |
+
"""Whisper 추출 텍스트를 LLM으로 마크다운으로 정리한다.
|
| 597 |
+
|
| 598 |
+
keyframe_descriptions가 주어지면 키프레임 설명을 MD에 통합한다.
|
| 599 |
+
"""
|
| 600 |
+
def _notify(percent: int, detail: str):
|
| 601 |
+
if on_progress:
|
| 602 |
+
on_progress({"step": "format", "percent": percent, "detail": detail})
|
| 603 |
+
|
| 604 |
+
model_info = LLM_MODELS.get(llm_provider, {})
|
| 605 |
+
model_name = model_info.get("name", llm_provider)
|
| 606 |
+
|
| 607 |
+
has_keyframes = bool(keyframe_descriptions and keyframe_descriptions.strip())
|
| 608 |
+
extra = " + 키프레임 컨텍스트" if has_keyframes else ""
|
| 609 |
+
_notify(10, f"{model_name}에 텍스트{extra} 전송 중...")
|
| 610 |
+
|
| 611 |
+
text, truncated = _truncate_text(text)
|
| 612 |
+
if truncated:
|
| 613 |
+
_notify(15, f"텍스트가 길어서 앞부분만 정리합니다 ({len(text)}자)")
|
| 614 |
+
|
| 615 |
+
_notify(30, f"{model_name} 처리 중...")
|
| 616 |
+
|
| 617 |
+
# 키프레임 설명이 있으면 통합 프롬프트 사용
|
| 618 |
+
if has_keyframes:
|
| 619 |
+
sys_prompt = FORMAT_SYSTEM_PROMPT_WITH_KEYFRAMES
|
| 620 |
+
usr_prompt = FORMAT_USER_TEMPLATE_WITH_KEYFRAMES.format(
|
| 621 |
+
title=title, text=text, keyframe_descriptions=keyframe_descriptions,
|
| 622 |
+
)
|
| 623 |
+
else:
|
| 624 |
+
sys_prompt = FORMAT_SYSTEM_PROMPT
|
| 625 |
+
usr_prompt = FORMAT_USER_TEMPLATE.format(title=title, text=text)
|
| 626 |
+
|
| 627 |
+
try:
|
| 628 |
+
result = _call_llm(
|
| 629 |
+
system_prompt=sys_prompt,
|
| 630 |
+
user_prompt=usr_prompt,
|
| 631 |
+
llm_provider=llm_provider,
|
| 632 |
+
api_key=api_key,
|
| 633 |
+
ollama_model=ollama_model,
|
| 634 |
+
)
|
| 635 |
+
except Exception as e:
|
| 636 |
+
_notify(0, f"LLM 오류: {str(e)}")
|
| 637 |
+
raise
|
| 638 |
+
|
| 639 |
+
if truncated:
|
| 640 |
+
result += "\n\n---\n> ⚠️ 원본 텍스트가 길어서 일부만 정리되었습니다.\n"
|
| 641 |
+
|
| 642 |
+
_notify(100, f"{model_name} 정리 완료")
|
| 643 |
+
return result
|
| 644 |
+
|
| 645 |
+
|
| 646 |
+
def translate_text(
|
| 647 |
+
text: str,
|
| 648 |
+
target_lang: str,
|
| 649 |
+
llm_provider: str = "gpt-4o-mini",
|
| 650 |
+
api_key: Optional[str] = None,
|
| 651 |
+
ollama_model: str = "llama3.2",
|
| 652 |
+
on_progress: Optional[Callable] = None,
|
| 653 |
+
) -> str:
|
| 654 |
+
"""텍스트를 지정 언어로 번역한다."""
|
| 655 |
+
def _notify(percent: int, detail: str):
|
| 656 |
+
if on_progress:
|
| 657 |
+
on_progress({"step": "translate", "percent": percent, "detail": detail})
|
| 658 |
+
|
| 659 |
+
lang_name = LANGUAGES.get(target_lang, target_lang)
|
| 660 |
+
model_info = LLM_MODELS.get(llm_provider, {})
|
| 661 |
+
model_name = model_info.get("name", llm_provider)
|
| 662 |
+
|
| 663 |
+
_notify(10, f"{lang_name}로 번역 준비 중...")
|
| 664 |
+
|
| 665 |
+
text, truncated = _truncate_text(text)
|
| 666 |
+
if truncated:
|
| 667 |
+
_notify(15, f"텍스트가 길어서 앞부분만 번역합니다 ({len(text)}자)")
|
| 668 |
+
|
| 669 |
+
_notify(30, f"{model_name}으로 {lang_name} 번역 중...")
|
| 670 |
+
|
| 671 |
+
try:
|
| 672 |
+
result = _call_llm(
|
| 673 |
+
system_prompt=TRANSLATE_SYSTEM_PROMPT.format(target_lang=lang_name),
|
| 674 |
+
user_prompt=TRANSLATE_USER_TEMPLATE.format(target_lang=lang_name, text=text),
|
| 675 |
+
llm_provider=llm_provider,
|
| 676 |
+
api_key=api_key,
|
| 677 |
+
ollama_model=ollama_model,
|
| 678 |
+
)
|
| 679 |
+
except Exception as e:
|
| 680 |
+
_notify(0, f"번역 오류: {str(e)}")
|
| 681 |
+
raise
|
| 682 |
+
|
| 683 |
+
if truncated:
|
| 684 |
+
result += f"\n\n---\n> ⚠️ 원본 텍스트가 길어서 일부만 번역되었습니다.\n"
|
| 685 |
+
|
| 686 |
+
_notify(100, f"{lang_name} 번역 완료")
|
| 687 |
+
return result
|
requirements.txt
ADDED
|
@@ -0,0 +1,15 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
pytubefix>=6.0.0
|
| 2 |
+
openai-whisper>=20231117
|
| 3 |
+
openai>=1.0.0
|
| 4 |
+
fastapi>=0.104.0
|
| 5 |
+
uvicorn>=0.24.0
|
| 6 |
+
jinja2>=3.1.0
|
| 7 |
+
python-multipart>=0.0.6
|
| 8 |
+
pydub>=0.25.1
|
| 9 |
+
python-dotenv>=1.0.0
|
| 10 |
+
# MD 포맷팅용 (선택적, 사용하는 LLM만 설치하면 됨)
|
| 11 |
+
google-generativeai>=0.3.0
|
| 12 |
+
anthropic>=0.18.0
|
| 13 |
+
# 키프레임 분석용 (선택적)
|
| 14 |
+
opencv-python>=4.8.0
|
| 15 |
+
Pillow>=10.0.0
|
run
ADDED
|
@@ -0,0 +1,68 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
#!/bin/bash
|
| 2 |
+
# YouTube Script Extractor 실행 스크립트
|
| 3 |
+
|
| 4 |
+
DIR="$(cd "$(dirname "$0")" && pwd)"
|
| 5 |
+
cd "$DIR"
|
| 6 |
+
|
| 7 |
+
# 첫 실행: .env 파일이 없으면 자동 설정
|
| 8 |
+
if [ ! -f .env ]; then
|
| 9 |
+
echo "=== YouTube Script Extractor 초기 설정 ==="
|
| 10 |
+
echo ""
|
| 11 |
+
|
| 12 |
+
read -p "OpenAI API Key를 입력하세요 (sk-...): " api_key
|
| 13 |
+
if [ -n "$api_key" ]; then
|
| 14 |
+
cat > .env << EOF
|
| 15 |
+
OPENAI_API_KEY=$api_key
|
| 16 |
+
DEFAULT_MODE=api
|
| 17 |
+
DEFAULT_MODEL=base
|
| 18 |
+
PORT=8000
|
| 19 |
+
EOF
|
| 20 |
+
echo ""
|
| 21 |
+
echo ".env 파일이 생성되었습니다."
|
| 22 |
+
else
|
| 23 |
+
cp .env.example .env
|
| 24 |
+
echo ".env.example을 복사했습니다. 나중에 .env 파일을 직접 수정해주세요."
|
| 25 |
+
fi
|
| 26 |
+
echo ""
|
| 27 |
+
fi
|
| 28 |
+
|
| 29 |
+
# ffmpeg 체크
|
| 30 |
+
if ! command -v ffmpeg &>/dev/null; then
|
| 31 |
+
echo "ffmpeg이 필요합니다. 설치합니다..."
|
| 32 |
+
brew install ffmpeg || { echo "ffmpeg 설치 실패. 'brew install ffmpeg'을 직접 실행해주세요."; exit 1; }
|
| 33 |
+
echo ""
|
| 34 |
+
fi
|
| 35 |
+
|
| 36 |
+
# Python 의존성 체크
|
| 37 |
+
if ! python3 -c "import yt_dlp" 2>/dev/null; then
|
| 38 |
+
echo "Python 의존성을 설치합니다..."
|
| 39 |
+
python3 -m pip install -r requirements.txt || { echo "설치 실패. python3와 pip이 설치되어 있는지 확인해주세요."; exit 1; }
|
| 40 |
+
echo ""
|
| 41 |
+
fi
|
| 42 |
+
|
| 43 |
+
# 실행 모드 선택
|
| 44 |
+
case "${1:-}" in
|
| 45 |
+
web|w)
|
| 46 |
+
echo "웹 서버를 시작합니다..."
|
| 47 |
+
python3 web.py
|
| 48 |
+
;;
|
| 49 |
+
"")
|
| 50 |
+
# URL 없이 실행하면 웹 서버
|
| 51 |
+
echo "웹 서버를 시작합니다..."
|
| 52 |
+
python3 web.py
|
| 53 |
+
;;
|
| 54 |
+
http*|www.*)
|
| 55 |
+
# URL이 주어지면 CLI 모드
|
| 56 |
+
python3 cli.py "$@"
|
| 57 |
+
;;
|
| 58 |
+
*)
|
| 59 |
+
echo "사용법:"
|
| 60 |
+
echo " ./run 웹 서버 시작"
|
| 61 |
+
echo " ./run web 웹 서버 시작"
|
| 62 |
+
echo " ./run <YouTube URL> CLI로 스크립트 추출"
|
| 63 |
+
echo ""
|
| 64 |
+
echo "CLI 옵션:"
|
| 65 |
+
echo " ./run <URL> --mode local 로컬 Whisper 사용"
|
| 66 |
+
echo " ./run <URL> --format srt SRT만 생성"
|
| 67 |
+
;;
|
| 68 |
+
esac
|
templates/index.html
ADDED
|
@@ -0,0 +1,1448 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
<!DOCTYPE html>
|
| 2 |
+
<html lang="ko">
|
| 3 |
+
<head>
|
| 4 |
+
<meta charset="UTF-8">
|
| 5 |
+
<meta name="viewport" content="width=device-width, initial-scale=1.0">
|
| 6 |
+
<title>YouTube Script Extractor</title>
|
| 7 |
+
<style>
|
| 8 |
+
* { margin: 0; padding: 0; box-sizing: border-box; }
|
| 9 |
+
|
| 10 |
+
body {
|
| 11 |
+
font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif;
|
| 12 |
+
background: #0f0f0f;
|
| 13 |
+
color: #e1e1e1;
|
| 14 |
+
min-height: 100vh;
|
| 15 |
+
}
|
| 16 |
+
|
| 17 |
+
.container {
|
| 18 |
+
max-width: 800px;
|
| 19 |
+
margin: 0 auto;
|
| 20 |
+
padding: 40px 20px;
|
| 21 |
+
}
|
| 22 |
+
|
| 23 |
+
h1 {
|
| 24 |
+
text-align: center;
|
| 25 |
+
font-size: 28px;
|
| 26 |
+
margin-bottom: 8px;
|
| 27 |
+
color: #fff;
|
| 28 |
+
}
|
| 29 |
+
|
| 30 |
+
.subtitle {
|
| 31 |
+
text-align: center;
|
| 32 |
+
color: #888;
|
| 33 |
+
margin-bottom: 40px;
|
| 34 |
+
font-size: 14px;
|
| 35 |
+
}
|
| 36 |
+
|
| 37 |
+
.card {
|
| 38 |
+
background: #1a1a1a;
|
| 39 |
+
border: 1px solid #333;
|
| 40 |
+
border-radius: 12px;
|
| 41 |
+
padding: 24px;
|
| 42 |
+
margin-bottom: 20px;
|
| 43 |
+
}
|
| 44 |
+
|
| 45 |
+
.form-group { margin-bottom: 16px; }
|
| 46 |
+
|
| 47 |
+
label {
|
| 48 |
+
display: block;
|
| 49 |
+
font-size: 13px;
|
| 50 |
+
color: #aaa;
|
| 51 |
+
margin-bottom: 6px;
|
| 52 |
+
font-weight: 500;
|
| 53 |
+
}
|
| 54 |
+
|
| 55 |
+
input[type="text"], select {
|
| 56 |
+
width: 100%;
|
| 57 |
+
padding: 10px 14px;
|
| 58 |
+
background: #0f0f0f;
|
| 59 |
+
border: 1px solid #333;
|
| 60 |
+
border-radius: 8px;
|
| 61 |
+
color: #fff;
|
| 62 |
+
font-size: 14px;
|
| 63 |
+
outline: none;
|
| 64 |
+
transition: border-color 0.2s;
|
| 65 |
+
}
|
| 66 |
+
|
| 67 |
+
input[type="text"]:focus, select:focus { border-color: #ff4444; }
|
| 68 |
+
|
| 69 |
+
.row { display: flex; gap: 12px; }
|
| 70 |
+
.row .form-group { flex: 1; }
|
| 71 |
+
|
| 72 |
+
.checkbox-group {
|
| 73 |
+
display: flex;
|
| 74 |
+
gap: 16px;
|
| 75 |
+
align-items: center;
|
| 76 |
+
}
|
| 77 |
+
|
| 78 |
+
.checkbox-group label {
|
| 79 |
+
display: flex;
|
| 80 |
+
align-items: center;
|
| 81 |
+
gap: 6px;
|
| 82 |
+
cursor: pointer;
|
| 83 |
+
color: #e1e1e1;
|
| 84 |
+
font-size: 14px;
|
| 85 |
+
}
|
| 86 |
+
|
| 87 |
+
input[type="checkbox"] {
|
| 88 |
+
accent-color: #ff4444;
|
| 89 |
+
width: 16px;
|
| 90 |
+
height: 16px;
|
| 91 |
+
}
|
| 92 |
+
|
| 93 |
+
.btn {
|
| 94 |
+
width: 100%;
|
| 95 |
+
padding: 12px;
|
| 96 |
+
background: #ff4444;
|
| 97 |
+
color: #fff;
|
| 98 |
+
border: none;
|
| 99 |
+
border-radius: 8px;
|
| 100 |
+
font-size: 16px;
|
| 101 |
+
font-weight: 600;
|
| 102 |
+
cursor: pointer;
|
| 103 |
+
transition: background 0.2s;
|
| 104 |
+
}
|
| 105 |
+
|
| 106 |
+
.btn:hover { background: #e03030; }
|
| 107 |
+
.btn:disabled {
|
| 108 |
+
background: #333;
|
| 109 |
+
color: #666;
|
| 110 |
+
cursor: not-allowed;
|
| 111 |
+
}
|
| 112 |
+
|
| 113 |
+
.mode-note {
|
| 114 |
+
font-size: 12px;
|
| 115 |
+
color: #666;
|
| 116 |
+
margin-top: 4px;
|
| 117 |
+
}
|
| 118 |
+
|
| 119 |
+
/* ========= 진행 상태 ========= */
|
| 120 |
+
#progress-section { display: none; }
|
| 121 |
+
|
| 122 |
+
.steps { display: flex; flex-direction: column; gap: 0; }
|
| 123 |
+
|
| 124 |
+
.step {
|
| 125 |
+
padding: 16px 0;
|
| 126 |
+
border-bottom: 1px solid #222;
|
| 127 |
+
color: #555;
|
| 128 |
+
transition: color 0.3s;
|
| 129 |
+
}
|
| 130 |
+
|
| 131 |
+
.step:last-child { border-bottom: none; }
|
| 132 |
+
.step.active { color: #fff; }
|
| 133 |
+
.step.done { color: #4caf50; }
|
| 134 |
+
.step.error { color: #ff4444; }
|
| 135 |
+
|
| 136 |
+
.step-header {
|
| 137 |
+
display: flex;
|
| 138 |
+
align-items: center;
|
| 139 |
+
gap: 12px;
|
| 140 |
+
}
|
| 141 |
+
|
| 142 |
+
.step-icon {
|
| 143 |
+
width: 28px;
|
| 144 |
+
height: 28px;
|
| 145 |
+
border-radius: 50%;
|
| 146 |
+
display: flex;
|
| 147 |
+
align-items: center;
|
| 148 |
+
justify-content: center;
|
| 149 |
+
font-size: 14px;
|
| 150 |
+
flex-shrink: 0;
|
| 151 |
+
background: #222;
|
| 152 |
+
color: #555;
|
| 153 |
+
transition: all 0.3s;
|
| 154 |
+
}
|
| 155 |
+
|
| 156 |
+
.step.active .step-icon {
|
| 157 |
+
background: #ff4444;
|
| 158 |
+
color: #fff;
|
| 159 |
+
}
|
| 160 |
+
|
| 161 |
+
.step.done .step-icon {
|
| 162 |
+
background: #4caf50;
|
| 163 |
+
color: #fff;
|
| 164 |
+
}
|
| 165 |
+
|
| 166 |
+
.step-title { font-size: 14px; font-weight: 500; }
|
| 167 |
+
|
| 168 |
+
.step-percent {
|
| 169 |
+
margin-left: auto;
|
| 170 |
+
font-size: 13px;
|
| 171 |
+
font-weight: 600;
|
| 172 |
+
font-family: 'SF Mono', Menlo, monospace;
|
| 173 |
+
color: #555;
|
| 174 |
+
min-width: 40px;
|
| 175 |
+
text-align: right;
|
| 176 |
+
}
|
| 177 |
+
|
| 178 |
+
.step.active .step-percent { color: #ff4444; }
|
| 179 |
+
.step.done .step-percent { color: #4caf50; }
|
| 180 |
+
|
| 181 |
+
/* 프로그레스 바 */
|
| 182 |
+
.step-progress {
|
| 183 |
+
height: 4px;
|
| 184 |
+
background: #222;
|
| 185 |
+
border-radius: 2px;
|
| 186 |
+
margin: 10px 0 0 40px;
|
| 187 |
+
overflow: hidden;
|
| 188 |
+
}
|
| 189 |
+
|
| 190 |
+
.step-progress-bar {
|
| 191 |
+
height: 100%;
|
| 192 |
+
background: #ff4444;
|
| 193 |
+
border-radius: 2px;
|
| 194 |
+
width: 0%;
|
| 195 |
+
transition: width 0.4s ease;
|
| 196 |
+
}
|
| 197 |
+
|
| 198 |
+
.step.done .step-progress-bar {
|
| 199 |
+
background: #4caf50;
|
| 200 |
+
width: 100%;
|
| 201 |
+
}
|
| 202 |
+
|
| 203 |
+
.step-progress-bar.indeterminate {
|
| 204 |
+
width: 30%;
|
| 205 |
+
animation: indeterminate 1.5s ease-in-out infinite;
|
| 206 |
+
}
|
| 207 |
+
|
| 208 |
+
@keyframes indeterminate {
|
| 209 |
+
0% { margin-left: 0%; }
|
| 210 |
+
50% { margin-left: 70%; }
|
| 211 |
+
100% { margin-left: 0%; }
|
| 212 |
+
}
|
| 213 |
+
|
| 214 |
+
/* 상세 정보 */
|
| 215 |
+
.step-detail {
|
| 216 |
+
font-size: 12px;
|
| 217 |
+
color: #555;
|
| 218 |
+
margin: 6px 0 0 40px;
|
| 219 |
+
min-height: 16px;
|
| 220 |
+
transition: color 0.3s;
|
| 221 |
+
}
|
| 222 |
+
|
| 223 |
+
.step.active .step-detail { color: #999; }
|
| 224 |
+
.step.done .step-detail { color: #4caf5099; }
|
| 225 |
+
|
| 226 |
+
/* 타이머 + 팁 */
|
| 227 |
+
.progress-footer {
|
| 228 |
+
display: flex;
|
| 229 |
+
justify-content: space-between;
|
| 230 |
+
align-items: center;
|
| 231 |
+
margin-top: 20px;
|
| 232 |
+
padding-top: 16px;
|
| 233 |
+
border-top: 1px solid #222;
|
| 234 |
+
}
|
| 235 |
+
|
| 236 |
+
.elapsed {
|
| 237 |
+
font-size: 13px;
|
| 238 |
+
color: #888;
|
| 239 |
+
font-family: 'SF Mono', Menlo, monospace;
|
| 240 |
+
}
|
| 241 |
+
|
| 242 |
+
.tip {
|
| 243 |
+
font-size: 12px;
|
| 244 |
+
color: #666;
|
| 245 |
+
font-style: italic;
|
| 246 |
+
max-width: 60%;
|
| 247 |
+
text-align: right;
|
| 248 |
+
animation: fadeInOut 6s ease-in-out infinite;
|
| 249 |
+
}
|
| 250 |
+
|
| 251 |
+
@keyframes fadeInOut {
|
| 252 |
+
0%, 100% { opacity: 0.4; }
|
| 253 |
+
50% { opacity: 1; }
|
| 254 |
+
}
|
| 255 |
+
|
| 256 |
+
/* ========= 결과 ========= */
|
| 257 |
+
#result-section { display: none; }
|
| 258 |
+
|
| 259 |
+
.result-banner {
|
| 260 |
+
background: #1a2e1a;
|
| 261 |
+
border: 1px solid #2d5a2d;
|
| 262 |
+
border-radius: 8px;
|
| 263 |
+
padding: 16px;
|
| 264 |
+
margin-bottom: 16px;
|
| 265 |
+
display: flex;
|
| 266 |
+
align-items: center;
|
| 267 |
+
gap: 12px;
|
| 268 |
+
}
|
| 269 |
+
|
| 270 |
+
.result-banner-icon { font-size: 24px; flex-shrink: 0; }
|
| 271 |
+
|
| 272 |
+
.result-banner-text { flex: 1; }
|
| 273 |
+
|
| 274 |
+
.result-banner-title {
|
| 275 |
+
font-size: 15px;
|
| 276 |
+
font-weight: 600;
|
| 277 |
+
color: #4caf50;
|
| 278 |
+
margin-bottom: 4px;
|
| 279 |
+
}
|
| 280 |
+
|
| 281 |
+
.result-banner-info { font-size: 13px; color: #888; }
|
| 282 |
+
|
| 283 |
+
.result-header {
|
| 284 |
+
display: flex;
|
| 285 |
+
justify-content: space-between;
|
| 286 |
+
align-items: center;
|
| 287 |
+
margin-bottom: 12px;
|
| 288 |
+
}
|
| 289 |
+
|
| 290 |
+
.result-title { font-size: 16px; font-weight: 600; color: #fff; }
|
| 291 |
+
|
| 292 |
+
.result-lang {
|
| 293 |
+
font-size: 12px;
|
| 294 |
+
background: #333;
|
| 295 |
+
color: #aaa;
|
| 296 |
+
padding: 4px 10px;
|
| 297 |
+
border-radius: 12px;
|
| 298 |
+
}
|
| 299 |
+
|
| 300 |
+
.result-text {
|
| 301 |
+
background: #0f0f0f;
|
| 302 |
+
border-radius: 8px;
|
| 303 |
+
padding: 16px;
|
| 304 |
+
font-size: 14px;
|
| 305 |
+
line-height: 1.7;
|
| 306 |
+
max-height: 400px;
|
| 307 |
+
overflow-y: auto;
|
| 308 |
+
white-space: pre-wrap;
|
| 309 |
+
word-break: break-word;
|
| 310 |
+
color: #ccc;
|
| 311 |
+
}
|
| 312 |
+
|
| 313 |
+
.download-buttons { display: flex; gap: 10px; margin-top: 16px; }
|
| 314 |
+
|
| 315 |
+
.btn-download {
|
| 316 |
+
flex: 1;
|
| 317 |
+
padding: 12px;
|
| 318 |
+
background: #1a3a1a;
|
| 319 |
+
border: 1px solid #2d5a2d;
|
| 320 |
+
border-radius: 8px;
|
| 321 |
+
color: #4caf50;
|
| 322 |
+
font-size: 14px;
|
| 323 |
+
font-weight: 600;
|
| 324 |
+
cursor: pointer;
|
| 325 |
+
text-align: center;
|
| 326 |
+
text-decoration: none;
|
| 327 |
+
transition: background 0.2s;
|
| 328 |
+
}
|
| 329 |
+
|
| 330 |
+
.btn-download:hover { background: #2d5a2d; color: #fff; }
|
| 331 |
+
|
| 332 |
+
.save-path {
|
| 333 |
+
font-size: 12px;
|
| 334 |
+
color: #666;
|
| 335 |
+
margin-top: 12px;
|
| 336 |
+
font-family: 'SF Mono', Menlo, monospace;
|
| 337 |
+
padding: 8px 12px;
|
| 338 |
+
background: #111;
|
| 339 |
+
border-radius: 6px;
|
| 340 |
+
}
|
| 341 |
+
|
| 342 |
+
.error-msg {
|
| 343 |
+
color: #ff4444;
|
| 344 |
+
padding: 12px;
|
| 345 |
+
font-size: 14px;
|
| 346 |
+
background: #2a1515;
|
| 347 |
+
border-radius: 8px;
|
| 348 |
+
border: 1px solid #4a2020;
|
| 349 |
+
}
|
| 350 |
+
|
| 351 |
+
.btn-new {
|
| 352 |
+
width: 100%;
|
| 353 |
+
padding: 10px;
|
| 354 |
+
background: transparent;
|
| 355 |
+
border: 1px solid #444;
|
| 356 |
+
border-radius: 8px;
|
| 357 |
+
color: #aaa;
|
| 358 |
+
font-size: 14px;
|
| 359 |
+
cursor: pointer;
|
| 360 |
+
margin-top: 12px;
|
| 361 |
+
transition: all 0.2s;
|
| 362 |
+
}
|
| 363 |
+
|
| 364 |
+
.btn-new:hover { border-color: #888; color: #fff; }
|
| 365 |
+
|
| 366 |
+
/* ========= LLM 옵션 (MD/번역 공용) ========= */
|
| 367 |
+
.llm-options {
|
| 368 |
+
display: none;
|
| 369 |
+
margin-top: 12px;
|
| 370 |
+
padding: 16px;
|
| 371 |
+
background: #111;
|
| 372 |
+
border-radius: 8px;
|
| 373 |
+
border: 1px solid #2a2a2a;
|
| 374 |
+
}
|
| 375 |
+
|
| 376 |
+
.llm-options.visible { display: block; }
|
| 377 |
+
|
| 378 |
+
.llm-options .form-group { margin-bottom: 12px; }
|
| 379 |
+
.llm-options .form-group:last-child { margin-bottom: 0; }
|
| 380 |
+
|
| 381 |
+
.llm-select-desc {
|
| 382 |
+
font-size: 12px;
|
| 383 |
+
color: #888;
|
| 384 |
+
margin-top: 4px;
|
| 385 |
+
}
|
| 386 |
+
|
| 387 |
+
/* ========= 설정 버튼 ========= */
|
| 388 |
+
.llm-header { display: flex; gap: 8px; align-items: stretch; }
|
| 389 |
+
.llm-header > div { flex: 1; }
|
| 390 |
+
|
| 391 |
+
.btn-settings {
|
| 392 |
+
padding: 8px 14px;
|
| 393 |
+
background: #222;
|
| 394 |
+
border: 1px solid #333;
|
| 395 |
+
border-radius: 8px;
|
| 396 |
+
color: #aaa;
|
| 397 |
+
font-size: 16px;
|
| 398 |
+
cursor: pointer;
|
| 399 |
+
transition: all 0.2s;
|
| 400 |
+
flex-shrink: 0;
|
| 401 |
+
position: relative;
|
| 402 |
+
}
|
| 403 |
+
.btn-settings:hover { background: #333; color: #fff; border-color: #555; }
|
| 404 |
+
.settings-dot {
|
| 405 |
+
display: none;
|
| 406 |
+
position: absolute;
|
| 407 |
+
top: 4px; right: 4px;
|
| 408 |
+
width: 8px; height: 8px;
|
| 409 |
+
border-radius: 50%;
|
| 410 |
+
background: #4caf50;
|
| 411 |
+
}
|
| 412 |
+
.settings-dot.active { display: block; }
|
| 413 |
+
.workflow-summary {
|
| 414 |
+
font-size: 12px;
|
| 415 |
+
color: #4caf50;
|
| 416 |
+
margin-top: 6px;
|
| 417 |
+
line-height: 1.5;
|
| 418 |
+
}
|
| 419 |
+
|
| 420 |
+
/* ========= 워크플로우 설정 모달 ========= */
|
| 421 |
+
.modal-overlay {
|
| 422 |
+
position: fixed;
|
| 423 |
+
top: 0; left: 0; right: 0; bottom: 0;
|
| 424 |
+
background: rgba(0,0,0,0.7);
|
| 425 |
+
z-index: 1000;
|
| 426 |
+
display: flex;
|
| 427 |
+
align-items: center;
|
| 428 |
+
justify-content: center;
|
| 429 |
+
backdrop-filter: blur(4px);
|
| 430 |
+
}
|
| 431 |
+
.modal {
|
| 432 |
+
background: #1a1a1a;
|
| 433 |
+
border: 1px solid #333;
|
| 434 |
+
border-radius: 16px;
|
| 435 |
+
padding: 28px;
|
| 436 |
+
width: 540px;
|
| 437 |
+
max-width: 90vw;
|
| 438 |
+
max-height: 85vh;
|
| 439 |
+
overflow-y: auto;
|
| 440 |
+
}
|
| 441 |
+
.modal-header {
|
| 442 |
+
display: flex;
|
| 443 |
+
justify-content: space-between;
|
| 444 |
+
align-items: center;
|
| 445 |
+
margin-bottom: 8px;
|
| 446 |
+
}
|
| 447 |
+
.modal-header h3 { color: #fff; font-size: 18px; margin: 0; }
|
| 448 |
+
.modal-close {
|
| 449 |
+
background: none;
|
| 450 |
+
border: none;
|
| 451 |
+
color: #666;
|
| 452 |
+
font-size: 24px;
|
| 453 |
+
cursor: pointer;
|
| 454 |
+
padding: 4px 8px;
|
| 455 |
+
border-radius: 6px;
|
| 456 |
+
transition: all 0.2s;
|
| 457 |
+
line-height: 1;
|
| 458 |
+
}
|
| 459 |
+
.modal-close:hover { background: #333; color: #fff; }
|
| 460 |
+
.modal-desc {
|
| 461 |
+
color: #888;
|
| 462 |
+
font-size: 13px;
|
| 463 |
+
margin-bottom: 20px;
|
| 464 |
+
line-height: 1.5;
|
| 465 |
+
}
|
| 466 |
+
.modal-step-row {
|
| 467 |
+
border: 1px solid #2a2a2a;
|
| 468 |
+
border-radius: 10px;
|
| 469 |
+
padding: 16px;
|
| 470 |
+
margin-bottom: 12px;
|
| 471 |
+
background: #111;
|
| 472 |
+
}
|
| 473 |
+
.modal-step-label {
|
| 474 |
+
display: flex;
|
| 475 |
+
align-items: center;
|
| 476 |
+
gap: 10px;
|
| 477 |
+
margin-bottom: 10px;
|
| 478 |
+
}
|
| 479 |
+
.modal-step-name {
|
| 480 |
+
font-size: 14px;
|
| 481 |
+
font-weight: 600;
|
| 482 |
+
color: #e1e1e1;
|
| 483 |
+
}
|
| 484 |
+
.modal-step-badge {
|
| 485 |
+
font-size: 11px;
|
| 486 |
+
padding: 2px 8px;
|
| 487 |
+
border-radius: 10px;
|
| 488 |
+
background: #222;
|
| 489 |
+
color: #666;
|
| 490 |
+
}
|
| 491 |
+
.modal-step-badge.custom {
|
| 492 |
+
background: #1a3a1a;
|
| 493 |
+
color: #4caf50;
|
| 494 |
+
}
|
| 495 |
+
.modal-step-fields { display: flex; flex-direction: column; gap: 8px; }
|
| 496 |
+
.modal-step-fields select,
|
| 497 |
+
.modal-step-fields input[type="text"] {
|
| 498 |
+
width: 100%;
|
| 499 |
+
padding: 8px 12px;
|
| 500 |
+
background: #0f0f0f;
|
| 501 |
+
border: 1px solid #333;
|
| 502 |
+
border-radius: 6px;
|
| 503 |
+
color: #fff;
|
| 504 |
+
font-size: 13px;
|
| 505 |
+
outline: none;
|
| 506 |
+
}
|
| 507 |
+
.modal-step-fields select:focus,
|
| 508 |
+
.modal-step-fields input[type="text"]:focus { border-color: #ff4444; }
|
| 509 |
+
.modal-step-sub {
|
| 510 |
+
margin-top: 10px;
|
| 511 |
+
padding-top: 10px;
|
| 512 |
+
border-top: 1px solid #222;
|
| 513 |
+
}
|
| 514 |
+
.modal-step-sub label {
|
| 515 |
+
font-size: 12px;
|
| 516 |
+
color: #888;
|
| 517 |
+
margin-bottom: 4px;
|
| 518 |
+
}
|
| 519 |
+
.modal-step-sub select,
|
| 520 |
+
.modal-step-sub input[type="text"] {
|
| 521 |
+
width: 100%;
|
| 522 |
+
padding: 6px 10px;
|
| 523 |
+
background: #0f0f0f;
|
| 524 |
+
border: 1px solid #333;
|
| 525 |
+
border-radius: 6px;
|
| 526 |
+
color: #fff;
|
| 527 |
+
font-size: 12px;
|
| 528 |
+
outline: none;
|
| 529 |
+
margin-bottom: 6px;
|
| 530 |
+
}
|
| 531 |
+
.modal-save-btn {
|
| 532 |
+
margin-top: 8px;
|
| 533 |
+
padding: 10px;
|
| 534 |
+
font-size: 14px;
|
| 535 |
+
}
|
| 536 |
+
</style>
|
| 537 |
+
</head>
|
| 538 |
+
<body>
|
| 539 |
+
<div class="container">
|
| 540 |
+
<h1>YouTube Script Extractor</h1>
|
| 541 |
+
<p class="subtitle">YouTube 영상의 음성을 텍스트로 변환합니다</p>
|
| 542 |
+
|
| 543 |
+
<div class="card" id="form-section">
|
| 544 |
+
<div class="form-group">
|
| 545 |
+
<label>YouTube URL</label>
|
| 546 |
+
<input type="text" id="url" placeholder="https://www.youtube.com/watch?v=..." onkeydown="if(event.key==='Enter') startTranscription()">
|
| 547 |
+
</div>
|
| 548 |
+
|
| 549 |
+
<div class="row">
|
| 550 |
+
<div class="form-group">
|
| 551 |
+
<label>변환 모드</label>
|
| 552 |
+
<select id="mode" onchange="toggleApiKey()">
|
| 553 |
+
<option value="api">Whisper API (빠르고 정확)</option>
|
| 554 |
+
<option value="local">로컬 Whisper (무료)</option>
|
| 555 |
+
</select>
|
| 556 |
+
</div>
|
| 557 |
+
<div class="form-group" id="model-group" style="display:none;">
|
| 558 |
+
<label>모델 크기</label>
|
| 559 |
+
<select id="model_size">
|
| 560 |
+
<option value="tiny">tiny (39MB, 가장 빠름)</option>
|
| 561 |
+
<option value="base" selected>base (74MB, 균형)</option>
|
| 562 |
+
<option value="small">small (244MB, 좋음)</option>
|
| 563 |
+
<option value="medium">medium (769MB, 높음)</option>
|
| 564 |
+
<option value="large">large (1.5GB, 최고)</option>
|
| 565 |
+
</select>
|
| 566 |
+
</div>
|
| 567 |
+
</div>
|
| 568 |
+
|
| 569 |
+
<div class="form-group" id="api-key-group">
|
| 570 |
+
<label>OpenAI API Key</label>
|
| 571 |
+
<input type="text" id="api_key" placeholder="sk-...">
|
| 572 |
+
<p class="mode-note">서버에 OPENAI_API_KEY 환경변수가 설정되어 있으면 비워둬도 됩니다.</p>
|
| 573 |
+
</div>
|
| 574 |
+
|
| 575 |
+
<div class="form-group">
|
| 576 |
+
<label>출력 형식</label>
|
| 577 |
+
<div class="checkbox-group">
|
| 578 |
+
<label><input type="checkbox" id="fmt_txt" checked> TXT (텍스트)</label>
|
| 579 |
+
<label><input type="checkbox" id="fmt_srt" checked> SRT (자막)</label>
|
| 580 |
+
<label><input type="checkbox" id="fmt_md" onchange="toggleLlmOptions()"> MD (AI 정리)</label>
|
| 581 |
+
<label><input type="checkbox" id="fmt_translate" onchange="toggleLlmOptions()"> 번역</label>
|
| 582 |
+
<label><input type="checkbox" id="fmt_keyframes" onchange="toggleLlmOptions()"> 키프레임 분석</label>
|
| 583 |
+
</div>
|
| 584 |
+
</div>
|
| 585 |
+
|
| 586 |
+
<!-- LLM 옵션 (MD/번역 공용) -->
|
| 587 |
+
<div class="llm-options" id="llm-options">
|
| 588 |
+
<div class="form-group">
|
| 589 |
+
<label>기본 LLM <span style="color:#666; font-weight:400;">(모든 AI 단계에 적용)</span></label>
|
| 590 |
+
<div class="llm-header">
|
| 591 |
+
<div>
|
| 592 |
+
<select id="llm-select" onchange="onLLMChange()">
|
| 593 |
+
<option value="">로딩 중...</option>
|
| 594 |
+
</select>
|
| 595 |
+
</div>
|
| 596 |
+
<button type="button" class="btn-settings" id="btn-settings" onclick="openSettingsModal()" title="워크플로우 설정">⚙<span class="settings-dot" id="settings-dot"></span></button>
|
| 597 |
+
</div>
|
| 598 |
+
<p class="llm-select-desc" id="llm-select-desc"></p>
|
| 599 |
+
<p class="workflow-summary" id="workflow-summary" style="display:none;"></p>
|
| 600 |
+
</div>
|
| 601 |
+
<div class="form-group" id="md-api-key-group" style="display:none;">
|
| 602 |
+
<label id="md-api-key-label">LLM API Key</label>
|
| 603 |
+
<input type="text" id="md_api_key" placeholder="API 키 입력...">
|
| 604 |
+
<p class="mode-note" id="md-api-key-note"></p>
|
| 605 |
+
</div>
|
| 606 |
+
<div class="form-group" id="md-ollama-group" style="display:none;">
|
| 607 |
+
<label>Ollama 모델</label>
|
| 608 |
+
<input type="text" id="md_ollama_model" value="llama3.2" placeholder="llama3.2">
|
| 609 |
+
<p class="mode-note">ollama pull <모델명> 으로 먼저 다운로드 필요</p>
|
| 610 |
+
</div>
|
| 611 |
+
<div class="form-group" id="translate-lang-group" style="display:none;">
|
| 612 |
+
<label>번역 언어</label>
|
| 613 |
+
<select id="translate-lang">
|
| 614 |
+
<option value="">로딩 중...</option>
|
| 615 |
+
</select>
|
| 616 |
+
</div>
|
| 617 |
+
</div>
|
| 618 |
+
|
| 619 |
+
<div class="form-group">
|
| 620 |
+
<label>저장 경로</label>
|
| 621 |
+
<input type="text" id="output_dir" placeholder="./output" value="./output">
|
| 622 |
+
<p class="mode-note">비워두면 기본 경로(./output)에 저장됩니다. 절대 경로도 사용 가능합니다 (예: /Users/이름/Desktop)</p>
|
| 623 |
+
</div>
|
| 624 |
+
|
| 625 |
+
<button class="btn" id="submit-btn" onclick="startTranscription()">스크립트 추출 시작</button>
|
| 626 |
+
</div>
|
| 627 |
+
|
| 628 |
+
<!-- 진행 상태 -->
|
| 629 |
+
<div class="card" id="progress-section">
|
| 630 |
+
<div class="steps">
|
| 631 |
+
<div class="step" id="step-download" data-step="download">
|
| 632 |
+
<div class="step-header">
|
| 633 |
+
<div class="step-icon">1</div>
|
| 634 |
+
<div class="step-title">오디오 다운로드</div>
|
| 635 |
+
<div class="step-percent" id="pct-download"></div>
|
| 636 |
+
</div>
|
| 637 |
+
<div class="step-progress"><div class="step-progress-bar" id="bar-download"></div></div>
|
| 638 |
+
<div class="step-detail" id="detail-download">YouTube에서 오디오를 추출합니다</div>
|
| 639 |
+
</div>
|
| 640 |
+
<div class="step" id="step-keyframe_extract" data-step="keyframe_extract" style="display:none;">
|
| 641 |
+
<div class="step-header">
|
| 642 |
+
<div class="step-icon">2</div>
|
| 643 |
+
<div class="step-title">키프레임 추출</div>
|
| 644 |
+
<div class="step-percent" id="pct-keyframe_extract"></div>
|
| 645 |
+
</div>
|
| 646 |
+
<div class="step-progress"><div class="step-progress-bar" id="bar-keyframe_extract"></div></div>
|
| 647 |
+
<div class="step-detail" id="detail-keyframe_extract">영상에서 주요 장면을 추출합니다</div>
|
| 648 |
+
</div>
|
| 649 |
+
<div class="step" id="step-keyframe_analysis" data-step="keyframe_analysis" style="display:none;">
|
| 650 |
+
<div class="step-header">
|
| 651 |
+
<div class="step-icon">3</div>
|
| 652 |
+
<div class="step-title">키프레임 분석</div>
|
| 653 |
+
<div class="step-percent" id="pct-keyframe_analysis"></div>
|
| 654 |
+
</div>
|
| 655 |
+
<div class="step-progress"><div class="step-progress-bar" id="bar-keyframe_analysis"></div></div>
|
| 656 |
+
<div class="step-detail" id="detail-keyframe_analysis">Vision AI로 화면을 분석합니다</div>
|
| 657 |
+
</div>
|
| 658 |
+
<div class="step" id="step-transcribe" data-step="transcribe">
|
| 659 |
+
<div class="step-header">
|
| 660 |
+
<div class="step-icon">2</div>
|
| 661 |
+
<div class="step-title">음성 인식</div>
|
| 662 |
+
<div class="step-percent" id="pct-transcribe"></div>
|
| 663 |
+
</div>
|
| 664 |
+
<div class="step-progress"><div class="step-progress-bar" id="bar-transcribe"></div></div>
|
| 665 |
+
<div class="step-detail" id="detail-transcribe">Whisper로 음성을 텍스트로 변환합니다</div>
|
| 666 |
+
</div>
|
| 667 |
+
<div class="step" id="step-format" data-step="format" style="display:none;">
|
| 668 |
+
<div class="step-header">
|
| 669 |
+
<div class="step-icon">3</div>
|
| 670 |
+
<div class="step-title">AI 마크다운 정리</div>
|
| 671 |
+
<div class="step-percent" id="pct-format"></div>
|
| 672 |
+
</div>
|
| 673 |
+
<div class="step-progress"><div class="step-progress-bar" id="bar-format"></div></div>
|
| 674 |
+
<div class="step-detail" id="detail-format">LLM으로 텍스트를 구조화합니다</div>
|
| 675 |
+
</div>
|
| 676 |
+
<div class="step" id="step-translate" data-step="translate" style="display:none;">
|
| 677 |
+
<div class="step-header">
|
| 678 |
+
<div class="step-icon">4</div>
|
| 679 |
+
<div class="step-title">번역</div>
|
| 680 |
+
<div class="step-percent" id="pct-translate"></div>
|
| 681 |
+
</div>
|
| 682 |
+
<div class="step-progress"><div class="step-progress-bar" id="bar-translate"></div></div>
|
| 683 |
+
<div class="step-detail" id="detail-translate">선택한 언어로 번역합니다</div>
|
| 684 |
+
</div>
|
| 685 |
+
<div class="step" id="step-save" data-step="save">
|
| 686 |
+
<div class="step-header">
|
| 687 |
+
<div class="step-icon" id="icon-save">3</div>
|
| 688 |
+
<div class="step-title">파일 저장</div>
|
| 689 |
+
<div class="step-percent" id="pct-save"></div>
|
| 690 |
+
</div>
|
| 691 |
+
<div class="step-progress"><div class="step-progress-bar" id="bar-save"></div></div>
|
| 692 |
+
<div class="step-detail" id="detail-save">결과 파일을 생성합니다</div>
|
| 693 |
+
</div>
|
| 694 |
+
</div>
|
| 695 |
+
<div class="progress-footer">
|
| 696 |
+
<div class="elapsed" id="elapsed"></div>
|
| 697 |
+
<div class="tip" id="tip"></div>
|
| 698 |
+
</div>
|
| 699 |
+
</div>
|
| 700 |
+
|
| 701 |
+
<!-- 결과 -->
|
| 702 |
+
<div class="card" id="result-section">
|
| 703 |
+
<div class="result-banner">
|
| 704 |
+
<div class="result-banner-icon">✓</div>
|
| 705 |
+
<div class="result-banner-text">
|
| 706 |
+
<div class="result-banner-title">스크립트 추출 완료</div>
|
| 707 |
+
<div class="result-banner-info" id="result-info"></div>
|
| 708 |
+
</div>
|
| 709 |
+
</div>
|
| 710 |
+
<div class="result-header">
|
| 711 |
+
<span class="result-title" id="result-title"></span>
|
| 712 |
+
<span class="result-lang" id="result-lang"></span>
|
| 713 |
+
</div>
|
| 714 |
+
<div class="result-text" id="result-text"></div>
|
| 715 |
+
<div class="download-buttons" id="download-buttons"></div>
|
| 716 |
+
<div class="save-path" id="save-path"></div>
|
| 717 |
+
<button class="btn-new" onclick="resetForm()">새 영상 추출하기</button>
|
| 718 |
+
</div>
|
| 719 |
+
|
| 720 |
+
<!-- 에러 -->
|
| 721 |
+
<div class="card" id="error-section" style="display:none;">
|
| 722 |
+
<div class="error-msg" id="error-msg"></div>
|
| 723 |
+
<button class="btn-new" onclick="resetForm()">다시 시도하기</button>
|
| 724 |
+
</div>
|
| 725 |
+
</div>
|
| 726 |
+
|
| 727 |
+
<!-- 워크플로우 설정 모달 -->
|
| 728 |
+
<div class="modal-overlay" id="settings-modal" style="display:none;" onclick="if(event.target===this)closeSettingsModal()">
|
| 729 |
+
<div class="modal">
|
| 730 |
+
<div class="modal-header">
|
| 731 |
+
<h3>워크플로우 설정</h3>
|
| 732 |
+
<button class="modal-close" onclick="closeSettingsModal()">×</button>
|
| 733 |
+
</div>
|
| 734 |
+
<p class="modal-desc">각 단계별로 다른 LLM을 지정할 수 있습니다.<br>비워두면 기본 LLM(<strong id="modal-global-llm-name" style="color:#e1e1e1;">—</strong>)을 사용합니다.</p>
|
| 735 |
+
|
| 736 |
+
<!-- MD 정리 -->
|
| 737 |
+
<div class="modal-step-row">
|
| 738 |
+
<div class="modal-step-label">
|
| 739 |
+
<span class="modal-step-name">MD 정리</span>
|
| 740 |
+
<span class="modal-step-badge" id="badge-format">기본 LLM</span>
|
| 741 |
+
</div>
|
| 742 |
+
<div class="modal-step-fields">
|
| 743 |
+
<select id="modal-format-llm" onchange="onModalLLMChange('format')">
|
| 744 |
+
<option value="">기본 LLM 사용</option>
|
| 745 |
+
</select>
|
| 746 |
+
<input type="text" id="modal-format-api-key" placeholder="API Key (비워두면 환경변수 사용)" style="display:none;">
|
| 747 |
+
</div>
|
| 748 |
+
</div>
|
| 749 |
+
|
| 750 |
+
<!-- 번역 -->
|
| 751 |
+
<div class="modal-step-row">
|
| 752 |
+
<div class="modal-step-label">
|
| 753 |
+
<span class="modal-step-name">번역</span>
|
| 754 |
+
<span class="modal-step-badge" id="badge-translate">기본 LLM</span>
|
| 755 |
+
</div>
|
| 756 |
+
<div class="modal-step-fields">
|
| 757 |
+
<select id="modal-translate-llm" onchange="onModalLLMChange('translate')">
|
| 758 |
+
<option value="">기본 LLM 사용</option>
|
| 759 |
+
</select>
|
| 760 |
+
<input type="text" id="modal-translate-api-key" placeholder="API Key (비워두면 환경변수 사용)" style="display:none;">
|
| 761 |
+
</div>
|
| 762 |
+
</div>
|
| 763 |
+
|
| 764 |
+
<!-- 키프레임 분석 -->
|
| 765 |
+
<div class="modal-step-row">
|
| 766 |
+
<div class="modal-step-label">
|
| 767 |
+
<span class="modal-step-name">키프레임 분석</span>
|
| 768 |
+
<span class="modal-step-badge" id="badge-keyframe">기본 LLM</span>
|
| 769 |
+
</div>
|
| 770 |
+
<div class="modal-step-fields">
|
| 771 |
+
<select id="modal-keyframe-llm" onchange="onModalLLMChange('keyframe')">
|
| 772 |
+
<option value="">기본 LLM 사용 (Vision 지원 필요)</option>
|
| 773 |
+
</select>
|
| 774 |
+
<input type="text" id="modal-keyframe-api-key" placeholder="API Key (비워두면 환경변수 사용)" style="display:none;">
|
| 775 |
+
</div>
|
| 776 |
+
<div class="modal-step-sub">
|
| 777 |
+
<label>추출 방식</label>
|
| 778 |
+
<select id="keyframe-method" onchange="onKeyframeMethodChange()">
|
| 779 |
+
<option value="scene">장면 전환 감지 (추천)</option>
|
| 780 |
+
<option value="interval">고정 간격</option>
|
| 781 |
+
</select>
|
| 782 |
+
<div id="keyframe-interval-group" style="display:none;">
|
| 783 |
+
<label>간격 (초)</label>
|
| 784 |
+
<input type="text" id="keyframe-interval" value="30" placeholder="30">
|
| 785 |
+
</div>
|
| 786 |
+
</div>
|
| 787 |
+
</div>
|
| 788 |
+
|
| 789 |
+
<button class="btn modal-save-btn" onclick="saveModalSettings()">설정 저장</button>
|
| 790 |
+
</div>
|
| 791 |
+
</div>
|
| 792 |
+
|
| 793 |
+
<script>
|
| 794 |
+
let timerInterval = null;
|
| 795 |
+
let startTime = null;
|
| 796 |
+
let tipInterval = null;
|
| 797 |
+
let selectedLLM = 'gpt-4o-mini';
|
| 798 |
+
let llmModels = [];
|
| 799 |
+
let languages = [];
|
| 800 |
+
let llmLoaded = false;
|
| 801 |
+
let langLoaded = false;
|
| 802 |
+
let visionModels = [];
|
| 803 |
+
let visionLoaded = false;
|
| 804 |
+
|
| 805 |
+
// 단계별 LLM 설정 (모달에서 관리)
|
| 806 |
+
let perStepConfig = {
|
| 807 |
+
format: { llm: '', apiKey: '' },
|
| 808 |
+
translate: { llm: '', apiKey: '' },
|
| 809 |
+
keyframe: { llm: '', apiKey: '', method: 'scene', interval: 30 },
|
| 810 |
+
};
|
| 811 |
+
|
| 812 |
+
const TIPS = [
|
| 813 |
+
"Whisper는 90개 이상의 언어를 자동으로 감지합니다",
|
| 814 |
+
"SRT 파일은 대부분의 동영상 플레이어에서 자막으로 사용할 수 있습니다",
|
| 815 |
+
"영상이 길수록 변환 시간이 오래 걸립니다",
|
| 816 |
+
"API 모드는 로컬 모드보다 보통 5~10배 빠릅니다",
|
| 817 |
+
"TXT 파일은 텍스트 편집기에서 바로 열 수 있습니다",
|
| 818 |
+
"추출된 스크립트는 ./output 폴더에 저장됩니다",
|
| 819 |
+
"로컬 모드는 인터넷 없이도 사용 가능합니다",
|
| 820 |
+
"Whisper API 비용은 1분당 약 $0.006입니다",
|
| 821 |
+
"MD 출력은 LLM이 텍스트를 깔끔하게 구조화합니다",
|
| 822 |
+
"Ollama는 로컬에서 무료로 LLM을 실행할 수 있습니다",
|
| 823 |
+
"번역 기능으로 원본과 번역본을 동시에 생성할 수 있습니다",
|
| 824 |
+
"키프레임 분석으로 발표 슬라이드나 화면 내용을 MD에 포함할 수 있습니다",
|
| 825 |
+
"워크플로우 설정에서 단계별로 다른 LLM을 지정할 수 있습니다",
|
| 826 |
+
"키프레임 분석에 Gemini Flash를 사용하면 무료입니다",
|
| 827 |
+
];
|
| 828 |
+
|
| 829 |
+
// 다운로드 버튼 라벨 매핑
|
| 830 |
+
const FMT_LABELS = {
|
| 831 |
+
'txt': 'TXT',
|
| 832 |
+
'srt': 'SRT',
|
| 833 |
+
'md': 'MD',
|
| 834 |
+
};
|
| 835 |
+
|
| 836 |
+
let STEP_ORDER = ['download', 'transcribe', 'save'];
|
| 837 |
+
|
| 838 |
+
function getStepOrder() {
|
| 839 |
+
const useMd = document.getElementById('fmt_md').checked;
|
| 840 |
+
const useTranslate = document.getElementById('fmt_translate').checked;
|
| 841 |
+
const useKeyframes = document.getElementById('fmt_keyframes').checked;
|
| 842 |
+
const order = ['download'];
|
| 843 |
+
if (useKeyframes) {
|
| 844 |
+
order.push('keyframe_extract');
|
| 845 |
+
order.push('keyframe_analysis');
|
| 846 |
+
}
|
| 847 |
+
order.push('transcribe');
|
| 848 |
+
if (useMd) order.push('format');
|
| 849 |
+
if (useTranslate) order.push('translate');
|
| 850 |
+
order.push('save');
|
| 851 |
+
return order;
|
| 852 |
+
}
|
| 853 |
+
|
| 854 |
+
function needsLLM() {
|
| 855 |
+
return document.getElementById('fmt_md').checked ||
|
| 856 |
+
document.getElementById('fmt_translate').checked ||
|
| 857 |
+
document.getElementById('fmt_keyframes').checked;
|
| 858 |
+
}
|
| 859 |
+
|
| 860 |
+
function toggleApiKey() {
|
| 861 |
+
const mode = document.getElementById('mode').value;
|
| 862 |
+
document.getElementById('api-key-group').style.display = mode === 'api' ? '' : 'none';
|
| 863 |
+
document.getElementById('model-group').style.display = mode === 'local' ? '' : 'none';
|
| 864 |
+
}
|
| 865 |
+
|
| 866 |
+
function toggleLlmOptions() {
|
| 867 |
+
const opts = document.getElementById('llm-options');
|
| 868 |
+
const translateChecked = document.getElementById('fmt_translate').checked;
|
| 869 |
+
const keyframesChecked = document.getElementById('fmt_keyframes').checked;
|
| 870 |
+
|
| 871 |
+
if (needsLLM()) {
|
| 872 |
+
opts.classList.add('visible');
|
| 873 |
+
if (!llmLoaded) loadLLMModels();
|
| 874 |
+
if (translateChecked && !langLoaded) loadLanguages();
|
| 875 |
+
if (keyframesChecked && !visionLoaded) loadVisionModels();
|
| 876 |
+
} else {
|
| 877 |
+
opts.classList.remove('visible');
|
| 878 |
+
}
|
| 879 |
+
|
| 880 |
+
// 번역 언어 선택 표시/숨기기
|
| 881 |
+
document.getElementById('translate-lang-group').style.display =
|
| 882 |
+
translateChecked ? '' : 'none';
|
| 883 |
+
}
|
| 884 |
+
|
| 885 |
+
async function loadLLMModels() {
|
| 886 |
+
try {
|
| 887 |
+
const res = await fetch('/api/llm-models?sort=price');
|
| 888 |
+
const data = await res.json();
|
| 889 |
+
llmModels = data.models;
|
| 890 |
+
llmLoaded = true;
|
| 891 |
+
renderLLMDropdown();
|
| 892 |
+
} catch (e) {
|
| 893 |
+
console.error('LLM 모델 로딩 실패:', e);
|
| 894 |
+
}
|
| 895 |
+
}
|
| 896 |
+
|
| 897 |
+
async function loadLanguages() {
|
| 898 |
+
try {
|
| 899 |
+
const res = await fetch('/api/languages');
|
| 900 |
+
const data = await res.json();
|
| 901 |
+
languages = data.languages;
|
| 902 |
+
langLoaded = true;
|
| 903 |
+
renderLanguageDropdown();
|
| 904 |
+
} catch (e) {
|
| 905 |
+
console.error('언어 목록 로딩 실패:', e);
|
| 906 |
+
}
|
| 907 |
+
}
|
| 908 |
+
|
| 909 |
+
function renderLLMDropdown() {
|
| 910 |
+
const select = document.getElementById('llm-select');
|
| 911 |
+
select.innerHTML = '';
|
| 912 |
+
llmModels.forEach(model => {
|
| 913 |
+
const opt = document.createElement('option');
|
| 914 |
+
opt.value = model.id;
|
| 915 |
+
opt.textContent = model.version
|
| 916 |
+
? `${model.name} [${model.version}]`
|
| 917 |
+
: model.name;
|
| 918 |
+
if (model.id === selectedLLM) opt.selected = true;
|
| 919 |
+
select.appendChild(opt);
|
| 920 |
+
});
|
| 921 |
+
onLLMChange();
|
| 922 |
+
}
|
| 923 |
+
|
| 924 |
+
function renderLanguageDropdown() {
|
| 925 |
+
const select = document.getElementById('translate-lang');
|
| 926 |
+
select.innerHTML = '';
|
| 927 |
+
languages.forEach(lang => {
|
| 928 |
+
const opt = document.createElement('option');
|
| 929 |
+
opt.value = lang.code;
|
| 930 |
+
opt.textContent = `${lang.name} (${lang.code})`;
|
| 931 |
+
select.appendChild(opt);
|
| 932 |
+
});
|
| 933 |
+
}
|
| 934 |
+
|
| 935 |
+
function onLLMChange() {
|
| 936 |
+
const select = document.getElementById('llm-select');
|
| 937 |
+
selectedLLM = select.value;
|
| 938 |
+
const model = llmModels.find(m => m.id === selectedLLM);
|
| 939 |
+
if (!model) return;
|
| 940 |
+
|
| 941 |
+
// 설명 텍스트 업데이트
|
| 942 |
+
document.getElementById('llm-select-desc').textContent = model.description;
|
| 943 |
+
|
| 944 |
+
// API 키 / Ollama 모델 UI 업데이트
|
| 945 |
+
const keyGroup = document.getElementById('md-api-key-group');
|
| 946 |
+
const ollamaGroup = document.getElementById('md-ollama-group');
|
| 947 |
+
|
| 948 |
+
if (selectedLLM === 'ollama') {
|
| 949 |
+
keyGroup.style.display = 'none';
|
| 950 |
+
ollamaGroup.style.display = '';
|
| 951 |
+
} else if (model.needs_key) {
|
| 952 |
+
keyGroup.style.display = '';
|
| 953 |
+
ollamaGroup.style.display = 'none';
|
| 954 |
+
const envKey = model.env_key || '';
|
| 955 |
+
document.getElementById('md-api-key-label').textContent = model.name + ' API Key';
|
| 956 |
+
document.getElementById('md-api-key-note').textContent =
|
| 957 |
+
envKey ? `서버에 ${envKey} 환경변수가 설정되어 있으면 비워둬도 됩니다.` : '';
|
| 958 |
+
} else {
|
| 959 |
+
keyGroup.style.display = 'none';
|
| 960 |
+
ollamaGroup.style.display = 'none';
|
| 961 |
+
}
|
| 962 |
+
|
| 963 |
+
// 워크플로우 요약 업데이트 (기본 LLM 변경 시 "기본 LLM 사용" 뱃지도 갱신)
|
| 964 |
+
updateWorkflowSummary();
|
| 965 |
+
}
|
| 966 |
+
|
| 967 |
+
function startTimer() {
|
| 968 |
+
startTime = Date.now();
|
| 969 |
+
timerInterval = setInterval(() => {
|
| 970 |
+
const sec = Math.floor((Date.now() - startTime) / 1000);
|
| 971 |
+
const min = Math.floor(sec / 60);
|
| 972 |
+
const s = sec % 60;
|
| 973 |
+
document.getElementById('elapsed').textContent =
|
| 974 |
+
min > 0 ? `${min}분 ${s}초 경과` : `${s}초 경과`;
|
| 975 |
+
}, 1000);
|
| 976 |
+
}
|
| 977 |
+
|
| 978 |
+
function stopTimer() {
|
| 979 |
+
if (timerInterval) clearInterval(timerInterval);
|
| 980 |
+
if (tipInterval) clearInterval(tipInterval);
|
| 981 |
+
}
|
| 982 |
+
|
| 983 |
+
function startTips() {
|
| 984 |
+
let tipIdx = Math.floor(Math.random() * TIPS.length);
|
| 985 |
+
const tipEl = document.getElementById('tip');
|
| 986 |
+
tipEl.textContent = TIPS[tipIdx];
|
| 987 |
+
tipInterval = setInterval(() => {
|
| 988 |
+
tipIdx = (tipIdx + 1) % TIPS.length;
|
| 989 |
+
tipEl.style.opacity = '0';
|
| 990 |
+
setTimeout(() => {
|
| 991 |
+
tipEl.textContent = TIPS[tipIdx];
|
| 992 |
+
tipEl.style.opacity = '';
|
| 993 |
+
}, 300);
|
| 994 |
+
}, 6000);
|
| 995 |
+
}
|
| 996 |
+
|
| 997 |
+
function updateStep(step, percent, detail) {
|
| 998 |
+
const stepEl = document.getElementById('step-' + step);
|
| 999 |
+
const barEl = document.getElementById('bar-' + step);
|
| 1000 |
+
const pctEl = document.getElementById('pct-' + step);
|
| 1001 |
+
const detailEl = document.getElementById('detail-' + step);
|
| 1002 |
+
|
| 1003 |
+
if (!stepEl) return;
|
| 1004 |
+
|
| 1005 |
+
// 현재 스텝을 active로
|
| 1006 |
+
const stepIdx = STEP_ORDER.indexOf(step);
|
| 1007 |
+
STEP_ORDER.forEach((s, i) => {
|
| 1008 |
+
const el = document.getElementById('step-' + s);
|
| 1009 |
+
if (!el) return;
|
| 1010 |
+
const icon = el.querySelector('.step-icon');
|
| 1011 |
+
if (i < stepIdx) {
|
| 1012 |
+
el.className = 'step done';
|
| 1013 |
+
icon.innerHTML = '✓';
|
| 1014 |
+
} else if (i === stepIdx) {
|
| 1015 |
+
el.className = 'step active';
|
| 1016 |
+
icon.textContent = String(i + 1);
|
| 1017 |
+
}
|
| 1018 |
+
});
|
| 1019 |
+
|
| 1020 |
+
// 프로그레스 바
|
| 1021 |
+
if (percent < 0) {
|
| 1022 |
+
barEl.className = 'step-progress-bar indeterminate';
|
| 1023 |
+
barEl.style.width = '';
|
| 1024 |
+
pctEl.textContent = '';
|
| 1025 |
+
} else {
|
| 1026 |
+
barEl.className = 'step-progress-bar';
|
| 1027 |
+
barEl.style.width = percent + '%';
|
| 1028 |
+
pctEl.textContent = percent + '%';
|
| 1029 |
+
}
|
| 1030 |
+
|
| 1031 |
+
if (detail) {
|
| 1032 |
+
detailEl.textContent = detail;
|
| 1033 |
+
}
|
| 1034 |
+
|
| 1035 |
+
if (percent >= 100) {
|
| 1036 |
+
stepEl.className = 'step done';
|
| 1037 |
+
stepEl.querySelector('.step-icon').innerHTML = '✓';
|
| 1038 |
+
pctEl.textContent = '100%';
|
| 1039 |
+
}
|
| 1040 |
+
}
|
| 1041 |
+
|
| 1042 |
+
function setAllStepsDone() {
|
| 1043 |
+
STEP_ORDER.forEach(step => {
|
| 1044 |
+
const el = document.getElementById('step-' + step);
|
| 1045 |
+
if (!el) return;
|
| 1046 |
+
el.className = 'step done';
|
| 1047 |
+
el.querySelector('.step-icon').innerHTML = '✓';
|
| 1048 |
+
document.getElementById('bar-' + step).style.width = '100%';
|
| 1049 |
+
document.getElementById('bar-' + step).className = 'step-progress-bar';
|
| 1050 |
+
document.getElementById('pct-' + step).textContent = '100%';
|
| 1051 |
+
});
|
| 1052 |
+
}
|
| 1053 |
+
|
| 1054 |
+
function resetForm() {
|
| 1055 |
+
document.getElementById('form-section').style.display = 'block';
|
| 1056 |
+
document.getElementById('progress-section').style.display = 'none';
|
| 1057 |
+
document.getElementById('result-section').style.display = 'none';
|
| 1058 |
+
document.getElementById('error-section').style.display = 'none';
|
| 1059 |
+
document.getElementById('submit-btn').disabled = false;
|
| 1060 |
+
document.getElementById('submit-btn').textContent = '스크립트 추출 시작';
|
| 1061 |
+
stopTimer();
|
| 1062 |
+
document.getElementById('elapsed').textContent = '';
|
| 1063 |
+
document.getElementById('tip').textContent = '';
|
| 1064 |
+
|
| 1065 |
+
// 선택적 스텝 숨기기
|
| 1066 |
+
document.getElementById('step-format').style.display = 'none';
|
| 1067 |
+
document.getElementById('step-translate').style.display = 'none';
|
| 1068 |
+
document.getElementById('step-keyframe_extract').style.display = 'none';
|
| 1069 |
+
document.getElementById('step-keyframe_analysis').style.display = 'none';
|
| 1070 |
+
|
| 1071 |
+
const allSteps = ['download', 'keyframe_extract', 'keyframe_analysis', 'transcribe', 'format', 'translate', 'save'];
|
| 1072 |
+
allSteps.forEach((step, i) => {
|
| 1073 |
+
const el = document.getElementById('step-' + step);
|
| 1074 |
+
if (!el) return;
|
| 1075 |
+
el.className = 'step';
|
| 1076 |
+
el.querySelector('.step-icon').textContent = String(i + 1);
|
| 1077 |
+
document.getElementById('bar-' + step).style.width = '0%';
|
| 1078 |
+
document.getElementById('bar-' + step).className = 'step-progress-bar';
|
| 1079 |
+
document.getElementById('pct-' + step).textContent = '';
|
| 1080 |
+
});
|
| 1081 |
+
document.getElementById('detail-download').textContent = 'YouTube에서 오디오를 추출합니다';
|
| 1082 |
+
document.getElementById('detail-keyframe_extract').textContent = '영상에서 주요 장면을 추출합니다';
|
| 1083 |
+
document.getElementById('detail-keyframe_analysis').textContent = 'Vision AI로 화면을 분석합니다';
|
| 1084 |
+
document.getElementById('detail-transcribe').textContent = 'Whisper로 음성을 텍스트로 변환합니다';
|
| 1085 |
+
document.getElementById('detail-format').textContent = 'LLM으로 텍스트를 구조화합니다';
|
| 1086 |
+
document.getElementById('detail-translate').textContent = '선택한 언어로 번역합니다';
|
| 1087 |
+
document.getElementById('detail-save').textContent = '결과 파일을 생성합니다';
|
| 1088 |
+
}
|
| 1089 |
+
|
| 1090 |
+
function getDownloadLabel(fmt) {
|
| 1091 |
+
// txt_ko → "TXT (한국어)", md_zhcn → "MD (中文(简体))"
|
| 1092 |
+
if (fmt.includes('_')) {
|
| 1093 |
+
const parts = fmt.split('_');
|
| 1094 |
+
const baseFmt = parts[0].toUpperCase();
|
| 1095 |
+
const langCode = parts.slice(1).join('_');
|
| 1096 |
+
// 언어 코드를 원래 형태로 복원 시도 (zhcn → zh-CN 등)
|
| 1097 |
+
const langEntry = languages.find(l =>
|
| 1098 |
+
l.code.replace('-', '').toLowerCase() === langCode.toLowerCase()
|
| 1099 |
+
);
|
| 1100 |
+
const langName = langEntry ? langEntry.name : langCode.toUpperCase();
|
| 1101 |
+
return `${baseFmt} (${langName})`;
|
| 1102 |
+
}
|
| 1103 |
+
return (FMT_LABELS[fmt] || fmt.toUpperCase());
|
| 1104 |
+
}
|
| 1105 |
+
|
| 1106 |
+
async function startTranscription() {
|
| 1107 |
+
const url = document.getElementById('url').value.trim();
|
| 1108 |
+
if (!url) { alert('YouTube URL을 입력해주세요.'); return; }
|
| 1109 |
+
|
| 1110 |
+
const mode = document.getElementById('mode').value;
|
| 1111 |
+
const apiKey = document.getElementById('api_key').value.trim();
|
| 1112 |
+
const modelSize = document.getElementById('model_size').value;
|
| 1113 |
+
const outputDir = document.getElementById('output_dir').value.trim() || './output';
|
| 1114 |
+
const formats = [];
|
| 1115 |
+
if (document.getElementById('fmt_txt').checked) formats.push('txt');
|
| 1116 |
+
if (document.getElementById('fmt_srt').checked) formats.push('srt');
|
| 1117 |
+
if (document.getElementById('fmt_md').checked) formats.push('md');
|
| 1118 |
+
|
| 1119 |
+
const useTranslate = document.getElementById('fmt_translate').checked;
|
| 1120 |
+
const translateLang = useTranslate ? document.getElementById('translate-lang').value : '';
|
| 1121 |
+
const enableKeyframes = document.getElementById('fmt_keyframes').checked;
|
| 1122 |
+
|
| 1123 |
+
if (formats.length === 0 && !useTranslate) {
|
| 1124 |
+
alert('출력 형식을 하나 이상 선택해주세요.');
|
| 1125 |
+
return;
|
| 1126 |
+
}
|
| 1127 |
+
|
| 1128 |
+
// LLM 파라미터 (MD 또는 번역 사용 시)
|
| 1129 |
+
const useLlm = needsLLM();
|
| 1130 |
+
const mdLlm = useLlm ? selectedLLM : '';
|
| 1131 |
+
const mdApiKey = useLlm ? document.getElementById('md_api_key').value.trim() : '';
|
| 1132 |
+
const mdOllamaModel = useLlm ? document.getElementById('md_ollama_model').value.trim() || 'llama3.2' : 'llama3.2';
|
| 1133 |
+
|
| 1134 |
+
// STEP_ORDER를 동적으로 설정
|
| 1135 |
+
STEP_ORDER = getStepOrder();
|
| 1136 |
+
|
| 1137 |
+
// 선택적 스텝 표시/숨기기 + 아이콘 번호 조정
|
| 1138 |
+
const formatStep = document.getElementById('step-format');
|
| 1139 |
+
const translateStep = document.getElementById('step-translate');
|
| 1140 |
+
const keyExtStep = document.getElementById('step-keyframe_extract');
|
| 1141 |
+
const keyAnalStep = document.getElementById('step-keyframe_analysis');
|
| 1142 |
+
const useMd = formats.includes('md');
|
| 1143 |
+
|
| 1144 |
+
formatStep.style.display = useMd ? '' : 'none';
|
| 1145 |
+
translateStep.style.display = useTranslate ? '' : 'none';
|
| 1146 |
+
keyExtStep.style.display = enableKeyframes ? '' : 'none';
|
| 1147 |
+
keyAnalStep.style.display = enableKeyframes ? '' : 'none';
|
| 1148 |
+
|
| 1149 |
+
// 스텝 번호 재배정
|
| 1150 |
+
STEP_ORDER.forEach((s, i) => {
|
| 1151 |
+
const el = document.getElementById('step-' + s);
|
| 1152 |
+
if (el) el.querySelector('.step-icon').textContent = String(i + 1);
|
| 1153 |
+
});
|
| 1154 |
+
|
| 1155 |
+
const btn = document.getElementById('submit-btn');
|
| 1156 |
+
btn.disabled = true;
|
| 1157 |
+
btn.textContent = '처리 중...';
|
| 1158 |
+
|
| 1159 |
+
document.getElementById('form-section').style.display = 'none';
|
| 1160 |
+
document.getElementById('progress-section').style.display = 'block';
|
| 1161 |
+
document.getElementById('result-section').style.display = 'none';
|
| 1162 |
+
document.getElementById('error-section').style.display = 'none';
|
| 1163 |
+
|
| 1164 |
+
startTimer();
|
| 1165 |
+
startTips();
|
| 1166 |
+
|
| 1167 |
+
try {
|
| 1168 |
+
const res = await fetch('/api/transcribe', {
|
| 1169 |
+
method: 'POST',
|
| 1170 |
+
headers: { 'Content-Type': 'application/json' },
|
| 1171 |
+
body: JSON.stringify({
|
| 1172 |
+
url, mode, api_key: apiKey, model_size: modelSize,
|
| 1173 |
+
formats, output_dir: outputDir,
|
| 1174 |
+
md_llm: mdLlm, md_api_key: mdApiKey, md_ollama_model: mdOllamaModel,
|
| 1175 |
+
translate_lang: translateLang,
|
| 1176 |
+
// Per-step LLM overrides (모달 설정)
|
| 1177 |
+
format_llm: perStepConfig.format.llm,
|
| 1178 |
+
format_api_key: perStepConfig.format.apiKey,
|
| 1179 |
+
translate_llm: perStepConfig.translate.llm,
|
| 1180 |
+
translate_api_key: perStepConfig.translate.apiKey,
|
| 1181 |
+
keyframe_llm: perStepConfig.keyframe.llm,
|
| 1182 |
+
keyframe_api_key: perStepConfig.keyframe.apiKey,
|
| 1183 |
+
// 키프레임 옵션
|
| 1184 |
+
enable_keyframes: enableKeyframes,
|
| 1185 |
+
keyframe_method: perStepConfig.keyframe.method,
|
| 1186 |
+
keyframe_interval: parseInt(perStepConfig.keyframe.interval) || 30,
|
| 1187 |
+
}),
|
| 1188 |
+
});
|
| 1189 |
+
const data = await res.json();
|
| 1190 |
+
|
| 1191 |
+
if (data.error) {
|
| 1192 |
+
showError(data.error);
|
| 1193 |
+
return;
|
| 1194 |
+
}
|
| 1195 |
+
|
| 1196 |
+
const jobId = data.job_id;
|
| 1197 |
+
const eventSource = new EventSource(`/api/stream/${jobId}`);
|
| 1198 |
+
|
| 1199 |
+
eventSource.onmessage = (event) => {
|
| 1200 |
+
const msg = JSON.parse(event.data);
|
| 1201 |
+
|
| 1202 |
+
if (msg.type === 'progress') {
|
| 1203 |
+
if (msg.step && msg.percent !== undefined) {
|
| 1204 |
+
updateStep(msg.step, msg.percent, msg.detail || '');
|
| 1205 |
+
}
|
| 1206 |
+
} else if (msg.type === 'completed') {
|
| 1207 |
+
eventSource.close();
|
| 1208 |
+
stopTimer();
|
| 1209 |
+
setAllStepsDone();
|
| 1210 |
+
setTimeout(() => showResult(msg.result), 500);
|
| 1211 |
+
} else if (msg.type === 'error') {
|
| 1212 |
+
eventSource.close();
|
| 1213 |
+
stopTimer();
|
| 1214 |
+
showError(msg.message);
|
| 1215 |
+
}
|
| 1216 |
+
};
|
| 1217 |
+
|
| 1218 |
+
eventSource.onerror = () => {
|
| 1219 |
+
eventSource.close();
|
| 1220 |
+
stopTimer();
|
| 1221 |
+
showError('서버 연결이 끊어졌습니다. 서버가 실행 중인지 확인해주세요.');
|
| 1222 |
+
};
|
| 1223 |
+
} catch (err) {
|
| 1224 |
+
stopTimer();
|
| 1225 |
+
showError(`요청 실패: ${err.message}`);
|
| 1226 |
+
}
|
| 1227 |
+
}
|
| 1228 |
+
|
| 1229 |
+
function showError(message) {
|
| 1230 |
+
document.getElementById('progress-section').style.display = 'none';
|
| 1231 |
+
document.getElementById('error-section').style.display = 'block';
|
| 1232 |
+
document.getElementById('error-msg').textContent = message;
|
| 1233 |
+
}
|
| 1234 |
+
|
| 1235 |
+
function showResult(result) {
|
| 1236 |
+
document.getElementById('progress-section').style.display = 'none';
|
| 1237 |
+
document.getElementById('result-section').style.display = 'block';
|
| 1238 |
+
|
| 1239 |
+
const elapsed = Math.floor((Date.now() - startTime) / 1000);
|
| 1240 |
+
const min = Math.floor(elapsed / 60);
|
| 1241 |
+
const sec = elapsed % 60;
|
| 1242 |
+
const timeStr = min > 0 ? `${min}분 ${sec}초` : `${sec}초`;
|
| 1243 |
+
|
| 1244 |
+
document.getElementById('result-info').textContent =
|
| 1245 |
+
`감지 언어: ${result.language} | 소요 시간: ${timeStr}`;
|
| 1246 |
+
document.getElementById('result-title').textContent = result.title;
|
| 1247 |
+
document.getElementById('result-lang').textContent = result.language;
|
| 1248 |
+
document.getElementById('result-text').textContent = result.text;
|
| 1249 |
+
|
| 1250 |
+
const downloadBtns = document.getElementById('download-buttons');
|
| 1251 |
+
downloadBtns.innerHTML = '';
|
| 1252 |
+
const filenames = [];
|
| 1253 |
+
const dirParam = result.output_dir_raw ? `?dir=${encodeURIComponent(result.output_dir_raw)}` : '';
|
| 1254 |
+
for (const [fmt, filename] of Object.entries(result.files)) {
|
| 1255 |
+
filenames.push(filename);
|
| 1256 |
+
const a = document.createElement('a');
|
| 1257 |
+
a.href = `/api/download/${encodeURIComponent(filename)}${dirParam}`;
|
| 1258 |
+
a.className = 'btn-download';
|
| 1259 |
+
a.download = filename;
|
| 1260 |
+
a.textContent = `${getDownloadLabel(fmt)} 다운로드`;
|
| 1261 |
+
downloadBtns.appendChild(a);
|
| 1262 |
+
}
|
| 1263 |
+
|
| 1264 |
+
const dir = result.output_dir || './output';
|
| 1265 |
+
document.getElementById('save-path').textContent =
|
| 1266 |
+
'저장 위치: ' + filenames.map(f => dir + '/' + f).join(', ');
|
| 1267 |
+
}
|
| 1268 |
+
// ========= Vision 모델 로딩 =========
|
| 1269 |
+
async function loadVisionModels() {
|
| 1270 |
+
try {
|
| 1271 |
+
const res = await fetch('/api/vision-models?sort=price');
|
| 1272 |
+
const data = await res.json();
|
| 1273 |
+
visionModels = data.models;
|
| 1274 |
+
visionLoaded = true;
|
| 1275 |
+
} catch (e) {
|
| 1276 |
+
console.error('Vision 모델 로딩 실패:', e);
|
| 1277 |
+
}
|
| 1278 |
+
}
|
| 1279 |
+
|
| 1280 |
+
// ========= 워크플로우 설정 모달 =========
|
| 1281 |
+
function openSettingsModal() {
|
| 1282 |
+
// 모달 열기 전 모델 데이터 로딩 확인
|
| 1283 |
+
if (!llmLoaded) loadLLMModels().then(() => populateModalDropdowns());
|
| 1284 |
+
if (!visionLoaded) loadVisionModels().then(() => populateModalDropdowns());
|
| 1285 |
+
populateModalDropdowns();
|
| 1286 |
+
restoreModalFromConfig();
|
| 1287 |
+
// 기본 LLM 이름 표시
|
| 1288 |
+
const globalModel = llmModels.find(m => m.id === selectedLLM);
|
| 1289 |
+
document.getElementById('modal-global-llm-name').textContent =
|
| 1290 |
+
globalModel ? globalModel.name : selectedLLM;
|
| 1291 |
+
document.getElementById('settings-modal').style.display = 'flex';
|
| 1292 |
+
}
|
| 1293 |
+
|
| 1294 |
+
function closeSettingsModal() {
|
| 1295 |
+
document.getElementById('settings-modal').style.display = 'none';
|
| 1296 |
+
}
|
| 1297 |
+
|
| 1298 |
+
function populateModalDropdowns() {
|
| 1299 |
+
const globalModel = llmModels.find(m => m.id === selectedLLM);
|
| 1300 |
+
const globalName = globalModel ? globalModel.name : '기본 LLM';
|
| 1301 |
+
|
| 1302 |
+
// MD 정리 / 번역 드롭다운 (전체 LLM)
|
| 1303 |
+
['format', 'translate'].forEach(step => {
|
| 1304 |
+
const select = document.getElementById(`modal-${step}-llm`);
|
| 1305 |
+
const currentVal = select.value;
|
| 1306 |
+
select.innerHTML = `<option value="">기본 LLM 사용 (${globalName})</option>`;
|
| 1307 |
+
llmModels.forEach(model => {
|
| 1308 |
+
const opt = document.createElement('option');
|
| 1309 |
+
opt.value = model.id;
|
| 1310 |
+
opt.textContent = model.version
|
| 1311 |
+
? `${model.name} [${model.version}]`
|
| 1312 |
+
: model.name;
|
| 1313 |
+
select.appendChild(opt);
|
| 1314 |
+
});
|
| 1315 |
+
select.value = currentVal || '';
|
| 1316 |
+
});
|
| 1317 |
+
|
| 1318 |
+
// 키프레임 드롭다운 (Vision 지원 모델만)
|
| 1319 |
+
const kfSelect = document.getElementById('modal-keyframe-llm');
|
| 1320 |
+
const kfVal = kfSelect.value;
|
| 1321 |
+
const isGlobalVision = globalModel && globalModel.supports_vision !== false;
|
| 1322 |
+
kfSelect.innerHTML = isGlobalVision
|
| 1323 |
+
? `<option value="">기본 LLM 사용 (${globalName})</option>`
|
| 1324 |
+
: `<option value="">선택 필요 (기본 LLM이 Vision 미지원)</option>`;
|
| 1325 |
+
visionModels.forEach(model => {
|
| 1326 |
+
const opt = document.createElement('option');
|
| 1327 |
+
opt.value = model.id;
|
| 1328 |
+
opt.textContent = model.version
|
| 1329 |
+
? `${model.name} [${model.version}]`
|
| 1330 |
+
: model.name;
|
| 1331 |
+
kfSelect.appendChild(opt);
|
| 1332 |
+
});
|
| 1333 |
+
kfSelect.value = kfVal || '';
|
| 1334 |
+
|
| 1335 |
+
updateModalBadges();
|
| 1336 |
+
}
|
| 1337 |
+
|
| 1338 |
+
function restoreModalFromConfig() {
|
| 1339 |
+
document.getElementById('modal-format-llm').value = perStepConfig.format.llm;
|
| 1340 |
+
document.getElementById('modal-format-api-key').value = perStepConfig.format.apiKey;
|
| 1341 |
+
document.getElementById('modal-translate-llm').value = perStepConfig.translate.llm;
|
| 1342 |
+
document.getElementById('modal-translate-api-key').value = perStepConfig.translate.apiKey;
|
| 1343 |
+
document.getElementById('modal-keyframe-llm').value = perStepConfig.keyframe.llm;
|
| 1344 |
+
document.getElementById('modal-keyframe-api-key').value = perStepConfig.keyframe.apiKey;
|
| 1345 |
+
document.getElementById('keyframe-method').value = perStepConfig.keyframe.method;
|
| 1346 |
+
document.getElementById('keyframe-interval').value = perStepConfig.keyframe.interval;
|
| 1347 |
+
|
| 1348 |
+
onKeyframeMethodChange();
|
| 1349 |
+
['format', 'translate', 'keyframe'].forEach(s => onModalLLMChange(s));
|
| 1350 |
+
updateModalBadges();
|
| 1351 |
+
}
|
| 1352 |
+
|
| 1353 |
+
function onModalLLMChange(step) {
|
| 1354 |
+
const select = document.getElementById(`modal-${step}-llm`);
|
| 1355 |
+
const apiKeyInput = document.getElementById(`modal-${step}-api-key`);
|
| 1356 |
+
const val = select.value;
|
| 1357 |
+
|
| 1358 |
+
if (val && val !== 'ollama') {
|
| 1359 |
+
// 선택한 모델의 needs_key 확인
|
| 1360 |
+
const allModels = step === 'keyframe' ? visionModels : llmModels;
|
| 1361 |
+
const model = allModels.find(m => m.id === val);
|
| 1362 |
+
if (model && model.needs_key) {
|
| 1363 |
+
apiKeyInput.style.display = '';
|
| 1364 |
+
apiKeyInput.placeholder = model.env_key
|
| 1365 |
+
? `API Key (${model.env_key} 환경변수가 있으면 생략 가능)`
|
| 1366 |
+
: 'API Key 입력...';
|
| 1367 |
+
} else {
|
| 1368 |
+
apiKeyInput.style.display = 'none';
|
| 1369 |
+
}
|
| 1370 |
+
} else {
|
| 1371 |
+
apiKeyInput.style.display = 'none';
|
| 1372 |
+
}
|
| 1373 |
+
updateModalBadges();
|
| 1374 |
+
}
|
| 1375 |
+
|
| 1376 |
+
function updateModalBadges() {
|
| 1377 |
+
const globalModel = llmModels.find(m => m.id === selectedLLM);
|
| 1378 |
+
const globalName = globalModel ? globalModel.name : '기본';
|
| 1379 |
+
|
| 1380 |
+
['format', 'translate', 'keyframe'].forEach(step => {
|
| 1381 |
+
const select = document.getElementById(`modal-${step}-llm`);
|
| 1382 |
+
const badge = document.getElementById(`badge-${step}`);
|
| 1383 |
+
if (select.value) {
|
| 1384 |
+
const allModels = step === 'keyframe' ? visionModels : llmModels;
|
| 1385 |
+
const model = allModels.find(m => m.id === select.value);
|
| 1386 |
+
badge.textContent = model ? model.name : select.value;
|
| 1387 |
+
badge.className = 'modal-step-badge custom';
|
| 1388 |
+
} else {
|
| 1389 |
+
badge.textContent = globalName;
|
| 1390 |
+
badge.className = 'modal-step-badge';
|
| 1391 |
+
}
|
| 1392 |
+
});
|
| 1393 |
+
}
|
| 1394 |
+
|
| 1395 |
+
function onKeyframeMethodChange() {
|
| 1396 |
+
const method = document.getElementById('keyframe-method').value;
|
| 1397 |
+
document.getElementById('keyframe-interval-group').style.display =
|
| 1398 |
+
method === 'interval' ? '' : 'none';
|
| 1399 |
+
}
|
| 1400 |
+
|
| 1401 |
+
function saveModalSettings() {
|
| 1402 |
+
perStepConfig.format.llm = document.getElementById('modal-format-llm').value;
|
| 1403 |
+
perStepConfig.format.apiKey = document.getElementById('modal-format-api-key').value.trim();
|
| 1404 |
+
perStepConfig.translate.llm = document.getElementById('modal-translate-llm').value;
|
| 1405 |
+
perStepConfig.translate.apiKey = document.getElementById('modal-translate-api-key').value.trim();
|
| 1406 |
+
perStepConfig.keyframe.llm = document.getElementById('modal-keyframe-llm').value;
|
| 1407 |
+
perStepConfig.keyframe.apiKey = document.getElementById('modal-keyframe-api-key').value.trim();
|
| 1408 |
+
perStepConfig.keyframe.method = document.getElementById('keyframe-method').value;
|
| 1409 |
+
perStepConfig.keyframe.interval = parseInt(document.getElementById('keyframe-interval').value) || 30;
|
| 1410 |
+
|
| 1411 |
+
updateWorkflowSummary();
|
| 1412 |
+
closeSettingsModal();
|
| 1413 |
+
}
|
| 1414 |
+
|
| 1415 |
+
function updateWorkflowSummary() {
|
| 1416 |
+
const dot = document.getElementById('settings-dot');
|
| 1417 |
+
const summary = document.getElementById('workflow-summary');
|
| 1418 |
+
|
| 1419 |
+
// 오버라이드가 있는지 확인
|
| 1420 |
+
const overrides = [];
|
| 1421 |
+
const getModelName = (id, models) => {
|
| 1422 |
+
const m = models.find(x => x.id === id);
|
| 1423 |
+
return m ? m.name : id;
|
| 1424 |
+
};
|
| 1425 |
+
|
| 1426 |
+
if (perStepConfig.format.llm) {
|
| 1427 |
+
overrides.push(`MD정리 → ${getModelName(perStepConfig.format.llm, llmModels)}`);
|
| 1428 |
+
}
|
| 1429 |
+
if (perStepConfig.translate.llm) {
|
| 1430 |
+
overrides.push(`번역 → ${getModelName(perStepConfig.translate.llm, llmModels)}`);
|
| 1431 |
+
}
|
| 1432 |
+
if (perStepConfig.keyframe.llm) {
|
| 1433 |
+
overrides.push(`키프레임 → ${getModelName(perStepConfig.keyframe.llm, visionModels)}`);
|
| 1434 |
+
}
|
| 1435 |
+
|
| 1436 |
+
if (overrides.length > 0) {
|
| 1437 |
+
dot.className = 'settings-dot active';
|
| 1438 |
+
summary.style.display = '';
|
| 1439 |
+
summary.textContent = '⚙ 단계별 설정: ' + overrides.join(' | ');
|
| 1440 |
+
} else {
|
| 1441 |
+
dot.className = 'settings-dot';
|
| 1442 |
+
summary.style.display = 'none';
|
| 1443 |
+
summary.textContent = '';
|
| 1444 |
+
}
|
| 1445 |
+
}
|
| 1446 |
+
</script>
|
| 1447 |
+
</body>
|
| 1448 |
+
</html>
|
transcriber.py
ADDED
|
@@ -0,0 +1,565 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""YouTube 영상에서 음성을 추출하고 텍스트로 변환하는 핵심 모듈."""
|
| 2 |
+
|
| 3 |
+
from __future__ import annotations
|
| 4 |
+
|
| 5 |
+
import os
|
| 6 |
+
import re
|
| 7 |
+
import shutil
|
| 8 |
+
import tempfile
|
| 9 |
+
from typing import Callable, List, Optional
|
| 10 |
+
from pathlib import Path
|
| 11 |
+
|
| 12 |
+
from pytubefix import YouTube
|
| 13 |
+
from pydub import AudioSegment
|
| 14 |
+
|
| 15 |
+
|
| 16 |
+
def download_audio(
|
| 17 |
+
url: str,
|
| 18 |
+
output_dir: Optional[str] = None,
|
| 19 |
+
on_progress: Optional[Callable] = None,
|
| 20 |
+
) -> dict:
|
| 21 |
+
"""YouTube URL에서 오디오를 다운로드한다."""
|
| 22 |
+
if output_dir is None:
|
| 23 |
+
output_dir = tempfile.mkdtemp()
|
| 24 |
+
|
| 25 |
+
def _notify(percent: int, detail: str):
|
| 26 |
+
if on_progress:
|
| 27 |
+
on_progress({"step": "download", "percent": percent, "detail": detail})
|
| 28 |
+
|
| 29 |
+
_notify(0, "영상 정보를 가져오는 중...")
|
| 30 |
+
|
| 31 |
+
# pytubefix 다운로드 진행 콜백
|
| 32 |
+
def _download_cb(stream, chunk, bytes_remaining):
|
| 33 |
+
total = stream.filesize
|
| 34 |
+
downloaded = total - bytes_remaining
|
| 35 |
+
pct = min(int(downloaded / total * 80), 80) # 0~80%: 다운로드
|
| 36 |
+
dl_mb = downloaded / (1024 * 1024)
|
| 37 |
+
total_mb = total / (1024 * 1024)
|
| 38 |
+
_notify(pct, f"다운로드 중... {dl_mb:.1f}MB / {total_mb:.1f}MB")
|
| 39 |
+
|
| 40 |
+
yt = YouTube(url, on_progress_callback=_download_cb)
|
| 41 |
+
title = yt.title
|
| 42 |
+
duration = yt.length
|
| 43 |
+
|
| 44 |
+
_notify(5, f"'{title}' 오디오 스트림 선택 중...")
|
| 45 |
+
|
| 46 |
+
audio_stream = yt.streams.filter(only_audio=True).order_by("abr").desc().first()
|
| 47 |
+
if not audio_stream:
|
| 48 |
+
raise RuntimeError("오디오 스트림을 찾을 수 없습니다.")
|
| 49 |
+
|
| 50 |
+
# 다운로드 (m4a/webm)
|
| 51 |
+
downloaded = audio_stream.download(output_path=output_dir)
|
| 52 |
+
|
| 53 |
+
# mp3로 변환 (80~100%)
|
| 54 |
+
_notify(85, "MP3로 변환 중...")
|
| 55 |
+
mp3_path = os.path.join(output_dir, Path(downloaded).stem + ".mp3")
|
| 56 |
+
audio = AudioSegment.from_file(downloaded)
|
| 57 |
+
audio.export(mp3_path, format="mp3", bitrate="192k")
|
| 58 |
+
|
| 59 |
+
if downloaded != mp3_path and os.path.exists(downloaded):
|
| 60 |
+
os.remove(downloaded)
|
| 61 |
+
|
| 62 |
+
file_mb = os.path.getsize(mp3_path) / (1024 * 1024)
|
| 63 |
+
_notify(100, f"다운로드 완료 ({file_mb:.1f}MB, {duration}초)")
|
| 64 |
+
|
| 65 |
+
return {"audio_path": mp3_path, "title": title, "duration": duration}
|
| 66 |
+
|
| 67 |
+
|
| 68 |
+
def download_video(
|
| 69 |
+
url: str,
|
| 70 |
+
output_dir: Optional[str] = None,
|
| 71 |
+
on_progress: Optional[Callable] = None,
|
| 72 |
+
max_resolution: int = 720,
|
| 73 |
+
) -> str:
|
| 74 |
+
"""키프레임 추출을 위해 YouTube 영상을 다운로드한다 (720p 이하)."""
|
| 75 |
+
if output_dir is None:
|
| 76 |
+
output_dir = tempfile.mkdtemp()
|
| 77 |
+
|
| 78 |
+
def _notify(percent: int, detail: str):
|
| 79 |
+
if on_progress:
|
| 80 |
+
on_progress({"step": "download", "percent": percent, "detail": detail})
|
| 81 |
+
|
| 82 |
+
_notify(50, "키프레임용 영상 다운로드 중...")
|
| 83 |
+
|
| 84 |
+
def _download_cb(stream, chunk, bytes_remaining):
|
| 85 |
+
total = stream.filesize
|
| 86 |
+
downloaded = total - bytes_remaining
|
| 87 |
+
pct = min(50 + int(downloaded / total * 45), 95)
|
| 88 |
+
dl_mb = downloaded / (1024 * 1024)
|
| 89 |
+
total_mb = total / (1024 * 1024)
|
| 90 |
+
_notify(pct, f"영상 다운로드 중... {dl_mb:.1f}MB / {total_mb:.1f}MB")
|
| 91 |
+
|
| 92 |
+
yt = YouTube(url, on_progress_callback=_download_cb)
|
| 93 |
+
|
| 94 |
+
# progressive 스트림 (오디오+비디오 합본, 작은 파일)
|
| 95 |
+
stream = (
|
| 96 |
+
yt.streams.filter(progressive=True, file_extension="mp4")
|
| 97 |
+
.filter(res=f"{max_resolution}p")
|
| 98 |
+
.first()
|
| 99 |
+
)
|
| 100 |
+
if not stream:
|
| 101 |
+
# 해상도 제한 없이 가장 낮은 progressive 시도
|
| 102 |
+
stream = (
|
| 103 |
+
yt.streams.filter(progressive=True, file_extension="mp4")
|
| 104 |
+
.order_by("resolution")
|
| 105 |
+
.first()
|
| 106 |
+
)
|
| 107 |
+
if not stream:
|
| 108 |
+
raise RuntimeError("영상 스트림을 찾을 수 없습니다.")
|
| 109 |
+
|
| 110 |
+
video_path = stream.download(output_path=output_dir, filename_prefix="video_")
|
| 111 |
+
_notify(98, "영상 다운로드 완료")
|
| 112 |
+
return video_path
|
| 113 |
+
|
| 114 |
+
|
| 115 |
+
def extract_keyframes(
|
| 116 |
+
video_path: str,
|
| 117 |
+
output_dir: Optional[str] = None,
|
| 118 |
+
method: str = "scene",
|
| 119 |
+
interval_seconds: int = 30,
|
| 120 |
+
max_frames: int = 50,
|
| 121 |
+
on_progress: Optional[Callable] = None,
|
| 122 |
+
) -> List[dict]:
|
| 123 |
+
"""영상에서 키프레임을 추출한다 (OpenCV).
|
| 124 |
+
|
| 125 |
+
method="scene": 히스토그램 비교로 장면 전환 감지.
|
| 126 |
+
method="interval": N초 간격으로 추출.
|
| 127 |
+
반환: [{"path": str, "timestamp": "MM:SS"}, ...]
|
| 128 |
+
"""
|
| 129 |
+
import cv2
|
| 130 |
+
|
| 131 |
+
def _notify(percent: int, detail: str):
|
| 132 |
+
if on_progress:
|
| 133 |
+
on_progress({"step": "keyframe_extract", "percent": percent, "detail": detail})
|
| 134 |
+
|
| 135 |
+
if output_dir is None:
|
| 136 |
+
output_dir = tempfile.mkdtemp()
|
| 137 |
+
os.makedirs(output_dir, exist_ok=True)
|
| 138 |
+
|
| 139 |
+
cap = cv2.VideoCapture(video_path)
|
| 140 |
+
if not cap.isOpened():
|
| 141 |
+
raise RuntimeError(f"영상 파일을 열 수 없습니다: {video_path}")
|
| 142 |
+
|
| 143 |
+
fps = cap.get(cv2.CAP_PROP_FPS) or 30.0
|
| 144 |
+
total_frames = int(cap.get(cv2.CAP_PROP_FRAME_COUNT))
|
| 145 |
+
total_seconds = total_frames / fps
|
| 146 |
+
|
| 147 |
+
_notify(5, f"영상 분석 시작 ({total_seconds:.0f}초, {fps:.0f}fps)")
|
| 148 |
+
|
| 149 |
+
keyframes = []
|
| 150 |
+
prev_hist = None
|
| 151 |
+
frame_idx = 0
|
| 152 |
+
threshold = 0.6 # 장면 전환 감지 임계값
|
| 153 |
+
|
| 154 |
+
while True:
|
| 155 |
+
ret, frame = cap.read()
|
| 156 |
+
if not ret:
|
| 157 |
+
break
|
| 158 |
+
|
| 159 |
+
current_seconds = frame_idx / fps
|
| 160 |
+
pct = min(int(5 + (frame_idx / total_frames) * 85), 90)
|
| 161 |
+
|
| 162 |
+
should_save = False
|
| 163 |
+
|
| 164 |
+
if method == "scene":
|
| 165 |
+
# HSV 히스토그램 비교로 장면 전환 감지
|
| 166 |
+
hsv = cv2.cvtColor(frame, cv2.COLOR_BGR2HSV)
|
| 167 |
+
hist = cv2.calcHist([hsv], [0, 1], None, [50, 60], [0, 180, 0, 256])
|
| 168 |
+
cv2.normalize(hist, hist, 0, 1, cv2.NORM_MINMAX)
|
| 169 |
+
|
| 170 |
+
if prev_hist is not None:
|
| 171 |
+
score = cv2.compareHist(prev_hist, hist, cv2.HISTCMP_CORREL)
|
| 172 |
+
if score < threshold:
|
| 173 |
+
should_save = True
|
| 174 |
+
else:
|
| 175 |
+
# 첫 프레임은 항상 저장
|
| 176 |
+
should_save = True
|
| 177 |
+
|
| 178 |
+
prev_hist = hist
|
| 179 |
+
else:
|
| 180 |
+
# interval 모드: N초 간격
|
| 181 |
+
if frame_idx == 0 or (frame_idx % int(fps * interval_seconds) == 0):
|
| 182 |
+
should_save = True
|
| 183 |
+
|
| 184 |
+
if should_save and len(keyframes) < max_frames:
|
| 185 |
+
mins = int(current_seconds // 60)
|
| 186 |
+
secs = int(current_seconds % 60)
|
| 187 |
+
timestamp = f"{mins:02d}:{secs:02d}"
|
| 188 |
+
frame_path = os.path.join(output_dir, f"keyframe_{len(keyframes) + 1:03d}.jpg")
|
| 189 |
+
cv2.imwrite(frame_path, frame, [cv2.IMWRITE_JPEG_QUALITY, 85])
|
| 190 |
+
keyframes.append({"path": frame_path, "timestamp": timestamp})
|
| 191 |
+
|
| 192 |
+
if len(keyframes) % 5 == 0:
|
| 193 |
+
_notify(pct, f"{len(keyframes)}개 키프레임 추출됨 ({timestamp})")
|
| 194 |
+
|
| 195 |
+
if len(keyframes) >= max_frames:
|
| 196 |
+
break
|
| 197 |
+
|
| 198 |
+
frame_idx += 1
|
| 199 |
+
|
| 200 |
+
cap.release()
|
| 201 |
+
_notify(100, f"키프레임 추출 완료: {len(keyframes)}개")
|
| 202 |
+
return keyframes
|
| 203 |
+
|
| 204 |
+
|
| 205 |
+
def _split_audio(audio_path: str, max_size_mb: int = 24) -> List[str]:
|
| 206 |
+
"""오디오 파일을 Whisper API 제한(25MB)에 맞게 분할한다."""
|
| 207 |
+
file_size = os.path.getsize(audio_path)
|
| 208 |
+
max_bytes = max_size_mb * 1024 * 1024
|
| 209 |
+
|
| 210 |
+
if file_size <= max_bytes:
|
| 211 |
+
return [audio_path]
|
| 212 |
+
|
| 213 |
+
audio = AudioSegment.from_mp3(audio_path)
|
| 214 |
+
total_ms = len(audio)
|
| 215 |
+
|
| 216 |
+
num_chunks = (file_size // max_bytes) + 1
|
| 217 |
+
chunk_ms = total_ms // num_chunks
|
| 218 |
+
|
| 219 |
+
chunks = []
|
| 220 |
+
base = Path(audio_path)
|
| 221 |
+
for i in range(num_chunks):
|
| 222 |
+
start = i * chunk_ms
|
| 223 |
+
end = min((i + 1) * chunk_ms, total_ms)
|
| 224 |
+
chunk = audio[start:end]
|
| 225 |
+
chunk_path = str(base.parent / f"{base.stem}_part{i}{base.suffix}")
|
| 226 |
+
chunk.export(chunk_path, format="mp3")
|
| 227 |
+
chunks.append(chunk_path)
|
| 228 |
+
|
| 229 |
+
return chunks
|
| 230 |
+
|
| 231 |
+
|
| 232 |
+
def transcribe_api(
|
| 233 |
+
audio_path: str,
|
| 234 |
+
api_key: str,
|
| 235 |
+
on_progress: Optional[Callable] = None,
|
| 236 |
+
) -> dict:
|
| 237 |
+
"""OpenAI Whisper API로 음성을 텍스트로 변환한다."""
|
| 238 |
+
from openai import OpenAI
|
| 239 |
+
|
| 240 |
+
def _notify(percent: int, detail: str):
|
| 241 |
+
if on_progress:
|
| 242 |
+
on_progress({"step": "transcribe", "percent": percent, "detail": detail})
|
| 243 |
+
|
| 244 |
+
_notify(0, "오디오 파일 분석 중...")
|
| 245 |
+
|
| 246 |
+
client = OpenAI(api_key=api_key)
|
| 247 |
+
chunks = _split_audio(audio_path)
|
| 248 |
+
total_chunks = len(chunks)
|
| 249 |
+
all_segments = []
|
| 250 |
+
full_text_parts = []
|
| 251 |
+
detected_language = None
|
| 252 |
+
time_offset = 0.0
|
| 253 |
+
|
| 254 |
+
if total_chunks > 1:
|
| 255 |
+
_notify(5, f"파일이 커서 {total_chunks}개로 분할하여 처리합니다")
|
| 256 |
+
|
| 257 |
+
for i, chunk_path in enumerate(chunks):
|
| 258 |
+
chunk_pct_start = int(i / total_chunks * 90)
|
| 259 |
+
chunk_pct_end = int((i + 1) / total_chunks * 90)
|
| 260 |
+
|
| 261 |
+
if total_chunks > 1:
|
| 262 |
+
_notify(chunk_pct_start, f"청크 {i + 1}/{total_chunks} Whisper API 전송 중...")
|
| 263 |
+
else:
|
| 264 |
+
_notify(10, "Whisper API로 전송 중...")
|
| 265 |
+
|
| 266 |
+
with open(chunk_path, "rb") as f:
|
| 267 |
+
response = client.audio.transcriptions.create(
|
| 268 |
+
model="whisper-1",
|
| 269 |
+
file=f,
|
| 270 |
+
response_format="verbose_json",
|
| 271 |
+
timestamp_granularities=["segment"],
|
| 272 |
+
)
|
| 273 |
+
|
| 274 |
+
if detected_language is None:
|
| 275 |
+
detected_language = response.language
|
| 276 |
+
|
| 277 |
+
full_text_parts.append(response.text)
|
| 278 |
+
|
| 279 |
+
if total_chunks > 1:
|
| 280 |
+
_notify(chunk_pct_end, f"청크 {i + 1}/{total_chunks} 완료")
|
| 281 |
+
else:
|
| 282 |
+
_notify(85, "응답 처리 중...")
|
| 283 |
+
|
| 284 |
+
if response.segments:
|
| 285 |
+
for seg in response.segments:
|
| 286 |
+
_get = (lambda k: seg[k]) if isinstance(seg, dict) else (lambda k: getattr(seg, k))
|
| 287 |
+
all_segments.append(
|
| 288 |
+
{
|
| 289 |
+
"start": _get("start") + time_offset,
|
| 290 |
+
"end": _get("end") + time_offset,
|
| 291 |
+
"text": _get("text"),
|
| 292 |
+
}
|
| 293 |
+
)
|
| 294 |
+
|
| 295 |
+
if chunks[-1] != chunk_path:
|
| 296 |
+
audio = AudioSegment.from_mp3(chunk_path)
|
| 297 |
+
time_offset += len(audio) / 1000.0
|
| 298 |
+
|
| 299 |
+
# 분할된 임시 파일 정리
|
| 300 |
+
for chunk_path in chunks:
|
| 301 |
+
if chunk_path != audio_path and os.path.exists(chunk_path):
|
| 302 |
+
os.remove(chunk_path)
|
| 303 |
+
|
| 304 |
+
_notify(100, f"변환 완료 (언어: {detected_language}, 세그먼트: {len(all_segments)}개)")
|
| 305 |
+
|
| 306 |
+
return {
|
| 307 |
+
"text": " ".join(full_text_parts),
|
| 308 |
+
"segments": all_segments,
|
| 309 |
+
"language": detected_language or "unknown",
|
| 310 |
+
}
|
| 311 |
+
|
| 312 |
+
|
| 313 |
+
def transcribe_local(
|
| 314 |
+
audio_path: str,
|
| 315 |
+
model_size: str = "base",
|
| 316 |
+
on_progress: Optional[Callable] = None,
|
| 317 |
+
) -> dict:
|
| 318 |
+
"""로컬 Whisper 모델로 음성을 텍스트로 변환한다."""
|
| 319 |
+
import whisper
|
| 320 |
+
|
| 321 |
+
def _notify(percent: int, detail: str):
|
| 322 |
+
if on_progress:
|
| 323 |
+
on_progress({"step": "transcribe", "percent": percent, "detail": detail})
|
| 324 |
+
|
| 325 |
+
_notify(0, f"Whisper {model_size} 모델 로딩 중...")
|
| 326 |
+
model = whisper.load_model(model_size)
|
| 327 |
+
|
| 328 |
+
_notify(20, "음성 분석 중... (영상 길이에 따라 시간이 걸립니다)")
|
| 329 |
+
result = model.transcribe(audio_path)
|
| 330 |
+
|
| 331 |
+
_notify(90, "결과 정리 중...")
|
| 332 |
+
segments = []
|
| 333 |
+
for seg in result.get("segments", []):
|
| 334 |
+
segments.append(
|
| 335 |
+
{"start": seg["start"], "end": seg["end"], "text": seg["text"]}
|
| 336 |
+
)
|
| 337 |
+
|
| 338 |
+
lang = result.get("language", "unknown")
|
| 339 |
+
_notify(100, f"변환 완료 (언어: {lang}, 세그먼트: {len(segments)}개)")
|
| 340 |
+
|
| 341 |
+
return {
|
| 342 |
+
"text": result["text"],
|
| 343 |
+
"segments": segments,
|
| 344 |
+
"language": lang,
|
| 345 |
+
}
|
| 346 |
+
|
| 347 |
+
|
| 348 |
+
def _format_srt_time(seconds: float) -> str:
|
| 349 |
+
"""초를 SRT 타임스탬프 형식으로 변환한다."""
|
| 350 |
+
hours = int(seconds // 3600)
|
| 351 |
+
minutes = int((seconds % 3600) // 60)
|
| 352 |
+
secs = int(seconds % 60)
|
| 353 |
+
millis = int((seconds % 1) * 1000)
|
| 354 |
+
return f"{hours:02d}:{minutes:02d}:{secs:02d},{millis:03d}"
|
| 355 |
+
|
| 356 |
+
|
| 357 |
+
def save_txt(result: dict, output_path: str) -> str:
|
| 358 |
+
"""텍스트 파일로 저장한다."""
|
| 359 |
+
with open(output_path, "w", encoding="utf-8") as f:
|
| 360 |
+
f.write(result["text"].strip())
|
| 361 |
+
return output_path
|
| 362 |
+
|
| 363 |
+
|
| 364 |
+
def save_srt(result: dict, output_path: str) -> str:
|
| 365 |
+
"""SRT 자막 파일로 저장한다."""
|
| 366 |
+
with open(output_path, "w", encoding="utf-8") as f:
|
| 367 |
+
for i, seg in enumerate(result["segments"], 1):
|
| 368 |
+
f.write(f"{i}\n")
|
| 369 |
+
f.write(
|
| 370 |
+
f"{_format_srt_time(seg['start'])} --> {_format_srt_time(seg['end'])}\n"
|
| 371 |
+
)
|
| 372 |
+
f.write(f"{seg['text'].strip()}\n\n")
|
| 373 |
+
return output_path
|
| 374 |
+
|
| 375 |
+
|
| 376 |
+
def _sanitize_filename(title: str) -> str:
|
| 377 |
+
"""파일명에 사용할 수 없는 문자를 제거한다."""
|
| 378 |
+
sanitized = re.sub(r'[<>:"/\\|?*]', "", title)
|
| 379 |
+
sanitized = sanitized.strip(". ")
|
| 380 |
+
return sanitized[:100] if sanitized else "untitled"
|
| 381 |
+
|
| 382 |
+
|
| 383 |
+
def process_video(
|
| 384 |
+
url: str,
|
| 385 |
+
mode: str = "api",
|
| 386 |
+
api_key: Optional[str] = None,
|
| 387 |
+
model_size: str = "base",
|
| 388 |
+
output_dir: str = "./output",
|
| 389 |
+
formats: Optional[List[str]] = None,
|
| 390 |
+
on_progress: Optional[Callable] = None,
|
| 391 |
+
# LLM 옵션 (전역 기본값, 하위 호환)
|
| 392 |
+
md_llm: Optional[str] = None,
|
| 393 |
+
md_api_key: Optional[str] = None,
|
| 394 |
+
md_ollama_model: str = "llama3.2",
|
| 395 |
+
# 번역 옵션
|
| 396 |
+
translate_lang: Optional[str] = None,
|
| 397 |
+
# 단계별 LLM 오버라이드
|
| 398 |
+
format_llm: Optional[str] = None,
|
| 399 |
+
format_api_key: Optional[str] = None,
|
| 400 |
+
translate_llm: Optional[str] = None,
|
| 401 |
+
translate_api_key: Optional[str] = None,
|
| 402 |
+
keyframe_llm: Optional[str] = None,
|
| 403 |
+
keyframe_api_key: Optional[str] = None,
|
| 404 |
+
# 키프레임 옵션
|
| 405 |
+
enable_keyframes: bool = False,
|
| 406 |
+
keyframe_method: str = "scene",
|
| 407 |
+
keyframe_interval: int = 30,
|
| 408 |
+
keyframe_max_frames: int = 50,
|
| 409 |
+
) -> dict:
|
| 410 |
+
"""전체 파이프라인: 다운로드 → (키프레임) → 변환 → (AI 정리) → (번역) → 저장."""
|
| 411 |
+
from formatter import make_llm_config, format_as_markdown, translate_text, analyze_keyframes
|
| 412 |
+
|
| 413 |
+
if formats is None:
|
| 414 |
+
formats = ["txt", "srt"]
|
| 415 |
+
|
| 416 |
+
os.makedirs(output_dir, exist_ok=True)
|
| 417 |
+
|
| 418 |
+
def progress(data):
|
| 419 |
+
if on_progress:
|
| 420 |
+
on_progress(data)
|
| 421 |
+
|
| 422 |
+
# 단계별 LLM 설정 빌드
|
| 423 |
+
llm_cfg = make_llm_config(
|
| 424 |
+
global_llm=md_llm, global_api_key=md_api_key,
|
| 425 |
+
global_ollama_model=md_ollama_model,
|
| 426 |
+
format_llm=format_llm, format_api_key=format_api_key,
|
| 427 |
+
translate_llm=translate_llm, translate_api_key=translate_api_key,
|
| 428 |
+
keyframe_llm=keyframe_llm, keyframe_api_key=keyframe_api_key,
|
| 429 |
+
)
|
| 430 |
+
|
| 431 |
+
# 1. 오디오 다운로드
|
| 432 |
+
audio_info = download_audio(url, on_progress=on_progress)
|
| 433 |
+
audio_path = audio_info["audio_path"]
|
| 434 |
+
title = audio_info["title"]
|
| 435 |
+
|
| 436 |
+
video_path = None
|
| 437 |
+
keyframe_dir = None
|
| 438 |
+
|
| 439 |
+
try:
|
| 440 |
+
# 2. 키프레임용 영상 다운로드 (선택 시)
|
| 441 |
+
if enable_keyframes:
|
| 442 |
+
video_path = download_video(url, on_progress=on_progress)
|
| 443 |
+
|
| 444 |
+
# 3. 키프레임 추출 (선택 시)
|
| 445 |
+
keyframe_descriptions = None
|
| 446 |
+
if enable_keyframes and video_path:
|
| 447 |
+
keyframe_dir = tempfile.mkdtemp()
|
| 448 |
+
keyframes = extract_keyframes(
|
| 449 |
+
video_path, output_dir=keyframe_dir,
|
| 450 |
+
method=keyframe_method,
|
| 451 |
+
interval_seconds=keyframe_interval,
|
| 452 |
+
max_frames=keyframe_max_frames,
|
| 453 |
+
on_progress=on_progress,
|
| 454 |
+
)
|
| 455 |
+
|
| 456 |
+
# 4. 키프레임 Vision LLM 분석
|
| 457 |
+
if keyframes:
|
| 458 |
+
kf_cfg = llm_cfg["keyframe"]
|
| 459 |
+
keyframe_descriptions = analyze_keyframes(
|
| 460 |
+
keyframe_paths=keyframes,
|
| 461 |
+
llm_provider=kf_cfg["llm"],
|
| 462 |
+
api_key=kf_cfg["api_key"],
|
| 463 |
+
on_progress=on_progress,
|
| 464 |
+
)
|
| 465 |
+
|
| 466 |
+
# 5. 음성 → 텍스트 변환
|
| 467 |
+
if mode == "api":
|
| 468 |
+
if not api_key:
|
| 469 |
+
raise ValueError("API 모드에서는 OpenAI API 키가 필요합니다.")
|
| 470 |
+
result = transcribe_api(audio_path, api_key, on_progress=on_progress)
|
| 471 |
+
else:
|
| 472 |
+
result = transcribe_local(audio_path, model_size, on_progress=on_progress)
|
| 473 |
+
|
| 474 |
+
# 6. AI 마크다운 정리 (MD 선택 시)
|
| 475 |
+
md_content = None
|
| 476 |
+
fmt_cfg = llm_cfg["format"]
|
| 477 |
+
if "md" in formats and fmt_cfg["llm"]:
|
| 478 |
+
md_content = format_as_markdown(
|
| 479 |
+
text=result["text"],
|
| 480 |
+
title=title,
|
| 481 |
+
llm_provider=fmt_cfg["llm"],
|
| 482 |
+
api_key=fmt_cfg["api_key"],
|
| 483 |
+
ollama_model=fmt_cfg["ollama_model"],
|
| 484 |
+
on_progress=on_progress,
|
| 485 |
+
keyframe_descriptions=keyframe_descriptions,
|
| 486 |
+
)
|
| 487 |
+
|
| 488 |
+
# 7. 번역 (선택 시)
|
| 489 |
+
translated_txt = None
|
| 490 |
+
translated_md = None
|
| 491 |
+
tr_cfg = llm_cfg["translate"]
|
| 492 |
+
if translate_lang and tr_cfg["llm"]:
|
| 493 |
+
translated_txt = translate_text(
|
| 494 |
+
text=result["text"],
|
| 495 |
+
target_lang=translate_lang,
|
| 496 |
+
llm_provider=tr_cfg["llm"],
|
| 497 |
+
api_key=tr_cfg["api_key"],
|
| 498 |
+
ollama_model=tr_cfg["ollama_model"],
|
| 499 |
+
on_progress=on_progress,
|
| 500 |
+
)
|
| 501 |
+
|
| 502 |
+
if md_content:
|
| 503 |
+
translated_md = translate_text(
|
| 504 |
+
text=md_content,
|
| 505 |
+
target_lang=translate_lang,
|
| 506 |
+
llm_provider=tr_cfg["llm"],
|
| 507 |
+
api_key=tr_cfg["api_key"],
|
| 508 |
+
ollama_model=tr_cfg["ollama_model"],
|
| 509 |
+
on_progress=on_progress,
|
| 510 |
+
)
|
| 511 |
+
|
| 512 |
+
# 8. 파일 저장
|
| 513 |
+
progress({"step": "save", "percent": 0, "detail": "파일 저장 중..."})
|
| 514 |
+
safe_title = _sanitize_filename(title)
|
| 515 |
+
saved_files = {}
|
| 516 |
+
|
| 517 |
+
if "txt" in formats:
|
| 518 |
+
txt_path = os.path.join(output_dir, f"{safe_title}.txt")
|
| 519 |
+
save_txt(result, txt_path)
|
| 520 |
+
saved_files["txt"] = txt_path
|
| 521 |
+
|
| 522 |
+
if "srt" in formats:
|
| 523 |
+
srt_path = os.path.join(output_dir, f"{safe_title}.srt")
|
| 524 |
+
save_srt(result, srt_path)
|
| 525 |
+
saved_files["srt"] = srt_path
|
| 526 |
+
|
| 527 |
+
if "md" in formats and md_content:
|
| 528 |
+
md_path = os.path.join(output_dir, f"{safe_title}.md")
|
| 529 |
+
with open(md_path, "w", encoding="utf-8") as f:
|
| 530 |
+
f.write(md_content)
|
| 531 |
+
saved_files["md"] = md_path
|
| 532 |
+
|
| 533 |
+
# 번역 파일 저장
|
| 534 |
+
if translated_txt:
|
| 535 |
+
lang_suffix = translate_lang.replace("-", "").lower()
|
| 536 |
+
tr_txt_path = os.path.join(output_dir, f"{safe_title}_{lang_suffix}.txt")
|
| 537 |
+
with open(tr_txt_path, "w", encoding="utf-8") as f:
|
| 538 |
+
f.write(translated_txt)
|
| 539 |
+
saved_files[f"txt_{lang_suffix}"] = tr_txt_path
|
| 540 |
+
|
| 541 |
+
if translated_md:
|
| 542 |
+
lang_suffix = translate_lang.replace("-", "").lower()
|
| 543 |
+
tr_md_path = os.path.join(output_dir, f"{safe_title}_{lang_suffix}.md")
|
| 544 |
+
with open(tr_md_path, "w", encoding="utf-8") as f:
|
| 545 |
+
f.write(translated_md)
|
| 546 |
+
saved_files[f"md_{lang_suffix}"] = tr_md_path
|
| 547 |
+
|
| 548 |
+
fmt_str = ", ".join(f.upper() for f in saved_files.keys())
|
| 549 |
+
progress({"step": "save", "percent": 100, "detail": f"{fmt_str} 파일 저장 완료"})
|
| 550 |
+
|
| 551 |
+
return {
|
| 552 |
+
"title": title,
|
| 553 |
+
"language": result["language"],
|
| 554 |
+
"text": result["text"],
|
| 555 |
+
"segments": result["segments"],
|
| 556 |
+
"files": saved_files,
|
| 557 |
+
}
|
| 558 |
+
finally:
|
| 559 |
+
# 임시 파일 정리
|
| 560 |
+
if os.path.exists(audio_path):
|
| 561 |
+
os.remove(audio_path)
|
| 562 |
+
if video_path and os.path.exists(video_path):
|
| 563 |
+
os.remove(video_path)
|
| 564 |
+
if keyframe_dir and os.path.exists(keyframe_dir):
|
| 565 |
+
shutil.rmtree(keyframe_dir, ignore_errors=True)
|
web.py
ADDED
|
@@ -0,0 +1,237 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
#!/usr/bin/env python3
|
| 2 |
+
"""YouTube Script Extractor - Web Server."""
|
| 3 |
+
|
| 4 |
+
import asyncio
|
| 5 |
+
import json
|
| 6 |
+
import os
|
| 7 |
+
import uuid
|
| 8 |
+
from pathlib import Path
|
| 9 |
+
|
| 10 |
+
from dotenv import load_dotenv
|
| 11 |
+
|
| 12 |
+
load_dotenv()
|
| 13 |
+
|
| 14 |
+
from fastapi import FastAPI, Request
|
| 15 |
+
from fastapi.responses import FileResponse, HTMLResponse, StreamingResponse
|
| 16 |
+
from fastapi.staticfiles import StaticFiles
|
| 17 |
+
from fastapi.templating import Jinja2Templates
|
| 18 |
+
|
| 19 |
+
from transcriber import process_video
|
| 20 |
+
from formatter import LLM_MODELS, get_models_sorted, get_languages
|
| 21 |
+
|
| 22 |
+
app = FastAPI(title="YouTube Script Extractor")
|
| 23 |
+
templates = Jinja2Templates(directory="templates")
|
| 24 |
+
|
| 25 |
+
OUTPUT_DIR = "./output"
|
| 26 |
+
os.makedirs(OUTPUT_DIR, exist_ok=True)
|
| 27 |
+
|
| 28 |
+
# 진행 중인 작업 추적
|
| 29 |
+
jobs: dict[str, dict] = {}
|
| 30 |
+
|
| 31 |
+
|
| 32 |
+
@app.get("/", response_class=HTMLResponse)
|
| 33 |
+
async def index(request: Request):
|
| 34 |
+
return templates.TemplateResponse("index.html", {"request": request})
|
| 35 |
+
|
| 36 |
+
|
| 37 |
+
@app.get("/api/llm-models")
|
| 38 |
+
async def get_llm_models(sort: str = "price"):
|
| 39 |
+
"""LLM 모델 목록을 정렬하여 반환한다."""
|
| 40 |
+
models = get_models_sorted(sort_by=sort)
|
| 41 |
+
return {"models": models}
|
| 42 |
+
|
| 43 |
+
|
| 44 |
+
@app.get("/api/languages")
|
| 45 |
+
async def get_supported_languages():
|
| 46 |
+
"""번역 지원 언어 목록을 반환한다."""
|
| 47 |
+
return {"languages": get_languages()}
|
| 48 |
+
|
| 49 |
+
|
| 50 |
+
@app.get("/api/vision-models")
|
| 51 |
+
async def get_vision_models(sort: str = "price"):
|
| 52 |
+
"""Vision 지원 LLM 모델 목록을 반환한다."""
|
| 53 |
+
key = "price_rank" if sort == "price" else "quality_rank"
|
| 54 |
+
models = [
|
| 55 |
+
{"id": k, **v} for k, v in LLM_MODELS.items()
|
| 56 |
+
if v.get("supports_vision", False)
|
| 57 |
+
]
|
| 58 |
+
return {"models": sorted(models, key=lambda x: x[key])}
|
| 59 |
+
|
| 60 |
+
|
| 61 |
+
@app.post("/api/transcribe")
|
| 62 |
+
async def start_transcription(request: Request):
|
| 63 |
+
body = await request.json()
|
| 64 |
+
url = body.get("url", "").strip()
|
| 65 |
+
mode = body.get("mode", "api")
|
| 66 |
+
api_key = body.get("api_key", "") or os.environ.get("OPENAI_API_KEY", "")
|
| 67 |
+
model_size = body.get("model_size", "base")
|
| 68 |
+
formats = body.get("formats", ["txt", "srt"])
|
| 69 |
+
output_dir = body.get("output_dir", "").strip() or OUTPUT_DIR
|
| 70 |
+
|
| 71 |
+
# LLM 옵션 (전역 기본값)
|
| 72 |
+
md_llm = body.get("md_llm", "")
|
| 73 |
+
md_api_key = body.get("md_api_key", "")
|
| 74 |
+
md_ollama_model = body.get("md_ollama_model", "llama3.2")
|
| 75 |
+
translate_lang = body.get("translate_lang", "")
|
| 76 |
+
|
| 77 |
+
# 단계별 LLM 오버라이드
|
| 78 |
+
format_llm = body.get("format_llm", "")
|
| 79 |
+
format_api_key = body.get("format_api_key", "")
|
| 80 |
+
translate_llm = body.get("translate_llm", "")
|
| 81 |
+
translate_api_key = body.get("translate_api_key", "")
|
| 82 |
+
keyframe_llm = body.get("keyframe_llm", "")
|
| 83 |
+
keyframe_api_key = body.get("keyframe_api_key", "")
|
| 84 |
+
|
| 85 |
+
# 키프레임 옵션
|
| 86 |
+
enable_keyframes = body.get("enable_keyframes", False)
|
| 87 |
+
keyframe_method = body.get("keyframe_method", "scene")
|
| 88 |
+
keyframe_interval = body.get("keyframe_interval", 30)
|
| 89 |
+
|
| 90 |
+
if not url:
|
| 91 |
+
return {"error": "YouTube URL을 입력해주세요."}
|
| 92 |
+
|
| 93 |
+
if mode == "api" and not api_key:
|
| 94 |
+
return {"error": "API 모드에서는 OpenAI API 키가 필요합니다."}
|
| 95 |
+
|
| 96 |
+
job_id = str(uuid.uuid4())[:8]
|
| 97 |
+
jobs[job_id] = {"status": "started", "progress": [], "result": None, "error": None}
|
| 98 |
+
|
| 99 |
+
asyncio.create_task(
|
| 100 |
+
_run_transcription(
|
| 101 |
+
job_id, url, mode, api_key, model_size, formats, output_dir,
|
| 102 |
+
md_llm=md_llm, md_api_key=md_api_key, md_ollama_model=md_ollama_model,
|
| 103 |
+
translate_lang=translate_lang,
|
| 104 |
+
format_llm=format_llm, format_api_key=format_api_key,
|
| 105 |
+
translate_llm=translate_llm, translate_api_key=translate_api_key,
|
| 106 |
+
keyframe_llm=keyframe_llm, keyframe_api_key=keyframe_api_key,
|
| 107 |
+
enable_keyframes=enable_keyframes,
|
| 108 |
+
keyframe_method=keyframe_method, keyframe_interval=keyframe_interval,
|
| 109 |
+
)
|
| 110 |
+
)
|
| 111 |
+
|
| 112 |
+
return {"job_id": job_id}
|
| 113 |
+
|
| 114 |
+
|
| 115 |
+
async def _run_transcription(
|
| 116 |
+
job_id: str,
|
| 117 |
+
url: str,
|
| 118 |
+
mode: str,
|
| 119 |
+
api_key: str,
|
| 120 |
+
model_size: str,
|
| 121 |
+
formats: list[str],
|
| 122 |
+
output_dir: str = OUTPUT_DIR,
|
| 123 |
+
md_llm: str = "",
|
| 124 |
+
md_api_key: str = "",
|
| 125 |
+
md_ollama_model: str = "llama3.2",
|
| 126 |
+
translate_lang: str = "",
|
| 127 |
+
format_llm: str = "",
|
| 128 |
+
format_api_key: str = "",
|
| 129 |
+
translate_llm: str = "",
|
| 130 |
+
translate_api_key: str = "",
|
| 131 |
+
keyframe_llm: str = "",
|
| 132 |
+
keyframe_api_key: str = "",
|
| 133 |
+
enable_keyframes: bool = False,
|
| 134 |
+
keyframe_method: str = "scene",
|
| 135 |
+
keyframe_interval: int = 30,
|
| 136 |
+
):
|
| 137 |
+
def on_progress(msg: str):
|
| 138 |
+
jobs[job_id]["progress"].append(msg)
|
| 139 |
+
|
| 140 |
+
try:
|
| 141 |
+
result = await asyncio.to_thread(
|
| 142 |
+
process_video,
|
| 143 |
+
url=url,
|
| 144 |
+
mode=mode,
|
| 145 |
+
api_key=api_key,
|
| 146 |
+
model_size=model_size,
|
| 147 |
+
output_dir=output_dir,
|
| 148 |
+
formats=formats,
|
| 149 |
+
on_progress=on_progress,
|
| 150 |
+
md_llm=md_llm or None,
|
| 151 |
+
md_api_key=md_api_key or None,
|
| 152 |
+
md_ollama_model=md_ollama_model,
|
| 153 |
+
translate_lang=translate_lang or None,
|
| 154 |
+
format_llm=format_llm or None,
|
| 155 |
+
format_api_key=format_api_key or None,
|
| 156 |
+
translate_llm=translate_llm or None,
|
| 157 |
+
translate_api_key=translate_api_key or None,
|
| 158 |
+
keyframe_llm=keyframe_llm or None,
|
| 159 |
+
keyframe_api_key=keyframe_api_key or None,
|
| 160 |
+
enable_keyframes=enable_keyframes,
|
| 161 |
+
keyframe_method=keyframe_method,
|
| 162 |
+
keyframe_interval=keyframe_interval,
|
| 163 |
+
)
|
| 164 |
+
jobs[job_id]["status"] = "completed"
|
| 165 |
+
# 절대 경로로 변환하여 저장 위치를 정확히 표시
|
| 166 |
+
abs_output = os.path.abspath(output_dir)
|
| 167 |
+
jobs[job_id]["result"] = {
|
| 168 |
+
"title": result["title"],
|
| 169 |
+
"language": result["language"],
|
| 170 |
+
"text": result["text"],
|
| 171 |
+
"files": {
|
| 172 |
+
fmt: os.path.basename(path) for fmt, path in result["files"].items()
|
| 173 |
+
},
|
| 174 |
+
"output_dir": abs_output,
|
| 175 |
+
"output_dir_raw": output_dir,
|
| 176 |
+
}
|
| 177 |
+
except Exception as e:
|
| 178 |
+
jobs[job_id]["status"] = "error"
|
| 179 |
+
jobs[job_id]["error"] = str(e)
|
| 180 |
+
|
| 181 |
+
|
| 182 |
+
@app.get("/api/status/{job_id}")
|
| 183 |
+
async def get_status(job_id: str):
|
| 184 |
+
job = jobs.get(job_id)
|
| 185 |
+
if not job:
|
| 186 |
+
return {"error": "작업을 찾을 수 없습니다."}
|
| 187 |
+
return job
|
| 188 |
+
|
| 189 |
+
|
| 190 |
+
@app.get("/api/stream/{job_id}")
|
| 191 |
+
async def stream_status(job_id: str):
|
| 192 |
+
async def event_generator():
|
| 193 |
+
seen = 0
|
| 194 |
+
while True:
|
| 195 |
+
job = jobs.get(job_id)
|
| 196 |
+
if not job:
|
| 197 |
+
yield f"data: {json.dumps({'type': 'error', 'message': '작업을 찾을 수 없습니다.'})}\n\n"
|
| 198 |
+
break
|
| 199 |
+
|
| 200 |
+
# 새 진행 메시지 전송
|
| 201 |
+
while seen < len(job["progress"]):
|
| 202 |
+
msg = job["progress"][seen]
|
| 203 |
+
if isinstance(msg, dict):
|
| 204 |
+
yield f"data: {json.dumps({'type': 'progress', **msg})}\n\n"
|
| 205 |
+
else:
|
| 206 |
+
yield f"data: {json.dumps({'type': 'progress', 'message': msg})}\n\n"
|
| 207 |
+
seen += 1
|
| 208 |
+
|
| 209 |
+
if job["status"] == "completed":
|
| 210 |
+
yield f"data: {json.dumps({'type': 'completed', 'result': job['result']})}\n\n"
|
| 211 |
+
break
|
| 212 |
+
elif job["status"] == "error":
|
| 213 |
+
yield f"data: {json.dumps({'type': 'error', 'message': job['error']})}\n\n"
|
| 214 |
+
break
|
| 215 |
+
|
| 216 |
+
await asyncio.sleep(0.5)
|
| 217 |
+
|
| 218 |
+
return StreamingResponse(event_generator(), media_type="text/event-stream")
|
| 219 |
+
|
| 220 |
+
|
| 221 |
+
@app.get("/api/download/{filename}")
|
| 222 |
+
async def download_file(filename: str, dir: str = ""):
|
| 223 |
+
base_dir = dir if dir else OUTPUT_DIR
|
| 224 |
+
file_path = os.path.join(base_dir, filename)
|
| 225 |
+
if not os.path.exists(file_path):
|
| 226 |
+
return {"error": "파일을 찾을 수 없습니다."}
|
| 227 |
+
return FileResponse(file_path, filename=filename)
|
| 228 |
+
|
| 229 |
+
|
| 230 |
+
if __name__ == "__main__":
|
| 231 |
+
import uvicorn
|
| 232 |
+
|
| 233 |
+
port = int(os.environ.get("PORT", 8000))
|
| 234 |
+
print("YouTube Script Extractor 웹 서버")
|
| 235 |
+
print(f" http://localhost:{port}")
|
| 236 |
+
print()
|
| 237 |
+
uvicorn.run(app, host="0.0.0.0", port=port)
|