Heebin Moon Claude Opus 4.6 commited on
Commit
a2749f3
·
0 Parent(s):

Initial commit: YouTube transcript formatter

Browse files

YouTube 영상의 자막을 추출하고 포맷팅하는 도구

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

Files changed (10) hide show
  1. .env.example +17 -0
  2. .gitignore +27 -0
  3. README.md +116 -0
  4. cli.py +111 -0
  5. formatter.py +687 -0
  6. requirements.txt +15 -0
  7. run +68 -0
  8. templates/index.html +1448 -0
  9. transcriber.py +565 -0
  10. web.py +237 -0
.env.example ADDED
@@ -0,0 +1,17 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # OpenAI API 키 (https://platform.openai.com/api-keys 에서 발급)
2
+ OPENAI_API_KEY=sk-proj-여기에_키를_입력하세요
3
+
4
+ # 기본 변환 모드 (api 또는 local)
5
+ DEFAULT_MODE=api
6
+
7
+ # 로컬 Whisper 모델 크기 (tiny, base, small, medium, large)
8
+ DEFAULT_MODEL=base
9
+
10
+ # MD 포맷팅용 API 키 (사용하는 LLM만 설정하면 됩니다)
11
+ # Google Gemini (https://aistudio.google.com/app/apikey)
12
+ GOOGLE_API_KEY=여기에_키를_입력하세요
13
+ # Anthropic Claude (https://console.anthropic.com/)
14
+ ANTHROPIC_API_KEY=여기에_키를_입력하세요
15
+
16
+ # 웹 서버 포트
17
+ PORT=8000
.gitignore ADDED
@@ -0,0 +1,27 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # 환경 변수 (API 키 포함)
2
+ .env
3
+
4
+ # 출력물
5
+ output/
6
+
7
+ # Python
8
+ __pycache__/
9
+ *.pyc
10
+ *.pyo
11
+ *.egg-info/
12
+ dist/
13
+ build/
14
+
15
+ # Claude Code 로컬 설정
16
+ .claude/
17
+
18
+ # OS
19
+ .DS_Store
20
+ Thumbs.db
21
+
22
+ # Whisper 모델 캐시
23
+ *.pt
24
+
25
+ # 임시 파일
26
+ *.tmp
27
+ temp/
README.md ADDED
@@ -0,0 +1,116 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # 🎬 YouTube Script Extractor
2
+
3
+ YouTube 영상에서 스크립트를 추출하고, AI로 깔끔하게 정리하는 도구입니다.
4
+
5
+ ## ✨ 주요 기능
6
+
7
+ - **자동 스크립트 추출** — YouTube URL만 입력하면 음성을 텍스트로 변환
8
+ - **AI 마크다운 정리** — LLM이 스크립트를 구조화된 마크다운으로 정리
9
+ - **다국어 번역** — 14개 언어 지원 (한국어, 영어, 일본어 등)
10
+ - **키프레임 분석 (CV)** — 영상의 핵심 장면을 추출하고 Vision AI로 분석
11
+ - **Smart Mix** — 파이프라인 단계별로 다른 LLM을 지정 가능
12
+ - **웹 UI + CLI** — 브라우저 기반 UI와 커맨드라인 모두 지원
13
+
14
+ ## 🤖 지원 LLM
15
+
16
+ | 모델 | 가격 | Vision | 비고 |
17
+ |------|------|--------|------|
18
+ | Ollama (로컬) | 무료 | ❌ | 인터넷 불필요 |
19
+ | Gemini 2.5 Flash-Lite | $0.10/1M토큰 | ✅ | ⭐ 최저가 |
20
+ | Gemini 3.1 Flash-Lite | $0.25/1M토큰 | ✅ | 최신 모델 |
21
+ | GPT-4o-mini | $0.15/1M토큰 | ✅ | OpenAI |
22
+ | Claude Sonnet 4.6 | $3/1M토큰 | ✅ | 고품질 |
23
+ | Claude Opus 4.6 | $5/1M토큰 | ✅ | 최고 품질 |
24
+
25
+ ## 🚀 빠른 시작
26
+
27
+ ### 1. 설치
28
+
29
+ ```bash
30
+ git clone https://huggingface.co/spaces/YOUR_USERNAME/youtube-script-extractor
31
+ cd youtube-script-extractor
32
+
33
+ # ffmpeg 설치 (macOS)
34
+ brew install ffmpeg
35
+
36
+ # Python 의존성 설치
37
+ pip install -r requirements.txt
38
+ ```
39
+
40
+ ### 2. API 키 설정
41
+
42
+ ```bash
43
+ cp .env.example .env
44
+ # .env 파일을 열어 사용할 LLM의 API 키를 입력하세요
45
+ ```
46
+
47
+ ### 3. 실행
48
+
49
+ ```bash
50
+ # 웹 UI (추천)
51
+ chmod +x run
52
+ ./run
53
+
54
+ # 또는 직접 실행
55
+ python3 web.py
56
+ ```
57
+
58
+ 브라우저에서 `http://localhost:8000` 접속
59
+
60
+ ### CLI 모드
61
+
62
+ ```bash
63
+ ./run "https://www.youtube.com/watch?v=VIDEO_ID"
64
+ ./run "https://www.youtube.com/watch?v=VIDEO_ID" --mode local --format srt
65
+ ```
66
+
67
+ ## ⚙️ 설정
68
+
69
+ `.env` 파일에서 설정합니다:
70
+
71
+ ```env
72
+ # 필수: Whisper API용
73
+ OPENAI_API_KEY=sk-proj-...
74
+
75
+ # 선택: 사용하는 LLM만 설정
76
+ GOOGLE_API_KEY=... # Gemini
77
+ ANTHROPIC_API_KEY=... # Claude
78
+
79
+ # 기본 설정
80
+ DEFAULT_MODE=api # api 또는 local
81
+ DEFAULT_MODEL=base # Whisper 모델 (tiny/base/small/medium/large)
82
+ PORT=8000
83
+ ```
84
+
85
+ ## 📁 프로젝트 구조
86
+
87
+ ```
88
+ ├── run # 실행 스크립트 (설치 + 실행)
89
+ ├── web.py # FastAPI 웹 서버
90
+ ├── cli.py # CLI 인터페이스
91
+ ├── transcriber.py # 음성 추출 + STT + 키프레임 추출
92
+ ├── formatter.py # LLM 호출 + 마크다운 정리 + 번역
93
+ ├── templates/
94
+ │ └── index.html # 웹 UI
95
+ ├── requirements.txt # Python 의존성
96
+ ├── .env.example # 환경 변수 템플릿
97
+ └── output/ # 추출 결과물 (git 제외)
98
+ ```
99
+
100
+ ## 🔧 Smart Mix (단계별 LLM 선택)
101
+
102
+ 웹 UI의 ⚙ 버튼으로 파이프라인 단계별 LLM을 설정할 수 있습니다:
103
+
104
+ - **MD 정리** — 스크립트를 마크다운으로 구조화 (저렴한 모델 OK)
105
+ - **번역** — 다국어 번역 (저렴한 모델 OK)
106
+ - **키프레임 분석** — Vision AI로 화면 분석 (Vision 지원 모델 필요)
107
+
108
+ ## 📋 요구사항
109
+
110
+ - Python 3.9+
111
+ - ffmpeg
112
+ - API 키 (사용하려는 LLM에 따라)
113
+
114
+ ## 📄 License
115
+
116
+ MIT
cli.py ADDED
@@ -0,0 +1,111 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ #!/usr/bin/env python3
2
+ """YouTube Script Extractor - CLI."""
3
+
4
+ import argparse
5
+ import os
6
+ import sys
7
+
8
+ from dotenv import load_dotenv
9
+
10
+ load_dotenv()
11
+
12
+ from transcriber import process_video
13
+
14
+
15
+ def main():
16
+ parser = argparse.ArgumentParser(
17
+ description="YouTube 영상에서 음성인식으로 스크립트를 추출합니다.",
18
+ formatter_class=argparse.RawDescriptionHelpFormatter,
19
+ epilog="""
20
+ 사용 예시:
21
+ # Whisper API 사용 (빠르고 정확)
22
+ python cli.py "https://www.youtube.com/watch?v=..." --mode api
23
+
24
+ # 로컬 Whisper 사용 (무료)
25
+ python cli.py "https://www.youtube.com/watch?v=..." --mode local --model base
26
+
27
+ # SRT 자막만 생성
28
+ python cli.py "https://www.youtube.com/watch?v=..." --format srt
29
+
30
+ # 출력 디렉토리 지정
31
+ python cli.py "https://www.youtube.com/watch?v=..." --output ./my_scripts
32
+ """,
33
+ )
34
+
35
+ parser.add_argument("url", help="YouTube 영상 URL")
36
+ parser.add_argument(
37
+ "--mode",
38
+ choices=["api", "local"],
39
+ default=os.environ.get("DEFAULT_MODE", "api"),
40
+ help="변환 모드 (기본: api)",
41
+ )
42
+ parser.add_argument(
43
+ "--api-key",
44
+ default=None,
45
+ help="OpenAI API 키 (미지정 시 OPENAI_API_KEY 환경변수 사용)",
46
+ )
47
+ parser.add_argument(
48
+ "--model",
49
+ default=os.environ.get("DEFAULT_MODEL", "base"),
50
+ choices=["tiny", "base", "small", "medium", "large"],
51
+ help="로컬 Whisper 모델 크기 (기본: base)",
52
+ )
53
+ parser.add_argument(
54
+ "--format",
55
+ default="txt,srt",
56
+ help="출력 형식, 쉼표로 구분 (기본: txt,srt)",
57
+ )
58
+ parser.add_argument(
59
+ "--output",
60
+ default="./output",
61
+ help="출력 디렉토리 (기본: ./output)",
62
+ )
63
+
64
+ args = parser.parse_args()
65
+
66
+ # API 키 처리
67
+ api_key = args.api_key or os.environ.get("OPENAI_API_KEY")
68
+ if args.mode == "api" and not api_key:
69
+ print("오류: API 모드에서는 OpenAI API 키가 필요합니다.")
70
+ print(" --api-key 인자를 사용하거나 OPENAI_API_KEY 환경변수를 설정하세요.")
71
+ sys.exit(1)
72
+
73
+ formats = [f.strip() for f in args.format.split(",")]
74
+
75
+ def on_progress(msg: str):
76
+ print(f" → {msg}")
77
+
78
+ print(f"YouTube Script Extractor")
79
+ print(f" URL: {args.url}")
80
+ print(f" 모드: {'Whisper API' if args.mode == 'api' else f'로컬 Whisper ({args.model})'}")
81
+ print(f" 출력: {', '.join(formats)}")
82
+ print()
83
+
84
+ try:
85
+ result = process_video(
86
+ url=args.url,
87
+ mode=args.mode,
88
+ api_key=api_key,
89
+ model_size=args.model,
90
+ output_dir=args.output,
91
+ formats=formats,
92
+ on_progress=on_progress,
93
+ )
94
+
95
+ print()
96
+ print(f"제목: {result['title']}")
97
+ print(f"감지된 언어: {result['language']}")
98
+ print(f"생성된 파일:")
99
+ for fmt, path in result["files"].items():
100
+ print(f" [{fmt.upper()}] {path}")
101
+
102
+ except KeyboardInterrupt:
103
+ print("\n중단되었습니다.")
104
+ sys.exit(130)
105
+ except Exception as e:
106
+ print(f"\n오류 발생: {e}")
107
+ sys.exit(1)
108
+
109
+
110
+ if __name__ == "__main__":
111
+ main()
formatter.py ADDED
@@ -0,0 +1,687 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """LLM을 사용하여 Whisper 추출 텍스트를 마크다운 정리/번역하는 모듈."""
2
+
3
+ from __future__ import annotations
4
+
5
+ import base64
6
+ import json
7
+ import os
8
+ import re
9
+ import time
10
+ from typing import Callable, List, Optional
11
+
12
+ # ── LLM 정보 (가격순 정렬) ──────────────────────────────────
13
+ LLM_MODELS = {
14
+ "ollama": {
15
+ "name": "Ollama (로컬, 무료)",
16
+ "model": "llama3.2",
17
+ "version": "llama3.2",
18
+ "price_rank": 0,
19
+ "quality_rank": 5,
20
+ "needs_key": False,
21
+ "env_key": None,
22
+ "supports_vision": False,
23
+ "description": "로컬 실행, 무료, 인터넷 불필요",
24
+ },
25
+ "gemini-flash-lite": {
26
+ "name": "Gemini 2.5 Flash-Lite",
27
+ "model": "gemini-2.5-flash-lite",
28
+ "version": "2.5-flash-lite",
29
+ "price_rank": 1,
30
+ "quality_rank": 4,
31
+ "needs_key": True,
32
+ "env_key": "GOOGLE_API_KEY",
33
+ "supports_vision": True,
34
+ "description": "⭐ 최저가, $0.10/1M토큰, 빠르고 안정적",
35
+ },
36
+ "gemini-flash": {
37
+ "name": "Gemini 3.1 Flash-Lite (Preview)",
38
+ "model": "gemini-3.1-flash-lite-preview",
39
+ "version": "3.1-flash-lite",
40
+ "price_rank": 2,
41
+ "quality_rank": 3,
42
+ "needs_key": True,
43
+ "env_key": "GOOGLE_API_KEY",
44
+ "supports_vision": True,
45
+ "description": "최신 모델, $0.25/1M토큰, 향상된 성능",
46
+ },
47
+ "gpt-4o-mini": {
48
+ "name": "GPT-4o-mini (OpenAI)",
49
+ "model": "gpt-4o-mini",
50
+ "version": "4o-mini",
51
+ "price_rank": 3,
52
+ "quality_rank": 2,
53
+ "needs_key": True,
54
+ "env_key": "OPENAI_API_KEY",
55
+ "supports_vision": True,
56
+ "description": "빠르고 저렴, $0.15/1M토큰",
57
+ },
58
+ "claude-sonnet": {
59
+ "name": "Claude Sonnet 4.6 (Anthropic)",
60
+ "model": "claude-sonnet-4-6",
61
+ "version": "sonnet-4.6",
62
+ "price_rank": 4,
63
+ "quality_rank": 1,
64
+ "needs_key": True,
65
+ "env_key": "ANTHROPIC_API_KEY",
66
+ "supports_vision": True,
67
+ "description": "Anthropic, 빠르고 고품질 ($3/1M토큰)",
68
+ },
69
+ "claude-opus": {
70
+ "name": "Claude Opus 4.6 (Anthropic)",
71
+ "model": "claude-opus-4-6",
72
+ "version": "opus-4.6",
73
+ "price_rank": 5,
74
+ "quality_rank": 0,
75
+ "needs_key": True,
76
+ "env_key": "ANTHROPIC_API_KEY",
77
+ "supports_vision": True,
78
+ "description": "Anthropic 최고 모델, 구조화 능력 최고 ($5/1M토큰)",
79
+ },
80
+ }
81
+
82
+ # ── 번역 지원 언어 ────────────────────────────────────────
83
+ LANGUAGES = {
84
+ "ko": "한국어",
85
+ "en": "English",
86
+ "ja": "日本語",
87
+ "zh-CN": "中文(简体)",
88
+ "zh-TW": "中文(繁體)",
89
+ "es": "Español",
90
+ "fr": "Français",
91
+ "de": "Deutsch",
92
+ "pt": "Português",
93
+ "ru": "Русский",
94
+ "vi": "Tiếng Việt",
95
+ "th": "ภาษาไทย",
96
+ "ar": "العربية",
97
+ "hi": "हिन्दी",
98
+ }
99
+
100
+
101
+ def get_models_sorted(sort_by: str = "price") -> list:
102
+ """정렬된 LLM 모델 리스트를 반환한다."""
103
+ key = "price_rank" if sort_by == "price" else "quality_rank"
104
+ return sorted(
105
+ [{"id": k, **v} for k, v in LLM_MODELS.items()],
106
+ key=lambda x: x[key],
107
+ )
108
+
109
+
110
+ def get_languages() -> list:
111
+ """지원되는 번역 언어 목록을 반환한다."""
112
+ return [{"code": k, "name": v} for k, v in LANGUAGES.items()]
113
+
114
+
115
+ def make_llm_config(
116
+ global_llm: Optional[str] = None,
117
+ global_api_key: Optional[str] = None,
118
+ global_ollama_model: str = "llama3.2",
119
+ format_llm: Optional[str] = None,
120
+ format_api_key: Optional[str] = None,
121
+ translate_llm: Optional[str] = None,
122
+ translate_api_key: Optional[str] = None,
123
+ keyframe_llm: Optional[str] = None,
124
+ keyframe_api_key: Optional[str] = None,
125
+ ) -> dict:
126
+ """단계별 LLM 설정 딕셔너리를 생성한다.
127
+
128
+ per-step 값이 없으면 global 값으로 fallback.
129
+ 반환: {"format": {...}, "translate": {...}, "keyframe": {...}}
130
+ """
131
+ def _resolve(step_llm, step_key):
132
+ llm = step_llm or global_llm
133
+ api_key = step_key or global_api_key
134
+ return {
135
+ "llm": llm,
136
+ "api_key": api_key,
137
+ "ollama_model": global_ollama_model,
138
+ }
139
+
140
+ return {
141
+ "format": _resolve(format_llm, format_api_key),
142
+ "translate": _resolve(translate_llm, translate_api_key),
143
+ "keyframe": _resolve(
144
+ keyframe_llm or "gemini-flash-lite",
145
+ keyframe_api_key,
146
+ ),
147
+ }
148
+
149
+
150
+ # ── 프롬프트 ──────────────────────────────────────────────
151
+
152
+ FORMAT_SYSTEM_PROMPT = """You are an expert document formatter. Your task is to transform raw speech-to-text transcriptions into clean, well-structured Markdown documents.
153
+
154
+ Rules:
155
+ - Detect the content's language and write the output in the SAME language
156
+ - Add a clear title as # heading
157
+ - Write a brief 2-3 sentence summary at the top
158
+ - Divide content into logical sections with ## headings
159
+ - Clean up filler words, repetitions, and stutters
160
+ - Fix obvious grammar/punctuation errors
161
+ - Keep the original meaning and tone intact
162
+ - Use bullet points or numbered lists where appropriate
163
+ - Add --- horizontal rules between major sections
164
+ - Do NOT add information that wasn't in the original text
165
+ - Output ONLY the formatted Markdown, no explanations"""
166
+
167
+ FORMAT_USER_TEMPLATE = """Here is a raw speech-to-text transcription from a YouTube video titled "{title}".
168
+ Please format it into a clean, readable Markdown document.
169
+
170
+ ---
171
+ {text}
172
+ ---"""
173
+
174
+ FORMAT_SYSTEM_PROMPT_WITH_KEYFRAMES = """You are an expert document formatter. Your task is to transform raw speech-to-text transcriptions into clean, well-structured Markdown documents, enhanced with visual context from video keyframes.
175
+
176
+ Rules:
177
+ - Detect the content's language and write the output in the SAME language
178
+ - Add a clear title as # heading
179
+ - Write a brief 2-3 sentence summary at the top
180
+ - Divide content into logical sections with ## headings
181
+ - Clean up filler words, repetitions, and stutters
182
+ - Fix obvious grammar/punctuation errors
183
+ - Keep the original meaning and tone intact
184
+ - Use bullet points or numbered lists where appropriate
185
+ - Add --- horizontal rules between major sections
186
+ - KEYFRAME CONTEXT: You are also given timestamped descriptions of visual keyframes.
187
+ Insert relevant visual descriptions as blockquotes (> 🖼 [timestamp] description) at
188
+ appropriate positions in the transcript where they add context.
189
+ - Only include keyframe descriptions that add meaningful value (skip redundant ones)
190
+ - Do NOT add information that wasn't in the original text or keyframes
191
+ - Output ONLY the formatted Markdown, no explanations"""
192
+
193
+ FORMAT_USER_TEMPLATE_WITH_KEYFRAMES = """Here is a raw speech-to-text transcription from a YouTube video titled "{title}".
194
+ Please format it into a clean, readable Markdown document.
195
+
196
+ ---
197
+ TRANSCRIPT:
198
+ {text}
199
+ ---
200
+
201
+ VISUAL KEYFRAME DESCRIPTIONS (timestamped):
202
+ {keyframe_descriptions}
203
+ ---"""
204
+
205
+ KEYFRAME_ANALYSIS_SYSTEM_PROMPT = """You are a visual content analyst. Analyze the provided video keyframes and describe what is shown.
206
+
207
+ Rules:
208
+ - Describe each frame concisely (1-3 sentences)
209
+ - Include any visible text (OCR) exactly as shown
210
+ - Note visual elements: diagrams, charts, code, slides, people, scenes
211
+ - Focus on informational content, not aesthetic quality
212
+ - For each frame, output one line in the format: [MM:SS] description
213
+ - Keep descriptions factual and relevant to the video content
214
+ - Output ONLY the descriptions, no extra commentary"""
215
+
216
+ KEYFRAME_ANALYSIS_USER_TEMPLATE = """Analyze these keyframes from a video. For each image, describe what is shown and extract any visible text.
217
+ The timestamps for each frame are provided as labels."""
218
+
219
+ TRANSLATE_SYSTEM_PROMPT = """You are a professional translator. Translate the given text accurately into {target_lang}.
220
+
221
+ Rules:
222
+ - Maintain the original meaning, tone, and nuance
223
+ - If the text uses Markdown formatting, preserve the Markdown structure
224
+ - Translate naturally and idiomatically, not word-by-word
225
+ - Keep proper nouns, brand names, and technical terms appropriately
226
+ - Do NOT add explanations, notes, or commentary
227
+ - Output ONLY the translated text"""
228
+
229
+ TRANSLATE_USER_TEMPLATE = """Translate the following text into {target_lang}:
230
+
231
+ ---
232
+ {text}
233
+ ---"""
234
+
235
+
236
+ # ── LLM 호출 함수들 (범용) ────────────────────────────────
237
+
238
+ def _call_openai(system_prompt: str, user_prompt: str, api_key: str) -> str:
239
+ """OpenAI GPT-4o-mini 호출."""
240
+ from openai import OpenAI
241
+
242
+ client = OpenAI(api_key=api_key)
243
+ response = client.chat.completions.create(
244
+ model="gpt-4o-mini",
245
+ messages=[
246
+ {"role": "system", "content": system_prompt},
247
+ {"role": "user", "content": user_prompt},
248
+ ],
249
+ temperature=0.3,
250
+ )
251
+ return response.choices[0].message.content
252
+
253
+
254
+ def _call_gemini(system_prompt: str, user_prompt: str, api_key: str,
255
+ model: str = "gemini-2.5-flash-lite") -> str:
256
+ """Google Gemini 호출."""
257
+ import google.generativeai as genai
258
+
259
+ genai.configure(api_key=api_key)
260
+ gmodel = genai.GenerativeModel(model)
261
+
262
+ prompt = system_prompt + "\n\n" + user_prompt
263
+ response = gmodel.generate_content(prompt)
264
+ return response.text
265
+
266
+
267
+ def _call_ollama(system_prompt: str, user_prompt: str, model_name: str = "llama3.2") -> str:
268
+ """Ollama 로컬 모델 호출."""
269
+ import urllib.request
270
+
271
+ payload = json.dumps({
272
+ "model": model_name,
273
+ "messages": [
274
+ {"role": "system", "content": system_prompt},
275
+ {"role": "user", "content": user_prompt},
276
+ ],
277
+ "stream": False,
278
+ "options": {"temperature": 0.3},
279
+ }).encode("utf-8")
280
+
281
+ req = urllib.request.Request(
282
+ "http://localhost:11434/api/chat",
283
+ data=payload,
284
+ headers={"Content-Type": "application/json"},
285
+ )
286
+
287
+ try:
288
+ with urllib.request.urlopen(req, timeout=300) as resp:
289
+ data = json.loads(resp.read().decode("utf-8"))
290
+ return data["message"]["content"]
291
+ except Exception as e:
292
+ if "Connection refused" in str(e):
293
+ raise RuntimeError(
294
+ "Ollama가 실행 중이 아닙니다. "
295
+ "'ollama serve' 명령으로 먼저 시작해주세요. "
296
+ "(설치: brew install ollama && ollama pull llama3.2)"
297
+ )
298
+ raise
299
+
300
+
301
+ def _call_claude(system_prompt: str, user_prompt: str, api_key: str, model: str = "claude-sonnet-4-6") -> str:
302
+ """Anthropic Claude 호출."""
303
+ import anthropic
304
+
305
+ client = anthropic.Anthropic(api_key=api_key)
306
+ response = client.messages.create(
307
+ model=model,
308
+ max_tokens=8192,
309
+ system=system_prompt,
310
+ messages=[
311
+ {"role": "user", "content": user_prompt},
312
+ ],
313
+ )
314
+ return response.content[0].text
315
+
316
+
317
+ # ── Vision LLM 호출 함수들 ───────────────────────────────
318
+
319
+ def _call_gemini_vision(
320
+ system_prompt: str,
321
+ user_prompt: str,
322
+ images: List[dict],
323
+ api_key: str,
324
+ model: str = "gemini-2.5-flash-lite",
325
+ ) -> str:
326
+ """Google Gemini Vision 호출."""
327
+ import google.generativeai as genai
328
+ from PIL import Image
329
+
330
+ genai.configure(api_key=api_key)
331
+ gmodel = genai.GenerativeModel(model)
332
+
333
+ parts = [system_prompt + "\n\n" + user_prompt]
334
+ for img_info in images:
335
+ img = Image.open(img_info["path"])
336
+ parts.append(img)
337
+ parts.append(f"[Timestamp: {img_info['timestamp']}]")
338
+
339
+ response = gmodel.generate_content(parts)
340
+ return response.text
341
+
342
+
343
+ def _call_openai_vision(
344
+ system_prompt: str,
345
+ user_prompt: str,
346
+ images: List[dict],
347
+ api_key: str,
348
+ ) -> str:
349
+ """OpenAI GPT-4o-mini Vision 호출."""
350
+ from openai import OpenAI
351
+
352
+ client = OpenAI(api_key=api_key)
353
+ content = [{"type": "text", "text": user_prompt}]
354
+ for img_info in images:
355
+ with open(img_info["path"], "rb") as f:
356
+ b64 = base64.b64encode(f.read()).decode()
357
+ content.append({
358
+ "type": "image_url",
359
+ "image_url": {"url": f"data:image/jpeg;base64,{b64}"},
360
+ })
361
+ content.append({"type": "text", "text": f"[Timestamp: {img_info['timestamp']}]"})
362
+
363
+ response = client.chat.completions.create(
364
+ model="gpt-4o-mini",
365
+ messages=[
366
+ {"role": "system", "content": system_prompt},
367
+ {"role": "user", "content": content},
368
+ ],
369
+ temperature=0.3,
370
+ )
371
+ return response.choices[0].message.content
372
+
373
+
374
+ def _call_claude_vision(
375
+ system_prompt: str,
376
+ user_prompt: str,
377
+ images: List[dict],
378
+ api_key: str,
379
+ model: str = "claude-sonnet-4-6",
380
+ ) -> str:
381
+ """Anthropic Claude Vision 호출."""
382
+ import anthropic
383
+
384
+ client = anthropic.Anthropic(api_key=api_key)
385
+ content = []
386
+ for img_info in images:
387
+ with open(img_info["path"], "rb") as f:
388
+ b64 = base64.b64encode(f.read()).decode()
389
+ content.append({
390
+ "type": "image",
391
+ "source": {"type": "base64", "media_type": "image/jpeg", "data": b64},
392
+ })
393
+ content.append({"type": "text", "text": f"[Timestamp: {img_info['timestamp']}]"})
394
+ content.append({"type": "text", "text": user_prompt})
395
+
396
+ response = client.messages.create(
397
+ model=model,
398
+ max_tokens=8192,
399
+ system=system_prompt,
400
+ messages=[{"role": "user", "content": content}],
401
+ )
402
+ return response.content[0].text
403
+
404
+
405
+ def _call_vision_llm(
406
+ system_prompt: str,
407
+ user_prompt: str,
408
+ images: List[dict],
409
+ llm_provider: str,
410
+ api_key: Optional[str] = None,
411
+ ) -> str:
412
+ """Vision LLM 호출 디스패처 (Rate limit 자동 재시도 포함)."""
413
+ def _do_call():
414
+ if llm_provider in ("gemini-flash", "gemini-flash-lite"):
415
+ key = api_key or os.environ.get("GOOGLE_API_KEY", "")
416
+ if not key:
417
+ raise ValueError("Google API 키가 필요합니다. (GOOGLE_API_KEY)")
418
+ model_id = LLM_MODELS[llm_provider]["model"]
419
+ return _call_gemini_vision(system_prompt, user_prompt, images, key, model=model_id)
420
+
421
+ elif llm_provider == "gpt-4o-mini":
422
+ key = api_key or os.environ.get("OPENAI_API_KEY", "")
423
+ if not key:
424
+ raise ValueError("OpenAI API 키가 필요합니다.")
425
+ return _call_openai_vision(system_prompt, user_prompt, images, key)
426
+
427
+ elif llm_provider in ("claude-sonnet", "claude-opus"):
428
+ key = api_key or os.environ.get("ANTHROPIC_API_KEY", "")
429
+ if not key:
430
+ raise ValueError("Anthropic API 키가 필요합니다. (ANTHROPIC_API_KEY)")
431
+ model_id = LLM_MODELS[llm_provider]["model"]
432
+ return _call_claude_vision(system_prompt, user_prompt, images, key, model=model_id)
433
+
434
+ elif llm_provider == "ollama":
435
+ raise ValueError("Ollama는 Vision(이미지 분석)을 지원하지 않습니다.")
436
+
437
+ else:
438
+ raise ValueError(f"지원하지 않는 Vision LLM: {llm_provider}")
439
+
440
+ return _retry_on_rate_limit(_do_call)
441
+
442
+
443
+ # ── Rate Limit 재시도 로직 ────────────────────────────────
444
+
445
+ def _is_rate_limit_error(error: Exception) -> bool:
446
+ """429 Rate Limit 에러인지 확인한다."""
447
+ err_str = str(error).lower()
448
+ err_type = type(error).__name__
449
+ return (
450
+ "rate_limit" in err_str
451
+ or "rate limit" in err_str
452
+ or "429" in err_str
453
+ or "resource_exhausted" in err_str
454
+ or "quota" in err_str
455
+ or err_type == "RateLimitError"
456
+ )
457
+
458
+
459
+ def _parse_retry_after(error: Exception) -> float:
460
+ """에러 메시지에서 대기 시간(초)을 추출한다."""
461
+ err_str = str(error)
462
+ # "Please try again in 2.129s" 같은 패턴
463
+ match = re.search(r"try again in (\d+\.?\d*)s", err_str)
464
+ if match:
465
+ return float(match.group(1))
466
+ # "Retry-After: 5" 헤더 패턴
467
+ match = re.search(r"retry.?after:?\s*(\d+)", err_str, re.IGNORECASE)
468
+ if match:
469
+ return float(match.group(1))
470
+ return 0.0
471
+
472
+
473
+ def _retry_on_rate_limit(func, *args, max_retries: int = 3, **kwargs):
474
+ """Rate limit 에러 시 exponential backoff으로 재시도한다."""
475
+ for attempt in range(max_retries + 1):
476
+ try:
477
+ return func(*args, **kwargs)
478
+ except Exception as e:
479
+ if not _is_rate_limit_error(e) or attempt >= max_retries:
480
+ raise
481
+ # 에러에서 대기 시간 추출, 없으면 exponential backoff
482
+ wait = _parse_retry_after(e)
483
+ if wait <= 0:
484
+ wait = (2 ** attempt) * 2 # 2초, 4초, 8초
485
+ wait = min(wait + 0.5, 60) # 여유 0.5초 추가, 최대 60초
486
+ print(f"⏳ Rate limit 초과, {wait:.1f}초 후 재시도... ({attempt + 1}/{max_retries})")
487
+ time.sleep(wait)
488
+
489
+
490
+ # ── 텍스트 LLM 호출 디스패처 ─────────────────────────────
491
+
492
+ def _call_llm(
493
+ system_prompt: str,
494
+ user_prompt: str,
495
+ llm_provider: str,
496
+ api_key: Optional[str] = None,
497
+ ollama_model: str = "llama3.2",
498
+ ) -> str:
499
+ """텍스트 LLM 호출 디스패처 (Rate limit 자동 재시도 포함)."""
500
+ def _do_call():
501
+ if llm_provider == "gpt-4o-mini":
502
+ key = api_key or os.environ.get("OPENAI_API_KEY", "")
503
+ if not key:
504
+ raise ValueError("OpenAI API 키가 필요합니다.")
505
+ return _call_openai(system_prompt, user_prompt, key)
506
+
507
+ elif llm_provider in ("gemini-flash", "gemini-flash-lite"):
508
+ key = api_key or os.environ.get("GOOGLE_API_KEY", "")
509
+ if not key:
510
+ raise ValueError("Google API 키가 필요합니다. (GOOGLE_API_KEY)")
511
+ model_id = LLM_MODELS[llm_provider]["model"]
512
+ return _call_gemini(system_prompt, user_prompt, key, model=model_id)
513
+
514
+ elif llm_provider == "ollama":
515
+ return _call_ollama(system_prompt, user_prompt, ollama_model)
516
+
517
+ elif llm_provider in ("claude-sonnet", "claude-opus"):
518
+ key = api_key or os.environ.get("ANTHROPIC_API_KEY", "")
519
+ if not key:
520
+ raise ValueError("Anthropic API 키가 필요합니다. (ANTHROPIC_API_KEY)")
521
+ model_id = LLM_MODELS[llm_provider]["model"]
522
+ return _call_claude(system_prompt, user_prompt, key, model=model_id)
523
+
524
+ else:
525
+ raise ValueError(f"지원하지 않는 LLM: {llm_provider}")
526
+
527
+ return _retry_on_rate_limit(_do_call)
528
+
529
+
530
+ # ── 키프레임 분석 ─────────────────────────────────────────
531
+
532
+ def analyze_keyframes(
533
+ keyframe_paths: List[dict],
534
+ llm_provider: str = "gemini-flash-lite",
535
+ api_key: Optional[str] = None,
536
+ on_progress: Optional[Callable] = None,
537
+ batch_size: int = 10,
538
+ ) -> str:
539
+ """Vision LLM으로 키프레임 이미지를 분석한다.
540
+
541
+ keyframe_paths: [{"path": str, "timestamp": str}, ...]
542
+ 반환: 타임스탬프별 설명 텍스트
543
+ """
544
+ def _notify(percent: int, detail: str):
545
+ if on_progress:
546
+ on_progress({"step": "keyframe_analysis", "percent": percent, "detail": detail})
547
+
548
+ model_info = LLM_MODELS.get(llm_provider, {})
549
+ model_name = model_info.get("name", llm_provider)
550
+
551
+ total = len(keyframe_paths)
552
+ _notify(5, f"{total}개 키프레임을 {model_name}으로 분석 준비 중...")
553
+
554
+ all_descriptions = []
555
+ batches = [keyframe_paths[i:i + batch_size] for i in range(0, total, batch_size)]
556
+
557
+ for idx, batch in enumerate(batches):
558
+ pct = int(10 + (idx / len(batches)) * 80)
559
+ _notify(pct, f"배치 {idx + 1}/{len(batches)} 분석 중 ({len(batch)}프레임)...")
560
+
561
+ try:
562
+ result = _call_vision_llm(
563
+ system_prompt=KEYFRAME_ANALYSIS_SYSTEM_PROMPT,
564
+ user_prompt=KEYFRAME_ANALYSIS_USER_TEMPLATE,
565
+ images=batch,
566
+ llm_provider=llm_provider,
567
+ api_key=api_key,
568
+ )
569
+ all_descriptions.append(result.strip())
570
+ except Exception as e:
571
+ _notify(pct, f"배치 {idx + 1} 분석 오류: {str(e)}")
572
+ raise
573
+
574
+ _notify(100, f"{total}개 키프레임 분석 완료")
575
+ return "\n".join(all_descriptions)
576
+
577
+
578
+ # ── 메인 함수들 ───────────────────────────────────────────
579
+
580
+ def _truncate_text(text: str, max_chars: int = 100000) -> tuple:
581
+ """텍스트가 너무 길면 잘라낸다. (text, truncated) 반환."""
582
+ if len(text) > max_chars:
583
+ return text[:max_chars], True
584
+ return text, False
585
+
586
+
587
+ def format_as_markdown(
588
+ text: str,
589
+ title: str,
590
+ llm_provider: str = "gpt-4o-mini",
591
+ api_key: Optional[str] = None,
592
+ ollama_model: str = "llama3.2",
593
+ on_progress: Optional[Callable] = None,
594
+ keyframe_descriptions: Optional[str] = None,
595
+ ) -> str:
596
+ """Whisper 추출 텍스트를 LLM으로 마크다운으로 정리한다.
597
+
598
+ keyframe_descriptions가 주어지면 키프레임 설명을 MD에 통합한다.
599
+ """
600
+ def _notify(percent: int, detail: str):
601
+ if on_progress:
602
+ on_progress({"step": "format", "percent": percent, "detail": detail})
603
+
604
+ model_info = LLM_MODELS.get(llm_provider, {})
605
+ model_name = model_info.get("name", llm_provider)
606
+
607
+ has_keyframes = bool(keyframe_descriptions and keyframe_descriptions.strip())
608
+ extra = " + 키프레임 컨텍스트" if has_keyframes else ""
609
+ _notify(10, f"{model_name}에 텍스트{extra} 전송 중...")
610
+
611
+ text, truncated = _truncate_text(text)
612
+ if truncated:
613
+ _notify(15, f"텍스트가 길어서 앞부분만 정리합니다 ({len(text)}자)")
614
+
615
+ _notify(30, f"{model_name} 처리 중...")
616
+
617
+ # 키프레임 설명이 있으면 통합 프롬프트 사용
618
+ if has_keyframes:
619
+ sys_prompt = FORMAT_SYSTEM_PROMPT_WITH_KEYFRAMES
620
+ usr_prompt = FORMAT_USER_TEMPLATE_WITH_KEYFRAMES.format(
621
+ title=title, text=text, keyframe_descriptions=keyframe_descriptions,
622
+ )
623
+ else:
624
+ sys_prompt = FORMAT_SYSTEM_PROMPT
625
+ usr_prompt = FORMAT_USER_TEMPLATE.format(title=title, text=text)
626
+
627
+ try:
628
+ result = _call_llm(
629
+ system_prompt=sys_prompt,
630
+ user_prompt=usr_prompt,
631
+ llm_provider=llm_provider,
632
+ api_key=api_key,
633
+ ollama_model=ollama_model,
634
+ )
635
+ except Exception as e:
636
+ _notify(0, f"LLM 오류: {str(e)}")
637
+ raise
638
+
639
+ if truncated:
640
+ result += "\n\n---\n> ⚠️ 원본 텍스트가 길어서 일부만 정리되었습니다.\n"
641
+
642
+ _notify(100, f"{model_name} 정리 완료")
643
+ return result
644
+
645
+
646
+ def translate_text(
647
+ text: str,
648
+ target_lang: str,
649
+ llm_provider: str = "gpt-4o-mini",
650
+ api_key: Optional[str] = None,
651
+ ollama_model: str = "llama3.2",
652
+ on_progress: Optional[Callable] = None,
653
+ ) -> str:
654
+ """텍스트를 지정 언어로 번역한다."""
655
+ def _notify(percent: int, detail: str):
656
+ if on_progress:
657
+ on_progress({"step": "translate", "percent": percent, "detail": detail})
658
+
659
+ lang_name = LANGUAGES.get(target_lang, target_lang)
660
+ model_info = LLM_MODELS.get(llm_provider, {})
661
+ model_name = model_info.get("name", llm_provider)
662
+
663
+ _notify(10, f"{lang_name}로 번역 준비 중...")
664
+
665
+ text, truncated = _truncate_text(text)
666
+ if truncated:
667
+ _notify(15, f"텍스트가 길어서 앞부분만 번역합니다 ({len(text)}자)")
668
+
669
+ _notify(30, f"{model_name}으로 {lang_name} 번역 중...")
670
+
671
+ try:
672
+ result = _call_llm(
673
+ system_prompt=TRANSLATE_SYSTEM_PROMPT.format(target_lang=lang_name),
674
+ user_prompt=TRANSLATE_USER_TEMPLATE.format(target_lang=lang_name, text=text),
675
+ llm_provider=llm_provider,
676
+ api_key=api_key,
677
+ ollama_model=ollama_model,
678
+ )
679
+ except Exception as e:
680
+ _notify(0, f"번역 오류: {str(e)}")
681
+ raise
682
+
683
+ if truncated:
684
+ result += f"\n\n---\n> ⚠️ 원본 텍스트가 길어서 일부만 번역되었습니다.\n"
685
+
686
+ _notify(100, f"{lang_name} 번역 완료")
687
+ return result
requirements.txt ADDED
@@ -0,0 +1,15 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ pytubefix>=6.0.0
2
+ openai-whisper>=20231117
3
+ openai>=1.0.0
4
+ fastapi>=0.104.0
5
+ uvicorn>=0.24.0
6
+ jinja2>=3.1.0
7
+ python-multipart>=0.0.6
8
+ pydub>=0.25.1
9
+ python-dotenv>=1.0.0
10
+ # MD 포맷팅용 (선택적, 사용하는 LLM만 설치하면 됨)
11
+ google-generativeai>=0.3.0
12
+ anthropic>=0.18.0
13
+ # 키프레임 분석용 (선택적)
14
+ opencv-python>=4.8.0
15
+ Pillow>=10.0.0
run ADDED
@@ -0,0 +1,68 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ #!/bin/bash
2
+ # YouTube Script Extractor 실행 스크립트
3
+
4
+ DIR="$(cd "$(dirname "$0")" && pwd)"
5
+ cd "$DIR"
6
+
7
+ # 첫 실행: .env 파일이 없으면 자동 설정
8
+ if [ ! -f .env ]; then
9
+ echo "=== YouTube Script Extractor 초기 설정 ==="
10
+ echo ""
11
+
12
+ read -p "OpenAI API Key를 입력하세요 (sk-...): " api_key
13
+ if [ -n "$api_key" ]; then
14
+ cat > .env << EOF
15
+ OPENAI_API_KEY=$api_key
16
+ DEFAULT_MODE=api
17
+ DEFAULT_MODEL=base
18
+ PORT=8000
19
+ EOF
20
+ echo ""
21
+ echo ".env 파일이 생성되었습니다."
22
+ else
23
+ cp .env.example .env
24
+ echo ".env.example을 복사했습니다. 나중에 .env 파일을 직접 수정해주세요."
25
+ fi
26
+ echo ""
27
+ fi
28
+
29
+ # ffmpeg 체크
30
+ if ! command -v ffmpeg &>/dev/null; then
31
+ echo "ffmpeg이 필요합니다. 설치합니다..."
32
+ brew install ffmpeg || { echo "ffmpeg 설치 실패. 'brew install ffmpeg'을 직접 실행해주세요."; exit 1; }
33
+ echo ""
34
+ fi
35
+
36
+ # Python 의존성 체크
37
+ if ! python3 -c "import yt_dlp" 2>/dev/null; then
38
+ echo "Python 의존성을 설치합니다..."
39
+ python3 -m pip install -r requirements.txt || { echo "설치 실패. python3와 pip이 설치되어 있는지 확인해주세요."; exit 1; }
40
+ echo ""
41
+ fi
42
+
43
+ # 실행 모드 선택
44
+ case "${1:-}" in
45
+ web|w)
46
+ echo "웹 서버를 시작합니다..."
47
+ python3 web.py
48
+ ;;
49
+ "")
50
+ # URL 없이 실행하면 웹 서버
51
+ echo "웹 서버를 시작합니다..."
52
+ python3 web.py
53
+ ;;
54
+ http*|www.*)
55
+ # URL이 주어지면 CLI 모드
56
+ python3 cli.py "$@"
57
+ ;;
58
+ *)
59
+ echo "사용법:"
60
+ echo " ./run 웹 서버 시작"
61
+ echo " ./run web 웹 서버 시작"
62
+ echo " ./run <YouTube URL> CLI로 스크립트 추출"
63
+ echo ""
64
+ echo "CLI 옵션:"
65
+ echo " ./run <URL> --mode local 로컬 Whisper 사용"
66
+ echo " ./run <URL> --format srt SRT만 생성"
67
+ ;;
68
+ esac
templates/index.html ADDED
@@ -0,0 +1,1448 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ <!DOCTYPE html>
2
+ <html lang="ko">
3
+ <head>
4
+ <meta charset="UTF-8">
5
+ <meta name="viewport" content="width=device-width, initial-scale=1.0">
6
+ <title>YouTube Script Extractor</title>
7
+ <style>
8
+ * { margin: 0; padding: 0; box-sizing: border-box; }
9
+
10
+ body {
11
+ font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif;
12
+ background: #0f0f0f;
13
+ color: #e1e1e1;
14
+ min-height: 100vh;
15
+ }
16
+
17
+ .container {
18
+ max-width: 800px;
19
+ margin: 0 auto;
20
+ padding: 40px 20px;
21
+ }
22
+
23
+ h1 {
24
+ text-align: center;
25
+ font-size: 28px;
26
+ margin-bottom: 8px;
27
+ color: #fff;
28
+ }
29
+
30
+ .subtitle {
31
+ text-align: center;
32
+ color: #888;
33
+ margin-bottom: 40px;
34
+ font-size: 14px;
35
+ }
36
+
37
+ .card {
38
+ background: #1a1a1a;
39
+ border: 1px solid #333;
40
+ border-radius: 12px;
41
+ padding: 24px;
42
+ margin-bottom: 20px;
43
+ }
44
+
45
+ .form-group { margin-bottom: 16px; }
46
+
47
+ label {
48
+ display: block;
49
+ font-size: 13px;
50
+ color: #aaa;
51
+ margin-bottom: 6px;
52
+ font-weight: 500;
53
+ }
54
+
55
+ input[type="text"], select {
56
+ width: 100%;
57
+ padding: 10px 14px;
58
+ background: #0f0f0f;
59
+ border: 1px solid #333;
60
+ border-radius: 8px;
61
+ color: #fff;
62
+ font-size: 14px;
63
+ outline: none;
64
+ transition: border-color 0.2s;
65
+ }
66
+
67
+ input[type="text"]:focus, select:focus { border-color: #ff4444; }
68
+
69
+ .row { display: flex; gap: 12px; }
70
+ .row .form-group { flex: 1; }
71
+
72
+ .checkbox-group {
73
+ display: flex;
74
+ gap: 16px;
75
+ align-items: center;
76
+ }
77
+
78
+ .checkbox-group label {
79
+ display: flex;
80
+ align-items: center;
81
+ gap: 6px;
82
+ cursor: pointer;
83
+ color: #e1e1e1;
84
+ font-size: 14px;
85
+ }
86
+
87
+ input[type="checkbox"] {
88
+ accent-color: #ff4444;
89
+ width: 16px;
90
+ height: 16px;
91
+ }
92
+
93
+ .btn {
94
+ width: 100%;
95
+ padding: 12px;
96
+ background: #ff4444;
97
+ color: #fff;
98
+ border: none;
99
+ border-radius: 8px;
100
+ font-size: 16px;
101
+ font-weight: 600;
102
+ cursor: pointer;
103
+ transition: background 0.2s;
104
+ }
105
+
106
+ .btn:hover { background: #e03030; }
107
+ .btn:disabled {
108
+ background: #333;
109
+ color: #666;
110
+ cursor: not-allowed;
111
+ }
112
+
113
+ .mode-note {
114
+ font-size: 12px;
115
+ color: #666;
116
+ margin-top: 4px;
117
+ }
118
+
119
+ /* ========= 진행 상태 ========= */
120
+ #progress-section { display: none; }
121
+
122
+ .steps { display: flex; flex-direction: column; gap: 0; }
123
+
124
+ .step {
125
+ padding: 16px 0;
126
+ border-bottom: 1px solid #222;
127
+ color: #555;
128
+ transition: color 0.3s;
129
+ }
130
+
131
+ .step:last-child { border-bottom: none; }
132
+ .step.active { color: #fff; }
133
+ .step.done { color: #4caf50; }
134
+ .step.error { color: #ff4444; }
135
+
136
+ .step-header {
137
+ display: flex;
138
+ align-items: center;
139
+ gap: 12px;
140
+ }
141
+
142
+ .step-icon {
143
+ width: 28px;
144
+ height: 28px;
145
+ border-radius: 50%;
146
+ display: flex;
147
+ align-items: center;
148
+ justify-content: center;
149
+ font-size: 14px;
150
+ flex-shrink: 0;
151
+ background: #222;
152
+ color: #555;
153
+ transition: all 0.3s;
154
+ }
155
+
156
+ .step.active .step-icon {
157
+ background: #ff4444;
158
+ color: #fff;
159
+ }
160
+
161
+ .step.done .step-icon {
162
+ background: #4caf50;
163
+ color: #fff;
164
+ }
165
+
166
+ .step-title { font-size: 14px; font-weight: 500; }
167
+
168
+ .step-percent {
169
+ margin-left: auto;
170
+ font-size: 13px;
171
+ font-weight: 600;
172
+ font-family: 'SF Mono', Menlo, monospace;
173
+ color: #555;
174
+ min-width: 40px;
175
+ text-align: right;
176
+ }
177
+
178
+ .step.active .step-percent { color: #ff4444; }
179
+ .step.done .step-percent { color: #4caf50; }
180
+
181
+ /* 프로그레스 바 */
182
+ .step-progress {
183
+ height: 4px;
184
+ background: #222;
185
+ border-radius: 2px;
186
+ margin: 10px 0 0 40px;
187
+ overflow: hidden;
188
+ }
189
+
190
+ .step-progress-bar {
191
+ height: 100%;
192
+ background: #ff4444;
193
+ border-radius: 2px;
194
+ width: 0%;
195
+ transition: width 0.4s ease;
196
+ }
197
+
198
+ .step.done .step-progress-bar {
199
+ background: #4caf50;
200
+ width: 100%;
201
+ }
202
+
203
+ .step-progress-bar.indeterminate {
204
+ width: 30%;
205
+ animation: indeterminate 1.5s ease-in-out infinite;
206
+ }
207
+
208
+ @keyframes indeterminate {
209
+ 0% { margin-left: 0%; }
210
+ 50% { margin-left: 70%; }
211
+ 100% { margin-left: 0%; }
212
+ }
213
+
214
+ /* 상세 정보 */
215
+ .step-detail {
216
+ font-size: 12px;
217
+ color: #555;
218
+ margin: 6px 0 0 40px;
219
+ min-height: 16px;
220
+ transition: color 0.3s;
221
+ }
222
+
223
+ .step.active .step-detail { color: #999; }
224
+ .step.done .step-detail { color: #4caf5099; }
225
+
226
+ /* 타이머 + 팁 */
227
+ .progress-footer {
228
+ display: flex;
229
+ justify-content: space-between;
230
+ align-items: center;
231
+ margin-top: 20px;
232
+ padding-top: 16px;
233
+ border-top: 1px solid #222;
234
+ }
235
+
236
+ .elapsed {
237
+ font-size: 13px;
238
+ color: #888;
239
+ font-family: 'SF Mono', Menlo, monospace;
240
+ }
241
+
242
+ .tip {
243
+ font-size: 12px;
244
+ color: #666;
245
+ font-style: italic;
246
+ max-width: 60%;
247
+ text-align: right;
248
+ animation: fadeInOut 6s ease-in-out infinite;
249
+ }
250
+
251
+ @keyframes fadeInOut {
252
+ 0%, 100% { opacity: 0.4; }
253
+ 50% { opacity: 1; }
254
+ }
255
+
256
+ /* ========= 결과 ========= */
257
+ #result-section { display: none; }
258
+
259
+ .result-banner {
260
+ background: #1a2e1a;
261
+ border: 1px solid #2d5a2d;
262
+ border-radius: 8px;
263
+ padding: 16px;
264
+ margin-bottom: 16px;
265
+ display: flex;
266
+ align-items: center;
267
+ gap: 12px;
268
+ }
269
+
270
+ .result-banner-icon { font-size: 24px; flex-shrink: 0; }
271
+
272
+ .result-banner-text { flex: 1; }
273
+
274
+ .result-banner-title {
275
+ font-size: 15px;
276
+ font-weight: 600;
277
+ color: #4caf50;
278
+ margin-bottom: 4px;
279
+ }
280
+
281
+ .result-banner-info { font-size: 13px; color: #888; }
282
+
283
+ .result-header {
284
+ display: flex;
285
+ justify-content: space-between;
286
+ align-items: center;
287
+ margin-bottom: 12px;
288
+ }
289
+
290
+ .result-title { font-size: 16px; font-weight: 600; color: #fff; }
291
+
292
+ .result-lang {
293
+ font-size: 12px;
294
+ background: #333;
295
+ color: #aaa;
296
+ padding: 4px 10px;
297
+ border-radius: 12px;
298
+ }
299
+
300
+ .result-text {
301
+ background: #0f0f0f;
302
+ border-radius: 8px;
303
+ padding: 16px;
304
+ font-size: 14px;
305
+ line-height: 1.7;
306
+ max-height: 400px;
307
+ overflow-y: auto;
308
+ white-space: pre-wrap;
309
+ word-break: break-word;
310
+ color: #ccc;
311
+ }
312
+
313
+ .download-buttons { display: flex; gap: 10px; margin-top: 16px; }
314
+
315
+ .btn-download {
316
+ flex: 1;
317
+ padding: 12px;
318
+ background: #1a3a1a;
319
+ border: 1px solid #2d5a2d;
320
+ border-radius: 8px;
321
+ color: #4caf50;
322
+ font-size: 14px;
323
+ font-weight: 600;
324
+ cursor: pointer;
325
+ text-align: center;
326
+ text-decoration: none;
327
+ transition: background 0.2s;
328
+ }
329
+
330
+ .btn-download:hover { background: #2d5a2d; color: #fff; }
331
+
332
+ .save-path {
333
+ font-size: 12px;
334
+ color: #666;
335
+ margin-top: 12px;
336
+ font-family: 'SF Mono', Menlo, monospace;
337
+ padding: 8px 12px;
338
+ background: #111;
339
+ border-radius: 6px;
340
+ }
341
+
342
+ .error-msg {
343
+ color: #ff4444;
344
+ padding: 12px;
345
+ font-size: 14px;
346
+ background: #2a1515;
347
+ border-radius: 8px;
348
+ border: 1px solid #4a2020;
349
+ }
350
+
351
+ .btn-new {
352
+ width: 100%;
353
+ padding: 10px;
354
+ background: transparent;
355
+ border: 1px solid #444;
356
+ border-radius: 8px;
357
+ color: #aaa;
358
+ font-size: 14px;
359
+ cursor: pointer;
360
+ margin-top: 12px;
361
+ transition: all 0.2s;
362
+ }
363
+
364
+ .btn-new:hover { border-color: #888; color: #fff; }
365
+
366
+ /* ========= LLM 옵션 (MD/번역 공용) ========= */
367
+ .llm-options {
368
+ display: none;
369
+ margin-top: 12px;
370
+ padding: 16px;
371
+ background: #111;
372
+ border-radius: 8px;
373
+ border: 1px solid #2a2a2a;
374
+ }
375
+
376
+ .llm-options.visible { display: block; }
377
+
378
+ .llm-options .form-group { margin-bottom: 12px; }
379
+ .llm-options .form-group:last-child { margin-bottom: 0; }
380
+
381
+ .llm-select-desc {
382
+ font-size: 12px;
383
+ color: #888;
384
+ margin-top: 4px;
385
+ }
386
+
387
+ /* ========= 설정 버튼 ========= */
388
+ .llm-header { display: flex; gap: 8px; align-items: stretch; }
389
+ .llm-header > div { flex: 1; }
390
+
391
+ .btn-settings {
392
+ padding: 8px 14px;
393
+ background: #222;
394
+ border: 1px solid #333;
395
+ border-radius: 8px;
396
+ color: #aaa;
397
+ font-size: 16px;
398
+ cursor: pointer;
399
+ transition: all 0.2s;
400
+ flex-shrink: 0;
401
+ position: relative;
402
+ }
403
+ .btn-settings:hover { background: #333; color: #fff; border-color: #555; }
404
+ .settings-dot {
405
+ display: none;
406
+ position: absolute;
407
+ top: 4px; right: 4px;
408
+ width: 8px; height: 8px;
409
+ border-radius: 50%;
410
+ background: #4caf50;
411
+ }
412
+ .settings-dot.active { display: block; }
413
+ .workflow-summary {
414
+ font-size: 12px;
415
+ color: #4caf50;
416
+ margin-top: 6px;
417
+ line-height: 1.5;
418
+ }
419
+
420
+ /* ========= 워크플로우 설정 모달 ========= */
421
+ .modal-overlay {
422
+ position: fixed;
423
+ top: 0; left: 0; right: 0; bottom: 0;
424
+ background: rgba(0,0,0,0.7);
425
+ z-index: 1000;
426
+ display: flex;
427
+ align-items: center;
428
+ justify-content: center;
429
+ backdrop-filter: blur(4px);
430
+ }
431
+ .modal {
432
+ background: #1a1a1a;
433
+ border: 1px solid #333;
434
+ border-radius: 16px;
435
+ padding: 28px;
436
+ width: 540px;
437
+ max-width: 90vw;
438
+ max-height: 85vh;
439
+ overflow-y: auto;
440
+ }
441
+ .modal-header {
442
+ display: flex;
443
+ justify-content: space-between;
444
+ align-items: center;
445
+ margin-bottom: 8px;
446
+ }
447
+ .modal-header h3 { color: #fff; font-size: 18px; margin: 0; }
448
+ .modal-close {
449
+ background: none;
450
+ border: none;
451
+ color: #666;
452
+ font-size: 24px;
453
+ cursor: pointer;
454
+ padding: 4px 8px;
455
+ border-radius: 6px;
456
+ transition: all 0.2s;
457
+ line-height: 1;
458
+ }
459
+ .modal-close:hover { background: #333; color: #fff; }
460
+ .modal-desc {
461
+ color: #888;
462
+ font-size: 13px;
463
+ margin-bottom: 20px;
464
+ line-height: 1.5;
465
+ }
466
+ .modal-step-row {
467
+ border: 1px solid #2a2a2a;
468
+ border-radius: 10px;
469
+ padding: 16px;
470
+ margin-bottom: 12px;
471
+ background: #111;
472
+ }
473
+ .modal-step-label {
474
+ display: flex;
475
+ align-items: center;
476
+ gap: 10px;
477
+ margin-bottom: 10px;
478
+ }
479
+ .modal-step-name {
480
+ font-size: 14px;
481
+ font-weight: 600;
482
+ color: #e1e1e1;
483
+ }
484
+ .modal-step-badge {
485
+ font-size: 11px;
486
+ padding: 2px 8px;
487
+ border-radius: 10px;
488
+ background: #222;
489
+ color: #666;
490
+ }
491
+ .modal-step-badge.custom {
492
+ background: #1a3a1a;
493
+ color: #4caf50;
494
+ }
495
+ .modal-step-fields { display: flex; flex-direction: column; gap: 8px; }
496
+ .modal-step-fields select,
497
+ .modal-step-fields input[type="text"] {
498
+ width: 100%;
499
+ padding: 8px 12px;
500
+ background: #0f0f0f;
501
+ border: 1px solid #333;
502
+ border-radius: 6px;
503
+ color: #fff;
504
+ font-size: 13px;
505
+ outline: none;
506
+ }
507
+ .modal-step-fields select:focus,
508
+ .modal-step-fields input[type="text"]:focus { border-color: #ff4444; }
509
+ .modal-step-sub {
510
+ margin-top: 10px;
511
+ padding-top: 10px;
512
+ border-top: 1px solid #222;
513
+ }
514
+ .modal-step-sub label {
515
+ font-size: 12px;
516
+ color: #888;
517
+ margin-bottom: 4px;
518
+ }
519
+ .modal-step-sub select,
520
+ .modal-step-sub input[type="text"] {
521
+ width: 100%;
522
+ padding: 6px 10px;
523
+ background: #0f0f0f;
524
+ border: 1px solid #333;
525
+ border-radius: 6px;
526
+ color: #fff;
527
+ font-size: 12px;
528
+ outline: none;
529
+ margin-bottom: 6px;
530
+ }
531
+ .modal-save-btn {
532
+ margin-top: 8px;
533
+ padding: 10px;
534
+ font-size: 14px;
535
+ }
536
+ </style>
537
+ </head>
538
+ <body>
539
+ <div class="container">
540
+ <h1>YouTube Script Extractor</h1>
541
+ <p class="subtitle">YouTube 영상의 음성을 텍스트로 변환합니다</p>
542
+
543
+ <div class="card" id="form-section">
544
+ <div class="form-group">
545
+ <label>YouTube URL</label>
546
+ <input type="text" id="url" placeholder="https://www.youtube.com/watch?v=..." onkeydown="if(event.key==='Enter') startTranscription()">
547
+ </div>
548
+
549
+ <div class="row">
550
+ <div class="form-group">
551
+ <label>변환 모드</label>
552
+ <select id="mode" onchange="toggleApiKey()">
553
+ <option value="api">Whisper API (빠르고 정확)</option>
554
+ <option value="local">로컬 Whisper (무료)</option>
555
+ </select>
556
+ </div>
557
+ <div class="form-group" id="model-group" style="display:none;">
558
+ <label>모델 크기</label>
559
+ <select id="model_size">
560
+ <option value="tiny">tiny (39MB, 가장 빠름)</option>
561
+ <option value="base" selected>base (74MB, 균형)</option>
562
+ <option value="small">small (244MB, 좋음)</option>
563
+ <option value="medium">medium (769MB, 높음)</option>
564
+ <option value="large">large (1.5GB, 최고)</option>
565
+ </select>
566
+ </div>
567
+ </div>
568
+
569
+ <div class="form-group" id="api-key-group">
570
+ <label>OpenAI API Key</label>
571
+ <input type="text" id="api_key" placeholder="sk-...">
572
+ <p class="mode-note">서버에 OPENAI_API_KEY 환경변수가 설정되어 있으면 비워둬도 됩니다.</p>
573
+ </div>
574
+
575
+ <div class="form-group">
576
+ <label>출력 형식</label>
577
+ <div class="checkbox-group">
578
+ <label><input type="checkbox" id="fmt_txt" checked> TXT (텍스트)</label>
579
+ <label><input type="checkbox" id="fmt_srt" checked> SRT (자막)</label>
580
+ <label><input type="checkbox" id="fmt_md" onchange="toggleLlmOptions()"> MD (AI 정리)</label>
581
+ <label><input type="checkbox" id="fmt_translate" onchange="toggleLlmOptions()"> 번역</label>
582
+ <label><input type="checkbox" id="fmt_keyframes" onchange="toggleLlmOptions()"> 키프레임 분석</label>
583
+ </div>
584
+ </div>
585
+
586
+ <!-- LLM 옵션 (MD/번역 공용) -->
587
+ <div class="llm-options" id="llm-options">
588
+ <div class="form-group">
589
+ <label>기본 LLM <span style="color:#666; font-weight:400;">(모든 AI 단계에 적용)</span></label>
590
+ <div class="llm-header">
591
+ <div>
592
+ <select id="llm-select" onchange="onLLMChange()">
593
+ <option value="">로딩 중...</option>
594
+ </select>
595
+ </div>
596
+ <button type="button" class="btn-settings" id="btn-settings" onclick="openSettingsModal()" title="워크플로우 설정">⚙<span class="settings-dot" id="settings-dot"></span></button>
597
+ </div>
598
+ <p class="llm-select-desc" id="llm-select-desc"></p>
599
+ <p class="workflow-summary" id="workflow-summary" style="display:none;"></p>
600
+ </div>
601
+ <div class="form-group" id="md-api-key-group" style="display:none;">
602
+ <label id="md-api-key-label">LLM API Key</label>
603
+ <input type="text" id="md_api_key" placeholder="API 키 입력...">
604
+ <p class="mode-note" id="md-api-key-note"></p>
605
+ </div>
606
+ <div class="form-group" id="md-ollama-group" style="display:none;">
607
+ <label>Ollama 모델</label>
608
+ <input type="text" id="md_ollama_model" value="llama3.2" placeholder="llama3.2">
609
+ <p class="mode-note">ollama pull &lt;모델명&gt; 으로 먼저 다운로드 필요</p>
610
+ </div>
611
+ <div class="form-group" id="translate-lang-group" style="display:none;">
612
+ <label>번역 언어</label>
613
+ <select id="translate-lang">
614
+ <option value="">로딩 중...</option>
615
+ </select>
616
+ </div>
617
+ </div>
618
+
619
+ <div class="form-group">
620
+ <label>저장 경로</label>
621
+ <input type="text" id="output_dir" placeholder="./output" value="./output">
622
+ <p class="mode-note">비워두면 기본 경로(./output)에 저장됩니다. 절대 경로도 사용 가능합니다 (예: /Users/이름/Desktop)</p>
623
+ </div>
624
+
625
+ <button class="btn" id="submit-btn" onclick="startTranscription()">스크립트 추출 시작</button>
626
+ </div>
627
+
628
+ <!-- 진행 상태 -->
629
+ <div class="card" id="progress-section">
630
+ <div class="steps">
631
+ <div class="step" id="step-download" data-step="download">
632
+ <div class="step-header">
633
+ <div class="step-icon">1</div>
634
+ <div class="step-title">오디오 다운로드</div>
635
+ <div class="step-percent" id="pct-download"></div>
636
+ </div>
637
+ <div class="step-progress"><div class="step-progress-bar" id="bar-download"></div></div>
638
+ <div class="step-detail" id="detail-download">YouTube에서 오디오를 추출합니다</div>
639
+ </div>
640
+ <div class="step" id="step-keyframe_extract" data-step="keyframe_extract" style="display:none;">
641
+ <div class="step-header">
642
+ <div class="step-icon">2</div>
643
+ <div class="step-title">키프레임 추출</div>
644
+ <div class="step-percent" id="pct-keyframe_extract"></div>
645
+ </div>
646
+ <div class="step-progress"><div class="step-progress-bar" id="bar-keyframe_extract"></div></div>
647
+ <div class="step-detail" id="detail-keyframe_extract">영상에서 주요 장면을 추출합니다</div>
648
+ </div>
649
+ <div class="step" id="step-keyframe_analysis" data-step="keyframe_analysis" style="display:none;">
650
+ <div class="step-header">
651
+ <div class="step-icon">3</div>
652
+ <div class="step-title">키프레임 분석</div>
653
+ <div class="step-percent" id="pct-keyframe_analysis"></div>
654
+ </div>
655
+ <div class="step-progress"><div class="step-progress-bar" id="bar-keyframe_analysis"></div></div>
656
+ <div class="step-detail" id="detail-keyframe_analysis">Vision AI로 화면을 분석합니다</div>
657
+ </div>
658
+ <div class="step" id="step-transcribe" data-step="transcribe">
659
+ <div class="step-header">
660
+ <div class="step-icon">2</div>
661
+ <div class="step-title">음성 인식</div>
662
+ <div class="step-percent" id="pct-transcribe"></div>
663
+ </div>
664
+ <div class="step-progress"><div class="step-progress-bar" id="bar-transcribe"></div></div>
665
+ <div class="step-detail" id="detail-transcribe">Whisper로 음성을 텍스트로 변환합니다</div>
666
+ </div>
667
+ <div class="step" id="step-format" data-step="format" style="display:none;">
668
+ <div class="step-header">
669
+ <div class="step-icon">3</div>
670
+ <div class="step-title">AI 마크다운 정리</div>
671
+ <div class="step-percent" id="pct-format"></div>
672
+ </div>
673
+ <div class="step-progress"><div class="step-progress-bar" id="bar-format"></div></div>
674
+ <div class="step-detail" id="detail-format">LLM으로 텍스트를 구조화합니다</div>
675
+ </div>
676
+ <div class="step" id="step-translate" data-step="translate" style="display:none;">
677
+ <div class="step-header">
678
+ <div class="step-icon">4</div>
679
+ <div class="step-title">번역</div>
680
+ <div class="step-percent" id="pct-translate"></div>
681
+ </div>
682
+ <div class="step-progress"><div class="step-progress-bar" id="bar-translate"></div></div>
683
+ <div class="step-detail" id="detail-translate">선택한 언어로 번역합니다</div>
684
+ </div>
685
+ <div class="step" id="step-save" data-step="save">
686
+ <div class="step-header">
687
+ <div class="step-icon" id="icon-save">3</div>
688
+ <div class="step-title">파일 저장</div>
689
+ <div class="step-percent" id="pct-save"></div>
690
+ </div>
691
+ <div class="step-progress"><div class="step-progress-bar" id="bar-save"></div></div>
692
+ <div class="step-detail" id="detail-save">결과 파일을 생성합니다</div>
693
+ </div>
694
+ </div>
695
+ <div class="progress-footer">
696
+ <div class="elapsed" id="elapsed"></div>
697
+ <div class="tip" id="tip"></div>
698
+ </div>
699
+ </div>
700
+
701
+ <!-- 결과 -->
702
+ <div class="card" id="result-section">
703
+ <div class="result-banner">
704
+ <div class="result-banner-icon">&#10003;</div>
705
+ <div class="result-banner-text">
706
+ <div class="result-banner-title">스크립트 추출 완료</div>
707
+ <div class="result-banner-info" id="result-info"></div>
708
+ </div>
709
+ </div>
710
+ <div class="result-header">
711
+ <span class="result-title" id="result-title"></span>
712
+ <span class="result-lang" id="result-lang"></span>
713
+ </div>
714
+ <div class="result-text" id="result-text"></div>
715
+ <div class="download-buttons" id="download-buttons"></div>
716
+ <div class="save-path" id="save-path"></div>
717
+ <button class="btn-new" onclick="resetForm()">새 영상 추출하기</button>
718
+ </div>
719
+
720
+ <!-- 에러 -->
721
+ <div class="card" id="error-section" style="display:none;">
722
+ <div class="error-msg" id="error-msg"></div>
723
+ <button class="btn-new" onclick="resetForm()">다시 시도하기</button>
724
+ </div>
725
+ </div>
726
+
727
+ <!-- 워크플로우 설정 모달 -->
728
+ <div class="modal-overlay" id="settings-modal" style="display:none;" onclick="if(event.target===this)closeSettingsModal()">
729
+ <div class="modal">
730
+ <div class="modal-header">
731
+ <h3>워크플로우 설정</h3>
732
+ <button class="modal-close" onclick="closeSettingsModal()">&times;</button>
733
+ </div>
734
+ <p class="modal-desc">각 단계별로 다른 LLM을 지정할 수 있습니다.<br>비워두면 기본 LLM(<strong id="modal-global-llm-name" style="color:#e1e1e1;">—</strong>)을 사용합니다.</p>
735
+
736
+ <!-- MD 정리 -->
737
+ <div class="modal-step-row">
738
+ <div class="modal-step-label">
739
+ <span class="modal-step-name">MD 정리</span>
740
+ <span class="modal-step-badge" id="badge-format">기본 LLM</span>
741
+ </div>
742
+ <div class="modal-step-fields">
743
+ <select id="modal-format-llm" onchange="onModalLLMChange('format')">
744
+ <option value="">기본 LLM 사용</option>
745
+ </select>
746
+ <input type="text" id="modal-format-api-key" placeholder="API Key (비워두면 환경변수 사용)" style="display:none;">
747
+ </div>
748
+ </div>
749
+
750
+ <!-- 번역 -->
751
+ <div class="modal-step-row">
752
+ <div class="modal-step-label">
753
+ <span class="modal-step-name">번역</span>
754
+ <span class="modal-step-badge" id="badge-translate">기본 LLM</span>
755
+ </div>
756
+ <div class="modal-step-fields">
757
+ <select id="modal-translate-llm" onchange="onModalLLMChange('translate')">
758
+ <option value="">기본 LLM 사용</option>
759
+ </select>
760
+ <input type="text" id="modal-translate-api-key" placeholder="API Key (비워두면 환경변수 사용)" style="display:none;">
761
+ </div>
762
+ </div>
763
+
764
+ <!-- 키프레임 분석 -->
765
+ <div class="modal-step-row">
766
+ <div class="modal-step-label">
767
+ <span class="modal-step-name">키프레임 분석</span>
768
+ <span class="modal-step-badge" id="badge-keyframe">기본 LLM</span>
769
+ </div>
770
+ <div class="modal-step-fields">
771
+ <select id="modal-keyframe-llm" onchange="onModalLLMChange('keyframe')">
772
+ <option value="">기본 LLM 사용 (Vision 지원 필요)</option>
773
+ </select>
774
+ <input type="text" id="modal-keyframe-api-key" placeholder="API Key (비워두면 환경변수 사용)" style="display:none;">
775
+ </div>
776
+ <div class="modal-step-sub">
777
+ <label>추출 방식</label>
778
+ <select id="keyframe-method" onchange="onKeyframeMethodChange()">
779
+ <option value="scene">장면 전환 감지 (추천)</option>
780
+ <option value="interval">고정 간격</option>
781
+ </select>
782
+ <div id="keyframe-interval-group" style="display:none;">
783
+ <label>간격 (초)</label>
784
+ <input type="text" id="keyframe-interval" value="30" placeholder="30">
785
+ </div>
786
+ </div>
787
+ </div>
788
+
789
+ <button class="btn modal-save-btn" onclick="saveModalSettings()">설정 저장</button>
790
+ </div>
791
+ </div>
792
+
793
+ <script>
794
+ let timerInterval = null;
795
+ let startTime = null;
796
+ let tipInterval = null;
797
+ let selectedLLM = 'gpt-4o-mini';
798
+ let llmModels = [];
799
+ let languages = [];
800
+ let llmLoaded = false;
801
+ let langLoaded = false;
802
+ let visionModels = [];
803
+ let visionLoaded = false;
804
+
805
+ // 단계별 LLM 설정 (모달에서 관리)
806
+ let perStepConfig = {
807
+ format: { llm: '', apiKey: '' },
808
+ translate: { llm: '', apiKey: '' },
809
+ keyframe: { llm: '', apiKey: '', method: 'scene', interval: 30 },
810
+ };
811
+
812
+ const TIPS = [
813
+ "Whisper는 90개 이상의 언어를 자동으로 감지합니다",
814
+ "SRT 파일은 대부분의 동영상 플레이어에서 자막으로 사용할 수 있습니다",
815
+ "영상이 길수록 변환 시간이 오래 걸립니다",
816
+ "API 모드는 로컬 모드보다 보통 5~10배 빠릅니다",
817
+ "TXT 파일은 텍스트 편집기에서 바로 열 수 있습니다",
818
+ "추출된 스크립트는 ./output 폴더에 저장됩니다",
819
+ "로컬 모드는 인터넷 없이도 사용 가능합니다",
820
+ "Whisper API 비용은 1분당 약 $0.006입니다",
821
+ "MD 출력은 LLM이 텍스트를 깔끔하게 구조화합니다",
822
+ "Ollama는 로컬에서 무료로 LLM을 실행할 수 있습니다",
823
+ "번역 기능으로 원본과 번역본을 동시에 생성할 수 있습니다",
824
+ "키프레임 분석으로 발표 슬라이드나 화면 내용을 MD에 포함할 수 있습니다",
825
+ "워크플로우 설정에서 단계별로 다른 LLM을 지정할 수 있습니다",
826
+ "키프레임 분석에 Gemini Flash를 사용하면 무료입니다",
827
+ ];
828
+
829
+ // 다운로드 버튼 라벨 매핑
830
+ const FMT_LABELS = {
831
+ 'txt': 'TXT',
832
+ 'srt': 'SRT',
833
+ 'md': 'MD',
834
+ };
835
+
836
+ let STEP_ORDER = ['download', 'transcribe', 'save'];
837
+
838
+ function getStepOrder() {
839
+ const useMd = document.getElementById('fmt_md').checked;
840
+ const useTranslate = document.getElementById('fmt_translate').checked;
841
+ const useKeyframes = document.getElementById('fmt_keyframes').checked;
842
+ const order = ['download'];
843
+ if (useKeyframes) {
844
+ order.push('keyframe_extract');
845
+ order.push('keyframe_analysis');
846
+ }
847
+ order.push('transcribe');
848
+ if (useMd) order.push('format');
849
+ if (useTranslate) order.push('translate');
850
+ order.push('save');
851
+ return order;
852
+ }
853
+
854
+ function needsLLM() {
855
+ return document.getElementById('fmt_md').checked ||
856
+ document.getElementById('fmt_translate').checked ||
857
+ document.getElementById('fmt_keyframes').checked;
858
+ }
859
+
860
+ function toggleApiKey() {
861
+ const mode = document.getElementById('mode').value;
862
+ document.getElementById('api-key-group').style.display = mode === 'api' ? '' : 'none';
863
+ document.getElementById('model-group').style.display = mode === 'local' ? '' : 'none';
864
+ }
865
+
866
+ function toggleLlmOptions() {
867
+ const opts = document.getElementById('llm-options');
868
+ const translateChecked = document.getElementById('fmt_translate').checked;
869
+ const keyframesChecked = document.getElementById('fmt_keyframes').checked;
870
+
871
+ if (needsLLM()) {
872
+ opts.classList.add('visible');
873
+ if (!llmLoaded) loadLLMModels();
874
+ if (translateChecked && !langLoaded) loadLanguages();
875
+ if (keyframesChecked && !visionLoaded) loadVisionModels();
876
+ } else {
877
+ opts.classList.remove('visible');
878
+ }
879
+
880
+ // 번역 언어 선택 표시/숨기기
881
+ document.getElementById('translate-lang-group').style.display =
882
+ translateChecked ? '' : 'none';
883
+ }
884
+
885
+ async function loadLLMModels() {
886
+ try {
887
+ const res = await fetch('/api/llm-models?sort=price');
888
+ const data = await res.json();
889
+ llmModels = data.models;
890
+ llmLoaded = true;
891
+ renderLLMDropdown();
892
+ } catch (e) {
893
+ console.error('LLM 모델 로딩 실패:', e);
894
+ }
895
+ }
896
+
897
+ async function loadLanguages() {
898
+ try {
899
+ const res = await fetch('/api/languages');
900
+ const data = await res.json();
901
+ languages = data.languages;
902
+ langLoaded = true;
903
+ renderLanguageDropdown();
904
+ } catch (e) {
905
+ console.error('언어 목록 로딩 실패:', e);
906
+ }
907
+ }
908
+
909
+ function renderLLMDropdown() {
910
+ const select = document.getElementById('llm-select');
911
+ select.innerHTML = '';
912
+ llmModels.forEach(model => {
913
+ const opt = document.createElement('option');
914
+ opt.value = model.id;
915
+ opt.textContent = model.version
916
+ ? `${model.name} [${model.version}]`
917
+ : model.name;
918
+ if (model.id === selectedLLM) opt.selected = true;
919
+ select.appendChild(opt);
920
+ });
921
+ onLLMChange();
922
+ }
923
+
924
+ function renderLanguageDropdown() {
925
+ const select = document.getElementById('translate-lang');
926
+ select.innerHTML = '';
927
+ languages.forEach(lang => {
928
+ const opt = document.createElement('option');
929
+ opt.value = lang.code;
930
+ opt.textContent = `${lang.name} (${lang.code})`;
931
+ select.appendChild(opt);
932
+ });
933
+ }
934
+
935
+ function onLLMChange() {
936
+ const select = document.getElementById('llm-select');
937
+ selectedLLM = select.value;
938
+ const model = llmModels.find(m => m.id === selectedLLM);
939
+ if (!model) return;
940
+
941
+ // 설명 텍스트 업데이트
942
+ document.getElementById('llm-select-desc').textContent = model.description;
943
+
944
+ // API 키 / Ollama 모델 UI 업데이트
945
+ const keyGroup = document.getElementById('md-api-key-group');
946
+ const ollamaGroup = document.getElementById('md-ollama-group');
947
+
948
+ if (selectedLLM === 'ollama') {
949
+ keyGroup.style.display = 'none';
950
+ ollamaGroup.style.display = '';
951
+ } else if (model.needs_key) {
952
+ keyGroup.style.display = '';
953
+ ollamaGroup.style.display = 'none';
954
+ const envKey = model.env_key || '';
955
+ document.getElementById('md-api-key-label').textContent = model.name + ' API Key';
956
+ document.getElementById('md-api-key-note').textContent =
957
+ envKey ? `서버에 ${envKey} 환경변수가 설정되어 있으면 비워둬도 됩니다.` : '';
958
+ } else {
959
+ keyGroup.style.display = 'none';
960
+ ollamaGroup.style.display = 'none';
961
+ }
962
+
963
+ // 워크플로우 요약 업데이트 (기본 LLM 변경 시 "기본 LLM 사용" 뱃지도 갱신)
964
+ updateWorkflowSummary();
965
+ }
966
+
967
+ function startTimer() {
968
+ startTime = Date.now();
969
+ timerInterval = setInterval(() => {
970
+ const sec = Math.floor((Date.now() - startTime) / 1000);
971
+ const min = Math.floor(sec / 60);
972
+ const s = sec % 60;
973
+ document.getElementById('elapsed').textContent =
974
+ min > 0 ? `${min}분 ${s}초 경과` : `${s}초 경과`;
975
+ }, 1000);
976
+ }
977
+
978
+ function stopTimer() {
979
+ if (timerInterval) clearInterval(timerInterval);
980
+ if (tipInterval) clearInterval(tipInterval);
981
+ }
982
+
983
+ function startTips() {
984
+ let tipIdx = Math.floor(Math.random() * TIPS.length);
985
+ const tipEl = document.getElementById('tip');
986
+ tipEl.textContent = TIPS[tipIdx];
987
+ tipInterval = setInterval(() => {
988
+ tipIdx = (tipIdx + 1) % TIPS.length;
989
+ tipEl.style.opacity = '0';
990
+ setTimeout(() => {
991
+ tipEl.textContent = TIPS[tipIdx];
992
+ tipEl.style.opacity = '';
993
+ }, 300);
994
+ }, 6000);
995
+ }
996
+
997
+ function updateStep(step, percent, detail) {
998
+ const stepEl = document.getElementById('step-' + step);
999
+ const barEl = document.getElementById('bar-' + step);
1000
+ const pctEl = document.getElementById('pct-' + step);
1001
+ const detailEl = document.getElementById('detail-' + step);
1002
+
1003
+ if (!stepEl) return;
1004
+
1005
+ // 현재 스텝을 active로
1006
+ const stepIdx = STEP_ORDER.indexOf(step);
1007
+ STEP_ORDER.forEach((s, i) => {
1008
+ const el = document.getElementById('step-' + s);
1009
+ if (!el) return;
1010
+ const icon = el.querySelector('.step-icon');
1011
+ if (i < stepIdx) {
1012
+ el.className = 'step done';
1013
+ icon.innerHTML = '&#10003;';
1014
+ } else if (i === stepIdx) {
1015
+ el.className = 'step active';
1016
+ icon.textContent = String(i + 1);
1017
+ }
1018
+ });
1019
+
1020
+ // 프로그레스 바
1021
+ if (percent < 0) {
1022
+ barEl.className = 'step-progress-bar indeterminate';
1023
+ barEl.style.width = '';
1024
+ pctEl.textContent = '';
1025
+ } else {
1026
+ barEl.className = 'step-progress-bar';
1027
+ barEl.style.width = percent + '%';
1028
+ pctEl.textContent = percent + '%';
1029
+ }
1030
+
1031
+ if (detail) {
1032
+ detailEl.textContent = detail;
1033
+ }
1034
+
1035
+ if (percent >= 100) {
1036
+ stepEl.className = 'step done';
1037
+ stepEl.querySelector('.step-icon').innerHTML = '&#10003;';
1038
+ pctEl.textContent = '100%';
1039
+ }
1040
+ }
1041
+
1042
+ function setAllStepsDone() {
1043
+ STEP_ORDER.forEach(step => {
1044
+ const el = document.getElementById('step-' + step);
1045
+ if (!el) return;
1046
+ el.className = 'step done';
1047
+ el.querySelector('.step-icon').innerHTML = '&#10003;';
1048
+ document.getElementById('bar-' + step).style.width = '100%';
1049
+ document.getElementById('bar-' + step).className = 'step-progress-bar';
1050
+ document.getElementById('pct-' + step).textContent = '100%';
1051
+ });
1052
+ }
1053
+
1054
+ function resetForm() {
1055
+ document.getElementById('form-section').style.display = 'block';
1056
+ document.getElementById('progress-section').style.display = 'none';
1057
+ document.getElementById('result-section').style.display = 'none';
1058
+ document.getElementById('error-section').style.display = 'none';
1059
+ document.getElementById('submit-btn').disabled = false;
1060
+ document.getElementById('submit-btn').textContent = '스크립트 추출 시작';
1061
+ stopTimer();
1062
+ document.getElementById('elapsed').textContent = '';
1063
+ document.getElementById('tip').textContent = '';
1064
+
1065
+ // 선택적 스텝 숨기기
1066
+ document.getElementById('step-format').style.display = 'none';
1067
+ document.getElementById('step-translate').style.display = 'none';
1068
+ document.getElementById('step-keyframe_extract').style.display = 'none';
1069
+ document.getElementById('step-keyframe_analysis').style.display = 'none';
1070
+
1071
+ const allSteps = ['download', 'keyframe_extract', 'keyframe_analysis', 'transcribe', 'format', 'translate', 'save'];
1072
+ allSteps.forEach((step, i) => {
1073
+ const el = document.getElementById('step-' + step);
1074
+ if (!el) return;
1075
+ el.className = 'step';
1076
+ el.querySelector('.step-icon').textContent = String(i + 1);
1077
+ document.getElementById('bar-' + step).style.width = '0%';
1078
+ document.getElementById('bar-' + step).className = 'step-progress-bar';
1079
+ document.getElementById('pct-' + step).textContent = '';
1080
+ });
1081
+ document.getElementById('detail-download').textContent = 'YouTube에서 오디오를 추출합니다';
1082
+ document.getElementById('detail-keyframe_extract').textContent = '영상에서 주요 장면을 추출합니다';
1083
+ document.getElementById('detail-keyframe_analysis').textContent = 'Vision AI로 화면을 분석합니다';
1084
+ document.getElementById('detail-transcribe').textContent = 'Whisper로 음성을 텍스트로 변환합니다';
1085
+ document.getElementById('detail-format').textContent = 'LLM으로 텍스트를 구조화합니다';
1086
+ document.getElementById('detail-translate').textContent = '선택한 언어로 번역합니다';
1087
+ document.getElementById('detail-save').textContent = '결과 파일을 생성합니다';
1088
+ }
1089
+
1090
+ function getDownloadLabel(fmt) {
1091
+ // txt_ko → "TXT (한국어)", md_zhcn → "MD (中文(简体))"
1092
+ if (fmt.includes('_')) {
1093
+ const parts = fmt.split('_');
1094
+ const baseFmt = parts[0].toUpperCase();
1095
+ const langCode = parts.slice(1).join('_');
1096
+ // 언어 코드를 원래 형태로 복원 시도 (zhcn → zh-CN 등)
1097
+ const langEntry = languages.find(l =>
1098
+ l.code.replace('-', '').toLowerCase() === langCode.toLowerCase()
1099
+ );
1100
+ const langName = langEntry ? langEntry.name : langCode.toUpperCase();
1101
+ return `${baseFmt} (${langName})`;
1102
+ }
1103
+ return (FMT_LABELS[fmt] || fmt.toUpperCase());
1104
+ }
1105
+
1106
+ async function startTranscription() {
1107
+ const url = document.getElementById('url').value.trim();
1108
+ if (!url) { alert('YouTube URL을 입력해주세요.'); return; }
1109
+
1110
+ const mode = document.getElementById('mode').value;
1111
+ const apiKey = document.getElementById('api_key').value.trim();
1112
+ const modelSize = document.getElementById('model_size').value;
1113
+ const outputDir = document.getElementById('output_dir').value.trim() || './output';
1114
+ const formats = [];
1115
+ if (document.getElementById('fmt_txt').checked) formats.push('txt');
1116
+ if (document.getElementById('fmt_srt').checked) formats.push('srt');
1117
+ if (document.getElementById('fmt_md').checked) formats.push('md');
1118
+
1119
+ const useTranslate = document.getElementById('fmt_translate').checked;
1120
+ const translateLang = useTranslate ? document.getElementById('translate-lang').value : '';
1121
+ const enableKeyframes = document.getElementById('fmt_keyframes').checked;
1122
+
1123
+ if (formats.length === 0 && !useTranslate) {
1124
+ alert('출력 형식을 하나 이상 선택해주세요.');
1125
+ return;
1126
+ }
1127
+
1128
+ // LLM 파라미터 (MD 또는 번역 사용 시)
1129
+ const useLlm = needsLLM();
1130
+ const mdLlm = useLlm ? selectedLLM : '';
1131
+ const mdApiKey = useLlm ? document.getElementById('md_api_key').value.trim() : '';
1132
+ const mdOllamaModel = useLlm ? document.getElementById('md_ollama_model').value.trim() || 'llama3.2' : 'llama3.2';
1133
+
1134
+ // STEP_ORDER를 동적으로 설정
1135
+ STEP_ORDER = getStepOrder();
1136
+
1137
+ // 선택적 스텝 표시/숨기기 + 아이콘 번호 조정
1138
+ const formatStep = document.getElementById('step-format');
1139
+ const translateStep = document.getElementById('step-translate');
1140
+ const keyExtStep = document.getElementById('step-keyframe_extract');
1141
+ const keyAnalStep = document.getElementById('step-keyframe_analysis');
1142
+ const useMd = formats.includes('md');
1143
+
1144
+ formatStep.style.display = useMd ? '' : 'none';
1145
+ translateStep.style.display = useTranslate ? '' : 'none';
1146
+ keyExtStep.style.display = enableKeyframes ? '' : 'none';
1147
+ keyAnalStep.style.display = enableKeyframes ? '' : 'none';
1148
+
1149
+ // 스텝 번호 재배정
1150
+ STEP_ORDER.forEach((s, i) => {
1151
+ const el = document.getElementById('step-' + s);
1152
+ if (el) el.querySelector('.step-icon').textContent = String(i + 1);
1153
+ });
1154
+
1155
+ const btn = document.getElementById('submit-btn');
1156
+ btn.disabled = true;
1157
+ btn.textContent = '처리 중...';
1158
+
1159
+ document.getElementById('form-section').style.display = 'none';
1160
+ document.getElementById('progress-section').style.display = 'block';
1161
+ document.getElementById('result-section').style.display = 'none';
1162
+ document.getElementById('error-section').style.display = 'none';
1163
+
1164
+ startTimer();
1165
+ startTips();
1166
+
1167
+ try {
1168
+ const res = await fetch('/api/transcribe', {
1169
+ method: 'POST',
1170
+ headers: { 'Content-Type': 'application/json' },
1171
+ body: JSON.stringify({
1172
+ url, mode, api_key: apiKey, model_size: modelSize,
1173
+ formats, output_dir: outputDir,
1174
+ md_llm: mdLlm, md_api_key: mdApiKey, md_ollama_model: mdOllamaModel,
1175
+ translate_lang: translateLang,
1176
+ // Per-step LLM overrides (모달 설정)
1177
+ format_llm: perStepConfig.format.llm,
1178
+ format_api_key: perStepConfig.format.apiKey,
1179
+ translate_llm: perStepConfig.translate.llm,
1180
+ translate_api_key: perStepConfig.translate.apiKey,
1181
+ keyframe_llm: perStepConfig.keyframe.llm,
1182
+ keyframe_api_key: perStepConfig.keyframe.apiKey,
1183
+ // 키프레임 옵션
1184
+ enable_keyframes: enableKeyframes,
1185
+ keyframe_method: perStepConfig.keyframe.method,
1186
+ keyframe_interval: parseInt(perStepConfig.keyframe.interval) || 30,
1187
+ }),
1188
+ });
1189
+ const data = await res.json();
1190
+
1191
+ if (data.error) {
1192
+ showError(data.error);
1193
+ return;
1194
+ }
1195
+
1196
+ const jobId = data.job_id;
1197
+ const eventSource = new EventSource(`/api/stream/${jobId}`);
1198
+
1199
+ eventSource.onmessage = (event) => {
1200
+ const msg = JSON.parse(event.data);
1201
+
1202
+ if (msg.type === 'progress') {
1203
+ if (msg.step && msg.percent !== undefined) {
1204
+ updateStep(msg.step, msg.percent, msg.detail || '');
1205
+ }
1206
+ } else if (msg.type === 'completed') {
1207
+ eventSource.close();
1208
+ stopTimer();
1209
+ setAllStepsDone();
1210
+ setTimeout(() => showResult(msg.result), 500);
1211
+ } else if (msg.type === 'error') {
1212
+ eventSource.close();
1213
+ stopTimer();
1214
+ showError(msg.message);
1215
+ }
1216
+ };
1217
+
1218
+ eventSource.onerror = () => {
1219
+ eventSource.close();
1220
+ stopTimer();
1221
+ showError('서버 연결이 끊어졌습니다. 서버가 실행 중인지 확인해주세요.');
1222
+ };
1223
+ } catch (err) {
1224
+ stopTimer();
1225
+ showError(`요청 실패: ${err.message}`);
1226
+ }
1227
+ }
1228
+
1229
+ function showError(message) {
1230
+ document.getElementById('progress-section').style.display = 'none';
1231
+ document.getElementById('error-section').style.display = 'block';
1232
+ document.getElementById('error-msg').textContent = message;
1233
+ }
1234
+
1235
+ function showResult(result) {
1236
+ document.getElementById('progress-section').style.display = 'none';
1237
+ document.getElementById('result-section').style.display = 'block';
1238
+
1239
+ const elapsed = Math.floor((Date.now() - startTime) / 1000);
1240
+ const min = Math.floor(elapsed / 60);
1241
+ const sec = elapsed % 60;
1242
+ const timeStr = min > 0 ? `${min}분 ${sec}초` : `${sec}초`;
1243
+
1244
+ document.getElementById('result-info').textContent =
1245
+ `감지 언어: ${result.language} | 소요 시간: ${timeStr}`;
1246
+ document.getElementById('result-title').textContent = result.title;
1247
+ document.getElementById('result-lang').textContent = result.language;
1248
+ document.getElementById('result-text').textContent = result.text;
1249
+
1250
+ const downloadBtns = document.getElementById('download-buttons');
1251
+ downloadBtns.innerHTML = '';
1252
+ const filenames = [];
1253
+ const dirParam = result.output_dir_raw ? `?dir=${encodeURIComponent(result.output_dir_raw)}` : '';
1254
+ for (const [fmt, filename] of Object.entries(result.files)) {
1255
+ filenames.push(filename);
1256
+ const a = document.createElement('a');
1257
+ a.href = `/api/download/${encodeURIComponent(filename)}${dirParam}`;
1258
+ a.className = 'btn-download';
1259
+ a.download = filename;
1260
+ a.textContent = `${getDownloadLabel(fmt)} 다운로드`;
1261
+ downloadBtns.appendChild(a);
1262
+ }
1263
+
1264
+ const dir = result.output_dir || './output';
1265
+ document.getElementById('save-path').textContent =
1266
+ '저장 위치: ' + filenames.map(f => dir + '/' + f).join(', ');
1267
+ }
1268
+ // ========= Vision 모델 로딩 =========
1269
+ async function loadVisionModels() {
1270
+ try {
1271
+ const res = await fetch('/api/vision-models?sort=price');
1272
+ const data = await res.json();
1273
+ visionModels = data.models;
1274
+ visionLoaded = true;
1275
+ } catch (e) {
1276
+ console.error('Vision 모델 로딩 실패:', e);
1277
+ }
1278
+ }
1279
+
1280
+ // ========= 워크플로우 설정 모달 =========
1281
+ function openSettingsModal() {
1282
+ // 모달 열기 전 모델 데이터 로딩 확인
1283
+ if (!llmLoaded) loadLLMModels().then(() => populateModalDropdowns());
1284
+ if (!visionLoaded) loadVisionModels().then(() => populateModalDropdowns());
1285
+ populateModalDropdowns();
1286
+ restoreModalFromConfig();
1287
+ // 기본 LLM 이름 표시
1288
+ const globalModel = llmModels.find(m => m.id === selectedLLM);
1289
+ document.getElementById('modal-global-llm-name').textContent =
1290
+ globalModel ? globalModel.name : selectedLLM;
1291
+ document.getElementById('settings-modal').style.display = 'flex';
1292
+ }
1293
+
1294
+ function closeSettingsModal() {
1295
+ document.getElementById('settings-modal').style.display = 'none';
1296
+ }
1297
+
1298
+ function populateModalDropdowns() {
1299
+ const globalModel = llmModels.find(m => m.id === selectedLLM);
1300
+ const globalName = globalModel ? globalModel.name : '기본 LLM';
1301
+
1302
+ // MD 정리 / 번역 드롭다운 (전체 LLM)
1303
+ ['format', 'translate'].forEach(step => {
1304
+ const select = document.getElementById(`modal-${step}-llm`);
1305
+ const currentVal = select.value;
1306
+ select.innerHTML = `<option value="">기본 LLM 사용 (${globalName})</option>`;
1307
+ llmModels.forEach(model => {
1308
+ const opt = document.createElement('option');
1309
+ opt.value = model.id;
1310
+ opt.textContent = model.version
1311
+ ? `${model.name} [${model.version}]`
1312
+ : model.name;
1313
+ select.appendChild(opt);
1314
+ });
1315
+ select.value = currentVal || '';
1316
+ });
1317
+
1318
+ // 키프레임 드롭다운 (Vision 지원 모델만)
1319
+ const kfSelect = document.getElementById('modal-keyframe-llm');
1320
+ const kfVal = kfSelect.value;
1321
+ const isGlobalVision = globalModel && globalModel.supports_vision !== false;
1322
+ kfSelect.innerHTML = isGlobalVision
1323
+ ? `<option value="">기본 LLM 사용 (${globalName})</option>`
1324
+ : `<option value="">선택 필요 (기본 LLM이 Vision 미지원)</option>`;
1325
+ visionModels.forEach(model => {
1326
+ const opt = document.createElement('option');
1327
+ opt.value = model.id;
1328
+ opt.textContent = model.version
1329
+ ? `${model.name} [${model.version}]`
1330
+ : model.name;
1331
+ kfSelect.appendChild(opt);
1332
+ });
1333
+ kfSelect.value = kfVal || '';
1334
+
1335
+ updateModalBadges();
1336
+ }
1337
+
1338
+ function restoreModalFromConfig() {
1339
+ document.getElementById('modal-format-llm').value = perStepConfig.format.llm;
1340
+ document.getElementById('modal-format-api-key').value = perStepConfig.format.apiKey;
1341
+ document.getElementById('modal-translate-llm').value = perStepConfig.translate.llm;
1342
+ document.getElementById('modal-translate-api-key').value = perStepConfig.translate.apiKey;
1343
+ document.getElementById('modal-keyframe-llm').value = perStepConfig.keyframe.llm;
1344
+ document.getElementById('modal-keyframe-api-key').value = perStepConfig.keyframe.apiKey;
1345
+ document.getElementById('keyframe-method').value = perStepConfig.keyframe.method;
1346
+ document.getElementById('keyframe-interval').value = perStepConfig.keyframe.interval;
1347
+
1348
+ onKeyframeMethodChange();
1349
+ ['format', 'translate', 'keyframe'].forEach(s => onModalLLMChange(s));
1350
+ updateModalBadges();
1351
+ }
1352
+
1353
+ function onModalLLMChange(step) {
1354
+ const select = document.getElementById(`modal-${step}-llm`);
1355
+ const apiKeyInput = document.getElementById(`modal-${step}-api-key`);
1356
+ const val = select.value;
1357
+
1358
+ if (val && val !== 'ollama') {
1359
+ // 선택한 모델의 needs_key 확인
1360
+ const allModels = step === 'keyframe' ? visionModels : llmModels;
1361
+ const model = allModels.find(m => m.id === val);
1362
+ if (model && model.needs_key) {
1363
+ apiKeyInput.style.display = '';
1364
+ apiKeyInput.placeholder = model.env_key
1365
+ ? `API Key (${model.env_key} 환경변수가 있으면 생략 가능)`
1366
+ : 'API Key 입력...';
1367
+ } else {
1368
+ apiKeyInput.style.display = 'none';
1369
+ }
1370
+ } else {
1371
+ apiKeyInput.style.display = 'none';
1372
+ }
1373
+ updateModalBadges();
1374
+ }
1375
+
1376
+ function updateModalBadges() {
1377
+ const globalModel = llmModels.find(m => m.id === selectedLLM);
1378
+ const globalName = globalModel ? globalModel.name : '기본';
1379
+
1380
+ ['format', 'translate', 'keyframe'].forEach(step => {
1381
+ const select = document.getElementById(`modal-${step}-llm`);
1382
+ const badge = document.getElementById(`badge-${step}`);
1383
+ if (select.value) {
1384
+ const allModels = step === 'keyframe' ? visionModels : llmModels;
1385
+ const model = allModels.find(m => m.id === select.value);
1386
+ badge.textContent = model ? model.name : select.value;
1387
+ badge.className = 'modal-step-badge custom';
1388
+ } else {
1389
+ badge.textContent = globalName;
1390
+ badge.className = 'modal-step-badge';
1391
+ }
1392
+ });
1393
+ }
1394
+
1395
+ function onKeyframeMethodChange() {
1396
+ const method = document.getElementById('keyframe-method').value;
1397
+ document.getElementById('keyframe-interval-group').style.display =
1398
+ method === 'interval' ? '' : 'none';
1399
+ }
1400
+
1401
+ function saveModalSettings() {
1402
+ perStepConfig.format.llm = document.getElementById('modal-format-llm').value;
1403
+ perStepConfig.format.apiKey = document.getElementById('modal-format-api-key').value.trim();
1404
+ perStepConfig.translate.llm = document.getElementById('modal-translate-llm').value;
1405
+ perStepConfig.translate.apiKey = document.getElementById('modal-translate-api-key').value.trim();
1406
+ perStepConfig.keyframe.llm = document.getElementById('modal-keyframe-llm').value;
1407
+ perStepConfig.keyframe.apiKey = document.getElementById('modal-keyframe-api-key').value.trim();
1408
+ perStepConfig.keyframe.method = document.getElementById('keyframe-method').value;
1409
+ perStepConfig.keyframe.interval = parseInt(document.getElementById('keyframe-interval').value) || 30;
1410
+
1411
+ updateWorkflowSummary();
1412
+ closeSettingsModal();
1413
+ }
1414
+
1415
+ function updateWorkflowSummary() {
1416
+ const dot = document.getElementById('settings-dot');
1417
+ const summary = document.getElementById('workflow-summary');
1418
+
1419
+ // 오버라이드가 있는지 확인
1420
+ const overrides = [];
1421
+ const getModelName = (id, models) => {
1422
+ const m = models.find(x => x.id === id);
1423
+ return m ? m.name : id;
1424
+ };
1425
+
1426
+ if (perStepConfig.format.llm) {
1427
+ overrides.push(`MD정리 → ${getModelName(perStepConfig.format.llm, llmModels)}`);
1428
+ }
1429
+ if (perStepConfig.translate.llm) {
1430
+ overrides.push(`번역 → ${getModelName(perStepConfig.translate.llm, llmModels)}`);
1431
+ }
1432
+ if (perStepConfig.keyframe.llm) {
1433
+ overrides.push(`키프레임 → ${getModelName(perStepConfig.keyframe.llm, visionModels)}`);
1434
+ }
1435
+
1436
+ if (overrides.length > 0) {
1437
+ dot.className = 'settings-dot active';
1438
+ summary.style.display = '';
1439
+ summary.textContent = '⚙ 단계별 설정: ' + overrides.join(' | ');
1440
+ } else {
1441
+ dot.className = 'settings-dot';
1442
+ summary.style.display = 'none';
1443
+ summary.textContent = '';
1444
+ }
1445
+ }
1446
+ </script>
1447
+ </body>
1448
+ </html>
transcriber.py ADDED
@@ -0,0 +1,565 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """YouTube 영상에서 음성을 추출하고 텍스트로 변환하는 핵심 모듈."""
2
+
3
+ from __future__ import annotations
4
+
5
+ import os
6
+ import re
7
+ import shutil
8
+ import tempfile
9
+ from typing import Callable, List, Optional
10
+ from pathlib import Path
11
+
12
+ from pytubefix import YouTube
13
+ from pydub import AudioSegment
14
+
15
+
16
+ def download_audio(
17
+ url: str,
18
+ output_dir: Optional[str] = None,
19
+ on_progress: Optional[Callable] = None,
20
+ ) -> dict:
21
+ """YouTube URL에서 오디오를 다운로드한다."""
22
+ if output_dir is None:
23
+ output_dir = tempfile.mkdtemp()
24
+
25
+ def _notify(percent: int, detail: str):
26
+ if on_progress:
27
+ on_progress({"step": "download", "percent": percent, "detail": detail})
28
+
29
+ _notify(0, "영상 정보를 가져오는 중...")
30
+
31
+ # pytubefix 다운로드 진행 콜백
32
+ def _download_cb(stream, chunk, bytes_remaining):
33
+ total = stream.filesize
34
+ downloaded = total - bytes_remaining
35
+ pct = min(int(downloaded / total * 80), 80) # 0~80%: 다운로드
36
+ dl_mb = downloaded / (1024 * 1024)
37
+ total_mb = total / (1024 * 1024)
38
+ _notify(pct, f"다운로드 중... {dl_mb:.1f}MB / {total_mb:.1f}MB")
39
+
40
+ yt = YouTube(url, on_progress_callback=_download_cb)
41
+ title = yt.title
42
+ duration = yt.length
43
+
44
+ _notify(5, f"'{title}' 오디오 스트림 선택 중...")
45
+
46
+ audio_stream = yt.streams.filter(only_audio=True).order_by("abr").desc().first()
47
+ if not audio_stream:
48
+ raise RuntimeError("오디오 스트림을 찾을 수 없습니다.")
49
+
50
+ # 다운로드 (m4a/webm)
51
+ downloaded = audio_stream.download(output_path=output_dir)
52
+
53
+ # mp3로 변환 (80~100%)
54
+ _notify(85, "MP3로 변환 중...")
55
+ mp3_path = os.path.join(output_dir, Path(downloaded).stem + ".mp3")
56
+ audio = AudioSegment.from_file(downloaded)
57
+ audio.export(mp3_path, format="mp3", bitrate="192k")
58
+
59
+ if downloaded != mp3_path and os.path.exists(downloaded):
60
+ os.remove(downloaded)
61
+
62
+ file_mb = os.path.getsize(mp3_path) / (1024 * 1024)
63
+ _notify(100, f"다운로드 완료 ({file_mb:.1f}MB, {duration}초)")
64
+
65
+ return {"audio_path": mp3_path, "title": title, "duration": duration}
66
+
67
+
68
+ def download_video(
69
+ url: str,
70
+ output_dir: Optional[str] = None,
71
+ on_progress: Optional[Callable] = None,
72
+ max_resolution: int = 720,
73
+ ) -> str:
74
+ """키프레임 추출을 위해 YouTube 영상을 다운로드한다 (720p 이하)."""
75
+ if output_dir is None:
76
+ output_dir = tempfile.mkdtemp()
77
+
78
+ def _notify(percent: int, detail: str):
79
+ if on_progress:
80
+ on_progress({"step": "download", "percent": percent, "detail": detail})
81
+
82
+ _notify(50, "키프레임용 영상 다운로드 중...")
83
+
84
+ def _download_cb(stream, chunk, bytes_remaining):
85
+ total = stream.filesize
86
+ downloaded = total - bytes_remaining
87
+ pct = min(50 + int(downloaded / total * 45), 95)
88
+ dl_mb = downloaded / (1024 * 1024)
89
+ total_mb = total / (1024 * 1024)
90
+ _notify(pct, f"영상 다운로드 중... {dl_mb:.1f}MB / {total_mb:.1f}MB")
91
+
92
+ yt = YouTube(url, on_progress_callback=_download_cb)
93
+
94
+ # progressive 스트림 (오디오+비디오 합본, 작은 파일)
95
+ stream = (
96
+ yt.streams.filter(progressive=True, file_extension="mp4")
97
+ .filter(res=f"{max_resolution}p")
98
+ .first()
99
+ )
100
+ if not stream:
101
+ # 해상도 제한 없이 가장 낮은 progressive 시도
102
+ stream = (
103
+ yt.streams.filter(progressive=True, file_extension="mp4")
104
+ .order_by("resolution")
105
+ .first()
106
+ )
107
+ if not stream:
108
+ raise RuntimeError("영상 스트림을 찾을 수 없습니다.")
109
+
110
+ video_path = stream.download(output_path=output_dir, filename_prefix="video_")
111
+ _notify(98, "영상 다운로드 완료")
112
+ return video_path
113
+
114
+
115
+ def extract_keyframes(
116
+ video_path: str,
117
+ output_dir: Optional[str] = None,
118
+ method: str = "scene",
119
+ interval_seconds: int = 30,
120
+ max_frames: int = 50,
121
+ on_progress: Optional[Callable] = None,
122
+ ) -> List[dict]:
123
+ """영상에서 키프레임을 추출한다 (OpenCV).
124
+
125
+ method="scene": 히스토그램 비교로 장면 전환 감지.
126
+ method="interval": N초 간격으로 추출.
127
+ 반환: [{"path": str, "timestamp": "MM:SS"}, ...]
128
+ """
129
+ import cv2
130
+
131
+ def _notify(percent: int, detail: str):
132
+ if on_progress:
133
+ on_progress({"step": "keyframe_extract", "percent": percent, "detail": detail})
134
+
135
+ if output_dir is None:
136
+ output_dir = tempfile.mkdtemp()
137
+ os.makedirs(output_dir, exist_ok=True)
138
+
139
+ cap = cv2.VideoCapture(video_path)
140
+ if not cap.isOpened():
141
+ raise RuntimeError(f"영상 파일을 열 수 없습니다: {video_path}")
142
+
143
+ fps = cap.get(cv2.CAP_PROP_FPS) or 30.0
144
+ total_frames = int(cap.get(cv2.CAP_PROP_FRAME_COUNT))
145
+ total_seconds = total_frames / fps
146
+
147
+ _notify(5, f"영상 분석 시작 ({total_seconds:.0f}초, {fps:.0f}fps)")
148
+
149
+ keyframes = []
150
+ prev_hist = None
151
+ frame_idx = 0
152
+ threshold = 0.6 # 장면 전환 감지 임계값
153
+
154
+ while True:
155
+ ret, frame = cap.read()
156
+ if not ret:
157
+ break
158
+
159
+ current_seconds = frame_idx / fps
160
+ pct = min(int(5 + (frame_idx / total_frames) * 85), 90)
161
+
162
+ should_save = False
163
+
164
+ if method == "scene":
165
+ # HSV 히스토그램 비교로 장면 전환 감지
166
+ hsv = cv2.cvtColor(frame, cv2.COLOR_BGR2HSV)
167
+ hist = cv2.calcHist([hsv], [0, 1], None, [50, 60], [0, 180, 0, 256])
168
+ cv2.normalize(hist, hist, 0, 1, cv2.NORM_MINMAX)
169
+
170
+ if prev_hist is not None:
171
+ score = cv2.compareHist(prev_hist, hist, cv2.HISTCMP_CORREL)
172
+ if score < threshold:
173
+ should_save = True
174
+ else:
175
+ # 첫 프레임은 항상 저장
176
+ should_save = True
177
+
178
+ prev_hist = hist
179
+ else:
180
+ # interval 모드: N초 간격
181
+ if frame_idx == 0 or (frame_idx % int(fps * interval_seconds) == 0):
182
+ should_save = True
183
+
184
+ if should_save and len(keyframes) < max_frames:
185
+ mins = int(current_seconds // 60)
186
+ secs = int(current_seconds % 60)
187
+ timestamp = f"{mins:02d}:{secs:02d}"
188
+ frame_path = os.path.join(output_dir, f"keyframe_{len(keyframes) + 1:03d}.jpg")
189
+ cv2.imwrite(frame_path, frame, [cv2.IMWRITE_JPEG_QUALITY, 85])
190
+ keyframes.append({"path": frame_path, "timestamp": timestamp})
191
+
192
+ if len(keyframes) % 5 == 0:
193
+ _notify(pct, f"{len(keyframes)}개 키프레임 추출됨 ({timestamp})")
194
+
195
+ if len(keyframes) >= max_frames:
196
+ break
197
+
198
+ frame_idx += 1
199
+
200
+ cap.release()
201
+ _notify(100, f"키프레임 추출 완료: {len(keyframes)}개")
202
+ return keyframes
203
+
204
+
205
+ def _split_audio(audio_path: str, max_size_mb: int = 24) -> List[str]:
206
+ """오디오 파일을 Whisper API 제한(25MB)에 맞게 분할한다."""
207
+ file_size = os.path.getsize(audio_path)
208
+ max_bytes = max_size_mb * 1024 * 1024
209
+
210
+ if file_size <= max_bytes:
211
+ return [audio_path]
212
+
213
+ audio = AudioSegment.from_mp3(audio_path)
214
+ total_ms = len(audio)
215
+
216
+ num_chunks = (file_size // max_bytes) + 1
217
+ chunk_ms = total_ms // num_chunks
218
+
219
+ chunks = []
220
+ base = Path(audio_path)
221
+ for i in range(num_chunks):
222
+ start = i * chunk_ms
223
+ end = min((i + 1) * chunk_ms, total_ms)
224
+ chunk = audio[start:end]
225
+ chunk_path = str(base.parent / f"{base.stem}_part{i}{base.suffix}")
226
+ chunk.export(chunk_path, format="mp3")
227
+ chunks.append(chunk_path)
228
+
229
+ return chunks
230
+
231
+
232
+ def transcribe_api(
233
+ audio_path: str,
234
+ api_key: str,
235
+ on_progress: Optional[Callable] = None,
236
+ ) -> dict:
237
+ """OpenAI Whisper API로 음성을 텍스트로 변환한다."""
238
+ from openai import OpenAI
239
+
240
+ def _notify(percent: int, detail: str):
241
+ if on_progress:
242
+ on_progress({"step": "transcribe", "percent": percent, "detail": detail})
243
+
244
+ _notify(0, "오디오 파일 분석 중...")
245
+
246
+ client = OpenAI(api_key=api_key)
247
+ chunks = _split_audio(audio_path)
248
+ total_chunks = len(chunks)
249
+ all_segments = []
250
+ full_text_parts = []
251
+ detected_language = None
252
+ time_offset = 0.0
253
+
254
+ if total_chunks > 1:
255
+ _notify(5, f"파일이 커서 {total_chunks}개로 분할하여 처리합니다")
256
+
257
+ for i, chunk_path in enumerate(chunks):
258
+ chunk_pct_start = int(i / total_chunks * 90)
259
+ chunk_pct_end = int((i + 1) / total_chunks * 90)
260
+
261
+ if total_chunks > 1:
262
+ _notify(chunk_pct_start, f"청크 {i + 1}/{total_chunks} Whisper API 전송 중...")
263
+ else:
264
+ _notify(10, "Whisper API로 전송 중...")
265
+
266
+ with open(chunk_path, "rb") as f:
267
+ response = client.audio.transcriptions.create(
268
+ model="whisper-1",
269
+ file=f,
270
+ response_format="verbose_json",
271
+ timestamp_granularities=["segment"],
272
+ )
273
+
274
+ if detected_language is None:
275
+ detected_language = response.language
276
+
277
+ full_text_parts.append(response.text)
278
+
279
+ if total_chunks > 1:
280
+ _notify(chunk_pct_end, f"청크 {i + 1}/{total_chunks} 완료")
281
+ else:
282
+ _notify(85, "응답 처리 중...")
283
+
284
+ if response.segments:
285
+ for seg in response.segments:
286
+ _get = (lambda k: seg[k]) if isinstance(seg, dict) else (lambda k: getattr(seg, k))
287
+ all_segments.append(
288
+ {
289
+ "start": _get("start") + time_offset,
290
+ "end": _get("end") + time_offset,
291
+ "text": _get("text"),
292
+ }
293
+ )
294
+
295
+ if chunks[-1] != chunk_path:
296
+ audio = AudioSegment.from_mp3(chunk_path)
297
+ time_offset += len(audio) / 1000.0
298
+
299
+ # 분할된 임시 파일 정리
300
+ for chunk_path in chunks:
301
+ if chunk_path != audio_path and os.path.exists(chunk_path):
302
+ os.remove(chunk_path)
303
+
304
+ _notify(100, f"변환 완료 (언어: {detected_language}, 세그먼트: {len(all_segments)}개)")
305
+
306
+ return {
307
+ "text": " ".join(full_text_parts),
308
+ "segments": all_segments,
309
+ "language": detected_language or "unknown",
310
+ }
311
+
312
+
313
+ def transcribe_local(
314
+ audio_path: str,
315
+ model_size: str = "base",
316
+ on_progress: Optional[Callable] = None,
317
+ ) -> dict:
318
+ """로컬 Whisper 모델로 음성을 텍스트로 변환한다."""
319
+ import whisper
320
+
321
+ def _notify(percent: int, detail: str):
322
+ if on_progress:
323
+ on_progress({"step": "transcribe", "percent": percent, "detail": detail})
324
+
325
+ _notify(0, f"Whisper {model_size} 모델 로딩 중...")
326
+ model = whisper.load_model(model_size)
327
+
328
+ _notify(20, "음성 분석 중... (영상 길이에 따라 시간이 걸립니다)")
329
+ result = model.transcribe(audio_path)
330
+
331
+ _notify(90, "결과 정리 중...")
332
+ segments = []
333
+ for seg in result.get("segments", []):
334
+ segments.append(
335
+ {"start": seg["start"], "end": seg["end"], "text": seg["text"]}
336
+ )
337
+
338
+ lang = result.get("language", "unknown")
339
+ _notify(100, f"변환 완료 (언어: {lang}, 세그먼트: {len(segments)}개)")
340
+
341
+ return {
342
+ "text": result["text"],
343
+ "segments": segments,
344
+ "language": lang,
345
+ }
346
+
347
+
348
+ def _format_srt_time(seconds: float) -> str:
349
+ """초를 SRT 타임스탬프 형식으로 변환한다."""
350
+ hours = int(seconds // 3600)
351
+ minutes = int((seconds % 3600) // 60)
352
+ secs = int(seconds % 60)
353
+ millis = int((seconds % 1) * 1000)
354
+ return f"{hours:02d}:{minutes:02d}:{secs:02d},{millis:03d}"
355
+
356
+
357
+ def save_txt(result: dict, output_path: str) -> str:
358
+ """텍스트 파일로 저장한다."""
359
+ with open(output_path, "w", encoding="utf-8") as f:
360
+ f.write(result["text"].strip())
361
+ return output_path
362
+
363
+
364
+ def save_srt(result: dict, output_path: str) -> str:
365
+ """SRT 자막 파일로 저장한다."""
366
+ with open(output_path, "w", encoding="utf-8") as f:
367
+ for i, seg in enumerate(result["segments"], 1):
368
+ f.write(f"{i}\n")
369
+ f.write(
370
+ f"{_format_srt_time(seg['start'])} --> {_format_srt_time(seg['end'])}\n"
371
+ )
372
+ f.write(f"{seg['text'].strip()}\n\n")
373
+ return output_path
374
+
375
+
376
+ def _sanitize_filename(title: str) -> str:
377
+ """파일명에 사용할 수 없는 문자를 제거한다."""
378
+ sanitized = re.sub(r'[<>:"/\\|?*]', "", title)
379
+ sanitized = sanitized.strip(". ")
380
+ return sanitized[:100] if sanitized else "untitled"
381
+
382
+
383
+ def process_video(
384
+ url: str,
385
+ mode: str = "api",
386
+ api_key: Optional[str] = None,
387
+ model_size: str = "base",
388
+ output_dir: str = "./output",
389
+ formats: Optional[List[str]] = None,
390
+ on_progress: Optional[Callable] = None,
391
+ # LLM 옵션 (전역 기본값, 하위 호환)
392
+ md_llm: Optional[str] = None,
393
+ md_api_key: Optional[str] = None,
394
+ md_ollama_model: str = "llama3.2",
395
+ # 번역 옵션
396
+ translate_lang: Optional[str] = None,
397
+ # 단계별 LLM 오버라이드
398
+ format_llm: Optional[str] = None,
399
+ format_api_key: Optional[str] = None,
400
+ translate_llm: Optional[str] = None,
401
+ translate_api_key: Optional[str] = None,
402
+ keyframe_llm: Optional[str] = None,
403
+ keyframe_api_key: Optional[str] = None,
404
+ # 키프레임 옵션
405
+ enable_keyframes: bool = False,
406
+ keyframe_method: str = "scene",
407
+ keyframe_interval: int = 30,
408
+ keyframe_max_frames: int = 50,
409
+ ) -> dict:
410
+ """전체 파이프라인: 다운로드 → (키프레임) → 변환 → (AI 정리) → (번역) → 저장."""
411
+ from formatter import make_llm_config, format_as_markdown, translate_text, analyze_keyframes
412
+
413
+ if formats is None:
414
+ formats = ["txt", "srt"]
415
+
416
+ os.makedirs(output_dir, exist_ok=True)
417
+
418
+ def progress(data):
419
+ if on_progress:
420
+ on_progress(data)
421
+
422
+ # 단계별 LLM 설정 빌드
423
+ llm_cfg = make_llm_config(
424
+ global_llm=md_llm, global_api_key=md_api_key,
425
+ global_ollama_model=md_ollama_model,
426
+ format_llm=format_llm, format_api_key=format_api_key,
427
+ translate_llm=translate_llm, translate_api_key=translate_api_key,
428
+ keyframe_llm=keyframe_llm, keyframe_api_key=keyframe_api_key,
429
+ )
430
+
431
+ # 1. 오디오 다운로드
432
+ audio_info = download_audio(url, on_progress=on_progress)
433
+ audio_path = audio_info["audio_path"]
434
+ title = audio_info["title"]
435
+
436
+ video_path = None
437
+ keyframe_dir = None
438
+
439
+ try:
440
+ # 2. 키프레임용 영상 다운로드 (선택 시)
441
+ if enable_keyframes:
442
+ video_path = download_video(url, on_progress=on_progress)
443
+
444
+ # 3. 키프레임 추출 (선택 시)
445
+ keyframe_descriptions = None
446
+ if enable_keyframes and video_path:
447
+ keyframe_dir = tempfile.mkdtemp()
448
+ keyframes = extract_keyframes(
449
+ video_path, output_dir=keyframe_dir,
450
+ method=keyframe_method,
451
+ interval_seconds=keyframe_interval,
452
+ max_frames=keyframe_max_frames,
453
+ on_progress=on_progress,
454
+ )
455
+
456
+ # 4. 키프레임 Vision LLM 분석
457
+ if keyframes:
458
+ kf_cfg = llm_cfg["keyframe"]
459
+ keyframe_descriptions = analyze_keyframes(
460
+ keyframe_paths=keyframes,
461
+ llm_provider=kf_cfg["llm"],
462
+ api_key=kf_cfg["api_key"],
463
+ on_progress=on_progress,
464
+ )
465
+
466
+ # 5. 음성 → 텍스트 변환
467
+ if mode == "api":
468
+ if not api_key:
469
+ raise ValueError("API 모드에서는 OpenAI API 키가 필요합니다.")
470
+ result = transcribe_api(audio_path, api_key, on_progress=on_progress)
471
+ else:
472
+ result = transcribe_local(audio_path, model_size, on_progress=on_progress)
473
+
474
+ # 6. AI 마크다운 정리 (MD 선택 시)
475
+ md_content = None
476
+ fmt_cfg = llm_cfg["format"]
477
+ if "md" in formats and fmt_cfg["llm"]:
478
+ md_content = format_as_markdown(
479
+ text=result["text"],
480
+ title=title,
481
+ llm_provider=fmt_cfg["llm"],
482
+ api_key=fmt_cfg["api_key"],
483
+ ollama_model=fmt_cfg["ollama_model"],
484
+ on_progress=on_progress,
485
+ keyframe_descriptions=keyframe_descriptions,
486
+ )
487
+
488
+ # 7. 번역 (선택 시)
489
+ translated_txt = None
490
+ translated_md = None
491
+ tr_cfg = llm_cfg["translate"]
492
+ if translate_lang and tr_cfg["llm"]:
493
+ translated_txt = translate_text(
494
+ text=result["text"],
495
+ target_lang=translate_lang,
496
+ llm_provider=tr_cfg["llm"],
497
+ api_key=tr_cfg["api_key"],
498
+ ollama_model=tr_cfg["ollama_model"],
499
+ on_progress=on_progress,
500
+ )
501
+
502
+ if md_content:
503
+ translated_md = translate_text(
504
+ text=md_content,
505
+ target_lang=translate_lang,
506
+ llm_provider=tr_cfg["llm"],
507
+ api_key=tr_cfg["api_key"],
508
+ ollama_model=tr_cfg["ollama_model"],
509
+ on_progress=on_progress,
510
+ )
511
+
512
+ # 8. 파일 저장
513
+ progress({"step": "save", "percent": 0, "detail": "파일 저장 중..."})
514
+ safe_title = _sanitize_filename(title)
515
+ saved_files = {}
516
+
517
+ if "txt" in formats:
518
+ txt_path = os.path.join(output_dir, f"{safe_title}.txt")
519
+ save_txt(result, txt_path)
520
+ saved_files["txt"] = txt_path
521
+
522
+ if "srt" in formats:
523
+ srt_path = os.path.join(output_dir, f"{safe_title}.srt")
524
+ save_srt(result, srt_path)
525
+ saved_files["srt"] = srt_path
526
+
527
+ if "md" in formats and md_content:
528
+ md_path = os.path.join(output_dir, f"{safe_title}.md")
529
+ with open(md_path, "w", encoding="utf-8") as f:
530
+ f.write(md_content)
531
+ saved_files["md"] = md_path
532
+
533
+ # 번역 파일 저장
534
+ if translated_txt:
535
+ lang_suffix = translate_lang.replace("-", "").lower()
536
+ tr_txt_path = os.path.join(output_dir, f"{safe_title}_{lang_suffix}.txt")
537
+ with open(tr_txt_path, "w", encoding="utf-8") as f:
538
+ f.write(translated_txt)
539
+ saved_files[f"txt_{lang_suffix}"] = tr_txt_path
540
+
541
+ if translated_md:
542
+ lang_suffix = translate_lang.replace("-", "").lower()
543
+ tr_md_path = os.path.join(output_dir, f"{safe_title}_{lang_suffix}.md")
544
+ with open(tr_md_path, "w", encoding="utf-8") as f:
545
+ f.write(translated_md)
546
+ saved_files[f"md_{lang_suffix}"] = tr_md_path
547
+
548
+ fmt_str = ", ".join(f.upper() for f in saved_files.keys())
549
+ progress({"step": "save", "percent": 100, "detail": f"{fmt_str} 파일 저장 완료"})
550
+
551
+ return {
552
+ "title": title,
553
+ "language": result["language"],
554
+ "text": result["text"],
555
+ "segments": result["segments"],
556
+ "files": saved_files,
557
+ }
558
+ finally:
559
+ # 임시 파일 정리
560
+ if os.path.exists(audio_path):
561
+ os.remove(audio_path)
562
+ if video_path and os.path.exists(video_path):
563
+ os.remove(video_path)
564
+ if keyframe_dir and os.path.exists(keyframe_dir):
565
+ shutil.rmtree(keyframe_dir, ignore_errors=True)
web.py ADDED
@@ -0,0 +1,237 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ #!/usr/bin/env python3
2
+ """YouTube Script Extractor - Web Server."""
3
+
4
+ import asyncio
5
+ import json
6
+ import os
7
+ import uuid
8
+ from pathlib import Path
9
+
10
+ from dotenv import load_dotenv
11
+
12
+ load_dotenv()
13
+
14
+ from fastapi import FastAPI, Request
15
+ from fastapi.responses import FileResponse, HTMLResponse, StreamingResponse
16
+ from fastapi.staticfiles import StaticFiles
17
+ from fastapi.templating import Jinja2Templates
18
+
19
+ from transcriber import process_video
20
+ from formatter import LLM_MODELS, get_models_sorted, get_languages
21
+
22
+ app = FastAPI(title="YouTube Script Extractor")
23
+ templates = Jinja2Templates(directory="templates")
24
+
25
+ OUTPUT_DIR = "./output"
26
+ os.makedirs(OUTPUT_DIR, exist_ok=True)
27
+
28
+ # 진행 중인 작업 추적
29
+ jobs: dict[str, dict] = {}
30
+
31
+
32
+ @app.get("/", response_class=HTMLResponse)
33
+ async def index(request: Request):
34
+ return templates.TemplateResponse("index.html", {"request": request})
35
+
36
+
37
+ @app.get("/api/llm-models")
38
+ async def get_llm_models(sort: str = "price"):
39
+ """LLM 모델 목록을 정렬하여 반환한다."""
40
+ models = get_models_sorted(sort_by=sort)
41
+ return {"models": models}
42
+
43
+
44
+ @app.get("/api/languages")
45
+ async def get_supported_languages():
46
+ """번역 지원 언어 목록을 반환한다."""
47
+ return {"languages": get_languages()}
48
+
49
+
50
+ @app.get("/api/vision-models")
51
+ async def get_vision_models(sort: str = "price"):
52
+ """Vision 지원 LLM 모델 목록을 반환한다."""
53
+ key = "price_rank" if sort == "price" else "quality_rank"
54
+ models = [
55
+ {"id": k, **v} for k, v in LLM_MODELS.items()
56
+ if v.get("supports_vision", False)
57
+ ]
58
+ return {"models": sorted(models, key=lambda x: x[key])}
59
+
60
+
61
+ @app.post("/api/transcribe")
62
+ async def start_transcription(request: Request):
63
+ body = await request.json()
64
+ url = body.get("url", "").strip()
65
+ mode = body.get("mode", "api")
66
+ api_key = body.get("api_key", "") or os.environ.get("OPENAI_API_KEY", "")
67
+ model_size = body.get("model_size", "base")
68
+ formats = body.get("formats", ["txt", "srt"])
69
+ output_dir = body.get("output_dir", "").strip() or OUTPUT_DIR
70
+
71
+ # LLM 옵션 (전역 기본값)
72
+ md_llm = body.get("md_llm", "")
73
+ md_api_key = body.get("md_api_key", "")
74
+ md_ollama_model = body.get("md_ollama_model", "llama3.2")
75
+ translate_lang = body.get("translate_lang", "")
76
+
77
+ # 단계별 LLM 오버라이드
78
+ format_llm = body.get("format_llm", "")
79
+ format_api_key = body.get("format_api_key", "")
80
+ translate_llm = body.get("translate_llm", "")
81
+ translate_api_key = body.get("translate_api_key", "")
82
+ keyframe_llm = body.get("keyframe_llm", "")
83
+ keyframe_api_key = body.get("keyframe_api_key", "")
84
+
85
+ # 키프레임 옵션
86
+ enable_keyframes = body.get("enable_keyframes", False)
87
+ keyframe_method = body.get("keyframe_method", "scene")
88
+ keyframe_interval = body.get("keyframe_interval", 30)
89
+
90
+ if not url:
91
+ return {"error": "YouTube URL을 입력해주세요."}
92
+
93
+ if mode == "api" and not api_key:
94
+ return {"error": "API 모드에서는 OpenAI API 키가 필요합니다."}
95
+
96
+ job_id = str(uuid.uuid4())[:8]
97
+ jobs[job_id] = {"status": "started", "progress": [], "result": None, "error": None}
98
+
99
+ asyncio.create_task(
100
+ _run_transcription(
101
+ job_id, url, mode, api_key, model_size, formats, output_dir,
102
+ md_llm=md_llm, md_api_key=md_api_key, md_ollama_model=md_ollama_model,
103
+ translate_lang=translate_lang,
104
+ format_llm=format_llm, format_api_key=format_api_key,
105
+ translate_llm=translate_llm, translate_api_key=translate_api_key,
106
+ keyframe_llm=keyframe_llm, keyframe_api_key=keyframe_api_key,
107
+ enable_keyframes=enable_keyframes,
108
+ keyframe_method=keyframe_method, keyframe_interval=keyframe_interval,
109
+ )
110
+ )
111
+
112
+ return {"job_id": job_id}
113
+
114
+
115
+ async def _run_transcription(
116
+ job_id: str,
117
+ url: str,
118
+ mode: str,
119
+ api_key: str,
120
+ model_size: str,
121
+ formats: list[str],
122
+ output_dir: str = OUTPUT_DIR,
123
+ md_llm: str = "",
124
+ md_api_key: str = "",
125
+ md_ollama_model: str = "llama3.2",
126
+ translate_lang: str = "",
127
+ format_llm: str = "",
128
+ format_api_key: str = "",
129
+ translate_llm: str = "",
130
+ translate_api_key: str = "",
131
+ keyframe_llm: str = "",
132
+ keyframe_api_key: str = "",
133
+ enable_keyframes: bool = False,
134
+ keyframe_method: str = "scene",
135
+ keyframe_interval: int = 30,
136
+ ):
137
+ def on_progress(msg: str):
138
+ jobs[job_id]["progress"].append(msg)
139
+
140
+ try:
141
+ result = await asyncio.to_thread(
142
+ process_video,
143
+ url=url,
144
+ mode=mode,
145
+ api_key=api_key,
146
+ model_size=model_size,
147
+ output_dir=output_dir,
148
+ formats=formats,
149
+ on_progress=on_progress,
150
+ md_llm=md_llm or None,
151
+ md_api_key=md_api_key or None,
152
+ md_ollama_model=md_ollama_model,
153
+ translate_lang=translate_lang or None,
154
+ format_llm=format_llm or None,
155
+ format_api_key=format_api_key or None,
156
+ translate_llm=translate_llm or None,
157
+ translate_api_key=translate_api_key or None,
158
+ keyframe_llm=keyframe_llm or None,
159
+ keyframe_api_key=keyframe_api_key or None,
160
+ enable_keyframes=enable_keyframes,
161
+ keyframe_method=keyframe_method,
162
+ keyframe_interval=keyframe_interval,
163
+ )
164
+ jobs[job_id]["status"] = "completed"
165
+ # 절대 경로로 변환하여 저장 위치를 정확히 표시
166
+ abs_output = os.path.abspath(output_dir)
167
+ jobs[job_id]["result"] = {
168
+ "title": result["title"],
169
+ "language": result["language"],
170
+ "text": result["text"],
171
+ "files": {
172
+ fmt: os.path.basename(path) for fmt, path in result["files"].items()
173
+ },
174
+ "output_dir": abs_output,
175
+ "output_dir_raw": output_dir,
176
+ }
177
+ except Exception as e:
178
+ jobs[job_id]["status"] = "error"
179
+ jobs[job_id]["error"] = str(e)
180
+
181
+
182
+ @app.get("/api/status/{job_id}")
183
+ async def get_status(job_id: str):
184
+ job = jobs.get(job_id)
185
+ if not job:
186
+ return {"error": "작업을 찾을 수 없습니다."}
187
+ return job
188
+
189
+
190
+ @app.get("/api/stream/{job_id}")
191
+ async def stream_status(job_id: str):
192
+ async def event_generator():
193
+ seen = 0
194
+ while True:
195
+ job = jobs.get(job_id)
196
+ if not job:
197
+ yield f"data: {json.dumps({'type': 'error', 'message': '작업을 찾을 수 없습니다.'})}\n\n"
198
+ break
199
+
200
+ # 새 진행 메시지 전송
201
+ while seen < len(job["progress"]):
202
+ msg = job["progress"][seen]
203
+ if isinstance(msg, dict):
204
+ yield f"data: {json.dumps({'type': 'progress', **msg})}\n\n"
205
+ else:
206
+ yield f"data: {json.dumps({'type': 'progress', 'message': msg})}\n\n"
207
+ seen += 1
208
+
209
+ if job["status"] == "completed":
210
+ yield f"data: {json.dumps({'type': 'completed', 'result': job['result']})}\n\n"
211
+ break
212
+ elif job["status"] == "error":
213
+ yield f"data: {json.dumps({'type': 'error', 'message': job['error']})}\n\n"
214
+ break
215
+
216
+ await asyncio.sleep(0.5)
217
+
218
+ return StreamingResponse(event_generator(), media_type="text/event-stream")
219
+
220
+
221
+ @app.get("/api/download/{filename}")
222
+ async def download_file(filename: str, dir: str = ""):
223
+ base_dir = dir if dir else OUTPUT_DIR
224
+ file_path = os.path.join(base_dir, filename)
225
+ if not os.path.exists(file_path):
226
+ return {"error": "파일을 찾을 수 없습니다."}
227
+ return FileResponse(file_path, filename=filename)
228
+
229
+
230
+ if __name__ == "__main__":
231
+ import uvicorn
232
+
233
+ port = int(os.environ.get("PORT", 8000))
234
+ print("YouTube Script Extractor 웹 서버")
235
+ print(f" http://localhost:{port}")
236
+ print()
237
+ uvicorn.run(app, host="0.0.0.0", port=port)