Spaces:
Runtime error
Runtime error
| title: PDF Extractor | |
| emoji: ๐ | |
| colorFrom: blue | |
| colorTo: indigo | |
| sdk: docker | |
| app_port: 7860 | |
| pinned: false | |
| license: mit | |
| short_description: Extract text from typed, scanned, and handwritten PDFs. | |
| # PDF Extractor | |
| A live demo of a local-first PDF text extractor. Drop in a PDF and it returns a | |
| clean, searchable, exportable transcript. Pages that already have a text layer | |
| are read directly; scanned image pages are run through OCR. | |
| This Space runs on a free shared CPU, so OCR is slower than it would be on your | |
| own machine and there can be a short cold-start delay when the Space wakes up. | |
| The first request after a rebuild may also be slow while model weights load. | |
| Source code and instructions for running it locally: | |
| https://github.com/Archie0099/PDF-Extractor | |
| ## Notes | |
| - Typed or born-digital pages are extracted instantly and exactly, with no OCR. | |
| - Scanned pages use PaddleOCR (English and Hindi). | |
| - For handwriting, the optional online OCR path (Google Gemini) is far more | |
| accurate. It is off by default and needs your own free API key, which stays in | |
| your browser. | |
| - Nothing is stored. Uploaded files are held in memory only while a job runs and | |
| are dropped afterward. | |