Buckets:
| # OCR demo files: Food for Space Flight | |
| Files for trying OCR on NASA's *Food for Space Flight* booklet. | |
| The [matching dataset](https://huggingface.co/datasets/uv-scripts/ocr-demo) | |
| contains seven page images with source-page metadata. | |
| | Path | Contents | | |
| |---|---| | |
| | `original/food-for-space-flight.pdf` | The unmodified nine-page NASA PDF. | | |
| | `demo/food-for-space-flight.pdf` | PDF pages 3-9, matching the seven dataset rows. | | |
| | `pages/` | The same seven page images and metadata. | | |
| Use one prefix as your input: `hf://buckets/uv-scripts/ocr-demo/demo` for the | |
| seven-page PDF, `/original` for the complete document, or `/pages` for images. | |
| Pointing a recursive OCR recipe at the Bucket root would process the same pages | |
| in multiple formats. | |
| ## Source and licence | |
| Source: NASA, *NASA Facts: Food for Space Flight*, NTRS document 19830005543. | |
| - [NASA source record](https://ntrs.nasa.gov/citations/19830005543) | |
| - [Original PDF](https://ntrs.nasa.gov/api/citations/19830005543/downloads/19830005543.pdf) | |
| - [Document metadata](https://ntrs.nasa.gov/api/citations/19830005543) | |
| - [NASA usage guidance](https://www.nasa.gov/nasa-brand-center/images-and-media/) | |
| NASA's document record states **Work of the US Gov. Public Use Permitted**. | |
| The source pages retain that status; the repository's code licence does not | |
| relicense the document. Attribute the source pages to NASA. NASA does not endorse | |
| this demo or verify any generated OCR output. | |
| Pages 3-9 (printed pages 2-8) were rendered at 300 DPI and saved as JPEG at quality | |
| 85. `manifest.json` records the source URLs, file sizes and SHA-256 checksums. | |
| The archival scan has damaged lettering and an imperfect existing OCR layer; | |
| this is demo input, with no verified ground-truth transcription. | |
Xet Storage Details
- Size:
- 1.75 kB
- Xet hash:
- c91b3255200aaba08a16b99446f97ce9a6b03935d5ae4fd7379b5102b2d14ccd
·
Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.