Spaces:
Sleeping
title: GLM-OCR REST API
emoji: 📄
colorFrom: blue
colorTo: indigo
sdk: docker
sdk_version: latest
app_file: app.py
pinned: false
Check out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference
GLM-OCR REST API
A high-performance OCR API built with FastAPI and GLM-4V for extracting text, tables, and formulas from documents.
Features
✨ Multi-format Support
- Images: JPG, JPEG, PNG, GIF, BMP, WebP
- Documents: PDF, DOCX
- Automatic conversion to images for processing
🎯 Extraction Types (in order: Table → Text → Formula)
- Table Recognition: Extract structured table data
- Text Recognition: Extract all text content
- Formula Recognition: Extract mathematical formulas
🚀 Advanced Features
- Batch processing (up to 10 files)
- Structured information extraction with JSON schema
- Page-by-page processing for multi-page documents
- CORS enabled for web integration
🔗 Interactive API Documentation
Once deployed, access the interactive Swagger UI to test all endpoints:
Swagger UI (Recommended)
- Local:
http://localhost:7860/docs - Deployed Space:
https://huggingface.co/spaces/your-username/glm-ocr-api/api/docs
ReDoc (Alternative)
- Local:
http://localhost:7860/redoc - Deployed Space:
https://huggingface.co/spaces/your-username/glm-ocr-api/api/redoc
The Swagger UI allows you to:
- ✅ Browse all available endpoints
- ✅ View request/response schemas
- ✅ Test endpoints directly in browser
- ✅ Download OpenAPI specification
Note: This application uses the new Hugging Face router endpoint (
https://router.huggingface.co). The oldapi-inference.huggingface.coendpoint is deprecated.
Setup
Prerequisites
- Python 3.8+
- Hugging Face API Token (for GLM-OCR model access)
Installation
- Clone the repository
git clone <your-repo-url>
cd <repo-name>
- Install dependencies
pip install -r requirements.txt
- Configure environment
# Copy the example env file
cp .env.example .env
# Edit .env and add your Hugging Face token
# HF_API_TOKEN=your_token_here
- Run the server
python app.py
The API will be available at: http://localhost:7860
API Documentation
Base URL
http://localhost:7860
Endpoints
1. Health Check
Check if the API is running.
Endpoint:
GET /health
Response:
{
"status": "healthy",
"timestamp": "2024-01-15T10:30:45.123456",
"model": "PrathameshRaut/DocumentOCR"
}
2. Extract Text/Tables/Formulas
Extract all content from a document (table, text, formula in order).
Endpoint:
POST /api/ocr/extract
Request:
- Content-Type:
multipart/form-data - Parameter:
file(UploadFile)
Supported Formats:
- Images:
.jpg,.jpeg,.png,.gif,.bmp,.webp - PDFs:
.pdf - Documents:
.docx,.doc
cURL Example:
curl -X POST "http://localhost:7860/api/ocr/extract" \
-F "file=@path/to/your/document.pdf"
Python Example:
import requests
url = "http://localhost:7860/api/ocr/extract"
files = {"file": open("document.pdf", "rb")}
response = requests.post(url, files=files)
data = response.json()
# Process results
for page in data["pages"]:
print(f"Page {page['page_number']}:")
print(f"Tables: {page['table']}")
print(f"Text: {page['text']}")
print(f"Formulas: {page['formula']}")
Response Example:
{
"filename": "document.pdf",
"file_type": ".pdf",
"total_pages": 1,
"extraction_timestamp": "2024-01-15T10:30:45.123456",
"pages": [
{
"page_number": 1,
"table": "Extracted table content here...",
"text": "Extracted text content here...",
"formula": "Extracted formula content here..."
}
]
}
3. Extract Structured Information
Extract information based on a custom JSON schema.
Endpoint:
POST /api/ocr/extract-structured
Request:
- Content-Type:
multipart/form-data - Parameters:
file(UploadFile): Document to processschema(string): JSON schema defining what to extract
Schema Example:
{
"id_number": "",
"name": "",
"date_of_birth": "",
"address": {
"street": "",
"city": "",
"state": "",
"zip_code": ""
}
}
cURL Example:
SCHEMA='{"id_number":"","name":"","date_of_birth":""}'
curl -X POST "http://localhost:7860/api/ocr/extract-structured" \
-F "file=@path/to/id_card.jpg" \
-F "schema=$SCHEMA"
Python Example:
import requests
import json
url = "http://localhost:7860/api/ocr/extract-structured"
schema = {
"id_number": "",
"name": "",
"date_of_birth": "",
"address": {
"street": "",
"city": "",
"state": ""
}
}
files = {"file": open("id_card.jpg", "rb")}
data = {"schema": json.dumps(schema)}
response = requests.post(url, files=files, data=data)
result = response.json()
# Access extracted structured data
for page in result["pages"]:
print(page["extracted_data"])
Response Example:
{
"filename": "id_card.jpg",
"file_type": ".jpg",
"total_pages": 1,
"extraction_timestamp": "2024-01-15T10:30:45.123456",
"pages": [
{
"page_number": 1,
"extracted_data": {
"id_number": "12345678",
"name": "John Doe",
"date_of_birth": "1990-01-15",
"address": {
"street": "123 Main St",
"city": "New York",
"state": "NY"
}
}
}
]
}
4. Batch Processing
Process multiple files at once (up to 10 files per request).
Endpoint:
POST /api/ocr/batch
Request:
- Content-Type:
multipart/form-data - Parameter:
files(List[UploadFile])
cURL Example:
curl -X POST "http://localhost:7860/api/ocr/batch" \
-F "files=@document1.pdf" \
-F "files=@document2.jpg" \
-F "files=@document3.docx"
Python Example:
import requests
url = "http://localhost:7860/api/ocr/batch"
files = [
("files", open("document1.pdf", "rb")),
("files", open("document2.jpg", "rb")),
("files", open("document3.docx", "rb"))
]
response = requests.post(url, files=files)
batch_results = response.json()
# Process results for each file
for file_result in batch_results["results"]:
print(f"File: {file_result['filename']}")
print(f"Status: {file_result['status']}")
if file_result['status'] == 'success':
for page in file_result['pages']:
print(f" Page {page['page_number']}:")
print(f" Table: {page['table']}")
print(f" Text: {page['text']}")
print(f" Formula: {page['formula']}")
Response Example:
{
"batch_timestamp": "2024-01-15T10:30:45.123456",
"total_files": 3,
"results": [
{
"filename": "document1.pdf",
"file_type": ".pdf",
"status": "success",
"total_pages": 2,
"pages": [
{
"page_number": 1,
"table": "...",
"text": "...",
"formula": "..."
}
]
},
{
"filename": "document2.jpg",
"status": "error",
"error": "Unsupported file type"
}
]
}
5. API Information
Get details about the API and supported formats.
Endpoint:
GET /api/info
Response:
{
"api_name": "GLM-OCR API",
"version": "1.0.0",
"model": "PrathameshRaut/DocumentOCR",
"supported_formats": {
"images": [".jpg", ".jpeg", ".png", ".gif", ".bmp", ".webp"],
"documents": [".docx", ".doc"],
"pdfs": [".pdf"]
},
"endpoints": {
"extract": "/api/ocr/extract",
"extract_structured": "/api/ocr/extract-structured",
"batch": "/api/ocr/batch",
"health": "/health",
"info": "/api/info"
},
"extraction_order": ["table", "text", "formula"]
}
Extraction Order
The API extracts content in the following order for each page:
- Table - Structured table data
- Text - All text content
- Formula - Mathematical formulas
This order is maintained in all responses.
JavaScript/Frontend Integration
Fetch Example
async function extractOCR(file) {
const formData = new FormData();
formData.append('file', file);
const response = await fetch('http://localhost:7860/api/ocr/extract', {
method: 'POST',
body: formData
});
const data = await response.json();
return data;
}
// Usage
document.getElementById('fileInput').addEventListener('change', async (e) => {
const file = e.target.files[0];
const result = await extractOCR(file);
console.log(result);
});
Axios Example
import axios from 'axios';
async function extractOCR(file) {
const formData = new FormData();
formData.append('file', file);
try {
const response = await axios.post(
'http://localhost:7860/api/ocr/extract',
formData,
{
headers: {
'Content-Type': 'multipart/form-data'
}
}
);
return response.data;
} catch (error) {
console.error('OCR extraction failed:', error);
}
}
Error Handling
The API returns appropriate HTTP status codes:
- 200 - Successful extraction
- 400 - Bad request (invalid file, unsupported format, etc.)
- 500 - Server error
Error Response Example:
{
"detail": "Unsupported file type. Supported: .jpg, .jpeg, .png, .gif, .bmp, .webp, .pdf, .docx, .doc"
}
🔧 Troubleshooting & Debugging
Issue: Empty extraction results (all fields are "")
Solution:
- Test model connection:
curl http://localhost:7860/debug/test-model
This will test if your HF_API_TOKEN is set correctly and show the actual model response format.
- Check HF_API_TOKEN:
# Local development
echo $HF_API_TOKEN
# Should not be empty. If empty, set it:
export HF_API_TOKEN="your_actual_token_here"
- View debug logs:
The API prints
[DEBUG]messages that show:
- Actual model response status codes
- Response JSON structure
- Error messages from the model
Look for these logs to understand what the model is returning.
- Verify token has model access:
- Go to: https://huggingface.co/settings/tokens
- Make sure your token has read access
- Check if you can access the model: https://huggingface.co/PrathameshRaut/DocumentOCR
Issue: Model endpoint not responding
Solution:
- Check the endpoint is correct:
https://api-inference.huggingface.co/models/PrathameshRaut/DocumentOCR
- Test with curl:
curl -X POST "https://api-inference.huggingface.co/models/PrathameshRaut/DocumentOCR" \
-H "Authorization: Bearer YOUR_HF_TOKEN" \
-H "Content-Type: application/json" \
-d '{"inputs": {"image": "data:image/png;base64,...", "text": "Extract text"}}'
- Possible reasons:
- Model is not deployed/not available
- Token doesn't have access to this model
- Model requires a specific request format
Issue: Timeout during extraction
Solution:
- Increase timeout in client code:
response = requests.post(
"http://localhost:7860/api/ocr/extract",
files=files,
timeout=180 # Increased from default 120
)
- File size:
- Large PDFs (100+ pages) take longer
- Keep files under 50MB
- Try with a smaller test file first
Issue: "HF_API_TOKEN not set"
Local development:
export HF_API_TOKEN="hf_xxxxxxxxxxxx"
python app.py
Hugging Face Space:
- Go to Space Settings
- Click "Secrets"
- Add secret:
- Name:
HF_API_TOKEN - Value: Your actual token from https://huggingface.co/settings/tokens
- Name:
Debug Endpoint
Use this endpoint to test and debug:
curl http://localhost:7860/debug/test-model
Response shows:
- If HF_API_TOKEN is set
- Test extraction result
- Actual model response format
- Debugging suggestions
This is very helpful to understand what the model is returning.
Deployment on Hugging Face Spaces
Create a new Space on Hugging Face
Choose "Docker" as the runtime (or "Python" with requirements.txt)
Upload these files:
app.pyrequirements.txt.env.example
In Space settings, add your secret:
- Add
HF_API_TOKENwith your Hugging Face API token
- Add
The Space will auto-deploy and be accessible at:
https://huggingface.co/spaces/your-username/your-space-name
Performance Tips
- For large PDFs: Process page-by-page using the single extract endpoint
- For many files: Use batch processing for efficiency
- For production: Use appropriate timeout values in client code
Troubleshooting
Issue: "CUDA out of memory"
- The model might need GPU. Ensure your Space has GPU enabled.
Issue: "HF_API_TOKEN not found"
- Make sure to set the environment variable in Space secrets
Issue: Slow extraction
- Model inference time depends on document complexity
- PDF extraction is slower than image extraction
- For large batches, consider increasing timeout values
Model Information
Model: PrathameshRaut/DocumentOCR
- Built on GLM-4V
- Supports document parsing and information extraction
- Recognizes text, tables, and formulas
License
MIT License - Feel free to use and modify as needed.
Support
For issues or questions:
- Check the troubleshooting section
- Review error messages in API responses
- Check Hugging Face model documentation
Version History
v1.0.0
- Initial release
- Support for image, PDF, and DOCX files
- Table, text, and formula extraction
- Batch processing
- Structured information extraction