--- library_name: transformers tags: - text-classification - korean - ocr - screenshot-classification - klue-roberta - capturemate --- # CaptureMate Category Classifier v2 4-Class OCR text-based screenshot category classifier for CaptureMate. This model classifies OCR-extracted screenshot text into one of four app categories: `schedule`, `shopping`, `place`, and `memo`. `unknown` samples are excluded from training. Ambiguous or low-confidence predictions should be handled by the application using a confidence threshold. ## Model Details ### Model Description This model is a fine-tuned KLUE-RoBERTa based text classification model for CaptureMate, an iOS screenshot organization and action recommendation app. The model receives OCR text extracted from screenshots and predicts the most relevant screenshot category. - **Developed by:** CaptureMate - **Model type:** Text classification - **Language(s):** Korean, English, mixed OCR text - **License:** Not specified - **Finetuned from model:** KLUE-RoBERTa base - **Number of labels:** 4 ### Labels | ID | Label | Description | |---:|---|---| | 0 | schedule | Schedule, reservation, ticket, event, or date-related screenshots | | 1 | shopping | Shopping, product, price, payment, or commerce-related screenshots | | 2 | place | Place, map, restaurant, store, travel, or location-related screenshots | | 3 | memo | Text, article, note, content, or general information screenshots | ## Uses ### Direct Use Use this model to classify OCR-extracted screenshot text into CaptureMate categories. ### Downstream Use This model can be used inside a screenshot processing pipeline: ```txt Screenshot -> OCR text extraction -> Text preprocessing -> Category classification -> Category-specific action recommendation ``` ## Out-of-Scope Use This model is not designed to: - Classify original images directly - Perform OCR - Extract structured fields such as dates, prices, addresses, or product names - Predict an `unknown` category directly - Classify content outside the CaptureMate screenshot domain ## Bias, Risks, and Limitations The model is trained on a small CaptureMate-specific screenshot OCR dataset. It may not generalize well to unrelated domains or OCR text from different user behavior patterns. Known limitations: - The model always predicts one of the four labels. - There is no direct `unknown` output. - Very short or noisy OCR text can lead to unstable predictions. - Ambiguous cases such as ticket screenshots, food blog screenshots, or shopping-like memo content may be confused. ### Recommendations Use a confidence threshold at inference time. Example: ```txt if confidence < 0.6: category = "unknown" else: category = predicted_label ``` ## How to Get Started with the Model ```python from transformers import AutoTokenizer, AutoModelForSequenceClassification import torch model_name = "YOUR_USERNAME/capturemate-category-classifier-v2-4class" tokenizer = AutoTokenizer.from_pretrained(model_name) model = AutoModelForSequenceClassification.from_pretrained(model_name) text = "2025.4.7 티켓 수령 공연 예약 정보" inputs = tokenizer( text, return_tensors="pt", truncation=True, max_length=256 ) with torch.no_grad(): outputs = model(**inputs) probs = torch.softmax(outputs.logits, dim=-1) pred_id = probs.argmax(dim=-1).item() confidence = probs[0][pred_id].item() label = model.config.id2label[pred_id] print(label, confidence) ``` ## Training Details ### Training Data The model was trained on OCR text extracted from screenshot images collected for the CaptureMate project. Dataset split: | Split | Samples | |---|---:| | Train | 340 | | Validation | 73 | | Test | 74 | `unknown` samples were removed before creating the final training, validation, and test splits. ### Training Procedure #### Preprocessing - OCR text was extracted from screenshots. - Empty text samples were removed. - Labels outside the four target categories were excluded. - Text was tokenized with the KLUE-RoBERTa tokenizer. - Maximum sequence length: 256 tokens. #### Training Hyperparameters | Hyperparameter | Value | |---|---:| | Learning rate | 1e-5 | | Epochs | 8 | | Train batch size | 16 | | Eval batch size | 16 | | Warmup ratio | 0.1 | | Weight decay | 0.01 | | Best model metric | macro F1 | | Seed | 42 | ## Evaluation ### Testing Data, Factors & Metrics #### Testing Data The test set contains 74 OCR text samples across four categories. #### Metrics The model was evaluated using: - Accuracy - Macro F1 - Macro precision - Macro recall Macro metrics are important because the dataset is not perfectly balanced across categories. ### Results | Metric | Validation | Test | |---|---:|---:| | Accuracy | 98.63% | 94.59% | | Macro F1 | 98.89% | 94.55% | | Macro Precision | 99.07% | 93.91% | | Macro Recall | 98.75% | 95.47% | ### Test Set Per-Class Results | Class | Precision | Recall | F1-score | Support | |---|---:|---:|---:|---:| | schedule | 90.00% | 100.00% | 94.74% | 9 | | shopping | 96.15% | 96.15% | 96.15% | 26 | | place | 94.74% | 100.00% | 97.30% | 18 | | memo | 94.74% | 85.71% | 90.00% | 21 | ### Confusion Matrix Rows are true labels, columns are predicted labels. | True \ Pred | schedule | shopping | place | memo | |---|---:|---:|---:|---:| | schedule | 9 | 0 | 0 | 0 | | shopping | 0 | 25 | 0 | 1 | | place | 0 | 0 | 18 | 0 | | memo | 1 | 1 | 1 | 18 | ## Technical Specifications ### Model Architecture and Objective - Architecture: KLUE-RoBERTa sequence classification - Objective: Single-label multi-class classification - Input: OCR text - Output: One of four CaptureMate categories ### Software - transformers - datasets - torch - scikit-learn ## Model Card Authors CaptureMate ## Model Card Contact CaptureMate project maintainer