Text Classification
Transformers
PyTorch
Safetensors
English
bert
nlp
code
classification
text-embeddings-inference
Instructions to use TigerByteCyber/codenlbert-sm with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use TigerByteCyber/codenlbert-sm with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="TigerByteCyber/codenlbert-sm")# Load model directly from transformers import AutoTokenizer, AutoModelForSequenceClassification tokenizer = AutoTokenizer.from_pretrained("TigerByteCyber/codenlbert-sm") model = AutoModelForSequenceClassification.from_pretrained("TigerByteCyber/codenlbert-sm", device_map="auto") - Notebooks
- Google Colab
- Kaggle
| datasets: | |
| - vishnun/CodevsNL | |
| language: | |
| - en | |
| metrics: | |
| - accuracy | |
| library_name: transformers | |
| tags: | |
| - nlp | |
| - code | |
| - classification | |
| ## PreFace | |
| Code vs Natural language classification using bert-small from prajwall, below are the metrics achieved | |
| ## Training Metrics | |
| | Epoch | Training Loss | Validation Loss | Accuracy | | | |
| |------:|--------------:|----------------:|---------:|---| | |
| | 1 | 0.022500 | 0.012705 | 0.997203 | | | |
| | 2 | 0.008700 | 0.013107 | 0.996880 | | | |
| | 3 | 0.002700 | 0.014081 | 0.997633 | | | |
| | 4 | 0.001800 | 0.010666 | 0.997526 | | | |
| | 5 | 0.000900 | 0.010800 | 0.998063 | | | |
| ## More | |
| - Github repo for installable python package: https://github.com/Vishnunkumar | |
| - Space on the extraction of code blocks from screenshots: https://huggingface.co/spaces/vishnun/SnapCode |