Nexar Quantum Language Classifier

A high-performance programming language classification model developed for the Nexar Quantum Code Analysis Engine.

The model predicts the programming language of a source code snippet using a hybrid ensemble architecture combining a fine-tuned CodeBERT transformer with classical machine learning models.


Model Overview

The classifier combines multiple models to improve prediction accuracy:

  • CodeBERT Transformer
  • XGBoost
  • Random Forest
  • Gradient Boosting
  • Weighted Ensemble

The ensemble leverages transformer-based semantic understanding together with statistical TF-IDF features for robust language identification.


Supported Languages

The model is trained to classify:

  • Python
  • Q#
  • Qiskit
  • OpenQASM
  • Cirq

(The exact list depends on the training dataset.)


Architecture

Source Code
      β”‚
      β–Ό
 β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
 β”‚  CodeBERT    β”‚
 β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
        β”‚
        β–Ό

TF-IDF Features
        β”‚
        β–Ό
 β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
 β”‚  XGBoost     β”‚
 β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
 β”‚ RandomForest β”‚
 β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
 β”‚GradientBoost β”‚
 β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
        β”‚
        β–Ό
 Weighted Ensemble
        β”‚
        β–Ό
 Predicted Language

Training

The model was trained using:

  • Hugging Face Transformers
  • Microsoft CodeBERT
  • Scikit-learn
  • XGBoost
  • PyTorch

Feature extraction:

  • TF-IDF
  • 1–3 gram features
  • 5,000 maximum features

Files

File Description
codebert/ Fine-tuned CodeBERT model
tfidf.pkl TF-IDF Vectorizer
xgboost.pkl XGBoost classifier
random_forest.pkl Random Forest classifier
gradient_boosting.pkl Gradient Boosting classifier
label_encoder.pkl Label encoder
ensemble_weights.json Ensemble weights

Usage

from transformers import AutoTokenizer, AutoModelForSequenceClassification

tokenizer = AutoTokenizer.from_pretrained("yashodhajayasinghe/nexar-quantum-language-classifier")
model = AutoModelForSequenceClassification.from_pretrained("yashodhajayasinghe/nexar-quantum-language-classifier")

Intended Uses

Suitable for:

  • Programming language detection
  • Static code analysis
  • Source code indexing
  • Repository analytics
  • Intelligent developer tools
  • AI-assisted code understanding

Limitations

Performance may decrease for:

  • Extremely short code snippets
  • Mixed-language files
  • Obfuscated or minified code
  • Languages not included in the training dataset

Security Notice

This repository contains several serialized Scikit-learn models (.pkl files).

These files were generated using joblib and contain standard Scikit-learn model objects only.

Python pickle files are inherently executable during deserialization, therefore Hugging Face's automated malware scanner may display heuristic warnings. Users should only load pickle files from trusted sources.


Citation

If you use this model in your research or project, please cite:

Nexar Quantum Language Classifier
Nexar Quantum Code Analysis Engine
2026

License

This project is released under the MIT License.


Author

Yashodha Lasith Jayasinghe

Software Engineer | AI & Machine Learning Developer

GitHub: https://github.com/yashodalasith

Hugging Face: https://huggingface.co/yashodhajayasinghe

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support