Sentence Similarity
sentence-transformers
Safetensors
English
modernbert
feature-extraction
dense
code
code-search
code-retrieval
text-embeddings-inference
Instructions to use Shuu12121/NightJar-large-CodeSearch-Embedding with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- sentence-transformers
How to use Shuu12121/NightJar-large-CodeSearch-Embedding with sentence-transformers:
from sentence_transformers import SentenceTransformer model = SentenceTransformer("Shuu12121/NightJar-large-CodeSearch-Embedding") sentences = [ "Extract json from string with support for '' and None.", "def json_loads(data: Any) -> dict[str, Any]:\n \"\"\"Extract json from string with support for '' and None.\"\"\"\n if not data:\n return {}\n try:\n return json_loads_util(data)\n except json.JSONDecodeError as err:\n raise APIError(\"Invalid json\") from err", "def str2json(v):\n \"\"\"\n convert str to json data\n :param v:\n :return:\n \"\"\"\n try:\n return json.loads(v)\n except:\n return None", "def extract_json_from_string(response_msg: str) -> str:\n \"\"\"\n Attempts to extract JSON (object or array) from within a larger string, not specific to markdown.\n \"\"\"\n json_pattern = re.compile(r\"\\{.*\\}|\\[.*\\]\")\n match = json_pattern.search(response_msg)\n if match:\n return match.group(0)\n\n return response_msg", "def parse_json(val: str):\n \"\"\"\n Parses json if string else return\n \"\"\"\n if isinstance(val, str):\n val = json.loads(val)\n if isinstance(val, dict):\n val = frappe._dict(val)\n return val" ] embeddings = model.encode(sentences) similarities = model.similarity(embeddings, embeddings) print(similarities.shape) # [5, 5] - Notebooks
- Google Colab
- Kaggle
Download tokenizer.json from Shuu12121/NightJar-large-CodeSearch-Embedding: direct link, hf CLI and curl.
- Browser
- Download file 3.52 MB
-
https://huggingface.co/Shuu12121/NightJar-large-CodeSearch-Embedding/resolve/main/tokenizer.json
- Command line
-
hf download hf://Shuu12121/NightJar-large-CodeSearch-Embedding/tokenizer.json
-
curl -L -o tokenizer.json https://huggingface.co/Shuu12121/NightJar-large-CodeSearch-Embedding/resolve/main/tokenizer.json
3.52 MB
File too large to display, you can check the raw version instead.