Initial publish via td-embeddings

Browse files

Files changed (7) hide show

README.md +186 -0
config.json +25 -0
generation_config.json +6 -0
onnx/model-fp32.onnx +3 -0
special_tokens_map.json +28 -0
tokenizer.json +0 -0
tokenizer_config.json +35 -0

README.md ADDED Viewed

	@@ -0,0 +1,186 @@

+---
+license: apache-2.0
+language:
+  - code
+library_name: transformers
+pipeline_tag: feature-extraction
+base_model: codesage/codesage-small
+tags:
+  - onnx
+  - teradata
+  - byom
+  - embeddings
+  - feature-extraction
+---
+> Read the disclaimer below before using this model.
+----
+# codesage-small -- ONNX for Teradata BYOM
+This repository hosts an **ONNX-converted** version of the upstream
+model [`codesage/codesage-small`](https://huggingface.co/codesage/codesage-small),
+packaged for the Teradata Vantage `mldb.ONNXEmbeddings` BYOM
+function. It is **not** the original PyTorch model -- only the
+inference graph and tokenizer needed for in-database embedding
+generation.
+What's different from upstream:
+- **Format**: ONNX (opset 14, IR version 8 -- BYOM 6+ compatible),
+  produced from the upstream weights with architecture-aware
+  post-processing baked in.
+- **Precision**: dynamic int8 quantization. See the variants table
+  below for what is shipped for this model.
+- **Pooling and post-processing**: this graph emits the raw
+  `sentence_embedding` tensor. Pooling rule is
+  **mean**.
+- **Verification**: every variant's cosine fidelity vs. the
+  upstream PyTorch reference is recorded on a fixed
+  CodeSearchNet sample. Numbers may not generalize
+  to your data.
+## Model details
+| | |
+|---|---|
+| Upstream repo | [`codesage/codesage-small`](https://huggingface.co/codesage/codesage-small) |
+| Architecture | `CodeSage` (encoder) |
+| Parameters | 128,010,240 |
+| Output dimensions | 1024 |
+| Pooling | `mean` |
+| Instruction prefix | no |
+| Max input tokens (advertised) | 2048 |
+| Languages | 9 |
+| License | apache-2.0 |
+| ONNX opset | 14 |
+| ONNX IR version | 8 (BYOM 6+ compatible) |
+<details>
+<summary>Full language list (9)</summary>
+- `c`
+- `c-sharp`
+- `go`
+- `java`
+- `javascript`
+- `typescript`
+- `php`
+- `python`
+- `ruby`
+</details>
+## Quantization variants
+This repository ships the following variants. Quality numbers are
+measured against the upstream PyTorch reference on a fixed
+CodeSearchNet sample. The **Size** column is the
+on-disk size of the ONNX weight file in megabytes (MB, 10^6 bytes).
+| Variant | Size (MB) | p50 cosine | R@1 |
+|---|---|---|---|
+| `fp32` | 512.2 | 1.000000 | — |
+How to read the quality columns:
+- **p50 cosine** is the median cosine similarity between this
+  variant's embeddings and the fp32 ONNX reference, computed over
+  a fixed evaluation set. Higher means closer to the unquantized
+  model; **1.0** is identical.
+- **R@1** is top-1 retrieval consistency: if you use this variant
+  as a search index, R@1 is the fraction of queries that get the
+  same nearest neighbor as the fp32 reference would. Higher is
+  better.
+Notes:
+- **fp32**: full-precision reference. Useful for an accuracy ceiling,
+  but BYOM users almost always want one of the int8 variants for
+  in-database scoring -- they are 3-4x smaller and load much faster.
+## Quickstart: using this model with Teradata BYOM
+Requires Teradata Vantage with **BYOM 6+** (`mldb.ONNXEmbeddings`).
+```python
+import getpass
+import teradataml as tdml
+from huggingface_hub import hf_hub_download
+repo_id   = "Teradata/codesage-small"
+model_id  = "codesage-small"        # arbitrary, used as the BYOM model_id
+onnx_file = "onnx/model-fp32.onnx"
+# 1. Download the ONNX + tokenizer for the chosen variant.
+hf_hub_download(repo_id=repo_id, filename=onnx_file,       local_dir="./")
+hf_hub_download(repo_id=repo_id, filename="tokenizer.json", local_dir="./")
+# 2. Connect to Vantage.
+tdml.create_context(
+    host=input("host: "),
+    username=input("user: "),
+    password=getpass.getpass("password: "),
+)
+# 3. Load model + tokenizer into BYOM tables (one-time per model_id).
+tdml.save_byom(model_id=model_id, model_file=onnx_file,
+               table_name="embeddings_models")
+tdml.save_byom(model_id=model_id, model_file="tokenizer.json",
+               table_name="embeddings_tokenizers")
+```
+Then call `mldb.ONNXEmbeddings` against an input table whose
+`txt` column carries the strings to embed:
+```sql
+SELECT *
+FROM mldb.ONNXEmbeddings(
+    ON (SELECT id, txt FROM your_input_table) AS InputTable
+    ON (SELECT model_id, model FROM embeddings_models
+         WHERE model_id = 'codesage-small') AS ModelTable DIMENSION
+    ON (SELECT model_id, tokenizer FROM embeddings_tokenizers
+         WHERE model_id = 'codesage-small') AS TokenizerTable DIMENSION
+    USING
+        Accumulate('id')
+        ModelOutputTensor('sentence_embedding')
+        OutputFormat('FLOAT32(1024)')
+        OverwriteCachedModel('*')
+) AS t
+ORDER BY id;
+```
+Pooling rule **`mean`** is applied **inside** the converted
+ONNX graph -- the output tensor named above already contains the
+pooled, post-processed embedding vector.
+## Original model attribution
+The original weights and training methodology belong to
+**the CodeSage authors**. Please cite their work, not this
+repository, in academic contexts. The canonical upstream model card
+is at
+[`codesage/codesage-small`](https://huggingface.co/codesage/codesage-small);
+refer to it for benchmarks, training details, intended use, and
+citation information.
+## Reporting issues
+For ONNX-conversion or BYOM-compatibility issues specific to this
+Teradata-converted artifact, please open a **Discussion** on this
+model's Hugging Face page. Questions about the underlying model
+quality, training, or intended use should go to the upstream
+maintainer's model card.
+----
+DISCLAIMER: The content herein ("Content") is provided "AS IS" and is not covered by any Teradata Operations, Inc. and its affiliates ("Teradata") agreements. Its listing here does not constitute certification or endorsement by Teradata.
+To the extent any of the Content contains or is related to any artificial intelligence ("AI") or other language learning models ("Models") that interoperate with the products and services of Teradata, by accessing, bringing, deploying or using such Models, you acknowledge and agree that you are solely responsible for ensuring compliance with all applicable laws, regulations, and restrictions governing the use, deployment, and distribution of AI technologies. This includes, but is not limited to, AI Diffusion Rules, European Union AI Act, AI-related laws and regulations, privacy laws, export controls, and financial or sector-specific regulations.
+While Teradata may provide support, guidance, or assistance in the deployment or implementation of Models to interoperate with Teradata's products and/or services, you remain fully responsible for ensuring that your Models, data, and applications comply with all relevant legal and regulatory obligations. Our assistance does not constitute legal or regulatory approval, and Teradata disclaims any liability arising from non-compliance with applicable laws.
+You must determine the suitability of the Models for any purpose. Given the probabilistic nature of machine learning and modeling, the use of the Models may in some situations result in incorrect output that does not accurately reflect the action generated. You should evaluate the accuracy of any output as appropriate for your use case, including by using human review of the output.

config.json ADDED Viewed

	@@ -0,0 +1,25 @@

+{
+    "_name_or_path": "codesage/codesage-small",
+    "architectures": [
+        "CodeSage"
+    ],
+    "auto_map": {
+        "AutoConfig": "config_codesage.CodeSageConfig",
+        "AutoTokenizer": "tokenization_codesage.CodeSageTokenizer",
+        "AutoModel": "modeling_codesage.CodeSageModel",
+        "AutoModelForMaskedLM": "modeling_codesage.CodeSageForMaskedLM",
+        "AutoModelForSequenceClassification": "modeling_codesage.CodeSageForSequenceClassification"
+    },
+    "activation_function": "gelu_new",
+    "attention_dropout_prob": 0.1,
+    "embedding_dropout_prob": 0.1,
+    "initializer_range": 0.02,
+    "layer_norm_epsilon": 1e-05,
+    "hidden_size": 1024,
+    "num_attention_heads": 8,
+    "num_hidden_layers": 6,
+    "intermediate_size": 4096,
+    "max_position_embeddings": 2048,
+    "residual_dropout_prob": 0.1,
+    "vocab_size": 49154
+}

generation_config.json ADDED Viewed

	@@ -0,0 +1,6 @@

+{
+  "_from_model_config": true,
+  "bos_token_id": 50256,
+  "eos_token_id": 50256,
+  "transformers_version": "4.28.1"
+}

onnx/model-fp32.onnx ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:598d2277f44d871de5c1baa9f553e5541e3b8cfe72385bd2535fa954ca4f06c0
+size 512235748

special_tokens_map.json ADDED Viewed

	@@ -0,0 +1,28 @@

+{
+  "additional_special_tokens": [
+    "<|endoftext|>",
+    "<fim_prefix>",
+    "<fim_middle>",
+    "<fim_suffix>",
+    "<fim_pad>",
+    "<filename>",
+    "<gh_stars>",
+    "<issue_start>",
+    "<issue_comment>",
+    "<issue_closed>",
+    "<jupyter_start>",
+    "<jupyter_text>",
+    "<jupyter_code>",
+    "<jupyter_output>",
+    "<empty_output>",
+    "<commit_before>",
+    "<commit_msg>",
+    "<commit_after>",
+    "<reponame>"
+  ],
+  "bos_token": "<|endoftext|>",
+  "eos_token": "<|endoftext|>",
+  "mask_token": "<mask>",
+  "pad_token": "<pad>",
+  "unk_token": "<|endoftext|>"
+}

tokenizer.json ADDED Viewed

The diff for this file is too large to render. See raw diff

tokenizer_config.json ADDED Viewed

	@@ -0,0 +1,35 @@

+{
+  "add_prefix_space": false,
+  "additional_special_tokens": [
+    "<|endoftext|>",
+    "<fim_prefix>",
+    "<fim_middle>",
+    "<fim_suffix>",
+    "<fim_pad>",
+    "<filename>",
+    "<gh_stars>",
+    "<issue_start>",
+    "<issue_comment>",
+    "<issue_closed>",
+    "<jupyter_start>",
+    "<jupyter_text>",
+    "<jupyter_code>",
+    "<jupyter_output>",
+    "<empty_output>",
+    "<commit_before>",
+    "<commit_msg>",
+    "<commit_after>",
+    "<reponame>"
+  ],
+  "bos_token": "<|endoftext|>",
+  "clean_up_tokenization_spaces": true,
+  "eos_token": "<|endoftext|>",
+  "add_eos_token": true,
+  "model_max_length": 1000000000000000019884624838656,
+  "unk_token": "<|endoftext|>",
+  "vocab_size": 49152,
+  "tokenizer_class": "CodeSageTokenizer",
+  "auto_map": {
+    "AutoTokenizer": ["tokenization_codesage.CodeSageTokenizer", null]
+  }
+}