plamo-embedding-1b-coreml
Model Summary
This is an unofficial Core ML conversion of pfnet/plamo-embedding-1b, a Japanese text embedding model developed by Preferred Networks, Inc., for running directly on iOS/macOS. All credit for the original model goes to Preferred Networks.
- Base model: pfnet/plamo-embedding-1b (1B params, hidden=2048)
- Output: a sentence embedding vector (2048 dims), already mean-pooled and L2-normalized. You can use a dot product instead of cosine similarity.
- Models are split by fixed sequence length.
Files
| File | Seq len | Precision | Approx. size |
|---|---|---|---|
plamo-embedding-1b_seq128_fp16.mlpackage |
128 | fp16 | ~2.0GB |
plamo-embedding-1b_seq256_fp16.mlpackage |
256 | fp16 | ~2.0GB |
plamo-embedding-1b_seq512_fp16.mlpackage |
512 | fp16 | ~2.0GB |
Usage (about the masks)
This model takes two kinds of masks as input:
attention_mask: the usual Transformer padding mask.embed_mask: a mask used only for mean pooling. When encoding a document, this can be the same asattention_mask. When encoding a query (when prepending the instruction prefix"次の文章に対して、関連する文章を検索してください: "), exclude the instruction tokens and set only the sentence's own tokens to 1.
Swift (Core ML) usage example
import CoreML
let configuration = MLModelConfiguration()
configuration.computeUnits = .all
let model = try MLModel(
contentsOf: Bundle.main.url(forResource: "plamo-embedding-1b_seq256_fp16", withExtension: "mlmodelc")!,
configuration: configuration
)
let inputIDs: MLMultiArray = ... // shape [1, 256], Int32
let attentionMask: MLMultiArray = ... // shape [1, 256], Int32
let embedMask: MLMultiArray = ... // shape [1, 256], Int32 (same as attentionMask for document encoding)
let input = try MLDictionaryFeatureProvider(dictionary: [
"input_ids": MLFeatureValue(multiArray: inputIDs),
"attention_mask": MLFeatureValue(multiArray: attentionMask),
"embed_mask": MLFeatureValue(multiArray: embedMask),
])
let output = try model.prediction(from: input)
let embedding = output.featureValue(for: "sentence_embedding")!.multiArrayValue!
// [1, 2048] L2-normalized vector (dtype: float16)
You must use the original model's own tokenizer.model (SentencePiece) as-is.
Python verification example
import numpy as np
import coremltools as ct
from transformers import AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("pfnet/plamo-embedding-1b", trust_remote_code=True)
tokenizer.pad_token = tokenizer.eos_token
tokenizer.padding_side = "left"
mlmodel = ct.models.MLModel("plamo-embedding-1b_seq256_fp16.mlpackage")
doc = "PLaMo-Embedding-1Bは、Preferred Networks, Inc. によって開発された日本語テキスト埋め込みモデルです。"
feats = tokenizer([doc], return_tensors="np", truncation=True, padding="max_length", max_length=256, add_special_tokens=True)
out = mlmodel.predict({
"input_ids": feats["input_ids"].astype(np.int32),
"attention_mask": feats["attention_mask"].astype(np.int32),
"embed_mask": feats["attention_mask"].astype(np.int32),
})
embedding = out["sentence_embedding"] # shape (1, 2048), L2-normalized
Notes
- This is a community conversion, not an official release from Preferred Networks.
Security
Audited against its upstream with model-audit-lite: weight format, bundled code, and a machine-readable lineage (ML-BOM). Details, checksums and how to reproduce: SECURITY.md.
モデルの概要
pfnet/plamo-embedding-1b(Preferred Networks製の日本語テキスト埋め込みモデル)を、iOS/macOS (Core ML) で直接動かせるように変換したものです。元モデルの著作権はPreferred Networksに帰属します。
- ベースモデル: pfnet/plamo-embedding-1b(1B params, hidden=2048)
- 出力: mean pooling + L2 normalize 済みの文埋め込みベクトル(2048次元)。コサイン類似度の代わりに内積(dot product)で類似度計算できます。
- 系列長ごとに固定長でモデルを分けています。
ファイル一覧
| ファイル | 系列長 | 精度 | サイズ目安 |
|---|---|---|---|
plamo-embedding-1b_seq128_fp16.mlpackage |
128 | fp16 | 約2.0GB |
plamo-embedding-1b_seq256_fp16.mlpackage |
256 | fp16 | 約2.0GB |
plamo-embedding-1b_seq512_fp16.mlpackage |
512 | fp16 | 約2.0GB |
使い方(マスクについて)
本モデルは2種類のマスクを入力に取ります:
attention_mask: Transformer内の通常のパディングマスクembed_mask: mean pooling専用のマスク。文書エンコード時はattention_maskと同じでOKです。クエリ エンコード時("次の文章に対して、関連する文章を検索してください: "というinstructionプレフィックス を付与する場合)は、instructionトークンを除外し、文そのもののトークンのみを1にしてください。
Swift (Core ML) での使用例
import CoreML
let configuration = MLModelConfiguration()
configuration.computeUnits = .all
let model = try MLModel(
contentsOf: Bundle.main.url(forResource: "plamo-embedding-1b_seq256_fp16", withExtension: "mlmodelc")!,
configuration: configuration
)
let inputIDs: MLMultiArray = ... // shape [1, 256], Int32
let attentionMask: MLMultiArray = ... // shape [1, 256], Int32
let embedMask: MLMultiArray = ... // shape [1, 256], Int32(文書エンコード時はattentionMaskと同じでOK)
let input = try MLDictionaryFeatureProvider(dictionary: [
"input_ids": MLFeatureValue(multiArray: inputIDs),
"attention_mask": MLFeatureValue(multiArray: attentionMask),
"embed_mask": MLFeatureValue(multiArray: embedMask),
])
let output = try model.prediction(from: input)
let embedding = output.featureValue(for: "sentence_embedding")!.multiArrayValue!
// [1, 2048] の L2 正規化済みベクトル(dtype: float16)
トークナイザは元モデルの tokenizer.model(SentencePiece)をそのまま使う必要があります。
Pythonでの検証例
import numpy as np
import coremltools as ct
from transformers import AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("pfnet/plamo-embedding-1b", trust_remote_code=True)
tokenizer.pad_token = tokenizer.eos_token
tokenizer.padding_side = "left"
mlmodel = ct.models.MLModel("plamo-embedding-1b_seq256_fp16.mlpackage")
doc = "PLaMo-Embedding-1Bは、Preferred Networks, Inc. によって開発された日本語テキスト埋め込みモデルです。"
feats = tokenizer([doc], return_tensors="np", truncation=True, padding="max_length", max_length=256, add_special_tokens=True)
out = mlmodel.predict({
"input_ids": feats["input_ids"].astype(np.int32),
"attention_mask": feats["attention_mask"].astype(np.int32),
"embed_mask": feats["attention_mask"].astype(np.int32),
})
embedding = out["sentence_embedding"] # shape (1, 2048), L2-normalized
備考
- 本変換は非公式のコミュニティ版です。Preferred Networksによる公式リリースではありません。
セキュリティー
model-audit-lite で変換元と突き合わせて監査済みです(重みの形式、同梱コード、機械可読な系譜=ML-BOM)。詳細・チェックサム・再現方法は SECURITY.md をご覧ください。
- Downloads last month
- 26
Model tree for masahiroid/plamo-embedding-1b-coreml
Base model
pfnet/plamo-embedding-1b