plamo-embedding-1b-coreml

English | 日本語

Model Summary

This is an unofficial Core ML conversion of pfnet/plamo-embedding-1b, a Japanese text embedding model developed by Preferred Networks, Inc., for running directly on iOS/macOS. All credit for the original model goes to Preferred Networks.

  • Base model: pfnet/plamo-embedding-1b (1B params, hidden=2048)
  • Output: a sentence embedding vector (2048 dims), already mean-pooled and L2-normalized. You can use a dot product instead of cosine similarity.
  • Models are split by fixed sequence length.

Files

File Seq len Precision Approx. size
plamo-embedding-1b_seq128_fp16.mlpackage 128 fp16 ~2.0GB
plamo-embedding-1b_seq256_fp16.mlpackage 256 fp16 ~2.0GB
plamo-embedding-1b_seq512_fp16.mlpackage 512 fp16 ~2.0GB

Usage (about the masks)

This model takes two kinds of masks as input:

  • attention_mask: the usual Transformer padding mask.
  • embed_mask: a mask used only for mean pooling. When encoding a document, this can be the same as attention_mask. When encoding a query (when prepending the instruction prefix "次の文章に対して、関連する文章を検索してください: "), exclude the instruction tokens and set only the sentence's own tokens to 1.

Swift (Core ML) usage example

import CoreML

let configuration = MLModelConfiguration()
configuration.computeUnits = .all
let model = try MLModel(
    contentsOf: Bundle.main.url(forResource: "plamo-embedding-1b_seq256_fp16", withExtension: "mlmodelc")!,
    configuration: configuration
)

let inputIDs: MLMultiArray = ...       // shape [1, 256], Int32
let attentionMask: MLMultiArray = ...  // shape [1, 256], Int32
let embedMask: MLMultiArray = ...      // shape [1, 256], Int32 (same as attentionMask for document encoding)

let input = try MLDictionaryFeatureProvider(dictionary: [
    "input_ids": MLFeatureValue(multiArray: inputIDs),
    "attention_mask": MLFeatureValue(multiArray: attentionMask),
    "embed_mask": MLFeatureValue(multiArray: embedMask),
])
let output = try model.prediction(from: input)
let embedding = output.featureValue(for: "sentence_embedding")!.multiArrayValue!
// [1, 2048] L2-normalized vector (dtype: float16)

You must use the original model's own tokenizer.model (SentencePiece) as-is.

Python verification example

import numpy as np
import coremltools as ct
from transformers import AutoTokenizer

tokenizer = AutoTokenizer.from_pretrained("pfnet/plamo-embedding-1b", trust_remote_code=True)
tokenizer.pad_token = tokenizer.eos_token
tokenizer.padding_side = "left"

mlmodel = ct.models.MLModel("plamo-embedding-1b_seq256_fp16.mlpackage")

doc = "PLaMo-Embedding-1Bは、Preferred Networks, Inc. によって開発された日本語テキスト埋め込みモデルです。"
feats = tokenizer([doc], return_tensors="np", truncation=True, padding="max_length", max_length=256, add_special_tokens=True)

out = mlmodel.predict({
    "input_ids": feats["input_ids"].astype(np.int32),
    "attention_mask": feats["attention_mask"].astype(np.int32),
    "embed_mask": feats["attention_mask"].astype(np.int32),
})
embedding = out["sentence_embedding"]  # shape (1, 2048), L2-normalized

Notes

  • This is a community conversion, not an official release from Preferred Networks.

Security

Audited against its upstream with model-audit-lite: weight format, bundled code, and a machine-readable lineage (ML-BOM). Details, checksums and how to reproduce: SECURITY.md.


モデルの概要

pfnet/plamo-embedding-1b(Preferred Networks製の日本語テキスト埋め込みモデル)を、iOS/macOS (Core ML) で直接動かせるように変換したものです。元モデルの著作権はPreferred Networksに帰属します。

  • ベースモデル: pfnet/plamo-embedding-1b(1B params, hidden=2048)
  • 出力: mean pooling + L2 normalize 済みの文埋め込みベクトル(2048次元)。コサイン類似度の代わりに内積(dot product)で類似度計算できます。
  • 系列長ごとに固定長でモデルを分けています。

ファイル一覧

ファイル 系列長 精度 サイズ目安
plamo-embedding-1b_seq128_fp16.mlpackage 128 fp16 約2.0GB
plamo-embedding-1b_seq256_fp16.mlpackage 256 fp16 約2.0GB
plamo-embedding-1b_seq512_fp16.mlpackage 512 fp16 約2.0GB

使い方(マスクについて)

本モデルは2種類のマスクを入力に取ります:

  • attention_mask: Transformer内の通常のパディングマスク
  • embed_mask: mean pooling専用のマスク。文書エンコード時はattention_maskと同じでOKです。クエリ エンコード時("次の文章に対して、関連する文章を検索してください: " というinstructionプレフィックス を付与する場合)は、instructionトークンを除外し、文そのもののトークンのみを1にしてください。

Swift (Core ML) での使用例

import CoreML

let configuration = MLModelConfiguration()
configuration.computeUnits = .all
let model = try MLModel(
    contentsOf: Bundle.main.url(forResource: "plamo-embedding-1b_seq256_fp16", withExtension: "mlmodelc")!,
    configuration: configuration
)

let inputIDs: MLMultiArray = ...       // shape [1, 256], Int32
let attentionMask: MLMultiArray = ...  // shape [1, 256], Int32
let embedMask: MLMultiArray = ...      // shape [1, 256], Int32(文書エンコード時はattentionMaskと同じでOK)

let input = try MLDictionaryFeatureProvider(dictionary: [
    "input_ids": MLFeatureValue(multiArray: inputIDs),
    "attention_mask": MLFeatureValue(multiArray: attentionMask),
    "embed_mask": MLFeatureValue(multiArray: embedMask),
])
let output = try model.prediction(from: input)
let embedding = output.featureValue(for: "sentence_embedding")!.multiArrayValue!
// [1, 2048] の L2 正規化済みベクトル(dtype: float16)

トークナイザは元モデルの tokenizer.model(SentencePiece)をそのまま使う必要があります。

Pythonでの検証例

import numpy as np
import coremltools as ct
from transformers import AutoTokenizer

tokenizer = AutoTokenizer.from_pretrained("pfnet/plamo-embedding-1b", trust_remote_code=True)
tokenizer.pad_token = tokenizer.eos_token
tokenizer.padding_side = "left"

mlmodel = ct.models.MLModel("plamo-embedding-1b_seq256_fp16.mlpackage")

doc = "PLaMo-Embedding-1Bは、Preferred Networks, Inc. によって開発された日本語テキスト埋め込みモデルです。"
feats = tokenizer([doc], return_tensors="np", truncation=True, padding="max_length", max_length=256, add_special_tokens=True)

out = mlmodel.predict({
    "input_ids": feats["input_ids"].astype(np.int32),
    "attention_mask": feats["attention_mask"].astype(np.int32),
    "embed_mask": feats["attention_mask"].astype(np.int32),
})
embedding = out["sentence_embedding"]  # shape (1, 2048), L2-normalized

備考

  • 本変換は非公式のコミュニティ版です。Preferred Networksによる公式リリースではありません。

セキュリティー

model-audit-lite で変換元と突き合わせて監査済みです(重みの形式、同梱コード、機械可読な系譜=ML-BOM)。詳細・チェックサム・再現方法は SECURITY.md をご覧ください。

Downloads last month
26
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for masahiroid/plamo-embedding-1b-coreml

Quantized
(1)
this model