moda-ner-v-catalog / README.md
ArkidMitra's picture
Add machine-readable eval metadata; resolve the placeholder score; correct library_name
64702f7 verified
|
Raw History Blame Contribute Delete
3.99 kB
metadata
license: cc-by-nc-4.0
library_name: open_clip
pipeline_tag: image-classification
base_model: HopitAI/moda-fashion-distilled
language:
  - en
metrics:
  - f1
  - accuracy
tags:
  - fashion
  - attribute-extraction
  - product-attributes
  - catalog-enrichment
  - e-commerce
  - computer-vision
  - image-classification
  - multi-label-classification
  - siglip
  - moda-ner
  - benchmark
  - reproducibility
  - non-commercial
  - color-detection
  - fit-prediction
model-index:
  - name: MODA_NER(V) Catalog
    results:
      - task:
          type: image-classification
          name: Fashion attribute extraction
        dataset:
          type: moda-general-attribute-suite
          name: MODA General Attribute Suite (catalog track)
        metrics:
          - type: f1
            value: 0.8292
            name: Field-macro set F1

MODA_NER(V) - Catalog

Tier * - open code, open weights. Licence: CC BY-NC 4.0.

Tier * — open code, open weights. Weights: CC BY-NC 4.0.

Ten field-specific supervised heads on our own frozen encoder (see Provenance below).

Input contract: one clean catalogue product image. Output: category, collar, colour, fabric, fastening, fit, neckline, pattern, pocket, sleeve length.

Field-macro set F1 (catalog track)
This released checkpoint 0.8292
Strongest external open baseline 0.6657

Selected best-internal fields: colour 0.7053, fit 0.6708, fabric 0.6693, pocket 0.9491.

Why non-commercial. This track is evaluated against a research-only corpus whose terms do not permit commercial use of models trained on it. We honour those terms, and they bind us as well: these weights are not part of Hopit's hosted product. For commercial deployment we fine-tune on the customer's own catalogue, which raises accuracy on their taxonomy and produces a model with no dependency on research-licensed data.

Provenance

The encoder these heads run on is ours: HopitAI/moda-fashion-distilled, MIT, already public. Nothing from another vendor is loaded at inference time. (Recorded in the programme documentation; we re-confirm it against this track's frozen artifacts before release.)

That is worth stating plainly, because the comparator on this track is a FashionSigLIP-based system and it would be easy to assume this model is that system with heads attached. It is not. FashionSigLIP appears in two other roles:

  • As the distillation teacher. An earlier ladder of checkpoints put conditional heads on frozen Marqo-FashionSigLIP. We distilled that system into our own encoder; the teacher is used during training and is not needed to serve.
  • As the baseline we measure against. The comparator figure quoted above is that same FashionSigLIP-based system.

Lineage, stated once rather than implied: moda-fashion-distilled is itself a distilled student built on ViT-B/16-SigLIP, from a teacher ensemble that included our own DeepFashion2 fine-tune. Marqo-FashionSigLIP is Apache-2.0. The DeepFashion2 corpus is research-only, so we do not describe this pipeline as provenance-clean end to end.

Credit for this model. CC BY-NC requires attribution. Cite the MODA General Attribute Suite (CITATION.cff) when reporting numbers from this track.

Links

Usage

The heads are not a transformers architecture, so load them through the suite repository rather than AutoModel:

git clone https://github.com/hopit-ai/Moda_ner && cd Moda_ner
pip install -r requirements-inference.txt
huggingface-cli download HopitAI/moda-ner-v-catalog --local-dir ./moda-ner-v-catalog
python models/inference.py --route catalog --model-dir ./moda-ner-v-catalog --images photo.jpg