Instructions to use kerasformers/sam2_hiera_base_plus with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- KerasFormers
How to use kerasformers/sam2_hiera_base_plus with KerasFormers:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Keras
How to use kerasformers/sam2_hiera_base_plus with Keras:
# Available backend options are: "jax", "torch", "tensorflow". import os os.environ["KERAS_BACKEND"] = "jax" import keras model = keras.saving.load_model("hf://kerasformers/sam2_hiera_base_plus") - sam2
How to use kerasformers/sam2_hiera_base_plus with sam2:
# Use SAM2 with images import torch from sam2.sam2_image_predictor import SAM2ImagePredictor predictor = SAM2ImagePredictor.from_pretrained(kerasformers/sam2_hiera_base_plus) with torch.inference_mode(), torch.autocast("cuda", dtype=torch.bfloat16): predictor.set_image(<your_image>) masks, _, _ = predictor.predict(<input_prompts>)# Use SAM2 with videos import torch from sam2.sam2_video_predictor import SAM2VideoPredictor predictor = SAM2VideoPredictor.from_pretrained(kerasformers/sam2_hiera_base_plus) with torch.inference_mode(), torch.autocast("cuda", dtype=torch.bfloat16): state = predictor.init_state(<your_video>) # add new prompts and instantly get the output on the same frame frame_idx, object_ids, masks = predictor.add_new_points(state, <your_prompts>): # propagate the prompts to get masklets throughout the video for frame_idx, object_ids, masks in predictor.propagate_in_video(state): ... - Notebooks
- Google Colab
- Kaggle
| { | |
| "library_name": "kerasformers", | |
| "kerasformers_version": "1.1.3", | |
| "model_module": "kerasformers.models.sam2", | |
| "model_class": "SAM2PromptableSegment", | |
| "variant": "sam2_hiera_base_plus", | |
| "weights": "model.weights.h5", | |
| "model_type": "sam2", | |
| "hidden_dim": 112, | |
| "blocks_per_stage": [ | |
| 2, | |
| 3, | |
| 16, | |
| 3 | |
| ], | |
| "embed_dim_per_stage": [ | |
| 112, | |
| 224, | |
| 448, | |
| 896 | |
| ], | |
| "num_attention_heads_per_stage": [ | |
| 2, | |
| 4, | |
| 8, | |
| 16 | |
| ], | |
| "window_size_per_stage": [ | |
| 8, | |
| 4, | |
| 14, | |
| 7 | |
| ], | |
| "global_attention_blocks": [ | |
| 12, | |
| 16, | |
| 20 | |
| ], | |
| "backbone_channel_list": [ | |
| 896, | |
| 448, | |
| 224, | |
| 112 | |
| ], | |
| "window_pos_embed_bg_size": [ | |
| 14, | |
| 14 | |
| ], | |
| "num_multimask_outputs": 3, | |
| "include_box_input": false, | |
| "include_mask_input": false, | |
| "multimask_output": true, | |
| "image_size": 1024 | |
| } |