allenai
/

MolmoPoint-8B

Image-Text-to-Text

Model card Files Files and versions

chrisc36 commited on about 1 month ago

Commit

2c280e9

·

verified ·

1 Parent(s): 74a5036

Update README.md

Files changed (1) hide show

README.md +8 -2

README.md CHANGED Viewed

@@ -16,7 +16,7 @@ tags:
 # MolmoPoint-8B
 MolmoPoint-8B is a fully-open VLM developed by the Allen Institute for AI (Ai2) that support image, video and multi-image understanding and grounding.
-It has novel pointing mechansim that improves image pointing, video pointing, and video tracking, see our technical report for details.
 Note the huggingface MolmoPoint model does not support training, see our github repo for the training code.
@@ -42,7 +42,7 @@ We recommend running MolmoPoint with `logits_processor=model.build_logit_process
 to enforce points tokens are generated in a valid way.
 In MolmoPoint, instead of coordinates points will be generated as a series of special
-tokens, to decode the tokens back into points requires some additional
 metadata from the preprocessor.
 The metadata is returned by the preprocessor using the `return_pointing_metadata` flag.
 Then `model.extract_image_points` and `model.extract_video_points` do the decoding, they
@@ -111,6 +111,9 @@ points = model.extract_image_points(
     metadata["subpatch_mapping"],
     metadata["image_sizes"]
 )
 print(points)
 ```
@@ -156,6 +159,9 @@ points = model.extract_video_points(
     metadata["timestamps"],
     metadata["video_size"]
 )
 print(points)
 ```

 # MolmoPoint-8B
 MolmoPoint-8B is a fully-open VLM developed by the Allen Institute for AI (Ai2) that support image, video and multi-image understanding and grounding.
+It has new pointing mechansim that improves image pointing, video pointing, and video tracking, see our technical report for details.
 Note the huggingface MolmoPoint model does not support training, see our github repo for the training code.
 to enforce points tokens are generated in a valid way.
 In MolmoPoint, instead of coordinates points will be generated as a series of special
+tokens, decoding the tokens back into points requires some additional
 metadata from the preprocessor.
 The metadata is returned by the preprocessor using the `return_pointing_metadata` flag.
 Then `model.extract_image_points` and `model.extract_video_points` do the decoding, they
     metadata["subpatch_mapping"],
     metadata["image_sizes"]
 )
+# points as a list of [object_id, image_num, x, y]
+# For multiple images, `image_num` is the index of the image the point is in
 print(points)
 ```
     metadata["timestamps"],
     metadata["video_size"]
 )
+# points as a list of [object_id, image_num, x, y]
+# For tracking, object_id uniquely identifies objects that might appear multiple frames.
 print(points)
 ```