Prompt48 commited on
Commit
958ab38
·
verified ·
1 Parent(s): 8807c75

Upload edit\Qwen3-TTS-test\.venv\Lib\site-packages\transformers\models\grounding_dino\image_processing_grounding_dino.py with huggingface_hub

Browse files
edit//Qwen3-TTS-test//.venv//Lib//site-packages//transformers//models//grounding_dino//image_processing_grounding_dino.py ADDED
@@ -0,0 +1,1621 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # coding=utf-8
2
+ # Copyright 2024 The HuggingFace Inc. team. All rights reserved.
3
+ #
4
+ # Licensed under the Apache License, Version 2.0 (the "License");
5
+ # you may not use this file except in compliance with the License.
6
+ # You may obtain a copy of the License at
7
+ #
8
+ # http://www.apache.org/licenses/LICENSE-2.0
9
+ #
10
+ # Unless required by applicable law or agreed to in writing, software
11
+ # distributed under the License is distributed on an "AS IS" BASIS,
12
+ # WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
13
+ # See the License for the specific language governing permissions and
14
+ # limitations under the License.
15
+ """Image processor class for Deformable DETR."""
16
+
17
+ import io
18
+ import pathlib
19
+ from collections import defaultdict
20
+ from collections.abc import Iterable
21
+ from typing import TYPE_CHECKING, Any, Callable, Optional, Union
22
+
23
+ import numpy as np
24
+
25
+ from ...feature_extraction_utils import BatchFeature
26
+ from ...image_processing_utils import BaseImageProcessor, get_size_dict
27
+ from ...image_transforms import (
28
+ PaddingMode,
29
+ center_to_corners_format,
30
+ corners_to_center_format,
31
+ id_to_rgb,
32
+ pad,
33
+ rescale,
34
+ resize,
35
+ rgb_to_id,
36
+ to_channel_dimension_format,
37
+ )
38
+ from ...image_utils import (
39
+ IMAGENET_DEFAULT_MEAN,
40
+ IMAGENET_DEFAULT_STD,
41
+ ChannelDimension,
42
+ ImageInput,
43
+ PILImageResampling,
44
+ get_image_size,
45
+ infer_channel_dimension_format,
46
+ is_scaled_image,
47
+ make_flat_list_of_images,
48
+ to_numpy_array,
49
+ valid_images,
50
+ validate_annotations,
51
+ validate_kwargs,
52
+ validate_preprocess_arguments,
53
+ )
54
+ from ...utils import (
55
+ ExplicitEnum,
56
+ TensorType,
57
+ is_flax_available,
58
+ is_jax_tensor,
59
+ is_scipy_available,
60
+ is_tf_available,
61
+ is_tf_tensor,
62
+ is_torch_available,
63
+ is_torch_tensor,
64
+ is_vision_available,
65
+ logging,
66
+ )
67
+
68
+
69
+ if is_torch_available():
70
+ import torch
71
+ from torch import nn
72
+
73
+
74
+ if is_vision_available():
75
+ import PIL
76
+
77
+ if is_scipy_available():
78
+ import scipy.special
79
+ import scipy.stats
80
+
81
+ if TYPE_CHECKING:
82
+ from .modeling_grounding_dino import GroundingDinoObjectDetectionOutput
83
+
84
+
85
+ logger = logging.get_logger(__name__) # pylint: disable=invalid-name
86
+
87
+ AnnotationType = dict[str, Union[int, str, list[dict]]]
88
+
89
+
90
+ class AnnotationFormat(ExplicitEnum):
91
+ COCO_DETECTION = "coco_detection"
92
+ COCO_PANOPTIC = "coco_panoptic"
93
+
94
+
95
+ SUPPORTED_ANNOTATION_FORMATS = (AnnotationFormat.COCO_DETECTION, AnnotationFormat.COCO_PANOPTIC)
96
+
97
+
98
+ # Copied from transformers.models.detr.image_processing_detr.get_size_with_aspect_ratio
99
+ def get_size_with_aspect_ratio(image_size, size, max_size=None) -> tuple[int, int]:
100
+ """
101
+ Computes the output image size given the input image size and the desired output size.
102
+
103
+ Args:
104
+ image_size (`tuple[int, int]`):
105
+ The input image size.
106
+ size (`int`):
107
+ The desired output size.
108
+ max_size (`int`, *optional*):
109
+ The maximum allowed output size.
110
+ """
111
+ height, width = image_size
112
+ raw_size = None
113
+ if max_size is not None:
114
+ min_original_size = float(min((height, width)))
115
+ max_original_size = float(max((height, width)))
116
+ if max_original_size / min_original_size * size > max_size:
117
+ raw_size = max_size * min_original_size / max_original_size
118
+ size = int(round(raw_size))
119
+
120
+ if (height <= width and height == size) or (width <= height and width == size):
121
+ oh, ow = height, width
122
+ elif width < height:
123
+ ow = size
124
+ if max_size is not None and raw_size is not None:
125
+ oh = int(raw_size * height / width)
126
+ else:
127
+ oh = int(size * height / width)
128
+ else:
129
+ oh = size
130
+ if max_size is not None and raw_size is not None:
131
+ ow = int(raw_size * width / height)
132
+ else:
133
+ ow = int(size * width / height)
134
+
135
+ return (oh, ow)
136
+
137
+
138
+ # Copied from transformers.models.detr.image_processing_detr.get_resize_output_image_size
139
+ def get_resize_output_image_size(
140
+ input_image: np.ndarray,
141
+ size: Union[int, tuple[int, int], list[int]],
142
+ max_size: Optional[int] = None,
143
+ input_data_format: Optional[Union[str, ChannelDimension]] = None,
144
+ ) -> tuple[int, int]:
145
+ """
146
+ Computes the output image size given the input image size and the desired output size. If the desired output size
147
+ is a tuple or list, the output image size is returned as is. If the desired output size is an integer, the output
148
+ image size is computed by keeping the aspect ratio of the input image size.
149
+
150
+ Args:
151
+ input_image (`np.ndarray`):
152
+ The image to resize.
153
+ size (`int` or `tuple[int, int]` or `list[int]`):
154
+ The desired output size.
155
+ max_size (`int`, *optional*):
156
+ The maximum allowed output size.
157
+ input_data_format (`ChannelDimension` or `str`, *optional*):
158
+ The channel dimension format of the input image. If not provided, it will be inferred from the input image.
159
+ """
160
+ image_size = get_image_size(input_image, input_data_format)
161
+ if isinstance(size, (list, tuple)):
162
+ return size
163
+
164
+ return get_size_with_aspect_ratio(image_size, size, max_size)
165
+
166
+
167
+ # Copied from transformers.models.detr.image_processing_detr.get_image_size_for_max_height_width
168
+ def get_image_size_for_max_height_width(
169
+ input_image: np.ndarray,
170
+ max_height: int,
171
+ max_width: int,
172
+ input_data_format: Optional[Union[str, ChannelDimension]] = None,
173
+ ) -> tuple[int, int]:
174
+ """
175
+ Computes the output image size given the input image and the maximum allowed height and width. Keep aspect ratio.
176
+ Important, even if image_height < max_height and image_width < max_width, the image will be resized
177
+ to at least one of the edges be equal to max_height or max_width.
178
+
179
+ For example:
180
+ - input_size: (100, 200), max_height: 50, max_width: 50 -> output_size: (25, 50)
181
+ - input_size: (100, 200), max_height: 200, max_width: 500 -> output_size: (200, 400)
182
+
183
+ Args:
184
+ input_image (`np.ndarray`):
185
+ The image to resize.
186
+ max_height (`int`):
187
+ The maximum allowed height.
188
+ max_width (`int`):
189
+ The maximum allowed width.
190
+ input_data_format (`ChannelDimension` or `str`, *optional*):
191
+ The channel dimension format of the input image. If not provided, it will be inferred from the input image.
192
+ """
193
+ image_size = get_image_size(input_image, input_data_format)
194
+ height, width = image_size
195
+ height_scale = max_height / height
196
+ width_scale = max_width / width
197
+ min_scale = min(height_scale, width_scale)
198
+ new_height = int(height * min_scale)
199
+ new_width = int(width * min_scale)
200
+ return new_height, new_width
201
+
202
+
203
+ # Copied from transformers.models.detr.image_processing_detr.get_numpy_to_framework_fn
204
+ def get_numpy_to_framework_fn(arr) -> Callable:
205
+ """
206
+ Returns a function that converts a numpy array to the framework of the input array.
207
+
208
+ Args:
209
+ arr (`np.ndarray`): The array to convert.
210
+ """
211
+ if isinstance(arr, np.ndarray):
212
+ return np.array
213
+ if is_tf_available() and is_tf_tensor(arr):
214
+ import tensorflow as tf
215
+
216
+ return tf.convert_to_tensor
217
+ if is_torch_available() and is_torch_tensor(arr):
218
+ import torch
219
+
220
+ return torch.tensor
221
+ if is_flax_available() and is_jax_tensor(arr):
222
+ import jax.numpy as jnp
223
+
224
+ return jnp.array
225
+ raise ValueError(f"Cannot convert arrays of type {type(arr)}")
226
+
227
+
228
+ # Copied from transformers.models.detr.image_processing_detr.safe_squeeze
229
+ def safe_squeeze(arr: np.ndarray, axis: Optional[int] = None) -> np.ndarray:
230
+ """
231
+ Squeezes an array, but only if the axis specified has dim 1.
232
+ """
233
+ if axis is None:
234
+ return arr.squeeze()
235
+
236
+ try:
237
+ return arr.squeeze(axis=axis)
238
+ except ValueError:
239
+ return arr
240
+
241
+
242
+ # Copied from transformers.models.detr.image_processing_detr.normalize_annotation
243
+ def normalize_annotation(annotation: dict, image_size: tuple[int, int]) -> dict:
244
+ image_height, image_width = image_size
245
+ norm_annotation = {}
246
+ for key, value in annotation.items():
247
+ if key == "boxes":
248
+ boxes = value
249
+ boxes = corners_to_center_format(boxes)
250
+ boxes /= np.asarray([image_width, image_height, image_width, image_height], dtype=np.float32)
251
+ norm_annotation[key] = boxes
252
+ else:
253
+ norm_annotation[key] = value
254
+ return norm_annotation
255
+
256
+
257
+ # Copied from transformers.models.detr.image_processing_detr.max_across_indices
258
+ def max_across_indices(values: Iterable[Any]) -> list[Any]:
259
+ """
260
+ Return the maximum value across all indices of an iterable of values.
261
+ """
262
+ return [max(values_i) for values_i in zip(*values)]
263
+
264
+
265
+ # Copied from transformers.models.detr.image_processing_detr.get_max_height_width
266
+ def get_max_height_width(
267
+ images: list[np.ndarray], input_data_format: Optional[Union[str, ChannelDimension]] = None
268
+ ) -> list[int]:
269
+ """
270
+ Get the maximum height and width across all images in a batch.
271
+ """
272
+ if input_data_format is None:
273
+ input_data_format = infer_channel_dimension_format(images[0])
274
+
275
+ if input_data_format == ChannelDimension.FIRST:
276
+ _, max_height, max_width = max_across_indices([img.shape for img in images])
277
+ elif input_data_format == ChannelDimension.LAST:
278
+ max_height, max_width, _ = max_across_indices([img.shape for img in images])
279
+ else:
280
+ raise ValueError(f"Invalid channel dimension format: {input_data_format}")
281
+ return (max_height, max_width)
282
+
283
+
284
+ # Copied from transformers.models.detr.image_processing_detr.make_pixel_mask
285
+ def make_pixel_mask(
286
+ image: np.ndarray, output_size: tuple[int, int], input_data_format: Optional[Union[str, ChannelDimension]] = None
287
+ ) -> np.ndarray:
288
+ """
289
+ Make a pixel mask for the image, where 1 indicates a valid pixel and 0 indicates padding.
290
+
291
+ Args:
292
+ image (`np.ndarray`):
293
+ Image to make the pixel mask for.
294
+ output_size (`tuple[int, int]`):
295
+ Output size of the mask.
296
+ """
297
+ input_height, input_width = get_image_size(image, channel_dim=input_data_format)
298
+ mask = np.zeros(output_size, dtype=np.int64)
299
+ mask[:input_height, :input_width] = 1
300
+ return mask
301
+
302
+
303
+ # Copied from transformers.models.detr.image_processing_detr.convert_coco_poly_to_mask
304
+ def convert_coco_poly_to_mask(segmentations, height: int, width: int) -> np.ndarray:
305
+ """
306
+ Convert a COCO polygon annotation to a mask.
307
+
308
+ Args:
309
+ segmentations (`list[list[float]]`):
310
+ List of polygons, each polygon represented by a list of x-y coordinates.
311
+ height (`int`):
312
+ Height of the mask.
313
+ width (`int`):
314
+ Width of the mask.
315
+ """
316
+ try:
317
+ from pycocotools import mask as coco_mask
318
+ except ImportError:
319
+ raise ImportError("Pycocotools is not installed in your environment.")
320
+
321
+ masks = []
322
+ for polygons in segmentations:
323
+ rles = coco_mask.frPyObjects(polygons, height, width)
324
+ mask = coco_mask.decode(rles)
325
+ if len(mask.shape) < 3:
326
+ mask = mask[..., None]
327
+ mask = np.asarray(mask, dtype=np.uint8)
328
+ mask = np.any(mask, axis=2)
329
+ masks.append(mask)
330
+ if masks:
331
+ masks = np.stack(masks, axis=0)
332
+ else:
333
+ masks = np.zeros((0, height, width), dtype=np.uint8)
334
+
335
+ return masks
336
+
337
+
338
+ # Copied from transformers.models.detr.image_processing_detr.prepare_coco_detection_annotation with DETR->GroundingDino
339
+ def prepare_coco_detection_annotation(
340
+ image,
341
+ target,
342
+ return_segmentation_masks: bool = False,
343
+ input_data_format: Optional[Union[ChannelDimension, str]] = None,
344
+ ):
345
+ """
346
+ Convert the target in COCO format into the format expected by GroundingDino.
347
+ """
348
+ image_height, image_width = get_image_size(image, channel_dim=input_data_format)
349
+
350
+ image_id = target["image_id"]
351
+ image_id = np.asarray([image_id], dtype=np.int64)
352
+
353
+ # Get all COCO annotations for the given image.
354
+ annotations = target["annotations"]
355
+ annotations = [obj for obj in annotations if "iscrowd" not in obj or obj["iscrowd"] == 0]
356
+
357
+ classes = [obj["category_id"] for obj in annotations]
358
+ classes = np.asarray(classes, dtype=np.int64)
359
+
360
+ # for conversion to coco api
361
+ area = np.asarray([obj["area"] for obj in annotations], dtype=np.float32)
362
+ iscrowd = np.asarray([obj.get("iscrowd", 0) for obj in annotations], dtype=np.int64)
363
+
364
+ boxes = [obj["bbox"] for obj in annotations]
365
+ # guard against no boxes via resizing
366
+ boxes = np.asarray(boxes, dtype=np.float32).reshape(-1, 4)
367
+ boxes[:, 2:] += boxes[:, :2]
368
+ boxes[:, 0::2] = boxes[:, 0::2].clip(min=0, max=image_width)
369
+ boxes[:, 1::2] = boxes[:, 1::2].clip(min=0, max=image_height)
370
+
371
+ keep = (boxes[:, 3] > boxes[:, 1]) & (boxes[:, 2] > boxes[:, 0])
372
+
373
+ new_target = {}
374
+ new_target["image_id"] = image_id
375
+ new_target["class_labels"] = classes[keep]
376
+ new_target["boxes"] = boxes[keep]
377
+ new_target["area"] = area[keep]
378
+ new_target["iscrowd"] = iscrowd[keep]
379
+ new_target["orig_size"] = np.asarray([int(image_height), int(image_width)], dtype=np.int64)
380
+
381
+ if annotations and "keypoints" in annotations[0]:
382
+ keypoints = [obj["keypoints"] for obj in annotations]
383
+ # Converting the filtered keypoints list to a numpy array
384
+ keypoints = np.asarray(keypoints, dtype=np.float32)
385
+ # Apply the keep mask here to filter the relevant annotations
386
+ keypoints = keypoints[keep]
387
+ num_keypoints = keypoints.shape[0]
388
+ keypoints = keypoints.reshape((-1, 3)) if num_keypoints else keypoints
389
+ new_target["keypoints"] = keypoints
390
+
391
+ if return_segmentation_masks:
392
+ segmentation_masks = [obj["segmentation"] for obj in annotations]
393
+ masks = convert_coco_poly_to_mask(segmentation_masks, image_height, image_width)
394
+ new_target["masks"] = masks[keep]
395
+
396
+ return new_target
397
+
398
+
399
+ # Copied from transformers.models.detr.image_processing_detr.masks_to_boxes
400
+ def masks_to_boxes(masks: np.ndarray) -> np.ndarray:
401
+ """
402
+ Compute the bounding boxes around the provided panoptic segmentation masks.
403
+
404
+ Args:
405
+ masks: masks in format `[number_masks, height, width]` where N is the number of masks
406
+
407
+ Returns:
408
+ boxes: bounding boxes in format `[number_masks, 4]` in xyxy format
409
+ """
410
+ if masks.size == 0:
411
+ return np.zeros((0, 4))
412
+
413
+ h, w = masks.shape[-2:]
414
+ y = np.arange(0, h, dtype=np.float32)
415
+ x = np.arange(0, w, dtype=np.float32)
416
+ # see https://github.com/pytorch/pytorch/issues/50276
417
+ y, x = np.meshgrid(y, x, indexing="ij")
418
+
419
+ x_mask = masks * np.expand_dims(x, axis=0)
420
+ x_max = x_mask.reshape(x_mask.shape[0], -1).max(-1)
421
+ x = np.ma.array(x_mask, mask=~(np.array(masks, dtype=bool)))
422
+ x_min = x.filled(fill_value=1e8)
423
+ x_min = x_min.reshape(x_min.shape[0], -1).min(-1)
424
+
425
+ y_mask = masks * np.expand_dims(y, axis=0)
426
+ y_max = y_mask.reshape(x_mask.shape[0], -1).max(-1)
427
+ y = np.ma.array(y_mask, mask=~(np.array(masks, dtype=bool)))
428
+ y_min = y.filled(fill_value=1e8)
429
+ y_min = y_min.reshape(y_min.shape[0], -1).min(-1)
430
+
431
+ return np.stack([x_min, y_min, x_max, y_max], 1)
432
+
433
+
434
+ # Copied from transformers.models.detr.image_processing_detr.prepare_coco_panoptic_annotation with DETR->GroundingDino
435
+ def prepare_coco_panoptic_annotation(
436
+ image: np.ndarray,
437
+ target: dict,
438
+ masks_path: Union[str, pathlib.Path],
439
+ return_masks: bool = True,
440
+ input_data_format: Union[ChannelDimension, str] = None,
441
+ ) -> dict:
442
+ """
443
+ Prepare a coco panoptic annotation for GroundingDino.
444
+ """
445
+ image_height, image_width = get_image_size(image, channel_dim=input_data_format)
446
+ annotation_path = pathlib.Path(masks_path) / target["file_name"]
447
+
448
+ new_target = {}
449
+ new_target["image_id"] = np.asarray([target["image_id"] if "image_id" in target else target["id"]], dtype=np.int64)
450
+ new_target["size"] = np.asarray([image_height, image_width], dtype=np.int64)
451
+ new_target["orig_size"] = np.asarray([image_height, image_width], dtype=np.int64)
452
+
453
+ if "segments_info" in target:
454
+ masks = np.asarray(PIL.Image.open(annotation_path), dtype=np.uint32)
455
+ masks = rgb_to_id(masks)
456
+
457
+ ids = np.array([segment_info["id"] for segment_info in target["segments_info"]])
458
+ masks = masks == ids[:, None, None]
459
+ masks = masks.astype(np.uint8)
460
+ if return_masks:
461
+ new_target["masks"] = masks
462
+ new_target["boxes"] = masks_to_boxes(masks)
463
+ new_target["class_labels"] = np.array(
464
+ [segment_info["category_id"] for segment_info in target["segments_info"]], dtype=np.int64
465
+ )
466
+ new_target["iscrowd"] = np.asarray(
467
+ [segment_info["iscrowd"] for segment_info in target["segments_info"]], dtype=np.int64
468
+ )
469
+ new_target["area"] = np.asarray(
470
+ [segment_info["area"] for segment_info in target["segments_info"]], dtype=np.float32
471
+ )
472
+
473
+ return new_target
474
+
475
+
476
+ # Copied from transformers.models.detr.image_processing_detr.get_segmentation_image
477
+ def get_segmentation_image(
478
+ masks: np.ndarray, input_size: tuple, target_size: tuple, stuff_equiv_classes, deduplicate=False
479
+ ):
480
+ h, w = input_size
481
+ final_h, final_w = target_size
482
+
483
+ m_id = scipy.special.softmax(masks.transpose(0, 1), -1)
484
+
485
+ if m_id.shape[-1] == 0:
486
+ # We didn't detect any mask :(
487
+ m_id = np.zeros((h, w), dtype=np.int64)
488
+ else:
489
+ m_id = m_id.argmax(-1).reshape(h, w)
490
+
491
+ if deduplicate:
492
+ # Merge the masks corresponding to the same stuff class
493
+ for equiv in stuff_equiv_classes.values():
494
+ for eq_id in equiv:
495
+ m_id[m_id == eq_id] = equiv[0]
496
+
497
+ seg_img = id_to_rgb(m_id)
498
+ seg_img = resize(seg_img, (final_w, final_h), resample=PILImageResampling.NEAREST)
499
+ return seg_img
500
+
501
+
502
+ # Copied from transformers.models.detr.image_processing_detr.get_mask_area
503
+ def get_mask_area(seg_img: np.ndarray, target_size: tuple[int, int], n_classes: int) -> np.ndarray:
504
+ final_h, final_w = target_size
505
+ np_seg_img = seg_img.astype(np.uint8)
506
+ np_seg_img = np_seg_img.reshape(final_h, final_w, 3)
507
+ m_id = rgb_to_id(np_seg_img)
508
+ area = [(m_id == i).sum() for i in range(n_classes)]
509
+ return area
510
+
511
+
512
+ # Copied from transformers.models.detr.image_processing_detr.score_labels_from_class_probabilities
513
+ def score_labels_from_class_probabilities(logits: np.ndarray) -> tuple[np.ndarray, np.ndarray]:
514
+ probs = scipy.special.softmax(logits, axis=-1)
515
+ labels = probs.argmax(-1, keepdims=True)
516
+ scores = np.take_along_axis(probs, labels, axis=-1)
517
+ scores, labels = scores.squeeze(-1), labels.squeeze(-1)
518
+ return scores, labels
519
+
520
+
521
+ # Copied from transformers.models.detr.image_processing_detr.post_process_panoptic_sample
522
+ def post_process_panoptic_sample(
523
+ out_logits: np.ndarray,
524
+ masks: np.ndarray,
525
+ boxes: np.ndarray,
526
+ processed_size: tuple[int, int],
527
+ target_size: tuple[int, int],
528
+ is_thing_map: dict,
529
+ threshold=0.85,
530
+ ) -> dict:
531
+ """
532
+ Converts the output of [`DetrForSegmentation`] into panoptic segmentation predictions for a single sample.
533
+
534
+ Args:
535
+ out_logits (`torch.Tensor`):
536
+ The logits for this sample.
537
+ masks (`torch.Tensor`):
538
+ The predicted segmentation masks for this sample.
539
+ boxes (`torch.Tensor`):
540
+ The predicted bounding boxes for this sample. The boxes are in the normalized format `(center_x, center_y,
541
+ width, height)` and values between `[0, 1]`, relative to the size the image (disregarding padding).
542
+ processed_size (`tuple[int, int]`):
543
+ The processed size of the image `(height, width)`, as returned by the preprocessing step i.e. the size
544
+ after data augmentation but before batching.
545
+ target_size (`tuple[int, int]`):
546
+ The target size of the image, `(height, width)` corresponding to the requested final size of the
547
+ prediction.
548
+ is_thing_map (`Dict`):
549
+ A dictionary mapping class indices to a boolean value indicating whether the class is a thing or not.
550
+ threshold (`float`, *optional*, defaults to 0.85):
551
+ The threshold used to binarize the segmentation masks.
552
+ """
553
+ # we filter empty queries and detection below threshold
554
+ scores, labels = score_labels_from_class_probabilities(out_logits)
555
+ keep = (labels != out_logits.shape[-1] - 1) & (scores > threshold)
556
+
557
+ cur_scores = scores[keep]
558
+ cur_classes = labels[keep]
559
+ cur_boxes = center_to_corners_format(boxes[keep])
560
+
561
+ if len(cur_boxes) != len(cur_classes):
562
+ raise ValueError("Not as many boxes as there are classes")
563
+
564
+ cur_masks = masks[keep]
565
+ cur_masks = resize(cur_masks[:, None], processed_size, resample=PILImageResampling.BILINEAR)
566
+ cur_masks = safe_squeeze(cur_masks, 1)
567
+ b, h, w = cur_masks.shape
568
+
569
+ # It may be that we have several predicted masks for the same stuff class.
570
+ # In the following, we track the list of masks ids for each stuff class (they are merged later on)
571
+ cur_masks = cur_masks.reshape(b, -1)
572
+ stuff_equiv_classes = defaultdict(list)
573
+ for k, label in enumerate(cur_classes):
574
+ if not is_thing_map[label]:
575
+ stuff_equiv_classes[label].append(k)
576
+
577
+ seg_img = get_segmentation_image(cur_masks, processed_size, target_size, stuff_equiv_classes, deduplicate=True)
578
+ area = get_mask_area(cur_masks, processed_size, n_classes=len(cur_scores))
579
+
580
+ # We filter out any mask that is too small
581
+ if cur_classes.size() > 0:
582
+ # We know filter empty masks as long as we find some
583
+ filtered_small = np.array([a <= 4 for a in area], dtype=bool)
584
+ while filtered_small.any():
585
+ cur_masks = cur_masks[~filtered_small]
586
+ cur_scores = cur_scores[~filtered_small]
587
+ cur_classes = cur_classes[~filtered_small]
588
+ seg_img = get_segmentation_image(cur_masks, (h, w), target_size, stuff_equiv_classes, deduplicate=True)
589
+ area = get_mask_area(seg_img, target_size, n_classes=len(cur_scores))
590
+ filtered_small = np.array([a <= 4 for a in area], dtype=bool)
591
+ else:
592
+ cur_classes = np.ones((1, 1), dtype=np.int64)
593
+
594
+ segments_info = [
595
+ {"id": i, "isthing": is_thing_map[cat], "category_id": int(cat), "area": a}
596
+ for i, (cat, a) in enumerate(zip(cur_classes, area))
597
+ ]
598
+ del cur_classes
599
+
600
+ with io.BytesIO() as out:
601
+ PIL.Image.fromarray(seg_img).save(out, format="PNG")
602
+ predictions = {"png_string": out.getvalue(), "segments_info": segments_info}
603
+
604
+ return predictions
605
+
606
+
607
+ # Copied from transformers.models.detr.image_processing_detr.resize_annotation
608
+ def resize_annotation(
609
+ annotation: dict[str, Any],
610
+ orig_size: tuple[int, int],
611
+ target_size: tuple[int, int],
612
+ threshold: float = 0.5,
613
+ resample: PILImageResampling = PILImageResampling.NEAREST,
614
+ ):
615
+ """
616
+ Resizes an annotation to a target size.
617
+
618
+ Args:
619
+ annotation (`dict[str, Any]`):
620
+ The annotation dictionary.
621
+ orig_size (`tuple[int, int]`):
622
+ The original size of the input image.
623
+ target_size (`tuple[int, int]`):
624
+ The target size of the image, as returned by the preprocessing `resize` step.
625
+ threshold (`float`, *optional*, defaults to 0.5):
626
+ The threshold used to binarize the segmentation masks.
627
+ resample (`PILImageResampling`, defaults to `PILImageResampling.NEAREST`):
628
+ The resampling filter to use when resizing the masks.
629
+ """
630
+ ratios = tuple(float(s) / float(s_orig) for s, s_orig in zip(target_size, orig_size))
631
+ ratio_height, ratio_width = ratios
632
+
633
+ new_annotation = {}
634
+ new_annotation["size"] = target_size
635
+
636
+ for key, value in annotation.items():
637
+ if key == "boxes":
638
+ boxes = value
639
+ scaled_boxes = boxes * np.asarray([ratio_width, ratio_height, ratio_width, ratio_height], dtype=np.float32)
640
+ new_annotation["boxes"] = scaled_boxes
641
+ elif key == "area":
642
+ area = value
643
+ scaled_area = area * (ratio_width * ratio_height)
644
+ new_annotation["area"] = scaled_area
645
+ elif key == "masks":
646
+ masks = value[:, None]
647
+ masks = np.array([resize(mask, target_size, resample=resample) for mask in masks])
648
+ masks = masks.astype(np.float32)
649
+ masks = masks[:, 0] > threshold
650
+ new_annotation["masks"] = masks
651
+ elif key == "size":
652
+ new_annotation["size"] = target_size
653
+ else:
654
+ new_annotation[key] = value
655
+
656
+ return new_annotation
657
+
658
+
659
+ # Copied from transformers.models.detr.image_processing_detr.binary_mask_to_rle
660
+ def binary_mask_to_rle(mask):
661
+ """
662
+ Converts given binary mask of shape `(height, width)` to the run-length encoding (RLE) format.
663
+
664
+ Args:
665
+ mask (`torch.Tensor` or `numpy.array`):
666
+ A binary mask tensor of shape `(height, width)` where 0 denotes background and 1 denotes the target
667
+ segment_id or class_id.
668
+ Returns:
669
+ `List`: Run-length encoded list of the binary mask. Refer to COCO API for more information about the RLE
670
+ format.
671
+ """
672
+ if is_torch_tensor(mask):
673
+ mask = mask.numpy()
674
+
675
+ pixels = mask.flatten()
676
+ pixels = np.concatenate([[0], pixels, [0]])
677
+ runs = np.where(pixels[1:] != pixels[:-1])[0] + 1
678
+ runs[1::2] -= runs[::2]
679
+ return list(runs)
680
+
681
+
682
+ # Copied from transformers.models.detr.image_processing_detr.convert_segmentation_to_rle
683
+ def convert_segmentation_to_rle(segmentation):
684
+ """
685
+ Converts given segmentation map of shape `(height, width)` to the run-length encoding (RLE) format.
686
+
687
+ Args:
688
+ segmentation (`torch.Tensor` or `numpy.array`):
689
+ A segmentation map of shape `(height, width)` where each value denotes a segment or class id.
690
+ Returns:
691
+ `list[List]`: A list of lists, where each list is the run-length encoding of a segment / class id.
692
+ """
693
+ segment_ids = torch.unique(segmentation)
694
+
695
+ run_length_encodings = []
696
+ for idx in segment_ids:
697
+ mask = torch.where(segmentation == idx, 1, 0)
698
+ rle = binary_mask_to_rle(mask)
699
+ run_length_encodings.append(rle)
700
+
701
+ return run_length_encodings
702
+
703
+
704
+ # Copied from transformers.models.detr.image_processing_detr.remove_low_and_no_objects
705
+ def remove_low_and_no_objects(masks, scores, labels, object_mask_threshold, num_labels):
706
+ """
707
+ Binarize the given masks using `object_mask_threshold`, it returns the associated values of `masks`, `scores` and
708
+ `labels`.
709
+
710
+ Args:
711
+ masks (`torch.Tensor`):
712
+ A tensor of shape `(num_queries, height, width)`.
713
+ scores (`torch.Tensor`):
714
+ A tensor of shape `(num_queries)`.
715
+ labels (`torch.Tensor`):
716
+ A tensor of shape `(num_queries)`.
717
+ object_mask_threshold (`float`):
718
+ A number between 0 and 1 used to binarize the masks.
719
+ Raises:
720
+ `ValueError`: Raised when the first dimension doesn't match in all input tensors.
721
+ Returns:
722
+ `tuple[`torch.Tensor`, `torch.Tensor`, `torch.Tensor`]`: The `masks`, `scores` and `labels` without the region
723
+ < `object_mask_threshold`.
724
+ """
725
+ if not (masks.shape[0] == scores.shape[0] == labels.shape[0]):
726
+ raise ValueError("mask, scores and labels must have the same shape!")
727
+
728
+ to_keep = labels.ne(num_labels) & (scores > object_mask_threshold)
729
+
730
+ return masks[to_keep], scores[to_keep], labels[to_keep]
731
+
732
+
733
+ # Copied from transformers.models.detr.image_processing_detr.check_segment_validity
734
+ def check_segment_validity(mask_labels, mask_probs, k, mask_threshold=0.5, overlap_mask_area_threshold=0.8):
735
+ # Get the mask associated with the k class
736
+ mask_k = mask_labels == k
737
+ mask_k_area = mask_k.sum()
738
+
739
+ # Compute the area of all the stuff in query k
740
+ original_area = (mask_probs[k] >= mask_threshold).sum()
741
+ mask_exists = mask_k_area > 0 and original_area > 0
742
+
743
+ # Eliminate disconnected tiny segments
744
+ if mask_exists:
745
+ area_ratio = mask_k_area / original_area
746
+ if not area_ratio.item() > overlap_mask_area_threshold:
747
+ mask_exists = False
748
+
749
+ return mask_exists, mask_k
750
+
751
+
752
+ # Copied from transformers.models.detr.image_processing_detr.compute_segments
753
+ def compute_segments(
754
+ mask_probs,
755
+ pred_scores,
756
+ pred_labels,
757
+ mask_threshold: float = 0.5,
758
+ overlap_mask_area_threshold: float = 0.8,
759
+ label_ids_to_fuse: Optional[set[int]] = None,
760
+ target_size: Optional[tuple[int, int]] = None,
761
+ ):
762
+ height = mask_probs.shape[1] if target_size is None else target_size[0]
763
+ width = mask_probs.shape[2] if target_size is None else target_size[1]
764
+
765
+ segmentation = torch.zeros((height, width), dtype=torch.int32, device=mask_probs.device)
766
+ segments: list[dict] = []
767
+
768
+ if target_size is not None:
769
+ mask_probs = nn.functional.interpolate(
770
+ mask_probs.unsqueeze(0), size=target_size, mode="bilinear", align_corners=False
771
+ )[0]
772
+
773
+ current_segment_id = 0
774
+
775
+ # Weigh each mask by its prediction score
776
+ mask_probs *= pred_scores.view(-1, 1, 1)
777
+ mask_labels = mask_probs.argmax(0) # [height, width]
778
+
779
+ # Keep track of instances of each class
780
+ stuff_memory_list: dict[str, int] = {}
781
+ for k in range(pred_labels.shape[0]):
782
+ pred_class = pred_labels[k].item()
783
+ should_fuse = pred_class in label_ids_to_fuse
784
+
785
+ # Check if mask exists and large enough to be a segment
786
+ mask_exists, mask_k = check_segment_validity(
787
+ mask_labels, mask_probs, k, mask_threshold, overlap_mask_area_threshold
788
+ )
789
+
790
+ if mask_exists:
791
+ if pred_class in stuff_memory_list:
792
+ current_segment_id = stuff_memory_list[pred_class]
793
+ else:
794
+ current_segment_id += 1
795
+
796
+ # Add current object segment to final segmentation map
797
+ segmentation[mask_k] = current_segment_id
798
+ segment_score = round(pred_scores[k].item(), 6)
799
+ segments.append(
800
+ {
801
+ "id": current_segment_id,
802
+ "label_id": pred_class,
803
+ "was_fused": should_fuse,
804
+ "score": segment_score,
805
+ }
806
+ )
807
+ if should_fuse:
808
+ stuff_memory_list[pred_class] = current_segment_id
809
+
810
+ return segmentation, segments
811
+
812
+
813
+ # Copied from transformers.models.owlvit.image_processing_owlvit._scale_boxes
814
+ def _scale_boxes(boxes, target_sizes):
815
+ """
816
+ Scale batch of bounding boxes to the target sizes.
817
+
818
+ Args:
819
+ boxes (`torch.Tensor` of shape `(batch_size, num_boxes, 4)`):
820
+ Bounding boxes to scale. Each box is expected to be in (x1, y1, x2, y2) format.
821
+ target_sizes (`list[tuple[int, int]]` or `torch.Tensor` of shape `(batch_size, 2)`):
822
+ Target sizes to scale the boxes to. Each target size is expected to be in (height, width) format.
823
+
824
+ Returns:
825
+ `torch.Tensor` of shape `(batch_size, num_boxes, 4)`: Scaled bounding boxes.
826
+ """
827
+
828
+ if isinstance(target_sizes, (list, tuple)):
829
+ image_height = torch.tensor([i[0] for i in target_sizes])
830
+ image_width = torch.tensor([i[1] for i in target_sizes])
831
+ elif isinstance(target_sizes, torch.Tensor):
832
+ image_height, image_width = target_sizes.unbind(1)
833
+ else:
834
+ raise TypeError("`target_sizes` must be a list, tuple or torch.Tensor")
835
+
836
+ scale_factor = torch.stack([image_width, image_height, image_width, image_height], dim=1)
837
+ scale_factor = scale_factor.unsqueeze(1).to(boxes.device)
838
+ boxes = boxes * scale_factor
839
+ return boxes
840
+
841
+
842
+ class GroundingDinoImageProcessor(BaseImageProcessor):
843
+ r"""
844
+ Constructs a Grounding DINO image processor.
845
+
846
+ Args:
847
+ format (`str`, *optional*, defaults to `AnnotationFormat.COCO_DETECTION`):
848
+ Data format of the annotations. One of "coco_detection" or "coco_panoptic".
849
+ do_resize (`bool`, *optional*, defaults to `True`):
850
+ Controls whether to resize the image's (height, width) dimensions to the specified `size`. Can be
851
+ overridden by the `do_resize` parameter in the `preprocess` method.
852
+ size (`dict[str, int]` *optional*, defaults to `{"shortest_edge": 800, "longest_edge": 1333}`):
853
+ Size of the image's `(height, width)` dimensions after resizing. Can be overridden by the `size` parameter
854
+ in the `preprocess` method. Available options are:
855
+ - `{"height": int, "width": int}`: The image will be resized to the exact size `(height, width)`.
856
+ Do NOT keep the aspect ratio.
857
+ - `{"shortest_edge": int, "longest_edge": int}`: The image will be resized to a maximum size respecting
858
+ the aspect ratio and keeping the shortest edge less or equal to `shortest_edge` and the longest edge
859
+ less or equal to `longest_edge`.
860
+ - `{"max_height": int, "max_width": int}`: The image will be resized to the maximum size respecting the
861
+ aspect ratio and keeping the height less or equal to `max_height` and the width less or equal to
862
+ `max_width`.
863
+ resample (`PILImageResampling`, *optional*, defaults to `Resampling.BILINEAR`):
864
+ Resampling filter to use if resizing the image.
865
+ do_rescale (`bool`, *optional*, defaults to `True`):
866
+ Controls whether to rescale the image by the specified scale `rescale_factor`. Can be overridden by the
867
+ `do_rescale` parameter in the `preprocess` method.
868
+ rescale_factor (`int` or `float`, *optional*, defaults to `1/255`):
869
+ Scale factor to use if rescaling the image. Can be overridden by the `rescale_factor` parameter in the
870
+ `preprocess` method. Controls whether to normalize the image. Can be overridden by the `do_normalize`
871
+ parameter in the `preprocess` method.
872
+ do_normalize (`bool`, *optional*, defaults to `True`):
873
+ Whether to normalize the image. Can be overridden by the `do_normalize` parameter in the `preprocess`
874
+ method.
875
+ image_mean (`float` or `list[float]`, *optional*, defaults to `IMAGENET_DEFAULT_MEAN`):
876
+ Mean values to use when normalizing the image. Can be a single value or a list of values, one for each
877
+ channel. Can be overridden by the `image_mean` parameter in the `preprocess` method.
878
+ image_std (`float` or `list[float]`, *optional*, defaults to `IMAGENET_DEFAULT_STD`):
879
+ Standard deviation values to use when normalizing the image. Can be a single value or a list of values, one
880
+ for each channel. Can be overridden by the `image_std` parameter in the `preprocess` method.
881
+ do_convert_annotations (`bool`, *optional*, defaults to `True`):
882
+ Controls whether to convert the annotations to the format expected by the DETR model. Converts the
883
+ bounding boxes to the format `(center_x, center_y, width, height)` and in the range `[0, 1]`.
884
+ Can be overridden by the `do_convert_annotations` parameter in the `preprocess` method.
885
+ do_pad (`bool`, *optional*, defaults to `True`):
886
+ Controls whether to pad the image. Can be overridden by the `do_pad` parameter in the `preprocess`
887
+ method. If `True`, padding will be applied to the bottom and right of the image with zeros.
888
+ If `pad_size` is provided, the image will be padded to the specified dimensions.
889
+ Otherwise, the image will be padded to the maximum height and width of the batch.
890
+ pad_size (`dict[str, int]`, *optional*):
891
+ The size `{"height": int, "width" int}` to pad the images to. Must be larger than any image size
892
+ provided for preprocessing. If `pad_size` is not provided, images will be padded to the largest
893
+ height and width in the batch.
894
+ """
895
+
896
+ model_input_names = ["pixel_values", "pixel_mask"]
897
+
898
+ # Copied from transformers.models.detr.image_processing_detr.DetrImageProcessor.__init__
899
+ def __init__(
900
+ self,
901
+ format: Union[str, AnnotationFormat] = AnnotationFormat.COCO_DETECTION,
902
+ do_resize: bool = True,
903
+ size: Optional[dict[str, int]] = None,
904
+ resample: PILImageResampling = PILImageResampling.BILINEAR,
905
+ do_rescale: bool = True,
906
+ rescale_factor: Union[int, float] = 1 / 255,
907
+ do_normalize: bool = True,
908
+ image_mean: Optional[Union[float, list[float]]] = None,
909
+ image_std: Optional[Union[float, list[float]]] = None,
910
+ do_convert_annotations: Optional[bool] = None,
911
+ do_pad: bool = True,
912
+ pad_size: Optional[dict[str, int]] = None,
913
+ **kwargs,
914
+ ) -> None:
915
+ if "pad_and_return_pixel_mask" in kwargs:
916
+ do_pad = kwargs.pop("pad_and_return_pixel_mask")
917
+
918
+ if "max_size" in kwargs:
919
+ logger.warning_once(
920
+ "The `max_size` parameter is deprecated and will be removed in v4.26. "
921
+ "Please specify in `size['longest_edge'] instead`.",
922
+ )
923
+ max_size = kwargs.pop("max_size")
924
+ else:
925
+ max_size = None if size is None else 1333
926
+
927
+ size = size if size is not None else {"shortest_edge": 800, "longest_edge": 1333}
928
+ size = get_size_dict(size, max_size=max_size, default_to_square=False)
929
+
930
+ # Backwards compatibility
931
+ if do_convert_annotations is None:
932
+ do_convert_annotations = do_normalize
933
+
934
+ super().__init__(**kwargs)
935
+ self.format = format
936
+ self.do_resize = do_resize
937
+ self.size = size
938
+ self.resample = resample
939
+ self.do_rescale = do_rescale
940
+ self.rescale_factor = rescale_factor
941
+ self.do_normalize = do_normalize
942
+ self.do_convert_annotations = do_convert_annotations
943
+ self.image_mean = image_mean if image_mean is not None else IMAGENET_DEFAULT_MEAN
944
+ self.image_std = image_std if image_std is not None else IMAGENET_DEFAULT_STD
945
+ self.do_pad = do_pad
946
+ self.pad_size = pad_size
947
+ self._valid_processor_keys = [
948
+ "images",
949
+ "annotations",
950
+ "return_segmentation_masks",
951
+ "masks_path",
952
+ "do_resize",
953
+ "size",
954
+ "resample",
955
+ "do_rescale",
956
+ "rescale_factor",
957
+ "do_normalize",
958
+ "do_convert_annotations",
959
+ "image_mean",
960
+ "image_std",
961
+ "do_pad",
962
+ "pad_size",
963
+ "format",
964
+ "return_tensors",
965
+ "data_format",
966
+ "input_data_format",
967
+ ]
968
+
969
+ @classmethod
970
+ # Copied from transformers.models.detr.image_processing_detr.DetrImageProcessor.from_dict with Detr->GroundingDino
971
+ def from_dict(cls, image_processor_dict: dict[str, Any], **kwargs):
972
+ """
973
+ Overrides the `from_dict` method from the base class to make sure parameters are updated if image processor is
974
+ created using from_dict and kwargs e.g. `GroundingDinoImageProcessor.from_pretrained(checkpoint, size=600,
975
+ max_size=800)`
976
+ """
977
+ image_processor_dict = image_processor_dict.copy()
978
+ if "max_size" in kwargs:
979
+ image_processor_dict["max_size"] = kwargs.pop("max_size")
980
+ if "pad_and_return_pixel_mask" in kwargs:
981
+ image_processor_dict["pad_and_return_pixel_mask"] = kwargs.pop("pad_and_return_pixel_mask")
982
+ return super().from_dict(image_processor_dict, **kwargs)
983
+
984
+ # Copied from transformers.models.detr.image_processing_detr.DetrImageProcessor.prepare_annotation with DETR->GroundingDino
985
+ def prepare_annotation(
986
+ self,
987
+ image: np.ndarray,
988
+ target: dict,
989
+ format: Optional[AnnotationFormat] = None,
990
+ return_segmentation_masks: Optional[bool] = None,
991
+ masks_path: Optional[Union[str, pathlib.Path]] = None,
992
+ input_data_format: Optional[Union[str, ChannelDimension]] = None,
993
+ ) -> dict:
994
+ """
995
+ Prepare an annotation for feeding into GroundingDino model.
996
+ """
997
+ format = format if format is not None else self.format
998
+
999
+ if format == AnnotationFormat.COCO_DETECTION:
1000
+ return_segmentation_masks = False if return_segmentation_masks is None else return_segmentation_masks
1001
+ target = prepare_coco_detection_annotation(
1002
+ image, target, return_segmentation_masks, input_data_format=input_data_format
1003
+ )
1004
+ elif format == AnnotationFormat.COCO_PANOPTIC:
1005
+ return_segmentation_masks = True if return_segmentation_masks is None else return_segmentation_masks
1006
+ target = prepare_coco_panoptic_annotation(
1007
+ image,
1008
+ target,
1009
+ masks_path=masks_path,
1010
+ return_masks=return_segmentation_masks,
1011
+ input_data_format=input_data_format,
1012
+ )
1013
+ else:
1014
+ raise ValueError(f"Format {format} is not supported.")
1015
+ return target
1016
+
1017
+ # Copied from transformers.models.detr.image_processing_detr.DetrImageProcessor.resize
1018
+ def resize(
1019
+ self,
1020
+ image: np.ndarray,
1021
+ size: dict[str, int],
1022
+ resample: PILImageResampling = PILImageResampling.BILINEAR,
1023
+ data_format: Optional[ChannelDimension] = None,
1024
+ input_data_format: Optional[Union[str, ChannelDimension]] = None,
1025
+ **kwargs,
1026
+ ) -> np.ndarray:
1027
+ """
1028
+ Resize the image to the given size. Size can be `min_size` (scalar) or `(height, width)` tuple. If size is an
1029
+ int, smaller edge of the image will be matched to this number.
1030
+
1031
+ Args:
1032
+ image (`np.ndarray`):
1033
+ Image to resize.
1034
+ size (`dict[str, int]`):
1035
+ Size of the image's `(height, width)` dimensions after resizing. Available options are:
1036
+ - `{"height": int, "width": int}`: The image will be resized to the exact size `(height, width)`.
1037
+ Do NOT keep the aspect ratio.
1038
+ - `{"shortest_edge": int, "longest_edge": int}`: The image will be resized to a maximum size respecting
1039
+ the aspect ratio and keeping the shortest edge less or equal to `shortest_edge` and the longest edge
1040
+ less or equal to `longest_edge`.
1041
+ - `{"max_height": int, "max_width": int}`: The image will be resized to the maximum size respecting the
1042
+ aspect ratio and keeping the height less or equal to `max_height` and the width less or equal to
1043
+ `max_width`.
1044
+ resample (`PILImageResampling`, *optional*, defaults to `PILImageResampling.BILINEAR`):
1045
+ Resampling filter to use if resizing the image.
1046
+ data_format (`str` or `ChannelDimension`, *optional*):
1047
+ The channel dimension format for the output image. If unset, the channel dimension format of the input
1048
+ image is used.
1049
+ input_data_format (`ChannelDimension` or `str`, *optional*):
1050
+ The channel dimension format of the input image. If not provided, it will be inferred.
1051
+ """
1052
+ if "max_size" in kwargs:
1053
+ logger.warning_once(
1054
+ "The `max_size` parameter is deprecated and will be removed in v4.26. "
1055
+ "Please specify in `size['longest_edge'] instead`.",
1056
+ )
1057
+ max_size = kwargs.pop("max_size")
1058
+ else:
1059
+ max_size = None
1060
+ size = get_size_dict(size, max_size=max_size, default_to_square=False)
1061
+ if "shortest_edge" in size and "longest_edge" in size:
1062
+ new_size = get_resize_output_image_size(
1063
+ image, size["shortest_edge"], size["longest_edge"], input_data_format=input_data_format
1064
+ )
1065
+ elif "max_height" in size and "max_width" in size:
1066
+ new_size = get_image_size_for_max_height_width(
1067
+ image, size["max_height"], size["max_width"], input_data_format=input_data_format
1068
+ )
1069
+ elif "height" in size and "width" in size:
1070
+ new_size = (size["height"], size["width"])
1071
+ else:
1072
+ raise ValueError(
1073
+ "Size must contain 'height' and 'width' keys or 'shortest_edge' and 'longest_edge' keys. Got"
1074
+ f" {size.keys()}."
1075
+ )
1076
+ image = resize(
1077
+ image,
1078
+ size=new_size,
1079
+ resample=resample,
1080
+ data_format=data_format,
1081
+ input_data_format=input_data_format,
1082
+ **kwargs,
1083
+ )
1084
+ return image
1085
+
1086
+ # Copied from transformers.models.detr.image_processing_detr.DetrImageProcessor.resize_annotation
1087
+ def resize_annotation(
1088
+ self,
1089
+ annotation,
1090
+ orig_size,
1091
+ size,
1092
+ resample: PILImageResampling = PILImageResampling.NEAREST,
1093
+ ) -> dict:
1094
+ """
1095
+ Resize the annotation to match the resized image. If size is an int, smaller edge of the mask will be matched
1096
+ to this number.
1097
+ """
1098
+ return resize_annotation(annotation, orig_size=orig_size, target_size=size, resample=resample)
1099
+
1100
+ # Copied from transformers.models.detr.image_processing_detr.DetrImageProcessor.rescale
1101
+ def rescale(
1102
+ self,
1103
+ image: np.ndarray,
1104
+ rescale_factor: float,
1105
+ data_format: Optional[Union[str, ChannelDimension]] = None,
1106
+ input_data_format: Optional[Union[str, ChannelDimension]] = None,
1107
+ ) -> np.ndarray:
1108
+ """
1109
+ Rescale the image by the given factor. image = image * rescale_factor.
1110
+
1111
+ Args:
1112
+ image (`np.ndarray`):
1113
+ Image to rescale.
1114
+ rescale_factor (`float`):
1115
+ The value to use for rescaling.
1116
+ data_format (`str` or `ChannelDimension`, *optional*):
1117
+ The channel dimension format for the output image. If unset, the channel dimension format of the input
1118
+ image is used. Can be one of:
1119
+ - `"channels_first"` or `ChannelDimension.FIRST`: image in (num_channels, height, width) format.
1120
+ - `"channels_last"` or `ChannelDimension.LAST`: image in (height, width, num_channels) format.
1121
+ input_data_format (`str` or `ChannelDimension`, *optional*):
1122
+ The channel dimension format for the input image. If unset, is inferred from the input image. Can be
1123
+ one of:
1124
+ - `"channels_first"` or `ChannelDimension.FIRST`: image in (num_channels, height, width) format.
1125
+ - `"channels_last"` or `ChannelDimension.LAST`: image in (height, width, num_channels) format.
1126
+ """
1127
+ return rescale(image, rescale_factor, data_format=data_format, input_data_format=input_data_format)
1128
+
1129
+ # Copied from transformers.models.detr.image_processing_detr.DetrImageProcessor.normalize_annotation
1130
+ def normalize_annotation(self, annotation: dict, image_size: tuple[int, int]) -> dict:
1131
+ """
1132
+ Normalize the boxes in the annotation from `[top_left_x, top_left_y, bottom_right_x, bottom_right_y]` to
1133
+ `[center_x, center_y, width, height]` format and from absolute to relative pixel values.
1134
+ """
1135
+ return normalize_annotation(annotation, image_size=image_size)
1136
+
1137
+ # Copied from transformers.models.detr.image_processing_detr.DetrImageProcessor._update_annotation_for_padded_image
1138
+ def _update_annotation_for_padded_image(
1139
+ self,
1140
+ annotation: dict,
1141
+ input_image_size: tuple[int, int],
1142
+ output_image_size: tuple[int, int],
1143
+ padding,
1144
+ update_bboxes,
1145
+ ) -> dict:
1146
+ """
1147
+ Update the annotation for a padded image.
1148
+ """
1149
+ new_annotation = {}
1150
+ new_annotation["size"] = output_image_size
1151
+
1152
+ for key, value in annotation.items():
1153
+ if key == "masks":
1154
+ masks = value
1155
+ masks = pad(
1156
+ masks,
1157
+ padding,
1158
+ mode=PaddingMode.CONSTANT,
1159
+ constant_values=0,
1160
+ input_data_format=ChannelDimension.FIRST,
1161
+ )
1162
+ masks = safe_squeeze(masks, 1)
1163
+ new_annotation["masks"] = masks
1164
+ elif key == "boxes" and update_bboxes:
1165
+ boxes = value
1166
+ boxes *= np.asarray(
1167
+ [
1168
+ input_image_size[1] / output_image_size[1],
1169
+ input_image_size[0] / output_image_size[0],
1170
+ input_image_size[1] / output_image_size[1],
1171
+ input_image_size[0] / output_image_size[0],
1172
+ ]
1173
+ )
1174
+ new_annotation["boxes"] = boxes
1175
+ elif key == "size":
1176
+ new_annotation["size"] = output_image_size
1177
+ else:
1178
+ new_annotation[key] = value
1179
+ return new_annotation
1180
+
1181
+ # Copied from transformers.models.detr.image_processing_detr.DetrImageProcessor._pad_image
1182
+ def _pad_image(
1183
+ self,
1184
+ image: np.ndarray,
1185
+ output_size: tuple[int, int],
1186
+ annotation: Optional[dict[str, Any]] = None,
1187
+ constant_values: Union[float, Iterable[float]] = 0,
1188
+ data_format: Optional[ChannelDimension] = None,
1189
+ input_data_format: Optional[Union[str, ChannelDimension]] = None,
1190
+ update_bboxes: bool = True,
1191
+ ) -> np.ndarray:
1192
+ """
1193
+ Pad an image with zeros to the given size.
1194
+ """
1195
+ input_height, input_width = get_image_size(image, channel_dim=input_data_format)
1196
+ output_height, output_width = output_size
1197
+
1198
+ pad_bottom = output_height - input_height
1199
+ pad_right = output_width - input_width
1200
+ padding = ((0, pad_bottom), (0, pad_right))
1201
+ padded_image = pad(
1202
+ image,
1203
+ padding,
1204
+ mode=PaddingMode.CONSTANT,
1205
+ constant_values=constant_values,
1206
+ data_format=data_format,
1207
+ input_data_format=input_data_format,
1208
+ )
1209
+ if annotation is not None:
1210
+ annotation = self._update_annotation_for_padded_image(
1211
+ annotation, (input_height, input_width), (output_height, output_width), padding, update_bboxes
1212
+ )
1213
+ return padded_image, annotation
1214
+
1215
+ # Copied from transformers.models.detr.image_processing_detr.DetrImageProcessor.pad
1216
+ def pad(
1217
+ self,
1218
+ images: list[np.ndarray],
1219
+ annotations: Optional[Union[AnnotationType, list[AnnotationType]]] = None,
1220
+ constant_values: Union[float, Iterable[float]] = 0,
1221
+ return_pixel_mask: bool = True,
1222
+ return_tensors: Optional[Union[str, TensorType]] = None,
1223
+ data_format: Optional[ChannelDimension] = None,
1224
+ input_data_format: Optional[Union[str, ChannelDimension]] = None,
1225
+ update_bboxes: bool = True,
1226
+ pad_size: Optional[dict[str, int]] = None,
1227
+ ) -> BatchFeature:
1228
+ """
1229
+ Pads a batch of images to the bottom and right of the image with zeros to the size of largest height and width
1230
+ in the batch and optionally returns their corresponding pixel mask.
1231
+
1232
+ Args:
1233
+ images (list[`np.ndarray`]):
1234
+ Images to pad.
1235
+ annotations (`AnnotationType` or `list[AnnotationType]`, *optional*):
1236
+ Annotations to transform according to the padding that is applied to the images.
1237
+ constant_values (`float` or `Iterable[float]`, *optional*):
1238
+ The value to use for the padding if `mode` is `"constant"`.
1239
+ return_pixel_mask (`bool`, *optional*, defaults to `True`):
1240
+ Whether to return a pixel mask.
1241
+ return_tensors (`str` or `TensorType`, *optional*):
1242
+ The type of tensors to return. Can be one of:
1243
+ - Unset: Return a list of `np.ndarray`.
1244
+ - `TensorType.TENSORFLOW` or `'tf'`: Return a batch of type `tf.Tensor`.
1245
+ - `TensorType.PYTORCH` or `'pt'`: Return a batch of type `torch.Tensor`.
1246
+ - `TensorType.NUMPY` or `'np'`: Return a batch of type `np.ndarray`.
1247
+ - `TensorType.JAX` or `'jax'`: Return a batch of type `jax.numpy.ndarray`.
1248
+ data_format (`str` or `ChannelDimension`, *optional*):
1249
+ The channel dimension format of the image. If not provided, it will be the same as the input image.
1250
+ input_data_format (`ChannelDimension` or `str`, *optional*):
1251
+ The channel dimension format of the input image. If not provided, it will be inferred.
1252
+ update_bboxes (`bool`, *optional*, defaults to `True`):
1253
+ Whether to update the bounding boxes in the annotations to match the padded images. If the
1254
+ bounding boxes have not been converted to relative coordinates and `(centre_x, centre_y, width, height)`
1255
+ format, the bounding boxes will not be updated.
1256
+ pad_size (`dict[str, int]`, *optional*):
1257
+ The size `{"height": int, "width" int}` to pad the images to. Must be larger than any image size
1258
+ provided for preprocessing. If `pad_size` is not provided, images will be padded to the largest
1259
+ height and width in the batch.
1260
+ """
1261
+ pad_size = pad_size if pad_size is not None else self.pad_size
1262
+ if pad_size is not None:
1263
+ padded_size = (pad_size["height"], pad_size["width"])
1264
+ else:
1265
+ padded_size = get_max_height_width(images, input_data_format=input_data_format)
1266
+
1267
+ annotation_list = annotations if annotations is not None else [None] * len(images)
1268
+ padded_images = []
1269
+ padded_annotations = []
1270
+ for image, annotation in zip(images, annotation_list):
1271
+ padded_image, padded_annotation = self._pad_image(
1272
+ image,
1273
+ padded_size,
1274
+ annotation,
1275
+ constant_values=constant_values,
1276
+ data_format=data_format,
1277
+ input_data_format=input_data_format,
1278
+ update_bboxes=update_bboxes,
1279
+ )
1280
+ padded_images.append(padded_image)
1281
+ padded_annotations.append(padded_annotation)
1282
+
1283
+ data = {"pixel_values": padded_images}
1284
+
1285
+ if return_pixel_mask:
1286
+ masks = [
1287
+ make_pixel_mask(image=image, output_size=padded_size, input_data_format=input_data_format)
1288
+ for image in images
1289
+ ]
1290
+ data["pixel_mask"] = masks
1291
+
1292
+ encoded_inputs = BatchFeature(data=data, tensor_type=return_tensors)
1293
+
1294
+ if annotations is not None:
1295
+ encoded_inputs["labels"] = [
1296
+ BatchFeature(annotation, tensor_type=return_tensors) for annotation in padded_annotations
1297
+ ]
1298
+
1299
+ return encoded_inputs
1300
+
1301
+ # Copied from transformers.models.detr.image_processing_detr.DetrImageProcessor.preprocess
1302
+ def preprocess(
1303
+ self,
1304
+ images: ImageInput,
1305
+ annotations: Optional[Union[AnnotationType, list[AnnotationType]]] = None,
1306
+ return_segmentation_masks: Optional[bool] = None,
1307
+ masks_path: Optional[Union[str, pathlib.Path]] = None,
1308
+ do_resize: Optional[bool] = None,
1309
+ size: Optional[dict[str, int]] = None,
1310
+ resample=None, # PILImageResampling
1311
+ do_rescale: Optional[bool] = None,
1312
+ rescale_factor: Optional[Union[int, float]] = None,
1313
+ do_normalize: Optional[bool] = None,
1314
+ do_convert_annotations: Optional[bool] = None,
1315
+ image_mean: Optional[Union[float, list[float]]] = None,
1316
+ image_std: Optional[Union[float, list[float]]] = None,
1317
+ do_pad: Optional[bool] = None,
1318
+ format: Optional[Union[str, AnnotationFormat]] = None,
1319
+ return_tensors: Optional[Union[TensorType, str]] = None,
1320
+ data_format: Union[str, ChannelDimension] = ChannelDimension.FIRST,
1321
+ input_data_format: Optional[Union[str, ChannelDimension]] = None,
1322
+ pad_size: Optional[dict[str, int]] = None,
1323
+ **kwargs,
1324
+ ) -> BatchFeature:
1325
+ """
1326
+ Preprocess an image or a batch of images so that it can be used by the model.
1327
+
1328
+ Args:
1329
+ images (`ImageInput`):
1330
+ Image or batch of images to preprocess. Expects a single or batch of images with pixel values ranging
1331
+ from 0 to 255. If passing in images with pixel values between 0 and 1, set `do_rescale=False`.
1332
+ annotations (`AnnotationType` or `list[AnnotationType]`, *optional*):
1333
+ List of annotations associated with the image or batch of images. If annotation is for object
1334
+ detection, the annotations should be a dictionary with the following keys:
1335
+ - "image_id" (`int`): The image id.
1336
+ - "annotations" (`list[Dict]`): List of annotations for an image. Each annotation should be a
1337
+ dictionary. An image can have no annotations, in which case the list should be empty.
1338
+ If annotation is for segmentation, the annotations should be a dictionary with the following keys:
1339
+ - "image_id" (`int`): The image id.
1340
+ - "segments_info" (`list[Dict]`): List of segments for an image. Each segment should be a dictionary.
1341
+ An image can have no segments, in which case the list should be empty.
1342
+ - "file_name" (`str`): The file name of the image.
1343
+ return_segmentation_masks (`bool`, *optional*, defaults to self.return_segmentation_masks):
1344
+ Whether to return segmentation masks.
1345
+ masks_path (`str` or `pathlib.Path`, *optional*):
1346
+ Path to the directory containing the segmentation masks.
1347
+ do_resize (`bool`, *optional*, defaults to self.do_resize):
1348
+ Whether to resize the image.
1349
+ size (`dict[str, int]`, *optional*, defaults to self.size):
1350
+ Size of the image's `(height, width)` dimensions after resizing. Available options are:
1351
+ - `{"height": int, "width": int}`: The image will be resized to the exact size `(height, width)`.
1352
+ Do NOT keep the aspect ratio.
1353
+ - `{"shortest_edge": int, "longest_edge": int}`: The image will be resized to a maximum size respecting
1354
+ the aspect ratio and keeping the shortest edge less or equal to `shortest_edge` and the longest edge
1355
+ less or equal to `longest_edge`.
1356
+ - `{"max_height": int, "max_width": int}`: The image will be resized to the maximum size respecting the
1357
+ aspect ratio and keeping the height less or equal to `max_height` and the width less or equal to
1358
+ `max_width`.
1359
+ resample (`PILImageResampling`, *optional*, defaults to self.resample):
1360
+ Resampling filter to use when resizing the image.
1361
+ do_rescale (`bool`, *optional*, defaults to self.do_rescale):
1362
+ Whether to rescale the image.
1363
+ rescale_factor (`float`, *optional*, defaults to self.rescale_factor):
1364
+ Rescale factor to use when rescaling the image.
1365
+ do_normalize (`bool`, *optional*, defaults to self.do_normalize):
1366
+ Whether to normalize the image.
1367
+ do_convert_annotations (`bool`, *optional*, defaults to self.do_convert_annotations):
1368
+ Whether to convert the annotations to the format expected by the model. Converts the bounding
1369
+ boxes from the format `(top_left_x, top_left_y, width, height)` to `(center_x, center_y, width, height)`
1370
+ and in relative coordinates.
1371
+ image_mean (`float` or `list[float]`, *optional*, defaults to self.image_mean):
1372
+ Mean to use when normalizing the image.
1373
+ image_std (`float` or `list[float]`, *optional*, defaults to self.image_std):
1374
+ Standard deviation to use when normalizing the image.
1375
+ do_pad (`bool`, *optional*, defaults to self.do_pad):
1376
+ Whether to pad the image. If `True`, padding will be applied to the bottom and right of
1377
+ the image with zeros. If `pad_size` is provided, the image will be padded to the specified
1378
+ dimensions. Otherwise, the image will be padded to the maximum height and width of the batch.
1379
+ format (`str` or `AnnotationFormat`, *optional*, defaults to self.format):
1380
+ Format of the annotations.
1381
+ return_tensors (`str` or `TensorType`, *optional*, defaults to self.return_tensors):
1382
+ Type of tensors to return. If `None`, will return the list of images.
1383
+ data_format (`ChannelDimension` or `str`, *optional*, defaults to `ChannelDimension.FIRST`):
1384
+ The channel dimension format for the output image. Can be one of:
1385
+ - `"channels_first"` or `ChannelDimension.FIRST`: image in (num_channels, height, width) format.
1386
+ - `"channels_last"` or `ChannelDimension.LAST`: image in (height, width, num_channels) format.
1387
+ - Unset: Use the channel dimension format of the input image.
1388
+ input_data_format (`ChannelDimension` or `str`, *optional*):
1389
+ The channel dimension format for the input image. If unset, the channel dimension format is inferred
1390
+ from the input image. Can be one of:
1391
+ - `"channels_first"` or `ChannelDimension.FIRST`: image in (num_channels, height, width) format.
1392
+ - `"channels_last"` or `ChannelDimension.LAST`: image in (height, width, num_channels) format.
1393
+ - `"none"` or `ChannelDimension.NONE`: image in (height, width) format.
1394
+ pad_size (`dict[str, int]`, *optional*):
1395
+ The size `{"height": int, "width" int}` to pad the images to. Must be larger than any image size
1396
+ provided for preprocessing. If `pad_size` is not provided, images will be padded to the largest
1397
+ height and width in the batch.
1398
+ """
1399
+ if "pad_and_return_pixel_mask" in kwargs:
1400
+ logger.warning_once(
1401
+ "The `pad_and_return_pixel_mask` argument is deprecated and will be removed in a future version, "
1402
+ "use `do_pad` instead."
1403
+ )
1404
+ do_pad = kwargs.pop("pad_and_return_pixel_mask")
1405
+
1406
+ if "max_size" in kwargs:
1407
+ logger.warning_once(
1408
+ "The `max_size` argument is deprecated and will be removed in a future version, use"
1409
+ " `size['longest_edge']` instead."
1410
+ )
1411
+ size = kwargs.pop("max_size")
1412
+
1413
+ do_resize = self.do_resize if do_resize is None else do_resize
1414
+ size = self.size if size is None else size
1415
+ size = get_size_dict(size=size, default_to_square=False)
1416
+ resample = self.resample if resample is None else resample
1417
+ do_rescale = self.do_rescale if do_rescale is None else do_rescale
1418
+ rescale_factor = self.rescale_factor if rescale_factor is None else rescale_factor
1419
+ do_normalize = self.do_normalize if do_normalize is None else do_normalize
1420
+ image_mean = self.image_mean if image_mean is None else image_mean
1421
+ image_std = self.image_std if image_std is None else image_std
1422
+ do_convert_annotations = (
1423
+ self.do_convert_annotations if do_convert_annotations is None else do_convert_annotations
1424
+ )
1425
+ do_pad = self.do_pad if do_pad is None else do_pad
1426
+ pad_size = self.pad_size if pad_size is None else pad_size
1427
+ format = self.format if format is None else format
1428
+
1429
+ images = make_flat_list_of_images(images)
1430
+
1431
+ if not valid_images(images):
1432
+ raise ValueError(
1433
+ "Invalid image type. Must be of type PIL.Image.Image, numpy.ndarray, "
1434
+ "torch.Tensor, tf.Tensor or jax.ndarray."
1435
+ )
1436
+ validate_kwargs(captured_kwargs=kwargs.keys(), valid_processor_keys=self._valid_processor_keys)
1437
+
1438
+ # Here, the pad() method pads to the maximum of (width, height). It does not need to be validated.
1439
+ validate_preprocess_arguments(
1440
+ do_rescale=do_rescale,
1441
+ rescale_factor=rescale_factor,
1442
+ do_normalize=do_normalize,
1443
+ image_mean=image_mean,
1444
+ image_std=image_std,
1445
+ do_resize=do_resize,
1446
+ size=size,
1447
+ resample=resample,
1448
+ )
1449
+
1450
+ if annotations is not None and isinstance(annotations, dict):
1451
+ annotations = [annotations]
1452
+
1453
+ if annotations is not None and len(images) != len(annotations):
1454
+ raise ValueError(
1455
+ f"The number of images ({len(images)}) and annotations ({len(annotations)}) do not match."
1456
+ )
1457
+
1458
+ format = AnnotationFormat(format)
1459
+ if annotations is not None:
1460
+ validate_annotations(format, SUPPORTED_ANNOTATION_FORMATS, annotations)
1461
+
1462
+ if (
1463
+ masks_path is not None
1464
+ and format == AnnotationFormat.COCO_PANOPTIC
1465
+ and not isinstance(masks_path, (pathlib.Path, str))
1466
+ ):
1467
+ raise ValueError(
1468
+ "The path to the directory containing the mask PNG files should be provided as a"
1469
+ f" `pathlib.Path` or string object, but is {type(masks_path)} instead."
1470
+ )
1471
+
1472
+ # All transformations expect numpy arrays
1473
+ images = [to_numpy_array(image) for image in images]
1474
+
1475
+ if do_rescale and is_scaled_image(images[0]):
1476
+ logger.warning_once(
1477
+ "It looks like you are trying to rescale already rescaled images. If the input"
1478
+ " images have pixel values between 0 and 1, set `do_rescale=False` to avoid rescaling them again."
1479
+ )
1480
+
1481
+ if input_data_format is None:
1482
+ # We assume that all images have the same channel dimension format.
1483
+ input_data_format = infer_channel_dimension_format(images[0])
1484
+
1485
+ # prepare (COCO annotations as a list of Dict -> DETR target as a single Dict per image)
1486
+ if annotations is not None:
1487
+ prepared_images = []
1488
+ prepared_annotations = []
1489
+ for image, target in zip(images, annotations):
1490
+ target = self.prepare_annotation(
1491
+ image,
1492
+ target,
1493
+ format,
1494
+ return_segmentation_masks=return_segmentation_masks,
1495
+ masks_path=masks_path,
1496
+ input_data_format=input_data_format,
1497
+ )
1498
+ prepared_images.append(image)
1499
+ prepared_annotations.append(target)
1500
+ images = prepared_images
1501
+ annotations = prepared_annotations
1502
+ del prepared_images, prepared_annotations
1503
+
1504
+ # transformations
1505
+ if do_resize:
1506
+ if annotations is not None:
1507
+ resized_images, resized_annotations = [], []
1508
+ for image, target in zip(images, annotations):
1509
+ orig_size = get_image_size(image, input_data_format)
1510
+ resized_image = self.resize(
1511
+ image, size=size, resample=resample, input_data_format=input_data_format
1512
+ )
1513
+ resized_annotation = self.resize_annotation(
1514
+ target, orig_size, get_image_size(resized_image, input_data_format)
1515
+ )
1516
+ resized_images.append(resized_image)
1517
+ resized_annotations.append(resized_annotation)
1518
+ images = resized_images
1519
+ annotations = resized_annotations
1520
+ del resized_images, resized_annotations
1521
+ else:
1522
+ images = [
1523
+ self.resize(image, size=size, resample=resample, input_data_format=input_data_format)
1524
+ for image in images
1525
+ ]
1526
+
1527
+ if do_rescale:
1528
+ images = [self.rescale(image, rescale_factor, input_data_format=input_data_format) for image in images]
1529
+
1530
+ if do_normalize:
1531
+ images = [
1532
+ self.normalize(image, image_mean, image_std, input_data_format=input_data_format) for image in images
1533
+ ]
1534
+
1535
+ if do_convert_annotations and annotations is not None:
1536
+ annotations = [
1537
+ self.normalize_annotation(annotation, get_image_size(image, input_data_format))
1538
+ for annotation, image in zip(annotations, images)
1539
+ ]
1540
+
1541
+ if do_pad:
1542
+ # Pads images and returns their mask: {'pixel_values': ..., 'pixel_mask': ...}
1543
+ encoded_inputs = self.pad(
1544
+ images,
1545
+ annotations=annotations,
1546
+ return_pixel_mask=True,
1547
+ data_format=data_format,
1548
+ input_data_format=input_data_format,
1549
+ update_bboxes=do_convert_annotations,
1550
+ return_tensors=return_tensors,
1551
+ pad_size=pad_size,
1552
+ )
1553
+ else:
1554
+ images = [
1555
+ to_channel_dimension_format(image, data_format, input_channel_dim=input_data_format)
1556
+ for image in images
1557
+ ]
1558
+ encoded_inputs = BatchFeature(data={"pixel_values": images}, tensor_type=return_tensors)
1559
+ if annotations is not None:
1560
+ encoded_inputs["labels"] = [
1561
+ BatchFeature(annotation, tensor_type=return_tensors) for annotation in annotations
1562
+ ]
1563
+
1564
+ return encoded_inputs
1565
+
1566
+ # Copied from transformers.models.owlvit.image_processing_owlvit.OwlViTImageProcessor.post_process_object_detection with OwlViT->GroundingDino
1567
+ def post_process_object_detection(
1568
+ self,
1569
+ outputs: "GroundingDinoObjectDetectionOutput",
1570
+ threshold: float = 0.1,
1571
+ target_sizes: Optional[Union[TensorType, list[tuple]]] = None,
1572
+ ):
1573
+ """
1574
+ Converts the raw output of [`GroundingDinoForObjectDetection`] into final bounding boxes in (top_left_x, top_left_y,
1575
+ bottom_right_x, bottom_right_y) format.
1576
+
1577
+ Args:
1578
+ outputs ([`GroundingDinoObjectDetectionOutput`]):
1579
+ Raw outputs of the model.
1580
+ threshold (`float`, *optional*, defaults to 0.1):
1581
+ Score threshold to keep object detection predictions.
1582
+ target_sizes (`torch.Tensor` or `list[tuple[int, int]]`, *optional*):
1583
+ Tensor of shape `(batch_size, 2)` or list of tuples (`tuple[int, int]`) containing the target size
1584
+ `(height, width)` of each image in the batch. If unset, predictions will not be resized.
1585
+
1586
+ Returns:
1587
+ `list[Dict]`: A list of dictionaries, each dictionary containing the following keys:
1588
+ - "scores": The confidence scores for each predicted box on the image.
1589
+ - "labels": Indexes of the classes predicted by the model on the image.
1590
+ - "boxes": Image bounding boxes in (top_left_x, top_left_y, bottom_right_x, bottom_right_y) format.
1591
+ """
1592
+ batch_logits, batch_boxes = outputs.logits, outputs.pred_boxes
1593
+ batch_size = len(batch_logits)
1594
+
1595
+ if target_sizes is not None and len(target_sizes) != batch_size:
1596
+ raise ValueError("Make sure that you pass in as many target sizes as images")
1597
+
1598
+ # batch_logits of shape (batch_size, num_queries, num_classes)
1599
+ batch_class_logits = torch.max(batch_logits, dim=-1)
1600
+ batch_scores = torch.sigmoid(batch_class_logits.values)
1601
+ batch_labels = batch_class_logits.indices
1602
+
1603
+ # Convert to [x0, y0, x1, y1] format
1604
+ batch_boxes = center_to_corners_format(batch_boxes)
1605
+
1606
+ # Convert from relative [0, 1] to absolute [0, height] coordinates
1607
+ if target_sizes is not None:
1608
+ batch_boxes = _scale_boxes(batch_boxes, target_sizes)
1609
+
1610
+ results = []
1611
+ for scores, labels, boxes in zip(batch_scores, batch_labels, batch_boxes):
1612
+ keep = scores > threshold
1613
+ scores = scores[keep]
1614
+ labels = labels[keep]
1615
+ boxes = boxes[keep]
1616
+ results.append({"scores": scores, "labels": labels, "boxes": boxes})
1617
+
1618
+ return results
1619
+
1620
+
1621
+ __all__ = ["GroundingDinoImageProcessor"]