Spaces:
Running on Zero
Running on Zero
| """Counting, and knowing when not to. | |
| This is the layer that turns boxes into a claim. It exists separately from | |
| `detection.py` because the hard part of counting livestock is not detection — it | |
| is being honest about the frames where detection stops working. | |
| Two capabilities share every line of it: | |
| **`cattle_detection`** counts cattle in a paddock. Cattle are large, separated, | |
| and there are tens of them. A COCO detector already has a `cow` class and this | |
| is genuinely the right tool. | |
| **`poultry_count`** counts birds in a frame. For a backyard flock or a yard of | |
| fifty layers, the same detector works. For 12,000 broilers in a shed it does | |
| not, and no threshold tuning will make it: the birds overlap, each one covers a | |
| few hundred pixels, and non-maximum suppression merges the ones that remain. The | |
| answer is a density head, not a better detector (ADR 0014). | |
| So this module measures whether it is still in the regime it was validated for, | |
| and when it is not, **it reports no count at all**. A sample presented as a count | |
| is the failure mode that costs the product its credibility: a farmer shown "18" | |
| for a shed of several hundred does not conclude that the number means something | |
| narrower than they thought. | |
| **How it knows.** It counts the frame three ways — whole, 2x2, 3x3 — and watches | |
| what the count does (`app/tiling.py`). In a frame the detector can read, the | |
| count stops moving, because there was nothing left to find. In a shed it never | |
| stops, because there are always more birds behind the ones in front. That is a | |
| measurement of what the detector is *missing*, which is the thing a saturation | |
| guard has to know and the thing box sizes cannot tell it. | |
| The first version of this file guarded on box size instead, and the evaluation | |
| set caught it: on a broiler house of a thousand birds it found a handful of large | |
| foreground birds, concluded the frame was sparse, and published the handful as a | |
| count. It withheld a number on three of twenty uncountable frames; the tiled test | |
| withholds on twenty of twenty (ADR 0018). | |
| **The thresholds belong to the detector, not to the problem.** They were first | |
| derived on YOLO11m and then inherited unchanged when the shipped model became | |
| YOLOX-m, which cost 22 points of coverage for no gain in safety — the grid a | |
| frame settles at depends on how much the detector's whole-frame pass resolves, | |
| and that is precisely what differs between detectors. Re-deriving them removed a | |
| rule entirely and restored the coverage. If the artefact changes again, re-run | |
| `evaluation/run.py` before trusting a number in this file. | |
| """ | |
| from __future__ import annotations | |
| from dataclasses import dataclass | |
| from statistics import median | |
| from uuid import UUID, uuid4 | |
| from app.capabilities import Capability, FORBIDDEN_CLAIMS | |
| from app.detectors import Detection, build | |
| from app.media import MediaRef, MediaStore | |
| from app.providers import ModelArtefact | |
| from app.quality import QualityVerdict, assess | |
| from app.schemas import ( | |
| ConfidenceLabel, | |
| InferenceLocation, | |
| InferenceRequest, | |
| InferenceResult, | |
| Observation, | |
| QualityCheck, | |
| ) | |
| from app.tiling import Level, converged, pyramid, subject_count | |
| #: How much the count may grow when the frame is cut finer before the frame is | |
| #: called unreadable. | |
| #: | |
| #: **Every number below was re-measured on the shipped YOLOX-m artefact** over | |
| #: the 61-image set (ADR 0018). They previously came from YOLO11m, and carrying | |
| #: them across backends was wrong: the grid at which a frame settles depends on | |
| #: how much the detector's whole-frame pass can resolve, which is exactly what | |
| #: differs between detectors. Re-deriving them moved coverage from 0.71 to 0.935 | |
| #: without letting a single dense frame through. | |
| #: | |
| #: **The two classes overlap here, and the honest reading is that this rule does | |
| #: not separate them on its own.** Last-refinement growth runs 0%–75% on frames a | |
| #: human could count and 7%–162% on frames a human could not. At 20% it catches | |
| #: seventeen of the twenty uncountable frames and wrongly withholds two countable | |
| #: ones; the two dense frames that slip past it are caught by | |
| #: `MAX_VALIDATED_COUNT`. Neither rule is sufficient alone, which is why they are | |
| #: OR-ed rather than tuned against each other. | |
| COUNT_GROWTH_TOLERANCE = 0.20 | |
| #: Median share of the frame a subject covers, below which no count is reported | |
| #: *even after the count has settled*. A frame can settle simply because every | |
| #: animal in it is a smudge the detector resolves the same way at every grid. | |
| #: | |
| #: **This rule never fires on the evaluation set, so it carries no evidence.** | |
| #: Measured at the grid that ran: the smallest median on a countable frame is | |
| #: 0.00077 (a hillside herd in the Turkish Eğribel pass) and the smallest on an | |
| #: uncountable one is 0.00042 — but every frame it would have caught was already | |
| #: withheld by convergence. 0.0006 sits in a gap two frames wide. | |
| #: | |
| #: It is kept as a backstop for a frame type this set does not contain, and it is | |
| #: labelled unexercised rather than described as if it were doing work. Do not | |
| #: cite it as a reason the guard is safe. | |
| SATURATION_MEDIAN_AREA_FRACTION = 0.0006 | |
| #: Median detection score below which the count is reported but capped at `low` | |
| #: confidence and a recapture is asked for. Measured on the fixtures: cattle | |
| #: 0.72, sparse hens 0.49, dense hens 0.47. Sparse and dense are only 0.02 | |
| #: apart, so this cannot decide whether to publish a number — it can only decide | |
| #: how much to trust one. | |
| LOW_CONFIDENCE_MEDIAN_SCORE = 0.50 | |
| #: Confusable detections per subject, above which the frame is *reported* as | |
| #: class-confused. **This is an observation, not a suppression**, and it used to | |
| #: be the latter. | |
| #: | |
| #: As a suppression rule it was measured net-harmful: on the 61-image set it | |
| #: withheld exactly one count — `cattle_ng_red_bororo`, three Red Bororo cattle, | |
| #: where the detector found all three and also called some of them `horse` — and | |
| #: it caught none of the twenty uncountable frames, because convergence and the | |
| #: count ceiling had already taken all twenty. It cost a correct answer on a | |
| #: Nigerian frame and bought nothing. | |
| #: | |
| #: The signal is still worth recording: it was real when the small backend called | |
| #: seven of thirteen hens `sheep`, and a future artefact may bring it back. So it | |
| #: rides along as `confusable_detections` and a warning, where it informs a | |
| #: reader without silently deleting a number. | |
| CONFUSION_RATIO = 1.0 | |
| #: Minimum subjects before that ratio means anything. Below this a single | |
| #: mislabelled animal would trip it. | |
| CONFUSION_MIN_SUBJECTS = 3 | |
| #: The largest count that has been checked against ground truth. Above it the | |
| #: service reports no number — not because a larger count is necessarily wrong, | |
| #: but because nobody has ever verified one. | |
| #: | |
| #: **Load-bearing, and tight.** It is what catches the two dense frames whose | |
| #: counts settle anyway: a Swedish free-range yard that stabilises at 32 birds | |
| #: and a Karamoja kraal that stabilises at 30, both holding many times that. The | |
| #: largest correctly published count on the set is 20 animals against a human | |
| #: count of 19, so the threshold sits directly on the edge of the evidence | |
| #: rather than at a comfortable distance from it. | |
| #: | |
| #: Raise it by measuring more frames, not by deciding the detector is probably | |
| #: fine up there. | |
| MAX_VALIDATED_COUNT = 20 | |
| HIGH_CONFIDENCE_SCORE = 0.70 | |
| MEDIUM_CONFIDENCE_SCORE = 0.50 | |
| #: Share of the frame one subject must cover before the per-animal capabilities | |
| #: — weight, body condition, skin — have something they could work with. At 15% | |
| #: of a 1080-line capture the animal is roughly 400 px across, which is the point | |
| #: below which a girth measured off it is noise rather than a measurement. | |
| ISOLATION_MIN_AREA_FRACTION = 0.15 | |
| class CountingProfile: | |
| """Everything that differs between counting cattle and counting birds.""" | |
| capability_key: str | |
| subject_noun: str | |
| count_observation: str | |
| #: COCO classes that are the subject. Only these are counted. | |
| subject_classes: tuple[str, ...] | |
| #: COCO classes the detector reaches for when it can no longer tell what it | |
| #: is looking at. **Never counted** — they are a signal, not a subject. | |
| #: The shipped YOLOX-m does not confuse them on any of the 61 evaluation | |
| #: images, so the signal is unexercised: see `CONFUSION_MIN_SUBJECTS`. | |
| confusable_classes: tuple[str, ...] | |
| unit: str | |
| #: Said on every run, whatever the result. These are the claims the brief | |
| #: forbids, stated before anyone can misread the number. | |
| standing_warning: str | |
| saturation_warning: str | |
| #: Whether a single well-framed subject is what this capability is for. | |
| reports_isolation: bool = False | |
| CATTLE_DETECTION = CountingProfile( | |
| capability_key="cattle_detection", | |
| subject_noun="cattle", | |
| count_observation="cattle_visible", | |
| subject_classes=("cow",), | |
| confusable_classes=("horse", "sheep"), | |
| unit="animals", | |
| standing_warning=( | |
| "A count of the animals visible in this frame. It is not the herd size — " | |
| "animals behind others, behind cover, or out of frame are not in it." | |
| ), | |
| saturation_warning=( | |
| "There are more cattle in this frame than can be counted from it. No " | |
| "number is reported. Move closer, or frame a smaller part of the herd." | |
| ), | |
| reports_isolation=True, | |
| ) | |
| POULTRY_COUNT = CountingProfile( | |
| capability_key="poultry_count", | |
| subject_noun="birds", | |
| count_observation="birds_visible", | |
| subject_classes=("bird",), | |
| confusable_classes=("sheep", "cat", "dog"), | |
| unit="birds", | |
| standing_warning=( | |
| "A count of the birds visible in this frame. It is never the flock " | |
| "population, and it must not be used to reconcile a house." | |
| ), | |
| saturation_warning=( | |
| "This flock is denser than a detector can count. No number is reported, " | |
| "because a detector undercounts a crowded shed by an amount nobody can " | |
| "estimate. Frame a smaller section, closer in." | |
| ), | |
| ) | |
| class DetectionCountRunner: | |
| """Runs one counting capability against one frame.""" | |
| def __init__(self, profile: CountingProfile) -> None: | |
| self.profile = profile | |
| def run( | |
| self, | |
| *, | |
| request: InferenceRequest, | |
| capability: Capability, | |
| artefact: ModelArtefact, | |
| store: MediaStore, | |
| request_id: UUID | None = None, | |
| ) -> InferenceResult: | |
| profile = self.profile | |
| request_id = request_id or uuid4() | |
| warnings: list[str] = [profile.standing_warning] | |
| # These capabilities read one frame. Saying so beats silently ignoring | |
| # the rest, and beats pretending a count was aggregated across them. | |
| if len(request.media_ids) > capability.frames_required: | |
| warnings.append( | |
| f"{len(request.media_ids)} frames were supplied; this capability " | |
| f"reads {capability.frames_required}." | |
| ) | |
| # `farm_id` and `captured_at` are what let a blob store find the object | |
| # in one request instead of scanning a prefix (`app/media.py`). A local | |
| # store ignores them. | |
| image = store.open_image(MediaRef( | |
| media_id=request.media_ids[0], | |
| farm_id=request.farm_id, | |
| captured_at=request.captured_at, | |
| object_path=request.path_for(request.media_ids[0]), | |
| )) | |
| verdict = assess(image) | |
| if verdict.blocked: | |
| return self._blocked(request, capability, artefact, verdict, warnings, request_id) | |
| detector = build(artefact) | |
| levels = pyramid( | |
| detector, image, profile.subject_classes, COUNT_GROWTH_TOLERANCE, | |
| ) | |
| final = levels[-1] | |
| settled = len(levels) < 2 or converged( | |
| levels[-2], levels[-1], profile.subject_classes, COUNT_GROWTH_TOLERANCE, | |
| ) | |
| subjects = [d for d in final.detections if d.label in profile.subject_classes] | |
| confusable = [d for d in final.detections if d.label in profile.confusable_classes] | |
| checks = list(verdict.checks) | |
| checks.append(self._framing_check(subjects)) | |
| checks.append(self._convergence_check(levels, profile, settled)) | |
| if profile.reports_isolation: | |
| checks.append(self._isolation_check(subjects)) | |
| saturation = self._withhold_reason(subjects, confusable, settled) | |
| observations: list[Observation] = [ | |
| # Emitted on every path, including the ones that publish no count, | |
| # because these three are what let a threshold be re-derived later | |
| # from stored results instead of re-run from photographs nobody kept. | |
| Observation(type="counting_grid", value=float(final.grid), confidence=None), | |
| Observation( | |
| type="subjects_detected", value=float(len(subjects)), | |
| unit=profile.unit, confidence=None, | |
| ), | |
| ] | |
| if subjects: | |
| observations.append(Observation( | |
| type="median_subject_frame_fraction", | |
| value=round(median(d.area_fraction for d in subjects), 5), | |
| unit="fraction", confidence=None, | |
| )) | |
| observations.append(Observation( | |
| type="largest_subject_frame_fraction", | |
| value=round(max(d.area_fraction for d in subjects), 5), | |
| unit="fraction", confidence=None, | |
| )) | |
| observations.append(Observation( | |
| type="confusable_detections", | |
| value=float(len(confusable)), | |
| confidence=None, | |
| )) | |
| if self._is_class_confused(subjects, confusable): | |
| # Reported, never suppressing. Measured net-harmful as a guard; see | |
| # `CONFUSION_RATIO`. | |
| warnings.append( | |
| f"The detector also labelled {len(confusable)} things in this " | |
| f"frame as another animal, which is as many as it called " | |
| f"{profile.subject_noun}. It may be struggling to tell what it " | |
| f"is looking at, so treat the number as a rough indication." | |
| ) | |
| if not subjects: | |
| # **Not a count of zero.** "We could not find any birds" and "there | |
| # are no birds" are different claims, and on a packed broiler house | |
| # the detector produces the first while the second would be absurd. | |
| # Emitting no count observation is what keeps the app from rendering | |
| # a zero it would have to defend. | |
| warnings.append( | |
| f"No {profile.subject_noun} were found in this frame. That is not " | |
| f"a count of zero — it means nothing recognisable was detected. " | |
| f"Capture again, closer in and better lit." | |
| ) | |
| confidence = ConfidenceLabel.LOW | |
| recapture = True | |
| elif saturation is not None: | |
| warnings.append(profile.saturation_warning) | |
| observations.append(Observation( | |
| type="count_withheld", | |
| value=saturation, | |
| confidence=None, | |
| )) | |
| # `subjects_detected` above is already the floor. It is deliberately | |
| # not named as a count anywhere on this path. | |
| confidence = ConfidenceLabel.LOW | |
| recapture = True | |
| else: | |
| mean_score = sum(d.score for d in subjects) / len(subjects) | |
| observations.append(Observation( | |
| type=profile.count_observation, | |
| value=float(len(subjects)), | |
| unit=profile.unit, | |
| confidence=round(mean_score, 3), | |
| )) | |
| confidence = _label(mean_score) | |
| recapture = False | |
| if median(d.score for d in subjects) < LOW_CONFIDENCE_MEDIAN_SCORE: | |
| warnings.append( | |
| f"The detector was unsure about most of these {profile.subject_noun}. " | |
| f"Treat the number as a rough indication and capture again closer in." | |
| ) | |
| confidence = ConfidenceLabel.LOW | |
| recapture = True | |
| if verdict.degraded: | |
| # A frame the gate flagged cannot produce a high-confidence claim, | |
| # whatever the detector's own scores say about it. | |
| confidence = ConfidenceLabel.LOW | |
| recapture = True | |
| return self._result( | |
| request=request, | |
| capability=capability, | |
| artefact=artefact, | |
| request_id=request_id, | |
| observations=observations, | |
| confidence=confidence, | |
| checks=checks, | |
| warnings=warnings, | |
| recapture=recapture, | |
| ) | |
| def _framing_check(self, subjects: list[Detection]) -> QualityCheck: | |
| if subjects: | |
| return QualityCheck(check="framing", passed=True) | |
| return QualityCheck( | |
| check="framing", passed=False, | |
| detail=f"No {self.profile.subject_noun} found in this frame.", | |
| ) | |
| def _isolation_check(self, subjects: list[Detection]) -> QualityCheck: | |
| """Whether the per-animal capabilities could use this frame. | |
| `cattle_detection` runs before weight, body condition and skin, and each | |
| of those needs one animal filling the frame. Reporting that here saves a | |
| second capture attempt later. | |
| """ | |
| large = [d for d in subjects if d.area_fraction >= ISOLATION_MIN_AREA_FRACTION] | |
| if len(large) == 1: | |
| return QualityCheck(check="subject_isolation", passed=True) | |
| return QualityCheck( | |
| check="subject_isolation", passed=False, | |
| detail=( | |
| f"{len(subjects)} animals in frame and {len(large)} close enough to " | |
| f"assess individually. Per-animal capabilities need one animal, " | |
| f"filling the frame." | |
| ), | |
| ) | |
| def _convergence_check( | |
| self, levels: list[Level], profile: CountingProfile, settled: bool | |
| ) -> QualityCheck: | |
| counts = [subject_count(level, profile.subject_classes) for level in levels] | |
| trail = " → ".join( | |
| f"{level.grid}x{level.grid}: {count}" for level, count in zip(levels, counts) | |
| ) | |
| if settled: | |
| return QualityCheck(check="count_convergence", passed=True, detail=trail) | |
| return QualityCheck( | |
| check="count_convergence", passed=False, | |
| detail=( | |
| f"The count kept rising as the frame was read more finely " | |
| f"({trail}), so animals are still hidden behind other animals." | |
| ), | |
| ) | |
| def _is_class_confused( | |
| self, subjects: list[Detection], confusable: list[Detection] | |
| ) -> bool: | |
| """Whether the detector is reaching for neighbouring classes as often as | |
| the right one. A reported signal, not a reason to withhold a count.""" | |
| return ( | |
| len(subjects) >= CONFUSION_MIN_SUBJECTS | |
| and len(confusable) >= len(subjects) * CONFUSION_RATIO | |
| ) | |
| def _withhold_reason( | |
| self, | |
| subjects: list[Detection], | |
| confusable: list[Detection], | |
| settled: bool, | |
| ) -> str | None: | |
| """Whether this frame has left the regime the detector was validated in. | |
| Any one signal is enough. They are OR-ed rather than AND-ed on purpose: a | |
| guard that needs every signal to agree is a guard that goes quiet as soon | |
| as one of them drifts, and going quiet here means publishing a number | |
| that is wrong by an unknown factor. | |
| On the 61-image set the two working signals withhold every one of the | |
| twenty uncountable frames and wrongly withhold two of the thirty-one | |
| countable ones. The last two signals never fire; their comments say so. | |
| """ | |
| if not subjects: | |
| return None | |
| # Seventeen of the twenty uncountable frames stop here. | |
| if not settled: | |
| return "count_did_not_converge" | |
| # The other two. A frame can settle and still be a shed: the detector | |
| # runs out of things it can resolve, so the count stops moving for the | |
| # wrong reason. A count larger than anything ever checked is the signal. | |
| if len(subjects) > MAX_VALIDATED_COUNT: | |
| return "beyond_validated_range" | |
| if median(d.area_fraction for d in subjects) < SATURATION_MEDIAN_AREA_FRACTION: | |
| return "subjects_too_small" | |
| return None | |
| def _blocked( | |
| self, request, capability, artefact, verdict: QualityVerdict, warnings, request_id | |
| ) -> InferenceResult: | |
| failure = verdict.first_failure | |
| warnings.append( | |
| failure.detail if failure and failure.detail | |
| else "The capture was not usable." | |
| ) | |
| return self._result( | |
| request=request, | |
| capability=capability, | |
| artefact=artefact, | |
| request_id=request_id, | |
| observations=[], | |
| confidence=None, | |
| checks=list(verdict.checks), | |
| warnings=warnings, | |
| recapture=True, | |
| ) | |
| def _result( | |
| self, *, request, capability, artefact, request_id, observations, | |
| confidence, checks, warnings, recapture, | |
| ) -> InferenceResult: | |
| forbidden = [o.type for o in observations if o.type in FORBIDDEN_CLAIMS] | |
| if forbidden: | |
| # Belt and braces. The registry holds these as data precisely so a | |
| # runner can be stopped by them rather than reviewed against them. | |
| raise ValueError(f"{capability.key} tried to emit a forbidden claim: {forbidden}") | |
| return InferenceResult( | |
| request_id=request_id, | |
| capability_key=capability.key, | |
| model_id=artefact.model_id, | |
| model_version=artefact.version, | |
| inference_location=InferenceLocation.REMOTE, | |
| subject_type=request.subject_type, | |
| subject_id=request.subject_id, | |
| observations=observations, | |
| # A measurement carries no interpretation. A count of animals is a | |
| # fact about the frame; what a farmer should do about it is not | |
| # something this model knows (ADR 0006). | |
| interpretations=[], | |
| observation_confidence=confidence, | |
| interpretation_confidence=None, | |
| quality_checks=checks, | |
| warnings=warnings, | |
| recommended_recapture=recapture, | |
| ) | |
| def _label(score: float) -> ConfidenceLabel: | |
| if score >= HIGH_CONFIDENCE_SCORE: | |
| return ConfidenceLabel.HIGH | |
| if score >= MEDIUM_CONFIDENCE_SCORE: | |
| return ConfidenceLabel.MEDIUM | |
| return ConfidenceLabel.LOW | |
| #: Capabilities with an implemented adapter. A capability that has a validated | |
| #: artefact but is absent from here returns 501 rather than a plausible result — | |
| #: see `models/README.md`. | |
| RUNNERS: dict[str, DetectionCountRunner] = { | |
| CATTLE_DETECTION.capability_key: DetectionCountRunner(CATTLE_DETECTION), | |
| POULTRY_COUNT.capability_key: DetectionCountRunner(POULTRY_COUNT), | |
| } | |