Thank you very much for the analysis and the ideas - I am grateful and will take them into account.
I propose to build it like this: for each cell, take the maximum recall (for specific 1% FPR), if we know which type (cell) of injection we are detecting. Then we get the integral recall of such an ensemble - integration metric of the dataset x our detector pool. Against which every detector can be compared. Did I understand your idea correctly?
