mlboydaisuke commited on
Commit
c9d5301
·
verified ·
1 Parent(s): 41ca64f

Replace dynamic-range int8 with weight-only int8

Browse files

Same change as efficientnet_b1 (merged earlier today), applied for consistency across the family: this repo's dynamic-range int8 is only mildly affected, but weight-only int8 is strictly closer to the float model at the same size reduction.

Measured against the torch float reference on real photos plus fixed random tensors at native resolution: dynamic min logit correlation 0.971 (top-1 4/6) vs weight-only 1.000 (top-1 6/6). The replacement file is quantized directly from this repository's published float32 file (litert quantize, recipe weight_only_wi8_afp32), so the weights are identical.

The model card is updated in the same commit: the metric entries for the removed dynamic file are dropped and a short quantized-variant note with the measured numbers is added.

README.md CHANGED
@@ -28,12 +28,6 @@ model-index:
28
  - name: Top 5 Accuracy (Full Precision)
29
  type: accuracy
30
  value: 0.9691
31
- - name: Top 1 Accuracy (Dynamic Quantized wi8 afp32)
32
- type: accuracy
33
- value: 0.8383
34
- - name: Top 5 Accuracy (Dynamic Quantized wi8 afp32)
35
- type: accuracy
36
- value: 0.9677
37
  ---
38
 
39
  # EfficientNet B6
@@ -49,6 +43,16 @@ acc@1 (on ImageNet-1K): 84.008%
49
  acc@5 (on ImageNet-1K): 96.916%
50
  num_params: 43,040,704
51
 
 
 
 
 
 
 
 
 
 
 
52
  ## Intended uses & limitations
53
 
54
  The model files were converted from pretrained weights from PyTorch Vision. The models may have their own licenses or terms and conditions derived from PyTorch Vision and the dataset used for training. It is your responsibility to determine whether you have permission to use the models for your use case.
 
28
  - name: Top 5 Accuracy (Full Precision)
29
  type: accuracy
30
  value: 0.9691
 
 
 
 
 
 
31
  ---
32
 
33
  # EfficientNet B6
 
43
  acc@5 (on ImageNet-1K): 96.916%
44
  num_params: 43,040,704
45
 
46
+ ### Quantized variant
47
+
48
+ `efficientnet_b6_weight_only_wi8_afp32.tflite` is a weight-only int8
49
+ quantization of the same weights (about 3.7x smaller than float32).
50
+ Weight-only quantization is used instead of dynamic-range quantization
51
+ because EfficientNet's SE and SiLU layers are sensitive to activation
52
+ quantization; in a spot check against the float model the weight-only
53
+ file keeps the top-1 predictions on real photos with a minimum logit
54
+ correlation of 1.000.
55
+
56
  ## Intended uses & limitations
57
 
58
  The model files were converted from pretrained weights from PyTorch Vision. The models may have their own licenses or terms and conditions derived from PyTorch Vision and the dataset used for training. It is your responsibility to determine whether you have permission to use the models for your use case.
efficientnet_b6_dynamic_wi8_afp32.tflite → efficientnet_b6_weight_only_wi8_afp32.tflite RENAMED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:a652ddf1695371800b210de68c36a46aa7eebd8828d2ea1dbf6a790d2f193223
3
- size 46132584
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:ca7a772019bcd4817d776902b4a7d41f4db154dcaa13ec5b75d9a92dc6a91dfc
3
+ size 46252896