mlboydaisuke commited on
Commit
c99338c
·
verified ·
1 Parent(s): 8111e0a

Add measured performance section (M4 Max)

Browse files

Thanks for publishing this bundle — it runs cleanly on both backends. This PR adds a measured Performance section so the card carries numbers alongside the install steps. Measured with the litert-lm CLI (`benchmark -p 256 -d 256 --runs 3 --cache no`) on an idle Apple M4 Max, generation-gated first (the backend produced correct text before any number was recorded). Happy to adjust the format if you'd like these to read differently.

Files changed (1) hide show
  1. README.md +12 -1
README.md CHANGED
@@ -29,4 +29,15 @@ deployment on Android using the
29
  * Follow the instructions in the app.
30
 
31
  To build the demo app from source, please follow the [instructions](https://github.com/google-ai-edge/gallery/blob/main/README.md)
32
- from the GitHub repository.
 
 
 
 
 
 
 
 
 
 
 
 
29
  * Follow the instructions in the app.
30
 
31
  To build the demo app from source, please follow the [instructions](https://github.com/google-ai-edge/gallery/blob/main/README.md)
32
+ from the GitHub repository.
33
+
34
+ ## Performance (Apple M4 Max, measured)
35
+
36
+ Measured with the LiteRT-LM CLI: `litert-lm benchmark -p 256 -d 256 --runs 3 --cache no`
37
+ (litert-lm 0.15.0) on an idle Apple M4 Max (macOS); 256 prefill / 256 decode tokens, 3 iterations
38
+ averaged by the tool. A desktop reference point — phone-side figures vary by SoC and backend.
39
+
40
+ | Backend | Prefill (tokens/s) | Decode (tokens/s) | Time-to-first-token (s) |
41
+ |---|---|---|---|
42
+ | CPU | 1,698 | 104.9 | 0.22 |
43
+ | GPU | 7,571 | 259.6 | 0.04 |