|
Download README.md from pphamlong/assignment-retrieval: direct link, hf CLI and curl.
- Browser
- Download file 2.41 kB
-
https://huggingface.co/pphamlong/assignment-retrieval/resolve/main/README.md
- Command line
-
hf download hf://pphamlong/assignment-retrieval/README.md
-
curl -L -o README.md https://huggingface.co/pphamlong/assignment-retrieval/resolve/main/README.md
2.41 kB
| license: apache-2.0 | |
| tags: | |
| - pytorch | |
| - mae | |
| - retrieval | |
| # Mae for Retrieval | |
| ## Overview | |
| Working implementation of **Mae** for **Retrieval** using a **large** configuration. The repository focuses on transparent code and repeatable smoke tests; benchmark claims are deliberately omitted. | |
| ## Repository status | |
| - The Python file contains the model and runnable example or training entry point. | |
| - `config.json` records the generated architecture settings. | |
| - `training_args.json` records the default experiment recipe. | |
| - `model.safetensors` is a valid initialization checkpoint for smoke tests; it is **not** presented as a trained benchmark checkpoint. | |
| - No benchmark score is claimed in this repository. | |
| ## Architecture | |
| | Item | Value | | |
| |---|---| | |
| | Architecture | Mae | | |
| | Scale | large | | |
| | Attention | multi query | | |
| | Fusion | co attention | | |
| | Activation | approx gelu | | |
| | Normalization | layernorm | | |
| ## Default experiment recipe | |
| The included configuration uses **rmsprop** with a **step** schedule. These are starting values in the script, not evidence of a completed run. For a meaningful evaluation, train all baselines with the same data exposure, tuning budget, and random seeds. | |
| ## Quick check | |
| ```bash | |
| python run.py --help | |
| ``` | |
| Inspect the script's `__main__` block for its generated smoke-test example. Because this is a custom implementation, generic automatic loading APIs require an explicit adapter before use. | |
| ## Evaluation guidance | |
| A useful first evaluation would use **Flickr30k**, report the task metric across at least three seeds, and include a matched-capacity baseline. Keep training logs and environment versions with any published result. | |
| ## Limitations | |
| The initialization checkpoint has not been trained or audited for robustness, fairness, or domain transfer. The implementation should be treated as an experimental starting point. Results from a future trained checkpoint must be documented separately from the defaults shipped here. | |
| ## Files | |
| - `run.py` — primary artifact | |
| - `README.md` — this documentation | |
| - `config.json` — architecture configuration | |
| - `training_args.json` — default experiment settings | |
| - `model.safetensors` — initialization checkpoint | |
| ## License | |
| Released under **apache-2.0**. Review the source-data terms separately when this repository is used with external datasets. | |