File size: 2,457 Bytes
434ccac | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 | ---
license: bsd-3-clause
tags:
- pytorch
- beit
- retrieval
---
# Beit for Retrieval
## Overview
This repository is a compact, custom PyTorch implementation of **Beit** for **Retrieval**. The **nano** configuration is intended for code review, smoke tests, and small controlled experiments rather than as a production-ready pretrained release.
## Repository status
- The Python file contains the model and runnable example or training entry point.
- `config.json` records the generated architecture settings.
- `training_args.json` records the default experiment recipe.
- `model.safetensors` is a valid initialization checkpoint for smoke tests; it is **not** presented as a trained benchmark checkpoint.
- No benchmark score is claimed in this repository.
## Architecture
| Item | Value |
|---|---|
| Architecture | Beit |
| Scale | nano |
| Attention | sparse |
| Fusion | co attention |
| Activation | swish |
| Normalization | batchnorm |
## Default experiment recipe
The included configuration uses **lion** with a **cosine** schedule. These are starting values in the script, not evidence of a completed run. For a meaningful evaluation, train all baselines with the same data exposure, tuning budget, and random seeds.
## Quick check
```bash
python model.py --help
```
Inspect the script's `__main__` block for its generated smoke-test example. Because this is a custom implementation, generic automatic loading APIs require an explicit adapter before use.
## Evaluation guidance
A useful first evaluation would use **Flickr30k**, report the task metric across at least three seeds, and include a matched-capacity baseline. Keep training logs and environment versions with any published result.
## Limitations
The initialization checkpoint has not been trained or audited for robustness, fairness, or domain transfer. The implementation should be treated as an experimental starting point. Results from a future trained checkpoint must be documented separately from the defaults shipped here.
## Files
- `model.py` — primary artifact
- `README.md` — this documentation
- `config.json` — architecture configuration
- `training_args.json` — default experiment settings
- `model.safetensors` — initialization checkpoint
## License
Released under **bsd-3-clause**. Review the source-data terms separately when this repository is used with external datasets.
|