AlphaMon

AlphaMon is a transformer-based battle policy model for Pokemon Showdown doubles formats, trained on large-scale replay logs and used with a Monte Carlo rollout search layer for decision support.

Current repository: https://github.com/davidemodolo/alphamon

Model summary

AlphaMon is a decoder-only transformer trained on tokenized Showdown battle logs.
It predicts:

  1. Next token (policy head)
  2. Win probability (value head)

At inference time, model outputs are combined with structured rollout search to produce ranked battle actions.

What is currently released

This model repo includes:

  1. Pretrained checkpoint:
    • checkpoints/alphamon_pretrained_best.pt
  2. Finetuned on gen9vgc2026regf checkpoint:
    • checkpoints/alphamon_finetuned_best.pt
  3. Training replay list used for dataset construction:
    • release/training_games_list.txt

Finetuned checkpoints are still in progress and may be added later.

Architecture

Main model config (current defaults from training code):

  1. Context length: 2048
  2. Hidden size: 512
  3. Attention heads: 8
  4. Layers: 6
  5. Dropout: 0.1

Training objective:

  1. Cross-entropy loss for policy tokens
  2. MSE loss for value prediction
  3. Combined loss: policy + 0.5 * value

Data

Training data is derived from publicly available Pokemon Showdown replay logs, harvested by format and filtered/cleaned before tokenization.

Data pipeline includes:

  1. Replay harvesting from Showdown search endpoint
  2. Metadata extraction (winner, ratings, format)
  3. Quality and format filtering
  4. Tokenization and mmap cache construction

Intended use

AlphaMon is intended for:

  1. Research and experimentation in game AI
  2. Decision support for Pokemon Showdown doubles battle states
  3. Study of sequence modeling + search-based action selection

It is not intended for:

  1. Fully autonomous ladder abuse or platform policy violations
  2. Safety-critical applications
  3. Legal advice or compliance decisions

Limitations

  1. Strongly domain-specific to Showdown-style tokenized logs
  2. Quality depends on training replay distribution and filtering
  3. Not guaranteed to choose legal-optimal actions in all edge cases
  4. Performance can degrade on unseen formats or mechanics

Inference notes

In the main project, inference uses:

  1. Monte Carlo rollout search for TURN, TEAM_PREVIEW, and FAINT objectives
  2. Optional team-aware logit bias to nudge generation toward known:
    • species tokens
    • move tokens
    • item tokens
    • tera type tokens

This bias is additive and does not hard-force outputs.

How to use with the project

  1. Clone the main repo:
  2. Download checkpoint from this HF model repo into:
    • checkpoints/alphamon_finetuned_best.pt
  3. Run server inference from the repo:
    • python server/server.py

License and third-party content

This repository is currently marked AGPL-3.0.

Important note: Pokemon-related names, mechanics, and source replay content may be subject to their own rights and terms by their respective owners/platforms.
This model card and repository license apply to the released project assets to the extent permitted.

Citation

If you use AlphaMon in research, please cite:

Modolo, Davide. AlphaMon: Transformer-based Showdown Battle Policy Model.

You can also cite the GitHub repository: https://github.com/davidemodolo/alphamon

Contact

Maintainer: Davide Modolo
HF: https://huggingface.co/davidemodolo
GitHub: https://github.com/davidemodolo

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support