--- tags: - font-generation - image-generation - autoregressive - multimodal - few-shot-learning - computer-vision library_name: pytorch --- # GAR-Font: Beyond Patches: Global-aware Autoregressive Model for Multimodal Few-Shot Font Generation

[![arXiv](https://img.shields.io/badge/arXiv-2601.01593-b31b1b?logo=arxiv)](https://arxiv.org/abs/2601.01593) [![Homepage](https://img.shields.io/badge/Homepage-GAR--Font-orange)](https://xtryer-s.github.io/projects_pages/GAR_Font/) [![Code](https://img.shields.io/badge/GitHub-Repository-blue?logo=github)](https://github.com/xTryer-s/GAR-Font)
--- ## 📌 Overview GAR-Font is a **global-aware autoregressive model for multimodal few-shot font generation**. It aims to generate high-quality glyphs with limited style references by modeling both global font characteristics and local glyph details. GAR-Font supports two generation settings: - **Vision-Only GAR-Font** Generate glyphs using only visual style references. - **Vision-Language GAR-Font** Generate glyphs using both visual references and natural language style descriptions. This repository provides a pretrained checkpoint of Vision-Only GAR-Font. --- ## 📦 Model Weights We provide the pretrained checkpoint: **generator_ckpt_pruned.pt** The checkpoint contains: - `vq_model`: pretrained VQ tokenizer for glyph representation - `model`: pretrained GAR-Font autoregressive generator The checkpoint can be directly used for inference and further research. --- ## 🚀 Usage Download: **generator_ckpt_pruned.pt** and use it as the `--generator-ckpt` argument in the inference scripts. For installation instructions, inference examples, training details, and complete implementation, please refer to our official GitHub repository: 🔗 https://github.com/xTryer-s/GAR-Font --- ## 📖 Paper **Beyond Patches: Global-aware Autoregressive Model for Multimodal Few-Shot Font Generation** arXiv: https://arxiv.org/abs/2601.01593 --- ## Citation If you find GAR-Font useful, please cite our paper: ```bibtex @misc{cai2026patchesglobalawareautoregressivemodel, title={Beyond Patches: Global-aware Autoregressive Model for Multimodal Few-Shot Font Generation}, author={Haonan Cai and Yuxuan Luo and Zhouhui Lian}, year={2026}, eprint={2601.01593}, archivePrefix={arXiv}, primaryClass={cs.CV}, url={https://arxiv.org/abs/2601.01593}, }