Improve model card for MALA model

This PR significantly improves the model card for the MALA model. It adds:

- The `pipeline_tag: image-feature-extraction` to ensure discoverability for relevant tasks on the Hub.
- The `library_name: transformers` for better integration and usage guidance.
- The `license: apache-2.0` for clarity on usage terms.
- A direct link to the paper on Hugging Face Papers.
- The paper's abstract for a comprehensive overview.
- A link to the inferred official GitHub repository for easy access to the code and further resources.
- The BibTeX citation for proper academic attribution.

This will greatly enhance the model's information and usability on the Hugging Face Hub.

Files changed (1) hide show

README.md +32 -1

README.md CHANGED Viewed

	@@ -1 +1,32 @@
1	- ~~MALA's Models~~

+---
+pipeline_tag: image-feature-extraction
+library_name: transformers
+license: apache-2.0
+---
+# MALA: Magnitude-Aware Linear Attention
+This repository contains models for **MALA (Magnitude-Aware Linear Attention)**, a novel attention mechanism introduced in the paper [Rectifying Magnitude Neglect in Linear Attention](https://huggingface.co/papers/2507.00698) (ICCV 2025 Highlight).
+MALA addresses a critical issue in Linear Attention by fully incorporating the magnitude information of the Query, enabling the attention score distribution to dynamically adapt and closely resemble that of Softmax Attention, while maintaining linear complexity. This leads to strong performance across a wide range of tasks.
+## Abstract
+As the core operator of Transformers, Softmax Attention exhibits excellent global modeling capabilities. However, its quadratic complexity limits its applicability to vision tasks. In contrast, Linear Attention shares a similar formulation with Softmax Attention while achieving linear complexity, enabling efficient global information modeling. Nevertheless, Linear Attention suffers from a significant performance degradation compared to standard Softmax Attention. In this paper, we analyze the underlying causes of this issue based on the formulation of Linear Attention. We find that, unlike Softmax Attention, Linear Attention entirely disregards the magnitude information of the Query. This prevents the attention score distribution from dynamically adapting as the Query scales. As a result, despite its structural similarity to Softmax Attention, Linear Attention exhibits a significantly different attention score distribution. Based on this observation, we propose Magnitude-Aware Linear Attention (MALA), which modifies the computation of Linear Attention to fully incorporate the Query’s magnitude. This adjustment allows MALA to generate an attention score distribution that closely resembles Softmax Attention while exhibiting a more well-balanced structure. We evaluate the effectiveness of MALA on multiple tasks, including image classification, object detection, instance segmentation, semantic segmentation, natural language processing, speech recognition, and image generation. Our MALA achieves strong results on all of these tasks.
+## Code
+The official code and other resources for MALA can be found on the [GitHub repository](https://github.com/aldjalkdf/MAViT).
+## Citation
+If you find this work useful, please cite the paper:
+```bibtex
+@inproceedings{fan2024rect,
+      title={Rectifying Magnitude Neglect in Linear Attention},
+      author={Qihang Fan and Huaibo Huang and Yuang Ai and Ran He },
+      year={2025},
+      booktitle={ICCV},
+}
+```