Hugging Face's logo Hugging Face
  • Models
  • Datasets
  • Spaces
  • Buckets new
  • Docs
  • Enterprise
  • Pricing
    • Website
      • Tasks
      • HuggingChat
      • Collections
      • Languages
      • Organizations
    • Community
      • Blog
      • Posts
      • Daily Papers
      • Hardware
      • Learn
      • Discord
      • Forum
      • GitHub
    • Solutions
      • Team & Enterprise
      • Hugging Face PRO
      • Enterprise Support
      • Inference Providers
      • Inference Endpoints
      • Storage Buckets

  • Log In
  • Sign Up

facebook
/
MobileMoE-S-QAT

Text Generation
Transformers
Safetensors
PyTorch
English
mobilemoe
facebook
meta
mixture-of-experts
MoE
on-device
quantization
int4
conversational
custom_code
8-bit precision
Model card Files Files and versions
xet
Community

Instructions to use facebook/MobileMoE-S-QAT with libraries, inference providers, notebooks, and local apps. Follow these links to get started.

  • Libraries
  • Transformers

    How to use facebook/MobileMoE-S-QAT with Transformers:

    # Use a pipeline as a high-level helper
    from transformers import pipeline
    
    pipe = pipeline("text-generation", model="facebook/MobileMoE-S-QAT", trust_remote_code=True)
    messages = [
        {"role": "user", "content": "Who are you?"},
    ]
    pipe(messages)
    # Load model directly
    from transformers import AutoModelForCausalLM
    model = AutoModelForCausalLM.from_pretrained("facebook/MobileMoE-S-QAT", trust_remote_code=True, device_map="auto")
  • Notebooks
  • Google Colab
  • Kaggle
  • Local Apps Settings
  • vLLM

    How to use facebook/MobileMoE-S-QAT with vLLM:

    Install from pip and serve model
    # Install vLLM from pip:
    pip install vllm
    # Start the vLLM server:
    vllm serve "facebook/MobileMoE-S-QAT"
    # Call the server using curl (OpenAI-compatible API):
    curl -X POST "http://localhost:8000/v1/chat/completions" \
    	-H "Content-Type: application/json" \
    	--data '{
    		"model": "facebook/MobileMoE-S-QAT",
    		"messages": [
    			{
    				"role": "user",
    				"content": "What is the capital of France?"
    			}
    		]
    	}'
    Use Docker
    docker model run hf.co/facebook/MobileMoE-S-QAT
  • SGLang

    How to use facebook/MobileMoE-S-QAT with SGLang:

    Install from pip and serve model
    # Install SGLang from pip:
    pip install sglang
    # Start the SGLang server:
    python3 -m sglang.launch_server \
        --model-path "facebook/MobileMoE-S-QAT" \
        --host 0.0.0.0 \
        --port 30000
    # Call the server using curl (OpenAI-compatible API):
    curl -X POST "http://localhost:30000/v1/chat/completions" \
    	-H "Content-Type: application/json" \
    	--data '{
    		"model": "facebook/MobileMoE-S-QAT",
    		"messages": [
    			{
    				"role": "user",
    				"content": "What is the capital of France?"
    			}
    		]
    	}'
    Use Docker images
    docker run --gpus all \
        --shm-size 32g \
        -p 30000:30000 \
        -v ~/.cache/huggingface:/root/.cache/huggingface \
        --env "HF_TOKEN=<secret>" \
        --ipc=host \
        lmsysorg/sglang:latest \
        python3 -m sglang.launch_server \
            --model-path "facebook/MobileMoE-S-QAT" \
            --host 0.0.0.0 \
            --port 30000
    # Call the server using curl (OpenAI-compatible API):
    curl -X POST "http://localhost:30000/v1/chat/completions" \
    	-H "Content-Type: application/json" \
    	--data '{
    		"model": "facebook/MobileMoE-S-QAT",
    		"messages": [
    			{
    				"role": "user",
    				"content": "What is the capital of France?"
    			}
    		]
    	}'
  • Docker Model Runner

    How to use facebook/MobileMoE-S-QAT with Docker Model Runner:

    docker model run hf.co/facebook/MobileMoE-S-QAT

You need to agree to share your contact information to access this model

The information you provide will be collected, stored, processed and shared in accordance with the Meta Privacy Policy.

Log in or Sign Up to review the conditions and access this model content.

Gated model
You can list files but not access them

Preview of files found in this repository
  • .gitattributes
    1.58 kB
    Add MobileMoE-S-QAT model card and figures about 23 hours ago
  • LICENSE
    11.5 kB
    Add FAIR Noncommercial Research License 6 days ago
  • README.md
    17 kB
    Update README.md about 21 hours ago
  • config.json
    2.79 kB
    Add config, generation config and tokenizer about 23 hours ago
  • configuration_mobilemoe.py
    3.88 kB
    Add MobileMoE INT4 modeling code about 23 hours ago
  • generation_config.json
    239 Bytes
    Add config, generation config and tokenizer about 23 hours ago
  • mobilemoe_pareto.png
    190 kB
    xet
    Add MobileMoE-S-QAT model card and figures about 23 hours ago
  • mobilemoe_recipe.png
    55.3 kB
    Add MobileMoE-S-QAT model card and figures about 23 hours ago
  • model.safetensors
    714 MB
    xet
    Add MobileMoE-S-QAT INT4 weights about 23 hours ago
  • modeling_mobilemoe.py
    35.7 kB
    Add MobileMoE INT4 modeling code about 23 hours ago
  • special_tokens_map.json
    99 Bytes
    Add config, generation config and tokenizer about 23 hours ago
  • tokenizer.json
    4.25 MB
    Add config, generation config and tokenizer about 23 hours ago
  • tokenizer_config.json
    886 Bytes
    Add config, generation config and tokenizer about 23 hours ago