Text Generation
Transformers
Safetensors
qwen3_moe
Neura Tech AI
Neuron
instruct
llm
transformer
mixture-of-experts
Mixture of Experts
multilingual
24B
Qwen3
Neuron-6x4B-Instruct
conversational
Instructions to use Neura-Tech-AI/Neuron-6x4B-Instruct with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Neura-Tech-AI/Neuron-6x4B-Instruct with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Neura-Tech-AI/Neuron-6x4B-Instruct") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("Neura-Tech-AI/Neuron-6x4B-Instruct") model = AutoModelForCausalLM.from_pretrained("Neura-Tech-AI/Neuron-6x4B-Instruct", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Neura-Tech-AI/Neuron-6x4B-Instruct with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Neura-Tech-AI/Neuron-6x4B-Instruct" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Neura-Tech-AI/Neuron-6x4B-Instruct", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Neura-Tech-AI/Neuron-6x4B-Instruct
- SGLang
How to use Neura-Tech-AI/Neuron-6x4B-Instruct with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Neura-Tech-AI/Neuron-6x4B-Instruct" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Neura-Tech-AI/Neuron-6x4B-Instruct", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Neura-Tech-AI/Neuron-6x4B-Instruct" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Neura-Tech-AI/Neuron-6x4B-Instruct", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use Neura-Tech-AI/Neuron-6x4B-Instruct with Docker Model Runner:
docker model run hf.co/Neura-Tech-AI/Neuron-6x4B-Instruct
File size: 2,483 Bytes
94a4e8e | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 | Neuron-6x4B-Instruct
Copyright (c) 2026 Neura Tech AI
Neuron-6x4B-Instruct is an open-source Mixture of Experts (MoE) large language
model developed by Neura Tech AI.
This model is built upon the following base models:
- Neura-Tech-AI/Neuron-4B-Instruct
Copyright (c) 2026 Neura Tech AI.
- Qwen/Qwen3-4B-Instruct-2507
Copyright (c) 2024 Alibaba Cloud and the Qwen Team.
- Qwen/Qwen3-4B-Thinking-2507
Copyright (c) 2024 Alibaba Cloud and the Qwen Team.
Neuron-4B-Instruct is based on the Qwen3 architecture and includes additional
modifications and improvements developed by Neura Tech AI.
The original Qwen3 models are licensed under the Apache License,
Version 2.0. A copy of the Apache License is included in the LICENSE file.
Neuron-6x4B-Instruct introduces additional modifications made by
Neura Tech AI, including but not limited to:
- Mixture of Experts (MoE) architecture
- Six-expert sparse routing design
- Dynamic expert routing configuration
- Expert composition and parameter merging
- Instruction tuning improvements
- Identity customization
- Chat template customization
- Alignment improvements
- Reasoning enhancements
- Multilingual capability improvements
- Coding and software engineering optimization
- Tool-calling optimization
- Long-context support optimization
- Dataset improvements
- Branding and documentation
- Model packaging and distribution
These modifications are Copyright (c) 2026
Neura Tech AI.
Developer
Neura Tech AI
Official AI Research & Development Organization
Project
Neuron
Model
Neuron-6x4B-Instruct
Architecture
Sparse Transformer Decoder
Mixture of Experts (MoE)
Qwen3 Architecture
Base Models
- Neura-Tech-AI/Neuron-4B-Instruct
- Qwen/Qwen3-4B-Instruct-2507
- Qwen/Qwen3-4B-Thinking-2507
Model Highlights
- Approximately 24 Billion Total Parameters
- Six Specialized Experts
- Dynamic Sparse Expert Routing
- Instruction-Tuned
- Multilingual Language Model
- Optimized for Reasoning, Coding, Mathematics, Tool Calling,
Agentic Workflows, and Long-Context Understanding
Acknowledgment
We sincerely thank Alibaba Cloud and the Qwen Team for releasing the
Qwen3 model family under the Apache License, Version 2.0. Their work
served as the architectural foundation that enabled the development of
Neuron-4B-Instruct and, subsequently, Neuron-6x4B-Instruct.
This NOTICE file is provided solely for attribution purposes and does
not modify, replace, or supersede the terms of the Apache License,
Version 2.0. |