Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
EschaLabs
/
escha-runtime-qwen3dense
like
7
Follow
Escha Labs
185
sglang
quantization
cuda
inference
escha
qwen3
dense
License:
apache-2.0
Model card
Files
Files and versions
xet
Community
1
Copy to bucket
new
main
escha-runtime-qwen3dense
16.9 MB
Ctrl+K
Ctrl+K
1 contributor
History:
9 commits
yzhwang
Clarify that 64k context and 8-9 streams are not simultaneous
36bb7bf
verified
about 9 hours ago
THIRD_PARTY_LICENSES
escha 1.2.0: docs + changelog for tensor parallelism
about 15 hours ago
sglang
Correct RADIX guidance (prefix cache is exact-match only); add long-context sizing
about 9 hours ago
.gitattributes
1.83 kB
escha 1.2.0 wheel: tensor parallelism (gated on world_size>1)
about 15 hours ago
LICENSE
Safe
11.3 kB
Escha Runtime qwen3dense — escha 1.1.0+qwen3dense SGLang wheel, serve.sh, cookbook
1 day ago
README.md
11.9 kB
Clarify that 64k context and 8-9 streams are not simultaneous
about 9 hours ago